Two-party privacy protection decision tree evaluation using additive homomorphic encryption
Through secure multi-party computing technology and decision tree learning algorithm, the problem that a classifier for continuous authentication requires a large amount of sensitive user data is solved, and the method of classifying user identity without publicizing the data is realized, which enhances privacy protection and security.
Patent Information
- Application Number
- CN202411607516.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-24
- Filing Date
- 2024-11-12
- Publication Date
- 2025-05-13
AI Technical Summary
Prior art requires a large amount of user data when creating classifiers for continuous authentication, which are highly sensitive and publicly used invade privacy and security.
The secure multi-party computing (MPC) technology is adopted, and the three-party protocol includes a client and two servers. The decision tree learning algorithm is used as the underlying classifier. The classification results of the encrypted input are calculated without publicizing user data.
This implements a method of classifying user identities without publicizing user data, protects sensitive user data and trained classifiers, and enhances privacy protection and security.
Smart Images

Figure CN119989162A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 548,289, filed on November 13, 2023, which is incorporated herein by reference in its entirety. Technical Field
[0003] The subject matter disclosed herein generally relates to privacy protection. More specifically, but not exclusively, the subject matter relates to methods for evaluating decision trees in a privacy-preserving manner. Background Art
[0004] The increasing sophistication of security measures due to the emergence of different types of attacks requires a variety of security techniques to effectively protect data and applications. One such technique is continuous authentication, which verifies the identity of the user throughout the interaction with the system by continuously monitoring and evaluating behavioral user data and biometric user data. However, creating classifiers for continuous authentication requires a large amount of user data, which is in most cases highly sensitive, leading to its disclosure infringing privacy and security. Summary of the invention
[0005] An example aspect of the present disclosure relates to a computer-implemented method for evaluating a decision tree in a privacy-preserving manner, comprising: receiving, by a first server, a first partial share of a decision tree and encrypted input data from a client, wherein the encrypted input data comprises a set of attributes; receiving, by a second server, a second partial share of the decision tree; by each of the first server and the second server: communicating with another server to calculate a classification result of the encrypted input data using the corresponding partial share of the decision tree received thereby and a secure multi-party computing method, the secure multi-party computing method comprising additive homomorphic encryption and oblivious transfer, such that the classification result is calculated without decrypting the encrypted input data, and such that the classification result is ultimately received by the first server; and sending, by the first server, the classification result to the client. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] The accompanying drawings are incorporated herein and constitute a part of the specification.
[0007] Figure 1 An example biometric recognition system is depicted in accordance with an embodiment.
[0008] Figure 2 An example of a full binary tree is shown.
[0009] Figure 3 An example of decision tree evaluation using arbitrary data is shown.
[0010] Figure 4 The operation mode of the 2-to-1 oblivious transfer (OT) protocol is shown.
[0011] Figure 5 is a block diagram of an example system that performs privacy-preserving decision tree evaluation in accordance with some embodiments.
[0012] Figure 6 Depicted is an overview of decision tree sharing during the initialization phase of a process for performing privacy-preserving decision tree evaluation.
[0013] Figure 7 An overview of an integer comparison protocol applied for performing privacy-preserving decision tree evaluation is shown in accordance with some embodiments.
[0014] Figure 8 An application of a secure integer comparison protocol for component-wise comparison of attribute vectors and threshold vectors to perform privacy-preserving decision tree evaluation is shown in accordance with some embodiments.
[0015] Fig.9A Evaluating decision nodes as true and assigning comparison results to corresponding edges to determine traversal directions as part of performing a privacy-preserving decision tree evaluation in accordance with some embodiments is shown.
[0016] Fig. 9B Evaluating a decision node as false and assigning the comparison result to a corresponding edge to determine a traversal direction as part of performing a privacy preserving decision tree evaluation in accordance with some embodiments is shown.
[0017] Fig.10 Shown is an example of edge labeling of three decision tree nodes with encrypted comparison results as part of performing a privacy-preserving decision tree evaluation in accordance with some embodiments.
[0018] Fig.11 Depicted is an m+1-choose-1 oblivious transfer process performed as part of performing a privacy-preserving decision tree evaluation in accordance with some embodiments.
[0019] Fig.12 A brief overview of an example protocol for performing privacy-preserving decision tree evaluation in accordance with some embodiments is shown.
[0020] Fig.13A , Fig. 13B , Fig. 13C and Fig.13D Depicted together are flow diagrams of example methods for performing privacy-preserving decision tree evaluation in accordance with some embodiments.
[0021] Fig.14Depicted is a flow diagram of an example method for performing privacy-preserving decision tree evaluation in accordance with some embodiments.
[0022] Fig.15 is a block diagram of an example computer system for implementing various embodiments.
[0023] In the drawings, like reference numbers generally refer to the same or similar elements. Additionally, generally, the left-most digit(s) of a reference number identifies the drawing in which the reference number first appears. DETAILED DESCRIPTION
[0024] 1. Introduction
[0025] As mentioned in the "Background" section above, security measures are becoming increasingly complex due to the emergence of different types of attacks, requiring multiple security technologies to effectively protect data and applications. One such technology is continuous authentication, which verifies the identity of a user throughout their interaction with the system by continuously monitoring and evaluating behavioral user data and biometric user data. However, creating classifiers for continuous authentication requires a large amount of user data, which is in most cases highly sensitive, resulting in its disclosure infringing privacy and security.
[0026] This article describes a protocol for privacy-preserving continuous authentication based on secure multi-party computation (MPC) that solves the aforementioned problems. The protocol allows multiple parties to compute joint results on their input data without disclosing their input data. In an embodiment, a three-party protocol is implemented using a decision tree learning algorithm as the underlying classifier, including a client with its user data and two servers responsible for computing the classification. Such an embodiment can adopt various MPC methods to compute classification results on encrypted inputs by interacting with each other, while protecting sensitive user data as well as trained classifiers.
[0027] As technology continues to develop, security measures become more and more complex. This is due to the emergence of different types of attackers and attacks. There are two main types of attackers: attackers from the inside and attackers from the outside. Attackers from the inside are usually individuals with legitimate access to the system or network, such as employees or contractors. They may intentionally or unintentionally cause damage to the system through malicious activities or through negligence. Attackers from the outside are individuals or groups who try to gain unauthorized access to the system or network from the outside. These attackers may use various strategies (such as phishing, social engineering, or brute force attacks) to gain access to the system. In order to effectively protect data and applications from these threats, a variety of security technologies are required. One of the security technologies is to authenticate the user to verify the true identity before granting access to the system or resources. This is usually done by applying an initial authentication mechanism (such as a password, multi-factor authentication, digital certificates, or even biometric data). However, in various situations, the initial authentication may not be enough. For example, a user may choose a weak password, forget to log out of the system, or even forget to lock his personal device, thereby allowing access to the adversary. Therefore, a mechanism for authenticating the user during the entire session may be considered preferred.
[0028] Individuals have unique behavioral patterns that can be identified and authenticated by analysis while using a device or service. For computers and mobile devices, keystroke dynamics, mouse dynamics, touch and slide patterns, phone orientation and gait recognition, and physical location and time zone are some of the factors that contribute to such patterns. Continuous authentication exploits these patterns and provides a novel security mechanism that verifies the identity of the user throughout the interaction with the system by continuously monitoring and evaluating behavioral user data and biometric user data to create a unique biometric data profile. The process typically involves obtaining behavioral user data, deriving suitable features from it, and creating a classification model for each user based on a machine learning algorithm. The trained classification model is then used to authenticate the user during the session, and depending on the corresponding results, appropriate measures are taken based on the criticality of the action being performed. Continuous authentication has a wide range of applications, ensuring authorized access to critical infrastructure in areas such as online banking and financial services, e-commerce transactions, medical applications that access sensitive health data, smart homes and mobile devices, and identity access management in enterprise applications.
[0029] Creating a classifier for continuous authentication assumes the application of machine learning, which requires a large amount of user data to train the classifier. However, behavioral user data and biometric user data are highly sensitive, and the disclosure of such data is a serious violation of privacy and security. Moreover, if such data falls into the wrong hands, it may be used to forge the identity of other users to bypass the security mechanisms of continuous authentication, thereby enabling the bad guys to infiltrate and cause damage to the system (e.g., by stealing, altering or destroying files and data, installing malware, rendering software or hardware inoperable, etc.). Similarly, the disclosure of trained classification models also puts the identity of users at risk. The sensitivity of machine learning models may come from two different sources. First, the trained model parameters themselves may contain sensitive information. Second, the model may have been created using sensitive data. It is well known that gaining access to machine learning models, whether through white-box or black-box access methods, can lead to model reversal attacks, thereby compromising the privacy of sensitive data. Therefore, it is crucial to find a way to classify the identity of users while protecting behavioral user data and biometric user data as well as the trained classifier.
[0030] To protect both user data and trained classifiers, certain embodiments described herein apply secure multi-party computation (MPC), a cryptographic concept that allows multiple parties to compute joint results on their input data without disclosing their input data to each other. The proposed implementation adopts a three-party protocol, including a client with user data that needs to be classified for authenticity and two servers responsible for computing the classification. For the classification process, the embodiment utilizes a decision tree learning algorithm as the underlying classifier. First, the client encrypts the data before the classification process to ensure its confidentiality. To protect sensitive user data used for model training, each server obtains a partial share of the pre-trained decision tree model that does not disclose the complete model. Then, the servers utilize various MPC methods (more precisely, additive homomorphic encryption and oblivious transfer) to interact with each other to compute the classification results on the encrypted input. Finally, the classification result is sent to the client, and depending on the result, appropriate actions are taken.
[0031] 2. Preliminary knowledge
[0032] Mathematical notations used in this disclosure are defined in Table 3 below.
[0033] 2.1 Continuous Certification
[0034] When interacting with a device or service, an individual has a unique behavior pattern that can be detected by various means. This pattern can be analyzed to help identify and authenticate the user. For computers and notebooks, keystroke dynamics and mouse dynamics are factors that contribute to unique behavior patterns. Keystroke dynamics involve analyzing an individual's typing speed, rhythm, and time between keystrokes to create a unique profile. Mouse dynamics similarly evaluates an individual's mouse movement and click patterns to create such a profile. For mobile devices, touch and slide patterns, phone orientation, and gait recognition are typical factors that contribute to unique behavior patterns. Touch and slide patterns and phone orientation are unique to each person and refer to the way a user interacts with his or her mobile device. Gait recognition involves analyzing the way an individual walks and moves, which can be detected using accelerometers and gyroscopes in mobile devices. In addition to these device-specific factors (presented as examples only, not limitations), location, time zone, and other factors can also be used as additional parameters to evaluate the credibility of the identity provided by the user.
[0035] This unique behavioral pattern can be used to implement an additional security mechanism known as continuous authentication, which continuously verifies the user's identity throughout their interaction with the system, not just at the initial login. To implement this security mechanism, behavioral and physiological biometrics are continuously monitored, tracked, and analyzed by a background process unknown to the user. Based on the collected data, a unique biometric data profile is created, which is then used to authenticate the user. Figure 1 A process 100 of an example biometric recognition system is depicted, including six stages that are executed as background processes during interaction with a user interface. The stages are data acquisition 104, feature extraction 106, model training 108, user registration 110, reasoning 114, and decision making 116. The model training stage and the user registration stage can only be performed once.
[0036] like Figure 1 As shown, process 100 begins with a user interacting with user interface 102 to obtain raw user data for creating and applying a data profile.
[0037] It will be clear to those skilled in the relevant art that there are a wide variety of behavioral data that such a system can collect. For example, it has been demonstrated that using mouse dynamics as the sole basis for continuous authentication has been sufficient to obtain accurate results. In order to form the basis of the continuous authentication concept, the Balabit Mouse Dynamics Challenge was launched to provide a behavioral dataset for cursor movement. Table 1 shows a summary of the raw mouse dynamics data from the Balabit Mouse Dynamics Challenge dataset:
[0038]
[0039] Table 1
[0040] In the next stage, the time steps are grouped into segments, each describing an action. These actions are then used to derive suitable features for characterizing the user. These features should preferably be universal, unique, invariant, measurable, difficult to spoof, and reasonable in the respective application or environment. The extracted features are then used to create a data profile, specifically a classification model based on a supervised machine learning algorithm. This user-specific classification model contains unique features specific to the user, which is finally stored in a dedicated database of user models 112 and will be used for authentication during the session.
[0041] After user registration 110, the system can apply continuous authentication during future sessions. The trained classification model is used to perform reasoning 114 on the monitored data. This allows the classifier to make decisions 116 based on the current user's behavior, determining whether it corresponds to a valid user or an impostor. Depending on the criticality of the action being performed, appropriate action is taken, such as requesting another login or taking additional security measures.
[0042] 2.2 Decision Tree
[0043] There are various supervised learning based classifiers. Possible examples in the context of continuous authentication are k-nearest neighbors, decision trees, random forests or even neural networks. In the embodiments described herein, a decision tree (DT) (also called a regression tree) is used as the underlying classification algorithm. A decision tree is a supervised learning algorithm for classification tasks. It has a hierarchical binary tree structure, which consists of internal or decision nodes, leaf nodes and edges connecting different nodes to each other. Each internal node is marked with a test condition and each leaf node is marked with a classification label. The test condition is a comparison of a given threshold with the value of a feature attribute of the input to be classified. Applying a decision tree algorithm (also called an inference process) means traversing the path within the tree based on a given input and reaching one of the possible classification labels of the corresponding leaf node.
[0044] In the following description, the structure of the DT model is introduced, including its various components and the overall application of the model. The decision tree is composed of a binary tree as the underlying data structure. Binary trees appear in many types, such as full or true binary trees, perfect binary trees, complete binary trees, balanced binary trees, etc.
[0045] In the embodiments described herein, the binary tree is represented as a full binary tree, where each tree node has two or zero child nodes. A full binary tree is a mathematical graph G = (V, E) with a hierarchical structure consisting of nodes, edges, and paths.
[0046] In particular, a full binary tree consists of 2m + 1 nodes (v 1 , …, v 2m+1 ) ∈ V, where there are m decision nodes and m + 1 leaf nodes (l 1 ,,…,l m+1 ) , so that Node indices are assigned via a breadth-first search (BFS) traversal.
[0047] A full binary tree also consists of a set of (e 1 ,…,e 2m )∈E refers to the 2m connecting edges, that is, for each e i ∈E, there exists a pair of nodes , so that This gives G a coherent structure. Each decision node v i = d j Must be connected to two child nodes, the first of which is called the left child node v 2i , the second child node is called the right child node v 2i+1 Except for the root node v 1 = d 1 In addition, each node has a link to its parent node The connections to child nodes are called left and right edges, and the connection to the parent node is called the parent edge. A leaf node has only one connection to its parent node and does not have any child nodes.
[0048] A full binary tree also consists of m + 1 paths, where each path starts from the root node v 1 Starts and leads to leaf node l k The sequence P k = (d 1 , …, d j , l k ). Each path additionally has a depth δ defined as the number of nodes in the path k = |P k |. The depth δ of a tree is defined as the depth of the longest path in the tree, i.e., .
[0049] exist Figure 2 , an example of a full binary tree 200 with m = 4 decision nodes and a tree depth of δ = 4 is shown.
[0050] To arrive at the definition of the DT model utilized by the embodiments described herein, the foregoing definition of a full binary tree may be extended with a few more properties.
[0051] In particular, a decision tree (DT) or decision tree model can be defined as mapping an n-dimensional attribute vector x to a classification label c using a given binary tree G. k ∈ {c 1 , …, c m+1} multivariable function: n → {0, 1}. In addition to these components of the binary tree G, the DT model also includes:
[0052] m thresholds т = (t 1 ,…,t m ), each of the thresholds is calculated using the allocation function THR: D → The allocation function assigned to a decision node adopts the decision node d j And return the corresponding threshold t j ,
[0053] m attribute indexes , each of the attribute indices is assigned to a decision node using another assignment function ATT∶D → [1, n], which takes the decision node d j and returns the attribute index i, and
[0054] m + 1 classification labels c = (c 1 , …, c m+1 ), each of the classification labels is obtained by labeling function LAB: L→ {c 1 , …, c m+1} and assigned to a leaf node, the labeling function uses the given leaf node l k And return the corresponding classification label c k .
[0055] Given input data, the DT model can be evaluated. That is, inference can be run on the attributes or feature vectors to retrieve the classification of the input. While the functions THR and ATT are required to establish the test conditions and traverse G correctly, LAB is used to finally retrieve the corresponding classification results.
[0056] Next, we explain the application of the DT model as a function for classifying arbitrary input data (also known as decision tree evaluation).
[0057] Given a DT model for DT evaluation, a bivariate comparison function is defined : 2 → {0,1}, which takes the input attribute x i ∈ {x 1 , …, x n} and threshold t j ∈ {t 1 ,…,t m}, and evaluates the greater-or-equal (GEQ) test condition [x i ≥ t j ] ∈ {0, 1}.
[0058] To run inference on a given input x and a DT model, the process starts from the root node v 1 Start. For each node v reached j , the process evaluates To obtain the decision bit b. If b j = 0, the process moves to the left child; or if b j = 1, the process moves to the right child node. Repeat this process until the leaf node l is reached. k , so that the leaf node LAB(l k ) is taken as the evaluation result. The evaluation result is expressed as (x).
[0059] Figure 3 An example of decision tree evaluation 300 using arbitrary data is shown, where given attribute x = (2, 3,0), threshold T = (1, 1, 4, 7), classification labels C = (0, 1, 0, 1, 0), and attribute assignments I = (1, 3, 2, 3). First, the process will and t 1 is compared to the GEQ test condition, which evaluates to true. Therefore, the process moves to the right child and With t 3 The comparison is made, which evaluates to false. Therefore, the process moves to the left child node (i.e., the leaf node) and retrieves the assigned classification label LAB(l 2 ) = 1. As a result, the process traverses the path P 2 , so that the input feature x is classified as label 1 and is therefore evaluated as valid. The entire process is represented as (x) = 1.
[0060] 2.3 Secure Multi-Party Computation
[0061] Both the input data and the trained parameters of the DT classifier contain behavioral user data. Since behavioral data is highly sensitive and the disclosure of such data is a serious violation of privacy and security, privacy-preserving machine learning techniques can be applied as described herein. In particular, secure multi-party computation (MPC) can be applied as described herein, allowing the owners of the model and the data to perform reasoning separately without either party knowing additional information other than their own. MPC techniques for privacy-preserving reasoning using the DT described herein include additive homomorphic encryption (AHE) and oblivious transfer (OT).
[0062] Additive Homomorphic Encryption
[0063] AHE is an encryption scheme that allows arithmetic addition and scalar multiplication to be performed on ciphertext to produce an encrypted result that is identical to the result of performing these operations on the plaintext.
[0064] An encryption scheme can be defined as a collection of three algorithms:
[0065] • pk,sk ← The probabilistic algorithm uses the security parameter , and outputs a key pair consisting of a public key pk and a private key sk.
[0066] • c ← This probabilistic algorithm takes a public key pk and a plaintext message m as input, and outputs a ciphertext c. Here, Will be used as The abbreviation of .
[0067] • m′ ← This deterministic algorithm takes sk and a ciphertext c, and outputs a plaintext message m', ie, the decryption of the encrypted message m.
[0068] If the public key and the private key are equal, i.e., pk = sk, then the encryption scheme can be called symmetric encryption or secret key encryption. Otherwise, it can be called asymmetric encryption or public key encryption.
[0069] An encryption scheme whose ciphertext is homomorphic with respect to addition is called an additively homomorphic encryption scheme. More precisely, if we have an encryption function and two plaintext values m 1 and m 2 , then adding two values and then encrypting their sum will lead to the same result as first encrypting the two values using the AHE scheme and then adding their ciphertexts. Therefore, in addition to the above algorithms associated with the definition of the encryption scheme, the AHE scheme also provides additional operations, as described below.
[0070] An additive homomorphic encryption scheme may be an encryption scheme that additionally provides the following methods:
[0071] • c 3 ← This operation uses two ciphertexts c 1 = and c 2 = , and perform homomorphic addition using the public key pk to calculate the encrypted result c 3 Due to the homomorphic nature of the scheme, the addition of ciphertexts produces the same output as adding two corresponding plaintexts (m 1 and m 2 ) and then encrypting the resulting sum. As a shorthand notation, the following notation for homomorphic addition of ciphertext is introduced:
[0072]
[0073] • c 3 ← This method uses the ciphertext c 1 = and the plaintext value m 2 , and perform arithmetic addition using pk to calculate the encrypted result c 3 . The same operator can be used for homomorphic addition of ciphertext and plaintext values:
[0074]
[0075] • c 3 ← This method uses the ciphertext c 1 = and the plaintext value m 2 , and perform homomorphic scalar multiplication using pk to compute the encrypted result c 3 As a shorthand notation, we introduce the following notation for homomorphic scalar multiplication:
[0076]
[0077] • c 3 ← This method uses the ciphertext c = , and uses pk to return the ciphertext encrypted with the additive inverse of m:
[0078]
[0079] • c ← This operation uses encrypted bits. and plaintext bits b, and use pk to execute the binary (XOR) operation, where c is The encrypted result bits of the operation. This is done by and This is accomplished by combining methods:
[0080]
[0081] We can use the ⊕ operator to describe the operation, and can be used Operators are used to describe the relationship between ciphertext and plaintext values. Operation.
[0082] In an embodiment, for all plaintext m 1 、m 2 For plaintext bits a and b, IND-CPA security and the following correctness conditions are required:
[0083] (1)
[0084] (2)
[0085] (3)
[0086] (4)
[0087] (5)
[0088] Common examples of AHE schemes known to those skilled in the relevant art are the Paillier scheme or the lifted Elgamal with elliptic curve (ECLE) scheme. However, such AHE schemes are provided herein as examples and are not intended to be limiting.
[0089] Oblivious transfer
[0090] Oblivious transfer (OT) is a cryptographic protocol that enables a sender to transmit one of several fragments of a data element to a receiver without revealing which fragment was chosen. There are different variants of OT, however, the basic variant is composed of Indicates 2-choice-1 OT. Figure 4 The operation mode 400 of 2-to-1 OT is shown.
[0091] In OT In the protocol, the sender 406 holds two values of bit length μ in pairs , the receiver 402 is interested in obtaining the value at index i ∈ {1, 2}. As a result of the OT protocol 404, the receiver 402 only learns that without receiving any information about other values. , while the sender 406 does not learn any information about i, that is, which element the receiver 402 requested.
[0092] OT Can be extended to n-choose-1 OT protocol (OT ), where the sender holds a set of n values , the receiver is interested in obtaining the value at index i ∈ {1,…,n}. OT It can be done by executing log n times OT operations or use more advanced techniques to implement them. Although OT generally applies public key cryptography, there is a variant based on symmetric key operations called OT extension (OTX). More precisely, OTX splits the transport protocol into an offline phase and an online phase.
[0093] The offline component relies on asymmetric cryptography to generate a finite set of oblivious transfers known as base or seed OTs. These base OTs can then be expanded from the online phase to an unlimited number of oblivious transfers done only through operations involving symmetric cryptography. This can be considered the OT equivalent of hybrid encryption.
[0094] 3. Privacy-preserving decision tree evaluation
[0095] Figure 5 is a block diagram of an example system 500 that performs privacy-preserving decision tree evaluation, according to an embodiment. Figure 5 The example system 500 implements a three-party MPC protocol. In particular, the example system 500 includes a client 502 that wants to classify its input and two servers 504 and 506 that calculate the classification. Each of the client 502, the server 504, and the server 506 can be a corresponding computing device or computer system (such as the following reference Fig.15 The computer system 1500 described above may be implemented in a corresponding computing device or computer system (such as the computer system 1500 described below). Fig.15 In addition, each of the client 502, the server 504, and the server 506 may be implemented by a corresponding physical machine or a virtual machine, or implemented on a corresponding physical machine or a virtual machine.
[0096] The input of client 502 is encrypted using AHE and remains hidden from both server 504 and server 506 during the overall inference process. Both server 504 and server 506 receive a portion of the previously trained DT model 512, so that neither of them knows the structure and thresholds of the complete DT model. In particular, server 504 receives DT share 508 and server 506 receives DT share 510. Both servers then communicate with each other to calculate the classification result for the encrypted input using their model shares and MPC techniques (i.e., AHE and OT). Finally, the classification result is sent to client 502.
[0097] In the following, each step of the protocol is explained in detail. In addition, in the following, client 502 is referred to as "client", server 504 is referred to as "server 1", and server 505 is referred to as "server 2".
[0098] 3.1 Initialization
[0099] First, each server creates its own key pair (pk i , sk i ), where i represents the corresponding server index. The client can access the public key pk 1 and pk 2 Both, and use pk 2 Encrypt its input. Let x = (x 1 , …,x n ) is the client input consisting of n feature attributes, and let each attribute x i ∈ is encoded as a tuple of μ bits. Then the client input is defined as
[0100]
[0101] The rows of the matrix represent individual attributes x. i ∈ x, the column represents the attribute x i The bit encoding of the individual bit x at index j within ij .
[0102] Assume that a trained decision tree model is given, i.e., a binary tree consisting of decision nodes and leaf nodes, the classification label of each leaf node, and the threshold of each decision node determined during the training phase of the DT model. Then, the DT is divided into two separate shares sent by the management entity to each server, i.e. Figure 6 Process 600 is depicted in FIG.
[0103] The first DT share delivered to server 1 consists of two elements. The first element is a tuple of m thresholds, where each threshold is encrypted bit by bit with a given bit length µ, i.e.,
[0104]
[0105] The second element is a tuple of attribute indices represented as follows
[0106]
[0107] Given x and T, the protocol uses I to define the relationship between their components such that for each threshold t j Through GEQ test conditions [ ≥ t j ] and the input attribute Related, the GEQ test condition is applied at each internal node d during decision tree evaluation. j Execute at.
[0108] As mentioned above, in the embodiment, the underlying binary tree is a full tree, and accordingly not necessarily a complete binary tree. Therefore, the given tuple of encryption thresholds and attribute indices does not indicate the actual alignment between nodes. In addition, all thresholds are in the public key pk of server 2. 2 The first DT share is encrypted and thus indecipherable to the server 1. Therefore, the first DT share does not reveal any sensitive information about the trained DT model, and thus does not reveal any user data that can be inferred.
[0109] The second share sent to server 2 includes the description of the tree structure given by G and the 1 Encrypted classification label tuple. The tree structure can be represented as a set of all internal nodes and leaf nodes including node edges but without associated thresholds and classification labels. Since the tree structure itself is not sufficient to determine sensitive information of the trained model, the second share similarly does not reveal any sensitive user data, and the actual DT model remains hidden from both servers. Figure 6 An overview of the DT sharing process including the components of each share is given in and Table 2. In addition, it is assumed that both servers are under separate management so that no supplementary information is shared except the information explicitly outlined in the protocol.
[0110] Table 2 below provides an overview of the share of decision tree models, which consist of a binary tree structure G, a threshold T, an input feature assignment I, and a classification label C:
[0111]
[0112] Table 2
[0113] 3.2 Comparative Calculation
[0114] As mentioned earlier, the client's input is encrypted before being sent to Server 1 through the network channel. In order to make it difficult for Server 1 to decipher the input, pk 2 , so that the encrypted input is defined as
[0115]
[0116] The input vector is encrypted bit by bit, i.e., for n features with μ bits per feature, there are a total of n · µ ciphertexts contained in middle.
[0117] Server 1 receives , and compute several integer comparisons of the values encrypted using AHE. For example, according to the techniques described in Tueno et al., "A Method for Securely Comparing Integers Using Binary Trees," Cryptography ePrintArchive (2021), ("Tueno et al.", which is incorporated herein by reference in its entirety), the encrypted input attribute The corresponding encryption threshold For comparison, as shown in Algorithm 1:
[0118]
[0119] At each decision node, compare the results b j A Boolean value that is defined as a greater than or equal comparison, that is, b := [ ≥ t j ], so that if ≥ t j , then b j is evaluated to 1, otherwise, b j is evaluated to 0.
[0120] Figure 7 The applied method of Tueno et al. for comparing two integers x each consisting of a tuple of μ bits is shown. (1) and x (2) An overview of the integer comparison protocol 700. Figure 5 In the context of a decision tree evaluation performed by a system, x (1) Represents the corresponding input attribute , x (2) Indicates the corresponding threshold t j .
[0121] To apply integer comparison of two integers between two parties, the first party creates a key pair (sk, pk) and shares its public key pk with the second party. Both parties have integers consisting of tuples of μ bits and want to know the result of the comparison without disclosing the actual values, i.e., whether x (1) ≥ x (2) . In its encrypted input After the transmission of , the second party performs an integer comparison on the data encrypted using AHE. The detailed comparison process is depicted as a black box for simplicity and can be found in Tueno et al. Essentially, inside the black box, the comparison result a is calculated and split into two parts that satisfy b. (1) ⊕ b (2) = two shares of b (1) and b (2) However, the interior of the black box is not accessible to the second party, but as a result of the black box functionality, only these shares are revealed, thereby affecting the first share b (1) Finally, the encrypted first share is passed to the first party so that both parties now hold a share of the result of their comparison. The individual shares do not reveal the actual value, but only the combination of the two values does.
[0122] exist Figure 5 In the context of a system with The first party is represented by the client and the second party is represented by server 1. As shown in Algorithm 1, the comparison result of each decision node is contained in в ∈ {0, 1} m Each component at index j is split into its shares b and b , thus generating two share tuples в (1) and (2) , where the superscript index indicates the subordinate relationship with the corresponding server. After executing the functions contained in the black box, server 1 obtains its share as plaintext and the second share as ciphertext. It encrypts the first share using its public key and passes both shares to server 2, which decrypts its share and uses the component-wise homomorphism of the two shares to encrypt the first share and the second share. Operation to derive the encrypted comparison result ,Right now,
[0123] For each j∈[1,m] , b j (1) pk1 ⊗ b j 2 = b j pk1
[0124] Figure 8 An application 800 of the secure integer comparison protocol of Tueno et al. for component-wise comparison of an attribute vector x and a threshold vector T is shown.
[0125] The operations performed by Server 1 in this section are summarized as follows
[0126]
[0127] in, is a tuple of comparison bit shares for server 1, is a tuple of Server 2’s encrypted comparison bit shares. Both shares are sent to Server 2, which compares its share tuples to (2) Decryption is performed and component-wise homomorphism is used Operation to derive the overall comparison result ,Right now,
[0128]
[0129]
[0130] However, the actual comparison result is encrypted and indecipherable to the server 2 so that sensitive information is not revealed. In summary, the process of deriving the overall comparison result performed by the server 2 is represented as
[0131]
[0132] 3.3 Edge marking
[0133] Now, as Server 2 has obtained the encrypted comparison result described by tuple B, the decision tree is applied to evaluate the client input in order to determine the corresponding classification label. This is done by traversing the tree and assigning the corresponding components of vector B to the tree edges. Fig.9A and Fig. 9B Outlines the operations performed at each node during the traversal of the tree.
[0134] In particular, Fig.9A and Fig. 9B Depicts the evaluation of decision nodes and the assignment of comparison results to corresponding edges to determine the traversal direction. Note that all values are represented by pk 1 encrypted and therefore illegible to the server 2. Furthermore, the server 2 receives only the result of the calculation, but neither the input x nor the threshold T shown for clarity.
[0135] The process starts at the first internal node (i.e., the root node). If the assigned comparison evaluates to true, the tree is traversed to the right child, otherwise to the left child, and the process is repeated until a leaf node is reached. Since the actual traversed path depends on the input x, the process uses the comparison result B to mark the traversed edge. Further according to the process, the value of the traversed edge is set to 0, otherwise it is set to 1, the reason will be explained in the subsequent chapters. Fig.9A In example 900, we see that the comparison is evaluated to true, so that the tree needs to be traversed to the right child. Therefore, the process is performed by flipping the result bit b j The right output edge is set to 0, and the left output edge is set to 1, that is, b j .exist Fig. 9B In the opposite example 902, where x i < t j , the process performs the same operation, so that the left edge is set to the comparison result bit 0, and the right edge is set to the inverted result 1. This process can be called edge marking.
[0136] The edge labeling process can be defined as follows. Each decision node v j Comparison operations with integers Therefore, for each v j , which processes the decision bits called comparison results Calculate and use Mark the left output edge with Label the right outgoing edge.
[0137] Fig.10 The edge labels of the three decision nodes of the tree 1000 are shown in the case of m = 3, wherein the encrypted comparison result (using pk 1 As mentioned before, it can be observed that server 2 simply assigns each left outgoing edge to the comparison result b j , and assign each right-hand outgoing edge to its inverse ¬ .
[0138] The functionality introduced in this section can be expressed as:
[0139]
[0140] Using decision tree model and the comparison result share as input, and output the updated binary tree G′ of the DT model, where each tuple Through the third value b j (respectively ) to expand.
[0141] 3.4 Tree Evaluation
[0142] After the edge labeling process, Server 2 must classify the label c along the communication appropriately k The tree is navigated by a path where all edges are marked with zeros, because the path represents a traversal route for the provided input x. This path will be referred to as the zero path. However, since b and c are encrypted, each path must be evaluated to identify the zero path. The evaluation of the path can be referred to as path aggregation.
[0143] Path aggregation can be defined as follows. For the path P defined above k , where the edges are labeled as described in the previous section, and the path aggregation can be defined as the path along the path leading to the corresponding leaf node c k The resulting value is called the associated path cost p k .
[0144] The path cost of the entire tree is described by the tuple p of all individual path costs,
[0145]
[0146] Among them, each component Along the path P k The sum of the edge labels of i is with P k The index of the decision bit associated with the ith edge in . Fig.10 In the example, p is evaluated as
[0147]
[0148] and depending on the client input, possible examples of p are
[0149]
[0150] In the case of m = 3 decision nodes, there are m + 1 = 4 different paths, of which only one sum is evaluated as 0 + 0 = 0, one sum is evaluated as 1 + 1 = 2, and the others are evaluated as 0 + 1 = 1 + 0 = 1. In the current state, if p is sent to server 1 for classification determination in the current state, p still reveals the tree structure that allows the reconstruction of the main part of the tree. Therefore, after permuting p by using the random permutation function π, each component is multiplied by a separate random scalar r k , thus generating
[0151]
[0152] The permutation function π defines a fixed permutation, or takes an index i ∈ and returns the new index (i) after the permutation, or takes an arbitrary tuple of elements N and returns the permuted tuple π(n).
[0153] Further according to the above example, the conversion function is set to
[0154]
[0155] Make The component is evaluated as
[0156]
[0157]
[0158]
[0159]
[0160] Server 2 will The encrypted, permuted, and randomized path evaluations in are sent to Server 1, which does not learn any information about the tree structure or the computational evaluations except that there is a zero path at index i (i = 1 in this example).
[0161] In summary, the process aggregates all paths of the updated binary tree G′, thereby producing a tuple of associated costs for each path, where each component is individually randomized. The process then selects a random permutation π to permute the path cost tuple. Finally, the resulting tuple is sent to server 1, and the applied permutation function is retained for later operations. The overall process can be expressed as
[0162]
[0163] 3.5 Classification determination
[0164] The final phase of the agreement includes First, according to the previous example, server 1 The components of are decrypted and the zero path is identified, that is, where the index , where j = π(i) = 1. Fig.11 As depicted, both servers then participate in Fig.11 OT 1102 Server 1 assumes the role of the receiver now holding the index π(i), while server 2 assumes the role of the sender with the encrypted bit list (i.e., the permuted classification label tuple as follows,
[0165]
[0166] This is based on the previous example ( As a result, server 1 obtains the requested classification label , and use sk 1 It is decrypted without any other information about the sender's data. Likewise, the server 2 will not learn any information about the receiver's data.
[0167] If c (i) = 1, the client is considered valid, otherwise, c π(i) = 0 and the client is invalid. This information is communicated to the client accordingly and appropriate action is taken.
[0168] 4. Summary
[0169] Table 3, shown below, provides a list of notations, variables, and symbols used herein to describe an example protocol for performing privacy-preserving decision tree evaluation:
[0170]
[0171]
[0172]
[0173] Table 3
[0174] also, Fig.12 A brief overview 1200 of an example protocol implementation is provided, including the five general phases of the protocol, the three parties involved, and the communications between the parties.
[0175] Now refer to Fig.13A , Fig. 13B , Fig. 13C and Fig.13D To describe the Fig.12 In particular, Fig.13A , Fig. 13B , Fig. 13C and Fig.13D 1300 for evaluating a decision tree in a privacy-preserving manner according to some embodiments. The method 1300 may be performed by processing logic that may include hardware (e.g., circuits, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executed on a processing device), or a combination thereof. It should be understood that not all steps may be required to perform the disclosure provided herein. Furthermore, as will be appreciated by one of ordinary skill in the art, some of the steps may be performed simultaneously or in parallel. Fig.13A, Fig. 13B , Fig. 13C and Fig.13D The steps are performed in different orders as shown.
[0176] Reference Figure 5 The method 1300 is described with reference to a system of FIG. 1301. However, the method 1300 is not limited to this example embodiment.
[0177] During (or before) the input phase, in 1302, server 1 is provided with: (a) a first public key-private key pair consisting of a first public key and a first private key; (b) a second public key, the second public key comprising a portion of a second public key-private key pair provided to server 2; and (c) a first portion of a share of a decision tree classifier, the first portion of the decision tree classifier consisting of: (i) threshold tuples respectively associated with decision tree nodes of the decision tree classifier, wherein each threshold in the threshold tuple is bit-encrypted (using the AHE scheme) using the second public key; and (ii) attribute index tuples respectively associated with decision tree nodes of the decision tree classifier, each attribute index in the attribute index tuple indicating an input data attribute associated with the corresponding decision tree node.
[0178] During (or before) the input phase, in 1304, the server 2 is provided with: (a) a second public-private key pair consisting of a second public key and a second private key; (b) the first public key; and (c) a second part of a decision tree classifier, the second part of the decision tree classifier consisting of: (i) a description of the tree structure of the decision tree classifier; and (ii) classification label tuples respectively associated with leaf nodes of the decision tree classifier, wherein each classification label is encrypted using the first public key (using the AHE scheme).
[0179] During (or before) the input phase, in 1306, the first public key and the second public key are provided to the client.
[0180] During the input phase, in 1308, the client obtains input data including a set of attributes (e.g., user attributes), and encrypts each attribute of the input data bit by bit (using the AHE scheme) using the second public key, thereby generating encrypted input data. During the input phase, in 1310, the client also sends the encrypted input data to the server 1 (which cannot be decrypted by the server 1 because it is encrypted with the second public key).
[0181] During the comparison calculation phase, in 1312, server 1 uses AHE to compare the encrypted input data with the encrypted threshold to generate a comparison result for each decision node of the decision tree classifier (which means that the encrypted input data and the encrypted threshold are never decrypted into plaintext), where each comparison result is represented by a first comparison result share in plaintext and a second comparison result share encrypted with a second public key (so that it cannot be decrypted by server 1).
[0182] In 1314 , Server 1 then encrypts each of the first comparison result shares using the first public key (using AHE).
[0183] In 1316 , Server 1 then sends the comparison results to Server 2 , wherein each sent comparison result is represented by a first comparison result share encrypted with the first public key (such that it cannot be decrypted by Server 2 ) and a second comparison result share encrypted with the second public key.
[0184] In 1318, for each comparison result, server 2 decrypts the second comparison result share using the second private key, and then performs homomorphic encryption on the first comparison result share encrypted with the first public key and the second comparison result decrypted into plain text. Operate to generate a comparison result encrypted with the first public key (such that it cannot be decrypted by server 2), wherein each comparison result is 1 if the comparison is evaluated to true, or each comparison result is 0 if the comparison is evaluated to false.
[0185] During the edge marking phase, in 1320, for each decision node in the decision tree classifier, server 2 marks the first output edge of the decision node (the edge corresponding to false) using the corresponding comparison result encrypted with the first public key, and marks the second output edge of the decision node (the edge corresponding to true) using the inverse of the corresponding comparison result encrypted with the first public key.
[0186] During the tree evaluation phase, in 1322, for each path from the root node of the decision tree to the corresponding leaf node, server 2 calculates a path cost representing the sum of the edge labels along the path, thereby generating a path cost tuple, each path cost tuple is encrypted with the first public key (so that it cannot be decrypted by server 2), where the path leading to the classification result is the only path with zero path cost (zero path).
[0187] In 1324, server 2 multiplies each path cost in the path cost tuple by the corresponding random scalar value, and then applies a permutation function to rearrange the order of the path costs in the path cost tuple, thereby generating a permuted and randomized path cost tuple encrypted with the first public key. In 1326, server 2 also uses the same permutation function to permute the classification label tuple encrypted with the first public key.
[0188] In 1328, Server 2 then sends the permuted and randomized path cost tuple encrypted with the first public key to Server 1.
[0189] In the classification determination phase, in 1330, server 1 decrypts the permuted and randomized path cost tuple using the first private key to identify the index associated with the zero path.
[0190] In 1332, Server 1 and Server 2 then engage in an oblivious transfer operation in which Server 1 obtains from Server 2 a classification label associated with the index associated with the zero path and encrypted with a first public key without Server 2 knowing the index associated with the zero path and without Server 1 knowing any classification labels associated with any other path of the decision tree classifier.
[0191] In 1334, Server 1 decrypts the classification label associated with the index associated with the zero path using the first private key.
[0192] In 1336, Server 1 then sends the classification label associated with the index associated with the zero path to the client.
[0193] Fig.14 is a flow chart of a method 1400 for evaluating a decision tree in a privacy-preserving manner according to some embodiments. The method 1400 may be performed by processing logic that may include hardware (e.g., circuits, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executed on a processing device), or a combination thereof. It should be understood that not all steps may be required to perform the disclosure provided herein. Furthermore, as will be appreciated by one of ordinary skill in the art, some of the steps may be performed simultaneously or in parallel. Fig.14 Execute in different orders as shown
[0194] Reference Figure 5 However, the method 1400 is not limited to this example embodiment.
[0195] In 1402 , server 1 receives a first portion of a decision tree and encrypted input data from a client, wherein the encrypted input data includes a set of attributes.
[0196] At 1404, server 2 receives a second portion of the share of the decision tree.
[0197] In 1406, each of Server 1 and Server 2 communicates with the other server to calculate a classification result for the encrypted input data using the corresponding partial shares of the decision tree received thereby and a secure multi-party computing method, the secure multi-party computing method including additive homomorphic encryption and oblivious transfer, so that the classification result is determined without decrypting the encrypted input data, and the classification result is ultimately received by Server 1.
[0198] In 1408, the server 1 sends the classification result to the client.
[0199] In an embodiment of the method 1400, the decision tree comprises a full binary decision tree.
[0200] In another embodiment of method 1400, the decision tree includes multiple decision nodes and multiple leaf nodes, the first portion of the decision tree includes threshold tuples respectively associated with the decision nodes of the decision tree and attribute index tuples respectively associated with the decision nodes of the decision tree, and the second portion of the decision tree includes a structural description of the decision tree and classification label tuples respectively associated with the leaf nodes of the decision tree.
[0201] Further in accordance with such an embodiment, the classification label tuple can be encrypted with a first public key, which is paired with a first private key accessible to the first server but not accessible to the second server, and the threshold tuple can be encrypted with a second public key, which is paired with a second private key accessible to the second server but not accessible to the first server.
[0202] In another embodiment of method 1400, a set of attributes is associated with a user, the decision tree comprises a classification model trained on one or more previously obtained sets of attributes associated with the user, and the classification result indicates whether the user should be authenticated.
[0203] In further accordance with such embodiments, the set of attributes associated with the user may include one or more of behavioral attributes associated with the user or biometric attributes associated with the user.
[0204] Further in accordance with such embodiments, method 1400 may be performed as part of an ongoing authentication process for determining whether a user should have ongoing access to a system or resource.
[0205] In another embodiment of method 1400, the first server and the second server communicate directly with each other. In an alternative embodiment, the first server and the second server communicate indirectly with each other via one or more intermediaries (eg, intermediate physical or virtual machines, intermediate computing devices, etc.).
[0206] For example, various embodiments may use one or more well-known computer systems (such as Fig.15For example, one or more computer systems 1500 may be used to implement any of the embodiments discussed herein, as well as combinations and sub-combinations thereof.
[0207] Computer system 1500 may include one or more processors (also referred to as central processing units or CPUs), such as processor 1504. Processor 1504 may be connected to a communication infrastructure or bus 1506.
[0208] The computer system 1500 may also include user input / output devices 1503 , such as a monitor, a keyboard, a pointing device, etc. The user input / output devices 1503 may communicate with the communication infrastructure 1506 through the user input / output interface 1502 .
[0209] One or more processors 1504 may be a graphics processing unit (GPU). In an embodiment, a GPU may be a processor that is a specialized electronic circuit designed to process mathematically intensive applications. A GPU may have a parallel structure that is effective for processing large blocks of data in parallel, such as mathematically intensive data common in computer graphics applications, images, videos, etc.
[0210] The computer system 1500 may also include a main memory or primary storage 1508, such as random access memory (RAM). The main memory 1508 may include one or more levels of cache. The main memory 1508 may store control logic (ie, computer software) and / or data therein.
[0211] The computer system 1500 may also include one or more secondary storage devices or secondary memories 1510. The secondary memories 1510 may include, for example, a hard disk drive 1512 and / or a removable storage device or drive 1514. The removable storage drive 1514 may be a floppy disk drive, a tape drive, an optical drive, an optical storage device, a tape backup device, and / or any other storage device / drive.
[0212] The removable storage drive 1514 can interact with the removable storage unit 1518. The removable storage unit 1518 may include a computer usable or readable storage device having computer software (control logic) and / or data stored thereon. The removable storage unit 1518 may be a floppy disk, a magnetic tape, an optical disk, a DVD, an optical storage disk, and / or any other computer data storage device. The removable storage drive 1514 can read from and / or write to the removable storage unit 1518.
[0213] Secondary memory 1510 may include other components, devices, assemblies, tools, or other methods for allowing computer programs and / or other instructions and / or data to be accessed by computer system 1500. Such components, devices, assemblies, tools, or other methods may include, for example, a removable storage unit 1522 and an interface 1520. Examples of removable storage unit 1522 and interface 1520 may include a program cartridge and a cartridge interface (such as found in video game devices), a removable memory chip (such as an EPROM or PROM) and an associated socket, a memory stick and USB port, a memory card and an associated memory card slot, and / or any other removable storage unit and associated interface.
[0214] The computer system 1500 may also include a communication or network interface 1524. The communication interface 1524 may enable the computer system 1500 to communicate and interact with any combination of external devices, external networks, external entities, etc. (referenced individually and collectively by reference numeral 1528). For example, the communication interface 1524 may allow the computer system 1500 to communicate with an external or remote device 1528 via a communication path 1526, which may be wired and / or wireless (or a combination thereof) and may include any combination of a LAN, a WAN, the Internet, etc. Control logic and / or data may be sent to / from the computer system 1500 via the communication path 1526.
[0215] Computer system 1500 may also be any one of a personal digital assistant (PDA), a desktop workstation, a laptop or notebook computer, a netbook, a tablet computer, a smartphone, a smart watch or other wearable device, an appliance, part of the Internet of Things, and / or an embedded system, to name a few non-limiting examples, or any combination thereof.
[0216] Computer system 1500 can be a client or server accessing or hosting any application and / or data through any delivery paradigm, including but not limited to: remote or distributed cloud computing solutions; local or on-premises software ("on-premises" cloud-based solutions); "as a service" models (e.g., content as a service (CaaS), digital content as a service (DCaaS), software as a service (SaaS), managed software as a service (MSaaS), platform as a service (PaaS), desktop as a service (DaaS), framework as a service (FaaS), backend as a service (BaaS), mobile backend as a service (MBaaS), infrastructure as a service (IaaS), etc.); and / or a hybrid model including any combination of the foregoing examples or other services or delivery paradigms.
[0217] Any applicable data structures, file formats, and modalities in computer system 1500 may be derived from standards including, but not limited to, JavaScript Object Notation (JSON), Extensible Markup Language (XML), Yet Another Markup Language (YAML), Extensible Hypertext Markup Language (XHTML), Wireless Markup Language (WML), MessagePack, XML User Interface Language (XUL), or any other functionally similar representations, either alone or in combination. Alternatively, proprietary data structures, formats, or modalities may be used, either exclusively or in combination with known or open standards.
[0218] In some embodiments, a tangible non-transitory apparatus or article of manufacture including a tangible non-transitory computer-usable or readable medium having control logic (software) stored thereon may also be referred to herein as a computer program product or program storage device. This includes, but is not limited to, computer system 1500, main memory 1508, secondary memory 1510, removable storage units 1518 and 1522, and tangible articles of manufacture embodying any combination of the foregoing. When executed by one or more data processing devices (such as computer system 1500), such control logic may cause such data processing devices to operate as described herein.
[0219] Based on the teachings contained in this disclosure, it is obvious to a person skilled in the relevant art how to use the Fig.15 It will be clear that data processing devices, computer systems, and / or computer architectures other than those shown can be used to make and use embodiments of the present disclosure. In particular, the embodiments can be operated with software, hardware, and / or operating system implementations other than those described herein.
[0220] It should be understood that the "Detailed Description" section (but not any other section) is intended to be used to interpret the claims. Other sections may set forth one or more but not all exemplary embodiments contemplated by the inventors, and are therefore not intended to limit the present disclosure or the appended claims in any way.
[0221] Although the present disclosure describes exemplary embodiments of exemplary fields and applications, it should be understood that the present disclosure is not limited thereto. Other embodiments and modifications thereto are possible and within the scope and spirit of the present disclosure. For example, without limiting the generality of this paragraph, the embodiments are not limited to the software, hardware, firmware, and / or entities shown in the drawings and / or described herein. In addition, the embodiments (whether or not explicitly described herein) also have significant utility for fields and applications beyond the examples described herein.
[0222] Embodiments are described herein with the aid of functional building blocks that illustrate implementations of specified functions and relationships thereof. For ease of description, the boundaries of these functional building blocks are arbitrarily defined herein. Alternative boundaries may be defined as long as the specified functions and relationships (or their equivalents) are properly performed. In addition, alternative embodiments may perform functional blocks, steps, operations, methods, etc., in a different order than described herein.
[0223] References to "one embodiment", "embodiment", "example embodiment" or similar phrases herein indicate that the described embodiment may include specific features, structures or variables, but not every embodiment necessarily includes the specific features, structures or variables. In addition, such phrases do not necessarily refer to the same embodiment. In addition, when describing specific features, structures or variables (whether or not explicitly mentioned or described in this article) in conjunction with an embodiment, incorporating such features, structures or variables into other embodiments will be within the knowledge of (multiple) technical personnel in the relevant field. Additionally, some embodiments may be described using the expressions "coupled" and "connected" and their derivatives. These terms are not necessarily intended to be synonyms of each other. For example, some embodiments may be described using the terms "connected" and / or "coupled" to indicate that two or more elements are in direct physical or electrical contact with each other. However, the term "coupled" may also mean that two or more elements are not in direct contact with each other, but still collaborate or interact with each other.
[0224] The breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents.
Claims
1. A computer-implemented method for evaluating a decision tree in a privacy-preserving manner, comprising: Receiving, by a first server, a first portion of a decision tree and encrypted input data from a client, wherein the encrypted input data includes a set of attributes; Receiving, by a second server, a second portion of the share of the decision tree; Each of the first server and the second server: communicates with the other server to calculate a classification result for the encrypted input data using the corresponding partial share of the decision tree received thereby and a secure multi-party computing method, wherein the secure multi-party computing method includes additive homomorphic encryption and oblivious transfer, so that the classification result is calculated without decrypting the encrypted input data, and so that the classification result is finally received by the first server; and The first server sends the classification result to the client.
2. The computer-implemented method of claim 1 , wherein: Decision trees include full binary decision trees.
3. The computer-implemented method of claim 1 , wherein: A decision tree consists of multiple decision nodes and multiple leaf nodes. The first part of the decision tree includes threshold tuples respectively associated with decision nodes of the decision tree and attribute index tuples respectively associated with decision nodes of the decision tree, and The second part of the decision tree includes a structural description of the decision tree and classification label tuples respectively associated with leaf nodes of the decision tree.
4. The computer-implemented method of claim 3, wherein: The classification label tuple is encrypted with a first public key that is paired with a first private key that is accessible to the first server but not to the second server, and Wherein the threshold tuple is encrypted using a second public key, wherein the second public key is paired with a second private key accessible to the second server but not to the first server.
5. The computer-implemented method of claim 1 , wherein: The attribute sets are associated with the user, wherein the decision tree comprises a classification model trained on one or more previously obtained attribute sets associated with the user, and wherein the classification result indicates whether the user should be authenticated.
6. The computer-implemented method of claim 5, wherein: The set of attributes associated with a user includes one or more of the following: Behavioral attributes associated with the user; or Biometric attributes associated with the user.
7. The computer-implemented method of claim 5, wherein: The computer-implemented method is performed as part of an ongoing authentication process for determining whether a user should have ongoing access to a system or resource.
8. The computer-implemented method of claim 1 , wherein: The first server and the second server communicate with each other directly or indirectly via one or more intermediaries.
9. A system for evaluating a decision tree in a privacy-preserving manner, comprising: A first server configured to receive a first portion of a decision tree and encrypted input data from a client, wherein the encrypted input data includes a set of attributes; and a second server configured to receive a second portion of the share of the decision tree; wherein each of the first server and the second server is further configured to communicate with the other server to calculate a classification result for the encrypted input data using the corresponding partial share of the decision tree received thereby and a secure multi-party computing method, wherein the secure multi-party computing method includes additive homomorphic encryption and oblivious transfer, so that the classification result is calculated without decrypting the encrypted input data, and the classification result is ultimately received by the first server, and The first server is further configured to send the classification result to the client.
10. The system according to claim 9, wherein: Decision trees include full binary decision trees.
11. The system according to claim 9, wherein: A decision tree consists of multiple decision nodes and multiple leaf nodes. The first part of the decision tree includes threshold tuples respectively associated with decision nodes of the decision tree and attribute index tuples respectively associated with decision nodes of the decision tree, and The second part of the decision tree includes a structural description of the decision tree and classification label tuples respectively associated with leaf nodes of the decision tree.
12. The system according to claim 11, wherein: The classification label tuple is encrypted with a first public key that is paired with a first private key that is accessible to the first server but not to the second server, and Wherein the threshold tuple is encrypted using a second public key, wherein the second public key is paired with a second private key accessible to the second server but not to the first server.
13. The system according to claim 9, wherein: The set of attributes is associated with the user, wherein the decision tree comprises a classification model trained on one or more previously obtained sets of attributes associated with the user, and wherein the classification result indicates whether the user should be authenticated.
14. The system according to claim 13, wherein: The set of attributes associated with a user includes one or more of the following: Behavioral attributes associated with the user; or Biometric attributes associated with the user.
15. The system according to claim 9, wherein: The first server and the second server are configured to communicate with each other directly or indirectly via one or more intermediaries.
16. A non-transitory computer-readable device having stored thereon instructions that, when executed by a first server, cause the first server to perform operations for evaluating a decision tree in a privacy-preserving manner, the operations comprising: receiving a first portion of shares of the decision tree and encrypted input data from a client, wherein the encrypted input data includes a set of attributes; communicating with a second server that has received a second partial share of the decision tree to calculate a classification result for the encrypted input data, wherein the first server and the second server are each configured to calculate the classification result for the encrypted input data using the corresponding partial share of the decision tree received thereby and a secure multi-party computation method, the secure multi-party computation method including additive homomorphic encryption and oblivious transfer, such that the classification result is determined without decrypting the encrypted input data, and such that the classification result is ultimately received by the first server; and Send the classification results to the client.
17. The non-transitory computer readable device of claim 16, wherein: Decision trees include full binary decision trees.
18. The non-transitory computer readable device of claim 16, wherein: A decision tree consists of multiple decision nodes and multiple leaf nodes. The first part of the decision tree includes threshold tuples respectively associated with decision nodes of the decision tree and attribute index tuples respectively associated with decision nodes of the decision tree, and The second part of the decision tree includes a structural description of the decision tree and classification label tuples respectively associated with leaf nodes of the decision tree.
19. The non-transitory computer readable device of claim 18, wherein: The classification label tuple is encrypted with a first public key that is paired with a first private key that is accessible to the first server but not to the second server, and Wherein the threshold tuple is encrypted using a second public key, wherein the second public key is paired with a second private key accessible to the second server but not to the first server.
20. The non-transitory computer readable device of claim 16, wherein: The set of attributes is associated with the user, wherein the decision tree comprises a classification model trained on one or more previously obtained sets of attributes associated with the user, and wherein the classification result indicates whether the user should be authenticated.