Two-party privacy preserving decision tree evaluation using additively homomorphic encryption
The method employs secure multi-party computation techniques to evaluate decision trees on encrypted data, addressing the challenge of creating classifiers for continuous authentication while maintaining privacy and security by ensuring sensitive user data remains confidential.
Patent Information
- Application Number
- JP2024196664
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-24
- Filing Date
- 2024-11-11
- Publication Date
- 2025-05-23
AI Technical Summary
The increasing complexity of security measures due to various types of attacks requires effective continuous authentication techniques, but creating classifiers for continuous authentication necessitates large amounts of sensitive user data, which poses privacy and security risks if disclosed.
A computer-implemented method for evaluating a decision tree in a privacy-preserving manner using secure multi-party computation (MPC) techniques, including additive homomorphic encryption and oblivious transfer, to compute classification results on encrypted input data without decrypting the data, ensuring that sensitive user data and the trained classifier remain confidential.
This approach allows for secure and privacy-preserving continuous authentication by enabling the computation of classification results without exposing sensitive user data, thereby protecting privacy and security while effectively authenticating users throughout their system interactions.
Smart Images

Figure 2025080236000001_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 548,289, filed November 13, 2023, which is incorporated herein by reference in its entirety. [Background technology]
[0002] The increasing complexity of security measures due to the emergence of various types of attacks requires multiple security techniques to effectively protect data and applications. One such technique is continuous authentication, which confirms a user's identity throughout their interactions with the system by continuously monitoring and evaluating behavioral and biometric user data. However, creating classifiers for continuous authentication requires large amounts of user data, which are in most cases highly sensitive and their disclosure would violate privacy and security. Summary of the Invention [Means for solving the problem]
[0003] 1. A computer-implemented method for evaluating a decision tree in a privacy-preserving manner, comprising: receiving, by a first server, a first partial share of the decision tree and encrypted input data from a client, the encrypted input data including a set of attributes; receiving, by a second server, a second partial share of the decision tree; computing, by each of the first server and the second server with the other server, the classification result on the encrypted input data using the respective partial shares of the decision tree received by the first server and the second server and secure multi-party computation methods including additive homomorphic encryption and oblivious transfer, such that a classification result is computed without decrypting the encrypted input data and such that the classification result is ultimately received by the first server; and transmitting, by the first server, the classification result to the client.
[0004] The accompanying drawings are incorporated in and form a part of this specification. [Brief description of the drawings]
[0005] [Figure 1] FIG. 1 illustrates an exemplary biometric recognition system according to an embodiment. [Diagram 2] FIG. 2 is a diagram showing an example of a full binary tree. [Diagram 3] FIG. 1 illustrates an example of decision tree evaluation using arbitrary data. [Figure 4] FIG. 1 illustrates a method of operation of the half-oblivious transfer (OT) protocol. [Diagram 5] FIG. 1 is a block diagram of an example system for performing privacy-preserving decision tree evaluation in accordance with some embodiments. [Figure 6] FIG. 1 illustrates an overview of decision tree contributions in the initialization phase of a process for performing privacy-preserving decision tree evaluation. [Figure 7]FIG. 13 illustrates an overview of the applied integer comparison protocol used to perform privacy-preserving decision tree evaluation according to some embodiments. [Figure 8] FIG. 13 illustrates the application of a secure integer comparison protocol for component-wise comparison of an attribute vector with a threshold vector to perform privacy-preserving decision tree evaluation according to some embodiments. [Figure 9A] FIG. 1 illustrates evaluating a decision node as true and allocating the comparison result to the corresponding edge for determining the traversal direction as part of performing a privacy-preserving decision tree evaluation according to some embodiments. [Figure 9B] FIG. 1 illustrates evaluating a decision node as false and allocating the comparison result to the corresponding edge for determining the traversal direction as part of performing a privacy-preserving decision tree evaluation according to some embodiments. [Figure 10] FIG. 13 illustrates an example of edge labeling of three decision tree nodes with encrypted comparison results as part of performing privacy-preserving decision tree evaluation according to some embodiments. [Figure 11] FIG. 13 illustrates a 1 in (m+1) oblivious transfer process that is performed as part of performing a privacy-preserving decision tree evaluation according to some embodiments. [Figure 12] FIG. 1 illustrates a brief overview of an exemplary protocol for performing privacy-preserving decision tree evaluation according to some embodiments. [Figure 13A] 1 is a flow diagram of an example method for performing privacy-preserving decision tree evaluation according to some embodiments. [Figure 13B] 1 is a flow diagram of an example method for performing privacy-preserving decision tree evaluation according to some embodiments. [Figure 13C] 1 is a flow diagram of an example method for performing privacy-preserving decision tree evaluation according to some embodiments. [Figure 13D] 1 is a flow diagram of an example method for performing privacy-preserving decision tree evaluation according to some embodiments. [Figure 14] 1 is a flow diagram of an example method for performing privacy-preserving decision tree evaluation according to some embodiments. [Figure 15] FIG. 1 is a block diagram of an example computer system useful for implementing various embodiments. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0006] In the drawings, like reference numbers generally indicate the same or similar elements. Further, the left-most digit of a reference number generally identifies the drawing in which the reference number first appears.
[0007] 1. Introduction As mentioned above in the Background section, the increasing complexity of security measures due to the emergence of various types of attacks requires multiple security techniques to effectively protect data and applications. One such technique is continuous authentication, which confirms a user's identity throughout their interactions with the system by continuously monitoring and evaluating behavioral and biometric user data. However, creating classifiers for continuous authentication requires large amounts of user data, which in most cases are highly sensitive and their disclosure would violate privacy and security.
[0008] A protocol for privacy-preserving continuous authentication based on secure multi-party computation (MPC) that addresses the aforementioned problems is described herein. The protocol allows multiple parties to compute a joint result on their input data without disclosing their input data. In an embodiment, a three-party protocol is implemented that includes a client with its own user data and two servers responsible for computing the classification using a decision tree learning algorithm as the underlying classifier. Such an embodiment may employ various methods of MPC to compute the classification result on the encrypted input through mutual interaction while protecting the sensitive user data and the trained classifier.
[0009] As technology continues to evolve, security measures are becoming more and more complex. This is due to the emergence of various types of attackers and attacks. There are two main types of attackers: internal attackers and external attackers. An internal attacker is an individual who has legitimate access to a system or network, usually an employee or contractor. An internal attacker may intentionally or unintentionally damage a system, either through malicious activity or negligence. An external attacker is an individual or group that attempts to gain unauthorized access to a system or network from the outside. These attackers may use various methods such as phishing, social engineering, or brute force attacks to gain access to a system. To effectively protect data and applications from these threats, multiple security techniques are required. One of these is the authentication of users to verify their true identity before granting access to a system or resource. This is usually done by applying initial authentication mechanisms such as passwords, multi-factor authentication, digital certificates, and even biometric data. However, in various scenarios, the initial authentication may not be sufficient. For example, users may choose weak passwords, forget to sign out of the system, or even forget to lock their devices, allowing access to adversaries. Therefore, a mechanism that authenticates the user for the entire session may be considered preferable.
[0010] Individuals have distinctive patterns of behavior that can be analyzed to identify and authenticate them while using a device or service. For computers and mobile devices, keystroke dynamics, mouse dynamics, touch and swipe patterns, phone orientation, and gait recognition, as well as physical location and time zone, are some of the factors that contribute to such patterns. Continuous authentication leverages these patterns to provide a novel security mechanism that confirms the identity of the user throughout their interactions with the system by continuously monitoring and evaluating behavioral and biometric user data and creating a distinctive biometric data profile. The process typically involves acquiring behavioral user data, deriving appropriate features from the behavioral user data, and creating a classification model for each user based on machine learning algorithms. The trained classification model is then used to authenticate the user during the session, and depending on the respective outcome, appropriate measures are taken based on the severity of the action being performed. Continuous authentication has broad applications in areas such as ensuring authorized access to critical infrastructure, online banking and financial services, e-commerce, medical applications accessing sensitive healthcare data, smart home and mobile devices, and identity access management in enterprise applications.
[0011] Creating a classifier for continuous authentication presupposes the application of machine learning, which requires a large amount of user data to train the classifier. However, behavioral and biometric user data are highly sensitive, and disclosure of such data is a gross violation of privacy and security. Furthermore, if such data falls into the wrong hands, it can be used to forge the identities of other users to circumvent the security mechanisms of continuous authentication, thereby allowing bad actors to infiltrate and damage the system (e.g., by stealing, modifying, or destroying files and data, installing malware, rendering software or hardware inoperable, etc.). Similarly, disclosure of a trained classification model puts the identity of a user at risk. Confidentiality of a machine learning model can result from two different sources. First, the trained model parameters themselves may contain sensitive information. Second, the model may have been created using sensitive data. It is well documented that accessing machine learning models, whether through white-box or black-box access methods, can lead to model inversion attacks and undermine the privacy of sensitive data. Therefore, it is crucial to find a way to classify user identities while protecting behavioral and biometric user data and trained classifiers.
[0012] To protect both the user data and the trained classifier, certain embodiments described herein apply Secure Multi-Party Computation (MPC), a cryptographic concept that allows multiple parties to compute a joint result on their input data without disclosing their input data to each other. The proposed implementation employs a three-party protocol that includes a client with user data that needs to be classified regarding authenticity, and two servers in charge of computing the classification. For the classification process, the embodiment utilizes a decision tree learning algorithm as the underlying classifier. First, to ensure the confidentiality of the data, the data is encrypted by the client before the classification process. To protect the sensitive user data utilized for training the model, each server obtains a partial share of the pre-trained decision tree model and does not disclose the complete model. Then, the servers interact with each other to compute the classification result on the encrypted input, utilizing various methods of MPC, more precisely, additive homomorphic encryption and oblivious transfer. Finally, the classification result is sent to the client, and appropriate measures are taken depending on the result.
[0013] 2. Background information The mathematical notations used in this disclosure are defined in Table 3 below.
[0014] 2.1. Continuous Authentication When interacting with a device or service, an individual has unique behavioral patterns that can be detected through various means. This pattern can be analyzed to help identify and authenticate the user. For computers and laptops, keystroke dynamics and mouse dynamics are factors that contribute to unique behavioral patterns. Keystroke dynamics involves analyzing an individual's typing speed, rhythm, and the time between keystrokes to create a unique profile. Mouse dynamics, similarly, evaluates an individual's mouse movement and click patterns to create such a profile. For mobile devices, touch and swipe patterns, phone orientation, and walk recognition are typical factors that contribute to unique behavioral patterns. Touch and swipe patterns and phone orientation are unique to each individual and indicate the way the user interacts with their mobile device. Walk recognition involves analyzing an individual's walking and movement patterns that can be detected through the use of the mobile device's accelerometer and gyroscope. Apart from these device-specific factors (which are presented only as examples and not as limitations), location, time zone, and other factors can serve as additional parameters for evaluating the credibility of the identity provided by the user.
[0015] This unique behavioral pattern can be used to implement an additional security mechanism called continuous authentication, which constantly verifies the user's identity throughout their interaction with the system, not just at the initial login. To achieve this, behavioral and physiological biometrics are continuously monitored, tracked, and analyzed through background processes that the user is unaware of. Based on the collected data, a unique biometric data profile is created and then used to authenticate the user. In this regard, FIG. 1 shows an exemplary biometric recognition system procedure 100 that includes six phases that are executed as background processes during interaction with the user interface. These phases are data acquisition 104, feature extraction 106, model training 108, user enrollment 110, inference 114, and decision making 116. The model training and user enrollment phases can be executed only once.
[0016] As shown in FIG. 1, the procedure 100 begins with a user interacting with a user interface 102 that captures raw user data for the creation and application of a data profile.
[0017] As will be apparent to one skilled in the art, there are a wide variety of behavioral data that can be collected by such a system. For example, using mouse dynamics as the sole basis for continuous authentication has already been demonstrated to be sufficient to achieve accurate results. The Balabit Mouse Dynamics Challenge was launched to provide a behavioral dataset regarding cursor movements to form the basis for the continuous authentication initiative. Table 1 shows an excerpt of raw mouse dynamics data from the Balabit Mouse Dynamics Challenge dataset.
[0018] [Table 1]
[0019] In the next phase, the time steps are grouped into segments, each describing one action. The actions are then used to derive appropriate features that are used to characterize the user. Preferably, these features should be universal, specific, immutable, measurable, difficult to spoof, and relevant in each application or environment. The extracted features are then used to create a data profile, in particular a classification model based on a supervised machine learning algorithm. This user-specific classification model, containing the characteristic features that are unique to the user, is finally stored in a dedicated database of the user model 112 and used for authentication during the session.
[0020] After user registration 110, the system can apply continuous authentication during subsequent sessions. The trained classification model is used to perform inferences 114 on the monitored data. This allows the classifier to make decisions 116 based on the current user's behavior to determine whether the behavior corresponds to a legitimate user or an impostor. Based on the severity of the action being performed, appropriate measures are taken, such as requesting the user to sign in again or taking additional security measures.
[0021] 2.2 Decision trees There are various classifiers based on supervised learning. Possible examples in the context of continuous authentication are k-nearest neighbors, decision trees, random forests, and even neural networks. In the embodiment described herein, decision trees (DT), also called regression trees, are used as the underlying classification algorithm. Decision trees are supervised learning algorithms used for classification tasks. Decision trees have a hierarchical binary tree structure consisting of internal or decision nodes, leaf nodes, and edges connecting the different nodes to each other. Each internal node is marked with a test condition and each leaf node is marked with a classification label. The test condition is a comparison of the value of a feature attribute of the input to be classified with a given threshold. Applying a decision tree algorithm, also called an inference process, means traversing a path in the tree based on a given input and arriving at one of the possible classification labels for each leaf node.
[0022] In the following description, the structure of the DT model is introduced, including its individual components and the overall application of the model. Decision trees are constructed from binary trees as the underlying data structure. Binary trees come in many types, such as full or proper binary trees, perfect binary trees, complete binary trees, and balanced binary trees.
[0023] In the embodiments described herein, binary trees are represented as full binary trees, where every tree node has either two or zero children. A full binary tree is a mathematical graph G = (V, E) with a hierarchical structure consisting of nodes, edges, and paths.
[0024] In particular, a complete binary tree is a tree with m decision nodes (d 1 , ..., d m )∈D⊂V and m + 1 leaf nodes (l 1 , ..., l m+1 )∈L⊂V, there are 2m + 1 nodes (v 1,..., v 2m+1 ) is included in V. The node index is assigned by a breadth-first search (BFS) scan.
[0025] The complete binary tree is a subset of the Cartesian product V×V (e 1 ,..., e 2m ) also includes the 2m connecting edges implied by ∈E, that is, for each e i ∈E, there exists a pair of nodes v i = (v a , v b ) such that ∈V. This gives G a consistent structure. Each decision node v a = d j must be connected to two child nodes. The first child is called the left child node v 2i , and the second child is called the right child node v 2i+1 . All nodes other than the root node v 1 = d 1 have an additional connection to their parent node
Number
[0026] The complete binary tree also includes m + 1 paths. Each path is a sequence P
[0027] starting from the root node v 1 and reaching a leaf node l k = (d k ,..., d 1 , l j ). Each path further has a depth δ k defined as the number of nodes in the path, i.e., δ k = |P k |. The depth δ of the tree is defined as the depth of the longest path in the tree, i.e., δ := ma(δ k ), where k ∈ [1, m + 1].
[0027] In FIG. 2, an example of a full binary tree 200 with m=4 decision nodes and tree depth δ=4 is shown.
[0028] To arrive at the definition of the DT model utilized by the embodiments described herein, the above definition of a full binary tree may be extended with several further attributes.
[0029] In particular, a decision tree (DT) or decision tree model is a method for determining an n-dimensional attribute vector x into classification labels c using a given binary tree G. k ∈{c 1 , ..., c m+1 Multivariate function that maps to
number
[0030] Decision node d j and the corresponding threshold t j The allocation function THR, which returns:
number
[0031] Decision node d j There are m attribute indices I = (i 1 , ..., i m )
[0032] For a given leaf node l k Take the corresponding classification label c k A labeling function LAB that returns L→{c 1 , ..., c m+1}, each of which is assigned to a leaf node. 1 , ..., c m+1 )
[0033] Given input data, it can be evaluated using the DT model, i.e., inference can be performed on the attributes or feature vectors to derive a classification of the input. The functions THR and ATT are required to establish the test conditions and correctly traverse G, while LAB is used to finally derive the corresponding classification result.
[0034] Next, the application of the DT model (also known as decision tree evaluation) as a function to classify arbitrary input data is described.
[0035] Given a DT model for DT evaluation, the input attribute x i ∈{x 1 , ..., x n} and threshold t j ∈{t 1 , ..., t m} and test the greater-or-equal condition [x i ≧t j ]∈{0, 1}, a two-variable comparison function COMP:
number
[0036] To perform inference for a given input x and a DT model, the process starts at the root node v 1 Start with each reached node v j In contrast, the process is
number
[0037] FIG. 3 shows an example of decision tree evaluation 300 using arbitrary data, given attributes x = (2, 3, 0), thresholds T = (1, 1, 4, 7), classification labels C = (0, 1, 0, 1, 0), and attribute allocation I = (1, 3, 2, 3). First, the process evaluates the GEQ test condition
number
number
[0038] 2.3 Secure Multiparty Computation Both the input data and the trained parameters of the DT classifier include behavioral user data. Since behavioral data is highly sensitive and the disclosure of such data is a gross violation of privacy and security, privacy-preserving machine learning techniques may be applied as described herein. In particular, secure multi-party computation (MPC) may be applied as described herein, allowing the model owner and the data owner to perform inference, respectively, without any party learning additional information other than itself. MPC techniques for privacy-preserving inference with DT described herein include additive homomorphic encryption (AHE) and oblivious transfer (OT).
[0039] additive homomorphic encryption AHE is a type of encryption scheme that allows one to perform arithmetic additions and scalar multiplications of ciphertext, producing an encrypted result that yields the same result as if the operations were performed on the plaintext.
[0040] An encryption method may be defined as a set of three algorithms: ·pk, sk←K EY G EN (λ). This probabilistic algorithm takes a security parameter λ and outputs a key pair containing a public key pk and a private key sk. c←E NC (m, pk). This probabilistic algorithm takes as input a public key pk and a plaintext message m, and outputs a ciphertext c.
number
[0041] If the public and private keys are equal, i.e., pk = sk, then the encryption scheme may be called symmetric or private-key encryption. Otherwise, the encryption scheme may be called asymmetric or public-key encryption.
[0042] An encryption scheme whose ciphertext is additively homomorphic is called an additively homomorphic encryption scheme. More precisely, for an encryption function and two plaintext values m 1 and m 2 If we have x, y, and y, then adding both values and later encrypting their sum will yield the same result as first encrypting both values and then adding their ciphertexts using the AHE scheme. Thus, in addition to the above algorithms related to the definition of the encryption scheme, the AHE scheme provides additional operations, as discussed below.
[0043] An additively homomorphic encryption scheme may be an encryption scheme that additionally provides the following methods:
[0044] ·c 3 ←A DD (c 1 , c 2 , pk). This operation is
number
number
number
[0045] ·c 3 ←A DD C ONST (c 1 , m 2 , pk). This method
number
number
[0046] ·c 3 ←M ULTI C ONST (c 1 , m 2 , pk). This method
number
number
[0047] ·c 3 ←I NV (c, pk). This method
number
number
[0048] c←X OR (a, b, pk). This operation is
number
number
[0049]
number
number
[0050] In an embodiment, IND-CPA security and the following correctness condition are met for all plaintexts m 1 , m 2 , and plaintext bits a, b are needed.
number
[0051] Common examples of AHE schemes known to those skilled in the art are the Paillier or Elliptic Curve Lifted Elgamal (ECLE) schemes, however such AHE schemes are provided herein as examples and are not intended to be limiting.
[0052] Oblivious Transfer Oblivious Transfer (OT) is a cryptographic protocol that allows a sender to transfer one of several data elements to a receiver without revealing which data element was selected. Many variations of OT exist, but the basic variant is:
number
[0053]
number
[0054]
number
number
number
number
[0055] The offline component relies on asymmetric cryptography to generate a finite set of what are called base or seed oblivious transfers. These base OTs can then be expanded in the online phase to an unlimited number of oblivious transfers, achieved solely by operations involving symmetric cryptography. This can be considered as the hybrid cryptography equivalent of OT.
[0056] 3. Privacy-preserving decision tree evaluation FIG. 5 is a block diagram of an exemplary system 500 for performing privacy-preserving decision tree evaluation according to an embodiment. The exemplary system 500 of FIG. 5 implements a three-party MPC protocol. In particular, the exemplary system 500 includes a client 502 that wishes to classify its input and two servers 504 and 506 that compute the classification. Each of the client 502, the server 504, and the server 506 may be implemented by or on a corresponding computing device or computer system (such as computer system 1500 described below with reference to FIG. 15). Furthermore, each of the client 502, the server 504, and the server 506 may be implemented by or on a corresponding physical or virtual machine.
[0057] The input of client 502 is encrypted using AHE and remains hidden to both server 504 and server 506 throughout the inference process. Both server 504 and server 506 receive a share of the previously trained DT model 512, and therefore neither of them knows the structure and thresholds of the complete DT model. In particular, server 504 receives DT share 508 and server 506 receives DT share 510. Both servers then communicate with each other to compute classification results on the encrypted inputs using their model shares and MPC techniques, i.e., AHE and OT. Finally, the classification results are sent to client 502.
[0058] In the following, each step of the protocol is described in detail, and in the following, client 502 is referred to as "CLIENT", server 504 is referred to as "SERVER 1", and server 506 is referred to as "SERVER 2".
[0059] 3.1 Initialization First, each server sends its individual key pair (pk i , sk i ), where i represents the server index. CLIENT creates both public keys pk 1 and pk 2 You can access the pk 2 Encrypt the CLIENT's input using x = (x 1 , ..., x n ) is a CLIENT input consisting of n feature attributes, and each attribute
number
number
[0060] It is assumed that a trained decision tree model is given, i.e., a binary tree consisting of decision nodes and leaf nodes, classification labels for each leaf node, and thresholds for each decision node that are determined during the training phase of the DT model. The DT is then split into two separate shares that are sent to each server respectively by a management entity, process 600 shown in Fig. 6.
[0061] The first DT share passed to SERVER 1 consists of two elements. The first element is a tuple of m thresholds, each of which is bit-wise encrypted to be of length μ, i.e.
number
[0062] The second element is I = (i 1 , ..., i m ), i j ∈[0, n] It is a tuple of attribute indexes expressed as follows:
[0063] Given x and T, the protocol evaluates each internal node d j GEQ test conditions performed in
number
number
[0064] As already discussed, in an embodiment, the underlying binary tree is a full tree, and therefore need not necessarily be a complete binary tree. Thus, a given encrypted threshold and attribute index tuple does not indicate the actual alignment between the nodes. Furthermore, all thresholds are encrypted using the public key pk of SERVER 2. 2 , and is therefore unreadable to SERVER 1. As a result, the first DT share does not reveal any sensitive information about the trained DT model, and therefore does not reveal any inferable user data.
[0065] The second share sent to SERVER 2 contains a description of the tree structure given by G and pk 1 The tree structure may be represented as the set of all internal and leaf nodes, including the edges of the nodes, but not including the associated thresholds and classification labels. Since the tree structure alone is not sufficient to determine the sensitive information of the trained model, the second share similarly does not reveal any sensitive user data, and the actual DT model remains hidden to both servers. Figure 6 and Table 2 give an overview of the DT sharing process, including the components of each share. Moreover, it is assumed that both servers are under separate administration, and therefore no supplementary information is shared other than that explicitly outlined in the protocol.
[0066] Table 2 below provides an overview of the shares of decision tree models consisting of binary tree structures G, thresholds T, input feature allocations I, and classification labels C. [Table 2]
[0067] 3.2 Comparative calculation As mentioned above, CLIENT's input is encrypted before being sent over the network channel to SERVER 1. To make the input unreadable to SERVER 1, pk 2 is applied, resulting in the encrypted input being
number
number
[0068] SERVER 1:
number
number
number
[0069] At each decision node, the comparison result b j is defined as a Boolean greater than or equal comparison, i.e.
number
number
[0070] Figure 7 shows the relationship between two integers x, each consisting of a tuple of μ bits. (1) and x (2) 5 shows an overview 700 of the Tueno et al. integer comparison protocol applied to compare x (1) are the respective input attributes
number
[0071] To apply integer comparison of two integers between two parties, the first party creates a key pair (sk, pk) and shares its public key pk with the second party. Both parties have an integer consisting of a μ-bit tuple and can know the result of the comparison, i.e., x, without disclosing its actual value. (1) ≧x (2) I want to know if the encrypted input is
number
number
[0072] In the context of the system in Fig. 5, this procedure is applied when multiple attributes of x are compared with multiple thresholds of T together with decision nodes of T, where party 1 is represented by CLIENT and party 2 is represented by SERVER 1. As shown in Algorithm 1, the comparison result of each decision node is stored in B∈{0, 1} m Then, each component of B with index j is its share
number
number
number
number
[0073] FIG. 8 shows an application 800 of the secure integer comparison protocol of Tueno et al. for component - unit comparison between an attribute vector x and a threshold vector T.
[0074] The operations of this section executed by SERVER 1 are
Number
Number
Number
Number
Number
Number
[0075] However, the actual comparison result is encrypted and remains unreadable by SERVER 2. Thus, highly confidential information is not revealed. In summary, the process of deriving the overall comparison result executed by SERVER 2 is
Number
[0076] 3.3 Edge Labeling Now that SERVER 2 has obtained the encrypted comparison result described by tuple B, the decision tree is applied to evaluate the CLIENT input to determine the matched classification label. This is done by traversing the tree and assigning the corresponding components of vector B to the edges of the tree. Figures 9A and 9B outline the operations performed at each node during the traversal of the tree.
[0077] In particular, Figures 9A and 9B show the evaluation of decision nodes and the allocation of comparison results to corresponding edges to determine the traversal direction. All values are 1 Note that the input x is encrypted by x = 0 and is therefore unreadable to SERVER 2. Furthermore, SERVER 2 only receives the result of the computation, and not the input x or the threshold T, which are shown for clarity.
[0078] The process starts with the first internal node, i.e., the root node. If the assigned comparison evaluates to true, then the tree is traversed to the right child node, otherwise it is traversed to the left child node, and the process is repeated until a leaf node is reached. Since the actual path traversed depends on the input x, the process uses the comparison result B to mark the traversed edge. Further according to the process, for reasons explained in the subsequent sections, the value of the traversed edge is set to 0, otherwise it is set to 1. In the example 900 of FIG. 9A, it can be seen that the comparison evaluates to true, and as a result, the tree needs to be traversed to the right child node. Thus, the process marks the result bit b j By inverting b, we set the outgoing right edge to 0 and the outgoing left edge to 1, i.e., b j Set to x i In the inverse example 902 of Figure 9B where t < t, the process does the same such that the left edge is set to the comparison result bit 0 and the right edge is set to the inverted result 1. This procedure may be referred to as edge labeling.
[0079] The process of edge labeling may be defined as follows: For each decision node v jBut integer comparison operations
number
number
number
number
[0080] Figure 10 shows the (pk 1 ,(encrypted with ,) the edge labeling of the three decision nodes in tree 1000 with m = 3. As mentioned above, SERVER 2 simply labels each outgoing left edge with the comparison result b j and each outgoing right edge is its inverse
number
[0081] The functions introduced in this section take as input a decision tree model T and the shares of the computation results, and for each tuple e i = (v a , v b )∈E is the third value b j (1 - b j ) to output the updated binary tree G' of the DT model extended by
number
[0082] 3.4 Tree evaluation After the edge labeling step, SERVER 2 determines that a path with all edges marked with 0 has the appropriate classification label c k Since ,B, ,denotes the traversal route of the provided input x that leads to,C,, we need to navigate the tree along this path. This path is called the zero path. However, because B and C are encrypted, every path must be evaluated to identify the zero path. The evaluation of the paths can be denoted as path aggregation.
[0083] A path aggregation may be defined as follows: Given a previously defined path P whose edges are labeled in the manner described in the previous section, k For each leaf node c k The resulting value may be defined as the sum of all edge labels along the path leading to p k It is expressed as:
[0084] The path cost of the whole tree is the tuple p of all individual path costs.
number
number
number
number
[0085] With m = 3 decision nodes, there are m + 1 = 4 distinct paths, with only one sum evaluating to 0 + 0 = 0, one evaluating to 1 + 1 = 2, and the others evaluating to 0 + 1 = 1 + 0 = 1. In the current state, if P was sent to SERVER 1 for a classification decision in the current state, P would still reveal the structure of the tree, allowing most of the tree to be reconstructed. Thus, after each component permutes P by using the random permutation function π, the individual random scalar r k is multiplied by, and therefore,
number
[0086] The sorting function π defines a fixed sorting and the index
number
[0087] Further following the previous example, the sorting function is
number
number
number
[0088] SERVER 2 is
number
[0089] In summary, the process aggregates all paths in the updated binary tree G', resulting in a tuple of associated costs for each path, where each component is randomized individually. Then, the process chooses a random permutation π to permute the path cost tuples. Finally, the resulting tuples are sent to SERVER 1, where the applied permutation function is retained for later operations. The overall procedure is as follows:
number
[0090] 3.5 Classification determination The final phase of the protocol is
number
number
number
number
number
number
number
[0091] c (i) If c = 1, then the CLIENT is considered valid, otherwise π(i) = 0, the CLIENT is considered invalid. This information is communicated to the client accordingly and appropriate action is taken.
[0092] 4. Summary Table 3 shown below (Table 4A, Table 4B) provides a list of notations, variables, and symbols used herein to describe an exemplary protocol for performing privacy-preserving decision tree evaluation. [Table 4A] [Table 4B]
[0093] Additionally, FIG. 12 provides a brief overview 1200 of an exemplary protocol implementation, including five broad phases of the protocol, three participating parties, and the communications between the parties.
[0094] The protocol represented by FIG. 12 will now be described with reference to FIG. 13A, FIG. 13B, FIG. 13C, and FIG. 13D. In particular, FIG. 13A, FIG. 13B, FIG. 13C, and FIG. 13D collectively illustrate a flow diagram of a method 1300 for evaluating a decision tree in a privacy-preserving manner according to some embodiments. The method 1300 may be performed by processing logic that may include hardware (e.g., circuits, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executed in a processing device), or a combination thereof. It should be understood that not all steps may be required to carry out the disclosure provided herein. Furthermore, some of the steps may be performed simultaneously or in a different order than shown in FIG. 13A, FIG. 13B, FIG. 13C, and FIG. 13D, as will be understood by those skilled in the art.
[0095] The method 1300 will be described with reference to the system of Figure 5. However, the method 1300 is not limited to that example embodiment.
[0096] During the input phase (or prior to the input phase), at 1302, SERVER 1 is provided with (a) a first public-private key pair consisting of a first public key and a first private key, (b) a second public key comprising a portion of the second public-private key pair provided to SERVER 2, and (c) a first partial share of a decision tree classifier, the first portion of the decision tree classifier consisting of (i) threshold tuples respectively associated with nodes of a decision tree of the decision tree classifier, where each threshold in the threshold tuple is bit-wise encrypted (using an AHE scheme) using the second public key, and (ii) attribute index tuples respectively associated with nodes of a decision tree of the decision tree classifier, where each attribute index in the attribute index tuple indicates an input data attribute associated with a corresponding decision tree node.
[0097] During the input phase (or prior to the input phase), at 1304, SERVER 2 is provided with (a) a second public-private key pair consisting of a second public key and a second private key, (b) the first public key, and (c) a second portion of the decision tree classifier, the second portion of the decision tree classifier consisting of (i) a description of the tree structure of the decision tree classifier, and (ii) tuples of classification labels respectively associated with leaf nodes of the decision tree classifier, where each classification label is encrypted (using an AHE scheme) using the first public key.
[0098] During the input phase (or prior to the input phase), at 1306, the CLIENT is provided with a first public key and a second public key.
[0099] During the input phase, at 1308, CLIENT obtains input data including a set of attributes (e.g., user attributes) and bitwise encrypts each attribute of the input data (using the AHE scheme) using the second public key, thereby generating encrypted input data. During the input phase, at 1310, CLIENT also sends the encrypted input data to SERVER 1 (the encrypted input data cannot be decrypted by SERVER 1 because it is encrypted with the second public key).
[0100] During the comparison calculation phase, at 1312, SERVER 1 uses AHE to compare the encrypted input data with the encrypted thresholds to generate comparison results for each of the decision nodes of the decision tree classifier (this means that the encrypted input data and the encrypted thresholds are never decrypted to plaintext), with each comparison result represented by a share of the first comparison result in plaintext and a share of the second comparison result encrypted with the second public key (hence the share of the second comparison result cannot be decrypted by SERVER 1).
[0101] At 1314, SERVER 1 then encrypts (using AHE) each of the shares of the first comparison result with the first public key.
[0102] At 1316, SERVER 1 then transmits the comparison results to SERVER 2, with each transmitted comparison result represented by a share of the first comparison result encrypted with the first public key (so that the share of the first comparison result cannot be decrypted by SERVER 2) and a share of the second comparison result encrypted with the second public key.
[0103] At 1318, for each comparison result, SERVER 2 decrypts a share of the second comparison result using the second private key, and then performs a homomorphic XOR operation on the share of the first comparison result encrypted with the first public key and the decrypted second comparison result to produce a comparison result encrypted with the first public key (and thus, the comparison result cannot be decrypted by SERVER 2), where each comparison result is either 1 if the comparison evaluates to true or 0 if the comparison evaluates to false.
[0104] During the edge labeling phase, at 1320, for each decision node of the decision tree classifier, SERVER 2 labels the first outgoing edge of the decision node (the edge corresponding to false) with the corresponding comparison result encrypted with the first public key and labels the second outgoing edge of the decision node (the edge corresponding to true) with the inverse of the corresponding comparison result encrypted with the first public key.
[0105] During the tree evaluation phase, at 1322, for each path from the root node of the decision tree to each leaf node, SERVER 2 calculates a path cost that represents the sum of the edge labels along the path, thereby generating a tuple of path costs, each encrypted with the first public key (so that the tuple cannot be decrypted by SERVER 2), and the path leading to the classification result is the only path whose path cost is zero (zero path).
[0106] At 1324, SERVER 2 multiplies each path cost in the tuple of path costs by a corresponding random scalar value, and then applies a sorting function to sort the order of the path costs in the tuple of path costs, thereby generating a sorted and randomized tuple of path costs encrypted with the first public key. At 1326, SERVER 2 further sorts the tuple of classification labels encrypted with the first public key using the same sorting function.
[0107] At 1328, SERVER 2 sends the sorted and randomized tuple of path costs encrypted with the first public key to SERVER 1.
[0108] In the classification decision phase, at 1330, SERVER 1 decrypts the sorted and randomized tuple of path costs using the first private key to identify the index associated with the zero path.
[0109] At 1332, SERVER 1 and SERVER 2 then participate in a lost communication operation where SERVER 1 obtains from SERVER 2 the classification label encrypted with the first public key that is related to the index associated with the zero path, without SERVER 2 knowing the index associated with the zero path and without SERVER 1 knowing any of the classification labels related to any of the other paths of the decision tree classifier.
[0110] At 1334, SERVER 1 decrypts the classification label related to the index associated with the zero path using the first private key.
[0111] At 1336, SERVER 1 sends the classification label related to the index associated with the zero path to the CLIENT.
[0112] Figure 14 is a flow diagram of a method 1400 for evaluating a decision tree in a privacy-preserving manner according to some embodiments. The method 1400 may be performed by processing logic that may include hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executed in a processing device), or a combination thereof. It should be understood that not all steps may be required to perform the disclosure provided herein. Furthermore, some of the steps may be performed simultaneously or in a different order than shown in Figure 14, as will be understood by those skilled in the art.
[0113] The method 1400 will be described with reference to the system of Figure 5. However, the method 1400 is not limited to that example embodiment.
[0114] At 1402, SERVER 1 receives a first partial share of the decision tree and encrypted input data from CLIENT, where the encrypted input data includes a set of attributes.
[0115] At 1404, SERVER 2 receives a second partial share of the decision tree.
[0116] At 1406, each of SERVER 1 and SERVER 2 communicates with the other server to compute a classification result on the encrypted input data using their respective partial shares of the decision trees received by those servers and secure multi-party computation methods including additive homomorphic encryption and oblivious transfer, such that the classification result is determined without decrypting the encrypted input data, and such that the classification result is ultimately received by SERVER 1.
[0117] At 1408, SERVER 1 sends the classification results to CLIENT.
[0118] In an embodiment of the method 1400, the decision tree comprises a full binary decision tree.
[0119] In another embodiment of method 1400, the decision tree includes a plurality of decision nodes and a plurality of leaf nodes, a first partial share of the decision tree includes a tuple of thresholds respectively associated with the decision nodes of the decision tree and a tuple of attribute indices respectively associated with the decision nodes of the decision tree, and a second partial share of the decision tree includes a description of the structure of the decision tree and a tuple of classification labels respectively associated with the leaf nodes of the decision tree.
[0120] Further according to such an embodiment, the classification label tuple may be encrypted with a first public key paired with a first private key accessible to the first server but not to the second server, and the threshold tuple may be encrypted with a second public key paired with a second private key accessible to the second server but not to the first server.
[0121] In a further embodiment of method 1400, a set of attributes is associated with a user, and the decision tree includes a classification model trained on a set of one or more previously obtained attributes associated with the user, and the classification result indicates whether the user should be authenticated or not authenticated.
[0122] Further according to such an embodiment, the set of attributes associated with the user may include one or more of behavioral attributes associated with the user or biometric attributes associated with the user.
[0123] Furthermore, according to such an embodiment, the method 1400 may be performed as part of a continuing authentication process used to determine whether a user should continue to access a system or resource.
[0124] In a further embodiment of method 1400, the first server and the second server communicate with each other directly. In an alternative embodiment, the first server and the second server communicate with each other indirectly through one or more intermediaries (e.g., an intermediate physical or virtual machine, an intermediate computing device, etc.).
[0125] Various embodiments may be implemented using one or more well-known computer systems, such as, for example, computer system 1500 shown in Figure 15. One or more computer systems 1500 may be used, for example, to implement any of the embodiments discussed herein, as well as combinations and subcombinations of those embodiments.
[0126] Computer system 1500 may include one or more processors (also referred to as central processing units or CPUs), such as processor 1504. Processor 1504 may be connected to a communications infrastructure or bus 1506.
[0127] The computer system 1500 may also include user input / output devices 1503 , such as a monitor, keyboard, pointing device, etc., that may communicate with a communications infrastructure 1506 through a user input / output interface 1502 .
[0128] One or more of the processors 1504 may be a graphics processing unit (GPU). In an embodiment, a GPU may be a processor that is a specialized electronic circuit designed to process mathematically intensive applications. A GPU may have a parallel structure that is efficient for parallel processing of large blocks of data, such as the mathematically intensive data common in computer graphics applications, images, videos, and the like.
[0129] The computer system 1500 may also include a main or primary memory 1508, such as a random access memory (RAM). The main memory 1508 may include one or more levels of cache. The main memory 1508 may store control logic (i.e., computer software) and / or data.
[0130] The system 1500 may also include one or more secondary storage devices or memories 1510. The secondary memory 1510 may include, for example, a hard disk drive 1512 and / or a removable storage device or drive 1514. The removable storage drive 1514 may be a floppy disk drive, a magnetic tape drive, a compact disk drive, an optical storage device, a tape backup device, and / or any other storage device / drive.
[0131] The removable storage drive 1514 may interact with a removable storage unit 1518. The removable storage unit 1518 may include a computer usable or readable storage device that stores computer software (control logic) and / or data. The removable storage device 1518 may be a floppy disk, magnetic tape, compact disk, DVD, optical storage disk, and / or any other computer data storage device. The removable storage drive 1514 may read from and / or write to the removable storage unit 1518.
[0132] Secondary memory 1510 may include other means, devices, components, equipment, or other techniques for allowing computer programs and / or other instructions and / or data to be accessed by computer system 1500. Such means, devices, components, equipment, or other techniques may include, for example, a removable storage unit 1522 and an interface 1520. Examples of removable storage units 1522 and interfaces 1520 may include a program cartridge and cartridge interface (such as found in a video game device), a removable memory chip (such as an EPROM or PROM) and associated socket, a memory stick and USB port, a memory card and associated memory card slot, and / or any other removable storage unit and associated interface.
[0133] Computer system 1500 may further include a communication or network interface 1524. Communication interface 1524 may enable computer system 1500 to communicate and interact with any combination of external devices, external networks, external entities, etc. (individually and collectively referred to by reference number 1528). For example, communication interface 1524 may enable computer system 1500 to communicate with external or remote devices 1528 via communication paths 1526, which may be wired and / or wireless (or a combination thereof) and may include any combination of a LAN, a WAN, the Internet, etc. Control logic and / or data may be transmitted to and from computer system 1500 via communication paths 1526.
[0134] Computer system 1500 may be any one or any combination of a personal digital assistant (PDA), a desktop workstation, a laptop or notebook computer, a netbook, a tablet, a smartphone, a smart watch or other wearable, an appliance, part of the Internet of Things, and / or an embedded system, just to name a few non-limiting examples.
[0135] The computer system 1500 may be a client or server that accesses or hosts any application and / or data via any delivery paradigm, including, but not limited to, remote or distributed cloud computing solutions, local or on-premise software ("on-premise" cloud-based solutions), an "as a service" model (e.g., content as a service (CaaS), digital content as a service (DCaaS), software as a service (SaaS), managed software as a service (MSaaS), platform as a service (PaaS), desktop as a service (DaaS), framework as a service (FaaS), backend as a service (BaaS), mobile backend as a service (MBaaS), infrastructure as a service (IaaS), etc.), and / or a hybrid model including any combination of the foregoing examples or other service or delivery paradigms.
[0136] Any applicable data structures, file formats, and schemas of computer system 1500 may be derived from standards including, but not limited to, JavaScript Object Notation (JSON), Extensible Markup Language (XML), Yet Another Markup Language (YAML), Extensible Hypertext Markup Language (XHTML), Wireless Markup Language (WML), MessagePack, XML User Interface Language (XUL), or any other functionally similar representations, alone or in combination. Alternatively, proprietary data structures, formats, or schemas may be used exclusively or in combination with known or open standards.
[0137] In some embodiments, a tangible, non-transitory apparatus or article of manufacture that includes a tangible, non-transitory computer usable or readable medium that stores control logic (software) may also be referred to herein as a computer program product or program storage device. This includes, but is not limited to, computer system 1500, main memory 1508, secondary memory 1510, removable storage units 1518 and 1522, and tangible articles of manufacture that embody any combination of the foregoing. Such control logic, when executed by one or more data processing devices (such as computer system 1500), may cause such data processing devices to operate as described herein.
[0138] It will be apparent to one of ordinary skill in the art, based on the teachings contained herein, how to make and use embodiments of the present disclosure using data processing devices, computer systems, and / or computer architectures other than those shown in Figure 15. In particular, embodiments may operate with software, hardware, and / or operating system implementations other than those described herein.
[0139] It should be understood that the Detailed Description section is intended to be used to interpret the claims, and that none of the other sections are intended to be used to interpret the claims. The other sections may set forth one or more, but not all, example embodiments contemplated by the inventor(s), and thus are not intended to limit the disclosure or the appended claims in any way.
[0140] Although the present disclosure describes exemplary embodiments for exemplary fields and applications, it should be understood that the present disclosure is not limited thereto. Other embodiments and modifications thereto are possible and are within the scope and spirit of the present disclosure. For example, without limiting the generality of this paragraph, the embodiments are not limited to the software, hardware, firmware, and / or entities illustrated in the figures and / or described herein. Moreover, the embodiments (whether or not explicitly described herein) have great utility for fields and applications beyond the examples described herein.
[0141] The embodiments have been described herein with the aid of functional building blocks that illustrate implementations of specified functions and relationships. The boundaries of these functional building blocks have been arbitrarily defined herein for convenience of description. Alternative boundaries may be defined so long as the specified functions and relationships (or their equivalents) are appropriately performed. Also, alternative embodiments may execute functional blocks, steps, operations, methods, etc. using an order different from that described herein.
[0142] Reference herein to "one embodiment," "embodiment," "exemplary embodiment," or similar phrases indicates that the described embodiment may include a particular feature, structure, or variable, but not every embodiment may necessarily include the particular feature, structure, or variable. Moreover, such phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or variable is described in connection with an embodiment, it would be within the knowledge of one of ordinary skill in the art to incorporate such feature, structure, or variable in other embodiments, whether or not explicitly mentioned or described herein. In addition, some embodiments may be described using the phrases "coupled" and "connected," along with their derivatives. These terms are not necessarily intended as synonyms for each other. For example, some embodiments may be described using the terms "connected" and / or "coupled" to indicate that two or more elements are in direct physical or electrical contact with each other. However, the term "coupled" may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.
[0143] The breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents. [Explanation of symbols]
[0144] 100 steps 102 User Interface 104 Data Acquisition 106 Feature Extraction 108 Model Training 110 User Registration 112 User Model 114 Reasoning 116 Decision making 200 Complete binary tree 300 Decision Tree Evaluation 400 How it works 402 Recipient 404 OT Protocol 406 Sender 500 System 502 Client 504 Server 506 Server 508 DT Share 510 DT Share 512 Trained DT Model 600 Process 700 Summary 800 Application 900 Example 902 Counterexample 1000 Tree 1200 Summary 1300 Method 1400 Method 1500 Computer System 1502 User Input / Output Interface 1503 User Input / Output Device 1504 Processor 1506 Communication Infrastructure or Bus 1508 Main Memory or Primary Memory 1510 Secondary Storage Device or Memory 1512 Hard Disk Drive 1514 Removable Storage Device or Drive 1518 Removable Storage Unit 1520 Interface 1522 Removable Storage Unit 1524 Communication or Network Interface 1526 Communication Path 1528 External or Remote Device
Claims
1. 1. A computer-implemented method for evaluating a decision tree in a privacy-preserving manner, comprising: receiving, by a first server, a first partial share of the decision tree and encrypted input data from a client, the encrypted input data including a set of attributes; receiving, by a second server, a second partial share of the decision tree; communicating by each of the first server and the second server with the other server to calculate the classification result on the encrypted input data using the respective partial shares of the decision trees received by the first server and the second server and secure multi-party computation methods including additive homomorphic encryption and oblivious transfer, such that the classification result is calculated without decrypting the encrypted input data and such that the classification result is ultimately received by the first server; and transmitting, by the first server, the classification result to the client.
2. 10. The computer-implemented method of claim 1, wherein the decision tree comprises a full binary decision tree.
3. the decision tree includes a plurality of decision nodes and a plurality of leaf nodes; the first partial share of the decision tree includes a tuple of thresholds respectively associated with decision nodes of the decision tree and a tuple of attribute indices respectively associated with the decision nodes of the decision tree; 2. The computer-implemented method of claim 1, wherein the second partial share of the decision tree includes a description of a structure of the decision tree and tuples of classification labels associated with each of the leaf nodes of the decision tree.
4. the classification label tuple is encrypted with a first public key paired with a first private key accessible to the first server but not to the second server; 4. The computer-implemented method of claim 3, wherein the threshold tuple is encrypted with a second public key paired with a second private key that is accessible to the second server but not to the first server.
5. 2. The computer-implemented method of claim 1, wherein the set of attributes is associated with a user, the decision tree includes a classification model trained on a set of one or more previously obtained attributes associated with the user, and the classification result indicates whether the user should be authenticated or not authenticated.
6. The set of attributes associated with the user may include: Behavioral attributes associated with the user; or The computer-implemented method of claim 5 , further comprising: determining whether the user is associated with one or more of the biometric attributes.
7. The computer-implemented method of claim 5 , wherein the method is performed as part of a continuing authentication process used to determine whether the user should continue to have access to a system or resource.
8. 10. The computer-implemented method of claim 1, wherein the first server and the second server communicate with each other directly or indirectly through one or more intermediaries.
9. 1. A system for evaluating decision trees in a privacy-preserving manner, comprising: a first server configured to receive a first partial share of the decision tree and encrypted input data from a client, the encrypted input data including a set of attributes; a second server configured to receive a second partial share of the decision tree; each of the first server and the second server is further configured to communicate with the other server to compute the classification result on the encrypted input data using the respective partial shares of the decision trees received by the first server and the second server and secure multi-party computation methods including additive homomorphic encryption and oblivious transfer, such that the classification result is computed without decrypting the encrypted input data and such that the classification result is eventually received by the first server; The first server is further configured to transmit the classification result to the client.
10. The system of claim 9 , wherein the decision tree comprises a full binary decision tree.
11. the decision tree includes a plurality of decision nodes and a plurality of leaf nodes; the first partial share of the decision tree includes a tuple of thresholds associated with each decision node of the decision tree and a tuple of attribute indices associated with each decision node of the decision tree; 10. The system of claim 9, wherein the second partial share of the decision tree includes a description of a structure of the decision tree and tuples of classification labels associated with each of the leaf nodes of the decision tree.
12. the classification label tuple is encrypted with a first public key paired with a first private key accessible to the first server but not to the second server; 12. The system of claim 11, wherein the threshold tuple is encrypted with a second public key paired with a second private key that is accessible to the second server but not to the first server.
13. 10. The system of claim 9, wherein the set of attributes is associated with a user, the decision tree includes a classification model trained on a set of one or more previously obtained attributes associated with the user, and the classification result indicates whether the user should be authenticated or not authenticated.
14. The set of attributes associated with the user may include: Behavioral attributes associated with the user; or The system of claim 13 , further comprising one or more of the biometric attributes associated with the user.
15. 10. The system of claim 9, wherein the first server and the second server are configured to communicate with each other directly or indirectly through one or more intermediaries.
16. 1. A non-transitory computer-readable device having stored thereon instructions that, when executed by a first server, cause the first server to perform operations for evaluating a decision tree in a privacy-preserving manner, the operations comprising: receiving from a client a first partial share of the decision tree and encrypted input data, the encrypted input data including a set of attributes; communicating with a second server that received a second partial share of the decision tree to compute a classification result for the encrypted input data, the first server and the second server each being configured to compute the classification result for the encrypted input data using the respective partial shares of the decision tree received by the first server and the second server and secure multi-party computation methods including additive homomorphic encryption and oblivious transfer, such that the classification result is determined without decrypting the encrypted input data and such that the classification result is ultimately received by the first server; and transmitting the classification result to the client.
17. 20. The non-transitory computer readable device of claim 16, wherein the decision tree comprises a full binary decision tree.
18. the decision tree includes a plurality of decision nodes and a plurality of leaf nodes; the first partial share of the decision tree includes a tuple of thresholds associated with each decision node of the decision tree and a tuple of attribute indices associated with each decision node of the decision tree; 17. The non-transitory computer-readable device of claim 16, wherein the second partial share of the decision tree includes a description of a structure of the decision tree and tuples of classification labels respectively associated with the leaf nodes of the decision tree.
19. the classification label tuple is encrypted with a first public key paired with a first private key accessible to the first server but not to the second server; 20. The non-transitory computer-readable device of claim 18, wherein the threshold tuple is encrypted with a second public key paired with a second private key that is accessible to the second server but not to the first server.
20. 17. The non-transitory computer-readable device of claim 16, wherein the set of attributes is associated with a user, the decision tree includes a classification model trained on a set of one or more previously obtained attributes associated with the user, and the classification result indicates whether the user should be authenticated or not authenticated.