Tree model two-party joint reasoning method and computer equipment
By homomorphically encrypting the feature numbers of non-leaf nodes of the tree model and combining with the secret sharing protocol, the joint inference problem between the two parties of the tree model under the condition of anonymized primary key is solved, and a safe and efficient inference process is achieved.
Patent Information
- Application Number
- CN202510201328.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The prior art lacks a method for performing joint inference between the two parties of tree model under the condition of anonymized primary keys, especially in the scenario where the feature number of split features of non-leave nodes is held by non-featured parties and the split threshold is held by both parties, and it is difficult to achieve a safe and efficient inference process.
By homomorphically encrypting the feature number vectors of non-leaf nodes of the tree model, and obtaining the eigenvalue fragmentation of the target object split feature through homomorphic operation with the target object's eigenvalue. Then, the secret comparison module and the leaf node array determine the module, combined with the preset leaf node prediction value, perform operations based on secret sharing, and obtain the target object's prediction result sharding.
It realizes the joint inference of the two-party tree model under the condition of anonymized primary key, ensuring the improvement of data privacy protection and inference efficiency.
Smart Images

Figure CN119990328A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification belong to the field of privacy computing technology, and in particular, relate to a tree model two-party joint reasoning method and computer equipment. Background Art
[0002] A decision tree model is a tree structure. Through the decision tree model, the attributes of the target object can be scored or classified according to the input feature values of the target object to obtain the final reasoning result. With the requirements of data security, privacy computing methods that take into account privacy protection requirements and data utilization have been widely used. In related technologies, multiple participants use their respective data to perform joint privacy computing, such as jointly modeling a tree model, jointly reasoning on a tree model, etc.
[0003] As the requirements for data security in various industries continue to increase, the requirements for fully anonymous privacy-preserving machine learning are further required, so that a more secure tree model can be built under the conditions of anonymous data primary key information and full data process encryption. Among them, the primary key refers to the identification information that can distinguish the data. However, the tree model for anonymous primary key modeling in related technologies lacks a two-party joint reasoning method. Summary of the invention
[0004] The purpose of this specification is to provide a tree model two-party joint reasoning method and computer device.
[0005] In a first aspect, the present specification provides a two-party joint reasoning method for a tree model, which is applied to a first party. The method also involves a second party. The first party and the second party respectively hold feature values of different features of a target object. The first party and the second party jointly hold a tree model. The feature number of the split feature of a non-leaf node of the tree model is held by the party that does not hold the feature value of the split feature. The feature threshold of the split feature is held by the first party and the second party in the form of fragments. The method includes:
[0006] For any non-leaf node, obtain the homomorphic ciphertext of the feature number vector of the non-leaf node, and jointly with the second party obtain the feature value shards of the split feature of the target object based on the homomorphic ciphertext of the feature number vector and the feature values of each feature of the target object; and jointly with the second party, perform secret comparison based on the shards of the feature value of the split feature and the shards of the feature threshold to obtain the shards of the secret comparison result; the secret comparison result is used to indicate the node into which the target object falls after the non-leaf node; the feature number vector is a one-hot vector with a value of 1 at the corresponding position of the feature number of the non-leaf node;
[0007] Determine, in conjunction with the second party, a slice of a leaf node array corresponding to the target object based on the slices of the secret comparison results of each non-leaf node; the leaf node array is used to indicate the leaf node into which the target object falls;
[0008] The second party jointly performs a secret sharing-based operation based on the slices of the leaf node array and the slices of the preset prediction values of each leaf node to obtain the slices of the prediction results of the target object.
[0009] The second aspect of the present specification provides a two-party joint reasoning device for a tree model, wherein a first party and a second party respectively hold feature values of different features of a target object; the first party and the second party jointly hold a tree model, the feature number of the split feature of a non-leaf node of the tree model is held by the party that does not hold the feature value of the split feature, and the feature threshold of the split feature is held by the first party and the second party in the form of slices; the device comprises:
[0010] A secret comparison module is used to obtain the homomorphic ciphertext of the feature number vector of any non-leaf node, and jointly with the second party obtain the feature value slices of the split feature of the target object based on the homomorphic ciphertext of the feature number vector and the feature values of each feature of the target object; and jointly with the second party, perform secret comparison based on the slices of the feature value of the split feature and the slices of the feature threshold to obtain the slices of the secret comparison result; the secret comparison result is used to indicate the node into which the target object falls after the non-leaf node; the feature number vector is a one-hot vector with a value of 1 at the corresponding position of the feature number of the non-leaf node;
[0011] A leaf node array determination module, used to determine, in conjunction with the second party, a slice of a leaf node array corresponding to the target object based on the slice of the secret comparison result of each non-leaf node; the leaf node array is used to indicate the leaf node into which the target object falls;
[0012] The prediction result acquisition module is used to jointly with the second party perform a secret sharing-based operation according to the slices of the leaf node array and the slices of the preset prediction values of each leaf node to obtain the slices of the prediction results of the target object.
[0013] A third aspect of the present specification provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the above-mentioned tree model two-party joint reasoning method.
[0014] A fourth aspect of the present specification provides a computing device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the above-mentioned tree model two-party joint reasoning method is implemented.
[0015] A fifth aspect of the present specification provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned tree model two-party joint reasoning method.
[0016] Based on this, the present specification provides a two-party joint reasoning method for a tree model. The method involves a first party and a second party, and the first party and the second party respectively hold different types of features of the target object. The first party and the second party also jointly hold a tree model, in which the feature number of the split feature of the non-leaf node is held by the non-feature party, and the split threshold is held by the two parties in the form of secret sharing shards. During the reasoning process, for each non-leaf node, according to the homomorphic ciphertext of the non-leaf node feature number vector and the homomorphic operation between the feature value of the target object, and through the homomorphic encryption to secret sharing protocol, the shard of the feature value of the split feature of the target object is obtained. And through the secret comparison jointly performed by the two parties, according to the shard of the split threshold of the non-leaf node, the shard of the secret comparison result indicating which child node the target object falls into is determined. The two parties then jointly determine the shard of the leaf node array of the leaf node indicating the target object falls into according to the shard of the secret comparison result of each non-leaf node. Finally, the two parties jointly determine the prediction result of the target object according to the shard of the leaf node array and the prediction value of each leaf node.
[0017] The above method targets the scenario where the feature number of the split feature of the non-leaf node is held by the non-feature party and the split threshold is held by both parties. By homomorphically encrypting the feature number vector of the non-leaf node of the tree model and performing homomorphic operations with the feature value of the target object, the ciphertext of the feature value of the split feature is obtained. Thus, the reasoning of the non-leaf node is completed based on the ciphertext of the feature value and the ciphertext of the split threshold, and the predicted value acquisition for the target object is completed based on this. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of this specification, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0019] Figure 1 is a schematic diagram of a tree model in one embodiment;
[0020] Figure 2 is a flow chart of a two-party joint reasoning method of a tree model in one embodiment;
[0021] Figure 3 is a flow chart of a two-party joint reasoning method of a tree model in another embodiment;
[0022] Figure 4 is a schematic diagram of a tree model prediction path in one embodiment;
[0023] Figure 5 It is a block diagram of a two-party joint reasoning device of a tree model in one embodiment. DETAILED DESCRIPTION
[0024] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.
[0025] The decision tree model generally exists in the form of a binary tree. The tree model includes leaf nodes and non-leaf nodes. When making inferences, the leaf node where the sample falls is determined based on the splitting conditions of the non-leaf nodes, and the score or classification of the sample is determined based on the weight value of the leaf node where the sample falls. The splitting conditions include the splitting features and the splitting thresholds corresponding to the splitting features. For example, Figure 1 As shown, Figure 1 It is a simple decision tree. The content in the non-leaf node represents the splitting condition of the node, A is the splitting feature, and the number is the corresponding splitting threshold. The content in the leaf node represents the weight value of the leaf node. When the sample is smaller than the corresponding splitting threshold, it will fall into the child node on the left side of the node (hereinafter referred to as the left child node), otherwise it will fall into the child node on the right side (hereinafter referred to as the right node), until the sample falls into a leaf node. For example, if the feature value of feature A of a sample is 1, then according to the splitting condition judgment of each non-leaf node, it can be determined that the sample falls into the leftmost leaf node, then the score value of the sample is 1. It should be noted that Figure 1 What is shown is a simple decision tree model. The decision tree model actually used may not generate the judgment conditions of non-leaf nodes based on only one feature, but may also complete the classification or scoring of samples through multiple trees at the same time.
[0026] There are many training methods for tree models, generally including Gradient Boosting DecisionTree (GBDT), Extreme Gradient Boosting (XGBoost), etc. Here we will briefly introduce GBDT. GBDT is an iterative decision tree algorithm. By iteratively training multiple decision trees, each decision tree is trained on the residual (the difference between the predicted value and the actual value) of the previous tree. During the training process, the structure and parameters of each decision tree are continuously optimized by the gradient boosting method to minimize the loss function of the model. In the prediction stage, the prediction results of multiple decision trees are summed or weighted averaged to obtain the final prediction result. Gradient decision trees are usually used to solve regression and classification problems. For example, a tree model can be trained to classify the financial credit risk of users, so as to quickly identify users with credit risks (such as fraud risks).
[0027] In the scenario of jointly training tree models with privacy protection, the participants (two examples are used here) are generally required to determine the common samples, that is, the intersection samples. Specifically, each participant generally holds the primary key and feature value corresponding to each training sample, and different participants hold different feature values. The two participants can determine the intersection samples with the same primary key based on the primary key.
[0028] Among them, the primary key is the identification information that can distinguish different training samples. For example, in the scenario of classifying the financial credit risk of users, the primary key can be the user's identity information, mobile phone number, and other identification information that can uniquely distinguish the user. Different participants hold different types of features. For example, in the scenario of classifying the financial credit risk of users, one participant can hold the transaction data between the user and other users, and another participant can hold the user's friend information, such as which user the user has a friend relationship with.
[0029] The method for determining the intersection samples during the tree model training process is generally determined through private set intersection (PSI). PSI can ensure that the primary key information and feature value information of non-intersection samples are not leaked. After PSI is executed, the two participants can determine the primary key information of the intersection samples, and the two participants can further jointly complete the tree model training based on the intersection samples.
[0030] As data privacy continues to strengthen, data holders have more stringent data demands, hoping to protect their intersection primary key information from being leaked during the modeling process.
[0031] Based on this, this specification provides a method for training a tree model with an anonymous primary key. Next, the implementation process of this method will be described through a specific embodiment.
[0032] This method is applicable to the decision tree model training method in which the calculation process data is encrypted and the primary key information cannot be inferred after the primary key information has been subjected to the circuit-PSI and the anonymous primary key privacy bucketing. Both parties hold the secret shared ciphertext output by the full anonymous bucketing protocol: feature value sharding, feature bucket number sharding, and bucket boundary marker value sharding.
[0033] Now assume that there are two participants P holding different feature data sets A and P B , where P A Own A Features, P B Own m B The two parties use their own data sets to train the decision tree model. The two parties use the existing technologies of anonymous key privacy intersection and anonymous key privacy bucketing protocol to obtain the secret share matrix of data alignment. <d>=< B1 ||U1||X1||...||B m ||U m ||X m ||Y>, the matrix size is n×(3m+1), where n is the number of data intersections, and m=m A +m B The total number of features. The first 3m columns indicate that each feature contains bucket number fragments. i >, bucket boundary mark value sharding i >(the plaintext corresponding to the minimum value in each bucket is 1, and the plaintext of the remaining values is 0), feature value sharding <X i >These three columns, the last column is the label value fragment <y>,|| represents the connector of each column of data in the matrix. A Features belong to P A , after m B Features belong to P B . Each feature will be divided into nb buckets.
[0034] Note: In this manual, angle brackets are used <x>Indicates that the current variable X is encrypted by secret sharing, and P A and P B Hold secret sharing shards separately <x> A and <x> B In this manual, brackets [X] are used to indicate A Indicates that the current variable X is homomorphically encrypted, and the A in the upper right corner indicates that the private key of the ciphertext belongs to P A , only P A The ciphertext can be decrypted.
[0035] 1. Extracting data by bucket
[0036] The purpose of this step is to determine the corresponding bucket thresholds when the primary key is anonymized. Specifically, A needs to obtain the ciphertext of the bucket thresholds of each bucket of B. Correspondingly, B needs to obtain the ciphertext of the bucket thresholds of each bucket of A. The specific implementation is detailed below.
[0037] (1)P A from <d> A Extract the fragments and send it to P B .P B Use the received fragments and the corresponding fragments you hold locally Perform secret sharing decryption to make P B Get the top m A The bucket number of the intersection data features and bucket boundary marker values
[0038] (2)P A from <d> A Extract the fragments Obtained using homomorphic encryption and send it to P B .P B Use the received fragments and the corresponding fragments you hold locally Perform accumulation operation to obtain the ciphertext under homomorphic encryption only
[0039] (3)P B according to and You can get the first m A features, the bucket boundary value position corresponding to each bucket number, and then extract The ciphertext value at the corresponding position in is used as the secret share of each bucket threshold Each array S i The size of is nb.
[0040] (4) Similar to (1-3), P B send and Give P A , P A Get the secret share of the bucket threshold
[0041] (5)P B Initialize the prediction array And encrypt it through secret sharing into a secret shard held by both parties
[0042] 2. Decision tree model training
[0043] (1) For the kth decision tree, both parties perform the following operations:
[0044] (1.1) Both parties use And the gradient derivative formula g = σ(p k )-y,h=(1-σ(p k ))·σ(p k ), in the secret sharing scenario, the derivative value is calculated { <g i >, <h i >} i∈[1,n] .
[0045] (1.2)P A Initialize the sample distribution of the first node (root node) locally is a vector of size n containing all 1s. Similarly, P B Initialize all 1 vectors of size n locally
[0046] (1.3) For the lth non-leaf node in the tree model, both parties perform the following operations:
[0047] (1.3.1)P B In local <g i > B , <h i > B } i∈[1,n] Perform homomorphic encryption and obtain {[ <g i > B ] B ,[ <h i > B ] B } i∈[1,n] and send it to P A .P A The local { <g i > A , <h i > A } i∈[1,n] and{[ <g i > B ] B ,[ <h i > B ] B } i∈[1,n] Add together to obtain the ciphertext array of homomorphic encryption only {[g i ] B ,[h i ] B } i∈[1,n]
[0048] (1.3.2)P A Initialize all 0 gradient bucket accumulation matrix Then, P A According to the bucket information and sample distribution information Will satisfy Corresponding [ <g j > B ] B ,[ <h j > B ] B Add to Among them, B i,j The bucket number corresponding to the i-th feature of the j-th sample is bucket.
[0049] (1.3.3)P A Using the method of converting homomorphic ciphertext to secret sharing ciphertext, Become a secret sharing shard held by both parties
[0050] (1.3.4) Similar to (1.3.1-1.3.3), P A In local <g i > B , <h i > B } i∈[1,n] Perform homomorphic encryption, P B Through the gradient bucket accumulation and homomorphic ciphertext to secret sharing ciphertext, both parties can obtain
[0051] (1.3.5) Merge the two gradient bucket accumulation matrices to obtain { <G i,q >, <H i,q >} i∈[1,m],q∈[1,nb] .
[0052] (1.3.6) Both parties use the decision tree calculation rules and the tree model gain value calculation formula under the MPC scenario Get the corresponding { <Gain i,q >} i∈[1,m],q∈[1,nb] ,in
[0053] (1.3.7) Both parties use the privacy extreme value solution method to obtain { <Gain i,q >} i∈[1,m],q∈[1,nb] The secret shard of the bucket number corresponding to the maximum value<max_idx> , and through privacy comparison, both parties can obtain<max_idx> and m A ×nb, if<max_idx> Greater than m A ×nb, it means that the sample distribution information corresponding to the feature and the bucket threshold are determined by P A Management requires<max_idx> Exclusively for P A Otherwise,<max_idx> Exclusively for P B .
[0054] (1.3.8) Assume P B Get the value max_idx, then P B We can know that the splitting boundary of the lth node is the jth bucket of the ith feature through max_idx=i×nb+j. B The distribution of n samples with bucket number j as the boundary Update and determine the sample distribution information of the left and right child nodes 2l and 2l+1 of the lth node, that is, Then extract [S i,j ] A , and convert it into a secret sharing shard value <T k,l >, as the splitting threshold of the current node.
[0055] (1.4) For the lth leaf node in the tree model, both parties perform the following operations:
[0056] (1.4.1) Assume that the split information of the parent node of the lth leaf node is given by P B Hold, then P B The formula for calculating leaf node weight can be obtained directly through max_idx in (1.3.8) In the required <g>and <h>exist{ <Gl i,q >, <Hl i,q >} i∈[1,m],q∈[1,nb] The corresponding position d in .
[0057] (1.4.2) However, since P A The feature information corresponding to the parent node and the location information of the maximum gain value are unknown, so both parties need to obtain it through an nb choose 1 oblivious transfer protocol OT. <g>and <h>Specifically, first P A Generate two random numbers rg and rh and calculate locally { <Gl i,q -rg> A , <Hl i,q -rh> A } i∈[1,m],q∈[1,nb] , and then use the array as the input of the OT protocol; at the same time, P B Input d in the OT protocol and get the output of the OT protocol <Gl d -rg> A , <Hl d -rh> A Then, P B Compute locally <g> B ←<Gl d -rg> A +<Gl d > B , <h> B ← <Hl d -rg> A + <Hl d > B ;P A Directly use rg and rh as secret shards <g> A ←rg, <h> A ←rh.
[0058] (1.4.3) Both parties use the formula Get the weight value of the leaf node of the tree model.
[0059] (1.5) For each sample i, both parties pass Update the prediction results, where f k (i) is the leaf node number corresponding to the i-th sample, which can be obtained by bit-multiplying each leaf node and Get f k (i) is a one-hot encoded vector, and then the vector is compared with <w>Dot product As the weight value of the leaf node.
[0060] (2) Finally, both parties output the feature information of the splitting point of each tree and the splitting threshold fragmentation as the model output result of this tree model training.
[0061] In the tree model obtained by the above training process, the same as the tree model jointly trained by two parties in the non-anonymous primary key scenario in the related art is that the predicted value of the leaf node is held by the two parties in the form of shards. However, it is different from the tree model jointly trained by two parties in the related art in that the feature number of the split feature of the non-leaf node is held by the non-feature party, and the split threshold of the split feature is jointly held by the two participating parties in the form of secret sharing shards.
[0062] Among them, secret sharing is also called secret splitting and secret sharing. The basic principle is to split the secret into multiple shares (shares), which are also called shards, and give them to different people for safekeeping. Only when more than a threshold number of people merge their shares can the secret be recovered; no information about the secret can be recovered from shares less than the threshold number. In privacy computing, the threshold number is usually the same as the number of participants. For two participants P0 and P1, arithmetic secret sharing is performed by splitting the plaintext value x of length l bits into two random number shares <x>0∈2 l and <x>1∈2 l Stored by P0 and P1 respectively, and satisfying x=( <x> 0+ <x>1) mod 2 l The multiple data obtained through secret sharing can be called secret sharing slices, which are also referred to as slices in this specification. Secret sharing also supports two parties to jointly perform four arithmetic operations of addition, subtraction, multiplication and division based on the secret sharing slices.
[0063] A non-feature party is a party that does not hold a feature of the split feature type. For example, if the split feature is the user's account registration time, which is held by the first party, and the second party does not hold the user's account registration time, the second party is a non-feature party. The feature number of the split feature is the number of the split feature among all features owned by the first party and the second party.
[0064] In the tree model of the related art, the feature number of the non-leaf node is generally held by the feature party, and the corresponding feature party also holds the split threshold of the non-leaf node. The difference between the tree model trained in this specification and the tree model in the related art makes the tree model reasoning method in the related art unable to be applied to the scenario targeted by this specification.
[0065] Based on this, the present specification provides a two-party joint reasoning method for a tree model. The method involves a first party and a second party, and the first party and the second party respectively hold different types of features of the target object. The first party and the second party also jointly hold a tree model, in which the feature number of the split feature of the non-leaf node is held by the non-feature party, and the split threshold is held by the two parties in the form of secret sharing shards. During the reasoning process, for each non-leaf node, according to the homomorphic ciphertext of the non-leaf node feature number vector and the homomorphic operation between the feature value of the target object, and through the homomorphic encryption to secret sharing protocol, the shard of the feature value of the split feature of the target object is obtained. And through the secret comparison jointly performed by the two parties, according to the shard of the split threshold of the non-leaf node, the shard of the secret comparison result indicating which child node the target object falls into is determined. The two parties then jointly determine the shard of the leaf node array of the leaf node indicating the target object falls into according to the shard of the secret comparison result of each non-leaf node. Finally, the two parties jointly determine the prediction result of the target object according to the shard of the leaf node array and the prediction value of each leaf node.
[0066] The above method targets the scenario where the feature number of the split feature of the non-leaf node is held by the non-feature party and the split threshold is held by both parties. By homomorphically encrypting the feature number vector of the non-leaf node of the tree model and performing homomorphic operations with the feature value of the target object, the ciphertext of the feature value of the split feature is obtained. Thus, the reasoning of the non-leaf node is completed based on the ciphertext of the feature value and the ciphertext of the split threshold, and the predicted value acquisition for the target object is completed based on this.
[0067] Next, the tree model two-party joint reasoning method shown in this specification will be described through an embodiment.
[0068] The method in this specification is applied to the scenario of joint reasoning between two parties. For the convenience of explanation, the two parties are respectively referred to as the first party and the second party. The first party and the second party respectively hold feature values of different features of the target object. In other words, for the target object, the first party and the second party respectively hold feature values of different types of features of the target object. The target object here is an object held by both parties. For example, if the target object is a user, the first party can hold feature values of features such as the user's account registration time and the number of interactions with high-risk accounts, and the second party can hold feature values of features such as the number of transactions of the user and whether the user has traded with high-risk users.
[0069] The two parties also jointly hold a tree model. The tree model is the same as the tree model jointly trained by the other two parties in that the predicted values of the leaf nodes of the tree model are held by the two parties in the form of secret sharing shards. Unlike the tree model jointly trained by the other two parties in the non-anonymous primary key scenario in the related art, for any non-leaf node of the tree model, the feature number of the split feature of the non-leaf node of the tree model is held by the party that does not hold the feature value of the split feature, and the feature threshold of the split feature is held by the first party and the second party in the form of shards.
[0070] For example, if the split feature of a non-leaf node is the user's account registration time, the user's account registration time is numbered 3 among all the features of the target object held by the first party and the second party. The feature value of the target object's account registration time is held by the first party, and the feature number 3 of the non-leaf node split feature is held by the second party. At the same time, both parties hold the secret shared shard of the split threshold of the split feature.
[0071] In addition, the two parties can jointly hold multiple tree models. For the sake of convenience, only the reasoning process of one tree model is described here. It is easy to understand that the reasoning process thresholds of other tree models are the same. After obtaining the prediction results of multiple tree models, the prediction results of multiple tree models can be added together to obtain the final prediction result.
[0072] In addition, the two participants can also use the tree model to infer the prediction results of multiple target objects at the same time. Here, the reasoning process of one target object is taken as an example. The reasoning process of other target objects is the same and will not be repeated.
[0073] In an optional embodiment, the tree model is obtained by training in conjunction with the second party based on a target training sample set, and the target training sample set is a training sample set obtained by executing a privacy intersection protocol of an anonymous primary key on the data of the object held in conjunction with the second party. In other words, the above-mentioned tree model can be obtained by a joint modeling method of a tree model with an anonymous primary key, for example, it can be obtained by training using the method mentioned above. This scenario does not limit this specification, and it is understood that if a tree model with the above-mentioned characteristics can be obtained by other modeling methods, the method provided in this specification can also be used to perform joint reasoning between the two parties.
[0074] Next, we will combine Figure 2 Here, the steps of the method shown in this specification are described. Here, the steps performed by the first party are described. The steps performed by the second party are the same as the steps performed by the first party, and are not repeated here.
[0075] Step 201, for any non-leaf node, obtain the homomorphic ciphertext of the feature number vector of the non-leaf node, and jointly with the second party obtain the feature value shards of the split feature of the target object based on the homomorphic ciphertext of the feature number vector and the feature values of each feature of the target object; and jointly with the second party perform secret comparison based on the shards of the feature value of the split feature and the shards of the feature threshold to obtain the shards of the secret comparison results.
[0076] Among them, the secret comparison result is used to indicate the node where the target object falls after the non-leaf node; the feature number vector is a unique hot vector with a value of 1 at the corresponding position of the feature number of the non-leaf node.
[0077] Specifically, for any non-leaf node, the feature number of the split feature of the non-leaf node is held by the non-feature party, and the party that owns the feature value of the split feature cannot obtain which type of feature is the split feature. Therefore, in order to complete the reasoning of non-leaf nodes, it is first necessary to obtain the ciphertext of the feature value of the split feature of the target object for each non-leaf node.
[0078] In order to solve the above problem, a method of homomorphically encrypting the feature numbers of non-leaf nodes of the tree model is adopted in step 201. First, the feature party mentioned below is the party that owns the split feature of the non-leaf node, and the non-feature party is the party that does not own the split feature of the non-leaf node. Specifically, first, the non-feature party constructs a feature number vector according to the feature number of the split feature of the non-leaf node held by the non-feature party. The number of elements of the feature number vector is the same as the total number of features. The vector is a unique hot vector in which the element at the corresponding position of the feature number is 1 and the other elements are 0. The non-feature party can homomorphically encrypt each element of the feature number vector and send it to the feature party. The feature party performs homomorphic multiplication operations on each element in the feature number vector with each feature held by the party, and performs homomorphic addition operations on each homomorphic multiplication result to obtain a homomorphic ciphertext of the target object feature value. The two parties then jointly execute the homomorphic encryption to secret sharing protocol to obtain the shards of the target object feature value respectively.
[0079] For ease of understanding, let's first explain homomorphic encryption. Homomorphic encryption is an encryption method commonly used in privacy computing. Compared with ordinary encryption, homomorphic encryption has the property of homomorphism, that is, it can perform calculations on data in an encrypted state. For example, by performing a homomorphic addition operation on the homomorphic encrypted ciphertext [a] of a and the unencrypted data b, the ciphertext [a+b] of a+b can be obtained, and a, b or a+b will not be leaked during the operation, and the decryption key does not need to be used. Moreover, homomorphic encryption also supports [a]*b=[a*b], and the fully homomorphic encryption algorithm also supports [a]*[b]=[a*b].
[0080] The homomorphic multiplication operation between the homomorphic ciphertext of the feature number vector and the eigenvalue of each feature in step 201 is to multiply the elements at each position of the feature number vector by the eigenvalue corresponding to the element. Since in the plaintext of the feature number vector, only the element at the position corresponding to the feature number is 1, and the other elements are 0, in the vector obtained by the homomorphic multiplication operation, the plaintext at the position corresponding to the feature number is the eigenvalue corresponding to the feature number, and the plaintext corresponding to the elements at other positions is 0. Then, the elements at each position of the vector are subjected to homomorphic addition operation, that is, the eigenvalue of the split feature can be obtained while ensuring that the information such as the eigenvalue and the feature number is not leaked.
[0081] It should also be noted that the reason why the above needs to be converted into the form of secret sharing shards and then secret comparison is performed to complete the reasoning of non-leaf nodes, rather than comparison in the form of homomorphic ciphertext, is because the secret comparison scheme of secret sharing in the relevant technology is more efficient than the comparison scheme of homomorphic ciphertext. Converting it into the form of secret sharing shards can improve the overall reasoning efficiency.
[0082] In addition, it should be noted that in order to ensure data security, the non-feature party uses the non-feature party's public key to encrypt. Correspondingly, the ciphertext can only be decrypted with the non-feature party's private key, and the feature party cannot obtain the non-feature party's private key and cannot decrypt it, ensuring data security.
[0083] In different executions, the feature party and the non-feature party may be different. For example, for the first non-leaf node, the first party may be the feature party, and for the second non-leaf node, the second party may be the feature party.
[0084] Next, step 201 will be described in detail. Step 201 includes two steps: obtaining feature value fragments of the split feature and secret comparison.
[0085] Next, we will take a non-leaf node as an example, and take the first party as the feature party and the first party as the non-feature party as examples to illustrate the process of how the first party and the second party jointly obtain the feature value fragmentation of the split feature of the target object.
[0086] First, the first party is a non-characteristic party. Step 201 specifically includes: for any non-leaf node, when the first party holds the characteristic number of the non-leaf node, constructing a characteristic number vector according to the characteristic number, and homomorphically encrypting the characteristic number vector.
[0087] The second party obtains the homomorphic ciphertext of the feature number vector to the second party; the second party obtains the homomorphic ciphertext of the feature value of the split feature of the target object according to the homomorphic operation between the homomorphic ciphertext of the feature number vector and the feature value of each feature of the target object held by the second party;
[0088] Cooperate with the second party to execute a homomorphic encryption-to-secret sharing protocol based on the homomorphic ciphertext of the characteristic value of the split characteristic of the target object to obtain fragments of the characteristic value of the split characteristic of the target object.
[0089] Among them, when performing homomorphic encryption, each element in the feature number vector is homomorphically encrypted separately, rather than homomorphically encrypting the feature number vector as a whole. The first step is the process of constructing the homomorphic ciphertext of the feature number vector, the second step is the process of the second party obtaining the homomorphic ciphertext of the eigenvalue of the split feature, and the third step is the process of converting the homomorphic ciphertext of the eigenvalue into a secret sharing shard.
[0090] Second, the first party is a feature party. Step 201 specifically includes: for any non-leaf node, when the second party holds the feature number of the non-leaf node, receiving the homomorphic ciphertext of the feature number vector sent by the second party;
[0091] Obtaining the homomorphic ciphertext of the eigenvalue of the split feature of the target object according to the homomorphic operation between the homomorphic ciphertext of the feature number vector and the eigenvalue of each feature of the target object held by the party;
[0092] A homomorphic ciphertext of the characteristic value of the split characteristic of the target object is sent to the second party, and a homomorphic encryption-to-secret sharing protocol is jointly executed with the second party to obtain a shard of the characteristic value of the split characteristic of the target object.
[0093] The first step is the process of receiving the feature number vector for homomorphic encryption. The second step is the process of calculating the homomorphic ciphertext of the split feature eigenvalue. The third step is the process of the two parties jointly converting the homomorphic ciphertext of the eigenvalue into a secret sharing shard.
[0094] Next, a possible implementation of a homomorphic encryption to secret sharing protocol will be described. Specifically, the feature party obtains the homomorphic ciphertext of the eigenvalue of the split feature by calculation, and can perform a homomorphic addition operation on the homomorphic ciphertext of the eigenvalue and a preset random number to obtain the first element. The first element is sent to the non-feature party, and the non-feature party decrypts the first element to obtain a shard of the eigenvalue of the split feature. Correspondingly, the feature party uses the opposite number of the above random number as the shard of the eigenvalue of the split feature.
[0095] After obtaining the eigenvalue shards of the target object's split feature, the two parties can jointly perform a secret comparison. Based on the relationship between the eigenvalue and the split threshold of the split feature, the two parties can determine whether the target object falls on the left child node or the right node of the non-leaf node, and then obtain the shards of the secret comparison results.
[0096] In an optional implementation, the secret comparison result can be implemented in a binary manner. For example, if the secret comparison result is a first value (such as 0), the target object falls on the left child node of the non-leaf node, and if the secret comparison result is a second value (such as 1), the target object falls on the right node of the non-leaf node. It should be noted that the first value here and the first value of the path value later may have different meanings.
[0097] As for the specific implementation method of secret comparison, it can be implemented based on secret comparison protocols such as obfuscation circuits and bit decomposition schemes.
[0098] Next, we will combine Figure 3 The flowchart shown is used to explain step 201 in detail.
[0099] Figure 3 First, the two participants P A and P B Both parties hold a set of prediction data D with aligned primary key identifiers and different features. A and D B It can be understood that each sample in the prediction data set can be regarded as the target object mentioned above.
[0100] Where D A Contains n samples and m A Features, D B Contains n samples and m B Features. A [idx][q] represents the data set D A The qth eigenvalue of the idxth sample in P. A Held m A Features and P B Held m B The characteristics are different.
[0101] In addition, both parties hold a privacy-protected decision tree model Tree A and Tree B Specifically, the decision tree has a total of t trees, and the depth of each tree is d, that is, the number of non-leaf nodes in each tree is p = 2 d -1, the number of leaf nodes is l = 2 d .
[0102] For the i-th non-leaf node (i∈[p]), it contains the ciphertext fragment of the feature number f[i] and the feature threshold<h[i]> , where the feature number is held by the non-feature party and the threshold is a secret share encrypted with a secret share<h[i]> A ,<h[i]> B For example, the feature number corresponding to the feature belongs to P A , then only P B Knowing the feature number f[i], in order to distinguish the owner of the number, let f B [i]=f[i],f A [i] = 0; otherwise, if it is not a feature, it is P A , then let f A [i]=f[i],f B [i]=0.
[0103] For the jth leaf node (j∈[l]), the corresponding leaf node weight information<w[j]> , also with a secret share<w[j]> A ,<w[j]> B The data is stored locally on both parties.
[0104] As before, this specification will use angle brackets when describing a basic secret sharing operation. <x>Indicates that the current variable X is encrypted by secret sharing, and P A and P B Hold secret sharing shards separately <x> A and <x> B , both parties need to <x> A and <x> B Perform the corresponding secret sharing operation. This article uses square brackets [X] A Indicates that the current variable X is P A Homomorphic public key encryption, the A in the upper right corner indicates that the private key of the ciphertext belongs to P A , only P A The ciphertext can be decrypted.
[0105] Figure 3 The content in the blue dotted box is the reasoning process for the k-th tree model, and idx is the identifier of the target object. Specifically, in step 201, the following is executed:
[0106] 1. Both parties generate their own homomorphically encrypted public and private keys locally and send the public key to each other. Specifically, P A Generate homomorphic encryption public key pk A and private key sk A , and the public key pk A Send to P B Similarly, P B Generate homomorphic encryption public key pk B and private key sk B , and the public key pk B Send to P A .
[0107] 2. Both parties use homomorphic encryption technology to perform one-hot encoding homomorphic encryption on the feature number information f[i] of all non-leaf nodes in the tree model, and send the ciphertext to the feature party. Specifically, for the i-th non-leaf node (i∈[p]), assume that the node feature number information f[i] is obtained by the non-feature party P B Hold, that is, f B [i] = f[i]. First, P B f B [i] becomes a one-hot encoded array U[i] (U[i] is a length of m A , and only U[i][f B [i]]=1, and the rest of the values are 0), and use the public key pk B The encrypted array U[i] is After completing the encryption operations on all leaf nodes, P A Holds the ciphertext matrix P B Holds the ciphertext matrix Both parties will and For the idx-th test data (i.e., target object), the k-th decision tree (idx∈[n],k∈[t]), both parties perform the following operations.
[0108] 3. For the i-th non-leaf node (i∈[p]), the feature square uses the ciphertext and the characteristics of our test sample D A(B) [idx] Calculate the ciphertext directly locally Then, the homomorphic ciphertext-to-secret sharing technique is used to obtain <D A(B) [idx][f[i]]>.
[0109] Specifically, assume that the feature square of the current i-th non-leaf node is P A , P A hold and D A [idx], first P A Using homomorphic multiplication and homomorphic addition of plaintext and ciphertext, we can calculate the homomorphic ciphertext Then, P A Generate a random number r locally and use the public key pk B Encrypted as Then, perform homomorphic addition to obtain Then, P A Will Send to P B , P B Using the private key sk B Decrypt the pseudo-random number D A [idx][f B [i]]+r; finally, P A Directly use the random number -r as the secret sharing ciphertext of the eigenvalue corresponding to the i-th non-leaf node, denoted as<v[i]> A = -r, and P B D A [idx][f B [i]]+r is used as the secret shared ciphertext of the eigenvalue of the i-th non-leaf node, denoted as<v[i]> B =D A [idx][f B [i]]+r.
[0110] 4. For the i-th non-leaf node (i∈[p]), both parties need to protect the ciphertext feature value<v[i]> and ciphertext threshold<h[i]> Without being leaked, the secret decision of the target object in the non-leaf node is obtained by secret comparison, and the ciphertext comparison result is secretly shared in the form of bits.<b[i]> bit express.
[0111] Specifically, for the i-th secret comparison, P A enter<v[i]> A and<h[i]> A , P B enter<v[i]> B and<h[i]> B After both parties have executed the agreement, P A get P B get And meet<v[i]> ≤<h[i]> (that is, it falls into the left child node), This means that the correct decision path is to move forward to the left child node.<v[i]> ><h[i]> ,but This means that the correct decision path is to move forward to the right child node. After p times of the above calculations, both parties can finally obtain the ciphertext array of the decision results of all non-leaf nodes in the kth tree. bit Here and This is the shard of the secret comparison result.
[0112] in, Represents XOR operation. The calculation logic of XOR operation is that when two binary bits are the same, the result is 0, and when two binary bits are different, the result is 1. The bit in the upper right corner of the above angle bracket represents that it is a bit-type secret sharing shard. The difference between bit-type secret sharing shard and numeric secret sharing shard is that bit-type secret sharing shard ensures that the XOR operation result of two shards is the corresponding plaintext.
[0113] Step 203, jointly with the second party, based on the sharding of the secret comparison results of each non-leaf node, and jointly with the second party, determine the sharding of the leaf node array corresponding to the target object.
[0114] The leaf node array is used to indicate the leaf node into which the target object falls.
[0115] Specifically, after obtaining the shards of the secret comparison results of each non-leaf node, it is possible to jointly determine with the second party the leaf node array of which leaf node the target object falls according to the connection relationship between the nodes.
[0116] The number of elements in the leaf node array can be the same as the number of leaf nodes. In the leaf node array, the value of the element corresponding to the leaf node where the target object falls is the first value, and the value of the element corresponding to other leaf nodes is the second value. For example, if the target object falls into the third leaf node, the first value is 1, and the second value is 0, then the third element in the leaf node array is 1, and the remaining elements are 0.
[0117] It should be noted that the leaf node array here is in ciphertext form, which can prevent one party from obtaining the specific prediction path, thereby preventing one party from inferring the splitting threshold of the tree model based on the prediction results of multiple target objects, thereby better protecting data security.
[0118] Next, a possible implementation of step 203 will be described. Step 203 may include: for each non-leaf node, the product of the path value calculated jointly with the second party and the secret comparison result is obtained to obtain the slice of the path value of the child node of the non-leaf node; when the path value is a first value, it indicates that the target object falls into the node corresponding to the path value, and when it is a second value, it indicates that the target object does not fall into the node corresponding to the path value; the slices of the path value of each leaf node are organized into an array according to the order of the leaf nodes to obtain the slices of the leaf node array; the leaf node array is a binary array, and in the leaf node array, the value of the element corresponding to the leaf node where the target object falls is the first value, and the value of the element corresponding to other leaf nodes is the second value. In an optional implementation, the first value is 1 and the second value is 0.
[0119] The plain text of the path value is used to indicate whether the target object falls into the current node. For each non-leaf node, the path value of the child node of the non-leaf node can be determined based on the secret sharing multiplication between the shard of the path value and the shard of the secret comparison result. In this way, the shard of the path value of the leaf node can be obtained through iterative calculation. Then, the shard of the path value of each leaf node can be combined into the shard of the leaf node array.
[0120] Next, we will combine Figure 3 The steps shown and Figure 4 Step 203 will be further described.
[0121] Specifically, both parties need to protect the ciphertext array secret bit The prediction results of each leaf node are obtained without being leaked. bit Privacy computing.
[0122] Specifically, both parties are initialized to size 2 d -1 secret path array <c> bit , and follow Figure 4 In the structure of the tree, c[i] indicates whether the i-th node is on the real predicted path (that is, the path value corresponding to the i-th node mentioned above). Since the root node is a must-go point, it satisfies<c[1]> bit = <1> , that is, holding<c[1]> bit The two parties traverse each non-leaf node in the order of the tree model. For the i-th node, the two parties use bit-based secret multiplication to calculate the secret path of the left and right leaf nodes. Since the secret comparison result is represented by 0 and 1 respectively to fall into the left child node and the right child node,
[0123] It should be noted that although the above formula describes the calculation process in the form of slices, it can be understood that the above process is a joint calculation of two parties. The plaintext of c[i] and b[i] is not disclosed during the calculation process, and the result obtained is a slice of the corresponding plaintext calculation result.
[0124] in, Represents AND calculation. AND calculation can make the result 1 when the two binary bits involved in the calculation are 1, and the result is 0 in other cases. For the calculation formula of the left child node,<c[i]> bit The shard representing the path value of the non-leaf node,<b[i]> bit The shard representing the secret comparison result.
[0125] Since b[i] is 0, it means it falls into the left child node, so we need to calculate the XOR result of b[i] and 1 to make the calculation formula consistent with<c[i]> bit The part that performs the AND operation When it falls into the left leaf node, the value of the shard is 1, and when it falls into the right leaf node, the value of the shard is 0. This is the only way to ensure the accuracy of the final operation result. For example, when b[i] is 0, the secret comparison result is that it falls into the left child node, and the XOR operation result of 0 and 1 is 1, which makes The bit-type secret sharing fragment is 1. Finally, this value is combined with<c[i]> bit Perform the AND operation of the joint calculation of the two parties to ensure that when the secret comparison result falls into the left child node and the path value c[i] is 1 (indicating that the target object falls into the current non-leaf node), the path value of the left child node is sharded<c[2i]> bit The bit-type secret sharing shard with a value of 1 ensures the accuracy of the result. The implementation method when b[i] and c[i] take other values is similar and will not be repeated here.
[0126] Correspondingly, since b[i] takes a value of 1 to indicate that the target object falls into the right child node, and takes a value of 0 to indicate that the target object falls into the left child node, it is possible to determine whether the target object falls into the right leaf node directly based on the AND operation result of the path value c[i] and b[i] indicating whether the target object falls into the current non-leaf node. For example, when both c[i] and b[i] take a value of 1, it indicates that the target object falls into the current non-leaf node and the right child node of the non-leaf node. Then the AND operation result of the two is 1, indicating that the target object falls into the right leaf node. If either c[i] or b[i] takes a value of 0, it can be determined that the target object will not fall into the right child node of the non-leaf node. Then the plain text of the AND operation result of c[i] and b[i] is 0, which also indicates that the target object will not fall into the right child node of the non-leaf node.
[0127] Finally, take the array <c> bit No.2 d-1 to 2 d -1 point ciphertext value is directly used as the leaf node array mentioned above, recorded as<e[j]> bit = <c[2 d-1 +j-1]> bit ,j∈[l].
[0128] Step 205, the second party jointly performs a secret sharing-based operation based on the slices of the leaf node array and the slices of the preset prediction values of each leaf node, to obtain the slices of the prediction results corresponding to the target object.
[0129] Specifically, the slices of the preset prediction values of each leaf node can be first formed into slices of the prediction value array according to the order of the leaf nodes. The two parties jointly perform the secret sharing multiplication operation, calculate the result of the bitwise multiplication of the leaf node array and the prediction value array, and obtain the slices of the first array, which represents the result of the bitwise multiplication of the leaf node array and the prediction value array. The two parties then jointly calculate the secret sharing addition operation between the bits of the first array to obtain the slices of the prediction results of the current tree model for the target object.
[0130] Still used here Figure 3 For the kth decision tree, both parties use secret multiplication and addition to calculate the prediction result of this tree. <e[j]> bit is a slice of the leaf node array,<w[j]> A slice of the prediction array.
[0131] In addition, when there are multiple tree models, the prediction results of each tree model can be added together to obtain the final prediction result for the target object. In other words, after the above steps, both parties obtain the ciphertext array consisting of the prediction results of each tree. <res>Then both parties use secret addition to accumulate all prediction results to obtain the final ciphertext prediction result of the idx-th test data.
[0132] Finally, the prediction results can be adjusted according to business needs. <pred>Decryption is performed so that the first party or the second party obtains the final prediction result pred.
[0133] In other words, the prediction result of the target object can be obtained by performing secret sharing decryption in conjunction with the second party according to the fragmentation of the prediction result. Here, the prediction result of the target object can be finally obtained by at least one of the first party and the second party.
[0134] The above embodiment introduces the reasoning method in the scenario where the intersection primary key of the predicted data set is public, that is, the primary key of the intersection of the objects owned by both parties (that is, the target object) is in plain text. The above embodiment requires the ability to protect the model information and training data set from being leaked while realizing the reasoning of the specified target object and publicizing the prediction results, that is, the participants cannot infer the model information and training data set information through the additional information obtained in the reasoning stage.
[0135] In the related technology, there is also a demand for anonymous reasoning of tree models with hidden primary keys that cannot associate the prediction results with the primary key information of a single object, that is, the prediction results can only reflect the overall trend, or can only obtain compliance scenario requirements for statistical information.
[0136] Based on this, this specification also provides a two-party joint reasoning method for a tree model with an anonymous primary key. The method is described by taking the execution steps of the first party as an example. It can be understood that the steps executed by the second party are the same as those of the first party and will not be repeated. The method includes:
[0137] Step 501, for multiple objects, jointly execute the privacy intersection protocol of the anonymized primary key with the second party to obtain the first slice of the feature values of all features of the target object; the second party obtains the second slice of the feature values of all features of the target object, and the target object is an object jointly held by the first party and the second party.
[0138] Specifically, both parties need to first execute the privacy intersection protocol of the anonymized primary key for the objects they each hold, so that both parties can obtain the shards of the feature values of the target object whose primary key identifiers are anonymized and whose feature values are all in ciphertext form. In the case of multiple target objects (the two parties do not know which specific target objects they are), the privacy intersection protocol of the anonymized primary key can be used to obtain the full ciphertext matrix <D A > and <D B >. Correspondingly, the data of the two input models are: P A hold <D A > A and <D B > A , P B hold <D A > B and <D B > B That is, both parties hold: A and the second party P B The slicing of the eigenvalues of each feature.
[0139] Step 502, for any non-leaf node, obtain the homomorphic ciphertext of the feature number vector of the non-leaf node, and jointly with the second party obtain the feature value fragmentation of the split feature of the target object based on the homomorphic ciphertext of the feature number vector and the feature values of each feature of the target object.
[0140] Different from step 201, after obtaining the homomorphic ciphertext of the shards of the eigenvalues of the obfuscated split feature sent by the feature party, it is necessary to sum the ciphertext with the eigenvalue shards of the obfuscated split feature held locally to obtain the final shards of the eigenvalues of the split feature.
[0141] Here, a specific example in which the first party is a feature party and the second party is a feature party is used to illustrate the implementation of step 502.
[0142] First, for any non-leaf node, when the first party holds the feature number of the non-leaf node, construct a feature number vector according to the feature number, and homomorphically encrypt the feature number vector;
[0143] The second party obtains the homomorphic ciphertext of the feature number vector to the second party; the second party obtains the homomorphic ciphertext of the feature value second shard of the target object split feature according to the homomorphic operation between the homomorphic ciphertext of the feature number vector and the second shard of the feature value of each feature of the target object held by the second party;
[0144] receiving a first value sent by a second party; the first value is calculated based on a sum of a homomorphic operation result between a homomorphic ciphertext of the feature number vector and a second shard of each feature value held by the second party and a preset random number;
[0145] The sum of the homomorphic decryption result of the first value and the first shard of the characteristic value of the target object's split characteristic is used as the characteristic value shard of the processed target object's split characteristic; the second party uses the opposite of the random number as the characteristic value shard of the processed target object's split characteristic.
[0146] The first and second in the first shard and the second shard here have similar meanings to the above, and are used to refer to the party to which the shard belongs. The first shard represents that the shard belongs to the first party, and the second shard represents that the shard belongs to the second party.
[0147] Specifically, the calculation process of the first value is to perform homomorphic multiplication operations on the homomorphic ciphertext of the feature number vector and the second slice of each feature value, and then perform homomorphic addition operations on the results of each homomorphic multiplication operation to obtain the first value. After the first party obtains the first value, it decrypts the first value and sums the first value with the first slice of the feature value of the split feature held locally to obtain the feature value slice of the split feature of the target object held by the first party.
[0148] Second, for any non-leaf node, if the second party holds the feature number of the non-leaf node, receiving the homomorphic ciphertext of the feature number vector sent by the second party;
[0149] Obtaining a homomorphic ciphertext of the first shard of the feature value of the target object split feature according to a homomorphic operation between the homomorphic ciphertext of the feature number vector and the first shard of the feature value of each feature of the target object held by the party;
[0150] Calculate the sum of the homomorphic ciphertext of the first shard of the characteristic value of the split characteristic of the target object and a preset random number to obtain a second value, and send the second value to the second party;
[0151] The opposite of the random number is used as a slice of the feature value of the target object split feature after processing; the second party uses the sum of the second value and the second slice of the feature value of the target object split feature held by the second party as the slice of the feature value of the target object split feature after processing.
[0152] Most of the implementation methods in the above steps can be found in the description of step 201 and will not be repeated here.
[0153] Still used here Figure 3 The symbols in the examples are explained in detail.
[0154] First, execute step 201 Figure 3 Follow steps 1 and 2 in the instructions. Follow step 3' below.
[0155] 3', for the i-th non-leaf node (i∈[p]), the feature square uses the ciphertext and the characteristics of our test samples <D A(B) [idx]> A(B) Calculate the ciphertext directly locally Then, the homomorphic ciphertext-to-secret sharing technique is used to obtain <D A(B) [idx][f[i]]>. Specifically, assume that the feature square of the current i-th non-leaf node is P A , P A hold and <D A [idx]> A First, P A Using homomorphic multiplication and homomorphic addition of plaintext and ciphertext, we can calculate the homomorphic ciphertext Then, P A Generate a random number r locally and use the public key pk B Encrypted as Then, perform homomorphic addition to obtain Then, P A Will Send to P B , P B Using the private key sk B Decrypt the pseudo-random number <D A [idx][f B [i]]> A +r; Finally, P A Directly use the random number -r as the secret sharing ciphertext of the eigenvalue corresponding to the i-th non-leaf node, denoted as<v[i]> A = -r, and P B Directly use f locally B [i]In <D A [idx]> B Select <D A [idx][f B [i]]> B , and ( <D A [idx][f B [i]]> A +r)+ <D A [idx][f B [i]]> B As the secret shared ciphertext of the eigenvalue of the i-th non-leaf node, denoted as<v[i]> B =D A [idx][f B [i]]+r.
[0156] The subsequent steps are the same as the secret comparison mentioned in step 201 and steps 203 to 205, and will not be repeated here.
[0157] like Figure 5 As shown, the specification also provides a two-party joint reasoning device for a tree model, wherein a first party and a second party respectively hold feature values of different features of a target object; the first party and the second party jointly hold a tree model, the feature number of the split feature of a non-leaf node of the tree model is held by the party that does not hold the feature value of the split feature, and the feature threshold of the split feature is held by the first party and the second party in the form of slices; the device includes:
[0158] The secret comparison module 510 is used to obtain the homomorphic ciphertext of the feature number vector of any non-leaf node, and obtain the feature value slices of the split feature of the target object in conjunction with the second party based on the homomorphic ciphertext of the feature number vector and the feature values of each feature of the target object; and perform secret comparison with the second party based on the slices of the feature value of the split feature and the slices of the feature threshold to obtain the slices of the secret comparison result; the secret comparison result is used to indicate the node into which the target object falls after the non-leaf node; the feature number vector is a one-hot vector with a value of 1 at the corresponding position of the feature number of the non-leaf node;
[0159] The leaf node array determination module 520 is used to determine the slice of the leaf node array corresponding to the target object in conjunction with the second party based on the slice of the secret comparison result of each non-leaf node; the leaf node array is used to indicate the leaf node into which the target object falls;
[0160] The prediction result acquisition module 530 is used to jointly with the second party perform a secret sharing-based operation according to the slices of the leaf node array and the slices of the preset prediction values of each leaf node to obtain the slices of the prediction results of the target object.
[0161] In an optional embodiment, the secret comparison module 510 is specifically used to: for any non-leaf node, when the first party holds the feature number of the non-leaf node, construct a feature number vector according to the feature number, and homomorphically encrypt the feature number vector; send the homomorphic ciphertext of the feature number vector to the second party; the second party obtains the homomorphic ciphertext of the feature value of the split feature of the target object based on the homomorphic operation between the homomorphic ciphertext of the feature number vector and the feature values of each feature of the target object held by the second party; and jointly execute the homomorphic encryption to secret sharing protocol with the second party based on the homomorphic ciphertext of the feature value of the split feature of the target object to obtain the sharding of the feature value of the split feature of the target object.
[0162] In an optional embodiment, the secret comparison module 510 is specifically used to: for any non-leaf node, when the second party holds the feature number of the non-leaf node, receive the homomorphic ciphertext of the feature number vector sent by the second party; obtain the homomorphic ciphertext of the feature value of the target object split feature based on the homomorphic operation between the homomorphic ciphertext of the feature number vector and the feature value of each feature of the target object held by this party; send the homomorphic ciphertext of the feature value of the target object split feature to the second party, and jointly execute the homomorphic encryption to secret sharing protocol with the second party to obtain the sharding of the feature value of the target object split feature.
[0163] In an optional embodiment, the leaf node array determination module 520 is specifically used to obtain, for each non-leaf node, a slice of the path value of the left child node of the non-leaf node by multiplying the path value calculated jointly with the second party and the secret comparison result; when the path value is a first value, it indicates that the target object falls into the node corresponding to the path value, and when it is a second value, it indicates that the target object does not fall into the node corresponding to the path value; the slices of the path value of each leaf node are organized into an array according to the order of the leaf nodes to obtain a slice of the leaf node array; the leaf node array is a binary array, and in the leaf node array, the value of the element corresponding to the leaf node where the target object falls is the first value, and the value of the element corresponding to other leaf nodes is the second value.
[0164] In an optional embodiment, it also includes a privacy intersection module 500 (not shown in the figure), which is used to jointly execute the privacy intersection protocol of the anonymous primary key with the second party for multiple objects to obtain the first slice of the feature values of all features of the target object; the second party obtains the second slice of the feature values of all features of the target object, and the target object is an object jointly held by the first party and the second party.
[0165] In an optional embodiment, the secret comparison module 510 is specifically used to: for any non-leaf node, when the first party holds the feature number of the non-leaf node, construct a feature number vector according to the feature number, and homomorphically encrypt the feature number vector; send the homomorphic ciphertext of the feature number vector to the second party; the second party obtains the homomorphic ciphertext of the second shard of the feature value of the target object split feature according to the homomorphic operation between the homomorphic ciphertext of the feature number vector and the second shard of the feature value of each feature of the target object held by the second party; receive the first value sent by the second party; the first value is calculated based on the sum of the homomorphic operation result between the homomorphic ciphertext of the feature number vector and the second shard of each feature value held by the second party and a preset random number; use the sum of the homomorphic decryption result of the first value and the first shard of the feature value of the target object split feature as the feature value shard of the split feature of the processed target object; the second party uses the opposite of the random number as the feature value shard of the split feature of the processed target object.
[0166] In an optional embodiment, the secret comparison module 510 is specifically used to: for any non-leaf node, when the second party holds the feature number of the non-leaf node, receive the homomorphic ciphertext of the feature number vector sent by the second party; obtain the homomorphic ciphertext of the first shard of the feature value of the target object split feature based on the homomorphic operation between the homomorphic ciphertext of the feature number vector and the first shard of the feature value of each feature of the target object held by this party; calculate the sum of the homomorphic ciphertext of the first shard of the feature value of the target object split feature and a preset random number to obtain a second value, and send the second value to the second party; use the opposite of the random number as the shard of the feature value of the target object split feature after processing; the second party uses the sum of the second value and the second shard of the feature value of the target object split feature held by the second party as the shard of the feature value of the target object split feature after processing.
[0167] In an optional implementation, a decryption module 540 (not shown) is further included, which is used to perform secret sharing decryption in conjunction with the second party according to the slices of the prediction results to obtain the prediction results of the target object.
[0168] In an optional implementation, the tree model is trained in conjunction with the second party based on a target training sample set, and the target training sample set is a training sample set obtained by executing a privacy intersection protocol of anonymized primary keys in conjunction with the second party on data of the held object.
[0169] In the 1990s, it was very clear whether the improvement of a technology was a hardware improvement (for example, improvements to the circuit structure of diodes, transistors, switches, etc.) or a software improvement (improvement of the method flow). However, with the development of technology, many improvements in the method flow today can be regarded as direct improvements in the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to ask a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs. The original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.
[0170] The controller can be implemented in any appropriate manner, for example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (such as software or firmware) that can be executed by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in a purely computer-readable program code manner, the controller can be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, this controller can be considered as a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as both software modules for implementing the method and structures within the hardware component.
[0171] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, this specification does not exclude that with the development of computer technology in the future, the computer that implements the functions of the above embodiments may be, for example, a personal computer, a laptop computer, a vehicle-mounted human-computer interaction device, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0172] Although one or more embodiments of the present specification provide method operation steps as described in the embodiments or flow charts, more or less operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way of executing the order of many steps, and does not represent the only execution order. When the device or terminal product in practice is executed, it can be executed in sequence or in parallel according to the method shown in the embodiments or the drawings (for example, a parallel processor or a multi-threaded processing environment, or even a distributed data processing environment). The term "include", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, product or equipment including a series of elements includes not only those elements, but also includes other elements that are not explicitly listed, or also includes elements inherent to such process, method, product or equipment. In the absence of more restrictions, it is not excluded that there are other identical or equivalent elements in the process, method, product or equipment including the elements. For example, if the words first, second, etc. are used to represent the name, they do not represent any specific order.
[0173] For the convenience of description, the above devices are described in various modules according to their functions. Of course, when implementing one or more of the present specification, the functions of each module can be implemented in the same or more software and / or hardware, or the module implementing the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0174] This specification is described with reference to the flowcharts and / or block diagrams of the methods, apparatus (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0175] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0176] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0177] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0178] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0179] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage, graphene storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0180] It should be understood by those skilled in the art that one or more embodiments of the present specification may be provided as a method, system or computer program product. Therefore, one or more embodiments of the present specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, one or more embodiments of the present specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0181] One or more embodiments of the present specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0182] Each embodiment in this specification is described in a progressive manner, and the same and similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. In the description of this specification, the description of the reference term "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representation of the above terms does not necessarily target the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, in the absence of mutual contradiction, a person skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0183] The above description is only an example of one or more embodiments of the present specification and is not intended to limit one or more embodiments of the present specification. For those skilled in the art, one or more embodiments of the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims.< / pred> < / res> < / c> < / c> < / x> < / x> < / x> < / x> < / x> < / x> < / x> < / x> < / x> < / w> < / h> < / g> < / h> < / g> < / h> < / g> < / h> < / g> < / d> < / d> < / x> < / x> < / x> < / y> < / d>
Claims
1. A two-party joint reasoning method for a tree model, applied to a first party, the method also involves a second party, the first party and the second party respectively hold feature values of different features of a target object; the first party and the second party jointly hold a tree model, the feature number of the split feature of a non-leaf node of the tree model is held by the party that does not hold the feature value of the split feature, and the feature threshold of the split feature is held by the first party and the second party in the form of fragments; the method comprises: For any non-leaf node, obtain the homomorphic ciphertext of the feature number vector of the non-leaf node, and jointly with the second party obtain the feature value shards of the split feature of the target object based on the homomorphic ciphertext of the feature number vector and the feature values of each feature of the target object; and jointly with the second party, perform secret comparison based on the shards of the feature value of the split feature and the shards of the feature threshold to obtain the shards of the secret comparison result; the secret comparison result is used to indicate the node into which the target object falls after the non-leaf node; the feature number vector is a one-hot vector with a value of 1 at the corresponding position of the feature number of the non-leaf node; Determine, in conjunction with the second party, a shard of the leaf node array corresponding to the target object based on the shard of the secret comparison result of each non-leaf node; The leaf node array is used to indicate the leaf node into which the target object falls; The second party jointly performs a secret sharing-based operation based on the slices of the leaf node array and the slices of the preset prediction values of each leaf node to obtain the slices of the prediction results of the target object.
2. According to the method of claim 1, for any non-leaf node, obtaining the homomorphic ciphertext of the feature number vector of the non-leaf node, and jointly obtaining the feature value fragmentation of the split feature of the target object based on the homomorphic ciphertext of the feature number vector and the feature values of each feature of the target object with the second party, including: For any non-leaf node, if the first party holds the feature number of the non-leaf node, construct a feature number vector according to the feature number, and homomorphically encrypt the feature number vector; The second party obtains the homomorphic ciphertext of the feature number vector to the second party; the second party obtains the homomorphic ciphertext of the feature value of the split feature of the target object according to the homomorphic operation between the homomorphic ciphertext of the feature number vector and the feature value of each feature of the target object held by the second party; Cooperate with the second party to execute a homomorphic encryption-to-secret sharing protocol based on the homomorphic ciphertext of the characteristic value of the split characteristic of the target object to obtain fragments of the characteristic value of the split characteristic of the target object.
3. According to the method of claim 1, for any non-leaf node, obtaining the homomorphic ciphertext of the feature number vector of the non-leaf node, and jointly obtaining the feature value fragmentation of the split feature of the target object based on the homomorphic ciphertext of the feature number vector and the feature values of each feature of the target object with the second party, comprising: For any non-leaf node, if the second party holds the feature number of the non-leaf node, receiving the homomorphic ciphertext of the feature number vector sent by the second party; Obtaining the homomorphic ciphertext of the eigenvalue of the split feature of the target object according to the homomorphic operation between the homomorphic ciphertext of the feature number vector and the eigenvalue of each feature of the target object held by the party; A homomorphic ciphertext of the characteristic value of the split characteristic of the target object is sent to the second party, and a homomorphic encryption-to-secret sharing protocol is jointly executed with the second party to obtain a shard of the characteristic value of the split characteristic of the target object.
4. According to the method of claim 1, the sharding of the secret comparison results of each non-leaf node and the sharding of the leaf node array corresponding to the target object jointly determined by the second party include: For each non-leaf node, obtain the shards of the path values of the left node and the right node of the non-leaf node according to the product of the path value jointly calculated by the two parties and the secret comparison result; When the path value is the first value, it indicates that the target object falls into the node corresponding to the path value; When the path value is the second value, it indicates that the target object does not fall into the node corresponding to the path value; The slices of the path value of each leaf node are arrayed according to the order of the leaf nodes to obtain the slices of the leaf node array; The leaf node array is a binary array. In the leaf node array, the value of the element corresponding to the leaf node where the target object falls is a first value, and the value of the element corresponding to other leaf nodes is a second value.
5. The method according to claim 1, further comprising: For multiple objects, jointly execute the privacy intersection protocol of the anonymized primary key with the second party to obtain the first slice of the feature values of all features of the target object; The second party obtains a second slice of feature values of all features of a target object, where the target object is an object held jointly by the first party and the second party.
6. The method according to claim 5, wherein for any non-leaf node, obtaining the homomorphic ciphertext of the feature number vector of the non-leaf node, and jointly obtaining the feature value fragmentation of the split feature of the target object based on the homomorphic ciphertext of the feature number vector and the feature values of each feature of the target object with the second party, comprises: For any non-leaf node, if the first party holds the feature number of the non-leaf node, construct a feature number vector according to the feature number, and homomorphically encrypt the feature number vector; The second party obtains the homomorphic ciphertext of the feature number vector to the second party; the second party obtains the homomorphic ciphertext of the feature value second shard of the target object split feature according to the homomorphic operation between the homomorphic ciphertext of the feature number vector and the second shard of the feature value of each feature of the target object held by the second party; receiving a first value sent by a second party; the first value is calculated based on a sum of a homomorphic operation result between a homomorphic ciphertext of the feature number vector and a second shard of each feature value held by the second party and a preset random number; The sum of the homomorphic decryption result of the first value and the first shard of the characteristic value of the target object's split characteristic is used as the characteristic value shard of the processed target object's split characteristic; the second party uses the opposite of the random number as the characteristic value shard of the processed target object's split characteristic.
7. According to the method of claim 5, for any non-leaf node, the homomorphic ciphertext of the feature number vector of the non-leaf node is obtained, and the feature value fragmentation of the split feature of the target object is obtained in conjunction with the second party based on the homomorphic ciphertext of the feature number vector and the feature values of each feature of the target object. For any non-leaf node, if the second party holds the feature number of the non-leaf node, receiving the homomorphic ciphertext of the feature number vector sent by the second party; Obtaining a homomorphic ciphertext of the first shard of the feature value of the target object split feature according to a homomorphic operation between the homomorphic ciphertext of the feature number vector and the first shard of the feature value of each feature of the target object held by the party; Calculate the sum of the homomorphic ciphertext of the first shard of the characteristic value of the split characteristic of the target object and a preset random number to obtain a second value, and send the second value to the second party; The opposite of the random number is used as a slice of the feature value of the target object split feature after processing; the second party uses the sum of the second value and the second slice of the feature value of the target object split feature held by the second party as the slice of the feature value of the target object split feature after processing.
8. The method according to claim 1, further comprising: According to the sharding of the prediction results, secret sharing decryption is performed jointly with the second party to obtain the prediction results of the target object.
9. According to the method of claim 1, the tree model is obtained by training in conjunction with the second party based on a target training sample set, and the target training sample set is a training sample set obtained by executing a privacy intersection protocol of anonymized primary keys on data of objects held in conjunction with the second party.
10. A computing device comprising a memory and a processor, wherein the memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Tree model-based two-party joint classification method, device, equipment and medium
CN111639367A
Federal learning method and system for longitudinal xgboost decision tree
CN114169537A
Gradient boosted tree model construction method and device, equipment and medium
CN114925853A
Multi-party joint cause fruit tree model construction system and method based on federal learning
CN116757286A
Method and device for training tree model
US20220036250A1