Tree model two-party joint reasoning method and computer equipment

By using homomorphic encryption and secret sharing protocols, the problem of lacking two-party joint reasoning in anonymous primary key modeling is solved, enabling secure tree model reasoning under anonymous data conditions, protecting primary key information from being leaked, and improving the security and privacy protection of data utilization.

CN119990328BActive Publication Date: 2025-12-26ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510201328.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-12-26
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

Existing technologies lack tree model reasoning methods for two-party joint covert primary key modeling, making it impossible to safely perform tree modeling and reasoning under covert data conditions.

Method used

Using homomorphic encryption and secret sharing protocol, the tree model is jointly held by the first and second parties. The split feature number of non-leaf nodes is held by the non-feature party, and the split feature threshold is held by both parties in the form of secret sharing fragments. Homomorphic operations and secret comparisons of feature values ​​are performed to determine the node of the target object in the tree model and the prediction result.

Benefits of technology

It enables secure and efficient tree model reasoning under anonymous data conditions, protects primary key information from being leaked, and improves the security and privacy protection of data utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990328B_ABST
    Figure CN119990328B_ABST
Patent Text Reader

Abstract

The specification provides a tree model two-party joint reasoning method. A first party and a second party respectively hold different types of characteristics of a target object. The first party and the second party also jointly hold a tree model, in which the feature numbers of the split features of the non-leaf nodes are held by the non-feature party, and the split thresholds are held by the two parties in the form of secret sharing fragments. For each non-leaf node, according to the homomorphic operation between the homomorphic ciphertext of the non-leaf node feature number vector and the characteristic value of the target object, the fragment of the characteristic value of the split feature of the target object is obtained. And through the secret comparison jointly executed by the two parties, according to the split threshold fragment of the non-leaf node, the fragment of the secret comparison result is determined. The two parties jointly determine the fragment of the leaf node array according to the fragment of the secret comparison result of each non-leaf node. Finally, the two parties jointly determine the prediction result of the target object according to the leaf node array and the fragment of the prediction value of each leaf node.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present specification belong to the technical field of privacy computing, and particularly relate to a tree model two-party joint inference method and a computer device. BACKGROUND

[0002] A decision tree model is a tree structure, and through the decision tree model, the attributes of a target object can be scored or classified according to the characteristic values of the characteristics of the target object, to obtain a final inference result. With the requirement of data security, privacy computing methods that take into account privacy protection requirements and data utilization are widely used. In related technologies, multiple participants use the data they hold to perform joint privacy computing, such as jointly modeling a tree model and jointly inferring a tree model.

[0003] With the continuous improvement of data security requirements in various industries, further requirements for full-trace privacy protection machine learning are required, which enables more secure tree modeling under the conditions of traceable data primary key information and data full-process encryption. The primary key refers to identification information that can distinguish the data. However, the tree model for modeling the primary key in related technologies lacks a two-party joint inference method. SUMMARY

[0004] The purpose of the present specification is to provide a tree model two-party joint inference method and a computer device.

[0005] The first aspect of the present specification provides a tree model two-party joint inference method applied to a first party, the method also involving a second party, the first party and the second party respectively holding characteristic values of different characteristics of a target object; the first party and the second party jointly hold a tree model, the characteristic number of the split characteristic of the non-leaf node of the tree model being held by the party that does not hold the characteristic value of the split characteristic, and the characteristic threshold of the split characteristic being held in the form of a slice by the first party and the second party; the method comprises:

[0006] For any non-leaf node, obtain the homomorphic ciphertext of the characteristic number vector of the non-leaf node, and jointly with the second party, according to the homomorphic ciphertext of the characteristic number vector and the characteristic values of the characteristics of the target object, obtain the characteristic value slice of the split characteristic of the target object; and jointly with the second party, according to the slice of the characteristic value of the split characteristic and the slice of the characteristic threshold, perform a secret comparison to obtain a slice of a secret comparison result; the secret comparison result is used to indicate the node into which the target object falls after the non-leaf node; the characteristic number vector is a one-hot vector whose value at the position corresponding to the characteristic number of the non-leaf node is 1;

[0007] determine, jointly with the second party, a leaf node array corresponding to the target object according to the secret comparison result of each non-leaf node; the leaf node array is used to indicate a leaf node in which the target object falls;

[0008] The second party jointly performs a secret sharing-based operation according to the leaf node array and a preset leaf node prediction value, to obtain a prediction result of the target object.

[0009] The second aspect of the specification provides a tree model two-party joint reasoning device, a first party and a second party respectively holding feature values of different features of a target object; the first party and the second party jointly hold a tree model, and a feature number of a split feature of a non-leaf node of the tree model is held by a party that does not hold a feature value of the split feature; and a feature threshold of the split feature is held in the form of a fragment by the first party and the second party; the device comprises:

[0010] A secret comparison module is configured to, for any non-leaf node, obtain homomorphic ciphertext of a feature number vector of the non-leaf node, and jointly with the second party, obtain a feature value fragment of a split feature of the target object according to the homomorphic ciphertext of the feature number vector and the feature values of each feature of the target object; and jointly with the second party, perform secret comparison according to the feature value fragment of the split feature and a feature threshold fragment, to obtain a secret comparison result fragment; the secret comparison result is used to indicate a node in which the target object falls after the non-leaf node; the feature number vector is a one-hot vector in which a value corresponding to a feature number of the non-leaf node is 1;

[0011] A leaf node array determination module is configured to, jointly with the second party, determine a leaf node array fragment corresponding to the target object according to the secret comparison result fragment of each non-leaf node; the leaf node array is used to indicate a leaf node in which the target object falls;

[0012] A prediction result acquisition module is configured to, jointly with the second party, perform a secret sharing-based operation according to the leaf node array and a preset leaf node prediction value, to obtain a prediction result of the target object.

[0013] The third aspect of the specification provides a computer-readable storage medium having a computer program stored thereon, when the computer program is executed in a computer, the computer program causes the computer to execute the tree model two-party joint reasoning method.

[0014] The fourth aspect of the specification provides a computing device comprising a memory and a processor, the memory having executable code stored therein, and the processor implementing the tree model two-party joint reasoning method when executing the executable code.

[0015] The fifth aspect of the specification provides a computer program product comprising computer programs / instructions which, when executed by a processor, implement the steps of the tree model two-party joint inference method described above.

[0016] Based on this, the specification provides a tree model two-party joint inference method. This method involves a first party and a second party, and the first party and the second party respectively hold different types of features of a target object. The first party and the second party also jointly hold a tree model, in which the feature numbers of the split features of the non-leaf nodes are held by the non-feature party, and the split thresholds are held by the two parties in the form of secret sharing fragments. In the inference process, for each non-leaf node, the homomorphic operation between the homomorphic ciphertext of the non-leaf node feature number vector and the feature value of the target object is performed, and through the homomorphic encryption to secret sharing protocol, the fragment of the feature value of the split feature of the target object is obtained. And through the secret comparison jointly executed by the two parties, according to the fragment of the split threshold of the non-leaf node, the fragment of the secret comparison result indicating which child node the target object falls into is determined. The two parties jointly determine the fragment of the leaf node array indicating the leaf node where the target object falls into according to the fragments of the secret comparison results of each non-leaf node. Finally, the two parties jointly determine the prediction result of the target object according to the leaf node array and the fragments of the prediction values of each leaf node.

[0017] The above method is for the scenario where the feature numbers of the split features of the non-leaf nodes are held by the non-feature party and the split thresholds are held by the two parties. Through homomorphic encryption of the feature number vector of the non-leaf node of the tree model and homomorphic operation between the feature value of the target object, the ciphertext of the feature value of the split feature is obtained. Thus, the inference of the non-leaf node is completed according to the ciphertext of the feature value and the ciphertext of the split threshold, and the prediction value for the target object is obtained based on this. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the specification, the drawings needed in the embodiment description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments described in the specification, and other drawings can also be obtained by those skilled in the art without creative labor.

[0019] Figure 1 FIG. 1 is a schematic diagram of a tree model in an embodiment;

[0020] Figure 2 FIG. 2 is a flowchart of a tree model two-party joint inference method in an embodiment;

[0021] Figure 3 FIG. 3 is a flowchart of a tree model two-party joint inference method in another embodiment;

[0022] Figure 4 This is a schematic diagram of a tree model predicting a path in one embodiment;

[0023] Figure 5 This is a block diagram of a tree model two-party joint reasoning device in one embodiment. Detailed Implementation

[0024] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0025] Decision tree models typically exist in the form of binary trees, including leaf nodes and non-leaf nodes. During inference, the splitting conditions of non-leaf nodes determine which leaf node a sample falls into, and the score or classification of the sample is determined based on the weight value of the leaf node it falls into. The splitting conditions include splitting features and the corresponding splitting thresholds. For example... Figure 1 As shown, Figure 1 It's a simple decision tree. The content of non-leaf nodes represents the splitting condition for that node, where A is the splitting feature and the number is the corresponding splitting threshold. The content of leaf nodes represents the weight value of that leaf node. If a sample's weight is less than the corresponding splitting threshold, it will fall into the left child node of that node (hereinafter referred to as the left child node); otherwise, it will fall into the right child node (hereinafter referred to as the right child node), until the sample falls into a leaf node. For example, if the feature value of feature A for a sample is 1, then based on the splitting conditions of each non-leaf node, it can be determined that the sample falls into the leftmost leaf node, and the score value of the sample is 1. It should be noted that... Figure 1 The example shown is a simple decision tree model. In practice, decision tree models may not generate judgment conditions for non-leaf nodes based on only one feature, and may also use multiple trees to classify or score samples.

[0026] There are various training methods for tree models, generally including Gradient Boosting Decision Tree (GBDT), Extreme Gradient Boosting (XGBoost), etc. Here, GBDT will be briefly introduced. GBDT is an iterative decision tree algorithm. By iteratively training multiple decision trees, each decision tree is trained on the residual (difference between predicted value and actual value) of the previous tree. During training, the structure and parameters of each decision tree are continuously optimized through gradient boosting to minimize the loss function of the model. In the prediction phase, the prediction results of multiple decision trees are summed or weighted averaged to obtain the final prediction result. Gradient decision trees are commonly used to solve regression and classification problems. For example, a tree model can be trained to classify users' financial credit risk, thereby quickly identifying users with credit risk (such as fraud risk).

[0027] In the scenario of jointly training tree models through privacy protection, the participants (here taking two as an example) generally need to determine common samples, i.e. intersection samples. Specifically, each participant generally holds the primary key and feature value corresponding to each training sample, and different participants hold different feature values. Two participants can determine intersection samples based on the primary key, which have the same primary key held by both participants.

[0028] Among them, the primary key is identification information that can distinguish different training samples, such as in the scenario of classifying users' financial credit risk, the primary key can be the user's identity information identifier, mobile phone number, etc. Different participants hold different types of features, such as in the scenario of classifying users' financial credit risk, one participant can hold the user's transaction data with other users, and another participant can hold the user's friend information, such as which user has a friend relationship, etc.

[0029] For the method of determining intersection samples in the tree model training process, generally, private set intersection (PSI) is used to determine. PSI can ensure that the primary key information and feature value information of non-intersection samples are not leaked, and after PSI is executed, the two participants can determine the primary key information of the intersection samples, and the two participants can further jointly complete the training of the tree model based on the intersection samples.

[0030] With the continuous strengthening of data privacy, data holders have more stringent data demands and hope to protect their intersection primary key information from being leaked during modeling.

[0031] Based on this, the present specification provides a tree model training method for anonymizing primary keys. Next, the implementation process of the method will be described through a specific embodiment.

[0032] The method is suitable for the case where, after the primary key information is completed by the primary key privacy intersection (Circuit-PSI) and the primary key privacy bucketing, both parties hold the secret share ciphertext output of the full-anonymized bucketing protocol: the feature value slice, the feature bucket number slice, and the bucket boundary marker value slice. The method completes the calculation process data encryption and the decision tree model training method that the primary key information cannot be speculated.

[0033] Now assume that two participants P A and P B hold different feature data sets, where P A has m A features, and P B has m B features and label information Y. Both parties want to use their respective data sets to train a decision tree model. Both parties use the existing technology of primary key privacy intersection and primary key privacy bucketing protocol to obtain the data-aligned secret share matrix <d>=< B1 ||U1||X1||...||B m ||U m ||X m ||Y>, the matrix size is n×(3m+1), where n is the number of data intersections, m=m A +m B The first 3m columns represent the total number of features, indicating the bucket number and shards for each feature. i > Bucket boundary marker value fragmentation i (The plaintext of the position corresponding to the minimum value in each bucket is 1, and the plaintext of the remaining values ​​is 0), Feature value partitioning <X i These three columns, the last one is the label value slice. <y>, || denotes the concatenation of the columns of the matrix. The first m A features in the matrix belong to P A , the last m B features belong to P B . Each feature will be divided into nb buckets.

[0034] Note: In this specification, angle brackets <x>denotes that the current variable X is secret share encrypted, and P A and P B respectively hold secret share shards <x> A and <x> B . In this specification, the current variable X is homomorphic encryption, and the A in the upper right corner indicates that the private key of the ciphertext belongs to P A . Only P A can decrypt the ciphertext. A .

[0035] 1. Bucket data extraction

[0036] The purpose of this step is to determine the corresponding bucket threshold under the condition of anonymous primary key. Specifically, A needs to obtain the ciphertext of the bucket threshold of each bucket of B. Correspondingly, B needs to obtain the ciphertext of the bucket threshold of each bucket of A. For specific implementation, see the following.

[0037] (1) P A from <d> A extracting the fragments and sending them to P B . P B decrypts the fragments locally using the received fragments and the corresponding fragments held by P to obtain the first m B intersection data features A bucket numbers and bucket boundary marker values and sends them to P

[0038] (2) P A from <d> A Extracting shards Obtained by homomorphic encryption And send it to P B . P B Locally use the received shard and the corresponding shard held by oneself Addition operation, obtain only homomorphic encryption under ciphertext

[0039] (3) P B According to And The first m A characteristics can be obtained, and then the ciphertext value at the corresponding position in The secret share of each bucket threshold is extracted The size of each array S i is nb.

[0040] (4) Similar to (1-3), P B Send And To P A , P A Obtain the secret share of the bucket threshold

[0041] (5) P B Initialize the prediction value array And encrypt it into a secret shard held by both parties through secret sharing

[0042] 2, Decision tree model training

[0043] (1) For the kth decision tree, both parties perform the following operations:

[0044] (1.1) Both parties use And the gradient derivation formula g = σ(p k )-y, h = (1-σ(p k ))·σ(p k ), in the secret sharing scenario, the derivative value {<g i >, <h i >} i∈[1,n] .

[0045] (1.2) P A Locally initialize the sample distribution of the first node (root node) Is a full 1 vector of size n. Similarly, P B Locally initialize a full 1 vector of size n

[0046] (1.3) For the l-th non-leaf node in the tree model, both sides perform the following operations:

[0047] (1.3.1)P B Locally for { <g i > B , <h i > B } i∈[1,n] Perform homomorphic encryption to obtain {[ <g i > B ] B ,[ <h i > B ] B } i∈[1,n] and send it to P A P A The local { <g i > A , <h i > A } i∈[1,n] and{[ <g i > B ] B ,[ <h i > B ] B } i∈[1,n] Add them together to obtain a homomorphically encrypted ciphertext array {[g i ] B ,[h i ] B } i∈[1,n]

[0048] (1.3.2)P A Initialize all-zero gradient bucketing cumulative matrix Then, P A According to the bucket information and sample distribution information Will satisfy The corresponding [ <g j > B ] B ,[ <h j > B ] B Accumulated to In the middle, B i,j The bucket number is the bucket number corresponding to the i-th feature of the j-th sample.

[0049] (1.3.3)P A Using homomorphic ciphertext to secretly share ciphertext, Transformed into a secret shared shard held by both parties

[0050] (1.3.4) Similar to (1.3.1-1.3.3), P A In the local {<g i > B ,<h i > B} i∈[1,n] Perform homomorphic encryption, P B Through the way of gradient bucket accumulation and homomorphic ciphertext to secret sharing ciphertext, both parties obtain

[0051] (1.3.5) Merge two gradient bucket accumulation matrices to obtain {<G i,q >,<H i,q >} i∈[1,m],q∈[1,nb] .

[0052] (1.3.6) Both parties use decision tree calculation rules and tree model gain value calculation formula in MPC scenario Obtain the {<Gain i,q >} corresponding to each bucket boundary value i∈[1,m],q∈[1,nb] , Where

[0053] (1.3.7) Both parties use the private extreme value solving method to obtain {<Gain i,q >} The maximum value of i∈[1,m],q∈[1,nb] Corresponding to the secret fragment <max_idx> of the bucket number, and through the private comparison, both parties obtain the size relationship of <max_idx> and m A ×nb, if <max_idx> is greater than m A ×nb, it means that the sample distribution information and the bucket threshold corresponding to the feature are managed by P A , Need to be disclosed to P A <max_idx> alone; Conversely, <max_idx> is disclosed to P B .

[0054] (1.3.8) Assuming P B Obtain the value max_idx, then P B Can know that the split boundary of the lth node is the jth bucket of the ith feature through max_idx=i×nb+j. P B With the bucket number j as the boundary, update the distribution of n samples Determine the sample distribution information of the left and right child nodes 2l and 2l+1 of the lth node, that is Then extract [S i,j ] A and convert it to a secret share shard value <T k,l > as the split threshold of the current node.

[0055] (1.4) For the lth leaf node in the tree model, both parties perform the following operations:

[0056] (1.4.1) Suppose the split information of the parent node of the lth leaf node is held by P B , then P B can directly obtain the calculation formula of the leaf node weight through the max_idx of (1.3.8) , where the required <g>and <h>In {<Gl i,q >,<Hl i,q >} i∈[1,m],q∈[1,nb] corresponding position d.

[0057] (1.4.2) However, since P A does not know the position information of the feature information and the maximum gain value corresponding to the parent node, both parties need to obtain <g>and <h>. Specifically, first P A Two random numbers rg and rh are generated and {<Gl i,q -rg> A ,<Hl i,q -rh> A} i∈[1,m],q∈[1,nb] are computed locally; subsequently, the array is input to the OT protocol; at the same time, P B inputs d in the OT protocol and obtains <Gl d -rg> A ,<Hl d -rh> A from the output of the OT protocol. Then, P B locally computes <g> B <Gl d -rg> A +<Gl d B <h> B ←<Hl d -rg> A +<Hl d > B ;P A directly as secret shards <g> A ←rg, <h> A ← rh.

[0058] (1.4.3) Both parties update the prediction result by formula to obtain the weight value of the leaf node of the tree model.

[0059] (1.5) For each sample i, both parties update the prediction result by formula , where f k (i) is the leaf node number corresponding to the i-th sample, which can be obtained by multiplying the bit of each leaf node by and to obtain the one-hot encoding vector of f k (i), and then the vector is multiplied by <w>Dot product obtains as the weight value of the leaf node.

[0060] (2) Finally, both parties output the feature information of the split point of each tree and the split threshold piece as the model output result of this tree model training.

[0061] In the tree model trained by the above training process, the same as the tree model trained by both parties in the non-anonymous primary key scene in the related art is that the prediction value of the leaf node is held by both parties in the form of a piece. Different from the tree model trained by both parties in the related art is that the feature number of the split feature of the non-leaf node is held by the non-feature party, and the split threshold of the split feature is held by the two parties in the form of a secret share.

[0062] Among them, secret sharing (secret sharing) is also called secret division and secret sharing. The basic principle is to split the secret into multiple shares (share), also known as a piece, and give it to different people to keep. Only more than the threshold number of people will combine their shares to recover the secret; from less than the threshold number of shares, no information can be recovered. In privacy computing, the threshold number is usually the same as the number of participants. For two participants P0, P1, arithmetic secret sharing divides the plaintext value x of length l bits into two random number shares <x>0 e 2 l and <x>1 < 2 l are stored by P0, P1, respectively, and satisfy x = ( <x> 0+ <x>1) mod 2 l The multiple data obtained by the secret sharing can be referred to as shards of the secret sharing, which are also simply referred to as shards in this specification. The secret sharing is based on the fact that the four arithmetic operations of addition, subtraction, multiplication and division can also be performed by the two parties jointly according to the shards of the secret sharing.

[0063] The non-feature party is a party that does not hold a feature of the type of the split feature. For example, the split feature is the account registration time of a user, which is held by the first party, and the second party does not hold the account registration time of the user, and the second party is the non-feature party. The feature number of the split feature is the number of the split feature in all features possessed by the first party and the second party.

[0064] In the tree model of the related art, the feature number of the non-leaf node is generally held by the feature party, and the corresponding feature party also holds the split threshold of the non-leaf node. The difference between the tree model trained in this specification and the tree model in the related art makes the tree model inference method in the related art unable to be applied to the scenario targeted by this specification.

[0065] Based on this, this specification provides a tree model two-party joint inference method. The method involves a first party and a second party, and the first party and the second party hold different types of features of a target object. The first party and the second party also jointly hold a tree model, in which the feature number of the split feature of the non-leaf node is held by the non-feature party, and the split threshold is held by the two parties in the form of secret sharing shards. In the inference process, for each non-leaf node, the shards of the feature values of the split feature of the target object are obtained according to the homomorphic operation between the homomorphic ciphertext of the feature number vector of the non-leaf node and the feature values of the target object, and through the homomorphic encryption to secret sharing protocol. And through the secret comparison jointly executed by the two parties, according to the shards of the split threshold of the non-leaf node, the shards of the secret comparison result indicating which sub-node the target object falls into are determined. The two parties jointly determine the shards of the leaf node array indicating the leaf node into which the target object falls according to the shards of the secret comparison result of each non-leaf node. Finally, the two parties jointly determine the prediction result of the target object according to the leaf node array and the shards of the prediction value of each leaf node.

[0066] The above method is for the scenario in which the feature number of the split feature of the non-leaf node is held by the non-feature party and the split threshold is held by the two parties. The feature value ciphertext of the split feature is obtained by homomorphic encryption of the feature number vector of the non-leaf node of the tree model and homomorphic operation between the feature value and the target object. Thus, the inference of the non-leaf node is completed according to the ciphertext of the feature value and the ciphertext of the split threshold, and the prediction value for the target object is obtained based on this.

[0067] Next, the tree model two-party joint inference method shown in this specification will be described through an embodiment.

[0068] The method in the specification is applied to a two-party joint inference scenario. For ease of illustration, the two parties are referred to as the first party and the second party. The first party and the second party hold feature values of different features of a target object. In other words, for a target object, the first party and the second party hold feature values of different types of features of the target object. The target object here is an object held by both parties. For example, the target object is a user, the first party can hold feature values of the user's account registration time, and the number of interactions with high-risk accounts, and the second party can hold feature values of the user's transaction frequency, whether or not to trade with a high-risk user.

[0069] The two parties also jointly hold a tree model. The tree model is the same as the tree model trained by the other two parties. The prediction value of the leaf node of the tree model is held by the two parties in the form of a secret share. Unlike the tree model obtained by the other two parties in the non-anonymous primary key scene, for any non-leaf node of the tree model, the feature number of the split feature of the non-leaf node of the tree model is held by the party that does not hold the feature value of the split feature, and the feature threshold of the split feature is held by the first party and the second party in the form of a split.

[0070] For example, the split feature of a certain non-leaf node is the user's account registration time, and the user's account registration time is numbered 3 among all the features of the target object held by the first party and the second party. The feature value of the account registration time of the target object is held by the first party, and then the feature number 3 of the split feature of the non-leaf node is held by the second party. At the same time, both parties hold a secret share of the split threshold of the split feature.

[0071] In addition, the two parties can jointly hold multiple tree models. Here, for ease of illustration, only the inference process of one tree model is described. It is easy to understand that the inference processes of other tree models are the same. After obtaining the prediction results of multiple tree models, the prediction results of multiple tree models can be added to obtain the final prediction result.

[0072] In addition, the two parties can also use the tree model to infer the prediction results of multiple target objects at the same time. Here, the inference process of one target object is described as an example, and the inference processes of other target objects are the same and will not be described here.

[0073] In an optional implementation, the tree model is trained by the second party according to a target training sample set, and the target training sample set is obtained by performing a privacy intersection protocol of an anonymous primary key on data of the object held by the second party. In other words, the tree model can be obtained by the tree model joint modeling method of an anonymous primary key, for example, can be trained by the method mentioned above. This scenario does not represent a limitation of the present specification, and it can be understood that if a tree model with the above characteristics can be obtained by other modeling methods, the joint inference of the two parties can also be performed by the method provided in the present specification.

[0074] Next, the steps of the method shown in the present specification will be described in combination with Figure 2 Here, the steps performed by the first party will be described, and the steps performed by the second party are the same as those performed by the first party, which will not be described here.

[0075] In step 201, for any non-leaf node, the homomorphic ciphertext of the feature number vector of the non-leaf node is obtained, and the feature value slice of the split feature of the target object is obtained by the second party in combination with the homomorphic ciphertext of the feature number vector and the feature value of each feature of the target object; and the second party is combined to perform a secret comparison on the slice of the feature value of the split feature and the slice of the feature threshold to obtain the slice of the secret comparison result.

[0076] The secret comparison result is used to indicate the node into which the target object falls after the non-leaf node; and the feature number vector is a one-hot vector whose value at the position corresponding to the feature number of the non-leaf node is 1.

[0077] Specifically, for any non-leaf node, the feature number of the split feature of the non-leaf node is held by the non-feature party, and the party that has the feature value of the split feature cannot obtain which type of feature is the split feature. Therefore, in order to complete the inference of the non-leaf node, first, the ciphertext of the feature value of the split feature of the target object needs to be obtained for each non-leaf node.

[0078] To solve the above problems, a method of homomorphic encryption of feature numbers of non-leaf nodes of a tree model is adopted in step 201. First, the feature party mentioned below is a party that has the splitting feature of the non-leaf node, and the non-feature party is a party that does not have the splitting feature of the non-leaf node. Specifically, first, the non-feature party constructs a feature number vector according to the feature numbers of the splitting features of the non-leaf node that it holds. The feature number vector has the same number of elements as the total number of features, and the elements of the vector are 1 at positions corresponding to the feature numbers and 0 at other positions. The non-feature party can homomorphically encrypt the elements of the feature number vector and send them to the feature party. The feature party performs homomorphic multiplication of each element in the feature number vector with each feature that it holds, respectively, and performs homomorphic addition of the results of the homomorphic multiplication, to obtain a homomorphic ciphertext of the feature value of the target object. The two parties jointly perform a homomorphic encryption-to-secret sharing protocol to obtain a secret sharing of the feature value of the target object, respectively.

[0079] For the convenience of understanding, first, homomorphic encryption is described. Homomorphic encryption algorithm is an encryption method commonly used in privacy calculation. Compared with ordinary encryption, homomorphic encryption has the property of homomorphism, that is, it can perform calculation on data in the encrypted state. For example, homomorphic addition operation of the homomorphic encryption ciphertext [a] of a and the unencrypted data b can obtain the ciphertext [a+b] of a+b, and the operation process does not leak a, b or a+b, nor does it need to use the decryption key. Moreover, homomorphic encryption also supports [a]*b=[a*b], and full homomorphic encryption algorithm also supports [a]*[b]=[a*b].

[0080] The homomorphic multiplication operation between the homomorphic ciphertext of the feature number vector and the feature values of each feature in step 201 is to multiply the elements at each position of the feature number vector with the feature values corresponding to the elements, respectively. Since in the plaintext of the feature number vector, only the elements at positions corresponding to the feature numbers are 1 and the other elements are 0, through the homomorphic multiplication operation, the plaintext at the position corresponding to the feature number in the vector obtained is the feature value corresponding to the feature number, and the plaintext corresponding to the elements at other positions is 0. Further, performing homomorphic addition operation on the elements at each position of the vector can obtain the feature value of the splitting feature while ensuring that the feature values, feature numbers, etc. are not leaked.

[0081] It should be further noted that the reason why the conversion into the form of secret sharing fragments is needed before secret comparison to complete the reasoning of the non-leaf node, instead of comparing in the form of homomorphic ciphertext, is that the secret comparison scheme of secret sharing in the related art is more efficient than the comparison scheme of homomorphic ciphertext, and the conversion into the form of secret sharing fragments can improve the overall reasoning efficiency.

[0082] It should be further noted that, in order to ensure data security, the non-feature party encrypts by using the public key of the non-feature party when encrypting. Correspondingly, the ciphertext can only be decrypted by the private key of the non-feature party, and the feature party cannot obtain the private key of the non-feature party and cannot decrypt it, thereby ensuring data security.

[0083] The feature party and the non-feature party can be different in different executions. For example, for the first non-leaf node, the first party can be the feature party, and for the second non-leaf node, the second party can be the feature party.

[0084] Next, step 201 will be described in detail. Step 201 includes two steps of obtaining the feature value slice of the split feature and secret comparison.

[0085] Next, taking a non-leaf node as an example, taking the first party as the feature party and the first party as the non-feature party respectively, the process of how the first party and the second party jointly obtain the feature value slice of the split feature of the target object will be described.

[0086] First, the first party is a non-feature party. Step 201 specifically includes: for any non-leaf node, in the case that the first party holds the feature number of the non-leaf node, a feature number vector is constructed according to the feature number, and the feature number vector is homomorphically encrypted.

[0087] The homomorphic ciphertext of the feature number vector is sent to the second party; the second party obtains the homomorphic ciphertext of the feature value of the split feature of the target object according to the homomorphic operation between the homomorphic ciphertext of the feature number vector and the feature value of each feature of the target object held by the second party;

[0088] The homomorphic ciphertext of the feature value of the split feature of the target object is jointly obtained by the second party according to the homomorphic ciphertext of the feature value of the split feature of the target object, and a homomorphic encryption to secret sharing protocol is executed to obtain the slice of the feature value of the split feature of the target object.

[0089] Wherein, when performing homomorphic encryption, each element in the feature number vector is homomorphically encrypted respectively, instead of homomorphically encrypting the feature number vector as a whole. The first step is the process of constructing the homomorphic ciphertext of the feature number vector, the second step is the process of the second party obtaining the homomorphic ciphertext of the feature value of the split feature, and the third step is the process of converting the homomorphic ciphertext of the feature value into a secret sharing slice.

[0090] Second, the first party is a feature party. Step 201 specifically includes: for any non-leaf node, in the case that the second party holds the feature number of the non-leaf node, receiving the homomorphic ciphertext of the feature number vector sent by the second party;

[0091] According to the homomorphic operation between the homomorphic ciphertext of the feature number vector and the feature values of each feature of the target object held by the first party, the homomorphic ciphertext of the feature value of the split feature of the target object is obtained;

[0092] The homomorphic ciphertext of the feature value of the split feature of the target object is sent to the second party, and a homomorphic encryption secret sharing protocol is jointly executed with the second party to obtain the split of the feature value of the split feature of the target object.

[0093] The first step is to receive the feature number vector subjected to homomorphic encryption. The second step is to calculate the homomorphic ciphertext of the feature value of the split feature, and the third step is to convert the homomorphic ciphertext of the feature value into a secret sharing split by the two parties.

[0094] Next, a possible implementation of a homomorphic encryption secret sharing protocol will be described. Specifically, the feature party obtains the homomorphic ciphertext of the feature value of the split feature by calculation, and can perform homomorphic addition operation on the homomorphic ciphertext of the feature value and a preset random number to obtain a first element. The first element is sent to the non-feature party, and the non-feature party decrypts the first element to obtain the split of the feature value of the split feature. Correspondingly, the feature party takes the opposite number of the random number as the split of the feature value of the split feature.

[0095] After obtaining the split of the feature value of the split feature of the target object, the two parties can jointly perform secret comparison. According to the size relationship between the feature value and the split threshold of the split feature, it can be determined whether the target object falls in the left child node or the right node of the non-leaf node, and then the split of the secret comparison result is obtained.

[0096] In an optional embodiment, the secret comparison result can be realized in a binary manner. For example, if the secret comparison result is a first value (such as 0), the target object falls in the left child node of the non-leaf node, and if the secret comparison result is a second value (such as 1), the target object falls in the right node of the non-leaf node. It should be noted that the first value of the first value of the path value in the following text can be different.

[0097] For the specific implementation of secret comparison, secret comparison protocols based on confusion circuit, bit decomposition scheme, etc. can be used.

[0098] Next, the steps of the secret comparison protocol will be described in detail with reference to the flowchart shown in FIG. 2. Figure 3

[0099] Figure 3 First, two participants P A and P B are described. Both parties hold a copy of the primary key identifier, which has been aligned, and the feature of the to-be-predicted data set D A and D​ B It can be understood that each sample in the prediction dataset can be regarded as the target object mentioned above.

[0100] where D A contains n samples and m A features, D B contains n samples and m B features. D A [idx][q] represents the qth feature value of the idxth sample in the dataset D A . P A holds m A features and P B holds m B features are different.

[0101] In addition, both parties hold a privacy-protected decision tree model Tree A and Tree B , specifically, the decision tree has a total of t trees, and the depth of each tree is d, that is, the number of non-leaf nodes of each tree is p = 2 d -1, and the number of leaf nodes is l = 2 d .

[0102] For the ith non-leaf node (i ∈ [p]), it contains the ciphertext fragments of the feature number f[i] and the feature threshold <h[i]>, wherein the feature number is held by the non-feature party, and the threshold is stored in the local of both parties in the form of secret shares <h[i]> A , <h[i]> B . For example, the feature corresponding to the feature number belongs to P A , then only P B knows the feature number f[i], in order to distinguish the ownership of the number, let f B [i] = f[i], f A [i] = 0; on the contrary, if the non-feature party is P A , then let f A [i] = f[i], f B [i] = 0.

[0103] For the jth leaf node (j ∈ [l]), the corresponding leaf node weight information <w[j]> is also stored in the local of both parties in the form of secret shares <w[j]> A , <w[j]> B .

[0104] As before, in this specification, when describing a certain secret sharing basic operation, the angle bracket <x>denotes that the current variable X is secret share encrypted, and P A and P B respectively hold secret share shards <x> A and <x> B , both parties need to be separately <x> A and <x> B The corresponding secret sharing operation is performed. In this paper, the current variable X is represented by square brackets [X] A , which means that X is homomorphic public key encryption of P A , and A in the upper right corner means that the private key of the ciphertext belongs to P A , only P A can decrypt the ciphertext.

[0105] Figure 3 The content in the blue dashed box is the reasoning process for the kth tree model, and idx is the identification of the target object. Specifically, in step 201, the following is performed:

[0106] 1. Both parties generate their own homomorphic encryption public and private keys locally, and send the public key to the other party. Specifically, P A generates a homomorphic encryption public key pk A and a private key sk A , and sends the public key pk A to P B ; similarly, P B generates a homomorphic encryption public key pk B and a private key sk B , and sends the public key pk B to P A .

[0107] 2. Both parties use homomorphic encryption technology to perform one-hot encoding of the feature number information f[i] of all non-leaf nodes of the tree model, and send the ciphertext to the feature party. Specifically, for the ith non-leaf node (i∈[p]), assume that the node feature number information f[i] is held by the non-feature party P B , i.e. f B [i] = f[i]. First, P B changes f B [i] to an one-hot encoded array U[i] (U[i] is an array of length m A , and only U[i][f B [i]] = 1, and the rest are 0), and encrypts the array U[i] using the public key pk B to get After completing all leaf node encryption operations, P A holds the ciphertext matrix P B holds the ciphertext matrix Both parties send and to the other party. For the idxth test data (i.e. the target object), the kth decision tree (idx∈[n], k∈[t]), both parties perform the operation.

[0108] 3. For the i-th non-leaf node (i∈[p]), the feature square uses the ciphertext and the characteristics D of the test sample A(B) [idx] calculates the ciphertext directly on the local machine. Then, using homomorphic ciphertext to secret sharing technology, we can obtain... <D A(B) [idx][f[i]]>.

[0109] Specifically, let's assume that the feature expression of the current i-th non-leaf node is P. A P A hold and D A [idx], first P A Homomorphic ciphertext can be calculated using homomorphic multiplication and homomorphic addition of plaintext and ciphertext. Then, P A Generate a random number r locally and use the public key pk. B Encryption Then, homomorphic addition is performed to obtain Subsequently, P A Will Send to P B P B Using private key sk B Decrypt the pseudo-random number D A [idx][f B [i]]+r; Finally, P A The random number - r is directly used as the secret sharing ciphertext for the feature value corresponding to the i-th non-leaf node, denoted as<v[i]> A =-r, while P B D A [idx][f B [i]]+r is the secret shared ciphertext of the feature value of the i-th non-leaf node, denoted as<v[i]> B =D A [idx][f B [i]]+r.

[0110] 4. For the i-th non-leaf node (i∈[p]), both parties need to protect the ciphertext feature value.<v[i]> and ciphertext threshold<h[i]> Without revealing the details, the secret decision-making process of the target object in the non-leaf node is obtained through secret comparison, and the ciphertext comparison result is secretly shared in bit form.<b[i]> bit express.

[0111] Specifically, for the i-th secret comparison, P A enter<v[i]> A and<h[i]> A , P B input <v[i]> B and <h[i]> B After both parties execute the agreement, P A obtains P B obtains and satisfies <v[i]>≤<h[i]> (i.e. falls into the left child node), which indicates that the correct decision path is to proceed to the left child node. Conversely, if <v[i]>><h[i]>, then which indicates that the correct decision path is to proceed to the right child node. After p times of the above calculation, the final parties can obtain the decision result ciphertext array of all non-leaf nodes of the kth tree bit Herein and are the shards of the secret comparison result.

[0112] wherein, represents the exclusive or operation. The calculation logic of the exclusive or operation is that the result is 0 when the two binary bits are the same, and the result is 1 when the two binary bits are different. The bit in the upper right corner of the above-mentioned angle bracket represents that it is a bit type secret sharing shard. The bit type secret sharing shard is different from the numerical value type secret sharing shard in that the bit type secret sharing shard guarantees that the exclusive or operation result of two shards is the corresponding plaintext.

[0113] Step 203, in combination with the second party, determines the shard of the leaf node array corresponding to the target object according to the shards of the secret comparison results of each non-leaf node and in combination with the second party.

[0114] wherein, the leaf node array is used to indicate the leaf node in which the target object falls.

[0115] Specifically, after obtaining the shards of the secret comparison results of each non-leaf node, the leaf node array in which the target object falls can be determined in combination with the second party according to the connection relationship between the nodes.

[0116] wherein, the number of elements in the leaf node array can be the same as the number of leaf nodes. In the leaf node array, the value of the element corresponding to the leaf node in which the target object falls is the first value, and the value of the element corresponding to the other leaf nodes is the second value. For example, the target object falls into the third leaf node, the first value is 1, and the second value is 0, so the third element in the leaf node array is 1, and the other elements are 0.

[0117] It should be noted that the leaf node array here is in the form of ciphertext, which can prevent a party from obtaining the specific prediction path, thereby avoiding a party from deducing the split threshold of the tree model according to the prediction results of multiple target objects, and better protecting data security.

[0118] Next, a possible implementation of step 203 will be described. Step 203 can include: for each non-leaf node, obtaining a shard of path values of child nodes of the non-leaf node from a product of a path value and a secret comparison result calculated jointly with the second party, the path value being a first value indicating that the target object falls into a node corresponding to the path value or a second value indicating that the target object does not fall into the node corresponding to the path value; grouping the shards of path values of the leaf nodes into an array according to an order of the leaf nodes to obtain a shard of a leaf node array, the leaf node array being a binary array in which a leaf node to which the target object falls corresponds to an element of a first value and other leaf nodes correspond to elements of a second value; and the first value being 1 and the second value being 0 in an optional implementation.

[0119] The plaintext of the path value is used to indicate whether the target object falls into the current node. For each non-leaf node, the path value of the child nodes of the non-leaf node can be determined according to the multiplication operation of the secret sharing between the shard of the path value and the shard of the secret comparison result. In this way, the shard of the path value of the leaf node can be obtained through iterative calculation. Further, the shards of the path values of the leaf nodes can be grouped into the shard of the leaf node array.

[0120] Next, step 203 will be further described in combination with Figure 3 and Figure 4 .

[0121] Specifically, both parties need to protect the secret of the ciphertext array bit Without being disclosed, the prediction result of each leaf node is performed Privacy computation of bit

[0122] In particular, both parties initialize a secret path array of size 2 d -1 <c> bit and in accordance with Figure 4 The structure of the tree is given by c[i], which indicates whether the i-th node is on the actual predicted path (i.e., the path value corresponding to the i-th node mentioned earlier). Since the root node is a necessary point, it satisfies the following condition.<c[1]> bit = <1> That is, holding<c[1]> bit The secret is shared using a 1-bit secret. Both parties traverse each non-leaf node in the order of the tree model. For the i-th node, they use bit-based secret multiplication to calculate the secret paths to the left and right leaf nodes. Since the secret comparison results use 0 and 1 to represent falling into the left and right child nodes respectively,

[0123] It should be noted that although the above formula describes the operation process in the form of fragments, it can be understood that the above process is a joint calculation by two parties. The plaintext of c[i] and b[i] is not disclosed during the calculation process, and the result obtained is a fragment of the corresponding plaintext operation result.

[0124] in, Representation and computation. The operation results in a value of 1 if both bits involved in the computation are 1, and 0 otherwise. For the formula involving the left child node,<c[i]> bit The partition representing the path value of the non-leaf node.<b[i]> bit A fragment representing the result of a secret comparison.

[0125] Since a value of 0 for b[i] indicates that the node falls into the left child node, it is necessary to calculate the XOR operation result of b[i] and 1 to make the expression equal to the value of b[i] in the expression.<c[i]> bit The part that performs AND operations A slice is assigned a value of 1 when it falls into a left leaf node and a value of 0 when it falls into a right leaf node. This ensures the accuracy of the final calculation result. For example, when b[i] is 0, it indicates that the secret comparison result is falling into the left child node, and the XOR operation of 0 and 1 results in 1, which makes... Share the 1-bit secret as a fragment. Finally, combine this value with...<c[i]> bit Perform a joint AND operation between the two parties to ensure that, if the secret comparison result indicates that the target object falls into the left child node and the path value c[i] is 1 (indicating that the target object falls into the current non-leaf node), the path value of the left child node is sliced.<c[2i]> bit A bit-type secret of 1 is used to share the slice, ensuring the accuracy of the result. The implementation is similar when b[i] and c[i] take other values, and will not be described in detail here.

[0126] Correspondingly, since the value of b[i] is 1 when the target object falls into the right child node, and the value of b[i] is 0 when the target object falls into the left child node, therefore, according to the result of the AND operation of the path value c[i] and b[i] indicating whether the target object falls into the current non-leaf node, it can be determined whether the target object falls into the right leaf node. For example, when the values of c[i] and b[i] are both 1, it indicates that the target object falls into the current non-leaf node, and falls into the right child node of the non-leaf node, and the result of the AND operation of the two is 1, indicating that the target object falls into the right leaf node. If the value of any one of c[i] and b[i] is 0, it can be determined that the target object will not fall into the right child node of the non-leaf node, and the plaintext of the result of the AND operation of c[i] and b[i] is 0, which also indicates that the target object will not fall into the right child node of the non-leaf node.

[0127] Final taking array <c> bit 2th d-1 to 2 d - a ciphertext value of one point, directly as the leaf node array mentioned above, denoted as <e[j]> bit <e[j]> d-1 + j - 1] bit , j e [l].

[0128] In step 205, the second party jointly performs a secret sharing-based operation according to the shards of the leaf node array and the shards of the preset leaf node prediction values, and the second party jointly obtains the shard of the prediction result corresponding to the target object.

[0129] Specifically, the shards of the preset leaf node prediction values can be first arranged in the order of the leaf nodes to form a shard of a prediction value array. The two parties jointly perform a secret sharing-based multiplication operation to calculate the result of multiplying the leaf node array and the prediction value array, and obtain a shard of a first array, which represents the result of multiplying the leaf node array and the prediction value array. The two parties then jointly perform a secret sharing-based addition operation between the bits of the first array to obtain the shard of the prediction result of the target object of the current tree model.

[0130] Here, the steps shown in Figure 3 Fig. 2 are still used for illustration. For the kth decision tree, the two parties use secret multiplication and addition to calculate the prediction result of the tree <e[j]> bit is a shard of the leaf node array, and <w[j]> is a shard of the prediction value array.

[0131] In addition, in the case of multiple tree models, the shards of the prediction results of the tree models can also be added to obtain the shard of the final prediction result of the target object. In other words, after the operation in the above steps, the two parties obtain a ciphertext array <res>Then both sides use secret addition to accumulate all the prediction results, and obtain the final ciphertext prediction result of the idxth test data

[0132] Finally, according to the business requirements, the prediction results <pred>The decryption is performed to enable the first party or the second party to obtain the final prediction result pred.

[0133] In other words, the secret sharing decryption can be performed in combination with the second party according to the shards of the prediction result to obtain the prediction result of the target object. Here, the party that finally obtains the prediction result of the target object can be at least one of the first party and the second party.

[0134] The above embodiments introduce the inference method of the two parties in the intersection primary key disclosure scenario of the prediction data set, that is, the primary keys of the intersection of each object (that is, the target object) owned by the two parties are plaintexts. The above embodiments require the ability to protect the model information and the training data set from being leaked while implementing the inference on the specified target object and disclosing the prediction result, that is, the participating party cannot infer the model information and the training data set information through the additional information obtained in the inference stage.

[0135] In the related art, there is also a demand for anonymized inference of a tree model with anonymized primary keys that cannot associate the prediction result with the primary key information of a single object, that is, a compliance scenario demand that can only reflect the overall trend or can only obtain statistical information.

[0136] Based on this, the present specification also provides a two-party joint inference method of a tree model with anonymized primary keys. This method takes the execution steps of the first party as an example for illustration. It can be understood that the steps executed by the second party are the same as those of the first party, and will not be described again. The method comprises:

[0137] Step 501: For a plurality of objects, a privacy intersection protocol of a tree model with anonymized primary keys is executed in combination with the second party to obtain a first shard of feature values of all features of a target object; the second party obtains a second shard of feature values of all features of the target object, and the target object is an object commonly held by the first party and the second party.

[0138] Specifically, the two parties first need to execute the privacy intersection protocol of the tree model with anonymized primary keys for the objects held by the two parties respectively, so that the two parties obtain the shards of the feature values of the target objects whose primary key identifiers are anonymized and whose feature values are in the form of ciphertexts. In the case of multiple target objects (the two parties do not know which objects are the target objects), the all-ciphertext matrix <D A > and <D B > can be obtained through the privacy intersection protocol of the tree model with anonymized primary keys. A holds <D A > A and <D B > A holds <D B > A and <D B > B holds <D B >That is, both parties hold the feature value fragments of the features of the first party P A and the features of the second party P B .

[0139] In step 502, for any non-leaf node, the homomorphic ciphertext of the feature number vector of the non-leaf node is obtained, and the feature value fragments of the split features of the target object are obtained by the second party in combination with the homomorphic ciphertext of the feature number vector and the feature values of the features of the target object.

[0140] Different from step 201, after obtaining the homomorphic ciphertext of the feature value fragments of the split features sent by the feature party, the ciphertext needs to be summed with the locally held feature value fragments of the split features to obtain the final feature value fragments of the split features.

[0141] Here, the implementation of step 502 is described by taking a specific example of a first party as a feature party and a second party as a feature party.

[0142] First, for any non-leaf node, in the case that the first party holds the feature number of the non-leaf node, a feature number vector is constructed according to the feature number, and the feature number vector is homomorphically encrypted;

[0143] The homomorphic ciphertext of the feature number vector is sent to the second party. The second party obtains the homomorphic ciphertext of the second feature value fragment of the split features of the target object by homomorphic operation between the homomorphic ciphertext of the feature number vector and the second feature value fragment of the features of the target object held by the second party;

[0144] The first value sent by the second party is received. The first value is calculated according to the homomorphic operation result between the homomorphic ciphertext of the feature number vector and the second feature value fragment of the features held by the second party, and the sum of a preset random number;

[0145] The homomorphic decryption result of the first value and the sum of the first feature value fragment of the split features of the target object are taken as the processed feature value fragments of the split features of the target object. The second party takes the inverse of the random number as the processed feature value fragments of the split features of the target object.

[0146] Here, the first and second in the first and second above have similar meanings, and are used to refer to the party to which the fragments belong. The first fragment means that the fragment belongs to the first party, and the second fragment means that the fragment belongs to the second party.

[0147] Specifically, the first value is calculated by performing homomorphic multiplication operation on the homomorphic ciphertext of the feature number vector and the second split of each feature value, and then performing homomorphic addition operation on the results of each homomorphic multiplication operation to obtain the first value. After obtaining the first value, the first party decrypts the first value and sums the first value with the first split of the feature value of the split feature held by the first party to obtain the feature value split of the target object held by the first party.

[0148] Second, for any non-leaf node, in the case that the second party holds the feature number of the non-leaf node, the homomorphic ciphertext of the feature number vector sent by the second party is received;

[0149] According to the homomorphic operation between the homomorphic ciphertext of the feature number vector and the first split of the feature value of each feature of the target object held by the first party, the homomorphic ciphertext of the first split of the feature value of the split feature of the target object is obtained.

[0150] The sum of the homomorphic ciphertext of the first split of the feature value of the split feature of the target object and a preset random number is calculated to obtain a second value, and the second value is sent to the second party;

[0151] The inverse number of the random number is taken as the split of the feature value of the processed split feature of the target object; the second party takes the sum of the second value and the second split of the feature value of the split feature of the target object held by the second party as the split of the feature value of the processed split feature of the target object.

[0152] Most of the implementation manners in the above steps can be referred to the description in step 201, which will not be repeated here.

[0153] Here still use Figure 3 the identification in the example for specific description.

[0154] First, the steps 1 and 2 in the description in step 201 are executed. Figure 3 The steps 1 and 2 in the description in step 201 are executed.

[0155] 3', for the i-th non-leaf node (i∈[p]), the feature party uses the ciphertext and the feature <D A(B) [idx]> A(B) The ciphertext is directly calculated at the local, and then the <D A(B) [idx][f[i]>] is obtained by using the homomorphic ciphertext to secret sharing technology. Specifically, it is assumed that the feature party of the current i-th non-leaf node is P A , P A holds and <D A [idx]> A , first P A Homomorphic ciphertext can be calculated using homomorphic multiplication and homomorphic addition of plaintext and ciphertext. Then, P A Generate a random number r locally and use the public key pk. B Encryption Then, homomorphic addition is performed to obtain Subsequently, P A Will Send to P B P B Using private key sk B Decrypt the pseudo-random number <D A [idx][f B [i]]> A +r; Finally, P A The random number - r is directly used as the secret sharing ciphertext for the feature value corresponding to the i-th non-leaf node, denoted as<v[i]> A =-r, while P B Use f directly on the local machine B [i]In <D A [idx]> B Selected <D A [idx][f B [i]]> B and will ( <D A [idx][f B [i]]> A +r)+ <D A [idx][f B [i]]> B The ciphertext shared as the feature value secret of the i-th non-leaf node is denoted as .<v[i]> B =D A [idx][f B [i]]+r.

[0156] The subsequent steps are the same as the secret comparison mentioned in step 201 and steps 203-205, and will not be repeated here.

[0157] like Figure 5 As shown, this specification also provides a tree model two-party joint reasoning device, in which the first party and the second party respectively hold feature values ​​of different features of the target object; the first party and the second party jointly hold the tree model, the feature number of the split feature of the non-leaf node of the tree model is held by the party that does not hold the feature value of the split feature, and the feature threshold of the split feature is held by the first party and the second party in the form of fragments; the device includes:

[0158] The secret comparison module 510 is configured to, for any non-leaf node, obtain homomorphic ciphertext of a feature number vector of the non-leaf node, and jointly with the second party, obtain a feature value slice of a split feature of the target object according to the homomorphic ciphertext of the feature number vector and feature values of each feature of the target object; and jointly with the second party, perform secret comparison according to the feature value slice of the split feature and a feature threshold slice to obtain a feature value slice of a secret comparison result; the feature value slice of the secret comparison result is used to indicate a node into which the target object falls after the non-leaf node; the feature number vector is a one-hot vector in which values corresponding to feature numbers of the non-leaf node are 1;

[0159] The leaf node array determination module 520 is configured to jointly with the second party, determine a feature value slice of a leaf node array corresponding to the target object according to the feature value slice of the secret comparison result of each non-leaf node; the leaf node array is used to indicate a leaf node into which the target object falls.

[0160] The prediction result acquisition module 530 is configured to jointly with the second party, perform secret sharing-based operation according to the feature value slice of the leaf node array and a feature value slice of each preset leaf node prediction value to obtain a feature value slice of a prediction result of the target object.

[0161] In an optional implementation, the secret comparison module 510 is specifically configured to, for any non-leaf node, in a case where the first party holds a feature number of the non-leaf node, construct a feature number vector according to the feature number, and perform homomorphic encryption on the feature number vector; and send the homomorphic ciphertext of the feature number vector to the second party; the second party obtains homomorphic ciphertext of feature values of a split feature of the target object according to homomorphic operation between the homomorphic ciphertext of the feature number vector and feature values of each feature of the target object held by the second party; and jointly with the first party, performs homomorphic encryption to secret sharing protocol according to the homomorphic ciphertext of the feature values of the split feature of the target object to obtain a feature value slice of the split feature of the target object.

[0162] In an optional implementation, the secret comparison module 510 is specifically configured to, for any non-leaf node, in a case where the second party holds a feature number of the non-leaf node, receive the homomorphic ciphertext of the feature number vector sent by the second party; obtain homomorphic ciphertext of feature values of a split feature of the target object according to homomorphic operation between the homomorphic ciphertext of the feature number vector and feature values of each feature of the target object held by the first party; send the homomorphic ciphertext of the feature values of the split feature of the target object to the second party, and jointly with the second party, perform homomorphic encryption to secret sharing protocol to obtain a feature value slice of the split feature of the target object.

[0163] In an optional implementation, the leaf node array determination module 520 is specifically configured to, for each non-leaf node, calculate a product of a path value and a secret comparison result jointly calculated with the second party, obtain a shard of a path value of a left child node of the non-leaf node; the path value is a first value when the target object falls into a node corresponding to the path value, and is a second value when the target object does not fall into the node corresponding to the path value; the shards of the path values of the leaf nodes are arranged into an array according to an order of the leaf nodes to obtain a shard of a leaf node array; the leaf node array is a binary array, and in the leaf node array, a value of an element corresponding to a leaf node in which the target object falls is the first value, and values of elements corresponding to other leaf nodes are the second value.

[0164] In an optional implementation, the privacy intersection module 500 (not shown in the figure) is further included, and is configured to, for a plurality of objects, jointly execute a privacy intersection protocol of an anonymous primary key with the second party to obtain a first shard of a feature value of each feature of the target object; the second party obtains a second shard of a feature value of each feature of the target object, and the target object is an object jointly held by the first party and the second party.

[0165] In an optional implementation, the secret comparison module 510 is specifically configured to: for any non-leaf node, in a case where the first party holds a feature number of the non-leaf node, construct a feature number vector according to the feature number, and homomorphically encrypt the feature number vector; send the homomorphic ciphertext of the feature number vector to the second party; the second party obtains homomorphic ciphertext of the second shard of the feature value of the split feature of the target object according to a homomorphic operation between the homomorphic ciphertext of the feature number vector and the second shard of the feature value of each feature of the target object held by the second party; receive the first value sent by the second party; the first value is calculated according to a sum of a homomorphic operation result between the homomorphic ciphertext of the feature number vector and the second shard of each feature value held by the second party and a preset random number; and a sum of a homomorphic decryption result of the first value and the first shard of the feature value of the split feature of the target object is taken as a processed feature value shard of the split feature of the target object; the second party takes an inverse number of the random number as the processed feature value shard of the split feature of the target object.

[0166] In an optional implementation, the secret comparison module 510 is specifically configured to: for any non-leaf node, in a case where the second party holds the feature number of the non-leaf node, receive homomorphic ciphertext of the feature number vector sent by the second party; obtain homomorphic ciphertext of the first feature value shard of the split feature of the target object according to homomorphic operation between the homomorphic ciphertext of the feature number vector and the first feature value shard of each feature of the target object held by the first party; calculate a sum of the homomorphic ciphertext of the first feature value shard of the split feature of the target object and a preset random number to obtain a second value, and send the second value to the second party; and take an inverse number of the random number as the shard of the feature value of the processed split feature of the target object. The second party takes a sum of the second value and a second shard of the feature value of the split feature of the target object held by the second party as the shard of the feature value of the processed split feature of the target object.

[0167] In an optional implementation, the method further includes a decryption module 540 (not shown in the figure) configured to jointly perform secret sharing decryption with the second party according to the shard of the prediction result to obtain the prediction result of the target object.

[0168] In an optional implementation, the tree model is trained by the first party and the second party jointly according to a target training sample set, and the target training sample set is obtained by performing a privacy intersection protocol of an anonymous primary key on data of the held object jointly with the second party.

[0169] In the 1990s, it was quite obvious to distinguish whether an improvement in a technology was in hardware (e.g., improvement in circuit structures of diodes, transistors, switches, etc.) or in software (improvement in method flow). However, as technology has evolved, many improvements in method flow today can be considered as direct improvements in hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structures by programming the improved method flow into hardware circuits. Therefore, it cannot be said that an improvement in a method flow cannot be implemented by hardware entity modules. For example, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)) is an integrated circuit whose logic function is determined by user programming of the device. A digital system is "integrated" on a PLD by the designer programming it themselves, without having to ask a chip manufacturer to design and fabricate a custom integrated circuit chip. Moreover, instead of manually fabricating integrated circuit chips, this programming is now mostly implemented using "logic compiler" software, which is similar to software compilers used in program development, and the original code to be compiled is also written in a specific programming language, which is called a hardware description language (HDL), and there are many such languages, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc., and the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should be aware that, as long as the method flow is logically programmed in the above-mentioned hardware description languages and programmed into an integrated circuit, a hardware circuit that implements the logical method flow can be easily obtained.

[0170] The controller can be implemented in any suitable way, for example, the controller can take the form of, for example, a microprocessor or processor and a computer readable medium storing computer readable program code, such as software or firmware, executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20 and Silicone Labs C8051F320, the memory controller can also be implemented as part of the control logic of the memory. The skilled person will also appreciate that, in addition to implementing the controller in pure computer readable program code, it is possible to implement the controller in the form of logic gates, switches, an application specific integrated circuit, a programmable logic controller and an embedded microcontroller, etc. to perform the same functions by logically programming the method steps. Such a controller can therefore be considered to be a hardware component, and the means included therein to perform the various functions can also be considered to be structures within the hardware component. Alternatively, or even additionally, the means to perform the various functions can be considered to be both a software module implementing the method and a structure within a hardware component.

[0171] The systems, apparatuses, modules or units illustrated by the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, the present specification does not exclude that with the development of computer technology in the future, computers implementing the functions of the above embodiments can be personal computers, laptop computers, vehicle human-computer interaction devices, cellular phones, camera phones, smart phones, personal digital assistants, media players, navigation devices, electronic mail devices, game consoles, tablet computers, wearable devices, or combinations of any of these devices.

[0172] Although the method operations of the embodiments of the present disclosure are described in a particular, sequential order, one or more of the method operations can be omitted, or the method operations can be performed in an order other than the described order. Additionally, one or more of the method operations can be performed concurrently, or with partial concurrence. Furthermore, one or more of the method operations can be performed by different entities, or over different time periods. The term "include" is used to mean "including but not limited to." Where only one of an item is illustrated is described as comprising, includes, or consists of, the term "comprising," "including," or "consists of" is intended to mean the item only. Where more than one item is illustrated as comprising, including, or consisting of, the terms "comprising," "including," or "consists of" are intended to mean the items in the alternative (i.e. the items are one of the items, but the other of the items are not the items). Where only one item is illustrated as comprising, including, or consisting of, the term "comprising," "including," or "consists of" is intended to mean the item only. As used herein, the article "a" is intended to include one or more items. Where only one item is illustrated as comprising, including, or consisting of, the term "comprising," "including," or "consists of" is intended to mean the item only. Where a combination is illustrated, the combination can be achieved by means of an exemplary combination or by means of components of a system or apparatus, where the components can or can not be physically combined with one another. The term "and / or" means "and" or "or". The term "about" means ±10% of the value of the Uncleared.

[0173] For ease of description, the above apparatuses are described in various modules with their respective functions. Of course, when implementing one or more of the present disclosure, the functions of the modules can be implemented in one or more software and / or hardware, or the modules implementing the same function can be implemented by a combination of sub-modules or sub-units. The above-described apparatus embodiments are only illustrative, for example, the division of the units is only a logical functional division, and in actual implementation, there can be another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed ones can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0174] The present disclosure is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus produce a means for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions of one or more flows and / or blocks Figure 1 The functions of one or more flows and / or blocks

[0175] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.

[0176] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.

[0177] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0178] The memory can include non-persistent memory and / or volatile memory, such as random access memory (RAM) and / or cache memory, non-volatile memory, such as read-only memory (ROM), EPROM, and / or flash memory. The memory is an example of computer-readable media.

[0179] Computer-readable media includes permanent and non-permanent, moveable and non- moveable media that can be implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile discs (DVDs) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, graphene storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to computing devices. According to the definition herein, computer-readable media does not include transitory media, such as modulated data signals and carrier waves.

[0180] Those skilled in the art will appreciate that the one or more embodiments described herein can be provided as a method, a system or a computer program product. Accordingly, the one or more embodiments described herein can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the one or more embodiments described herein can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable code.

[0181] The one or more embodiments described herein can be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform particular tasks or implement particular abstract data types. The one or more embodiments described herein can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules can be located in both local and remote computer storage media including memory storage devices.

[0182] The various embodiments described in this specification can be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform particular tasks or implement particular abstract data types. The various embodiments described in this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules can be located in both local and remote computer storage media including memory storage devices.

[0183] The above description is only some embodiments of the one or more embodiments described in this specification and is not intended to limit the one or more embodiments described in this specification. The one or more embodiments described in this specification can have various modifications and changes for those skilled in the art. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the one or more embodiments described in this specification shall be included in the scope of claims.< / pred> < / res> < / c> < / c> < / x> < / x> < / x> < / x> < / x> < / x> < / x> < / x> < / x> < / w> < / h> < / g> < / h> ​​< / g> < / h> < / g> < / h> < / g> < / d> < / d> < / x> < / x> < / x> < / y> ​​< / d>

Claims

1. A tree model two-party joint inference method applied to a first party, the method also involving a second party, the first party and the second party respectively holding feature values of different features of a target object; the first party and the second party jointly holding a tree model, the feature number of a split feature of a non-leaf node of the tree model being held by a party that does not hold the feature value of the split feature, and the feature threshold of the split feature being held in the form of a shard by the first party and the second party; the method comprising: for any non-leaf node, obtaining homomorphic ciphertext of a feature number vector of the non-leaf node, and jointly with the second party obtaining a feature value shard of a split feature of the target object according to the homomorphic ciphertext of the feature number vector and the feature values of each feature of the target object; and jointly with the second party performing a secret comparison according to the feature value shard of the split feature and a feature threshold shard to obtain a feature value shard of a secret comparison result; the secret comparison result being used to indicate a node into which the target object falls after the non-leaf node; the feature number vector being a one-hot vector in which a value of 1 is taken at a position corresponding to the feature number of the non-leaf node; jointly with the second party determining a feature value shard of an array of leaf nodes corresponding to the target object according to the feature value shard of the secret comparison result of each non-leaf node; the array of leaf nodes being used to indicate a leaf node into which the target object falls; and jointly with the second party performing an operation based on secret sharing according to the feature value shard of the array of leaf nodes and a feature value shard of each leaf node prediction value predetermined in advance to obtain a feature value shard of a prediction result of the target object. 2.The method of claim 1, wherein the obtaining, for any non-leaf node, homomorphic ciphertext of a feature number vector of the non-leaf node, and jointly with the second party obtaining a feature value shard of a split feature of the target object according to the homomorphic ciphertext of the feature number vector and the feature values of each feature of the target object comprises: for any non-leaf node, in a case where the first party holds the feature number of the non-leaf node, constructing a feature number vector according to the feature number, and homomorphically encrypting the feature number vector; and sending the homomorphic ciphertext of the feature number vector to the second party; the second party obtaining homomorphic ciphertext of the feature values of the split feature of the target object according to a homomorphic operation between the homomorphic ciphertext of the feature number vector and the feature values of each feature of the target object held by the second party; and jointly with the first party performing a homomorphic encryption to secret sharing protocol according to the homomorphic ciphertext of the feature values of the split feature of the target object to obtain a feature value shard of the split feature of the target object. 3.The method of claim 1, wherein the obtaining, for any non-leaf node, homomorphic ciphertext of a feature number vector of the non-leaf node, and jointly with the second party obtaining a feature value shard of a split feature of the target object according to the homomorphic ciphertext of the feature number vector and the feature values of each feature of the target object comprises: for any non-leaf node, in a case where the second party holds the feature number of the non-leaf node, receiving homomorphic ciphertext of a feature number vector sent by the second party. ​ ​ ​ ​ ​ ​ ​ ​ performing homomorphic operation between the homomorphic ciphertext of the feature number vector and the feature values of each feature of the target object held by the second party, to obtain homomorphic ciphertext of the feature values of the split features of the target object; sending the homomorphic ciphertext of the feature values of the split features of the target object to the second party, and jointly performing a homomorphic encryption secret sharing protocol with the second party to obtain the split of the feature values of the split features of the target object.

4. The method of claim 1, wherein the split of the secret comparison result of each non-leaf node and the split of the array of leaf nodes corresponding to the target object are determined jointly by the first party and the second party, comprising: for each non-leaf node, obtaining the split of the path values of the left node and the right node of the non-leaf node according to the product of the path value calculated jointly by the two parties and the secret comparison result; the path value being the first value indicates that the target object falls into the node corresponding to the path value; the path value being the second value indicates that the target object does not fall into the node corresponding to the path value; the split of the path values of each leaf node is arranged in the order of the leaf nodes to obtain the split of the array of leaf nodes; the array of leaf nodes is a binary array, and the value of the element corresponding to the leaf node where the target object falls into is the first value, and the values of the elements corresponding to the other leaf nodes are the second value.

5. The method of claim 1, further comprising: for a plurality of objects, performing a privacy intersection protocol of the anonymous primary key jointly with the second party to obtain the first split of the feature values of all features of the target object; the second party obtains the second split of the feature values of all features of the target object, and the target object is an object held by the first party and the second party.

6. The method of claim 5, wherein the homomorphic ciphertext of the feature number vector of any non-leaf node is obtained, and the split of the feature values of the split features of the target object is obtained jointly by the first party and the second party according to the homomorphic ciphertext of the feature number vector and the feature values of each feature of the target object, comprising: for any non-leaf node, the feature number of the non-leaf node is held by the first party, the feature number vector is constructed according to the feature number, and the feature number vector is homomorphic encrypted; sending the homomorphic ciphertext of the feature number vector to the second party; the second party obtains the homomorphic ciphertext of the second split of the feature values of the split features of the target object by performing homomorphic operation between the homomorphic ciphertext of the feature number vector and the second split of the feature values of each feature of the target object held by the second party; receiving the first value sent by the second party; the first value is calculated according to the homomorphic operation result between the homomorphic ciphertext of the feature number vector and the second split of each feature value held by the second party, and the sum of the preset random number; the sum of the homomorphic decryption result of the first value and the first split of the feature values of the split features of the target object is taken as the split of the feature values of the split features of the processed target object; the second party takes the opposite number of the random number as the split of the feature values of the split features of the processed target object.

7. The method of claim 5, wherein the homomorphic ciphertext of the feature number vector of the arbitrary non-leaf node is obtained, and the feature value slice of the split feature of the target object is obtained according to the homomorphic ciphertext of the feature number vector and the feature values of the features of the target object held by the second party. The homomorphic ciphertext of the feature number vector of the arbitrary non-leaf node is received, in the case that the second party holds the feature number of the non-leaf node. The homomorphic ciphertext of the first feature value slice of the split feature of the target object is obtained according to the homomorphic operation between the homomorphic ciphertext of the feature number vector and the first feature value slice of the features of the target object held by the first party. The sum of the homomorphic ciphertext of the first feature value slice of the split feature of the target object and a preset random number is calculated to obtain a second value, and the second value is sent to the second party. The inverse of the random number is taken as the slice of the feature values of the processed split feature of the target object. The second party takes the sum of the second value and the second slice of the feature values of the split feature of the target object held by the second party as the slice of the feature values of the processed split feature of the target object.

8. The method of claim 1, further comprising: performing secret sharing decryption according to the slice of the prediction result and the second party to obtain the prediction result of the target object.

9. The method of claim 1, wherein the tree model is trained by the second party according to a target training sample set, and the target training sample set is obtained by performing a privacy intersection protocol on the data of the held object and the second party.

10. A computing device comprising a memory and a processor, wherein the memory stores executable code, and the processor executes the executable code to implement the method of any one of claims 1-9.

Citation Information

Patent Citations

  • Tree model-based two-party joint classification method, device, equipment and medium

    CN111639367A

  • Multi-party joint cause fruit tree model construction system and method based on federal learning

    CN116757286A