Tree model two-party joint reasoning method and computer equipment

Through homomorphic encryption and secret sharing protocols, the data privacy leakage problem of the two-party joint tree model inference under the anonymous primary key is solved, and a safe and efficient tree model prediction is achieved.

CN119990329APending Publication Date: 2025-05-13ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510201522.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing technology lacks a method for two parties to jointly conduct tree model inference under the protection of anonymous primary key information, resulting in a high risk of data privacy leakage.

Method used

Through homomorphic encryption and secret sharing protocols, we ensure that the split feature number and threshold of non-leaf nodes are held by non-feature parties. The shards of split feature values are obtained by using homomorphic encryption to secret sharing protocol, secret comparison is carried out to determine the falling node of the target object, and prediction is made based on the leaf node array and the shards of predicted values.

Benefits of technology

Without leaking the eigenvalue plaintext, secure tree model inference and prediction is realized, data privacy is protected, and inference efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990329A_ABST
    Figure CN119990329A_ABST
Patent Text Reader

Abstract

The invention provides a tree model two-party joint reasoning method. The first party holds the feature value of the first feature of the target object, and the second party holds the feature value of the second feature of the target object. The first party and the second party also jointly hold a tree model, the feature number of the split feature of the non-leaf node in the tree model is held by the non-feature party, and the split threshold is held by the two parties in the form of secret sharing fragments. And for each non-leaf node, obtaining homomorphic ciphertext of each feature value of the opposite side, and for each non-leaf node, obtaining fragments of the feature values of the splitting features of the target object through a homomorphic encryption-to-secret sharing protocol. And through secret comparison jointly executed by the two parties, according to the fragment of the split threshold value of the non-leaf node, the ciphertext result of which sub-node the target object falls into is determined. And the two parties jointly determine the ciphertext result of which leaf node the target object falls into according to the fragment of the secret comparison result of each non-leaf node. And finally, the two parties jointly determine a prediction result of the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification belong to the field of privacy computing technology, and in particular, relate to a tree model two-party joint reasoning method and computer equipment. Background Art

[0002] A decision tree model is a tree structure. Through the decision tree model, the attributes of the target object can be scored or classified according to the input feature values ​​of the target object to obtain the final reasoning result. With the requirements of data security, privacy computing methods that take into account privacy protection requirements and data utilization have been widely used. In related technologies, multiple participants use their respective data to perform joint privacy computing, such as jointly modeling a tree model, jointly reasoning on a tree model, etc.

[0003] As the requirements for data security in various industries continue to increase, the requirements for fully anonymous privacy-preserving machine learning are further required, so that a more secure tree model can be built under the conditions of anonymous data primary key information and full data process encryption. Among them, the primary key refers to the identification information that can distinguish the data. However, the tree model for anonymous primary key modeling in related technologies lacks a two-party joint reasoning method. Summary of the invention

[0004] The purpose of this specification is to provide a tree model two-party joint reasoning method and computer device.

[0005] In a first aspect, the present specification provides a two-party joint reasoning method for a tree model, which is applied to a first party. The method also involves a second party, wherein the first party holds a feature value of a first feature of a target object, and the second party holds a feature value of a second feature of the target object; the first party and the second party jointly hold a tree model, and the feature number of the split feature of a non-leaf node of the tree model is held by the party that does not hold the feature value of the split feature, and the feature threshold of the split feature is held by the first party and the second party in the form of fragments; the method includes:

[0006] For each second feature, obtaining a homomorphic ciphertext of a second feature value of each second feature of the target object;

[0007] For any non-leaf node, jointly with the second party, according to the homomorphic ciphertext of each second characteristic value, execute the homomorphic encryption to secret sharing protocol to obtain the sharding of the characteristic value of the split characteristic of the target object; and jointly with the second party, perform secret comparison based on the sharding of the characteristic value of the split characteristic and the sharding of the characteristic threshold to obtain the sharding of the secret comparison result; the secret comparison result is used to indicate the node into which the target object falls after the non-leaf node;

[0008] Determine, in conjunction with the second party, a slice of a leaf node array corresponding to the target object based on the slices of the secret comparison results of each non-leaf node; the leaf node array is used to indicate the leaf node into which the target object falls;

[0009] The second party jointly performs a secret sharing-based operation based on the slices of the leaf node array and the slices of the preset prediction values ​​of each leaf node to obtain the slices of the prediction results of the target object.

[0010] The second aspect of the present specification provides a two-party joint reasoning device for a tree model, wherein a first party holds a feature value of a first feature of a target object, and a second party holds a feature value of a second feature of the target object; the first party and the second party jointly hold a tree model, the feature number of the split feature of a non-leaf node of the tree model is held by the party that does not hold the feature value of the split feature, and the feature threshold of the split feature is held by the first party and the second party in the form of slices; the device comprises:

[0011] A ciphertext acquisition module, used for acquiring, for each second feature, a homomorphic ciphertext of a second feature value of each second feature of the target object;

[0012] The secret comparison module is used to execute the homomorphic encryption to secret sharing protocol for any non-leaf node in conjunction with the second party based on the homomorphic ciphertext of each second characteristic value, and obtain the sharding of the characteristic value of the split characteristic of the target object; and to perform secret comparison with the second party based on the sharding of the characteristic value of the split characteristic and the sharding of the characteristic threshold to obtain the sharding of the secret comparison result; the secret comparison result is used to indicate the node into which the target object falls after the non-leaf node;

[0013] A leaf node array determination module, used to determine, in conjunction with the second party, a slice of a leaf node array corresponding to the target object based on the slice of the secret comparison result of each non-leaf node; the leaf node array is used to indicate the leaf node into which the target object falls;

[0014] The prediction result acquisition module is used for the second party to jointly perform a secret sharing-based operation based on the slices of the leaf node array and the slices of the preset prediction values ​​of each leaf node to obtain the slices of the prediction results of the target object.

[0015] A third aspect of the present specification provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the above-mentioned tree model two-party joint reasoning method.

[0016] A fourth aspect of the present specification provides a computing device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the above-mentioned tree model two-party joint reasoning method is implemented.

[0017] A fifth aspect of the present specification provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned tree model two-party joint reasoning method.

[0018] The present specification provides a two-party joint reasoning method for a tree model. The method involves a first party and a second party, wherein the first party holds a feature value of a first feature of a target object, and the second party holds a feature value of a second feature of the target object. The first party and the second party also jointly hold a tree model, in which the feature number of the split feature of a non-leaf node is held by the non-feature party, and the split threshold is held by both parties in the form of a secret sharing shard. During the reasoning process, the homomorphic ciphertext of each feature value of the other party is obtained, and for each non-leaf node, the second party is jointly obtained according to the homomorphic ciphertext of each feature value, through a homomorphic encryption to secret sharing protocol, to obtain the shard of the feature value of the split feature of the target object. And through the secret comparison jointly performed by the two parties, according to the shard of the split threshold of the non-leaf node, the shard of the secret comparison result indicating which child node the target object falls into is determined. The two parties then jointly determine the shard of the leaf node array indicating the leaf node into which the target object falls according to the shard of the secret comparison result of each non-leaf node. Finally, the two parties jointly determine the prediction result of the target object according to the shard of the leaf node array and the prediction value of each leaf node.

[0019] The above method aims at the scenario where the feature number of the split feature of the non-leaf node is held by the non-feature party and the split threshold is held by both parties. The feature value is encrypting by homomorphic encryption, ensuring that the plaintext of the feature value is not leaked. The homomorphic encryption is converted to a secret sharing protocol to obtain the shards of the feature value of the split feature, thereby completing the reasoning of the non-leaf node by using secret comparison, and based on this, the predicted value acquisition for the target object is completed. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of this specification, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0021] Figure 1 is a schematic diagram of a tree model in one embodiment;

[0022] Figure 2 is a flow chart of a two-party joint reasoning method of a tree model in one embodiment;

[0023] Figure 3 is a flow chart of a two-party joint reasoning method of a tree model in another embodiment;

[0024] Figure 4 is a schematic diagram of a tree model prediction path in one embodiment;

[0025] Figure 5 It is a block diagram of a two-party joint reasoning device of a tree model in one embodiment. DETAILED DESCRIPTION

[0026] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.

[0027] The decision tree model generally exists in the form of a binary tree. The tree model includes leaf nodes and non-leaf nodes. When making inferences, the leaf node where the sample falls is determined based on the splitting conditions of the non-leaf nodes, and the score or classification of the sample is determined based on the weight value of the leaf node where the sample falls. The splitting conditions include the splitting features and the splitting thresholds corresponding to the splitting features. For example, Figure 1 As shown, Figure 1 It is a simple decision tree. The content in the non-leaf node represents the splitting condition of the node, A is the splitting feature, and the number is the corresponding splitting threshold. The content in the leaf node represents the weight value of the leaf node. When the sample is smaller than the corresponding splitting threshold, it will fall into the child node on the left side of the node (hereinafter referred to as the left child node), otherwise it will fall into the child node on the right side (hereinafter referred to as the right node), until the sample falls into a leaf node. For example, if the feature value of feature A of a sample is 1, then according to the splitting condition judgment of each non-leaf node, it can be determined that the sample falls into the leftmost leaf node, then the score value of the sample is 1. It should be noted that Figure 1 What is shown is a simple decision tree model. The decision tree model actually used may not generate the judgment conditions of non-leaf nodes based on only one feature, but may also complete the classification or scoring of samples through multiple trees at the same time.

[0028] There are many training methods for tree models, generally including Gradient Boosting DecisionTree (GBDT), Extreme Gradient Boosting (XGBoost), etc. Here we will briefly introduce GBDT. GBDT is an iterative decision tree algorithm. By iteratively training multiple decision trees, each decision tree is trained on the residual (the difference between the predicted value and the actual value) of the previous tree. During the training process, the structure and parameters of each decision tree are continuously optimized by the gradient boosting method to minimize the loss function of the model. In the prediction stage, the prediction results of multiple decision trees are summed or weighted averaged to obtain the final prediction result. Gradient decision trees are usually used to solve regression and classification problems. For example, a tree model can be trained to classify the financial credit risk of users, so as to quickly identify users with credit risks (such as fraud risks).

[0029] In the scenario of jointly training tree models with privacy protection, the participants (two examples are used here) are generally required to determine the common samples, that is, the intersection samples. Specifically, each participant generally holds the primary key and feature value corresponding to each training sample, and different participants hold different feature values. The two participants can determine the intersection samples with the same primary key based on the primary key.

[0030] Among them, the primary key is the identification information that can distinguish different training samples. For example, in the scenario of classifying the financial credit risk of users, the primary key can be the user's identity information, mobile phone number, and other identification information that can uniquely distinguish the user. Different participants hold different types of features. For example, in the scenario of classifying the financial credit risk of users, one participant can hold the transaction data between the user and other users, and another participant can hold the user's friend information, such as which user the user has a friend relationship with.

[0031] The method for determining the intersection samples during the tree model training process is generally determined through private set intersection (PSI). PSI can ensure that the primary key information and feature value information of non-intersection samples are not leaked. After PSI is executed, the two participants can determine the primary key information of the intersection samples, and the two participants can further jointly complete the tree model training based on the intersection samples.

[0032] As data privacy continues to strengthen, data holders have more stringent data demands, hoping to protect their intersection primary key information from being leaked during the modeling process.

[0033] Based on this, this specification provides a method for training a tree model with an anonymous primary key. Next, the implementation process of this method will be described through a specific embodiment.

[0034] This method is applicable to the decision tree model training method in which the calculation process data is encrypted and the primary key information cannot be inferred after the primary key information has been subjected to the circuit-PSI and the anonymous primary key privacy bucketing. Both parties hold the secret shared ciphertext output by the full anonymous bucketing protocol: feature value sharding, feature bucket number sharding, and bucket boundary marker value sharding.

[0035] Now assume that there are two participants P holding different feature data sets A and P B , where P A Own A Features, P B Own m B The two parties use their own data sets to train the decision tree model. The two parties use the existing technologies of anonymous key privacy intersection and anonymous key privacy bucketing protocol to obtain the secret share matrix of data alignment. <d>= <B1||U1||X1||...||B m ||U m ||X m ||Y>, the matrix size is n×(3m+1), where n is the number of data intersections, and m=m A +m B The total number of features. The first 3m columns indicate that each feature contains bucket number fragments. i >, bucket boundary mark value sharding i >(the plaintext corresponding to the minimum value in each bucket is 1, and the plaintext of the remaining values ​​is 0), feature value sharding <X i >These three columns, the last column is the label value fragment <y>,|| represents the connector of each column of data in the matrix. A Features belong to P A , after m B Features belong to P B . Each feature will be divided into nb buckets.

[0036] Note: In this manual, angle brackets are used <x>Indicates that the current variable X is encrypted by secret sharing, and P A and P B Hold secret sharing shards separately <x> A and <x> B In this manual, brackets [X] are used to indicate A Indicates that the current variable X is homomorphically encrypted, and the A in the upper right corner indicates that the private key of the ciphertext belongs to P A , only P A The ciphertext can be decrypted.

[0037] 1. Extracting data by bucket

[0038] The purpose of this step is to determine the corresponding bucket thresholds when the primary key is anonymized. Specifically, A needs to obtain the ciphertext of the bucket thresholds of each bucket of B. Correspondingly, B needs to obtain the ciphertext of the bucket thresholds of each bucket of A. The specific implementation is detailed below.

[0039] (1)P A from <d> A Extract the fragments and send it to P B .P B Use the received fragments and the corresponding fragments you hold locally Perform secret sharing decryption to make P B Get the top m A The bucket number of the intersection data features and bucket boundary marker values

[0040] (2)P A from <d> A Extract the fragments Obtained using homomorphic encryption and send it to P B .P B Use the received fragments and the corresponding fragments you hold locally Perform accumulation operation to obtain the ciphertext under homomorphic encryption only

[0041] (3)P B according to and You can get the first m A features, the bucket boundary value position corresponding to each bucket number, and then extract The ciphertext value at the corresponding position in is used as the secret share of each bucket threshold Each array S i The size of is nb.

[0042] (4) Similar to (1-3), P B send and Give P A , P A Get the secret share of the bucket threshold

[0043] (5)P B Initialize the prediction array And encrypt it through secret sharing into a secret shard held by both parties

[0044] 2. Decision tree model training

[0045] (1) For the kth decision tree, both parties perform the following operations:

[0046] (1.1) Both parties use And the gradient derivative formula g = σ(p k )-y,h=(1-σ(p k ))·σ(p k ), in the secret sharing scenario, the derivative value is calculated { <g i >, <h i >} i∈[1,n] .

[0047] (1.2)P A Initialize the sample distribution of the first node (root node) locally is a vector of size n containing all 1s. Similarly, P B Initialize all 1 vectors of size n locally

[0048] (1.3) For the lth non-leaf node in the tree model, both parties perform the following operations:

[0049] (1.3.1)P B In local <g i > B , <h i > B } i∈[1,n] Perform homomorphic encryption and obtain {[ <g i > B ] B ,[ <h i > B ] B } i∈[1,n] and send it to P A .P A The local { <g i > A , <h i > A } i∈[1,n] and{[ <g i > B ] B ,[ <h i > B ] B } i∈[1,n] Add together to obtain the ciphertext array of homomorphic encryption only {[g i ] B ,[h i ] B } i∈[1,n]

[0050] (1.3.2)P A Initialize all 0 gradient bucket accumulation matrix Then, P A According to the bucket information and sample distribution information Will satisfy Corresponding [ <g j > B ] B ,[ <h j > B ] B Add to Among them, B i,j The bucket number corresponding to the i-th feature of the j-th sample is bucket.

[0051] (1.3.3)P A Using the method of converting homomorphic ciphertext to secret sharing ciphertext, Become a secret sharing shard held by both parties

[0052] (1.3.4) Similar to (1.3.1-1.3.3), P A In local <g i > B , <h i > B } i∈[1,n] Perform homomorphic encryption, P B Through the gradient bucket accumulation and homomorphic ciphertext to secret sharing ciphertext, both parties can obtain

[0053] (1.3.5) Merge the two gradient bucket accumulation matrices to obtain { <G i,q >, <H i,q >} i∈[1,m],q∈[1,nb] .

[0054] (1.3.6) Both parties use the decision tree calculation rules and the tree model gain value calculation formula under the MPC scenario Get the corresponding { <Gain i,q >} i∈[1,m],q∈[1,nb] ,in

[0055] (1.3.7) Both parties use the privacy extreme value solution method to obtain { <Gain i,q >} i∈[1,m],q∈[1,nb] The secret shard of the bucket number corresponding to the maximum value<max_idx> , and through privacy comparison, both parties can obtain<max_idx> and m A ×nb, if<max_idx> Greater than m A ×nb, it means that the sample distribution information corresponding to the feature and the bucket threshold are determined by P A Management requires<max_idx> Exclusively for P A Otherwise,<max_idx> Exclusively for P B .

[0056] (1.3.8) Assume P B Get the value max_idx, then P B We can know that the splitting boundary of the lth node is the jth bucket of the ith feature through max_idx=i×nb+j. B The distribution of n samples with bucket number j as the boundary Update and determine the sample distribution information of the left and right child nodes 2l and 2l+1 of the lth node, that is, Then extract [S i,j ] A , and convert it into a secret sharing shard value <T k,l >, as the splitting threshold of the current node.

[0057] (1.4) For the lth leaf node in the tree model, both parties perform the following operations:

[0058] (1.4.1) Assume that the split information of the parent node of the lth leaf node is given by P B Hold, then P B The formula for calculating leaf node weight can be obtained directly through max_idx in (1.3.8) In the required <g>and <h>exist{ <Gl i,q >, <Hl i,q >} i∈[1,m],q∈[1,nb] The corresponding position d in .

[0059] (1.4.2) However, since P A The feature information corresponding to the parent node and the location information of the maximum gain value are unknown, so both parties need to obtain it through an nb choose 1 oblivious transfer protocol OT. <g>and <h>Specifically, first P A Generate two random numbers rg and rh and calculate locally { <Gl i,q -rg> A , <Hl i,q -rh> A } i∈[1,m],q∈[1,nb] , and then use the array as the input of the OT protocol; at the same time, P B Input d in the OT protocol and get the output of the OT protocol <Gl d -rg> A , <Gl d -rh> A Then, P B Compute locally <g> B ←<Gl d -rg> A +<Gl d > B , <h> B ← <Hl d -rg> A + <Hl d > B ;P A Directly use rg and rh as secret shards <g> A ←rg, <h> A ←rh.

[0060] (1.4.3) Both parties use the formula Get the weight value of the leaf node of the tree model.

[0061] (1.5) For each sample i, both parties pass Update the prediction results, where f k (i) is the leaf node number corresponding to the i-th sample, which can be obtained by bit-multiplying each leaf node and Get f k (i) is a one-hot encoded vector, and then the vector is compared with <w>Dot product As the weight value of the leaf node.

[0062] (2) Finally, both parties output the feature information of the splitting point of each tree and the splitting threshold fragmentation as the model output result of this tree model training.

[0063] In the tree model obtained by the above training process, the same as the tree model jointly trained by two parties in the non-anonymous primary key scenario in the related art is that the predicted value of the leaf node is held by the two parties in the form of shards. However, it is different from the tree model jointly trained by two parties in the related art in that the feature number of the split feature of the non-leaf node is held by the non-feature party, and the split threshold of the split feature is jointly held by the two participating parties in the form of secret sharing shards.

[0064] For any non-leaf node, the feature side refers to the side with the feature value of the split feature, and the non-feature side refers to the side without the feature value of the split feature.

[0065] Among them, secret sharing is also called secret splitting and secret sharing. The basic principle is to split the secret into multiple shares (shares), which are also called shards, and give them to different people for safekeeping. Only when more than a threshold number of people merge their shares can the secret be recovered; no information about the secret can be recovered from shares less than the threshold number. In privacy computing, the threshold number is usually the same as the number of participants. For two participants P0 and P1, arithmetic secret sharing is performed by splitting the plaintext value x of length l bits into two random number shares <x>0∈2 l and <x>1∈2 l Stored by P0 and P1 respectively, and satisfying x=( <x> 0+ <x>1) mod2 l The multiple data obtained through secret sharing can be called secret sharing slices, which are also referred to as slices in this specification. Secret sharing also supports two parties to jointly perform four arithmetic operations of addition, subtraction, multiplication and division based on the secret sharing slices.

[0066] A non-feature party is a party that does not hold a feature of the split feature type. For example, if the split feature is the user's account registration time, which is held by the first party, and the second party does not hold the user's account registration time, the second party is a non-feature party. The feature number of the split feature is the number of the split feature among all features owned by the first party and the second party.

[0067] In the tree model of the related art, the feature number of the non-leaf node is generally held by the feature party, and the corresponding feature party also holds the split threshold of the non-leaf node. The difference between the tree model trained in this specification and the tree model in the related art makes the tree model reasoning method in the related art unable to be applied to the scenario targeted by this specification.

[0068] Based on this, the present specification provides a two-party joint reasoning method for a tree model. The method involves a first party and a second party, the first party holds a feature value of a first feature of a target object, and the second party holds a feature value of a second feature of the target object. The first party and the second party also jointly hold a tree model, in which the feature number of the split feature of a non-leaf node is held by the non-feature party, and the split threshold is held by the two parties in the form of a secret sharing shard. During the reasoning process, the homomorphic ciphertext of each feature value of the other party is obtained, and for each non-leaf node, the second party is jointly obtained according to the homomorphic ciphertext of each feature value, through a homomorphic encryption to secret sharing protocol, to obtain the shard of the feature value of the split feature of the target object. And through the secret comparison jointly performed by the two parties, according to the shard of the split threshold of the non-leaf node, the shard of the secret comparison result indicating which child node the target object falls into is determined. The two parties then jointly determine the shard of the leaf node array indicating the leaf node where the target object falls according to the shard of the secret comparison result of each non-leaf node. Finally, the two parties jointly determine the prediction result of the target object according to the shard of the leaf node array and the prediction value of each leaf node.

[0069] The above method aims at the scenario where the feature number of the split feature of the non-leaf node is held by the non-feature party and the split threshold is held by both parties. The feature value is encrypting by homomorphic encryption, ensuring that the plaintext of the feature value is not leaked. The homomorphic encryption is converted to a secret sharing protocol to obtain the shards of the feature value of the split feature, thereby completing the reasoning of the non-leaf node by using secret comparison, and based on this, the predicted value acquisition for the target object is completed.

[0070] Next, the tree model two-party joint reasoning method shown in this specification will be explained through an embodiment.

[0071] The method in this specification is applied to the scenario of joint reasoning between two parties. For the convenience of explanation, the two parties are respectively referred to as the first party and the second party. The first party and the second party respectively hold feature values ​​of different features of the target object. For the convenience of distinction here, the feature held by the first party is referred to as the first feature, and the feature held by the second party is referred to as the second feature. The first feature may include multiple features, and the second feature may also include multiple features. In other words, the first party holds the feature value of the first feature of the target object, and the second party holds the feature value of the second feature of the target object. The target object here is an object held by both parties. For example, if the target object is a user, the first party may hold feature values ​​of features such as the user's account registration time and the number of interactions with high-risk accounts, and the second party may hold feature values ​​of features such as the number of transactions of the user and whether it has traded with high-risk users.

[0072] The two parties also jointly hold a tree model. The tree model is the same as the tree model jointly trained by the other two parties in that the predicted values ​​of the leaf nodes of the tree model are held by the two parties in the form of secret sharing shards. Unlike the tree model jointly trained by the other two parties in the non-anonymous primary key scenario in the related art, for any non-leaf node of the tree model, the feature number of the split feature of the non-leaf node of the tree model is held by the party that does not hold the feature value of the split feature, and the feature threshold of the split feature is held by the first party and the second party in the form of shards.

[0073] For example, if the split feature of a non-leaf node is the user's account registration time, the user's account registration time is numbered 3 among all the features of the target object held by the first party and the second party. The feature value of the target object's account registration time is held by the first party, and the feature number 3 of the non-leaf node split feature is held by the second party. At the same time, both parties hold the secret shared shard of the split threshold of the split feature.

[0074] In addition, the two parties can jointly hold multiple tree models. For the sake of convenience, only the reasoning process of one tree model is described here. It is easy to understand that the reasoning process thresholds of other tree models are the same. After obtaining the prediction results of multiple tree models, the prediction results of multiple tree models can be added together to obtain the final prediction result.

[0075] In addition, the two participants can also use the tree model to infer the prediction results of multiple target objects at the same time. Here, the reasoning process of one target object is taken as an example. The reasoning process of other target objects is the same and will not be repeated.

[0076] In an optional embodiment, the tree model is obtained by training in conjunction with the second party based on a target training sample set, and the target training sample set is a training sample set obtained by executing a privacy intersection protocol of an anonymous primary key on the data of the object held in conjunction with the second party. In other words, the above-mentioned tree model can be obtained by a joint modeling method of a tree model with an anonymous primary key, for example, it can be obtained by training using the method mentioned above. This scenario does not limit this specification, and it is understood that if a tree model with the above-mentioned characteristics can be obtained by other modeling methods, the method provided in this specification can also be used to perform joint reasoning between the two parties.

[0077] Next, we will combine Figure 2 Here, the steps of the method shown in this specification are described. Here, the steps performed by the first party are described. The steps performed by the second party are the same as the steps performed by the first party, and are not repeated here.

[0078] Step 200: For each second feature, obtain a homomorphic ciphertext of the feature value of each second feature of the target object.

[0079] Correspondingly, the second party also obtains the homomorphic ciphertext of the characteristic value of each first characteristic of the target object for each first characteristic.

[0080] This step is specifically for encrypting the feature value. Since for any non-leaf node, the feature value and feature number of the split feature are held by different parties, in order to determine the feature value of the split feature without leaking the feature number and feature value, the feature party encrypts the feature value and sends it to the non-feature party, so that the non-feature party can determine the ciphertext of the feature value of the split feature.

[0081] The specific implementation of step 200 may be: receiving each feature value of the target object held by the second party after homomorphic encryption and sent by the second party. Correspondingly, in order to enable the second party to obtain the feature value of each first feature, the method further includes: homomorphically encrypting each feature value of the target object held by the party and sending it to the second party.

[0082] For the keys used for encryption, the first-party encryption can use the first-party's public key, and correspondingly, the ciphertext can only be decrypted by the first-party's private key. The second-party encryption can use the second-party's public key, and correspondingly, the ciphertext can only be decrypted by the second-party's private key. This ensures data security.

[0083] When performing homomorphic encryption, homomorphic encryption is performed on each eigenvalue separately, rather than treating all eigenvalues ​​as a whole for homomorphic encryption.

[0084] Here we will combine Figure 3 Step 200 is described with reference to an embodiment of the present invention.

[0085] Figure 3 First, the two participants P A and P B Both parties hold a set of prediction data D with aligned primary key identifiers and different features. A and D B It can be understood that each sample in the prediction data set can be regarded as the target object mentioned above.

[0086] Where D A Contains n samples and m A Features, D B Contains n samples and m B Features. A [idx][q] represents the data set D A The qth eigenvalue of the idxth sample in P. A Held m A Features and P B Held m B The characteristics are different.

[0087] In addition, both parties hold a privacy-protected decision tree model Tree A and Tree B Specifically, the decision tree has a total of t trees, and the depth of each tree is d, that is, the number of non-leaf nodes in each tree is p = 2 d -1, the number of leaf nodes is l = 2 d .

[0088] For the i-th non-leaf node (i∈[p]), it contains the ciphertext fragment of the feature number f[i] and the feature threshold<h[i]> , where the feature number is held by the non-feature party and the threshold is a secret share encrypted with a secret share<h[i]> A ,<h[i]> B For example, the feature number corresponding to the feature belongs to P A , then only P B Knowing the feature number f[i], in order to distinguish the owner of the number, let f B [i]=f[i],f A [i] = 0; otherwise, if it is not a feature, it is P A , then let f A [i]=f[i],f B [i]=0.

[0089] For the jth leaf node (j∈[l]), the corresponding leaf node weight information<w[j]> , also with a secret share<w[j]> A ,<w[j]> B The data is stored locally on both parties.

[0090] As before, this specification will use angle brackets when describing a basic secret sharing operation. <x>Indicates that the current variable X is encrypted by secret sharing, and P A and P B Hold secret sharing shards separately <x> A and <x> B , both parties need to <x> A and <x> B Perform the corresponding secret sharing operation. This article uses square brackets [X] A Indicates that the current variable X is P A Homomorphic public key encryption, the A in the upper right corner indicates that the private key of the ciphertext belongs to P A , only P A The ciphertext can be decrypted.

[0091] Figure 3 The content in the blue dotted box is the reasoning process for the k-th tree model, and idx is the identifier of the target object. For the implementation of this step, first, both parties generate their own homomorphic encrypted public and private keys locally and send the public key to each other. Specifically, P A Generate homomorphic encryption public key pk A and private key sk A , and the public key pk A Send to P B Similarly, P B Generate homomorphic encryption public key pk B and private key sk B , and the public key pk B Send to P A .

[0092] Secondly, for the idxth test data (i.e., the target object), both parties use the homomorphic encryption public key to encrypt their own feature arrays, and then send the ciphertext to the other party. Specifically, P A Using public key pk A Encrypted test sample array D A [idx], get the homomorphic ciphertext And send it to P B Similarly, P B Using public key pk B Encrypted test sample array D B [idx], get the homomorphic ciphertext And send it to P A .

[0093] Step 201, for any non-leaf node, jointly with the second party, execute the homomorphic encryption to secret sharing protocol based on the homomorphic ciphertext of each second eigenvalue, and obtain the shards of the eigenvalues ​​of the split features of the target object; and jointly with the second party, perform secret comparison based on the shards of the eigenvalues ​​of the split features and the shards of the feature thresholds to obtain the shards of the secret comparison results.

[0094] The secret comparison result is used to indicate the node into which the target object falls after the non-leaf node.

[0095] Specifically, based on the homomorphic ciphertext of the eigenvalue of the split feature in each second eigenvalue, the homomorphic ciphertext of the eigenvalue corresponding to the feature number of the split feature is determined. For example, if the feature number is 2, then the homomorphic ciphertext of the eigenvalue with feature number 2 is taken out from each second eigenvalue. Then, according to the homomorphic encryption to secret sharing protocol, the eigenvalue sharding of the split feature is obtained. Then, the two parties jointly perform a secret comparison based on the sharding of the eigenvalue of the split feature of the target object and the sharding of the split threshold to determine the inference result of the non-leaf node.

[0096] For ease of understanding, let's first explain homomorphic encryption. Homomorphic encryption is an encryption method commonly used in privacy computing. Compared with ordinary encryption, homomorphic encryption has the property of homomorphism, that is, it can perform calculations on data in an encrypted state. For example, by performing a homomorphic addition operation on the homomorphic encrypted ciphertext [a] of a and the unencrypted data b, the ciphertext [a+b] of a+b can be obtained, and a, b or a+b will not be leaked during the operation, and the decryption key does not need to be used. Moreover, homomorphic encryption also supports [a]*b=[a*b], and the fully homomorphic encryption algorithm also supports [a]*[b]=[a*b].

[0097] It should also be noted that the reason why the above needs to be converted into the form of secret sharing shards and then secret comparison is performed to complete the reasoning of non-leaf nodes, rather than comparison in the form of homomorphic ciphertext, is because the secret comparison scheme of secret sharing in the relevant technology is more efficient than the comparison scheme of homomorphic ciphertext. Converting it into the form of secret sharing shards can improve the overall reasoning efficiency.

[0098] In different executions, the feature party and the non-feature party may be different. For example, for the first non-leaf node, the first party may be the feature party, and for the second non-leaf node, the second party may be the feature party.

[0099] Next, step 201 will be described in detail. Step 201 includes two steps: obtaining feature value fragments of the split feature and secret comparison.

[0100] Next, we will take a non-leaf node as an example, and take the first party as the feature party and the first party as the non-feature party as examples to illustrate the process of how the first party and the second party jointly obtain the feature value fragmentation of the split feature of the target object.

[0101] First, the first party is the feature party. Step 201 specifically includes: for any non-leaf node, when the feature number of the split feature of the non-leaf node is held by the second party, receiving a third value sent by the second party, the third value being a homomorphic ciphertext of the feature value of the split feature of the target object after random number obfuscation; performing homomorphic decryption processing on the third value to obtain a shard of the feature value of the split feature of the target object held by the party; the second party uses the random number used for obfuscation processing as the shard of the feature value of the split feature of the target object held by the second party.

[0102] Specifically, the two parties jointly execute the homomorphic encryption to secret sharing protocol to obtain the eigenvalue shards of the split feature. During the calculation process, the second party first confuses the homomorphic ciphertext of the eigenvalue of the split feature. Then the obfuscated homomorphic ciphertext is sent to the first party. The first party decrypts the obfuscated data to obtain the eigenvalue shards of the split feature. The second party determines the eigenvalue shards of the split feature based on the random number used for the obfuscation process.

[0103] Second, the second party is the feature party. Step 201 specifically includes: for any non-leaf node, when the feature number of the split feature of the non-leaf node is held by the first party, obtain a first value from the homomorphic ciphertext of each received second feature value, the first value being the homomorphic ciphertext of the feature value of the split feature of the target object; calculate the homomorphic sum of the first value and a preset random number to obtain a second value, and send the second value to the second party; use the opposite of the preset random number as a shard of the feature value of the split feature of the target object held by this party; the second party uses the homomorphic decryption result of the second value as a shard of the feature value of the split feature of the target object held by the second party.

[0104] Specifically, the first party first obtains the homomorphic ciphertext of the eigenvalue corresponding to the split feature from the homomorphic ciphertext of the received second eigenvalue, which is the first value mentioned above. Then the homomorphic addition operation of the first value and the preset random number is calculated to obtain the second value, which is the eigenvalue of the split feature after the obfuscation process. The second value is sent to the second party, and the second party decrypts the second value to obtain the shard of the eigenvalue of the split feature. The first party uses the opposite number of the above random number as the shard of the eigenvalue of the split feature.

[0105] After obtaining the eigenvalue shards of the target object's split feature, the two parties can jointly perform a secret comparison. Based on the relationship between the eigenvalue and the split threshold of the split feature, the two parties can determine whether the target object falls on the left child node or the right node of the non-leaf node, and then obtain the shards of the secret comparison results.

[0106] In an optional implementation, the secret comparison result can be implemented in a binary manner. For example, if the secret comparison result is a first value (such as 0), the target object falls on the left child node of the non-leaf node, and if the secret comparison result is a second value (such as 1), the target object falls on the right node of the non-leaf node. It should be noted that the first value here and the first value of the path value later may have different meanings.

[0107] As for the specific implementation method of secret comparison, it can be implemented based on secret comparison protocols such as obfuscation circuits and bit decomposition schemes.

[0108] Next, we will combine Figure 3 The flowchart shown is used to explain step 201 in detail.

[0109] Specifically, in step 201, the following is performed:

[0110] 1. For the i-th non-leaf node (i∈[p]), both parties need to protect the eigenvalue array D A [idx],D B Under the condition that [idx] and the node feature number information f[i] are not leaked, the ciphertext feature value corresponding to the feature number information f[i] is obtained through the privacy ciphertext extraction method based on homomorphic encryption and secret sharing ciphertext conversion. <D A(B) [idx][f[i]]>.

[0111] Specifically, it is assumed that the node feature number information f[i] is obtained by the non-feature method P B Hold, that is, f B [i] = f[i]. First, P B Directly select the ciphertext array The f B [i] The ciphertext of the eigenvalues ​​of the features Then, P B Generate a random number r locally and use the public key pk A Encrypted as Then, perform homomorphic addition to obtain

[0112] It should be noted that according to the properties of homomorphic encryption, the homomorphic addition here can also be You can also get

[0113] Then, P B Will Send to P A , P A Using the private key sk A Decrypt the pseudo-random number D A [idx][f B [i]]+r.

[0114] Finally, P B Directly use the random number -r as the secret sharing ciphertext of the eigenvalue corresponding to the i-th non-leaf node, denoted as<v[i]> B =-r, similarly, P A D A [idx][f B [i]]+r is used as the secret shared ciphertext of the eigenvalue of the i-th non-leaf node, denoted as<v[i]> A =D A [idx][f B [i]]+r. After p times of the above calculation, both parties can finally obtain the characteristic ciphertext array of all non-leaf nodes of the kth tree <v>.

[0115] 2. For the i-th non-leaf node (i∈[p]), both parties need to protect the ciphertext feature value<v[i]> and ciphertext threshold<h[i]> Without being leaked, the secret decision of the dxth data in the non-leaf node is obtained by secret comparison, and the ciphertext comparison result secretly shared in bit form is obtained.<b[i]> bit Specifically, for the i-th secret comparison, P A enter<v[i]> A and<h[i]> A , P B enter<v[i]> B and<h[i]> B After both parties have executed the agreement, P A get P B get And meet<v[i]> ≤<h[i]> hour, This means that the correct decision path is to move forward to the left child node. On the contrary, if<v[i]> ><h[i]> ,but This means that the correct decision path is to move forward to the right child node. After p times of the above calculations, both parties can finally obtain the ciphertext array of the decision results of all non-leaf nodes in the kth tree. bit .

[0116] For the i-th non-leaf node (i∈[p]), both parties need to protect the ciphertext feature value.<v[i]> and ciphertext threshold<h[i]> Without being leaked, the secret decision of the target object in the non-leaf node is obtained by secret comparison, and the ciphertext comparison result secretly shared in bit form is obtained.<b[i]> bit express.

[0117] Specifically, for the i-th secret comparison, P A enter<v[i]> A and<h[i]> A , P B enter<v[i]> B and<h[i]> B After both parties have executed the agreement, P A get P B get And meet<v[i]> ≤<h[i]> (that is, it falls into the left child node), This means that the correct decision path is to move forward to the left child node.<v[i]> ><h[i]> ,but This means that the correct decision path is to move forward to the right child node. After p times of the above calculations, both parties can finally obtain the ciphertext array of the decision results of all non-leaf nodes in the kth tree. bit Here and This is the shard of the secret comparison result.

[0118] in, Represents XOR operation. The calculation logic of XOR operation is that when two binary bits are the same, the result is 0, and when two binary bits are different, the result is 1. The bit in the upper right corner of the above angle bracket represents that it is a bit-type secret sharing shard. The difference between bit-type secret sharing shard and numeric secret sharing shard is that bit-type secret sharing shard ensures that the XOR operation result of two shards is the corresponding plaintext.

[0119] Step 203, jointly with the second party, based on the sharding of the secret comparison results of each non-leaf node, and jointly with the second party, determine the sharding of the leaf node array corresponding to the target object.

[0120] The leaf node array is used to indicate the leaf node into which the target object falls.

[0121] Specifically, after obtaining the shards of the secret comparison results of each non-leaf node, it is possible to jointly determine with the second party the leaf node array of which leaf node the target object falls according to the connection relationship between the nodes.

[0122] The number of elements in the leaf node array can be the same as the number of leaf nodes. In the leaf node array, the value of the element corresponding to the leaf node where the target object falls is the first value, and the value of the element corresponding to other leaf nodes is the second value. For example, if the target object falls into the third leaf node, the first value is 1, and the second value is 0, then the third element in the leaf node array is 1, and the remaining elements are 0.

[0123] It should be noted that the leaf node array here is in ciphertext form, which can prevent one party from obtaining the specific prediction path, thereby preventing one party from inferring the splitting threshold of the tree model based on the prediction results of multiple target objects, thereby better protecting data security.

[0124] Next, a possible implementation of step 203 will be described. Step 203 may include: for each non-leaf node, the product of the path value calculated jointly with the second party and the secret comparison result is obtained to obtain the slice of the path value of the child node of the non-leaf node; when the path value is a first value, it indicates that the target object falls into the node corresponding to the path value, and when it is a second value, it indicates that the target object does not fall into the node corresponding to the path value; the slices of the path value of each leaf node are organized into an array according to the order of the leaf nodes to obtain the slices of the leaf node array; the leaf node array is a binary array, and in the leaf node array, the value of the element corresponding to the leaf node where the target object falls is the first value, and the value of the element corresponding to other leaf nodes is the second value. In an optional implementation, the first value is 1 and the second value is 0.

[0125] The plain text of the path value is used to indicate whether the target object falls into the current node. For each non-leaf node, the path value of the child node of the non-leaf node can be determined based on the secret sharing multiplication between the shard of the path value and the shard of the secret comparison result. In this way, the shard of the path value of the leaf node can be obtained through iterative calculation. Then, the shard of the path value of each leaf node can be combined into the shard of the leaf node array.

[0126] Next, we will combine Figure 3 The steps shown and Figure 4 Step 203 will be further described.

[0127] Specifically, both parties need to protect the ciphertext array secret bit The prediction results of each leaf node are obtained without being leaked. bit Privacy computing.

[0128] Specifically, both parties are initialized to size 2 d -1 secret path array <c> bit , and follow Figure 4 In the structure of the tree, c[i] indicates whether the i-th node is on the real predicted path (that is, the path value corresponding to the i-th node mentioned above). Since the root node is a must-go point, it satisfies<c[1]> bit = <1> , that is, holding<c[1]> bit The two parties traverse each non-leaf node in the order of the tree model. For the i-th node, the two parties use bit-based secret multiplication to calculate the secret path of the left and right leaf nodes. Since the secret comparison result is represented by 0 and 1 respectively to fall into the left child node and the right child node,

[0129] It should be noted that although the above formula describes the calculation process in the form of slices, it can be understood that the above process is a joint calculation of two parties. The plaintext of c[i] and b[i] is not disclosed during the calculation process, and the result obtained is a slice of the corresponding plaintext calculation result.

[0130] in, Represents AND calculation. AND calculation can make the result 1 when the two binary bits involved in the calculation are 1, and the result is 0 in other cases. For the calculation formula of the left child node,<c[i]> bit The shard representing the path value of the non-leaf node,<b[i]> bit The shard representing the secret comparison result.

[0131] Since b[i] is 0, it means it falls into the left child node, so we need to calculate the XOR result of b[i] and 1 to make the calculation formula consistent with<c[i]> bit The part that performs the AND operation When it falls into the left leaf node, the value of the shard is 1, and when it falls into the right leaf node, the value of the shard is 0. This is the only way to ensure the accuracy of the final operation result. For example, when b[i] is 0, the secret comparison result is that it falls into the left child node, and the XOR operation result of 0 and 1 is 1, which makes The bit-type secret sharing fragment is 1. Finally, this value is combined with<c[i]> bit Perform the AND operation of the joint calculation of the two parties to ensure that when the secret comparison result falls into the left child node and the path value c[i] is 1 (indicating that the target object falls into the current non-leaf node), the path value of the left child node is sharded<c[2i]> bit The bit-type secret sharing shard with a value of 1 ensures the accuracy of the result. The implementation method when b[i] and c[i] take other values ​​is similar and will not be repeated here.

[0132] Correspondingly, since b[i] takes a value of 1 to indicate that the target object falls into the right child node, and takes a value of 0 to indicate that the target object falls into the left child node, it is possible to determine whether the target object falls into the right leaf node directly based on the AND operation result of the path value c[i] and b[i] indicating whether the target object falls into the current non-leaf node. For example, when both c[i] and b[i] take a value of 1, it indicates that the target object falls into the current non-leaf node and the right child node of the non-leaf node. Then the AND operation result of the two is 1, indicating that the target object falls into the right leaf node. If either c[i] or b[i] takes a value of 0, it can be determined that the target object will not fall into the right child node of the non-leaf node. Then the plain text of the AND operation result of c[i] and b[i] is 0, which also indicates that the target object will not fall into the right child node of the non-leaf node.

[0133] Finally, take the array <c> bit No.2 d-1 to 2 d -1 point ciphertext value is directly used as the leaf node array mentioned above, recorded as<e[j]> bit = <c[2 d-1 +j-1]> bit ,j∈[l].

[0134] Step 205, the second party jointly performs a secret sharing-based operation based on the slices of the leaf node array and the slices of the preset prediction values ​​of each leaf node, to obtain the slices of the prediction results corresponding to the target object.

[0135] Specifically, the slices of the preset prediction values ​​of each leaf node can be first formed into slices of the prediction value array according to the order of the leaf nodes. The two parties jointly perform the secret sharing multiplication operation, calculate the result of the bitwise multiplication of the leaf node array and the prediction value array, and obtain the slices of the first array, which represents the result of the bitwise multiplication of the leaf node array and the prediction value array. The two parties then jointly calculate the secret sharing addition operation between the bits of the first array to obtain the slices of the prediction results of the current tree model for the target object.

[0136] Still used here Figure 3 For the kth decision tree, both parties use secret multiplication and addition to calculate the prediction result of this tree. <e[j]> bit is a slice of the leaf node array,<w[j]> A slice of the prediction array.

[0137] In addition, when there are multiple tree models, the prediction results of each tree model can be added together to obtain the final prediction result for the target object. In other words, after the above steps, both parties obtain the ciphertext array consisting of the prediction results of each tree. <res>Then both parties use secret addition to accumulate all prediction results and obtain the final ciphertext prediction result of the idxth test data.

[0138] Finally, the prediction results can be adjusted according to business needs. <pred>Decryption is performed so that the first party or the second party obtains the final prediction result pred.

[0139] In other words, the prediction result of the target object can be obtained by performing secret sharing decryption in conjunction with the second party according to the fragmentation of the prediction result. Here, the prediction result of the target object can be finally obtained by at least one of the first party and the second party.

[0140] The above embodiment introduces the reasoning method in the scenario where the intersection primary key of the predicted data set is public, that is, the primary key of the intersection of the objects owned by both parties (that is, the target object) is in plain text. The above embodiment requires the ability to protect the model information and training data set from being leaked while realizing the reasoning of the specified target object and publicizing the prediction results, that is, the participants cannot infer the model information and training data set information through the additional information obtained in the reasoning stage.

[0141] In the related technology, there is also a demand for anonymous reasoning of tree models with hidden primary keys that cannot associate the prediction results with the primary key information of a single object, that is, the prediction results can only reflect the overall trend, or can only obtain compliance scenario requirements for statistical information.

[0142] Based on this, this specification also provides a two-party joint reasoning method for a tree model with an anonymous primary key. The method is described by taking the execution steps of the first party as an example. It can be understood that the steps executed by the second party are the same as those of the first party and will not be repeated. The method includes:

[0143] Step 501, for multiple objects, jointly execute the privacy intersection protocol of the anonymized primary key with the second party to obtain the first slice of the feature values ​​of all features of the target object; the second party obtains the second slice of the feature values ​​of all features of the target object, and the target object is an object jointly held by the first party and the second party.

[0144] Specifically, both parties need to first execute the privacy intersection protocol of the anonymized primary key for the objects they each hold, so that both parties can obtain the shards of the feature values ​​of the target object whose primary key identifiers are anonymized and whose feature values ​​are all in ciphertext form. In the case of multiple target objects (the two parties do not know which specific target objects they are), the privacy intersection protocol of the anonymized primary key can be used to obtain the full ciphertext matrix <D A > and <D B >. Correspondingly, the data of the two input models are: P A hold <D A > A and <D B > A , P B hold <D A > B and <D B > B That is, both parties hold: A and the second party P B The slicing of the eigenvalues ​​of each feature.

[0145] Step 502, for each second feature, jointly with the second party, based on the shards of the feature values ​​of each second feature held by both parties, execute a secret sharing to homomorphic encryption protocol to obtain the homomorphic ciphertext of each second feature value of the target object.

[0146] This step corresponds to step 200 above. The difference is that in this step, since the eigenvalue fragments of the split feature are held by two parties, the first party obtains the homomorphic ciphertext of the second eigenvalue through the secret sharing to homomorphic encryption protocol. Correspondingly, the two parties can also use the secret sharing to homomorphic encryption protocol to enable the second party to obtain the homomorphic ciphertext of each first eigenvalue.

[0147] A possible secret sharing to homomorphic encryption protocol is described here, and this embodiment does not limit this specification.

[0148] For the second characteristic, the second party can homomorphically encrypt the slices of each second characteristic value and send the homomorphically encrypted ciphertext to the first party. The first party performs homomorphic addition operations on the homomorphic ciphertext of the slices of the second characteristic value and the slices of each second characteristic value held locally to obtain the homomorphic ciphertext of each second characteristic value.

[0149] Here, the second party may use the second party's public key for encryption, so that the first party cannot decrypt the second characteristic value and disclose the private content.

[0150] It can be understood that, for the first feature, the methods executed by the two parties are similar and will not be repeated here.

[0151] Still used here Figure 3 The symbols in the examples are explained in detail.

[0152] For the idx-th test data (i.e., the target object), both parties need to use the secret shared ciphertext conversion homomorphic ciphertext method to encrypt their own characteristic secret shared values ​​with the homomorphic encryption public key respectively, and then send the ciphertext to the other party. The other party completes the decryption operation of the secret shared ciphertext value under the protection of the homomorphic ciphertext.

[0153] Specifically, for the encrypted feature array <D A [idx]>, P A Using public key pk A Encrypted test sample array <D A [idx]> A , get the homomorphic ciphertext And send it to P B , then, P B Using public key pk A Encryption of the party's holdings <D A [idx]> B get And use homomorphic ciphertext addition to calculate Where l is the number of bits of the secret shared ciphertext, 2 l is the scope of secret sharing, c∈{0,1}, which satisfies the calculation conditions of secret sharing plaintext and can be directly regarded as Similarly, P B Using public key pk B Encrypted test sample array <D B [idx]> B , get the homomorphic ciphertext And send it to P A , P A Obtained by calculation

[0154] After executing step 502, you can continue to execute Figure 2 Steps 201-205 shown complete the reasoning process of the tree model.

[0155] like Figure 5 As shown, the specification also provides a two-party joint reasoning device for a tree model, wherein a first party holds a feature value of a first feature of a target object, and a second party holds a feature value of a second feature of the target object; the first party and the second party jointly hold a tree model, and a feature number of a split feature of a non-leaf node of the tree model is held by the party that does not hold the feature value of the split feature, and a feature threshold of the split feature is held by the first party and the second party in the form of slices; the device comprises:

[0156] The ciphertext acquisition module 510 is used to acquire the homomorphic ciphertext of the second feature value of each second feature of the target object for each second feature;

[0157] The secret comparison module 520 is used to execute the homomorphic encryption to secret sharing protocol for any non-leaf node in conjunction with the second party based on the homomorphic ciphertext of each second characteristic value, and obtain the sharding of the characteristic value of the split characteristic of the target object; and to perform secret comparison with the second party based on the sharding of the characteristic value of the split characteristic and the sharding of the characteristic threshold to obtain the sharding of the secret comparison result; the secret comparison result is used to indicate the node into which the target object falls after the non-leaf node;

[0158] The leaf node array determination module 530 is used to determine the slice of the leaf node array corresponding to the target object in conjunction with the second party based on the slice of the secret comparison result of each non-leaf node; the leaf node array is used to indicate the leaf node into which the target object falls;

[0159] The prediction result acquisition module 540 is used for the second party to jointly perform a secret sharing-based operation based on the slices of the leaf node array and the slices of the preset prediction values ​​of each leaf node to obtain the slices of the prediction results of the target object.

[0160] In an optional embodiment, the secret comparison module 520 is specifically used to: for any non-leaf node, when the feature number of the split feature of the non-leaf node is held by the first party, obtain a first value from the received homomorphic ciphertext of each second feature value, the first value being the homomorphic ciphertext of the feature value of the split feature of the target object; calculate the homomorphic sum of the first value and a preset random number to obtain a second value, and send the second value to the second party; use the opposite of the preset random number as a shard of the feature value of the split feature of the target object held by this party; the second party uses the homomorphic decryption result of the second value as the shard of the feature value of the split feature of the target object held by the second party.

[0161] In an optional embodiment, the secret comparison module 520 is specifically used to: for any non-leaf node, when the feature number of the split feature of the non-leaf node is held by the second party, receive a third value sent by the second party, the third value being a homomorphic ciphertext of the feature value of the split feature of the target object after random number obfuscation processing; perform homomorphic decryption processing on the third value to obtain a shard of the feature value of the split feature of the target object held by this party; the second party uses the random number used for obfuscation processing as the shard of the feature value of the split feature of the target object held by the second party.

[0162] In an optional embodiment, the leaf node array determination module 530 is specifically used to obtain, for each non-leaf node, the slices of the path values ​​of the left node and the right node of the non-leaf node based on the product of the path value jointly calculated by the two parties and the secret comparison result; when the path value is a first value, it indicates that the target object falls into the node corresponding to the path value; when the path value is a second value, it indicates that the target object does not fall into the node corresponding to the path value; the slices of the path value of each leaf node are organized into an array according to the order of the leaf nodes to obtain the slices of the leaf node array; the leaf node array is a binary array, in which the value of the element corresponding to the leaf node where the target object falls is the first value, and the value of the element corresponding to other leaf nodes is the second value.

[0163] In an optional embodiment, it also includes a privacy intersection module 550 (not shown in the figure), which is used to jointly execute the privacy intersection protocol of the anonymous primary key with the second party for multiple objects to obtain the first slice of the feature values ​​of all features of the target object; the second party obtains the second slice of the feature values ​​of all features of the target object, and the target object is an object jointly held by the first party and the second party.

[0164] In an optional implementation, the ciphertext acquisition module 510 is specifically used to execute a secret sharing to homomorphic encryption protocol for each second feature in conjunction with the second party based on the shards of the feature values ​​of each second feature held by both parties, so as to obtain the homomorphic ciphertext of each second feature value of the target object.

[0165] In an optional implementation, the ciphertext acquisition module 510 is specifically used to receive the characteristic values ​​of the target object held by the second party and sent by the second party after homomorphic encryption; and homomorphically encrypt the characteristic values ​​of the target object held by the party and send them to the second party.

[0166] In an optional implementation, a decryption module 560 (not shown) is further included, which is used to perform secret sharing decryption in conjunction with the second party according to the slices of the prediction results to obtain the prediction results of the target object.

[0167] In an optional implementation, the tree model is trained in conjunction with the second party based on a target training sample set, and the target training sample set is a training sample set obtained by executing a privacy intersection protocol of anonymized primary keys in conjunction with the second party on data of the held object.

[0168] In the 1990s, it was very clear whether the improvement of a technology was a hardware improvement (for example, improvements to the circuit structure of diodes, transistors, switches, etc.) or a software improvement (improvement of the method flow). However, with the development of technology, many improvements in the method flow today can be regarded as direct improvements in the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to ask a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs. The original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages ​​and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.

[0169] The controller can be implemented in any appropriate manner, for example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (such as software or firmware) that can be executed by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in a purely computer-readable program code manner, the controller can be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, this controller can be considered as a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as both software modules for implementing the method and structures within the hardware component.

[0170] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, this specification does not exclude that with the development of computer technology in the future, the computer that implements the functions of the above embodiments may be, for example, a personal computer, a laptop computer, a vehicle-mounted human-computer interaction device, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0171] Although one or more embodiments of the present specification provide method operation steps as described in the embodiments or flow charts, more or less operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way of executing the order of many steps, and does not represent the only execution order. When the device or terminal product in practice is executed, it can be executed in sequence or in parallel according to the method shown in the embodiments or the drawings (for example, a parallel processor or a multi-threaded processing environment, or even a distributed data processing environment). The term "include", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, product or equipment including a series of elements includes not only those elements, but also includes other elements that are not explicitly listed, or also includes elements inherent to such process, method, product or equipment. In the absence of more restrictions, it is not excluded that there are other identical or equivalent elements in the process, method, product or equipment including the elements. For example, if the words first, second, etc. are used to represent the name, they do not represent any specific order.

[0172] For the convenience of description, the above devices are described in various modules according to their functions. Of course, when implementing one or more of the present specification, the functions of each module can be implemented in the same or more software and / or hardware, or the module implementing the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0173] This specification is described with reference to the flowcharts and / or block diagrams of the methods, apparatus (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0174] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0175] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0176] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0177] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0178] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage, graphene storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0179] It should be understood by those skilled in the art that one or more embodiments of the present specification may be provided as a method, system or computer program product. Therefore, one or more embodiments of the present specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, one or more embodiments of the present specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0180] One or more embodiments of the present specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0181] Each embodiment in this specification is described in a progressive manner, and the same and similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. In the description of this specification, the description of the reference term "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representation of the above terms does not necessarily target the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, in the absence of mutual contradiction, a person skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0182] The above description is only an example of one or more embodiments of the present specification and is not intended to limit one or more embodiments of the present specification. For those skilled in the art, one or more embodiments of the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims.< / pred> < / res> < / c> < / c> < / v> < / x> < / x> < / x> < / x> < / x> < / x> < / x> < / x> < / x> < / w> < / h> < / g> < / h> < / g> < / h> < / g> < / h> < / g> < / d> < / d> < / x> < / x> < / x> < / y> ​​< / d>

Claims

1. A two-party joint reasoning method for a tree model, applied to a first party, the method also involves a second party, the first party holds a feature value of a first feature of a target object, and the second party holds a feature value of a second feature of the target object; the first party and the second party jointly hold a tree model, the feature number of the split feature of a non-leaf node of the tree model is held by the party that does not hold the feature value of the split feature, and the feature threshold of the split feature is held by the first party and the second party in the form of slices; the method comprises: For each second feature, obtaining a homomorphic ciphertext of a second feature value of each second feature of the target object; For any non-leaf node, jointly with the second party, according to the homomorphic ciphertext of each second characteristic value, execute the homomorphic encryption to secret sharing protocol to obtain the sharding of the characteristic value of the split characteristic of the target object; and jointly with the second party, perform secret comparison based on the sharding of the characteristic value of the split characteristic and the sharding of the characteristic threshold to obtain the sharding of the secret comparison result; the secret comparison result is used to indicate the node into which the target object falls after the non-leaf node; Determine, in conjunction with the second party, a shard of the leaf node array corresponding to the target object based on the shard of the secret comparison result of each non-leaf node; The leaf node array is used to indicate the leaf node into which the target object falls; The second party jointly performs a secret sharing-based operation based on the slices of the leaf node array and the slices of the preset prediction values ​​of each leaf node to obtain the slices of the prediction results of the target object.

2. The method according to claim 1, wherein for any non-leaf node, jointly with the second party, according to the homomorphic ciphertext of each second eigenvalue, a homomorphic encryption-to-secret sharing protocol is executed to obtain the shards of the eigenvalues ​​of the splitting characteristics of the target object, including: For any non-leaf node, when the feature number of the split feature of the non-leaf node is held by the first party, obtaining a first value from the homomorphic ciphertext of each received second feature value, where the first value is the homomorphic ciphertext of the feature value of the split feature of the target object; Calculate a homomorphic sum of the first value and a preset random number to obtain a second value, and send the second value to the second party; The opposite number of the preset random number is used as a shard of the characteristic value of the split characteristic of the target object held by this party; the second party uses the homomorphic decryption result of the second value as the shard of the characteristic value of the split characteristic of the target object held by the second party.

3. The method according to claim 1, wherein for any non-leaf node, jointly with the second party, according to the homomorphic ciphertext of each second eigenvalue, a homomorphic encryption-to-secret sharing protocol is executed to obtain the shards of the eigenvalues ​​of the splitting characteristics of the target object, including: For any non-leaf node, when the feature number of the split feature of the non-leaf node is held by the second party, receiving a third value sent by the second party, wherein the third value is a homomorphic ciphertext of the feature value of the split feature of the target object after random number obfuscation processing; The third value is homomorphically decrypted to obtain a shard of the characteristic value of the split characteristic of the target object held by the party; the second party uses the random number used for obfuscation processing as the shard of the characteristic value of the split characteristic of the target object held by the second party.

4. According to the method of claim 1, the sharding of the secret comparison results of each non-leaf node and the sharding of the leaf node array corresponding to the target object jointly determined by the second party include: For each non-leaf node, obtain the shards of the path values ​​of the left node and the right node of the non-leaf node according to the product of the path value jointly calculated by the two parties and the secret comparison result; When the path value is the first value, it indicates that the target object falls into the node corresponding to the path value; When the path value is the second value, it indicates that the target object does not fall into the node corresponding to the path value; The slices of the path value of each leaf node are arrayed according to the order of the leaf nodes to obtain the slices of the leaf node array; The leaf node array is a binary array. In the leaf node array, the value of the element corresponding to the leaf node where the target object falls is a first value, and the value of the element corresponding to other leaf nodes is a second value.

5. The method according to claim 1, further comprising: For multiple objects, jointly execute the privacy intersection protocol of the anonymized primary key with the second party to obtain the first slice of the feature values ​​of all features of the target object; The second party obtains a second slice of feature values ​​of all features of a target object, where the target object is an object held jointly by the first party and the second party.

6. The method according to claim 5, wherein for each second feature, obtaining the homomorphic ciphertext of the second feature value of each second feature of the target object comprises: For each second feature, jointly with the second party, based on the shards of the feature values ​​of each second feature held by both parties, a secret sharing to homomorphic encryption protocol is executed to obtain the homomorphic ciphertext of each second feature value of the target object.

7. According to the method of claim 1, the step of obtaining the homomorphic ciphertext of the second feature value of each second feature of the target object for each second feature comprises: Receiving each characteristic value of the target object held by the second party and sent by the second party after homomorphic encryption; The method further comprises: Each characteristic value of the target object held by this party is homomorphically encrypted and sent to the second party.

8. The method according to claim 1, further comprising: According to the sharding of the prediction results, secret sharing decryption is performed jointly with the second party to obtain the prediction results of the target object.

9. According to the method of claim 1, the tree model is obtained by training in conjunction with the second party based on a target training sample set, and the target training sample set is a training sample set obtained by executing a privacy intersection protocol of anonymized primary keys on data of objects held in conjunction with the second party.

10. A computing device comprising a memory and a processor, wherein the memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 9 is implemented.