Method and device for multi-party joint training of tree model for protecting privacy data

Through homomorphic encryption and computing technology, the security and privacy protection of the multi-party data joint training tree model in machine learning is achieved, and the comprehensive data utilization problem under the data island phenomenon is solved, ensuring that the data is not leaked while model training is completed.

CN114547684BActive Publication Date: 2025-07-08ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210140436.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-15
Publication Date
2025-07-08
Estimated Expiration
2042-02-15

AI Technical Summary

Technical Problem

In machine learning, how to safely carry out multi-party data joint training tree model while ensuring data is not leaked, especially when facing data island phenomena, comprehensive multi-party data for joint training.

Method used

By using homomorphic encryption and homomorphic operations, data encryption is used to encrypt the intermediate party, and data interaction and calculation are performed in the encrypted state, ensuring that all parties do not disclose private data during the training process. The specific steps include the first party homomorphic encryption sample gradient and sending it to the second party. After the second party performs homomorphic operations, the intermediate party decrypts and sends it to the first party. The first party splits the node based on the encryption and plaintext results.

Benefits of technology

This realizes that multi-party joint training of risk identification models without leaking private data, enhancing data security and privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114547684B_ABST
    Figure CN114547684B_ABST
Patent Text Reader

Abstract

An embodiment of this specification provides a method and apparatus for protecting a risk identification model for multi-party joint training of privacy data. The method includes: The first party obtains the sample gradients of each user in the first node to be split in the current tree, and uses the target public key of the intermediate party to homomorphically encrypt the sample gradients of each user, and sends the obtained ciphertexts of each gradient to the second party, where the sample gradients are determined at least according to the label values of the corresponding users; The second party performs a homomorphic operation based on the second features of each user and the ciphertexts of each gradient to determine the ciphertext of the intermediate result, and sends it to the intermediate party, where the ciphertext of the intermediate result is related to the second gain for splitting according to the second features; The intermediate party decrypts the ciphertext of the intermediate result into the plaintext of the intermediate result and sends it to the first party; The first party splits each user in the first node based on the first gain and the plaintext of the intermediate result, where the first gain is the gain calculated using the sample gradients and for splitting according to the first features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of privacy security technologies, and particularly to a method and apparatus for multi-party joint training of a tree model for protecting privacy data. Background Art

[0002] The data required for machine learning may involve multiple platforms. For example, in the scenario of merchant classification based on machine learning, the electronic payment platform has transaction flow (bill) data of merchants, and the banking institution has bill settlement data of merchants and labels of merchant types. In this scenario, the electronic payment platform does not have merchant labels and cannot train a model alone, while the merchants owned by the banking institution have fewer features. Even if a model can be trained alone, the accuracy of the model classification result is not high enough. In the case where data exists in the form of islands, due to issues such as industry competition, data security, and user privacy, data integration faces great resistance, and it is difficult to integrate data scattered on various platforms to train a machine learning model. Under the premise of ensuring that data is not leaked, using multi-party data to jointly train a machine learning model has become a major challenge.

[0003] In machine learning models, tree models are widely used in classification scenarios. When facing the phenomenon of data islands, how to comprehensively integrate multi-party data and securely perform multi-party joint training of tree models has become a problem to be solved. Summary of the Invention

[0004] One or more embodiments of this specification provide a method and apparatus for multi-party joint training of a risk identification model for protecting privacy data, so as to securely perform joint training to obtain a risk identification model.

[0005] According to a first aspect, a method for multi-party joint training of a risk identification model for protecting privacy data is provided. The multi-party includes a first party, a second party, and an intermediate party. The first party holds a number of first features of a user and the label value of whether the user is a risk user. The second party holds a number of second features of the user. The risk identification model is a tree model. The method includes:

[0006] The first party obtains the sample gradients of each user in the first node to be split in the current tree, homomorphically encrypts the sample gradients of each user using the target public key of the intermediate party, and sends the obtained ciphertexts of each gradient to the second party, where the sample gradients are determined at least based on the label value of the corresponding user.

[0007] The second party performs a homomorphic operation based on the second features of each user and the ciphertexts of each gradient to determine the ciphertext of the intermediate result, and sends it to the intermediate party, where the ciphertext of the intermediate result is related to the second gain obtained by splitting based on the second features.

[0008] The intermediate party decrypts the ciphertext of the intermediate result into the plaintext of the intermediate result and sends it to the first party;

[0009] Based on the first gain and the plaintext of the intermediate result, the first party splits each user in the first node, where the first gain is a gain calculated using the sample gradient and split according to the first feature.

[0010] In an implementable manner, splitting each user in the first node includes:

[0011] The first party determines a target gain based on the first gain and the plaintext of the intermediate result;

[0012] The first party splits each user in the first node based on the target gain.

[0013] In an implementable manner, the first gain includes multiple first gain values split according to respective first alternative feature values of each first feature; the plaintext of the intermediate result is a second gain, and the second gain includes multiple second gain values split according to respective second alternative feature values of each second feature;

[0014] Determining the target gain includes:

[0015] The first party determines the gain value with the largest numerical value among the multiple first gain values and the multiple second gain values as the target gain.

[0016] In an implementable manner, the first gain includes multiple first gain values split according to respective first alternative feature values of each first feature; the plaintext of the intermediate result includes: a gradient histogram corresponding to each second feature, and the gradient histogram of any second feature includes the sum of the sample gradients corresponding to the respective second alternative feature values of this second feature;

[0017] Determining the target gain includes:

[0018] The first party determines the second gain based on the gradient histograms corresponding to each second feature, where the second gain includes multiple second gain values split according to respective second alternative feature values of each second feature;

[0019] The first party determines the gain value with the largest numerical value among the multiple first gain values and the multiple second gain values as the target gain.

[0020] In an implementable manner, splitting each user in the first node based on the target gain includes:

[0021] If the target gain belongs to the first gain, the first party divides each user into the left child node or the right child node of the first node based on the characteristics and split value corresponding to the target gain, and the characteristic values of this characteristic corresponding to each user, to obtain a first division result;

[0022] The first party sends the first division result to the second party;

[0023] Based on the first division result, the second party divides the users corresponding to the second node corresponding to the first node into the left child node or the right child node of the second node.

[0024] In an implementable manner, it further includes:

[0025] If the target gain belongs to the second gain, the first party sends the ciphertext of the characteristic corresponding to the target gain and the ciphertext of the split value to the second party;

[0026] Based on the corresponding plaintext of the characteristic, the split value plaintext, and the characteristic values of this characteristic plaintext of each user, the second party divides each user into the left child node or the right child node of the second node, to obtain a second division result;

[0027] The first party receives the second division result from the second party, and based on the second division result, divides each user into the left child node or the right child node of the first node.

[0028] In an implementable manner, the determining of the second gain includes:

[0029] For any target second characteristic, the first party uses the ciphertext of any second alternative characteristic value corresponding to the target second characteristic as the target value, and based on the gradient histogram of this target second characteristic, calculates the first sum value of the sum of all sample gradients on the left side of this target value, and the second sum value of the sum of all sample gradients on the right side of this target value; based on the first sum value, the second sum value, and the third sum value of the sample gradients of each user, determine the second gain value for splitting according to the any second alternative characteristic value of this target second characteristic.

[0030] In an implementable manner, the sample gradient includes a first-order gradient and a second-order gradient; both the first sum value and the second sum value include the sum of the corresponding first-order gradient sums and the sum of the second-order gradient sums.

[0031] In an implementable manner, the user belongs to the intersection of the user sets held by the first party and the second party respectively.

[0032] In an implementable manner, it further includes:

[0033] If the current tree reaches a preset generation condition, the first party determines the weight values of each leaf node based on the sample gradients in the current leaf nodes of the current tree.

[0034] In an implementable manner, the first party obtains the sample gradients of each user in the first node to be split in the current tree, including:

[0035] If the first node is the root node of the current tree, determine the sample gradients of each user based on the predicted values and label values of each user; the predicted values are determined according to the generated tree;

[0036] If the first node is not the root node, read the calculated sample gradients of each user falling into the first node.

[0037] In an implementable manner, the first party is used to perform the bill settlement service, and the first feature is related to the bill settlement service; the second party is used to perform the bill generation service, and the second feature is related to the bill generation service.

[0038] In an implementable manner, the first party and the second party are located in different geographical regions.

[0039] According to a second aspect, a method for protecting privacy data in a multi-party joint training risk identification model is provided. The multi-party includes a first party, a second party, and an intermediate party. The first party holds a number of first features of users and the label values of whether they are risk users. The second party holds a number of second features of users. The risk identification model is a tree model. The method is executed by the first party and includes:

[0040] Obtain the sample gradients of each user in the first node to be split in the current tree, and use the target public key of the intermediate party to homomorphically encrypt the sample gradients of each user, and send the obtained ciphertexts of each gradient to the second party, where the sample gradients are determined at least according to the label values of the corresponding users;

[0041] Receive the intermediate result plaintext from the intermediate party. The intermediate result plaintext is the result of the intermediate party decrypting the intermediate result ciphertext sent by the second party. The intermediate result ciphertext is determined by the second party based on the second features of each user and the ciphertexts of each gradient through homomorphic operation, and is related to the second gain obtained by splitting according to the second features;

[0042] Split each user in the first node based on the first gain and the intermediate result plaintext, where the first gain is the gain calculated using the sample gradients and obtained by splitting according to the first features.

[0043] According to a third aspect, there is provided an apparatus for protecting privacy data in a multi-party joint training risk identification model. The multi-party includes a first party, a second party, and an intermediate party. The first party holds a number of first features of users and the label value of whether the user is a risk user. The second party holds a number of second features of users. The risk identification model is a tree model. The apparatus is deployed at the first party and includes:

[0044] An acquisition encryption module, configured to acquire the sample gradients of each user in the first node to be split in the current tree, and homomorphically encrypt the sample gradients of each user by using the target public key of the intermediate party, and send the obtained ciphertexts of each gradient to the second party, where the sample gradients are determined at least according to the label value of the corresponding user;

[0045] A first receiving module, configured to receive the intermediate result plaintext from the intermediate party. The intermediate result plaintext is the result of the intermediate party decrypting the intermediate result ciphertext sent by the second party. The intermediate result ciphertext is determined by the second party through homomorphic operations based on the second features of each user and the ciphertexts of each gradient, and is related to the second gain for splitting according to the second features;

[0046] A splitting module, configured to split each user in the first node based on the first gain and the intermediate result plaintext, where the first gain is the gain calculated by using the sample gradients and for splitting according to the first features.

[0047] According to a fourth aspect, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method described in the second aspect.

[0048] According to a fifth aspect, there is provided a computing device, including a memory and a processor. Among them, an executable code is stored in the memory. When the processor executes the executable code, the method described in the second aspect is implemented.

[0049] According to the method and apparatus provided in the embodiments of the present specification, during the process of multi-party joint training of a risk identification model (tree model), when the first party (holding a number of first features of users and the label value of whether the user is a risk user) and the second party (holding a number of second features of users) that hold privacy data interact with each other for training-related data, first use the target public key provided by the intermediate party for homomorphic encryption or perform homomorphic operations and then send them to the other party, so that the other party cannot obtain the privacy data plaintext of its own party, and protect their respective privacy data during the process of jointly training the model. Description of the Drawings

[0050] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0051] Figure 1A and Figure 1B is a schematic diagram of the implementation framework of an embodiment disclosed in this specification;

[0052] Figure 2 is a schematic flowchart of a method for protecting privacy data in a multi-party joint training risk identification model provided in the embodiment;

[0053] Figure 3 is a schematic structural diagram of the current tree provided in the embodiment;

[0054] Figure 4 is a schematic flowchart of a method for protecting privacy data in a multi-party joint training risk identification model provided in the embodiment;

[0055] Figure 5 is a schematic block diagram of a device for protecting privacy data in a multi-party joint training risk identification model provided in the embodiment. Detailed implementation manners

[0056] The following will describe in detail the technical solutions of the embodiments of this specification in conjunction with the drawings.

[0057] This specification discloses a method and a device for protecting privacy data in a multi-party joint training risk identification model. First, the application scenarios and technical concepts of the method will be introduced as follows:

[0058] Data often exists in the form of isolated islands. Due to issues such as industry competition, data security, and user privacy, there are great obstacles to data integration. Under the premise of ensuring data non-disclosure, using multi-party data for joint training of machine learning models has become a major challenge. How to jointly train a risk identification model (such as a tree model) with multiple parties has become a problem to be solved.

[0059] In view of this, the inventor proposes a method for protecting privacy data in a multi-party joint training risk identification model, Figure 1A and Figure 1B shows a schematic diagram of the implementation framework of an embodiment disclosed in this specification. As Figure 1A and Figure 1BAs shown in the figure, the scenario of a multi-party joint training risk identification model for protecting privacy data involves parties that can include: the first party A, the second party B, and the intermediate party C, where the intermediate party C can be a party trusted by both the first party A and the second party B. Each of the participating parties can be implemented through any device, platform, server, or cluster of devices with computing and processing capabilities. The above-mentioned multiple parties need to jointly train a risk identification model while protecting data privacy, where the risk identification model is a tree model.

[0060] Among them, the first party A and the second party B can pre-set a control application corresponding to the intermediate party C, so that the first party A and the second party B have the ability to jointly train the model, that is, based on the control application they set, the first party A and the second party B can, with the assistance of the intermediate party C, achieve the joint training of the risk identification model (tree model).

[0061] The first party A and the second party B are holders of the data required for model training. The first party A holds a number of first features XA of each first user in the first user set and the label value indicating whether the user is a risk user. For example, the first features XA include x1, x2, x3, and the label value of the i-th user is y. i . The second party B holds a number of second features XB of each second user in the second user set. For example, the second features XB include x4, x5, x6. The intermediate party C generates a pair of public and private key pairs, namely the target public key PK and the target private key SK, for data encryption and decryption during the process of jointly training the risk identification model.

[0062] In an exemplary scenario, the first party A is used to perform bill settlement services, such as a bank site. Correspondingly, the first features are at least related to bill settlement services. For example, they can include but are not limited to: whether the first user has opened multiple accounts with the same name, whether the addresses of its remittance parties are scattered, whether it has been risk mitigated within a preset time, etc. The second party B is used to perform bill generation services, such as an e-commerce site. Correspondingly, the second features are at least related to bill generation services. For example, they can include but are not limited to: the number of inflow / outflow parties of the transactions generated by the second user, the amount of inflow / outflow of the transactions generated by the second user, whether the transactions generated by the second user are early morning transactions, etc. In one case, the first features and / or the second features also include basic attribute features of the user, such as but not limited to: user identification, hobbies, occupation, population labels, and the number of accounts created (including accounts created on the side of the first party A and / or the second party B) by the user, etc.

[0063] In an exemplary scenario, the first party A and the second party B are located in different geographical regions. For example, the first party A and the second party B can be located in different countries, such as in an international card acceptance business scenario; the first party A and the second party B can also be located in different cities, different provinces, etc.

[0064] It should be understood that the user characteristics maintained by each participating party and the label values indicating whether a user is a risky user belong to private data and cannot be exchanged in plaintext during the joint training process to protect the security of private data. And finally, Party A (the first party) and Party B (the second party) hope to train a risk identification model (a tree model) for risk identification, and both Party A and Party B have the risk identification model on their respective local sides.

[0065] To perform joint training of the model without disclosing private data, according to the embodiments of this specification, as Figure 1A shown, during the joint training process of the risk identification model (tree model), Party A (the first party) obtains the sample gradients G of each user (who belongs to both the first users and the second users) in the first node to be split in the current tree (the sample gradient of user i is denoted as G i ), and uses the target public key PK of the intermediate party C to homomorphically encrypt the sample gradients G of each user, and sends the obtained encrypted gradients [G] PK to Party B (the second party), where the sample gradient G is determined at least according to the label value y of the corresponding user (the sample gradient G i is determined by the label value y i of user i). Here, [] represents encryption, and the subscript represents the key used for encryption.

[0066] After Party B (the second party) obtains the encrypted gradients [G] PK , based on the second feature XB of each user and the encrypted gradients [G] PK , it performs homomorphic operations to determine the encrypted intermediate result [Z] PK , and sends the encrypted intermediate result [Z] PK to the intermediate party C, where the encrypted intermediate result [Z] PK is related to the second gain IB for splitting according to the second feature XB. In one implementation, the encrypted intermediate result [Z] PK can be the encrypted second gain values for splitting according to the second alternative feature values of each second feature XB; or it can be the gradient histograms corresponding to each second feature XB, where each gradient histogram includes the encrypted sum of the sample gradients G of the second alternative feature values of the corresponding second feature XB.

[0067] After that, the intermediate party C uses its target private key SK to decrypt the encrypted intermediate result [Z] PK into the intermediate result plaintext Z, and sends the intermediate result plaintext Z to Party A (the first party). Party A (the first party) splits each user in the first node based on its first gain IA and the intermediate result plaintext Z, where the first gain IA is the gain calculated using the sample gradient G and for splitting according to the first feature XA.

[0068] It is understandable that the first party A and the second party B maintain a tree model with synchronous splitting, so that when the tree model training is completed, both the first party A and the second party B locally obtain the tree model (risk identification model).

[0069] During the entire training process, there is no clear text exchange of data between the first party A and the second party B. All communication data is encrypted data or the intermediate results after homomorphic operations. In this way, it is ensured that private data will not be leaked during the joint training process, enhancing the security of the data.

[0070] The following combines specific embodiments to elaborate in detail on the method for jointly training a risk identification model that protects private data provided in this specification.

[0071] Figure 2 The flowchart of the method for jointly training a risk identification model that protects private data in an embodiment of this specification is shown. Among them, the multiple parties include the first party A, the second party B, and the intermediate party C. It should be understood that before performing model iterative training (that is, iteratively training the risk identification model, namely the tree model), first is the initialization phase. In this initialization phase, the intermediate party C generates an asymmetric key pair for homomorphic encryption, that is, the target public key PK and the target private key SK. Then, the target public key PK is sent to the first party A and the second party B respectively, and the first party A and the second party B store the target public key PK. The intermediate party C keeps the target private key SK private.

[0072] In addition, to ensure the effective training of the model, the first party A and the second party B need to perform an intersection of private data, and subsequently, jointly train the risk identification model, that is, the tree model, through the characteristics of the users in the intersection and the label values indicating whether they are risk users. Among them, the first party A holds a number of first characteristics XA of each first user in the first user set TA and the label value y indicating whether they are risk users, and the second party B holds a number of second characteristics XB of each second user in the second user set TB.

[0073] Among them, the label value indicating whether it is a risk user may include a first numerical value and a second numerical value. The first numerical value indicates a risk user, for example, 1, and the second numerical value indicates a non-risk user (that is, a non-risk user), for example, 0.

[0074] In one implementation, the process of the first party A and the second party B finding the intersection of private data can be that the first party A and the second party B send the user identifiers of each user in their respective user sets to the intermediate party C. After the intermediate party C obtains the user identifiers 1 of each first user sent by the first party A and the user identifiers 2 of each second user sent by the second party B, based on the user identifiers 1 and the user identifiers 2, it determines the user identifiers of the users in the intersection between the first party A and the second party B (subsequently referred to as the third users, and the third users are the users shared between the first users and the second users). Then, the intermediate party C feeds back the user identifiers of the third users to the first party A and the second party B respectively, so that the first party A and the second party B can determine the user identifiers of their respective users in the intersection, that is, determine the third users, and further determine the characteristics of the corresponding users (the first party A also determines the label values of the corresponding users) for joint training of the model.

[0075] In another implementation, the process of the first party A and the second party B finding the intersection of private data can be that the first party A uses the target public key PK to homomorphically encrypt the user identifiers 1 of each first user to obtain each first identifier ciphertext and sends it to the second party B. The second party B uses the target public key PK to homomorphically encrypt the user identifiers 2 of each second user to obtain each second identifier ciphertext and sends it to the first party A. The first party A and the second party B respectively compare each first identifier ciphertext and each second identifier ciphertext to determine the ciphertexts of the identifiers in the intersection. Furthermore, the first party A and the second party B determine the user identifiers of the third users in the intersection based on the ciphertexts of the identifiers in the intersection they respectively determine, and further determine the characteristics of the corresponding third users (the first party A also determines the label values of the corresponding users) for joint training of the model.

[0076] Then, it enters Figure 2 the model iterative training process shown. Among them, the first party A, the second party B, and the intermediate party C can all be implemented by any device, platform, server, or device cluster with computing and processing capabilities. The method is executed through the cooperation of the first party A, the second party B, and the intermediate party C, and the method includes the following steps S210 - S250:

[0077] In step S210, the first party A obtains the sample gradients G of each user in the first node to be split in the current tree. The sample gradients are determined at least according to the label values y of the corresponding users. Among them, these users belong to the intersection of the user sets held by the first party A and the second party B (all or part of the aforementioned third users). In one implementation manner, the sample gradients G of each user include the first - order gradient g and the second - order gradient h, and the first - order gradient g and the second - order gradient h are respectively the first - order partial derivative and the second - order partial derivative of a preset objective function. Among them, the objective function is used to calculate the loss value of the tree model, including the loss function and the regularization term, and can be expressed by the following formula:

[0078]

[0079] Among them, L (t) represents the loss value determined by using the currently generated trees (the first t - 1 trees) and the current tree (the t-th tree to be generated), n represents the number of users participating in the calculation of the loss value, for example, it is the number of the third user as described above, y i represents the label value of the i-th user, represents the predicted value of the i-th user determined based on the first t - 1 trees and the t-th tree, represents the predicted value of the i-th user determined based on the first t - 1 trees, f t (x i ) represents the predicted value of the i-th user determined based on the t-th tree, x i represents the feature of the i-th user, represents the sum of the regularization terms of t trees, Ω(f t ) represents the regularization term of the t-th tree, C1 is a constant, representing the sum of the regularization terms of the first t - 1 trees.

[0080] After Taylor expansion and transformation to be represented by the weight values of the leaf nodes of the current tree, i.e., the t-th tree, formula (1) can be deformed and sorted out as:

[0081]

[0082]

[0083] Among them, g i represents at the first-order partial derivative, that is, the first-order gradient of the i-th user, h i represents at the second-order partial derivative, that is, the second-order gradient of the i-th user. The first-order gradient and the second-order gradient of the i-th user constitute the sample gradient of the i-th user. I j represents each user in the first node, T represents the number of leaf nodes of the t-th tree, w j represents the weight value of the j-th leaf node, is a constant, γ and λ are preset parameters, and is a constant.

[0084] In one embodiment, in step S210, it can be specifically set as: if the first node is the root node of the current tree, determine the sample gradients of each user based on the predicted values and label values of each user. The predicted values are determined according to the generated trees.

[0085] In one implementation, the current tree can be the first tree. Correspondingly, the predicted value can be determined based on the 0th tree. The predicted value determined based on the 0th tree is the predicted value preset for each user. The preset predicted value can be set to any value between the second value and the first value, for example, 0.5. In another implementation, the current tree is not the first tree. For example, the current tree is the tth tree. Correspondingly, the predicted values of each user are determined based on the generated trees, that is, the first t - 1 trees. Party A determines the label value y and the predicted value of each user in the first node (root node). (Determined based on the generated trees, that is, the first t - 1 trees) After that, based on the label value y and the predicted value of each user determine the sample gradient G of each user.

[0086] In another embodiment, in step S210, it can be specifically set as: if the first node is not the root node, read the calculated sample gradient G of each user falling into the first node.

[0087] In one embodiment, Party A is used to perform the bill settlement service. Correspondingly, the first feature it holds is at least related to the bill settlement service; Party B is used to perform the bill generation service. Correspondingly, the second feature it holds is at least related to the bill generation service. In one embodiment, Party A and Party B are located in different geographical regions. For example, in different countries, or in different cities, provinces, etc.

[0088] After that, in step S220, Party A uses the target public key of the intermediate party C to homomorphically encrypt the sample gradient G of each user, and sends the obtained gradient ciphertexts [G] PK to Party B. Among them, the ith gradient ciphertext [G i PK includes the first - order gradient ciphertext [g i PK and the second - order gradient ciphertext [h i PK .

[0089] Then, in step S230, Party B performs homomorphic operations based on the second feature of each user and the gradient ciphertexts [G] PK to determine the intermediate result ciphertext and send it to the intermediate party C. Among them, the intermediate result ciphertext is related to the second gain split according to the second feature.

[0090] In this step, for each second feature held by Party B, based on the feature value of the second feature of each user, several second alternative feature values of the second feature, and the gradient ciphertexts [G] PK ​​​, determine the ciphertext of the histogram of gradients corresponding to the second feature, where the several second alternative feature values can be determined based on the feature values of the second feature of each user, and each ciphertext of the histogram of gradients includes the ciphertext of the sum of the sample gradients of each second alternative feature value corresponding to the second feature.

[0091] For example, the minimum value of the feature value of the second feature XB1 of each user is 1, and the maximum value is 50. Correspondingly, the several second alternative feature values can be set to 10, 20, 30, 40, and 50 respectively. Among them, the sum of the sample gradients corresponding to the second alternative feature value 10 is the cumulative sum of the gradient ciphertexts of the users whose feature values of the corresponding second feature are less than or equal to 10 (which is a ciphertext obtained through homomorphic operations); the sum of the sample gradients corresponding to the second alternative feature value 20 is the cumulative sum of the gradient ciphertexts of the users whose feature values of the corresponding second feature are greater than 10 and less than or equal to 20; and so on. The sum of the sample gradients corresponding to the second alternative feature value 50 is the cumulative sum of the gradient ciphertexts of the users whose feature values of the corresponding second feature are greater than 40 and less than or equal to 50.

[0092] It can be understood that the above example is only an example and does not limit the setting of the second alternative feature values. Specifically, the several second alternative feature values of the second feature can also be set in combination with the specific distribution of the feature values of the second feature. And for the setting of the first alternative feature values, reference can be made to the setting of the second alternative feature values.

[0093] In one embodiment, the second party B can directly use the ciphertexts of the histograms of gradients corresponding to each second feature as the intermediate result ciphertext [Z] PK , and send it to the intermediate party C.

[0094] In another embodiment, after the second party B determines the ciphertexts of the histograms of gradients corresponding to each second feature, it can also perform homomorphic operations on the ciphertexts of the histograms of gradients corresponding to each second feature to determine the ciphertexts of the second gain values corresponding to each second alternative feature value of the second feature, that is, the second gain ciphertext [IB] corresponding to the second feature. PK . Then, the first party A uses the second gain ciphertexts [IB] corresponding to each second feature PK as the intermediate result ciphertext [Z] PK , and sends it to the intermediate party C. Among them, the process of determining the second gain (ciphertext) corresponding to the second feature will be introduced later and will not be elaborated here.

[0095] It can be understood that the intermediate result ciphertext also includes the ciphertexts of each second alternative feature value of each second feature. In addition, the intermediate result ciphertext can also include the ciphertexts of each second feature.

[0096] Next, after the intermediate party C obtains the ciphertext of the intermediate result sent by the second party B, in step S240: the intermediate party C decrypts the ciphertext of the intermediate result [Z] PK into the plaintext Z of the intermediate result and sends it to the first party A. In this step, the intermediate party C uses the target private key SK it holds to decrypt the ciphertext of the intermediate result [Z] PK to obtain the plaintext Z of the intermediate result, and then sends the plaintext Z of the intermediate result to the first party A.

[0097] It can be understood that when the intermediate party C decrypts the ciphertext of the intermediate result, it decrypts the sum of the gradients of each second alternative eigenvalue in the ciphertext of the intermediate result (which is a ciphertext), or decrypts the ciphertext of the second gain value of each second alternative eigenvalue in the ciphertext of the intermediate result, and does not decrypt the ciphertext of each second alternative eigenvalue of each second feature in the ciphertext of the intermediate result. Correspondingly, in the subsequent process, the plaintext of the intermediate result obtained by the first party A does not include the plaintext of each second alternative eigenvalue of each second feature, achieving better protection of the private data of the second party B.

[0098] After that, in step S250, the first party A splits each user in the first node based on the first gain IA and the plaintext Z of the intermediate result. The first gain IA is the gain calculated using the sample gradient G and splitting according to the first feature XA.

[0099] Specifically, in step S250, in one case, the first party A can first determine the gradient histogram corresponding to each first feature based on the sample gradient G of each user, the eigenvalue of each first feature of each user, and the first alternative eigenvalue of each first feature. The gradient histogram corresponding to each first feature includes the sum of the sample gradients corresponding to each first alternative eigenvalue of this first feature. After that, the first party A determines the first gain corresponding to each first feature based on the gradient histogram corresponding to each first feature. Among them, the first gain includes multiple first gain values obtained by splitting according to each first alternative eigenvalue of each first feature respectively.

[0100] In another case, the first party A can directly determine the first intermediate sum and the second intermediate sum corresponding to each first alternative eigenvalue of each first feature based on the sample gradient of each user, the eigenvalue of each first feature of each user, and the first alternative eigenvalue of each first feature. Among them, the first intermediate sum corresponding to the first alternative eigenvalue is the sum of the sample gradients of the users whose eigenvalues of the corresponding first feature are less than or equal to this first alternative eigenvalue; the second intermediate sum corresponding to the first alternative eigenvalue is the sum of the sample gradients of the users whose eigenvalues of the corresponding first feature are greater than this first alternative eigenvalue.

[0101] Furthermore, the first party A determines the first gain value corresponding to each first alternative feature value of each first feature based on the first intermediate sum and the second intermediate sum corresponding to each first alternative feature value of each first feature, that is, multiple first gain values obtained by splitting according to each first alternative feature value of each first feature.

[0102] For example, the minimum value of the feature value of the first feature XA1 in the first node is 11, and the maximum value is 29. Several first alternative feature values of the first feature XA1 can be set to 15, 20, 25, and 30 respectively; among them, the first intermediate sum corresponding to the first alternative feature value 15 is the accumulated sum of the sample gradients of the users whose corresponding feature values of the first feature XA1 are less than or equal to 15, and the corresponding second intermediate sum is the accumulated sum of the sample gradients of the users whose corresponding feature values of the first feature XA1 are greater than 15; and so on, the first intermediate sum corresponding to the first alternative feature value 30 is the accumulated sum of the sample gradients of the users whose corresponding feature values of the first feature XA1 are less than or equal to 30, and the corresponding second intermediate sum is the accumulated sum of the sample gradients of the users whose corresponding feature values of the first feature XA1 are greater than 30.

[0103] It can be understood that each sample gradient includes the corresponding first-order gradient and second-order gradient. The first intermediate sum corresponding to the first alternative feature value includes the accumulated sum of the first-order gradient and the accumulated sum of the second-order gradient. Similarly, the corresponding second intermediate sum includes the accumulated sum of the first-order gradient and the accumulated sum of the second-order gradient. Subsequently, for each first alternative feature value of each first feature, the first party A determines the first gain value corresponding to the first alternative feature value based on the accumulated sum of the first-order gradient and the accumulated sum of the second-order gradient included in the first intermediate sum corresponding to it, the accumulated sum of the first-order gradient and the accumulated sum of the second-order gradient included in the second intermediate sum corresponding to it, and the accumulated sum of the first-order gradient and the accumulated sum of the second-order gradient of all users.

[0104] Specifically, the gain value corresponding to any alternative feature value (including each first alternative feature value of the first feature and each second alternative feature value of the second feature) can be determined by the following formula;

[0105]

[0106] where Gain represents the gain value corresponding to the alternative feature value, G L and H L respectively represent the accumulated sum of the first-order gradient and the accumulated sum of the second-order gradient in the first intermediate sum corresponding to the alternative feature value; G R and H R respectively represent the accumulated sum of the first-order gradient and the accumulated sum of the second-order gradient in the second intermediate sum corresponding to the alternative feature value, and λ represents a preset parameter, which is a constant.

[0107] Subsequently, in one implementation manner, in step S250, it is specifically set as follows: in step 01, the first party A determines the target gain based on the first gain and the plaintext of the intermediate result. Generally, the target gain is the gain with the largest value among the first gain and the second gain. In step 02, the first party A splits each user in the first node based on the target gain. That is, based on the target gain, each user in the first node is divided into the left child node or the right child node of the first node. Among them, the specific splitting strategy used by the first party A when splitting each user in the first node is not limited in this embodiment.

[0108] Correspondingly, in one embodiment, the first gain includes multiple first gain values obtained by splitting according to each first alternative feature value of each first feature, that is, the first gain values corresponding to each first alternative feature value of each of the foregoing first features; the aforementioned plaintext of the intermediate result is the second gain, and the second gain includes multiple second gain values obtained by splitting according to each second alternative feature value of each second feature, that is, the second gain values corresponding to each second alternative feature value of each of the second features; at this time, in step 01, the first party A determines the gain with the largest value among the multiple first gain values and the multiple second gain values as the target gain. Specifically, the first party A directly compares the magnitudes of the multiple first gain values and the multiple second gain values, and determines the gain value with the largest value among them as the target gain.

[0109] In another embodiment, the first gain includes multiple first gain values obtained by splitting according to each first alternative feature value of each first feature; the plaintext of the intermediate result includes: the gradient histogram corresponding to each second feature, and the gradient histogram of any second feature includes the sum of the sample gradients corresponding to each second alternative feature value of this second feature; at this time, in step 01, it is specifically set as follows: in step 011, the first party A determines the second gain based on the gradient histogram corresponding to each second feature, where the second gain includes multiple second gain values obtained by splitting according to each second alternative feature value of each second feature; in step 012, the first party A determines the gain value with the largest value among the multiple first gain values and the multiple second gain values as the target gain.

[0110] It can be understood that the sum of the sample gradients of each alternative feature value of each feature in the gradient histogram is sorted in ascending order of each alternative feature value, that is, for the sum of the sample gradients from left to right in the gradient histogram, the corresponding alternative feature values increase in sequence. In view of this, for any target second feature, on the left side of the ciphertext of each second alternative feature value in the gradient histogram corresponding to this target second feature is the sum of the sample gradients of the users whose feature values of the corresponding target second feature are less than or equal to the plaintext of this second alternative feature value, and on the right side is the sum of the sample gradients of the users whose feature values of the corresponding target second feature are greater than the plaintext of this second alternative feature value.

[0111] Correspondingly, in one implementation, in step 011, for any target second feature, Party A takes the ciphertext of any second alternative feature value corresponding to the target second feature as the target value, and based on the gradient histogram of the target second feature, calculates the first sum value of the sum of all sample gradients on the left side of the target value, and the second sum value of the sum of all sample gradients on the right side of the target value; based on the first sum value, the second sum value, and the third sum value of the sample gradients of each user, determine the second gain value for splitting according to any second alternative feature value of the target second feature.

[0112] In one implementation, the foregoing sample gradients include first-order gradients and second-order gradients; correspondingly, both the first sum value and the second sum value include the sum of the corresponding first-order gradient sums and the sum of the second-order gradient sums. Among them, the sum of the corresponding first-order gradient sums included in the first sum value is the sum of all first-order gradients on the left side of the target value, and the sum of the corresponding second-order gradient sums included is the sum of all second-order gradients on the left side of the target value. The sum of the corresponding first-order gradient sums included in the second sum value is the sum of all first-order gradients on the right side of the target value, and the sum of the corresponding second-order gradient sums included is the sum of all second-order gradients on the right side of the target value.

[0113] Furthermore, Party A determines each second gain value based on the foregoing formula (3), where G in formula (3) L Substitute the sum of the corresponding first-order gradient sums included in the first sum value, and H in formula (3) L Substitute the sum of the corresponding second-order gradient sums included in the first sum value; G in formula (3) R Substitute the sum of the corresponding first-order gradient sums included in the second sum value, and H in formula (3) R Substitute the sum of the corresponding second-order gradient sums included in the second sum value.

[0114] After Party A determines the first gain and the second gain, it can determine the gain value with the largest numerical value from the multiple first gain values included in the first gain and the multiple second gain values included in the second gain as the target gain. Furthermore, based on the target gain, split each user in the first node.

[0115] Specifically, given that the determined target gain may be the gain value corresponding to the first feature or the gain value corresponding to the second feature, and given that Party A cannot obtain the plaintext of each second alternative feature value of each second feature, nor can it obtain the plaintext of each second feature (i.e., Party A does not know which features of the user Party B specifically holds and the specific feature values). Therefore, when the target gain is the gain value of the first feature held by Party A itself, the users of the first node can be directly split based on the feature corresponding to the target gain (a certain first feature corresponding thereto) and the split value (a certain first alternative feature value corresponding thereto), and then the split result (i.e., the subsequent first division result) is synchronized to Party B, so that Party B performs synchronous splitting.

[0116] As Figure 3 shown, the first node includes users a, b, c, d, e, f, g, and h. Among them, Party A determines that users a, b, c, and d are divided into the left child node of the first node, and users e, f, g, and h are divided into the left child node of the first node based on the feature corresponding to the target gain and the split value. Subsequently, Party B divides the users (users a, b, c, d, e, f, g, and h) in the second node corresponding to the first node into the left child node (users a, b, c, and d) or the right child node (users e, f, g, and h) of the second node based on the first division result sent by Party A.

[0117] When the target gain is the gain value of the second feature held by Party B, the ciphertext of the feature corresponding to the target gain (the ciphertext of a certain second feature) and the ciphertext of the split value (the ciphertext of a certain second alternative feature value) can be sent to Party B. Then, Party B splits the users in the second node corresponding to the first node and synchronizes the split result (i.e., the subsequent second division result) to Party A, so that Party A synchronously splits each user in its first node to achieve the splitting of the first node in this iteration.

[0118] Correspondingly, in one embodiment, in step 02, it is specifically set as follows: in step 021, if the target gain belongs to the first gain, Party A divides each user into the left child node or the right child node of the first node based on the feature corresponding to the target gain, the split value, and the feature value of this feature corresponding to each user, to obtain the first division result.

[0119] In step 022, the first party A sends the first partitioning result to the second party A; wherein, the first partitioning result at least includes information indicating which users are partitioned to the left child node of the first node and which users are partitioned to the right child node of the first node. In one case, the first partitioning result may further include the feature ciphertext corresponding to the target gain and the split value ciphertext obtained by homomorphic encryption using the target public key PK.

[0120] In step 023, based on the first partitioning result, the second party B partitions the users of the second node corresponding to the first node into the left child node or the right child node of the second node. The users in the second node are the same as those in the first node.

[0121] Correspondingly, in another embodiment, in step 02, it is further specifically set as: in step 024, if the target gain belongs to the second gain, the first party A sends the feature ciphertext and the split value ciphertext corresponding to the target gain to the second party B.

[0122] In step 025, based on the corresponding feature plaintext, split value plaintext, and the feature values of the feature plaintext of each user, the second party B partitions each user into the left child node or the right child node of the second node to obtain the second partitioning result. Wherein, the second partitioning result at least includes information indicating which users are partitioned to the left child node of the second node and which users are partitioned to the right child node of the second node.

[0123] In step 026, the first party A receives the second partitioning result from the second party B and, based on the second partitioning result, partitions each user into the left child node or the right child node of the first node.

[0124] Splitting each user in the first node, that is, splitting the first node, realizes an update of the current tree, which is also equivalent to an update of the tree model. The current tree is iteratively updated multiple times until the current tree reaches the preset generation condition. The first party A determines the weight values of each leaf node based on the sample gradients in the current leaf nodes of the current tree. Then, the first party A sends the weight values of each leaf node of the current tree to the second party B, and the second party B obtains the weight values to achieve synchronous iterative update of the current tree between the first party A and the second party B, that is, synchronously obtaining the tree model.

[0125] Wherein, the preset generation condition includes: the users in the current leaf nodes of the current tree cannot be further split, or the depth of the current tree reaches the preset depth threshold.

[0126] After determining the current tree, the first party A or the second party B uses the tree model including the current tree and the eigenvalue of the features (including the first feature and the second feature) of each user (the aforementioned third user) within the intersection of the first party A and the second party B to determine the current prediction value of each user. Furthermore, based on the aforementioned objective function, the current prediction value and the label value of each user, the current loss value is determined; it is judged whether the current loss value is less than the preset loss threshold; if the current loss value is less than the preset loss threshold, it is determined that the tree model meets the preset convergence condition, and it is determined that the training of the tree model is completed.

[0127] After that, the tree model can continue to be tested and verified. If the test and verification pass, the tree model can be used for risk prediction. If the current loss value is not less than the preset loss threshold, it is determined that the tree model does not meet the preset convergence condition, and the tree model continues to be trained (for example, continue to generate a new tree) until the tree model meets the preset convergence condition.

[0128] In another implementation, if the test and verification of the tree model fails, the features required for training the tree model can be adjusted according to expert experience. For example, the first feature x1 is removed and a new first feature x7 is added. Another example is to remove the first feature x5 and add a new first feature x8, etc. The learning rate and other information required for training the tree model can also be adjusted according to expert experience.

[0129] In one implementation, when the current tree does not meet the preset generation condition, if the loss value calculated using the label value and the target prediction value of each third user is less than the preset loss threshold, it can also be considered that the tree model including the current tree meets the preset convergence condition, where the target prediction value is determined based on the weight value of the leaf nodes included in the current tree, the weight value of each leaf node in the generated tree, and the objective function for each third user.

[0130] In this embodiment, during the process of the multi-party joint training of the risk identification model (tree model), when the first party (holding several first features of the user and the label value of whether it is a risk user) and the second party (holding several second features of the user) who hold private data interact with each other regarding the training-related data, first, after performing homomorphic encryption or homomorphic operation using the target public key provided by the intermediate party, it is sent to the other party, so that the other party cannot obtain the plaintext of its own private data, and the protection of their respective private data is realized during the process of jointly training the model.

[0131] Among them, in the tree models generated locally by the first party A and the second party B in this embodiment, when the splitting feature and splitting value corresponding to a node are not of one's own party, the splitting feature and splitting value corresponding to this node both exist in ciphertext form. For example, taking the first party A as an example, for the tree model (after training) of the first party A locally, if the splitting feature corresponding to a node therein belongs to the second feature, then the splitting feature and splitting value corresponding to this node are both in ciphertext form.

[0132] Corresponding to the above method embodiment, an embodiment of this specification also provides a method for jointly training a risk identification model by multiple parties to protect private data. The multiple parties include a first party, a second party, and an intermediate party. The first party holds several first features of a user and the label value of whether the user is a risk user. The second party holds several second features of the user. The risk identification model is a tree model. The method is executed by the first party, as Figure 4 shown, and includes:

[0133] S410. Obtain the sample gradients of each user in the first node to be split in the current tree, and use the target public key of the intermediate party to homomorphically encrypt the sample gradients of each user, and send the obtained ciphertexts of each gradient to the second party, where the sample gradients are determined at least according to the label value of the corresponding user;

[0134] S420. Receive the intermediate result plaintext from the intermediate party. The intermediate result plaintext is the result of the intermediate party decrypting the intermediate result ciphertext sent by the second party. The intermediate result ciphertext is determined by the second party through homomorphic operations based on the second features of each user and the ciphertexts of each gradient, and is related to the second gain for splitting according to the second features;

[0135] S430. Split each user in the first node based on the first gain and the intermediate result plaintext, where the first gain is the gain calculated using the sample gradients and for splitting according to the first features.

[0136] In an implementable manner, the S430 includes:

[0137] Determine the target gain based on the first gain and the intermediate result plaintext;

[0138] Split each user in the first node based on the target gain.

[0139] In an implementable manner, the first gain includes multiple first gain values obtained by splitting according to each first alternative feature value of each first feature; the intermediate result plaintext is the second gain, and the second gain includes multiple second gain values obtained by splitting according to each second alternative feature value of each second feature;

[0140] The determining of the target gain includes:

[0141] Determine the gain value with the largest numerical value among the multiple first gain values and the multiple second gain values as the target gain.

[0142] In an implementable manner, the first gain includes multiple first gain values obtained by splitting according to the first alternative feature values of each first feature respectively; the intermediate result plaintext includes: the gradient histogram corresponding to each second feature, and the gradient histogram of any second feature includes the sum of the sample gradients corresponding to the second alternative feature values of this second feature.

[0143] The determination of the target gain includes:

[0144] Based on the gradient histogram corresponding to each second feature, determine the second gain, which includes multiple second gain values obtained by splitting according to the second alternative feature values of each second feature respectively.

[0145] Determine the gain value with the largest numerical value among the multiple first gain values and the multiple second gain values as the target gain.

[0146] In an implementable manner, splitting each user in the first node based on the target gain includes:

[0147] If the target gain belongs to the first gain, based on the feature and split value corresponding to the target gain, and the feature value of this feature corresponding to each user, divide each user into the left child node or the right child node of the first node to obtain a first division result.

[0148] Send the first division result to the second party; so that the second party divides each user of the second node corresponding to the first node into the left child node or the right child node of the second node based on the first division result.

[0149] In an implementable manner, it further includes:

[0150] If the target gain belongs to the second gain, send the feature ciphertext and split value ciphertext corresponding to the target gain to the second party; so that the second party divides each user into the left child node or the right child node of the second node based on the corresponding feature plaintext, split value plaintext, and the feature value of this feature plaintext of each user to obtain a second division result.

[0151] Receive the second division result from the second party, and divide each user into the left child node or the right child node of the first node based on the second division result.

[0152] In an implementable manner, the determination of the second gain includes:

[0153] For any target second feature, use the ciphertext of any second alternative feature value corresponding to the target second feature as the target value. Based on the histogram of gradients of the target second feature, calculate the first sum value of the sum of all sample gradients on the left side of the target value, and the second sum value of the sum of all sample gradients on the right side of the target value; Based on the first sum value, the second sum value, and the third sum value of the sample gradients of each user, determine the second gain value for splitting according to the any second alternative feature value of the target second feature.

[0154] In an implementable manner, the sample gradient includes a first-order gradient and a second-order gradient; both the first sum value and the second sum value include the sum of the corresponding first-order gradients and the sum of the second-order gradients.

[0155] In an implementable manner, the users belong to the intersection of the user sets held by the first party and the second party respectively.

[0156] In an implementable manner, it further includes:

[0157] If the current tree reaches the preset generation condition, based on the sample gradients in the current leaf nodes of the current tree, determine the weight values of each leaf node.

[0158] In an implementable manner, the S410 includes:

[0159] If the first node is the root node of the current tree, determine the sample gradients of each user based on the predicted values and label values of each user; the predicted values are determined according to the generated tree;

[0160] If the first node is not the root node, read the calculated sample gradients of each user falling into the first node.

[0161] In an implementable manner, the first party is used to perform the bill settlement service, and the first feature is related to the bill settlement service; the second party is used to perform the bill generation service, and the second feature is related to the bill generation service.

[0162] In an implementable manner, the first party and the second party are located in different geographical regions.

[0163] The above content describes specific embodiments of this specification, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments, and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily have to be performed in the specific order or continuous order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0164] Corresponding to the above method embodiments, an embodiment of this specification provides an apparatus 500 for identifying risks in a multi-party joint training of a privacy data protection risk identification model. The multi-party includes a first party, a second party, and an intermediate party. The first party holds a number of first features of a user and a label value indicating whether the user is a risk user. The second party holds a number of second features of the user. The risk identification model is a tree model. The apparatus is deployed on the first party, and its schematic block diagram is as Figure 5 shown and includes:

[0165] An encryption acquisition module 510, configured to acquire the sample gradients of each user in the first node to be split in the current tree, and homomorphically encrypt the sample gradients of each user by using the target public key of the intermediate party, and send the obtained ciphertexts of each gradient to the second party, where the sample gradients are determined at least according to the label values of the corresponding users;

[0166] A first receiving module 520, configured to receive the intermediate result plaintext from the intermediate party. The intermediate result plaintext is the result of the intermediate party decrypting the intermediate result ciphertext sent by the second party. The intermediate result ciphertext is determined by the second party through homomorphic operations based on the second features of each user and the ciphertexts of each gradient, and is related to the second gain for splitting according to the second features;

[0167] A splitting module 530, configured to split each user in the first node based on the first gain and the intermediate result plaintext, where the first gain is the gain calculated by using the sample gradients and for splitting according to the first features.

[0168] In an implementable manner, the splitting module 530 includes:

[0169] A first determination unit (not shown in the figure), configured to determine a target gain based on the first gain and the intermediate result plaintext;

[0170] A splitting unit (not shown in the figure), configured to split each user in the first node based on the target gain.

[0171] In an implementable manner, the first gain includes a plurality of first gain values obtained by splitting according to each first alternative feature value of each first feature; the intermediate result plaintext is the second gain, and the second gain includes a plurality of second gain values obtained by splitting according to each second alternative feature value of each second feature;

[0172] The first determination unit is specifically configured to determine the gain value with the largest numerical value among the plurality of first gain values and the plurality of second gain values as the target gain.

[0173] In one implementable manner, the first gain includes a plurality of first gain values that are respectively split according to the first alternative feature values of the respective first features; the intermediate result plaintext includes: the histogram of gradients corresponding to the respective second features, and the histogram of gradients of any second feature includes the sum of the sample gradients corresponding to the respective second alternative feature values of the second feature;

[0174] The first determination unit is specifically configured to determine the second gain based on the histogram of gradients corresponding to the respective second features, including a plurality of second gain values that are respectively split according to the second alternative feature values of the respective second features;

[0175] Determine the gain value with the largest value among the plurality of first gain values and the plurality of second gain values as the target gain.

[0176] In one implementable manner, the splitting unit is specifically configured to, if the target gain belongs to the first gain, divide each user into the left child node or the right child node of the first node based on the feature and the splitting value corresponding to the target gain, and the feature value of this feature corresponding to each user, to obtain a first division result;

[0177] Send the first division result to the second party; so that the second party divides the users of the second node corresponding to the first node into the left child node or the right child node of the second node based on the first division result.

[0178] In one implementable manner, the splitting unit is further specifically configured to:

[0179] If the target gain belongs to the second gain, send the feature ciphertext and the splitting value ciphertext corresponding to the target gain to the second party; so that the second party divides each user into the left child node or the right child node of the second node based on the corresponding feature plaintext, the splitting value plaintext, and the feature value of this feature plaintext of each user, to obtain a second division result;

[0180] Receive the second division result from the second party, and divide each user into the left child node or the right child node of the first node based on the second division result.

[0181] In one implementable manner, the first determination unit is specifically configured to, for any target second feature, use the ciphertext of any second alternative feature value corresponding to the target second feature as the target value, and based on the histogram of gradients of the target second feature, calculate a first sum value of the sums of the sample gradients of all samples to the left of the target value, and a second sum value of the sums of the sample gradients of all samples to the right of the target value; based on the first sum value, the second sum value, and a third sum value of the sample gradients of each user, determine a second gain value for splitting according to the any second alternative feature value of the target second feature.

[0182] In one implementable manner, the sample gradients include first-order gradients and second-order gradients; both the first sum value and the second sum value include the sum of the corresponding first-order gradients and the sum of the second-order gradients.

[0183] In one implementable manner, the users belong to the intersection of the user sets respectively held by the first party and the second party.

[0184] In one implementable manner, the apparatus further includes:

[0185] A determination module, configured to, if the current tree reaches a preset generation condition, determine the weight values of each leaf node based on the sample gradients in each current leaf node of the current tree.

[0186] In one implementable manner, the obtaining encryption module 510 is specifically configured to, if the first node is the root node of the current tree, determine the sample gradients of each user based on the predicted values and label values of each user; the predicted values are determined according to the generated tree;

[0187] If the first node is not the root node, read the calculated sample gradients of each user falling into the first node.

[0188] In one implementable manner, the first party is used to perform bill settlement services, and the first feature is related to bill settlement services; the second party is used to perform bill generation services, and the second feature is related to bill generation services.

[0189] In one implementable manner, the first party and the second party are located in different geographical regions.

[0190] The above apparatus embodiments correspond to the method embodiments. For specific descriptions, reference can be made to the description in the method embodiment part, which will not be elaborated here. The apparatus embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For specific descriptions, reference can be made to the corresponding method embodiments.

[0191] The embodiments of this specification also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method for asset transfer in the payment platform provided in this specification.

[0192] The embodiments of this specification also provide a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method for asset transfer in the payment platform provided in this specification is implemented.

[0193] The various embodiments in this specification are all described in a progressive manner. For the parts that are the same or similar among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the storage medium and the computing device, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiments.

[0194] Those skilled in the art should be able to realize that, in the above one or more examples, the functions described in the embodiments of the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0195] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the embodiments of the present invention. It should be understood that the above is only the specific embodiments of the embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for a multi - party joint training risk identification model to protect privacy data. The multi - parties include a first party, a second party, and an intermediate party. The first party holds several first features of users and the label values of whether the users are risk users. The second party holds several second features of users. The risk identification model is a tree model. The method includes: The first party obtains the sample gradients of each user in the first node to be split in the current tree, and uses the target public key of the intermediate party to homomorphically encrypt the sample gradients of each user, and sends the obtained ciphertexts of each gradient to the second party, where the sample gradients are determined at least according to the label values of the corresponding users; The second party performs a homomorphic operation based on the second features of each user and the ciphertexts of each gradient to determine the intermediate result ciphertext, and sends it to the intermediate party, where the intermediate result ciphertext is related to the second gain obtained by splitting according to the second features; The intermediate party decrypts the intermediate result ciphertext into the intermediate result plaintext and sends it to the first party; The first party splits each user in the first node based on the first gain and the intermediate result plaintext, where the first gain is the gain calculated using the sample gradients and obtained by splitting according to the first features.

2. The method according to claim 1, wherein The splitting of each user in the first node includes: The first party determines the target gain based on the first gain and the intermediate result plaintext; The first party splits each user in the first node based on the target gain.

3. The method according to claim 2, wherein, The first gain includes multiple first gain values obtained by splitting according to each first alternative feature value of each first feature; the intermediate result plaintext is the second gain, and the second gain includes multiple second gain values obtained by splitting according to each second alternative feature value of each second feature; The determining of the target gain includes: The first party determines the gain value with the largest numerical value among the multiple first gain values and the multiple second gain values as the target gain.

4. The method according to claim 2, wherein The first gain includes multiple first gain values obtained by splitting according to each first alternative feature value of each first feature; The intermediate result plaintext includes: the gradient histograms corresponding to each second feature. The gradient histogram of any second feature includes the sum of the sample gradients corresponding to each second alternative feature value of this second feature; The determining of the target gain includes: The first party determines the second gain based on the gradient histograms corresponding to each second feature, where it includes multiple second gain values obtained by splitting according to each second alternative feature value of each second feature; The first party determines the gain value with the largest numerical value among the multiple first gain values and the multiple second gain values as the target gain.

5. The method according to claim 2, wherein, The splitting of each user in the first node based on the target gain includes: If the target gain belongs to the first gain, the first party divides each user into the left child node or the right child node of the first node based on the feature and split value corresponding to the target gain, and the feature value of this feature corresponding to each user, to obtain the first division result; The first party sends the first division result to the second party; Based on the first partitioning result, the second party partitions the users corresponding to the second node corresponding to the first node into the left child node or the right child node of the second node.

6. The method according to claim 5, further comprising: If the target gain belongs to the second gain, the first party sends the ciphertext of the feature corresponding to the target gain and the ciphertext of the split value to the second party; Based on the corresponding plaintext of the feature, the plaintext of the split value, and the feature value of the plaintext of the feature of each user, the second party partitions each user into the left child node or the right child node of the second node to obtain a second partitioning result; The first party receives the second partitioning result from the second party and, based on the second partitioning result, partitions each user into the left child node or the right child node of the first node.

7. The method according to claim 4, wherein The determination of the second gain includes: For any target second feature, the first party uses the ciphertext of any second alternative feature value corresponding to the target second feature as the target value, and based on the gradient histogram of the target second feature, calculates the first sum value of the sum of all sample gradients on the left side of the target value and the second sum value of the sum of all sample gradients on the right side of the target value; based on the first sum value, the second sum value, and the third sum value of the sample gradients of each user, determine the second gain value for splitting according to the any second alternative feature value of the target second feature.

8. The method according to claim 7, wherein The sample gradient includes a first-order gradient and a second-order gradient; both the first sum value and the second sum value include the sum of the corresponding first-order gradients and the sum of the second-order gradients.

9. The method according to any one of claims 1 to 7, wherein The users belong to the intersection of the user sets held by the first party and the second party respectively.

10. The method according to any one of claims 1-7, further comprising: If the current tree reaches a preset generation condition, the first party determines the weight values of each leaf node based on the sample gradients in the current leaf nodes of the current tree.

11. The method according to any one of claims 1-7, wherein, The first party obtains the sample gradients of the users in the first node to be split in the current tree, including: If the first node is the root node of the current tree, determine the sample gradients of each user based on the predicted values and label values of each user; the predicted values are determined according to the generated tree; If the first node is not the root node, read the calculated sample gradients of the users falling into the first node.

12. The method according to any one of claims 1-7, wherein, The first party is used to perform the bill settlement service, and the first feature is related to the bill settlement service; the second party is used to perform the bill generation service, and the second feature is related to the bill generation service.

13. The method according to any one of claims 1-7, wherein, The first party and the second party are located in different geographical regions.

14. A method for jointly training a risk identification model for protecting privacy data by multiple parties, the multiple parties including a first party, a second party, and an intermediate party, the first party holds a number of first features of users and the label values of whether they are risk users, the second party holds a number of second features of users, the risk identification model is a tree model, and the method is executed by the first party and includes: Obtain the sample gradients of each user in the first node to be split in the current tree, and use the target public key of the intermediate party to homomorphically encrypt the sample gradients of each user, and send the obtained ciphertexts of each gradient to the second party, where the sample gradients are determined at least according to the label values of the corresponding users; Receive the intermediate result plaintext from the intermediate party, where the intermediate result plaintext is the result of the intermediate party decrypting the intermediate result ciphertext sent by the second party, and the intermediate result ciphertext is determined by the second party through homomorphic operations based on the second features of each user and the ciphertexts of each gradient, and is related to the second gain for splitting according to the second features; Split each user in the first node based on the first gain and the intermediate result plaintext, where the first gain is the gain calculated using the sample gradients and for splitting according to the first features.

15. An apparatus for a multi-party joint training risk identification model that protects privacy data, where the multi-party includes a first party, a second party, and an intermediate party. The first party holds several first features of a user and the label value of whether the user is a risk user. The second party holds several second features of the user. The risk identification model is a tree model. The apparatus is deployed on the first party and includes: An acquisition and encryption module configured to obtain the sample gradients of each user in the first node to be split in the current tree, and use the target public key of the intermediate party to homomorphically encrypt the sample gradients of each user, and send the obtained ciphertexts of each gradient to the second party, where the sample gradients are determined at least according to the label values of the corresponding users; A first receiving module configured to receive the intermediate result plaintext from the intermediate party, where the intermediate result plaintext is the result of the intermediate party decrypting the intermediate result ciphertext sent by the second party, and the intermediate result ciphertext is determined by the second party through homomorphic operations based on the second features of each user and the ciphertexts of each gradient, and is related to the second gain for splitting according to the second features; A splitting module configured to split each user in the first node based on the first gain and the intermediate result plaintext, where the first gain is the gain calculated using the sample gradients and for splitting according to the first features.

16. A computing device, comprising a memory and a processor, wherein, The memory stores executable code, and when the processor executes the executable code, the method described in claim 14 is implemented.

Citation Information

Patent Citations

  • Two-party decision tree training method and system

    CN111738359A

  • Multi-party joint decision tree construction method, device and readable storage medium

    WO2021249086A1