Vertical Federated XGBoost Training Method, System, Device, Medium and Product
Through the method of index matrix and secret sharing, the training problem of the federated XGBoost algorithm in the scenario of separation of the computing party and data party and distrust between each other is solved, and a safe and efficient model training process is realized, ensuring the accuracy and security of the model.
Patent Information
- Application Number
- CN202211593740.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-13
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-12-13
AI Technical Summary
The existing federated XGBoost algorithm cannot be applied to scenarios where the computing party and data party are separated and distrust each other, resulting in the inability to perform model training effectively.
The index matrix and secret sharing are used to replace the interaction of original data and tags, and the accuracy of node splitting and node weight calculation is ensured through the addition homomorphism of secret sharing. The index matrix and secret sharing are used to replace the interaction of original data and tags, so that neither the Guest party nor the Host party can know each other's data.
While ensuring security, the lossless training of the model is realized, and the model training can be efficiently completed with only a small number of interactions, ensuring the accuracy and security of the model.
Smart Images

Figure CN118194029B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of machine learning technology, and in particular, to a vertical federated XGBoost training method, system, device, medium and product. Background Art
[0002] Federated learning is a new paradigm of machine learning, aiming to jointly perform machine learning algorithm modeling with multiple parties while meeting the requirements of user privacy protection, data security, and legal compliance. As a typical representative of machine learning algorithms, XGBoost [1] has less hyperparameter tuning, excellent interpretability, and powerful performance, and is widely used in fields such as recommendation systems and data mining.
[0003] In federated learning, there are generally two role parties, Guest and Host. Currently, in the case where the computing party and the data party are separated in the federated XGBoost algorithm, usually a certain role party is both the computing party and the data party, that is, this role party provides both data and computing services. In the scenario where the role party is separated into a data party and a computing party, that is, the role party separates data resources and computing resources, and at the same time the two resource services do not trust each other, the existing federated learning XGBoost is not applicable. Summary of the Invention
[0004] To overcome the problems in the related art, the present disclosure provides a vertical federated XGBoost training method, system, device, medium and product.
[0005] According to the first aspect of the embodiments of the present disclosure, a vertical federated XGBoost training method is provided. The Guest party includes a first data party and a first computing party, and the Host party includes a second computing party and a second data party, including:
[0006] Sending both the first index matrix of the first data party and the second index matrix of the second data party to the first computing party and the second computing party, where the first index matrix and the second index matrix are obtained by sorting each feature according to their respective samples;
[0007] Dividing the sample labels into normal labels and encrypted labels based on the secret sharing method, sending the normal labels to the first computing party, and sending the encrypted labels to the second computing party, where the second computing party has a private key for decrypting the encrypted labels;
[0008] Obtaining a first intersection sample of the first computing party's current sample in the first index matrix of the normal labels, and obtaining a second intersection sample of the second computing party's current sample in the second index matrix of the encrypted labels;
[0009] Determine candidate split points for the first intersection sample and the second intersection sample based on the number of buckets, divide the first intersection sample and the second intersection sample into left and right subsets, calculate the information gain of each split point, and determine the optimal split point based on the information gain;
[0010] Split the nodes at the positions of the optimal split points of the first intersection sample and the second intersection sample until the iteration condition is met or the sample set cannot be divided any further, obtain multiple decision trees of the Guest party and the Host party, and obtain a vertical federated XGBoost model based on the multiple decision trees.
[0011] In some embodiments, divide the sample labels into normal labels and encrypted labels based on the secret sharing method, send the normal labels to the first computing party, and send the encrypted labels to the second computing party. The second computing party has a private key for decrypting the encrypted labels, including:
[0012] Generate a public key and a private key by the second computing party, and send the public key to the first data party;
[0013] Encrypt some of the sample labels with the public key to generate encrypted labels, and send the encrypted labels to the second computing party.
[0014] In some embodiments, split the nodes at the positions of the optimal split points of the first intersection sample and the second intersection sample until the iteration condition is met or the sample set cannot be divided any further, obtain multiple decision trees of the Guest party and the Host party, and obtain a vertical federated XGBoost model based on the multiple decision trees, including:
[0015] After iterating for a preset number of training rounds, determine whether the first decision tree and the second decision tree reach the maximum number of trees. The first decision tree and the second decision tree are obtained by splitting the first intersection sample and the second intersection sample;
[0016] When the first decision tree and the second decision tree reach the maximum number of trees, index the thresholds of each non-leaf node of each based on the features and split point ids recorded by each node, and obtain the vertical federated XGBoost model.
[0017] In some embodiments, before sending both the first index matrix of the first data party and the second index matrix of the second data party to the first computing party and the second computing party, where the first index matrix and the second index matrix are obtained by sorting each feature according to their respective samples, further include:
[0018] Perform interactive security verification on the Guest party and the Host party.
[0019] According to the second aspect of the embodiments of the present disclosure, a vertical federated XGBoost training system is provided, including:
[0020] A sending module that sends both the first index matrix of the first data party and the second index matrix of the second data party to the first computing party and the second computing party. The first index matrix and the second index matrix are obtained by sorting each feature according to their respective samples.
[0021] An encryption module that divides the sample labels into normal labels and encrypted labels based on secret sharing, sends the normal labels to the first computing party, and sends the encrypted labels to the second computing party. The second computing party has a private key for decrypting the encrypted labels.
[0022] An obtaining module that obtains the first intersection samples of the first computing party's current samples in the first index matrix of the normal labels, and obtains the second intersection samples of the second computing party's current samples in the second index matrix of the encrypted labels.
[0023] A determining module that determines candidate split points for the first intersection samples and the second intersection samples through the number of buckets, divides the first intersection samples and the second intersection samples into left and right subsets, calculates the information gain of each split point, and determines the optimal split point based on the information gain.
[0024] A splitting module that splits the nodes at the positions of the optimal split points of the first intersection samples and the second intersection samples until the iteration condition is met or the sample set cannot be divided any further, obtains multiple decision trees of the Guest party and the Host party, and obtains a vertical federated XGBoost model based on the multiple decision trees.
[0025] In some embodiments, the encryption module is further configured to:
[0026] Generate a public key and a private key through the second computing party, and send the public key to the first data party;
[0027] Encrypt some of the sample labels with the public key to generate encrypted labels, and send the encrypted labels to the second computing party.
[0028] In some embodiments, the splitting module is further configured to,
[0029] After the iteration of the preset number of training rounds, it is determined whether the first decision tree and the second decision tree reach the maximum number of trees, where the first decision tree and the second decision tree are obtained by splitting according to the first intersection sample and the second intersection sample;
[0030] When the first decision tree and the second decision tree reach the maximum number of trees, based on the features and split point ids recorded by each node, the thresholds of each non-leaf node of each are indexed to obtain the vertical federated XGBoost model.
[0031] An embodiment of the third aspect of the present application provides an electronic device, including a processor and a memory, where at least one instruction, at least one segment of program, code set or instruction set is stored in the memory, and the instruction, the program, the code set or the instruction set is loaded and executed by the processor to implement the steps of the vertical federated XGBoost training method provided by the embodiment of the first aspect of the present application.
[0032] An embodiment of the fourth aspect of the present application provides a non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of a mobile terminal, enabling the mobile terminal to execute the steps of the vertical federated XGBoost training method provided by the embodiment of the first aspect of the present application.
[0033] An embodiment of the fifth aspect of the present application provides a computer program product, when the instructions in the computer program product are executed by a processor of a mobile terminal, enabling the mobile terminal to execute the steps of the vertical federated XGBoost training method provided by the embodiment of the first aspect of the present application.
[0034] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: By using two forms of index matrix and secret sharing to replace the interaction of the original data and labels, neither the Guest party nor the Host party can know the data of the other party; while ensuring security, the additive homomorphism of secret sharing is used to ensure the accuracy of node splitting and node weight calculation, so as to ensure that the model is lossless. At the same time, the computing party and the data party can efficiently complete the training of the model with only a few interaction times.
[0035] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.
[0037] Figure 1It is a flowchart of a vertical federated XGBoost training method shown according to an exemplary embodiment.
[0038] Figure 2 It is a block diagram of a vertical federated XGBoost training system shown according to an exemplary embodiment.
[0039] Figure 3 It is an internal structure diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0040] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all the implementation manners consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0041] In the following description, Secret Share (SS): A technology in cryptography for splitting and storing secrets, aiming to disperse the risk of secret leakage and tolerate intrusion to a certain extent.
[0042] Semi-honest Model: A security model. Under this security model, when a party contacts and processes the privacy data of other parties, on the basis of strictly observing the protocol specifications, it tries its best to extract effective information from the contacted and processed data.
[0043] Extreme Gradient Boosting (XGBoost): An ensemble machine learning algorithm based on decision trees, which reduces the model bias by iteratively learning multiple weak decision trees and continuously fitting the data.
[0044] Vertical Federated Learning (VFL): Refers to federated learning in which different parties have the same sample space and different features.
[0045] Data role: The role party that provides samples.
[0046] Compute role: The role party that provides computing.
[0047] Guest role: The role party that provides sample features and labels.
[0048] Host party (host role): The party that only provides sample features.
[0049] Figure 1 It is a flowchart of a vertical federated XGBoost training method shown according to an exemplary embodiment, as Figure 1 shown, and includes the following steps:
[0050] In step S101, the first index matrix of the first data party and the second index matrix of the second data party are both sent to the first computing party and the second computing party. The first index matrix and the second index matrix are obtained by sorting each feature according to their respective samples.
[0051] Specifically, the first computing party is denoted as G_C, the first data party is denoted as G_D, the second computing party is denoted as H_C, and the second data party is denoted as H_D.
[0052] First step: G_D sorts each feature of the samples it owns to obtain the sorted first index matrix T1. Each column represents the sorting result of each sample on this feature, and T1 is sent to H_C and G_C.
[0053] Second step: H_D sorts each feature of the samples it owns to obtain the sorted second index matrix T2, and sends T2 to H_C and G_C.
[0054] In step S102, the sample labels are divided into normal labels and encrypted labels based on the secret sharing method. The normal labels are sent to the first computing party, and the encrypted labels are sent to the second computing party. The second computing party has the private key for decrypting the encrypted labels.
[0055] In some embodiments, dividing the sample labels into normal labels and encrypted labels based on the secret sharing method, sending the normal labels to the first computing party, and sending the encrypted labels to the second computing party. The second computing party has the private key for decrypting the encrypted labels, includes:
[0056] Generating a public key and a private key by the second computing party, and sending the public key to the first data party;
[0057] Encrypting some of the sample labels with the public key to generate encrypted labels, and sending the encrypted labels to the second computing party.
[0058] Specifically:
[0059] Third step: H_C generates a public-private key pair (k, p), and sends the public key k to G_D.
[0060] Step 4: G_D divides the label y corresponding to the sample i , into <y i1 >, <y i2 > in two parts by the SS method, with y i = <y i1 > + <y i2 >, and encrypts <y i2 > with the public key to obtain [[<y i2 >]], sends <y i1 > to G_C, and sends [[<y i2 >]] to H_C.
[0061] In step S103, obtain the first intersection samples of the current sample of the first computing party in the first index matrix of the normal label, and obtain the second intersection samples of the current sample of the second computing party in the second index matrix of the encrypted label.
[0062] Specifically:
[0063] Step 5: H_C decrypts [[<y i2 >]] with the private key p.
[0064] Step 6: G_C calculates g i1 , h i1 , where g i1 = y_hat i - <y i1 >, h i1 = 1, i ∈ [0, m), where g i1 is the first-order gradient and h i1 is the second-order gradient.
[0065] Step 7: H_C calculates g i2 , h i2 , where g i2 = - <y i2 >, h i2 = 0, i ∈ [0, m).
[0066] Step 8: For the current sample set D and feature j, G_C extracts the sample intersection with D in T1j and maintains the original order.
[0067] In step S104, determine the candidate split points for the first intersection samples and the second intersection samples through the number of buckets, divide the first intersection samples and the second intersection samples into left and right subsets, calculate the information gain of each split point, and determine the optimal split point based on the information gain.
[0068] Specifically:
[0069] Step 8: Determine candidate split points based on the number of buckets for the intersection, divide the sample set into two subsets, the left and the right, and calculate the Gain1 (information gain) for each split point. Among them,
[0070] G1 and H1 are calculated by the Guest. G2 and H2 are calculated by the Host and sent to the Guest. If n = 1, the parent node of this node is set as a leaf node and no node splitting is performed.
[0071] Step 9: Similarly to Step 8, H_C calculates the Gain2 for each split point and sends the Gain2 to G_C.
[0072] Step 10: G_C calculates the Gain of the current feature based on the Gain2 sent by H_C in Step 8, where Gain = Gain1 + Gain2.
[0073] Step 11: Repeat Steps 8 to 10 to calculate the Gain values of all candidate split points for all features of the Guest and the Host.
[0074] Step 12: The G_C party finds the feature and split point corresponding to the maximum Gain value in the Guest and the Host, and records them as the optimal feature and split point.
[0075] In step S105, split the node at the position of the optimal split point of the first intersection sample and the second intersection sample until the iteration condition is met or the sample set cannot be further divided, obtain multiple decision trees of the Guest party and the Host party, and obtain a vertical federated XGBoost model based on the multiple decision trees.
[0076] In some embodiments, splitting the node at the position of the optimal split point of the first intersection sample and the second intersection sample until the iteration condition is met or the sample set cannot be further divided, obtaining multiple decision trees of the Guest party and the Host party, and obtaining a vertical federated XGBoost model based on the multiple decision trees includes:
[0077] After iterating for a preset number of training rounds, determine whether the first decision tree and the second decision tree reach the maximum number of trees. The first decision tree and the second decision tree are obtained by splitting the first intersection sample and the second intersection sample;
[0078] When the first decision tree and the second decision tree reach the maximum number of trees, index the thresholds of each non-leaf node of each based on the features and split point ids recorded by each node, and obtain the vertical federated XGBoost model.
[0079] Specifically:
[0080] Step 13: G_C determines whether the optimal split point belongs to Host. If so, G_C sends the optimal feature and split point to H_C; otherwise, G_C directly completes node splitting according to the optimal split point, records the split point id and feature, and then sends the id of the left subset to H_C.
[0081] Step 14: H_C determines whether the optimal split point belongs to Host according to the information sent by G_C in Step 13. If so, it completes node splitting according to the received optimal feature and split point, and sends the id of the left subset to G_C; otherwise, it completes node splitting according to the received id of the left subset.
[0082] Step 15: For G_C, if the optimal split point belongs to Host, it completes node splitting according to the id of the left subset sent by H_C in Step 14. Otherwise, it skips.
[0083] Step 16: G_C recursively executes Step 6 for the divided left and right subsets to complete the construction of the tree, and then proceeds to Step 18. Among them, if the iteration condition is met or the sample set cannot be further divided, the current node is marked as a leaf node, and the current weight w is calculated, where, Otherwise, it is recorded as a non-leaf node.
[0084] Step 17: H_C recursively executes Step 6 for the divided left and right subsets to complete the construction of the tree, and then proceeds to Step 19. Among them, if the iteration condition is met or the sample set cannot be further divided, the current node is marked as a leaf node; otherwise, it is recorded as a non-leaf node.
[0085] Step 18: G_C obtains the prediction results y_pred of all sample sets based on all leaf nodes i , calculates y_hat i = y_hat i + y_pred i . If the current tree is the last tree, execute Step 20; otherwise, iteratively execute Step 6.
[0086] Step 19: H_C determines that if the current tree is the last tree, execute Step 21; otherwise, iteratively execute Step 7.
[0087] Step 20: G_C returns the recorded features and split point ids to G_D. G_D indexes the thresholds of each non-leaf node belonging to the Guest side to obtain the vertical federated XGBoost model of the Guest side. Among them, for each split of each tree, the current feature and id need to be recorded. After all the tree splits are completed, then index the thresholds of each split point of each tree.
[0088] The twenty - first step: H_C returns the recorded features and split point ids to H_D, and H_D indexes each non - leaf node threshold belonging to the Host party to obtain the vertical federated XGBoost model of the Host party.
[0089] In some embodiments, before sending both the first index matrix of the first data party and the second index matrix of the second data party to the first computing party and the second computing party, where the first index matrix and the second index matrix are obtained by sorting each feature according to their respective samples, it further includes:
[0090] Performing interactive security verification on the Guest party and the Host party.
[0091] Specifically, this application ensures data security by completing the contact and processing of the privacy data of other participating parties under the semi - honest model.
[0092] In summary, through the above - mentioned training interaction method, this application has
[0093] Security:
[0094] Party G_D: G_D has interactions with G_C and H_C respectively; for party G_C, the information sent by G_D to G_C includes T1, <y i1 >, [<y i2 >]]; T1 is the index matrix after the feature sorting of party G_D, which only contains sequential indexes, and party G_C cannot restore the original feature values; <y i 1 >, <y i 2 > is the label information split by SS, and there is y i =<y i1 >+<y i2 >, but what party G_C obtains is [[<y i2 >]] after encryption and cannot restore the real y. For party H_C, H_C only gets <y i2 > and cannot restore the real y.
[0095] Party G_C: G_C has interactions with G_D and H_C respectively. Among them, the interactive security between G_C and G_D has been verified; during the calculation process of G_C and H_C, only G and H are interacted. For G, G is the gradient accumulation sum of multiple samples. When the number of samples is greater than 1, the gradient of each sample cannot be restored; when the number of samples is equal to 1, pruning operations are performed and no nodes are divided, so the security of the gradient of each sample is guaranteed; the same is true for H.
[0096] Party H_C: H_C interacts with G_D, G_C, and H_D respectively. Among them, the interaction security between H_C and G_D, G_C has been verified; in the interaction with H_D, H_C obtained the feature sorting index matrix of H_D and cannot restore the original features.
[0097] Party H_D: H_D interacts with H_C, and the security has been verified in Party H_C.
[0098] Accuracy:
[0099] Feature sorting: During the tree construction process, node splitting requires sorting and bucketing features to determine the splitting point and the optimal feature. In the whole process, only the sorting information of feature values is used for features. The T1 and T2 matrices sent by G_D and H_D are equivalent to the original matrix and are lossless in feature sorting.
[0100] Selection of the optimal feature and splitting point: Since there is y i = <y i1 >+ <y i2 >, the calculation results of G and H are lossless, and thus the result of Gain is also lossless, which is equivalent to the original result in the selection of the optimal feature and splitting point.
[0101] Prediction result: The sample prediction result is obtained from w of the leaf node. Since G and H are lossless as described in 2, w is also lossless.
[0102] Figure 2 It is a block diagram of a vertical federated XGBoost training system shown according to an exemplary embodiment. Referring to Figure 2 , the device includes a sending module 201, an encryption module 202, an acquisition module 203, a determination module 204, and a splitting module 205.
[0103] The sending module 201 sends both the first index matrix of the first data party and the second index matrix of the second data party to the first calculation party and the second calculation party. The first index matrix and the second index matrix are obtained by sorting each feature according to their respective samples;
[0104] The encryption module 202 divides the sample labels into normal labels and encrypted labels based on the secret sharing method, sends the normal labels to the first calculation party, and sends the encrypted labels to the second calculation party. The second calculation party has the private key for decrypting the encrypted labels;
[0105] The acquisition module 203 acquires the first intersection samples of the first calculation party's current samples in the first index matrix of the normal labels, and acquires the second intersection samples of the second calculation party's current samples in the second index matrix of the encrypted labels;
[0106] The determination module 204 determines candidate split points for the first intersection sample and the second intersection sample based on the number of buckets, divides the first intersection sample and the second intersection sample into left and right subsets, calculates the information gain of each split point, and determines the optimal split point based on the information gain;
[0107] The splitting module 205 splits the nodes at the positions of the optimal split points of the first intersection sample and the second intersection sample until the iteration condition is satisfied or the sample set cannot be divided any further, obtains multiple decision trees of the Guest party and the Host party, and obtains a vertical federated XGBoost model based on the multiple decision trees.
[0108] In some embodiments, the encryption module is further configured to:
[0109] Generate a public key and a private key through the second computing party, and send the public key to the first data party;
[0110] Encrypt some of the sample labels through the public key to generate encrypted labels, and send the encrypted labels to the second computing party.
[0111] In some embodiments, the splitting module is further configured to,
[0112] After iterating for a preset number of training rounds, determine whether the first decision tree and the second decision tree reach the maximum number of trees, where the first decision tree and the second decision tree are obtained by splitting the first intersection sample and the second intersection sample;
[0113] When the first decision tree and the second decision tree reach the maximum number of trees, index the thresholds of each non-leaf node of each based on the features and split point ids recorded by each node, and obtain the vertical federated XGBoost model.
[0114] Regarding the system in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0115] In one embodiment, an electronic device is provided. The electronic device may be a terminal, and its internal structural diagram may be as Figure 3As shown in the figure. The electronic device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, near-field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a vertical federated XGBoost training method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.
[0116] Those skilled in the art can understand that Figure 3 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0117] In one embodiment, the vertical federated XGBoost training system provided by this application can be implemented in the form of a computer program, and the computer program can run on an electronic device such as Figure 3 shown. Each program module constituting the vertical federated XGBoost training system can be stored in the memory of the electronic device.
[0118] At least one instruction, at least one program, a code set, or an instruction set is stored in the memory of the electronic device, and the instruction, the program, the code set, or the instruction set is loaded and executed by the processor to implement the vertical federated XGBoost training system according to any one of the above embodiments. For example, to implement a vertical federated XGBoost training system, it includes: sending both the first index matrix of the first data party and the second index matrix of the second data party to the first computing party and the second computing party, where the first index matrix and the second index matrix are obtained by sorting each feature according to their respective samples; dividing the sample labels into normal labels and encrypted labels based on the secret sharing method, sending the normal labels to the first computing party, and sending the encrypted labels to the second computing party, where the second computing party has a private key for decrypting the encrypted labels; obtaining a first intersection sample of the first computing party's current sample in the first index matrix of the normal labels, and obtaining a second intersection sample of the second computing party's current sample in the second index matrix of the encrypted labels; determining candidate split points for the first intersection sample and the second intersection sample through the number of buckets, dividing the first intersection sample and the second intersection sample into left and right subsets, calculating the information gain of each split point, and determining the optimal split point based on the information gain; splitting the nodes at the position of the optimal split point of the first intersection sample and the second intersection sample until the iteration condition is met or the sample set cannot be divided any further, obtaining decision trees of multiple Guest parties and Host parties, and obtaining a vertical federated XGBoost model based on the multiple decision trees.
[0119] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: sending both the first index matrix of the first data party and the second index matrix of the second data party to the first computing party and the second computing party, where the first index matrix and the second index matrix are obtained by sorting each feature according to their respective samples; dividing the sample labels into normal labels and encrypted labels based on the secret sharing method, sending the normal labels to the first computing party, and sending the encrypted labels to the second computing party, where the second computing party has a private key for decrypting the encrypted labels; obtaining a first intersection sample of the first computing party's current sample in the first index matrix of the normal labels, and obtaining a second intersection sample of the second computing party's current sample in the second index matrix of the encrypted labels; determining candidate split points for the first intersection sample and the second intersection sample through the number of buckets, dividing the first intersection sample and the second intersection sample into left and right subsets, calculating the information gain of each split point, and determining the optimal split point based on the information gain; splitting the nodes at the position of the optimal split point of the first intersection sample and the second intersection sample until the iteration condition is met or the sample set cannot be divided any further, obtaining multiple decision trees of the Guest party and the Host party, and obtaining a vertical federated XGBoost model based on the multiple decision trees.
[0120] In one embodiment, a computer program product is provided. When the instructions in the computer program product are executed by a processor of a mobile terminal, the mobile terminal is enabled to perform the following steps: sending both the first index matrix of the first data party and the second index matrix of the second data party to the first computing party and the second computing party, where the first index matrix and the second index matrix are obtained by sorting each feature according to their respective samples; dividing the sample labels into normal labels and encrypted labels based on the secret sharing method, sending the normal labels to the first computing party, and sending the encrypted labels to the second computing party, where the second computing party has a private key for decrypting the encrypted labels; obtaining a first intersection sample of the first computing party's current sample in the first index matrix of the normal labels, and obtaining a second intersection sample of the second computing party's current sample in the second index matrix of the encrypted labels; determining candidate split points for the first intersection sample and the second intersection sample through the number of buckets, dividing the first intersection sample and the second intersection sample into left and right subsets, calculating the information gain of each split point, and determining the optimal split point based on the information gain; splitting the nodes at the positions of the optimal split points of the first intersection sample and the second intersection sample until the iteration condition is met or the sample set cannot be further divided, obtaining multiple decision trees of the Guest party and the Host party, and obtaining a vertical federated XGBoost model based on the multiple decision trees.
[0121] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static random access memory (SRAM) and dynamic random access memory (DRAM), etc.
[0122] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0123] The above embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed. However, it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A vertical federated XGBoost training method, characterized in that, The Guest party includes a first data party and a first computing party, and the Host party includes a second computing party and a second data party, including: Send the first index matrix of the first data party and the second index matrix of the second data party to the first computing party, and send the first index matrix of the first data party and the second index matrix of the second data party to the second computing party. The first index matrix and the second index matrix are obtained by sorting each feature according to their respective samples; Divide the sample labels into normal labels and encrypted labels based on the secret sharing method, send the normal labels to the first computing party, and send the encrypted labels to the second computing party. The second computing party has a private key for decrypting the encrypted labels; Obtain the first intersection samples of the first computing party's current samples in the first index matrix of the normal labels, and obtain the second intersection samples of the second computing party's current samples in the second index matrix of the encrypted labels; Determine candidate split points for the first intersection samples and the second intersection samples through the number of buckets, divide the first intersection samples and the second intersection samples into left and right subsets, calculate the information gain of each split point, and determine the optimal split point based on the information gain; Split the nodes at the positions of the optimal split points of the first intersection samples and the second intersection samples until the iteration condition is met or the sample set cannot be divided any further, obtain multiple decision trees of the Guest party and the Host party, and obtain a vertical federated XGBoost model based on the multiple decision trees.
2. The vertical federated XGBoost training method according to claim 1, wherein, Divide the sample labels into normal labels and encrypted labels based on the secret sharing method, send the normal labels to the first computing party, and send the encrypted labels to the second computing party. The second computing party has a private key for decrypting the encrypted labels, including: Generate a public key and a private key by the second computing party, and send the public key to the first data party; Encrypt some of the sample labels with the public key to generate encrypted labels, and send the encrypted labels to the second computing party.
3. The vertical federated XGBoost training method according to claim 1, wherein, Split the nodes at the positions of the optimal split points of the first intersection samples and the second intersection samples until the iteration condition is met or the sample set cannot be divided any further, obtain multiple decision trees of the Guest party and the Host party, and obtain a vertical federated XGBoost model based on the multiple decision trees, including: After iterating through a preset number of training rounds, determine whether the first decision tree and the second decision tree reach the maximum number of trees. The first decision tree and the second decision tree are obtained by splitting the first intersection samples and the second intersection samples; When the first decision tree and the second decision tree reach the maximum number of trees, index the thresholds of each non-leaf node of each according to the features and split point ids recorded by each node, and obtain the vertical federated XGBoost model.
4. The vertical federated XGBoost training method according to claim 1, wherein Before sending the first index matrix of the first data party and the second index matrix of the second data party to the first computing party and the second computing party, where the first index matrix and the second index matrix are obtained by sorting each feature according to their respective samples, it further includes: Performing interactive security verification on the Guest party and the Host party.
5. A vertical federated XGBoost training system, characterized in that, It includes: A sending module that sends both the first index matrix of the first data party and the second index matrix of the second data party to the first computing party, and sends both the first index matrix of the first data party and the second index matrix of the second data party to the second computing party, where the first index matrix and the second index matrix are obtained by sorting each feature according to their respective samples; An encryption module that divides the sample labels into normal labels and encrypted labels based on secret sharing, sends the normal labels to the first computing party, and sends the encrypted labels to the second computing party, where the second computing party has a private key for decrypting the encrypted labels; An acquisition module that acquires the first intersection samples of the first computing party's current samples in the first index matrix of the normal labels, and acquires the second intersection samples of the second computing party's current samples in the second index matrix of the encrypted labels; A determination module that determines candidate split points for the first intersection samples and the second intersection samples through the number of buckets, divides the first intersection samples and the second intersection samples into left and right subsets, calculates the information gain of each split point, and determines the optimal split point based on the information gain; A splitting module that splits the nodes at the positions of the optimal split points of the first intersection samples and the second intersection samples until the iteration condition is met or the sample set cannot be further divided, obtains multiple decision trees for the Guest party and the Host party, and obtains a vertical federated XGBoost model based on the multiple decision trees.
6. The vertical federated XGBoost training system according to claim 5, wherein The encryption module is further used for: Generating a public key and a private key through the second computing party, and sending the public key to the first data party; Encrypting some of the sample labels with the public key to generate encrypted labels, and sending the encrypted labels to the second computing party.
7. The vertical federated XGBoost training system according to claim 5, wherein The splitting module is further used for, After iterating through a preset number of training rounds, determining whether the first decision tree and the second decision tree have reached the maximum number of trees, where the first decision tree and the second decision tree are obtained by splitting the first intersection samples and the second intersection samples; When the first decision tree and the second decision tree reach the maximum number of trees, indexing the thresholds of each non-leaf node of each based on the features and split point ids recorded in each node, and obtaining the vertical federated XGBoost model.
8. An electronic device, characterized in that, It includes a processor and a memory, where at least one instruction, at least one program, a code set, or an instruction set is stored in the memory, and the instruction, the program, the code set, or the instruction set is loaded and executed by the processor to implement the vertical federated XGBoost training method according to any one of claims 1-4.
9. A non-transitory computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the mobile terminal, the mobile terminal is enabled to execute the vertical federated XGBoost training method according to any one of claims 1-4.
10. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor of the mobile terminal, the mobile terminal is enabled to execute the vertical federated XGBoost training method according to any one of claims 1-4.
Citation Information
Patent Citations
Multi-party XGBoost security prediction model training method based on secret sharing and federated learning
CN112464287A
Longitudinal federal modeling method based on LightGBM algorithm
CN113591152A