Tree model longitudinal federated learning feature importance calculation method and system based on TreeSHAP

By optimizing the TreeSHAP method in a vertical federated learning environment, calculating the sample proportion and shunt characteristics of each node in the tree model, and recursively updating the proportion coefficient array, the problems of high communication overhead and computational complexity in the existing methods are solved, and efficient feature importance calculation is achieved.

CN120105047APending Publication Date: 2025-06-06BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510167020.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06

Smart Images

  • Figure CN120105047A_ABST
    Figure CN120105047A_ABST
Patent Text Reader

Abstract

The invention provides a tree model longitudinal federated learning feature importance calculation method and system based on TreeSHAP, and the method comprises the steps: predicting a data set, calculating the sample proportion of each node in a tree model, and recording the shunting features used by each sub-tree; traversing each node of the tree by adopting a recursive method, and updating the proportionality coefficient array of each node according to the shunting correctness of the sample and the sample proportion to obtain the proportionality coefficient of each leaf node under the condition of different shunting correctness; and recursively inputting the data set into each node, calculating a shunting correctness vector and a target characteristic shunting correctness vector of each sample according to the shunting characteristic of each node, querying the proportionality coefficient according to the shunting correctness vector of each sample, and calculating the SHAP value of each sample. The calculation method provided by the invention can reduce communication overhead and calculation complexity, improve calculation efficiency and is more suitable for a federated learning environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer software technology, and in particular to a method and system for calculating feature importance of vertical federated learning of a tree model based on TreeSHAP. Background Art

[0002] Federated Learning was proposed by Google in 2016, allowing multiple participants to collaborate on training models without sharing data. According to the distribution of samples or feature spaces among the participants, federated learning is divided into horizontal federated learning (HFL) and vertical federated learning (VFL). In horizontal federated learning, each participant has the same feature space but the sample space does not overlap; while in vertical federated learning, each participant has different features of the same batch of samples, and only one party has the sample labels, resulting in different sub-model structures of each participant, and the training process requires multi-party communication.

[0003] In vertical federated learning, evaluating the contribution of each participant is an important issue. Traditional horizontal federated learning can evaluate the contribution by the accuracy of the sub-model submitted by the participant, but in vertical federated learning, due to the different sub-model structures of each participant, it is impossible to directly use the accuracy for evaluation. Therefore, evaluating the contribution of the participants by feature importance becomes a feasible method.

[0004] SHAP (Shapley Additive Explanations) value is a comprehensive model interpretation method that evaluates the importance of features by calculating their marginal contribution to the model output. Lundberg proposed TreeSHAP based on SHAP, which greatly improved the calculation efficiency of SHAP value by using the structural characteristics of the tree model. TreeSHAP calculates the proportional coefficient through dynamic programming and recursive functions (Extend and Unwind), avoiding the high computational complexity of traversing all feature subsets in the traditional SHAP method.

[0005] In the prior art, TreeSHAP has been applied to multiple fields, such as the bank wealth management product recommendation model (patent publication number: CN111882441A) and the transformer health status assessment model (patent publication number: CN118427723A), to explain model decisions and improve the interpretability of the model.

[0006] Although TreeSHAP performs well in feature importance calculation of tree models, it has the following drawbacks in the vertical federated learning environment:

[0007] 1. High communication overhead: TreeSHAP calculates the SHAP value according to the sample dimension. In vertical federated learning, the distribution of samples at each node needs to be determined through communication between participants, which leads to a large amount of communication overhead and seriously affects the computing efficiency.

[0008] 2. High computational complexity: The Unwind function in TreeSHAP is complex to implement, involving multiple upward and downward recursive calls, and the time complexity is O(d 2 ) is computationally inefficient when the tree depth is large.

[0009] 3. Not suitable for vertical federated learning: The existing TreeSHAP method is not optimized for the vertical federated learning environment and cannot effectively deal with the problems of uneven distribution of participants’ characteristics and high communication overhead.

[0010] Therefore, there is an urgent need for a feature importance calculation method suitable for the vertical federated learning environment, which can reduce communication overhead and computational complexity while ensuring calculation accuracy. Summary of the invention

[0011] In view of this, an embodiment of the present invention provides a method and system for calculating feature importance of vertical federated learning of a tree model based on TreeSHAP to eliminate or improve one or more defects existing in the prior art.

[0012] On the one hand, the present invention provides a method for calculating feature importance of vertical federated learning of a tree model based on TreeSHAP, the method comprising the following steps:

[0013] Input the data set into each node of the tree model, calculate the sample ratio of each node, and record the diversion features used by the subtree with each node as the root;

[0014] A recursive method is used to traverse each node of the tree, and according to the diversion correctness of the sample and the sample proportion, the proportional coefficient array of each node is updated to obtain the proportional coefficient of each leaf node under different diversion correctness conditions;

[0015] The data set is recursively input into each node, a diversion correctness vector of each sample and a target feature diversion correctness vector are calculated according to the diversion feature of each node, the proportional coefficient is queried according to the diversion correctness vector of each sample, and a SHAP value of each sample is calculated.

[0016] In some embodiments of the present invention, a data set is input into each node of a tree model, and the sample ratio of each node is calculated, including:

[0017] Traverse each node of the tree model and count the number of samples passing through the corresponding node;

[0018] Then the sample ratio calculation formula for the left subtree and the right subtree is:

[0019]

[0020] Among them, r l Represents the sample ratio of the left subtree; r r represents the sample ratio of the right subtree; |N| represents the total number of samples passing through the current node; |N l | represents the number of samples of the left subtree; |N r | represents the number of samples of the right subtree.

[0021] In some embodiments of the present invention, the proportional coefficient array of each node is updated according to the correctness of the sample diversion and the sample ratio, and the updating formula of the proportional coefficient array is:

[0022] c i =c i-1 ×g+c i ×r;

[0023] Among them, c i represents the proportional coefficient when i features are introduced; g represents the correctness of the diversion; r represents the sample ratio.

[0024] In some embodiments of the present invention, after obtaining the proportionality coefficient of each leaf node under different diversion correctness conditions, the method further includes:

[0025] When the path traverses to the leaf node, the proportional coefficient is stored with the ID of the leaf node and the label of the diversion correctness as subscripts.

[0026] In some embodiments of the present invention, the method further comprises:

[0027] For a target feature, each node on a path is traversed twice according to the correct and incorrect diversions;

[0028] The storage space complexity is calculated based on the depth of the tree, the number of nodes on each path, and the label of the diversion correctness.

[0029] In some embodiments of the present invention, the data set is recursively input into each node, and the diversion correctness vector of each sample and the target feature diversion correctness vector are calculated according to the diversion feature of each node, including:

[0030] When traversing to an internal node, the target feature introduction coefficient is updated according to whether the parent node uses the target feature diversion;

[0031] Communicate with the participants to obtain a vector of the diversion of the data set sample at the current node;

[0032] The left and right subtrees are recursively traversed respectively, the diversion direction and the traversal direction are compared, and according to whether the current node uses the target feature diversion, the diversion correctness vector and the target feature diversion correctness vector are updated.

[0033] In some embodiments of the present invention, querying the proportionality coefficient according to the split correctness vector of each sample and calculating the SHAP value of each sample includes:

[0034] When traversing to the leaf node, the SHAP value of the target feature for all samples under the corresponding path is calculated according to the diversion correctness vector, the target feature diversion correctness vector, the target feature introduction coefficient and the proportional coefficient. The calculation formula is:

[0035] φ N,f =C[(N,G)]×(Tx)×v N ;

[0036] Wherein, f represents the target feature; C represents the proportional coefficient; N represents the number of nodes; G represents the diversion correctness vector; T represents the target feature diversion correctness vector; x represents the target feature introduction coefficient, which is used to indicate whether the target feature is introduced; v N Indicates the leaf node input value.

[0037] In some embodiments of the present invention, a pruning strategy is introduced into the method, including:

[0038] When calculating the target feature, if the target feature does not appear in the preset node after traversal, and the subtree of the preset node does not contain the target feature, the recursion stops;

[0039] If none of the trees in the tree model contains the target feature, the calculation of the tree is skipped.

[0040] In some embodiments of the present invention, the method further comprises:

[0041] When targeting a participant, consider all characteristics of that participant as a whole;

[0042] During the calculation of the proportional coefficient, when a feature belonging to the participant is encountered, the calculation of the proportional coefficient array is skipped; during the calculation of the SHAP value, the target feature diversion correctness vector and the target feature introduction coefficient are updated multiple times.

[0043] On the other hand, the present invention also provides a tree model longitudinal federated learning feature importance calculation system based on TreeSHAP, which, when executed, implements the steps of any of the methods mentioned above.

[0044] The present invention provides a method and system for calculating the importance of features of vertical federated learning of a tree model based on TreeSHAP, comprising: predicting a data set, calculating the sample ratio of each node in the tree model and recording the diversion features used by each subtree; traversing each node of the tree by a recursive method, updating the proportional coefficient array of each node according to the diversion correctness and sample ratio of the sample, and obtaining the proportional coefficient of each leaf node under different diversion correctness conditions; recursively inputting the data set into each node, calculating the diversion correctness vector of each sample and the target feature diversion correctness vector according to the diversion feature of each node, querying the proportional coefficient according to the diversion correctness vector of each sample, and calculating the SHAP value of each sample. The calculation process provided by the present invention is performed according to the dimension of the data set, and the diversion situation of each node only needs to be communicated by both parties once, thereby greatly reducing the overhead required for the communication process, and being more suitable for a federated learning environment. At the same time, the calculation process is modified to be two-step top-down, and the process of reversely adjusting the proportional coefficient in the Unwind function is avoided by traversing the features on the path and the possibility of all diversion correctness. At the same time, in the present invention, the overall time complexity is O(tld 2 2 d ), where t represents the number of trees, l represents the number of leaf nodes, and d represents the average tree depth. Since the average tree depth is a hyperparameter set manually during model training, and the average tree depth is often not set too large under the premise of controlling overfitting, that is, d 2 <<s, s represents the number of samples, which greatly improves the calculation efficiency.

[0045] Additional advantages, purposes, and features of the present invention will be described in part in the following description, and will become apparent to those skilled in the art after studying the following, or may be learned from the practice of the present invention. The purposes and other advantages of the present invention may be achieved and obtained by the structures specifically indicated in the specification and the accompanying drawings.

[0046] Those skilled in the art will appreciate that the objectives and advantages that can be achieved with the present invention are not limited to the above specific description, and the above and other objectives that can be achieved by the present invention will be more clearly understood from the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of the present application, and do not constitute a limitation of the present invention. In the drawings:

[0048] Figure 1 The figure is a schematic diagram of the steps of a method for calculating feature importance of vertical federated learning of a tree model based on TreeSHAP in one embodiment of the present invention.

[0049] Figure 2Schematic diagram of the Coeff process flow in one embodiment of the present invention.

[0050] Figure 3 FIG. 1 is a flow chart of a Divert process in one embodiment of the present invention.

[0051] Figure 4 The figure is a schematic diagram of an exemplary tree model structure in one embodiment of the present invention. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments and the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.

[0053] It should also be noted that, in order to avoid obscuring the present invention due to unnecessary details, only structures and / or processing steps closely related to the solutions according to the present invention are shown in the accompanying drawings, while other details that are not closely related to the present invention are omitted.

[0054] It should be emphasized that the term “include / comprises” when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.

[0055] It should also be noted that, unless otherwise specified, the term “connection” herein may refer not only to a direct connection but also to an indirect connection involving an intermediate.

[0056] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0057] It should be emphasized here that the step marks mentioned below are not intended to limit the order of the steps, but it should be understood that the steps can be executed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be executed simultaneously.

[0058] In order to solve the problems of large communication overhead, high computational complexity and inapplicability of the existing TreeSHAP technology in vertical federated learning, the present invention provides a method for calculating feature importance in vertical federated learning based on a tree model of TreeSHAP, such as Figure 1 As shown, the method includes the following steps S101 to S103:

[0059] Step S101: input the data set into each node of the tree model, calculate the sample ratio of each node, and record the diversion features used by the subtree with each node as the root.

[0060] Step S102: traverse each node of the tree recursively, update the proportional coefficient array of each node according to the diversion correctness of the sample and the sample proportion, and obtain the proportional coefficient of each leaf node under different diversion correctness conditions.

[0061] Step S103: recursively input the data set into each node, calculate the diversion correctness vector of each sample and the target feature diversion correctness vector according to the diversion feature of each node, query the proportionality coefficient according to the diversion correctness vector of each sample, and calculate the SHAP value of each sample.

[0062] In step S101, the data set is first predicted, the sample ratio of each node is calculated, and the diversion features used by each subtree are recorded.

[0063] In some embodiments, during the prediction process, each node of the tree model is traversed, the number of samples passing through each node is counted, and the sample ratio of the left subtree and the right subtree is calculated. The calculation formula is shown in formula (1):

[0064]

[0065] Among them, r l Represents the sample ratio of the left subtree; r r represents the sample ratio of the right subtree; |N| represents the total number of samples passing through the current node; |N l | represents the number of samples in the left subtree; |N r | represents the number of samples in the right subtree. The sample ratio can be used for subsequent proportional coefficient calculations to reflect the probability of sample diversion on the node.

[0066] At the same time, in the process of traversing the tree model, the diversion features used by each subtree are recorded. For each internal node, its diversion features and diversion threshold are recorded; for each subtree, all diversion features on the path from the root node to the leaf node are recorded. Recording diversion features can be used in subsequent pruning strategies to avoid repeated calculations of irrelevant features.

[0067] In step S102, the proportional coefficient of each leaf node under different diversion correctness conditions is calculated. In one embodiment of the present invention, this step is called the Coeff process, and its steps are as follows: Figure 2 shown.

[0068] In the existing TreeSHAP technology, the update method of the scale factor rolling array is shown in formula (2):

[0069] c i =c i-1 +c i ×r; (2)

[0070] Among them, c i It represents the proportional coefficient when i features are introduced; r represents the sample ratio.

[0071] At this time, the feature is introduced. If the sample is diverted correctly (i.e., diverted along the current calculation path), the diversion weight is 100%; if it is diverted incorrectly, that is, the sample is diverted into another path at this node, this part of the proportional coefficient will be adjusted to 0 during the Unwind process.

[0072] To address this problem, the present invention adds the correctness of the diversion at each step into the update formula of the proportional coefficient rolling array. If the diversion is correct, the correctness of the diversion is 1, otherwise it is 0. Therefore, the update formula of the proportional coefficient array is as follows:

[0073] c i =c i-1 ×g+c i ×r; (3)

[0074] Among them, c i It represents the proportional coefficient when i features are introduced; g represents the correctness of the diversion. If the diversion is correct, g = 1, otherwise g = 0; r represents the sample ratio.

[0075] Step S102 The Coeff process does not access the specific data set, but calculates in advance the SHAP value proportional coefficient corresponding to each leaf node under the possibility of all diversion correctness.

[0076] In some embodiments, Figure 2 As shown, when the path traverses to a leaf node, the proportional coefficient is stored with the leaf node ID and the diversion correctness label as the subscript. The leaf node ID is a unique number that identifies the current leaf node in the tree model; the diversion correctness label indicates whether the diversion of the sample on the current path is correct. Exemplarily, the diversion correctness label is a binary vector with a length of the number of nodes on the path, and each bit indicates whether the diversion of a node is correct (1 for correct, 0 for wrong).

[0077] In some embodiments, for a target feature, each node on a path is traversed twice according to the diversion correctness and diversion error, that is, each node has two possible diversion situations: correct or wrong. Assuming that there are l leaf nodes in the tree model and the average depth of the tree is d, the number of nodes on each path is d, and the number of possible combinations of diversion correctness labels is 2 d , the space complexity of storing the scale factor is O(l·2 d ).

[0078] In step S103, the data set is input into the tree model, the diversion correctness vector and the target feature diversion correctness vector of each sample are recursively calculated, and finally the SHAP value of each sample is calculated. In one embodiment of the present invention, this step is called the Divert process. In the Divert process, the data set is input into the node in a manner similar to the prediction process in step S101. The difference is that the features of some nodes will be missing when calculating the SHAP value, and all samples need to be input to calculate the weighted sum of the child nodes. Relatively speaking, samples that are diverted to other paths during the prediction process will not be input into the current path. The Divert process flow is as follows: Figure 3 shown.

[0079] In some embodiments, when traversing to an internal node, the target feature introduction coefficient is updated according to whether the parent node uses the target feature diversion. Specifically: If the parent node uses the target feature diversion, the sample ratio of the current node needs to be multiplied into the target feature introduction coefficient. The participating parties then communicate to obtain the diversion situation vector of the data set sample at the current node. Recursively traverse the left and right subtrees respectively, compare the diversion direction and the traversal direction, and update the diversion correctness vector and the target feature diversion correctness vector according to whether the current node uses the target feature diversion. Specifically: Whether the diversion direction and the traversal direction are the same is recorded as D, such as if the diversion direction and the traversal direction are the same, then D = 1, otherwise D = 0. If the current node uses the target feature diversion, use D to update the target feature diversion correctness vector, otherwise update the diversion correctness vector.

[0080] In some embodiments, when the recursive method traverses to a leaf node, the SHAP value of the target feature for all samples under the path is calculated according to the diversion correctness vector, the target feature diversion correctness vector, the target feature introduction coefficient and the proportional coefficient calculated in step S102. The calculation formula is shown in formula (4):

[0081] φ N,f =C[(N,G)]×(Tx)×v N ; (4)

[0082] Where f represents the target feature; C represents the scale factor; N represents the number of nodes; G represents the diversion correctness vector; T represents the target feature diversion correctness vector; x represents the target feature introduction coefficient, which is used to indicate whether the target feature is introduced; v N Indicates the leaf node input value.

[0083] In some embodiments, the target feature needs to be specified in step S102 (Coeff process) and step S103 (Divert process), so it is necessary to traverse in the feature dimension. Under this design, since the features that do not exist in the path will lead to useless repeated calculations, a pruning strategy is introduced. By calculating the sample ratio and diversion features in step S101, basic data is provided for the pruning strategy. Specifically, when calculating the target feature f, if no node with f as the diversion feature appears when traversing to node N, and the subtree of node N does not contain the target feature f, the recursive process will stop here. Furthermore, if none of the trees in the tree model contains the target feature f, the calculation of the tree will be skipped.

[0084] In some embodiments, for a vertical federated learning environment, communication overhead occurs in the process of calculating the diversion of a data set sample at the current node in step S103. At the same time, in order to measure the overall contribution of the participants rather than the importance of a single feature of the participant, in step S102 (Coeff process) and step S103 (Divert process), when a certain participant is targeted, all features of the participant will be regarded as the same feature. That is, in the Coeff process, when encountering features belonging to the participant, the calculation of the proportional coefficient array will be skipped, and in the Divert process, the target correctness label and the target feature introduction coefficient will be updated multiple times. Specifically, the sample needs to be correctly diverted in all the features belonging to the participant, and the target feature diversion correctness label is 1, that is, the label is updated using an AND operation.

[0085] In order to better understand the content of the present invention, the present invention is further described below in conjunction with a specific embodiment.

[0086] The following will demonstrate the calculation process of this method using an example tree model and data set.

[0087] The example dataset is shown in Table 1.

[0088] Table 1 Example dataset

[0089]

[0090] Example tree model Figure 4 As shown in the figure, the information of internal nodes includes the shunt feature and sample ratio, and the information of leaf nodes includes the node output value and sample ratio. For simplicity, the shunt threshold of all internal nodes is set to 0, that is, when the corresponding feature value of the sample is non-positive, it will be shunted to the left subtree, otherwise it will be shunted to the right subtree.

[0091] In this embodiment, the path A→C→D→v 4 , target feature D as an example to further illustrate the present invention.

[0092] During initialization, the proportional coefficient array C is set to [1, 0, 0, ...], which means that the proportional coefficient is 1 when 0 non-target features are introduced, the sample's shunt correctness vector is all 0, and the target feature shunt correctness vector is all 1.

[0093] Corresponding to step S102 (Coeff process):

[0094] Coeff will be calculated based on the correctness of the shunting corresponding to each feature value of sample 0. Since the feature values ​​of sample 0 at A, C, and D are negative, negative, and positive, respectively, the shunting will be to the left, left, and right at the corresponding nodes. 4 , right, left, and right), the diversion correctness vector G = [0,1,1].

[0095] Initially, the scale factor array and various parameters are shown in Table 2.

[0096] Table 2 Initial Coeff calculation table

[0097]

[0098] In each round of traversal, the proportional coefficient array will be updated in reverse order according to the update formula as shown in formula (3). 0 , when all features are missing, the calculation formula is as follows:

[0099] c 0 =c i ×r; (5)

[0100] After traversing to the D node, the proportional coefficient array is shown in Table 3.

[0101] Table 3 Coeff calculation table after accessing D node

[0102]

[0103]

[0104] In this embodiment, the target feature D is used for calculation, so Coeff will skip the diversion process based on the D node. 4 When , the binomial coefficient is calculated based on the number of introduced features and the number of path features, and the proportional coefficient array is weighted summed. The purpose of the binomial coefficient is to eliminate the deviation of the introduction of different numbers of features.

[0105] The final scale factor array and results are shown in Table 4.

[0106] Table 4 Final Coeff calculation table

[0107]

[0108] For better understanding, here is another example.

[0109] Similarly, the diversion correctness vector G = [1, 1, 0] can be calculated in the same way, and the target feature is the proportional coefficient of A, as shown in Table 5.

[0110] Table 5 Target feature ACoeff calculation table

[0111]

[0112] The diversion of the target features in the diversion correctness vector G will not affect the calculation of the proportional coefficient in the Coeff process. Therefore, part of the target features in the diversion correctness vector G is deleted, and the proportional coefficient of each combination can be stored with {leaf node ID, G} as the tuple subscript.

[0113] In actual operation, the deleted diversion correctness vector G is converted into a decimal number to reduce storage space.

[0114] The stored results are shown in Table 6. Among them, the deleted part of the diversion correctness vector G is bolded. To facilitate the calculation of the SHAP value in step S103 (Divert process), D is used as the target feature, and the proportional coefficient of the remaining diversion correctness is also calculated and displayed.

[0115] Table 6 Calculation results of target feature D proportional coefficient

[0116]

[0117] Corresponding to step S103 (Divert process):

[0118] During the Divert process, the data set will be recursively input into each node, and the sample diversion correctness vector and the target feature diversion correctness vector will be iteratively calculated based on the node's diversion features. When input into a leaf node, the proportional coefficient will be queried based on each label of each sample, and the impact of the current target feature on the output value of each sample will be calculated.

[0119] Next, follow the example path (A→C→D→v 4 ) shows the changes in different label values ​​of each traversal stage of the dataset sample:

[0120] Initially, as shown in Table 7, the diversion correctness label of all samples is 0, and the target feature diversion correctness label is 1. The former will be shifted and accumulated during the traversal process, and the latter means that the sample will be diverted according to the traversal direction at the target feature node. The target feature introduction coefficient is 1, which will be multiplied when traversing to the target feature.

[0121] Table 7 Initial Divert calculation table

[0122]

[0123] The data set is input to the root node and split by feature A. Samples 1 and 2 are split to the right subtree, in the same direction as the path; sample 0 is split in the wrong direction. The split correctness label is updated, as shown in Table 8.

[0124] Table 8 Update Divert calculation table

[0125]

[0126] Input to the right subtree and divert with feature C. Samples 0 and 1 are diverted to the left subtree in the same direction as the path; sample 2 is diverted in the wrong direction. Update the diversion correctness label, as shown in Table 9.

[0127] Table 9 Update Divert calculation table

[0128]

[0129] Input to the left subtree and divert with the target feature D. Samples 0 and 1 are diverted to the right subtree in the same direction as the path; sample 2 is diverted in the wrong direction. Update the target feature diversion correctness label, as shown in Table 10.

[0130] Table 10 Update Divert calculation table

[0131]

[0132] Input to leaf node L 4 , the parent node is the target feature, so the sample ratio of this node is multiplied into the target feature introduction coefficient. According to the shunt correctness vector of each sample, the ratio coefficient is queried, and the SHAP value of the target feature D for each sample is calculated, as shown in Table 11.

[0133] Table 11 Final Divert Calculation Table

[0134]

[0135] Corresponding to the above method, the present invention also provides a tree model longitudinal federated learning feature importance calculation system based on TreeSHAP, which, when executed, implements the steps of the tree model longitudinal federated learning feature importance calculation method based on TreeSHAP described above.

[0136] Corresponding to the above method, the present invention also provides an electronic device, which includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, the processor is used to execute the computer instructions stored in the memory, and when the computer instructions are executed by the processor, the electronic device implements the steps of the method described above.

[0137] The embodiment of the present invention also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps of the aforementioned edge computing server deployment method are implemented. The computer-readable storage medium can be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the technical field.

[0138] It should be understood by those skilled in the art that the exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software or a combination of the two. Whether it is performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier.

[0139] It should be clear that the present invention is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present invention.

[0140] In the present invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or replace features of other embodiments.

[0141] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the embodiments of the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for calculating feature importance of vertical federated learning in a tree model based on TreeSHAP, characterized in that: The method comprises the following steps: Input the data set into each node of the tree model, calculate the sample ratio of each node, and record the diversion features used by the subtree with each node as the root; A recursive method is used to traverse each node of the tree, and according to the diversion correctness of the sample and the sample proportion, the proportional coefficient array of each node is updated to obtain the proportional coefficient of each leaf node under different diversion correctness conditions; The data set is recursively input into each node, a diversion correctness vector of each sample and a target feature diversion correctness vector are calculated according to the diversion feature of each node, the proportional coefficient is queried according to the diversion correctness vector of each sample, and a SHAP value of each sample is calculated.

2. The method for calculating feature importance of vertical federated learning based on TreeSHAP in a tree model according to claim 1, characterized in that: Input the data set into each node of the tree model and calculate the sample proportion of each node, including: Traverse each node of the tree model and count the number of samples passing through the corresponding node; Then the sample ratio calculation formula for the left subtree and the right subtree is: Among them, r l Represents the sample ratio of the left subtree; r r represents the sample ratio of the right subtree; |N| represents the total number of samples passing through the current node; |N l | represents the number of samples of the left subtree; |N r | represents the number of samples of the right subtree.

3. The method for calculating feature importance of vertical federated learning based on TreeSHAP in a tree model according to claim 1, characterized in that: According to the correctness of the sample diversion and the sample ratio, the proportional coefficient array of each node is updated. The updating formula of the proportional coefficient array is: c i =c i-1 ×g+c i ×r; Among them, c i represents the proportional coefficient when i features are introduced; g represents the correctness of the diversion; r represents the sample ratio.

4. The method for calculating feature importance of vertical federated learning based on TreeSHAP in a tree model according to claim 3, characterized in that: After obtaining the proportional coefficient of each leaf node under different diversion correctness conditions, it also includes: When the path traverses to the leaf node, the proportional coefficient is stored with the ID of the leaf node and the label of the diversion correctness as subscripts.

5. The method for calculating feature importance of vertical federated learning based on TreeSHAP in a tree model according to claim 1, characterized in that: The method further comprises: For a target feature, each node on a path is traversed twice according to the correct and incorrect diversions; The storage space complexity is calculated based on the depth of the tree, the number of nodes on each path, and the label of the diversion correctness.

6. The method for calculating feature importance of vertical federated learning based on TreeSHAP in a tree model according to claim 1, characterized in that: The data set is recursively input into each node, and the diversion correctness vector of each sample and the target feature diversion correctness vector are calculated according to the diversion feature of each node, including: When traversing to an internal node, the target feature introduction coefficient is updated according to whether the parent node uses the target feature diversion; Communicate with the participants to obtain a vector of the diversion of the data set sample at the current node; The left and right subtrees are recursively traversed respectively, the diversion direction and the traversal direction are compared, and according to whether the current node uses the target feature diversion, the diversion correctness vector and the target feature diversion correctness vector are updated.

7. The method for calculating feature importance of vertical federated learning based on TreeSHAP in a tree model according to claim 1, characterized in that: The proportionality coefficient is queried according to the diversion correctness vector of each sample, and the SHAP value of each sample is calculated, including: When traversing to the leaf node, the SHAP value of the target feature for all samples under the corresponding path is calculated according to the diversion correctness vector, the target feature diversion correctness vector, the target feature introduction coefficient and the proportional coefficient. The calculation formula is: φ N,f =C[(N,G)]×(T-x)×v N ; Wherein, f represents the target feature; C represents the proportional coefficient; N represents the number of nodes; G represents the diversion correctness vector; T represents the target feature diversion correctness vector; x represents the target feature introduction coefficient, which is used to indicate whether the target feature is introduced; v N Indicates the leaf node input value.

8. The method for calculating feature importance of vertical federated learning based on TreeSHAP in a tree model according to claim 1, characterized in that: A pruning strategy is introduced into the method, including: When calculating the target feature, if the target feature does not appear in the preset node after traversal, and the subtree of the preset node does not contain the target feature, the recursion stops; If none of the trees in the tree model contains the target feature, the calculation of the tree is skipped.

9. The method for calculating feature importance of vertical federated learning based on TreeSHAP in a tree model according to claim 1, characterized in that: The method further comprises: When targeting a participant, consider all characteristics of that participant as a whole; During the calculation of the proportional coefficient, when a feature belonging to the participant is encountered, the calculation of the proportional coefficient array is skipped; during the calculation of the SHAP value, the target feature diversion correctness vector and the target feature introduction coefficient are updated multiple times.

10. A tree model longitudinal federated learning feature importance calculation system based on TreeSHAP, characterized in that: When the system is executed, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • User prediction interpretation Treecap method based on financial product recommendation scene

    CN111882441A

  • Transformer health state assessment method based on interpretable tree model

    CN118427723A