Feature evaluation method and device, computer readable storage medium, and electronic device
By filtering the features provided by the data provider in federated learning, establishing a decision tree and converting it into a feature network, and calculating the scores of feature nodes, the problems of low efficiency and poor accuracy of feature evaluation in the prior art are solved, and more efficient and accurate feature evaluation is achieved.
Patent Information
- Application Number
- CN202111217559.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-10-19
AI Technical Summary
The feature evaluation method in existing federated learning modeling considers a single dimension, resulting in low evaluation efficiency and poor accuracy.
By obtaining the features provided by the data provider, strongly related features are filtered out, a decision tree is established, and the decision tree is converted into a feature network, and the scores of feature nodes are calculated to complete feature evaluation.
Improve the accuracy and efficiency of feature evaluation, reduce modeling costs, and only one modeling is required to obtain scores for each feature.
Smart Images

Figure CN113935633B_ABST
Abstract
Description
Background Art
[0002] With the gradual development and application of emerging technologies such as artificial intelligence, user data security and privacy leakage issues have gradually attracted attention from all walks of life at home and abroad. Countries have introduced corresponding laws and regulations to limit the scope of use of user data, thus forming data islands.
[0003] In order to solve the problem of data silos and ensure the data security of all parties, federated learning has emerged to achieve collaborative modeling based on multi-party data. However, in actual business implementation, how to obtain appropriate features from massive feature sets for modeling has become a research focus.
[0004] Although there are feature evaluation methods in existing federated learning modeling, the feature evaluation scheme only considers a single dimension, which not only reduces the efficiency of the evaluation, but also affects the accuracy of the feature evaluation.
[0005] Therefore, a new feature evaluation method needs to be provided.
[0006] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0007] The purpose of the present disclosure is to provide a feature evaluation method, a feature evaluation device, a computer-readable storage medium and an electronic device, thereby at least to a certain extent overcoming the problems of low feature evaluation efficiency and poor evaluation accuracy caused by the limitations and defects of related technologies.
[0008] According to one aspect of the present disclosure, there is provided a feature evaluation method, comprising:
[0009] Acquire features provided by a data provider, obtain strongly correlated features included in the features, and filter the strongly correlated features to obtain valid features;
[0010] Establish a decision tree based on the effective features, convert the current sub-decision tree into a current feature network, and determine the scores of the feature nodes included in the current feature network;
[0011] When it is determined that the feature node exists in a non-current sub-decision tree included in the decision tree, the score of the feature node in the decision tree is obtained according to the score of the feature node to complete the evaluation of the features included in the decision tree.
[0012] In an exemplary embodiment of the present disclosure, the features provided by the data provider are acquired, the strongly correlated features included in the features are obtained, and the strongly correlated features are filtered to obtain valid features, including:
[0013] Calculating the correlation between different features among the features provided by the data provider;
[0014] When the correlation between multiple features is strongly correlated, the information value of the multiple features is calculated, and the multiple features are filtered according to the information value to obtain the feature with the highest information value among the strongly correlated features;
[0015] The effective features of the data provider are obtained according to the features with higher information value among the strongly correlated features.
[0016] In an exemplary embodiment of the present disclosure, establishing a decision tree based on the effective features includes:
[0017] Acquire the valid features, perform anonymization on the feature names of the valid features, and establish a first sub-decision tree based on the feature names after the anonymization;
[0018] Calculating the residual of the first sub-decision tree, and using the residual of the first sub-decision tree to establish a subsequent sub-decision tree;
[0019] The iteration is continued until the number of iterations reaches a preset number of iterations or the residual of the sub-decision tree is less than a preset threshold, thereby obtaining a decision tree established based on the effective features.
[0020] In an exemplary embodiment of the present disclosure, establishing a decision tree based on the effective features includes:
[0021] Acquire sample data corresponding to the feature provided by the data, calculate the information gain of the feature according to the sample data, and use the feature with the largest information gain as the root node of the first sub-decision tree;
[0022] Classify the sample data according to the root node to obtain first sample data and second sample data;
[0023] According to the first sample data and the second sample data, the first sub-decision tree is calculated and constructed according to the information gain of the feature corresponding to the first sample data and the information gain corresponding to the second sample data.
[0024] In an exemplary embodiment of the present disclosure, converting a current sub-decision tree into a current feature network, and determining scores of feature nodes included in the current feature network, includes:
[0025] Obtaining a current sub-decision tree, using a split node in the current sub-decision tree as a feature node in the feature network, and converting a directional relationship between the split nodes into a directed edge in the feature network;
[0026] The current feature network is established according to the feature nodes and the directed edges, and the scores of the feature nodes included in the current feature network are determined according to the information gains of the feature nodes in the current feature network.
[0027] In an exemplary embodiment of the present disclosure, determining the scores of the feature nodes included in the current feature network according to the information gains of the feature nodes in the current feature network includes:
[0028] Normalizing the information gain of the feature node to obtain a standardized information gain, and obtaining the weight of the directed edge through the standardized information gain;
[0029] The score of the feature node is iteratively calculated according to the weight of the directed edge until the score of the feature node converges to obtain the score of the feature node.
[0030] In an exemplary embodiment of the present disclosure, obtaining the score of the feature node in the decision tree according to the score of the feature node includes:
[0031] Calculating a first score of a feature node included in the current feature network in a feature network corresponding to the non-current sub-decision tree;
[0032] According to the first score, the score of the feature node included in the current network in the decision tree is obtained.
[0033] According to one aspect of the present disclosure, there is provided a feature evaluation device, comprising:
[0034] An effective feature acquisition module is used to acquire features provided by a data provider, obtain strongly correlated features included in the features, and filter the strongly correlated features to obtain effective features;
[0035] A feature node score calculation module, used to establish a decision tree based on the effective features, convert the current sub-decision tree into a current feature network, and determine the scores of the feature nodes included in the current feature network;
[0036] The feature evaluation module is used to obtain the score of the feature node in the decision tree according to the score of the feature node when it is determined that the feature node exists in a non-current sub-decision tree included in the decision tree, so as to complete the evaluation of the features included in the decision tree.
[0037] According to one aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the feature evaluation method described in any of the above exemplary embodiments is implemented.
[0038] According to one aspect of the present disclosure, there is provided an electronic device, including:
[0039] Processor; and
[0040] A memory, configured to store executable instructions of the processor;
[0041] The processor is configured to perform the feature evaluation method described in any of the above exemplary embodiments by executing the executable instructions.
[0042] The embodiment of the present disclosure provides a feature evaluation method, which obtains features provided by a data provider, obtains strongly correlated features included in the features, filters the strongly correlated features to obtain valid features; establishes a decision tree based on the valid features, converts the current sub-decision tree into a current feature network, and determines the scores of feature nodes included in the current feature network; when it is determined that the feature node exists in a non-current sub-decision tree included in the decision tree, obtains the score of the feature node in the decision tree according to the score of the feature node, so as to complete the evaluation of the features included in the decision tree; on the one hand, since the invalid features included in the features provided by the data provider are first filtered to obtain valid features, a decision tree is established based on the valid features, and the decision tree is converted into a feature network, and each feature node is scored based on the feature network, the problem of considering a single dimension when evaluating features in the prior art is solved, and the accuracy of feature evaluation is improved; on the other hand, converting the current sub-decision tree into a feature network and calculating the scores of feature nodes can be performed in parallel, and only one modeling is required to obtain the scores of each feature. Compared with the communication cost problem caused by multiple modeling in the prior art, the present disclosure only requires one modeling, which not only reduces the cost consumption of modeling, but also improves the efficiency of feature evaluation.
[0043] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.
[0045] Figure 1 The flowchart of a feature evaluation method that can be applied to the embodiment of the present disclosure is schematically shown.
[0046] Figure 2A block diagram of a feature evaluation system according to an exemplary embodiment of the present disclosure is schematically shown.
[0047] Figure 3 A flowchart of a method for filtering strongly correlated features to obtain effective features according to an exemplary embodiment of the present disclosure is schematically shown.
[0048] Figure 4 A flowchart of a method for establishing a decision tree based on effective features according to an exemplary embodiment of the present disclosure is schematically shown.
[0049] Figure 5 A schematic diagram schematically illustrates a sub-decision tree according to an exemplary embodiment of the present disclosure.
[0050] Figure 6 A flowchart of a method for establishing a decision tree based on effective features according to an exemplary embodiment of the present disclosure is schematically shown.
[0051] Figure 7 A flowchart of a method for converting a current sub-decision tree into a current feature network and determining scores of feature nodes included in the current feature network according to an exemplary embodiment of the present disclosure is schematically shown.
[0052] Figure 8 A flowchart of a method for determining scores of feature nodes included in a current feature network according to gain information of feature nodes in the current feature network according to an exemplary embodiment of the present disclosure is schematically shown.
[0053] Fig. 9 A schematic diagram schematically illustrates a scenario of converting a sub-decision tree into a feature network according to an exemplary embodiment of the present disclosure.
[0054] Fig.10 A flowchart of a method for obtaining a score of a feature node in a decision tree according to the score of the feature node according to an exemplary embodiment of the present disclosure is schematically shown.
[0055] Fig.11 The flowchart schematically shows a feature evaluation method according to an exemplary embodiment of the present disclosure.
[0056] Fig.12 A block diagram schematically shows a feature evaluation device according to an exemplary embodiment of the present disclosure.
[0057] Fig.13 An electronic device for implementing the above-mentioned feature evaluation method according to an exemplary embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0058] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as being limited to the examples set forth herein; on the contrary, these embodiments are provided so that the present disclosure will be more comprehensive and complete, and the concepts of the example embodiments are fully conveyed to those skilled in the art. The described features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present disclosure.
[0059] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0060] With the gradual development and application of emerging technologies such as artificial intelligence, the data security and privacy leakage of users have gradually attracted the attention of all sectors of society at home and abroad. Therefore, governments of various countries have introduced relevant laws and regulations to restrict enterprises from releasing user data, and clearly stipulate that the ownership of these data belongs to the users, and the enterprises that provide services to users only have the right to use these data, and cannot share or trade user data with others, which may expose user privacy. Under such strict legal restrictions, the phenomenon of data islands has been formed. In order to solve the problem of "data islands" and ensure the data security of all parties, federated learning has emerged. It is based on differential privacy, homomorphic encryption and other confidentiality technologies, and is committed to protecting the privacy of user data and realizing the joint modeling of multi-party data, thereby giving full play to the value of social data resources and strengthening orderly data sharing.
[0061] Although many scholars and companies have been conducting research in the field of federated learning in recent years, many of these studies are aimed at how to convert traditional modeling algorithms for safe use in federated scenarios. Some studies are also aimed at the problem that processes such as communication and encryption in federated learning consume too much time, resulting in low efficiency in federated modeling. However, in actual business implementation, there is also a key issue: how to evaluate whether the features provided by each participant are valuable. This is because the number of features involved in modeling in federated learning is not the more the better. Some redundant features will not only increase the cost of modeling, such as storage, modeling time, and usage fees, but also affect the overall accuracy and reduce the overall effect of the model. Therefore, how to evaluate the importance of the features provided by each participant to the final model, and select the appropriate feature set to enter the subsequent modeling stage, so as to ensure the model effect while reducing the cost of modeling.
[0062] Although there are many classic solutions for feature evaluation in the traditional machine learning field, such as SHAP (SHapley Additive exPlanation, explaining machine learning model output) and LIME (Local Interpretable Model-Agnostic Explanations), these solutions do not meet the security requirements of federated learning and may expose the data of participants during the feature evaluation process. Therefore, they are not suitable for federated scenarios.
[0063] Even though feature evaluation exists in existing federated learning modeling, the existing feature evaluation scheme has a single evaluation dimension and does not comprehensively measure the comprehensive contribution of different features to the model from multiple perspectives, which reduces the efficiency of feature evaluation and the accuracy of evaluation results.
[0064] Based on one or more of the above problems, this exemplary embodiment first provides a feature evaluation method, which can be run on a server, a server cluster or a cloud server, etc. Of course, those skilled in the art can also run the method of the present invention on other platforms as required, and this exemplary embodiment does not specifically limit this. Figure 1 As shown, the feature evaluation method may include the following steps:
[0065] Step S110. Acquire features provided by a data provider, obtain strongly correlated features included in the features, filter the strongly correlated features, and obtain valid features;
[0066] Step S120. Establish a decision tree based on the effective features, convert the current sub-decision tree into a current feature network, and determine the scores of the feature nodes included in the current feature network;
[0067] Step S130. When it is determined that the feature node exists in a non-current sub-decision tree included in the decision tree, the score of the feature node in the decision tree is obtained according to the score of the feature node to complete the evaluation of the features included in the decision tree.
[0068] The feature evaluation method described above obtains the features provided by the data provider, obtains the strongly correlated features included in the features, filters the strongly correlated features, and obtains valid features; establishes a decision tree based on the valid features, converts the current sub-decision tree into the current feature network, and determines the scores of the feature nodes included in the current feature network; when it is determined that the feature node exists in the non-current sub-decision tree included in the decision tree, the score of the feature node in the decision tree is obtained according to the score of the feature node, so as to complete the evaluation of the features included in the decision tree; on the one hand, since the invalid features included in the features provided by the data provider are first filtered to obtain valid features, a decision tree is established based on the valid features, and the decision tree is converted into a feature network, and each feature node is scored based on the feature network, the problem of considering a single dimension when evaluating features in the prior art is solved, and the accuracy of feature evaluation is improved; on the other hand, converting the current sub-decision tree into a feature network and calculating the scores of feature nodes can be performed in parallel, and only one modeling is required to obtain the scores of each feature. Compared with the communication cost problem caused by multiple modeling in the prior art, the present disclosure only requires one modeling, which not only reduces the cost consumption of modeling, but also improves the efficiency of feature evaluation.
[0069] Hereinafter, each step involved in the feature evaluation method according to the exemplary embodiment of the present disclosure is explained and illustrated in detail.
[0070] First, the application scenario and purpose of the example embodiment of the present disclosure are explained and illustrated. Specifically, the example embodiment of the present disclosure can be applied to evaluate features from multiple dimensions in federated modeling. In the present disclosure, first, the features provided by each data provider are obtained, and the provided features are screened to obtain valid features; then a decision tree is established using the valid features, and the decision tree is converted into a feature network, and the score of each feature node is calculated through the feature network to complete the evaluation of the features, thereby achieving multi-angle evaluation of the features and improving the accuracy of the feature evaluation.
[0071] Next, the feature evaluation system involved in the exemplary embodiment of the present disclosure is explained and described. Figure 2As shown, the task processing system may include an invalid feature filtering module 210, a multi-party federated modeling module 220, a weighted feature network generation module 230, a feature contribution quantification module 240, and a weighted feature evaluation module 250. Among them, the invalid feature filtering module 210 is used to calculate the correlation between the features after obtaining the features provided by each data provider, obtain the strongly correlated features, and calculate the IV (Information Value) value of the strongly correlated features, retain the features with the highest IV value, and obtain the valid features of the data provider; the multi-party federated modeling module 220 is connected to the invalid feature filtering module 210 network, and is used to anonymize the valid features provided by the data provider, and establish a decision tree based on the anonymously processed features. Each decision tree is established based on the residual of the previous decision tree, and each split node in the decision tree is the feature with the largest information gain in the current sample data; the weighted feature network generation module 230 is connected to the multi-party federated modeling module 220 network, and is used to use the split nodes in the decision tree as feature nodes in the feature network, and the split nodes are The directed relationship is converted into a directed edge between feature nodes in the feature network, the information gain of the feature node is obtained, and the weight of the directed edge is calculated based on the information gain of the feature node; the feature contribution quantification module 240 is used to obtain the score of the feature node in the feature network by performing multiple rounds of iterations on the weight of the directed edge; the weighted feature evaluation module 250 is connected to the feature contribution quantification module 240 network, and is used to determine the non-current feature network containing the feature node, calculate the score of the feature node in the non-current feature network, and obtain the score of the feature network in the decision tree based on the score of the feature node in the non-current feature network.
[0072] The following will be combined Figure 2 Steps S110 to S130 are explained and illustrated in detail.
[0073] In step S110, the features provided by the data provider are acquired, the strongly correlated features included in the features are obtained, and the strongly correlated features are filtered to obtain valid features.
[0074] Among them, the data provider is a plurality of data providers, each of which can provide one or more features. In this example embodiment, there is no specific limitation on the number of features provided by the data provider. In this example embodiment, after obtaining the features provided by the data provider, in order to avoid invalid features affecting the evaluation effect and efficiency of other features, it is necessary to delete the invalid features before performing feature evaluation. An invalid feature is a feature that has low discrimination for sample data. When there is a null value in the sample data under a certain feature, it is possible to determine whether the feature is an invalid feature by calculating the coverage of the sample data under the feature. When there is the same value in the sample data under a certain feature, it is possible to determine whether the feature is an invalid feature by calculating the variance of the sample data under the feature. In addition to considering the validity of the feature itself, it is also necessary to consider the correlation between features.
[0075] In this example embodiment, reference Figure 3 As shown, acquiring the features provided by the data provider, obtaining the strongly correlated features included in the features, filtering the strongly correlated features, and obtaining valid features may include steps S310 to S330:
[0076] Step S310. Calculate the correlation between different features among the features provided by the data provider;
[0077] Step S320. When the correlation between the multiple features is strongly correlated, the information value of the multiple features is calculated, and the multiple features are filtered according to the information value to obtain the feature with the highest information value among the strongly correlated features;
[0078] Step S330: Obtain the effective features of the data provider according to the features with the highest information value among the strongly correlated features.
[0079] Steps S310 to S330 are explained and illustrated below. Specifically, after obtaining the features provided by the data provider and the sample data corresponding to the features, first, calculate the coverage of the sample data corresponding to the features or the variance of the sample data corresponding to the features. When the coverage of the sample data does not meet the preset threshold or the variance is close to 0, the feature is identified as an invalid feature and deleted; when the coverage of the sample data meets the preset threshold or the variance is greater than 0, calculate the correlation between different features in the features provided by the data provider. First, calculate the correlation between the features based on the sample data corresponding to the features, wherein the correlation between the features can be determined by calculating the Pearson correlation coefficient. In the Pearson coefficient, the larger the absolute value of the correlation coefficient, the stronger the correlation. When the correlation coefficient between the features is greater than 0.6, it can be considered that the correlation between the features is strongly correlated. When the correlation between the features is strongly correlated, the IV value of the feature can be calculated, and the strongly correlated features can be filtered by the IV value of the feature. Specifically, the feature with the highest IV value among the strongly correlated features can be retained as a valid feature.
[0080] For example, the features provided by the data provider include six features: A, B, C, D, E, and F. Among them, the sample data corresponding to feature A includes null values, but the coverage rate of the sample data corresponding to feature A is 0.95, which is greater than the preset threshold value of 0.8. Therefore, feature A can be considered as a valid feature; the preset threshold value can be 0.8 or 0.7, and the preset threshold value is not specifically limited in this example embodiment; features B, C, and D are strongly correlated features. By calculating the IV values of features B, C, and D, when the IV value of feature B is the highest, features C and feature D can be deleted, and feature B can be used as a valid feature; the sample data corresponding to features E and F include neither null values nor identical values. Therefore, features E and F can also be identified as valid features. The valid features included in the data provided by the data provider are: four features: A, B, E, and F.
[0081] In step S120, a decision tree is established based on the effective features, the current sub-decision tree is converted into a current feature network, and the scores of the feature nodes included in the current feature network are determined.
[0082] In this example embodiment, after removing invalid features from the features provided by the data provider, modeling can be performed using valid features based on vertical federated learning. During the modeling process, valid features are all embodied in an anonymous manner, and the interaction process is encrypted and protected in a homomorphic encryption manner to avoid leakage of sample data corresponding to the features during the modeling process. The vertical federated learning method can be a federated random forest or a federated gradient boosting model, which is not specifically limited in this example embodiment.
[0083] In this example embodiment, reference Figure 4 As shown, establishing a decision tree based on effective features may include steps S410 to S430:
[0084] Step S410: Acquire the valid features, anonymize the feature names of the valid features, and establish a first sub-decision tree based on the feature names after the anonymization process;
[0085] Step S420. Calculate the residual of the first sub-decision tree, and use the residual of the first sub-decision tree to establish a subsequent sub-decision tree;
[0086] Step S430. Continue iterating until the number of iterations reaches a preset number of iterations or the residual of the sub-decision tree is less than a preset threshold, and obtain a decision tree established based on the effective features.
[0087] Steps S410 to S430 will be explained and illustrated below. Specifically, first, valid data provided by the data provider is obtained. In order to ensure the security of the features provided by each data provider and the sample data corresponding to the features during the modeling process, the feature names of the valid features can be anonymized. For example, the data provider's number and the number of the feature in the features provided by the data provider can be used to represent it. Other numbers can also be used to represent the valid features. In this example embodiment, there is no specific limitation on the anonymization of the valid features. The first sub-decision tree can be constructed based on the anonymously processed feature name. When the vertical federated learning method is a federated gradient boosting model, the model is composed of multiple sub-decision trees, and each sub-decision tree is in a serial relationship. And the next sub-decision tree can be established based on the residual of the previous sub-decision tree. Therefore, after obtaining the first sub-decision tree, the residual of the first sub-decision tree can be calculated, and the second sub-decision tree can be established based on the residual of the first sub-decision tree. The sub-decision trees are established by continuous iteration until the number of iterations reaches the preset number of iterations or the residual of the last sub-decision tree is less than the preset threshold. Based on this, a decision tree established by effective features can be obtained, wherein the decision tree includes multiple sub-decision trees, and the preset number of iterations can be 5 times or 10 times, which is not specifically limited in this example embodiment; the preset threshold can be 0.1 or 0.2, which is not specifically limited in this example embodiment.
[0088] Refer to the schematic diagram of the sub-decision tree shown in 5, wherein the box 501 represents a leaf node, wherein the leaf node can be represented by leaf, and the circle 502 represents the split node where the valid features provided by different data providers are located. In each split node, P represents the data provider, and F represents the feature, wherein the first feature of the first data provider can be represented as: P1F1, and the fourth feature of the second data provider can be represented as: P2F4.
[0089] In this exemplary embodiment, in the process of establishing a sub-decision tree, in addition to establishing a subsequent sub-decision tree based on the residual of the previous sub-decision tree, when establishing the current sub-decision tree, it is also necessary to consider the information gain of different features, refer to Figure 6 As shown, establishing a decision tree based on effective features includes steps S610 to S630:
[0090] S610. Obtain sample data corresponding to the feature provided by the data, calculate the information gain of the feature according to the sample data, and use the feature with the largest information gain as the root node of the first sub-decision tree;
[0091] S620. Classify the sample data according to the root node to obtain first sample data and second sample data;
[0092] S630. Based on the first sample data and the second sample data, calculate and construct the first sub-decision tree based on the information gain of the feature corresponding to the first sample data and the information gain corresponding to the second sample data.
[0093] Steps S610 to S630 will be explained and illustrated below. Specifically: First, sample data corresponding to effective features are obtained, and the information gain of effective features is calculated based on the sample data, and the split node where the feature with the largest information gain among the effective features is located is used as the root node of the first sub-decision tree; after obtaining the root node, a classification condition corresponding to the root node can be generated, and the sample data can be classified according to the classification condition to obtain first sample data that meets the classification condition and second sample data that does not meet the classification condition; after obtaining the first sample data and the second sample data, the information gain of the feature corresponding to the first sample data can be calculated based on the first sample data, and the information gain of the feature corresponding to the second sample data can be calculated based on the second sample data, and the feature with the largest information gain can be selected as the two child nodes of the root node, and the above process is repeated until the first sub-decision tree is constructed.
[0094] For example, when the feature with the highest information gain is age, the age feature can be anonymized and the anonymized age feature can be used as the root node of the first sub-decision tree; then, the classification condition corresponding to the age is determined, wherein the classification condition can be whether the age is over 18 years old. According to the classification condition, the sample data can be divided into first sample data not exceeding 18 years old and second sample data over 18 years old. For the first sample data, the information gain of the feature corresponding to the first sample data can be calculated, and the feature with the largest information gain is selected as the first node of the second layer of the first sub-decision tree, wherein the feature with the largest information gain in the first sample data may be age or other features; the processing method of the second sample data is the same as that of the first sample data, and the second node of the second layer of the first sub-decision tree is obtained; then the first node of the second layer of the first sub-decision tree and the classification condition corresponding to the second node are determined, and the above steps are repeated to obtain the first sub-decision tree.
[0095] After obtaining the decision tree, the decision tree can be converted into a feature network. Figure 7 As shown, converting the current sub-decision tree into the current feature network and determining the scores of the feature nodes included in the current feature network may include step S710 and step S720:
[0096] Step S710. Obtain the current sub-decision tree, use the split nodes in the current sub-decision tree as feature nodes in the feature network, and convert the pointing relationship between the split nodes into directed edges in the feature network;
[0097] Step S720: Establish the current feature network according to the feature nodes and the directed edges, and determine the scores of the feature nodes included in the current feature network according to the information gains of the feature nodes in the current feature network.
[0098] Step S710 and step S720 are explained and illustrated below. Specifically, when converting a sub-decision tree into a feature network, the split nodes in the sub-decision tree can be used as feature nodes of the feature network, and the pointing relationship between the split nodes can be converted into directed edges between the feature nodes, wherein the weight of each feature node in the feature network is the number of times the feature node is used in the feature network; the weight of the directed edge in the feature network can be determined based on the information gain of the two feature nodes connected to the directed edge, and after obtaining the feature network corresponding to the current sub-decision tree, the multiple sub-decision trees included in the decision tree can be converted in the same manner as the current sub-decision tree is converted, and the obtained feature network can be numbered according to the order of the sub-decision trees in the decision tree.
[0099] Furthermore, after converting the sub-decision tree into a feature network, each feature node in the feature network can be evaluated. A variety of network node sorting algorithms can be used in the evaluation process. In this example embodiment, the network node sorting algorithm is not specifically limited. When the network sorting algorithm is the classic PageRank algorithm, in traditional social networks, the greater the out-degree of a user, the stronger the communication ability of the user. When calculating, the receiver's score should be distributed to different senders according to the weight in the opposite direction of the directed edge. Similarly, in the feature network, the greater the out-degree of a feature node, that is, the more times it is a parent node, the more important the feature node is.
[0100] Specifically, when evaluating a feature node, it can be evaluated through the feature quantity hypothesis and the feature quality hypothesis, where the feature quantity hypothesis is: the larger the out-degree of a feature node, the higher the score of the feature node, which can be understood as the more times the feature node is used, that is, the closer the feature node is to the root node, the more important it is, because the out-degree of the feature node close to the root node is greater than the feature node close to the leaf node. In extreme cases, the leaf node, that is, the non-feature node, has no out-degree, and the score of the node is lower; the feature quality hypothesis is: the larger the in-degree of a feature node, the higher the score of the feature node.
[0101] refer to Figure 8 As shown, determining the scores of the feature nodes included in the current feature network according to the gain information of the feature nodes in the current feature network may include step S810 and step S820:
[0102] Step S810: normalizing the information gain of the feature node to obtain a normalized information gain, and obtaining the weight of the directed edge through the normalized information gain;
[0103] Step S820: Iteratively calculate the score of the feature node according to the weight of the directed edge until the score of the feature node converges to obtain the score of the feature node.
[0104] Step S810 and step S820 are explained and illustrated below. Specifically, when calculating the score of a feature node, on the one hand, the weight of the feature node, that is, the number of times the feature node is used in the feature network, can be considered. On the other hand, the score of the feature node can be calculated by the feature quality hypothesis. When calculating according to the feature quality hypothesis, first, the information gain of the two feature nodes connected to the directed edge is standardized to obtain the standardized information gain; then, the weight of the directed edge is obtained by summing the standardized information gains of the two feature nodes; then, the score of any feature node in the current feature network in any round of iteration is calculated according to the weight of the directed edge, and the iteration is continued until the score of any feature node converges.
[0105] The information gain of the feature node can be standardized according to expression (1):
[0106]
[0107] Among them, χ * is the standardized value of the information gain χ of the feature node, max is the maximum value of information gain in the current feature network, χ min It is the value with the minimum information gain in the current feature network.
[0108] The score of the feature node in the k+1th round can be obtained by expression (2):
[0109]
[0110] Among them, α is the damping coefficient of PageRank, usually 0.85, The weights for assigning the score of feature node F1 to feature nodes F2, F3, and F4 are respectively: are the scores of feature nodes F1, F2, F3 and F4 in the kth round respectively.
[0111] After obtaining the score of the feature node in the k+1th round, the score of each node can be stored to determine whether the score of the feature node obtained in the k+1th round is converged. Specifically, the determination can be based on expression (3):
[0112]
[0113] Among them, it can be Store the scores of each feature node in the current feature network in round k, T is the state transfer matrix of the current feature network, which can be determined based on the directed edges between different feature nodes in the feature network.
[0114] refer to Fig. 9 The sub-decision tree and the feature network corresponding to the sub-decision tree are shown, wherein the sub-decision tree includes split nodes: F2, F4, F1, F3, F3 and F4, and the information gain corresponding to each split node is 3, 2, 4, 1, 2, 3. After standardizing the information gain of the split node, the standardized information gain of the split node can be obtained as follows: 1, 0, After converting the sub-decision tree into a feature network, the weight of the directed edge between feature node F2 and feature node F4 is The weight of the directed edge between feature node F1 and feature node F2 is The weight of the directed edge between feature node F1 and feature node F3 is The weight of the directed edge between feature node F1 and feature node F4 is The weight of the directed edge between feature node F4 and feature node F3 is Since the transfer relationship between split nodes in the child decision tree is to assign the weight of the child split node to the parent split node associated with the child split node, the feature node F1 assigns its score to Assigned to feature node F2,; feature node F3 assigns its score Assigned to feature node F1, and its score Assigned to feature node F4; feature node F4 assigns its score Assigned to feature node F1, and its score Assigned to feature node F2. Therefore, we can get the score of feature node F1 at the k+1th iteration:
[0115]
[0116] The matrix of feature nodes in round k+1 can be expressed as:
[0117]
[0118] Among them, the score of the feature node can be assigned to the weight of any feature node according to its score, and the state transfer matrix corresponding to the score of the feature node can be obtained:
[0119]
[0120] By comparing the matrix of the feature node in the k+1th round with the matrix of the feature node in the kth round, it is determined whether the score of the feature node in the current k+1th round has converged. When converged, the score of the feature node in the k+1th round can be used as the score of each feature node in the current feature network.
[0121] In step S130, when it is determined that the feature node exists in a non-current sub-decision tree included in the decision tree, the score of the feature node in the decision tree is obtained according to the score of the feature node to complete the evaluation of the features included in the decision tree.
[0122] In this example embodiment, the scores of the feature nodes included in the feature networks corresponding to all sub-decision trees included in the decision tree can be calculated according to the method of step S120. After calculating the score of each feature node in each feature network,.
[0123] In this example embodiment, reference Fig.10As shown, obtaining the score of the feature node in the decision tree according to the score of the feature node may include step S1010 and step S1020:
[0124] Step S1010. Calculate a first score of a feature node included in the current feature network in a feature network corresponding to the non-current sub-decision tree;
[0125] Step S1020: According to the first score, obtain the score of the feature node included in the current network in the decision tree.
[0126] Step S1010 and step S1020 are explained and illustrated below. Specifically, after calculating the scores of the feature nodes included in the feature networks corresponding to all sub-decision trees included in the decision tree, since a single feature node may exist in different feature networks, it is necessary to perform a weighted sum of the scores of the feature node in multiple feature networks to obtain the score of the feature node in the decision tree, that is, first, calculate the first score of the feature node included in the current feature network in the feature network corresponding to the non-current sub-decision tree; then, perform a weighted sum of the first score to obtain the score of the feature node in the decision tree.
[0127] For example, when there are n sub-decision trees in the decision tree, the standardized score of the j-th feature of data provider i in the k-th sub-decision tree can be recorded as x ijk , and the proportion of the improvement of the decision tree effect by the kth child decision tree to the entire decision tree effect is ρ i , so the score of the jth feature of data provider i in the decision tree is X ij It can be expressed as expression (4):
[0128]
[0129] After obtaining the score of the feature node in the decision tree, the features can be screened again according to the score of the feature node to optimize the model, or the benefits can be distributed to the data provider who provides the feature node according to the score of the feature node in the decision tree.
[0130] The feature evaluation method provided by the example embodiments of the present disclosure has at least the following advantages: on the one hand, compared with the traditional feature evaluation mechanism, the feature evaluation method can meet the security requirements of federated modeling through mechanisms such as anonymity and encryption, and the entire evaluation process does not require multiple modeling, nor does it require multiple analyses that affect the efficiency and security of federated modeling; on the other hand, compared with the split and gain schemes in the existing gradient boosting tree model, the feature evaluation method not only considers the number of times the feature is used in each sub-decision tree and the information gain, but also considers the order in which the feature nodes appear. The closer the feature node is to the root node, the more important it is, that is, the random walk model taking the PageRank algorithm as an example iterates the scores of the nodes in the network structure, thereby improving the accuracy of feature evaluation.
[0131] The following, combined Fig.11 The feature evaluation method of the exemplary embodiment of the present disclosure is further explained and illustrated. The feature evaluation method may include:
[0132] Step S1102. Acquire features provided by multiple data providers and sample data corresponding to the features;
[0133] Step S1104. When it is determined that the coverage of the sample data corresponding to the feature is higher than a preset threshold and the variance of the sample data is not close to 0, the correlation between the features is calculated through the sample data;
[0134] Step S1106. Acquire features with strong correlation, and calculate the information value of the strong correlation features;
[0135] Step S1108. retain the features with the highest information value among the strongly correlated features, and obtain the valid features provided by the data provider;
[0136] Step S1110. Establish a first sub-decision tree according to the effective features, and establish a second sub-decision tree according to the residual of the first sub-decision tree, and iterate the sub-decision tree through the residual of the previous sub-decision tree until the residual of the last sub-decision tree is 0, and obtain a decision tree corresponding to the effective features;
[0137] Step S1112. Convert the sub-decision trees included in the decision tree into a feature network, and determine the weights of directed edges between different feature nodes in the feature network according to the information gain of each split node in the sub-decision tree;
[0138] Step S1114. Calculate the weight of the feature node and assign its score to any feature node according to the weight of the directed edge;
[0139] Step S1116. Calculate the score of any feature node in the current feature network according to the weight of the feature node assigning its score to the feature node;
[0140] Step S1118. Determine a non-current feature network including the current feature node, and calculate the score of the feature node in the non-current feature network;
[0141] Step S1120. Perform a weighted summation on the scores of the feature node in different feature networks to obtain the score of the feature network in the decision tree;
[0142] Step S1122: Screen the features again to optimize the model according to the scores of the feature nodes in the feature network, or distribute benefits to different data providers according to the scores of the feature nodes.
[0143] The exemplary embodiment of the present disclosure also provides a feature evaluation device, referring to Fig.12 As shown, it may include: an effective feature acquisition module 1210, a feature node score calculation module 1220 and a feature evaluation module 1230. Among them:
[0144] The effective feature acquisition module 1210 is used to acquire features provided by a data provider, obtain strongly correlated features included in the features, and filter the strongly correlated features to obtain effective features;
[0145] A feature node score calculation module 1220 is used to establish a decision tree based on the effective features, convert the current sub-decision tree into a current feature network, and determine the scores of the feature nodes included in the current feature network;
[0146] The feature evaluation module 1230 is used to obtain the score of the feature node in the decision tree based on the score of the feature node when it is determined that the feature node exists in a non-current sub-decision tree included in the decision tree, so as to complete the evaluation of the features included in the decision tree.
[0147] The specific details of each module in the above-mentioned feature evaluation device have been described in detail in the corresponding feature evaluation method, so they will not be repeated here.
[0148] In an exemplary embodiment of the present disclosure, the features provided by the data provider are acquired, the strongly correlated features included in the features are obtained, and the strongly correlated features are filtered to obtain valid features, including:
[0149] Calculating the correlation between different features among the features provided by the data provider;
[0150] When the correlation between multiple features is strongly correlated, the information value of the multiple features is calculated, and the multiple features are filtered according to the information value to obtain the feature with the highest information value among the strongly correlated features;
[0151] The effective features of the data provider are obtained according to the features with higher information value among the strongly correlated features.
[0152] In an exemplary embodiment of the present disclosure, establishing a decision tree based on the effective features includes:
[0153] Acquire the valid features, perform anonymization on the feature names of the valid features, and establish a first sub-decision tree based on the feature names after the anonymization;
[0154] Calculating the residual of the first sub-decision tree, and using the residual of the first sub-decision tree to establish a subsequent sub-decision tree;
[0155] The iteration is continued until the number of iterations reaches a preset number of iterations or the residual of the sub-decision tree is less than a preset threshold, thereby obtaining a decision tree established based on the effective features.
[0156] In an exemplary embodiment of the present disclosure, establishing a decision tree based on the effective features includes:
[0157] Acquire sample data corresponding to the feature provided by the data, calculate the information gain of the feature according to the sample data, and use the feature with the largest information gain as the root node of the first sub-decision tree;
[0158] Classify the sample data according to the root node to obtain first sample data and second sample data;
[0159] According to the first sample data and the second sample data, the first sub-decision tree is calculated and constructed according to the information gain of the feature corresponding to the first sample data and the information gain corresponding to the second sample data.
[0160] In an exemplary embodiment of the present disclosure, converting a current sub-decision tree into a current feature network, and determining scores of feature nodes included in the current feature network, includes:
[0161] Obtaining a current sub-decision tree, using a split node in the current sub-decision tree as a feature node in the feature network, and converting a directional relationship between the split nodes into a directed edge in the feature network;
[0162] The current feature network is established according to the feature nodes and the directed edges, and the scores of the feature nodes included in the current feature network are determined according to the information gains of the feature nodes in the current feature network.
[0163] In an exemplary embodiment of the present disclosure, determining the scores of the feature nodes included in the current feature network according to the information gains of the feature nodes in the current feature network includes:
[0164] Normalizing the information gain of the feature node to obtain a standardized information gain, and obtaining the weight of the directed edge through the standardized information gain;
[0165] The score of the feature node is iteratively calculated according to the weight of the directed edge until the score of the feature node converges to obtain the score of the feature node.
[0166] In an exemplary embodiment of the present disclosure, obtaining the score of the feature node in the decision tree according to the score of the feature node includes:
[0167] Calculating a first score of a feature node included in the current feature network in a feature network corresponding to the non-current sub-decision tree;
[0168] According to the first score, the score of the feature node included in the current network in the decision tree is obtained.
[0169] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.
[0170] In addition, although the steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all the steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps, etc.
[0171] In an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided.
[0172] Those skilled in the art will appreciate that various aspects of the present disclosure may be implemented as systems, methods or program products. Therefore, various aspects of the present disclosure may be specifically implemented in the following forms, namely: complete hardware implementation, complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software, which may be collectively referred to herein as "circuits", "modules" or "systems".
[0173] Refer to the following Fig.13 1300 according to this embodiment of the present disclosure is described. Fig.13The electronic device 1300 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0174] like Fig.13 As shown, the electronic device 1300 is in the form of a general computing device. The components of the electronic device 1300 may include, but are not limited to: at least one processing unit 1310, at least one storage unit 1320, a bus 1330 connecting different system components (including the storage unit 1320 and the processing unit 1310), and a display unit 1340.
[0175] The storage unit stores program codes, which can be executed by the processing unit 1310, so that the processing unit 1310 performs the steps according to various exemplary embodiments of the present disclosure described in the above “Exemplary Method” section of this specification. For example, the processing unit 1310 can perform the following steps: Figure 1 The step S110 shown in the figure: acquiring the features provided by the data provider, obtaining the strongly correlated features included in the features, filtering the strongly correlated features, and obtaining effective features; S120: establishing a decision tree based on the effective features, converting the current sub-decision tree into the current feature network, and determining the scores of the feature nodes included in the current feature network; S130: when it is determined that the feature node exists in a non-current sub-decision tree included in the decision tree, obtaining the score of the feature node in the decision tree according to the score of the feature node, so as to complete the evaluation of the features included in the decision tree.
[0176] The storage unit 1320 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 13201 and / or a cache storage unit 13202 , and may further include a read-only storage unit (ROM) 13203 .
[0177] The storage unit 1320 may also include a program / utility 13204 having a set (at least one) of program modules 13205, such program modules 13205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0178] Bus 1330 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0179] The electronic device 1300 may also communicate with one or more external devices 1400 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 1300, and / or communicate with any device that enables the electronic device 1300 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 1350. Furthermore, the electronic device 1300 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 1360. As shown, the network adapter 1360 communicates with other modules of the electronic device 1300 via a bus 1330. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1300, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0180] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the implementation of the present disclosure.
[0181] In an exemplary embodiment of the present disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the above method of the present specification is stored. In some possible implementations, various aspects of the present disclosure may also be implemented in the form of a program product, which includes a program code, and when the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary implementations of the present disclosure described in the above "Exemplary Method" section of the present specification.
[0182] According to the program product for implementing the above method in the embodiment of the present disclosure, it can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, an apparatus or a device.
[0183] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0184] Computer readable signal media may include data signals propagated in baseband or as part of a carrier wave, in which readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Readable signal media may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0185] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.
[0186] Program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0187] In addition, the above-mentioned figures are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, and are not intended to be limiting. It is easy to understand that the processes shown in the above-mentioned figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be performed synchronously or asynchronously, for example, in multiple modules.
[0188] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The specification and examples are to be considered as exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims.
Claims
1. A feature evaluation method, characterized in that: include: Acquire features provided by a data provider, obtain strongly correlated features included in the features, and filter the strongly correlated features to obtain valid features; A decision tree is established based on the effective features, a current sub-decision tree is obtained, a split node in the current sub-decision tree is used as a feature node in a feature network, and a directional relationship between the split nodes is converted into a directed edge in the feature network; a current feature network is established based on the feature nodes and the directed edges, and scores of feature nodes included in the current feature network are determined based on information gains of feature nodes in the current feature network; When it is determined that the feature node exists in a non-current sub-decision tree included in the decision tree, the score of the feature node in the decision tree is obtained according to the score of the feature node to complete the evaluation of the features included in the decision tree.
2. The feature evaluation method according to claim 1, characterized in that: Acquire the features provided by the data provider, obtain the strongly correlated features included in the features, filter the strongly correlated features, and obtain valid features, including: Calculating the correlation between different features among the features provided by the data provider; When the correlation between multiple features is strongly correlated, the information value of the multiple features is calculated, and the multiple features are filtered according to the information value to obtain the feature with the highest information value among the strongly correlated features; The effective features of the data provider are obtained according to the features with higher information value among the strongly correlated features.
3. The feature evaluation method according to claim 1, characterized in that: Establishing a decision tree based on the effective features includes: Acquire the valid features, perform anonymization on the feature names of the valid features, and establish a first sub-decision tree based on the feature names after the anonymization; Calculating the residual of the first sub-decision tree, and using the residual of the first sub-decision tree to establish a subsequent sub-decision tree; The iteration is continued until the number of iterations reaches a preset number of iterations or the residual of the sub-decision tree is less than a preset threshold, thereby obtaining a decision tree established based on the effective features.
4. The feature evaluation method according to claim 3, characterized in that: Establishing a decision tree based on the effective features includes: Acquire sample data corresponding to the feature provided by the data, calculate the information gain of the feature according to the sample data, and use the feature with the largest information gain as the root node of the first sub-decision tree; Classify the sample data according to the root node to obtain first sample data and second sample data; According to the first sample data and the second sample data, the first sub-decision tree is calculated and constructed according to the information gain of the feature corresponding to the first sample data and the information gain corresponding to the second sample data.
5. The feature evaluation method according to claim 1, characterized in that: Determining scores of feature nodes included in the current feature network according to information gains of feature nodes in the current feature network includes: Normalizing the information gain of the feature node to obtain a standardized information gain, and obtaining the weight of the directed edge through the standardized information gain; The score of the feature node is iteratively calculated according to the weight of the directed edge until the score of the feature node converges to obtain the score of the feature node.
6. The feature evaluation method according to claim 5, characterized in that: Obtaining the score of the feature node in the decision tree according to the score of the feature node includes: Calculating a first score of a feature node included in the current feature network in a feature network corresponding to the non-current sub-decision tree; According to the first score, the score of the feature node included in the current network in the decision tree is obtained.
7. A feature evaluation device, characterized in that: include: An effective feature acquisition module is used to acquire features provided by a data provider, obtain strongly correlated features included in the features, and filter the strongly correlated features to obtain effective features; A feature node score calculation module is used to establish a decision tree based on the effective features, obtain a current sub-decision tree, use the split nodes in the current sub-decision tree as feature nodes in a feature network, and convert the directional relationship between the split nodes into directed edges in the feature network; establish a current feature network based on the feature nodes and the directed edges, and determine the scores of the feature nodes included in the current feature network based on the information gain of the feature nodes in the current feature network; The feature evaluation module is used to obtain the score of the feature node in the decision tree according to the score of the feature node when it is determined that the feature node exists in a non-current sub-decision tree included in the decision tree, so as to complete the evaluation of the features included in the decision tree.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the feature evaluation method according to any one of claims 1 to 6 is implemented.
9. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to perform the feature evaluation method described in any one of claims 1-6 by executing the executable instructions.
Citation Information
Patent Citations
Improved feature filtering method and device based on correlation feature selection
CN110135469A