Decision tree generation method and device and computer readable storage medium

By sorting the feature values ​​before the decision tree is constructed, the sorted multi-groups are generated, which solves the problem of low efficiency in decision tree construction and realizes a more efficient decision tree generation process.

CN120069014APending Publication Date: 2025-05-30JINGDONG TECH HLDG CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311575193.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, the construction efficiency of decision trees is low, and frequent traversal and sorting of feature values ​​is required, resulting in low computational efficiency.

Method used

Before building a decision tree, sort the multitudes corresponding to each feature according to the size of the feature value, generate the sorted multitudes, and use these multitudes to directly generate the decision tree to avoid repeated sorting during the construction process.

Benefits of technology

By pre-sorting and storing multiple groups, the sorting operations during the decision tree construction process are reduced, data access efficiency and computing efficiency are improved, and the speed of generating decision trees is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069014A_ABST
    Figure CN120069014A_ABST
Patent Text Reader

Abstract

The invention relates to a decision tree generation method and device and a computer readable storage medium, and relates to the field of data processing. The decision tree generation method comprises the steps that a plurality of samples are acquired, and each sample comprises a sample identifier, a label and a feature value of one or more features; for each feature value, generating a multi-tuple comprising the feature value, a sample identifier of a sample to which the feature value belongs and a label; for the multi-tuple corresponding to each feature, sorting the multi-tuple according to the feature value of the feature; and generating a decision tree by using the sequenced multi-tuples. In the process of constructing the decision tree, re-ordering is not needed, the multi-tuple comprises information needed for calculating information gain, determining segmentation points and dividing sample sets, and complete data of samples does not need to be accessed. Therefore, according to the embodiment of the invention, the calculation efficiency of generating the decision tree can be improved on the whole.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing, and in particular, to a method, an apparatus, and a computer-readable storage medium for generating a decision tree. Background Art

[0002] Data-driven artificial intelligence and its related technologies have played a huge role in all walks of life and brought high value. For example, a decision tree can be constructed through training samples, and then the trained decision tree can be used for prediction. During the process of constructing a decision tree, the feature values of samples are usually traversed to determine appropriate splitting points for generating child nodes. Summary of the Invention

[0003] One technical problem to be solved by the embodiments of the present disclosure is: how to improve the construction efficiency of a decision tree.

[0004] According to a first aspect of some embodiments of the present disclosure, there is provided a method for generating a decision tree, including: obtaining multiple samples, where each sample includes a sample identifier, a label, and feature values of one or more features; for each feature value, generating a tuple including the feature value, the sample identifier of the sample to which the feature value belongs, and the label; for the tuples corresponding to each feature, sorting the tuples according to the feature values of the features; and generating a decision tree by using the sorted tuples.

[0005] In some embodiments, generating a decision tree by using the sorted tuples includes: determining an alternative feature and an alternative feature value for splitting from one or more features and their feature values; when calculating the information gain corresponding to the alternative feature value of the alternative feature, according to the sorting result of the tuples corresponding to the alternative feature, dividing the samples in the tuples before the alternative feature value into a first sample set, and dividing the samples in the tuples after the alternative feature value into a second sample set; determining the information gain corresponding to the alternative feature value of the alternative feature according to the labels of the samples in each sample set; and generating a decision tree according to the information gain corresponding to the alternative feature value of the alternative feature.

[0006] In some embodiments, among the multi-tuples corresponding to the alternative features, there are multi-tuples with missing feature values. And according to the sorting result of the multi-tuples corresponding to the alternative features, pre-partitioning the samples in the multi-tuples before the multi-tuple with the alternative feature value and the samples in the multi-tuples after the multi-tuple with the alternative feature value into a first sample set and a second sample set respectively includes: according to the sorting result of the multi-tuples corresponding to the alternative features, partitioning the samples in the multi-tuples with non-missing feature values that are before the multi-tuple with the alternative feature value into the first sample set, and partitioning the samples in the multi-tuples with non-missing feature values that are after the multi-tuple with the alternative feature value into the second sample set; determining multiple allocation methods for allocating the multi-tuples with missing feature values to the first sample set and the second sample set to obtain the first sample set and the second sample set under multiple allocation methods.

[0007] In some embodiments, determining the information gain corresponding to the alternative feature value of the alternative feature according to the labels of the samples in each sample set includes: according to the labels of the samples in each sample set under each allocation method, determining the information gain corresponding to the alternative feature value of the alternative feature under the allocation method; determining the maximum information gain among multiple allocation methods as the information gain corresponding to the alternative feature value of the alternative feature.

[0008] In some embodiments, determining the alternative feature and the alternative feature value for splitting from one or more features and their feature values includes: selecting an alternative feature from one or more features; sampling the multi-tuples of the alternative feature; selecting an alternative feature value from the feature values of the multi-tuples obtained after sampling.

[0009] In some embodiments, sampling the multi-tuples of the alternative feature includes: determining the multi-tuples corresponding to the quantiles in the feature values of the alternative feature as the multi-tuples obtained after sampling.

[0010] In some embodiments, each sample further includes an intervention condition identifier of the intervention condition applied to the sample, and the multi-tuple further includes an intervention condition identifier. And when the intervention conditions applied to multiple samples include multiple ones, determining the information gain corresponding to the alternative feature value of the alternative feature according to the labels of the samples in each sample set includes: for each sample set, traversing the combinations of two intervention condition identifiers in the intervention condition identifiers of the multi-tuples corresponding to the sample set, calculating the difference in the means of the labels of the samples in the sample set under the two intervention conditions in each combination, and calculating the sum of these differences as the feedback information of the sample set; determining the information gain corresponding to the alternative feature value of the alternative feature according to the difference in the feedback information of the first sample set and the second sample set.

[0011] In some embodiments, the sorted multi-tuples corresponding to each feature are stored in memory, and the starting address in memory of the first multi-tuple corresponding to each feature is recorded.

[0012] In some embodiments, each sample further includes an intervention condition identifier of the intervention condition applied to the sample.

[0013] In some embodiments, the generation method further includes: for each leaf node of the decision tree, determining the prediction result corresponding to the leaf node according to the labels of the samples corresponding to the leaf node under multiple intervention conditions, so as to use the decision tree for predicting the intervention results of multiple intervention conditions.

[0014] In some embodiments, determining the prediction result corresponding to the leaf node according to the feedback values of the samples corresponding to the leaf node under multiple intervention conditions includes: for each intervention condition, determining the mean value of the labels of the samples corresponding to the leaf node under the intervention condition as the expected value under the intervention condition corresponding to the leaf node; determining the prediction result corresponding to the leaf node according to the expected values corresponding to the leaf node under each intervention condition.

[0015] In some embodiments, determining the prediction result corresponding to the leaf node according to the expected values corresponding to the leaf node under each intervention condition includes: for each intervention condition, comparing the expected value corresponding to the leaf node under the intervention condition with the corresponding intervention result threshold to determine the effectiveness information of the intervention condition; determining the effectiveness information of each intervention condition as the prediction result corresponding to the leaf node.

[0016] In some embodiments, the sample is a biological sample, the intervention condition includes at least one of the drug type or drug dose, and the feedback information represents the reaction of the biological sample after the intervention condition is applied; alternatively, the sample is the user of the application, the intervention condition includes at least one of the resources or information delivered to the user, and the feedback information represents the interaction operation of the user with the application after the intervention condition is applied.

[0017] According to a second aspect of some embodiments of the present disclosure, there is provided a decision tree generation device, including: an acquisition module configured to acquire multiple samples, where each sample includes a sample identifier, a label, and feature values of one or more features; a multi-tuple generation module configured to generate, for each feature value, a multi-tuple including the feature value, the sample identifier of the sample to which the feature value belongs, and the label; a sorting module configured to sort the multi-tuples corresponding to each feature according to the feature values of the features; and a decision tree generation module configured to generate a decision tree using the sorted multi-tuples.

[0018] According to a third aspect of some embodiments of the present disclosure, there is provided a decision tree generation device, including: a memory; and a processor coupled to the memory, the processor being configured to execute any one of the foregoing decision tree generation methods based on instructions stored in the memory.

[0019] According to a fourth aspect of some embodiments of the present disclosure, there is provided a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, it implements any one of the foregoing decision tree generation methods.

[0020] Before constructing the decision tree in the embodiments of the present disclosure, the multi-tuples corresponding to each feature are sorted according to the magnitude of the feature values. The multi-tuples include feature values, sample identifiers, and labels. Therefore, during the process of constructing the decision tree, there is no need to re-sort, and the multi-tuples include the information required for calculating the information gain, determining the splitting point, and dividing the sample set, without the need to access the complete data of the samples. Therefore, the above embodiments can improve the computational efficiency of generating the decision tree as a whole.

[0021] Other features and advantages of the present disclosure will become clear from the following detailed description of the exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0023] Figure 1 FIG. shows a schematic flow chart of a method for generating a decision tree according to some embodiments of the present disclosure.

[0024] Figure 2 FIG. shows a schematic flow chart of a method for constructing a decision tree according to some embodiments of the present disclosure.

[0025] Figure 3 FIG. shows a schematic flow chart of a method for generating an intervention result prediction model according to some embodiments of the present disclosure.

[0026] Figure 4 FIG. shows a schematic flow chart of a prediction method according to some embodiments of the present disclosure.

[0027] Figure 5 FIG. shows a schematic structural diagram of a decision tree generation device according to some embodiments of the present disclosure.

[0028] Figure 6 FIG. shows a schematic structural diagram of a decision tree generation device according to some other embodiments of the present disclosure.

[0029] Figure 7 FIG. shows a schematic structural diagram of a decision tree generation device according to some further embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] Next, in combination with the accompanying drawings in the embodiments of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and in no way limits the present disclosure and its application or use. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present disclosure.

[0031] Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.

[0032] At the same time, it should be understood that, for the sake of convenience of description, the dimensions of the various parts shown in the drawings are not drawn in accordance with the actual proportional relationship.

[0033] For technologies, methods, and devices known to those of ordinary skill in the relevant fields, they may not be discussed in detail, but under appropriate circumstances, the said technologies, methods, and devices should be regarded as a part of the specification.

[0034] In all the examples shown and discussed here, any specific value should be construed as merely exemplary, rather than as a limitation. Therefore, other examples of the exemplary embodiments may have different values.

[0035] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0036] In the related art, for the training samples used to generate decision trees, a data reading method with row-based indexing is usually adopted, and each row in the data table corresponds to a sample. Table 1 exemplarily shows the storage form of the training samples in the data table. As shown in Table 1, the sample identifier represents samples 0 to 10, and each sample further includes the label of the sample and the feature values of features 0 to 2. The symbol "-" represents a missing value. According to needs, the sample may further include intervention information, and the intervention information may indicate whether an intervention condition is imposed on the sample or the type of the intervention condition. In Table 1, it is schematically indicated that 0 represents that the sample is not imposed with an intervention condition, and 1 represents that the sample is imposed with an intervention condition. If there are multiple intervention conditions, 0 may be used to represent that the sample is not imposed with an intervention condition, 1 may be used to represent that the sample is imposed with intervention condition 1, 2 may be used to represent that the sample is imposed with intervention condition 2... and so on.

[0037] Table 1

[0038] Sample Identification Label Feature 0 Feature 1 Feature 2 Intervention Condition 0 0.85 4.2 9.0 9.2 1 1 0.31 3.1 0.0 - 0 2 0.40 4.2 4.2 4.1 0 3 0.76 - 2.0 3.6 1 4 0.63 5.9 5.3 5.0 1 5 0.98 0.7 8.9 8.8 0 6 0.50 6.2 - - 0 7 0.99 8.7 5.3 5.5 1 8 0.36 8.3 2.7 5.6 1 9 0.36 - 4.4 3.4 0 10 0.54 1.3 1.2 - 1

[0039] In the process of traversing the data table shown in Table 1 to construct a decision tree, in order to obtain different values of features, it is necessary to first read the samples and then read the feature values of the samples to complete this process. Therefore, this reading process is a random access of data, which affects the efficiency of feature traversal. Moreover, the construction process of the decision tree involves a recursive process of constructing child nodes. In the related art, it is necessary to sort the feature values during the process of generating child nodes for each leaf node, so as to continue the node division. Therefore, a large number of sorting operations exist in the production process. Moreover, in the newly generated leaf nodes, the samples are disordered, resulting in the non-reusability of the previous sorting, thus affecting the speed of generating the decision tree.

[0040] In order to solve at least the above problems, the embodiments of the present disclosure reprocess the storage form of the samples before constructing the decision tree. Figure 1 FIG. shows a schematic flowchart of a method for generating a decision tree according to some embodiments of the present disclosure. As Figure 1 shown, the method for generating the decision tree of this embodiment includes steps S102 to S108.

[0041] In step S102, multiple samples are obtained, where each sample includes a sample identifier, a label, and feature values of one or more features. Different samples have the same one or more features. In each sample, the feature value corresponding to a certain feature is one.

[0042] In addition to the above information, the sample may further include intervention information on the intervention conditions applied to the sample. In some embodiments, the sample is an intervened object for an intervention experiment, such as a biological sample, a user, etc. The features of the sample are used to indicate information such as the attributes, categories, behaviors, etc. of the sample.

[0043] In some embodiments, the sample is a biological sample, the intervention conditions include at least one of the drug type or drug dose, and the label represents the reaction of the biological sample after applying the intervention conditions; or, the sample is an applied user, the intervention conditions include at least one of the resources or information delivered to the user, and the label represents the interaction operation of the user with the application after applying the intervention conditions. According to needs, the embodiments of the present disclosure can also be applied to other technical fields, which will not be elaborated here.

[0044] In some embodiments, the samples are divided into multiple groups corresponding to multiple intervention conditions, and the samples in each group are subjected to the corresponding intervention condition. Thus, for each sample, only one experiment needs to be conducted, that is, applying one intervention condition. After the intervention, the feedback value of the sample can be recorded as a label. The direct feedback result of each sample can be a numerical value or other types of information. If it is other types of information, it can be quantified and represented by a numerical value, which is convenient for subsequent calculation processes.

[0045] In step S104, for each eigenvalue, a multi-tuple including the eigenvalue, the sample identifier of the sample to which the eigenvalue belongs, and the label is generated.

[0046] For example, for the eigenvalue 9.0 of sample 0 under feature 1 in Table 1, a quadruple {9.0, 0, 0.85, 1} can be generated. The values in the quadruple represent the eigenvalue, the sample identifier, the label, and the intervention condition in sequence. The present disclosure has no limitation on the arrangement order of the contents in the multi-tuple. However, the arrangement order of various elements in different multi-tuples needs to be consistent.

[0047] In step S106, for the multi-tuples corresponding to each feature, the multi-tuples are sorted according to the eigenvalues of the feature, and the sorting result is stored. That is, each feature corresponds to multiple sorted multi-tuples with the same number as the number of samples, and these multi-tuples are sorted according to the magnitudes of the eigenvalues in the multi-tuples. If two multi-tuples have the same eigenvalue, they can be sorted according to the magnitudes of other elements in the multi-tuples or sorted according to other preset rules.

[0048] Taking feature 0 in Table 1 as an example, the sorting result of the eigenvalues of feature 0 of each sample is 0.7, 1.3, 3.1, 4.2, 4.2, 5.9, 6.2, 8.3, 8.7, -, -. In this sorting, the missing values are exemplarily placed at the end of the sorting result. According to requirements, these missing values can also be arranged at the starting position. According to the above sorting result, the arrangement order of the quadruples corresponding to feature 0 is {0.7, 5, 0.98, 0}, {1.3, 10, 0.54, 1}, {3.1, 1, 0.31, 0}, {4.2, 0, 0.85, 1}, {4.2, 2, 0.40, 0}, {5.9, 4, 0.63, 1}, {6.2, 6, 0.50, 0}, {8.3, 8, 0.36, 1}, {8.7, 7, 0.99, 1}, {NULL, 3, 0.76, 1}, {NULL, 9, 0.36, 0}. "NULL"

[0049] (empty) represents a missing value.

[0050] The above processing converts the row-major index method into the column-major index method where the eigenvalues are located, performs row-column conversion on the storage of the original sample data, and presents it in the form of multi-tuples. Moreover, the converted columns are pre-sorted and stored according to the eigenvalues. Therefore, the speed of data access in the subsequent decision tree construction process is improved.

[0051] In some embodiments, the sorted multi-tuples corresponding to each feature are stored in memory, and the starting address in memory of the first multi-tuple corresponding to each feature is recorded. Thus, in the stage of generating the decision tree, these eigenvalue values only need to be accessed as needed, and the process of accessing each feature is a continuous memory access, greatly improving the memory access efficiency.

[0052] In step S108, a decision tree is generated using the sorted multi-tuples. That is, in the process of constructing the decision tree, data is read from the sorted multi-tuples corresponding to each feature, so that there is no need to repeatedly sort each feature.

[0053] Before constructing the decision tree in the above embodiments, the multi-tuples corresponding to each feature are sorted according to the magnitude of the eigenvalue values. The multi-tuples include eigenvalue values, sample identifiers, and labels. Thus, in the process of constructing the decision tree, there is no need to re-sort, and the multi-tuples include the information required for calculating the information gain, determining the splitting point, and dividing the sample set, and there is no need to access the complete data of the samples again. Therefore, the above embodiments can improve the computational efficiency of generating the decision tree as a whole.

[0054] The embodiments of generating a decision tree are described below by way of example. In some embodiments, from one or more features and their eigenvalue values, alternative features and alternative eigenvalue values for splitting are determined; when calculating the information gain corresponding to the alternative eigenvalue values of the alternative features, according to the sorting result of the multi-tuples corresponding to the alternative features, the samples in the multi-tuples before the multi-tuple with the alternative eigenvalue value are divided into the first sample set, and the samples in the multi-tuples after the multi-tuple with the alternative eigenvalue value are divided into the second sample set, and the multi-tuple including the alternative eigenvalue value can be divided into one of the sample sets according to a preset rule; according to the labels of the samples in each sample set, the information gain corresponding to the alternative eigenvalue values of the alternative features is determined; a decision tree is generated according to the information gain corresponding to the alternative eigenvalue values of the alternative features. Thus, when calculating the information gain, there is no need to sort the eigenvalue values each time, but the pre-generated sorting result can be utilized to select the samples that meet each division condition and add them to the sample set. The following describes a complete decision tree construction process by way of example.

[0055] Figure 2 Shows a flowchart of a method for constructing a decision tree according to some embodiments of the present disclosure. As Figure 2As shown, the decision tree construction method of this embodiment includes S202 to S208.

[0056] In step S202, a root node of the decision tree is generated, where the root node corresponds to all samples.

[0057] Then, steps S204 to S208 are repeatedly executed until a stop condition is reached. Here, the node to be processed is the current leaf node of the decision tree, and the current alternative feature is a feature that has not been used to divide the nodes in the decision tree. In some embodiments, the number of samples corresponding to the node to be processed is greater than the sample threshold, thereby avoiding overfitting of the model and improving the accuracy of decision tree prediction.

[0058] In some embodiments, the stop condition is that the depth of the decision tree reaches the depth threshold; or, there is no longer a node with a corresponding number of samples greater than the sample threshold.

[0059] In step S204, from one or more features and their feature values, an alternative feature and an alternative feature value for splitting are determined, and the information gain of each current alternative feature and feature value is calculated using the difference in the feedback information of the samples corresponding to the node to be processed. For example, all features and all feature values of the samples in the node to be processed can be traversed, and each feature value of each feature is used as an alternative feature value in turn, and the information gain of the alternative feature value is calculated.

[0060] In some embodiments, each feature can be traversed and selected as an alternative feature one by one.

[0061] In some embodiments, in order to further improve the calculation efficiency, sampling can be performed on the multi-tuples of alternative features; the alternative feature values are selected from the feature values of the multi-tuples obtained after sampling. That is, for all feature values of a certain feature, instead of calculating their information gain one by one, only some of the feature values are selected to calculate the information gain. In some embodiments, the multi-tuples corresponding to the quantiles in the feature values of the alternative features are determined as the multi-tuples obtained after sampling, that is, the feature values at the quantiles are used as the alternative feature values, such as the feature values at 20%, 40%, 60%, and 80% after sorting. Thus, the amount of calculation is reduced and the calculation efficiency is improved without causing a great impact on the model accuracy.

[0062] In some embodiments, according to the sorting results of the multi-tuples corresponding to the alternative features, the samples in the multi-tuples before the multi-tuple with the alternative feature value are divided into the first sample set, and the samples in the multi-tuples after the multi-tuple with the alternative feature value are divided into the second sample set; according to the labels of the samples in each sample set, the information gain corresponding to the alternative feature value of the alternative feature is determined. For example, calculate the mean value of the labels of each sample set, and take the difference between the means of the two sample sets as the information gain. Thus, by reading the sorted multi-tuples, the calculation of the information gain can be completed without accessing the original data of the samples.

[0063] In some embodiments, when the samples and multi-tuples further include intervention identifiers and there are multiple intervention messages, the following method can be used to determine the information gain. For each of the first sample set and the second sample set corresponding to the alternative feature, traverse the combinations of two intervention conditions among the multiple intervention conditions, calculate the difference between the mean values of the labels of the samples in the set under the two intervention conditions in each combination, and calculate the sum of these differences as the feedback information of the set; determine the information gain of the alternative feature according to the difference between the feedback information of the first sample set and the second sample set. Since the multi-tuples carry intervention identifiers, when calculating the information gain under multiple intervention conditions, the sorted quadruples can also be directly used to complete the calculation.

[0064] For example, when dividing the samples corresponding to a node to be processed by whether feature 1 is greater than 5, the multi-tuples with a feature value of 5 and the multi-tuples before it in the multi-tuples of feature 1 can be directly divided into set 1, and the multi-tuples after the multi-tuple with a feature value of 5 can be divided into set 2. Suppose there are 4 intervention conditions, then the samples in set 1 and set 2 can be divided into 4 groups respectively according to the applied intervention conditions, and the mean value of the feedback values of the samples in each group can be calculated.

[0065] Then, for set 1, generate all combinations of 2 intervention conditions, such as {group 1, group 2}, {group 1, group 3}, {group 1, group 4}, {group 2, group 3}, {group 2, group 4}, {group 3, group 4}, calculate the difference between the mean values corresponding to each combination, and then sum these differences as the feedback information of set 1. Set 2 is calculated in a similar manner. Finally, take the absolute value of the difference between the feedback information of set 1 and set 2 as the information gain of feature 1.

[0066] In step S206, according to the information gain of each alternative feature, select the alternative feature for division. For example, select the alternative feature with the largest information gain.

[0067] Embodiments of the present disclosure can also specifically handle the case where a feature has missing values. In some embodiments, according to the sorting result of the multi-tuples corresponding to the alternative features, among the multi-tuples with non-missing feature values, the samples in the multi-tuples before the multi-tuples with alternative feature values are divided into the first sample set, and the samples in the multi-tuples after the multi-tuples with alternative feature values are divided into the second sample set; determine various allocation methods for allocating the multi-tuples with missing feature values to the first sample set and the second sample set, so as to obtain the first sample set and the second sample set under various allocation methods. For example, when the samples of the multi-tuples with non-missing feature values of a certain feature are respectively divided into set 1 and set 2, then traverse all the ways of continuously allocating the samples of the multi-tuples with missing feature values of this feature to set 1 and set 2, thereby generating all sample allocation methods corresponding to this feature.

[0068] Thus, in the case where the features of the samples have missing values, it is also possible to consider various possible partitioning methods of such samples and make decisions among these partitioning methods. For example, according to the labels of the samples in each sample set under each allocation method, determine the information gain corresponding to the alternative feature value of the alternative feature under the allocation method; determine the maximum information gain among multiple allocation methods as the information gain corresponding to the alternative feature value of the alternative feature.

[0069] In step S208, use the alternative features for partitioning to process the node to be processed and its corresponding samples, so as to generate child nodes of the node to be processed and determine the samples corresponding to each child node. When partitioning, according to the sample identifiers in the multi-tuples, the corresponding relationship between the samples and the newly generated child nodes can be determined.

[0070] In the process of constructing a decision tree in the above embodiments, it is possible to calculate the information gain and partition the samples based on the sorted multi-tuples corresponding to each feature, thereby improving the calculation efficiency of generating the decision tree.

[0071] After generating the decision tree, the decision tree can be used in the prediction process. For example, when the samples used to generate the decision tree include intervention identifiers and there are multiple intervention conditions applied to these samples, the decision tree can be used for multi-intervention condition prediction. The following refers to Figure 3 Describe an embodiment of a method for generating an intervention result prediction model.

[0072] Figure 3 FIG. shows a flowchart of a method for generating an intervention result prediction model according to some embodiments of the present disclosure. As Figure 3 shown, the method for generating the intervention result prediction model of this embodiment includes steps S302 to S304.

[0073] In step S302, after generating the decision tree, for each intervention condition, the mean value of the labels of the samples corresponding to the leaf nodes under this intervention condition is determined as the expected value of the leaf nodes under this intervention condition.

[0074] In step S304, according to the expected values of the leaf nodes corresponding to each intervention condition, the prediction results corresponding to the leaf nodes are determined.

[0075] In some embodiments, the expected values of the leaf nodes corresponding to each intervention condition are determined as the prediction results corresponding to the leaf nodes. This way can obtain quantitative and more refined prediction results.

[0076] In some embodiments, for each intervention condition, the expected value of the leaf nodes corresponding to the intervention condition is compared with the corresponding intervention result threshold to determine the effectiveness information of the intervention condition; the effectiveness information of each intervention condition is determined as the prediction result corresponding to the leaf nodes. This way can obtain qualitative prediction results with higher reliability.

[0077] For example, for the samples corresponding to a certain leaf node, the mean value of the labels for intervention condition 1 is 3.5, the mean value of the labels for intervention condition 2 is 4.1, the mean value of the labels for intervention condition 3 is 0.8, and the mean value of the labels for intervention condition 4 is -2.

[0078] One way is that {intervention condition 1: 3.5; intervention condition 2: 4.1; intervention condition 3: 0.8; intervention condition 4: -2} can be used as the prediction result corresponding to this leaf node.

[0079] Another way is that if the threshold for intervention condition 1 is 3 and being greater than the threshold is effective, the threshold for intervention condition 2 is 4 and being greater than the threshold is effective, the threshold for intervention condition 3 is 1 and being less than the threshold is effective, and the threshold for intervention condition 4 is 0 and being less than the threshold is effective, then {intervention condition 1: effective; intervention condition 2: effective; intervention condition 3: effective; intervention condition 4: effective} can be obtained.

[0080] As can be seen from the above embodiments, by applying an intervention condition to each sample and constructing a decision tree model, the obtained model can predict the feedback results of an individual under each intervention condition. Therefore, the feedback results of an individual under each intervention condition can be efficiently obtained.

[0081] After obtaining the intervention result prediction model, the model can also be used to predict the intervention results of the samples to be tested. The following refers to Figure 4 Describe the embodiments of the prediction method of the present disclosure.

[0082] Figure 4The flowchart shows a prediction method according to some embodiments of the present disclosure. As Figure 4 shown, the prediction method of this embodiment includes steps S402 to S406.

[0083] In step S402, the sample to be tested is input into the decision tree, and each leaf node of the decision tree corresponds to a prediction result under multiple intervention conditions.

[0084] In step S404, the leaf node to which the sample to be tested belongs in the decision tree is determined.

[0085] For example, starting from the root node, according to the division condition corresponding to each node, it is determined to which child node the sample to be tested should correspond until it corresponds to a certain leaf node.

[0086] In step S406, according to the prediction result corresponding to the leaf node, the intervention result of the sample to be tested under multiple intervention conditions is determined.

[0087] For example, if the prediction result of the leaf node is {Intervention condition 1: effective; Intervention condition 2: effective; Intervention condition 3: ineffective; Intervention condition 4: ineffective}, it means that intervention conditions 1 and 2 are effective for the sample to be tested, and intervention conditions 3 and 4 are ineffective for the sample to be tested.

[0088] Through the above embodiments, the experimental results can be obtained without conducting experiments on the sample to be tested, and moreover, the prediction of the intervention results for multiple intervention conditions can be obtained at one time. Thus, the embodiments of the present disclosure can improve the efficiency of information processing.

[0089] It should be noted that in the technical solution of the present disclosure, in terms of the collection, acquisition, update, analysis, processing, use, transmission, storage, etc. of the user's personal information, all comply with the provisions of relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. Necessary measures are taken for the user's personal information to prevent illegal access to the user's personal information data, and to safeguard the user's personal information security, network security, and national security.

[0090] Next, refer to Figures 5 to 7 to describe an embodiment of the decision tree generation device of the present disclosure.

[0091] Figure 5 The structural diagram shows a decision tree generation device according to some embodiments of the present disclosure. As Figure 5As shown in the figure, the generating device 50 of the intervention result prediction model of this embodiment includes: an obtaining module 510, configured to obtain multiple samples, where each sample includes a sample identifier, a label, and the eigenvalue of one or more features; a multi-tuple generating module 520, configured to generate, for each eigenvalue, a multi-tuple including the eigenvalue, the sample identifier of the sample to which the eigenvalue belongs, and the label; a sorting module 530, configured to sort the multi-tuples corresponding to each feature according to the eigenvalue of the feature; a decision tree generating module 540, configured to generate a decision tree by using the sorted multi-tuples.

[0092] In some embodiments, the decision tree generating module 540 is further configured to determine, from one or more features and their eigenvalues, an alternative feature and an alternative eigenvalue for splitting; when calculating the information gain corresponding to the alternative eigenvalue of the alternative feature, according to the sorting result of the multi-tuples corresponding to the alternative feature, divide the samples in the multi-tuples before the multi-tuple of the alternative eigenvalue into a first sample set, and divide the samples in the multi-tuples after the multi-tuple of the alternative eigenvalue into a second sample set; determine the information gain corresponding to the alternative eigenvalue of the alternative feature according to the labels of the samples in each sample set; generate a decision tree according to the information gain corresponding to the alternative eigenvalue of the alternative feature.

[0093] In some embodiments, among the multi-tuples corresponding to the alternative feature, there are multi-tuples with missing eigenvalues, and the decision tree generating module 540 is further configured to, according to the sorting result of the multi-tuples corresponding to the alternative feature, divide the samples in the multi-tuples with non-missing eigenvalues before the multi-tuple of the alternative eigenvalue into a first sample set, and divide the samples in the multi-tuples with non-missing eigenvalues after the multi-tuple of the alternative eigenvalue into a second sample set; determine multiple allocation methods for allocating the multi-tuples with missing eigenvalues to the first sample set and the second sample set, so as to obtain the first sample set and the second sample set under multiple allocation methods.

[0094] In some embodiments, the decision tree generating module 540 is further configured to determine the information gain corresponding to the alternative eigenvalue of the alternative feature under each allocation method according to the labels of the samples in each sample set under each allocation method; determine the maximum information gain among multiple allocation methods as the information gain corresponding to the alternative eigenvalue of the alternative feature.

[0095] In some embodiments, the decision tree generating module 540 is further configured to select an alternative feature from one or more features; sample the multi-tuples of the alternative feature; select an alternative eigenvalue from the eigenvalues of the sampled multi-tuples.

[0096] In some embodiments, the decision tree generation module 540 is further configured to determine, as the multi-tuples obtained after sampling, the multi-tuples corresponding to the quantiles in the feature values of the alternative features.

[0097] In some embodiments, each sample further includes an intervention condition identifier of the intervention condition applied to the sample, and the multi-tuples further include the intervention condition identifier. When the intervention conditions applied to multiple samples include multiple ones, the decision tree generation module 540 is further configured to, for each sample set, traverse the combinations of two intervention condition identifiers among the intervention condition identifiers of the multi-tuples corresponding to the sample set, calculate the difference between the means of the labels of the samples in the sample set under the two intervention conditions in each combination, and calculate the sum of these differences as the feedback information of the sample set; determine the information gain corresponding to the alternative feature value of the alternative feature according to the difference between the feedback information of the first sample set and the second sample set.

[0098] In some embodiments, the sorted multi-tuples corresponding to each feature are stored in the memory, and the starting address in the memory of the first multi-tuple corresponding to each feature is recorded.

[0099] In some embodiments, each sample further includes an intervention condition identifier of the intervention condition applied to the sample.

[0100] In some embodiments, the generating device 50 further includes: an intervention result prediction model generation module 550, configured to, for each leaf node of the decision tree, determine the prediction result corresponding to the leaf node according to the labels of the samples corresponding to the leaf node under multiple intervention conditions, so as to use the decision tree for predicting the intervention results of multiple intervention conditions.

[0101] In some embodiments, the intervention result prediction model generation module 550 is further configured to, for each intervention condition, determine the mean value of the labels of the samples corresponding to the leaf node under the intervention condition as the expected value under the intervention condition corresponding to the leaf node; determine the prediction result corresponding to the leaf node according to the expected value corresponding to the leaf node under each intervention condition.

[0102] In some embodiments, the intervention result prediction model generation module 550 is further configured to, for each intervention condition, compare the expected value under the intervention condition corresponding to the leaf node with the corresponding intervention result threshold to determine the effectiveness information of the intervention condition; determine the effectiveness information of each intervention condition as the prediction result corresponding to the leaf node.

[0103] In some embodiments, the sample is a biological sample, the intervention condition includes at least one of the type of drug or the dosage of the drug, and the feedback information represents the reaction of the biological sample after the intervention condition is applied; alternatively, the sample is the user of the application, the intervention condition includes at least one of the resources or information delivered to the user, and the feedback information represents the interaction operation of the user with the application after the intervention condition is applied.

[0104] Figure 6 FIG. shows a schematic structural diagram of a decision tree generation device according to some other embodiments of the present disclosure. As Figure 6 shown, the decision tree generation device 60 of this embodiment includes: a memory 610 and a processor 620 coupled to the memory 610. The processor 620 is configured to execute the decision tree generation method in any of the foregoing embodiments based on the instructions stored in the memory 610.

[0105] Among them, the memory 610 may include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory stores, for example, an operating system, application programs, a boot loader (Boot Loader), and other programs.

[0106] Figure 7 FIG. shows a schematic structural diagram of a decision tree generation device according to still some other embodiments of the present disclosure. As Figure 7 shown, the decision tree generation device 70 of this embodiment includes: a memory 710 and a processor 720, and may further include an input / output interface 730, a network interface 740, a storage interface 750, etc. These interfaces 730, 740, 750 and the memory 710 and the processor 720 may be connected through a bus 760, for example. Among them, the input / output interface 730 provides a connection interface for input / output devices such as a display, a mouse, a keyboard, and a touch screen. The network interface 740 provides a connection interface for various networking devices. The storage interface 750 provides a connection interface for external storage devices such as an SD card and a USB flash drive.

[0107] Embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, the decision tree generation method in any of the foregoing is implemented.

[0108] Those skilled in the art should understand that the embodiments of the present disclosure may be provided as a method, a system, or a computer program product. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure may take the form of a computer program product implemented on one or more computer-usable non-transitory storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0109] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more of the flows Figure 1 one or more flows and / or blocks Figure 1 or the functions specified in one or more of the blocks.

[0110] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more flows and / or blocks Figure 1 or the functions specified in one or more of the blocks.

[0111] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more flows and / or blocks Figure 1 or the functions specified in one or more of the blocks.

[0112] The above are only the preferred embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A method for generating a decision tree, comprising: obtaining multiple samples, where each sample includes a sample identifier, a label, and feature values of one or more features; for each feature value, generating a tuple including the feature value, the sample identifier of the sample to which the feature value belongs, and the label; for the tuples corresponding to each feature, sorting the tuples according to the feature values of the feature; generating a decision tree using the sorted tuples.

2. The generating method according to claim 1, wherein, the generating a decision tree using the sorted tuples includes: determining an alternative feature and an alternative feature value for splitting from the one or more features and their feature values; when calculating the information gain corresponding to the alternative feature value of the alternative feature, according to the sorting result of the tuples corresponding to the alternative feature, dividing the samples in the tuples before the tuple of the alternative feature value into a first sample set, and dividing the samples in the tuples after the tuple of the alternative feature value into a second sample set; determining the information gain corresponding to the alternative feature value of the alternative feature according to the labels of the samples in each sample set; generating the decision tree according to the information gain corresponding to the alternative feature value of the alternative feature.

3. The generating method according to claim 2, wherein, in the tuples corresponding to the alternative feature, there are tuples with missing feature values, and the dividing the samples in the tuples before the tuple of the alternative feature value and the samples in the tuples after the tuple of the alternative feature value into the first sample set and the second sample set respectively according to the sorting result of the tuples corresponding to the alternative feature includes: according to the sorting result of the tuples corresponding to the alternative feature, dividing the samples in the tuples with non-missing feature values that are before the tuple of the alternative feature value into the first sample set, and dividing the samples in the tuples with non-missing feature values that are after the tuple of the alternative feature value into the second sample set; determining multiple allocation methods for allocating the tuples with missing feature values to the first sample set and the second sample set to obtain the first sample set and the second sample set under the multiple allocation methods.

4. The generating method according to claim 3, wherein, the determining the information gain corresponding to the alternative feature value of the alternative feature according to the labels of the samples in each sample set includes: determining the information gain corresponding to the alternative feature value of the alternative feature under the allocation method according to the labels of the samples in each sample set under each allocation method; determining the maximum information gain among the multiple allocation methods as the information gain corresponding to the alternative feature value of the alternative feature.

5. The generating method according to claim 2, wherein, the determining an alternative feature and an alternative feature value for splitting from the one or more features and their feature values includes: selecting an alternative feature from the one or more features; sampling the tuples of the alternative feature; Select alternative feature values from the feature values of the multi-tuples obtained after sampling.

6. The generation method according to claim 5, wherein, the sampling of the multi-tuples of the alternative features includes: Determine the multi-tuples obtained after sampling as the multi-tuples corresponding to the quantiles in the feature values of the alternative features.

7. The generation method according to claim 2, wherein, Each sample further includes an intervention condition identifier of the intervention condition applied to the sample, and the multi-tuples further include the intervention condition identifier. When the intervention conditions applied to the multiple samples include multiple ones, the determining the information gain corresponding to the alternative feature value of the alternative feature according to the labels of the samples in each sample set includes: For each sample set, traverse the combinations of two intervention condition identifiers in the intervention condition identifiers of the multi-tuples corresponding to the sample set, calculate the difference in the means of the labels of the samples in the sample set under the two intervention conditions in each combination, and calculate the sum of these differences as the feedback information of the sample set; Determine the information gain corresponding to the alternative feature value of the alternative feature according to the difference in the feedback information of the first sample set and the second sample set.

8. The generation method according to claim 1, wherein, The sorted multi-tuples corresponding to each feature are stored in memory, and the starting address in memory of the first multi-tuple corresponding to each feature is recorded.

9. The generation method according to claim 1, wherein, Each sample further includes an intervention condition identifier of the intervention condition applied to the sample.

10. The generation method according to claim 9, further includes: For each leaf node of the decision tree, determine the prediction result corresponding to the leaf node according to the labels of the samples corresponding to the leaf node under the multiple intervention conditions, so as to use the decision tree for predicting the intervention results of the multiple intervention conditions.

11. The generation method according to claim 10, wherein, The determining the prediction result corresponding to the leaf node according to the feedback values of the samples corresponding to the leaf node under the multiple intervention conditions includes: For each intervention condition, determine the mean value of the labels of the samples corresponding to the leaf node under the intervention condition as the expected value of the leaf node under the intervention condition; Determine the prediction result corresponding to the leaf node according to the expected values of the leaf node under each intervention condition.

12. The generation method according to claim 11, wherein, The determining the prediction result corresponding to the leaf node according to the expected values of the leaf node under each intervention condition includes: For each intervention condition, compare the expected value of the leaf node under the intervention condition with the corresponding intervention result threshold to determine the effectiveness information of the intervention condition; Determine the effectiveness information of each intervention condition as the prediction result corresponding to the leaf node.

13. The generation method according to any one of claims 9 to 12, wherein: The sample is a biological sample, the intervention condition includes at least one of the drug type or drug dose, and the feedback information represents the reaction of the biological sample after the intervention condition is applied; or, The sample is the user of the application, the intervention condition includes at least one of the resources or information delivered to the user, and the feedback information represents the interaction operation of the user with the application after the intervention condition is applied.

14. A decision tree generation device, comprising: an acquisition module configured to acquire multiple samples, where each sample includes a sample identifier, a label, and feature values of one or more features; a multi-tuple generation module configured to generate, for each feature value, a multi-tuple including the feature value, the sample identifier of the sample to which the feature value belongs, and the label; a sorting module configured to sort the multi-tuples corresponding to each feature according to the feature values of the feature; a decision tree generation module configured to generate a decision tree by using the sorted multi-tuples.

15. A decision tree generation device, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute the decision tree generation method according to any one of claims 1 to 13 based on instructions stored in the memory.

16. A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the decision tree generation method according to any one of claims 1 to 13 is implemented.