PMML file editing method, device, equipment and medium based on decision tree model

By parsing the PMMML file, restoring the decision tree model allows manual editing of node information and generating a new decision tree model, solving the problem that the decision tree cannot be manually modified in the existing technology, and achieving flexible decision tree editing and adjustment.

CN114820157BActive Publication Date: 2025-05-16HANGZHOU FRAUDMETRIX TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210333555.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-30
Publication Date
2025-05-16
Estimated Expiration
2042-03-30

AI Technical Summary

Technical Problem

The modification of the manual participation decision tree model cannot be implemented in the prior art, resulting in the inability to edit decision tree based on manual experience and business needs in the fields of financial risk management.

Method used

By obtaining the pre-save pmmml file, all node information is parsed out to restore the decision tree model, allowing manual editing of the information of the specified node, and continuing to split based on the edited node, generating a new decision tree model, and saving it as a new pmmml file.

Benefits of technology

The modification of the manual participation decision tree model is realized, and the decision tree model can be flexibly adjusted according to different business needs, solving the problem of the inability to manually participate in the modification of the tree.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114820157B_ABST
    Figure CN114820157B_ABST
Patent Text Reader

Abstract

The present application relates to a pmml file editing method, device, equipment and medium based on a decision tree model, and belongs to the field of computer technology. The method includes: obtaining a pre-saved pmml file as a first pmml file; reading the first pmml file, parsing the information of all nodes to restore the first decision tree model; manually editing the node information of the specified node of the first decision tree model; continuing to split based on the edited node to generate a second decision tree model, and saving a second pmml file for the second decision tree model. According to the embodiment of the present application, some node information can be edited according to manual experience and business needs, which effectively solves the problem of being unable to manually participate in the modification of the tree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a PMML file editing method, device, equipment and medium based on a decision tree model. Background Art

[0002] In the field of financial risk management, problems such as credit default are becoming increasingly prominent, and decision tree models can be used to alleviate such problems. The decision tree model calculates the effective distinguishing points of user attributes through the ID3 algorithm, C4.5 (information gain rate) algorithm or gini index, and builds a tree from top to bottom according to the user's attributes. The tree can be regarded as multiple reasonable rules from the root node to the leaf node, so as to evaluate the credit of users and avoid risks.

[0003] During the training process, the decision tree model usually uses a fixed splitting rule, which is generally saved through a pmml (Predictive Model Markup Language) file. According to the pmml specification, only part of the node information of the tree is saved, and not all the node information of the entire tree. In addition, when the pmml file is parsed through Jpmml, the leaf nodes at the bottom of the tree are found, and the leaf node information is analyzed to generate the result. This process uses the decision tree model and does not involve tree modification. For some node information that needs to be edited based on manual experience and business needs, it cannot be realized, that is, manual editing of the tree is impossible. Summary of the invention

[0004] The embodiments of the present application provide a method, apparatus, device and medium for editing a PMML file based on a decision tree model, so as to at least solve the problem in the related art that it is impossible to manually participate in the modification of the tree.

[0005] In a first aspect, an embodiment of the present application provides a pmml file editing method based on a decision tree model, comprising: obtaining a pre-saved pmml file as a first pmml file; reading the first pmml file, parsing the information of all nodes to restore the first decision tree model; manually editing the node information of the specified nodes of the first decision tree model; continuing to split based on the edited nodes to generate a second decision tree model, and saving a second pmml file for the second decision tree model.

[0006] In some of the embodiments, the reading of the first pmml file and parsing of the information of all nodes to restore the first decision tree model include: traversing several Node element sets contained in the TreeModel element and saving them into an array, and parsing each Node element set in turn; when using pre-order traversal parsing, recursively finding the Node elements used to save leaf node information layer by layer, returning after saving the leaf node information, and judging whether the level of the next Node element set is consistent with the current level. If so, floating the node as a sibling node, and using the full set and the left subtree information as a difference set to obtain the right subtree information; when all Node element sets are parsed, the information of all nodes of the entire tree is restored.

[0007] In some embodiments, manually editing the node information of the designated node of the first decision tree model includes: acquiring all samples of the designated node, and clustering them through a one-dimensional clustering algorithm to generate a single-layer split tree.

[0008] In some embodiments, when the designated node is a leaf node, continuing to split based on the edited node to generate a second decision tree model includes: continuing to split based on each node of the single-layer split tree to generate the second decision tree model.

[0009] In some embodiments, saving the second pmml file for the second decision tree model includes: performing a post-order traversal of the tree for the second decision tree model to calibrate the left child node and the right child node of the tree; discarding the right child node information, and continuously sinking its child nodes, and retaining the leaf nodes at the bottom layer.

[0010] In some of the embodiments, before obtaining the pre-saved pmml file as the first pmml file, the method further includes: training a decision tree model through a pipeline and saving the pmml file.

[0011] In some of the embodiments, the pipeline training model is implemented by sklearn or pyspark coding.

[0012] In the second aspect, an embodiment of the present application provides a pmml file editing device based on a decision tree model, including: an acquisition module, a parsing module, an editing module and a generation module, the acquisition module is used to acquire a pre-saved pmml file as a first pmml file; the parsing module is used to read the first pmml file, parse out the information of all nodes to restore the first decision tree model; the editing module is used to manually edit the node information of the specified node of the first decision tree model; the generation module is used to continue splitting based on the edited node, generate a second decision tree model, and save a second pmml file for the second decision tree model.

[0013] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute any of the methods described above.

[0014] In a fourth aspect, an embodiment of the present application provides a storage medium, wherein a computer program is stored in the storage medium, wherein the computer program is configured to execute any of the above methods when run.

[0015] Compared with the related art, although the pre-saved pmml file, i.e., the first pmml file, only saves some nodes, the embodiment of the present application restores the entire tree, i.e., the first decision tree model; then, some designated nodes can be manually edited; then, based on the edited nodes, the nodes are further split to generate a new decision tree, i.e., the second decision tree model; finally, the new pmml file, i.e., the second pmml file, is saved, so the problem of not being able to manually participate in the modification of the tree can be effectively solved. Among them, for the samples of the designated nodes, the K-Means algorithm can be used for clustering to generate a single-layer binary tree or a multi-branch tree, which has high flexibility and can modify the decision tree model according to different business needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0017] Figure 1 It is a schematic diagram of expressing a decision tree structure according to a related art example;

[0018] Figure 2 It is a flow chart of a pmml file editing method based on a decision tree model according to an embodiment of the present application;

[0019] Figure 3 is a code schematic diagram of a DataDictionary in a pmml file according to an embodiment of the present application;

[0020] Figure 4 is a schematic diagram of a Node element in a pmml file according to an embodiment of the present application;

[0021] Figure 5 is a schematic diagram of a Node element set and a Node element included in a TreeModel element according to an embodiment of the present application;

[0022] Figure 6 is a schematic diagram of expressing a left subtree and a right subtree according to an embodiment of the present application;

[0023] Figure 7 is a schematic diagram of a node recovery order of a decision tree according to an embodiment of the present application;

[0024] Figure 8 is a schematic diagram of preserving a multitree variable delineation rule through a CompoundPredicate element according to an embodiment of the present application;

[0025] Fig. 9 is a schematic diagram of a sequence of saving nodes in a second decision tree model according to an embodiment of the present application;

[0026] Fig.10 It is a structural block diagram of a pmml file editing device based on a decision tree model according to an embodiment of the present application;

[0027] Fig.11 It is a schematic diagram of the internal structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments provided in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application. In addition, it can also be understood that although the efforts made in this development process may be complex and lengthy, for ordinary technicians in the field related to the contents disclosed in the present application, some changes such as design, manufacturing or production based on the technical contents disclosed in the present application are only conventional technical means, and should not be understood as insufficient contents disclosed in the present application.

[0029] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those of ordinary skill in the art that the embodiments described in this application may be combined with other embodiments without conflict.

[0030] Unless otherwise defined, the technical terms or scientific terms involved in this application should be understood by people with ordinary skills in the technical field to which this application belongs. The words "one", "a", "a", "the" and the like involved in this application do not indicate a quantitative limitation, and may represent the singular or plural. The terms "include", "comprise", "have" and any of their variations involved in this application are intended to cover non-exclusive inclusions; for example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units that are not listed, or may also include other steps or units inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in this application refers to greater than or equal to two. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships, for example, "A and / or B" can represent: A exists alone, A and B exist at the same time, and B exists alone. The terms "first", "second", "third" and the like involved in the present application are merely used to distinguish similar objects and do not represent a specific ordering of the objects.

[0031] The decision tree model, referred to as "decision tree", can make multi-level strategy selections based on data, with benefits as the basis for judgment. In short, the model is a tree with splitting conditions.

[0032] Figure 1 It is a schematic diagram of a decision tree structure expression according to a related technical example, such as Figure 1 As shown in the figure, the decision tree can be used to perform risk assessment for a financial institution. For example, the number of samples at the root node is 30,000, and the good-bad ratio is 3.5 (23,364 divided by 6,636, where the sample label is 0 for good and the sample label is 1 for bad); pay_0, pay_2, pay_3, and pay_6 are attribute variables, and the attribute values ​​of the samples are assigned according to the repayment situation in September 2005, such as -1 for on-time repayment, 1 for a delay of 1 month, 2 for a delay of 2 months...8 for a delay of 8 months, and 9 for a delay of at least 9 months; bill_amt1 and pay_amt3 are also attributes, indicating loan information. When the samples are split according to the attribute value of 1.5, the samples with pay_0≤1.5 are divided into the left child node, and the good-bad ratio is 5 (22411 divided by 4459); the samples with pay_0>1.5 are divided into the right child node, and the good-bad ratio is 0.44 (953 divided by 2177). It can be seen that the split point with the attribute value of 1.5 is significantly distinguished, and the following node splits are similar.

[0033] It is worth noting that the above Figure 1 The decision tree structure in is obtained by splitting according to a fixed splitting rule, and only some nodes (such as Figure 1 Therefore, as far as the whole tree is concerned, there are some nodes (such as Figure 1 In addition, when parsing the pmml file through Jpmml, the leaf nodes at the bottom of the tree (such as Figure 1 The middle mark c shows that the leaf node information is analyzed to generate the result. This process uses a decision tree model and does not involve tree modification. Therefore, it is impossible to manually participate in the tree modification.

[0034] In order to solve the above problems, the embodiment of the present application provides a pmml file editing method based on a decision tree model. Figure 2 is a flow chart of a pmml file editing method based on a decision tree model according to an embodiment of the present application, such as Figure 2 As shown, the method includes:

[0035] S201: Acquire a pre-saved pmml file as a first pmml file;

[0036] S202: Read the first pmml file, parse the information of all nodes to restore the first decision tree model;

[0037] S203: manually editing node information of a designated node of the first decision tree model;

[0038] S204: Continue splitting based on the edited node to generate a second decision tree model, and save a second pmml file for the second decision tree model.

[0039] According to the above content, although the pre-saved pmml file, i.e., the first pmml file, only saves some nodes, the embodiment of the present application restores the entire tree, i.e., the first decision tree model; then, some specified nodes can be manually edited; then, based on the edited nodes, continue to split to generate a new decision tree, i.e., the second decision tree model; finally, save the new pmml file, i.e., the second pmml file, so it can effectively solve the problem of being unable to manually participate in the modification of the tree.

[0040] In some embodiments, before step S201, a decision tree model is trained through pipeline and a pmml file is saved, specifically in the form of standard PMML 4.1-Tree Models, to obtain the pre-saved pmml file described in step S201. It should be noted that the pipeline training model can be implemented through sklearn or pyspark coding.

[0041] In step S201, as far as the first pmml file is concerned, according to the pmml specification, the important information of the node is stored in the TreeModel element; the attribute variables used by the model are stored in the DataDictionary element. Figure 3 is a code diagram of a DataDictionary in a pmml file according to an embodiment of the present application. For example content of the DataDictionary element, see Figure 3 In the Node element, SimplePredicate is used to save the variable delimitation rules, and ScoreDistribution is used to save the sample quantity and labels. Figure 4 is a schematic diagram of a Node element in a pmml file according to an embodiment of the present application. For example content of the Node element, see Figure 4 .

[0042] In the tree model, non-leaf nodes are saved as Node element sets, and leaf nodes are saved as Node elements. The Node element set is used to encapsulate the variable demarcation rules of the tree model, and the Node element is used to encapsulate the leaf node information. Each Node element set contains a SimplePredicate element to represent the variable demarcation rules, and multiple variable demarcation rules can be represented by CompoundPredicate.

[0043] Specifically, in the TreeModel element, the first Node element set is the left subtree of the entire tree, the second to Nth Node element sets are the right subtree of the entire tree, and the multiple Node element sets or Node elements contained in each Node element set delineate the left subtree and the right subtree in the above manner, and so on, wherein the parent node and its child nodes are stored in a nested form. In the Node element set, SimplePredicate is stored as the index object header, and the stored content is the variable delineation rule of the left subtree of the node. For example, Figure 5 is a schematic diagram of a Node element set and a Node element included in a TreeModel element according to an embodiment of the present application, Figure 6 is a schematic diagram of expressing a left subtree and a right subtree according to an embodiment of the present application. Specifically, Figure 5 In the nested order from outside to inside, there are the first Node element set, the second Node element set and the Node element. The first Node element set is divided into Node element set 1, Node element set 2, Node element set 3 and Node element set 4 from top to bottom. Node element set 1 expresses the left subtree of the whole tree ( Figure 6The number is 1), Node element set 2, Node element set 3 and Node element set 4 represent the right subtree of the whole tree, where Node element set 2, Node element set 3 and Node element set 4 correspond to Figure 6 The numbers 2, 3, and 4 are hidden in the TreeModel element. Figure 6 The expression of nodes b1, b2, and b3 in . And, if Figure 5 As shown, the Node element set 1 also includes the second Node element set. The second Node element set is divided into Node element set (1), Node element set (2) and Node element set (3) from top to bottom. The left subtree and the right subtree are still defined in the above manner. Specifically, Figure 5 In the Node element set (1), the left subtree ( Figure 6 The number is 1.1), Node element set (2) and Node element set (3) express the right subtree ( Figure 6 The number in the middle is 1.2), and so on. The Node element set (1) contains multiple Node elements, which are divided into Node element ① and Node element ②. Among them, Node element ① expresses that the left subtree is a leaf node ( Figure 6 The number in the label is 1.1.1), and the Node element ② expresses that the right subtree is a leaf node ( Figure 6 The winning number is 1.1.2).

[0044] In some of the embodiments, in step S202, jdom2 (jar package) is used to read the first pmml file, instantiate it into xml format, and parse the nodes. Specifically, traverse several Node element sets contained in the TreeModel element and save them in an array, and parse each Node element set in turn; when using pre-order traversal parsing, recursively find the Node element used to save the leaf node information layer by layer, save the leaf node information and return, and judge whether the level of the next Node element set is consistent with the current level. If so, float the node as a brother node, and use the full set and the left subtree information as a difference set to obtain the right subtree information; when all Node element sets are parsed, restore the information of all nodes of the entire tree.

[0045] For example, traversing from top to bottom Figure 5 The Node element set and Node element shown, Figure 7 is a schematic diagram of a node recovery order of a decision tree according to an embodiment of the present application, which can be Figure 7 Nodes ① are parsed in the order shown. . See Figure 5, when traversing the first Node element set (i.e., Node element set 1), we first encounter a SimplePredicate element (i.e., SimplePredicate field = "pay_0" operator = "lessOrEqual" value = "1.5", indicating that the variable delimitation rule is pay_0 ≤ 1.5), and save the node (corresponding to Figure 7 Node ① in the traversal table), continue parsing. If there is no Node element set or Node element, then determine that node ① is a leaf node. If a Node element set or Node element is encountered, recurse the above operation. In this embodiment, when continuing to traverse downward, a Node element set is encountered. At this time, it can be known that node ① is a non-leaf node. Therefore, continue to traverse downward and encounter a SimplePredicate element (i.e., SimplePredicate field = "pay_2" operator = "lessOrEqual" value = "1.5", indicating that the variable demarcation rule is pay_2 ≤ 1.5), save the node (corresponding to Figure 7 Node ② in the traversal, continue to traverse downwards, encounter a Node element, and now we know that node ② is a non-leaf node. Therefore, continue to traverse downwards and encounter a SimplePredicate element (i.e. SimplePredicate field = "pay_amt3" operator = "lessOrEqual" value = "622.5", indicating that the variable delimitation rule is pay_amt3 ≤ 622.5), save the node (corresponding to Figure 7 Node ③ in the traversal, continue to traverse downward, no Node element is encountered, therefore, node ③ is saved as a leaf node, and the label and sample number of leaf node ③ are parsed, that is, there are 5176 "good samples" with label 0, 1433 "bad samples" with label 1, and the total number of samples is 6609. Next, recursively return to determine whether the level of the next Node element is consistent with the current level. If consistent, the node is floated up as a sibling node. In this embodiment, the SimplePredicate field of the next Node element is consistent with the level of the Node element corresponding to node ③ (both are pay_amt3), so the node is floated up as a sibling node of node ③. The floated node is Figure 7 Node ④ in the diagram is also a leaf node, and its label and number of samples can be parsed. For the variable delimitation rule of node ④, since pay_amt3 of SimplePredicate field already exists, the whole set is used to make a difference with the variable delimitation rule of node ③ (pay_amt3≤622.5), and pay_amt3>622.5 is obtained. At this point, the parsing of Node element set (1) is completed.

[0046] then, Figure 5 In the Node element set (2), the node is parsed and the level information pay_6 is judged to be non-existent and different from the current level pay_2, so it is sunk to a leaf node (corresponding to Figure 7 Node ⑤ in the figure, its parent node sinks to become the right child node of node ① (i.e. corresponding to Figure 7 Then, the traversal method is the same as above, parsing the Node element set (3) (corresponding to Figure 7 When node ⑥ in the Node 1 is found, it is used as the right child of node ⑦, and the rule of ⑦ is obtained at the same time. At this point, the parsing of Node element set 1 is completed.

[0047] According to the above method, Figure 5 Node element set 2, Node element set 3 and Node element set 4 are parsed in sequence.

[0048] It should be noted that since the first pmml file does not store nodes ⑦, , The corresponding Node element set, so the embodiment of the present application needs to restore the node information based on the node information that has been parsed. For example, for node ⑦, the whole set and the left subtree information can be used as the difference set, and the variable demarcation rule is pay_2>1.5, and its sample number can be calculated by subtracting the corresponding sample number in node ① from the corresponding sample number in node ②, and the value is the cumulative sample number of its left child node and right child node; node , The information can be obtained in the same way.

[0049] Therefore, according to the above content, when all Node element sets are parsed, the information of all nodes of the entire tree can be restored, that is, the first decision tree model can be restored.

[0050] In some of the embodiments, in step S203, the node information of the designated node can be manually edited. This embodiment provides a method for artificially generating a single-layer split tree, which can cluster the samples of the designated node using a one-dimensional clustering algorithm (such as the K-Means algorithm or the Jenks Natural Breaks algorithm) to generate a single-layer split tree. For example, the K-Means algorithm is used here to first determine the number k of designated nodes, such as leaf nodes, that need to be divided. k = 2 means that a binary tree will be generated, and k greater than 2 means that a multi-branch tree will be generated; then, the IV (Information Value) values ​​of the samples in the leaf nodes are calculated. The IV value is mainly used to encode the input variables and evaluate the predictive ability. The size of the IV value of the characteristic variable indicates the strength of the predictive ability of the variable. Sorting to find the largest IV value can significantly modify the relevant node information of the tree. For a single variable, the K-Means algorithm is used to divide the samples into K categories by distance measurement. Taking the credit risk assessment of financial institutions as an example, pay_0 can be manually divided. For example, the K-Means algorithm is used to divide the sample into two categories. The classification point value is 1, and k=2. Each variable delineation rule can be saved by SimplePredicate, and the preservation of the multi-branch tree borrows the CompoundPredicate element in the pmml standard format. The CompoundPredicate element can save multiple variable delineation rules saved by multiple SimplePredicates. Figure 8 is a schematic diagram of preserving a multitree variable demarcation rule through a CompoundPredicate element according to an embodiment of the present application, such as Figure 8 As shown, the three rules of pay_0≤1.5, 1.5<pay_0≤3, and pay_0>3 are saved, among which the two rules of pay_0≤1.5 and pay_0>3 hide expressions. At this time, k=3, which means that the single-layer split tree has three nodes.

[0051] In some of the embodiments, in step S204, the nodes can be further split based on the edited nodes. For example, after a single-layer split tree is generated at the leaf node, a new tree is created again from each node of the single-layer split tree using the existing splitting algorithm. It should be noted that the C4.5 algorithm can split a binary tree or a multi-branch tree, while the ID3 algorithm or gini cannot split a multi-branch tree. The C4.5 algorithm is a computational method for selecting attributes and splitting in a decision tree. Therefore, a new decision tree model, i.e., the second decision tree model, can be generated. According to this method, the samples of the specified nodes can be clustered by the K-Means algorithm to generate a single-layer binary tree or a multi-branch tree, which is highly flexible and can meet different business needs.

[0052] Then, the second decision tree model is saved again to obtain a new pmml file, namely the second pmml file. In the process of resaving, the second decision tree model is traversed in post-order to mark the left and right child nodes of the tree, and the left child node information of the tree is stored in the SimplePredicate element, the right child node information is discarded, and its child nodes are continuously sunk, and the leaf nodes at the bottom layer are retained. Fig. 9 is a schematic diagram of the order of saving nodes of the second decision tree model according to an embodiment of the present application, such as Fig. 9 As shown, nodes H and I are separated by the attributes in the parent node D, so the new Node element set contains SimplePredicate and saves the attributes of node D. The labels and samples of nodes H and I are stored in ScoreDistribution. A new Node element is created to store nodes J and K (nodes J and K are the bottom nodes). It is determined that the attributes in node E have appeared in node D, and node E is the child node of node B, so node E is deleted as the right node, and then it is merged with the Node element set composed of nodes H and J, and the attribute splitting conditions contained in node B are stored. The right subtree traversal is similar. Finally, the second pmml file can be obtained.

[0053] In summary, the embodiment of the present application generates a tree structure by parsing the pmml text, which can then be manually participated in and edited to complete the adjustment of the decision tree node information, and finally the edited tree structure is saved as a pmml text, where some node information is edited according to manual experience and business needs, effectively solving the problem of being unable to manually participate in the modification of the tree. Moreover, for samples of specified nodes, the K-Means algorithm can be clustered to generate a single-layer binary tree or a multi-branch tree, which is highly flexible and can meet different business needs.

[0054] The embodiment of the present application also provides a pmml file editing device based on a decision tree model, Fig.10 is a structural block diagram of a pmml file editing device based on a decision tree model according to an embodiment of the present application, such as Fig.10 As shown, the device includes: an acquisition module 1, a parsing module 2, an editing module 3 and a generating module 4. The acquisition module 1 is used to acquire a pre-saved pmml file as a first pmml file; the parsing module 2 is used to read the first pmml file, parse out the information of all nodes to restore the first decision tree model; the editing module 3 is used to manually edit the node information of the designated node of the first decision tree model; the generating module 4 is used to continue splitting based on the edited node, generate a second decision tree model, and save a second pmml file for the second decision tree model.

[0055] It should be noted that the above modules can be functional modules or program modules, and can be implemented by software or hardware. For modules implemented by hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.

[0056] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.

[0057] In addition, in combination with the pmml file editing method based on the decision tree model in the above embodiment, the embodiment of the present application can provide a storage medium for implementation. The storage medium stores a computer program; when the computer program is executed by the processor, any one of the pmml file editing methods based on the decision tree model in the above embodiment is implemented.

[0058] In one embodiment of the present application, an electronic device is also provided, which may be a terminal. The electronic device includes a processor, a memory, a network interface, a display screen and an input device connected via a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a PMML file editing method based on a decision tree model is implemented. The display screen of the electronic device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device may be a touch layer covered on the display screen, or a key, trackball or touchpad provided on the housing of the electronic device, or an external keyboard, touchpad or mouse, etc.

[0059] In one embodiment, Fig.11 is a schematic diagram of the internal structure of an electronic device according to an embodiment of the present application, such as Fig.11 As shown, an electronic device is provided, which may be a server, and its internal structure diagram may be as shown in Fig.11 As shown. The electronic device includes a processor, a network interface, an internal memory and a non-volatile memory connected through an internal bus, wherein the non-volatile memory stores an operating system, a computer program and a database. The processor is used to provide computing and control capabilities, the network interface is used to communicate with an external terminal through a network connection, the internal memory is used to provide an environment for the operation of the operating system and the computer program, the computer program is executed by the processor to implement a PMML file editing method based on a decision tree model, and the database is used to store data.

[0060] Those skilled in the art will understand that Fig.11 The structure shown in the figure is merely a block diagram of a partial structure related to the scheme of the present application, and does not constitute a limitation on the electronic device to which the scheme of the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different arrangement of components.

[0061] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0062] Those skilled in the art should understand that the technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0063] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A pmml file editing method based on a decision tree model, characterized in that: include: Obtain a pre-saved pmml file as the first pmml file; Read the first pmml file, parse out information of all nodes to restore the first decision tree model; Manually editing node information of a designated node of the first decision tree model; Continue splitting based on the edited node to generate a second decision tree model, and save a second pmml file for the second decision tree model; Among them, the reading of the first pmml file and parsing the information of all nodes to restore the first decision tree model includes: traversing several Node element sets contained in the TreeModel element and saving them into an array, and parsing each Node element set in turn; when all Node element sets are parsed, restoring the information of all nodes of the entire tree.

2. The method according to claim 1, characterized in that The reading of the first pmml file and parsing the information of all nodes to restore the first decision tree model includes: When using pre-order traversal parsing, recursively find the Node element used to save the leaf node information layer by layer, save the leaf node information and return, and determine whether the level of the next Node element set is consistent with the current level. If so, float the node up as a sibling node, and use the full set and the left subtree information as the difference to get the right subtree information.

3. The method according to claim 1, characterized in that The manually editing the node information of the designated node of the first decision tree model includes: All samples of the designated node are obtained, and clustered by a one-dimensional clustering algorithm to generate a single-layer split tree.

4. The method according to claim 3, characterized in that In the case where the designated node is a leaf node, continuing to split based on the edited node to generate a second decision tree model includes: Based on each node of the single-layer splitting tree, the nodes are further split to generate the second decision tree model.

5. The method according to claim 1, characterized in that: Saving the second pmml file for the second decision tree model comprises: Using post-order traversal of the tree for the second decision tree model, calibrating the left child node and the right child node of the tree; The right child node information is discarded, and its child nodes continue to sink, and the leaf nodes at the bottom layer are retained.

6. The method according to any one of claims 1 to 5, characterized in that: Before obtaining the pre-saved pmml file as the first pmml file, the method further includes: Train the decision tree model through pipeline and save the pmml file.

7. The method according to claim 6, characterized in that The pipeline training process is implemented through sklearn or pyspark coding.

8. A pmml file editing device based on a decision tree model, characterized in that: include: An acquisition module, used for acquiring a pre-saved pmml file as a first pmml file; A parsing module, used for reading the first pmml file, parsing the information of all nodes to restore the first decision tree model; An editing module, used for manually editing node information of a designated node of the first decision tree model; A generation module, used to continue splitting based on the edited node, generate a second decision tree model, and save a second pmml file for the second decision tree model; Among them, when the parsing module reads the first pmml file and parses the information of all nodes to restore the first decision tree model, it is used to traverse several Node element sets contained in the TreeModel element and save them into an array, and parse each Node element set in turn; when all Node element sets are parsed, the information of all nodes of the entire tree is restored.

9. An electronic device, comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 7 when executed.

Citation Information

Patent Citations

  • Decision engine implementation method, device and apparatus and storage medium

    CN112148260A

  • Method and device for decision engine, machine readable storage medium and processor

    CN114217825A