Information processing program, method, and device

By optimizing the selection of determination rules and strategically adding candidate rules to maximize the objective function, the method addresses the challenge of high computational cost in generating rule lists for rule-based determination models, achieving efficient and accurate results.

WO2025126280A1PCT designated stage expired Publication Date: 2025-06-19FUJITSU LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/044283
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-11
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing methods for constructing rule-based determination models face challenges in achieving both high determination accuracy and interpretability while incurring a large computational cost when enumerating multiple rule lists.

Method used

The approach involves extracting an optimal rule list by selecting determination rules that maximize an objective function based on interpretability and accuracy, and then determining whether adding additional candidate rules increases the objective function value, thereby reducing the computational cost of enumerating multiple rule lists.

Benefits of technology

This method reduces the computational cost of enumerating multiple rule lists that achieve both determination accuracy and interpretability, allowing for the efficient generation of approximate solution rule lists close to the optimal solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023044283_19062025_PF_FP_ABST
    Figure JP2023044283_19062025_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device: selects one or more determination rules from among a plurality of determination rules so as to maximize the value of an objective function based on the determination accuracy and the interpretability of the determination rules, and extracts a first rule list; generates, when an additional-candidate determination rule which has been selected from among the plurality of determination rules is added to a candidate set, at least one candidate rule list by adding, if the value of an objective function about a candidate rule list comprising determination rules in the candidate set increases, the additional-candidate determination rule to the candidate set, but not adding, if the value of the objective function does not increase, the additional-candidate determination rule and a determination rule including one or more conditions included in the additional-candidate determination rule, to the candidate set; and outputs the first rule list and the candidate rule list in which the difference between the objective function values of the first rule list and the candidate rule list is equal to or less than a threshold.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing program, method, and device

[0001] The disclosed technology relates to an information processing program, an information processing method, and an information processing device.

[0002] Conventionally, a decision rule for constructing a rule-based decision model has been extracted from among a plurality of decision rules. For example, a device has been proposed that searches for rules with high or low evaluation values ​​by depth-first search among a plurality of rules in a hierarchical relationship. This device searches for candidate rules from a plurality of rules that have a tree structure (hierarchical relationship) in which child rules (rules at a lower level) and sibling rules (rules at the same level) are arranged according to the order of their condition items, in the reverse order of the order in which the tree structure is constructed. The device then compares the evaluation values ​​of the rules in the order found and outputs the rule with the highest evaluation value.

[0003] Furthermore, for example, an information processing device has been proposed that performs the following processing on tree-structured data representing a plurality of judgment rules each including one or more conditions, with the conditions included in each judgment rule as vertices, starting from the highest vertex as an evaluation target. This device evaluates an objective function for a judgment rule set when a judgment rule including a condition corresponding to the vertex to be evaluated is added to a judgment rule set to be extracted. Then, if the evaluation indicated by the objective function monotonically decreases, this device excludes from the evaluation targets judgment rules including conditions corresponding to the vertex to be evaluated and vertices lower than the vertex to be evaluated.

[0004] JP 2005-38073 A International Publication No. 2023 / 188411

[0005] In constructing a rule-based decision model, in order to achieve both the accuracy of the decision made by the decision model and the interpretability of the decision results, it is necessary to extract decision rules that can perform highly accurate decisions with as few and as short a number of decision rules as possible. Conventional techniques can extract a rule list, which is a set of decision rules that is an optimal solution that maximizes the value of an objective function for extracting the decision rules described above.

[0006] However, providing users with only one rule list that is the optimal solution may result in missing useful information. Also, listing and observing rule lists other than the optimal rule list may lead to the discovery of new knowledge. Taking this into consideration, it is envisioned to list multiple rule lists that can achieve both determination accuracy and interpretability.

[0007] However, in this case, it is necessary to repeatedly extract one optimal rule list, discard that rule list, and then extract the next optimal rule list, which poses the problem of high computational costs for listing multiple rule lists.

[0008] In one aspect, the disclosed technology aims to reduce the calculation cost for enumerating multiple rule lists that can achieve both determination accuracy and interpretability.

[0009] In one aspect, the disclosed technology extracts an optimal rule list from a plurality of judgment rules, each of which includes one or more conditions and a label for data that satisfies the one or more conditions. The optimal rule list is extracted by selecting one or more judgment rules from the plurality of judgment rules so as to maximize the value of an objective function based on the interpretability and judgment accuracy of the judgment rules. The disclosed technology also determines whether or not the value of the objective function for a candidate rule list made up of judgment rules in the candidate set increases when an additional candidate judgment rule selected from the plurality of judgment rules is added to a candidate set. If the value of the objective function increases, the disclosed technology adds the additional candidate judgment rule to the candidate set. If the value of the objective function does not increase, the disclosed technology generates one or more candidate rule lists by not adding the additional candidate judgment rule and a judgment rule that includes the one or more conditions included in the additional candidate judgment rule to the candidate set. The disclosed technology then outputs the candidate rule list and the first rule list such that the difference between the objective function value for the first rule list and the objective function value for the candidate rule list is equal to or less than a predetermined threshold.

[0010] One of the advantages is that it is possible to reduce the calculation cost for listing multiple rule lists that can achieve both high determination accuracy and interpretability.

[0011] 1 is a functional block diagram of an information processing device according to the present embodiment; FIG. 2 is a diagram showing an example of a dataset; FIG. 3 is a diagram for explaining a judgment rule; FIG. 4 is a diagram showing an example of a judgment model; FIG. 5 is a diagram for explaining a case where all possible combinations of feature quantities are candidates for a condition part; FIG. 6 is a diagram showing an example of an enumeration tree; FIG. 7 is a diagram for explaining the number of correct answers and the number of incorrect answers; FIG. 8 is a diagram for explaining an example of generating an enumeration tree candidate; FIG. 9 is a diagram showing an example of a plurality of enumeration tree candidates; FIG. 10 is a diagram showing an example of output of a plurality of judgment models; FIG. 11 is a block diagram showing a schematic configuration of a computer functioning as an information processing device; FIG. 12 is a flowchart showing an example of an information processing routine; FIG. 13 is a flowchart showing an example of an enumeration tree candidate generation process;

[0012] Hereinafter, an example of an embodiment of the disclosed technology will be described with reference to the drawings.

[0013] 1, a data set is input to an information processing device 10 according to this embodiment. The information processing device 10 extracts, from the input data set, determination rules for constructing a rule-based determination model, and outputs the extracted rules as a list of multiple rules.

[0014] A dataset includes a plurality of data items, each associated with one or more features and a label. FIG. 2 shows an example of a dataset. In the example of FIG. 2, each row (each record) is a single piece of data, and the items "gender" and their values, "age" and their values, and "past arrest history" and their values ​​are each features. Also, "whether or not the person has recidivism" is a label.

[0015] A rule-based decision model is a model expressed by a rule list, which is a set of decision rules in the form "if A then B" (if A, then prediction is B), and a rule definition. As shown in FIG. 3, the part A in the decision rule is called the "condition part" and the part B is called the "label." In other words, a decision rule r is a set (s, c)∈U×C of a condition part s and a label c. U and C will be described later.

[0016] The condition part s is a combination of one or more feature quantities included in the data set. Specifically, all feature quantities X={x 1 , x 2 , ..., x M} (M is the total number of features), the condition part s is s∈U. In this embodiment, the condition part s is the logical product of the combination of features. For example, the first feature x 1 , the fourth feature x 4 , and the fifth feature x 5 The condition part s consisting of a combination of x 1 ∧x 4 ∧x 5 As described above, a feature is a combination of an item included in a dataset and its value, such as "gender = 1 (male)." If the value of an item is a numeric value, a discretized value may be used, or the feature may be expressed using an inequality sign, such as "age > 40 years old."

[0017] The label c represents a predicted value for classifying data that satisfies the condition indicated by the condition part, and if the set of labels is C, then the label c is c∈C. For example, when predicting whether or not there will be a recidivism from data such as the dataset shown in FIG. 2 , one of the set of labels C={1 (recidivism has occurred), 0 (recidivism has not occurred)} indicating whether or not there will be a recidivism is set as the label c, similar to the labels of the dataset. Note that the label c of the judgment rule r does not need to be the same as the label of the dataset. For example, if the labels included in the dataset are E, F, and G, the label c may be set to "1" if the predicted value is E, and to "2" if the predicted value is F or G.

[0018] FIG. 4 shows an example of a judgment model. In the example of FIG. 4, the dashed line portion is a rule list, and the remaining portion is a rule definition. If the data to be judged satisfies the condition indicated by the condition part of a judgment rule, the judgment model acquires the label of the judgment rule as a predicted value and outputs a judgment result for the data to be judged based on the predicted value based on each judgment rule and the rule definition. For example, if the rule definition specifies that the judgment result is determined by majority vote of the predicted values ​​based on each judgment rule, the judgment model outputs the most popular predicted value ("0" or "1" in the example of FIG. 4) among the predicted values ​​based on each judgment rule as the judgment result. Note that in FIG. 4, "default 1" is an example of a rule definition, and indicates that if the data to be judged does not satisfy the condition indicated by the condition part of any judgment rule included in the rule list, the label "1" is output as the judgment result.

[0019] Here, when constructing a judgment model that is guaranteed to be optimal and has high interpretability, the rule list must (1) be composed of a small number of judgment rules that are short, and (2) be able to predict with high accuracy. There is a trade-off between (1) and (2). Therefore, it is desirable to construct a judgment model that optimizes the value (hereinafter referred to as the "objective function value") obtained by evaluating the rule list using an objective function that is integrated according to the ratio of importance given to (1) and (2).

[0020] Therefore, as shown in Figure 5, it is conceivable to set all possible combinations of feature quantities as candidates for condition parts based on a data set, evaluate all the candidate condition parts, extract the optimal condition part, and create a rule list as a set of decision rules including the extracted condition part. In this case, since all condition parts are evaluated, it is possible to construct an optimal decision model that satisfies the above (1) and (2). However, the number of all condition parts, i.e., the number of all combinations of feature quantities, is 2, where n is the number of feature quantities. nSince there is -1 condition, combinatorial explosion occurs when there are a large number of features. For example, if there are 30 features, there are approximately 1 billion condition parts, and if there are 50 features, there are approximately 1,000 trillion condition parts, and it is impossible to evaluate all of the condition parts in a realistic amount of time. Furthermore, among all the condition parts, there are several condition parts that have a low contribution to the actual judgment, so considering such condition parts also results in unnecessary processing load.

[0021] Furthermore, when multiple rule lists are output, it is necessary to repeatedly extract one optimal rule list, discard that rule list, and then extract the next optimal rule list, which increases the calculation cost.

[0022] Therefore, the information processing device 10 according to this embodiment quickly calculates a decision rule that includes an overwhelmingly small number of condition parts relative to the total number of possible condition parts, without missing any condition parts included in the rule list that is the optimal solution. Furthermore, the information processing device 10 constructs a rule list of objective function values ​​whose difference from the objective function value evaluated in the rule list that is the optimal solution is within a predetermined value ε, thereby outputting multiple rule lists that are approximate solutions to the optimal solution. The information processing device 10 according to this embodiment will be described in detail below.

[0023] As shown in FIG. 1 , the information processing device 10 functionally includes an extraction unit 12, a generation unit 14, and an output unit 16.

[0024] The extraction unit 12 acquires a dataset input to the information processing device 10. The extraction unit 12 selects one or more determination rules from a plurality of determination rules comprehensively generated from the dataset so as to maximize an objective function value based on the interpretability and determination accuracy of the determination rules, and extracts a first rule list (optimal rule list). For example, conventional techniques such as the method disclosed in Patent Document 2 may be applied to the extraction of the optimal rule list by the extraction unit 12.

[0025] In this embodiment, a condition part s is extracted from U so as to optimize an objective function value that indicates that (1) the rule list is composed of few and short decision rules, and (2) the prediction by each decision rule is highly accurate, and a decision rule r including the extracted condition part s is extracted. In this embodiment, the following four scores are defined as indicators that become larger the better the rule list composed of the extracted decision rule r is, and the sum of these scores is used as the objective function value, and a rule list that maximizes the objective function value is extracted.

[0026] 1. The fewer the number of decision rules included in the rule list, the higher the score. 2. The shorter the length of the decision rules included in the rule list, the higher the score. 3. The more data that the decision rule correctly predicts, the higher the score. 4. The fewer data that the decision rule incorrectly predicts, the higher the score.

[0027] The generation unit 14 calculates an objective function value for a candidate rule list consisting of the judgment rules in the candidate set when an additional candidate judgment rule selected from a plurality of judgment rules based on a data set input to the information processing device 10 is added to the candidate set. If the objective function value increases, the generation unit 14 adds the additional candidate judgment rule to the candidate set. On the other hand, if the objective function value does not increase, the generation unit 14 does not add the additional candidate judgment rule and a judgment rule that includes a judgment unit included in the additional candidate judgment rule to the candidate set. Then, the generation unit 14 generates a candidate rule list from the judgment rules in the candidate set.

[0028] Specifically, the generating unit 14 generates an enumeration tree in which each condition part included in the judgment rule is a vertex. As shown in Fig. 6, in the enumeration tree, the condition part corresponding to each vertex is a condition part obtained by adding one feature to the condition part corresponding to the parent vertex of that vertex. Note that Fig. 6 shows the enumeration tree in which X = {x 1 , x 2 , x 3}, where each square represents a vertex and the character string inside the square represents the condition part corresponding to that vertex. Note that a vertex with a blank inside the square is a root vertex, i.e., a vertex that does not have a parent vertex.

[0029] More specifically, the generation unit 14 determines whether the objective function value for the enumeration tree candidate increases when an additional candidate vertex is added to the enumeration tree candidate in the order of depth-first search relative to the root vertex. If the objective function value increases, the generation unit 14 adds the additional candidate vertex to the enumeration tree candidate. On the other hand, if the objective function value does not increase, the generation unit 14 does not add the additional candidate vertex and descendant vertices of the additional candidate vertex to the enumeration tree candidate. The generation unit 14 generates multiple enumeration tree candidates by adding different combinations of vertices to the enumeration tree candidates.

[0030] For example, the generation unit 14 calculates an increment B of the objective function value when a candidate vertex is added to the enumeration tree, for example, by the following equation (1).

[0031] B = -λ size -λ len ・|Length of decision rule|-λ prec ・|Number of incorrect answers|+λ recall ・|Number of correct answers| ... (1)

[0032] -λ size represents the penalty for adding one additional candidate vertex to the enumeration tree candidate. len |Length of the decision rule| represents a penalty proportional to the length of the decision rule corresponding to the added vertex. The length of the decision rule is the number of features included in the condition part of the decision rule. prec |Number of incorrect answers| represents a penalty proportional to the number of data items for which the prediction is incorrect due to the judgment rule corresponding to the added vertex. recall |Number of correct answers| represents the reward proportional to the number of data whose prediction is correct according to the decision rule corresponding to the added vertex. size , λ len , λ prec , and λ recall Each of the above is a parameter that indicates the importance of each item. The importance may be a predetermined value or may be received from the user.

[0033] The number of correct answers corresponds to the number of data included in "correct_cover(r)" shown in FIG. 7, and the number of incorrect answers corresponds to the number of data included in "incorrect_cover(r)" shown in FIG. 7. In FIG. 7, D is a set of all data included in the dataset, and cover(r) is a set of data that satisfies the condition part s of the determination rule r = (s, c) corresponding to the vertex of the additional candidate. Also, in FIG. 7, circles correspond to each piece of data, and differences in color (density) within the circles represent differences in the labels of the data. If the specific label to be determined in the determination model is y, the label of the data corresponding to the white circle is y, and the labels of the data corresponding to the hatched circle and the black circle are other than y. Therefore, when each piece of data is a pair (x, c) of a feature x and a label c, and the specific label is y, then correct_cover(r) = {(x, y) ∈ cover(r) | c = y}. Also, incorrect_cover(r)=cover(r)\correct_cover(r).

[0034] Furthermore, the generating unit 14 calculates B' shown in the following formula (2) as an evaluation value satisfying monotonicity that decreases as the length of the determination rule to be added to the rule list increases.

[0035] B' = -λ size -λ len ・|Length of decision rule| + λ recall ・|Number of correct answers| ... (2)

[0036] B' is the increment B of the objective function value shown in equation (1) minus λ prec - This is the value excluding the penalty for |number of incorrect answers|. Because B' has one less penalty than B, B'≧B always holds. Therefore, B' can be estimated as the upper bound of the increment in the objective function value, that is, the maximum value of the increment B in the objective function value when the judgment rule to be evaluated is added to the rule list. In other words, B' is the increment in the objective function value when it is assumed that adding the judgment rule to be evaluated to the rule list will result in the greatest improvement in the performance of the judgment model constructed using the rule list.

[0037] The longer the length of the judgment rule, the smaller the range covering the condition part of the judgment rule ("cover(r)" in FIG. 7) tends to be. Therefore, as the length of the judgment rule increases, the number of incorrect answers decreases, but the number of correct answers also decreases. In equation (2), -λ len ・|Length of decision rule| increases, and +λ recall The number of correct answers decreases. In other words, B' has a monotonicity in that it continues to decrease as the length of the judgment rule increases. Therefore, if B' for a judgment rule to be evaluated becomes a negative value, adding a longer judgment rule to the rule list will not improve the objective function value. Therefore, there is no need to add judgment rules longer than that judgment rule to the rule list.

[0038] The generation unit 14 excludes the vertex of the candidate to be added and any vertices lower than the vertex of the candidate to be added from the candidates to be added when the objective function value monotonically decreases, and adds the vertex of the candidate to be added to the enumeration tree candidates when the objective function value increases.

[0039] Specifically, when the upper bound B' of the increment for a vertex of a candidate for addition is equal to or less than 0, the generation unit 14 prunes the vertex of the candidate for addition and its descendant vertices from the enumeration tree candidates, thereby excluding the vertex from the candidates for addition. Furthermore, when the increment B of the objective function value for a vertex that has not been excluded from the candidates for addition is a positive value, the generation unit 14 adds the vertex to the enumeration tree candidates.

[0040] An example of generating an enumeration tree candidate will be described with reference to Fig. 8. In the example of Fig. 8, the order of search by depth-first search from the root vertex is written within each vertex. The increment of the objective function value for the i-th vertex (hereinafter referred to as "vertex i") in the depth-first search order is written as B(i), and the upper bound of the increment is written as B'(i).

[0041] The generation unit 14 selects vertex 1 as a candidate for addition as a child vertex of the root vertex, and determines whether or not to add vertex 1 based on the increment B(1) of the objective function value and the upper bound B'(1) of the increment. If B(1) > 0 and B'(1) > 0, vertex 1 is added to the enumeration tree candidates (H in FIG. 8 ). Next, the generation unit 14 selects vertex 2, a child vertex of vertex 1, as a candidate for addition. If B'(2) ≦ 0, vertex 2 is subject to pruning, and therefore the generation unit 14 does not add vertex 2 or any vertex that is a descendant of vertex 2 to the enumeration candidates (I in FIG. 8 ). Next, the generation unit 14 selects vertex 3, a child vertex of vertex 1, as a candidate for addition. If B(3) > 0 and B'(3) > 0, vertex 3 is added to the enumeration tree candidates (J in FIG. 8 ).

[0042] Next, if there is no child vertex of vertex 3, the generation unit 14 selects vertex 4, which is a child vertex of the root vertex, as an additional candidate. Here, if B'(4) ≤ 0, vertex 4 is subject to pruning, so the generation unit 14 does not add vertex 4 or its descendants to the enumeration candidates (K in FIG. 8 ). Next, the generation unit 14 selects vertex 5, which is a child vertex of the root vertex, as an additional candidate. Here, if B(5) > 0 and B'(5) > 0, vertex 5 is added to the enumeration tree candidates (L in FIG. 8 ). In this way, by pruning vertices that do not increase the objective function value, it is no longer necessary to determine whether or not to add their descendant vertices to the enumeration tree candidates, thereby reducing the processing load of generating enumeration tree candidates.

[0043] The shape of the enumeration tree candidate changes depending on which determination rule a vertex corresponding to the added vertex is to be added. Therefore, the generation unit 14 generates multiple enumeration tree candidates, for example, as shown in FIG. 9, by setting vertices corresponding to different determination rules as the first candidate vertex to be added to the root vertex (H in FIG. 8).

[0044] The output unit 16 outputs the candidate rule list and the optimal solution of the rule list such that the difference between the objective function value for the optimal solution of the rule list and the objective function value for the candidate rule list is equal to or less than a predetermined threshold ε.

[0045] Specifically, for each enumeration tree candidate generated by the generation unit 14, the output unit 16 calculates an objective function value g, which is the sum of the above-mentioned four scores, for a set of determination rules corresponding to each vertex included in the enumeration tree candidate. If the difference between the objective function value f calculated when the extraction unit 12 calculates the rule list of the optimal solution and the objective function value g of the enumeration tree candidate is equal to or less than ε, the output unit 16 generates a rule list of an approximate solution to the optimal solution from the enumeration tree candidate. The output unit 16 constructs a determination model by assigning rule definitions to each of the rule list of the optimal solution and the rule list of the approximate solution, and outputs a plurality of determination models as shown in FIG.

[0046] The information processing device 10 may be realized by, for example, a computer 40 shown in FIG. 11 . The computer 40 includes a CPU (Central Processing Unit) 41, a GPU (Graphics Processing Unit) 42, a memory 43 as a temporary storage area, and a non-volatile storage device 44. The computer 40 also includes an input / output device 45 such as an input device and a display device, and an R / W (Read / Write) device 46 that controls reading and writing of data from and to a storage medium 49. The computer 40 also includes a communication I / F (Interface) 47 that is connected to a network such as the Internet. The CPU 41, GPU 42, memory 43, storage device 44, input / output device 45, R / W device 46, and communication I / F 47 are connected to one another via a bus 48.

[0047] The storage device 44 is, for example, a hard disk drive (HDD), a solid state drive (SSD), a flash memory, etc. The storage device 44 serving as a storage medium stores an information processing program 50 for causing the computer 40 to function as the information processing device 10. The information processing program 50 includes an extraction process control instruction 52, a generation process control instruction 54, and an output process control instruction 56.

[0048] The CPU 41 reads the information processing program 50 from the storage device 44, expands it in the memory 43, and sequentially executes the control instructions contained in the information processing program 50. The CPU 41 operates as the extraction unit 12 shown in FIG. 1 by executing the extraction process control instruction 52. The CPU 41 also operates as the generation unit 14 shown in FIG. 1 by executing the generation process control instruction 54. The CPU 41 also operates as the output unit 16 shown in FIG. 1 by executing the output process control instruction 56. As a result, the computer 40 that has executed the information processing program 50 functions as the information processing device 10. The CPU 41 that executes the program is hardware. Part of the program may also be executed by the GPU 42.

[0049] The functions realized by the information processing program 50 may be realized by, for example, a semiconductor integrated circuit, more specifically, an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or the like.

[0050] Next, the operation of the information processing device 10 according to this embodiment will be described. When a data set is input to the information processing device 10 and an instruction to output a plurality of determination programs is given, the information processing device 10 executes an information processing routine shown in Fig. 12. Note that the information processing routine is an example of an information processing method of the disclosed technology.

[0051] In step S10, the extraction unit 12 acquires a data set input to the information processing device 10. Next, in step S12, the extraction unit 12 selects one or more determination rules from among a plurality of determination rules based on the data set so as to maximize an objective function value based on the interpretability and determination accuracy of the determination rule, and extracts an optimal rule list. Note that the objective function value evaluated for the optimal rule list is denoted by f.

[0052] Next, in step S20, an enumeration tree candidate generation process is executed. The enumeration tree candidate generation process will now be described with reference to FIG.

[0053] In step S22, the generation unit 14 prepares a vertex set storing vertices corresponding to the condition parts included in each of a plurality of determination rules based on the data set, and an enumeration tree candidate consisting only of the root vertex. Next, in step S24, the generation unit 14 selects one vertex as an additional candidate from the vertex set in the order of the depth search direction, and calculates an increment B and an upper bound B' of the increment in the objective function value when the additional candidate vertex is added to the enumeration tree candidate.

[0054] Next, in step S26, the generation unit 14 determines whether the upper bound B' of the increment of the objective function value is greater than 0. If B'>0, the process proceeds to step S30, and if B'≦0, the process proceeds to step S28. In step S28, the generation unit 14 excludes the vertices that are candidates for addition and the descendant vertices of the vertices that are candidates for addition from the vertex set, and the process proceeds to step S34.

[0055] In step S30, the generation unit 14 determines whether the increment B of the objective function is greater than 0. If B>0, the process proceeds to step S32, and if B≦0, the process proceeds to step S34. In step S32, the generation unit 14 adds the vertices that are candidates for addition to the enumeration tree candidate. Next, in step S34, the generation unit 14 determines whether the vertex set is empty. If the vertex set is empty, the process proceeds to step S36, and if unprocessed vertices remain in the vertex set, the process proceeds to step S24.

[0056] In step S36, the generation unit 14 stores the generated enumeration tree candidate in a predetermined storage area, and the enumeration tree candidate generation process ends. The generation unit 14 sets vertices with different corresponding determination rules as vertices to be initially selected as additional candidates in step S24, and executes the enumeration tree candidate generation process shown in Figure 13 to generate multiple enumeration tree candidates, and then returns to the information processing routine (Figure 12).

[0057] Next, in step S40, the output unit 16 selects one of the enumeration tree candidates generated in step S20. Next, in step S42, the output unit 16 calculates the objective function value g of the selected enumeration tree candidate. Next, in step S44, the output unit 16 calculates the difference f-g between the objective function value f of the optimal rule list calculated in step S12 and the objective function value g of the enumeration tree candidate calculated in step S42. Then, the output unit 16 determines whether the difference f-g is equal to or less than a threshold ε. If f-g≦ε, the process proceeds to step S46, and if f-g>ε, the process proceeds to step S48.

[0058] In step S46, the output unit 16 generates a rule list from the enumeration tree candidates. Next, in step S48, the output unit 16 determines whether all enumeration tree candidates have been selected. If there are unselected enumeration tree candidates, the process returns to step S40. If all enumeration tree candidates have been selected, the process proceeds to step S50.

[0059] In step S50, the output unit 16 constructs and outputs a plurality of determination models by adding rule definitions to the optimal rule list extracted in step S12 and to each of the plurality of rule lists generated in step S46. Then, the information processing routine ends.

[0060] As described above, the information processing device according to this embodiment selects one or more determination rules from a plurality of determination rules based on a dataset so as to maximize an objective function value based on the interpretability and determination accuracy of the determination rules, thereby extracting an optimal rule list. The information processing device also determines whether or not the objective function value of a candidate rule list made up of the determination rules in the candidate set increases when a candidate determination rule selected from the plurality of determination rules is added to a candidate set. If the objective function value increases, the information processing device adds the candidate determination rule to the candidate set. If the objective function value does not increase, the information processing device does not add the candidate determination rule and a determination rule that includes one or more conditions included in the candidate determination rule to the candidate set. The information processing device then executes a process of generating a candidate rule list from the determination rules included in the candidate set multiple times to generate multiple candidate rule lists. Furthermore, the information processing device outputs a candidate rule list and an optimal rule list in which the difference between the objective function value of the optimal rule list and the objective function value of the candidate rule list is equal to or less than a predetermined threshold. As a result, the information processing device according to this embodiment can output the optimal solution of the rule list and the rule list of an approximate solution in which the difference between the optimal solution and the objective function is at most ε, without evaluating all rule lists. In other words, it is possible to reduce the calculation cost for listing multiple rule lists that can achieve both determination accuracy and interpretability.

[0061] In the above embodiment, the increment B of the objective function includes four scores, but this is not limiting. It is sufficient to use a score that can be set to B', which is always larger than the increment B and has the monotony of continuing to decrease as the length of the added judgment rule increases. For example, from the above formulas (1) and (2), -λ size and +λ recall B and B' may be used with at least one of |number of correct answers| removed.

[0062] In the above embodiment, the information processing program is stored (installed) in advance in a storage device, but this is not limiting. The program according to the disclosed technology may be provided in a form stored in a storage medium such as a CD-ROM, a DVD-ROM, or a USB memory.

[0063] REFERENCE SIGNS LIST 10 Information processing device 12 Extraction unit 14 Generation unit 16 Output unit 40 Computer 41 CPU 42 GPU 43 Memory 44 Storage device 45 Input / output device 46 R / W device 47 Communication I / F 48 Bus 49 Storage medium 50 Information processing program 52 Extraction process control command 54 Generation process control command 56 Output process control command

Claims

1. Select one or more determination rules from a plurality of determination rules each including one or more conditions and a label for data corresponding to the one or more conditions so as to maximize the value of an objective function based on the interpretability and determination accuracy of the determination rules, extract a first rule list, and when an additional candidate determination rule selected from the plurality of determination rules is added to a candidate set, if the value of the objective function for a candidate rule list composed of the determination rules in the candidate set increases, add the additional candidate determination rule to the candidate set, and if the value of the objective function does not increase, do not add the additional candidate determination rule and a determination rule including the one or more conditions included in the additional candidate determination rule to the candidate set, generate one or more of the candidate rule lists, and output the candidate rule list and the first rule list when the difference between the value of the objective function for the first rule list and the value of the objective function for the candidate rule list is equal to or less than a predetermined threshold. An information processing program for causing a computer to execute a process including this.

2. The objective function includes a first term for making the value of the objective function smaller as the length of the conditions included in the determination rule is longer, and a second term for making the value of the objective function smaller as the number of incorrect answers satisfying the conditions of the determination rule is larger. When the value obtained by removing the second term from the objective function is greater than 0, it is determined that the value of the objective function increases. The information processing program according to claim 1.

3. The objective function further includes a third term for making the value of the objective function smaller as the number of determination rules included in the candidate set is larger, and a fourth term for making the value of the objective function larger as the number of correct answers satisfying the conditions of the determination rule is larger. The information processing program according to claim 2.

4. An enumeration tree candidate having each determination rule included in the candidate set as a vertex, and based on the inclusion relationship of the conditions of the determination rule corresponding to the vertex, when adding an additional candidate vertex to the enumeration tree candidate connected by a parent-child relationship between the vertices, if the value of the objective function for the enumeration tree candidate increases, add the vertex of the additional candidate to the enumeration tree candidate; if the value of the objective function does not increase, do not add the vertex of the additional candidate and the vertices of the descendants of the vertex of the additional candidate to the enumeration tree candidate, and generate a plurality of the enumeration tree candidates by making the combinations of the vertices added to the enumeration tree candidate different from each other. The information processing program according to any one of claims 1 to 3.

5. From a plurality of determination rules each including one or more conditions and a label for data corresponding to the one or more conditions, select one or more determination rules so as to maximize the value of the objective function based on the interpretability and determination accuracy of the determination rules, and extract a first rule list. When adding a determination rule of an additional candidate selected from the plurality of determination rules to a candidate set, if the value of the objective function for a candidate rule list composed of the determination rules in the candidate set increases, add the determination rule of the additional candidate to the candidate set; if the value of the objective function does not increase, do not add the determination rule of the additional candidate and the determination rules including the one or more conditions included in the determination rule of the additional candidate to the candidate set, and generate one or more of the candidate rule lists. Output the candidate rule list and the first rule list when the difference between the value of the objective function for the first rule list and the value of the objective function for the candidate rule list is less than or equal to a predetermined threshold. An information processing method in which a computer executes a process including this.

6. The objective function includes a first term for making the value of the objective function smaller as the length of the conditions included in the determination rule is longer, and a second term for making the value of the objective function smaller as the number of incorrect answers satisfying the conditions of the determination rule is larger. When the value obtained by removing the second term from the objective function is greater than 0, determine that the value of the objective function increases. The information processing method according to claim 5.

7. The information processing method according to claim 6, wherein the objective function further includes a third term for making the value of the objective function smaller as the number of determination rules included in the candidate set is larger, and a fourth term for making the value of the objective function larger as the number of correct answers satisfying the conditions of the determination rules is larger.

8. An enumeration tree candidate having each determination rule included in the candidate set as a vertex, wherein, based on the inclusion relationship of the conditions of the determination rules corresponding to the vertices, when adding an additional candidate vertex to the enumeration tree candidate in which the vertices are connected in a parent-child relationship, if the value of the objective function for the enumeration tree candidate increases, the additional candidate vertex is added to the enumeration tree candidate, and if the value of the objective function does not increase, the additional candidate vertex and the descendants of the additional candidate vertex are not added to the enumeration tree candidate, and by making the combinations of vertices added to the enumeration tree candidate different from each other, a plurality of the enumeration tree candidates are generated. The information processing method according to any one of claims 5 to 7.

9. An extraction unit that selects one or more determination rules from a plurality of determination rules each including one or more conditions and a label for data corresponding to the one or more conditions so as to maximize the value of an objective function based on the interpretability and determination accuracy of the determination rules, and extracts a first rule list; a generation unit that, when adding a determination rule of an additional candidate selected from the plurality of determination rules to a candidate set, if the value of the objective function for a candidate rule list composed of the determination rules in the candidate set increases, adds the determination rule of the additional candidate to the candidate set, and if the value of the objective function does not increase, does not add the determination rule of the additional candidate and the determination rules including the one or more conditions included in the determination rule of the additional candidate to the candidate set, and generates one or more of the candidate rule lists; and an output unit that outputs the candidate rule list and the first rule list when the difference between the value of the objective function for the first rule list and the value of the objective function for the candidate rule list is equal to or less than a predetermined threshold. An information processing apparatus including the above.

10. The objective function includes a first term for making the value of the objective function smaller as the length of the conditions included in the determination rule is longer, and a second term for making the value of the objective function smaller as the number of incorrect answers satisfying the conditions of the determination rule is larger. The generation unit determines that the value of the objective function increases when the value obtained by removing the second term from the objective function is greater than 0. The information processing apparatus according to claim 9.

11. The objective function further includes a third term for making the value of the objective function smaller as the number of determination rules included in the candidate set is larger, and a fourth term for making the value of the objective function larger as the number of correct answers satisfying the conditions of the determination rule is larger. The information processing apparatus according to claim 10.

12. The generation unit is an enumeration tree candidate having each determination rule included in the candidate set as a vertex. Based on the inclusion relationship of the conditions of the determination rule corresponding to the vertex, when the value of the objective function for the enumeration tree candidate increases by adding an additional candidate vertex by connecting the vertices in a parent-child relationship, the additional candidate vertex is added to the enumeration tree candidate. When the value of the objective function does not increase, the additional candidate vertex and the descendant vertices of the additional candidate vertex are not added to the enumeration tree candidate. By making the combinations of vertices added to the enumeration tree candidate different from each other, a plurality of the enumeration tree candidates are generated. The information processing apparatus according to any one of claims 9 to 11.

13. From a plurality of determination rules each including one or more conditions and a label for data corresponding to the one or more conditions, select one or more determination rules so as to maximize the value of an objective function based on the interpretability and determination accuracy of the determination rules, extract a first rule list, and when an additional candidate determination rule selected from the plurality of determination rules is added to a candidate set, if the value of the objective function for a candidate rule list composed of the determination rules in the candidate set increases, add the additional candidate determination rule to the candidate set, and if the value of the objective function does not increase, do not add the additional candidate determination rule and a determination rule including the one or more conditions included in the additional candidate determination rule to the candidate set, generate one or more of the candidate rule lists, and output the candidate rule list and the first rule list for which the difference between the value of the objective function for the first rule list and the value of the objective function for the candidate rule list is equal to or less than a predetermined threshold. A non-transitory storage medium storing an information processing program for causing a computer to execute a process including this.

Citation Information

Patent Citations

  • Feature rule finding method by width priority search

    JP2004326641A

  • Computer system and model learning method

    JP2024009471A

  • Accurate and interpretable rules for user segmentation

    US20220058503A1

  • Inference system, information processing system, inference method, and recording medium

    WO2018025288A1

  • Determination rule extraction program, device, and method

    WO2023188411A1