Information processing program, information processing method, and information processing device
By classifying the rules of the machine learning model into clusters and identifying the relationship between clusters, the problem of users' difficulty in understanding the multi-rule model is solved, and a more intuitive mastery of model content is achieved.
Patent Information
- Application Number
- CN202280102613.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-16
- Publication Date
- 2025-07-22
AI Technical Summary
It is difficult for users to easily master multiple rules contained in machine learning models, and as the number of rules increases, the workload and time requirements increase.
By calculating the eigenvalues of the rules, classifying the rules into multiple clusters based on the similarity of the eigenvalues, identifying the inclusion and hierarchical relationships between the clusters, and outputting information on these relationships to help users understand the model.
Make it easier for users to master the content of the model, understand the relationship between rules, and reduce the workload and time to understand complex models.
Smart Images

Figure CN120359528A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing program, an information processing method, and an information processing device. Background Art
[0002] Conventionally, there has been a technique for generating a model by machine learning. The model has a plurality of rules, each rule representing a condition for classifying data using one or more explanatory variables, and the model outputs a label representing the result of classifying data including a plurality of input explanatory variables.
[0003] As an existing technique, for example, there is a technique for generating rules describing the operation of a black-box machine learning model by combining conditions identified based on the output of a surrogate black-box model that simulates the operation of the black-box machine learning model.
[0004] Prior Art Documents
[0005] Document 1: U.S. Patent Application Publication No. 2019 / 0147369. Summary of the Invention
[0006] Technical Problem to be Solved by the Invention
[0007] However, in the prior art, it is difficult for a user to easily grasp the content of the model. For example, in the case of outputting a list of the plurality of rules that the model has for the user to refer to, the larger the number of rules included in the model, the more likely the workload and working time required for the user to grasp the content of the model will increase.
[0008] In one aspect, the present invention aims to make it easier to grasp the content of the model.
[0009] Means for Solving the Technical Problem
[0010] According to one aspect, an information processing program, an information processing method, and an information processing device are provided to perform processing including: obtaining a plurality of data including values of a plurality of explanatory variables; obtaining a rule set including rules, where a rule represents a condition for classifying the plurality of data using one or more of the plurality of explanatory variables; calculating a plurality of feature values, each of the plurality of feature values being calculated based on one or more of the obtained plurality of data for a corresponding rule in the rule set, and the one or more data satisfying the condition represented by the corresponding rule in the rule set; classifying the rules in the rule set into any of a plurality of clusters based on the similarity between the feature values among the calculated plurality of feature values; identifying an inclusion relationship between the clusters in the plurality of clusters based on the conditions represented by the rules classified into each of the plurality of clusters; identifying a hierarchical relationship between the clusters in the plurality of clusters based on the identified inclusion relationship between the clusters; and outputting information indicating the hierarchical relationship between the clusters in the identified plurality of clusters.
[0011] Technical effects of the present invention
[0012] According to one aspect, the content of the model can be grasped more easily. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is an explanatory diagram depicting an example of an information processing method according to an embodiment.
[0014] Figure 2 is an explanatory diagram depicting an example of an information processing system 200.
[0015] Figure 3 is a block diagram depicting an example of the hardware configuration of an information processing device 100.
[0016] Figure 4 is a block diagram depicting an example of the functional configuration of an information processing device 100.
[0017] Figure 5 is an explanatory diagram depicting an example of the operation flow of an information processing device 100.
[0018] Figure 6 is an explanatory diagram depicting an example of learning a model 600.
[0019] Figure 7 is an explanatory diagram depicting an example of generating a feature vector for each rule.
[0020] Figure 8 is an explanatory diagram depicting an example of identifying an inclusion relationship between clusters.
[0021] Figure 9 is an explanatory diagram depicting an example of identifying an inclusion relationship between clusters.
[0022] Figure 10 It is an explanatory diagram depicting an example showing the hierarchical relationship between clusters.
[0023] Figure 11 It is an explanatory diagram depicting a specific example of learning the model.
[0024] Figure 12 It is an explanatory diagram depicting a specific example of learning the model.
[0025] Figure 13 It is an explanatory diagram depicting a specific example of generating feature vectors related to rules respectively.
[0026] Figure 14 It is an explanatory diagram depicting a specific example of generating feature vectors related to rules respectively.
[0027] Figure 15 It is an explanatory diagram depicting a specific example of clustering a rule set.
[0028] Figure 16 It is an explanatory diagram depicting a specific example of clustering a rule set.
[0029] Figure 17 It is an explanatory diagram depicting a first specific example of generating explanation information related to clusters.
[0030] Figure 18 It is an explanatory diagram depicting a first specific example of generating explanation information related to clusters.
[0031] Figure 19 It is an explanatory diagram depicting a second specific example of generating explanation information for clusters.
[0032] Figure 20 It is an explanatory diagram depicting a second specific example of generating explanation information for clusters.
[0033] Figure 21 It is an explanatory diagram depicting a specific example of calculating the inclusion rate between clusters.
[0034] Figure 22 It is an explanatory diagram depicting a specific example of calculating the inclusion rate between clusters.
[0035] Figure 23 It is an explanatory diagram depicting a specific example of identifying the hierarchical relationship between clusters.
[0036] Figure 24 It is an explanatory diagram depicting a specific example of showing the hierarchical relationship between clusters.
[0037] Figure 25 It is a flowchart depicting an example of the overall processing procedure. Detailed implementation manners
[0038] Embodiments of an information processing program, an information processing method, and an information processing apparatus according to the present invention will be described in detail with reference to the accompanying drawings.
[0039] (An example of the information processing method according to the embodiment)
[0040] Figure 1 FIG. is an explanatory diagram depicting an example of the information processing method according to the embodiment. The information processing apparatus 100 is a computer for providing information that makes it easier to grasp the content of the model. The information processing apparatus 100 is, for example, a server or a personal computer (PC).
[0041] The model, for example, has a function of outputting a label indicating the result of classifying data including a plurality of input explanatory variables. For example, the model includes a plurality of rules, and each rule represents a condition for classifying data using one or more explanatory variables. For example, the model is of a type called a white box model. For example, the model is of a type called a rule-based model.
[0042] For example, the model is generated by machine learning. For example, the model includes rules generated by machine learning. More specifically, the model includes rules generated based on training data that associates samples including values of a plurality of explanatory variables with correct labels indicating the results of classifying the samples. For example, the model may also include manually generated rules.
[0043] Here, the user may wish to grasp the content of the model. The content of the model, for example, indicates how the model uses a plurality of explanatory variables included in the input data to classify the input data. For example, the user wishes to grasp the content of the model in order to grasp the reliability of the model, or to grasp the reliability of the label indicating the result of classifying the data output by the model.
[0044] Therefore, for example, the following method is conceived, in which a list of a plurality of rules included in the model is output so that the user can refer to the list, thereby making it easier for the user to grasp the content of the model. However, with this method, it may be difficult for the user to easily grasp the content of the model.
[0045] For example, the number of rules included in a model generated by machine learning tends to increase as the complexity of the format of the data to be input increases. For example, the number of rules included in a model generated by machine learning tends to increase as the number of explanatory variables in the data to be input increases. The more rules included in the model, the greater the workload and working time required when the user refers to the list of rules included in the model and tries to grasp the content of the model.
[0046] Therefore, in the present embodiment, an information processing method is described that can make it easier to grasp the content of a model.
[0047] The information processing device 100 obtains a plurality of data. The data includes, for example, values of a plurality of explanatory variables. Among the plurality of data, any two data may include values of different types of explanatory variables. The information processing device 100 obtains the plurality of data by receiving the input of the plurality of data based on, for example, a user's operation input. The information processing device 100 may obtain the plurality of data by receiving the plurality of data from another computer.
[0048] The information processing device 100 obtains a rule set 110 including a plurality of rules. A rule represents a condition for classifying data using one or more explanatory variables. The rule set 110 forms a model that outputs a label indicating a result of classifying data including the input explanatory variables. The information processing device 100 obtains the rule set 110, for example, by generating the rule set 110 through machine learning based on a plurality of training data. The training data is, for example, data that associates a sample including values of a plurality of explanatory variables with a correct label indicating a result of classifying the sample. The information processing device 100 may obtain the rule set 110 by receiving the rule set 110 from another computer, for example. The information processing device 100 may obtain the rule set 110 by receiving the input of the rule set 110 based on, for example, a user's operation input.
[0049] (1-1) The information processing device 100 calculates a feature value of a rule based on one or more data among the obtained data that satisfy the condition represented by the rule included in the obtained rule set 110. The feature value is, for example, a feature vector of the rule, and the feature vector has components that are statistical values in one or more data that satisfy the condition represented by the rule, and the statistical value is related to each of the plurality of explanatory variables. The statistical value is, for example, a ratio of data in which the explanatory variable value is a specific value to one or more data that satisfy the condition represented by the rule. The statistical value may be, for example, the maximum value, minimum value, average value, mode value, or median value of the explanatory variable values in one or more data that satisfy the condition represented by the rule. Therefore, the information processing device 100 can determine the similarity between rules and obtain a guideline for clustering a plurality of rules.
[0050] (1-2) The information processing device 100 classifies each rule into one of a plurality of clusters based on the similarity between the feature values calculated for each rule. The information processing device 100 classifies each rule into one of the plurality of clusters using the x-means method based on the similarity between the feature values. As described above, the information processing device 100 can group two or more rules that are similar to each other into one cluster.
[0051] (1-3) The information processing device 100 identifies the inclusion relationship between clusters among multiple clusters based on the conditions represented by the rules classified into each of the multiple clusters. For example, among the multiple data obtained for each cluster, the information processing device 100 identifies a data set that satisfies the conditions represented by one or more rules classified into the cluster. The data set is, for example, a group of data that satisfies at least any one of the conditions represented by one or more rules. The data set can be, for example, a group of data that satisfies the conditions obtained by combining the conditions represented by one or more rules.
[0052] For each combination of clusters among the multiple clusters, the information processing device 100, for example, calculates the inclusion rate of one cluster relative to another cluster. For example, the inclusion rate is a value representing the proportion of the overlapping data between a data set that satisfies the conditions represented by one or more rules classified into one cluster and another data set that satisfies the conditions represented by one or more rules classified into another cluster in the whole of the other data set.
[0053] For example, for each combination of clusters among the multiple clusters, the information processing device 100 determines whether the inclusion rate of one cluster relative to another cluster is at least equal to a threshold. For example, when for any cluster combination, the inclusion rate of one cluster relative to another cluster is at least equal to the threshold, the information processing device 100 determines that one cluster in the combination has an inclusion relationship including the other cluster.
[0054] As described above, the information processing device 100 specifies the inclusion relationship between clusters among the multiple clusters by, for example, determining whether one cluster has an inclusion relationship including the other cluster for each combination of clusters among the multiple clusters. Therefore, the information processing device 100 can specify the inclusion relationship between clusters among the multiple clusters, and can specify the relationship between the rules forming the model.
[0055] (1-4) The information processing device 100 specifies the hierarchical relationship between clusters among the multiple clusters based on the inclusion relationship between clusters in the specified clusters. For example, for each combination of the multiple clusters, when one cluster is in an inclusion relationship including the other cluster, the information processing device 100 determines that one cluster is in a higher hierarchical relationship than the other cluster.
[0056] As described, the information processing device 100 specifies the hierarchical relationship between clusters among the multiple clusters by determining whether one cluster is in a higher hierarchical relationship than the other cluster for each combination of the multiple clusters. Therefore, the information processing device 100 can specify the hierarchical relationship between clusters among the multiple clusters, and can specify the relationship between the rules forming the model, and can obtain a guide that makes it easier to intuitively grasp the relationship between the rules forming the model.
[0057] (1-5) The information processing device 100 outputs information indicating the hierarchical relationship between clusters in a specified cluster. For example, the information processing device 100 outputs FIG. 120 showing the hierarchical relationship between clusters in a plurality of clusters. Thus, the information processing device 100 can make it easier for the user to intuitively grasp the relationship between the rules forming the model and can make it easier for the user to understand the content of the model.
[0058] Here, although the case where the information processing device 100 operates independently has been described, the configuration is not limited thereto. For example, the information processing device 100 can cooperate with another computer. For example, the information processing device 100 can cooperate with another computer that learns rules representing conditions for classifying data using one or more explanatory variables. For example, a plurality of computers can implement the functions of the information processing device 100. For example, the functions of the information processing device 100 can be implemented on the cloud.
[0059] (Example of information processing system 200)
[0060] Next, refer to Figure 2 Describe an example of an information processing system 200 to which the Figure 1 depicted information processing device 100 is applied.
[0061] Figure 2 is an explanatory diagram depicting an example of the information processing system 200. In Figure 2 the information processing system 200 includes an information processing device 100, a learning device 201, and a client device 202.
[0062] In the information processing system 200, the information processing device 100 and the learning device 201 are connected via a wired or wireless network 210. The network 210 is, for example, a local area network (LAN), a wide area network (WAN), or the Internet. In the information processing system 200, the information processing device 100 and the client device 202 are connected via a wired or wireless network 210.
[0063] The information processing device 100 is a computer that makes it easier for a user of a model to grasp the content of the model. The information processing device 100 obtains a plurality of data. The data includes, for example, values of a plurality of explanatory variables. The information processing device 100 obtains a plurality of data, for example, by receiving an input of a plurality of data based on an operation input by a user of the device, or the information processing device 100 can obtain a plurality of data by receiving a plurality of data from another computer. The other computer is, for example, the learning device 201.
[0064] The information processing device 100 obtains a rule set including a plurality of rules that form a model. The rules represent conditions for classifying data using one or more explanatory variables. The rule set forms a model that outputs a label indicating the result of classifying input data including a plurality of explanatory variables. The information processing device 100 obtains the rule set, for example, by receiving the rule set from another computer. The other computer is, for example, the learning device 201.
[0065] The information processing device 100 calculates a feature vector of a rule based on one or more data among the obtained plurality of data that satisfy the conditions represented by the rules included in the obtained rule set. The feature vector includes, for example, statistical values related to each of the plurality of explanatory variables in the one or more data that satisfy the conditions represented by the rule as components.
[0066] The information processing device 100 classifies each rule into one of a plurality of clusters based on the similarity between the feature vectors in the calculated feature vectors of each rule. The information processing device 100 identifies an inclusion relationship between the clusters based on the conditions represented by the rules used to classify the clusters into each cluster. The information processing device 100 identifies a hierarchical relationship between the clusters based on the inclusion relationship between the clusters in the identified clusters.
[0067] The information processing device 100 outputs information indicating the hierarchical relationship between the clusters in the identified clusters. The information processing device 100 sends the information indicating the hierarchical relationship between the clusters in the identified clusters to another computer. The other computer is, for example, the client device 202. The information processing device 100 is, for example, a server or a PC.
[0068] The learning device 201 is a computer for generating rules that indicate conditions for classifying data using one or more explanatory variables. The learning device 201 obtains, for example, a plurality of training data. The training data is, for example, data that associates a sample including values of a plurality of explanatory variables with a correct label indicating the result of classifying the sample. For example, the learning device 201 obtains a plurality of training data by receiving an input of the plurality of training data based on an operation input of a user of the device. For example, the learning device 201 generates a rule set based on the plurality of training data by machine learning. For example, the learning device 201 sends the generated rule set to the information processing device 100. The learning device 201 is, for example, a server or a PC.
[0069] The client device 202 is a computer used by the user of the model. The client device 202 receives, for example, information indicating the hierarchical relationship between clusters among the identified multiple clusters from the information processing device 100. The client device 202 outputs the received information indicating the hierarchical relationship between clusters among the multiple clusters so that the user of the model can refer to the information. The client device 202 is, for example, a PC, a tablet terminal, or a smart phone.
[0070] Here, although the case where the information processing device 100 is a device different from the learning device 201 has been described, the configuration is not limited thereto. For example, the information processing device 100 may have the functions of the learning device 201 and may also operate as the learning device 201.
[0071] Here, although the case where the information processing device 100 is a device different from the client device 202 has been described, the configuration is not limited thereto. For example, the information processing device 100 may have the functions of the client device 202 and may also operate as the client device 202.
[0072] (Application example of the information processing system 200)
[0073] Next, an application example of the information processing system 200 is described. The information processing system 200 can be applied to the case of a learning model that has a function of outputting a label indicating the result of classifying an input data set using, for example, a data set publicly available on the Internet as training data. The information processing system 200 can be specifically applied to the case of a learning model that has the following function: based on a data set representing the characteristics of an animal, output a label indicating whether the input animal is a mammal.
[0074] Therefore, the information processing system 200 can make it easier for the user of the model to grasp the effectiveness or reliability of the learned model. The information processing system 200 can make it easier for the user of the model to accept the classification result of the model. The information processing system 200 can make it easier for the user of the model to use the model with confidence.
[0075] For example, the information processing system 200 can be applied to the case of a learning model that has the following function: based on a data set indicating the characteristics of a viewer, output a label predicting whether the viewer of the input website will purchase a product. As described above, the information processing system 200 can make it easier for the user of the model to grasp the effectiveness or reliability of the learned model in the marketing field. The information processing system 200 can make it easier for the user of the model to accept the classification result of the model. The information processing system 200 can make it easier for the user of the model to use the model with confidence.
[0076] (Hardware configuration example of the information processing device 100)
[0077] Next, refer to Figure 3 an example of the hardware configuration of the information processing device 100.
[0078] Figure 3 is a block diagram depicting an example of the hardware configuration of the information processing device 100. In Figure 3 it, the information processing device 100 has a central processing unit (CPU) 301, a memory 302, a network interface (I / F) 303, a recording medium I / F 304, and a recording medium 305. The components are connected to each other via a bus 300.
[0079] Here, the CPU 301 manages the overall control of the information processing device 100. The memory 302 includes, for example, a read-only memory (ROM), a random access memory (RAM), and a flash ROM. For example, the flash ROM and the ROM store various programs, and the RAM serves as a working area for the CPU 301. The programs stored in the memory 302 are loaded onto the CPU 301, and the encoded processing is executed by the CPU 301.
[0080] The network I / F 303 is connected to the network 210 via a communication line and is connected to other computers via the network 210. The network I / F 303 manages the internal interface with the network 210 and controls the input and output of data from other computers. The network I / F 303 is, for example, a modem or a LAN adapter.
[0081] The recording medium I / F 304 controls the reading and writing of data regarding the recording medium 305 under the control of the CPU 301. The recording medium I / F 304 is, for example, a disk drive, a solid state drive (SSD), a universal serial bus (USB) port, etc. The recording medium 305 is a non-volatile memory that stores the data written thereto under the control of the recording medium I / F 304. The recording medium 305 is, for example, a disk, a semiconductor memory, a USB memory, etc. The recording medium 305 can be detached from the information processing device 100.
[0082] In addition to the components mentioned above, the information processing device 100 may have, for example, a keyboard, a mouse, a display, a printer, a scanner, a microphone, a speaker, etc. The information processing device 100 may have multiple recording medium I / Fs 304 and recording media 305. The information processing device 100 may omit the recording medium I / F 304 and the recording medium 305.
[0083] (Example of the hardware configuration of the learning device 201)
[0084] The example of the hardware configuration of the learning device 201 is specifically similar to Figure 3An example of the hardware configuration of the information processing device 100 depicted therein is thus omitted from the description.
[0085] (Example of the hardware configuration of the client device 202)
[0086] The example of the hardware configuration of the client device 202 is specifically similar to Figure 3 the example of the hardware configuration of the information processing device 100 depicted therein, and thus the description thereof is omitted.
[0087] (Example of the functional configuration of the information processing device 100)
[0088] Next, with reference to Figure 4 an example of the functional configuration of the information processing device 100 will be described.
[0089] Figure 4 FIG. is a block diagram showing an example of the functional configuration of the information processing device 100. The information processing device 100 includes a storage unit 400, an acquisition unit 401, a learning unit 402, a calculation unit 403, a classification unit 404, an identification unit 405, a generation unit 406, and an output unit 407.
[0090] The storage unit 400 is implemented by a storage area such as Figure 3 the memory 302 or the recording medium 305 depicted therein. Hereinafter, although the case where the storage unit 400 is included in the information processing device 100 is described, the configuration is not limited thereto. For example, the storage unit 400 may be included in a device different from the information processing device 100, and the storage content of the storage unit 400 may be referred to from the information processing device 100.
[0091] The acquisition unit 401 to the output unit 407 serve as examples of a controller. For example, the functions of the acquisition unit 401 to the output unit 407 are implemented by causing the CPU 301 to execute a program stored in a storage area such as Figure 3 the memory 302 or the recording medium 305 depicted therein or through the network I / F 303. The processing result of the functional unit is stored, for example, in a storage area such as Figure 3 the memory 302 or the recording medium 305 depicted therein.
[0092] The storage unit 400 stores various types of information referred to or updated in the processing of the functional unit. The storage unit 400 stores, for example, a plurality of input data therein. Each of the plurality of input data includes, for example, values of a plurality of explanatory variables. Each of the plurality of input data is obtained by, for example, the acquisition unit 401.
[0093] The storage unit 400 stores, for example, a rule set. The rule set includes, for example, a plurality of rules representing conditions for classifying data using one or more explanatory variables. For example, the rule set includes rules generated by machine learning. For example, the rule set may further include manually generated rules. The rule set is generated, for example, by the learning unit 402. The rule set may be obtained, for example, by the obtaining unit 401.
[0094] The storage unit 400 stores, for example, a plurality of training data. The training data is data that associates a sample including values of a plurality of explanatory variables with a correct label indicating the result of classifying the sample. The training data is obtained, for example, by the obtaining unit 401.
[0095] The obtaining unit 401 obtains various types of information used by the functional units during processing. The obtaining unit 401 stores the obtained various types of information in the storage unit 400, or outputs the obtained information to the functional units. The obtaining unit 401 may also output the various types of information stored in the storage unit 400 to each functional unit. The obtaining unit 401 may obtain various types of information based on, for example, a user's operation input. The obtaining unit 401 may receive various information from a device other than, for example, the information processing device 100.
[0096] The obtaining unit 401 obtains a plurality of input data. The obtaining unit 401 obtains a plurality of input data by receiving the plurality of input data from another computer. The other computer is, for example, the learning device 201. The obtaining unit 401 may obtain a plurality of input data by receiving the input of the plurality of input data based on, for example, a user's operation input of the device.
[0097] The obtaining unit 401 obtains a rule set. The obtaining unit 401 obtains a rule set by receiving the rule set from another computer. The other computer is, for example, the learning device 201. The obtaining unit 401 may obtain a rule set by receiving the input of the rule set based on, for example, a user's operation input of the device.
[0098] The obtaining unit 401 obtains a plurality of training data. For example, the obtaining unit 401 obtains a plurality of training data by receiving the plurality of training data from another computer. The other computer is, for example, the learning device 201. For example, the obtaining unit 401 may obtain a plurality of training data by receiving the input of the plurality of training data based on a user's operation input of the device. For example, the training data may be used as input data.
[0099] The obtaining unit 401 may receive a start trigger for starting processing by any of the functional units. The start trigger may be, for example, a predetermined operation input by the user. The start trigger may be, for example, receiving predetermined information from another computer. The start trigger may be, for example, any of the functional units outputting predetermined information.
[0100] For example, the acquisition unit 401 may receive a plurality of pieces of training data as a start trigger for starting the processing of the learning unit 402. The acquisition unit 401 may receive, for example, a plurality of input data as a start trigger for starting the processing of the calculation unit 403, the classification unit 404, the recognition unit 405, and the generation unit 406.
[0101] The learning unit 402 generates a rule set. The rule set includes, for example, a plurality of rules representing conditions for classifying data using one or more explanatory variables. For example, the learning unit 402 generates a plurality of rules by machine learning based on the plurality of pieces of training data obtained by the acquisition unit 401, and generates a rule set including the plurality of rules. Therefore, the learning unit 402 can generate a rule set on the information processing device 100 and make it easier for the information processing device 100 to operate independently.
[0102] The calculation unit 403 calculates a feature value related to each rule included in the rule set obtained by the acquisition unit 401 or generated by the learning unit 402. The feature value is, for example, a feature vector related to the rule. The calculation unit 403 calculates the feature value of the rule based on one or more input data among the plurality of input data obtained by the acquisition unit 401 that satisfy the conditions represented by the rules included in the rule set obtained by the acquisition unit 401 or generated by the learning unit 402.
[0103] For example, among the plurality of input data obtained by the acquisition unit 401, the calculation unit 403 identifies one or more input data that satisfy the conditions represented by the rules included in the rule set obtained by the acquisition unit 401 or generated by the learning unit 402. For example, the calculation unit 403 calculates a feature vector for each rule included in the rule set, and the components thereof are statistical values related to each of the plurality of explanatory variables in the identified one or more input data. Therefore, the calculation unit 403 can evaluate the similarity between rules based on the comparison of the feature values.
[0104] The classification unit 404 classifies each rule included in the rule set obtained by the acquisition unit 401 or generated by the learning unit 402 into one of a plurality of clusters. The classification unit 404 classifies each rule into one of the plurality of clusters, for example, based on the similarity between the feature values calculated for the rules.
[0105] For example, the classification unit 404 calculates the sum of squares of the differences between the components of the explanatory variables among the calculated feature vectors of each rule as the similarity. For example, the classification unit 404 classifies each rule included in the rule set into one of a plurality of clusters based on the calculated similarity. Thus, the classification unit 404 can group two or more rules that are similar to each other into one cluster. The classification unit 404 can enable the user of the model to grasp two or more rules as a whole. The classification unit 404 can enable the user of the model to grasp two or more rules as a group, so that the user of the model can grasp the content of the model.
[0106] The identification unit 405 identifies the inclusion relationship between the clusters among the plurality of clusters based on the conditions expressed by the rules classified into the plurality of clusters. From the plurality of input data obtained by the acquisition unit 401, the identification unit 405, for example, identifies, for each cluster, a data set that satisfies the conditions represented by one or more rules classified into that cluster. For example, from the plurality of input data obtained by the acquisition unit 401, the identification unit 405, for each of the plurality of clusters, identifies a data set that includes input data that satisfies at least any one of the conditions represented by one or more rules classified into the cluster. For example, from the plurality of input data obtained by the acquisition unit 401, the identification unit 405 can identify, for each cluster, a data set that includes input data that satisfies the combined conditions represented by one or more rules classified into the cluster.
[0107] For each combination of the clusters among the plurality of clusters, the identification unit 405, for example, calculates the inclusion rate of one cluster with respect to another cluster. The inclusion rate is, for example, a value that indicates the proportion of the overlapping input data between a first data set that satisfies the conditions represented by one or more rules classified into one cluster and a second data set that satisfies the conditions represented by one or more rules classified into another cluster, that is, the proportion in the whole of the second data set.
[0108] The inclusion rate can be, for example, a value that represents the proportion of the overlapping conditions between the set of conditions represented by one or more rules classified into one cluster and the set of conditions represented by one or more rules classified into another cluster, that is, the proportion in the whole set of conditions represented by one or more rules classified into another cluster.
[0109] For example, for each combination of the clusters among the plurality of clusters, the identification unit 405 determines whether the inclusion rate of one cluster with respect to another cluster is at least equal to a threshold value. For example, when, for any cluster combination, the inclusion rate of one cluster with respect to another cluster is at least equal to the threshold value, the identification unit 405 determines that one cluster in the combination has an inclusion relationship including the other cluster.
[0110] For example, for each combination of clusters, the recognition unit 405 determines whether one cluster has an inclusion relationship including another cluster, and thereby recognizes the inclusion relationships among the clusters in the plurality of clusters. As described above, the recognition unit 405 can recognize the inclusion relationships among the clusters in the plurality of clusters and can recognize the relationships among the rules forming the model.
[0111] The generation unit 406 recognizes the hierarchical relationships among the clusters based on the inclusion relationships among the clusters in the recognized clusters. For example, for each combination of clusters in the plurality of clusters, when one cluster is in an inclusion relationship including another cluster, the generation unit 406 determines that one cluster is in a higher hierarchical relationship than the other cluster.
[0112] The generation unit 406 recognizes the hierarchical relationships among the clusters in the plurality of clusters by, for example, determining for each combination of clusters in the plurality of clusters whether one cluster is in a higher hierarchical relationship than the other cluster. As described above, the generation unit 406 can recognize the hierarchical relationships among the clusters in the plurality of clusters, can recognize the relationships among the rules forming the model, and can obtain guidelines that make it easier to intuitively grasp the relationships among the rules forming the model.
[0113] The generation unit 406 generates information indicating the hierarchical relationships among the clusters in the recognized clusters. For example, the generation unit 406 generates a graph indicating the hierarchical relationships among the clusters in the recognized clusters. The graph is formed, for example, by nodes representing the clusters and edges connecting the nodes representing the respective clusters having a hierarchical relationship. For example, the graph includes a node representing the highest cluster as the root node. For example, the graph is a tree structure. As described above, the generation unit 406 can obtain reference information that makes it easier for the user of the model to intuitively grasp the hierarchical relationships among the clusters.
[0114] The generation unit 406 generates information indicating the characteristics of one or more rules classified into a cluster for each cluster. For example, the generation unit 406 generates a graph indicating the characteristics of one or more rules classified into a cluster. For example, the graph is formed by nodes representing the rules and edges between the nodes representing the corresponding rules having a connection relationship. For example, the graph is a tree structure. As described above, the generation unit 406 can obtain reference information that makes it easier for the user of the model to intuitively grasp one or more rules classified into a cluster.
[0115] The output unit 407 outputs the processing result of at least one of the functional units. The output form is, for example, to display on a display, print out on a printer, send to an external device via the network I / F 303, or store in a storage area such as the memory 302 or the recording medium 305. As described above, the output unit 407 can notify the user of the processing result of at least any of the functional units, thereby improving the convenience of the information processing apparatus 100.
[0116] The output unit 407 outputs information indicating the hierarchical relationship between clusters in the multiple clusters generated by the generation unit 406. The output unit 407 outputs, for example, information indicating the hierarchical relationship between clusters in the multiple clusters, so that the user of the model can refer to this information. For example, the output unit 407 sends information indicating the hierarchical relationship between clusters in the multiple clusters to the client device 202. Therefore, the output unit 407 can enable the user of the model to more easily and intuitively grasp the hierarchical relationship between clusters, and can make it easier to intuitively grasp the relationship between the rules forming the model.
[0117] When the output unit 407 receives a designation of any one of the multiple clusters, the output unit 407 outputs the information generated by the generation unit 406, the information indicating the features of one or more rules classified into any one of the clusters that have received the designation. The output unit 407 outputs, for example, information indicating the features of one or more rules classified into any one of the clusters that have received the designation, so that the user of the model can refer to this information. For example, the output unit 407 sends information indicating the features of one or more rules classified into any one of the clusters that have received the designation to the client device 202. Therefore, the output unit 407 can enable the user of the model to more easily and intuitively grasp the connection relationship between the rules in one or more similar rules forming the model.
[0118] (Operation flow of the information processing device 100)
[0119] Next, refer to Figure 5 to describe the operation flow of the information processing device 100.
[0120] Figure 5 is an explanatory diagram depicting the operation flow of the information processing device 100. In Figure 5 , the information processing device 100 receives the input 500 of multiple pieces of training data. The training data is, for example, data that associates a sample including values of multiple explanatory variables with a label indicating the correct answer obtained by classifying the sample through the model.
[0121] The information processing device 100 executes the "rule creation" process 501, and thereby learns, through machine learning, a model including a rule set based on multiple pieces of training data. The model has a function of outputting a label that indicates the result of classifying data in response to the input of data including values of multiple explanatory variables. A rule is information indicating a condition for classifying data using one or more explanatory variables. The model is an artificial intelligence (AI) model. The model is, for example, an XAI (explainable AI) model. The model is, for example, a rule-based model.
[0122] The information processing device 100 performs a "vectorization" process 502 to generate feature vectors corresponding to the rules of the rule set included in the model, respectively. For example, for each rule, the information processing device 100 identifies one or more pieces of training data that satisfy the condition represented by the rule in a plurality of pieces of training data. For example, for each rule, the information processing device 100 generates a feature vector that includes statistical values for each explanatory variable in the identified one or more pieces of training data as components. The statistical value is, for example, the maximum value, the minimum value, the average value, the median value, or the mode value. The statistical value can be, for example, the variance.
[0123] The information processing device 100 performs a "clustering" process 503 to classify each rule included in the rule set into one of a plurality of clusters. For example, by the x-means method, the information processing device 100 classifies each rule included in the rule set into one of a plurality of clusters based on the Euclidean distance between the calculated feature vectors.
[0124] The information processing device 100 generates explanation information for each of the plurality of clusters by performing an "explanation generation" process 504. The explanation information is information about one or more rules classified into the cluster. The explanation information is, for example, information indicating what features the one or more rules classified into the cluster have. For example, for each cluster, the information processing device 100 identifies one or more pieces of training data that satisfy at least any of the conditions expressed by the one or more rules classified into the cluster from a plurality of pieces of training data. For example, the information processing device 100 generates new training data for each cluster by adding a label indicating a positive (correct) example to the values of a plurality of explanatory variables included in the identified one or more pieces of training data. For example, for each cluster, the information processing device 100 generates a decision tree for classifying data based on the newly generated training data as explanation information. For example, for each cluster, the information processing device 100 may generate a rule-based model for classifying data based on the newly generated training data as explanation information.
[0125] The information processing device 100 calculates the inclusion rate between clusters by executing the "inclusion rate calculation" process 505. For each combination of clusters, the information processing device 100 calculates, for example, the inclusion rate of one cluster relative to another cluster in the combination. For example, for each cluster, the information processing device 100 identifies, in a plurality of training data, a training data set that satisfies the conditions represented by one or more rules classified into the cluster. For example, for each combination of clusters, the information processing device 100 identifies the training data that overlaps between the training data set that satisfies the conditions represented by one or more rules classified into one cluster and the training data set that satisfies the conditions represented by one or more rules classified into another cluster in the combination. For example, for each combination of clusters, the information processing device 100 calculates the proportion of the identified overlapping training data in the entire training data set that satisfies the conditions represented by one or more rules classified into another cluster in the combination as the inclusion rate.
[0126] The information processing device 100 identifies the inclusion relationship between clusters by executing the "obtain hierarchy" process 506, and identifies the hierarchical relationship between clusters based on the inclusion relationship between clusters. For each combination of clusters, the information processing device 100 determines whether one cluster and another cluster in the combination are in a hierarchical relationship based on whether the inclusion rate of one cluster relative to another cluster is at least equal to a threshold. For any combination of clusters, when the inclusion rate of one cluster relative to another cluster is at least equal to the threshold, the information processing device 100 identifies their hierarchical relationship as one in which one cluster is at a higher hierarchical level than the other cluster. For each combination of clusters, the information processing device 100 identifies the hierarchical relationship between clusters by determining whether one cluster and another cluster are in a hierarchical relationship.
[0127] The information processing device 100 displays a graph representing the identified hierarchical relationship between clusters on a display by executing the "screen output" process 507. Thus, the information processing device 100 can enable the user of the model to more easily and intuitively grasp the relationship between the rules forming the model and more easily grasp the content of the model.
[0128] (Example of the operation of the information processing device 100)
[0129] Next, refer to Figures 6 to 10 Describe an example of the operation of the information processing device 100. First, refer to Figure 6 Describe an example of the information processing device 100 that learns the model 600.
[0130] Figure 6 is an explanatory diagram depicting an example of learning the model 600. In Figure 6In this case, the information processing device 100 obtains a plurality of training data indicating characteristics related to an animal. The training data indicates, for example, a plurality of characteristic values related to the animal, and each of the characteristic values is an explanatory variable associated with a correct label indicating whether the animal is a mammal. The characteristic values indicate, for example, the size of the animal, the presence or absence of body hair, the presence or absence of feathers, whether the animal is a carnivore, whether the animal breathes with lungs, or whether the animal is oviparous. The information processing device 100 obtains, for example, 93 pieces of training data. For example, there are 15 explanatory variables. The information processing device 100 learns the model 600 based on the 93 pieces of training data obtained, for example.
[0131] The model 600 is, for example, a rule-based model and includes a plurality of rules, and each rule represents a condition for classifying whether an animal is a mammal using one or more explanatory variables. The model 600 is implemented, for example, by a storage area such as Figure 3 the memory 302 or the recording medium 305 of the information processing device 100 as depicted. As Figure 6 depicted, the model 600 has fields: chunk, label, weight, len, npos, nneg, supp, and conf. The model 600 stores each rule as a record 600-a by setting information in each field for each rule.
[0132] "a" is an arbitrary integer.
[0133] In the chunk field, a rule is set. In the label field, flag information indicating whether the rule is a rule representing a mammal is set. For example, a value of 1 for the flag information indicates that the rule represents a mammal, and a value of 0 for the flag information indicates that the rule represents a non-mammal. In the weight field, the weight of the rule is set. In the len field, the number of conditions forming the rule is set.
[0134] In the npos field, the number of training data that satisfy the rule and match the flag information among the plurality of training data is set. For example, a value of 1 for the flag information indicates that: among the plurality of training data, the number of training data that satisfy the rule and include the correct label indicating that the object is a mammal is set. For example, a value of 0 for the flag information indicates that: among the plurality of training data, the number of training data that satisfy the rule and include the correct label indicating that the object is not a mammal is set.
[0135] In the nneg field, the number of training data that satisfy the rule and do not match the flag information among multiple training data is set. For example, a value of 1 for the flag information indicates that: among multiple training data, the number of training data that satisfy the rule and include the correct label indicating that the object is not a mammal is set. For example, in the nneg field, a value of 0 for the flag information indicates that: the number of training data that satisfy the rule and include the correct label indicating that the object is a mammal is set.
[0136] In the supp field, npos / total_npos is set. total_npos is the sum of npos of the rules. In the conf field, npos / (npos + nneg) is set. The model 600 represents, for example, 321 rules. For example, the information processing device 100 can delete from the model 600 the rules that do not satisfy at least x% or more of the conditions in 93 pieces of training data. "x" is, for example, 1. For example, the information processing device 100 can delete from the model 600 the rules where npos / (npos + nneg) is less than y. "y" is, for example, 0.9.
[0137] Next, with reference to Figure 7 An example of the information processing device 100 generating a feature vector for each rule represented by the model 600 based on multiple training data is described.
[0138] Figure 7 is an explanatory diagram depicting an example of generating a feature vector for each rule. In Figure 7 Among multiple training data, the information processing device 100 identifies, for each rule represented by the model 600, one or more training data that satisfy one or more conditions represented by the rule. The information processing device 100 generates a feature vector for each rule represented by the model 600, and the feature vector includes the statistical value of each explanatory variable in the identified one or more training data as a component. The information processing device 100 can generate a feature vector by replacing the values "yes" and "no" of the explanatory variable with values 1 and 0.
[0139] For example, the information processing device 100 generates a feature vector 711 that includes the statistical value of each explanatory variable in one or more training data that satisfy rule 701 among multiple training data as a component. For example, the information processing device 100 generates a feature vector 712 that includes the statistical value of each explanatory variable in one or more training data that satisfy rule 702 among multiple training data as a component.
[0140] The information processing device 100 classifies each rule represented by the model 600 into one of a plurality of clusters based on the Euclidean distance between the generated feature vectors using the x-mean method. For example, it is assumed that the information processing device 100 classifies each rule represented by the model 600 into one of 20 clusters. In the following description, the i-th cluster may be referred to as "cluster i". "i" is, for example, from 0 to 19.
[0141] Next, an example of the information processing device 100 identifying the inclusion relationship between clusters among the classified multiple clusters will be described with reference to Figure 8 and Figure 9 FIGS.
[0142] Figure 8 and Figure 9 are explanatory diagrams depicting an example of identifying the inclusion relationship between clusters. In Figure 8 the information processing device 100 calculates the inclusion rate between clusters. For each combination of clusters, the information processing device 100 calculates the inclusion rate of one cluster in another cluster and the inclusion rate of the other cluster in one cluster. The inclusion rate of one cluster in another cluster is defined, for example, as (A∩B) / A. "A" is a set of training data that satisfies at least any of the conditions represented by one or more rules classified into the other cluster. "B" is a set of training data that satisfies at least any of the conditions represented by one or more rules classified into one cluster.
[0143] The information processing device 100 calculates the inclusion rate between clusters as depicted, for example, in Table 800. For example, the information processing device 100 calculates the inclusion rate of cluster 1 with respect to cluster 2 and the inclusion rate of cluster 2 with respect to cluster 1 for the combination of cluster 1 and cluster 2. For example, the information processing device 100 calculates the inclusion rate of cluster 1 with respect to cluster 3 and the inclusion rate of cluster 3 with respect to cluster 1 for the combination of cluster 1 and cluster 3. For example, the information processing device 100 calculates the inclusion rate of cluster 2 with respect to cluster 3 and the inclusion rate of cluster 3 with respect to cluster 2 for the combination of cluster 2 and cluster 3. This will be described with reference to Figure 9 FIGS.
[0144] In Figure 9 it is assumed that the information processing device 100 calculates the inclusion rate between clusters among 20 clusters as depicted in Table 900. The horizontal axis indicates the cluster number. The vertical axis indicates the cluster number. Table 900 shows the inclusion rate of the cluster with the cluster number on the horizontal axis with respect to the cluster with the cluster number on the vertical axis. White indicates that among the multiple ranges where the inclusion rate is between 0 and 1, the inclusion rate belongs to the first range closest to 1. The upper end of the first range is, for example, 1.
[0145] The light grid shading indicates the inclusion ratio belonging to the second range among multiple ranges between the inclusion ratios of 0 and 1, where the second range is the next closest to 1 after the first range. The dotted shading indicates the inclusion ratio belonging to the third range among multiple ranges, where the third range is the next closest to 1 after the second range. The dense grid shading indicates the inclusion ratio belonging to the fourth range among multiple ranges, where the fourth range is the next closest to 1 and closest to 0 after the third range. The lower end of the fourth range is, for example, 0.
[0146] For each combination of clusters, the information processing device 100 determines whether one cluster includes another cluster based on whether the inclusion ratio of one cluster to another cluster is at least equal to a threshold value. The threshold value is, for example, 0.9. The threshold value is, for example, preset by the user of the model. For example, for any combination of clusters, when the inclusion ratio of one cluster in the combination to another cluster in the combination is at least equal to the threshold value, the information processing device 100 determines that one cluster includes another cluster. For each combination of clusters, the information processing device 100 determines whether one cluster includes another cluster, and thereby identifies the inclusion relationship between the clusters among the multiple clusters.
[0147] Next, with reference to Figure 10 , an example is described in which the information processing device 100 displays the hierarchical relationship between clusters based on the identified inclusion relationship between the clusters.
[0148] Figure 10 is an explanatory diagram depicting an example of displaying the hierarchical relationship between clusters. In Figure 10 , the information processing device 100 identifies the hierarchical relationship between the clusters among the multiple clusters based on the inclusion relationship between the clusters in the multiple clusters. For example, for each combination of clusters, the information processing device 100 determines whether one cluster is in a higher hierarchical relationship than another cluster, and whether the other cluster is in a higher hierarchical relationship than one cluster. For example, for any combination of clusters, when one cluster in the combination includes the other cluster in the combination, thus having an inclusion relationship, and the one cluster is a higher-level cluster than the other cluster, the information processing device 100 identifies the hierarchical relationship between the clusters.
[0149] The information processing device 100 generates a graph 1000 indicating the identified hierarchical relationship between the clusters. The graph 1000 includes nodes 1001, 1011, 1012, 1013, 1014, 1015, 1021, 1022, 1023, 1024, 1025, 1026, 1027, 1028, 1031, 1032, 1033, 1034, 1035, and 1041 representing the clusters. In Figure 10In the depicted example, nodes 1001, 1011 to 1015, 1021 to 1028, 1031 to 1035, and 1041 are indicated by circles, for example. The numbers attached to the bottoms of nodes 1001, 1011 to 1015, 1021 to 1028, 1031 to 1035, and 1041 are the cluster numbers represented by the nodes, respectively. FIG. 1000 includes edges connecting the nodes, and these nodes represent respective clusters of cluster combinations having a hierarchical relationship among nodes 1001, 1011 to 1015, 1021 to 1028, 1031 to 1035, and 1041.
[0150] The information processing device 100 displays FIG. 1000. At this time, the information processing device 100 can display FIG. 1000 such that the larger the number of rules classified into the clusters represented by nodes 1001, 1011 to 1015, 1021 to 1028, 1031 to 1035, and 1041, the larger the size of the circle.
[0151] The information processing device 100 receives a designation of any one of nodes 1001, 1011 to 1015, 1021 to 1028, 1031 to 1035, and 1041 in the displayed FIG. 1000. The information processing device 100 receives, for example, a click on any one of nodes 1001, 1011 to 1015, 1021 to 1028, 1031 to 1035, and 1041 in the displayed FIG. 1000 as the designation of the node.
[0152] In response to the designation of the node, the information processing device 100 can display explanation information 1010 about one or more rules classified into the cluster represented by any of the designated nodes. The explanation information 1010 is, for example, information indicating the features of one or more training data among the multiple training data that satisfy at least any of the conditions represented by one or more rules classified into the cluster.
[0153] The explanation information 1010 is, for example, a decision tree that can classify one or more training data among the multiple training data that satisfy at least any of the conditions represented by one or more rules classified into the cluster as positive examples. The explanation information 1010 can be, for example, a feature vector that includes, as components, statistical values of each explanatory variable in one or more training data among the multiple training data that satisfy at least any of the conditions represented by one or more rules classified into the cluster.
[0154] Therefore, the information processing device 100 can enable the user to more easily and intuitively grasp the relationships among the rules forming the model, and enable the user to easily grasp the content of the model. The information processing device 100 enables the user to easily grasp the content of one or more rules classified into the cluster.
[0155] (Specific Example of the Operation of Information Processing Apparatus 100)
[0156] Next, a specific example of the operation of information processing apparatus 100 will be described with reference to Figures 11 to 24 First, a specific example of information processing apparatus 100 obtaining a plurality of training data and learning a model will be described with reference to Figure 11 and Figure 12 FIGS.
[0157] Figure 11 and Figure 12 are explanatory diagrams depicting specific examples of learning the model. In Figure 11 and Figure 12 information processing apparatus 100 obtains the animal data management table 1100. For convenience, the content stored in the animal data management table 1100 is depicted and divided in Figure 11 and Figure 12 The animal data management table 1100 is implemented, for example, by a storage area of a storage such as the memory 302 or the recording medium 305 of the information processing apparatus 100 depicted in Figure 3 FIG.
[0158] As depicted in Figure 11 the animal data management table 1100 has fields for id, name, hair, feathers, oviparous, flight, aquatic, carnivorous, teeth, vertebrae, and lung respiration. As depicted in Figure 12 the animal data management table 1100 also has fields for poison, fins, legs, tail, captive breeding, approximately feline size, and label.
[0159] In the animal data management table 1100, information is set in each field for each animal, and the animal data is stored as a record 1100-b. "b" is an arbitrary integer. In the examples depicted in Figure 11 and Figure 12 for example, there are 93 pieces of animal data. Therefore, b is, for example, from 1 to 93.
[0160] In the id field, an id is set, which is a number assigned to the animal. In the name field, the name of the animal is set. In the hair field, information indicating whether the animal has hair / fur is set. In the feathers field, information indicating whether the animal has feathers is set. In the oviparous field, information indicating whether the animal is oviparous is set. In the flight field, information indicating whether the animal flies is set. In the aquatic field, information indicating whether the animal is aquatic is set.
[0161] In the carnivore field, information indicating whether the animal is a carnivore is set. In the teeth field, information indicating whether the animal has teeth is set. In the vertebra field, information indicating whether the animal has vertebrae is set. In the lung respiration field, information indicating whether the animal performs lung respiration is set.
[0162] In the poison field, information indicating whether the animal has poison is set. In the fin field, information indicating whether the animal has fins is set. In the leg field, information indicating the number of legs the animal has is set. In the tail field, information indicating whether the animal has a tail is set. In the artificially reared field, information indicating whether the animal is reared by humans is set. In the approximately feline size field, information indicating whether the animal is approximately the size of a feline is set.
[0163] In the label field, flag information indicating whether the animal is a mammal is set. For example, flag information with a value of 1 indicates that the animal is a mammal, while flag information with a value of 0 indicates that the animal is not a mammal. The information processing device 100 generates training data based on each animal data in the animal data management table 1100.
[0164] For example, the information processing device 100 generates training data for each animal data. The training data includes the values of the fields other than the id, name, and label of the animal data as explanatory variable values, and the training data also includes the value of the label field of the animal data as the correct label. For example, there are 15 explanatory variables. Therefore, the information processing device 100 can generate 93 pieces of training data.
[0165] The information processing device 100 trains the model 600 based on the 93 pieces of generated training data. The model 600 represents, for example, 321 rules. The rules represent, for example, one or more conditions for classifying an animal as a mammal or a non - mammal using one or more explanatory variables. The content of the model 600 is similar to Figure 6 the content depicted, and thus its description is omitted.
[0166] The information processing device 100 can delete from the model 600 the rules representing conditions that do not satisfy at least x% or more of the 93 pieces of training data. "x" is, for example, 1. The information processing device 100 can, for example, delete from the model 600 the rules where npos / (npos + nneg) is less than y. "y" is, for example, 0.9. Thus, the information processing device 100 can learn the model 600.
[0167] Next, with reference to Figure 13 and Figure 14 a specific example of the information processing device 100 generating a feature vector related to the rules represented by the model 600 is described.
[0168] Figure 13 and Figure 14 are explanatory diagrams depicting specific examples of generating feature vectors related to rules respectively. In Figure 13 and Figure 14 , the information processing device 100 generates feature vectors related to each rule represented by the model 600 based on a plurality of training data. The information processing device 100 can, for example, replace the values of the explanatory variables "yes" and "no" with the values 1 and 0, and then generate feature vectors. For example, for each rule represented by the model 600, the information processing device 100 identifies one or more training data that satisfy one or more conditions represented by the rule in the plurality of training data.
[0169] For example, the information processing device 100 generates, for each rule represented by the model 600, a feature vector that includes statistical values for each explanatory variable in the identified one or more training data as components. The statistical value is, for example, the proportion of training data in which a specific explanatory variable value is a specific value or range among one or more training data that satisfy one or more conditions represented by the rule. The statistical value can be, for example, the maximum value, minimum value, average value, mode, or median of the explanatory variable.
[0170] In Figure 13 and Figure 14 the depicted example, for each rule represented by the model 600, the information processing device 100 calculates the proportion of training data with hair = "yes" among one or more training data that satisfy one or more conditions represented by the rule as a component of the feature vector. In addition, for each rule represented by the model 600, the information processing device 100 calculates the proportion of training data with hair = "no" among one or more training data that satisfy one or more conditions represented by the rule as a component of the feature vector.
[0171] In addition, for each rule represented by the model 600, the information processing device 100 calculates the proportion of training data with feathers = "yes" among one or more training data that satisfy one or more conditions represented by the rule, and this proportion is calculated as a component of the feature vector. In addition, for each rule represented by the model 600, the information processing device 100 calculates the proportion of training data with feathers = "no" among one or more training data that satisfy one or more conditions represented by the rule, and this proportion is calculated as a component of the feature vector.
[0172] In addition, for each rule represented by the model 600, the information processing device 100 calculates the proportion of training data with oviparity = "no" among one or more training data that satisfy one or more conditions represented by the rule, and this proportion is calculated as a component of the feature vector. In addition, for each rule represented by the model 600, the information processing device 100 calculates the proportion of training data with oviparity = "yes" among one or more training data that satisfy one or more conditions represented by the rule, and this proportion is calculated as a component of the feature vector.
[0173] In addition, for each rule represented by the model 600, the information processing device 100 calculates the proportion of training data with the number of legs < 3 among one or more training data that satisfy one or more conditions represented by the rule, and this proportion is calculated as a component of the feature vector. In addition, for each rule represented by the model 600, the information processing device 100 calculates the proportion of training data with the number of legs ≥ 3 among one or more training data that satisfy one or more conditions represented by the rule, and this proportion is calculated as a component of the feature vector.
[0174] In addition, for each rule represented by the model 600, the information processing device 100 calculates the proportion of training data with the number of legs < 5 among one or more training data that satisfy one or more conditions represented by the rule, and this proportion is calculated as a component of the feature vector. In addition, for each rule represented by the model 600, the information processing device 100 calculates the proportion of training data with the number of legs ≥ 5 among one or more training data that satisfy one or more conditions represented by the rule, and this proportion is calculated as a component of the feature vector.
[0175] In addition, for each rule represented by the model 600, the information processing device 100 calculates the proportion of training data with tail = "yes" among one or more training data that satisfy one or more conditions represented by the rule, and this proportion is calculated as a component of the feature vector. In addition, for each rule represented by the model 600, the information processing device 100 calculates the proportion of training data with tail = "no" among one or more training data that satisfy one or more conditions represented by the rule, and this proportion is calculated as a component of the feature vector. For example, the information processing device 100 generates a feature vector including the calculated components and stores the feature vector using the feature vector management table 1300. The feature vector management table 1300 is implemented, for example, by a storage area such as Figure 3 the memory 302 of the information processing device 100 or the recording medium 305 as depicted.
[0176] As Figure 13 depicted, the feature vector management table 1300 has fields for rule, hair = "yes", hair = "no", feather = "yes", feather = "no", and oviparity = "no". AsFigure 14 As depicted, the feature vector management table 1300 also has fields for oviparous = "yes", legs < 3, legs ≥ 3, legs < 5, legs ≥ 5, tail = "yes", and tail = "no". The feature vector management table 1300 stores the feature vectors as records 1300-c by setting information in each field for each rule. "c" is an arbitrary integer.
[0177] In the rule field, the rules are set. In the field where hair = "yes", the proportion of the training data with hair = "yes" among one or more training data that satisfy one or more conditions represented by the rule is set as a component of the feature vector. In the field where hair = "no", the proportion of the training data with hair = "no" among one or more training data that satisfy one or more conditions represented by the rule is set as a component of the feature vector.
[0178] In the field where feathers = "yes", the proportion of the training data with feathers = "yes" among one or more training data that satisfy one or more conditions represented by the rule is set as a component of the feature vector. In the field where feathers = "no", the proportion of the training data with feathers = "no" among one or more training data that satisfy one or more conditions represented by the rule is set as a component of the feature vector.
[0179] In the field where oviparous = "no", the proportion of the training data with oviparous = "no" among one or more training data that satisfy one or more conditions represented by the rule is set as a component of the feature vector. In the field where oviparous = "yes", the proportion of the training data with oviparous = "yes" among one or more training data that satisfy one or more conditions represented by the above rule is set as a component of the feature vector.
[0180] In the field where legs < 3, the proportion of the training data with the number of legs < 3 among one or more training data that satisfy one or more conditions represented by the above rule is set as a component of the feature vector. In the field where legs ≥ 3, the proportion of the training data with the number of legs ≥ 3 among one or more training data that satisfy one or more conditions represented by the above rule is set as a component of the feature vector.
[0181] In the field where legs < 5, the proportion of the training data with the number of legs < 5 among one or more training data that satisfy one or more conditions represented by the above rule is set as a component of the feature vector. In the field where legs ≥ 5, the proportion of the training data with the number of legs ≥ 5 among one or more training data that satisfy one or more conditions represented by the above rule is set as a component of the feature vector.
[0182] In the field where the tail = "yes", the proportion of the training data with the tail = "yes" among one or more training data that satisfy one or more conditions represented by the rule is set as a component of the feature vector. In the field where the tail = "no", the proportion of the training data with the tail = "no" among one or more training data that satisfy one or more conditions represented by the rule is set as a component of the feature vector.
[0183] Next, with reference to Figure 15 and Figure 16 , a specific example of the information processing device 100 clustering the rule set with reference to the feature vector management table 1300 is described.
[0184] Figure 15 and 16 are explanatory diagrams depicting specific examples of clustering the rule set. In Figure 15 , the information processing device 100 refers to the feature vector management table 1300, and based on the Euclidean distance between the feature vectors, classifies each rule represented by the model 600 into one of a plurality of clusters by the x-means method. For example, the information processing device 100 classifies each rule represented by the model 600 into one of 20 clusters. In the following description, the i-th cluster may be represented as "cluster i". "i" is, for example, from 0 to 19.
[0185] Figure 15 FIG. 1520 in
[0186] depicts the result of arranging each rule represented by the model 600 in a two-dimensional space by principal component analysis. For example, the information processing device 100 classifies the rule group 1500 depicted in FIG. 1520 into cluster 0. For example, the information processing device 100 classifies the rule group 1501 depicted in FIG. 1520 into cluster 1. For example, the information processing device 100 classifies the rule group 1502 depicted in FIG. 1520 into cluster 2. For example, the information processing device 100 classifies the rule group 1503 depicted in FIG. 1520 into cluster 3. For example, the information processing device 100 classifies the rule group 1504 depicted in FIG. 1520 into cluster 4.
[0186] For example, the information processing device 100 classifies the rule group 1505 depicted in FIG. 1520 into cluster 5. For example, the information processing device 100 classifies the rule group 1506 depicted in FIG. 1520 into cluster 6. For example, the information processing device 100 classifies the rule group 1507 depicted in FIG. 1520 into cluster 7. For example, the information processing device 100 classifies the rule group 1508 depicted in FIG. 1520 into cluster 8. For example, the information processing device 100 classifies the rule group 1509 depicted in FIG. 1520 into cluster 9.
[0187] For example, the information processing device 100 classifies the rule group 1510 depicted in FIG. 1520 into cluster 10. For example, the information processing device 100 classifies the rule group 1511 depicted in FIG. 1520 into cluster 11. For example, the information processing device 100 classifies the rule group 1512 depicted in FIG. 1520 into cluster 12. For example, the information processing device 100 classifies the rule group 1513 depicted in FIG. 1520 into cluster 13. For example, the information processing device 100 classifies the rule group 1514 depicted in FIG. 1520 into cluster 14.
[0188] For example, the information processing device 100 classifies the rule group 1515 depicted in FIG. 1520 into cluster 15. For example, the information processing device 100 classifies the rule group 1516 depicted in FIG. 1520 into cluster 16. For example, the information processing device 100 classifies the rule group 1517 depicted in FIG. 1520 into cluster 17. For example, the information processing device 100 classifies the rule group 1518 depicted in FIG. 1520 into cluster 18. For example, the information processing device 100 classifies the rule group 1519 depicted in FIG. 1520 into cluster 19. For example, the information processing device 100 classifies the rule group 1520 depicted in FIG. 1520 into cluster 20. Refer to Figure 16 for the description.
[0189] Figure 16 Table 1600 in
[0190] Next, refer to Figure 17 and Figure 18 to describe the first specific example in which the information processing device 100 generates explanation information related to clusters.
[0191] Figure 17 and Figure 18 are explanatory diagrams depicting the first specific example of generating explanation information related to clusters. As Figure 17 depicted, when the information processing device 100 generates explanation information related to cluster 2, for example, it obtains the rule group 1502 classified into cluster 2. Refer to Figure 18 for the description.
[0192] In Figure 18In this case, the information processing device 100 identifies, for example, one or more pieces of training data that satisfy at least any of the conditions represented by one or more rules included in the obtained rule group 1502 from among a plurality of pieces of training data. The information processing device 100 can, for example, identify one or more pieces of training data that satisfy all of the conditions represented by one or more rules included in the obtained rule group 1502 from among a plurality of pieces of training data. The information processing device 100 can, for example, identify one or more pieces of training data that satisfy all of the conditions represented by z% or more of the rules included in the obtained rule group 1502 from among a plurality of pieces of training data. "z" is, for example, 10.
[0193] The information processing device 100 generates new training data, for example, by assigning labels indicating positive examples to values of a plurality of explanatory variables included in the identified one or more pieces of training data. For example, based on the generated new training data, the information processing device 100 generates a decision tree 1800 for classifying data as explanatory information for cluster 2.
[0194] Next, with reference to Figure 19 and Figure 20 a second specific example in which the information processing device 100 generates explanatory information for a cluster will be described.
[0195] Figure 19 and Figure 20 are explanatory diagrams depicting a second specific example of generating explanatory information for a cluster. As Figure 19 depicted, when generating explanatory information for cluster 11, for example, the information processing device 100 obtains a rule group 1511 classified into cluster 11. A description will be given with reference to Figure 20 this.
[0196] In Figure 20 this case, the information processing device 100 identifies, for example, one or more pieces of training data that satisfy at least any of the conditions represented by one or more rules included in the obtained rule group 1511 from among a plurality of pieces of training data. The information processing device 100 can, for example, identify one or more pieces of training data that satisfy all of the conditions represented by one or more rules included in the obtained rule group 1511 from among a plurality of pieces of training data. For example, the information processing device 100 can identify one or more pieces of training data that satisfy all of the conditions represented by z% or more of the rules included in the obtained rule group 1511 from among a plurality of pieces of training data. "z" is, for example, 10.
[0197] For example, the information processing device 100 can generate new training data by adding a label indicating a positive example to the values of multiple explanatory variables included in one or more pieces of training data that have been identified. For example, based on the generated new training data, the information processing device 100 generates a decision tree 2000 for classifying data as explanatory information for cluster 11. The information processing device 100 similarly generates explanatory information for other clusters.
[0198] Next, with reference to Figure 21 and Figure 22 , a specific example of the information processing device 100 calculating the inclusion rate between multiple clusters will be described.
[0199] Figure 21 and Figure 22 are explanatory diagrams depicting specific examples of calculating the inclusion rate between clusters. In Figure 21 and Figure 22 , the information processing device 100 calculates the inclusion rate between clusters. The information processing device 100 calculates, for each combination of clusters, the inclusion rate of one cluster in the combination with respect to the other cluster in the combination and the inclusion rate of the other cluster with respect to the one cluster. The inclusion rate of one cluster with respect to another cluster is defined as, for example, (A ∩ B) / A. "A" is a set of training data that satisfies at least any of the conditions represented by one or more rules classified into the other cluster. "B" is a set of training data that satisfies at least any of the conditions represented by one or more rules classified into one cluster.
[0200] The information processing device 100 calculates the inclusion rate between clusters, for example, as shown in Table 2100. For example, the rows of Table 2100 correspond to clusters. For example, the columns of Table 2100 correspond to clusters. The element in the i-th row and j-th column of Table 2100 indicates the inclusion rate of cluster j with respect to cluster i. The information processing device 100 can calculate the inclusion rate between clusters by comparing the rules classified into each cluster. For example, the information processing device 100 can calculate the inclusion rate between clusters for each combination of clusters based on the overlapping conditions between one or more conditions represented by one or more rules classified into one cluster and one or more conditions represented by one or more rules classified into another cluster.
[0201] Next, with reference to Figure 23 , a specific example of the information processing device 100 identifying the hierarchical relationship between clusters will be described.
[0202] Figure 23 is an explanatory diagram depicting a specific example of identifying the hierarchical relationship between clusters. In Figure 23In this case, for each combination of clusters, the information processing device 100 determines whether one cluster includes another cluster based on whether the inclusion rate of one cluster relative to another cluster is at least equal to a threshold value. The threshold value is, for example, 0.9. The threshold value is, for example, preset by the user of the model. For example, for any combination of clusters, when the inclusion rate of one cluster in the combination relative to another cluster in the combination is at least equal to the threshold value, the information processing device 100 determines that one cluster includes another cluster. For each combination of clusters, the information processing device 100 determines whether one cluster includes another cluster and determines whether the other cluster includes one cluster, so as to identify the inclusion relationship between the clusters in the plurality of clusters.
[0203] The information processing device 100 identifies the hierarchical relationship between the clusters in the plurality of clusters based on the inclusion relationship between the clusters in the plurality of clusters. For example, for each combination of clusters, the information processing device 100 determines whether one cluster in the combination has a hierarchical relationship in which one cluster is higher than the other cluster in the combination, and whether the other cluster has a hierarchical relationship in which the other cluster is higher than the one cluster. For example, in the case where any combination of clusters has an inclusion relationship in which one cluster includes the other cluster, the information processing device 100 identifies a hierarchical relationship in which one cluster is higher than the other cluster.
[0204] The information processing device 100 generates a graph 2300 showing the identified hierarchical relationship between the clusters. The graph 2300 includes nodes 2301, 2311, 2312, 2313, 2314, 2315, 2321, 2322, 2323, 2324, 2325, 2326, 2327, 2328, 2331, 2332, 2333, 2334, 2335, and 2341 representing the clusters. In Figure 23 the depicted example, the nodes 2301, 2311 to 2315, 2321 to 2328, 2331 to 2335, and 2341 are represented by circles, for example. The numbers attached to the bottoms of the nodes 2301, 2311 to 2315, 2321 to 2328, 2331 to 2335, and 2341 are the cluster numbers represented by the nodes. The graph 2300 includes edges connecting the nodes, and these nodes represent the respective clusters of the cluster combinations having a hierarchical relationship among the nodes 2301, 2311 to 2315, 2321 to 2328, 2331 to 2335, and 2341. For example, the graph 2300 shows that cluster 2 includes all other clusters. The graph 2300 shows, for example, that cluster 18 is included in clusters 2, 15, 17, 4, and 13.
[0205] Refer to Figure 24 for a specific example in which the information processing device 100 displays the hierarchical relationship between the clusters.
[0206] Figure 24 is an explanatory diagram depicting a specific example of displaying the hierarchical relationship between the clusters. InFigure 24 In this case, the information processing device 100 displays the generated Figure 2300.
[0207] The information processing device 100 receives a designation of any one of the nodes 2301, 2311 to 2315, 2321 to 2328, 2331 to 2335, and 2341 of the displayed Figure 2300. For example, the information processing device 100 receives a click on any one of the nodes 2301, 2311 to 2315, 2321 to 2328, 2331 to 2335, and 2341 of the displayed Figure 2300 as the designation of the node.
[0208] In response to the designation of any one of the nodes, the information processing device 100 displays explanation information about one or more rules classified into a cluster represented by any of the designated nodes. For example, in response to the designation of the node 2301, the information processing device 100 displays the decision tree 1800 that serves as the explanation information about the cluster 2 represented by the node 2301. In response to the designation of the node 2315, the information processing device 100 displays the decision tree 2000 that serves as the explanation information about the cluster 11 represented by the node 2315.
[0209] Therefore, the information processing device 100 can enable the user to more easily and intuitively grasp the relationships between the rules forming the model, and enable the user to easily grasp the content of the model. The information processing device 100 can enable the content of one or more rules classified into the cluster to be easily grasped. For example, the information processing device 100 can enable the user to confirm the content of one or more rules classified into each cluster represented by the node group 2350 including the nodes 2311 to 2315. Therefore, the information processing device 100 can, for example, enable the user to confirm the content of each cluster belonging to the same level and having a similar granularity in parallel. Therefore, the information processing device 100 can, for example, enable the user to grasp relatively large classifications of mammals such as herbivores, carnivores, or livestock.
[0210] The information processing device 100 enables the user to, for example, trace down the nodes 2301, 2311, 2321, 2331, and 2341 for confirmation, as shown by the arrow 2360. For example, the information processing device 100 can enable the user to grasp how the animals are classified hierarchically. As described above, the information processing device 100 enables the user to confirm and compare the clusters in the horizontal direction and the clusters in the vertical direction, and enables the user to confirm the content of the model with a desired granularity.
[0211] (Overall processing procedure)
[0212] Next, an example of the overall processing procedure executed by the information processing device 100 will be described with reference to Figure 25 An example of the overall processing executed by, for example,Figure 3 It is implemented by the depicted CPU 301, storage areas such as the memory 302 and the recording medium 305, and the network I / F 303.
[0213] Figure 25 It is a flowchart showing an example of the overall processing procedure. In Figure 25 it, the information processing device 100 obtains a plurality of training data (step S2501).
[0214] Next, the information processing device 100 generates a rule set based on the obtained plurality of training data (step S2502). Next, the information processing device 100 generates a corresponding feature vector for each rule in the generated rule set based on the obtained plurality of training data (step S2503).
[0215] Next, the information processing device 100 classifies each rule in the generated rule set into one of a plurality of clusters by clustering the generated rule set based on the generated feature vectors (step S2504). Then, the information processing device 100 generates corresponding explanation information for each of the plurality of clusters (step S2505).
[0216] Next, the information processing device 100 calculates the inclusion rate between the clusters in the plurality of clusters (step S2506). Next, the information processing device 100 identifies the hierarchical structure between the clusters in the plurality of clusters (step S2507).
[0217] Next, the information processing device 100 outputs information representing the hierarchical structure between the clusters (step S2508). Then, the information processing device 100 ends the overall processing.
[0218] As described above, according to the information processing device 100, a plurality of data including values of a plurality of explanatory variables can be obtained. According to the information processing device 100, a rule set including a plurality of rules can be obtained, and the plurality of rules represent conditions for classifying data using one or more explanatory variables. According to the information processing device 100, based on one or more data among the obtained plurality of data that satisfy the conditions represented by the rules included in the obtained rule set, a feature value related to the rules can be calculated. According to the information processing device 100, based on the similarity between the feature values of the calculated rules, each rule can be classified into any one of a plurality of clusters. According to the information processing device 100, based on the conditions expressed by the rules of the clusters classified into the clusters among the plurality of clusters, the inclusion relationship between the clusters among the plurality of clusters can be identified. According to the information processing device 100, based on the identified inclusion relationship between the clusters among the plurality of clusters, the hierarchical relationship between the clusters among the plurality of clusters can be identified. According to the information processing device 100, information indicating the hierarchical relationship between the clusters among the plurality of clusters can be output. Therefore, the information processing device 100 can make it easier for the user to grasp the content of the rule set.
[0219] According to the information processing device 100, a designation of any one of the plurality of clusters can be received. According to the information processing device 100, information indicating the features of one or more rules classified into any one of the clusters for which the designation has been received can be output. Therefore, the information processing device 100 can make it easier for the user to grasp the content of the cluster.
[0220] According to the information processing device 100, based on training data in which a sample including values of a plurality of explanatory variables is associated with a correct label indicating a classification result of the sample, rules included in the rule set can be generated by machine learning. Therefore, the information processing device 100 can appropriately generate a rule set. The information processing device 100 can generate a rule set by itself and can be easily operated independently.
[0221] According to the information processing device 100, one or more data that satisfy the conditions represented by each rule included in the obtained rule set can be identified from the obtained plurality of data. According to the information processing device 100, a feature vector related to the rules can be calculated, and statistical values related to each of the plurality of explanatory variables in the identified one or more data are components of the feature vector. Therefore, the information processing device 100 can accurately evaluate the similarity of the rules and cluster the rule set.
[0222] According to the information processing apparatus 100, for each cluster, a data set that satisfies a condition represented by one or more rules classified into the cluster can be identified among the obtained multiple data. According to the information processing apparatus 100, for each combination of clusters among the multiple clusters, an inclusion rate of one cluster with respect to another cluster can be calculated. According to the information processing apparatus 100, when the inclusion rate of a first cluster with respect to a second cluster among the multiple clusters is at least equal to a threshold value, it can be determined that the first cluster has an inclusion relationship including the second cluster. According to the information processing apparatus 100, an inclusion relationship between clusters among the multiple clusters can be identified. Therefore, the information processing apparatus 100 can identify the inclusion relationship between clusters with high accuracy. Even when the types of rules included in each cluster are different, the information processing apparatus 100 can enable the evaluation of the inclusion relationship between clusters.
[0223] According to the information processing apparatus 100, for each combination of clusters among the multiple clusters, an inclusion rate of one cluster in the combination with respect to another cluster in the combination can be calculated. According to the information processing apparatus 100, when the inclusion rate of a first cluster with respect to a second cluster among the multiple clusters is at least equal to a threshold value, it can be determined that the first cluster has an inclusion relationship including the second cluster. According to the information processing apparatus 100, an inclusion relationship between clusters among the multiple clusters can be identified. Therefore, the information processing apparatus 100 can identify the inclusion relationship between clusters with high accuracy. Even when the types of rules included in each cluster are different, the information processing apparatus 100 can enable the evaluation of the inclusion relationship between clusters.
[0224] According to the information processing apparatus 100, when a third cluster among the multiple clusters has an inclusion relationship including a fourth cluster, it can be determined that the third cluster has a hierarchical relationship in which the third cluster is higher than the fourth cluster. According to the information processing apparatus 100, a hierarchical relationship between clusters among the multiple clusters can be identified. Therefore, the information processing apparatus 100 can identify the hierarchical relationship between clusters with high accuracy.
[0225] The information processing method described in the present embodiment can be implemented by executing a prepared program on a computer such as a personal computer or a workstation. The information processing program described in the present embodiment is stored in a computer-readable recording medium, and is read out from the recording medium and executed by the computer. The recording medium is a hard disk, a floppy disk, a compact disc read-only memory (CD-ROM), a magneto-optical disc (MO), a digital versatile disc (DVD), or the like. Further, the information processing program described in the present embodiment can be distributed via a network such as the Internet.
[0226] List of Reference Numerals
[0227] 100 Information processing apparatus
[0228] 110 Rule set
[0229] Figures 120, 1000, 1520, 2300
[0230] 200 Information Processing System
[0231] 201 Learning Device
[0232] 202 Client Device
[0233] 210 Network
[0234] 300 Bus
[0235] 301 CPU
[0236] 302 Memory
[0237] 303 Network I / F 304 Recording Medium I / F 305 Recording Medium
[0238] 400 Storage Unit
[0239] 401 Acquisition Unit
[0240] 402 Learning Unit
[0241] 403 Calculation Unit
[0242] 404 Classification Unit
[0243] 405 Identification Unit
[0244] 406 Generation Unit
[0245] 407 Output Unit
[0246] 500 Input
[0247] 501, 502, 503, 504, 505, 506, 507 Processing
[0248] 600 Model
[0249] 701, 702 Rules
[0250] 711, 712 Feature Vectors
[0251] Tables 800, 900, 1600, 2100
[0252] Nodes 1001, 1011, 1012, 1013, 1014, 1015, 1021, 1022, 1023, 1024, 1025, 1026, 1027, 1028, 1031, 1032, 1033, 1034, 1035, 1041, 2301, 2311, 2312, 2313, 2314, 2315, 2321, 2322, 2323, 2324, 2325, 2326, 2327, 2328, 2331, 2332, 2333, 2334, 2335, 2341
[0253] 1010 Explanation Information
[0254] 1100 Animal Data Management Table
[0255] 1300 Feature Vector Management Table
[0256] 1500, 1501, 1502, 1503, 1504, 1505, 1506, 1507, 1508, 1509, 1510, 1511, 1512, 1513, 1514, 1515, 1516, 1517, 1518, 1519 Rule Group 1800, 2000 Decision Tree
Claims
1. An information processing program for causing a computer to perform processing, the processing including: Obtaining a plurality of data including values of a plurality of explanatory variables; Obtaining a rule set including rules, the rules representing conditions for classifying the plurality of data using one or more of the plurality of explanatory variables; Calculating a plurality of feature values, each of the plurality of feature values being calculated based on one or more of the obtained plurality of data for a corresponding rule in the rule set, the one or more data satisfying the conditions represented by the corresponding rule in the rule set; Classifying the rules in the rule set into any of a plurality of clusters based on the similarity between the feature values among the calculated plurality of feature values; Identifying an inclusion relationship between the clusters in the plurality of clusters based on the conditions represented by the rules classified into each of the plurality of clusters; Identifying a hierarchical relationship between the clusters in the plurality of clusters based on the identified inclusion relationship between the clusters in the plurality of clusters; And Outputting information indicating the hierarchical relationship between the clusters in the identified plurality of clusters.
2. The information processing program according to claim 1, the processing further including: Receiving a designation of any one of the plurality of clusters, wherein The output includes output information indicating features of one or more rules classified into any one of the clusters that received the designation.
3. The information processing program according to claim 1 or 2, wherein The rules included in the rule set are generated by machine learning based on training data, in which samples including values of the plurality of explanatory variables are associated with correct labels indicating classification results of the samples.
4. The information processing program according to claim 1 or 2, wherein The calculation includes: calculating a feature vector related to each rule included in the obtained rule set, the feature vector having statistical values related to each of the plurality of explanatory variables in one or more of the obtained plurality of data as components, the one or more data satisfying the conditions represented by each rule.
5. The information processing program according to claim 1 or 2, the processing further including: For each of the plurality of clusters, identifying a data set from the obtained plurality of data that satisfies the conditions represented by the one or more rules classified into each of the plurality of clusters; And For each combination of a first cluster and a second cluster among the plurality of clusters, calculating an inclusion rate of the second cluster in the first cluster, the inclusion rate representing the proportion of overlapping data between a first identified data set that satisfies the conditions represented by one or more rules classified into the first cluster and a second identified data set that satisfies the conditions represented by one or more rules classified into the second cluster in the entire second data set that satisfies the conditions represented by the one or more rules classified into the second cluster, wherein Identifying the inclusion relationship includes: when the inclusion rate of the first cluster in the second cluster is at least equal to a threshold value, determining that the first cluster includes the second cluster, thereby identifying the inclusion relationship.
6. The information processing program according to claim 1 or 2, wherein the processing further includes: For each combination of a first cluster and a second cluster among the multiple clusters, calculating the inclusion rate of the second cluster in the first cluster, where the inclusion rate represents the proportion of the condition of overlap between a first set of conditions represented by one or more rules classified into the first cluster and a second set of conditions represented by one or more rules classified into the second cluster in the entire second set of conditions represented by the one or more rules classified into the second cluster, where Identifying the inclusion relationship includes: when the inclusion rate of the second cluster in the first cluster is at least equal to a threshold value, identifying the inclusion relationship by determining that the first cluster includes the second cluster.
7. The information processing program according to claim 5, wherein Identifying the hierarchical relationship includes: when a third cluster among the multiple clusters includes a fourth cluster among the multiple clusters, identifying the hierarchical relationship between the third cluster and the fourth cluster by determining that the third cluster is higher than the fourth cluster in the hierarchical relationship.
8. An information processing method executed by a computer, the method including: Obtaining a plurality of data including values of a plurality of explanatory variables; Obtaining a rule set including rules, where the rules represent conditions for classifying the plurality of data using one or more of the plurality of explanatory variables; Calculating a plurality of feature values, each of the plurality of feature values being calculated based on one or more of the obtained plurality of data for a corresponding rule in the rule set, where the one or more data satisfy the conditions represented by the corresponding rule in the rule set; Classifying the rules in the rule set into any of a plurality of clusters based on the similarity between the feature values among the calculated plurality of feature values; Identifying the inclusion relationship between the clusters in the plurality of clusters based on the conditions represented by the rules classified into each of the plurality of clusters; Identifying the hierarchical relationship between the clusters in the plurality of clusters based on the identified inclusion relationship between the clusters in the plurality of clusters; And Outputting information indicating the hierarchical relationship between the clusters in the identified plurality of clusters.
9. An information processing apparatus, including a controller, wherein, The controller is configured to: Obtain a plurality of data including values of a plurality of explanatory variables; Obtain a rule set including rules, where the rules represent conditions for classifying the plurality of data using one or more of the plurality of explanatory variables; Calculating a plurality of feature values, each of the plurality of feature values being calculated based on one or more of the obtained plurality of data for a corresponding rule in the rule set, where the one or more data satisfy the conditions represented by the corresponding rule in the rule set; Classifying the rules in the rule set into any of a plurality of clusters based on the similarity between the feature values among the calculated plurality of feature values; Identify the inclusion relationships between clusters among the multiple clusters based on the conditions represented by the rules classified into each of the multiple clusters; Identify the hierarchical relationships between clusters among the multiple clusters based on the identified inclusion relationships between clusters among the multiple clusters; And Output information indicating the hierarchical relationships between clusters among the multiple clusters that have been identified.
Citation Information
Patent Citations
Rule Determination for Black-Box Machine-Learning Models
US20190147369A1