Explanation Information Output Program, Explanation Information Output Method, and Information Processing Apparatus
The information processing apparatus addresses the challenge of visualizing overall trends in XAI by clustering data based on factor contribution degrees, enhancing the efficiency of decision-making in machine learning model outputs.
Patent Information
- Application Number
- JP2021129880
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-08-06
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-08-06
AI Technical Summary
Existing explainable AI (XAI) technologies struggle to show the overall trend of factor contribution degrees for machine learning model predictions, making it difficult and time-consuming for users to formulate optimal measures.
An information processing apparatus that clusters data based on factor contribution degrees and outputs diagrams showing the magnitude of these contributions for each group, enabling visualization of the overall trend in machine learning model outputs.
Enables users to quickly grasp the overall trend of machine learning model predictions, facilitating timely and effective decision-making.
Smart Images

Figure 0007700565000001 
Figure 0007700565000002 
Figure 0007700565000003
Abstract
Description
Technical Field
[0001] The present invention relates to an explanatory information output program and the like.
Background Art
[0002] In recent years, machine learning models generated by machine learning (AI: Artificial Intelligence) have been used. Due to the nature of the mechanism, machine learning models are basically difficult to interpret in one aspect, and explainable AI (XAI: Explainable AI) is used to address this. XAI is a technology that outputs the factor contribution degree for each feature amount input to the machine learning model and presents to humans in an explainable manner which feature amounts led to the prediction result or estimation result.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Patent Document 3
Patent Document 4
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the above technology, although the factor contribution degree can be calculated as explanatory information for the prediction result, it is difficult to show the overall trend.
[0005] For example, since XAI outputs the factor contribution degree for each prediction result (instance) of AI, the user has to individually check the relationship between each prediction result and the factor contribution degree in order to grasp the overall trend. As a result, when the user takes measures against the prediction result based on the factor contribution degree, it takes time and it is difficult to formulate the optimal measures against the prediction result.
[0006] On one aspect, an object is to provide an explanation information output program, an explanation information output method, and an information processing apparatus that can show the overall trend with respect to the output result of a machine learning model.
Means for Solving the Problem
[0007] In the first aspect, the explanation information output program causes a computer to execute a process of acquiring the contribution degree of each of a plurality of factors included in each of the plurality of data with respect to the output result of the machine learning model when each of the plurality of data is input, clustering the plurality of data based on the contribution degree of each of the plurality of factors, and outputting explanation information including a diagram showing the magnitude of the contribution degree of each of the plurality of factors with respect to the output result when the data included in the group is input for each of the plurality of groups generated by the clustering.
Effect of the Invention
[0008] According to one embodiment, the overall trend with respect to the output result of the machine learning model can be shown.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
MODE FOR CARRYING OUT THE INVENTION
[0010] Hereinafter, embodiments of the explanatory information output program, explanatory information output method, and information processing apparatus disclosed in the present application will be described in detail with reference to the drawings. Note that the present invention is not limited to these embodiments. Also, the respective embodiments can be appropriately combined within a non - contradictory range.
[0011] FIG. 1 is a diagram for explaining an information processing apparatus 10 according to Embodiment 1. The information processing apparatus 10 shown in FIG. 1 is an example of a computer that uses XAI to which an algorithm such as LIME is applied to convert the prediction result of a machine learning model into explanatory information that can be visually understood by a user and output it.
[0012] Specifically, the information processing apparatus 10 acquires the contribution degree of each of a plurality of factors included in each of a plurality of data with respect to the output result of the machine learning model when each of the plurality of data is input. The information processing apparatus 10 clusters the plurality of data based on the contribution degree of each of the plurality of factors. The information processing apparatus 10 outputs explanatory information including a diagram representing the magnitude of the contribution degree of each of the plurality of factors with respect to the output result when the data included in the group is input, for each of the plurality of groups generated by clustering.
[0013] For example, as shown in FIG. 1, the information processing apparatus 10 inputs each input data having feature amount A, feature amount B, feature amount C, and feature amount D into a machine learning model to obtain each prediction result. Then, the information processing apparatus 10 uses the input data, the prediction result, and XAI to obtain the contribution degree of each of the factors A, B, C, and D included in each input data. Here, the contribution degree of a factor is information indicating the degree to which each feature amount contributes to the prediction result. Factor A indicates the contribution degree of feature amount A, factor B indicates the contribution degree of feature amount B, factor C indicates the contribution degree of feature amount C, and factor D indicates the contribution degree of feature amount D.
[0014] Subsequently, the information processing apparatus 10 clusters the contribution degrees of the factors corresponding to each input data. For example, the information processing apparatus 10 specifies each input data by each vector of factors A, B, C, and D in a feature space having feature amount a, feature amount b, feature amount c, and feature amount d as each dimension (4D), and clusters each input data.
[0015] Thereafter, the information processing apparatus 10 sorts and displays, for each cluster, the ratio of each factor to the whole of the cluster in terms of an area ratio. In this way, since the information processing apparatus 10 clusters the prediction result by the reason (factor vector) of the output of the machine learning model and displays it in a form such as an area ratio that is visually easy to understand, it can show the overall tendency with respect to the output result of the machine learning model.
[0016] FIG. 2 is a functional block diagram showing the functional configuration of the information processing apparatus 10 according to the first embodiment. As shown in FIG. 2, the information processing apparatus 10 includes a communication unit 11, an output unit 12, a storage unit 13, and a control unit 20.
[0017] The communication unit 11 controls communication with other devices. For example, the communication unit 11 receives an instruction to start processing and data to be determined (input data) from an administrator terminal or the like, and transmits the processing result by the control unit 20 to the administrator terminal.
[0018] The output unit 12 outputs various types of information. For example, the output unit 12 outputs the output results of the machine learning model 15 described later, the explanatory information generated by the control unit 20, and the like.
[0019] The storage unit 13 stores various types of data, programs executed by the control unit 20, and the like. This storage unit 13 stores the training data DB 14, the machine learning model 15, and the input data DB 16.
[0020] The training data DB 14 is a database that stores the training data used for the machine learning of the machine learning model 15. Specifically, the training data DB 14 stores a set of training data having a plurality of feature amounts and correct answer information (labels). As an example, the training data used for generating a machine learning model for determining the possibility of contract termination from contract information in a telecommunications carrier will be described.
[0021] FIG. 3 is a diagram showing an example of information stored in the training data DB 14. As shown in FIG. 3, each piece of training data stored in the training data DB 14 has "member ID, gender, age, contract period, monthly amount, annual income, average communication volume, label". Here, each of "gender, age, contract period, monthly amount, annual income, average communication volume" is a feature amount, and "label" is correct answer information. Note that a member identifier is set for the "member ID", the member's gender is set for the "gender", the member's age is set for the "age", and the member's contract period is set for the "contract period". The monthly charge for which the member has contracted is set for the "monthly amount", the member's annual income is set for the "annual income", the average value of the data communication volume used by the member per month is set for the "average communication volume", and the presence or absence of contract termination is set for the "label".
[0022] In the example of FIG. 3, for the member with member ID = 1, "male, in his 40s, contract is more than 2 years, monthly amount is 8000 yen, annual income is 8 million yen, uses an average of 5 GB of communication volume" is set, and it is set that this member is "continuing" without terminating the contract.
[0023] The machine learning model 15 is a machine learning model generated using the training data stored in the training data DB 14 so as to output a determination result according to the input data having a plurality of feature amounts. In the example of the above communication carrier, when the input data is input, the machine learning model 15 outputs the probability of cancellation and the probability of non-cancellation. Note that a neural network, deep learning, or the like can be adopted for the machine learning model 15.
[0024] The input data DB 16 is a database that stores the input data to be input to the machine learning model 15 and stores the input data to be determined. In the example of the above communication carrier, each input data stored in the input data DB 16 is data having the feature amounts of the member whose cancellation is to be determined.
[0025] FIG. 4 is a diagram showing an example of the information stored in the input data DB 16. As shown in FIG. 4, each input data stored in the input data DB 16 has "member ID, gender, age, contract period, monthly amount, annual income, average traffic volume". Here, each of "gender, age, contract period, monthly amount, annual income, average traffic volume" is a feature amount. Note that the description of each feature amount is the same as that in FIG. 3, and thus the detailed description is omitted. In the example of FIG. 4, for the data of the member with the member ID = 01, the feature amounts are set to "male, in his 50s, 5 years or more, monthly amount of 8,000 yen, annual income of 12 million yen, average 2 GB".
[0026] The control unit 20 is a processing unit that controls the entire information processing apparatus 10 and includes a machine learning unit 21, a prediction unit 22, an explanation execution unit 23, and a display control unit 24.
[0027] The machine learning unit 21 generates a machine learning model 15 using each piece of training data stored in the training data DB 14. Specifically, the machine learning unit 21 trains the machine learning model 15 by supervised learning using the training data. Explaining with the example in FIG. 3, the machine learning unit 21 acquires the training data of the member ID from the training data DB 14, and inputs feature quantities such as gender into the machine learning model 15. Then, the machine learning unit 21 executes the machine learning of the machine learning model 15 so that the error between the output value of the machine learning model 15 and the label "continuing (no cancellation)" becomes small.
[0028] The prediction unit 22 executes a prediction using the machine learning model 15 for each piece of input data stored in the input data DB 16. Explaining with the above example, the prediction unit 22 acquires the input data with the member ID of 01 from the input data DB 16, and inputs feature quantities such as gender into the machine learning model 15. Then, the prediction unit 22 predicts whether the member with the member ID = 01 will cancel the contract or not using the output result of the machine learning model 15. Note that the prediction unit 22 displays the prediction result on the output unit 12 and stores it in the storage unit 13.
[0029] The explanation execution unit 23 generates explanation information that can be confirmed by the user for each prediction result by the prediction unit 22. Specifically, the explanation execution unit 23 acquires the contribution degree of each of the plurality of factors included in each of the plurality of input data with respect to the output result of the machine learning model 15 when each of the plurality of input data is input. Then, the explanation execution unit 23 clusters the plurality of input data based on the contribution degree of each of the plurality of factors. After that, for each of the plurality of groups generated by the clustering, the explanation execution unit 23 generates explanation information including a diagram showing the magnitude (proportion) of the contribution degree of each of the plurality of factors with respect to the output result when the input data included in the group is input.
[0030] First, the explanation execution unit 23 obtains the factor contribution degrees using each input data, the prediction result, and XAI. FIG. 5 is a diagram for explaining the acquisition of the factor contribution degrees. As shown in FIG. 5, the explanation execution unit 23 inputs input data having feature amount a, feature amount b, feature amount c, and feature amount d into the machine learning model 15 and obtains a prediction result. Then, the explanation execution unit 23 generates each neighboring data with the feature amounts of the input data variously changed, inputs each neighboring data into the machine learning model 15, and obtains each prediction result.
[0031] Subsequently, the explanation execution unit 23 inputs the input data and the prediction result, each neighboring data and each prediction result into XAI, and generates an explainable model (linear regression model) that locally approximates using the input data and the neighboring data for the complex machine learning model 15. Then, the explanation execution unit 23 obtains the contribution degrees of factor A corresponding to feature amount a, factor B corresponding to feature amount b, factor C corresponding to feature amount c, and factor D corresponding to feature amount d by calculating the partial regression coefficients of the linear regression model.
[0032] In this way, the explanation execution unit 23 obtains the prediction result and the factor contribution degree for each of the N input data. Note that the acquisition of the factor contribution degree using XAI is not limited to the above-described processing, and known methods such as algorithms such as LIME can be adopted.
[0033] Next, the explanation execution unit 23 clusters the input data using the factor contribution degrees and calculates the proportion of the factors within each cluster for each cluster. FIG. 6 is a diagram for explaining the calculation of the proportion of the factors. As shown in FIG. 6, the explanation execution unit 23 maps N factor contribution groups for each of the N input data into a four-dimensional feature space having feature amount a, feature amount b, feature amount c, and feature amount d as each axis. That is, the explanation execution unit 23 maps each input data to the position in the feature space specified by a factor vector having the contribution degree of factor A, the contribution degree of factor B, the contribution degree of factor C, and the contribution degree of factor D as each vector.
[0034] After that, the explanation execution unit 23 clusters the input data in the feature space to generate clusters such as Cluster 1, Cluster 2, and Cluster 3. Then, the explanation execution unit 23 generates the proportion of factors within each cluster. For example, the explanation execution unit 23 obtains the factor contribution degrees of each input data belonging to Cluster 4, which is an example of the first group, and calculates the total contribution degree of Factor A, the total contribution degree of Factor B, the total contribution degree of Factor C, and the total contribution degree of Factor D. Then, the explanation execution unit 23 calculates the ratios of Factor A, Factor B, Factor C, and Factor D occupied by each factor in Cluster 4.
[0035] In this way, the explanation execution unit 23 represents the ratio (proportion) occupied by each factor in each cluster as an area ratio, generates explanation information including a diagram in which each area ratio is sorted, outputs it to the display control unit 24, stores it in the storage unit 13, and outputs it to the output unit 12.
[0036] The display control unit 24 visualizes the explanation information generated by the explanation execution unit 23 and displays and outputs it to the output unit 12. For example, the display control unit 24 displays and outputs information obtained by subdividing the cluster by mapping the instances within the cluster with the feature amount as the axis.
[0037] FIG. 7 is a diagram for explaining a display example of the explanation information. As shown in FIG. 7, the display control unit 24 displays together a diagram showing the proportion of factors within Cluster 4 generated by the explanation execution unit 23 (FIG. 7(a)) and a diagram obtained by subdividing Cluster 4 (FIG. 7(b)). Here, FIG. 7(a) is a diagram showing the proportion of factors generated by the method described with reference to FIG. 6. FIG. 7(b) is a diagram generated by the display control unit 24. For example, the display control unit 24 maps each input data within Cluster 4 using Factor A and Factor B of the input data within Cluster 4 as vectors (factor vectors) in a two-dimensional space with Factor A (age) and Factor B (annual income), which are the specified factors specified by the user, among the factors within Cluster 4 as the axes. Then, the display control unit 24 generates a pie chart representing the number of each of the plurality of factors included in each input data.
[0038] More specifically, the display control unit 24 calculates the total value of the factor contribution degrees of each factor for the input data corresponding to elderly people whose age is equal to or greater than the threshold among the input data in the cluster 4, and generates a pie chart showing the ratio of the factors using the total value of the factor contribution degrees of each factor. Similarly, the display control unit 24 calculates the total value of the factor contribution degrees of each factor for the input data corresponding to the young generation whose age is less than the threshold, and generates and displays a pie chart showing the ratio of the factors using the total value of the factor contribution degrees of each factor. In the pie chart, the number of instances in the cluster is represented by the area of the circle.
[0039] As a result, the display control unit 24 can present explanatory information obtained by subdividing the cluster by mapping the instances in the cluster with the feature amount as the axis for the cluster clustered from the factor contribution degrees. For example, since the clusters are clustered by the factor contribution degrees, there may be cases where the actual feature amounts are different even if the factor contribution degrees are close. Taking the prediction of those who withdraw from the membership as an example, when factor A with a high factor contribution degree is the age group, there may be no difference in the factor contribution degrees such as the monthly fee even though there are differences in the average annual income etc. between the young generation and the elderly. On the other hand, by generating and displaying information obtained by subdividing the cluster by the display control unit 24, the difference in the factor contribution degrees of the feature amounts with low factor contribution degrees becomes clear, and the user can visually confirm the difference.
[0040] That is, although users with similar factor contribution degrees are clustered within the same cluster, there may be differences in the actual feature amounts, so it is conceivable that this cannot be read only from (a) of FIG. 7. Also, when the selection of the number of clusters is not appropriate, there may be cases where the tendency of the factor contribution degrees is different even within the same cluster, such as factor A (age group) of user A being "factor contribution degree = 0.5, actual value of the feature amount = 60s" and factor A (age group) of user B being "factor contribution degree = 0.5, actual value of the feature amount = 20s".
[0041] Even in this case, assuming that the clusters are divided into two according to age as shown in Fig. 7(b) generated by the display control unit 24, if there are differences in average annual income and family composition between the younger generation and the elderly, subtle differences in trends can be visualized in the feature quantities with lower factor contribution degrees, and the user can subdivide and check the clusters. Also, the user can check the trend of the real values of the feature quantities of factor A (age), which cannot be confirmed only from the graph in Fig. 7(a).
[0042] As another example, the display control unit 24 visualizes the characteristics of the overall trend more specifically by mapping each cluster according to the axis of the feature quantity.
[0043] Fig. 8 is a diagram for explaining an example of displaying explanatory information. As shown in Fig. 8, the display control unit 24 displays together a diagram showing the proportion of factors in cluster 4 generated by the explanation execution unit 23 (Fig. 8(a)) and the mapping result of each cluster (Fig. 8(b)). Here, Fig. 8(a) is a diagram showing the proportion of factors generated by the method described with reference to Fig. 6. Fig. 8(b) is a diagram generated by the display control unit 24.
[0044] For example, the display control unit 24 generates a pie chart of the proportion of factors for each cluster. Then, the display control unit 24 maps the pie chart corresponding to each cluster into the two-dimensional space of feature quantity A and feature quantity B. Specifically, the display control unit 24 calculates the total value of the factor contribution degrees of each factor for the input data in cluster 1 where feature quantity A and feature quantity B are equal to or greater than the threshold, and generates a pie chart of the proportion of factors using the total value of the factor contribution degrees of each factor. Similarly, the display control unit 24 calculates the total value of the factor contribution degrees of each factor for the input data in cluster 2 where feature quantity A is equal to or greater than the threshold and feature quantity B is less than the threshold, and generates a pie chart of the proportion of factors using the total value of the factor contribution degrees of each factor.
[0045] As a result, the user can check clusters with a large number of feature quantities for which countermeasures are desired, such as the contract period, age, and gender in the case of predicting withdrawal, making it easier to formulate countermeasures. Regarding the feature quantities to be used as axes, those with a high factor contribution degree may be adopted, or the user can arbitrarily select them.
[0046] For example, in the graph shown in Fig. 8(a), the tendency of the factor contribution degree of each cluster can be confirmed, but it is difficult to visually confirm the real values of the feature quantities of each factor and how many users are included in each cluster. On the other hand, with the pie chart shown in Fig. 8(b) generated by the display control unit 24, the user can easily visually confirm that there are many users with a high age and a long contract period in cluster 4, that there are many users in cluster 4, and that there are few users in cluster 1.
[0047] Fig. 9 is a flowchart showing the processing flow according to the first embodiment. As shown in Fig. 9, when the control unit 20 of the information processing apparatus 10 is instructed to start processing (S101: Yes), it generates a machine learning model 15 using training data (S102).
[0048] Subsequently, the control unit 20 of the information processing apparatus 10 inputs the input data into the machine learning model 15 (S103), obtains the prediction result (S104), and obtains the factor contribution degree using XAI or the like (S105). Here, if there is unprocessed input data (S106: Yes), the control unit 20 returns to S103 and executes the subsequent processing for the next input data.
[0049] On the other hand, if there is no unprocessed input data (S106: No), the control unit 20 clusters the input data using the factor contribution degree (S107). Then, the control unit 20 calculates the factor proportion for each clustered cluster (S108), and generates and outputs an explanation screen for displaying the explanation information (S109).
[0050] As described above, the information processing apparatus 10 can classify the prediction results (instances) into clusters and output them according to the ratio of the factor contribution degrees for each cluster. As a result, when the user checks the prediction results, the user can check them in the order of the displayed factor contribution degrees, so that the tendency of the overall prediction results can be grasped.
[0051] Also, when clustering using the feature amounts that are inputs for prediction, the weights of the respective feature amounts become equal. However, since the information processing apparatus 10 clusters using the vector of the factor contribution degrees, it is possible to weight each feature amount according to the factor contribution degree. As a result, the information processing apparatus 10 can display the prediction results within the cluster in association with the tendency of the factor contribution degrees, and can improve the visibility for the user.
[0052] The data examples, number of clusters, feature amounts, number of feature amounts, factors, graph examples, screen examples, etc. used in the above embodiments are merely examples and can be arbitrarily changed. Note that a cluster is an example of a group. Also, as an example of the magnitude of the factor contribution degree, area, specific gravity, etc. were exemplified, but the present invention is not limited thereto, and for example, each index such as a numerical value, a total value within a cluster, an average value, etc. can also be used. Also, the axes of the feature amounts described with reference to FIGS. 7 and 8 can be arbitrarily set and changed.
[0053] Also, in the above embodiments, the case of the cancellation by a communications carrier was described as an example, but the present invention is not limited thereto. For example, the information processing apparatus 10 can be applied to various analyses such as detection of a suspicious person using voice data or image data.
[0054] Regarding the processing procedures, control procedures, specific names, information including various data and parameters shown in the above documents and drawings, they may be arbitrarily changed unless otherwise specified.
[0055] In addition, the specific forms of distribution and integration of the components of each device are not limited to those shown in the figures. For example, the explanation execution unit 23 and the display control unit 24 may be integrated. That is, all or part of the components may be functionally or physically distributed and integrated in any unit according to various loads, usage situations, etc. Furthermore, each processing function of each device may be realized in whole or in any part thereof by a CPU and a program analyzed and executed by the CPU, or may be realized as hardware by wired logic.
[0056] FIG. 10 is a diagram for explaining a hardware configuration example. As shown in FIG. 10, the information processing apparatus 10 includes a communication device 10a, an HDD (Hard Disk Drive) 10b, a memory 10c, and a processor 10d. In addition, each part shown in FIG. 10 is mutually connected by a bus or the like. Note that the information processing apparatus 10 may have a display, a touch panel, or the like.
[0057] The communication device 10a is a network interface card or the like and communicates with other devices. The HDD 10b stores programs and databases for operating the functions shown in FIG. 2.
[0058] The processor 10d reads a program for executing the same processing as each processing unit shown in FIG. 2 from the HDD 10b or the like and expands it in the memory 10c, thereby operating a process for executing each function described in FIG. 2 and the like. For example, this process executes the same functions as each processing unit included in the information processing apparatus 10. Specifically, the processor 10d reads a program having the same functions as the machine learning unit 21, the prediction unit 22, the explanation execution unit 23, the display control unit 24, etc. from the HDD 10b or the like. Then, the processor 10d executes a process for executing the same processing as the machine learning unit 21, the prediction unit 22, the explanation execution unit 23, the display control unit 24, etc.
[0059] In this way, the information processing apparatus 10 operates as an information processing apparatus that executes the explanation information output method by reading and executing a program. Also, the information processing apparatus 10 can read the program from a recording medium by a medium reading device and realize the same functions as those in the above-described embodiments by executing the read program. Note that the program in other embodiments is not limited to being executed by the information processing apparatus 10. For example, the above-described embodiments may be similarly applied when another computer or server executes the program, or when these cooperate to execute the program.
[0060] This program may be distributed via a network such as the Internet. Also, this program may be recorded on a computer-readable recording medium such as a hard disk, a flexible disk (FD), a CD-ROM, a MO (Magneto-Optical disk), a DVD (Digital Versatile Disc), and may be executed by being read from the recording medium by a computer.
Explanation of Signs
[0061] 10 Information processing apparatus 11 Communication unit 12 Output unit 13 Storage unit 14 Training data DB 15 Machine learning model 16 Input data DB 20 Control unit 21 Machine learning unit 22 Prediction unit 23 Explanation execution unit 24 Display control unit
Claims
1. When each of a plurality of data is input, obtain the contribution degree of each of a plurality of factors included in each of the plurality of data with respect to the output result of the machine learning model, Cluster the plurality of data based on the contribution degree of each of the plurality of factors, For each of the plurality of groups generated by the clustering, cause the computer to execute a process of outputting explanatory information including a diagram representing the magnitude of the contribution degree of each of the plurality of factors with respect to the output result when the data included in the group is input, The process of outputting is as follows: For each of the plurality of groups, calculate the total value of the contribution degree of each of the plurality of factors with respect to the output result of each data included in the group, For each of the plurality of groups, generate the diagram representing the magnitude of the contribution degree of each of the plurality of factors based on the total value of the contribution degree of each of the plurality of factors included in the group, Output the explanatory information including the diagram corresponding to each of the plurality of groups. An explanatory information output program characterized by this.
2. The process of outputting is as follows: For each of the plurality of groups, use the total value of the contribution degree of each of the plurality of factors included in the group to calculate the ratio that each of the plurality of factors occupies within the group, For each of the plurality of groups, generate the diagram in which the ratio of each of the plurality of factors is represented by an area ratio, Output the explanatory information including the diagram corresponding to each of the plurality of groups. The explanatory information output program according to claim 1, characterized by this.
3. The process of outputting is as follows: On a feature space with a plurality of specified factors, which are specified factors among the plurality of factors, as axes, use a factor vector with the plurality of specified factors as vectors to identify each data included in a first group among the plurality of groups, Output the explanatory information including a pie chart representing the number of each of the plurality of factors included in each data and the diagram corresponding to the first group. The explanatory information output program according to claim 2, characterized by this.
4. The process of outputting is as follows: For each of the plurality of groups, on a feature space with a plurality of specified factors, which are specified factors among the plurality of factors, as axes, use a factor vector with the plurality of specified factors as vectors to identify each data included in each group, For each of the plurality of groups, generate a pie chart representing the number of each of the plurality of factors included in each of the specified data, Output the explanatory information including the diagram and the pie chart corresponding to each of the plurality of groups, The explanatory information output program according to claim 2, characterized in that.
5. Obtain the contribution degree of each of the plurality of factors included in each of the plurality of data with respect to the output result of the machine learning model when each of the plurality of data is input, Cluster the plurality of data based on the contribution degree of each of the plurality of factors, For each of the plurality of groups generated by the clustering, the computer executes a process of outputting explanatory information including a diagram representing the magnitude of the contribution degree of each of the plurality of factors with respect to the output result when the data included in the group is input, The process of outputting is, For each of the plurality of groups, calculate the total value of the contribution degree of each of the plurality of factors with respect to the output result of each data included in the group, For each of the plurality of groups, generate the diagram representing the magnitude of the contribution degree of each of the plurality of factors based on the total value of the contribution degree of each of the plurality of factors included in the group, An explanatory information output method characterized by outputting the explanatory information including the diagram corresponding to each of the plurality of groups.
6. Obtain the contribution degree of each of the plurality of factors included in each of the plurality of data with respect to the output result of the machine learning model when each of the plurality of data is input, Cluster the plurality of data based on the contribution degree of each of the plurality of factors, Output explanatory information including a diagram representing the magnitude of the contribution degree of each of the plurality of factors with respect to the output result when the data included in the group is input for each of the plurality of groups generated by the clustering, Including a control unit, The control unit is, For each of the plurality of groups, calculate the total value of the contribution degree of each of the plurality of factors with respect to the output result of each data included in the group, For each of the plurality of groups, generate the diagram representing the magnitude of the contribution degree of each of the plurality of factors based on the total value of the contribution degree of each of the plurality of factors included in the group, An information processing apparatus characterized by outputting the explanatory information including the figure corresponding to each of the plurality of groups.
Citation Information
Patent Citations
Data analysis device and data analysis method
JP2020024542A
Data analysis device
JP2020135066A
Visually indicating contributions of clinical risk factors
US20180322955A1
US2021/27191
Explainable artificial intelligence-based sales maximization decision models
WO2021096564A1