Information Processing Program, Information Processing Apparatus, and Information Processing Method

The information processing system effectively identifies and outputs important causal relationships by clustering similar data and conditions, addressing the challenge of pinpointing effective measures in large datasets.

JP7705072B2Active Publication Date: 2025-07-09FUJITSU LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023579976
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-10
Publication Date
2025-07-09
Estimated Expiration
2042-02-10

AI Technical Summary

Technical Problem

Existing methods struggle to identify important causal relationships leading to problem solutions, especially when dealing with large datasets where numerous conditions apply, making it difficult to pinpoint effective measures for individual customers in marketing or other fields.

Method used

An information processing system that classifies data groups and causal graphs into clusters based on similarities, allowing for the identification of important causal relationships by grouping similar data and conditions, and specifying conditions under which these relationships occur.

Benefits of technology

Enables easy identification and output of important causal relationships, facilitating targeted problem-solving strategies by aggregating similar data and conditions, even in large datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007705072000001
    Figure 0007705072000001
  • Figure 0007705072000002
    Figure 0007705072000002
  • Figure 0007705072000003
    Figure 0007705072000003
Patent Text Reader

Abstract

The present invention involves: referring to a storage unit that stores a plurality of sets of data, each consisting of a combination of a plurality of feature quantities, and extracting, for each of a plurality of conditions, a data group for which the combination satisfies the condition; identifying, for each of the plurality of conditions, the relationship between the plurality of feature quantities included in the data group corresponding to the condition; classifying the relationships for the plurality of conditions into a plurality of first clusters; classifying the data groups for the plurality of conditions into a plurality of second clusters; classifying the data groups for the plurality of conditions into a plurality of third clusters in such a way that a plurality of data groups which are classified into the same second cluster and for which the relationships are classified into the same first cluster are classified into the same cluster; and identifying, for each of the plurality of third clusters, a first condition that can be used to classify the data groups classified in the cluster and the data groups classified in other clusters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing program, an information processing apparatus, and an information processing method.

Background Art

[0002] In recent years, in various fields such as marketing and medicine, for example, the formulation of measures for solving various problems has been carried out by AI (Artificial Intelligence). Specifically, the formulation of such measures is carried out by considering, for example, not only the correlation between causes and effects but also the causal relationships that express the relationship between causes and effects. Therefore, in recent years, for example, technologies for estimating causal relationships for the entire data have been studied (see, for example, Non-Patent Document 1).

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Here, for example, in the case of promotion in marketing, each customer who purchases a product has characteristics that lead to the purchase of the product. Therefore, in order to formulate appropriate measures for each customer, it is necessary to identify not only the causal relationships common to all customers, but also the causal relationships for each customer who meets multiple conditions.

[0005] Therefore, when formulating measures to solve various problems, for example, data conditions are obtained based on emerging pattern enumeration, and in addition to the causal relationships for the entire data, a method is used to identify the causal relationships for each data that meets each condition.

[0006] However, for example, when a large number of causal relationships corresponding to each condition are identified, it may not be possible to identify the causal relationships (hereinafter also referred to as important causal relationships) that lead to the solution of the problem.

[0007] Therefore, in one aspect, an object of the present invention is to provide an information processing program, an information processing apparatus, and an information processing method that enable identification of important causal relationships leading to the solution of a problem.

Means for Solving the Problem

[0008] In one aspect of the embodiment, referring to a storage unit that stores a plurality of data composed of combinations of a plurality of feature amounts, for each of the plurality of conditions, a data group in which the combination satisfies each condition is extracted, and for each of the plurality of conditions, the relationship between the plurality of feature amounts included in the data group corresponding to each condition is specified. Based on the first similarity between the relationships for each of the plurality of conditions, the relationships for each of the plurality of conditions are classified into a plurality of first clusters. Based on the second similarity between the data groups for each of the plurality of conditions, the data groups for each of the plurality of conditions are classified into a plurality of second clusters. A plurality of the data groups in which the classified first clusters of the relationships corresponding to each data group are the same and the classified second clusters of each data group are the same are classified into the same cluster. The data groups for each of the plurality of conditions are classified into a plurality of third clusters. For each of the plurality of third clusters, a first condition capable of classifying the data groups classified into each cluster and the data groups classified into other clusters is specified, and the specified first condition is output together with the classification results of the plurality of third clusters. The computer is caused to execute the process.

Effect of the Invention

[0009] According to one aspect, it is possible to identify an important causal relationship leading to the solution of the problem.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

[0011] [Configuration of Information Processing System] First, the configuration of the information processing system 10 will be described. FIG. 1 is a diagram for explaining the configuration of the information processing system 10.

[0012] As shown in FIG. 1, for example, the information processing system 10 includes an information processing apparatus 1 and an operation terminal 5.

[0013] The information processing apparatus 1 is, for example, a physical machine or a virtual machine, and performs a process (hereinafter referred to as a cause identification process) of identifying a causal relationship that causes a target variable from a data group 131 to be processed (hereinafter also referred to as a target data group 131). The target data group 131 is, for example, a data group including a plurality of data formed by combinations of a plurality of feature amounts.

[0014] The operation terminal 5 is, for example, one or more PCs (Personal Computers), and is a terminal for an operator to input necessary information. Specifically, the operation terminal 5 transmits, for example, the target data group 131 input by the operator to the information processing apparatus 1. Hereinafter, the cause identification process in the comparative example will be described.

[0015] [Cause Identification Process in Comparative Example] FIG. 2 is a flowchart for explaining the cause identification process in the comparative example. FIGS. 3 to 5 are diagrams for explaining the cause identification process in the comparative example.

[0016] The information processing apparatus 1 waits, for example, until the cause identification timing (NO in S1). The cause identification timing is, for example, the timing when the operator inputs to the information processing apparatus 1 that the cause identification process is to be started.

[0017] When the cause identification timing is reached (YES in S1), the information processing apparatus 1 refers to, for example, a storage unit storing the target data group 131, and for each of a plurality of conditions (hereinafter also simply referred to as a plurality of conditions) specified in advance by the operator, extracts a data group 131 (hereinafter also referred to as a partial data group 131) in which combinations of a plurality of feature amounts satisfy each condition (S2). The plurality of conditions are, for example, conditions for combinations of feature amounts determined based on emerging pattern enumeration. Hereinafter, a specific example of the process of S2 will be described.

[0018] [Specific Example of the Process of S2] FIG. 3 and FIG. 4 are diagrams for explaining a specific example of S2. Specifically, FIG. 3 and FIG. 4 are diagrams for explaining a specific example of the target data group 131. Hereinafter, it is assumed that the target data group 131 is a data group related to a plurality of students, and the objective variable is the grade of each student, and the explanation will be given. Also, hereinafter, it is assumed that the target data group 131 includes "name", "age", "gender", "weekday study time", "weekday free time", "number of absences", and "commute time" as feature amounts, and the explanation will be given.

[0019] In the target data group 131 shown in FIG. 3, in the data of the first row, for example, "A" is set as the "name", "20" is set as the "age", "male" is set as the "gender", "60 (minutes)" is set as the "weekday study time", "120 (minutes)" is set as the "weekday free time", "0 (days)" is set as the "number of absences", and "30 (minutes)" is set as the "commute time".

[0020] Also, in the target data group 131 shown in FIG. 3, in the data of the second row, for example, "B" is set as the "name", "18" is set as the "age", "female" is set as the "gender", "120 (minutes)" is set as the "weekday study time", "60 (minutes)" is set as the "weekday free time", "0 (days)" is set as the "number of absences", and "20 (minutes)" is set as the "commute time". Explanation of other information included in FIG. 3 is omitted.

[0021] Then, for example, when one condition (hereinafter also referred to as a specific condition) included in a plurality of predetermined conditions is "weekday free time > 60 ∧ commute time < 30", the information processing apparatus 1 specifies, as the partial data group 131 corresponding to the specific condition, the partial data group 131 including the data of the first row, the second row, and the fifth row, as shown in the underlined part of FIG. 4.

[0022] Returning to FIG. 2, for example, for each of a plurality of conditions, the information processing apparatus 1 generates a causal graph 132 showing the relationship among a plurality of feature amounts included in a partial data group 131 corresponding to each condition (S3). Hereinafter, a specific example of the causal graph will be described.

[0023] [Specific Example of Causal Graph] FIG. 5 is a diagram for explaining a specific example of the causal graph 132. Specifically, FIG. 5 is a diagram for explaining a specific example of the causal graph 132 showing the relationship among a plurality of feature amounts included in the partial data group 131 (the partial data group 131 corresponding to a specific condition) described in FIG. 4. Note that each node in the causal graph 132 shown in FIG. 4 corresponds to each of the feature amounts described in FIG. 4. In addition, the arrows and numerical values between the nodes indicate the presence or absence of a causal relationship and the strength of the causal relationship among the plurality of feature amounts described in FIG. 4.

[0024] Specifically, for example, in the causal graph 132 shown in FIG. 5, the arrow from the node corresponding to "father has a teaching profession" to the node corresponding to "grade in the first semester" indicates that when the father of the student has a teaching profession, the grade in the first semester drops by 1.34 (points). Also, the arrow from the node corresponding to "grade in the first semester" to the node corresponding to "grade in the second semester" indicates that when the grade in the first semester rises by 1 (point), the grade in the second semester rises by 0.88 (points). Explanation of other information included in FIG. 5 is omitted.

[0025] Returning to FIG. 2, the information processing apparatus 1, for example, specifies, from the causal graph 132 generated for each of a plurality of conditions, the causal relationship that causes the target variable (S4).

[0026] Specifically, for example, for each of the plurality of causal graphs 132 generated in the process of S3, the information processing apparatus 1 specifies, among the causal relationships included in each causal graph 132, the causal relationships that are not included in other causal graphs 132 (that is, unique causal relationships). Then, the information processing apparatus 1, for example, specifies the specified causal relationships as the conditions under which the causal relationship that causes the objective function appears.

[0027] As a result, the information processing apparatus 1 can, for example, identify a specific causal relationship that does not appear in the causal graph 132 corresponding to the entire target data group 131.

[0028] However, for example, when a large number of causal graphs 132 are generated in the process of S3, the information processing apparatus 1 cannot easily identify important causal relationships (important causal relationships leading to the solution of the purpose) from the generated large number of causal graphs 132.

[0029] Therefore, the information processing apparatus 1 in the present embodiment refers to a storage unit that stores a target data group 131 composed of a combination of a plurality of feature amounts, and extracts a partial data group 131 in which a combination of a plurality of feature amounts satisfies each condition for each of the plurality of conditions. Then, the information processing apparatus 1 generates, for example, a causal graph 132 indicating the relationship between a plurality of feature amounts included in the partial data group 131 corresponding to each condition for each of the plurality of conditions.

[0030] Subsequently, the information processing apparatus 1 classifies the causal graphs 132 for each of the plurality of conditions into a plurality of clusters (hereinafter, also referred to as a plurality of first clusters) based on, for example, the similarity (hereinafter, also referred to as the first similarity) between the causal graphs 132 for each of the plurality of conditions. Further, the information processing apparatus 1 classifies the partial data groups 131 for each of the plurality of conditions into a plurality of clusters (hereinafter, also referred to as a plurality of second clusters) based on, for example, the similarity (hereinafter, also referred to as the second similarity) between the partial data groups 131 for each of the plurality of conditions.

[0031] Thereafter, the information processing apparatus 1 classifies the partial data groups 131 for each of the plurality of conditions into a plurality of clusters (hereinafter, also referred to as a plurality of third clusters) so that a plurality of partial data groups 131 in which the first clusters in which the causal graphs 132 corresponding to the respective partial data groups 131 are classified are the same and the second clusters in which the respective partial data groups 131 are classified are the same are classified into the same cluster.

[0032] Then, for each of the plurality of third clusters, the information processing apparatus 1 specifies a condition (hereinafter also referred to as the first condition) that can classify the partial data group 131 classified into each cluster and the partial data group 131 classified into other clusters, and outputs the specified first condition together with the classification results of the plurality of third clusters (hereinafter also simply referred to as classification results).

[0033] Specifically, for example, for each of the plurality of third clusters, the information processing apparatus 1 outputs, as the classification result, information indicating a causal relationship included in the plurality of causal graphs 132 classified into each cluster and not included in the causal graph 132 corresponding to the entire target data group 131 among the causal relationships included in the plurality of causal graphs 132 classified into each cluster.

[0034] That is, the information processing apparatus 1 in the present embodiment classifies the partial data groups 131 extracted for each of the plurality of conditions into a plurality of third clusters so that, among the partial data groups 131 that can be determined to be essentially similar, the partial data groups 131 for which the corresponding causal graphs 132 can also be determined to be essentially similar are included in the same cluster. Then, for example, the information processing apparatus 1 specifies the first condition, which is a condition under which an important causal relationship leading to problem-solving appears, for each of the plurality of third clusters.

[0035] Thereby, the information processing apparatus 1 in the present embodiment can aggregate combinations that can be determined to be essentially similar even when there are a large number of combinations of the partial data group 131 and the causal graph 132, and can easily specify important causal relationships leading to problem-solving.

[0036] Further, the information processing apparatus 1 in the present embodiment can output, for example, by outputting the classification result corresponding to each cluster and the first condition together, in a form of associating the important causal relationship specified by the cause identification process with the first condition under which the important causal relationship appears. Therefore, for example, the operator can easily grasp the correspondence between the important causal relationship specified by the cause identification process and the first condition under which the important causal relationship appears by viewing each piece of information output by the information processing apparatus 1.

[0037] [Hardware Configuration of Information Processing System] Next, the hardware configuration of the information processing system 10 will be described. FIG. 6 is a diagram for explaining the hardware configuration of the information processing apparatus 1.

[0038] As shown in FIG. 6, the information processing apparatus 1 includes a CPU 101 as a processor, a memory 102, an I / O interface 103, and a storage medium 104. Each unit is connected to each other via a bus 105.

[0039] The storage medium 104 has, for example, a program storage area (not shown) for storing a program 110 for performing cause identification processing (hereinafter also referred to as the information processing program 110). Further, the storage medium 104 has an information storage area 130 for storing information used when performing cause identification processing. Note that the storage medium 104 may be, for example, an HDD or an SSD (Solid State Drive).

[0040] The CPU 101 executes the program 110 loaded from the storage medium 104 into the memory 102 to perform cause identification processing.

[0041] The I / O interface 103 is, for example, an interface device such as a network interface card, and can access the operation terminal 5.

[0042] [Functions of Information Processing System] Next, the functions of the information processing system 10 will be described. FIG. 7 is a block diagram of the functions of the information processing apparatus 1.

[0043] As shown in FIG. 7, for example, in the information processing apparatus 1, various functions including a data reception unit 111, a data extraction unit 112, a graph generation unit 113 (hereinafter also referred to as a relationship identification unit 113), a first similarity calculation unit 114, a second similarity calculation unit 115, a clustering unit 116, a condition identification unit 117, and a condition output unit 118 are realized by the organic cooperation of hardware such as a CPU 101 and a memory 102 and a program 110.

[0044] Further, the information processing apparatus 1 stores, for example, a target data group 131, a causal graph 132, first similarity information 133, second similarity information 134, and importance information 135 in an information storage area 130.

[0045] The data reception unit 111 receives, for example, the target data group 131 input by an operator via the operation terminal 5. Then, the data reception unit 111 stores, for example, the received target data group 131 in the information storage area 130.

[0046] The data extraction unit 112 refers to, for example, the target data group 131 stored in the information storage area 130, and extracts a partial data group 131 in which a combination of a plurality of feature amounts satisfies each of a plurality of conditions specified in advance by the operator for each of the plurality of conditions.

[0047] The graph generation unit 113 generates, for example, a causal graph 132 indicating the relationship between a plurality of feature amounts included in the partial data group 131 corresponding to each condition for each of the plurality of conditions.

[0048] The first similarity calculation unit 114 calculates, for example, first similarity information 133 between the causal graphs 132 for each of the plurality of conditions. Then, the first similarity calculation unit 114 stores, for example, the calculated first similarity information 133 in the information storage area 130.

[0049] The second similarity calculation unit 115 calculates, for example, second similarity information 134 between the partial data groups 131 for each of the plurality of conditions. Then, the second similarity calculation unit 115 stores, for example, the calculated second similarity information 134 in the information storage area 130.

[0050] The clustering unit 116 classifies causal graphs 132 for each of a plurality of conditions into a plurality of first clusters, for example, based on the first similarity calculated by the first similarity calculation unit 114.

[0051] Also, the clustering unit 116 classifies partial data groups 131 for each of a plurality of conditions into a plurality of second clusters, for example, based on the second similarity calculated by the second similarity calculation unit 115.

[0052] Furthermore, the clustering unit 116 classifies partial data groups 131 for each of a plurality of conditions into a plurality of third clusters such that, for example, a plurality of partial data groups 131 for which the first clusters in which the causal graphs 132 corresponding to the respective partial data groups 131 are classified are the same and the second clusters in which the respective partial data groups 131 are classified are the same are classified into the same cluster.

[0053] The condition specifying unit 117 specifies, for example, a first condition that can classify the partial data group 131 classified into each cluster and the partial data group 131 classified into other clusters for each of the plurality of third clusters.

[0054] The condition output unit 118 outputs, for example, the first condition specified by the condition specifying unit 117 to the operation terminal 5 together with the classification results of the plurality of third clusters by the clustering unit 116.

[0055] [Outline of the First Embodiment] Next, the outline of the first embodiment will be described. FIG. 8 is a flowchart for explaining the outline of the cause identification process in the first embodiment.

[0056] The information processing apparatus 1 waits, for example, until the cause identification timing (NO in S11).

[0057] When the cause identification timing is reached (YES in S11), the information processing apparatus 1 refers to, for example, the target data group 131 stored in the information storage area 130, and extracts, for each of a plurality of conditions, a partial data group 131 in which a combination of a plurality of feature amounts satisfies each condition (S12).

[0058] Subsequently, the information processing apparatus 1 generates, for example, a causal graph 132 indicating the relationship between a plurality of feature amounts included in the partial data group 131 corresponding to each condition for each of the plurality of conditions (S13).

[0059] Next, the information processing apparatus 1 classifies the causal graphs 132 for each of the plurality of conditions into a plurality of first clusters based on, for example, the first similarity between the causal graphs 132 for each of the plurality of conditions (S14).

[0060] Further, the information processing apparatus 1 classifies the partial data groups 131 for each of the plurality of conditions into a plurality of second clusters based on, for example, the second similarity between the partial data groups 131 for each of the plurality of conditions (S15).

[0061] Furthermore, the information processing apparatus 1 classifies the partial data groups 131 for each of the plurality of conditions into a plurality of third clusters so that, for example, a plurality of partial data groups 131 in which the first clusters in which the causal graphs 132 corresponding to each partial data group 131 are classified are the same and the second clusters in which each partial data group 131 is classified are the same are classified into the same cluster (S16).

[0062] Thereafter, the information processing apparatus 1 specifies and outputs, for example, a first condition that enables classification between the partial data group 131 classified into each cluster and the partial data group 131 classified into another cluster for each of the plurality of third clusters (S17).

[0063] As a result, the information processing apparatus 1 in the present embodiment can aggregate combinations that can be determined to be essentially close even when there are a large number of combinations of the partial data group 131 and the causal graph 132, and can easily specify important causal relationships leading to problem solving.

[0064] Also, the information processing apparatus 1 in the present embodiment can output, for example, the classification result corresponding to each cluster and the first condition together, so as to output in a form of associating the important causal relationship identified by the cause identification process with the first condition under which the important causal relationship appears. Therefore, for example, by browsing each piece of information output by the information processing apparatus 1, an operator can easily grasp the correspondence between the important causal relationship identified by the cause identification process and the first condition under which the important causal relationship appears.

[0065] Furthermore, the information processing apparatus 1 in the present embodiment can, for example, classify a plurality of data groups 131 that can be determined to be similar to each other into the same cluster among the plurality of data groups 131 that can be determined to be similar to each other, so as to output information indicating whether a plurality of causal graphs 132 that are similar to each other are essentially close, and information indicating whether the data groups 131 that are similar to each other are essentially close. Therefore, for example, by browsing each piece of information output by the information processing apparatus 1, an operator can easily make a judgment on whether a plurality of causal graphs 132 that are similar to each other are essentially close, and a judgment on whether the data groups 131 that are similar to each other are essentially close.

[0066] [Details of the First Embodiment] Next, the details of the first embodiment will be described. FIGS. 9 to 11 are flowchart diagrams for explaining the details of the cause identification process in the first embodiment. Also, FIGS. 12 to 19 are diagrams for explaining the details of the cause identification process in the first embodiment.

[0067] As shown in FIG. 9, the data reception unit 111 waits, for example, until receiving the target data group 131 transmitted from the operation terminal 5 (NO in S21).

[0068] When the target data group 131 is received (YES in S21), the data receiving unit 111 stores the received target data group 131 in the information storage area 130 (S22). Specifically, the data receiving unit 111 stores the target data group 131 described in FIG. 3 in the information storage area 130, for example.

[0069] Thereafter, as shown in FIG. 10, the data extraction unit 112 waits until the cause identification timing, for example (NO in S31).

[0070] When the cause identification timing is reached (YES in S31), the data extraction unit 112 refers to the target data group 131 stored in the information storage area 130 and extracts partial data groups 131 in which combinations of a plurality of feature amounts satisfy each of a plurality of conditions for each of the plurality of conditions (S32).

[0071] Specifically, as shown in FIG. 12, the data extraction unit 112 extracts, for example, a partial data group 131a corresponding to the student S corresponding to the previously specified condition A, a partial data group 131b corresponding to the student S corresponding to the previously specified condition B, and a partial data group 131c corresponding to the student S corresponding to the previously specified condition C from the target data group 131 for a plurality of students S.

[0072] Subsequently, the graph generation unit 113 generates a causal graph 132 showing the relationship between a plurality of feature amounts included in the partial data group 131 corresponding to each condition for each of the plurality of conditions, for example (S33).

[0073] Specifically, as shown in FIG. 13, the graph generation unit 113 generates, for example, a causal graph 132a showing the relationship between a plurality of feature amounts included in the partial data group 131a, a causal graph 132b showing the relationship between a plurality of feature amounts included in the partial data group 131b, and a causal graph 132c showing the relationship between a plurality of feature amounts included in the partial data group 131c.

[0074] Note that the information processing apparatus 1 may perform the processes of S12 and S13 by using, for example, Wide Learning (registered trademark), which is a machine learning technique for generating a learning model (white box type learning model) capable of explaining the reasons for evaluation.

[0075] Next, the first similarity calculation unit 114 calculates, for example, first similarity information 133 between causal graphs 132 for each of a plurality of conditions (S34).

[0076] Specifically, the first similarity calculation unit 114 may calculate, as the first similarity information 133, for example, the distance of the adjacency matrix for the causal graph 132 generated in the process of S32, the distance of the causal effect on the target variable for the causal graph 132 generated in the process of S32, and the like.

[0077] Also, the second similarity calculation unit 115 calculates, for example, second similarity information 134 between partial data groups 131 for each of a plurality of conditions (S35).

[0078] Specifically, the second similarity calculation unit 115 may calculate, as the second similarity information 134, for example, the Jaccard coefficient, Dice coefficient, Simpson coefficient, etc. for the partial data group 131 extracted in the process of S31.

[0079] Then, the clustering unit 116 classifies the causal graphs 132 for each of a plurality of conditions into a plurality of first clusters according to the first similarity information 133, for example (S36).

[0080] Specifically, as shown in FIG. 14, the clustering unit 116 classifies the causal graph 132 generated in the process of S33 into a plurality of first clusters including cluster CL11, cluster CL12, and cluster CL13 such that the causal graphs 132 with high first similarity information 133 calculated in the process of S34 are classified into the same first cluster.

[0081] Further, the clustering unit 116 classifies the partial data groups 131 for each of a plurality of conditions into a plurality of second clusters according to, for example, the second similarity information 134 (S37).

[0082] Specifically, as shown in FIG. 15, the clustering unit 116 classifies the partial data groups 131 extracted in the process of S32 into a plurality of second clusters including a cluster CL21, a cluster CL22, and a cluster CL23 such that, for example, the partial data groups 131 with high second similarity information 134 calculated in the process of S35 are classified into the same second cluster.

[0083] Thereafter, the clustering unit 116 classifies combinations of the partial data groups 131 for each of a plurality of conditions and the causal graphs 132 into a plurality of third clusters so that, for example, a plurality of causal graphs 132 in which the first clusters in which the respective causal graphs 132 are classified are the same and the second clusters in which the partial data groups 131 corresponding to the respective causal graphs 132 are classified are the same are classified into the same cluster (S38).

[0084] In other words, in the process of S38, the clustering unit 116 classifies combinations of the partial data groups 131 for each of a plurality of conditions and the causal graphs 132 into a plurality of third clusters so that, for example, a plurality of partial data groups 131 in which the first clusters in which the causal graphs 132 corresponding to the respective partial data groups 131 are classified are the same and the second clusters in which the respective partial data groups 131 are classified are the same are classified into the same cluster.

[0085] Specifically, as shown in FIG. 16, the clustering unit 116 classifies, for example, the causal graphs 132 corresponding to the partial data groups 131 classified into the cluster CL21 in the process of S37 among the causal graphs 132 classified into the cluster CL11 in the process of S36 into the cluster CL31.

[0086] Further, as shown in FIG. 16, for example, the clustering unit 116 classifies the causal graph 132 corresponding to the partial data group 131 classified into the cluster CL21 in the process of S37 among the causal graphs 132 classified into the cluster CL13 in the process of S36 into the cluster CL33.

[0087] Furthermore, as shown in FIG. 16, for example, the clustering unit 116 classifies the causal graph 132 corresponding to the partial data group 131 classified into the cluster CL22 in the process of S37 among the causal graphs 132 classified into the cluster CL12 in the process of S36 into the cluster CL35. Explanation of other information included in FIG. 16 is omitted.

[0088] That is, the clustering unit 116 performs clustering on the combination of the partial data group 131 and the causal graph 132 according to the similarity between the partial data groups 131 and the similarity between the causal graphs 132, so as to classify such that combinations that are essentially close are included in the same cluster.

[0089] Thereby, the information processing apparatus 1 can aggregate a plurality of partial data groups 131 and a plurality of causal graphs 132 that are essentially closely related.

[0090] Then, the clustering unit 116 excludes, for example, clusters in which the number of partial data groups 131 included in each cluster is equal to or less than a predetermined number from a plurality of third clusters (S39). The predetermined number may be, for example, 1.

[0091] That is, it is possible to determine that combinations not classified into clusters containing many combinations of the partial data group 131 and the causal graph 132 in the process of S38 are outliers. Therefore, the clustering unit 116 excludes, for example, combinations that can be determined as outliers (clusters including only combinations that can be determined as outliers).

[0092] Subsequently, as shown in FIG. 11, for example, the condition specifying unit 117 specifies a common partial data group 131d common to the partial data groups 131 included in each cluster for each of a plurality of third clusters (S41).

[0093] Specifically, as shown in FIG. 17, for example, among the data constituting the plurality of partial data groups 131 included in the cluster CL33, the condition specifying unit 117 specifies, as the common partial data group 131d, data included in partial data groups 131 that account for a predetermined ratio or more (for example, 80 (%) or more).

[0094] Further, for example, for each of a plurality of third clusters, the condition specifying unit 117 generates a common causal graph 132d common to the causal graphs 132 included in each cluster (S41).

[0095] Specifically, as shown in FIG. 18, for example, among the edges constituting the causal graphs 132 corresponding to the plurality of partial data groups 131 included in the cluster CL33, the condition specifying unit 117 generates, as the common causal graph 132d, a new causal graph 132 that includes a predetermined ratio or more (for example, 80 (%) or more) of the edges.

[0096] Thereafter, for example, for each of a plurality of third clusters, the condition specifying unit 117 generates a learning model by performing machine learning using the common partial data group 131d classified into each cluster as a positive example and the common partial data group 131d classified into other clusters as a negative example (S42).

[0097] That is, for example, for each of a plurality of third clusters, the condition specifying unit 117 learns to classify between the data group classified into each cluster and the data group classified into other clusters.

[0098] Specifically, the condition specifying unit 117 can evaluate the target (data) as positive or negative, explain the reason for the evaluation, and comprehensively list the conditions composed of all combinations of variables, for example, by using Wide Leaning (registered trademark). Furthermore, the condition specifying unit 117 generates a learning model that can assign importance (hereinafter also simply referred to as importance) to the conditions listed by using a method such as logistic regression.

[0099] Then, for example, for each of the plurality of third clusters, the condition specifying unit 117 specifies the conditions indicated by the learning model corresponding to each cluster as the first conditions (S43). Hereinafter, a specific example of the process of S43 will be described.

[0100] [Specific Example of the Process of S43] FIG. 19 is a diagram for explaining a specific example of the process of S43. Specifically, FIG. 19 is a diagram for explaining a specific example of importance information 135 indicating the importance of each condition output from the learning model generated in the process of S42.

[0101] The importance information 135 shown in FIG. 19 indicates, for example, that the "importance" of the condition "age < 20 ∧ no repeat year" is "0.9", and the "importance" of the condition "weekday study time > 30 minutes" is "0.6". Explanation of other information included in FIG. 19 is omitted.

[0102] Then, for example, when "0.9" in the importance information 135 shown in FIG. 19 is the maximum value of the "importance", the condition specifying unit 117 specifies, for example, "age < 20 ∧ no repeat year" as the first condition.

[0103] Returning to FIG. 11, the condition output unit 118 outputs, for example, the first conditions specified for each of the plurality of third clusters to the operation terminal 5 in association with the common causal graph 132d (S44).

[0104] Specifically, in this case, the condition output unit 118 may output, for example, the edges included in the common causal graph 132d in a state where the edges not included in the causal graph 132 corresponding to the entire target data group 131 (for example, the causal graph 132 generated in advance by the graph generation unit 113) are emphasized.

[0105] As a result, the information processing apparatus 1 can output, for example, the important causal relationships identified by the cause identification process in an emphasized form. Therefore, the information processing apparatus 1 can output, for example, in a form in which the important causal relationships identified by the cause identification process and the conditions (first conditions) of the partial data group 131 in which the causal relationships appear are associated with each other for each of a plurality of third clusters.

[0106] As described above, the information processing apparatus 1 in the present embodiment refers to the information storage area 130 storing the target data group 131 composed of a combination of a plurality of feature amounts, and extracts, for each of a plurality of conditions, the partial data group 131 in which the combination of the plurality of feature amounts satisfies each condition. Then, the information processing apparatus 1 identifies, for example, the causal graph 132 among the plurality of feature amounts included in the partial data group 131 corresponding to each condition for each of the plurality of conditions.

[0107] Subsequently, the information processing apparatus 1 classifies the causal graphs 132 for each of the plurality of conditions into a plurality of first clusters based on, for example, the first similarity among the causal graphs 132 for each of the plurality of conditions. Further, the information processing apparatus 1 classifies the partial data groups 131 for each of the plurality of conditions into a plurality of second clusters based on, for example, the second similarity among the partial data groups 131 for each of the plurality of conditions.

[0108] Thereafter, the information processing apparatus 1 classifies the partial data groups 131 for each of the plurality of conditions into a plurality of third clusters so that, for example, a plurality of partial data groups 131 in which the classified first clusters corresponding to the respective partial data groups 131 are the same and the classified second clusters of the respective partial data groups 131 are the same are classified into the same cluster.

[0109] Then, for each of the plurality of third clusters, the information processing apparatus 1 identifies a first condition that can classify the partial data group 131 classified into each cluster and the partial data group 131 classified into other clusters, and outputs the identified first condition together with the classification results of the plurality of third clusters.

[0110] That is, the information processing apparatus 1 in the present embodiment classifies the partial data groups 131 extracted for each of the plurality of conditions into a plurality of third clusters such that, among the partial data groups 131 that can be determined to be essentially close, the partial data groups 131 for which the corresponding causal graph 132 can also be determined to be essentially close are included in the same cluster. Then, the information processing apparatus 1 identifies, for example, the first condition, which is a condition under which an important causal relationship leading to problem solving appears, for each of the plurality of third clusters.

[0111] Thereby, the information processing apparatus 1 in the present embodiment can aggregate combinations that can be determined to be essentially close even when there are a large number of combinations of the partial data group 131 and the causal graph 132, and can easily identify important causal relationships leading to problem solving.

[0112] Further, the information processing apparatus 1 in the present embodiment can output, for example, by combining the classification result corresponding to each cluster and the first condition, in a form of associating the important causal relationship identified by the cause identification process with the first condition under which the important causal relationship appears. Therefore, for example, by viewing each piece of information output by the information processing apparatus 1, the operator can easily grasp the correspondence between the important causal relationship identified by the cause identification process and the first condition under which the important causal relationship appears.

[0113] Furthermore, the information processing apparatus 1 in the present embodiment can output information indicating whether a plurality of causal graphs 132 that can be determined to be similar to each other are essentially close, or information indicating whether a plurality of data groups 131 that can be determined to be similar to each other are essentially close, by classifying a plurality of data groups 131 corresponding to causal graphs 132 that can be determined to be similar to each other among a plurality of data groups 131 that can be determined to be similar to each other into the same cluster. Therefore, an operator can easily make a judgment as to whether a plurality of causal graphs 132 that are similar to each other are essentially close, or a judgment as to whether a plurality of data groups 131 that are similar to each other are essentially close, for example, by viewing each piece of information output by the information processing apparatus 1.

Explanation of Signs

[0114] 1: Information processing apparatus 5: Operation terminal 10: Information processing system NW: Network

Claims

1. Referring to a storage unit that stores a plurality of data consisting of combinations of a plurality of feature amounts, for each of the plurality of conditions, a data group in which the combination satisfies each condition is extracted, For each of the plurality of conditions, the relationship between the plurality of feature amounts included in the data group corresponding to each condition is specified, Based on the first similarity between the relationships for each of the plurality of conditions, the relationships for each of the plurality of conditions are classified into a plurality of first clusters, Based on the second similarity between the data groups for each of the plurality of conditions, the data groups for each of the plurality of conditions are classified into a plurality of second clusters, The data groups for each of the plurality of conditions are classified into a plurality of third clusters such that a plurality of the data groups in which the relationships corresponding to each data group are classified into the same first cluster and the second clusters into which each data group is classified are the same are classified into the same cluster, For each of the plurality of third clusters, a first condition that can classify the data group classified into each cluster and the data group classified into another cluster is specified, The specified first condition is output together with the classification results of the plurality of third clusters, An information processing program characterized by causing a computer to execute the process.

2. In Claim 1, In the process of specifying the first condition, For each cluster among the plurality of third clusters in which the number of data groups classified into each cluster is equal to or more than a predetermined number, the first condition is specified and output, An information processing program characterized by this.

3. In Claim 1, In the process of specifying the first condition, For each of the plurality of third clusters, a common data group common to one or more of the data groups classified into each cluster is specified, For each of the plurality of third clusters, a condition that can classify the common data group classified into each cluster and the common data group classified into another cluster is specified as the first condition, An information processing program characterized by this.

4. In Claim 3, In the process of specifying the first condition, For each of the plurality of third clusters, a learning model is generated by performing machine learning with the common data group classified into each cluster as a positive example and the common data group classified into another cluster as a negative example, For each of the plurality of third clusters, the condition indicated by the learning model corresponding to each cluster is specified as the first condition, An information processing program characterized by this.

5. In claim 1, in the process of specifying the first condition, for each of the plurality of third clusters, a common relationship common to one or more of the relationships corresponding to one or more of the data groups classified into each cluster is specified, in the process of outputting, information indicating the common relationship is output as the classification result of the plurality of third clusters, An information processing program characterized by the above.

6. In claim 5, in the process of outputting, information indicating a relationship not included in the relationship among the plurality of feature amounts included in each of the data stored in the storage unit is output from among the common relationships, An information processing program characterized by the above.

7. A data extraction unit that refers to a storage unit storing a plurality of data consisting of combinations of a plurality of feature amounts, and extracts a data group in which the combination satisfies each condition for each of the plurality of conditions; a relationship specifying unit that specifies a relationship among the plurality of feature amounts included in the data group corresponding to each condition for each of the plurality of conditions; Based on a first similarity among the relationships for each of the plurality of conditions, the relationships for each of the plurality of conditions are classified into a plurality of first clusters, and based on a second similarity among the data groups for each of the plurality of conditions, the data groups for each of the plurality of conditions are classified into a plurality of second clusters, and a plurality of the data groups in which the first clusters in which the relationships corresponding to each data group are classified are the same and the second clusters in which each data group is classified are the same are classified into the same cluster, a clustering unit that classifies the data groups for each of the plurality of conditions into a plurality of third clusters; a condition specifying unit that specifies a first condition that can classify the data group classified into each cluster and the data group classified into another cluster for each of the plurality of third clusters; An information processing apparatus, comprising: a condition output unit that outputs the specified first condition together with the classification result of the plurality of third clusters. An information processing apparatus characterized by the above.

8. Referring to a storage unit storing a plurality of data consisting of combinations of a plurality of feature amounts, for each of the plurality of conditions, a data group in which the combination satisfies each condition is extracted, for each of the plurality of conditions, a relationship among the plurality of feature amounts included in the data group corresponding to each condition is specified, Based on the first similarity among the relationships for each of the plurality of conditions, classify the relationships for each of the plurality of conditions into a plurality of first clusters. Based on the second similarity among the data groups for each of the plurality of conditions, classify the data groups for each of the plurality of conditions into a plurality of second clusters. Classify the data groups for each of the plurality of conditions into a plurality of third clusters such that a plurality of the data groups for which the first clusters in which the corresponding relationships are classified are the same and the second clusters in which the data groups are classified are the same are classified into the same cluster. For each of the plurality of third clusters, specify a first condition capable of classifying the data groups classified into each cluster and the data groups classified into other clusters. Output the specified first condition together with the classification results of the plurality of third clusters. An information processing method, characterized in that a computer executes the processing.

Citation Information

Patent Citations

  • Interaction detector, medium with program for interaction detection recorded therein, and interaction detection method

    JP2008003819A

  • Causal relation analyzing device, causal relation analyzing method and program

    JP2008203964A

  • Relation search system, information processing device, method, and program

    WO2018168580A1

  • Information processing device, information processing method, and information processing program

    WO2021014823A1