Causal search system and causal search program
The causal discovery system improves reliability and reduces user workload by creating sample datasets and enabling automated and manual selection modes to determine causal graphs, addressing issues in conventional systems with large datasets.
Patent Information
- Application Number
- JP2024003772
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-15
- Publication Date
- 2025-07-28
AI Technical Summary
Conventional causal discovery programs face challenges in improving reliability and reducing user burden when dealing with large datasets, as synthesizing multiple causal graphs often results in inappropriate edge formations like loops, increasing the workload for directed edge selection.
A causal discovery system that creates multiple sample datasets from an original dataset, performs causal discovery on each, and allows for graph and directed edge selection modes to determine a reliable causal graph, reducing user workload through automated and manual selection processes.
Enhances the reliability of causal graphs by automating the selection process, minimizing user effort in selecting edges and graphs, even with large datasets, while avoiding inappropriate formations like loops.
Smart Images

Figure 2025110058000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a causal discovery system and a causal discovery program for discovering causal relationships between variables based on a dataset including a plurality of variables.
Background Art
[0002] Conventionally, statistical causal discovery programs such as LiNGAM for analyzing causal relationships between variables based on a dataset including a plurality of variables are known. The causal discovery program creates a causal graph in which the causal relationships between variables are represented by directed edges by performing causal discovery on the input dataset. Conventional causal discovery programs are described in, for example, Patent Document 1.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In this type of causal discovery program, in order to improve the reliability of causal discovery, for example, it is conceivable to create a plurality of datasets by restoration extraction from the original dataset and perform causal discovery on each of the created plurality of datasets.
[0005] However, simply synthesizing a plurality of causal graphs obtained by causal discovery on a plurality of datasets results in an inappropriate state such as some directed edges forming loops. Therefore, it is necessary to select necessary directed edges from among the directed edges included in the plurality of causal graphs. However, when the number of variables is large, the burden on the user for the directed edge selection work increases.
[0006] Therefore, an object of the present invention is to provide a causal discovery system and a causal discovery program that can obtain a highly reliable causal graph by causal discovery using restoration extraction and can reduce the burden on the user for determining the causal graph.
Means for Solving the Problems
[0007] To solve the above problems, a first invention of the present application is a causal discovery system for discovering causal relationships between variables based on an original dataset including a plurality of variables, wherein a computer executes: (a) a process of creating a plurality of sample datasets by restoration extraction from the original dataset; (b) a process of obtaining a plurality of causal graphs indicating the causal relationships between the variables by directed edges by performing causal discovery for each of the plurality of sample datasets; and (c) a process of determining one causal graph based on the plurality of causal graphs. In the process (c), a graph selection mode for selecting a causal graph for use in the one causal graph from the plurality of causal graphs obtained in the process (b), and a directed edge selection mode for selecting directed edges to be used in the one causal graph from the plurality of directed edges obtained in the process (b) are switchable.
[0008] A second invention of the present application is the causal discovery system according to the first invention, wherein in the graph selection mode, the number of appearances of the same causal graph among the plurality of causal graphs obtained in the process (b) is displayed.
[0009] A third invention of the present application is the causal discovery system according to the first or second invention, wherein in the graph selection mode, the degree of fitness for the original dataset is displayed for each of the plurality of causal graphs obtained in the process (b).
[0010] A fourth invention of the present application is the causal discovery system according to any one of the first to third inventions, wherein in the process (b), the appearance probability of each directed edge in the plurality of causal graphs is calculated, and in the directed edge selection mode, a predetermined number of directed edges with a high appearance probability are selected.
[0011] The fifth invention of the present application is a causal exploration system for any one of the first to third inventions, wherein in the process (b), the appearance probability of each directed edge in a plurality of causal graphs is calculated, and in the directed edge selection mode, for any two variables, the appearance probability of the first directed edge from one variable to the other variable, the appearance probability of the second directed edge from the other variable to the one variable, and the appearance probability of the state without a directed edge are respectively displayed, and any one of the first directed edge, the second directed edge, and the state without a directed edge is selected.
[0012] The sixth invention of the present application is a causal exploration program, which, when installed in the computer of the causal exploration system of any one of the first to fifth inventions of the present application, causes the computer to execute the processes (a) to (c).
Advantages of the Invention
[0013] According to the first to sixth inventions of the present application, by creating a plurality of sample data sets from one original data set, a plurality of causal graphs are obtained. Then, one causal graph is determined from the plurality of causal graphs. Thereby, the reliability of the causal graph can be improved. Further, when determining one causal graph, the graph selection mode and the directed edge selection mode can be switched. Thereby, the burden on the user for the determination work of the causal graph can be reduced.
[0014] In particular, according to the second invention of the present application, one causal graph can be selected in consideration of the number of appearances of the same causal graph. Thereby, the work burden on the user for selecting the causal graph can be reduced.
[0015] In particular, according to the third invention of the present application, one causal graph can be selected in consideration of the fitness of the causal graph with respect to the original data set. Thereby, the work burden on the user for selecting the causal graph can be reduced.
[0016] In particular, according to the fourth invention of the present application, a predetermined number of directed edges can be automatically selected based on the appearance probability of the directed edges. Thereby, the work load of the user who selects the directed edge can be reduced.
[0017] In particular, according to the fifth invention of the present application, any one of a first directed edge, a second directed edge, and a state without a directed edge can be selected between two variables in consideration of the appearance probability. Thereby, the work load of the user who selects the directed edge can be reduced.
Brief Description of the Drawings
[0018]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Modes for Carrying Out the Invention
[0019] Hereinafter, embodiments of the present invention will be described with reference to the drawings.
[0020] <1. Configuration of the Causal Exploration System> FIG. 1 is a diagram showing the configuration of a causal relationship exploration system 1 according to an embodiment of the present invention. This causal relationship exploration system 1 is a system that explores causal relationships between variables based on a dataset including a plurality of variables and outputs a visualized causal graph G. As shown in FIG. 1, the causal relationship exploration system 1 includes a computer 10, a display unit 20, and an input unit 30.
[0021] The computer 10 is an information processing device for executing causal relationship exploration processing. As shown in FIG. 1, the computer 10 includes a processor 11 such as a CPU, a memory 12 such as a RAM, and a storage unit 13 such as a hard disk drive.
[0022] The storage unit 13 stores a causal relationship exploration program 131. The causal relationship exploration program 131 is application software for causing the computer 10 to execute the processing of steps S1 to S6 described later. The causal relationship exploration program 131 is read from a storage medium M such as a CD or a DVD and installed in the computer 10. However, the computer 10 may be downloaded to the computer 10 via a network N such as the Internet.
[0023] The display unit 20 is a device that displays various information output from the computer 10. For example, a liquid crystal display is used for the display unit 20. The input unit 30 is a device that inputs various information to the computer 10. For example, a keyboard or a mouse is used for the input unit 30. Note that the display unit 20 and the input unit 30 may be realized by a single device such as a touch panel display. The display unit 20 and the input unit 30 are electrically connected to the computer 10.
[0024] FIG. 2 is a block diagram conceptually showing the functions of the computer 10 described above. As shown in FIG. 2, the computer 10 includes a preprocessing unit 41, a resampling unit 42, a causal exploration unit 43, a causal graph determination unit 44, a causal inference unit 45, and an output unit 46. Each function of the preprocessing unit 41, the resampling unit 42, the causal exploration unit 43, the causal graph determination unit 44, the causal inference unit 45, and the output unit 46 is realized by the processor 11 of the computer 10 operating according to the causal exploration program 131.
[0025] FIG. 3 is a flowchart showing the flow of the causal exploration process by the causal exploration system 1. Hereinafter, the functions of each unit shown in FIG. 2 will be described along the flow of the process in FIG. 3.
[0026] When performing the causal exploration process, first, a data set to be the object of causal exploration is input to the computer 10 (step S1). In this step S1, the data set input to the computer 10 is hereinafter referred to as the "original data set D0". The original data set D0 is, for example, various measurement data in a manufacturing apparatus. However, the type of data in the original data set D0 is not limited. The original data set D0 is stored in the storage unit 13 of the computer 10.
[0027] FIG. 4 is a diagram showing an example of the original data set D0. As shown in FIG. 4, the original data set D0 input to the causal exploration system 1 is tabular numerical data. The original data set D0 includes a plurality of variables X1, X2, X3, ···. More specifically, the original data set D0 has a plurality of sets of numerical groups d1, d2, d3, ···. Each numerical group d1, d2, d3, ··· is composed of a plurality of variables X1, X2, X3, ···.
[0028] The output unit 46 displays the input original data set D0 on the display unit 20. At this time, the output unit 46 may display on the display unit 20, for each of the variables X1, X2, X3, ···, a histogram showing the numerical distribution, various statistical quantities such as the average value, and the presence or absence of missing values. Thereby, the user of the causal exploration system 1 can grasp the characteristics of the numerical distribution of each variable X1, X2, X3, ···. Also, the computer 10 may be able to display a scatter plot between two variables specified by the user on the display unit 20.
[0029] Next, the preprocessing unit 41 of the computer 10 performs preprocessing on the original data set D0 (step S2). For example, the preprocessing unit 41 interpolates the missing values in the original data set D0. Also, the preprocessing unit 41 may perform deletion of variables not used for causal exploration, etc.
[0030] Subsequently, the resampling unit 42 of the computer 10 creates a plurality of sample data sets D1, D2, D3, ··· from the original data set D0 (step S3). FIG. 5 is a diagram conceptually showing the state of the process of step S3. In FIG. 5, a case is shown where the original data set D0 contains five numerical groups d1 to d5. As shown in FIG. 5, the resampling unit 42 creates the sample data sets D1, D2, D3, ··· by randomly extracting the numerical groups d1, d2, d3, ··· a predetermined number of times from the original data set D0 as the population by the sampling with replacement method.
[0031] In the example of FIG. 5, the number of the numerical groups d1, d2, d3, ··· included in the original data set D0 and the number of the numerical groups included in each of the sample data sets D1, D2, D3, ··· are the same number (both five). However, the number of the numerical groups included in each of the sample data sets D1, D2, D3, ··· may be a number different from the number of the numerical groups d1, d2, d3, ··· included in the original data set D0.
[0032] In the resampling extraction method, the numerical groups once extracted from the original dataset D0 are targeted for extraction again without being deleted from the original dataset D0. Therefore, as shown in Fig. 5, there may be cases where the same numerical group is included two or more times in one sample dataset. Also, there may be cases where a part of the numerical groups d1, d2, d3, ··· included in the original dataset D0 is not included in one sample dataset.
[0033] In this way, the resampling unit 42 creates a plurality of sample datasets D1, D2, D3, ··· from one original dataset D0. Thereby, the number of datasets targeted for causal exploration can be increased.
[0034] Subsequently, the causal exploration unit 43 of the computer 10 performs causal exploration for each of the plurality of sample datasets D1, D2, D3, ··· (step S4). The causal exploration unit 43 analyzes the causal relationship between the variables X1, X2, X3, ··· for each of the sample datasets D1, D2, D3, ··· according to the statistical causal exploration algorithm. Then, the causal exploration unit 43 creates a causal graph G that visualizes the causal relationship between the variables X1, X2, X3, ··· for each of the sample datasets D1, D2, D3, ···. The output unit 46 displays the created causal graph G on the display unit 20.
[0035] As the statistical causal exploration algorithm for creating the causal graph G, for example, DirectLiNGAM, ICA-LiNGAM, BottomUpParceLiNGAM, RCD, CAM-UV, etc. can be used. In addition, when the presence or absence of the causal relationship between some variables is known, the user may input the relationship into the computer 10 as prior knowledge. In that case, the causal exploration unit 43 performs causal exploration while observing the constraints of the input prior knowledge.
[0036] FIG. 6 is a diagram showing an example of a causal graph G. As shown in FIG. 6, the causal graph G is an image showing the causal relationships between variables X1, X2, X3, ··· by directed edges A (arrows). In the causal graph G of FIG. 6, the causal relationships between six variables X1 to X6 are shown by the directed edges A. The base end side of the directed edge A is the variable (factor) that exerts an influence. The tip side of the directed edge A is the variable (result) that is affected.
[0037] The causal exploration unit 43 creates a causal graph G for each of a plurality of sample data sets D1, D2, D3, ···. The causal graph G created by causal exploration for one sample data set has at most one valid edge between two variables, and no loop is formed by a plurality of valid edges A. However, when all the valid edges A created by causal exploration for a plurality of sample data sets D1, D2, D3, ··· are displayed, as shown in FIG. 6, there may be two directed edges A between some two variables, or a loop may be formed by some directed edges A.
[0038] The causal exploration unit 43 calculates the appearance probability P of each directed edge A in the obtained plurality of causal graphs G. The appearance probability P is the ratio of the number of appearances of the directed edge A to the number of created causal graphs G (the number of sample data sets). Then, as shown in FIG. 6, the causal exploration unit 43 attaches the appearance probability P to each directed edge A in the causal graph G. The user of the causal exploration system 1 can infer the reliability of each directed edge A by referring to these appearance probabilities P.
[0039] Subsequently, the causal graph determination unit 44 determines one causal graph G based on the plurality of causal graphs G created in step S4 (step S5). The causal exploration system 1 of the present embodiment has three processing modes as the processing mode of step S5: "graph selection mode", "automatic directed edge selection mode", and "manual directed edge selection mode". The user of the causal exploration system 1 can switch between the "graph selection mode", "automatic directed edge selection mode", and "manual directed edge selection mode" by operating the input unit 30.
[0040] FIG. 7 is a diagram showing an example of a screen displayed on the display unit 20 when the "graph selection mode" is selected. The "graph selection mode" is a mode for selecting a causal graph G to be used as the one causal graph G from a plurality of causal graphs G obtained in step S4. The causal graph determination unit 44 counts the number of appearances of the same causal graph G among the plurality of causal graphs G. The same causal graph G is a causal graph G in which the number, position, and direction of the directed edges A all match.
[0041] As shown in FIG. 7, the causal graph determination unit 44 displays the number of appearances of the causal graph G in the form of a ranking table R in descending order. In the example of FIG. 7, the number of appearances of the most frequent causal graph G is 54 times. The causal graph G with a large number of appearances is presumed to have a high reliability. The user of the causal search system 1 selects one causal graph G in consideration of the number of appearances displayed in the ranking table R. This can reduce the workload of the user who selects the causal graph G.
[0042] In the example of FIG. 7, when the user selects an arbitrary row of the ranking table R, the causal graph G of that rank is displayed on the right side. In this way, the user can select the causal graph G that is considered to be the most appropriate while visually checking the causal graph G.
[0043] Also, in the example of FIG. 7, the fitness is displayed in the ranking table R together with the number of appearances. The fitness is an index indicating how well the causal graph G fits the original dataset D0. The fitness can be calculated using, for example, RMSEA (Root Mean Square Error of Approximation), likelihood, AIC (Akaike's Information Criterion), DIC (Deviance Information Criterion), etc. The user of the causal search system 1 selects one causal graph G in consideration of the fitness displayed in the ranking table R. This can further reduce the workload of the user who selects the causal graph G.
[0044] Note that the ranking table R may display the fitness of the causal graph G in descending order. Also, the ranking table R in the order of the number of occurrences and the ranking table R in the order of fitness may be switchable according to the user's selection.
[0045] FIG. 8 is a diagram showing an example of a screen displayed on the display unit 20 when the "automatic directed edge selection mode" is selected. The "automatic directed edge selection mode" is a mode in which, from a plurality of directed edges A obtained in step S4, the directed edge A to be used for one causal graph G is automatically selected based on conditions specified by the user.
[0046] As described above, in step S4, the occurrence probability P is calculated for each of the many directed edges A included in the plurality of causal graphs G. The causal graph determination unit 44 first displays the occurrence probability P for all the directed edges A between the plurality of variables X1, X2, X3,... as shown in the upper diagram of FIG. 8.
[0047] Next, the user operates the input unit 30 to input the number of directed edges A to be selected. Then, as shown in the lower diagram of FIG. 8, the causal graph determination unit 44 displays only the specified number of directed edges A with the highest occurrence probability P and makes the other directed edges A non-displayed. As a result, a predetermined number of directed edges A with a high occurrence probability P are automatically selected by the computer 10. Therefore, the work burden on the user for selecting the directed edge A can be reduced.
[0048] Note that the user of the causal exploration system 1 may specify a threshold value of the occurrence probability P instead of the number of directed edges A to be selected. In that case, the causal graph determination unit 44 displays only the directed edges A having an occurrence probability P equal to or higher than the specified threshold value and makes the other directed edges A non-displayed. For example, if 0.5 is specified as the threshold value of the occurrence probability P with respect to the upper diagram of FIG. 8, the causal graph G in the lower part of FIG. 8 can be obtained.
[0049] As shown in the upper part of FIG. 8, when all the directed edges A are displayed, there are two directed edges A between two variables, or some directed edges A form a loop. However, the user of the causal exploration system 1 can create a causal graph G without inappropriate situations such as loops as shown in the lower part of FIG. 8 by adjusting the number of directed edges A to be selected.
[0050] FIG. 9 is a diagram showing an example of a screen displayed on the display unit 20 when the "manual directed edge selection mode" is selected. The "manual directed edge selection mode" is a mode in which the user selects the directed edge A to be used in one causal graph G from among the plurality of directed edges A obtained in step S4 while confirming.
[0051] In the "manual directed edge selection mode", as shown in FIG. 9, for two variables X1 and X2, the appearance probability P1 of the directed edge (first directed edge) A from one variable X1 to the other variable X2, the appearance probability P2 of the directed edge (second directed edge) A from the other variable X2 to one variable X1, and the appearance probability P3 of the state where there is no directed edge A between the two variables X1 and X2 are respectively displayed.
[0052] The user of the causal exploration system 1 refers to these appearance probabilities P1, P2, and P3 and selects whether the directed edge A between the two variables X1 and X2 is the first directed edge, the second directed edge, or the state where there is no directed edge A. The causal graph determination unit 44 performs the same display for all combinations of two variables included in the original data set D0 and requests the user to select a directed edge A. As a result, the directed edge A between all variables can be determined in consideration of the appearance probability P. This can reduce the workload of the user who selects the directed edge A.
[0053] As described above, in this causal discovery system 1, a plurality of sample datasets D1, D2, D3, ··· are created from one original dataset D0. Then, one causal graph G is determined from a plurality of causal graphs G created based on the sample datasets D1, D2, D3, ···. Thereby, even when the amount of data in the original dataset D0 is small, a highly reliable causal graph G can be obtained.
[0054] Also, in this causal discovery system 1, when determining one causal graph G, the "graph selection mode", "automatic directed edge selection mode", and "manual directed edge selection mode" can be switched. Thereby, while reducing the user's workload, a causal graph G without inappropriate situations such as loops can be determined. The computer 10 stores the determined causal graph G in the storage unit 13 and displays it on the display unit 20.
[0055] When one causal graph G is determined, finally, causal inference is performed based on the causal graph G (step S6). In this step S6, the causal inference unit 45 performs an intervention process on the causal graph G. In the intervention process, for example, a certain variable among a plurality of variables X1, X2, X3, ··· is changed. Then, the causal inference unit 45 changes the values of other variables based on the causal relationship shown by the causal graph G. And the change of the variable is displayed on the display unit 20 via the output unit 46. Thereby, the user of the causal discovery system 1 can observe the influence of the change of one variable on other variables.
[0056] <2. Modification Example> As described above, an embodiment of the present invention has been described, but the present invention is not limited to the above embodiment.
[0057] The causal exploration system 1 of the above embodiment had three processing modes as the processing mode in step S5: "graph selection mode", "automatic directed edge selection mode", and "manual directed edge selection mode". That is, the causal exploration system 1 of the above embodiment had two types of "directed edge selection modes". However, the "directed edge selection mode" that the causal exploration system 1 has may be only any one of the "automatic directed edge selection mode" and the "manual directed edge selection mode".
[0058] Also, in the "graph selection mode", the causal exploration system 1 of the above embodiment displayed the number of occurrences or fitness of the causal graph G in the form of the ranking table R. However, the causal exploration system 1 may display the number of occurrences or fitness of the causal graph G in a form other than the ranking table R.
[0059] Moreover, each element that appeared in the above embodiment and modification examples may be appropriately combined within a range where no contradiction occurs. For example, in the process of step S5, a plurality of "graph selection mode", "automatic directed edge selection mode", and "manual directed edge selection mode" may be used in combination. Specifically, for the causal graph G created by the "graph selection mode" or the "automatic valid edge selection mode", modification may be performed by adding or deleting the valid edge A using the function of the "manual directed edge selection mode" for some variables, and the finally used causal graph G may be determined.
Explanation of Signs
[0060] 1: Causal exploration system 10: Computer 20: Display unit 30: Input unit 41: Preprocessing unit 42: Resampling unit 43: Causal exploration unit 44: Causal graph determination unit 45: Causal inference unit 46: Output unit 131: Causal exploration program A: Directed edge D0: Original dataset D1: Sample dataset D2: Sample dataset D3: Sample dataset G: Causal graph P: Appearance probability R: Ranking table X1: Variable X2: Variable X3: Variable X4: Variable X5: Variable X6: Variable d1: Numerical group d2: Numerical group d3: Numerical group d4: Numerical group d5: Numerical group
Claims
1. A causal relationship exploration system for exploring causal relationships between variables based on an original dataset including a plurality of variables, wherein a computer performs: (a) a process of creating a plurality of sample datasets by restoration extraction from the original dataset; (b) a process of obtaining a plurality of causal graphs indicating the causal relationships between the variables by directed edges by performing causal relationship exploration for each of the plurality of sample datasets; (c) a process of determining one causal graph based on the plurality of causal graphs; and in the process (c), a graph selection mode of selecting a causal graph for use in the one causal graph from the plurality of causal graphs obtained by the process (b); a directed edge selection mode of selecting directed edges to be used in the one causal graph from the plurality of directed edges obtained by the process (b); are switchable, the causal relationship exploration system.
2. The causal relationship exploration system according to claim 1, wherein in the graph selection mode, the number of occurrences of the same causal graph among the plurality of causal graphs obtained in the process (b) is displayed, the causal relationship exploration system.
3. The causal relationship exploration system according to claim 1, wherein in the graph selection mode, the fitness for the original dataset is displayed for each of the plurality of causal graphs obtained in the process (b), the causal relationship exploration system.
4. The causal relationship exploration system according to claim 1, wherein in the process (b), the appearance probability of each directed edge in the plurality of causal graphs is calculated, and in the directed edge selection mode, a predetermined number of directed edges with high appearance probabilities are selected, the causal relationship exploration system.
5. The causal relationship exploration system according to claim 1, wherein in the process (b), the appearance probability of each directed edge in the plurality of causal graphs is calculated, and in the directed edge selection mode, for any two variables, the appearance probability of a first directed edge from one variable to the other variable; the appearance probability of a second directed edge from the other variable to the one variable; the appearance probability of a state without a directed edge; are each displayed, and one of the first directed edge, the second directed edge, and the state without a directed edge is selected, the causal relationship exploration system.
6. A causal exploration program that, when installed on the computer of the causal exploration system according to any one of claims 1 to 5, causes the computer to execute the processes (a) to (c).
Citation Information
Patent Citations
Printer and management method
JP2023062325A