Program, information processing method, and information processing device.
By generating multiple second datasets, the information processing of causal relationships and causal effects is optimized, solving the problem of differences in causal effects between different datasets, and realizing the estimation of causal relationships and individual causal effects across datasets.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-04-08
AI Technical Summary
Existing technologies assume that causal effects are common across all datasets and cannot handle differences in causal effects between one dataset and other datasets.
By generating multiple second datasets, information processing devices and programs are used to optimize information on causal relationships and causal effects, in order to minimize the loss function values of multiple datasets and estimate integrated causal relationships and individual causal effects.
It enables the estimation of integrated causal relationships and individual causal effects across multiple datasets, applicable to datasets with different sets of variables and attributes.
Smart Images

Figure 2026060185000001_ABST
Abstract
Description
[Technical Field]
[0001] This invention relates to a program, an information processing method, and an information processing device. [Background technology]
[0002] Information processing techniques are sometimes used to estimate causal relationships between variables in multiple datasets. For example, a causal relationship in which one variable influences another can be represented as a causal graph of a linear causal model. A causal graph may include causal effects between variables. A causal effect is the degree to which one variable influences another.
[0003] Here, assuming that causal effects are common across all datasets, a method has been proposed to optimize causal effects so that pseudo-data closely resembling the input data can be generated. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] B. Huang, 3 others, “Causal Discovery from Multiple Data Sets with Non-Identical Variables”, [online], 2020, AAAI'20, [Retrieved September 2, 2020], Internet<URL:https: / / cdn.aaai.org / ojs / 6575 / 6575-13-9800-1-10-20200519.pdf> [Overview of the project] [Problems that the invention aims to solve]
[0005] The proposed method assumes that causal effects are common across all datasets. Therefore, it has the problem of being unable to handle cases where the causal effects in one dataset differ from those in other datasets.
[0006] In one aspect, the present invention aims to enable the estimation of integrated causal relationships between variables across multiple datasets and individual causal effects in each dataset. [Means for solving the problem]
[0007] In one embodiment, a program is provided that is executed by a computer. This program causes the computer to perform the following actions: The computer obtains a plurality of first data sets, each containing a plurality of variable values corresponding to a plurality of variables, which are used to estimate causal relationships and causal effects between variables based on the plurality of variable values. The computer sets up first information about the integrated causal relationships and causal effects in the plurality of first data sets as a whole, and second information about the weighting of the causal effects contained in the first information for each of the plurality of first data sets. Based on the first and second information, the computer generates a plurality of second data sets corresponding to the plurality of first data sets. The computer optimizes the first and second information to minimize the value of the loss function for the plurality of first data sets and the plurality of second data sets.
[0008] In one embodiment, a method for information processing performed by a computer is provided. In another embodiment, an information processing device having a storage unit and a processing unit is provided. [Effects of the Invention]
[0009] In one respect, it is possible to estimate the integrated causal relationships between variables across multiple datasets, as well as the individual causal effects within each dataset. [Brief explanation of the drawing]
[0010] [Figure 1] This is a diagram illustrating the information processing device of the first embodiment. [Figure 2] This figure shows an example of the hardware of the information processing device according to the second embodiment. [Figure 3] This is an example of a causal graph. [Figure 4] This figure shows examples of multiple datasets. [Figure 5] This figure shows an example of the functions of an information processing device. [Figure 6] This figure shows an example of the optimization process. [Figure 7] This figure shows an example of data stored in the data storage unit. [Figure 8] This figure shows an example of pseudo-data generation in an optimization loop. [Figure 9] This figure shows an example of loss calculation in an optimization loop. [Figure 10] This figure shows an example of gradient propagation in an optimization loop. [Figure 11] This figure shows an example of the output results. [Figure 12] This is a flowchart showing an example of processing performed by an information processing device. [Figure 13] This is a diagram showing a comparative example. [Figure 14] This figure shows an example of a causal graph that cannot be distinguished in the comparative example. [Modes for carrying out the invention]
[0011] This embodiment will be described below with reference to the drawings. [First Embodiment] A first embodiment will be described.
[0012] Figure 1 is a diagram illustrating the information processing device of the first embodiment. The information processing device 10 estimates causal relationships and causal effects between variables based on multiple datasets, each having multiple variable values corresponding to multiple variables. Specifically, the information processing device 10 estimates a causal graph of a linear causal model in the union of a set of variables, i.e., an integrated set of variables, from the multiple datasets. The multiple variables included in each of the multiple datasets may all be the same, or they may be different in at least two of the datasets. In the latter case, some of the multiple variables included in any first dataset will overlap with some of the multiple variables included in at least one other dataset. The information processing device 10 has a storage unit 11 and a processing unit 12.
[0013] The memory unit 11 may be a volatile semiconductor memory such as RAM (Random Access Memory), or a non-volatile storage such as an HDD (Hard Disk Drive) or flash memory. The processing unit 12 is a processor such as a CPU (Central Processing Unit), GPU (Graphics Processing Unit), or DSP (Digital Signal Processor). However, the processing unit 12 may also include application-specific electronic circuits such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array). The processor executes programs stored in memory such as RAM (which may also be the memory unit 11). A collection of multiple processors is sometimes called a "multiprocessor" or simply a "processor."
[0014] The causal relationships between variables are represented by a causal graph of a linear causal model. In this case, dataset X is represented as tabular data, with rows corresponding to samples and columns corresponding to variables. All variables are of numerical type. The linear causal model is expressed by equation (1).
[0015]
number
[0016] X is represented by formula (2). n is the number of samples. d is the number of variables x i 's number. The white "R" represents the entire set of real numbers. X is the transposed row-column representation of the table data of n samples with d items.
[0017]
Number
[0018] B is represented by formula (3). B is the adjacency matrix of the causal graph.
[0019]
Number
[0020] b ij represents the direct causal effect of x j →x i . That is, when b ij ≠0, b ij represents the causal relationship that x j affects x i , and the degree of that influence is b ij . When b ij =0, b ij indicates that x j does not affect x i , that is, there is no causal relationship from x j to x i .
[0021] E is represented by formula (4). E is the noise matrix.
[0022]
Number
[0023] In one example, e i is a uniform random number where -1 < e i < 1. The memory unit 11 stores multiple first data sets used for estimating causal relationships and causal effects. The total number of first data sets is m. The multiple first data sets are pre-input into the information processing device 10 and stored in the memory unit 11. The number of samples in each of the multiple first data sets may be different, and they do not have to be the same sample. One first data set corresponds to one input dataset containing multiple variable values. Also, the set of variables included in each first data set may be different.
[0024] Here, for example, in the analysis of hospital diagnostic records, examples of variables include the amount of medication administered to patients, their blood glucose levels, and their urination frequency. In this case, the patients represent the sample. Patients may differ in each hospital. In this case, one first data set contains measured values such as medication dosage, blood glucose levels, and urination frequency as variable values, corresponding to each of multiple patients in a given hospital. However, the content of medical procedures and patient examinations may differ between hospitals. For example, one hospital may measure a patient's blood pressure but not their body temperature, while another hospital may measure body temperature but not their blood pressure. Another example is in experimental records of chemical reactions, where one experiment may measure reaction energy but not reaction rate, while another experiment may measure reaction rate but not reaction energy. Furthermore, the number of experiments and experimental conditions may differ from experiment to experiment.
[0025] Thus, the set of variables in one first data set may differ from the set of variables in another first data set. However, as mentioned above, a portion of the set of variables in one first data set will overlap with a portion of the set of variables in at least one other first data set.
[0026] In this case, it is conceivable to estimate a linear causal model of each variable in the union of the variable sets of each first data set. This linear causal model can take the form of, for example, equation (5).
[0027]
number
[0028] A is called the composite causal effect matrix. A is represented by equation (6).
[0029]
number
[0030] a ij is, x j →x i This represents the overall causal effect. d is the total number of variables belonging to the union of the variable sets of each first data point. I is the identity matrix. The adjacency matrix B, which represents the direct causal effect, can be reconstructed by equation (7).
[0031]
number
[0032] In this case, the kth first data X (k) e in i Gaussian mixture distribution G i (k) This can be modeled as follows: k is an index that identifies the first data (dataset). k = 1, 2, ..., m. First data X (k) In contrast, pseudo-data X' (k) =AE (k) These are randomly generated.
[0033] And then, X' (k) The distribution of X (k) Each G i (k) And A can be estimated. The similarity of the distributions is X (k) and X' (k) It is measured by a predetermined loss function. The loss function can be said to be a function that shows the similarity or difference between the data. The closer the value of the loss function for two sets of data is to 0, the more similar the two sets of data are, or the smaller the difference between the two sets of data is. Examples of loss functions include MMD (Maximum Mean Discrepancy) and optimal transport distance. i(k) And in the estimation of A, each G is calculated using gradient descent to minimize the loss function. i (k) And A is optimized. As a gradient descent method, for example, stochastic gradient descent or batch gradient descent (steepest descent) can be used.
[0034] However, equation (5) assumes the same linear causal model for all datasets. This assumption has the problem of not being able to consider the possibility that causal effects may differ for each dataset. Specifically, in the example of multiple hospitals mentioned above, the effect of a drug may differ in hospitals where the patient trends regarding the administration of a certain drug differ. For example, the dosage and effect of a drug may differ between young people and the elderly, so the causal effect may differ between hospitals with many young patients and hospitals with many elderly patients. Also, for example, in the analysis of experimental results from multiple universities or research institutes, the results of experiments using water may differ if the laboratory is located in a different area due to differences in water quality in each area. In such cases, the linear causal model in equation (5) cannot obtain the causal effect specific to each dataset.
[0035] Therefore, the processing unit 12 estimates a linear causal model for each variable in the union of the variable sets of each first data set, as follows. It is assumed that the topology of the combined causal effect matrix A is the same for all first data sets. Here, the combined causal effect matrix A reflects the causal relationships between variables, shown in matrix B. These causal relationships determine the topology of the causal graph.
[0036] The processing unit 12 generates a multiplicative matrix F to represent different causal effects for each first data point. (k) Implement F (k) This is expressed by equation (8).
[0037]
number
[0038] The backslash and {0} in equation (8) indicate the exclusion of 0, that is, to ensure that the causal relationship between the variables shown in A does not disappear. (k) By making all the elements of A positive real numbers, we can add the assumption that the sign of the causal effect represented by A is not reversed.
[0039] Processing unit 12 processes the first data X (k) A linear causal model of equation (9) is generated for each step.
[0040]
number
[0041] A and F in the parentheses on the right-hand side of equation (9) (k) The symbol for this operation (a circle with a dot inside) indicates the Hadamard product. The Hadamard product is performed on each matrix A, F (k) This is an operation that calculates the product of corresponding elements in a given matrix. Therefore, the result of the Hadamard product is a d x d matrix.
[0042] Then, the processing unit 12 calculates the pseudo-data X' based on equation (9). (k) Randomly generate X' (k) The distribution of X (k) Similar to the distribution of each G i (k) and A and each F (k) We estimate that each G i (k) This is noise information based on a Gaussian mixture distribution, and E (k) These are the elements. As mentioned above, a loss function L based on MMD or optimal transport distance is used to measure the similarity between distributions. In addition, in this estimation, the processing unit 12 reduces the value of the loss function by gradient descent for each G i (k) and A and each F (k) And optimize. i (k) ,A,each F (k) The initial value is determined, for example, randomly.
[0043] Thus, the processing unit 12 sets up first information regarding the integrated causal relationships and causal effects between variables in the entire set of first data, and second information regarding the weighting of the causal effects included in the first information for each of the set of first data. The combined causal effect matrix A corresponds to the first information. The multiplicative variable matrix F (1) ,…,F (m) This corresponds to the second piece of information.
[0044] Then, the processing unit 12 generates multiple second data corresponding to multiple first data based on the first and second information. Multiple pseudo-data X' (1) ,…,X' (m) This corresponds to multiple second data points. The processing unit 12 optimizes the first and second information to reduce the value of the loss function for the multiple first data points and the multiple second data points.
[0045] This allows the information processing device 10 to estimate the integrated causal relationships between variables in multiple datasets (first data) and the individual causal effects in each dataset (first data). The integrated causal relationships in multiple first data are represented by the combined causal effect matrix A. The individual causal effects in the k-th first data are represented by A and the multiplicative variable matrix F. (k) It is expressed by the Hadamard product of . In a more specific example, the information processing device 10 can estimate integrated causal relationships and individual causal effects for multiple datasets that may have different sets of variables and properties, such as diagnostic records from hospitals with different patient tendencies or experimental records from laboratories with different equipment.
[0046] Furthermore, the processing unit 12 can individually calculate the causal graph for the k-th first data point as follows. First, the processing unit 12 calculates matrix A using equation (10). (k) Calculate.
[0047]
number
[0048] Processing unit 12 is A (k) Using this, equation (11) gives the adjacency matrix B representing the direct causal effect on each variable of the k-th first data point. (k) It can be calculated.
[0049]
number
[0050] As a result, the processing unit 12 is B (k) Based on this, a causal graph can be generated for each first data point. For example, the processing unit 12 may present the information to the user by displaying an image representing the causal graph for each first data point on a display device, or by transmitting information representing the causal graph for each first data point to another information processing device. In this way, the information processing device 10 can also help the user easily grasp the detailed causal effects for each first data point.
[0051] [Second Embodiment] Next, a second embodiment will be described. Figure 2 shows an example of the hardware of the information processing device according to the second embodiment.
[0052] The information processing device 100 estimates, for each dataset, a causal graph of a linear causal model in the union of the integrated set of variables, i.e., the union of the sets of variables, from multiple datasets. The information processing device 100 includes a processor 101, RAM 102, HDD 103, GPU 104, input interface 105, media reader 106, and communication interface 107. These units of the information processing device 100 are connected to a bus internally. The processor 101 corresponds to the processing unit 12 of the first embodiment. The RAM 102 or HDD 103 corresponds to the storage unit 11 of the first embodiment. The information processing device 100 may also be called a computer.
[0053] The processor 101 is an arithmetic unit that executes program instructions. The processor 101 is, for example, a CPU. The processor 101 loads at least a portion of the programs and data stored in the HDD 103 into the RAM 102 and executes the program. The processor 101 may include multiple processor cores. The information processing device 100 may have multiple processors. The processor that executes one of the multiple processes performed by the information processing device 100 may be different from the processor that executes a different process from the multiple processes. In addition, at least a portion of the processes described below may be executed in parallel using multiple processors or processor cores. A collection of multiple processors is sometimes called a "multiprocessor" or simply a "processor". A processor may also be called a "processor circuitry".
[0054] RAM 102 is a volatile semiconductor memory that temporarily stores programs executed by the processor 101 and data used by the processor 101 for calculations. The information processing device 100 may also be equipped with other types of memory, and may be equipped with multiple types of memory.
[0055] HDD103 is a non-volatile storage device that stores software programs such as the OS (Operating System), middleware, and application software, as well as data. The information processing device 100 may also be equipped with other types of storage devices such as flash memory or SSD (Solid State Drive), and may be equipped with multiple non-volatile storage devices.
[0056] The GPU 104 outputs an image to the display 111 connected to the information processing unit 100, according to instructions from the processor 101. Any type of display can be used as the display 111, such as a CRT (Cathode Ray Tube) display, an LCD (Liquid Crystal Display), a plasma display, or an organic electro-luminescence (OEL) display.
[0057] The input interface 105 acquires input signals from the input device 112 connected to the information processing device 100 and outputs them to the processor 101. The input device 112 can be a pointing device such as a mouse, touch panel, touchpad, or trackball, or a keyboard, remote controller, or button switch. Furthermore, multiple types of input devices may be connected to the information processing device 100.
[0058] The media reader 106 is a reading device that reads programs and data recorded on the recording medium 113. The recording medium 113 can be, for example, a magnetic disk, an optical disk, a magneto-optical disk (MO), or semiconductor memory. Magnetic disks include flexible disks (FD) and HDDs. Optical disks include CDs (Compact Discs) and DVDs (Digital Versatile Discs).
[0059] The media reader 106 copies programs and data read from the recording medium 113 to other recording media such as RAM 102 or HDD 103. The read programs are executed by the processor 101, for example. The recording medium 113 may be a portable recording medium and may be used for distributing programs and data. The recording medium 113 and HDD 103 are sometimes referred to as computer-readable recording media.
[0060] The communication interface 107 is connected to the network 114 and communicates with other information processing devices via the network 114. The communication interface 107 may be a wired communication interface connected to a wired communication device such as a switch or router, or a wireless communication interface connected to a wireless communication device such as a base station or access point.
[0061] Figure 3 shows an example of a causal graph. As mentioned above, the causal relationships between variables in a dataset are represented by a causal graph of a linear causal model. The linear causal model is expressed as X = BX + E, as shown in equation (1). The dataset X is represented as tabular data, with rows corresponding to samples and columns corresponding to variables.
[0062] Figure 3(A) illustrates matrix 21, which corresponds to dataset X, having variable values x1, x2, and x3 for multiple samples. Figure 3(B) illustrates matrix 22, which corresponds to adjacency matrix B. For variables x1, x2, and x3, B is a 3x3 matrix. Element b of B ij is, x j →x i This represents the direct causal effect. Figure 3(C) shows the causal graph 23 corresponding to the adjacency matrix B. Note that the noise matrix E = (e1, e2, e3). The elements of E are e i For example, this is a uniformly random number between (-1,1).
[0063] Here, the causal graph includes nodes corresponding to each variable and directed edges connecting pairs of nodes that have a causal relationship. The fact that variable x1 influences variable x2, that is, that there is a causal relationship from variable x1 to variable x2, is indicated by directed edges from the node of variable x1 to the node of variable x2.
[0064] For example, causal graph 23 shows that variable x1 influences variables x2 and x3, and variable x2 influences variable x3. The direct causal effect from variable x1 to variable x2 is 2.0. The direct causal effect from variable x1 to variable x3 is 0.5. The direct causal effect from variable x2 to variable x3 is -1.0.
[0065] Figure 4 shows an example of multiple datasets. The information processing device 100 has m different datasets X ~(1) ,X ~(2) ,…,X ~(m) We will perform a comprehensive causal search regarding "X". ~ " represents the letter X with a tilde symbol "~" above it.
[0066] each X ~(k) The sample sizes may differ for each X. ~(k) The samples included do not have to be identical. Also, each X ~(k) is the variable set V (k) It is linked to V. (k) and V (k’) They do not have to match. However, V (k) A portion of it is at least one other V (k’) This overlaps with part of the previous section.
[0067] One example of a dataset is a set of diagnostic records from a single hospital. For instance, the diagnostic records from one hospital constitute one dataset. For example, one hospital might measure a patient's blood pressure but not their temperature, while another hospital might measure their temperature but not their blood pressure. Furthermore, the patients themselves may differ from hospital to hospital.
[0068] Another example of a dataset is experimental records of chemical reactions. For instance, the experimental records of one researcher constitute one dataset. For example, one researcher might measure reaction energy but not reaction rate, while another researcher might measure reaction rate but not reaction energy. Furthermore, the number of experiments and experimental conditions may differ from one researcher to another.
[0069] Dataset group 30 is an example of multiple datasets. Dataset group 30 is a subset of dataset X. ~(1) ,X ~(2) ,X ~(3) It has. Dataset X ~(1) This includes the values of variables x1, x2, and x3. (Dataset)~(2) includes the respective values of variables x1 and x3. The dataset ~(3) includes the respective values of variables x2 and x4.
[0070] Dataset X ~(1) , X ~(2) , X ~(3) The union of each variable set of, X ~(1) , X ~(2) , X ~(3) is (x1, x2, x3, x4). The information processing device 100 performs integrated causal exploration on each of dataset X
[0071] In the causal exploration by the information processing device 100, X in Equation (9) is used as a linear causal model (k) =(A * F (k) ) E (k) to enable the estimation of the integrated causal relationship between variables across all datasets and the individual causal effects in each dataset. Note that the above asterisk symbol "*" indicates the Hadamard product of A and F (k) .
[0072] Figure 5 is a diagram showing a functional example of the information processing device. The information processing device 100 includes a data storage unit 120, a calculation processing unit 130, and a result output unit 140. The storage areas of the RAM 102 and the HDD 103 are used for the data storage unit 120. The calculation processing unit 130 and the result output unit 140 are realized by the program stored in the RAM 102 being executed by the processor 101.
[0073] The data storage unit 120 stores the input dataset X ~(1) , X ~(2) , …, X ~(m) and the comprehensive causal effect matrix A and the multiple variable matrix F in Equation (9) (1) , F (2) , …, F (m) . As described above, E (k) The noise value e included ini (k) The noise is generated based on a predetermined distribution function. For example, a Gaussian mixture distribution G is used as the distribution function that generates the noise value. In this case, E (k) =(e1 (k) ,e2 (k) ,…,e d (k) Each element of ) is a Gaussian mixture distribution G1 (k) G2 (k) ,…,G d (k) It is generated based on e i (k) is G i (k) It is generated based on G1. d is the total number of variables belonging to the union of the variable sets of each dataset. Therefore, the data storage unit 120 generates G1 (k) G2 (k) ,…,G d (k) It also remembers that information.
[0074] A, F (1) ,F (2) ,…,F (m) G1 (1) ,…,G d (m) The initial values of each element are determined, for example, randomly. However, specific values may be given as the initial values for each element.
[0075] Furthermore, for example, in the case where the aforementioned hospital-specific diagnostic records are used as the input dataset, X ~(1) ,X ~(2) ,…,X ~(m) For example, m diagnostic records (datasets) obtained from m hospitals are input into the information processing device 100 and stored in the data storage unit 120. In the example where the experimental records of the chemical reaction described above are used as the input dataset, X ~(1) ,X ~(2) ,…,X ~(m) For example, m experimental records (datasets) obtained from m researchers are input into the information processing device 100 and stored in the data storage unit 120.
[0076] The calculation processing unit 130 calculates the pseudo-dataset X^ based on equation (9). (1) ,X^ (2) ,…,X^ (m) This generates the following: Here, "X^" represents the letter X with a caret symbol "^" placed above it.
[0077] The calculation processing unit 130 processes the input dataset X ~(1) ,…,X ~(m) and pseudo-dataset X^ (1) ,…,X^ (m) Based on a predetermined loss function, gradient descent is used to obtain A and each F. (k) ,each G i (k) The loss function is optimized using methods such as MMD or optimal transport distance. In the optimization, gradient descent is used to minimize the loss function, and A and F are each optimized. (k) ,each G i (k) The values of the elements are estimated. For minimizing the sum of MMDs, equation (8) in Non-Patent Document 1 may be helpful, for example.
[0078] The calculation processing unit 130 processes A and each F. (k) ,each G i (k) Once the optimization of is complete, X is calculated based on equation (11). ~(1) ,…,X ~(m) Adjacent matrix B corresponding to (1) ,…,B (m) The calculation processing unit 130 calculates B (1) ,…,B (m) The data is stored in the data storage unit 120.
[0079] The result output unit 140 outputs the calculation results from the calculation processing unit 130. For example, the result output unit 140 outputs B (1) ,…,B (m) Based on this, m causal graphs may be displayed on the display 111, or the information of the m causal graphs may be transmitted to other information processing devices via the network 114.
[0080] Figure 6 shows an example of the optimization process. Diagram 40 illustrates the optimization process performed by the calculation processing unit 130. The calculation processing unit 130 is G1 (k) ,…,G d (k) Using E (k) The calculation processing unit 130 generates E (1) ,…,E (m) A and F (1) ,…,F (m) Using and, according to equation (9), the pseudo-dataset X^ (1) ,X^ (2) ,…,X^ (m) The calculation processing unit 130 generates the input dataset X. ~(1) ,…,X ~(m) and pseudo-dataset X^ (1) ,…,X^ (m) Minimize the value of the loss function with respect to (G1 (1) ,…,G d (1) ),…,(G1 (m) ,…,G d (m) ) and A and F (1) ,…,F (m) The process involves optimizing the following. This optimization is performed using gradient descent. Examples of gradient descent methods that can be used include stochastic gradient descent and batch gradient descent (steepest descent).
[0081] Here, the Gaussian mixture distribution G i (k) Further details are provided. The computation processing unit 130, through the above optimization, processes the dataset X ~(k) The variable x in i As a model to approximate the noise distribution, a Gaussian mixture distribution G i (k) Learn each G i (k) It has the following parameters:
[0082] The first parameter is the number of Gaussian distributions g included in the Gaussian mixture. g is a user parameter. All G i (k) We can also use g as a common value. The larger g is, the better the approximation.
[0083] The second parameter is the mean vector μ. i (k) =(μ i,1 (k) ,…,μ i,g (k) ) The third parameter is the dispersion vector ρ i (k) =(ρ i,1 (k) ,…,ρ i,g (k) )
[0084] The fourth parameter is the weight vector w i (k) =(w i,1 (k) ,…,w i,g (k) ) μ i (k) ρ i (k) ,w i (k) These are the parameters that are adjusted through optimization.
[0085] each G i (k) This generates random noise as a weighted sum of Gaussian distributions. Specifically, the calculation processing unit 130 generates a Gaussian distribution z for each t. i,t (k) ~N(μ i,t (k) ρ i,t (k) The calculation processing unit 130 generates random numbers (w'). The calculation processing unit 130 also transforms the weight vector using the Softmax function so that its sum is 1. The transformed weight vector (w' i,1 (k) ,…,w' i,g (k) ) = Softmax(w i,1 (k) ,…,w i,g (k) )
[0086] Then, the calculation processing unit 130 calculates the weighted sum v i(k) =Σ t w' i,t (k) z i,t (k) G i (k) Random value (noise value) e i (k) Let's assume that. The calculation processing unit 130 calculates each G based on the gradient descent method. i (k) The gradient ∇G i (k) Therefore, each G i (k) The parameter values are updated. The calculation processing unit 130 calculates the gradient ∇G i (k) The gradient ∇μ with respect to each parameter is i (k) ,∇ρ i (k) ,∇w i (k) Each of these is calculated. Then, the calculation processing unit 130 calculates μ i (k) ←μ i (k) -δ∇μ i (k) ρ i (k) ←ρ i (k) -δ∇ρ i (k) ,w i (k) ←w i (k) -δ∇w i (k) The update is performed for each parameter. Here, δ is the learning rate.
[0087] Note that the parameters to be adjusted in A are each a ij That is. Also, F (k) The parameters to be adjusted in each f ij That is the case. Figure 7 shows an example of data stored in the data storage unit.
[0088] The data storage unit 120 stores the input dataset group 150 and the variable group 160. The input dataset group 150 is a collection of input datasets. For example, the input dataset group 150 includes input datasets 151, 152, and 153. In one example, input dataset 151 is the diagnostic record of hospital (1). Input dataset 152 is the diagnostic record of hospital (2). Input dataset 153 is the diagnostic record of hospital (3).
[0089] Input dataset 151 contains the values for the variables "drug dosage," "blood glucose level," and "urination frequency" for multiple patients at hospital (1). Input dataset 152 contains the values for the variables "drug dosage" and "blood glucose level" for multiple patients at hospital (2). Input dataset 153 contains the values for the variables "drug dosage," "blood glucose level," and "blood pressure" for multiple patients at hospital (3). Input datasets 151, 152, and 153 are combined into input dataset X. ~(1) ,X ~(2) ,X ~(3) This is one example.
[0090] The variable group 160 is a set of values for each parameter representing the linear causal model of equation (9). The values of each parameter in the variable group 160 are the parameter values to be adjusted by optimization in the calculation processing unit 130.
[0091] For example, for input datasets 150, the variable set 160 is (G 薬剤投与量 (1) ,G 血糖値 (1) ,G 尿頻度 (1) ,G 血圧 (1) ), (G 薬剤投与量 (2) ,G 血糖値 (2) ,G 尿頻度 (2) ,G 血圧 (2) ), (G 薬剤投与量 (3) ,G 血糖値 (3) ,G 尿頻度 (3) ,G 血圧(3) ) includes.
[0092] Furthermore, for the input datasets 150, the variable group 160 consists of a composite causal effect matrix A and a multiple variable matrix F. (1) ,F (2) ,F (3) Includes. Next, we will explain the loop processing (optimization loop) in the gradient descent optimization performed by the calculation processing unit 130 on the input dataset 150 and the variable group 160.
[0093] Figure 8 shows an example of pseudo-data generation in an optimization loop. First, the calculation processing unit 130, (G 薬剤投与量 (1) ,G 血糖値 (1) ,G 尿頻度 (1) ,G 血圧 (1) ) and A and F (1) Based on this, a dummy dataset 161 is generated that corresponds to the input dataset 151 of hospital (1). The number of samples in dataset 161 is the same as the number of samples in the input dataset 151.
[0094] Furthermore, the calculation processing unit 130 is (G 薬剤投与量 (2) ,G 血糖値 (2) ,G 尿頻度 (2) ,G 血圧 (2) ) and A and F (2) Based on this, a dummy dataset 162 is generated that corresponds to the input dataset 152 of hospital (2). The number of samples in dataset 162 is the same as the number of samples in the input dataset 152.
[0095] Furthermore, the calculation processing unit 130 is (G 薬剤投与量 (3) ,G 血糖値 (3) ,G 尿頻度 (3) ,G 血圧 (3)) and A and F (3) Based on this, a dummy dataset 163 is generated that corresponds to the input dataset 153 of hospital (3). The number of samples in dataset 163 is the same as the number of samples in the input dataset 153.
[0096] Datasets 161, 162, and 163 are pseudo-datasets X^ (1) ,X^ (2) ,X^ (3) This is one example. Below, datasets 161, 162, and 163 will be referred to as pseudo-datasets 161, 162, and 163.
[0097] Figure 9 shows an example of loss calculation in an optimization loop. Figure 9 illustrates the case where MMD is used as the loss function. The calculation processing unit 130 uses an MMD function that takes the input dataset 151 and the pseudo-dataset 161 as input, and uses this function as Loss(1), which represents the difference between the input dataset 151 and the pseudo-dataset 161. Loss(1) can also be said to be a function that shows the similarity between the input dataset 151 and the pseudo-dataset 161.
[0098] The calculation processing unit 130 uses an MMD function that takes the input dataset 152 and the pseudo-dataset 162 as inputs, and defines the function Loss(2) as the difference between the input dataset 152 and the pseudo-dataset 162. Loss(2) can also be described as a function that shows the similarity between the input dataset 152 and the pseudo-dataset 162.
[0099] The calculation processing unit 130 uses an MMD function that takes the input dataset 153 and the pseudo-dataset 163 as inputs, and defines the function Loss(3) as the difference between the input dataset 153 and the pseudo-dataset 163. Loss(3) can also be described as a function that shows the similarity between the input dataset 153 and the pseudo-dataset 163.
[0100] The calculation processing unit 130 can, for example, make the loss function L a function that represents the sum of Loss(1), Loss(2), and Loss(3), such as L = Loss(1) + Loss(2) + Loss(3).
[0101] The value of the loss function can be said to be a similarity score, indicating the degree of similarity between the input datasets and the pseudo-datasets. Alternatively, it can be said to be a deviation score, indicating the degree of divergence between the input datasets and the pseudo-datasets. A smaller similarity score (or deviation score) indicates a higher degree of similarity between the input datasets and the pseudo-datasets. The minimum value of the similarity score (or deviation score) is 0.
[0102] The calculation processing unit 130 may calculate the values of Loss(1), Loss(2), Loss(3), and L in the optimization loop and save these values as history in the data storage unit 120.
[0103] Figure 10 shows an example of gradient propagation in an optimization loop. Here, the loss function L is given by each G i (k) ,A,each F (k) The parameters to be adjusted are stored as variables. The calculation processing unit 130 calculates the gradient ∇L of the loss function L with respect to these parameters to be adjusted. ∇ is a vector differential operator with respect to each parameter to be adjusted.
[0104] For example, among the components of ∇L, the vector that has a component corresponding to the parameter to be adjusted included in A is conveniently denoted as ∇A. i (k) ,F (k) Similarly, for ∇L, G i (k) ,F (k) For convenience, the vector with the corresponding component is called ∇G i (k) ,∇F (k) It can be written as shown. The explanation in Figure 6 also uses this convenient notation for the gradient. The calculation processing unit 130 processes each G i (k) ,A,each F(k) For the current value of i (k) , each ∇G (k) gradient information 170 including the values of ∇A and each ∇F is obtained.
[0105] Then, based on the gradient information 170, the calculation processing unit 130 updates each adjustment target parameter of the variable group 160. For example, G 薬剤投与量 (1) includes μ 薬剤投与量 (1) . Also, ∇G 薬剤投与量 (1) includes ∇μ 薬剤投与量 (1) . In this case, the calculation processing unit 130 uses the learning rate δ to update μ 薬剤投与量 (1) ←μ 薬剤投与量 (1) -δ∇μ 薬剤投与量 (1) in the following manner. The calculation processing unit 130 updates the value of each adjustment target parameter for each adjustment target parameter in the same way for A and F (k) as well.
[0106] The calculation processing unit 130 repeatedly executes a series of procedures exemplified in FIGS. 8 to 10 (optimization loop). Then, when the calculation processing unit 130 satisfies a predetermined end condition, it ends the optimization loop and determines each G i (k) , A, and each F (k) . The calculation processing unit 130 stores each G i (k) , A, and each F (k) after optimization in the data storage unit 120.
[0107] FIG. 11 is a diagram showing an example of output of results. The calculation processing unit 130 applies each of F (1) , F (2) , and F (3) to the optimized A (step ST1). Specifically, the calculation processing unit 130 calculates A (1) , A (2) , and A (3) based on equation (10) (step ST2).
[0108] Based on Equation (11), the calculation processing unit 130 calculates B corresponding to Hospital (1), Hospital (2), and Hospital (3) respectively. (1) , B (2) , B (3) (Step ST3). Based on B, the result output unit 140 generates a causal graph 181 corresponding to Hospital (1) and displays it on the display 111. Based on B, the result output unit 140 generates a causal graph 182 corresponding to Hospital (2) and displays it on the display 111. Based on B, the result output unit 140 generates a causal graph 183 corresponding to Hospital (3) and displays it on the display 111. The result output unit 140 may transmit the information of the causal graphs 181, 182, and 183 to other information processing devices via the network 114 and display the causal graphs 181, 182, and 183 on the displays of other information processing devices. (1) (2) (3)
[0109] Next, the processing procedure of the information processing device 100 will be described. FIG. 12 is a flowchart showing a processing example of the information processing device. (S10) The calculation processing unit 130 acquires the input data sets X ~(1) , X ~(2) , …, X ~(m) .
[0110] (S11) The calculation processing unit 130 initializes each variable G i (k) , A, F (k) . Let i = 1, …, d. Let k = 1, …, m. For example, the calculation processing unit 130 randomly determines the initial values of each G i (k) , A, F (k) .
[0111] (S12) The calculation processing unit 130 determines whether to complete the repetition of steps S13 to S16 below. If the repetition is completed, the process proceeds to step S17. If the repetition is not completed, the process proceeds to step S13. For example, the calculation processing unit 130 counts the number of repetitions of steps S13 to S16 and determines that the repetition is completed when the number of repetitions reaches a certain number. Also, the calculation processing unit 130 determines that the repetition is not completed if the number of repetitions has not reached a certain number. However, the termination condition for the optimization loop may be other conditions. For example, the calculation processing unit 130 may use each value of Loss(1) to Loss(m) or L=Σ k You may determine whether the value of Loss(k) has become smaller than the reference value, and if each of those values or the value of L has become smaller than the reference value, you may determine that the iteration is complete.
[0112] (S13) The calculation processing unit 130 processes each G i (k) ,A,each F (k) Using this, based on equation (9), the pseudo-dataset X^ (1) ,X^ (2) ,…,X^ (m) Generates. (S14) The calculation processing unit 130 is (X ~(k) ,X^ (k) The MaximumMeanDiscrepancy function (MMD function), which takes the input ) as input, is obtained as Loss(k). Note that the function used as Loss(k) can be any other function, such as the optimal transport distance.
[0113] (S15) The calculation processing unit 130 calculates each gradient ∇G from Loss(1) to Loss(m). i (k) ,∇A,∇F (k) The calculation is performed by the calculation processing unit 130 for each G. i=1,...,d. k=1,...,m. i (k) ,A,each F (k) For each parameter to be adjusted, each gradient ∇G i (k) ,∇A,∇F (k) The values of the elements can be calculated.
[0114] For example, the calculation processing unit 130 calculates the loss function L = Σ k It can be expressed as Loss(k). As mentioned above, ∇G i (k) G is one of the components of ∇L. i (k) This is a vector with a component corresponding to . ∇A is a vector with a component corresponding to A among all the components of ∇L. ∇F (k) F is one of the total components of ∇L. (k) This is a vector with components corresponding to [the specified value].
[0115] (S16) The calculation processing unit 130 processes each variable G i (k) ,A,F (k) Each gradient ∇G i (k) ,∇A,∇F (k) The update is performed based on the following: i=1,...,d. k=1,...,m. Then the process proceeds to step S12. Here, the calculation processing unit 130, as described above, sets the learning rate to δ and updates the value of p in the form p←p-δ∂L / ∂p for the change in the loss function L ∂L / ∂p corresponding to the parameter component p.
[0116] (S17) The calculation processing unit 130 calculates A and F after optimization. (k) Using equation (10), matrix A (k) The calculation processing unit 130 calculates A (k) Using equation (11), B (k) We calculate the following, where k=1,…,m.
[0117] (S18) The result output unit 140 outputs the calculation result B (1) ,B (2) ,…,B (m) The output unit 140 outputs the result B (k) Based on this, an image showing a causal graph for each k unit may be output. Then, the processing of the information processing device 100 ends.
[0118] Thus, the information processing device 100 can estimate integrated causal relationships across multiple datasets and individual causal effects within each dataset. Figure 13 shows a comparative example.
[0119] The comparative example illustrates the case where the linear causal model is represented by equation (5) X=AE. The optimization process of the comparative example is represented by diagram 50. Diagram 50 shows the multiplicative matrix F (1) ,F (2) ,…,F (m) Diagram 40 differs from this in that it does not include [a specific element]. Therefore, in the comparative example, only a common causal effect (matrix A) is obtained for all datasets.
[0120] Figure 14 shows an example of a causal graph that cannot be distinguished in the comparative example. The comparative example method assumes the same linear causal model X=AE for all datasets. This assumption has the problem of failing to consider the possibility that causal effects may differ from dataset to dataset. For example, the effect of a drug may differ in hospitals where patient tendencies regarding drug administration differ.
[0121] Causal graphs 61, 62, and 63 show the intrinsic causal effects between variables in data from diabetic patients at three hospitals (1), (2), and (3). Variable x1 is the drug dosage. Variable x2 is the blood glucose level. Variable x3 is the urination frequency. As shown in causal graphs 61, 62, and 63, the causal effects for variables x1, x2, and x3 differ depending on the patient's characteristics at the three hospitals (1), (2), and (3). However, the method in the comparative example only yields common causal effects across all datasets and cannot obtain different causal effects for each hospital. In other words, the method in the comparative example cannot distinguish between causal graphs 61, 62, and 63.
[0122] Other examples of situations where causal effects differ across datasets are also possible. For instance, the results of water-based experiments may vary due to differences in water quality if the laboratory is located in a different region. Therefore, the information processing device 100 expresses the linear causal model using equation (9), thereby determining the integrated causal relationship (A) across multiple datasets and the individual causal effects (A and F) in each dataset. (k) The Hadamard product of ( ) can be estimated. For this reason, the information processing device 100 calculates the matrix B representing the direct causal effect for hospitals (1), (2), and (3) in Figure 14. (1) ,B (2) ,B (3) It is possible to distinguish and calculate B. (1) ,B (2) ,B (3) It is possible to distinguish between the corresponding causal graphs 61, 62, and 63. In this way, the information processing device 100 can obtain detailed causal graphs that reflect the respective environments in which each input dataset was acquired. By distinguishing and presenting the causal graphs that reflect each environment to the user, the information processing device 100 can support the user in performing detailed data analysis.
[0123] As described above, the information processing device 100 performs the following processing, for example: The data storage unit 120 stores a plurality of first data. Each of the plurality of first data contains a plurality of variable values corresponding to a plurality of variables and is used to estimate causal relationships and causal effects between variables based on the plurality of variable values. The calculation processing unit 130 sets first information regarding the integrated causal relationships and causal effects in the plurality of first data as a whole, and second information regarding the weighting of the causal effects included in the first information for each of the plurality of first data. Based on the first information and the second information, the calculation processing unit 130 generates a plurality of second data corresponding to the plurality of first data. The calculation processing unit 130 optimizes the first information and the second information to reduce the value of the loss function for the plurality of first data and the plurality of second data.
[0124] This allows the information processing device 100 to estimate the integrated causal relationships between variables in multiple datasets (first data) and the individual causal effects in each dataset (first data). Here, the input dataset X ~(1) ,…,X ~(m)This is an example of multiple first data sets. Multiple pseudo-datasets (X^ (1) ,…,X^ (m) ) is an example of multiple second data points. The combined causal effect matrix A is an example of the first information. The multiplicative variable matrix F (1) ,…,F (m) This is an example of second-level information. Also, the sum of the Loss function Σ, which indicates the similarity (or deviation) between data such as MMD and optimal transport distance. k=1 m Loss(k) is an example of a loss function.
[0125] More specifically, the first information includes a composite causal effect matrix whose elements are multiple first parameters that show the causal effects between variables belonging to the union of multiple variables in each of the multiple first data. The second information is a multiplicative variable matrix whose elements are multiple second parameters that weight the multiple first parameters, and includes a multiplicative variable matrix corresponding to each of the multiple first data. This allows the information processing device 100 to estimate the individual causal effects for each of the first data in detail. Note that the multiplicative variable matrix F (k) The second parameter f included in ij This is the first parameter a, which represents the causal effect between variables in the composite causal effect matrix. ij This can be said to represent the weights in the k-th input dataset. Also, the first parameter a ij This can also be said to represent an integrated causal relationship between variables.
[0126] Furthermore, the calculation processing unit 130 can use gradient descent to determine the values of each of the multiple first parameters included in the overall causal effect matrix and each of the multiple second parameters included in each multiplicative variable matrix when optimizing the first and second information. This allows the information processing device 100 to efficiently determine the values of each parameter to be adjusted in the first and second information. ij This is an example of a parameter to be adjusted in the first information.
[0127] For example, among a plurality of first data, some of the variables included in any one of the data may overlap with some of the variables included in other data, and the parts other than some may not overlap. Therefore, the information processing apparatus 100 can estimate the integrated causal effect in a plurality of data sets (first data) and the individual causal effect in each data set (first data) with respect to the union of the variable sets of each of the plurality of first data.
[0128] Based on the optimized first information and second information, the result output unit 140 outputs information indicating the causal effect between variables in each of the plurality of first data. Thereby, the information processing apparatus 100 can assist the user in grasping the individual causal effect in each of the plurality of first data. For example, the result output unit 140 may output an image of a causal graph as the information indicating the causal effect. The user can grasp the causal relationship between variables and the individual causal effect of the first data from the causal graph.
[0129] In generating a plurality of second data, the calculation processing unit 130 generates a plurality of second data based on the first information, the second information, and noise information based on a predetermined distribution function. In optimization, the calculation processing unit 130 performs optimization of the distribution function together with the first information and the second information. Thereby, the calculation processing unit 130 can appropriately perform optimization of the first information and the second information. Here, the noise matrices E (1) ,…,E (m) are an example of noise information. The mixture Gaussian distribution G i (k) is an example of the distribution function used for generating the noise information.
[0130] Note that the information processing in the first embodiment can be realized by causing the processing unit 12 to execute a program. Also, the information processing in the second embodiment can be realized by causing the processor 101 to execute a program. The program can be recorded on a computer-readable recording medium 113.
[0131] For example, a program can be distributed by distributing a recording medium 113 on which the program is stored. Alternatively, the program may be stored on another computer and distributed via a network. A computer may, for example, store (install) a program stored on the recording medium 113 or a program received from another computer into a storage device such as RAM 102 or HDD 103, and then read and execute the program from that storage device. [Explanation of Symbols]
[0132] 10 Information Processing Devices 11 Storage section 12 Processing Units
Claims
1. On the computer, Multiple first data sets, each containing multiple variable values corresponding to multiple variables, are acquired for use in estimating causal relationships and causal effects between variables based on the multiple variable values. First information relating to the integrated causal relationships and causal effects in the entirety of the plurality of first data sets, and second information relating to the weighting of the causal effects included in the first information for each of the plurality of first data sets, are set. Based on the first and second information, a plurality of second data corresponding to the plurality of first data are generated. The first and second information are optimized to reduce the value of the loss function relating to the plurality of first data and the plurality of second data. A program that executes a process.
2. The first information includes a composite causal effect matrix whose elements are a plurality of first parameters that show the causal effects between variables belonging to the union of the plurality of variables in each of the plurality of first data, The second information is a multiplicative matrix whose elements are a plurality of second parameters that weight the plurality of first parameters, and includes the multiplicative matrix corresponding to each of the plurality of first data, The program according to claim 1.
3. To the aforementioned computer, In the optimization described above, the gradient descent method is used to determine the values of each of the plurality of first parameters and each of the plurality of second parameters. The program according to claim 2 that causes a process to be executed.
4. The variables included in one of the plurality of first data sets and the variables included in the other of the plurality of first data sets overlap in part, and the remaining parts do not overlap. The program according to claim 1.
5. To the aforementioned computer, Based on the first and second information after optimization, information indicating the causal effect in each of the plurality of first data is output. A program according to claim 1 that causes a process to be executed.
6. To the aforementioned computer, In the generation of the plurality of second data, the plurality of second data are generated based on the first information, the second information, and noise information based on a predetermined distribution function. In the optimization described above, the distribution function is optimized together with the first information and the second information. A program according to claim 1 that causes a process to be executed.
7. Computers Multiple first data sets, each containing multiple variable values corresponding to multiple variables, are acquired for use in estimating causal relationships and causal effects between variables based on the multiple variable values. First information relating to the integrated causal relationships and causal effects in the entirety of the plurality of first data sets, and second information relating to the weighting of the causal effects included in the first information for each of the plurality of first data sets, are set. Based on the first and second information, a plurality of second data corresponding to the plurality of first data are generated. The first and second information are optimized to reduce the value of the loss function relating to the plurality of first data and the plurality of second data. Information processing methods.
8. A storage unit that stores a plurality of first data sets, each containing a plurality of variable values corresponding to a plurality of variables, which are used to estimate causal relationships and causal effects between variables based on the plurality of variable values, A processing unit sets first information relating to the integrated causal relationships and causal effects in the entirety of the plurality of first data, and second information relating to the weighting of the causal effects included in the first information for each of the plurality of first data, generates a plurality of second data corresponding to the plurality of first data based on the first information and the second information, and optimizes the first information and the second information to reduce the value of the loss function relating to the plurality of first data and the plurality of second data, An information processing device having