Program, information processing method, and information processing device
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2026-04-02
AI Technical Summary
Existing technologies cannot effectively handle the differences in causal effects between different datasets when estimating causal relationships between multiple datasets, resulting in the inability to accurately estimate the causal effects of individual datasets.
By generating pseudo-data, information processing devices and methods are used to optimize the estimation of causal relationships and causal effects, generate integrated causal relationships that conform to multiple datasets and causal effects of individual datasets, and perform data processing and optimization using multiple processing units and storage units.
It enables accurate estimation of integrated causal relationships across multiple datasets and causal effects in individual datasets, applicable to datasets with different sets of variables and attributes, such as hospital diagnostic records and laboratory records.
Smart Images

Figure JP2025028557_02042026_PF_FP_ABST
Abstract
Description
Program, information processing method, and information processing device.
[0001] This invention relates to a program, an information processing method, and an information processing device.
[0002] Information processing techniques are sometimes used to estimate causal relationships between variables in multiple datasets. For example, a causal relationship in which one variable influences another can be represented as a causal graph in a linear causal model. A causal graph may include causal effects between variables. A causal effect is the degree to which one variable influences another.
[0003] Here, assuming that causal effects are common across all datasets, a method has been proposed to optimize causal effects so that pseudo-data closely resembling the input data can be generated.
[0004] B. Huang, and 3 others, “Causal Discovery from Multiple Data Sets with Non-Identical Variables”, [online], 2020, AAAI'20, [Retrieved September 2, 2020], Internet <URL: https: / / cdn. aaai. org / ojs / 6575 / 6575-13-9800-1-10-20200519. pdf>
[0005] The method proposed above assumes that causal effects are common across all datasets. Therefore, it has the problem of being unable to handle cases where the causal effects in one dataset differ from those in other datasets.
[0006] In one aspect, the present invention aims to enable the estimation of integrated causal relationships between variables across multiple datasets and individual causal effects in each dataset.
[0007] In one embodiment, a program is provided that is executed by a computer. This program causes the computer to perform the following operations: The computer obtains a plurality of first data sets, each containing a plurality of variable values corresponding to a plurality of variables, which are used to estimate causal relationships and causal effects between variables based on the plurality of variable values. The computer sets up first information regarding the integrated causal relationships and causal effects in the plurality of first data sets as a whole, and second information regarding the weighting of the causal effects contained in the first information for each of the plurality of first data sets. Based on the first and second information, the computer generates a plurality of second data sets corresponding to the plurality of first data sets. The computer optimizes the first and second information to minimize the value of the loss function for the plurality of first data sets and the plurality of second data sets.
[0008] In one embodiment, a method for information processing performed by a computer is provided. In another embodiment, an information processing device having a storage unit and a processing unit is provided.
[0009] In one aspect, it is possible to estimate the integrated causal relationships between variables across multiple datasets and the individual causal effects in each dataset. The above and other objects, features and advantages of the present invention will become apparent from the following description in conjunction with the accompanying drawings illustrating preferred embodiments of the invention.
[0010] This is a diagram illustrating the information processing device of the first embodiment. This is a diagram showing an example of the hardware of the information processing device of the second embodiment. This is a diagram showing an example of a causal graph. This is a diagram showing an example of multiple datasets. This is a diagram showing an example of the functions of the information processing device. This is a diagram showing an example of optimization processing. This is a diagram showing an example of data stored in the data storage unit. This is a diagram showing an example of pseudo-data generation in the optimization loop. This is a diagram showing an example of loss calculation in the optimization loop. This is a diagram showing an example of gradient propagation in the optimization loop. This is a diagram showing an example of result output. This is a flowchart showing an example of processing by the information processing device. This is a diagram showing a comparative example. This is a diagram showing an example of a causal graph that cannot be distinguished in the comparative example.
[0011] Hereinafter, this embodiment will be described with reference to the drawings. [First Embodiment] The first embodiment will be described.
[0012] Figure 1 is a diagram illustrating an information processing device of a first embodiment. The information processing device 10 estimates causal relationships and causal effects between variables based on multiple datasets, each having multiple variable values corresponding to multiple variables. Specifically, the information processing device 10 estimates a causal graph of a linear causal model in the union of a set of variables, an integrated set of variables, from the multiple datasets. The multiple variables included in each of the multiple datasets may all be the same, or they may be different in at least two of the datasets. In the latter case, some of the multiple variables included in any first dataset will overlap with some of the multiple variables included in at least one other dataset. The information processing device 10 has a storage unit 11 and a processing unit 12.
[0013] The memory unit 11 may be a volatile semiconductor memory such as RAM (Random Access Memory), or a non-volatile storage such as an HDD (Hard Disk Drive) or flash memory. The processing unit 12 is a processor such as a CPU (Central Processing Unit), GPU (Graphics Processing Unit), or DSP (Digital Signal Processor). However, the processing unit 12 may also include application-specific electronic circuits such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array). The processor executes programs stored in memory such as RAM (which may also be the memory unit 11). A collection of multiple processors is sometimes called a "multiprocessor" or simply a "processor."
[0014] The causal relationships between variables are represented by a causal graph of a linear causal model. In this case, dataset X is represented as tabular data, with rows corresponding to samples and columns corresponding to variables. All variables are of numerical type. The linear causal model is expressed by equation (1).
[0015]
[0016] X is represented by formula (2). n is the number of samples. d is the number of variables x i . The white "R" represents the entire set of real numbers. X is the transposed row-column representation of the table data of n samples with d items.
[0017]
[0018] B is represented by formula (3). B is the adjacency matrix of the causal graph.
[0019]
[0020] b ij represents the direct causal effect of x j → x i . That is, when b ij ≠ 0, b ij represents the causal relationship that x j affects x i , and the degree of that influence is b ij . When b ij = 0, b ij indicates that x j does not affect x i , that is, there is no causal relationship from x j to x i .
[0021] E is represented by formula (4). E is the noise matrix.
[0022]
[0023] In one example, e i is a uniform random number such that -1 < e i < 1. The storage unit 11 stores a plurality of first data used for estimating causal relationships and causal effects. Let the total number of the first data be m. The plurality of first data are pre-input to the information processing device 10 and stored in the storage unit 11. The number of samples in each of the plurality of first data may be different, and they may not be the same samples. One first data corresponds to one input data set including a plurality of variable values. Also, the variable sets included in each first data may be different.
[0024] Here, for example, in the analysis of hospital diagnostic records, examples of variables include the amount of medication administered to patients, their blood glucose levels, and their urination frequency. In this case, the patients represent the sample. Patients may differ from hospital to hospital. In this case, one first data set includes measured values such as medication dosage, blood glucose levels, and urination frequency as variable values, corresponding to each of multiple patients in a given hospital. However, the content of medical procedures and patient examinations may differ between hospitals. For example, one hospital may measure a patient's blood pressure but not their body temperature, while another hospital may measure body temperature but not their blood pressure. Another example is in experimental records of chemical reactions, where one experiment may measure reaction energy but not reaction rate, while another experiment may measure reaction rate but not reaction energy. Furthermore, the number of experiments and experimental conditions may differ from experiment to experiment.
[0025] Thus, the set of variables in one first data set may differ from the set of variables in another first data set. However, as mentioned above, a portion of the set of variables in one first data set will overlap with a portion of the set of variables in at least one other first data set.
[0026] In this case, it is conceivable to estimate a linear causal model of each variable in the union of the variable sets of each first data set. This linear causal model can take the form of, for example, equation (5).
[0027]
[0028] A is called the composite causal effect matrix. A is represented by equation (6).
[0029]
[0030] a ij is, x j →x i This represents the overall causal effect. d is the total number of variables belonging to the union of the variable sets of each first data point. I is the identity matrix. The adjacency matrix B, which represents the direct causal effect, can be reconstructed by equation (7).
[0031]
[0032] In this case, the kth first data X(k) e in i Gaussian mixture distribution G i (k) This can be modeled as follows: k is an index that identifies the first data (dataset). k = 1, 2, ..., m. First data X (k) In contrast, pseudo-data X' (k) = AE (k) These are randomly generated.
[0033] And then, X' (k) The distribution of X (k) Each G i (k) And A can be estimated. The similarity of the distributions is X (k) and X' (k) It is measured by a predetermined loss function. The loss function can be said to be a function that indicates the similarity or difference between the data. The closer the value of the loss function for two data sets is to 0, the more similar the two data sets are, or the smaller the difference between the two data sets is. Examples of loss functions used include MMD (Maximum Mean Discrepancy) and optimal transport distance. i (k) And in the estimation of A, each G is calculated using gradient descent to minimize the loss function. i (k) And A is optimized. As a gradient descent method, for example, stochastic gradient descent or batch gradient descent (steepest descent) can be used.
[0034] However, equation (5) assumes the same linear causal model for all datasets. This assumption has the problem of not being able to consider the possibility that causal effects may differ for each dataset. Specifically, in the example of multiple hospitals mentioned above, the effect of a drug may differ in hospitals where the patient trends regarding the administration of a certain drug differ. For example, the dosage and effect of a drug may differ between young people and the elderly, so the causal effect may differ between hospitals with many young patients and hospitals with many elderly patients. Also, for example, in the analysis of experimental results from multiple universities or research institutes, the results of experiments using water may differ if the laboratories are located in different regions due to differences in water quality in each region. In such cases, the linear causal model in equation (5) cannot obtain the causal effect specific to each dataset.
[0035] Therefore, the processing unit 12 estimates a linear causal model for each variable in the union of the variable sets of each first data set, as follows. It is assumed that the topology of the combined causal effect matrix A is the same for all first data sets. Here, the combined causal effect matrix A reflects the causal relationships between variables, shown in matrix B. These causal relationships determine the topology of the causal graph.
[0036] The processing unit 12 generates a multiplicative matrix F to represent different causal effects for each first data point. (k) We will introduce F. (k) This is expressed by equation (8).
[0037]
[0038] The backslash symbol and {0} in equation (8) indicate the exclusion of 0, that is, to ensure that the causal relationship between the variables shown in A does not disappear. (k) By making all the elements of A positive real numbers, we can add the assumption that the sign of the causal effect shown by A is not reversed.
[0039] Processing unit 12 processes the first data X (k) A linear causal model of equation (9) is generated for each step.
[0040]
[0041] A and F in the parentheses on the right-hand side of equation (9)(k) The symbol for this operation (a circle with a dot inside) indicates the Hadamard product. The Hadamard product is performed on each matrix A, F (k) This is an operation that calculates the product of corresponding elements in a given matrix. Therefore, the result of the Hadamard product is a d x d matrix.
[0042] Then, the processing unit 12 generates pseudo-data X' based on equation (9). (k) Randomly generate X' (k) The distribution of X (k) Each G i (k) and A and each F (k) We estimate that each G i (k) This is noise information based on a Gaussian mixture distribution, and E (k) These are the elements. As mentioned above, a loss function L based on MMD or optimal transport distance is used to measure the similarity between distributions. In addition, in this estimation, the processing unit 12 reduces the value of the loss function by gradient descent for each G i (k) and A and each F (k) And optimize. Note that each G i (k) , A, each F (k) The initial value is determined, for example, randomly.
[0043] Thus, the processing unit 12 sets up first information regarding the integrated causal relationships and causal effects between variables in the entire set of first data, and second information regarding the weighting of the causal effects included in the first information for each of the set of first data. The combined causal effect matrix A corresponds to the first information. The multiplicative variable matrix F (1) , ..., F (m) This corresponds to the second piece of information.
[0044] Then, the processing unit 12 generates a plurality of second data corresponding to a plurality of first data based on the first information and the second information. (1) , ..., X' (m)This corresponds to multiple second data points. The processing unit 12 optimizes the first and second information to reduce the value of the loss function related to the multiple first data points and the multiple second data points.
[0045] This allows the information processing device 10 to estimate the integrated causal relationships between variables in multiple datasets (first data) and the individual causal effects in each dataset (first data). The integrated causal relationships in multiple first data are represented by the combined causal effect matrix A. The individual causal effects in the kth first data are represented by A and the multiplicative variable matrix F. (k) It is expressed by the Hadamard product of . In a more specific example, the information processing device 10 can estimate integrated causal relationships and individual causal effects for multiple datasets that may have different sets of variables and properties, such as diagnostic records from hospitals with different patient tendencies or experimental records from laboratories with different equipment.
[0046] Furthermore, the processing unit 12 can individually obtain a causal graph for the k-th first data as follows. First, the processing unit 12 obtains matrix A by equation (10). (k) Calculate.
[0047]
[0048] Processing unit 12 is A (k) Using this, equation (11) gives the adjacency matrix B representing the direct causal effect on each variable of the k-th first data point. (k) It can be calculated.
[0049]
[0050] As a result, the processing unit 12 is B (k) Based on this, a causal graph can be generated for each first data point. For example, the processing unit 12 may present the information to the user by displaying an image representing the causal graph for each first data point on a display device, or by transmitting information representing the causal graph for each first data point to another information processing device. In this way, the information processing device 10 can also help the user easily grasp the detailed causal effects for each first data point.
[0051] [Second Embodiment] Next, a second embodiment will be described. Figure 2 is a diagram showing an example of the hardware of the information processing device according to the second embodiment.
[0052] The information processing device 100 estimates, for each dataset, a causal graph of a linear causal model in the union of variable sets, which is an integrated set of variables, from multiple datasets. The information processing device 100 includes a processor 101, RAM 102, HDD 103, GPU 104, input interface 105, media reader 106, and communication interface 107. These units of the information processing device 100 are connected to a bus internally. The processor 101 corresponds to the processing unit 12 of the first embodiment. The RAM 102 or HDD 103 corresponds to the storage unit 11 of the first embodiment. The information processing device 100 may also be called a computer.
[0053] The processor 101 is an arithmetic unit that executes program instructions. The processor 101 is, for example, a CPU. The processor 101 loads at least a portion of the program and data stored in the HDD 103 into the RAM 102 and executes the program. The processor 101 may include multiple processor cores. The information processing device 100 may have multiple processors. The processor that executes one of the multiple processes performed by the information processing device 100 may be different from the processor that executes a different process from the multiple processes. In addition, at least a portion of the processes described below may be executed in parallel using multiple processors or processor cores. A collection of multiple processors is sometimes called a "multiprocessor" or simply a "processor". A processor may also be called a "processor circuitry".
[0054] RAM 102 is a volatile semiconductor memory that temporarily stores programs executed by the processor 101 and data used by the processor 101 for calculations. The information processing device 100 may also be equipped with other types of memory, and may be equipped with multiple types of memory.
[0055] The HDD 103 is a non-volatile storage device that stores software programs such as the OS (Operating System), middleware, and application software, as well as data. The information processing device 100 may also be equipped with other types of storage devices such as flash memory or SSD (Solid State Drive), and may be equipped with multiple non-volatile storage devices.
[0056] The GPU 104 outputs an image to the display 111 connected to the information processing device 100, according to instructions from the processor 101. Any type of display can be used as the display 111, such as a CRT (Cathode Ray Tube) display, a liquid crystal display (LCD), a plasma display, or an organic electro-luminescence (OEL) display.
[0057] The input interface 105 acquires input signals from the input device 112 connected to the information processing device 100 and outputs them to the processor 101. The input device 112 can be a pointing device such as a mouse, touch panel, touchpad, or trackball, or a keyboard, remote controller, or button switch. Furthermore, multiple types of input devices may be connected to the information processing device 100.
[0058] The media reader 106 is a reading device that reads programs and data recorded on the recording medium 113. The recording medium 113 can be, for example, a magnetic disk, an optical disk, a magneto-optical disk (MO), or semiconductor memory. Magnetic disks include flexible disks (FD) and HDDs. Optical disks include CDs (Compact Discs) and DVDs (Digital Versatile Discs).
[0059] The media reader 106 copies programs and data read from the recording medium 113 to other recording media such as RAM 102 or HDD 103. The read programs are executed by the processor 101, for example. The recording medium 113 may be a portable recording medium and may be used for distributing programs and data. The recording medium 113 and HDD 103 are sometimes referred to as computer-readable recording media.
[0060] The communication interface 107 is connected to the network 114 and communicates with other information processing devices via the network 114. The communication interface 107 may be a wired communication interface connected to a wired communication device such as a switch or router, or a wireless communication interface connected to a wireless communication device such as a base station or access point.
[0061] Figure 3 shows an example of a causal graph. As mentioned above, the causal relationships between variables included in a dataset are represented by a causal graph of a linear causal model. The linear causal model is expressed as X = BX + E, as shown in equation (1). The dataset X is represented as tabular data, with rows corresponding to samples and columns corresponding to variables.
[0062] Figure 3(A) shows the variable x for multiple samples. 1 , x 2 , x 3 An example of matrix 21 corresponding to dataset X having the variable values x is shown. Figure 3(B) shows an example of matrix 22 corresponding to adjacency matrix B. 1 , x 2 , x 3 In contrast, B is a 3x3 matrix. Element b of B ij is, x j →x i This represents the direct causal effect. Figure 3(C) shows the causal graph 23 corresponding to the adjacency matrix B. Note that the noise matrix E = (e 1 , e 2 , e 3 ) is an element of E. i For example, this is a uniformly random number of the range (-1, 1).
[0063] Here, the causal graph includes nodes corresponding to each variable and directed edges connecting pairs of nodes having a causal relationship. Variable x 1 affecting variable x 2 , that is, there is a causal relationship from variable x 1 to variable x 2 is indicated by a directed edge from the node of variable x 1 to the node of variable x 2 .
[0064] For example, causal graph 23 shows that variable x 1 affects variable x 2 , x 3 , and variable x 2 affects variable x 3 . The direct causal effect from variable x 1 to variable x 2 [[ID=二十九]] is 2.0. The direct causal effect from variable x 1 to variable x 3 is 0.5. The direct causal effect from variable x 2 to variable x 3 is -1.0.
[0065] Figure 4 is a diagram showing examples of a plurality of datasets. Information processing apparatus 100 performs integrated causal exploration on m different datasets X ~(1) , X ~(2) , …, X ~(m) . Here, “X ~ ” represents a character with a tilde symbol “~” attached above X.
[0066] The number of samples in each X ~(k) may be different. The samples included in each X ~(k) do not have to be the same. Also, each X ~(k) is associated with a variable set V (k) . V (k) and V (k’) do not have to match. However, a part of V (k) overlaps with a part of at least one other V (k’) .
[0067] One example of a dataset is a set of diagnostic records from a single hospital. For instance, the diagnostic records from one hospital constitute one dataset. For example, one hospital might measure a patient's blood pressure but not their temperature, while another hospital might measure their temperature but not their blood pressure. Furthermore, the patients themselves may differ from hospital to hospital.
[0068] Another example of a dataset is experimental records of chemical reactions. For instance, the experimental records of one researcher constitute one dataset. For example, one researcher might measure reaction energy but not reaction rate, while another researcher might measure reaction rate but not reaction energy. Furthermore, the number of experiments and experimental conditions may differ from one researcher to another.
[0069] Dataset group 30 is an example of multiple datasets. Dataset group 30 is a subset of dataset X. ~(1) , X ~(2) , X ~(3) It has. Dataset X ~(1) is the variable x 1 , x 2 , x 3 The dataset includes each of the values. ~(2) is the variable x 1 , x 3 The dataset includes each of the values. ~(3) is the variable x 2 , x 4 Includes each of the values.
[0070] Dataset X ~(1) , X ~(2) , X ~(3) The union of each set of variables is (x 1 , x 2 , x 3 , x 4 The information processing device 100 processes the dataset X. ~(1) , X ~(2) , X ~(3) For each of these, the union of each set of variables (x 1 , x 2 , x 3 , x 4 We will conduct an integrated causal search regarding the variable x at the bottom of Figure 4. 1 ~x 4 The graph connecting the points with a dotted line shows the variable x 1 ~x4 This represents an unknown causal graph for [the given condition].
[0071] In the causal search performed by the information processing device 100, the linear causal model is X in equation (9). (k) = (A * F (k) ) E (k) By using this method, it becomes possible to estimate the integrated causal relationships between variables across the entire dataset, as well as the individual causal effects within each dataset. Note that the asterisk symbol "*" above represents A and F. (k) The Hadamard product is shown.
[0072] Figure 5 shows an example of the functions of an information processing device. The information processing device 100 has a data storage unit 120, a calculation processing unit 130, and a result output unit 140. The data storage unit 120 uses the storage area of RAM 102 or HDD 103. The calculation processing unit 130 and the result output unit 140 are realized by the execution of a program stored in RAM 102 by the processor 101.
[0073] The data storage unit 120 stores the input dataset X ~(1) , X ~(2) , ..., X ~(m) and the combined causal effect matrix A and the multiple variable matrix F in equation (9) (1) , F (2) , ..., F (m) Remember this. As mentioned above, E (k) The noise value e included i (k) The noise is generated based on a predetermined distribution function. For example, a Gaussian mixture distribution G is used as the distribution function that generates the noise value. In this case, E (k) = (e 1 (k) , e 2 (k) , ..., e d (k) Each element of ) is a Gaussian mixture distribution G 1 (k) , G 2 (k) ..., G d (k) It is generated based on e i (k) is G i(k) It is generated based on the following. d is the total number of variables belonging to the union of the variable sets of each dataset. Therefore, the data storage unit 120 is G 1 (k) , G 2 (k) ..., G d (k) It also remembers that information.
[0074] A, F (1) , F (2) , ..., F (m) , G 1 (1) ..., G d (m) The initial values of each element are determined, for example, randomly. However, specific values may be given as the initial values for each element.
[0075] Furthermore, for example, in the case where the aforementioned hospital-specific diagnostic records are used as the input dataset, X ~(1) , X ~(2) , ..., X ~(m) For example, information (datasets) of m diagnostic records obtained from m hospitals are input to the information processing device 100 and stored in the data storage unit 120. In the example where the experimental records of the chemical reaction described above are used as the input dataset, X ~(1) , X ~(2) , ..., X ~(m) Information (datasets) of m experimental records obtained from m researchers are input into the information processing device 100 and stored in the data storage unit 120.
[0076] The calculation processing unit 130 calculates the pseudo-dataset X^ based on equation (9). (1) ,X^ (2) ,...,X^ (m) This generates the following: Here, "X^" represents the character X with a caret symbol "^" placed above it.
[0077] The calculation processing unit 130 processes the input dataset X ~(1) , ..., X ~(m) and pseudo-dataset X^ (1) ,...,X^ (m) Based on a predetermined loss function relating to A, each F is obtained by gradient descent. (k) , each G i(k) The loss function is optimized. For example, MMD or optimal transport distance may be used. In optimization, gradient descent is used to minimize the loss function, and A and F are each optimized. (k) , each G i (k) The values of the elements are estimated. For minimizing the sum of MMD, for example, equation (8) in Non-Patent Document 1 may be helpful.
[0078] The calculation processing unit 130 is A, each F (k) , each G i (k) Once the optimization is complete, X is calculated based on equation (11). ~(1) , ..., X ~(m) Adjacent matrix B corresponding to (1) , ..., B (m) The calculation processing unit 130 calculates B (1) , ..., B (m) The data is stored in the data storage unit 120.
[0079] The result output unit 140 outputs the calculation results from the calculation processing unit 130. For example, the result output unit 140 outputs B (1) , ..., B (m) Based on this, m causal graphs may be displayed on the display 111, or the information of the m causal graphs may be transmitted to other information processing devices via the network 114.
[0080] Figure 6 shows an example of the optimization process. Diagram 40 illustrates the optimization process performed by the calculation processing unit 130. The calculation processing unit 130 is G 1 (k) ..., G d (k) Using E (k) The calculation processing unit 130 generates E (1) , ..., E (m) A and F (1) , ..., F (m) Using and , the pseudo-dataset X^ is obtained by equation (9). (1) ,X^ (2) ,...,X^ (m) The calculation processing unit 130 generates the input dataset X. ~(1) , ..., X ~(m) and pseudo-dataset X^ (1) ,...,X^(m) Minimize the value of the loss function with respect to (G 1 (1) ..., G d (1) ), ..., (G 1 (m) ..., G d (m) ) and A and F (1) , ..., F (m) The process involves optimizing the following. This optimization is performed using gradient descent. Examples of gradient descent methods that can be used include stochastic gradient descent and batch gradient descent (steepest descent).
[0081] Here, the Gaussian mixture distribution G i (k) To elaborate, the calculation processing unit 130, through the above optimization, processes the dataset X ~(k) The variable x i As a model to approximate the noise distribution, a Gaussian mixture distribution G i (k) Learn about each G. i (k) It has the following parameters:
[0082] The first parameter is the number of Gaussian distributions g included in the Gaussian mixture. g is a user parameter. All G i (k) We can also use g as a common value. The larger g is, the better the approximation.
[0083] The second parameter is the mean vector μ i (k) = (μ i,1 (k) , ..., μ i,g (k) ) The third parameter is the dispersion vector ρ i (k) = (ρ i,1 (k) , ..., ρ i,g (k) )
[0084] The fourth parameter is the weight vector w i (k) = (w i,1 (k) ... lol i,g(k) ) is μ i (k) , ρ i (k) ,w i (k) These are the parameters that are adjusted through optimization.
[0085] Each G i (k) This generates random noise as a weighted sum of Gaussian distributions. Specifically, the calculation processing unit 130 generates a Gaussian distribution z for each t. i,t (k) ~N(μ) i,t (k) , ρ i,t (k) The calculation processing unit 130 generates random numbers (w'). The calculation processing unit 130 also transforms the weight vector using the Softmax function so that its sum is 1. The transformed weight vector (w' i,1 (k) ,..., w' i,g (k) )=Softmax(w i,1 (k) ... lol i,g (k) )
[0086] Then, the calculation processing unit 130 calculates the weighted sum v i (k) = Σ t w' i,t (k) z i,t (k) G i (k) Random value (noise value) e i (k) The calculation processing unit 130 calculates each G based on the gradient descent method. i (k) The gradient ∇G i (k) Therefore, each G i (k) The parameter values are updated. The calculation processing unit 130 calculates the gradient ∇G i (k) The gradient ∇μ with respect to each parameter is i (k) ,∇ρ i (k) ,∇w i(k) Each of these is calculated. Then, the calculation processing unit 130 calculates μ i (k) ←μ i (k) -δ∇μ i (k) , ρ i (k) ←ρ i (k) -δ∇ρ i (k) ,w i (k) ←w i (k) -δ∇w i (k) The update is performed for each parameter. Here, δ is the learning rate.
[0087] Note that the parameters to be adjusted in A are each a ij That is. Also, F (k) The parameters to be adjusted in this case are each f ij Figure 7 shows an example of data stored in the data storage unit.
[0088] The data storage unit 120 stores an input dataset group 150 and a variable group 160. The input dataset group 150 is a collection of input datasets. For example, the input dataset group 150 includes input datasets 151, 152, and 153. In one example, input dataset 151 is a diagnostic record from hospital (1). Input dataset 152 is a diagnostic record from hospital (2). Input dataset 153 is a diagnostic record from hospital (3).
[0089] Input dataset 151 contains the values of the variables "drug dosage," "blood glucose level," and "urination frequency" for multiple patients at hospital (1). Input dataset 152 contains the values of the variables "drug dosage" and "blood glucose level" for multiple patients at hospital (2). Input dataset 153 contains the values of the variables "drug dosage," "blood glucose level," and "blood pressure" for multiple patients at hospital (3). Input datasets 151, 152, and 153 are input dataset X. ~(1) , X ~(2) , X ~(3) This is one example.
[0090] The variable group 160 is a set of values for each parameter representing the linear causal model of equation (9). The values of each parameter in the variable group 160 are the parameter values to be adjusted by optimization in the calculation processing unit 130.
[0091] For example, for input datasets 150, the variable group 160 is (G 薬剤投与量 (1) , G 血糖値 (1) , G 尿頻度 (1) , G 血圧 (1) ), (G 薬剤投与量 (2) , G 血糖値 (2) , G 尿頻度 (2) , G 血圧 (2) ), (G 薬剤投与量 (3) , G 血糖値 (3) , G 尿頻度 (3) , G 血圧 (3) ) includes.
[0092] Furthermore, for the input dataset 150, the variable group 160 consists of a composite causal effect matrix A and a multiple variable matrix F (1) , F (2) , F (3) This includes the following. Next, we will describe the loop processing (optimization loop) in the gradient descent optimization performed by the calculation processing unit 130 on the input dataset group 150 and the variable group 160.
[0093] Figure 8 shows an example of pseudo-data generation in the optimization loop. First, the calculation processing unit 130 (G 薬剤投与量 (1) , G 血糖値 (1) , G 尿頻度 (1) , G 血圧 (1) ) and A and F (1) Based on this, a dummy dataset 161 is generated that corresponds to the input dataset 151 of hospital (1). The number of samples in dataset 161 is the same as the number of samples in the input dataset 151.
[0094] Furthermore, the calculation processing unit 130 is (G 薬剤投与量 (2) , G 血糖値 (2) , G 尿頻度 (2) , G 血圧 (2) ) and A and F (2) Based on this, a dummy dataset 162 is generated that corresponds to the input dataset 152 of hospital (2). The number of samples in dataset 162 is the same as the number of samples in the input dataset 152.
[0095] Furthermore, the calculation processing unit 130 is (G 薬剤投与量 (3) , G 血糖値 (3) , G 尿頻度 (3) , G 血圧 (3) ) and A and F (3) Based on this, a dummy dataset 163 is generated that corresponds to the input dataset 153 of hospital (3). The number of samples in dataset 163 is the same as the number of samples in the input dataset 153.
[0096] Datasets 161, 162, and 163 are pseudo-datasets X^ (1) ,X^ (2) ,X^ (3) This is one example. Hereafter, datasets 161, 162, and 163 will be referred to as pseudo-datasets 161, 162, and 163.
[0097] Figure 9 shows an example of loss calculation in an optimization loop. Figure 9 illustrates the case where MMD is used as the loss function. The calculation processing unit 130 uses an MMD function that takes the input dataset 151 and the pseudo-dataset 161 as inputs, and sets the function Loss(1) to represent the difference between the input dataset 151 and the pseudo-dataset 161. Loss(1) can also be said to be a function that shows the similarity between the input dataset 151 and the pseudo-dataset 161.
[0098] The calculation processing unit 130 uses an MMD function that takes the input dataset 152 and the pseudo-dataset 162 as inputs, and sets the function Loss(2) to represent the difference between the input dataset 152 and the pseudo-dataset 162. Loss(2) can also be said to be a function that shows the similarity between the input dataset 152 and the pseudo-dataset 162.
[0099] The calculation processing unit 130 uses an MMD function that takes the input dataset 153 and the pseudo-dataset 163 as inputs, and sets the function Loss(3) to represent the difference between the input dataset 153 and the pseudo-dataset 163. Loss(3) can also be said to be a function that shows the similarity between the input dataset 153 and the pseudo-dataset 163.
[0100] The calculation processing unit 130 can, for example, make the loss function L a function that represents the sum of Loss(1), Loss(2), and Loss(3), such as L = Loss(1) + Loss(2) + Loss(3).
[0101] The value of the loss function can be said to be a similarity score, indicating the degree of similarity between the input datasets and the pseudo-datasets. Alternatively, the value of the loss function can be said to be a deviation score, indicating the degree of divergence between the input datasets and the pseudo-datasets. A smaller similarity score (or deviation score) indicates a higher degree of similarity between the input datasets and the pseudo-datasets. The minimum value of the similarity score (or deviation score) is 0.
[0102] The calculation processing unit 130 may calculate the values of Loss(1), Loss(2), Loss(3), and L in the optimization loop and store these values as history in the data storage unit 120.
[0103] Figure 10 shows an example of gradient propagation in an optimization loop. Here, the loss function L is defined as each G i (k) , A, each F (k) The parameters to be adjusted are stored as variables. The calculation processing unit 130 calculates the gradient ∇L of the loss function L with respect to these parameters to be adjusted. ∇ is a vector differential operator with respect to each parameter to be adjusted.
[0104] For example, among the components of ∇L, the vector that has a component corresponding to the parameter to be adjusted included in A is conveniently denoted as ∇A. i (k) , F (k) Similarly, for ∇L, G i (k) , F (k) For convenience, the vector with the corresponding component is called ∇G i (k) ,∇F (k) It can be written as shown. The explanation in Figure 6 also uses this convenient notation for the gradient. The calculation processing unit 130 processes each G i (k) , A, each F (k) For the current value of each ∇G i (k) , ∇A, each ∇F (k) Obtain gradient information 170 that includes the value.
[0105] Then, the calculation processing unit 130 updates each of the parameters to be adjusted in the variable group 160 based on the gradient information 170. For example, G 薬剤投与量 (1) μ 薬剤投与量 (1) It includes ∇G 薬剤投与量 (1) is, ∇μ 薬剤投与量 (1) This includes. In this case, the calculation processing unit 130 uses the learning rate δ to calculate μ 薬剤投与量 (1) ←μ 薬剤投与量 (1) -δ∇μ 薬剤投与量 (1) The calculation processing unit 130 updates as follows: A, F (k) Similarly, the value of each parameter to be adjusted is updated.
[0106] The calculation processing unit 130 repeatedly executes the series of steps illustrated in Figures 8 to 10 (optimization loop). Then, when the predetermined termination conditions are met, the calculation processing unit 130 terminates the optimization loop and each G i (k) , A, each F (k) The calculation processing unit 130 determines each G after optimization. i (k) , A, each F(k) The data is stored in the data storage unit 120.
[0107] Figure 11 shows an example of the output result. The calculation processing unit 130 calculates F for A after optimization. (1) , F (2) , F (3) Apply each of these (step ST1). Specifically, the calculation processing unit 130 calculates based on equation (10), A (1) , A (2) , A (3) Calculate (Step ST2).
[0108] The calculation processing unit 130 calculates B corresponding to hospital (1), hospital (2), and hospital (3) based on formula (11). (1) , B (2) , B (3) The result output unit 140 is B (1) Based on this, a causal graph 181 corresponding to hospital (1) is generated and displayed on the display 111. The result output unit 140 is B (2) Based on this, a causal graph 182 corresponding to hospital (2) is generated and displayed on the display 111. The result output unit 140 is B (3) Based on this, a causal graph 183 corresponding to hospital (3) is generated and displayed on the display 111. The result output unit 140 may transmit the information of causal graphs 181, 182, and 183 to other information processing devices via the network 114, and have the causal graphs 181, 182, and 183 displayed on the displays of the other information processing devices.
[0109] Next, the processing procedure of the information processing device 100 will be described. Figure 12 is a flowchart showing an example of processing by the information processing device. (S10) The calculation processing unit 130 processes the input dataset X ~(1) , X ~(2) , ..., X ~(m) Obtain it.
[0110] (S11) The calculation processing unit 130 processes each variable G i (k) , A, F (k) Initialize the following: i = 1, ..., d. k = 1, ..., m. For example, the calculation processing unit 130 initializes each G i(k) , A, F (k) The initial value is determined randomly.
[0111] (S12) The calculation processing unit 130 determines whether to complete the repetition of steps S13 to S16 below. If the repetition is to be completed, the process proceeds to step S17. If the repetition is not to be completed, the process proceeds to step S13. For example, the calculation processing unit 130 counts the number of repetitions of steps S13 to S16 and determines that the repetition is complete when the number of repetitions reaches a certain number. Also, the calculation processing unit 130 determines that the repetition is not to be completed if the number of repetitions has not reached a certain number. However, the termination condition for the optimization loop may be other conditions. For example, the calculation processing unit 130 may use each value of Loss(1) to Loss(m) or L = Σ k The system may determine whether the value of Loss(k) has become smaller than the reference value, and if each of these values or the value of L has become smaller than the reference value, it may be determined that the iteration is complete.
[0112] (S13) The calculation processing unit 130 controls each G i (k) , A, each F (k) Using this, based on equation (9), the pseudo-dataset X^ (1) ,X^ (2) ,...,X^ (m) (S14) The calculation processing unit 130 generates (X ~(k) ,X^ (k) The MaximumMeanDiscrepancy function (MMD function), which takes the input ) as input, is obtained as Loss(k). Note that the function used as Loss(k) may be other functions such as optimal transport distance.
[0113] (S15) The calculation processing unit 130 calculates each gradient ∇G from Loss(1) to Loss(m). i (k) ,∇A,∇F (k) The calculation is performed as follows: i = 1, ..., d. k = 1, ..., m. The calculation processing unit 130 calculates each G i (k) , A, each F (k) For each parameter to be adjusted, each gradient ∇G i (k),∇A,∇F (k) The values of the elements can be calculated.
[0114] For example, the calculation processing unit 130 calculates the loss function L = Σ k It can be expressed as Loss(k). As mentioned above, ∇G i (k) G is one of the components of ∇L i (k) This is a vector with a component corresponding to . ∇A is a vector with a component corresponding to A among all the components of ∇L. ∇F (k) F is one of the total components of ∇L (k) This is a vector with components corresponding to [the specified value].
[0115] (S16) The calculation processing unit 130 processes each variable G i (k) , A, F (k) Each gradient ∇G i (k) ,∇A,∇F (k) The update is performed based on the following: i = 1, ..., d. k = 1, ..., m. Then the process proceeds to step S12. Here, the calculation processing unit 130, as described above, sets the learning rate to δ and updates the value of p such that p ← p - δ∂L / ∂p for the change in the loss function L ∂L / ∂p corresponding to the parameter component p.
[0116] (S17) The calculation processing unit 130 calculates A and F after optimization. (k) Using this, matrix A is obtained from equation (10). (k) The calculation processing unit 130 calculates A (k) Using equation (11), B (k) We calculate the following, where k = 1, ..., m.
[0117] (S18) The result output unit 140 outputs the calculation result B (1) , B (2) , ..., B (m) The output unit 140 outputs the result B (k) Based on this, an image showing a causal graph for each k-unit may be output. Then, the processing of the information processing device 100 ends.
[0118] Thus, the information processing device 100 can estimate the integrated causal relationships in multiple datasets and the individual causal effects in each dataset. Figure 13 shows a comparative example.
[0119] The comparative example illustrates the case where the linear causal model is represented by equation (5) X = AE. The optimization process of the comparative example is represented by diagram 50. Diagram 50 is a multiplicative matrix F (1) , F (2) , ..., F (m) Diagram 40 differs from the comparative example in that it does not include [a specific element]. Therefore, the comparative example only yields a common causal effect (matrix A) for all datasets.
[0120] Figure 14 shows an example of a causal graph that cannot be distinguished in the comparative example. The comparative example method assumes the same linear causal model X = AE for all datasets. This assumption has the problem of not being able to consider the possibility that causal effects may differ for each dataset. For example, the effect of a drug may differ in hospitals where the patient trends regarding the administration of a certain drug differ.
[0121] Causal graphs 61, 62, and 63 show the inherent causal effects between variables in the data of diabetic patients at three hospitals (1), (2), and (3). Variable x 1 This represents the drug dosage. Variable x 2 This represents blood glucose level. Variable x 3 This is the frequency of urination. As shown in causal graphs 61, 62, and 63, in the three hospitals (1), (2), and (3), the variable x 1 , x 2 , x 3 The causal effect differs depending on the patient's tendencies. However, the comparative example method only yields a common causal effect across all datasets and cannot obtain different causal effects for each hospital. In other words, the comparative example method cannot distinguish between causal graphs 61, 62, and 63.
[0122] Other examples can be considered where causal effects differ for each dataset. For instance, if the laboratory is located in a different region, the results of experiments using water may differ due to differences in water quality. Therefore, the information processing device 100 expresses the linear causal model using equation (9), thereby showing the integrated causal relationship (A) across multiple datasets and the individual causal effects (A and F) in each dataset. (k) The Hadamard product of ( ) can be estimated. For this reason, the information processing device 100 calculates the matrix B representing the direct causal effect for hospitals (1), (2), and (3) in Figure 14. (1) , B (2) , B (3) It is possible to distinguish and calculate B. (1) , B (2) , B (3) It is possible to distinguish between the corresponding causal graphs 61, 62, and 63. In this way, the information processing device 100 can obtain detailed causal graphs that reflect the respective environments in which each input dataset was acquired. By distinguishing and presenting the causal graphs that reflect each environment to the user, the information processing device 100 can support detailed data analysis by the user.
[0123] As described above, the information processing device 100 performs the following processing, for example: The data storage unit 120 stores a plurality of first data. Each of the plurality of first data contains a plurality of variable values corresponding to a plurality of variables and is used to estimate causal relationships and causal effects between variables based on the plurality of variable values. The calculation processing unit 130 sets first information regarding the integrated causal relationships and causal effects in the plurality of first data as a whole, and second information regarding the weighting of the causal effects included in the first information for each of the plurality of first data. Based on the first information and the second information, the calculation processing unit 130 generates a plurality of second data corresponding to the plurality of first data. The calculation processing unit 130 optimizes the first information and the second information to reduce the value of the loss function for the plurality of first data and the plurality of second data.
[0124] As a result, the information processing device 100 can estimate the integrated causal relationships between variables in multiple datasets (first data) and the individual causal effects in each dataset (first data). Here, input dataset X ~(1) , ..., X ~(m) This is an example of multiple first data sets. Multiple pseudo-datasets (X^ (1) ,...,X^ (m) ) is an example of multiple second data points. The combined causal effect matrix A is an example of the first information. The multiplicative variable matrix F (1) , ..., F (m) This is an example of the second type of information. Also, the sum of the function Loss Σ, which indicates the similarity (or deviation) between data such as MMD and optimal transport distance. k=1 m Loss(k) is an example of a loss function.
[0125] More specifically, the first information includes a composite causal effect matrix whose elements are multiple first parameters that show the causal effects between variables belonging to the union of multiple variables in each of the multiple first data. The second information is a multiplicative variable matrix whose elements are multiple second parameters that weight the multiple first parameters, and includes a multiplicative variable matrix corresponding to each of the multiple first data. This allows the information processing device 100 to estimate the individual causal effects for each of the first data in detail. Note that the multiplicative variable matrix F (k) The second parameter f included in ij This is the first parameter a, which represents the causal effect between variables in the composite causal effect matrix. ij This can be said to represent the weight in the k-th input dataset. Also, the first parameter a ij This can also be said to represent an integrated causal relationship between variables.
[0126] Furthermore, the calculation processing unit 130 can use gradient descent to determine the values of each of the multiple first parameters included in the overall causal effect matrix and each of the multiple second parameters included in each multiplicative variable matrix when optimizing the first and second information. This allows the information processing device 100 to efficiently determine the values of each parameter to be adjusted in the first and second information. ijThis is an example of a parameter to be adjusted in the first information.
[0127] For example, multiple variables contained in one of the multiple first data sets may overlap with multiple variables contained in another of the multiple first data sets, but may not overlap in other parts. Therefore, the information processing device 100 can estimate the integrated causal effect across multiple datasets (first data) and the individual causal effect in each dataset (first data) for the union of each set of variables in the multiple first data sets.
[0128] The result output unit 140 outputs information showing the causal effects between variables in each of the multiple first data sets, based on the optimized first and second information. This allows the information processing device 100 to support the user in understanding the individual causal effects in each of the multiple first data sets. For example, the result output unit 140 may output an image of a causal graph as information showing the causal effects. The user can then understand the causal relationships between variables and the individual causal effects of the first data sets from the causal graph.
[0129] The calculation processing unit 130 generates multiple second data based on the first information, the second information, and noise information based on a predetermined distribution function. In the optimization process, the calculation processing unit 130 optimizes the distribution function along with the first and second information. This allows the calculation processing unit 130 to appropriately optimize the first and second information. Here, the noise matrix E (1) , ..., E (m) This is an example of noise information. Gaussian mixture distribution G i (k) This is an example of a distribution function used to generate noise information.
[0130] The information processing in the first embodiment can be achieved by having the processing unit 12 execute a program. The information processing in the second embodiment can be achieved by having the processor 101 execute a program. The program can be recorded on a computer-readable recording medium 113.
[0131] For example, a program can be distributed by distributing a recording medium 113 on which the program is stored. Alternatively, the program may be stored on another computer and distributed via a network. A computer may, for example, store (install) a program stored on the recording medium 113 or a program received from another computer into a storage device such as RAM 102 or HDD 103, and then read and execute the program from that storage device.
[0132] The above merely illustrates the principle of the present invention. Furthermore, numerous modifications and changes are possible for those skilled in the art, and the present invention is not limited to the exact configurations and applications shown and described above. All corresponding modifications and equivalents are considered to be within the scope of the present invention as defined by the appended claims and their equivalents.
[0133] 10 Information processing device 11 Storage unit 12 Processing unit
Claims
1. A program that causes a computer to perform the following operations: acquire a plurality of first data sets, each containing a plurality of variable values corresponding to a plurality of variables, which are used to estimate causal relationships and causal effects between variables based on the plurality of variable values; set first information relating to the integrated causal relationships and causal effects in the plurality of first data sets as a whole, and second information relating to the weighting of the causal effects included in the first information for each of the plurality of first data sets; generate a plurality of second data sets corresponding to the plurality of first data sets based on the first information and the second information; and optimize the first information and the second information so as to reduce the value of the loss function relating to the plurality of first data sets and the plurality of second data sets.
2. The program according to claim 1, wherein the first information includes a composite causal effect matrix whose elements are a plurality of first parameters that show the causal effects between variables belonging to the union of the plurality of variables in each of the plurality of first data, and the second information includes a multiplicative variable matrix whose elements are a plurality of second parameters that weight the plurality of first parameters, and which includes the multiplicative variable matrix corresponding to each of the plurality of first data.
3. The program according to claim 2, which causes the computer to perform a process in which it determines the value of each of the plurality of first parameters and the value of each of the plurality of second parameters using the gradient descent method in the optimization.
4. The program according to claim 1, wherein the plurality of variables included in any of the plurality of first data and the plurality of variables included in the other of the plurality of first data overlap in part, but do not overlap in the remaining parts.
5. The program according to claim 1, which causes the computer to perform a process that outputs information indicating the causal effect in each of the plurality of first data based on the first information and second information after optimization.
6. The program according to claim 1, which causes the computer to perform the following processes: in generating the plurality of second data, generate the plurality of second data based on the first information, the second information, and noise information based on a predetermined distribution function; and in the optimization, perform the optimization of the distribution function together with the first information and the second information.
7. An information processing method comprising: a computer acquiring a plurality of first data sets, each containing a plurality of variable values corresponding to a plurality of variables, which are used to estimate causal relationships and causal effects between variables based on the plurality of variable values; setting first information relating to the integrated causal relationships and causal effects in the plurality of first data sets as a whole, and second information relating to the weighting of the causal effects included in the first information for each of the plurality of first data sets; generating a plurality of second data sets corresponding to the plurality of first data sets based on the first information and the second information; and optimizing the first information and the second information to reduce the value of the loss function relating to the plurality of first data sets and the plurality of second data sets.
8. Information processing device comprising: a storage unit that stores a plurality of first data sets, each containing a plurality of variable values corresponding to a plurality of variables, which are used to estimate causal relationships and causal effects between variables based on the plurality of variable values; a processing unit that sets first information relating to the integrated causal relationships and causal effects in the plurality of first data sets as a whole, and second information relating to the weighting of the causal effects included in the first information for each of the plurality of first data sets, generates a plurality of second data sets corresponding to the plurality of first data sets based on the first information and the second information, and optimizes the first information and the second information so as to reduce the value of the loss function relating to the plurality of first data sets and the plurality of second data sets.
Citation Information
Patent Citations
Information processing method, information processing device and program
JP2022013844A
Medical learning device, medical learning method, and medical information processing system
JP2024066412A