Feature selection method, relevance assessment method, feature selection device, relevance assessment device, and program
The feature selection method addresses the limitation of conventional techniques by evaluating conditional distributions and performing statistical tests to identify features affecting treatment effect variance, enhancing the accuracy of treatment effect assessments.
Patent Information
- Application Number
- JP2024521498
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-05-19
- Publication Date
- 2025-10-01
- Estimated Expiration
- 2042-05-19
AI Technical Summary
Conventional feature selection techniques based on the mean of treatment effects fail to account for features affecting distribution parameters like variance, leading to erroneous conclusions about treatment effect heterogeneity.
A feature selection method that estimates the conditional distributions of outcomes with and without treatment, using distribution distance measures and propensity scores to evaluate feature relevance, followed by a statistical hypothesis testing procedure to select features related to the magnitude of treatment effects.
Accurately identifies features that impact the variance of treatment effects, improving the reliability of treatment effect assessments by correctly selecting features that correlate with the magnitude of treatment effects.
Smart Images

Figure 0007747191000022 
Figure 0007747191000023 
Figure 0007747191000024
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a feature selection method, a relevance assessment method, a feature selection device, a relevance assessment device, and a program. [Background technology]
[0002] Treatment effect (also known as causal effect) is a measure of the effect of some treatment (e.g., drug administration, participation in an educational program, etc.) on an outcome (e.g., cholesterol levels in the case of drug administration, grades in the case of participation in an educational program, etc.) when administered to an individual. Treatment effect often varies depending on the characteristics possessed by each individual. For example, recent medical research has investigated how the effect of administering a vaccine against the novel coronavirus on the immune system varies depending on individual characteristics such as age, gender, and race (Non-Patent Document 1).
[0003] Such a treatment effect is defined as the difference between the potential outcome with the treatment and the potential outcome without the treatment, but it is impossible to calculate the treatment effect for each individual because when an individual receives a treatment, only the potential outcome with the treatment is observed, not the potential outcome without the treatment.
[0004] For this reason, conventional techniques evaluate the treatment effect using some kind of evaluation index. For example, an evaluation index called the conditional average treatment effect (CATE), which can be estimated from observed data, is used. This is defined as the average amount of treatment effect for multiple individuals who have the same attribute with respect to a certain feature. Conventional feature selection techniques select features that have a significant impact on the magnitude of the average treatment effect by fitting a linear regression model to the observed data regarding the treatment variable (a variable that takes two values indicating whether or not the treatment has been administered) based on this conditional average treatment effect (Non-Patent Document 2). [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] Kamal Abu Jabal, Hila Ben-Amram, Karine Beiruti, Yunis Batheesh, Christian Sussan, Salman Zarka, and Michael Edelstein. "Impact of age, ethnicity, sex and prior infection status on immunogenicity following a single dose of the BNT162b2 mRNA COVID-19 vaccine: Realworld evidence from healthcare workers, Israel, December 2020 to January 2021". Eurosurveillance, 26(6), 2021. [Non-patent document 2] Qingyuan Zhao, Dylan S. Small, and Ashkan Ertefaie. "Selective inference for effect modification via the lasso". arXiv preprint arXiv:1705.08020, 2017. Summary of the Invention [Problem to be solved by the invention]
[0006] However, in conventional feature selection techniques based on the mean of treatment effects, if a feature affects a distribution parameter such as the variance of treatment exchange, the feature is treated as being unrelated to the magnitude of the treatment effect, which can lead to erroneous conclusions about the heterogeneity of treatment effects.
[0007] The present disclosure has been made in view of the above points, and provides a technique for selecting features that are related to the magnitude of treatment effects. [Means for solving the problem]
[0008] The feature selection device according to one aspect of the present disclosure selects an observed value a of a treatment variable A, which indicates whether a predetermined treatment has been administered to an individual i. i and the feature variable X=[X1, ,X d ] T (where d is the number of features) observed value x i and the observed value y of the outcome variable Y representing a predetermined outcome for the individual i. i The observed data D={(a i ,x i ,y i )} to estimate the potential outcome Y when the treatment is administered. 1 Features X m The conditional distribution P(Y 1 |X m = x) and the potential outcome Y if the treatment is not performed. 0 Features X m The conditional distribution P(Y 0 |X m = x) is the distribution distance between the feature X m The evaluation scale I shows how much the attribute value x of m Estimator of ~ I m an evaluation procedure for calculating the observed data D and the evaluation scale I m (m=1, , d) and a selection procedure to select features that are related to the magnitude of the treatment effect. [Effects of the Invention]
[0009] Techniques are provided for selecting features that correlate with the magnitude of treatment effect. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 10 is a diagram illustrating an example of algorithm 1 (CRT). [Figure 2] FIG. 1 shows an example of Algorithm 2 (Feature Selection for Distributional Treatment Effect Modifier Discovery). [Figure 3] FIG. 1 is a diagram illustrating an example of a hardware configuration of a feature selection device according to an embodiment of the present invention. [Figure 4] FIG. 1 is a diagram illustrating an example of the functional configuration of a feature selection device according to an embodiment of the present invention. [Figure 5] 10 is a flowchart illustrating an example of feature selection processing according to the present embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] An embodiment of the present invention will now be described. In the following embodiment, a feature selection device 10 capable of selecting features related to the magnitude of treatment effects will be described. To achieve this, the feature selection device 10 according to this embodiment (a) evaluates the relevance of individual features to the magnitude of treatment effects, and then (b) selects features based on this evaluation.
[0012] <Theoretical structure> The theoretical configurations of (a) and (b) above will be explained below.
[0013] <Evaluation of the relevance of individual characteristics to the magnitude of treatment effects> First, the theoretical structure of (a) above will be explained.
[0014] Let A∈{0,1} be a binary variable (treatment variable) that indicates whether or not a treatment was administered, with A=1 if the treatment was administered and A=0 if the treatment was not administered. Examples of treatments include administering medication and participating in an educational program.
[0015] Let us define the random variables (feature variables) that represent the d features that each individual has.
[0016]
number
[0017] In this embodiment, observed data consisting of observed values (a, x, y) of (A, X, Y) for n individuals is
[0018]
number
[0019] Given the variables A, X, and Y above, the potential outcome of a treatment is Y. 1 , the potential outcome of no treatment is Y 0 Then, the treatment effect for each individual is the difference between the potential outcomes Y 1 -Y 0 On the other hand, the outcome variable Y is defined as Y=(1-A)Y depending on the value of the treatment variable A. 0 +AY 1 Therefore, the outcome variable Y is defined as the potential outcome Y 0 or potential outcome Y 1 It is not possible to calculate the treatment effect, which is the difference in potential outcomes, for each individual.
[0020] Therefore, in this embodiment, an evaluation scale (evaluation index) that can be estimated from the observation data D and that represents the correlation between the individual characteristics and the magnitude of the treatment effect is formulated. m For (m∈{1, ,d}), the potential outcome Y when the action is taken 1 For each feature X m The conditional distribution P(Y 1 |Xm = x) and the potential outcome Y if no treatment is administered. 0 For each feature X m The conditional distribution P(Y 0 |X m = x) is considered. For this distance, we use the existing distribution distance measure MMD (maximum mean discrepancy, Reference 1).
[0021] The inter-distribution distance measure is a positive definite kernel function k Y : Using R × R → R, it is defined as the following formula (1).
[0022]
number
[0023]
number
[0024] Based on the above inter-distribution distance measures, we define the features related to the distribution parameters of the treatment effects as follows:
[0025] (Definition) Feature X m is called a distributional treatment effect modifier if it satisfies the following conditions:
[0026] Condition: Feature X m At least two of the values of x m and x m ★ (However, x m ≠x m ★ ) for which the distribution distance measure D m2 take different values, i.e., the following holds:
[0027]
number
[0028] In order to detect such fluctuations in the distribution distance measure, the variance of the distribution distance measure as shown in the following equation (2) is used as a new feature relevance evaluation measure.
[0029]
number
[0030]
number
[0031] The above propensity score e(X) is estimated by fitting a conditional distribution model, such as a logistic regression model, to the observed data D in advance.
[0032] Feature X m When takes a discrete value, the weight for individual i∈{1, ,n} is formulated as shown in the following equation (4) using the weight function of equation (3).
[0033]
number
[0034]
number
[0035]
number
[0036]
number
[0037]
number
[0038] Based on equation (6), the feature value X m = the inter-distribution distance measure D for x m 2 (x) is calculated as follows:
[0039]
number
[0040]
number
[0041]
number
[0042]
number
[0043] Feature selection based on relevance assessment Next, the theoretical configuration of (b) above will be explained.
[0044] In order to select features that are related to the magnitude of the treatment effect based on the evaluation scale estimator in equation (8), we perform a multiple test dealing with the null and alternative hypotheses as shown in equation (9) below for the evaluation scale in equation (2).
[0045]
number
[0046] Each feature X m For (m∈{1, ,d}), the evaluation measure I in Eq. (2) is used to determine whether to reject the null hypothesis. m is the test statistic, and the feature X m The test statistic I when follows the null hypothesis m The distribution of the test statistic is approximated and a p-value is calculated for the evaluation scale estimated by equation (8). To this end, in this embodiment, the distribution of the test statistic is approximated using the conditional randomization test (CRT) described in Reference 4. In the CRT, X in X m Features other than X -m =X\X m Then, the conditional distribution P(X m |X -m ), features that conform to the null hypothesis (i.e., in this embodiment, features that are not related to the magnitude of the treatment effect) are artificially generated, and the distribution of the test statistic is approximated based on these.
[0047] Algorithm 1, which is a procedure for calculating an approximate p-value based on CRT, is shown in Figure 1. As shown in Figure 1, Algorithm 1 takes the observed data D and the estimated value of the test statistic as input and outputs the p-value. First, by fitting the probabilistic generative model L to the observed data D, the conditional distribution P(X m |X -m) is estimated (line 1). Here, various conditional distribution models proposed in the prior art can be used as the probabilistic generative model L. Next, the following two steps are repeated for each b = 1, , B (line 2).
[0048] Step 1: For i=1, , n, we define the artificial feature as x m,i (b) ~L(X m |x -m,i ) and sample as x m,i (b) ∪x -m,i x i (b) (lines 3 to 6).
[0049] Step 2: Data {(a i ,x i (b) ,y i )|i=1, ,n} to obtain the test statistic (estimate) ~ I m (b) Calculate (line 7).
[0050] Then, after repeating the above steps 1 and 2, the p-value is approximately calculated using the following equation (10) (line 9), and the p-value is returned to the caller of algorithm 1 (line 10).
[0051]
number
[0052] In this embodiment, d X1,...,X d By performing the calculation of equation (10) for each of the d p-values ^p1, ,^p dIt is known that in a multiple testing problem where statistical hypothesis testing is performed multiple times using multiple p-values, the probability of erroneously rejecting the null hypothesis increases. Therefore, in this embodiment, the p-value is corrected using the p-value correction technique described in Reference 5, and the null hypothesis is rejected only when this corrected p-value is equal to or less than a predetermined significance level α, and the feature X m We conclude that has a significant impact on the magnitude of the distribution parameter of the treatment effect. In other words, the set of features whose adjusted p-values are less than the significance level α is the set of distribution treatment effect modifiers.
[0053] Feature Selection for Discovering Distributional Treatment Effect Modifiers Figure 2 shows Algorithm 2, a feature selection procedure for discovering distributional treatment effect modifiers, combining (a) and (b) above. As shown in Figure 2, Algorithm 2 takes the observed data D and the significance level α as input and outputs a set of distributional treatment effect modifiers (feature set). The features included in this feature set are those that have a significant impact on the magnitude of the distribution parameter of the treatment effect. First, for m = 1, , d, the test statistic (estimate) is calculated using Equation (8) using the observed data D, and then the p-value is calculated using Algorithm 1 (lines 1 to 4). Next, these d p-values are corrected using a p-value correction technique (line 5). The set of features corresponding to the corrected p-values below the significance level α is defined as a set of distributional treatment effect modifiers (line 6), and this set is returned to the caller of Algorithm 2 (line 7).
[0054] <Example of Hardware Configuration of Feature Selection Device 10> An example of the hardware configuration of a feature selection device 10 according to this embodiment is shown in Fig. 3. As shown in Fig. 3, the feature selection device 10 according to this embodiment includes an input device 101, a display device 102, an external I / F 103, a communication I / F 104, a RAM (Random Access Memory) 105, a ROM (Read Only Memory) 106, an auxiliary storage device 107, and a processor 108. Each of these pieces of hardware is connected to each other via a bus 109 so as to be able to communicate with each other.
[0055] The input device 101 is, for example, a keyboard, a mouse, a touch panel, a physical button, etc. The display device 102 is, for example, a display, a display panel, etc. Note that the feature selection device 10 does not necessarily have to include at least one of the input device 101 and the display device 102, for example.
[0056] The external I / F 103 is an interface with an external device such as a recording medium 103a. The feature selection device 10 can read from and write to the recording medium 103a via the external I / F 103. Examples of the recording medium 103a include a flexible disk, a CD (Compact Disc), a DVD (Digital Versatile Disk), an SD memory card (Secure Digital memory card), and a USB (Universal Serial Bus) memory card.
[0057] The communication I / F 104 is an interface for connecting the feature selection device 10 to a communication network. The RAM 105 is a volatile semiconductor memory (storage device) that temporarily stores programs and data. The ROM 106 is a non-volatile semiconductor memory (storage device) that can store programs and data even when the power is turned off. The auxiliary storage device 107 is a storage device (storage device) such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), or a flash memory. The processor 108 is an arithmetic device such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit).
[0058] The feature selection device 10 according to this embodiment has the hardware configuration shown in Fig. 3, and is therefore capable of implementing the feature selection process described below. Note that the hardware configuration shown in Fig. 3 is merely an example, and the hardware configuration of the feature selection device 10 is not limited to this. For example, the feature selection device 10 may have multiple auxiliary storage devices 107 or multiple processors 108, may not have some of the hardware shown in the figure, or may have various hardware components other than the hardware shown in the figure.
[0059] <Example of functional configuration of feature selection device 10> An example of the functional configuration of a feature selection device 10 according to this embodiment is shown in Fig. 4. As shown in Fig. 4, the feature selection device 10 according to this embodiment includes an input unit 201, a relevance assessment unit 202, a feature selection unit 203, and an output unit 204. The feature selection device 10 according to this embodiment also includes an observation data DB 205, a first parameter DB 206, and a second parameter DB 207.
[0060] The input unit 201, the relevance assessment unit 202, the feature selection unit 203, and the output unit 204 are realized, for example, by processing in which one or more programs installed in the feature selection device 10 are executed by the processor 108. Note that all or part of these one or more programs may be stored in advance in the auxiliary storage device 107, or may be downloaded from a predetermined server or the like via a communication network.
[0061] The observation data DB 205, the first parameter DB 206, and the second parameter DB 207 are realized, for example, by the auxiliary storage device 107. Note that all or part of these DBs may be realized, for example, by a database server connected to the feature selection device 10 via a communication network.
[0062] The input unit 201 receives input of various hyperparameters and observation data D. The input unit 201 stores the observation data D in an observation data DB 205, and also stores hyperparameters used to evaluate the relevance of individual features to the magnitude of treatment effect in a first parameter DB 206 and hyperparameters used to select features based on the evaluation of relevance in a second parameter DB 207. Here, examples of the hyperparameters used to evaluate the relevance of individual features to the magnitude of treatment effect include the hyperparameter r of the approximation algorithm RFFs. Examples of the hyperparameters used to select features based on the evaluation of relevance include the significance level α and hyperparameters for training the probabilistic generative model L (i.e., adapting it to the observation data D). These hyperparameters and observation data D may be input via the input device 101, or may be input or acquired from a terminal, a server, or the like connected to the feature selection device 10 via a communication network.
[0063] The relevance assessment unit 202 assesses the relevance of individual characteristics to the magnitude of the treatment effect using the observed data D stored in the observed data DB 205 and the hyperparameters stored in the first parameter DB 206. That is, the relevance assessment unit 202 uses the observed data D and the hyperparameter r to calculate the evaluation scale estimator of Equation (8) for m=1, , d. ~ I m This corresponds to the second line of Algorithm 2 shown in Figure 2.
[0064] The feature selection unit 203 selects features based on the evaluation of relevance using the observation data D stored in the observation data DB 205 and the hyperparameters stored in the second parameter DB 207. That is, the feature selection unit 203 selects p-values ^p1, , ^p by Algorithm 1 shown in FIG. dAfter calculating these p-values, these p-values are corrected using p-value correction techniques, and the set of features corresponding to the corrected p-values below the significance level α is estimated as the set of distribution treatment effect modifiers ^S. This corresponds to lines 3 and 5-6 of Algorithm 2 in Figure 2.
[0065] Here, the seventh line of Algorithm 1 shown in Figure 1 calculates the test statistic by equation (8). ~ I m (b) is calculated, this seventh line is executed by the relevance assessment unit 202. However, this seventh line may be executed by the feature selection unit 203, or, for example, lines 1 to 8 may be executed by the relevance assessment unit 202 and lines 9 to 10 may be executed by the feature selection unit 203.
[0066] The output unit 204 outputs the feature set ^S (a set of distribution treatment effect modifiers) estimated by the feature selection unit 203 to a predetermined output destination. Note that the predetermined output destination is not particularly limited, and examples thereof include the display device 102, the auxiliary storage device 107, and other terminals or servers connected via a communication network.
[0067] The observed data DB 205 stores observed data D. The first parameter DB 206 stores hyperparameters used to evaluate the relevance of individual characteristics to the magnitude of treatment effect. The second parameter DB 207 stores hyperparameters used to select features based on the evaluation of relevance.
[0068] 4 is an example, and the functional configuration of the feature selection device 10 is not limited to this. For example, the feature selection device 10 may not have the relevance assessment unit 202 or the first parameter DB 206, and another device connected to the feature selection device 10 via a communication network may have the relevance assessment unit 202 and the first parameter DB 206 (and in addition to them, the observation data DB 205, etc.). In this case, the feature selection device 10 may use the evaluation scale estimator transmitted from the other device (this other device may be called, for example, a "relevance assessment device") as the evaluation scale estimator.~ I m (m=1,···,d) to estimate the set of distributional treatment effect modifiers ^S.
[0069] <Feature selection processing> An example of feature selection processing according to this embodiment will be described below with reference to Fig. 5. In the following, it is assumed that observed data D is stored in the observed data DB 205, and that necessary hyperparameters are stored in the first parameter DB 206 and the second parameter DB 207, respectively.
[0070] First, the relevance assessment unit 202 calculates the evaluation scale estimator of Equation (8) using the observed data D stored in the observed data DB 205 and the hyperparameters stored in the first parameter DB 206. ~ I m (m=1, . . . , d) is calculated (step S101).
[0071] Next, the feature selection unit 203 uses the observed data D stored in the observed data DB 205 and the hyperparameters stored in the second parameter DB 207 to select the p-value ^p m After calculating (m=1,···,d), these p-values are corrected using p-value correction technology, and the set of features corresponding to the corrected p-values below the significance level α is estimated as the set of distribution treatment effect modifiers ^S (step S102).
[0072] Then, the output unit 204 outputs the set ^S of distribution treatment effect modifying factors to a predetermined output destination (step S103).
[0073] In the feature selection process shown in FIG. 5, the evaluation scale estimator ~ I m After calculating (m=1, , d), in step S102, the p-value ^p mAlthough the set of distributional treatment effect modifiers ^S was estimated by calculating and correcting (m = 1,...,d), this is not limited to this. For example, the set of distributional treatment effect modifiers ^S may be estimated by Algorithm 2 shown in Figure 2. That is, for m = 1,...,d, the evaluation scale estimator (estimated value of the test statistic) ~ I m and p-value^p m After iteratively calculating , the set of distributional treatment effect modifiers ^S may be estimated by correcting the p-values.
[0074] <Experiment> Below, an experiment conducted to evaluate the feature selection device 10 according to this embodiment will be described.
[0075] In this experiment, we artificially generated a dataset containing features that influence the variance of the treatment effect. This dataset was generated as follows:
[0076] The treatment variable A and the characteristic variable X were sampled from the following distributions.
[0077] A~Ber(0.5) X|A=0~N(-μ,Σ), and X|A=1~N(μ,Σ) Here, Ber is the Bernoulli distribution and N is the normal distribution.
[0078]
number
[0079] Result variable Y=(1-A)Y 0 +AY 1 To sample the potential outcome Y 0 and Y 1 was generated as follows:
[0080] Y 0 ~N(-5,1);Y 1 ~N(0,h(g(X1, ,X5)) 2 ) Here, g and h are functions such as:
[0081]
number
[0082] The true positive rate (TPR) and false positive rate (FPR) are used as measures for evaluating the performance of the feature selection device 10 according to this embodiment. TP / d T , d FP / (dd T ) where d T = 5 is the number of truly relevant features, and d TP and d FP are the number of truly relevant features that are correctly determined to be truly relevant, and the number of unrelated features that are incorrectly determined to be relevant, respectively.
[0083] By generating the above data while changing the random numbers, we prepared a dataset consisting of 50 different artificial data, and evaluated the mean and standard deviation values of TPT and FPR in 50 experiments when the sample size of the observed data was n = 2000.
[0084] A two-layer neural network was used as the propensity score e used in the relevance assessment unit 202. Here, in this neural network, a rectified linear function (ReLU) was used as the activation function, and each layer was a linear layer with 50 hidden neurons.
[0085] The probabilistic generative model L used in the feature selection unit 203 is the conditional variational autoencoder (CVAE) described in Reference 6. Here, a two-layer neural network is used as the encoder and decoder of the CVAE. In this neural network, the rectified linear unit (ReLU) is used as the activation function, and each layer is linear, with 128 hidden neurons.
[0086] Furthermore, the bandwidth of the kernel function h X_1 ,···,h X_d and h Y is given by a heuristic called the median heuristic (Reference 7).
[0087] Under the above settings, the evaluation results of the feature selection device 10 according to this embodiment are shown in Table 1.
[0088] [Table 1] Here, "Proposed" is the evaluation result of the feature selection device 10 according to this embodiment, and represents the mean and standard deviation values of TPR and FPR. On the other hand, "SI-EM" represents the mean and standard deviation values of TPR and FPR obtained by the existing method described in Non-Patent Document 2.
[0089] In Table 1 above, we can see that SI-EM only selects features related to the mean value of the treatment effect, resulting in an extremely low TPR. On the other hand, Proposed uses a relevance assessment measure based on the distribution parameters of the treatment effect, so it correctly selects features that affect the variance of the treatment effect, achieving a high TPR and a low FPR.
[0090] <Summary> As described above, the feature selection device 10 according to this embodiment selects the potential result Y 1 For each feature X mThe conditional distribution P(Y 1 |X m = x) and the potential outcome Y if no treatment is administered. 0 For each feature X m The conditional distribution P(Y 0 |X m = x) is the characteristic X m This allows us to evaluate the relevance of each individual's characteristics to the magnitude of the distribution parameters of the treatment effect.
[0091] Furthermore, the feature selection device 10 according to this embodiment calculates the p-value of each feature using the above evaluation scale and determines whether or not to reject the null hypothesis that the feature is not related to the magnitude of the treatment effect, thereby enabling the selection of features related to the magnitude of the treatment effect.
[0092] The present invention is not limited to the above-described specifically disclosed embodiments, and various modifications, changes, and combinations with known technologies are possible without departing from the scope of the claims.
[0093] [References] Reference 1: Arthur Gretton, Karsten M. Borgwardt, Malte J Rasch, Bernhard Scholkopf, and Alexander Smola. "A kernel two-sample test". JMLR, 13(1):723-773, 2012. Reference 2: Paul R. Rosenbaum and Donald B. Rubin. "The central role of the propensity score in observational studies for causal effects". Biometrika, 70(1):41-55, 1983. Reference 3: Ali Rahimi, Benjamin Recht, et al. "Random features for large-scale kernel machines". In NeurIPS, volume 3, page 5, 2007. Reference 4: Emmanuel Candes, Yingying Fan, Lucas Janson, and Jinchi Lv. "Panning for gold: 'Model-X' knockoffs for high dimensional controlled variable selection". Journal of the Royal Statistical Society: Series B (Statistical Methodology), 80 (3):551-577, 2018. Reference 5: Yoav Benjamini and Yosef Hochberg. "Controlling the false discovery rate: A practical and powerful approach to multiple testing". Journal of the Royal statistical society: series B (Methodology), 57(1):289-300, 1995. Reference 6: Kihyuk Sohn, Honglak Lee, and Xinchen Yan. "Learning structured output representation using deep conditional generative models". In NeurIPS, pages 3483-3491, 2015. Reference 7: Damien Garreau, Wittawat Jitkrittum, Motonobu Kanagawa. "Large sample analysis of the median heuristic". arXiv preprint arXiv:1707.07269, 2017. [Explanation of symbols]
[0094] 10 Feature Selection Device 101 Input Device 102 Display device 103 External I / F 103a Recording media 104 Communication I / F 105 RAM 106 ROM 107 Auxiliary storage 108 processors 109 Bus 201 Input section 202 Relevance Assessment Department 203 Feature Selection Unit 204 Output section 205 Observation Data DB 206 First Parameter DB 207 Second Parameter DB
Claims
1. The observed value a of the treatment variable A, which indicates whether or not a specific treatment was administered to the individual i. i and a feature variable X=[X 1 , ..., X d ] T (where d is the number of features) observed value x i and the observed value y of the outcome variable Y representing the predetermined outcome of the individual i. i Observation data D = {(a i , x i , y i )} to estimate the potential outcome Y when the treatment is administered. 1 Feature X m The conditional distribution P(Y 1 |X m = x) and the potential outcome Y if the treatment is not administered. 0 Feature X m The conditional distribution P(Y 0 |X m = x) is the distribution distance between the feature X m The evaluation scale I indicates how much the attribute value x of m Estimator of ~ I m an evaluation procedure for calculating The observation data D and the evaluation scale I m a selection procedure for selecting features related to the magnitude of the treatment effect using (m=1, ..., d); A computer-implemented feature selection method.
2. The selection procedure comprises: Feature x including artificial features generated from a probabilistic generative model L that fits the observed data D i (b) and the observed value a of the treatment variable A i and the observed value y of the outcome variable Y i Data consisting of {(a i , x i (b) , y i )}, and the evaluation scale I m Estimator of ~ I m (b) Calculate The estimator ~ I m and the estimator ~ I m (b) Using this, the approximate value of p^p m Calculate The approximation of the p-value ^p m and a predetermined significance level α, I m The null hypothesis H represents 0 0,m The feature selection method according to claim 1 , wherein features rejected by are selected as features related to the magnitude of the treatment effect.
3. The distribution distance is the maximum mean discrepancy (MMD), The evaluation procedure includes: The evaluation scale I is calculated by approximating a kernel function included in the inter-distribution distance using RFFs (random Fourier features) using a weighting function using a propensity score estimated from a conditional distribution model that fits the observed data D. m Estimator of ~ I m The feature selection method according to claim 1 or 2, further comprising the step of:
4. The observed value a of the treatment variable A, which indicates whether or not a specific treatment was administered to the individual i. i and a feature variable X=[X 1 , ..., X d ] T (where d is the number of features) observed value x i and the observed value y of the outcome variable Y representing the predetermined outcome of the individual i. i Observation data D = {(a i , x i , y i )} to estimate the potential outcome Y when the treatment is administered. 1 Feature X m The conditional distribution P(Y 1 |X m = x) and the potential outcome Y if the treatment is not administered. 0 Feature X m The conditional distribution P(Y 0 |X m = x) is the distribution distance between the feature X m The evaluation scale representing the degree to which the feature X varies depending on the attribute value x is m Scale I for assessing the association between the distribution parameters of the treatment effect and m As the evaluation scale I m Estimator of ~ I m An evaluation procedure to calculate A computer-implemented relevance assessment method.
5. The observed value a of the treatment variable A, which indicates whether or not a specific treatment was administered to the individual i. i and a feature variable X=[X 1 , ..., X d ] T (where d is the number of features) observed value x i and the observed value y of the outcome variable Y representing the predetermined outcome of the individual i. i Observation data D = {(a i , x i , y i )} to estimate the potential outcome Y when the treatment is administered. 1 Feature X m The conditional distribution P(Y 1 |X m = x) and the potential outcome Y if the treatment is not administered. 0 Feature X m The conditional distribution P(Y 0 |X m = x) is the distribution distance between the feature X m The evaluation scale I indicates how much the attribute value x of m Estimator of ~ I m an evaluator configured to calculate The observation data D and the evaluation scale I m a selection unit configured to select features related to the magnitude of the treatment effect using (m=1, ..., d); A feature selection device having:
6. The observed value a of the treatment variable A, which indicates whether or not a specific treatment was administered to the individual i. i and a feature variable X=[X 1 , ..., X d ] T (where d is the number of features) observed value x i and the observed value y of the outcome variable Y representing the predetermined outcome of the individual i. i Observation data D = {(a i , x i , y i )} to estimate the potential outcome Y when the treatment is administered. 1 Feature X m The conditional distribution P(Y 1 |X m = x) and the potential outcome Y if the treatment is not administered. 0 Feature X m The conditional distribution P(Y 0 |X m = x) is the distribution distance between the feature X m The evaluation scale representing the degree to which the feature X varies depending on the attribute value x is m Scale I for assessing the association between the distribution parameters of the treatment effect and m As the evaluation scale I m Estimator of ~ I m an evaluator configured to calculate A relevance assessment device having the following.
7. A program that causes a computer to function as the feature selection device according to claim 5 or the relevance assessment device according to claim 6.