Conditional independence test method
The proposed quantum kernel-based conditional independence testing method addresses the challenge of uncovering nonlinear causal relationships with small data sets, enhancing causal discovery capabilities.
Patent Information
- Application Number
- PCT/JP2025/010329
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-28
- Filing Date
- 2025-03-18
- Publication Date
- 2025-10-02
AI Technical Summary
Existing causal discovery methods struggle to uncover nonlinear causal relationships between data, especially with small sample sizes, limiting their applicability and accuracy.
A conditional independence testing method using quantum algorithms, specifically quantum kernel methods, to estimate conditional statistics and perform causal exploration with smaller data sets, enabling the identification of nonlinear relationships.
Enables causal discovery of nonlinear relationships with small sample sizes, expanding the range of applicable problems compared to conventional techniques and improving estimation accuracy.
Smart Images

Figure JP2025010329_02102025_PF_FP_ABST
Abstract
Description
Conditional independence test method
[0001] The present disclosure relates to a conditional independence test method.
[0002] Traditionally, large data samples were required to uncover nonlinear causal relationships between data using causal discovery methods based on statistical theory.
[0003] JP 2023-012480 JP 10-268906 JP 2022-530082 JP JP 2021-513391
[0004] In recent years, qLiNGAM (quantum Linear Non-Gaussian Model) has been proposed as an effective method even for small sample sizes. However, it assumes that the relationships between data are linear, and does not achieve sufficient accuracy for subjects that deviate from this assumption, limiting its scope of application.
[0005] Therefore, in this disclosure, we propose a conditional independence testing method that extends the conditional independence testing that can capture nonlinear features to a form that can be applied to quantum algorithms (particularly, quantum kernel methods).
[0006] In order to solve the above problem, one embodiment of a conditional independence testing method according to the present disclosure is a conditional independence testing method that estimates conditional statistics using a quantum algorithm for two variables that are the subject of an independence test and multiple condition variables.
[0007] FIG. 1 is a diagram showing the inclusion relationship of linear / nonlinear relational expressions. FIG. 2 is an example of a workflow for identifying a causal relationship between variables. FIG. 3 is a flowchart when the present method is applied to a PC algorithm. FIG. 4 is a detailed description (conceptual diagram) of S102. FIG. 5 is a detailed description (flowchart) of S102. FIG. 6 is a diagram showing an example of an independence test. FIG. 7 is a diagram showing an example of a quantum circuit used for calculating the quantum kernel. FIG. 8 is a diagram showing an example of a quantum circuit representing Vi(x). FIG. 9 is a diagram showing another example of a circuit used for calculating the quantum kernel.
[0008] Hereinafter, embodiments will be described in detail with reference to the drawings. In the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted.
[0009] (List of References) ・[1] Bach-Jordan (2002): Kernel independent component analysis, Journal of machine learning research, 2002, 3.Jul: 1-48. ・[2] Fukumizu et. al. (2007): Kernel measures of conditional dependence, Advances in neural information processing systems 20 (2007). ・[3] Kawaguchi (2023): Application of quantum computing to a linear non-gaussian acyclic model for novel medical knowledge discovery, Plos one 18.4 (2023): e0283933. ・[4] Maeda-Kawaguchi-Tezuka (2023): Estimation of mutual information via quantum kernel method, arXiv preprint arXiv:2310.12396 (2023). ・[5] Zhang et. al. (2012): Kernel-based Conditional Independence Test and Application in Causal Discovery, arXiv preprint arXiv:1202.3775 (2012). ・[6] Gretton et. al. (2007): A kernel statistical independence test, Advances in neural information processing systems 20 (2007). ・[7] Schuld et. al. (2001): The effect of data encoding on the expressive power of variational quantum machine learning models, Physical Revew A 103, 032430 (2021).・[8] Nakaji- Yamamoto(2021): Expressibility of the alternating layered ansatz for quantum computation, Quantum 5, 434 (2021).
[0010] (1. Embodiment) In recent years, the importance of efficiently extracting useful insights by collecting and analyzing data has increased in various fields. Furthermore, as AI becomes more complex and larger in scale, the development of so-called explainable AI, which visualizes the basis for AI output in a form that can be interpreted by humans, is also progressing. Causal discovery is one method for extracting complex data sets, in which the relationships between elements are not immediately apparent, in a form that can be interpreted by humans. This method identifies the causal relationships behind a given data set and provides deeper information than the correlations obtained by conventional statistical methods. Among the various causal discovery methods, those that can handle nonlinear relationships are generally called nonparametric methods and are characterized by their broad range of applicability. In fact, the relationships between events in social and natural phenomena are often expressed as nonlinear relationships rather than linear relationships.
[0011] However, in the past, there was no causal discovery method that could reveal nonlinear causal relationships using a small number of samples. The present method provides a means to make this possible. The embodiment described below relates to a nonlinear causal discovery method that applies a conditional independence test using a quantum algorithm.
[0012] For the sake of the discussion below, we will explain the elemental technologies of the causal discovery method. The following two steps are required to clarify the causal relationships between variables (see Figure 2).
[0013] Step 1. [Independence test] Identify correlations and independence between variables. Step 2. [Orientation rules] Based on Step 1, identify causal relationships between variables.
[0014] [Independence Test] Step 1, called an independence test, determines the independence of variables. Here, we define variable independence. Given a probability distribution Pxy(x,y) and its marginal distributions Px(x) and Py(y), the random variables X and Y given by these probability distributions are said to be independent if Pxy(x,y)=Px(x)Py(y). This is a nonlinear extension of uncorrelation and is a generalization of uncorrelation. According to the correspondence between the distributions of two variables and the correlation coefficient (see, for example, https: / / ja.wikipedia.org / wiki / %E7%9B%B8%E9%96%A2%E4%BF%82%E6%95%B0), variables with a linear relationship are assigned a nonzero correlation coefficient, while variables with a nonlinear relationship are displayed as a zero correlation coefficient. This demonstrates that the correlation coefficient can only express a linear relationship. In other words, in order to clarify the complex relationships between variables, it is necessary to select a method that can clarify not only linear relationships but also nonlinear relationships.
[0015] Independence tests are classified into the following two types: 1-A. Chi-square test: A method for determining linear independence. 1-B. Kernel estimation: A method for determining non-linear independence. Table 1 shows the advantages and disadvantages of the above methods, and Table 2 is a summary of Table 1 from the perspective of the shape of the data.
[0016]
[0017]
[0018] 1 is a diagram showing the inclusion relationship between linear and nonlinear relationships. Linear relationships can be distinguished using correlation coefficients and chi-square tests, while nonlinear relationships can be distinguished using kernel estimation.
[0019] This application proposes a method for calculating conditional statistics using a quantum algorithm, which makes it possible to clarify nonlinear relationships using a smaller amount of data than conventional methods.
[0020] [Conditional independence tests] The discussion so far has concerned methods for directly determining independence between variables. For later discussion, we will also touch on conditional independence tests. Given random variables X and Y, X and Y are said to be conditionally independent under another random variable Z if Pxy|z(x|y,z)=Px|z(x|z). When performing nonlinear causal exploration, which will be described later, it is also necessary to determine such higher-order conditional independence.
[0021] Table 3 shows a classification of conventional and proposed methods based on the independence test and the kernel used. Generally, when the relationship between variables is linear, a method using an unconditional independence test is used, and when the relationship between variables is nonlinear, a method using a conditional independence test is used. The proposed method according to the embodiment of the present application is a method using a conditional independence test and a quantum kernel, and is called QKCIT (Quantum Kernel Conditional Independence Test) and QHSCIC (Quantum Hilbert-Schmidt Conditional Independence Criterion), which will be described in detail later.
[0022]
[0023] Although it is not a common name, in this application the method described in [2] will be referred to as HSCIC (Hilbert-Schmidt Conditional Independence Criterion).
[0024]
[0025] [Causal Discovery Method] In step 2 of Figure 2, orientation rules are applied to determine the causal relationships between variables from state 2. These are steps to determine the direction of causality based on the independence between variables; for details, please refer to the flowchart in Figure 3. At this stage, it is necessary to select the "class (linearity)" to assume for the causal graph to be analyzed. In other words, the appropriate method will differ depending on the assumptions made about the causal graph. Typical algorithms are categorized as follows:
[0026] 2-A. Methods that assume linearity in the graph (e.g., LiNGAM, qLiNGAM) 2-B. Methods that assume nonlinearity in the graph (e.g., PC algorithm)
[0027] Table 4 shows the advantages and disadvantages of using or not using conditions for independence testing. As shown in Table 4, conditional independence testing and methods using classical kernels have the advantage of being able to handle any causal graph, but they tend to require large amounts of data (approximately 1,000 or more). This is due to the high degree of freedom that comes from not explicitly specifying the relationships between variables. In other words, nonparametric methods make no assumptions about the parameters of the causal graph they identify, so while they can handle nonlinear relationships, they also require less information to use as prior knowledge. However, when considering practical use, collecting large amounts of data is often difficult, and there are many cases where causal exploration is desired with a sample size of 100 or less.
[0028] [Method and Effects of the Present Application] This application provides a method that uses quantum computing to perform part of the causal discovery process (testing independence between variables) to reveal relationships between variables even with a smaller number of samples than conventional methods. In particular, by executing a quantum kernel method using a type of quantum transformation circuit (e.g., an IQP circuit) that is computationally difficult to imitate classically, causal discovery can be performed on data with shapes that are difficult to handle efficiently on classical computers. IQP circuits are considered difficult to imitate on classical computers due to the computation time required. As a result, quantum circuits can execute calculations in high-dimensional spaces that cannot be handled classically, and relationships between them can sometimes be clearly distinguished even with a small number of samples. The central limit theorem is one of the reasons why classical causal discovery methods require large amounts of data. In independence testing in classical causal discovery methods, a Gaussian kernel is theoretically adopted when using kernel methods. However, while the Gaussian kernel is suitable when the data distribution is close to a normal distribution, it becomes difficult to distinguish information when the input data has large variance (variance). This is because the probability density of a Gaussian decays exponentially in the tail regions away from the mean.
[0029] In this application, a quantum kernel is constructed using a quantum circuit called an IQP circuit for data encoding. This kernel is difficult to simulate using classical methods, and quantum computing may offer advantages unique to quantum computing. Additionally, to ensure a wide distribution of features projected onto the Bloch sphere, the design flexibility is expanded by using an activation function that appears in classical neural networks when inputting data. The quantum kernel method embeds quantum states into a high-dimensional Hilbert space using kernel mapping for calculation, enabling efficient nonlinear data transformation.
[0030] Method: Estimation of conditional statistics using quantum algorithms. Effect: It is now possible to perform causal exploration for nonlinear problems using small samples, expanding the range of applicable problems compared to conventional techniques.
[0031]
[0032] [Positioning of the proposed method] Table 5 lists the statistics to be estimated and the estimation methods. MI is an abbreviation for Mutual Information, SMI for Squared-loss MI, TCI for Test Conditional Independence statistics, and CSMI for Conditional Squared-loss MI, and will be abbreviated below. MI and SMI are unconditional statistics and are used in unconditional independence tests. TCI and CSMI are conditional statistics and are used in conditional independence tests.
[0033] Equations 1 to 4 are the mathematical definitions of MI, SMI, TCI, and CSMI, respectively. Here, Px and Py denote the probability density functions that generate samples of feature quantities X and Y. Kx|z and Ky|z represent the conditional kernel matrices.
[0034]
[0035]
[0036]
[0037]
[0038] As one method in this application, we have developed a new method for estimating TCI using a quantum kernel and a new method for performing independence testing using this method. We call this method QKCIT. As another method in this application, we have developed a new method for estimating CSMI using a quantum kernel. We call this method QHSCIC.
[0039] [Explanation of the Proposed Method] [Example 1] QKCIT (Quantum Kernel Conditional Independence Test) Example 1 is a method for performing the KCIT (Kernel Conditional Independence Test) [5] using a quantum computer or simulator. This method does not perform mutual information estimation or shuffle testing. Therefore, the associated testing method is described below. The procedure is as follows (see the flowchart in the example):
[0040] 1. Take samples X and Y for which you want to test conditional independence, and conditional variable Z, and take the proposition that X and Y are independent on Z as the null hypothesis. 2. Using a quantum kernel method that runs on a quantum computer or simulator, estimate the statistical quantity TCI according to the calculation method shown in Method 1. 3. Using the method described in [5, Proposition 6], approximate the TCI distribution with a gamma distribution from the information in the quantum kernel matrix, and find its variance and mean. 4. Estimate which region in 3 the TCI calculated in 2 falls into and calculate the p-value. 5. Perform a statistical test according to the p-value calculated in 4 to determine whether to reject the null hypothesis.
[0041] [Example 2] QHSCIC (Quantum Hilbert-Schmidt Conditional Independence Criterion) Example 2 is a method for performing HSCIC (Hilbert-Schmidt Conditional Independence Criterion) [2] using a quantum computer or simulator. This method involves estimating the mutual information (CSMI) and then performing a shuffle test. The procedure is as follows (see the flowchart in the example):
[0042] 1. Take samples X and Y for which you want to test conditional independence, and conditional variable Z, and take the proposition that X and Y are independent on Z as the null hypothesis. 2. Using a quantum kernel method that runs on a quantum computer or simulator, estimate the conditional mutual information (CSMI), a type of conditional statistic, according to the calculation method shown in Method 2. 3. As explained in the previous slide, shuffle X, Y, and Z, and estimate the CSMI again in step 2 for the shuffled variables. 4. Repeat step 3 100 times, and calculate the p-value by determining in what percentage the CSMI obtained in step 2 ranks. 5. Perform a statistical test based on the p-value calculated in step 4 to determine whether to reject the null hypothesis.
[0043] [Effects] Quantum kernel methods can now be applied to the estimation of conditional statistics. This enables causal exploration of nonlinear problems with small sample sizes, broadening the range of applicable problems compared to conventional techniques (qLiNGAM). Comparing Examples 1 and 2 with existing methods (KCIT, HSCIC), that is, using conditional independence tests with quantum kernels improves estimation accuracy with smaller data sets [3, 4]. Furthermore, quantum computing enables efficient nonlinear transformations. As a result, the range of suitable problems is expanded. In general, [5] shows that KCIT can determine conditional independence better than HSIC under the relationship shown in [5].
[0044] As explained above, if we can estimate the conditional statistics of the probability distribution that generates the variables, we can perform an independence test using the shuffle test. Therefore, we will explain how to estimate the conditional statistics using the quantum kernel method below.
[0045] In the following, we will denote the kernel matrices generated by a quantum circuit from variables X, Y, and Z as Kx, Ky, and Kz. Examples of quantum circuits include various PQC (Parametrized Quantum Circuits), such as HEA (Hardware Efficient Ansatz) and ALA (Alternate Layer Ansatz), IQP (Instantaneous Quantum Polynomial time), and DQC1, but are not limited to the ones explicitly mentioned.
[0046] [Method 1] [TCI estimation method using quantum kernel (QKCIT: Quantum Kernel Conditional Independence Test subroutine)]
[0047]
[0048] In the above equation 5, I is an identity matrix, and J is a matrix in which all elements are 1. A centered kernel matrix is placed for this n*n matrix as follows:
[0049]
[0050] Furthermore, for the hyperparameter μ, Tokki, For these matrices, Calculate.
[0051] [Method 2] [Estimation method using the quantum kernel of CSMI (subroutine of QHSCIC: Quantum Hilbert-Schmidt Conditional Independence Criterion test)]
[0052] For a sufficiently small hyperparameter δ, In response to these, In this case, if a quantum kernel that satisfies the characteristic property is selected, then for any precision ε, when the number of samples n is sufficiently large, this value will approximate CSMI(X,Y|Z) with precision ε.
[0053] [Differences from conventional techniques] There are four combinations of causal discovery methods: · Unconditional or conditional independence tests · Linear (unconditional independence tests) or nonlinear (conditional independence tests) orientation rules (= 2 × 2 combinations). Furthermore, existing classical kernel methods require large amounts of data to uncover nonlinear causal relationships. It has been shown that qLiNGAM, a causal discovery method based on quantum kernel methods for independence testing, can sometimes uncover causal relationships between data sets with smaller amounts of data than conventional methods [3]. However, because the independence tests used in this method are unconditional independence tests and use LiNGAM as the orientation rule, it can only uncover linear causal relationships.
[0054] In this study, we have developed a conditional independence test using quantum kernel methods, devising a means to perform conditional independence tests. This allows for the adoption of nonlinear methods as orientation rules. Furthermore, as suggested by qLiNGAM, the use of quantum kernels is expected to improve estimation accuracy with small amounts of data. As a result, nonlinear causal relationships between variables can be clarified even in situations where only a small amount of data is available.
[0055] Maeda, Kawaguchi, and Tezuka have devised an unconditional independence test using quantum kernel methods. These provide a method for estimating the MI and SMI mentioned above using quantum kernel methods (see Table 5). For comparison, the calculation method is explained below.
[0056] [MI estimation method (Bach-Jordan [1])] Let κ be the hyperparameter presented in [1]. First, create kernel matrices Kx and Ky using the selected kernel from variables X and Y, as described in [1]. When the size of these matrices is n*n, prepare matrices of the same size, tmp1 and tmp2. Following the definitions of these matrices, calculate Kk and Dk, each of size 2n*2n. The definitions of each matrix are as follows:
[0057]
[0058] Singular value decomposition is performed on Kk and Dk, and the resulting singular values are σk and σ D Finally, calculate the following quantities:
[0059]
[0060] In this case, if the number of samples n is large enough, this value will approximate MI(X,Y).
[0061] [Method for estimating SMI (Fukumizu, et. al. [2])] Calculate Tr(K´´y,K´´x) for matrices K´´x and K´´y in the CSMI shown on the left. In this case, if a quantum kernel that satisfies a property called characteristic is selected, then for any precision ε, when the number of samples n is sufficiently large, this value will approximate SMI(X,Y) with precision ε.
[0062] Example (Example of applying a quantum algorithm to a PC algorithm) Figure 3 is a flowchart showing a case where the quantum algorithm according to the present application is applied to a PC algorithm, which is a type of nonlinear causal search method. By using a quantum algorithm, it is expected that advantages unique to quantum computing can be obtained.
[0063] Each step of the flowchart (S101 to S105) will be explained below.
[0064] S101: Creating an undirected graph Draw edges between all N variables to create a complete undirected graph.
[0065] S102: Conditional independence test for two variables For N variables, a conditional independence test is performed for all pairs of two variables, with the other variable given, and branches are eliminated in the case of marginal independence (= unconditional independence) or conditional independence. A quantum algorithm is used to estimate conditional statistics for this independence test. At this stage, a conditional independence test is performed for any two variables and all other marginal variables. Note that S102 is a step that performs processing based on quantum computing and associated processing, and quantum kernel methods can be used as one quantum algorithm method, for example.
[0066] S103: Eliminate edges based on the results of the conditional independence test. In practice, the result of the conditional independence test at step S102 indicates that the nodes are independent, and so executing S103 at that point reduces the amount of calculation.
[0067] S104: Apply orientation rules for three variables. For three variables (X,Y,Z) where (X,Y) and (Y,Z) are adjacent but (X,Z) is not adjacent, if X and Z are dependent, given Y, then identify X → Y ← Z. Let F' be a partially directed graph.
[0068] S105: Apply orientation rules for four variables Repeat the following until the branches of F' can no longer be directed. a: In F', when X → Y exists, YZ exists, and X and Z are not adjacent, then Y → Z. b: When a path exists from X to Y and XY is in F', then X → Y. c: For four variables X, Y, Z, and W, when Y → Z and W → Z exist, and XY, XZ, and XW exist, then X → Z.
[0069] More preferably, steps S102 to S105 are performed assuming nonlinearity, which is expected to produce better results. In any of steps S102 to S105, the user may draw causal arrows between variables at their discretion, based on prior knowledge. As a result, a causal graph that identifies causal relationships more closely recognizable to the user can be obtained at the end of the process.
[0070] Note that S102 to S105 can be converted into operations for solving the SAT (satisfiability) problem, and the SAT problem can be efficiently solved using several quantum algorithms (including quantum-inspired algorithms), such as quantum annealing machines and Grover's algorithm, and these algorithms may also be used.
[0071] Figure 4 is a diagram detailing the steps shown in Figure 3, focusing on the environment in which they are executed (classical computer, quantum computer). In particular, it is a conceptual diagram detailing S102, which corresponds to the dashed line. Here, the part executed on the quantum computer may be executed by a simulator.
[0072] Fig. 4 shows only one flow from S501 to S505. First, a combination of two variables, x and y, is selected from N variables, k is selected from integers between 0 and N-2, and a combination of k variables is selected from N-2 variables other than x and y, and the independence test for all combinations may be performed by exhaustively repeating this process.
[0073] However, it is possible to omit unnecessary selections as appropriate. For example, if information that certain variables x and y are independent has already been obtained, it is possible to remove x and y, which are known to be independent, from the candidates when selecting two variables from N variables. As mentioned above, this pruning based on background knowledge can be performed at any stage from S102 to S105.
[0074] Examples of quantum circuits in S502-504 include, but are not limited to, various PQCs (Parametrized Quantum Circuits), such as HEA (Hardware Efficient Ansatz), ALA (Alternated Layer Ansatz), IQP (Instantaneous Quantum Polynomial time), DQC1, etc. Furthermore, when S503 is executed on a quantum computer simulator, the coefficient values are directly accessed and the float values are registered in classical registers.
[0075] Regarding S503 (reading values from the quantum circuit), in the present application, the target statistics are estimated using information on the coefficients obtained when the final state of the quantum computation is decomposed in the computational basis.
[0076] There are various methods for determining these coefficients using a quantum computer. The simplest method involves performing multiple measurements by projecting onto a measurement basis, and then identifying each coefficient using statistical estimation from the obtained observations. More generally, quantum state tomography can also be used. Furthermore, techniques called classical shadow or shadow tomography can be used to reduce the number of required observations and efficiently estimate the state. Furthermore, any method that can identify the desired quantum state is not limited to those listed here.
[0077] On the other hand, when using a simulator without a quantum computer, there is no need to insert a measurement step. In other words, in a classical computer, it corresponds to each calculation process after the calculation. Information representing the quantum state is stored in a register in the classical computer, and this information can be read directly (in a quantum computing simulator, this information is called a state vector or density matrix).
[0078] Finally, the final state of quantum computation can refer to various things. Generally, the entire Hilbert space spanned by all qubits is often used, but for various reasons, the state after projection into a subspace may also be treated. Quantities formed from observation results for such subspaces may also be treated as elements of a kernel matrix. Note that in this application, the term "final state" is not limited to those explicitly stated here, as long as it relates to statistics estimated using values calculated through quantum computation expressions.
[0079] Example (Example of specific steps of the proposed method; detailed description of S102 in Fig. 3) Fig. 5 is a flowchart that details step S102 described in Fig. 4, focusing on the algorithm flow. Unlike Fig. 4, Fig. 5 describes a flowchart for a case where a comprehensive independence test is performed.
[0080] In S1026, it is determined whether the selection in S1024 (selection of k combinations of variables that have not yet been selected from N-2 variables other than x and y) has covered all selection methods, i.e., all N-2Ck ways; if YES, the process proceeds to S1027; if NO, the process proceeds to S1024.
[0081] In S1028, it is determined whether or not all selection methods, i.e., all NC2 ways, have been covered in the selection in S1021 (selecting combinations of two variables x and y that have not yet been selected from N variables), and if YES, the process proceeds to S103, and if NO, the process proceeds to S1021.
[0082] Simulation / Implementation Example As an example of implementation, we will explain the case where 10 variables and 100 data (100 data numerical vectors with 10 features (variables)) are given as shown in Figure 6.
[0083] In S501, if necessary, the data is preprocessed using normalization or the activation function described in the next section, and then x and y are selected, and k variables are selected from the other N-2 variables. As described in the explanation of Figure 5, a test of independence can be performed for all combinations by comprehensively selecting a combination of two variables, x and y, from N variables, selecting k from integers between 0 and N-2, and selecting a combination of k variables from the N-2 variables other than x and y.
[0084] In S502, each x = (x1, ..., x100) is encoded into a quantum state using a quantum circuit, and in S503, The value of is read from the quantum circuit and a real value k(i,j) is calculated. Here, by calculating all combinations of integer values 0≦i, j≦100 for each of i and j, the kernel matrix Kx can be obtained. In other words, the matrix Kx is the matrix whose (i,j) component is K(i,j). Similarly, the kernel matrix Ky is calculated for y. The kernel matrix for k variables is also calculated in the same way. Here, the kernel matrix for multiple variables can be obtained by calculating the joint distribution. From the kernel matrix calculated above, the conditional kernel matrix shown in S504 is further calculated.
[0085] In S505, conditional statistics are estimated based on the conditional kernel matrix calculated in S504, and a shuffle test or a test based on an estimated gamma distribution is performed.
[0086] [Supplementary Information] Regarding the independence test using the shuffle test, we adopted the shuffle test as an independence test by estimating mutual information. (Specific implementation method) 1. Estimate the mutual information between the data (for example, MI, SMI, CMI, CSMI as shown on the left). 2. Shuffle (randomly rearrange) the data and calculate the above mutual information. 3. Repeat step 2 100 times, and if the mutual information calculated in step 1 is within the top 5%, determine that the data are dependent; if not, determine that the data are independent. 4. Perform steps 1-3 for each variable and each conditional variable.
[0087] (Data shuffling method) For X=(x_1,…,x_1000) and Y=(y_1,…,y_1000), generate (x_1,…,x_1000,y_1,…,y_1000), randomly rearrange the components of this vector, and split it into two vectors with 1000 components each.
[0088] Example of quantum circuit implementation Figure 7 shows an overall view of the quantum circuit used to extract the i-th and j-th components of each variable (feature) and encode them into a quantum state using a quantum circuit.
[0089] The quantum circuit shown by UIQP in Figure 7 is called an IQP circuit, and is known as a class of quantum circuits that are difficult to emulate efficiently on classical computers.
[0090] This embodiment is expected to obtain advantages unique to quantum computing by using quantum algorithms, and furthermore, by operating a specific class of quantum circuits, such as an IQP circuit, on a quantum computer, it is possible to reduce the amount of calculation compared to a classical computer.
[0091] Here, the unitary matrix corresponding to UIQP is given by:
[0092]
[0093] Here, H is a Hadamard matrix. Each Vi(x) has a certain configuration, and as an example, it can be expressed as a quantum circuit consisting of U(x) and controlled-U(x) as shown in Figure 7. Here, i in Vi(x) is an index value that takes an integer value from 1 to D, and D represents the value of a hyperparameter (the number of iterations of the circuit Vi(x)).
[0094] Also, rotation gates around various axes are candidates for U, for example, U1(x) is adopted. Here, U1(x) and controlled-U1(x) can be expressed as matrices as follows.
[0095]
[0096]
[0097] The input data x may be processed with the activation function shown below before being encoded using the IQP circuit. By processing it in this way, the symmetry of the function is expected to allow the data to be densely encoded on the Bloch sphere, which is the space of quantum states, and more accurate calculations are expected to be possible.
[0098]
[0099] 8, there is a degree of freedom in the embedding of arguments, rotation gates, and control gates, and they are not limited to those shown here. In this embodiment, a controlled-X gate acts on any two qubits.
[0100] Supplementary Material (Other Quantum Circuits) When using a kernel generated by a quantum circuit to estimate SMI and CSMI using the method of Fukumizu, et. al. in [2], the kernel had to satisfy a characteristic property. In general, it is known that a kernel is characteristic if it is universal, and it is known that if a specific quantum circuit is used, it is possible to generate a kernel that satisfies the universal property when the number of qubits is increased infinitely (for example, Theorem in [7]).
[0101] Therefore, in addition to the IQP circuit, we propose the following circuit as an example. Abstractly, we combine the ansatz Wi and the data encoding term S(x) to obtain W0S(x) W1S(x)...S(x)W L This quantum circuit is used.
[0102] Here, L is a hyperparameter that represents the number of iterations of the basic circuit. In Figure 9, as an example, data is encoded using RZ as described in [7] Schuld et. al. That is, given data x=(x0,...,x n-1 ) prepare n qubits and encode the data, and in this case, the 1-qubit local Hamiltonian RZ acting on the i-th qubit is Here, σz is the Pauli Z matrix and θ is a hyperparameter. A specific quantum circuit is shown in Figure 9. More specifically, Figure 9 shows an example using a controlled-x gate as Wi and Rz for data encoding.
[0103] For Wi, we use hardware efficient ansatz, alternating layered ansatz, all-to-all tensor product ansatz, etc. (see [8]), and for the data encoding layer S(x), we use circuits such as RX, RY, and RZ with embedded hyperparameters. Figure 9 shows an example of a hardware efficient ansatz.
[0104] Figure 9: Another example of a circuit used to calculate the quantum kernel. W0S(x) W1S(x)...S(x)W for L=1. L This is an example in which the quantum circuit is connected as a hardware efficient ansatz using the equation 19, Controlled-X as the ansatz. The data to be encoded here is x i =(x i1 ,...,x in ), x j =(x j1 ,...,x jn), and the (i, j) element K(i, j) of the kernel matrix is calculated from this. Here, n is the number of data, and in the example of FIG. 8, n=100.
[0105] There is a degree of freedom in how arguments are embedded, as well as in the rotation gate and control gate, and they are not limited to those shown here; general parameterized quantum circuits are also acceptable. In this example, a controlled-X gate is applied to any two qubits. θ is a hyperparameter that is similar to the width of the Gaussian kernel.
[0106] [Glossary] KCIT... Kernel Conditional Independence Test. Developed in a paper by Zhang et al. This refers to a method for calculating the p-value and performing conditional independence tests by using the kernel method to calculate the statistic (real-valued) TCI (explained below). Since its null distribution can be calculated from a gamma distribution, this method performs p-value calculations and conditional independence tests. There are no restrictions on the kernel selection; in other words, it is theoretically guaranteed that independence can be demonstrated by increasing the number of samples infinitely regardless of the classical or quantum kernel used. TCI... Test Conditional Independence statics. This is a real-valued value, as explained above. It is a quantity that can be calculated using only the given data and the selected kernel. It is used in the KCIT described above and the QKCIT described below. QKCIT... Quantum Kernel Conditional Independence Test. This is a method developed in this application that calculates TCI using the quantum kernel method and uses it to perform independence tests. The quantum kernels used in this application include the IQP circuit described in the examples, as well as the universal kernel circuit shown in Figure 7. HSCIC...Hilbert-Schmidt Conditional Independence Criterion test. This is the name given to the conditional independence test proposed in [2]. While Bach-Jordan [1] uses MI to perform independence testing, HSCIC uses CSMI. Note that the kernel used must be characteristic. QSHCIC...Quantum Hilbert-Schmidt Conditional Independence Criterion test. This method, developed in this application, performs independence testing by estimating SMI using a quantum kernel. While it is unclear whether the IQP circuit is characteristic, the circuit in Figure 7 approaches universality as the number of qubits is increased, and is therefore characteristic, making it applicable to this method. MI...Mutual Information.Represents a certain distance between two probability distributions. SMI...Abbreviation for Squared-loss Mutual Information. A value obtained by smoothly correcting the above MI for outliers. CSMI...Abbreviation for Conditional Squared-loss Mutual Information. A value obtained by generalizing the above SMI for conditional distributions.
[0107] (2. Other Embodiments) The processing according to each of the above-described embodiments may be implemented in various different forms other than the above-described embodiments.
[0108] Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using known methods. Furthermore, the information, including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings, can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.
[0109] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.
[0110] Furthermore, the above-described embodiments and modifications can be combined as appropriate within the scope of not causing any contradiction in the processing content.
[0111] Furthermore, the effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0112] (3. Effects of the present disclosure) Quantum kernel methods can now be applied to the estimation of conditional statistics. This makes it possible to perform causal exploration with a small number of samples for objects with nonlinear relationships, and enables highly accurate estimation of causal relationships in a wider range of problem settings than conventional technology (qLiNGAM).
[0113] Note that the present technology can also be configured as follows. (1) A conditional independence testing method for two variables to be tested for independence and multiple condition variables, the conditional independence testing method comprising: estimating conditional statistics using a quantum algorithm. (2) The conditional independence testing method according to (1), wherein the quantum algorithm encodes the values of the two variables and the values of the multiple variables into a quantum state using a quantum circuit, reads values of projected components with respect to a predetermined basis from the quantum state, calculates a kernel matrix based on the projected component values, calculates a conditional kernel matrix from the kernel matrix, and estimates conditional statistics from the conditional kernel matrix. (3) The conditional independence testing method according to (2), wherein the encoding is performed as the angle of a rotation gate. (4) The conditional independence testing method according to (2) or (3), wherein the kernel matrix is calculated using a quantum kernel. (5) The conditional independence testing method according to any one of (2) to (4), characterized in that the process of estimating the conditional statistics from the conditional kernel matrix is a process of: taking a proposition that the first sample and the second sample are independent with respect to the conditional variables as a null hypothesis for a first sample, a second sample, and a conditional variable for which conditional independence is to be determined; estimating the statistics by a calculation method using a quantum kernel method that runs on a quantum computer or a simulator; approximating the distribution of the statistics with a gamma distribution from information on the quantum kernel matrix, and calculating its variance and mean; estimating in which region in the distribution the statistics is located, calculating a p-value; and performing a statistical test according to the p-value to determine whether to reject the null hypothesis.(6) The conditional independence testing method according to any one of (2) to (4), characterized in that the process of estimating the conditional statistics from the conditional kernel matrix is a process of: taking a proposition that the first sample and the second sample are independent on the conditional variables as a null hypothesis for a first sample, a second sample, and a conditional variable for which conditional independence is to be determined; estimating a target information amount as conditional mutual information by a calculation method using a quantum kernel method that runs on a quantum computer or a simulator; shuffling the conditional variables for the first sample and the second sample; and re-estimating the conditional mutual information for the shuffled variables a predetermined number of times; calculating a p-value based on the ranking of the target information amount in the estimated conditional mutual information; and performing a statistical test according to the p-value to determine whether to reject the null hypothesis.
Claims
1. A conditional independence testing method for two variables to be tested for independence and multiple condition variables, characterized in that it estimates conditional statistics using a quantum algorithm.
2. The conditional independence testing method according to claim 1, characterized in that the quantum algorithm: encodes the values of the two variables and the values of the multiple variables into a quantum state using a quantum circuit; reads out values of projection components with respect to a predetermined basis from the quantum state; calculates a kernel matrix based on the values of the projection components; calculates a conditional kernel matrix from the kernel matrix; and estimates conditional statistics from the conditional kernel matrix.
3. The conditional independence testing method according to claim 2, wherein the encoding is performed as an angle of a rotation gate.
4. The conditional independence testing method according to claim 2, wherein the kernel matrix is calculated using a quantum kernel.
5. The conditional independence testing method according to claim 2, wherein the process of estimating the conditional statistic from the conditional kernel matrix is a process that, for a first sample, a second sample, and a conditional variable for which conditional independence is to be determined, the proposition that the first sample and the second sample are independent with respect to the conditional variable is treated as a null hypothesis, estimates the statistic by a calculation method using a quantum kernel method that operates on a quantum computer or simulator, approximates the distribution of the statistic with a gamma distribution from information on the quantum kernel matrix, calculates its variance and mean, estimates the region in the distribution in which the statistic is located, calculates a p-value, and performs a statistical test according to the p-value to determine whether to reject the null hypothesis.
6. The conditional independence testing method according to claim 2, wherein the process of estimating the conditional statistics from the conditional kernel matrix is a process that, for a first sample, a second sample, and a conditional variable for which conditional independence is to be determined, the proposition that the first sample and the second sample are independent with respect to the conditional variable is treated as a null hypothesis, estimates the target information amount as conditional mutual information by a calculation method using a quantum kernel method that runs on a quantum computer or simulator, shuffles the conditional variables for the first sample and the second sample, and estimates the conditional mutual information again a predetermined number of times for the shuffled variables, calculates a p-value based on the ranking of the target information amount in the estimated conditional mutual information, and performs a statistical test according to the p-value to determine whether to reject the null hypothesis.
Citation Information
Cited By
LCD display screen production quality collaborative management and control method based on detection data chain
CN121146296A