Information processing device, information processing method, and program

JPWO2022260129A5Pending Publication Date: 2025-06-06
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023527920
Authority / Receiving Office
JP · JP
Patent Type
Applications
Priority Date
2022-06-09
Filing Date
2022-06-09
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Current diagnostic techniques for multifactorial and sporadic diseases, such as ALS, rely heavily on clinical findings and electrophysiological tests, lacking molecular biomarkers, especially for sporadic ALS which accounts for 90-95% of cases, making digital diagnosis challenging.

Method used

An information processing device and method that analyzes gene expression levels using a nonlinear model to identify gene combinations with high HSIC scores, selecting causative and related genes to generate molecular biomarkers for digital diagnosis, and determine ALS onset by clustering genetic data on a feature space.

Benefits of technology

Accurately identifies genes for diagnosing multifactorial diseases, generates molecular biomarkers for digital diagnosis, and determines ALS onset by effectively classifying healthy and ALS patient data, demonstrating a statistically significant difference in ROC AUC values.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

This information processing device comprises a processing unit which, with respect to each combination of genes included in a gene data set, calculates a criterion of dependency on a causal gene of a multifactorial disorder or a sporadic disease and on a related gene to the multifactorial disorder or the sporadic disease, and selects a predetermined number of gene combinations from the data set on the basis of the criterion.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and program

[0001] The present invention relates to an information processing device, an information processing method, and a program. This application claims priority to U.S. Provisional Application No. 63 / 208,509, filed June 9, 2021, the contents of which are incorporated herein by reference.

[0002] Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disease caused by the loss of motor neurons, and the development of diagnostic techniques for ALS is urgently needed.

[0003] Patent application No. 2017-29116 published

[0004] Because ALS diagnosis is based on clinical findings and electrophysiological testing after clinical symptoms have progressed, molecular biomarkers are needed for digital diagnosis of ALS. However, for sporadic ALS, which accounts for 90-95% of ALS cases, genes that could serve as molecular biomarkers remain unknown. This issue is not limited to ALS, but is also true for other multifactorial or sporadic diseases, such as Alzheimer's disease and Parkinson's disease, the majority of which are sporadic.

[0005] An object of the present invention is to provide an information processing device, an information processing method, and a program that can identify genes that can be used to diagnose multifactorial diseases or sporadic diseases.

[0006] One aspect of the present invention is an information processing device that includes a processing unit that calculates a measure of dependency on causative genes of a multifactorial disease or a sporadic disease and on genes associated with the multifactorial disease or the sporadic disease for each combination of the genes included in a gene dataset, and selects a predetermined number of gene combinations from the dataset based on the measure.

[0007] According to one aspect of the present invention, genes capable of diagnosing multifactorial or sporadic diseases can be identified.

[0008] 1 is a diagram for explaining an overview of an embodiment. FIG. 1 is a diagram illustrating an example of the configuration of an information processing device according to an embodiment. FIG. 2 is a flowchart illustrating a series of processing steps performed by a processing unit related to identification of gene combinations according to an embodiment. FIG. 2 is a diagram illustrating a method for calculating an HSIC score. FIG. 3 is a flowchart illustrating a series of processing steps performed by a processing unit related to generation of molecular biomarkers for digital diagnosis according to an embodiment. FIG. 4 is a flowchart illustrating a series of processing steps performed by a processing unit related to determination of ALS onset according to an embodiment. FIG. 5 is a diagram illustrating a list of causative or associated genes of ALS. FIG. 6 is a diagram illustrating calculation results of HSIC scores for combinations of three causative genes. FIG. 7 is a diagram illustrating ROC evaluation results for combinations of causative genes. FIG. 8 is a diagram illustrating calculation results of HSIC scores for combinations of three associated genes. FIG. 9 is a diagram illustrating ROC evaluation results for combinations of associated genes. FIG. 10 is a list of gene combinations sorted in descending order of HSIC score. FIG. 11 is a list of gene combinations sorted in descending order of AUC calculated by logistic regression. FIG. 12 is a diagram illustrating the top 50 genes with the highest appearance frequency. FIG. 13 is a diagram illustrating calculation results of HSIC scores when the number of gene combinations is increased to four. FIG. 14 is a diagram illustrating the expression levels of PRKARIA in healthy individuals and ALS patients. 1 is a diagram showing the expression levels of QPCT in healthy subjects and ALS patients. FIG. 2 is a diagram showing the expression levels of TMEM71 in healthy subjects and ALS patients. FIG. 3 is a diagram showing a feature space. FIG. 4 is a diagram showing a feature space. FIG. 5 is a diagram showing the results of ROC evaluation for the combination of PRKAR1A, QPCT, and TMEM71. FIG. 6 is a diagram showing the correlation between the expression level of PRKAR1A and survival time. FIG. 7 is a diagram showing the correlation between the expression level of QPCT and survival time. FIG. 8 is a diagram showing the correlation between the expression level of TMEM71 and survival time. FIG. 9 is a diagram showing the correlation between the expression level of PRKAR1A and age at onset. FIG. 10 is a diagram showing the correlation between the expression level of QPCT and age at onset. FIG. 11 is a diagram showing the correlation between the expression level of TMEM71 and age at onset. FIG. 12 is a diagram showing the correlation between the expression level of PRKAR1A and bulbar paralysis type and systemic type. FIG. 13 is a diagram showing the correlation between the expression level of QPCT and bulbar paralysis type and systemic type. FIG. 14 is a diagram showing the correlation between the expression level of TMEM71 and bulbar paralysis type and systemic type. FIG. 1 is a diagram comparing the expression levels of PRKAR1A between healthy individuals and ALS patients.1 is a diagram comparing the expression levels of QPCT in healthy individuals and ALS patients. FIG. 2 is a diagram comparing the expression levels of TMEM71 in healthy individuals and ALS patients. FIG. 3 is a diagram showing the results of ROC evaluation for the combination of PRKAR1A, QPCT, and TMEM71 extracted from a small number of cases. FIG. 4 is an image of iPS cells and motor neurons obtained from the iPS cells. FIG. 5 is a diagram comparing the expression levels of PRKAR1A in motor neurons of healthy individuals and ALS patients. FIG. 6 is a diagram comparing the expression levels of QPCT in motor neurons of healthy individuals and ALS patients. FIG. 7 is a diagram comparing the gene expression levels of TMEM71 in motor neurons of healthy individuals and ALS patients. FIG. 8 is a diagram showing the results of ROC evaluation for the combination of PRKAR1A, QPCT, and TMEM71 extracted from motor neurons. FIG. 1 is a diagram showing an example of the relative expression levels of TDP-43 with respect to each of the PRKAR1A, QPCT, and TMEM71 genes. FIG. 1 is a graph showing the expression levels of PRKAR1A extracted from each of healthy individuals and ALS patients. FIG. 2 is a graph showing the expression levels of QPCT extracted from each of healthy individuals and ALS patients. FIG. 3 is a graph showing the expression levels of TMEM71 extracted from each of healthy individuals and ALS patients. FIG. 4 is a graph showing the expression levels of SPG11 extracted from ALS causative genes and ALS-associated genes. FIG. 5 is a graph showing the expression levels of CHMP2B extracted from ALS causative genes and ALS-associated genes. FIG. 6 is a graph showing the expression levels of CSNK1G3 extracted from ALS causative genes and ALS-associated genes. FIG. 7 is a graph showing the expression levels of DYNC1H1 extracted from ALS causative genes and ALS-associated genes.

[0009] Hereinafter, an information processing apparatus, an information processing method, and a program according to an embodiment will be described with reference to the drawings.

[0010] [Overview] Figure 1 is a diagram for explaining an overview of this embodiment. As shown in Figure 1, in this embodiment, gene expression levels in peripheral blood mononuclear cells (PBMCs) from healthy individuals and patients with a multifactorial disease or a sporadic disease are analyzed, and a high-dimensional nonlinear model is used to select a combination of genes for classifying healthy individuals and patients with a multifactorial disease or a sporadic disease based on the gene expression levels. The multifactorial disease or sporadic disease is, for example, ALS, but is not limited to this, and may be Alzheimer's disease, Parkinson's disease, or the like. Preferably, sporadic ALS is used. In the following, as an example, the multifactorial disease or sporadic disease will be described as "sporadic ALS."

[0011] In addition, "multifactorial disease" is defined as a disease that is thought to develop due to the interaction of genetic predisposition and environmental factors, and "sporadic disease" is defined as a disease with no family history. However, since there are many cases where the same disease corresponds to both "multifactorial disease" and "sporadic disease", in this field, "multifactorial disease" and "sporadic disease" are used almost synonymously. "Sporadic ALS" is also a multifactorial disease.

[0012] 2 is a diagram showing an example of the configuration of the information processing device 100 according to the embodiment. As shown in the figure, the information processing device 100 includes, for example, a communication interface 110, an input interface 120, an output interface 130, a processing unit 140, and a storage unit 150.

[0013] The communication interface 110 communicates with an external device via a network such as a wide area network (WAN) or a local area network (LAN). For example, the communication interface 110 includes a network interface card (NIC) or a wireless communication module. The external device may be, for example, a personal computer or a server installed in a facility (e.g., a research institute, a university, or a company) where research or drug development is conducted.

[0014] The input interface 120 accepts various input operations by the user and outputs an electrical signal corresponding to the accepted input operation to the processing unit 140. For example, the input interface 120 is a mouse, a keyboard, a touch panel, a drag ball, a switch, a button, or the like.

[0015] The output interface 130 is, for example, a display or a speaker. The display may be, for example, an LCD (Liquid Crystal Display) or an organic EL (Electro Luminescence) display. The display may be a touch panel that is integrated with the input interface 120.

[0016] The processing unit 140 is realized by a processor such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit) executing a program stored in the storage unit 150. Some or all of the functions of the processing unit 140 may be realized by hardware such as an LSI (Large Scale Integration), an ASIC (Application Specific Integrated Circuit), or an FPGA (Field-Programmable Gate Array), or may be realized by a combination of software and hardware. Each function of the processing unit 140 will be described later.

[0017] The storage unit 150 is realized by, for example, a hard disk drive (HDD), a flash memory, an electrically erasable programmable read-only memory (EEPROM), a read-only memory (ROM), a random access memory (RAM), etc. The storage unit 150 stores various programs such as firmware and application programs.

[0018] [Identification of Gene Combinations] FIG. 3 is a flowchart showing the flow of a series of processes performed by the processing unit 140 regarding identification of gene combinations according to this embodiment.

[0019] First, the processing unit 140 selects a combination of a predetermined number of genes from a group of causative genes for ALS (step S100). Thirty-three genes, such as SOD1, ALS2, ALS3, and SETX, are known as causative genes for ALS (see FIG. 7 for details, described later). The processing unit 140 selects a combination of a predetermined number of genes from the 33 causative genes. The predetermined number is preferably three, but is not limited thereto, and may be, for example, two or four or more. Hereinafter, the description will be given assuming that the predetermined number is "3" as an example. For example, when selecting any three genes from the 33 causative genes, the processing unit 140 may select 5,456 combinations.

[0020] Next, the processing unit 140 calculates a measure of the dependency (or independence) between the causative genes combined in S100 (step S102).

[0021] Gene expression analysis generally uses linear models such as linear logistic regression and Hotelling's t-test. However, life phenomena are considered to be nonlinear science, and the pathology of a disease cannot be explained by a single factor. Therefore, in this embodiment, gene expression is analyzed using a nonlinear model.

[0022] For example, the processing unit 140 uses HSIC (Hilbert-Schmidt Independence Criterion), which is a type of machine learning that can detect nonlinear structures in high-dimensional data, to calculate the HSIC score as a measure of dependency between combinations of causative genes.

[0023] 4 is a diagram illustrating a method for calculating the HSIC score. As shown in the figure, the processing unit 140 distributes combinations of ALS causative genes (vector data representing each of the three causative genes) in a (reproducing kernel) Hilbert space (the feature space in the figure), and calculates the HSIC scores of the causative genes in the Hilbert space. For example, the processing unit 140 calculates the HSIC scores for all 5,456 combinations.

[0024] Next, the processing unit 140 selects a combination of a predetermined number of genes from the group of ALS-related genes (step S104). In addition to the causative genes described above, 126 genes, such as APEX1, APOE, AR, and CCS, are known as ALS-related genes (see FIG. 7 for details, described later). The processing unit 140 selects a combination of a predetermined number of genes from the 126 related genes. As described above, the predetermined number is preferably three, but is not limited thereto and may be, for example, two or four or more. For example, when selecting any three genes from the 126 related genes, the processing unit 140 may select 325,500 combinations.

[0025] Next, the processing unit 140 calculates a measure of the dependency (or independence) between the related genes combined in S104 (step S106). As in the case of the causative genes, the processing unit 140 calculates the HSIC score for all 325,500 combinations.

[0026] Next, the processing unit 140 selects the gene combination with the highest HSIC score from among the gene combinations for which the HSIC scores have been calculated (step S108).

[0027] For example, the processing unit 140 performs a linear regression analysis such as logistic regression to eliminate the influence of multicollinearity, and selects or extracts specific combinations including genes with a high frequency of occurrence (number of occurrences) from a set of multiple combinations for which HSIC scores have been calculated (hereinafter referred to as a combination population). For example, the processing unit 140 may select or extract combinations including genes whose frequency of occurrence is equal to or greater than a threshold as specific combinations (in other words, combinations to be excluded). The threshold is, for example, 10, but is not limited to this and may be any other value.

[0028] The processing unit 140 excludes specific combinations containing genes with high occurrence frequencies from the combination population. The processing unit 140 selects the gene combination with the highest HSIC score from the combination population from which the specific combinations have been excluded. As described in the examples below, the combination of PRKAR1A, QPCT, and TMEM71 is selected from combinations of ALS causative genes or related genes as the gene combination with the highest HSIC score. This completes the series of processes related to identifying the gene combination.

[0029] [Generation of Molecular Biomarkers for Digital Diagnosis] FIG. 5 is a flowchart showing the flow of a series of processes performed by the processing unit 140 regarding generation of molecular biomarkers for digital diagnosis according to the embodiment.

[0030] First, the processing unit 140 distributes genetic data of each of a plurality of healthy individuals (hereinafter also referred to as a group of healthy individuals) in a three-dimensional feature space whose dimensions are the expression levels of these three genes, based on the expression levels of PRKAR1A, QPCT, and TMEM71 of each of these healthy individuals (step S200). For example, the genetic data of the healthy individuals distributed in the feature space may be represented as a three-dimensional vector (e1, e2, e3) in which the expression level of PRKAR1A is the first element e1, the expression level of QPCT is the second element e2, and the expression level of TMEM71 is the third element e3.

[0031] Next, the processing unit 140 distributes the genetic data of each ALS patient on a three-dimensional feature space, which has the expression levels of these three genes as dimensions, based on the expression levels of PRKAR1A, QPCT, and TMEM71 of each of the ALS patients (hereinafter also referred to as an ALS patient group) (step S202). The genetic data of the ALS patients distributed on the feature space may also be represented as a three-dimensional vector (e1, e2, e3), similar to the genetic data of healthy individuals.

[0032] Next, the processing unit 140 clusters the genetic data of healthy individuals and the genetic data of ALS patients in a three-dimensional feature space (step S204). For example, as shown in Figures 19A and 19B in the Examples described below, in a three-dimensional feature space in which the expression levels of RKAR1A, QPCT, and TMEM71 are defined as dimensions, the processing unit 140 classifies the genetic data of healthy individuals (Health Control in the figures) and the genetic data of ALS patients (ALS in the figures) into clusters.

[0033] Next, the processing unit 140 stores the clusters of genetic data of healthy individuals and the clusters of genetic data of ALS patients formed in the feature space in the storage unit 150 as molecular biomarkers for digital diagnosis (step S206), thereby completing a series of processes related to the generation of molecular biomarkers for digital diagnosis.

[0034] [Determination of ALS Onset] FIG. 6 is a flowchart showing the flow of a series of processes performed by the processing unit 140 regarding the determination of ALS onset according to this embodiment.

[0035] First, the processing unit 140 acquires genetic data of a subject to be diagnosed with ALS (step S300). The genetic data of the subject may be expressed as a three-dimensional vector (e1, e2, e3) as described above.

[0036] Next, the processing unit 140 distributes the genetic data of the subject on a feature space in which clusters (clusters of healthy subjects and clusters of ALS patients) that are molecular biomarkers are formed (step S302).

[0037] Next, the processing unit 140 calculates the distance D1 between the subject's genetic data and the cluster of healthy individuals in the feature space, and calculates the distance D2 between the subject's genetic data and the cluster of ALS patients (step S304).

[0038] Next, the processing unit 140 determines whether the subject will develop ALS at some point in the future, or whether the subject has already developed ALS at the present time, based on the respective distances between the two clusters (step S306).

[0039] For example, if the distance D2 to the cluster of ALS patients is shorter than the distance D1 to the cluster of healthy individuals (D1 > D2), that is, if the subject's genetic data is closer to the cluster of ALS patients than to the cluster of healthy individuals, the processing unit 140 may determine that the subject will develop ALS at some point in the future, or that the subject has already developed ALS at the present time.

[0040] On the other hand, if the distance D2 to the cluster of ALS patients is longer than the distance D1 to the cluster of healthy individuals (D1 < D2), that is, if the subject's genetic data is closer to the cluster of healthy individuals than to the cluster of ALS patients, the processing unit 140 may determine that the subject will not develop ALS at some point in the future and that the subject does not currently have ALS.

[0041] Next, the processing unit 140 outputs the determination result regarding whether or not the subject has developed ALS (step S308).

[0042] For example, the processing unit 140 may transmit the determination result to an external device via the communication interface 110, or may output the determination result via the output interface 130 (e.g., a display). This completes the series of processes related to determining the onset of ALS.

[0043] According to the embodiment described above, the information processing device 100 calculates a measure of mutual dependency (e.g., HSIC score) between the genes included in each combination of ALS causative or related genes, and selects the combination with the highest measure from among the multiple combinations for which the measure has been calculated (i.e., from among the combination population), thereby making it possible to identify genes capable of diagnosing ALS.

[0044] Furthermore, according to the above-described embodiment, the information processing device 100 distributes genetic data of healthy individuals in a feature space based on the expression levels of genes derived from healthy individuals (genes included in the combination with the highest scale described above), and further distributes genetic data of ALS patients in the same feature space based on the expression levels of genes derived from ALS patients (genes included in the combination with the highest scale described above).The information processing device 100 then clusters the genetic data of healthy individuals and the genetic data of ALS patients in the feature space.This makes it possible to generate molecular biomarkers for digital diagnosis in the feature space.

[0045] Furthermore, according to the above-described embodiment, the information processing device 100 acquires genetic data of a subject to be diagnosed with ALS and distributes the genetic data of the subject in a feature space in which clusters of healthy individuals and ALS patients are formed. The information processing device 100 calculates the distance between each cluster and the genetic data of the subject in the feature space, and based on the distance, determines whether the subject will develop ALS at some point in the future or whether the subject has already developed ALS at the present time. This allows for accurate determination of whether ALS has developed.

[0046] The above describes the form for carrying out the present invention using an embodiment, but the present invention is not limited to such an embodiment, and various modifications and substitutions can be made within the scope that does not deviate from the gist of the present invention.

[0047] The above-described embodiment can be expressed as follows: (Supplementary Note 1) An information processing device comprising a processing unit that calculates, for each combination of genes included in a gene dataset, a measure of dependency on causative genes of a multifactorial disease or a sporadic disease and on genes associated with the multifactorial disease or the sporadic disease, and selects a predetermined number of gene combinations from the dataset based on the measure.

[0048] (Supplementary Note 2) The information processing device according to Supplementary Note 2, wherein the processing unit distributes data of a first gene, which is a gene derived from a healthy individual, on a certain feature space based on the expression level of the first gene; distributes data of a second gene, which is a gene derived from a patient with the multifactorial disease or the sporadic disease, on the feature space based on the expression level of the second gene; clusters the data of the first gene on the feature space based on the expression level of the first gene; and clusters the data of the second gene on the feature space based on the expression level of the second gene.

[0049] (Supplementary Note 3) The information processing device described in Supplementary Note 2, wherein the processing unit distributes data of a third gene, which is a gene derived from a subject to be diagnosed with the multifactorial disease or the sporadic disease, on the feature space in which clusters of the healthy individuals and the patients are formed, calculates a distance between the data of the third gene and the cluster on the feature space, and determines whether the subject will develop the multifactorial disease or the sporadic disease, or whether the subject has developed the multifactorial disease or the sporadic disease, based on the distance.

[0050] (Appendix 4) The information processing device according to appendix 1 or 2, wherein the multifactorial disease or the sporadic disease includes amyotrophic lateral sclerosis, the predetermined number is three, and the combination of the predetermined number of genes includes at least PRKAR1A, QPCT, and TMEM71.

[0051] (Supplementary Note 5) The information processing device according to Supplementary Note 1 or 2, wherein the processing unit performs linear regression analysis to exclude specific combinations including genes whose occurrence frequency is equal to or greater than a threshold from a population which is a set of gene combinations for which the measure has been calculated, and selects the gene combination with the highest measure from the population from which the specific combination has been excluded.

[0052] (Supplementary Note 6) The information processing device according to Supplementary Note 1 or 2, wherein the processing unit distributes the dataset in a Hilbert space, calculates a Hilbert-Schmidt dependency measure as the measure for each combination of the genes included in the dataset distributed in the Hilbert space, and selects the combination with the highest Hilbert-Schmidt dependency measure from among the plurality of combinations for which the Hilbert-Schmidt dependency measure has been calculated.

[0053] (Supplementary Note 7) An information processing method using a computer, comprising: calculating, for each combination of genes included in a gene dataset, a measure of dependency on causative genes of a multifactorial disease or a sporadic disease and on genes associated with the multifactorial disease or the sporadic disease; and selecting a predetermined number of gene combinations from the dataset based on the measure.

[0054] (Appendix 8) The information processing method according to Appendix 7, further comprising: distributing data of a first gene, which is a gene derived from a healthy individual, on a feature space based on the expression level of the first gene; distributing data of a second gene, which is a gene derived from a patient with the multifactorial disease or the sporadic disease, on the feature space based on the expression level of the second gene; clustering the data of the first gene on the feature space based on the expression level of the first gene; and clustering the data of the second gene on the feature space based on the expression level of the second gene.

[0055] (Appendix 9) The information processing method according to Appendix 8, further comprising: distributing data of a third gene, which is a gene derived from a subject to be diagnosed with the multifactorial disease or the sporadic disease, on the feature space in which clusters of the healthy individuals and the patients are formed; calculating the distance between the data of the third gene and the cluster on the feature space; and determining whether the subject will develop the multifactorial disease or the sporadic disease based on the distance, or determining whether the subject has developed the multifactorial disease or the sporadic disease.

[0056] (Appendix 10) The information processing method according to appendix 7 or 8, wherein the multifactorial disease or the sporadic disease includes amyotrophic lateral sclerosis, the predetermined number is three, and the combination of the predetermined number of genes includes at least PRKAR1A, QPCT, and TMEM71.

[0057] (Supplementary Note 11) The information processing method according to Supplementary Note 7 or 8, further comprising: performing linear regression analysis to exclude specific combinations including genes whose occurrence frequency is equal to or greater than a threshold from a population that is a set of gene combinations for which the measure has been calculated; and selecting the gene combination with the highest measure from the population from which the specific combination has been excluded.

[0058] (Supplementary Note 12) A program to be executed by a computer, the program comprising: calculating, for each combination of genes included in a gene dataset, a measure of dependency on causative genes of a multifactorial disease or a sporadic disease and on genes associated with the multifactorial disease or the sporadic disease; and selecting a predetermined number of gene combinations from the dataset based on the measure.

[0059] (Appendix 13) The program described in Appendix 12, further comprising: distributing data of a first gene, which is a gene derived from a healthy individual, on a feature space based on the expression level of the first gene; distributing data of a second gene, which is a gene derived from a patient with the multifactorial disease or the sporadic disease, on the feature space based on the expression level of the second gene; clustering the data of the first gene on the feature space based on the expression level of the first gene; and clustering the data of the second gene on the feature space based on the expression level of the second gene.

[0060] (Appendix 14) The program described in Appendix 13, further comprising: distributing data of a third gene, which is a gene derived from a subject to be diagnosed with the multifactorial disease or the sporadic disease, on the feature space in which clusters of the healthy individuals and the patients are formed; calculating the distance between the data of the third gene and the cluster on the feature space; and determining whether the subject will develop the multifactorial disease or the sporadic disease, or determining whether the subject has developed the multifactorial disease or the sporadic disease, based on the distance.

[0061] (Appendix 15) The program according to appendix 12 or 13, wherein the multifactorial disease or the sporadic disease includes amyotrophic lateral sclerosis, the predetermined number is three, and the combination of the predetermined number of genes includes at least PRKAR1A, QPCT, and TMEM71.

[0062] (Supplementary Note 16) The program according to Supplementary Note 12 or 13, further comprising: performing linear regression analysis to exclude specific combinations including genes whose occurrence frequency is equal to or greater than a threshold from a population that is a set of gene combinations for which the measure has been calculated; and selecting the gene combination with the highest measure from the population from which the specific combination has been excluded.

[0063] (Supplementary Note 17) A computer-readable storage medium storing the program according to Supplementary Note 12 or 13.

[0064] Experimental Example 1 (Microarray Data and Normalization) Gene expression data (GSE112676, 233 ALS and 508 CTL) were used for HSIC analysis. Gene expression signals were normalized using raw expression intensities and detection p-values ​​(GSE112676_HT12_V3_preQC_nonnormalized.txt) downloaded from R using the limma package (v3.32.10) functions (backgroundCorrect and normalizeBetweenArrays). Batch effects were removed using the ComBat algorithm implemented in the sva package (v3.35.2) in R. One sample (GSM3077426) that remained outliers even after batch effect correction was excluded from further analysis. For the HSIC prediction shown in Figure 1 above, ALS-related genes from the ALS online database (ALSoD, https: / / alsod.ac.uk / ) were used. Unbiased HSIC prediction was performed using the top 1,000 differentially expressed genes among those detectable in 20% or more of the samples.

[0065] Experimental Example 2 (Preparation of Pluripotent Stem Cells) Pluripotent stem cells are prepared. Examples of pluripotent stem cells include embryonic stem cells (ES cells), induced pluripotent stem cells (iPS cells), cloned embryonic stem (ntES) cells obtained by nuclear transfer, spermatogonial stem cells (GS cells), and embryonic germ cells (EG cells). Preferred pluripotent stem cells are ES cells, iPS cells, and ntES cells. More preferred pluripotent stem cells are human pluripotent stem cells, with human ES cells and human iPS cells being particularly preferred. Furthermore, cells usable in the present invention may not only be pluripotent stem cells, but also cell populations induced by so-called "direct reprogramming," in which differentiation into desired cells is directly induced without going through pluripotent stem cells. Human iPS cells were used in this experiment. Hereinafter, unless otherwise specified, iPS cells are assumed to be human iPS cells.

[0066] iPS cells were generated from fibroblasts or PBMCs of healthy individuals and sporadic ALS patients using episomal vectors of OCT3 / 4, Sox2, Klf4, L-Myc, Lin28, and dominant-negative p53, or OCT3 / 4, Sox2, Klf4, L-Myc, Lin28, and shRNA for p53, respectively. They were cultured in a feeder-free, xeno-free culture system using StemFit (Ajinomoto) supplemented with penicillin / streptomycin.

[0067] [Experimental Example 3] (Differentiation of iPS cells into motor neurons) Motor neurons were differentiated from iPS cells. Specifically, iPS cells were dissociated into single cells and rapidly reaggregated in a low-cell-adhesion U-shaped 96-well plate (Lipidule-Coated Plate A-U96, NOF Corporation, Tokyo, Japan).

[0068] Aggregations were grown in 5% KSR (Invitrogen, Waltham, MA), minimal essential medium-non-essential amino acids (Invitrogen), L-glutamine (Sigma-Aldrich, St. Louis, MO), 2-mercaptoethanol (Wako, Osaka, Japan), 2 μM dorsomorphin (Sigma-Aldrich), 10 μM SB431542 (Cayman, Ann Arbor, MI), 3 μM CHIR99021 (Cayman), and 12.5 ng / mL fibroblast growth factor (Wako) for 11 days at the neural induction stage.

[0069] On day 4, 100 nM retinoic acid (Sigma-Aldrich) and 500 nM Smoothened Ligand (Enzo Life Sciences, Farmingdale, NY) were added. After patterning in Neurobasal Medium supplemented with B27 Supplement (Thermo Fisher Scientific), 100 nM retinoic acid, 500 nM smoothened ligand, and 10 μM DAPT (Selleck, Houston, TX), the cells were separated into clusters using Accumax (Innovative Cell Technologies, San Diego, CA) on day 16, dissociated into single cells, and attached to Matrigel (BD Biosciences, Franklin Lakes, NJ)-coated dishes.

[0070] Adherent cells were cultured for 8 days in neurobasal medium containing 10 ng / ml brain-derived neurotrophic factor (R&D Systems, Minneapolis, MN), 10 ng / ml glial cell line-derived neurotrophic factor (R&D Systems), and 10 ng / ml neurotrophin-3 (R&D Systems). On day 21, cells were dissociated into single cells using an Accumax and seeded at 2 × 10 cells / well onto iMatrix-coated 24-well plates (Corning).

[0071] Experimental Example 4 (Quantitative RT-PCR) Total RNA was extracted from cultured cells using an RNeasy Plus Mini kit (QIAGEN). 1 μg of RNA was reverse transcribed using ReverTra Ace (TOYOBO, Osaka, Japan). Quantitative PCR analysis was performed using SYBR Premix Ex TaqII (TAKARA) and StepOnePlus (Thermo Fisher Scientific) for reverse transcription.

[0072] [Experimental Example 5] (Statistical Analysis) The results were analyzed using Student's t-test to determine statistical significance. Differences of p<0.05 were considered significant. The analysis was performed using GraphPad Prism software version 8.0 for Windows (GraphPad Software, San Diego, CA).

[0073] (Results) [Experimental Example 6] Gene combinations for classifying healthy subjects and ALS patients were selected by analyzing gene expression levels in peripheral blood mononuclear cells (PBMCs). As described in the above embodiment, gene expression levels were analyzed using a nonlinear model, and HSIC was used for the analysis.

[0074] Gene combinations that show differences between healthy individuals and ALS patients result in high HSIC scores, while gene combinations with no differences result in HSIC scores close to 0. By identifying combinations that produce high HSIC scores, the researchers extracted genes that distinguish between healthy individuals and ALS patients.

[0075] First, the validity of the method described in this embodiment was verified using a group of genes known to be associated with ALS. Fig. 7 is a diagram showing a list of ALS causative genes or related genes. For example, 33 causative genes known to be causative genes of ALS were selected from the genes shown in Fig. 7, and further, combinations of three genes were selected from the 33 causative genes.

[0076] The HSIC score was calculated as a measure for classifying healthy subjects and ALS patients based on the expression levels of the three genes. Figure 8 shows the results of calculating the HSIC scores for combinations of three causative genes. Figure 8 shows only the top 15 combinations with the highest HSIC scores. The causative gene combination with the highest HSIC score is SPG11, CHMP2B, and VCP (HSIC score 0.0988).

[0077] Of the total combinations of 33 ALS causative genes (5,456 combinations), the combination with the highest HSIC score was SPG11, CHMP2B, and VCP (HSIC score 0.0988). These three causative gene combinations were evaluated using ROC (Receiver Operating Characteristics). Figure 9 shows the ROC evaluation results for the causative gene combinations. As shown in Figure 9, the combination of SPG11, CHMP2B, and VCP with the highest HSIC score had an AUC (area under the curve) of 0.75 in ROC. Therefore, the results showed a statistically significant difference in AUC.

[0078] Next, three gene combinations were similarly selected from the 126 ALS-related genes (see Figure 7). The HSIC scores for the three related gene combinations were calculated for the healthy control group, while the HSIC scores for the three related gene combinations were calculated for the ALS patient group. Figure 10 shows the results of calculating the HSIC scores for the three related gene combinations. Similar to Figure 8, Figure 10 shows only the top 15 combinations with the highest HSIC scores. The related gene combination with the highest HSIC score is CSNK1G3, CHMP2B, and DYNC1H1 (HSIC score 0.11365).

[0079] Of the total combinations of 126 related genes (325,500 combinations), the combination with the highest HSIC score was CSNK1G3, CHMP2B, and DYNC1H1 (HSIC score 0.11365). These three related gene combinations were evaluated using ROC. Figure 11 shows the ROC evaluation results for the related gene combinations. As shown in Figure 11, the combination of CSNK1G3, CHMP2B, and DYNC1H1 with the highest HSIC score had an AUC (AUC for classifying healthy subjects and ALS patients) of 0.75 in ROC. Therefore, the results showed a statistically significant difference in AUC.

[0080] These results demonstrate the validity of the method of this embodiment for finding a gene set that can classify a group of healthy subjects and a group of ALS patients.

[0081] [Experimental Example 7] To investigate unknown causes of ALS, HSIC scores were calculated for gene combinations from genes not known to be associated with ALS (non-associated genes). To avoid multicollinearity, which is a problem in multiple regression models, analysis was performed using linear regression, and genes that repeatedly appeared in the gene list extracted by the analysis (genes with high frequency of appearance) were excluded from the causative genes or associated genes of ALS.

[0082] On the other hand, we have listed gene combinations that distinguish between healthy individuals and ALS patients using logistic regression, a linear regression model. Figure 12 shows a list of gene combinations sorted in descending order of HSIC score.

[0083] When examining genes that frequently appear in logistic regression, it was found that there was a bias in the frequency of gene appearance. Figure 13 is a list of gene combinations arranged in descending order of AUC obtained by logistic regression. Figure 14 is a diagram showing the top 50 genes with the highest frequency of appearance. As shown in Figure 14, TPT1 appeared 25 times, ATP5I appeared 39 times, CAPZA2 appeared 17 times, and RPL22 appeared 11 times.

[0084] To eliminate the effects of multicollinearity, we excluded genes that appeared repeatedly 10 or more times in linear regression among gene combinations with high HSIC scores. In the results of Figure 14, TPT1, ATP5I, CAPZA2, and RPL22 all appeared 10 or more times. Therefore, we excluded combinations that included any one of these four genes. The gene combination with the highest HSIC score was PRKAR1A, QPCT, and TMEM71 (the ninth combination from the top in Figure 12).

[0085] [Experimental Example 8] Furthermore, we investigated whether the classification accuracy of ALS would improve if the number of gene combinations was increased to four (if the predetermined number was changed from three to four). Figure 15 shows the calculation results of the HSIC score when the number of gene combinations was increased to four. When combining four genes, the HSIC score did not change significantly compared to when combining three genes. Therefore, we selected a combination of three genes.

[0086] [Experimental Example 9] Next, the expression levels of PRKAR1A, QPCT, and TMEM71 in PBMCs of healthy subjects and ALS patients were compared. Figure 16 is a diagram showing the expression levels of PRKAR1A in healthy subjects and ALS patients, respectively. Figure 17 is a diagram showing the expression levels of QPCT in healthy subjects and ALS patients, respectively. Figure 18 is a diagram showing the expression levels of TMEM71 in healthy subjects and ALS patients, respectively. For all three genes, the expression levels were higher in ALS patients than in healthy subjects.

[0087] In addition, the genes of healthy individuals and the genes of ALS patients were distributed on a three-dimensional feature space, with the expression levels of PRKAR1A, QPCT, and TMEM71 as dimensions. Figures 19A and 19B are diagrams showing the feature space. As shown in Figures 19A and 19B, in the three-dimensional feature space, the genes of healthy individuals were classified into the same cluster, and the genes of ALS patients were classified into the same cluster.

[0088] The combination of PRKAR1A, QPCT, and TMEM71 was evaluated using ROC. Figure 20 shows the results of ROC evaluation for the combination of PRKAR1A, QPCT, and TMEM71. As shown in Figure 20, the combination of PRKAR1A, QPCT, and TMEM71 had an AUC (AUC for classifying healthy subjects and ALS patients) of 0.83 in ROC. Therefore, the results showed that there was a statistically significant difference in AUC.

[0089] [Experimental Example 10] Furthermore, the correlation between the expression levels of PRKAR1A, QPCT, and TMEM71 genes and clinical information of ALS obtained from published data was investigated. Figure 21A is a diagram showing the correlation between the expression level of PRKAR1A and survival time. Figure 21B is a diagram showing the correlation between the expression level of QPCT and survival time. Figure 21C is a diagram showing the correlation between the expression level of TMEM71 and survival time. Figure 22A is a diagram showing the correlation between the expression level of PRKAR1A and age of onset. Figure 22B is a diagram showing the correlation between the expression level of QPCT and age of onset. Figure 22C is a diagram showing the correlation between the expression level of TMEM71 and age of onset. Figure 23A is a diagram showing the correlation between the expression level of PRKAR1A and bulbar paralysis type and systemic paralysis type. Figure 23B is a diagram showing the correlation between the expression level of QPCT and bulbar paralysis type and systemic paralysis type. FIG. 23C is a diagram showing the correlation between the expression level of TMEM71 and bulbar and systemic types.

[0090] As shown in Figures 21A, 21B, and 21C, the gene expression levels of PRKAR1A and TMEM71 were correlated with survival time, although QPCT did not show a significant correlation. As shown in Figures 22A, 22B, and 22C, there was no correlation between the age of onset and the expression levels of the three genes. As shown in Figures 23A, 23B, and 23C, although there was no difference in QPCT, PRKAR1A and TMEM71 showed significantly higher expression levels in patients with generalized ALS than in patients with bulbar ALS.

[0091] [Experimental Example 11] Furthermore, the expression levels of three genes, PRKAR1A, QPCT, and TMEM71, were confirmed using PBMCs from healthy individuals and ALS patient PBMCs in our possession. PBMCs were collected from 12 ALS patients and 12 healthy individuals, and RNA was extracted. Figure 24A shows a comparison of the expression levels of PRKAR1A in healthy individuals with those in ALS patients. Figure 24B shows a comparison of the expression levels of QPCT in healthy individuals with those in ALS patients. Figure 24C shows a comparison of the expression levels of TMEM71 in healthy individuals with those in ALS patients. As shown in Figures 24A, 24B, and 24C, the expression levels of PRKAR1A and QPCT were significantly higher in ALS patients than in healthy individuals, and the expression level of TMEM71 tended to be higher in ALS patients.

[0092] Figure 25 shows the ROC evaluation results for the combination of PRKAR1A, QPCT, and TMEM71 extracted from a small number of cases. As shown in Figure 25, by combining the expression levels of the three genes PRKAR1A, QPCT, and TMEM71, even with a small number of cases (12 ALS patients and 12 healthy subjects), the AUC for classifying healthy subjects from ALS patients was 0.85. These results confirmed that examining the expression levels of the three genes in PBMCs can distinguish between ALS patient groups and healthy subjects.

[0093] [Experimental Example 12] The expression of three genes was examined in motor neurons derived from iPS cells established from 26 healthy individuals and 18 ALS patients. Figure 26 shows images of iPS cells and motor neurons derived from the iPS cells. Figure 27A shows a comparison of the expression levels of PRKAR1A between motor neurons from healthy individuals and motor neurons from ALS patients. Figure 27B shows a comparison of the expression levels of QPCT between motor neurons from healthy individuals and motor neurons from ALS patients. Figure 27C shows a comparison of the expression levels of TMEM71 between motor neurons from healthy individuals and motor neurons from ALS patients. Figure 28 shows the results of ROC evaluation of the combination of PRKAR1A, QPCT, and TMEM71 extracted from motor neurons. As shown in Figures 27A, 27B, and 27C, there was no difference in the expression levels of the PRKAR1A, QPCT, and TMEM71 genes between motor neurons of healthy subjects and motor neurons of ALS patients. As shown in Figure 28, when the expression levels of the three genes were compared as a set, the AUC was 0.79, and healthy subjects and ALS patients could be classified.

[0094] [Experimental Example 13] Furthermore, because the accumulation of TDP-43 is deeply involved in the pathology of ALS, the relationship between three genes and TDP-43 was investigated. Figure 29 is a diagram showing an example of the relative expression levels of TDP-43 for each of the PRKAR1A, QPCT, and TMEM71 genes. As shown in Figure 29, when PRKAR1A, QPCT, and TMEM71 were knocked down with siRNA, the expression levels of TDP-43 significantly increased in motor neurons derived from both healthy individuals and ALS patients. From these results, the gene group identified using HSIC in the method of this embodiment has not previously been known to be associated with the pathology of ALS, but this experiment revealed that it may be a new player in the pathology of ALS.

[0095] [Experimental Example 14] The expression levels of three genes extracted from each of healthy individuals and ALS patients were graphed. Figure 30A is a graph showing the expression levels of PRKARIA extracted from each of healthy individuals and ALS patients. Figure 30B is a graph showing the expression levels of QPCT extracted from each of healthy individuals and ALS patients. Figure 30C is a graph showing the expression levels of TMEM71 extracted from each of healthy individuals and ALS patients. The vertical axis of each graph represents the gene expression level, and the horizontal axis represents the case.

[0096] For each of PRKAR1A, QPCT, and TMEM71, there was a tendency for expression levels to be higher in ALS patients compared to healthy individuals, but there was large variation between samples, and it was not possible to distinguish between healthy individuals and ALS patients based on the expression levels of each gene alone.

[0097] On the other hand, as explained in the above embodiment, it is possible to classify healthy individuals from ALS patients by combining the expression levels of these three genes, PRKAR1A, QPCT, and TMEM71. Therefore, the usefulness of the combination of the three genes extracted by HSIC was demonstrated.

[0098] Similarly, the expression levels of genes extracted from ALS causative genes and ALS-associated genes were graphed. Figure 31A is a graph showing the expression levels of SPG11 extracted from ALS causative genes and ALS-associated genes. Figure 31B is a graph showing the expression levels of CHMP2B extracted from ALS causative genes and ALS-associated genes. Figure 31C is a graph showing the expression levels of CSNK1G3 extracted from ALS causative genes and ALS-associated genes. Figure 31D is a graph showing the expression levels of DYNC1H1 extracted from ALS causative genes and ALS-associated genes. The vertical axis of each figure represents gene expression levels, and the horizontal axis represents cases. Even with these genes SPG11, CHMP2B, CSNK1G3, and DYNC1H1, it was not possible to classify healthy individuals from ALS patients using each gene alone, demonstrating the usefulness of a method for separating healthy individuals from ALS patients using a combination of the three genes PRKAR1A, QPCT, and TMEM71.

[0099] As described above, we used a machine learning algorithm called HSIC, a high-dimensional nonlinear statistical model, to discover blood molecular biomarkers necessary for digital diagnosis of ALS from real data. The identified molecular biomarkers had not previously attracted attention in ALS. However, we found that controlling the expression of these genes with siRNA changed the expression level of TDP-43, an important key molecule in ALS, indicating that these markers may be related to ALS.

[0100] HSIC is used to measure the statistical dependence between two random vectors. It transforms the two random vectors into two recurrent kernel Hilbert spaces (RKHSs) and then measures their statistical dependence using the Hilbert-Schmidt (HS) operator of these two RKHSs. Because ALS is a heterogeneous disease exhibiting nonlinear biological phenomena and its pathology cannot be explained by a single factor, we applied this model to identify gene combinations for classifying ALS patients and healthy controls using blood sample data. By utilizing a nonlinear model, we successfully identified a novel gene combination: PRKAR1A, QPCT, and TMEM71.

[0101] The PRKAR1A gene encodes a serine / threonine kinase, cAMP-dependent protein kinase type I-alpha regulatory subunit, which is a major mediator of cAMP signaling in mammals. Phosphorylation via the cAMP / PKA signaling pathway is induced by various intracellular physiological ligands and is critically involved in the regulation of metabolism, cell proliferation, differentiation, and apoptosis. Deficiency of one or both alleles of this gene causes multiple sclerosis syndrome in humans and embryonic lethal defects in mice. Although the relationship between the PRKAR1A gene and ALS has not been clarified, it has been reported that PKA activity is elevated in the spinal cord of ALS patients and SOD1 mice, and that synaptic repair by cAMP / PKA promotes activity-dependent neuroprotection of motor neurons in ALS. From these findings, increased expression of the PRKAR1A gene is expected to have a preventive effect against ALS. The Gene ID of the PRKAR1A gene, an example of the NCBI Reference Sequence, and the address of the NCBI reference site are as follows:

[0102] PRKAR1A Gene ID: 5573 NM_001276289.2, NM_212472.1 https: / / www.ncbi.nlm.nih.gov / gene / 5573

[0103] The QPTC gene encodes glutaminyl peptide cyclotransferase. It has been reported that glutaminyl cyclase expression is increased in the peripheral blood of Alzheimer's disease patients, and glutaminyl cyclase inhibitors may be a potential treatment for Alzheimer's disease. QPCT has also been identified as a therapeutic target for Huntington's disease, and polymorphisms in the QPCT gene have been reported to be associated with susceptibility to schizophrenia. Although its association with the pathology of ALS is unclear, increased expression of the QPCT gene may be related to a common pathway in neurodegeneration. The Gene ID, an example NCBI Reference Sequence, and the address of the NCBI reference site for the QPTC gene are listed below.

[0104] OPCT Gene ID: 25797 NM_012413.4, NM_012413.3 https: / / www.ncbi.nlm.nih.gov / gene / 25797

[0105] TMEM71 encodes a transmembrane protein, but its function has not been clearly elucidated. Knockout mice show only slight hypothyroidism and no phenotype. TMEM71 expression is increased in glioblastoma and is associated with immune responses, and TMEM71 showed a high positive correlation with PD-1 and PD-L1. Knockout mice show only slight hypothyroidism and no phenotype. Increased TMEM71 gene expression may be associated with immune responses to ALS. The Gene ID for the TMEM71 gene, an example NCBI Reference Sequence, and the NCBI reference site address are listed below.

[0106] TMEM71 Gene ID: 137835 NM_001145153.2, NM_144649.1 https: / / www.ncbi.nlm.nih.gov / gene / 137835

[0107] We used the nonlinear model HSIC and real data to identify gene combinations for classifying ALS. This method will not only be useful for identifying molecular biomarkers for digital diagnosis of ALS, but may also lead to a completely new mechanism for ALS pathogenesis that goes beyond human idea-driven approaches. Furthermore, this method is not limited to ALS, but can also be applied to other multifactorial or sporadic diseases.

[0108] 100... Information processing device, 110... Communication interface, 120... Input interface, 130... Output interface, 140... Processing unit, 150... Storage unit

Claims

1. For each combination of genes in the gene dataset, a measure of dependency on causative genes of a multifactorial disease or a sporadic disease and on associated genes of the multifactorial disease or the sporadic disease is calculated; selecting a combination of a predetermined number of genes from the dataset based on the measure; A processing unit is provided, The processing unit includes: distributing data on a first gene, which is a gene derived from a healthy subject, in a certain feature space based on the expression amount of the first gene; distributing data on a second gene, the second gene being a gene derived from a patient with the multifactorial disease or the sporadic disease, on the feature space based on an expression level of the second gene; clustering the data of the first gene in the feature space based on the expression level of the first gene; clustering the data of the second gene in the feature space based on the expression level of the second gene; Information processing device.

2. The processing unit includes: Distributing data on a third gene, which is a gene derived from a subject to be diagnosed with the multifactorial disease or the sporadic disease, on the feature space in which the clusters of the healthy subjects and the patients are formed; Calculating a distance between the data of the third gene and the cluster in the feature space; determining whether the subject will develop the multifactorial disease or the sporadic disease based on the distance, or determining whether the subject has developed the multifactorial disease or the sporadic disease; The information processing device according to claim 1 .

3. The multifactorial disease or the sporadic disease comprises amyotrophic lateral sclerosis; the predetermined number is three, The combination of the predetermined number of genes includes at least PRKAR1A, QPCT, and TMEM71.

3. The information processing device according to claim 1 or 2.

4. The processing unit includes: By performing a linear regression analysis, a specific combination including a gene whose occurrence frequency is equal to or higher than a threshold is excluded from a population which is a set of the gene combinations for which the scale has been calculated; selecting the combination of genes with the highest measure from the population with the specific combination excluded; 3. The information processing device according to claim 1 or 2.

5. An information processing method using a computer, comprising: calculating, for each combination of genes in a gene dataset, a measure of dependency on causative genes of a multifactorial disease or a sporadic disease and on associated genes of the multifactorial disease or the sporadic disease; selecting a combination of a predetermined number of genes from the dataset based on the measure; distributing data on a first gene, which is a gene derived from a healthy subject, on a certain feature space based on an expression amount of the first gene; distributing data of a second gene, the second gene being a gene derived from a patient with the multifactorial disease or the sporadic disease, on the feature space based on an expression level of the second gene; clustering data of the first gene on the feature space based on the expression level of the first gene; The method further comprises: clustering data of the second gene on the feature space based on the expression level of the second gene. Information processing methods.

6. Distributing data on a third gene, which is a gene derived from a subject to be diagnosed with the multifactorial disease or the sporadic disease, on the feature space in which the clusters of the healthy subjects and the patients are formed; calculating a distance between the data of the third gene and the cluster in the feature space; determining whether the subject will develop the multifactorial disease or the sporadic disease based on the distance, or determining whether the subject has developed the multifactorial disease or the sporadic disease; The information processing method according to claim 5 , further comprising:

7. The multifactorial disease or the sporadic disease comprises amyotrophic lateral sclerosis; the predetermined number is three, The combination of the predetermined number of genes includes at least PRKAR1A, QPCT, and TMEM71.

7. The information processing method according to claim 5 or 6.

8. performing a linear regression analysis to exclude a specific combination including genes whose occurrence frequency is equal to or higher than a threshold from a population that is a set of the gene combinations for which the scale has been calculated; selecting the combination of genes for which the measure is highest from the population with the particular combination excluded; The information processing method according to claim 5 or 6, further comprising:

9. A program for causing a computer to execute the program, calculating, for each combination of genes in a gene dataset, a measure of dependency on causative genes of a multifactorial disease or a sporadic disease and on associated genes of the multifactorial disease or the sporadic disease; selecting a combination of a predetermined number of genes from the dataset based on the measure; distributing data on a first gene, which is a gene derived from a healthy subject, on a certain feature space based on an expression amount of the first gene; distributing data of a second gene, the second gene being a gene derived from a patient with the multifactorial disease or the sporadic disease, on the feature space based on an expression level of the second gene; clustering data of the first gene on the feature space based on the expression level of the first gene; The method further comprises: clustering data of the second gene on the feature space based on the expression level of the second gene. program.

10. Distributing data on a third gene, which is a gene derived from a subject to be diagnosed with the multifactorial disease or the sporadic disease, on the feature space in which the clusters of the healthy subjects and the patients are formed; calculating a distance between the data of the third gene and the cluster in the feature space; determining whether the subject will develop the multifactorial disease or the sporadic disease based on the distance, or determining whether the subject has developed the multifactorial disease or the sporadic disease; The program of claim 9 , further comprising:

11. The multifactorial disease or the sporadic disease comprises amyotrophic lateral sclerosis; the predetermined number is three, The combination of the predetermined number of genes includes at least PRKAR1A, QPCT, and TMEM71. The program according to claim 9 or 10.

12. performing a linear regression analysis to exclude a specific combination including genes whose occurrence frequency is equal to or higher than a threshold from a population that is a set of the gene combinations for which the scale has been calculated; selecting the combination of genes for which the measure is highest from the population with the particular combination excluded; The program according to claim 9 or 10, further comprising: