Methods for classifying cancer
Patent Information
- Application Number
- JP2024508310
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-08-12
- Filing Date
- 2022-08-12
- Publication Date
- 2025-08-21
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to a computer-implemented method for the classification of cancer, in particular a diagnostic classification of cancer based on the biological state of specific genomic DNA sites or transcripts and / or an in vitro method for the classification of cancer. The present disclosure provides a method that allows the classification of cancer samples, in particular tumor samples obtained from patients, by analyzing a large number of genetic sites, preferably genome-wide, and combining the biological states of the analyzed genetic sites into a biological state pattern, which is directly and / or indirectly compared with pre-determined biological state patterns belonging to different cancer types or tumor types. The present disclosure is particularly useful for classifying cancers of the central nervous system, i.e. brain tumor samples and / or spinal tumors, because these cancers need to be correctly identified from a wide variety of separate tumor types that have different prognostic values and require treatment regimes developed for each tumor type in the clinical context. However, the present disclosure may be beneficial for other cancers as well, such as sarcomas. [Background technology]
[0002] Considering brain tumors alone, the World Health Organization (WHO) classification describes over 100 different tumors. Many of these show complex patterns of potentially overlapping histological features. Moreover, histologically identical tumors may belong to different molecular groups with very different treatment requirements and prognosis. The same is true for tumors of the spinal cord and tumors originating from tissues outside the central nervous system. Therefore, more advanced diagnostic tools are needed.
[0003] Epigenetic patterns, for example the epigenetic state of different gene sites, play a key role in the development, differentiation and pathogenesis of diseases such as multiple sclerosis, diabetes, schizophrenia, aging and various forms of cancer, including tumors of the central nervous system. Tumors originate from different progenitor cell populations shaped by genetic and epigenetic alterations. It is now recognized that many tumors, including tumors of the central nervous system, belonging to different biological groups are not necessarily histologically distinct. Most tumors show a diverse histological spectrum without clear boundaries. Epigenetic modifications, such as methylation, preserve the information of the cell of origin, i.e., the original identity. Therefore, methylation data, such as DNA methylation patterns, have great potential in identifying molecular subgroups of tumors, such as tumors of the central nervous system. Similar results can be obtained by analyzing the transcripts of the respective genes of interest.
[0004] Nevertheless, treatment regimens, and in particular successful treatment, for many cancers, particularly those of the central nervous system, are highly dependent on early and accurate diagnosis and classification of tumors. In view of the above, new methods that overcome at least some of the problems in the art would be beneficial. Summary of the Invention [Problem to be solved by the invention]
[0005] (overview) The present disclosure aims to provide strategies and methods for diagnostic classification of cancer samples with greater efficiency, specificity and sensitivity. [Means for solving the problem]
[0006] This object of the invention is achieved by the features of the independent claims. Preferred embodiments are defined in the dependent claims. "Aspects", "examples" and "embodiments" herein which are not within the scope of the claims do not form part of the invention and are provided for illustrative purposes only.
[0007] According to an independent aspect of the present disclosure, a computer-implemented method for diagnostically classifying cancer is provided, the method comprising classifying the cancer using a classification algorithm based on a biological state or biological state pattern of a set of genetic sites of a cancer sample.
[0008] The classification algorithm is trained using biological data derived from classified cancer types, such as previously classified cancer types. In particular, the cancer types may be previously classified and / or new cancer types identified using the classification algorithm. For example, the classification algorithm may classify a cancer sample as unknown, and such unknown cancer sample may be further analyzed to determine its cancer type. Further analysis may be performed by various means, such as software and / or medical personnel.
[0009] The classification algorithm is trained with data relating to the biological state of at least the gene sites in Table 1 (SEQ ID NO: 1 to SEQ ID NO: 688). Training the classification algorithm with data for all gene sites in Table 1 can provide an efficient and flexible classification tool.
[0010] In particular, the cancer samples may include: (i) cancer sample data for all 688 gene sites in Table 1; or (ii) cancer sample data for a subset of the 688 gene sites in Table 1 (e.g., at least three gene sites in the cancer sample genome); can be classified using:
[0011] In other words, the classification algorithm is trained with biological data for all 688 gene sites in Table 1, but it may not be necessary to provide cancer sample data for all 688 gene sites to classify a cancer sample. The number of gene sites used to classify a cancer sample may be selected depending on circumstances such as the data available from the cancer sample (e.g., only data for a subset of the 688 gene sites may be available for analysis), time constraints (fewer gene sites are faster to analyze), and sensitivity requirements (more gene sites are more accurate to analyze).
[0012] In view of the above, the computer-implemented method for cancer diagnostic classification can reduce the processing resources used by the GPU and / or reduce the power consumed by the GPU. Furthermore, by using cancer sample data of a subset of the 688 gene sites in Table 1 (e.g., at least three gene sites in the cancer sample genome), the performance, power consumption, and / or programming flexibility of the GPU performing the cancer diagnostic classification method can be improved.
[0013] Preferably, said set of gene sites comprises at least three gene sites of the cancer sample genome selected from the list consisting of the gene sites in Table 1 herein.
[0014] Preferably, the biological state of said gene site includes the biological state of the gene site listed in Table 1 herein, and preferably includes the biological state of up to 20 kb (or 15 kb or 12 kb) upstream and / or downstream of each said gene site. For example, the biological state of said gene site includes the biological state of the gene site listed in Table 1 and the biological state of up to 10 kb, preferably up to 8 kb, or up to 6 kb, or up to 4 kb, or up to 2 kb upstream and / or downstream of said gene site.
[0015] According to some embodiments, which may be combined with other embodiments described herein, the method further comprises determining a biological state for at least three genetic sites in the genome of the cancer sample.
[0016] Additionally or alternatively, the method further comprises determining a biological state pattern of said set of gene sites based on the determined biological state of each of said at least three gene sites.
[0017] According to another independent aspect of the present disclosure, a method for diagnostic classification of cancer is provided.
[0018] According to some embodiments, which may be combined with other embodiments described herein, the method for diagnostic classification of cancer is an in-vitro method.
[0019] In a preferred embodiment, the method comprises: - providing a cancer sample; - determining a biological status for at least three gene sites of the genome of the cancer sample, the gene sites being selected from the list consisting of the gene sites in Table 1, - determining a biological state pattern based on the determined biological state of each of the at least three genetic sites; and - classifying a cancer type based on the determined biological state pattern and on predefined biological state patterns for different cancer types; Includes.
[0020] Preferably, the biological state of said gene sites includes the biological state of the gene sites listed in Table 1 herein, and preferably the biological state of up to 20 kb (or 15 kb or 12 kb) upstream and / or downstream of each said gene site. For example, the biological state of said gene sites includes the biological state of the gene sites listed in Table 1, and preferably the biological state of up to 10 kb, preferably up to 8 kb, or up to 6 kb, or up to 4 kb, or up to 2 kb upstream and / or downstream of said gene site.
[0021] Preferably, the step of determining said biological state pattern comprises combining the biological states of said gene sites into said biological state pattern.
[0022] Preferably, the step of classifying the cancer type comprises comparing the biological state pattern of the set of gene sites to predetermined biological state patterns derived from the biological state data for different cancer types.
[0023] Preferably, a cancer is classified as a particular cancer type if the biological state pattern of the set of gene sites differs from biological state data derived from a previously classified cancer type by at most 5%, preferably at most 4%, or at most 3%, or at most 2%, or at most 1%.
[0024] Preferably, said biological state is selected from the group comprising or consisting of epigenetic state, mutational state, copy number state and RNA expression.
[0025] Preferably, said epigenetic state is a methylation state.
[0026] Preferably, the set of gene sites includes at least 10, preferably at least 20, or at least 30, or at least 40, or at least 50, or at least 60, or at least 70, or at least 80, or at least 90, or at least 100, or all of the gene sites of the cancer sample genome in Table 1.
[0027] Preferably, at least one of said at least three gene sites is the one having the highest value of variable importance (imp_sum) in Tables 3 to 172 herein. Most preferably, at least one of the at least three gene sites is selected from the group comprising (or consisting of): PTPRN2 (SEQ ID NO: 491), PRDM16 (SEQ ID NO: 477), HDAC4 (SEQ ID NO: 249), PAX6 (SEQ ID NO: 431), and MAD1L1 (SEQ ID NO: 349).
[0028] Preferably, the biological states of said gene site include only the biological states of the gene sites listed in Table 1, not including any bases upstream and / or downstream of said gene site.
[0029] Preferably, said biological state is a methylation status and / or said biological state pattern is a methylation status pattern.
[0030] Preferably, the cancer is a cancer of the central nervous system or a sarcoma, however, the disclosure is not limited thereto and other cancer types, such as, for example, carcinoma, sarcoma, myeloma, neural crest tumor (e.g., melanoma), leukemia, lymphoma, and mixed types, may also be classified using the method according to the present invention.
[0031] Preferably, the cancer is a cancer listed in Table 2.
[0032] Preferably, the method further comprises determining a further (second) biological state which is different from the (first) biological state and which relates to at least one of said genetic sites in said cancer sample genome.
[0033] Preferably, the method further comprises the step of correlating said predetermined further (second) biological state of said at least one genetic site related to said cancer sample genome with said classified cancer type.
[0034] Preferably, the method further comprises the step of determining at least one genetic site having said further (second) biological condition determined as an alternative or additional biomarker in the diagnosis of said classified cancer type.
[0035] Preferably, said further (second) biological state is selected from the group comprising or consisting of epigenetic state, mutational state, RNA expression and copy number state.
[0036] According to another independent aspect of the present disclosure, there is provided a computer-readable storage medium having stored thereon computer-executable instructions that, when executed, cause a computer to perform the methods described herein.
[0037] The term "computer-readable storage medium" may refer to any storage device used to store data accessible by a computer, and any other means for providing access to data by a computer. Examples of storage-type computer-readable media include magnetic hard disks, floppy disks, optical disks such as CD-ROMs and DVDs, magnetic tapes, memory chips, and the like.
[0038] According to another independent aspect of the present disclosure, a system for diagnosing cancer is provided, the system including one or more processors and a memory coupled to the one or more processors and including instructions executable by the one or more processors to perform a method described herein.
[0039] The system may be a computer system. The term "computer system" may refer to a system having a computer, the computer including a computer-readable storage medium having software embedded therein for operating the computer.
[0040] As used herein, the term "software" is used interchangeably with "program" and refers to predetermined rules for operating a computer. Examples of software include software, code segments, instructions, computer programs, and programmed logic.
[0041] An embodiment of the present disclosure provides for classification of cancer samples in cancer diagnosis using a classification algorithm that is a machine learning (ML) algorithm.
[0042] The term "classification" refers to a procedure and / or algorithm that classifies individual items into groups or classes based on quantitative information about one or more item-specific characteristics (called traits, variables, features, etc.), based on statistical models and / or training sets of previously labeled items. Specifically, in the context of the present disclosure, classification preferably means determining which particular cancer type a cancer sample belongs to, e.g., as determined by epigenetic characteristics.
[0043] The term "machine learning algorithm" as used throughout this application refers to an algorithm that builds a model based on training data to make predictions or decisions without being explicitly programmed. In particular, the term "classification" refers to a machine learning algorithm that classifies individual items into groups or classes based on quantitative information about one or more features (called traits, variables, characteristics, features, etc.) specific to the items, based on a statistical model and / or a training dataset of previously labeled items. Specifically, in the context of the present invention, classification preferably means determining which specific cancer type a cancer sample belongs to (e.g., by its epigenetic state pattern).
[0044] The term "training dataset" in the context of the present invention refers to a set of biological state data, such as genomic methylation data, of a large number of tumors that have been classified by prior art methods and are therefore of known tumor type.
[0045] The classification algorithm may be any suitable algorithm for establishing correlation between the dataset, i.e., the biological state or biological state pattern of the cancer sample and the biological state data derived from the pre-classified cancer type (which may be a predetermined biological state or biological state pattern). Methods for establishing correlation between datasets include, but are not limited to, discriminant analysis (DA) (e.g., linear-, quadratic-, regularized-DA), discriminant function analysis (DFA), kernel methods (e.g., SVM), multidimensional scaling (MDS), non-parametric methods (e.g., k-nearest neighbor classifiers), PLS (partial least squares), tree-based methods (e.g., logistic regression, CART, random forests, boosting / bagging), generalized linear models (e.g., logistic regression), principal component-based methods (e.g., SIMCA), generalized additive models, fuzzy logic-based methods, neural networks, and genetic algorithm-based methods.
[0046] Those skilled in the art will have no problem in selecting a suitable method / algorithm for establishing the correlation between the biological state or biological state pattern of cancer sample and the biological state data derived from the pre-classified cancer type of the present invention.In one embodiment, the method / algorithm used for correlating the biological state or biological state pattern of cancer sample and the biological state data derived from the pre-classified cancer type of the present invention is selected (or composed) from the group including DA (e.g., linear, quadratic, regularized discriminant analysis), DFA, kernel method (e.g., SVM), MDS, non-parametric method (e.g., k-nearest neighbor classifier), PLS (partial least squares), tree-based method (e.g., logistic regression, CART, random forest method, boosting method), generalized linear model (e.g., logistic regression), and principal component analysis.
[0047] In one exemplary embodiment, the classification algorithm uses random forest analysis. As used herein, the term "random forest analysis" refers to a computational method based on the idea of using multiple different decision trees to calculate the overall most predicted class (mode). In a specific application, the mode is either a tumor type or a class, based on how many decision trees predict the sample to match a particular class. The class predicted by the majority is selected as the predicted class for the sample. The different decision trees used in this algorithm are trained on a randomly generated subset of the training dataset and a randomly selected set of the variables. This is why the algorithm depends on two hyperparameters: the number of random trees to use and the number of random variables to use to train the different trees.
[0048] The term "biological state" may refer to the epigenetic state, mutational state, RNA expression or copy number state of a gene or gene site.
[0049] The term "epigenetic state" refers to a measure of epigenetic changes (or functionally related changes of up-regulation and / or down-regulation) of gene activity of a particular gene site and / or gene in the genome of a biological sample. The epigenetic state includes epigenetic down-regulation and / or up-regulation of the activity of a gene site in the biological sample compared to the activity of the same gene site in a physiological tissue. Such down-regulation and / or up-regulation can be due to, for example, DNA methylation, histone modification, or other epigenetic effects.
[0050] The term "epigenetic state pattern" refers to a combination of epigenetic states of multiple gene sites and / or genes. It includes an overview of the epigenetic states of gene sites and / or genes. Thus, in its simplest form, an epigenetic state pattern can include information about which of said multiple gene sites and / or genes have epigenetically modified activity compared to physiological state and which do not. For example, if DNA methylation is used as a measure of epigenetic influence on the activity of genes or gene sites, said epigenetic state pattern can also include information about which of gene sites and / or genes are epigenetically upregulated and / or downregulated in terms of hypermethylation (leading to downregulation) or hypomethylation (leading to upregulation).
[0051] In some embodiments, the classification algorithm of the present disclosure may be trained with epigenetic data derived from a classified cancer type, such as a pre-classified cancer type. The epigenetic data may be provided in the form of a pre-defined pattern or pre-defined epigenetic state pattern. The term "pre-defined pattern" or "pre-defined epigenetic state pattern" refers to an epigenetic state pattern that is predetermined and typical of a particular cancer type, such as one of the cancer types listed in Table 2 (and e.g., Tables 3-172). An initial iteration of the pre-defined pattern is determined by the inventors and is used to train the classification algorithm.
[0052] In a preferred embodiment of the present disclosure, the predetermined epigenetic state pattern is associated with a cancer type listed in Table 2 (and, for example, Tables 3 to 172, respectively). Moreover, the predetermined epigenetic state pattern includes essentially the same gene sites as the set of gene sites of the cancer sample being analyzed. If, as will be explained in detail below, determining the epigenetic state of the set of gene sites results in an epigenetic state pattern corresponding to one of the predetermined patterns, the cancer associated with the sample can be classified as belonging to this cancer type. The predetermined epigenetic state pattern is preferably determined by a classification algorithm. That is, the predetermined epigenetic state pattern is not accessible to the user per se, but is included in the results of a classification algorithm that uses, for example, machine learning and continuously updates its own reference material. Thus, the predetermined epigenetic state pattern determined by the classification algorithm changes over time in order to further increase sensitivity and specificity. Thus, the predetermined patterns used in the present invention are continuously changing, so that it is not feasible or useful to give an example thereof. However, those skilled in the art are familiar with these aspects of machine learning and can readily derive classification algorithms for establishing the unique pre-defined epigenetic state patterns used herein.
[0053] Biological changes, such as epigenetic changes, in cancer tissues are known to be specific to a particular cancer type or subtype. The biological state of a gene site can be determined using various methods known to those skilled in the art. For example, the biological state, such as the epigenetic state of a gene site, can be determined by assaying histone modifications, proteomics, or transcriptomics. One approach may be based on, for example, assaying transposase-accessible chromatin using sequencing (ATAC-seq). Another approach is to assay DNA methylation. Determining the epigenetic state of a gene site through the determination of methylation is the preferred approach used in the present disclosure, since robust and reliable DNA methylation assays have been established and are readily available. However, a single data point measured by the aforementioned assay does not determine the type of cancer. The type of cancer is determined by the epigenetic downregulation or upregulation of its gene site, which in turn determines the metabolism and phenotype of the cancer. However, the activation of gene sites can be determined by many different epigenetic approaches, as outlined above. Therefore, to classify cancer types, it is more prudent to determine the effect of epigenetic changes on the activity of gene sites, rather than relying on specific epigenetic changes measured by a specific type of assay. In theory, all epigenetic approaches should ultimately show that the same gene sites have pathological activity, provided that all gene sites and their activities are equally accessible by the various assays. It is this pathological gene site activity that derives and defines the cancer type.
[0054] Taken together, the disclosed methods classify cancers based on the biological state of specific genomic DNA sites or transcripts.
[0055] In one embodiment of the present disclosure, the inventors use DNA methylation to find gene sites with pathological activity in cancer genome.Then, the epigenetic status of these gene sites is used to find typical patterns for different cancer types.Thus, the inventors find the set of gene sites that have the greatest impact on the differentiation of different cancer types.
[0056] To this end, we tested our approach using Illumina's methylation bead chip testing a large number of classically typed tumor specimens. Using Illumina's Human-manMethylation450 (450k) BeadChip, we are able to assay DNA methylation at 482,421 CpG dinucleotides. The platform measures DNA methylation by genotyping sodium bisulfite-treated DNA. Only small amounts of DNA are needed to run the assay, and both frozen and paraffin (FFPE) material can be used. To date, approximately 90,000 tumor samples have been profiled by us, allowing validation of the surprisingly successful approach disclosed herein.
[0057] As is readily apparent to those skilled in the art, classification according to the present disclosure also means that stratification and / or diagnosis of cancer is achieved. The term "stratification" refers to classifying or grouping patients according to one or more predetermined criteria. In certain embodiments, stratification is performed in a diagnostic setting to group patients according to the prognosis of disease progression with or without treatment. In certain embodiments, stratification is used to distribute patients enrolled for clinical studies according to their individual characteristics. In certain embodiments, stratification is used to identify the optimal treatment option for patients.
[0058] The term "diagnosis" or "diagnostic" as used herein refers to the identification or classification of a molecular or pathological state, disease or condition. For example, "diagnosis" may refer to the identification of a particular type of cancer, such as lung cancer. "Diagnosis" may also refer to the classification of a particular type of cancer, for example, by histological characteristics (e.g., non-small cell lung cancer), molecular characteristics (e.g., lung cancer characterized by nucleotide and / or amino acid mutations in a particular gene or protein), or both. However, it is important to note that the present disclosure is directed to strictly in vitro methods in all its embodiments. No method steps of any embodiment are performed in a human or animal body.
[0059] The terms "cancer type", "tumor type" or "tumor class" refer to a particular kind or subcategory of tumor that can be classified based on tissue origin, genetic makeup, histology, etc. In particular in the field of brain tumors, there are various different tumor types or classes of tumors of the central nervous system that can be differentiated, for example, by histopathology (1. Acta Neuropathol. 2007 Aug; 114(2): 97-109. Epub 2007 Jul6. "The 2007 WHO classification of tumours of the central nervous system.". Louis DN(1), Ohgaki H, Wiestler OD, Cavenee WK, Burger PC, Jouvet A, Scheithauer BW, Kleihues P.). In particular, the present disclosure relates to the cancer types listed in Table 2.
[0060] The term "cancer sample" or "tumor sample" as used herein refers to a sample obtained from a patient. The tumor sample can be obtained from a patient by routine means known to those skilled in the art, i.e., by biopsy (obtained by aspiration or puncture, excision, or any other surgical method that results in biopsy or excised cellular material). For sites that are difficult to reach with an open biopsy, a "closed" biopsy can be performed using a stereotaxic device through a small hole drilled by the surgeon in the skull. The stereotaxic device allows the surgeon to precisely position the biopsy probe in three-dimensional space, allowing access to almost anywhere in the brain. Thus, tissue can be obtained for the diagnostic method of the present disclosure. However, the actual removal of the sample from the patient is not part of the method of the present invention. Thus, "providing a cancer sample" simply relates to making the sample available for use in the laboratory without the step of first obtaining the sample from the patient.
[0061] The term "cancer" or "tumor" is not limited to the stage, grade, histomorphological characteristics, invasiveness, aggressiveness or malignancy of the affected tissue or cell mass, and specifically includes stage 0 cancer, stage I cancer, stage II cancer, stage III cancer, stage IV cancer, grade I cancer, grade II cancer, grade III cancer, malignant cancer, primary cancer, and all other types of cancer, malignant tumors, and the like.
[0062] As used herein, the term "gene site" refers to a region of DNA that includes or consists of a gene, particularly a gene or gene site listed in Table 1. In particular, the term "gene site" refers to a DNA sequence having a locus as defined in Table 1. A gene site may include, for example, up to 12 kb, preferably up to 10 kb, up to 8 kb, or up to 6 kb, or up to 4 kb, or up to 2 kb of additional base pairs upstream and / or downstream of the gene. Thus, the biological state, such as the epigenetic state, of a gene site may refer to the biological state of the gene itself, and also the biological state of the additional string of base pairs upstream and / or downstream of the gene. In a preferred embodiment of the present disclosure, the biological state of the gene sites in the set includes the biological state of the gene sites listed in Table 1, and the biological state of up to 10 kb, preferably up to 8 kb, or up to 6 kb, or up to 4 kb, or up to 2 kb upstream and / or downstream of the gene. In a further preferred embodiment, the biological state of the gene site includes only the biological state of the gene site listed in Table 1 without any bases upstream and / or downstream of the gene site. In this case, only the biological state of the gene site itself is used, and the gene site does not include any bases other than the gene site listed in Table 1.
[0063] The term "set of gene sites" refers to a number of gene sites grouped together, for example, it is the epigenetic state of this set of gene sites that is evaluated in the present disclosure and then combined into a pattern and analyzed by a classification algorithm.
[0064] As used herein, the term "CpG site" or "CpG position" refers to a region of DNA where a cytosine nucleotide is adjacent to a guanine nucleotide in the linear sequence of bases along its length, with the cytosine (C) being one phosphate (p) away from the guanine (G). Approximately 70% of human gene promoters are CpG-rich. Regions of the genome rich in CpG sites are known as "CpG islands." Cytosines in CpG dinucleotides can be methylated to form 5-methylcytosines. Methylation (i.e., introduction of a methyl group) at cytosines at CpG sites in the promoter of a gene can cause gene silencing, a feature found in many human cancers. In contrast, hypomethylation of CpG sites is generally associated with overexpression of cancer genes in cancer cells. In the context of this disclosure, the term "independent genomic CpG positions" means that each CpG position of a group of genomic CpG positions can be individually probed for its methylation status.
[0065] The term "methylation status" as used herein refers to the methylation status of a CpG site, and refers to the presence or absence of 5-methylcytosine at a CpG site in genomic DNA. If an individual's DNA is not methylated at a CpG site, the position is 0% methylated. If all of an individual's DNA is methylated at the CpG site, the position is 100% methylated. If only a portion of an individual's DNA, for example 50%, 75%, or 80%, is methylated at the CpG site, the CpG position is said to be 50%, 75%, or 80% methylated, respectively. The term "methylation status" reflects the relative or absolute amount of methylation at a CpG position. Methylation at a CpG position can be evaluated by any method used in the art. The terms "methylation" and "hypermethylation" are used interchangeably herein. When used in reference to a CpG position, the terms refer to a methylation state that corresponds to an increased abundance of 5-methylcytosine at a CpG site in the DNA of a biological sample obtained from a patient compared to the amount of 5-methylcytosine found at the CpG site within the same genomic location in a biological sample obtained from a healthy individual or from an individual suffering from a different class or type of tumor.
[0066] The term "biological sample" is used in the broadest sense. In the practice of the present disclosure, a biological sample is generally obtained from a subject. The sample may be any biological tissue or bodily fluid that can assay for the biological state of the present disclosure. In many cases, the sample is a "clinical sample" (i.e., a sample obtained from or derived from the patient being tested). The sample may also be an archived sample with a known history of diagnosis, treatment, and / or outcome. Examples of biological samples suitable for use in the practice of the present disclosure include, but are not limited to, bodily fluids, such as blood samples (e.g., blood smears), and cerebrospinal fluid, brain tissue samples, spinal cord tissue samples, or bone marrow tissue samples (such as tissue or fine needle biopsy samples). "Biological sample" may also include tissue sections, such as frozen sections taken for histological purposes. The term "biological sample" also encompasses any material derived from processing a biological sample. Derived material includes, but is not limited to, cells (or their progeny) isolated from the sample, and nucleic acid molecules (DNA and / or RNA) extracted from the sample. Processing of biological samples may include one or more of filtration, distillation, extraction, concentration, inactivation of interfering components, addition of reagents, and the like.
[0067] The method according to some embodiments of the present disclosure includes a step of "determining the epigenetic status" of a set of gene sites. This step can be accomplished by any means suitable for assaying the epigenetically altered activity of gene sites. In a preferred embodiment of the present disclosure, the epigenetic status of a set of gene sites is determined in a biological sample obtained from a patient by evaluating the DNA methylation status of a large number of independent genomic CpG positions, in particular the CpG positions within the gene sites as described above, preferably the CpG positions within the gene sites listed in Table 1. The determination of the methylation status can be performed using any method known in the art to be suitable for evaluating the methylation of cytosine residues in DNA. Such methods are known and described in the art. The skilled person will know how to select the most suitable method depending on the number of samples to be tested, the amount of sample available, etc.
[0068] Thus, the methylation status of a genomic CpG position or combination of genomic CpG positions according to the present disclosure can be determined using any of a wide variety of methods, generally classified into methylation-specific PCR (MSP)-based approaches and approaches employing PCR performed under methylation-independent conditions (MIP). Methylation-independent PCR (MIP) primers are used in most of the available PCR-based methods. They are designed to amplify methylated and unmethylated DNA in proportion. In contrast, methylation-specific PCR (MSP) primers are designed to amplify only methylated templates.
[0069] Examples of methylation-independent PCR-based techniques include, but are not limited to, direct bisulfite sequencing (Frommer et al., PNAS USA, 1992, 89:1827-1831), pyrosequencing (Collela et al., Biotechniques, 2003, 35:146-150; Uhlmann et al., J. Electrophoresis, 2002, 23:4072-4079; Tost et al., Biotechniques, 2003, 35:152-156), Combined Bisulfite Restriction Analysis or "COBRA" (Xiong et al., Nucleic Acids Res., 1997, 25:2532-2534), Methylation-Sensitive Single-Nucleotide Primer Extension (MSP), and ELISA (Methylation-Sensitive Single-Nucleotide Primer Extension). Extension or "MS-SNUPE" (Gonzalgo et al., Nucleic Acids Res, Nucleic Acids Res., 1997, 25:2529-2531), Methylation-Sensitive Melting Curve Analysis or "MS-MSA" (Worm et al., Clin. Chem., 2001, 47:1183-1189), Methylation-Sensitive High-Resolution Melting or "MS-HRM" (Wojdacz et al., Nucleic Acids Res., 2007, 35:e41), MALDI-TOF mass spectrometry with base-specific cleavage and primer extension (Ehrich et al., PNAS USA, 2005, 102:15785-15790), and HeavyMethyl (Cottrell et al., Nucleic Acids Res., 2005, 103:15785-15790). Res.,2004,32:e10).
[0070] Examples of methylation-specific PCR-based techniques include, for example, methylation-specific PCR or "MSP" (Herman et al., PNAS USA, 1996, 93:9821-9826; Mackay et al., Hum. Genet., 2006, 120:262-269; Mackay et al., Hum. Genet., 2005, 116:255-261; Palmisano et al., Cancer Res., 2000, 60:5954-5958; Voso et al., Blood, 2004, 103:698-700), MethylLight (Eads et al., Nucleic Acids Res., 2000, 28:e32; Eads et al., Cancer Res., 1999, 59:2302-2306; Lo et al., Cancer Res., 1999, 59:3899-3903), Melting curve Methylation Specific PCR or "McMSP" (Akey et al., Genomics, 2002, 80:376-384), Sensitive Melting Analysis after Real-Time MSP or "SMART-MSP" (Kristensen et al., Nucleic Acids Res., 2008, 36:e42), and Methylation-Specific Fluorescent Amplicon Generation or "MS-FLAG" (Bonanno et al., Clin. Chem., 2007, 53:2119-2127).
[0071] Many of these methods rely on prior treatment of DNA with sodium bisulfite. This treatment converts unmethylated cytosines to uracils, while methylated cytosines remain unchanged (Clark et al., Nucleic Acids Res., 1994, 22:2990-2997). This change in DNA sequence after bisulfite conversion can be detected using a variety of methods, including PCR amplification followed by DNA sequencing. The use of bisulfite-converted DNA for DNA methylation analysis has surpassed almost all other methodologies for DNA methylation analysis and has arguably become the gold standard for detecting changes in DNA methylation. The protocol described by Frommer et al. (PNAS USA, 1992, 89:1827-1831) is widely used for sodium bisulfite treatment of DNA, and a variety of commercial kits are now available for this purpose.
[0072] Therefore, in the method according to the present disclosure, the step of determining the epigenetic state can be achieved by determining the methylation state of a gene promoter or a combination of gene promoters of the present disclosure. It can be performed using any of the above techniques, or any combination of these techniques. Those skilled in the art will recognize that if the methylation state of a combination of gene promoters needs to be determined, the determination can be performed using the same DNA methylation analysis technique or a different DNA methylation analysis technique. Other methods include oligonucleotide methylation tiling arrays, BeadChip assays, HPLC / MS methods, methylation-specific multiplex ligation-dependent probe amplification (MS-MPLA), bisulfite sequencing, and assays using antibodies against DNA methylation (i.e., ELISA assays).
[0073] By using the statistical model described herein, the inventors found that gene sites including those listed in Table 1 are sufficient to classify cancer samples into a number of different cancer types. It may be possible to classify even more cancer types by analyzing the named gene sites, but this has been verified for the cancer types listed in Table 2. Thus, according to the present disclosure, to classify a cancer type, it is sufficient to determine the epigenetic status of these selected gene sites, in particular at least three gene sites. Thus, a complete analysis of the entire genome of a cancer type can be avoided. For a sufficiently specific classification, only the gene sites listed in Table 1 need to be analyzed, which results in a faster and less laborious diagnosis.
[0074] The inventors further found that a set of gene sites containing at least three of the gene sites listed in Table 1 is sufficient for classifying cancer samples. However, the larger the set, the higher the accuracy. Thus, in a preferred embodiment of the present disclosure, the set of gene sites contains at least 10, preferably at least 20, or at least 30, or at least 40, or at least 50, or at least 60, or at least 70, or at least 80, or at least 90, or at least 100 genes of the sample genome of the cancer to be classified. Preferably, a set of gene sites containing 100 or less, 90 or less, 80 or less, 70 or less, 60 or less, 50 or less, 40 or less, 30 or less, 20 or less, or 10 or less gene sites provides a good balance between accuracy and required work. The embodiments of the present disclosure are not limited thereto, and the set may contain more than 100 gene sites or may contain all gene sites listed in Table 1.
[0075] Although all of the gene sites or genes listed in Table 1 may be used to classify cancer types as described in Table 2 herein, the inventors have identified a subset of genes that are meant to be of higher importance, i.e., higher accuracy, when used to classify a particular cancer type. Thus, it is preferred that a pre-defined pattern for a cancer type as described in Table 2 includes at least three gene sites for that particular cancer type. It is even more preferred that a pre-defined pattern for a cancer type includes at least three gene sites for that particular cancer type selected from the gene sites listed in Tables 3 to 172, respectively. In a preferred embodiment, the set of gene sites of the cancer sample genome to be analyzed includes exactly the same gene sites or genes as the pre-defined pattern.
[0076] In one preferred embodiment, the statistical model we employ provides a measure of variable importance of gene sites for each cancer.
[0077] As can be seen from Tables 3-172, different gene sites have different importance to the classification. Therefore, to improve the accuracy of the classification, the epigenetic data for a cancer type preferably includes the gene sites listed in Tables 3-172 for that cancer type that have the highest variable importance for that cancer type.
[0078] As mentioned above, the set of gene sites of the analyzed cancer sample genome preferably includes the same genes or gene sites as the epigenetic data derived from the pre-classified cancer type (predetermined pattern).Therefore, the set of gene sites of the analyzed cancer sample genome and the epigenetic data derived from the pre-classified cancer type preferably also include the same number of genes or gene sites.
[0079] Analyzing gene sites of a set of genes containing three genes is beneficial in that it is less laborious, but the accuracy of classification increases with the number of genes analyzed per cancer type. Thus, it is preferred that the predetermined pattern for a cancer type listed in Table 2 includes at least 10, preferably at least 20, or at least 30, or at least 40, or at least 50, or at least 60, or at least 70, or at least 80, or at least 90, or at least 100 gene sites or genes listed in Table 1. The preferred gene sites or genes "for a cancer type" are those listed in Tables 3 to 172 for each cancer type, respectively. As explained above for the set, 80 to 100 genes provide a good balance between accuracy and workload.
[0080] This classification may include directly or indirectly comparing the epigenetic state pattern of the set of gene sites with the predetermined epigenetic state pattern, for example by determining the overlap of the two patterns, i.e., how similar or different the two patterns are from each other. This may be determined statistically, for example, and may be expressed as a numerical value. In particular, the difference between the patterns may be expressed as a percentage. The accuracy of the classification may be influenced by allowing patterns that are more or less different from the predetermined pattern to be classified as the cancer type to which the predetermined pattern belongs. For a reasonable accuracy, it is preferred that a cancer is classified as being related to the predetermined pattern if the epigenetic state pattern of the set of gene sites differs from the predetermined pattern by at most 5%, preferably at most 4%, or at most 3%, or at most 2%, or at most 1%. These values are practically useful and can be achieved by the method of the present invention.
[0081] As explained above, the predetermined epigenetic state patterns used for comparison are determined by the present inventors by analyzing more than 90,000 cancer samples obtained from various different sources. This process is also part of this disclosure, and will be described in detail below.
[0082] In all embodiments, the methods of the present disclosure are performed as ex-vivo or in-vitro methods.
[0083] In another preferred embodiment of the invention, then, the invention comprises carrying out a method according to the invention and providing said patient with an appropriate therapy, said therapy being based, at least in part, on the results of the method according to the invention.
[0084] In another preferred embodiment of the invention, the invention relates to a method for developing a treatment plan for a cancer (e.g. a tumor type) classified using the method according to the invention. Preferably, the method further comprises providing an appropriate treatment to a patient based on the developed treatment plan.
[0085] By "treatment" is meant the alleviation and / or amelioration of symptoms of a disease. An effective treatment, for example, reduces the mass of a tumor or the number of cancer cells. The treatment can also avoid (prevent) and reduce the spread of cancer, for example by influencing metastasis and / or its formation. The treatment can be a naive treatment (before other treatments of the disease are started) or a treatment after an initial treatment (for example after surgery or recurrence). The treatment can also be a combined treatment, including, for example, chemotherapy, surgery, and / or radiation therapy. The treatment can also modulate autoimmune responses, infections, and inflammation.
[0086] Most preferably, the method according to the present disclosure is used for classification of tumors of the central nervous system, therefore, said tumor is preferably a brain tumor or a spinal tumor, and said tumor type is a brain tumor type or a spinal tumor type. As already mentioned herein above, these tumors are characterized by a huge epigenetic diversity that has a significant impact on the development of therapeutic regimens to allow the best treatment of patients. When the tumor disease is a tumor of the central nervous system (CNS), it is preferred that said tumor type comprises at least 184 different classes of CNS tumors. Furthermore, the present disclosure is also applicable to sarcomas. In a preferred embodiment, said CNS tumor is selected from the list of cancer types or tumor types in Table 2.
[0087] The determination of DNA methylation level of the present disclosure is preferably carried out using a genome array or chip that contains probes specific for the methylation of at least 1000 CpG positions. It is preferable to test as many positions as possible to enable the generation of highly specific classifications. Therefore, genome-wide DNA methylation assays such as HumanMethylation450k-chip (Illumina®) are preferred.
[0088] The classification algorithm may be based on Random Forest (RF). Training of an RF-based classification algorithm according to some embodiments of the present disclosure may include the preceding steps of selecting, among all CpG positions used, the CpG positions that provide the purest splitting rule, and using said selected CpG positions as a training data optimization set for training the classification rule.
[0089] In another embodiment of the present disclosure, training the RF-based classification algorithm may include downsampling the number of bootstrap samples for each tumor type to a minority class, the minority class being the smallest sample size of the tumor type in the training dataset.
[0090] Another embodiment of the present disclosure provides the above method, further comprising: including the methylation data of the classified tumor samples in a training dataset to obtain an enriched training dataset; and calculating an enriched classification rule by random forest analysis based on the enriched training dataset. Optionally, the classification of the tumor samples can be repeated using the enriched classification rule. This embodiment is useful for the continuous development and refinement of the original training dataset. Each further classified tumor type has a genomic DNA methylation profile or epigenetic status pattern that can further enrich the classification for that tumor type and be used as a predetermined epigenetic status pattern in the present disclosure. Thus, the disclosure in a preferred embodiment provides a classification system that features a self-learning classification rule.
[0091] In order to provide classification rules with good specificity and sensitivity, the predetermined methylation data / epigenetic status pattern used in the context of the present disclosure includes the methylation status / level at said CpG positions of at least one, two, three, four, five, six or more independent samples for each pre-classified cancer type.
[0092] Another aspect of the present disclosure then relates to a method for stratifying the treatment of a tumor patient, comprising classifying the tumor type / cancer type of the patient's tumor according to the classification method of the present disclosure and stratifying the patient's treatment according to the diagnosed tumor type.
[0093] A further aspect of the present disclosure relates to a computer-implemented method for generating classification rules to assist in the classification of tumor samples in cancer diagnosis, comprising: providing DNA methylation data of CpG positions of multiple independent genomes of genomes of multiple diverse pre-classified tumor types of the same tumor type (e.g., brain tumor, lung cancer, leukemia, etc.); and calculating a random forest of binary decision trees from the DNA methylation data, where in each binary decision tree of the random forest, each node is a CpG position, each end leaves a specific tumor type, and each binary split rule is a methylation state at the CpG position. This method can be used to create a predetermined epigenetic state pattern as described above.
[0094] To learn classification rules that can predict the class assignment of future diagnosed cases, we applied the machine learning algorithm Random Forest (RF; Breiman, 2001). The RF algorithm is a so-called ensemble method that combines the predictions of several "weak" classifiers to improve the prediction accuracy. RF uses binary classification trees (Classification and Regression Trees (CART); Breiman et al., 1983) as "weak" classifiers. Each of these trees is a sequence of binary partitioning rules that are learned by recursive binary partitioning. The CART algorithm starts with all samples assigned to a "root" node and tries to find the variables, e.g., the measured CpG probes, and the corresponding cutoffs that result in the purest partition into different classes. To measure this gain in class "purity", the Gini index, a classical statistical measure for inequality, can be used. To fit the tree, the CART algorithm performs these steps iteratively until no further improvement is possible, i.e., only samples of the same class are assigned to the final "leaf" nodes or a prespecified node size is reached. To predict the class of a new diagnostic case, the binary splitting rules are compared to the new data, starting from the root node to one of the leaf nodes, and the tree then predicts, or votes, for the class that dominates that leaf node.
[0095] Decision trees have the advantage that they are non-parametric and do not rely on distributional assumptions. Moreover, the trees can learn complex non-linear relationships and interactions, are easy to interpret, and can be efficiently fitted to large datasets. The main drawbacks of decision trees are that they often tend to overfit the data and have weak predictive performance.
[0096] However, to improve the predictive accuracy of a single tree, the RF algorithm combines thousands of trees by bootstrap aggregation (bagging). Briefly, each tree is fitted with a training data set that is generated by drawing a bootstrap sample, i.e., by randomly selecting two-thirds of the data with replacement. Furthermore, at each node, only a random subset of the available variables is used to find the optimal splitting rule. This additional source of randomization allows the selection of variables with low predictive value that would otherwise be filtered out by the most salient variables. This feature ensures that the resulting trees are uncorrelated, i.e., different variables are used to find the optimal prediction rule. Taking a majority vote over thousands of bootstrap-aggregated and uncorrelated trees significantly improves the predictive accuracy of RF. The majority vote, i.e., the proportion of trees voting for a class, can be used as an empirical class probability or score and has proven to be a very useful tool for diagnostics.
[0097] To validate the resulting RF classifier, we repeat 5-fold cross-validation. In each cross-validation, the reference set is randomly split into 5 folds. Then, 4 / 5 of the data is used to train the RF classifier and 1 / 5 is used for prediction. Now, the estimated test error of this classifier is about 3.1%.
[0098] Alternatively, the resulting RF classifier is validated by repeated 3-fold cross-validation. In each cross-validation, the reference set is randomly split into 3 folds. Two-thirds of the data is then used to train the RF classifier and one-third is used for prediction. Currently, the estimated test error of the classifier is about 4.9%.
[0099] The classification scores produced by RF, i.e., the proportion of trees voting for a class, are usually unevenly distributed across classes. Moreover, when interpreted as class probabilities, the scores often fail to estimate the actual class probabilities and are therefore said to be not properly calibrated. However, to determine the classification of a single case in terms of clinical diagnosis, the uncertainty associated with the individual predictions, in terms of confidence scores, or estimated class probabilities, is necessary. To obtain recalibrated scores that are comparable across classes and have improved estimates of the certainty of the individual predictions, we fit a calibration model to the raw RF scores. This calibration model is a multinomial logistic regression model, with the tumor subclass as the response variable and the “raw” RF scores as the explanatory variables. Furthermore, the model is fitted by incorporating a small ridge penalty on the likelihood to stabilize the estimates in situations where the classes are completely separable, as well as to prevent overfitting of the model. The amount of this regularization, i.e., the penalty parameter, is determined by performing a 10-fold cross-validation and choosing the value that minimizes the misclassification error. To fit this model independently, the “raw” RF scores are required; that is, scores must be generated by an RF classifier that has not been trained with the same samples. To generate such independent "raw" scores, we apply 3-fold cross-validation.
[0100] To validate the class predictions generated using the recalibrated scores of the calibration model, 3-fold nested cross-validation is applied. In each cross-validation, the reference set is randomly split into 3, and 2 / 3 of the data is used to train the RF classifier and 1 / 3 is used for prediction. In each of these 3 cross-validations, 3-fold nested cross-validation is applied to generate independent RF scores, which are used to train the calibration model. The predicted RF scores obtained from the outer cross-validation loop are recalibrated using a suitable calibration model (i.e., a model fitted using the RF scores generated by using the remaining 2 / 3 of the data in the inner loop). Currently, the estimated test error of the classifier when using the recalibrated scores for prediction is about 3.2%.
[0101] Some embodiments of the present disclosure relate to methods, wherein the diverse tumor types are selected from metastatic tumors, tumors originating from a particular tissue, tumors at a particular stage, recurrent tumors, tumors with particular genetic mutations, tumors of patients with different genders, ages or genetic backgrounds.
[0102] The present disclosure will now be further described in the following non-limiting examples with reference to the accompanying figures and sequences. For purposes of this disclosure, all references cited herein are incorporated by reference in their entirety. [Brief description of the drawings]
[0103] [Figure 1] Heatmap representation of the reference set. Color codes indicate different tumor classes, FFPE samples, frozen samples, and samples misclassified by cross-validation. The heatmap shows the methylation profile of the 10,000 CpG probes most important for classification (highest average gain in Gini purity). [Diagram 2]Example of a binary decision tree. At each node, the CpG probes and corresponding cutoffs are used for the binary decision. The last leaf node displays the abbreviation of the tumor subclass, i.e., EPN_PFA means posterior fossa adenoma subtype A. [Diagram 3] Median test error estimated from three 5-fold cross-validation runs. [Figure 4] The left panel symbolically shows the histology of WNT and group 3 medulloblastomas, which are indistinguishable. The right panel is a multidimensional scaling (MDS) analysis of 107 medulloblastoma samples of all molecular subtypes using the 21,092 most variable CpG probes. WNT medulloblastomas are colored blue, SHH medulloblastomas are colored red, group 3 medulloblastomas are colored yellow, and group 4 medulloblastomas are colored green. [Figure 5A] A shows the histological results of the patients, B shows the classifier scores, with the highest scoring items highlighted. [Figure 5B] A shows the histological results of the patients, B shows the classifier scores, with the highest scoring items highlighted. [Figure 6A] Panels A and B show the histological examination results of the patients, and panel C shows the classifier scores, with the highest scoring item highlighted. [Figure 6B] Panels A and B show the histological examination results of the patients, and panel C shows the classifier scores, with the highest scoring item highlighted. [Figure 6C] Panels A and B show the histological examination results of the patients, and panel C shows the classifier scores, with the highest scoring item highlighted. [Figure 7]1 shows a schematic of how a classifier is trained and validated by three-fold nested cross-validation. In each outer cross-validation run, the training data is used for an inner three-fold nested cross-validation that generates independent RF scores. These scores can be used to fit a calibration model and then applied to recalibrate the RF scores generated by predicting the test data in the outer loop. The RF scores generated in the outer loop can be used to fit a calibration model using all data in the reference set that will later be used for new diagnostic cases. [Figure 8] Genomic plot showing the importance measures of the PTPRN2 gene, CpG sites, and RF variables. [Figure 9] Heatmap showing methylation values of 100 CpGs located in PAX6, PTPRN2, and OSTM1 with the highest standard deviation across 75 ATRT samples. Hierarchical clustering was applied with Euclidean distance as distance metric and complete linkage as linkage method to permute rows and columns. Color codes for class annotation indicate known molecular subtypes, and gene annotation indicates the gene in which the CpG is located. [Figure 10] A shows the projection of 75 ATRT tumor samples onto the first two PCs resulting from PCA, B shows the projection of 75 ATRT tumor samples onto the coordinates calculated by tSNE analysis, C shows the CART tree with two consecutive splitting rules, D shows the scatter plot of 75 ATRT tumor samples, where the x-axis and y-axis are the methylation values of the two CpG sites selected by the CART tree. The corresponding splitting rule cutoffs are marked with dashed lines. [Figure 11] tSNE of 1167 samples for which DNA methylation and gene expression data were available. tSNE coordinates were calculated based on gene expression data for the 688 most significant genes or gene sites. Class labels and colors correspond to the classes predicted by the methylation classifier. [Figure 12A]Confusion matrix showing the results of three-fold cross-validation for validating the RF and multinomial logistic regression models. Similar to the classifier trained on methylation data, most errors occur between closely related entities such as MB group 3 and 4 subtypes. [Figure 12B] Confusion matrix showing the results of three-fold cross-validation for validating the RF and multinomial logistic regression models. Similar to the classifier trained on methylation data, most errors occur between closely related entities such as MB group 3 and 4 subtypes. [Figure 13] A simulation study examining the brain tumor classifier performance of a classifier trained with CpG probes located on a random subset of signature genes and random hg19 genes. [Figure 14] tSNE dimensionality reduction of DNA methylation profiles of 9084 TCGA cases from 33 different projects, each focused on a specific tumor entity. [Figure 15] The left figure shows the confusion matrix showing the results of 3-fold cross-validation, and the right figure shows the tSNE dimensionality reduction highlighting the incorrectly predicted samples in cross-validation. [Figure 16] Confusion matrices for four different statistical or machine learning models trained on the TCGA cohort shown in Figure 14 . DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0104] Infinium Methylation Assay Genome-wide screening of DNA methylation patterns was performed using Infinium HumanMethylation450 BeadChips (Illumina, San Diego, USA). Combining Infinium I and Infinium II assay chemistry technologies, the BeadChips cover 99% of RefSeq genes and 96% of CpG islands.
[0105] DNA concentration was measured using PicoGreen (Life Technologies, Darmstadt, Germany). The quality of genomic DNA samples was checked by agarose gel analysis, and samples with an average fragment size >3 kb were selected for methylation analysis. Formalin-fixed, paraffin-embedded (FFPE) DNA samples were assessed for quality by real-time PCR analysis using the Infinium HD FFPE QC Kit (Illumina) on a Light Cycler 480 Real-Time PCR System (Roche, Mannheim, Germany). Laboratory work was performed at the Genomics and Proteomics Core Facility of the German Cancer Research Center (DKFZ), Heidelberg, Germany.
[0106] DNA from each sample (500 ng genomic DNA and 250 ng FFPE DNA, respectively) was subjected to bisulfite conversion using the EZ-96 DNA Methylation Kit (Zymo Research Corporation, Orange, USA) according to the manufacturer's recommendations. Bisulfite treatment deaminates unmethylated cytosines to uracil, whereas methylated cytosines are not susceptible to bisulfite and remain cytosines. After bisulfite conversion, FFPE samples were treated with the Infinium HD DNA Restoration Kit (Illumina) according to the manufacturer's recommendations. Enzymatic reactions were used to repair degraded FFPE DNA in preparation for whole genome amplification.
[0107] Each sample underwent whole genome amplification and enzymatic fragmentation according to the procedures in the Illumina Infinium HD Assay Methylation Protocol Guide (genomic DNA) or Infinium HD FFPE Methylation Guide (FFPE DNA), respectively. The DNA was applied to an Infinium HumanMethylation450 BeadChip and hybridized at 48°C for 16-24 hours. During hybridization, DNA molecules anneal to locus-specific DNA oligomers bound to individual bead types. Depending on the probe design for a particular CpG site, one or two probes are used to interrogate the CpG locus.
[0108] Allele-specific primer annealing is followed by single-base extension using DNP and biotin-labeled ddNTPs. Infinium I's assay design ensures that both bead types (one each for the methylated and unmethylated state) for the same CpG locus are detected in the same color channel because they incorporate the same type of labeled nucleotide determined by the base preceding the "C" being tested in the CpG locus. Infinium II uses only one type of bead with a unique type of probe, and therefore is able to detect both alleles. Methylated and unmethylated signals are generated in the green and red channels, respectively.
[0109] After extension, the arrays were fluorescently stained and scanned to measure the intensity of each CpG. Microarray scanning was performed using an iScan array scanner (Illumina). DNA methylation values are described as beta values and recorded for each locus in each sample. DNA methylation beta values are continuous variables between 0 and 1 and represent the percentage of methylation of a particular cytosine, corresponding to the ratio of methylated signals to the sum of methylated and unmethylated signals.
[0110] Data Preprocessing All data analyses were performed using the open source statistical programming language R (R Core Team, 2014). Raw data files generated by the iScan array scanner were read and preprocessed using functions in the Bioconductor package minfi (Aryee et al., 2014). The minfi package performed the same preprocessing steps as recommended by Illumina's BeadStudio software.
[0111] In addition, the following filtering criteria were applied: removal of probes targeting chromosomes X and Y (n=11,551), removal of probes containing single nucleotide polymorphisms (dbSNP132 Common) within 5 base pairs of the targeted CpG site (n=24,536), and probes that do not uniquely map to the human reference genome (hg19) but allow one mismatch (n=9,993). In total, 438,370 probes were kept for analysis.
[0112] Training the classifier The Random Forest (RF) algorithm implemented in RandomForest (Liaw and Wiener, 2002) was used to learn the classification of 1899 samples assigned to 72 different brain tumor subtypes. The RF algorithm is a so-called ensemble method that combines the predictions of several "weak" classifiers to achieve an improved prediction accuracy. RF uses binary classification trees (Classification and Regression Trees (CART); Breiman et al., 1983) as "weak" classifiers. Each of these trees represents a sequence of binary partitioning rules that are learned by recursive binary partitioning. The CART algorithm starts with all samples assigned to a "root" node and tries to find the variables, e.g. the measured CpG probes, and the corresponding cutoffs that result in the purest partition into different classes. To measure this gain in class "purity", the Gini index, a classical statistical measure of inequality, is used. To fit the tree, the CART algorithm performs these steps iteratively until no further improvement is possible, i.e. only samples of the same class are assigned to the final "leaf" nodes or a prespecified node size is reached. To predict the class of a new diagnostic case, the binary splitting rules are compared with the new data, from the root node to one of the leaf nodes. The tree then predicts or votes for the class that dominates its leaf node. However, to improve the predictive accuracy of a single tree, the RF algorithm combines thousands of trees by bootstrap aggregation (bagging). Briefly, each tree is fitted with a training dataset that is generated by drawing a bootstrap sample (i.e., randomly selecting two-thirds of the data with replacement). Furthermore, at each node, only a random subset of the available variables is used to find the optimal splitting rule. To predict the class of a diagnostic sample, RF takes a majority vote of all trees in the forest.
[0113] To learn the classification, the default parameter settings of the randomForest function were used and 10,000 decision trees were fitted. Furthermore, considering the highly unbalanced class sizes, a downsampling strategy was adopted, i.e., the number of bootstrap samples drawn from each class to fit the decision tree is equal to the number of samples from the minority class. To further improve the predictive performance of the classifier, variable selection was performed: in a first step, the algorithm is used to calculate the importance of variables (e.g., the average improvement in the Gini purity of CpG probes when used in the splitting rule). The final classifier was trained using only the 30,000 CpG probes with the highest variable importance.
[0114] An overview of the classifier training is shown in Figure 7.
[0115] Internal Validation To validate the resulting classifier and estimate its performance in predicting future diagnostic cases, we repeatedly applied 5-fold cross-validation. For example, in each cross-validation run, the reference set is randomly split into 5. Then, 4 / 5 of the data is used to train the RF classifier described above, and 1 / 5 is used for prediction. Now, the estimated test error of the classifier is about 3.1%.
[0116] Example 1: Differentiation of WNT and Group 3 medulloblastomas Medulloblastoma is the most common pediatric malignant brain tumor and contains four distinct molecular variants. These variants are known as WNT, SHH, group 3, and group 4. These variants are histologically indistinguishable but can be clearly distinguished by their DNA methylation patterns (see Figure 4). WNT tumors exhibit an activated Wnt signaling pathway and are associated with a good prognosis. SHH medulloblastomas exhibit activation of the Hedgehog signaling pathway and are associated with a moderate to good prognosis. Although both WNT and SSH variants have already been well characterized molecularly, the genetic programs that lead to the pathogenesis of group 3 and group 4 medulloblastomas remain largely unknown.
[0117] Example 2: Changing the diagnosis of anaplastic astrocytoma (WHO III) A female brain tumor patient born in 1944 was diagnosed with anaplastic astrocytoma (WHO III) based on histological findings (see FIG. 5A). Using the classification procedure of the present invention, the diagnosis was changed to glioblastoma (WHO IV) with a classification score of 0.335 (see FIG. 5B).
[0118] Example 3: Changing the diagnosis of Schwannoma A male patient born in 1969 was diagnosed with schwannoma based on histological examination (Figures 6A and 6B). However, the classification procedure of the present disclosure allowed the patient to be diagnosed with meningioma (WHO I) (see Figure 6C).
[0119] Example 4: DNA methylation-based classification of tumors using three gene sites Atypical teratoid rhabdoid tumors (ATRT) are rare pediatric brain tumors that can be classified into three molecular subgroups: ATRT-TYR, ATRT-SHH, and ATRT-MYC (Ho et al., 2020, PMID:31889194).
[0120] We identified genes containing CpG sites that are most important for the classification of brain tumors and molecular subtypes. The importance of these CpGs in classification was measured by applying a permutation-based variable importance measure of the Random Forest (RF) algorithm (Strobl et al. 2007, PMID: 17254353). Among them, three genes, PAX6, PTPRN2, and OSTM1, contain many CpGs that are important for classification. Figure 8 shows the PTPRN2 gene and the CpG sites on it. Most CpGs have a positive variable importance measure, indicating that these CpGs can predict the classification of brain tumors.
[0121] In the following, we show how CpGs located on three genes PAX6, PTPRN2 and OSTM1 can be used to classify ATRT into three molecular subtypes by applying different unsupervised and supervised statistical methods. After preprocessing, we identified 1022 CpGs located on three genes. We found that applying unsupervised hierarchical clustering to the methylation values of the 100 CpGs with the highest standard deviation across 75 ATRT samples resulted in a near-perfect separation into the three molecular subtypes of ATRT (Figure 9).
[0122] Next, we apply principal component analysis (PCA) to the methylation values of all 1022 CpGs as an example of a linear dimensionality reduction method. Projecting the samples onto the first two principal components (PCs), which explain most of the variation in the data, reveals a grouping into three molecular subtypes (Figure 10A). Furthermore, as an example of a nonlinear dimensionality reduction method, t-distributed stochastic neighbor embedding (t-SNE) is applied to the methylation data, and the resulting projection also shows a clustering into three ATRT subtypes (Figure 10B). Other linear and nonlinear dimensionality reduction methods that can be applied to achieve similar results include, for example, multidimensional scaling (MDS), factor analysis (FA), nonnegative matrix factorization (NMF), truncated singular value decomposition (SVD), stochastic neighbor embedding (SNE), uniform manifold approximation and projection (UMAP) for dimensionality reduction, and linear discriminant analysis (LDA).
[0123] To show how supervised statistical methods can be applied to fit models predicting ATRT subtypes, we applied classification and regression trees (CART) to the methylation data (Figure 10C). At each node, the CART algorithm automatically tries all 1022 available CpG probes and possible cutoffs and selects the CpG probe and the corresponding cutoff that leads to the purest partition into ATRT subtypes. The algorithm stops as soon as the purity of the classes, measured by the Gini coefficient, does not improve any further. Here, the CART algorithm found two consecutive partitioning rules (Figure 10D) involving only two CpG probes that lead to a near-perfect separation of the ATRT subclasses. Random forests usually combine hundreds or thousands of CART trees by bootstrap aggregation (bagging) to achieve improved prediction accuracy. Other supervised methods that can be applied to fit models with comparable predictive performance are, for example, gradient boosting machines (GBM), support vector machines (SVM), multinomial logistic regression models, and (deep) neural nets.
[0124] Example 5: Gene Expression Data Used to Classify Tumors Identified from DNA Methylation Data By analyzing DNA methylation data and training a machine learning model for brain tumor classification, we identified 688 genes containing CpG sites that are likely to be most important for molecular brain tumor type classification. To show that these brain tumor entities can also be recognized in gene expression data and that this data can be used to train machine learning models with similar performance, we analyzed 1167 brain tumor samples for which both DNA methylation and gene expression data were available. This paired gene expression and methylation dataset contains samples from 79 out of a total of 184 classes defined by DNA methylation levels.
[0125] Figure 11 shows the 1167 samples projected onto t-distributed stochastic neighbor embedding (tSNE) applied to gene expression data of the 688 most significant genes identified in the methylation data. The groups are colored and labeled according to the classes, and the samples are classified by a DNA methylation classifier. The general clustering of the classes is very similar to the tSNE performed on DNA methylation, and even new subentities such as medulloblastoma (MB) groups 3 and 4 subtypes I-VIII can be identified. This proves that the gene expression data of the 688 identified genes is highly predictive for the 184 classes.
[0126] To demonstrate that gene expression data can also be used to train supervised machine learning models, the gene expression dataset was reduced to 1057 samples belonging to 50 classes, with a minimum class sample size of 7 samples. We then trained a basic random forest (RF) model and a lasso-penalized multinomial logistic regression model on this dataset and validated the performance of both models by 3-fold cross-validation (CV). The CV results showed that the accuracy of RF was 0.788 (Figure 12B) and the accuracy of the logistic regression model was 0.766 (Figure 12A), proving that gene expression can be used to train similar classification models.
[0127] Thus, the inventors have shown that the biological state used to train the classification algorithm is not limited to methylation, but can also be another biological state, such as gene expression.
[0128] Example 6: Simulation studies to investigate the brain tumor classifier performance of classifiers trained with CpG probes located on a random subset of signature genes and random hg19 genes To show that a subset of the 688 signature genes already predicts the defined brain tumor methylation classes, we performed a simulation study. In this study, a random forest classifier was trained with CpG probes located on different random subsets of the 688 signature genes. The number of genes was varied from 3 to 688 in steps equal to 20, and training was repeated at least three times for each number of genes. In addition, we also trained classifiers with CpG probes located on genes randomly sampled from all known genes available in the hg19 genome. The performance of each trained classifier was measured by the overall accuracy and the number of classes with per-class accuracy greater than 0.8.
[0129] Figure 13 shows the results of this simulation test. For the three gene subset, the difference between the genes selected from the signature gene list in Table 1 and the randomly selected genes is most clear, i.e., the overall accuracy of the signature genes is about 0.8, while for the random gene classifier it is always below 0.5. Increasing the number of genes improves the overall accuracy of both the signature gene classifier and the random gene classifier to a level of about 0.90 accuracy or higher. The signature gene classifier always performs better than the classifier trained with random genes. Considering the number of classes for which a class accuracy above 0.8 is achieved, the simulation shows that the genes in Table 1 are important for reliably predicting more specific classes.
[0130] Example 7: Classifiers for other pan-cancer tumors To demonstrate that the signature gene list can also be used to train high-performance classification models to predict other cancer types, we trained the RF classifier on a large cohort of publicly available DNA methylation array samples from the Cancer Genome Atlas Project (TCGA).
[0131] Figure 14 shows tSNE of 9084 samples from 31 different TCGA projects investigating different cancer types (e.g., LUAD is an abbreviation for lung adenocarcinoma, BRCA for breast cancer, etc.). A complete list of TCGA projects and their abbreviations can be found at the following link: https: / / portal.gdc.cancer.gov / projects. For each project, we defined tumor and control tissue classes as much as possible, resulting in a total of 53 classes. For this dataset, we trained an RF classifier with all CpGs found in the genes listed in the signature list in Table 1, and the resulting classifier achieved an overall accuracy of 0.9226, as measured by three-fold statistical cross-validation (Figure 15: the confusion matrix on the left shows the results of the three-fold cross-validation; the plot on the right shows the tSNE dimensionality reduction highlighting the samples incorrectly predicted in the cross-validation). Errors usually occur between related entities such as lung squamous cell carcinoma (LUSC) and lung adenocarcinoma (LUAD).
[0132] Applying other statistical or machine learning algorithms suitable for multi-class classification tasks, predictive models with comparable accuracy can be fitted, as shown in Figure 16. Figure 16 shows the confusion matrices of four different statistical or machine learning models trained on the TCGA cohort shown in Figure 14. The regularized logistic regression model showed the highest overall accuracy of 0.9343, followed by the linear kernel support vector machine (SVM) with an accuracy of 0.9299, the extreme gradient boosted trees (XGBoost) classifier with an accuracy of 0.9239, and the radial basis function kernel SVM with an accuracy of 0.9101. More careful tuning of the hyperparameters could improve the performance of all presented predictive models.
[0133] (References) R Core Team (2014). R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. URL http: / / www.R-project.org / . MJ Aryee, AE Jaffe, H Corrada-Bravo, C Ladd-Acosta, AP Feinberg, KD Hansen, RA Irizarry. Minfi: A flexible and comprehensive Bioconductor package for the analysis of Infinium DNA Methylation microarrays. Bioinformatics 2014, in press. doi: 10.1093 / bioinformatics / btu049. A. Liaw and M. Wiener (2002). Classification and Regression by randomForest. R News 2(3), 18--22. Bioconductor: Open software development for computational biology and bioinformatics R. Gentleman, V. J. Carey, D. M. Bates, B. Bolstad, M. Dettling, S. Dudoit, B. Ellis, L. Gautier, Y. Ge, et al. 2004, Genome Biology, 5, R80.
[0134] [Table 1] TIFF2024537596000003.tif248168TIFF2024537596000004.tif248168TIFF2024537596000005.tif248168TIFF20245375960 00006.tif248168TIFF2024537596000007.tif248168TIFF2024537596000008.tif248168TIFF2024537596000009.tif248168 TIFF2024537596000010.tif248168TIFF2024537596000011.tif248168TIFF2024537596000012.tif248168TIFF20245375960 00013.tif248168TIFF2024537596000014.tif248168TIFF2024537596000015.tif248168TIFF2024537596000016.tif203168
[0135] [Table 2] TIFF2024537596000018.tif248168TIFF2024537596000019.tif248168TIFF2024537596 000020.tif248168TIFF2024537596000021.tif248168TIFF2024537596000022.tif24816 8TIFF2024537596000023.tif248168TIFF2024537596000024.tif248168TIFF2024537596 000025.tif248168TIFF2024537596000026.tif248168TIFF2024537596000027.tif74168
[0136] Tables 3 to 172: Classification of the cancer types listed in Table 2 according to the present disclosure. The classification data for each cancer type listed in Table 2 are presented in individual tables. Each table consists of the following columns: The first column shows the gene sites selected for cancer type classification. The second column shows the overall statistical importance (imp_sum) of a particular gene site in the classification of cancer type. The overall importance (imp_sum) of a particular gene site is calculated by multiplying the number of single measurement points (n_probes) in the fourth column by the mean variable importance (imp_mean) in the third column. Higher values represent more important gene sites. The third column shows the average variable importance (imp_mean) of all single measurement points (n_probes) for a given gene site according to the statistical model used (e.g., based on random forest). The fourth column indicates the number of single measurement points (n_probes; CpG site methylation probes located within the gene site).
[0137] [Table 3] TIFF2024537596000029.tif248168TIFF2024537596000030.tif248168TIFF2024537596000031.tif248168TIFF2024537596000032.tif248168TIFF2024537596000033.tif248168TIFF2024537596000034.tif248168TIFF2024537596000035.tif248168TIFF2024537596000036.tif248168TIFF2024537596000037.tif248168TIFF2024537596000038.tif248168TIFF2024537596000039.tif248168TIFF2024537596000040.tif248168TIFF2024537596000041.tif248168TIFF2024537596000042.tif248168TIFF2024537596000043.tif248168TIFF2024537596000044.tif248168TIFF2024537596000045.tif248168TIFF2024537596000046.tif248168TIFF2024537596000047.tif248168TIFF2024537596000048.tif248168TIFF2024537596000049.tif248168TIFF2024537596000050.tif248168TIFF2024537596000051.tif248168TIFF2024537596000052.tif248168TIFF2024537596000053.tif248168TIFF2024537596000054.tif248168TIFF2024537596000055.tif248168TIFF2024537596000056.tif248168TIFF2024537596000057.tif248168TIFF2024537596000058.tif248168TIFF2024537596000059.tif248168TIFF2024537596000060.tif248168TIFF2024537596000061.tif248168TIFF2024537596000062.tif248168TIFF2024537596000063.tif248168TIFF2024537596000064.tif248168TIFF2024537596000065.tif248168TIFF2024537596000066.tif248168TIFF2024537596000067.tif248168TIFF2024537596000068.tif248168TIFF2024537596000069.tif248168TIFF2024537596000070.tif248168TIFF2024537596000071.tif248168TIFF2024537596000072.tif248168TIFF2024537596000073.tif248168TIFF2024537596000074.tif248168TIFF2024537596000075.tif248168TIFF2024537596000076.tif248168TIFF2024537596000077.tif248168TIFF2024537596000078.tif248168TIFF2024537596000079.tif248168TIFF2024537596000080.tif248168TIFF2024537596000081.tif248168TIFF2024537596000082.tif248168TIFF2024537596000083.tif248168TIFF2024537596000084.tif248168TIFF2024537596000085.tif248168TIFF2024537596000086.tif248168TIFF2024537596000087.tif248168TIFF2024537596000088.tif248168TIFF2024537596000089.tif248168TIFF2024537596000090.tif248168TIFF2024537596000091.tif170168TIFF2024537596000092.tif248168TIFF2024537596000093.tif248168TIFF2024537596000094.tif248168TIFF2024537596000095.tif248168TIFF2024537596000096.tif248168TIFF2024537596000097.tif248168TIFF2024537596000098.tif248168TIFF2024537596000099.tif248168TIFF2024537596000100.tif248168TIFF2024537596000101.tif248168TIFF2024537596000102.tif248168TIFF2024537596000103.tif248168TIFF2024537596000104.tif248168TIFF2024537596000105.tif248168TIFF2024537596000106.tif248168TIFF2024537596000107.tif248168TIFF2024537596000108.tif248168TIFF2024537596000109.tif248168TIFF2024537596000110.tif248168TIFF2024537596000111.tif248168TIFF2024537596000112.tif248168TIFF2024537596000113.tif248168TIFF2024537596000114.tif248168TIFF2024537596000115.tif248168TIFF2024537596000116.tif248168TIFF2024537596000117.tif248168TIFF2024537596000118.tif248168TIFF2024537596000119.tif248168TIFF2024537596000120.tif248168TIFF2024537596000121.tif248168TIFF2024537596000122.tif248168TIFF2024537596000123.tif248168TIFF2024537596000124.tif248168TIFF2024537596000125.tif248168TIFF2024537596000126.tif248168TIFF2024537596000127.tif248168TIFF2024537596000128.tif248168TIFF2024537596000129.tif248168TIFF2024537596000130.tif248168TIFF2024537596000131.tif248168TIFF2024537596000132.tif248168TIFF2024537596000133.tif248168TIFF2024537596000134.tif248168TIFF2024537596000135.tif248168TIFF2024537596000136.tif248168TIFF2024537596000137.tif248168TIFF2024537596000138.tif248168TIFF2024537596000139.tif248168TIFF2024537596000140.tif248168TIFF2024537596000141.tif248168TIFF2024537596000142.tif248168TIFF2024537596000143.tif248168TIFF2024537596000144.tif248168TIFF2024537596000145.tif248168TIFF2024537596000146.tif248168TIFF2024537596000147.tif248168TIFF2024537596000148.tif248168TIFF2024537596000149.tif248168TIFF2024537596000150.tif248168TIFF2024537596000151.tif248168TIFF2024537596000152.tif248168TIFF2024537596000153.tif248168TIFF2024537596000154.tif248168TIFF2024537596000155.tif248168TIFF2024537596000156.tif248168TIFF2024537596000157.tif248168TIFF2024537596000158.tif248168TIFF2024537596000159.tif248168TIFF2024537596000160.tif248168TIFF2024537596000161.tif248168TIFF2024537596000162.tif248168TIFF2024537596000163.tif248168TIFF2024537596000164.tif248168TIFF2024537596000165.tif248168TIFF2024537596000166.tif248168TIFF2024537596000167.tif248168TIFF2024537596000168.tif248168TIFF2024537596000169.tif248168TIFF2024537596000170.tif248168TIFF2024537596000171.tif248168TIFF2024537596000172.tif248168TIFF2024537596000173.tif248168TIFF2024537596000174.tif248168TIFF2024537596000175.tif248168TIFF2024537596000176.tif248168TIFF2024537596000177.tif248168TIFF2024537596000178.tif248168TIFF2024537596000179.tif248168TIFF2024537596000180.tif248168TIFF2024537596000181.tif248168TIFF2024537596000182.tif248168TIFF2024537596000183.tif248168TIFF2024537596000184.tif248168TIFF2024537596000185.tif248168TIFF2024537596000186.tif248168TIFF2024537596000187.tif248168TIFF2024537596000188.tif248168TIFF2024537596000189.tif248168TIFF2024537596000190.tif248168TIFF2024537596000191.tif248168TIFF2024537596000192.tif248168TIFF2024537596000193.tif248168TIFF2024537596000194.tif248168TIFF2024537596000195.tif248168TIFF2024537596000196.tif248168TIFF2024537596000197.tif248168TIFF2024537596000198.tif248168TIFF2024537596000199.tif248168TIFF2024537596000200.tif248168TIFF2024537596000201.tif248168TIFF2024537596000202.tif248168TIFF2024537596000203.tif248168TIFF2024537596000204.tif248168TIFF2024537596000205.tif248168TIFF2024537596000206.tif248168TIFF2024537596000207.tif248168TIFF2024537596000208.tif248168TIFF2024537596000209.tif248168TIFF2024537596000210.tif248168TIFF2024537596000211.tif248168TIFF2024537596000212.tif248168TIFF2024537596000213.tif248168TIFF2024537596000214.tif248168TIFF2024537596000215.tif248168TIFF2024537596000216.tif54168.
Claims
1. A method for classifying a cancer sample taken from a patient, said method comprising:
1. A method comprising: classifying a cancer sample taken from a patient whose biological state is to be determined using a classification algorithm trained using at least data regarding the biological state of all gene sites in Table 1 (SEQ ID NO: 1 to SEQ ID NO: 688), wherein the biological state is derived from a classified cancer type, and wherein classifying the cancer comprises applying the classification algorithm to data regarding the biological state of a set of gene sites of the cancer sample, wherein the set of gene sites comprises at least three gene sites of the cancer sample genome selected from the gene sites in Table 1 (SEQ ID NO: 1 to SEQ ID NO: 688).
2. The method described in claim 1, wherein the classification algorithm is a machine learning (ML) algorithm.
3. The method described in claim 1, wherein the step of classifying the cancer includes a step of using the machine learning algorithm to build a model based on the trained data in Table 1.
4. The method of claim 1, wherein the machine learning algorithm is based on at least one of discriminant analysis, discriminant function analysis, kernel methods, multidimensional scaling, nonparametric methods, partial least squares, tree-based methods, generalized linear models, principal component-based methods, generalized additive models, fuzzy logic-based methods, neural networks, and genetic algorithm-based methods.
5. determining a biological state for each of the at least three genetic loci in the cancer sample genome; and 10. The method of claim 1, further comprising determining a biological state pattern for the set of gene sites based on the determined biological states of the at least three gene sites.
6. The method, providing data from a cancer sample; determining a biological state for at least three gene sites in the genome of the cancer sample, wherein the gene sites are selected from the list consisting of the gene sites in Table 1; determining a biological state pattern based on the determined biological states of each of the at least three genetic sites; and classifying the cancer type based on the determined biological state pattern and pre-determined biological states associated with different cancer types.
7. The method described in claim 1, wherein the cancer is classified as a particular cancer type if the biological state pattern of the set of gene sites differs from the biological state data derived from a previously classified cancer type by at most 5%, preferably at most 4%, or at most 3%, or at most 2%, or at most 1%.
8. 2. The method of claim 1, wherein the biological state is selected from the group consisting of epigenetic state, mutational state, copy number, and RNA expression.
9. 9. The method of claim 8, wherein the epigenetic state is a methylation state.
10. The method described in claim 9, wherein the biological state is selected from the epigenetic states, the epigenetic state is a methylation state, and the biological state pattern is a methylation state pattern.
11. The method described in claim 6, further comprising defining at least one genetic site having a determined further (second) biological state as an alternative or additional biomarker in determining the classified cancer type.
12. The method described in claim 11, wherein the further (second) biological state is selected from the group including epigenetic state, mutation state, RNA expression, and copy number.
13. The method described in claim 1, wherein the classification means determining whether a cancer sample belongs to a specific cancer type determined by its epigenetic state pattern, and the epigenetic state pattern is a methylation state pattern.
14. 2. The method of claim 1, wherein the set of gene sites comprises at least 10 gene sites, or all of the gene sites, in Table 1 (SEQ ID NO: 1 to SEQ ID NO: 688).
15. 15. The method of claim 14, wherein the set of gene sites comprises at least 20, or at least 30, or at least 40, or at least 50, or at least 60, or at least 70, or at least 80, or at least 90, or at least 100, or all of the gene sites in Table 1 (SEQ ID NO: 1 to SEQ ID NO: 688).
16. 2. The method of claim 1, wherein the gene site is the gene site with the highest variable importance value in Tables 3 to 172, respectively.
17. 2. The method of claim 1, wherein the biological states of the gene sites include the biological states of the gene sites listed in Table 1 (SEQ ID NO: 1 to SEQ ID NO: 688) and the biological states of up to 12 kb upstream and / or downstream of the gene.
18. 18. The method of claim 17, wherein the biological states of the gene site include the biological states of the gene site listed in Table 1 (SEQ ID NO: 1 to SEQ ID NO: 688) and the biological states of up to 10 kb, or up to 8 kb, or up to 6 kb, or up to 4 kb, or up to 2 kb upstream and / or downstream of the gene.
19. 2. The method of claim 1, wherein the biological states of the gene sites include only the biological states of the gene sites listed in Table 1 (SEQ ID NO: 1 to SEQ ID NO: 688) without any bases upstream and / or downstream of the gene site.
20. The method of claim 1 , wherein the biological state is a methylation state and / or the biological state pattern is a methylation state pattern.
21. 2. The method of claim 1, wherein the cancer is selected from the group consisting of carcinoma, sarcoma, myeloma, neural crest tumors including melanoma, leukemia, lymphoma and mixed types.
22. 10. The method of claim 1, wherein the cancer is a cancer listed in Table 2.
23. determining a further biological state for at least one of the genetic loci relative to the cancer sample genome that is different from the biological state, wherein the further biological state is selected from the group consisting of epigenetic state, mutation state, RNA expression, and copy number; 10. The method of claim 1, further comprising: correlating the additional biological state of the at least one genetic site with respect to the cancer sample genome with the classified cancer type.
24. 2. The method of claim 1, wherein the at least three gene sites comprise one or more of the following: PTPRN2 (SEQ ID NO:491), PRDM16 (SEQ ID NO:477), HDAC4 (SEQ ID NO:249), PAX6 (SEQ ID NO:431), and MAD1L1 (SEQ ID NO:349).
25. The method of claim 1, wherein the method is a computer-implemented method.
26. 1. A computer-readable storage medium having stored thereon computer-executable instructions that, when executed, cause a computer to perform a method for classifying a cancer sample obtained from a patient, the method comprising: The method comprises: inputting data from the cancer sample into a classification model, the data relating to the biological state of a set of gene sites in the cancer sample, the set of gene sites comprising at least three gene sites in the genome of the cancer sample selected from the gene sites in Table 1 (SEQ ID NO: 1 to SEQ ID NO: 688); applying the classification model to data relating to the biological state of the set of gene sites in the cancer sample to classify the cancer sample; wherein the classification model is trained based on a machine learning algorithm using at least data related to the biological states of the gene sites, particularly all gene sites, in Table 1 (SEQ ID NO: 1 to SEQ ID NO: 688) derived from the classified cancer types.
27. A trained model constructed from a machine learning algorithm based on the training data in Table 1 for causing a computer to output a classification of a cancer sample taken from a patient, comprising: the model is configured to receive data of the cancer sample, the data relating to a biological state of a set of gene sites of the cancer sample, the set of gene sites comprising at least three gene sites of the cancer sample genome selected from the gene sites in Table 1 (SEQ ID NO: 1 to SEQ ID NO: 688); the model causes the computer function to classify the cancer sample based on data relating to the biological state of a set of gene sites in the cancer sample; The model is trained based on a machine learning algorithm using at least data related to the biological status of the gene sites in Table 1 (SEQ ID NO: 1 to SEQ ID NO: 688), particularly all gene sites, derived from classified cancer types. A trained model.
28. A data structure for classifying cancer samples taken from a patient, comprising: the cancer sample data, the data relating to the biological state of a set of gene sites of the cancer sample, the set of gene sites comprising at least three gene sites of the cancer sample genome selected from the gene sites in Table 1 (SEQ ID NO: 1 to SEQ ID NO: 688); wherein the data structure is a computer-executed process of: inputting the cancer sample data into a classification model; applying the classification model to the data to classify the cancer sample; A data structure, wherein the classification model is trained based on a machine learning algorithm using at least data related to the biological states of the gene sites in Table 1 (SEQ ID NO: 1 to SEQ ID NO: 688), particularly all gene sites, derived from classified cancer types.
29. 1. A system for classifying cancer, comprising: one or more processors; a memory coupled to the one or more processors and including instructions executable by the one or more processors to perform the method of claim 1 for classifying a cancer sample taken from a patient; The method comprises:
1. A system comprising: classifying a cancer sample taken from a patient whose biological state is to be determined using a classification algorithm trained using at least data regarding the biological state of all gene sites in Table 1 (SEQ ID NO: 1 to SEQ ID NO: 688), wherein the biological state is derived from a classified cancer type, and wherein classifying the cancer comprises applying the classification algorithm to data regarding the biological state of a set of gene sites in the cancer sample, wherein the set of gene sites comprises at least three gene sites in the cancer sample genome selected from the gene sites in Table 1 (SEQ ID NO: 1 to SEQ ID NO: 688).
30. The system described in claim 29, wherein the system includes one or more processors and a memory connected to the one or more processors and containing instructions executable by the one or more processors to perform the method described in claim 1.
31. The system described in claim 29, wherein the system includes a computer, the computer including a computer-readable storage medium having embedded software for operating the computer.