Methods for characterizing cancer
Patent Information
- Application Number
- JP2024507116
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-08-07
- Filing Date
- 2022-08-04
- Publication Date
- 2025-08-13
AI Technical Summary
Current molecular profiling methods for cancer diagnosis, such as next-generation sequencing (NGS), are costly, require significant investment, and have long turnaround times, especially in laboratories with low specimen submissions, necessitating sample pooling over multiple weeks, and lack off-the-shelf solutions tailored for neuro-oncology.
A computer-implemented method using targeted nanopore sequencing and classification algorithms that selectively sequence specific gene sites, allowing immediate analysis of single samples, including frozen sections and liquid biopsies, by analyzing initial nucleotide sequences to determine cancer types through flexible target selection and efficient processing.
Enables rapid, cost-effective, and flexible cancer classification with high concordance rates for molecular markers, reducing turnaround time from days to hours, and providing comprehensive profiling of CNS tumors with minimal infrastructure requirements.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to cancer characterization, and more particularly to a computer-implemented method for classification and / or molecular characterization of cancer in a diagnostic setting. The present disclosure provides a method for characterizing tumor samples obtained from patients by analyzing biological states, such as methylation status, mutation or copy number status, or mutational status, of selected gene sites using targeted nanopore sequencing. Embodiments of the present disclosure are particularly useful for characterizing various tumors, such as central nervous system tumors and / or sarcomas, and identifying targets for therapeutic intervention. The term "characterization" as used throughout the present disclosure includes one or more of classification, diagnosis, treatment response prediction, and stratification. [Background technology]
[0002] Molecular markers can be used in cancer diagnosis, such as brain tumor diagnosis. For example, DNA methylation-based cancer classification provides a comprehensive molecular approach to diagnose tumors of the central nervous system (CNS). Indeed, DNA methylation profiling of human brain tumors has already had a major impact in clinical neuro-oncology. This is complemented by the assessment of mutations, gene fusions, and DNA copy number status of key driver genes.
[0003] Currently, multiple gene sequencing approaches are available for cancer diagnosis, but traditional sequencing setups for complete molecular profiling require significant investments, and batch processing of samples for sequencing and methylation profiling can result in long turnaround times.
[0004] Moreover, neuropathology laboratories cannot rely on off-the-shelf products, as they do not cover genes relevant to neuro-oncology. Therefore, custom assays are usually set up if equipment for next-generation sequencing (NGS) is available. Thus, the benefits of custom neuropathology NGS panels can only be efficiently utilized if the number of cases is sufficient. Therefore, in facilities with low sample submission volumes, samples must be pooled over multiple weeks.
[0005] For these reasons, there is a need for more advanced diagnostic tools. Summary of the Invention [Problem to be solved by the invention]
[0006] (overview) It is an object of the present disclosure to provide diagnostic techniques that are highly flexible, can be efficiently performed on a single sample, and / or can begin immediately upon receipt of, for example, a frozen section, a liquid biopsy sample, or cells from a biological sample. It is also an object of the present disclosure to improve the speed of analysis of tumor samples, such as frozen section analysis.
[0007] This object is achieved by the features of the independent claims. Preferred embodiments are defined in the dependent claims. "Aspects", "examples" and "embodiments" herein which are not included in the claims do not form part of the invention and are provided for illustrative purposes only. [Means for solving the problem]
[0008] According to an independent aspect of the present disclosure, there is provided a computer-implemented method for cancer diagnosis, the method comprising the steps of:
[0009] a) selectively sequencing a polymer of a biological sample according to at least one target gene site by translocating said polymer through a nanopore of a nanopore sequencing system, comprising: (i) analyzing an initial nucleotide sequence of a first polymer of the biological sample while the first polymer is moving through a nanopore of the nanopore sequencing system to determine whether the initial nucleotide sequence corresponds to the at least one target gene site; and (ii) continuing said sequencing of said first polymer only if said initial nucleotide sequence of said first polymer corresponds to said at least one target gene site to obtain measured data of said first polymer; and b) determining a biological state of a nucleotide sequence of the first polymer corresponding to the at least one target gene site based on the measurement data; and c) classifying the cancer based on the biological state of the nucleotide sequence of the first polymer using a classification algorithm, wherein the classification algorithm is trained based on the at least one target gene site and biological state data relating to a type of cancer.
[0010] An embodiment of the present disclosure combines targeted nanopore sequencing with a classification algorithm. Both the classification algorithm and the nanopore sequencing system use the same target gene site. That is, the nanopore sequencing system sequences only the target gene site required by the classification algorithm to classify the cancer. In other words, the nanopore sequencing system selectively sequences certain polymers in a pool and rejects other polymers. Thus, the cancer classification of the present disclosure is highly flexible in target selection, efficiently performed on a single sample, and can be started immediately upon receiving a biological sample such as a frozen section, liquid biopsy sample, cell or FFPE (formalin-fixed paraffin-embedded tissue) sample.
[0011] Furthermore, the targeted nanopore sequencing of the present disclosure analyzes the initial nucleotide sequence, i.e., nucleotide data rather than raw measurement signals, to determine whether the initial nucleotide sequence corresponds to the at least one target gene site. By using the nucleotide sequence rather than raw signals (e.g., ion current signals) to match the initial nucleotide sequence to the at least one target gene site, reasonable computational resources are utilized, resulting in target enrichment. Furthermore, since the embodiments of the present disclosure do not utilize raw signal comparison to determine whether the initial nucleotide sequence corresponds to the at least one target gene site, there is no need to convert the reference genome (i.e., at least one target gene site) into signal space.
[0012] Although the above has been described for clarity using only a first polymer of a biological sample, it should be understood that hundreds, thousands, or even tens of thousands of polymers may be sequenced and used by the methods of the present disclosure to classify a type of cancer.
[0013] Nanopore sequencing is a third generation approach used to sequence polymers such as polynucleotides in the form of DNA or RNA. With nanopore sequencing, single molecules of DNA or RNA can be sequenced without the need for PCR amplification or chemical labeling of the sample. This allows nanopore sequencing to offer low-cost sequencing, high mobility for testing, and rapid sample processing with results displayed in real time.
[0014] A nanopore sequencing system includes a biological or solid membrane in which one or more nanopores are located. The membrane is surrounded by an electrolyte solution, which divides the electrolyte solution into two chambers. A bias voltage is applied across the membrane, which induces an electric field that causes the movement of charged particles, such as ions, in the electrolyte solution. A voltage drop is concentrated near and within the nanopore, so that when a charged particle is near the pore region (the "trap region"), it experiences a force from the electric field. Within the trap region, the movement of ions produces a steady ionic current that can be measured with electrodes placed on the membrane.
[0015] Polymers such as DNA also have a net charge that is subjected to forces from an electric field. Once inside the nanopore, the polymer moves through the nanopore due to electrophoretic, electroosmotic, and thermophoretic forces. Inside the nanopore, the polymer partially restricts the flow of ions, resulting in a decrease in ionic current. Various factors, such as shape, size, and chemical composition, result in different changes in the magnitude of the ionic current and duration of movement. Based on this modulation of the ionic current, different polymers can be sensed and identified.
[0016] According to some embodiments, which may be combined with other embodiments described herein, the nanopore and corresponding electrodes constitute a sensor element. In some embodiments, the nanopore sequencing system includes an array of sensor elements. The array of sensor elements may be configured to provide sequential measurements of polymers from selected sensor elements in a multiplexed manner.
[0017] An embodiment of the present disclosure uses selective sequencing, which refers to the ability of a nanopore sequencing system to determine whether a polymer should be fully sequenced while it is being sequenced. This requires rapid classification of the amperometric signal from the first part of a read (i.e., the first nucleotide sequence) to determine whether the polymer should be fully sequenced or rejected and replaced with a new polymer.
[0018] The determination of whether the polymer should be fully sequenced is based on a comparison between the initial nucleotide sequence of the first polymer and a reference genome, i.e., at least one target gene site. In particular, the initial nucleotide sequence of the first polymer of the biological sample is analyzed while the first polymer is moving through the nanopore, and it is determined whether the initial nucleotide sequence corresponds to the at least one target gene site. Sequencing of the first polymer continues only if the initial nucleotide sequence of the first polymer matches the corresponding nucleotide sequence of the target gene site.
[0019] An exemplary software module that can be used to perform targeted sequencing is ReadFish. ReadFish is an open source software that enables targeted nanopore sequencing of gigabase-sized genomes. At least one target gene site can be included in the ReadFish-based target to enable targeted sequencing.
[0020] In an embodiment of the present disclosure, ReadFish operates on nucleotide data rather than raw measurement signals to determine whether an initial nucleotide sequence corresponds to at least one target gene site.
[0021] According to some embodiments, which may be combined with other embodiments described herein, the initial nucleotide sequence of the first polymer is a subset of the nucleotide sequence of the first polymer obtained as the first polymer translocates partially through the nanopore. In particular, the subset of nucleotide sequences is located at the beginning or front of the first polymer, where "head" and "front" refer to the tip of the polymer that first enters the nanopore.
[0022] In a typical application, the decision as to whether a polymer should be fully sequenced can be made with sufficient accuracy after measuring a few hundred nucleotides of the polymer, for example up to 500 nucleotides, up to 300 nucleotides, or even up to 100 nucleotides, This is in comparison to the ability of nanopore sequencing systems to perform measurements on sequences ranging in length from hundreds to tens of thousands of nucleotides.
[0023] In an embodiment of the present disclosure, sequencing of the first polymer continues only if the initial nucleotide sequence of the first polymer matches a corresponding nucleotide sequence of a target gene site, but if the initial nucleotide sequence of the first polymer does not correspond to (or matches a corresponding nucleotide sequence of) the at least one target gene site, the first polymer is rejected.
[0024] If the first polymer is rejected, a further (second) polymer can be measured without completing the measurement of the first polymer. This is done while the first polymer is being measured, thus saving measurement time. In particular, the analysis can identify at an early stage in such a lead that further measurement of the polymer being measured is not necessary.
[0025] Rejection of the first polymer can occur in a variety of ways.
[0026] In a first example, rejecting the first polymer includes excluding the first polymer from the nanopore. For example, a voltage across the nanopore can be reversed to exclude the first polymer from the nanopore. The voltage across the nanopore can then be reset to admit a second polymer into the nanopore.
[0027] In a second embodiment, rejecting the first polymer comprises stopping measurements from the nanopore in which the first polymer is located. For example, the nanopore sequencing system includes an array of sensor elements. The array of sensor elements may be configured to perform sequential measurements of polymers from selected sensor elements in a multiplexed manner. In that case, rejecting the first polymer may comprise stopping measurements from a currently selected sensor element and starting measurements from a newly selected sensor element.
[0028] The targeted nanopore sequencing of the present disclosure analyzes an initial nucleotide sequence (i.e., nucleotide data, not raw measurement signals) to determine whether the initial nucleotide sequence corresponds to at least one target gene site. Thus, the method may include determining the initial nucleotide sequence from raw measurement data, such as ion current data, measured by the nanopore sequencing system described above.
[0029] In some embodiments, the step of determining the initial nucleotide sequence from raw measurement data comprises: obtaining raw measurement data, such as ionic current data, from the nanopore; and performing (real-time) base calling to obtain at least the first nucleotide sequence of the first polymer; may include.
[0030] According to some embodiments, which may be combined with other embodiments described herein, further raw measurement data of the first polymer obtained where the initial nucleotide sequence of the first polymer corresponds to the at least one target gene site may also be processed using base calling to determine the nucleotide sequence of the first polymer for determining the biological state for the purpose of classifying cancer.
[0031] The term "base call" as used throughout this disclosure refers to the process of assigning a nucleic acid base to a current change resulting from a nucleotide passing through the nanopore.
[0032] According to some embodiments, which can be combined with other embodiments described herein, the base calling uses a neural network, in particular a recurrent neural network (RNN).
[0033] Real-time base-calling software is available, for example, from Oxford Nanopore Technologies (ONT). ONT has developed a number of base-callers for nanopore sequencing data, initially using hidden Markov models and available through the metrichor cloud service. They were then replaced by neural network models running on central processing units and then GPUs. For real-time base-calling, ONT offers a variety of computational platforms with built-in GPUs (minIT, Mk1C, GridION, and PromethION). These devices enable real-time base-calling sufficient to keep pace with the nanopores generating the data. More recently, these base-callers have acquired a server-client configuration, so that raw signals can be passed to a server and nucleotide sequences can be sent back.
[0034] In embodiments of the present disclosure, GPU base calling can be used to deliver a real-time stream of nucleotide data from nanopore sequencing (e.g., using up to 512 channels simultaneously). At the same time, the GPU can base call completed reads, allowing optimized tools such as minimap to map reads as they are generated, allowing dynamic updates of both the target and reference genomes as results change.
[0035] According to some embodiments, which can be combined with other embodiments described herein, the step of selectively sequencing the polymer of the biological sample according to at least one target gene site in step a) comprises sequencing (i) the nucleotides of the first polymer corresponding to the at least one target gene site and (ii) a predetermined number of nucleotides upstream and / or downstream of the at least one target gene site. Thus, to ensure optimal targeting, (a) one or two flanks can be added to the target gene site.
[0036] According to some embodiments, which may be combined with other embodiments described herein, the predetermined number of nucleotides upstream and / or downstream of the at least one target gene site is 25 kb or up to 25 kb. In some embodiments, the predetermined number of nucleotides upstream and / or downstream of the at least one target gene site is 10 kb or up to 10 kb, or preferably 20 kb or up to 20 kb, or preferably 30 kb or up to 30 kb, or preferably 50 kb or up to 50 kb, or preferably 75 kb or up to 75 kb, or preferably 100 kb or up to 100 kb. Preferably, a 10 kb upstream flank and a 10 kb downstream flank are provided for every target gene site.
[0037] According to some embodiments, which may be combined with other embodiments described herein, the biological state used to classify cancer is selected from the group comprising (or consisting of): the mutational state and the epigenetic state of the nucleotide sequence of the first polymer.
[0038] The biological state may be derived from measurement data, such as raw measurement data or ion current data, obtained by the nanopore sequencing system, In other words, the raw measurement data, such as the ion current data, may provide information about the nucleotide sequence and the biological state.
[0039] According to some embodiments, which may be combined with other embodiments described herein, the biological state may be derived from the nucleotide sequences and / or the raw measurement data using one or more neural networks, in particular deep learning based neural networks.
[0040] Said mutation state can be copy number variation (CNV). Said copy number variation is a phenomenon in which a part of genome is repeated and the number of repeats in genome varies among individuals. CNV can be derived directly from the nucleotide sequence obtained by base calling.
[0041] The term "epigenetic state" refers to a measure of epigenetic changes (or functionally related changes of up-regulation and / or down-regulation) of gene activity of a particular gene site and / or gene in the genome of a biological sample. The epigenetic state includes epigenetic down-regulation and / or up-regulation of the activity of a gene site in the biological sample compared to the activity of the same gene site in a physiological tissue. Such down-regulation and / or up-regulation can be due to, for example, DNA methylation, histone modification, or other epigenetic effects.
[0042] The epigenetic state can be a methylation state. Nanopore sequencing allows for the detection of the methylation state of bases in DNA directly from the reads without the need for additional experimental techniques. The methylation state can be determined by an appropriate software module such as megalodon from Oxford Nanopore Technologies (ONT). Megalodon is a research command line tool that extracts highly accurate modified base and sequence variant calls from raw nanopore reads by anchoring informative base calling neural network outputs to a reference genome / transcriptome. Megalodon is publicly available at https: / / github.com / nanoporetech / megalodon.
[0043] The term "methylation status" as used herein may refer to the methylation state of a CpG position, and thus may refer to the presence or absence of 5-methylcytosine at a CpG site in genomic DNA. If none of an individual's DNA is methylated at a CpG site, the position is 0% methylated. If all of an individual's DNA is methylated at a specified CpG site, the position is 100% methylated. If only a portion of an individual's DNA, for example 50%, 75%, or 80%, is methylated at the CpG site, the CpG position is said to be 50%, 75%, or 80% methylated, respectively.
[0044] The term "methylation status" reflects the relative or absolute amount of methylation at a gene site, particularly at a CpG position. The terms "methylation" and "hypermethylation" are used interchangeably herein. When used in reference to a CpG position, they refer to a methylation status that corresponds to an increased abundance of 5-methylcytosine at a CpG site in DNA of a biological sample obtained from a patient, compared to the amount of 5-methylcytosine found at the CpG site in the same genomic position of a biological sample obtained from a healthy individual or from an individual suffering from a different class or type of tumor.
[0045] According to some embodiments, which may be combined with other embodiments described herein, the at least one target gene site corresponds to at least one CpG position.
[0046] As used herein, the term "CpG site" or "CpG position" refers to a region of DNA where a cytosine nucleotide is adjacent to a guanine nucleotide in the linear sequence of bases along its length, with the cytosine (C) being one phosphate (p) away from the guanine (G). Approximately 70% of human gene promoters are CpG-rich. Regions of the genome rich in CpG sites are known as "CpG islands." Cytosines in CpG dinucleotides can be methylated to form 5-methylcytosines. Methylation (i.e., introduction of a methyl group) at cytosines at CpG sites in the promoter of a gene can cause gene silencing, a feature found in many human cancers. In contrast, hypomethylation of CpG sites is generally associated with overexpression of cancer genes in cancer cells. In the context of this disclosure, the term "independent genomic CpG positions" means that each CpG position of a group of genomic CpG positions can be individually probed for its methylation status.
[0047] According to some embodiments, which may be combined with other embodiments described herein, the at least one target gene site is a plurality (or set) of target gene sites, and a respective biological state is determined for each target gene site of the plurality (or set) of target gene sites.
[0048] Preferably, the set of target gene sites comprises at least 10, preferably at least 20, or at least 30, or at least 40, or at least 50, or at least 60, or at least 70, or at least 80, or at least 90, or at least 100 target gene sites. In some embodiments, a set of target gene sites can be defined that can be used to characterize multiple cancer types. In other words, the same set of target gene sites can be used to characterize different cancer types.
[0049] As used herein, the term "gene site" refers to a region of DNA that includes or consists of a gene, particularly a gene or gene site, suitable for identifying a cancer type. In particular, the term "gene site" refers to a DNA sequence having a gene locus. A gene site may include additional base pairs upstream and / or downstream of a gene (e.g., up to 12 kb, preferably up to 10 kb, up to 8 kb, or up to 6 kb, or up to 4 kb, or up to 2 kb, upstream and / or downstream of a gene). Thus, the biological state of a gene site may refer to the biological state of the gene itself as well as the biological state of additional base pair sequences upstream and / or downstream of a gene.
[0050] In one embodiment of the present disclosure, DNA methylation can be used to find gene sites with pathological activity in the cancer genome. Thus, a set of gene sites that have the greatest impact on distinguishing different cancer types can be defined and used for both targeted nanopore sequencing and cancer classification.
[0051] Once the biological state of the target gene site is determined, the cancer type of the biological sample can be determined using a classification algorithm.The classification algorithm is trained based on the biological state data related to the at least one target gene site and different cancer types.Therefore, in an embodiment of the present disclosure, the classification of biological sample in cancer diagnosis is provided, using a classification algorithm that can be a machine learning (ML) algorithm.
[0052] The term "classification" refers to a procedure and / or algorithm that classifies individual items into groups or classes based on quantitative information about one or more item-specific characteristics (called traits, variables, features, etc.), based on statistical models and / or training sets of previously labeled items. Specifically, in the context of the present disclosure, classification preferably means determining to which particular cancer type a biological sample belongs, e.g., as determined by epigenetic characteristics.
[0053] The term "machine learning algorithm" as used throughout this application refers to an algorithm that builds a model based on training data to make predictions or decisions without being explicitly programmed. In particular, the term "classification" refers to a machine learning algorithm that classifies individual items into groups or classes based on quantitative information about one or more features (called traits, variables, characteristics, features, etc.) specific to the items, based on a statistical model and / or a training dataset of previously labeled items. Specifically, in the context of the present invention, classification preferably means determining which specific cancer type a biological sample belongs to (e.g., by its epigenetic state pattern).
[0054] The term "training dataset" in the context of this disclosure refers to a set of biological state data, such as genomic methylation data, for a large number of tumors that have been classified by prior art methods and are therefore of known tumor type.
[0055] Said classification algorithm can be any suitable algorithm for establishing correlation between data sets, i.e., biological state of DNA polymers of biological samples and biological state data derived from pre-classified cancer types.Methods for establishing correlation between data sets include, but are not limited to, discriminant analysis (DA) (e.g., linear-, quadratic-, regularized-DA), discriminant function analysis (DFA), kernel methods (e.g., SVM), non-parametric methods (e.g., k-nearest neighbor classifier), PLS (partial least squares), tree-based methods (e.g., CART, random forest), gradient boosting, generalized linear models (e.g., logistic regression), principal component-based methods (e.g., SIMCA), generalized additive models, fuzzy logic-based methods, neural networks, and genetic algorithm-based methods.
[0056] A person skilled in the art will have no problem selecting an appropriate method / algorithm for establishing a correlation between the biological state data of a biological sample and the biological state data derived from a pre-classified cancer type.
[0057] In one exemplary embodiment, the classification algorithm uses random forest analysis. As used herein, the term "random forest analysis" refers to a computational method based on the idea of using multiple different decision trees to calculate the overall most predicted class (mode). In a specific application, the mode will be either a tumor type or a class, based on how many decision trees predict the sample to match a particular class. The class predicted by the majority is selected as the predicted class for the sample. The different decision trees used in this algorithm are trained on randomly generated subsets of the training dataset and on a randomly selected set of the variables.
[0058] Examples of correlations between target gene sites and cancer types, as well as suitable classification algorithms, are provided in EP 3 268 492 B1 and Capper, D., Jones, D., Sill, M. et al.; DNA methylation-based classification of central nervous system tumours; Nature 555, 469-474 (2018); https: / / doi.org / 10.1038 / nature26000.
[0059]
[0013] Embodiments of the present disclosure allow for the characterization of cancer. As used throughout this disclosure, the term "characterization" includes one or more of classification, diagnosis, and stratification.
[0060] The term "diagnosis" is used herein to refer to the identification or classification of a molecular or pathological state, disease or condition. For example, "diagnosis" may refer to the identification of a particular type of cancer, for example, of the central nervous system (CNS). It is important to note that the present disclosure, in all its embodiments, is directed to strictly in vitro methods. No method steps of any embodiment are performed on the human or animal body.
[0061] The term "stratification" refers to classifying or grouping patients according to one or more predetermined criteria. In certain embodiments, stratification is performed in a diagnostic setting to group patients according to the prognosis of disease progression with or without treatment. In certain embodiments, stratification is used to distribute patients enrolled in clinical studies according to their individual characteristics. Stratification can be used to identify optimal treatment options for patients.
[0062] As used herein, the term "sample" or "biological sample" is used in the broadest sense. In the practice of the present disclosure, a biological sample is generally obtained from a subject. The sample may be any biological tissue or bodily fluid that can assay the biological state of the present disclosure. In many cases, the sample is a "clinical sample" (i.e., a sample obtained from or derived from a patient to be tested). The sample may also be an archived sample with a known history of diagnosis, treatment, and / or outcome.
[0063] Examples of biological samples suitable for use in the practice of the present disclosure include, but are not limited to, bodily fluids, such as blood samples (e.g., blood smears), and cerebrospinal fluid, brain tissue samples, spinal cord tissue samples, or bone marrow tissue samples (such as tissue or fine needle biopsy samples). Biological samples include preserved samples, such as alcohol-preserved samples and formalin-fixed paraffin-embedded (FFPE) tissue samples. "Biological samples" may also include tissue sections, such as frozen sections taken for histological purposes. The term "biological sample" also encompasses any material derived from processing a biological sample. Derived materials include, but are not limited to, cells (or their progeny) isolated from the sample, and nucleic acid molecules (DNA and / or RNA) extracted from the sample. Processing a biological sample may include one or more of filtration, distillation, extraction, concentration, inactivation of interfering components, addition of reagents, and the like.
[0064] Preferably, said biological sample is a cancer sample.
[0065] The term "cancer" or "tumor" is not limited to the stage, grade, histomorphological characteristics, invasiveness, aggressiveness or malignancy of the affected tissue or cell mass, and specifically includes stage 0 cancer, stage I cancer, stage II cancer, stage III cancer, stage IV cancer, grade I cancer, grade II cancer, grade III cancer, malignant cancer, primary cancer, and all other types of cancer, malignant tumors, and the like.
[0066] The term "cancer sample" or "tumor sample" as used herein refers to a sample obtained from a patient. The tumor sample can be obtained from a patient by routine means known to those skilled in the art, i.e., by biopsy (obtained by aspiration or puncture, excision, or any other surgical method that results in biopsy or excised cellular material). For sites that are difficult to reach with an open biopsy, a "closed" biopsy can be performed using a stereotaxic device through a small hole drilled by the surgeon in the skull. The stereotaxic device allows the surgeon to precisely position the biopsy probe in three-dimensional space, allowing access to almost anywhere in the brain. Thus, tissue can be obtained for the diagnostic method of the present disclosure. However, the actual removal of the sample from the patient is not part of the method of the present invention. Thus, "providing a cancer sample" simply relates to making the sample available for use in the laboratory without the step of first obtaining the sample from the patient.
[0067] In some embodiments, biological samples include, but are not limited to, biological fluids, cells, tissues, and cell lines containing biomarkers. In some embodiments, biological samples include, but are not limited to, primary cells, induced pluripotent cells (IPCs), hybridomas, recombinant cells, whole blood, stem cells, cancer cells, bone cells, chondrocytes, neuronal cells, glial cells, epithelial cells, skin cells, scalp cells, lung cells, mucosal cells, muscle cells, skeletal muscle cells, striated muscle cells, smooth muscle cells, cardiac cells, secretory cells, adipocytes, blood cells, red blood cells, basophils, eosinophils, monocytes, lymphocytes, T cells, B cells, neutrophils, NK cells, regulatory T cells, dendritic cells, or the like. These include, but are not limited to, cells, Th17 cells, Th1 cells, Th2 cells, bone marrow cells, macrophages, monocyte-derived stromal cells, bone marrow cells, spleen cells, thymus cells, pancreatic cells, oocytes, sperm, kidney cells, fibroblasts, intestinal cells, cells of the female or male reproductive tract, prostate cells, bladder cells, eye cells, corneal cells, retinal cells, sensory cells, keratinocytes, liver cells, brain cells, kidney cells, and colon cells, as well as transformed counterparts of these cell types.
[0068] According to another independent aspect of the present disclosure, there is provided a computer-readable storage medium having stored thereon computer-executable instructions that, when executed, cause a computer to perform the methods described herein.
[0069] The term "computer-readable storage medium" may refer to any storage device used to store data accessible by a computer, and any other means for providing access to data by a computer. Examples of storage device-type computer-readable media include magnetic hard disks, floppy disks, optical disks such as CD-ROMs and DVDs, magnetic tapes, memory chips, and the like.
[0070] According to another independent aspect of the present disclosure, a system for diagnosing cancer is provided, the system including one or more processors and a memory coupled to the one or more processors and including instructions executable by the one or more processors to perform the methods described herein.
[0071] The system may be a computer system. The term "computer system" may refer to a system having a computer, the computer comprising a computer-readable storage medium having embedded software for operating the computer.
[0072] As used herein, the term "software" is used interchangeably with "program" and refers to predetermined rules for operating a computer. Examples of software include software, code segments, instructions, computer programs, and programmed logic.
[0073] Further aspects, advantages, and features of the disclosure will be apparent from the claims, the description, and the accompanying drawings.
[0074] So that the above-mentioned characteristic aspects of the present disclosure may be understood in detail, a more particular description of the present disclosure briefly summarized above can be understood by reference to the embodiments. The accompanying drawings relate to embodiments of the present disclosure and are described below. [Brief description of the drawings]
[0075] [Figure 1] FIG. 1 shows a flowchart of a computer-implemented method for cancer diagnosis according to an embodiment described herein. [Figure 2A] Figure 2A shows a timeline of a conventional NGS panel sequencing and analysis pipeline, as well as a timeline of a conventional EPIC array analysis pipeline for neuropathology diagnosis. [Figure 2B] FIG. 2B shows a timeline of the RAPID-CNS sequencing and analysis pipeline for a single sample according to an embodiment of the present disclosure. [Diagram 3] FIG. 3 shows the clinically relevant changes and the agreement of the classification. [Figure 4] FIG. 4 shows the CNV plots obtained using RAPID-CNS, panel sequencing, and EPIC array analysis. [Diagram 5] FIG. 5 shows the MGMT promoter methylation values averaged across the MGMT promoter region. [Figure 6] FIG. 6 shows the copy number variation (CNV) profile of Ewing's sarcoma samples. [Figure 7] FIG. 7 shows a schematic of the results of the methylation classifier. [Figure 8] FIG. 8 shows the MGMT promoter methylation status of Ewing sarcoma samples. [Figure 9] FIG. 9 shows Rapid-CNS2 data for Ewing's sarcoma samples. [Figure 10] FIG. 10 shows the confidence scores for methylation-based classification of Ewing's sarcoma samples. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0076] Reference will now be made in detail to various embodiments of the present disclosure, one or more examples of which are illustrated in the figures. In the following description of the drawings, like reference numerals refer to like components. Generally, only the differences with respect to individual embodiments will be described. Each example is provided to illustrate the present disclosure, and not to limit it. Furthermore, features illustrated or described as part of one embodiment may be used in other embodiments or in combination with other embodiments to produce yet another embodiment. It is intended that this specification include such modifications and variations.
[0077] Figure 1 shows the rapid and comprehensive adaptive nanopore sequencing (RAPID-CNS or RAPD-CNS) for cancer diagnosis, especially for CNS tumors. 2 1 shows a flowchart of a computer-implemented method 100 for
[0078] Method 100 combines targeted nanopore sequencing with a classification algorithm. Both the classification algorithm and the nanopore sequencing system use the same target gene sites. That is, the nanopore sequencing system sequences only the target gene sites that the classification algorithm needs to classify the cancer. In other words, the nanopore sequencing system selectively sequences certain polymers in the pool while rejecting other polymers. Thus, the cancer classification of the present disclosure is highly flexible in target selection, efficiently performed on a single sample, and can be started immediately after receipt of the frozen section.
[0079] Detailed Description
[0080] First, a biological sample is provided. The biological sample may be obtained from a patient by routine means known to those skilled in the art, such as by biopsy. However, the actual removal of the biological sample from the patient is not part of the method of the present invention. Thus, "providing a biological sample" simply relates to making the sample available to the laboratory without first obtaining the biological sample from the patient.
[0081] A nanopore sequencing system is used to selectively sequence a polymer of the biological sample. The selective sequencing is performed using a reference genome, i.e., at least one target gene site. During the selective sequencing, a first polymer of the biological sample is moved through a nanopore of a nanopore sequencing system to obtain raw measurement data (Block 110).
[0082] An initial nucleotide sequence of the first polymer is derived from the raw measurement data, e.g., using base calling, and analyzed while the first polymer is partially translocated through the nanopore, in block 120. The initial nucleotide sequence of the first polymer can have a length of several hundred nucleotides, such as 500 nucleotides or less, 300 nucleotides or less, or 100 nucleotides or less.
[0083] In block 130, it is determined whether the initial nucleotide sequence corresponds to at least one target gene site.
[0084] If it is determined that the initial nucleotide sequence does not correspond to at least one target gene site, the first polymer is rejected and further (second) polymers can be measured without completing measurement of the first polymer (block 140).
[0085] If it is determined that the initial nucleotide sequence corresponds to at least one target gene site, the method 100 proceeds to block 150, where sequencing of the first polymer continues to obtain (raw) measurement data of the first polymer. Based on these measurement data of the first polymer, the nucleotide sequence of the first polymer can be determined. This nucleotide sequence of the first polymer can include the initial nucleotide sequence previously determined.
[0086] Then, in block 160 , the biological state of the nucleotide sequence of the first polymer corresponding to the at least one target gene site is determined based on the measurement data obtained in block 150 .
[0087] Finally, in block 170, a cancer type of the biological sample is determined using a classification algorithm based on the biological state of the nucleotide sequence of the first polymer, the classification algorithm being trained based on the at least one target gene site and biological state data associated with a cancer type.
[0088] While the above example has been described using only a first polymer of a biological sample for clarity, it should be understood that hundreds, thousands, or even tens of thousands of polymers may be sequenced and used to classify the cancer type in block 170. EXAMPLES
[0089] Example: Glioma patient samples
[0090] Brief Overview of the Examples
[0091] In addition to tissue structure, the WHO classification 2021 includes multiple molecular markers for routine diagnosis. Sequencing set-up for full molecular profiling requires significant investment, while batch processing of samples for sequencing and methylation profiling can delay turnaround time. This disclosure introduces RAPID-CNS, a nanopore-adaptive sequencing pipeline that enables comprehensive mutational, methylation, and copy number profiling of CNS tumors in a single third-generation sequencing assay. It can be performed on a single sample, does not require additional library preparation, and allows for flexible target selection.
[0092] Utilizing ReadFish (https: / / github.com / LooseLab / readfish), a toolkit enabling targeted nanopore sequencing, we sequenced DNA from 22 diffuse glioma patient samples with a MinION device (Oxford Nanopore Technologies Ltd). Target regions included the Heidelberg Brain Tumor NGS Panel and preselected CpG sites for methylation classification with an adaptive random forest classifier. Pathological alterations, copy number profiles, and methylation classes were called using a custom bioinformatics pipeline. These were compared with the corresponding NGS panel sequencing and EPIC array results.
[0093] Copy number profiles were found to be in perfect agreement with the EPIC array, with 94% of pathological mutations concordant with the NGS panel-seq. MGMT promoter status was accurately identified in all samples. Methylation families were detected with 96% concordance. Of the alterations critical for a combined diagnosis that fit the classification, 97% were concordant across the cohort (19 / 22 cases with perfect concordance, all sufficient for a definite diagnosis).
[0094] RAPID-CNS offers a rapid and flexible alternative to traditional NGS and array-based methods for SNV / Indel analysis, copy number alteration detection, and methylation classification. By modifying the target size and enabling real-time analysis, the turnaround time of approximately 5 days can be further reduced to less than 24 hours. This results in a low-capital approach that is cost-effective in low-throughput settings and beneficial in cases where an immediate diagnosis is required.
[0095] Detailed Description of the Preferred Embodiments
[0096] Molecular markers are decisive for brain tumor diagnosis. The 2021 WHO classification of brain tumors significantly increased the gene set required for routine evaluation. Multiple sequencing approaches are available. However, neuropathology laboratories cannot rely on off-the-shelf assays, as they do not cover genes relevant for neuro-oncology and contain large target regions that are not always necessary. Therefore, custom assays are usually set up when equipment for next-generation sequencing (NGS) is available. On the other hand, the benefits of custom neuropathology NGS panels can only be efficiently exploited if the number of cases is sufficient. Therefore, laboratories with low specimen submission volumes must pool samples over multiple weeks.
[0097] Herein, we introduce a RAPID-CNS-custom neuro-oncology-based workflow that uses third-generation sequencing for copy number profiling, mutation and methylation analysis, which is flexible in target selection, efficiently performed on a single sample and can be started immediately after receipt of frozen sections.
[0098] Nanopore sequencing has advantages over current NGS methods in terms of longer read sizes, short and easy library preparation, ability to call base modifications, real-time analysis, and portable sequencing equipment, all at low cost. However, small instruments such as MinION (Oxford Nanopore Technologies Ltd) produce low coverage data, making it difficult to detect pathological changes or hard-to-map regions such as the TERT promoter. Nanopore offers a "ReadUntil" adaptive sampling toolkit that allows for real-time rejection of reads during sequencing. ReadFish exploits this capability to enable targeted adaptive sequencing without the need for an additional library preparation step. This allows for significantly improved coverage of "targeted" regions through enrichment during sequencing, ensuring the detection of clinically relevant changes.
[0099] RAPID-CNS is enabled by adaptive nanopore sequencing through ReadFish and is performed using a portable MinION device (Oxford Nanopore Technologies Ltd). We developed target regions covering the brain tumor NGS panel and CpG sites required for methylation classification and performed ReadFish-based sequencing on 22 diffuse glioma samples that had previously undergone brain tumor NGS panel and Infinium Methylation EPIC array (EPIC) analysis.
[0100] Samples were selected to cover pathological alterations (IDH, 1p / 19q co-deletion, chr7 gain / chr10 loss, TERT promoter, EGFR amplification, CDKN2A / B deletion, MGMT status) and associated methylation classes identified by conventional methods.
[0101] Cryopreserved brain tumor tissue was prepared for nanopore sequencing using ONT's SQK-LSK109 Ligation Sequencing Kit. Incubation times and other parameters were optimized to improve library quality, data generation, and on-target rates (see Supplemental Methods section).
[0102] A single sample was loaded onto a FLO-MIN106 R9.4.1 flow cell and run on a MinION 1B (Oxford Nanopore Technologies Ltd). ReadFish controlled the sequencing in real time and was run using a consumer notebook equipped with an 8GB NVIDIA RTX 2080 Ti GPU. Samples were sequenced for up to 72 hours and our targets covered 5.56% of the whole genome. By reducing the size of the target region, sequencing time could be reduced to 24 hours.
[0103] Sequencing data were analyzed using a bioinformatics pipeline customized for neuro-oncology targets. SNVs were filtered for clinical relevance by their 1000 genomes population frequency (<0.01) and COSMIC annotation. Copy number alterations were estimated using coverage depth of mapped reads.
[0104] Nanopore sequencing has the advantage of being able to estimate base modifications from a single DNA sequencing assay. Methylated bases were identified using megalodon, a deep neural network-based modified base caller. The output of megalodon was used to calculate methylation values across targeted CpG sites and to assess the methylation status of the MGMT promoter.
[0105] A random forest classifier based on a previously published reference set (Capper, D., Jones, D., Sill, M. et al.; DNA methylation-based classification of central nervous system tumours; Nature 555, 469-474 (2018); https: / / doi.org / 10.1038 / nature26000) was trained to predict the methylation class for samples. MGMT promoter methylation status was assigned by averaging methylation values across all CpG sites in the MGMT promoter region (see Supplementary Results section). The average run time of RAPID-CNS from tissue collection to reporting was less than 5 days.
[0106] Nanopore sequencing significantly reduced library preparation time to 3.5 h compared to 48 h for panel sequencing and 72 h for EPIC array (Figure 2A and Figure 2B). This allowed the integration of both data categories into one laboratory workflow. Despite differences in sequencing technologies and method-specific data analysis pipelines, the concordance rate of detected SNVs was 78%, regardless of clonality or clinical relevance. Importantly, diagnostically relevant pathogenic variants, such as IDH1 R132H / S and TERT promoter, were concordant in 22 / 22 and 19 / 22 samples, respectively (Figure 3: Colored blocks indicate the presence of a variant, and the concordance rate of detected variants is shown in the legend. Triangle indications for methylation classes indicate samples with concordance for methylation families, and blocks also indicate concordance for subclasses. The percentages on the left indicate concordance of alterations in all samples).
[0107] Furthermore, we derived copy number plots (CNPs) from copy number levels calculated for nanopore data (Supplementary Fig. 1a). Plots generated with nanopore data showed significantly better resolution than those obtained with panel sequencing data (Fig. 4 and Fig. 5). Complete concordance with EPIC array analysis was found for CNV levels in all samples. RAPID-CNS also enabled gene-level CNV detection. Of the changes critical for making an integrated molecular diagnosis, 217 / 220 were consistent across the entire cohort (complete concordance in 19 / 22 cases).
[0108] The inclusion of CpGs relevant for methylation classification in the ReadFish-based targets also enabled methylation class prediction. The ability of nanopore sequencing to robustly aid in methylation classification using low-pass whole-genome nanopore sequencing data has been previously demonstrated by nanoDx (https: / / www.medrxiv.org / content / 10.1101 / 2021.03.06.21252627v1.full). The methylation family predicted by RAPID-CNS was consistent with the corresponding EPIC array-based classification in 21 / 22 cases, and the methylation subclass was consistent in 14 cases. The MGMT promoter status was also consistent with the corresponding EPIC array analysis in all cases. Nanopore identified the MGMT promoter status as unmethylated in one sample consistent with the EPIC array, which was assigned as methylated by pyrosequencing.
[0109] The target regions in RAPID-CNS can be easily modified by editing the BED file, which in principle could reduce sequencing times over this study. As no additional library preparation steps are required, it is possible to change the target regions for each sample as needed. As the MinION is a portable handheld device, it also makes it a reasonable choice for smaller neuropathology laboratories.
[0110] We used GPUs to run ReadFish, but it can also be run using CPUs with full efficiency. Taken together, the RAPID-CNS approach can be set up with low capital costs, is cost-effective even in low throughput settings, and provides a rapid and flexible alternative to traditional NGS methods for SNV / Indel analysis, as well as for methylation classification and copy number change detection when corresponding methods (e.g. methylation arrays) are not feasible.
[0111] Supplementary method
[0112] 1. Optimized Nanopore Library Preparation for Adaptive Sampling
[0113] 40 × 10 μm sections were prepared from cryopreserved tumor tissues with established molecular markers (IRB approval 2018-614N-MA, 005 / 2003) if the tumor cell content (based on H&E staining) was greater than 60%. DNA was then extracted using the Promega Maxwell RSC Blood DNA Kit (catalog no. AS1400, Promega) on a Maxwell RSC 48 instrument (AS8500, Promega) according to the manufacturer's instructions. DNA concentration was measured with a microplate reader (FLUOStar Omega, BMG Labtech) using the Invitrogen Qubit DNA BR Assay Kit (Q32851, Thermo Fisher Scientific).
[0114] The DNA was then sheared to approximately 9–11 kb using g-TUBE (Covaris) for 120 s at 7200 rpm in a total volume of 50 μl. Fragment lengths were assessed on an Agilent 2100 Bioanalyzer (catalog no. G2939A, Agilent Technologies) using the Agilent DNA 12000 Kit (catalog no. 5067-1508, Agilent Technologies).
[0115] Sequencing libraries were prepared using the SQK-LSK109 Ligation Sequencing Kit with the following modifications: 48 μl of sheared DNA (2–2.5 μg) was introduced into the end-prep reaction, and control DNA was omitted. The end-prep reaction was modified to be incubated at 20°C for 30 min, 65°C for 30 min, and then cooled to 4°C in a thermal cycler. Washing was modified to use AMPure XP beads and 80% ethanol, and the elution time was changed to 5 min. Adaptor ligation was incubated at room temperature for 60 min. The ligation mix was then incubated with AMPure XP beads at 0.4× for 10 min, cleanup was performed with Long Fragment Buffer (LFB), and the final library was eluted in a total volume of 31 μl.
[0116] Library concentrations were measured on a benchtop Quantus fluorometer (Promega) using the Invitrogen Qubit DNA HS Assay Kit (Q32851, Thermo Fisher Scientific). Libraries were loaded (500–600 ng) onto a FLO-MIN106R9.4.1 flow cell with a minimum of 1100 pores available, following an FC Check prior to loading. Flow cells were flushed after approximately 24 h, a total of two times per sample, using a Flow Cell Wash Kit (EXP-WSH003) according to the manufacturer's instructions. All sequencing runs were performed on a MinION 1B (Oxford Nanopore Technologies).
[0117] 2. ReadFish
[0118] Targeted nanopore sequencing was performed in real time using a custom panel with ReadFish on an 8GB NVIDIA RTX 2080 Ti consumer notebook. Targets included regions from a neuropathology panel and CpG sites that were useful for classification with a random forest methylation classifier (available on GitHub). Sites were flanked with 25 kbp to ensure optimal targeting by ReadFish. ReadFish was run using Guppy 4.4.1 in fast base calling mode.
[0119] 3. Classification of methylation
[0120] To classify nanopore sequencing-derived DNA methylation profiles of central nervous system tumors, we trained a random forest classifier on the public 450k methylation array reference dataset from MNP classifier version 11 (GSE90496). The dataset was preprocessed as described in Capper, D., Jones, D., Sill, M. et al.; DNA methylation-based classification of central nervous system tumours; Nature 555, 469-474 (2018); https: / / doi.org / 10.1038 / nature26000.
[0121] For a batch of 22 nanopore sequencing samples, the intersection of CpG probes measured for all samples was selected to train the classifier. The methylation array dataset was reduced to these 3,285 probes.
[0122] Nanopore sequencing often measures CpG probes at low coverage, leading to discretely distributed methylation values, i.e., (0, 0.5, 1). Because nanopore sequencing often cannot detect finer methylation differences for all CpG probes, we trained the RF classifier on dichotomized methylation values. This follows the assumption that partitioning rules learned on binary data are more robust and can be applied to methylation signals from nanopore sequencing data.
[0123] After dichotomizing the reduced reference methylation dataset, an RF was trained on 1000 trees and the resulting permutation-based variable importance measure was applied to select the 1000 CpGs with the highest variable importance to train the final RF, again on 1000 trees. The out-of-bag error of this classifier was 4%.
[0124] Supplementary results
[0125] 1.RAPID-CNS analysis pipeline
[0126] The bioinformatics pipeline requires raw FAST5 files as input. Full instructions for setting up the analysis are available on GitHub. Once set up, RAPID-CNS runs the entire analysis with a single command. It can be run on an LSF cluster or a GPU workstation. Base calling and subsequent SNV and CNV detection are completed within 10 hours, while methylation calling and classification require an additional 12 hours.
[0127] 2.SNV detection
[0128] ANNOVAR annotated tables for all nanopore sequencing samples and the corresponding panel sequencing results are available to us.
[0129] 3.CNV detection
[0130] Copy number plots obtained using the RAPID-CNS pipeline show higher resolution and clearer visualization of copy number levels compared to NGS panel sequencing (Figure 4, left and center). When the mapped read depth is calculated, the detected copy number variations are comparable to the EPIC array results (Figure 4, left and right). The normalized read depth is shown on the Y-axis, with "2" indicating the average autosomal level. Furthermore, the genes covered by the copy number variation and their zygosity are annotated and output as an Excel file (available to us).
[0131] 4.MGMT promoter methylation
[0132] The two probes used by the MGMT-STP27 approach were not reliably covered in all samples analyzed. Methylation frequencies across all CpG sites covering the MGMT promoter region were averaged. Using pyrosequencing as the gold standard, we found a significant difference in the average methylation between methylated and unmethylated samples (Wilcoxon rank sum test p-value = 2.719e-06). As shown in Figure 5, a threshold of 10% was assigned to the MGMT promoter methylation status.
[0133] Additional Examples
[0134] We analyzed additional tumor types to show that the method, especially the classification, of this application is also valid for tumors. Sequencing was performed with the new SQK-LSK110 kit on the GridION device, demonstrating the versatility of the approach. Analysis was primarily performed simultaneously with base calling, methylation calling, and alignment. This reduced the analysis time from data to PDF report from up to 5 days to less than 8 hours compared to performing base calling, methylation calling, and alignment separately.
[0135] Ewing sarcoma samples
[0136] One Ewing sarcoma sample was analyzed. Figure 6 shows the copy number variation (CNV) profile of the Ewing sarcoma sample, Figure 7 shows the results of the array-based methylation classifier in a schematic manner, Figure 8 shows the MGMT promoter methylation status calculated using the array data, and Figure 9 shows the corresponding RAPID-CNS data. As can be seen, there is a perfect agreement between the CNV profile, the methylation classification, and the MGMT promoter methylation status. Figure 10 shows that the methylation-based classification achieved a confidence score of 99.5%.
[0137] Metastatic samples
[0138] Nine metastatic samples were analyzed. We found (i) perfect concordance in CNV profiles of all samples, (ii) perfect concordance in MGMT promoter status of all samples, with low SNV concordance under the current prioritization and filtering scheme (74 false negatives, 52 false positives, 22 true positives; more than five metastatic samples with concordant panel sequencing data). This could be improved by better filtering and prioritization of variants.
[0139] Glioma samples
[0140] 95 glioma samples were analyzed. We found that there was (i) perfect concordance of CNV profiles of all samples with EPIC array, (ii) perfect concordance of MGMT promoter status with the EPIC array, (iii) perfect concordance of IDH1 / TERTp mutation status detected by NGS panel sequencing, and (iv) concordance of methylation family levels in all samples with the EPIC array (except for three with very low input DNA concentration predicted as control group).
[0141] Meningioma samples
[0142] One meningioma sample was analyzed, and we found that (i) the CNV profile was perfectly concordant with the EPIC array, (ii) the MGMT promoter status was perfectly concordant with the EPIC array, and (iii) the methylation class was concordant with the EPIC array.
[0143] An embodiment of the present disclosure combines targeted nanopore sequencing with a classification algorithm. Both the classification algorithm and the nanopore sequencing system use the same target gene sites. That is, the nanopore sequencing system sequences only the target gene sites required by the classification algorithm to classify the cancer. In other words, the nanopore sequencing system selectively sequences certain polymers in a pool while rejecting other polymers. Thus, the cancer classification of the present disclosure is highly flexible in target selection, efficiently performed on a single sample, and can begin immediately upon receipt of a frozen section.
[0144] Furthermore, the targeted nanopore sequencing of the present disclosure analyzes the initial nucleotide sequence, i.e., nucleotide data, rather than raw measurement signals, to determine whether the initial nucleotide sequence corresponds to at least one target gene site. By using nucleotide sequence, rather than raw signals (e.g., ion current signals), to match the initial nucleotide sequence to at least one target gene site, reasonable computational resources are utilized, resulting in target enrichment. Furthermore, since the embodiment of the present disclosure does not utilize raw signal comparison to determine whether the initial nucleotide sequence corresponds to at least one target gene site, there is no need to convert the reference genome, i.e., the at least one target gene site, into signal space.
[0145] While the forgoing is directed to embodiments of the present disclosure, other and further embodiments of the present disclosure may be devised without departing from the basic scope thereof, which scope is determined by the appended claims.
Claims
1. A method for the detection of a tumor, comprising: obtaining measurement data for one or more target gene sites of a biological sample obtained from an organism, said measurement data comprising methylation data; inputting the measurement data into a classification model, wherein the classification model has been trained with a machine learning algorithm using a training dataset comprising methylation status data associated with one or more cancer types and one or more gene sites, including one or more target gene sites, and wherein the classification model is configured to output a classification of one or more cancer types; outputting, by the classification model, a specific cancer type detected at the one or more target gene sites from the biological sample based on the methylation data; A method comprising:
2. The method described in claim 1, wherein the one or more target gene sites are a set or a plurality of target gene sites, and wherein a respective biological state is determined for each target gene site of the set or a plurality of target gene sites.
3. The method described in claim 2, wherein the set or plurality of target gene sites comprises at least 10, preferably at least 20, or at least 30, or at least 40, or at least 50, or at least 60, or at least 70, or at least 80, or at least 90, or at least 100 target gene sites.
4. The method described in claim 2, wherein the set of target gene sites is used to characterize multiple cancer types, and the set of target gene sites is used to characterize different cancer types.
5. The method described in claim 1, wherein the methylation data defines an epigenetic state pattern including one or more target gene sites from a biological sample.
6. The method described in claim 5, wherein the epigenetic state pattern is a methylation state pattern of at least one or more CpG positions, and the methylation state of at least one or more CpG positions indicates the total or partial presence or absence of 5-methylcytosine at one CpG position in genomic DNA, respectively.
7. The method described in claim 5, wherein the classification of the specific cancer type by the classification model is based on epigenetic state patterns.
8. The method described in claim 1, wherein the measurement data from biological samples and / or the training dataset include data from samples obtained by sequencing and methylation profiling.
9. The method described in claim 2, wherein the output of the classification model is based on the biological state of the nucleotide sequence of the biological sample.
10. The method of claim 1, wherein the classification of the specific cancer type is output as a digital file or a printed document.
11. A method for determining whether a patient has a cancer type-specific epigenetic pattern, comprising: executing computer instructions for stratifying at least one patient to one or more treatment options, wherein the at least one patient has a disease caused by the epigenetic pattern; 6. The method of claim 5, further comprising outputting one or more treatment options for treating the particular cancer type in the at least one patient.
12. The method of claim 1, wherein the method is for detecting tumors such as central nervous system tumors and / or sarcomas.
13. The computer-implemented method (100) for detecting cancer of claim 1, further comprising: a) selectively sequencing a polymer of a biological sample according to at least one target gene site by translocating said polymer through a nanopore of a nanopore sequencing system; (i) analyzing an initial nucleotide sequence of a first polymer of the biological sample while the first polymer is moving through a nanopore of the nanopore sequencing system to determine whether the initial nucleotide sequence corresponds to the at least one target gene site (120); and (ii) continuing the sequencing of the first polymer to obtain measurement data for the first polymer only if the initial nucleotide sequence of the first polymer corresponds to the at least one target gene site (150); b) determining (160) the biological state of the nucleotide sequence of the first polymer corresponding to the at least one target gene site based on the measurement data; and c) classifying the cancer using a classification algorithm based on the biological state of the nucleotide sequence of the first polymer (170), wherein the classification algorithm is trained based on biological state data relating to the at least one target gene site and cancer type.
14. the first nucleotide sequence of the first polymer is 14. The computer-implemented method (100) of claim 13, wherein the subset of the nucleotide sequence of the first polymer obtained when the first polymer partially translocates through the nanopore.
15. Step a) 15. The computer-implemented method (100) of claim 13 or 14, further comprising: rejecting (140) the first polymer if the initial nucleotide sequence of the first polymer does not correspond to the at least one target gene site.
16. Rejecting (140) the first polymer excluding the first polymer from the nanopore; or The computer-implemented method (100) of claim 15, comprising ceasing measurements from the nanopore.
17. 16. The computer-implemented method (100) of claim 15, further comprising taking a measurement from a second polymer that translocates through the nanopore after the first polymer is rejected.
18. Step a) acquiring raw measurement data from the nanopore; and 14. The computer-implemented method (100) of claim 13, further comprising performing base calling to obtain at least the initial nucleotide sequence of the first polymer.
19. 20. The computer-implemented method (100) of claim 18, wherein the base calling uses a neural network, in particular a recurrent neural network.
20. Selectively sequencing the polymer of the biological sample according to at least one target gene site of step a) comprises: (i) nucleotides of the first polymer corresponding to the at least one target gene site; and 14. The computer-implemented method (100) of claim 13, comprising: (ii) sequencing a predetermined number of nucleotides upstream and / or downstream of the at least one target gene site.
21. 14. The computer-implemented method (100) of claim 13, wherein the biological state is selected from the group comprising a mutational state and an epigenetic state of the nucleotide sequence of the first polymer.
22. 22. The computer-implemented method (100) of claim 21, wherein the mutational state is a copy number variation and / or the epigenetic state is a methylation state.
23. 23. The computer-implemented method (100) of claim 21 or claim 22, wherein the biological state of the nucleotide sequence of the first polymer is derived from the measurement data using one or more neural networks, in particular deep learning-based neural networks.
24. 14. The computer-implemented method (100) of claim 13, wherein the at least one target gene site comprises at least one CpG position.
25. 14. The computer-implemented method (100) of claim 13, wherein the at least one target gene site is a plurality of target gene sites, and a respective biological state is determined for each of the plurality of target gene sites.
26. A computer-readable storage medium having stored thereon computer-executable instructions that, when executed, cause a computer to perform the method (100) of claim 13.
27. The computer-readable storage medium of claim 26 storing instructions for tumor detection, comprising: When executed by one or more processors, the one or more processors: obtaining measurement data for one or more target gene sites of a biological sample, said measurement data comprising methylation data; inputting the measurement data into a classification model, wherein the classification model has been trained with a machine learning algorithm using a training dataset comprising methylation status data associated with one or more cancer types and one or more gene sites, including one or more target gene sites, and wherein the classification model is configured to output a classification of one or more cancer types; and outputting, by the classification model, a particular cancer type detected in the one or more target gene site samples based on the methylation data.
28. The computer-readable storage medium of claim 26 storing a classification model for tumor detection, comprising: When executed by one or more processors, the one or more processors: receiving measurement data of one or more target gene sites of a biological sample, said measurement data comprising methylation data; inputting the measurement data into a classification model, wherein the classification model has been trained with a machine learning algorithm using a training dataset comprising methylation status data associated with one or more cancer types and one or more gene sites, including one or more target gene sites, and wherein the classification model is configured to output a classification of one or more cancer types; and outputting, by the classification model, a particular cancer type detected in the one or more target gene site samples based on the methylation data.
29. A method for the detection of a tumor, comprising: receiving, by a classification model, measurement data of CpGs of a biological sample, said measurement data including methylation data; wherein the classification model is trained with a machine learning algorithm using a training dataset comprising methylation status data associated with one or more cancer types and one or more gene sites including one or more target gene sites, and the classification model is configured to output a classification of one or more cancer types; and outputting, by the classification model, a particular cancer type detected in the one or more target gene site samples based on the methylation data.
30. The method of claim 29, wherein the classification of the specific cancer type is output as a digital file or a printed document.
31. 1. A system for detecting cancer, comprising: one or more processors; and 14. A system comprising: a memory coupled to said one or more processors and including instructions executable by said one or more processors to implement the method (100) of claim 13.
32. The system of claim 31, a classification model stored in a memory; one or more processors communicatively coupled to the memory and configured to access the classification model; Computer instructions stored on a computer-readable medium, which when executed by the one or more processors, cause the one or more processors to: obtaining measurement data for one or more target gene sites of a biological sample, said measurement data comprising methylation data; inputting the measurement data into a classification model, wherein the classification model has been trained with a machine learning algorithm using a training dataset comprising methylation status data associated with one or more cancer types and one or more gene sites, including one or more target gene sites, and wherein the classification model is configured to output a classification of one or more cancer types; and outputting, by the classification model, a particular cancer type detected in the one or more target gene site samples based on the methylation data.
33. The system described in claim 31, wherein the one or more target gene sites are a set or a plurality of target gene sites, and wherein a respective biological state is determined for each target gene site of the set or a plurality of target gene sites.
34. The system of claim 33, wherein the set or plurality of target gene sites comprises at least 10, preferably at least 20, or at least 30, or at least 40, or at least 50, or at least 60, or at least 70, or at least 80, or at least 90, or at least 100 target gene sites.
35. A system as described in claim 33, wherein the set of target gene sites is used to characterize multiple cancer types, and the set of target gene sites is used to characterize different cancer types.
36. The system described in claim 32, wherein the methylation data defines an epigenetic state pattern including one or more target gene sites from a biological sample.
37. The system described in claim 36, wherein the epigenetic state pattern is a methylation state pattern of at least one or more CpG positions, and a set of methylation states of at least one or more CpG positions each indicates the total or partial presence or absence of 5-methylcytosine at one CpG position in genomic DNA.
38. The system described in claim 36, wherein the classification of the specific cancer type by the classification model is based on epigenetic state patterns.
39. The system described in claim 32, wherein the measurement data from biological samples and / or the training dataset include data from samples obtained by sequencing and methylation profiling.
40. The system described in claim 33, wherein the output of the classification model is based on the biological state of nucleotide sequences obtained from the biological sample.
41. The system described in claim 32, wherein the classification of the specific cancer type is output as a digital file or a printed document.
42. The computer instructions, when executed by one or more processors, cause the one or more processors to: executing computer instructions for stratifying at least one patient to one or more treatment options, wherein the at least one patient has a disease caused by a cancer type-specific epigenetic status pattern; and outputting one or more treatment options for treating the particular cancer type in the at least one patient.
43. The computer instructions, when executed by one or more processors, cause the one or more processors to:
43. The system of claim 42, further comprising treating the at least one patient with at least one of the one or more treatment options.
44. The system of claim 32, wherein the output includes an output regarding classification of tumors such as central nervous system tumors and / or sarcomas.