Method and system for identifying tumor origin
By receiving and analyzing molecular phenotypic data from patient samples through a computer system, and using logistic regression and machine learning algorithms to classify cancer conditions and tissues of origin, the accuracy and cost issues of cancer screening in existing technologies have been resolved, enabling efficient early detection and accurate localization of various cancer types.
Patent Information
- Application Number
- CN202480050173.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-31
- Filing Date
- 2024-07-29
- Publication Date
- 2026-03-03
AI Technical Summary
Existing cancer screening technologies lack specificity and accuracy in early detection and identification of the tissue of origin for various cancer types, especially in asymptomatic patients, leading to a wide range of assessments, high costs, and increased complexity.
Using a computer-based approach, molecular phenotypic data from patient samples is received, and logistic regression models and machine learning algorithms are used to classify cancer status and tissue of origin. Combined with genetic and epigenetic information, including methylation status and histone modifications, binary and multi-category classifications of various cancer types are performed.
It achieves robust tissue assignment across multiple cancer types, improves the accuracy and efficiency of early cancer detection, reduces detection costs and complexity, and can guide more accurate diagnostic and treatment decisions.
Smart Images

Figure CN121605489A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims priority to U.S. Provisional Application No. 63 / 516,681, filed July 31, 2023, which is incorporated herein by reference in its entirety. background
[0002] Cancer is a leading cause of disease worldwide. Every year, tens of millions of people around the world are diagnosed with cancer, and more than half of those diagnosed eventually die from it. In many countries, cancer is the second leading cause of death after cardiovascular disease. Early detection is associated with improved outcomes for many cancers.
[0003] Several screening tests are available for detecting cancer. A physical examination and history check general health signs, including examination for signs of disease, such as lumps or other unusual physical symptoms. The patient's health habits and history of past illnesses and treatments are also collected. Laboratory tests are another type of screening test and may involve medical procedures to obtain samples of tissue, blood, urine, or other substances in the body before laboratory testing is performed. Imaging procedures screen for cancer by generating visual representations of areas inside the body. Genetic tests detect harmful mutations in certain genes associated with some types of cancer. Genetic tests are particularly useful for many diagnostic methods.
[0004] Despite these advances, there is a significant need for larger-scale screening technologies to improve the understanding of current findings, which focus on symptom indications and advanced cancer. To this end, multiple cancer detection methods with high specificity, clinically useful sensitivity, and highly accurate tissue of origin (TOO) identification will limit the scope, cost, and complexity of assessing patients, including asymptomatic individuals.
[0005] As an example, this paper describes the detection and localization of multiple cancer types using cell-free DNA (cfDNA) or other analytes. Using the methods and systems described herein, robust TOO allocation across a wide range of tumor types is achieved and can guide diagnostic assessment. In short, a binary classification is first performed to determine the cancer status, such as whether the sample has cancer. Samples classified as cancer then undergo a second model, which determines the TOO by classifying the sample as one of the cancer types. As shown in this paper, population-scale studies of cancer cfDNA demonstrate consistent performance across various representative screening cohorts for multiple cancer types. Invention Overview
[0006] This document describes a computer-implemented method comprising: receiving one or more datasets in a computer system including one or more hardware processors and one or more computer-readable storage media, wherein the datasets include molecular phenotypes obtained from test samples from a patient, and wherein the computer-readable media includes instructions that, when executed by the processor, cause one or more hardware processors to perform one or more classifications of the test samples. In other embodiments, the molecular phenotype includes methylation state, histone modifications, chromatin state, fragment length, or transcription factor occupancy of more than one genomic region. In other embodiments, the molecular phenotype includes epigenetic data, genomic sequence data, proteomic data, microbiome data, imaging data, histological data, and / or metadata. In other embodiments, the epigenetic data includes methylation, histone acetylation, chromatin state, or DNA circularization interaction data. In other embodiments, the method includes performing a first classification and a subsequent second classification, wherein the second classification is performed only if the first classification includes a predetermined category. In other embodiments, the first and second classifications are performed using a logistic regression model. In other embodiments, the first classification is performed using a logistic regression model, and the second classification is performed using Naive Bayes, decision trees, support vector machines (SVM), random forest classifiers, k-nearest neighbors (KNN), or neural networks. In other embodiments, the first classification is the cancer status, and the second classification is the cancer type. In other embodiments, the cancer status is a binary category including both cancer and non-cancer states. In other embodiments, for originating cancer signals and / or originating tissues, the cancer status and / or cancer type include classifications based on one or more of the following conditions (including methylation status): In some embodiments, sequencing data is analyzed to detect specific cancer-associated signals. These signals include genetic mutations, epigenetic alterations (such as DNA methylation), and gene expression patterns indicating the presence of cancer. Genetic and epigenetic information includes a variety of epigenetic markers or functional elements, such as TFB, CTCF binding sites, genetic variations such as CNV, SNV, insertions / deletions, fusions, mRNA expression, fragment omics patterns, fragment omics levels, fragment endpoint density, and histone acetylation or methylation markers associated with the following: balanced enhancers including H3K4me1, H3K27ac, H3K27me3; promoter regions including H3K4me3, H3 / H4ac, H3K4me1, H3K27me3, H3K9me3 and / or H3.3; and open chromatin including H3Ac and H4Ac, H3K4me1, H3K4me2, H3K4me3, H2BK120ub, H3.3, and H3S10ph.
[0007] In other embodiments, the cancer type includes breast cancer, colorectal cancer, lung cancer, bladder cancer, pancreatic cancer, ovarian cancer, liver cancer, stomach cancer, esophageal cancer, kidney cancer, melanoma, gallbladder cancer, or uterine cancer. In other embodiments, the cancer type also includes the tissue from which the cancer originated in the patient. In other embodiments, the cancer type also includes the tissue from which cancer of unknown primary (CUP) originated in the patient. In other embodiments, the classification is based on a set of cancer-specific models. In other embodiments, each cancer-specific model outputs a score for the cancer type. In other embodiments, a cancer type prediction is made when the score exceeds a threshold. In other embodiments, no cancer type prediction is made when the score is below a threshold. In other embodiments, a tumor score is estimated for the cancer type. In other embodiments, all scores below the threshold output a tumor score of zero (TF=0) and a cancer-free label. In other embodiments, the sample includes cell-free DNA (cfDNA). In other embodiments, the sample includes blood, plasma, saliva, or urine. In other embodiments, the sample includes biological fluids, biological solids, or biological tissue.
[0008] This document describes a method comprising: obtaining or having obtained a sample from a subject; detecting one or more features in the sample; and classifying the subject's cancer status. In other embodiments, the cancer status includes determining the tissue of origin for one or more cells in the sample. In other embodiments, the tissue of origin for one or more cells in the sample. In other embodiments, one or more features include methylation status, histone modifications, chromatin status, fragment length, or transcription factor occupancy of more than one genomic region. In other embodiments, one or more features include epigenetic data, genomic sequence data, proteomic data, microbiome data, imaging data, histological data, and / or metadata. In other embodiments, epigenetic data includes methylation, histone acetylation, chromatin status, or DNA circularization interaction data. In other embodiments, a first classification and a subsequent second classification are performed, wherein the second classification is performed only if the first classification includes a predetermined category. In other embodiments, the first and second classifications are performed using a logistic regression model. In other embodiments, the first classification is performed using a logistic regression model, and the second classification is performed using Naive Bayes, decision trees, support vector machines (SVM), random forest classifiers, k-nearest neighbors (KNN), or neural networks. In other embodiments, the first category is cancer status, and the second category is cancer type. In other embodiments, cancer status is a binary category including both cancer and non-cancer states. In other embodiments, for originating cancer signals and / or originating tissue, cancer status and / or cancer type include classifications based on one or more of the following conditions (including methylation status): In some embodiments, sequencing data is analyzed to detect specific cancer-associated signals. These signals include genetic mutations, epigenetic alterations (such as DNA methylation), and gene expression patterns indicating the presence of cancer. Genetic and epigenetic information includes a variety of epigenetic markers or functional elements, such as TFB, CTCF binding sites, genetic variations such as CNV, SNV, insertions / deletions, fusions, mRNA expression, fragment omics patterns, fragment omics levels, fragment endpoint density, and histone acetylation or methylation markers associated with the following: balanced enhancers including H3K4me1, H3K27ac, H3K27me3; promoter regions including H3K4me3, H3 / H4ac, H3K4me1, H3K27me3, H3K9me3 and / or H3.3; and open chromatin including H3Ac and H4Ac, H3K4me1, H3K4me2, H3K4me3, H2BK120ub, H3.3, and H3S10ph.
[0009] In other embodiments, cancer types include breast cancer, colorectal cancer, lung cancer, bladder cancer, pancreatic cancer, ovarian cancer, liver cancer, stomach cancer, esophageal cancer, kidney cancer, melanoma, gallbladder cancer, or uterine cancer. In other embodiments, cancer types also include the tissue in which the cancer originates in the patient. In other embodiments, cancer types also include tissue in which cancer of unknown primary origin (CUP) originates in the patient. In other embodiments, the classification is based on a set of cancer-specific models.
[0010] This document describes a method, a computer-implemented method, comprising receiving one or more datasets in a computer system including one or more hardware processors and one or more computer-readable storage media, wherein the datasets include molecular phenotypes obtained from test samples from a patient, and wherein the computer-readable media includes instructions that, when executed by the processor, cause one or more hardware processors to perform one or more classifications of the test samples, the classifications including a first classification and a subsequent second classification, the first classification including a cancer condition, the second classification including a cancer type, wherein the second classification is performed only if the first classification includes a predetermined category, wherein the first and / or second classifications are based on a set of cancer-specific models, wherein the molecular phenotypes include epigenetic data, and wherein the cancer type also includes the tissue from which the cancer originates in the patient.
[0011] This document describes a system capable of performing any of the foregoing embodiments. This document also describes a computer-readable medium containing instructions enabling any of the foregoing embodiments. Brief description of the attached diagram
[0012] Figure 1 Large-scale epigenomic assays for pan-cancer analysis. Detection loss can be estimated using maxMAF or similar epiMAF and genomic maxMAF from “driver” genes defined in the TVF. Group performance across cancer types confirms the ability to identify tissue of origin (TOO).
[0013] Figure 2 Example of a sample-level methylation prediction workflow.
[0014] Figure 3 A multi-cancer classification workflow, including origin cancer signals and origin tissue multi-class classifiers described as binary classifiers.
[0015] Figure 4 Identification is performed using a confusion matrix. Detailed description
[0016] While various embodiments of this disclosure have been shown and described herein, those skilled in the art will understand that such embodiments are provided by way of example only. Many variations, modifications, and substitutions will occur to those skilled in the art without departing from this disclosure. It should be understood that various alternatives may be adopted to the embodiments of this disclosure described herein.
[0017] The term “about” and its grammatical equivalents associated with a reference value can include a range of values that are plus or minus 10% of that value. For example, the quantity “about 10” can include quantities from 9 to 11. The term “about” associated with a reference value can include a range of values that are plus or minus 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1%.
[0018] The term “at least” and its grammatical equivalents associated with a reference value can include the reference value and values greater than that value. For example, a quantity “at least 10” can include the value 10 and any value higher than 10, such as 11, 100, and 1,000.
[0019] The term “at most” and its syntactic equivalents associated with a reference value can include the reference value and be less than that value. For example, a quantity “at most 10” can include any value of 10 and below 10, such as 9, 8, 5, 1, 0.5, and 0.1.
[0020] Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” as used herein can include plural referents. Thus, for example, reference to “cell” can include more than one such cell, and reference to “culture” can include reference to one or more cultures and their equivalents known to those skilled in the art, and so on. Unless otherwise clearly indicated, all technical and scientific terms used herein may have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0021] Cancer can be indicated by epigenetic variations such as methylation. Examples of methylation changes in cancer include localized increases in DNA methylation at CpG islands at transcription start sites (TSS) of genes involved in normal growth control, DNA repair, cell cycle regulation, and / or cell differentiation. This hypermethylation may be associated with an abnormal loss of transcriptional capacity of the genes involved and occurs at least as frequently as point mutations and deletions that cause altered gene expression. DNA methylation profiling can be used to detect regions in the genome with different levels of methylation (“differentially methylated regions” or “DMRs”) that have changed during development or been perturbed by disease (e.g., cancer or any cancer-related disease). The genome of cancer cells has an imbalance in the aforementioned DNA methylation patterns, and therefore an imbalance in the functional packaging of DNA. Thus, the combination of chromatin organization abnormalities with methylation changes, when analyzed together, may help enhance cancer profiling. Combining MBD partitioning with fragmentomics data (such as the start and end positions of fragment mappings (related to nucleosome position), fragment length, and associated nucleosome occupancy) can be used for chromatin structure analysis in hypermethylation studies to improve biomarker detection rates.
[0022] Methylation profiling can include identifying methylation patterns across different regions of the genome. For example, after partitioning and sequencing molecules based on their methylation levels (e.g., the relative number of methylation sites per molecule), the sequences of molecules in different partitions can be mapped to a reference genome. This can reveal regions of the genome that are more or less methylated compared to other regions. In this way, genomic regions can differ in their methylation levels in contrast to individual molecules.
[0023] The properties of nucleic acid molecules can be modified, which can include various chemical or protein modifications (i.e., epigenetic modifications). Non-limiting examples of chemical modifications may include, but are not limited to, covalent DNA modifications, including DNA methylation. In some embodiments, DNA methylation includes adding a methyl group to cytosine (the cytosine followed by guanine in a nucleic acid sequence) at a CpG site. In some embodiments, DNA methylation includes adding a methyl group to adenine, such as N6-methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the fifth carbon of the 6-carbon ring of cytosine). In some embodiments, 5-methylation includes adding a methyl group to the 5C position of cytosine to produce 5-methylcytosine (m5c). In some embodiments, methylation includes derivatives of m5c. Derivatives of m5c include, but are not limited to, 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5-carboxycytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the third carbon of the 6-carbon ring of cytosine). In some embodiments, 3C methylation involves adding a methyl group to the 3C position of cytosine to generate 3-methylcytosine (3mC). Other examples include N6-methyladenine or glycosylation. DNA methylation involves adding a methyl group to DNA (e.g., CpG) and can alter the expression of methylated DNA regions. Methylation can also occur at non-CpG sites; for example, methylation can occur at CpA, CpT, or CpC sites. DNA methylation can alter the activity of methylated DNA regions. For example, when DNA in a promoter region is methylated, gene transcription can be repressed. DNA methylation is essential for normal development, and abnormalities in methylation can disrupt epigenetic regulation. Disruptions in epigenetic regulation, such as repression, can lead to diseases such as cancer. DNA methylation of promoters may indicate cancer.
[0024] CpG dinucleotides are dinucleotides CpG (cytosine-phosphate-guanine) on the sense strand of a double-stranded DNA molecule, i.e., at the 5' end of the nucleic acid sequence. In the 3' direction, cytosine is followed by guanine and its complementary CpG on the antisense chain. CpG dyads can be fully methylated or hemimethylated (methylated on only one chain).
[0025] CpG dinucleotides are underrepresented in the normal human genome, where most CpG dinucleotide sequences are transcriptionally inert (e.g., near-centromere regions of chromosomes and heterochromatin regions of DNA in repetitive elements) and are methylated. However, many CpG islands are protected from such methylation, especially around transcription start sites (TSS).
[0026] Protein modifications include components that bind chromatin, particularly histones (including their modified forms), as well as components that bind other proteins, such as those involved in replication or transcription. This disclosure provides methods for processing and analyzing nucleic acids with varying degrees of modification, such that the nature of their original modifications is correlated with nucleic acid tags, which can be decoded during nucleic acid analysis by sequencing. Genetic variations in the nucleic acid modifications of a sample can then be correlated with the degree of modification (epigenetic variation) of that nucleic acid in the original sample, said nucleic acids including single-stranded (e.g., ssDNA or RNA) or double-stranded molecules (e.g., dsDNA).
[0027] DNA loss can reduce the presence of one or more types of DNA, making them difficult to detect (e.g., cfDNA). In one or more other cases, existing methods for measuring DNA methylation, such as enrichment or depletion methods, can have relatively high levels of resolution, such as from about 100 base pairs (bp) to about 200 bp, which can make it difficult to accurately determine the amount of DNA methylation. The accuracy of DNA methylation determination can affect the accuracy of tumor score estimation in a sample. Since tumor scores are used to determine whether a sample is derived from a subject with or without a tumor, the accuracy of tumor score estimation can influence individual diagnostic and / or treatment decisions.
[0028] Some aspects of this invention relate to diagnostic methods and systems for identifying the tissue of origin (TOO) or cancer origin signal (CSO) of cancer cells. Specifically, this invention relates to cancer origin signal (CSO) and / or tissue of origin (TOO) technologies that utilize genetic, epigenetic, proteomics, fragmentomics, and / or transcriptomic profiling and bioinformatics analyses to determine the origin of malignant cells in various cancer types.
[0029] Multiple Cancer Early Detection (MCED) tests can detect multiple cancer types with a single, minimally invasive test. These tests typically analyze biomarkers in blood (cfDNA to detect ctDNA) or other bodily fluids to identify the presence of cancer and provide information about the tissue of origin (TOO). MCED tests incorporate Cancer Origin Signal (CSO) and / or TOO detection technologies to enhance diagnostic capabilities by detecting cancer and accurately locating its origin. Further information can be found in Klein et al., Annals of Oncology (2021) and Liu et al., Annals of Oncology (2021), each of which is incorporated herein by reference in its entirety.
[0030] After detecting cancer signals, CSO and / or TOO technologies are implemented to precisely pinpoint its origin. The molecular (genetic and / or epigenetic) profile of the detected cancer is compared with a reference database containing tissue-specific molecular imprints. Bioinformatics algorithms incorporating machine learning, artificial intelligence, and deep learning methods are used to analyze the data to identify patterns matching the tissue-specific imprints. Typical steps include feature extraction, i.e., identifying key molecular features such as somatic variations and epigenetic variations (i.e., methylation patterns) based on sequencing data. Pattern matching, where these features are compared with the reference database to determine the most probable tissue of origin, and probability scores are assigned to each potential tissue of origin, with the highest score indicating the most likely primary site.
[0031] sample The sample can be any biological sample isolated from the subject. The sample can be a bodily sample. Samples can include body tissues such as known or suspected solid tumors, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells or leukocytes, endothelial cells, tissue biopsies, cerebrospinal fluid, synovial fluid, lymph, ascites, interstitial fluid or extracellular fluid, fluids in the intercellular spaces (including gingival crevicular fluid), bone marrow, pleural effusion, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. Samples are preferably bodily fluids, particularly blood and its fractions, and urine. Samples can be in the form initially isolated from the subject, or can be further processed to remove or add components, such as cells, or to enrich one component relative to other components. Therefore, preferred bodily fluids for analysis are plasma or serum containing cell-free nucleic acids. Samples can be isolated or obtained from the subject and transported to the sample analysis site. Samples can be stored and transported at desirable temperatures, such as room temperature, 4°C, -20°C, and / or -80°C. Samples may be isolated from or obtained from the subject at the sample analysis site. Subjects may be humans, mammals, animals, companion animals, service animals, or pets. Subjects may have cancer. Subjects may not have cancer or may have detectable symptoms of cancer. Subjects may have been treated with one or more cancer therapies, such as chemotherapy, antibodies, vaccines, or biologics. Subjects may be in remission. Subjects may be diagnosed or may not be diagnosed with a susceptibility to cancer or any cancer-related genetic mutations / disorders.
[0032] The volume of plasma can depend on the desired read depth for the sequenced region. Exemplary volumes are 0.4–40 ml, 5–20 ml, and 10–20 ml. For example, the volume can be 0.5 mL, 1 mL, 5 mL, 10 mL, 20 mL, 30 mL, or 40 mL. The volume of plasma sampled can be from 5 ml to 20 ml.
[0033] Samples can contain varying amounts of nucleic acids containing genomic equivalents. For example, a sample of approximately 30 ng DNA can contain approximately 10,000 (10^3) nucleotides. 4 The genome equivalent of 200 billion haploid human genomes, and in the case of cfDNA, it can contain approximately 200 billion (2 × 10⁻⁶) haploid human genomes. 11 ) individual polynucleotide molecules. Similarly, a sample of about 100 ng DNA can contain about 30,000 haploid human genome equivalents, and in the case of cfDNA, about 600 billion individual molecules.
[0034] The sample may contain nucleic acids from different sources, such as nucleic acids and cell-free nucleic acids from cells of the same subject, or nucleic acids and cell-free nucleic acids from cells of different subjects. The sample may contain nucleic acids carrying mutations. For example, the sample may contain DNA carrying germline mutations and / or somatic mutations. Germline mutations refer to mutations present in the germline DNA of the subject. Somatic mutations refer to mutations originating from the somatic cells of the subject, such as cancer cells. The sample may contain DNA carrying cancer-related mutations (e.g., cancer-related somatic mutations). The sample may contain epigenetic variations (i.e., chemical or protein modifications) that are associated with the presence of genetic variations (such as cancer-related mutations). In some embodiments, the sample contains epigenetic variations associated with the presence of genetic variations, wherein the sample does not contain said genetic variations.
[0035] Exemplary amounts of cell-free nucleic acid in the pre-amplification sample range from about 1 fg to about 1 µg, such as 1 pg to 200 ng, 1 ng to 100 ng, 10 ng to 1000 ng. For example, amounts can be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. Amounts can be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng of cell-free nucleic acid molecules. The quantity can be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or 200 ng of cell-free nucleic acid molecules. The method may include obtaining 1 femtogram (fg) to 200 ng.
[0036] Cell-free nucleic acids are nucleic acids that are not contained within cells or otherwise bound to cells, or in other words, nucleic acids that remain in a sample after the removal of intact cells. Cell-free nucleic acids include DNA, RNA, and their hybrids, including genomic DNA, mitochondrial DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), and / or fragments of any of these. Cell-free nucleic acids can be double-stranded, single-stranded, or hybrids thereof. Cell-free nucleic acids can be released into body fluids through secretion or cell death procedures such as cell necrosis and apoptosis. Some cell-free nucleic acids are released into body fluids from cancer cells, such as circulating tumor DNA (ctDNA). Others are released from healthy cells. In some embodiments, cfDNA is cell-free fetal DNA (cffDNA). In some embodiments, cell-free nucleic acids are produced by tumor cells. In some embodiments, cell-free nucleic acids are produced by a mixture of tumor cells and non-tumor cells.
[0037] Cell-free nucleic acids have an example size distribution of approximately 100-500 nucleotides, with molecules of 110 to 230 nucleotides representing approximately 90% of the molecules, a mode of approximately 168 nucleotides, and a second small peak in the range of 240 to 440 nucleotides. Cell-free nucleic acids can be separated from body fluids via a fractionation or partitioning step, in which the cell-free nucleic acids present in solution are separated from intact cells and other insoluble components in the body fluid. Partitioning can include techniques such as centrifugation or filtration. Optionally, cells in the body fluid can be lysed, and cell-free and cellular nucleic acids are processed together. Typically, nucleic acids can be precipitated with alcohol after adding buffer and washing steps. Further cleaning steps, such as silica-based columns, can be used to remove contaminants or salts. Nonspecific bulk carrier nucleic acids (such as Cot-1 DNA) or DNA or proteins for bisulfite sequencing, hybridization, and / or ligation can be added throughout the reaction to optimize certain aspects of the procedure, such as yield.
[0038] Following such processing, the sample can include various forms of nucleic acids, including double-stranded DNA, single-stranded DNA, and single-stranded RNA. In some embodiments, single-stranded DNA and RNA can be converted into double-stranded forms, and therefore included in subsequent processing and analytical steps.
[0039] Analytes Analytes may include nucleic acid analytes and non-nucleic acid analytes. This disclosure provides a method for detecting genetic variations in biological samples from a subject. Biological samples may include polynucleotides from cancer cells. Polynucleotides may be DNA (e.g., genomic DNA, cDNA), RNA (e.g., mRNA, small RNA), or any combination thereof. Biological samples may include, for example, tumor tissue from a biopsy. In some cases, biological samples may include blood or saliva. In specific cases, biological samples may contain cell-free DNA (“cfDNA”) or circulating tumor DNA (“ctDNA”). Cell-free DNA may be present, for example, in blood.
[0040] Examples of non-nucleic acid analytes include, but are not limited to, lipids, carbohydrates, peptides, proteins, glycoproteins (N-linked or O-linked), lipoproteins, phosphoproteins, specific phosphorylated or acetylated variants of proteins, amidated variants of proteins, hydroxylated variants of proteins, methylated variants of proteins, ubiquitinated variants of proteins, sulfated variants of proteins, viral proteins (e.g., viral capsids, viral envelopes, viral outer shells, viral appendages, viral glycoproteins, viral spikes, etc.), extracellular and intracellular proteins, antibodies, and antigen-binding fragments. This also includes receptors, antigens, surface proteins, transmembrane proteins, differentiation protein clusters, protein channels, protein pumps, carrier proteins, phospholipids, glycoproteins, glycolipids, cell-cell interaction protein complexes, antigen-presenting complexes, major histocompatibility complexes, engineered T-cell receptors, T-cell receptors, B-cell receptors, chimeric antigen receptors, extracellular matrix proteins, post-translational modifications (e.g., phosphorylation, glycosylation, ubiquitination, nitrosation, methylation, acetylation, or lipidation) of cell surface proteins, gap junctions, and adhesion junctions.
[0041] Typically, systems, apparatus, methods, and compositions can be used to analyze any number of analytes, further including both nucleic acid analytes and non-nucleic acid analytes. For example, the number of analytes being analyzed can be at least about 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 100, at least 1,000, at least 10,000, at least 100,000, or more different analytes present in a region of the sample or a single feature of the substrate.
[0042] One or more nucleic acid analytes and / or non-nucleic acid analytes constitute a set of molecular interactions in the biological system under study (e.g., a cell), which can be considered as an “interaction set”—molecular interactions occurring between molecules belonging to different biochemical families (proteins, nucleic acids, lipids, carbohydrates, etc.) and also within a given family. In various embodiments, the interaction set is a protein-DNA interaction set (a network formed by transcription factors (and DNA or chromatin regulatory proteins) and their target genes). In other embodiments, the interaction set refers to a protein-protein interaction network (PPI) or a protein-protein interaction network (PIN). The methods described herein allow for the study and analysis of interaction sets. Techniques such as proteogenomics (whole genome sequencing, whole exome sequencing, and RNA-seq, as well as mass spectrometry, as examples) can support the study of interaction sets.
[0043] analyze The methods of this invention can be used to diagnose the presence of a condition, particularly cancer, in a subject, to characterize the condition (e.g., to stage the cancer or determine its heterogeneity), monitor the condition's response to treatment, and achieve prognostic assessment of the risk of condition progression or subsequent disease development. This disclosure can also be used to determine the efficacy of a particular treatment option. If treatment is successful, a successful treatment option may increase the amount of copy number variations or rare mutations detected in the subject's blood as more cancer cells may die and shed DNA. In other instances, this may not occur. In yet another instance, some treatment options may be correlated with the genetic profile of the cancer over time. This correlation can be used to select a therapy. Additionally, if cancer is observed to be in remission after treatment, the methods of this invention can be used to monitor residual disease or disease recurrence.
[0044] The types and numbers of cancers that can be detected include leukemia, brain cancer, lung cancer, skin cancer, nasal cancer, laryngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, colorectal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, solid tumors, heterogeneous tumors, and homogeneous tumors. Cancer type and / or stage can be detected based on genetic variations including: mutations, rare mutations, insertions / deletions, copy number variations, transversions, translocations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structural alterations, gene fusions, gene truncation, gene amplification, gene duplication, chromosomal damage, DNA damage, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in 5-methylcytosine.
[0045] Genetic and other analyte data can also be used to characterize specific forms of cancer. Cancers are often heterogeneous in both composition and stage. Genetic profiling data can allow for the characterization of specific subtypes of cancer, which may be important in the diagnosis or treatment of that specific subtype. This information can also provide subjects or practitioners with clues about the prognosis of a specific type of cancer and allow them to adjust treatment options based on disease progression. Some cancers can progress and become more aggressive and genetically unstable. Other cancers may remain benign, inactive, or dormant. The systems and methods disclosed herein can be used to determine disease progression.
[0046] The analyses of this invention can also be used to determine the efficacy of a particular treatment option. If the treatment is successful, a successful treatment option may increase the amount of copy number variations or rare mutations detected in the subject's blood as more cancer cells may die and shed DNA. In other instances, this may not occur. In yet another instance, some treatment options may be correlated with the genetic profile of the cancer over time. This correlation can be used to select a therapy. Additionally, if cancer is observed to be in remission after treatment, the method of this invention can be used to monitor residual disease or recurrence of disease.
[0047] The methods of this invention can also be used to detect genetic variations in conditions other than cancer. Following the onset of certain diseases, immune cells, such as B cells, can undergo rapid clonal expansion. Copy number variation detection can be used to monitor clonal expansion and certain immune states. In this instance, copy number variation analysis can be performed over time to generate a spectrum of how a particular disease may progress. Copy number variation or even rare mutation detection can be used to determine how a pathogen population changes during the course of infection. This can be particularly important during chronic infections (such as HIV / AID or hepatitis infections), where the virus can alter its life cycle state and / or mutate into a more virulent form during the course of infection. When immune cells attempt to destroy transplanted tissue, the methods of this invention can be used to determine or analyze the host body's rejection activity to monitor the state of the transplanted tissue and to modify the process of rejection treatment or prevention.
[0048] Furthermore, the methods of this disclosure can be used to characterize the heterogeneity of anomalies in a subject. Such methods may include, for example, generating a genetic profile of extracellular polynucleotides derived from the subject, wherein the genetic profile includes more than one set of data obtained from copy number variation and rare mutation analysis. In some embodiments, the anomaly is cancer. In some embodiments, the anomaly may be a condition leading to a heterogeneous genomic population. In the example of cancer, some tumors are known to contain tumor cells at different stages of cancer. In other examples, heterogeneity may include multiple lesions of the disease. Again, in the example of cancer, multiple tumor lesions may be present, perhaps one or more of which are the result of metastases that have spread from the primary site.
[0049] The method of this invention can be used to generate or analyze a fingerprint or dataset that represents the sum of genetic information derived from different cells in heterogeneous diseases. This dataset can include individual or combined copy number variation and mutation analyses.
[0050] The methods of this invention can be used for the diagnosis, prognosis, monitoring, or observation of cancer or other diseases. In some embodiments, the methods herein do not involve the diagnosis, prognosis, or monitoring of the fetus, and therefore do not involve noninvasive prenatal testing. In other embodiments, these methods can be used in pregnant subjects to diagnose, prognose, monitor, or observe cancer or other diseases in unborn subjects whose DNA and other polynucleotides may co-circulate with maternal molecules.
[0051] Determination of the 5-methylcytosine pattern of nucleic acids Bisulfite-based sequencing and its variations provide a means of determining the methylation patterns of nucleic acids. In some embodiments, determining the methylation pattern includes distinguishing between 5-methylcytosine (5mC) and unmethylated cytosine. In some embodiments, determining the methylation pattern includes distinguishing between N6-methyladenine and unmethylated adenine. In some embodiments, determining the methylation pattern includes distinguishing between 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), and 5-carboxycytosine (5caC) and unmethylated cytosine. Examples of bisulfite sequencing include, but are not limited to, oxidized bisulfite sequencing (OX-BS-seq), Tet-assisted bisulfite sequencing (TAB-seq), and reduced bisulfite sequencing (redBS-seq).
[0052] Oxidized bisulfite sequencing (OX-BS-seq) is used to distinguish between 5mC and 5hmC by first converting 5hmC to 5fC, followed by bisulfite sequencing as described above. Tet-assisted bisulfite sequencing (TAB-seq) can also be used to distinguish between 5mC and 5hmC. In TAB-seq, 5hmC is protected by glycosylation. As mentioned earlier, 5mC is converted to 5caC using the Tet enzyme before bisulfite sequencing. Reduced bisulfite sequencing is used to distinguish 5fC from modified cytosine.
[0053] Typically, in bisulfite sequencing, nucleic acid samples are split into two aliquots, and one aliquot is treated with bisulfite. Bisulfite converts native cytosine and certain modified cytosine nucleotides (e.g., 5-formylcytosine or 5-carboxycytosine) to uracil, while other modified cytosines (e.g., 5-methylcytosine, 5-hydroxymethylcytosine) are not converted. Comparison of the nucleic acid sequences of molecules from the two aliquots indicates which cytosines were converted to uracil and which were not. Therefore, modified and unmodified cytosines can be identified. Initially splitting the sample into two aliquots is disadvantageous for samples containing only small amounts of nucleic acids and / or including heterogeneous cellular / tissue-derived samples such as body fluids containing cell-free DNA.
[0054] This disclosure provides methods for enabling bisulfite sequencing and its variations. These methods function by linking nucleic acids in a population to a capture motif (i.e., a tag that can be captured or immobilized). Capture motifs include, but are not limited to, biotin, avidin, streptavidin, nucleic acids containing a specific nucleotide sequence, haptens recognized by antibodies, and magnetically attractive particles. Extraction motifs can be members of binding pairs such as biotin / streptavidin or haptens / antibodies. In some embodiments, the capture motif attached to the analyte is captured by its binding pair, which is attached to a separable motif, such as magnetically attractive particles or large particles that can be precipitated by centrifugation. The capture motif can be any type of molecule that allows affinity separation of nucleic acids with the capture motif from nucleic acids lacking the capture motif. Example capture motifs are biotin or oligonucleotides, where biotin allows affinity separation by binding to streptavidin linked to or capable of being linked to a solid phase, and oligonucleotides allow affinity separation by binding to complementary oligonucleotides linked to or capable of being linked to a solid phase. After the capture motif is linked to the sample nucleic acid, the sample nucleic acid is used as an amplification template. After amplification, the original template remains connected to the captured portion, but the amplicon does not connect to the captured portion.
[0055] The capture portion can be attached to the sample nucleic acid as a component of an adaptor, which can also provide binding sites for amplification and / or sequencing primers. In some methods, the sample nucleic acid is attached to an adaptor at both ends, with both adaptors containing the capture portion. Preferably, any cytosine residues in the adaptor are modified, such as by 5-methylcytosine, to protect against bisulfite. In some cases, the capture portion is attached via a cleavable adapter (e.g., photocleavable dethiobiotin-TEG or uracil residues cleavable by the USER™ enzyme, Chem. Commun. (Camb. 2015 Feb21; 51(15): 3266-3269), in which case the capture portion can be removed if desired.
[0056] The amplicon is denatured and then contacted with an affinity reagent used to capture the tag. The original template binds to the affinity reagent, while the amplified nucleic acid molecules do not. Therefore, the original template can be separated from the amplified nucleic acid molecules.
[0057] After isolation or partitioning, the corresponding populations of nucleic acids (i.e., the original template and amplification products) can be subjected to bisulfite treatment, with the original template population receiving bisulfite treatment while the amplification products do not. Optionally, the amplification products can undergo bisulfite treatment while the original template population does not. After such treatment, the corresponding populations can be amplified (in the case of the original template population, this converts uracil to thymine). The populations can also undergo biotinylated probe hybridization for enrichment. The corresponding populations are then analyzed and sequences are compared to determine which cytosines are 5-methylated (or 5-hydroxymethylated) in the original sample. Detection of T nucleotides (corresponding to unmethylated cytosine converted to uracil) in the template population and C nucleotides at corresponding positions in the amplification population indicates unmodified C. The presence of C at corresponding positions in the original template and amplification population indicates the presence of modified C in the original sample.
[0058] In some embodiments, the method uses sequential DNA-seq and bisulfite-seq (BIS-seq) NGS library preparation with molecularly tagged DNA libraries. The process involves tagging of an adaptor (e.g., biotin), DNA-seq amplification of the entire library, parental molecule recovery (e.g., streptavidin bead pull-down), bisulfite conversion, and BIS-seq. In some embodiments, the method identifies 5-methylcytosine at single-base resolution through preparative sequential NGS amplification of parental library molecules with and without bisulfite treatment. This can be achieved by modifying the 5-methylated NGS adaptor used in BIS-seq (oriented adaptor; Y-shaped / forked, replaced with 5-methylcytosine) with a marker (e.g., biotin) on one of the two adaptor strands. Sample DNA molecules are linked adaptors and are amplified (e.g., by PCR). Since only parental molecules will have tagged adaptor ends, they can be selectively recovered from their amplified progeny using tag-specific capture methods (e.g., streptavidin magnetic beads). Because the parental molecules retain the 5-methylation marker, bisulfite conversion on the captured library will produce a 5-methylation state at single-base resolution during BIS-seq, thus preserving molecular information in the corresponding DNA-seq. In some embodiments, the bisulfite-treated library can be combined with the untreated library prior to enrichment / NGS by adding a sample-tagged DNA sequence in a standard multiplex NGS workflow. As with the BIS-seq workflow, bioinformatics analysis can be performed for genome alignment and 5-methylation base recognition. In summary, this method provides the ability to selectively recover parental, ligated molecules carrying the 5-methylcytosine marker after library amplification, allowing for parallel processing of bisulfite-converted DNA. This overcomes the detrimental nature of bisulfite treatment to the quality / sensitivity of DNA-seq information extracted from the workflow. With this method, the recovered ligated, parental DNA molecules (via tagged adaptors) allow for the amplification of a complete DNA library and the parallel application of treatments that induce epigenetic DNA modifications. This disclosure discusses the use of BIS-seq methods to identify 5-methylated cytosine (5-methylcytosine), but this should not be limiting. Variations of BIS-seq have been developed to identify hydroxymethylated cytosine (5hmC; OX-BS-seq, TAB-seq), formylcytosine (5fC; redBS-seq), and carboxycytosine. These methods can be implemented using sequential / parallel library preparation as described herein.
[0059] Alternative methods for analyzing modified nucleic acids This disclosure provides alternative methods for analyzing modified nucleic acids (e.g., methylated, histone-linked, and other modifications discussed above). In some such methods, a population of nucleic acids with varying degrees of modification (e.g., each nucleic acid molecule has 0, 1, 2, 3, 4, 5, or more methyl groups) is contacted with an adaptor, and the population is then stratified according to the degree of modification. The adaptor is attached to one or both ends of the nucleic acid molecules in the population. Preferably, the adaptor contains a sufficient number of different tags such that the number of tag combinations results in a high probability, e.g., 95%, 99%, or 99.9%, that two nucleic acids with the same start and end points receive different tag combinations. After attaching the adaptor, the nucleic acid is amplified from a primer that binds to a primer binding site within the adaptor. Adaptors with the same or different tags may contain the same or different primer binding sites, but preferably the adaptor contains the same primer binding site. After amplification, the nucleic acid is contacted with an agent that preferably binds the modified nucleic acid (such as those previously described). The nucleic acid is divided into at least two partitions, the difference between the at least two partitions being the degree to which the modified nucleic acid binds to the agent. For example, if the reagent has an affinity for the modified nucleic acid, the overrepresented modified nucleic acid (compared to the median representation in the population) preferentially binds to the reagent, while the underrepresented modified nucleic acid does not bind to the reagent or is more easily eluted from it. After separation, the different partitions can then undergo additional processing steps, which typically include parallel but separate additional amplification and sequence analysis. The sequence data from the different partitions can then be compared.
[0060] Nucleic acid molecules can be coupled to Y-adaptors containing primer binding sites and tags. The molecule is then amplified. The amplified molecule is then partitioned by contacting an antibody that preferentially binds to 5-methylcytosine to produce two partitions. One partition contains the unmethylated original molecule and the amplified copy that has lost methylation. The other partition contains the original DNA molecule with methylation. The two partitions are then processed and sequenced separately, and the methylated partition is further amplified. The sequence data of the two partitions can then be compared. In this example, the tag is not used to distinguish between methylated and unmethylated DNA, but rather to distinguish the different molecules within these partitions, allowing one to determine whether reads with the same start and end points are based on the same or different molecules.
[0061] This disclosure also provides methods for analyzing nucleic acid populations, wherein at least some nucleic acids contain one or more modified cytosine residues, such as 5-methylcytosine and any other modifications previously described. In these methods, the nucleic acid population is contacted with an adaptor comprising one or more cytosine residues, such as 5-methylcytosine, modified at the 5C position. Preferably, all cytosine residues in such an adaptor are also modified, or all such cytosine residues in the primer-binding region of the adaptor are modified. The adaptor is attached to both ends of the nucleic acid molecules in the population. Preferably, the adaptor contains a sufficient number of different tags such that the number of tag combinations results in a high probability, for example, 95%, 99%, or 99.9%, that two nucleic acids with the same start and end points receive different tag combinations. The primer-binding sites in such an adaptor may be the same or different, but are preferably the same. After the adaptor is attached, the nucleic acid is amplified by primers that bind to the primer-binding sites of the adaptor. The amplified nucleic acid is divided into a first aliquot and a second aliquot. The first aliquot is sequenced, with or without further processing. This determines the sequence data of molecules in the first aliquot regardless of the initial methylation state of the nucleic acid molecules. Nucleic acid molecules in the second aliquot are treated with bisulfite. This treatment converts unmodified cytosine to uracil. The bisulfite-treated nucleic acids then undergo amplification, initiated by primers targeting the original primer binding sites of the adaptor attached to the nucleic acid. Now only the nucleic acid molecules initially attached to the adaptor (unlike their amplification products) are amplified because these nucleic acids retain cytosine at the primer binding sites of the adaptor, while the amplification products have lost the methylation of these cytosine residues, which were converted to uracil during the bisulfite treatment. Therefore, only the original molecules in the population (at least some of which are methylated) undergo amplification. After amplification, these nucleic acids are sequenced. Comparison of the sequences determined from the first and second aliquots can particularly indicate which cytosine residues in the nucleic acid population have undergone methylation.
[0062] Divide the sample into more than one subsample; aspects of the sample; epigenetic properties. Analysis of Characteristics In some embodiments described herein, different forms of nucleic acid populations (e.g., hypermethylated and hypomethylated DNA in a sample, such as a capture set of cfDNA as described herein) can be physically partitioned based on one or more properties of the nucleic acids, followed by further analysis, such as differential modification or isolation of nucleotides, tagging, and / or sequencing. This approach can be used to determine, for example, whether certain sequences are hypermethylated or hypomethylated. In some embodiments, analysis is performed on hypermethylated variable epigenetic target regions to determine whether they exhibit the hypermethylation properties of tumor cells, and / or analysis is performed on hypomethylated variable epigenetic target regions to determine whether they exhibit the hypomethylation properties of tumor cells. Additionally, by partitioning heterogeneous nucleic acid populations, one can increase rare signals, for example, by enriching rare nucleic acid molecules that are more prevalent in one fraction (or partition) of the population. For example, by partitioning a sample into hypermethylated and hypomethylated nucleic acid molecules, it is easier to detect genetic variations that are present in hypermethylated DNA but less so (or absent) in hypomethylated DNA. By analyzing more than one fraction of a sample, multidimensional analysis of individual loci or nucleic acid species in the genome can be performed, thus enabling greater sensitivity.
[0063] In some cases, heterogeneous nucleic acid samples are partitioned into two or more partitions (e.g., at least three, four, five, six, or seven partitions). In some implementations, each partition is differentially tagged. The tagged partitions can then be pooled together for collective sample preparation and / or sequencing. The partitioning-tag-pooling step can occur more than once, with each round of partitioning occurring based on a different attribute (as exemplified in this paper) and tagged using differential tags that distinguish them from other partitioning and partitioning methods.
[0064] Examples of attributes that can be used for partitioning include sequence length, methylation level, nucleosome binding, sequence mismatch, immunoprecipitation, and / or proteins that bind to DNA. The resulting partitions can include one or more of the following nucleic acid forms: single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), shorter DNA fragments, and longer DNA fragments. In some embodiments, partitioning is typically performed based on cytosine modification (e.g., cytosine methylation) or methylation, and optionally combined with at least one additional partitioning step, which can be based on any of the aforementioned attributes or forms of DNA. In some embodiments, a heterogeneous population of nucleic acids is partitioned into nucleic acids having one or more epigenetic modifications and nucleic acids not having said one or more epigenetic modifications. Examples of epigenetic modifications include the presence or absence of methylation, methylation level, methylation type (e.g., 5-methylcytosine with other types of methylation, such as adenine methylation and / or cytosine hydroxymethylation), and association with one or more proteins (such as histones) and the level of association. Optionally or additionally, the heterogeneous nucleic acid population can be partitioned into nucleosome-associated nucleic acid molecules and nucleosome-free nucleic acid molecules. Optionally or additionally, the heterogeneous nucleic acid population can be partitioned into single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA). Optionally or additionally, the heterogeneous nucleic acid population can be partitioned based on nucleic acid length (e.g., molecules with a maximum length of 160 bp and molecules with a length greater than 160 bp).
[0065] In some cases, each partition (representing a different nucleic acid form) is differentially labeled, and the partitions are pooled together and then sequenced. In other cases, the different forms are sequenced separately. In some implementations, different nucleic acid populations are partitioned into two or more distinct partitions. Each partition represents a different nucleic acid form, and the first partition (also called a subsample) contains DNA with a larger proportion of cytosine modifications than the second subsample. Each partition is tagged differently. The first subsample undergoes a procedure that differently affects the first and second nucleotides in the DNA of the first subsample, where the first nucleotide is modified or unmodified, and the second nucleotide is modified or unmodified, different from the first nucleotide, and the first and second nucleotides have the same base-pairing specificity. The tagged nucleic acids are pooled together and then sequenced. Sequence reads are obtained and analyzed, including computer simulation (in silico) to distinguish the first and second nucleotides in the DNA of the first subsample. Tags are used to sort reads from different partitions. Analysis can be performed at the level of individual partitions and at the level of the entire nucleic acid population to detect genetic variation. For example, the analysis may include computer simulation analysis to determine genetic variations, such as CNVs, SNVs, insertions / deletions, and fusions in the nucleic acids of each partition. In some cases, computer simulation analysis may include determining chromatin structure. For example, the coverage of sequence reads can be used to determine the location of nucleosomes in chromatin. Higher coverage may be associated with higher nucleosome occupancy in a genomic region, while lower coverage may be associated with lower nucleosome occupancy or nucleosome depleted regions (NDRs).
[0066] Samples may include nucleic acids with various modifications, including post-replication modifications of nucleotides and binding to one or more proteins (typically non-covalent).
[0067] In the implementation scheme, the nucleic acid population is a population of nucleic acids obtained from serum, plasma, or blood samples of subjects suspected of having vegetations, tumors, or cancer, or previously diagnosed with vegetations, tumors, or cancer. The nucleic acid population includes nucleic acids with different levels of methylation. Methylation can occur by any one or more post-replication or post-transcriptional modifications. Post-replication modifications include modifications to nucleotide cytosine, particularly at the 5-position of the nucleobase, such as 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, and 5-carboxycytosine. The affinity agent can be an antibody with desired specificity, a natural binding partner or a variant thereof (Bock et al., Nat Biotech 28: 1106-1114 (2010); Song et al., Nat Biotech 29: 68-72 (2011)), or, for example, an artificial peptide selected for specificity to a given target via phage display.
[0068] Examples of capture fractions envisioned herein include methyl-binding domains (MBDs) and methyl-binding proteins (MBPs) as described herein, including proteins such as MeCP2 and antibodies that preferentially bind to 5-methylcytosine. Similarly, partitioning of different forms of nucleic acids can be performed using histone-binding proteins that can separate histone-bound nucleic acids from free or unbound nucleic acids. Examples of histone-binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48, and SANT domain peptides. For some affinity agents and modifications, although binding to the agent may occur substantially all-or-nothing depending on whether the nucleic acid is modified, separation may be to a certain extent. In such cases, nucleic acids overrepresented in a modification bind to the agent to a greater extent than nucleic acids underrepresented in the modification. Optionally, modified nucleic acids may bind in an all-or-nothing manner. However, various levels of modification can then be eluted sequentially from the binding agent.
[0069] For example, in some implementations, partitioning can be binary or based on the degree / level of modification. For instance, a methyl-binding domain protein (e.g., the MethylMiner methylated DNA enrichment kit (ThermoFisherScientific)) can be used to partition all methylated fragments with unmethylated fragments. Subsequently, additional partitioning can include eluting fragments with different methylation levels by adjusting the salt concentration of the solution containing the methyl-binding domain and the binding fragment. As the salt concentration increases, fragments with higher methylation levels are eluted. In some cases, the final partitioning represents nucleic acids with different degrees of modification (overrepresentation or underrepresentation). Overrepresentation and underrepresentation can be defined by the number of modifications a nucleic acid carries relative to the median number of modifications per strand in the population. For example, if the median number of 5-methylcytosine residues in the nucleic acids in a sample is 2, then nucleic acids containing more than two 5-methylcytosine residues are overrepresented, while nucleic acids with one or zero 5-methylcytosine residues are underrepresented. The purpose of affinity separation is to enrich over-represented nucleic acids in the binding phase and under-represented nucleic acids in the non-binding phase (i.e., in solution). The nucleic acids in the binding phase can be eluted before subsequent processing.
[0070] When using the MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific), sequential elution can be used to separate methylated regions with different levels of methylation. For example, low-methylated regions (e.g., unmethylated) can be separated from methylated regions by contacting a group of nucleic acids with MBDs attached to magnetic beads from the kit. The beads are used to isolate methylated nucleic acids from unmethylated nucleic acids. Subsequently, one or more elution steps are performed sequentially to elute nucleic acids with different methylation levels. For example, the first group of methylated nucleic acids can be eluted at a salt concentration of 160 mM or higher, such as at least 150 mM, at least 200 mM, at least 300 mM, at least 400 mM, at least 500 mM, at least 600 mM, at least 700 mM, at least 800 mM, at least 900 mM, at least 1000 mM, or at least 2000 mM. After such methylated nucleic acids have been eluted, magnetic separation is again used to separate nucleic acids with higher levels of methylation from those with lower levels of methylation. The elution and magnetic separation steps can be repeated to produce various partitions, such as hypomethylated partitions (representing no methylation), methylated partitions (representing low methylation levels), and hypermethylated partitions (representing high methylation levels).
[0071] In some methods, nucleic acids bound to an affinity separator undergo a washing step. This washing step removes nucleic acids that are weakly bound to the affinity agent. Such nucleic acids can be enriched with a degree of modification close to the mean or median (i.e., an intermediate value between nucleic acids that remain bound to the solid and those that do not upon initial contact with the agent). Affinity separation results in at least two, and sometimes three or more, partitions of nucleic acids with different degrees of modification. While the partitions remain separate, nucleic acids from at least one partition, and usually two or three (or more) partitions, are attached to nucleic acid tags, which are typically provided as part of an adaptor, and the nucleic acids in different partitions receive different tags that distinguish members of one partition from members of another. Tags attached to nucleic acid molecules in the same partition can be the same or different from each other. However, if they are different, the tags can have a shared encoding to identify the molecules to which they are attached as belonging to a particular partition. For more details on partitioning nucleic acid samples based on properties such as methylation, see WO2018 / 119452, which is incorporated herein by reference. In some implementations, nucleic acid molecules can be hierarchically separated into different partitions based on whether they bind to a specific protein or a fragment thereof and whether they do not bind to that specific protein or a fragment thereof.
[0072] Nucleic acid molecules can be fractionated based on DNA-protein binding. Protein-DNA complexes can be fractionated based on specific properties of the protein. Examples of such properties include various epitopes, modifications (e.g., histone methylation or acetylation), or enzymatic activities. Examples of proteins that can bind DNA and be used as the basis for fractionation can include, but are not limited to, protein A and protein G. Any suitable method can be used for fractionation of nucleic acid molecules based on protein-binding regions. Examples of methods for fractionation of nucleic acid molecules based on protein-binding regions include, but are not limited to, SDS-PAGE, chromatin immunoprecipitation (ChIP), heparin chromatography, and asymmetric field flow fractionation (AF4).
[0073] In some implementations, nucleic acid fractionation is performed by contacting the nucleic acid with the methylation-binding domain (“MBD”) of a methylation-binding protein (“MBP”). The MBD binds to 5-methylcytosine (5mC). The MBD is coupled to paramagnetic beads (such as Dynabeads® M-280 streptavidin) via a biotin linker. Fractionation into fractions with different degrees of methylation can be performed by eluting the fractions with increasing NaCl concentrations.
[0074] An exemplary method for molecular tag identification of MBD bead partition libraries using NGS is as follows: Extracted DNA samples (e.g., plasma DNA extracted from human samples) are physically partitioned using a methyl-binding domain protein-bead purification kit, retaining all eluents from the process for downstream processing.
[0075] Differential molecular tags and NGS-feasible linker sequences were applied in parallel to each partition. For example, hypermethylated, residual methylated ('washed'), and hypomethylated partitions were linked to NGS linkers with molecular tags.
[0076] All molecularly tagged partitions were reassembled and subsequently amplified using adaptor-specific DNA primer sequences.
[0077] Enrich / hybridize the recombined and amplified total library to target genomic regions of interest (e.g., cancer-specific genetic variations and differentially methylated regions).
[0078] The enriched total DNA library was re-amplified and tagged with samples. Different samples were pooled and subjected to multiplex assays on an NGS instrument.
[0079] Bioinformatics analysis of NGS data was performed, using molecular tags to identify unique molecules and deconvolving samples into molecules with differentially expressed molecular dividers (MBDs). This analysis can simultaneously produce information on the relative 5-methylcytosine content of genomic regions with standard gene sequencing / variation detection.
[0080] Examples of MBPs envisioned in this article include, but are not limited to: (a) MeCP2, a protein that preferentially binds to 5-methyl-cytosine compared to binding to unmodified cytosine.
[0081] (b) RPL26, PRP8 and DNA mismatch repair protein MHS6 preferentially bind to 5-hydroxymethyl-cytosine compared to binding to unmodified cytosine.
[0082] (c) FOXK1, FOXK2, FOXP1, FOXP4 and FOXI3 preferentially bind 5-formyl-cytosine compared to binding unmodified cytosine (Iurlaro et al., Genome Biol. 14: R119 (2013)).
[0083] (d) Antibodies specific to one or more methylated nucleotide bases.
[0084] Typically, elution varies with the number of methylation sites per molecule, with molecules having more methylation eluting at increasing salt concentrations. To elute DNA into different populations based on the degree of methylation, a series of elution buffers with increasing NaCl concentrations can be used. Salt concentrations can range from about 100 nM to about 2500 mM NaCl. In one embodiment, the process produces three (3) partitions. Molecules are contacted with a solution of a first salt concentration, and this solution contains molecules containing methyl-binding domains that can attach to capture moieties such as streptavidin. At the first salt concentration, one population of molecules will bind to MBD, and one population will remain unbound. The unbound population can be separated into a “hypomethylated” population. For example, the first partition representing hypomethylated DNA is the partition that remains unbound at low salt concentrations (e.g., 100 mM or 160 mM). The second partition representing moderately methylated DNA is eluted using a moderate salt concentration (e.g., between 100 mM and 2000 mM). This is also separated from the sample. The third partition, representing the highly methylated form of DNA, was eluted with a high salt concentration (e.g., at least about 2000 mM).
[0085] This disclosure also provides methods for analyzing nucleic acid populations, wherein at least some nucleic acids contain one or more modified cytosine residues, such as 5-methylcytosine and any other modifications previously described. In these methods, after partitioning, a sample of nucleic acid subsamples is contacted with an adaptor containing one or more cytosine residues modified at the 5C position (such as 5-methylcytosine). Preferably, all cytosine residues in such an adaptor are also modified, or all such cytosine residues in the primer-binding region of the adaptor are modified. The adaptor is attached to both ends of nucleic acid molecules in the population. Preferably, the adaptor contains a sufficient number of different tags such that the number of tag combinations results in a high probability, for example, 95%, 99%, or 99.9%, that two nucleic acids with the same start and end points receive different tag combinations. The primer-binding sites in such an adaptor can be the same or different, but are preferably the same. After the adaptor is attached, the nucleic acid is amplified by primers that bind to the primer-binding sites of the adaptor. The amplified nucleic acid is divided into a first aliquot and a second aliquot. With or without further processing, the first aliquot is sequenced. This determines the sequence data of molecules in the first aliquot regardless of the initial methylation state of the nucleic acid molecules. Nucleic acid molecules in the second aliquot undergo a procedure that differently affects the first and second nucleobases in the DNA, where the first nucleobase includes cytosine modified at position 5, and the second nucleobase includes unmodified cytosine. This procedure can be bisulfite treatment or another procedure that converts unmodified cytosine to uracil. The nucleic acids that have undergone this procedure are then amplified using primers targeting the original primer binding sites of the adaptor attached to the nucleic acid. Now only the nucleic acid molecules initially attached to the adaptor (unlike their amplification products) are amplified because these nucleic acids retain cytosine at the primer binding sites of the adaptor, while the amplification products have lost the methylation of these cytosine residues, which have been converted to uracil during bisulfite treatment. Therefore, only the original molecules in the population (at least some of which are methylated) undergo amplification. After amplification, these nucleic acids are sequenced. Comparison of sequences determined from the first and second aliquots can particularly indicate which cytosines in the nucleic acid population have undergone methylation.
[0086] Such analysis can be performed using the following exemplary procedure. After partitioning, the two ends of the methylated DNA are ligated to a Y-shaped adaptor containing primer binding sites and a tag. The cytosine in the adaptor is modified at position 5 (e.g., 5-methylated). This modification of the adaptor serves to protect the primer binding sites in subsequent transformation steps (e.g., bisulfite treatment, TAP transformation, or any other transformation that does not affect the modified cytosine but affects the unmodified cytosine). After adaptor attachment, the DNA molecule is amplified. The amplified product is divided into two aliquots for sequencing with and without transformation. The untransformed aliquot may undergo sequence analysis with or without further treatment. The other aliquot undergoes a procedure that differently affects the first and second nucleotides in the DNA, where the first nucleotide includes the cytosine modified at position 5, and the second nucleotide includes the unmodified cytosine. This procedure may be bisulfite treatment or another procedure to convert the unmodified cytosine to uracil. When contacted with primers specific to the original primer binding site, only primer binding sites protected by cytosine modification can support amplification. Therefore, only the original molecule, not copies from the first amplification, undergoes further amplification. The further amplified molecules then undergo sequence analysis. The sequences from the two aliquots can then be compared. In the isolation scheme discussed above, the nucleic acid tag in the adaptor is not used to distinguish between methylated and unmethylated DNA, but rather to distinguish nucleic acid molecules within the same partition.
[0087] This causes the first subsample to undergo different processes, affecting the first nucleotide and the second nucleotide in the DNA of the first subsample. The program of nucleobases The method disclosed herein includes steps of subjecting a first subsample to procedures that differently affect a first nucleotide and a second nucleotide in the DNA of the first subsample, wherein the first nucleotide is a modified or unmodified nucleotide, the second nucleotide is a modified or unmodified nucleotide different from the first nucleotide, and the first and second nucleotides have the same base-pairing specificity. In some embodiments, if the first nucleotide is modified or unmodified adenine, then the second nucleotide is modified or unmodified adenine; if the first nucleotide is modified or unmodified cytosine, then the second nucleotide is modified or unmodified cytosine; if the first nucleotide is modified or unmodified guanine, then the second nucleotide is modified or unmodified guanine; if the first nucleotide is modified or unmodified thymine, then the second nucleotide is modified or unmodified thymine (wherein, for the purposes of this step, modified and unmodified uracil are included in modified thymine).
[0088] In some embodiments, the first nucleobase is a modified or unmodified cytosine, and then the second nucleobase is a modified or unmodified cytosine. For example, the first nucleobase may comprise unmodified cytosine (C), and the second nucleobase may comprise one or more of 5-methylcytosine (mC) and 5-hydroxymethylcytosine (hmC). Alternatively, the second nucleobase may comprise C, and the first nucleobase may comprise one or more of mC and hmC. Other combinations are also possible, such as those indicated, for example, in the overview above and the discussion below, such as where one of the first and second nucleobases comprises mC and the other comprises hmC.
[0089] In some implementations, the procedure that differently affects the first and second nucleobases in the DNA of the first subsample includes bisulfite conversion. Bisulfite treatment converts unmodified cytosine and certain modified cytosine nucleotides (e.g., 5-formylcytosine (fC) or 5-carboxycytosine (caC)) into uracil, while other modified cytosines (e.g., 5-methylcytosine and 5-hydroxymethylcytosine) are not converted. Therefore, in the case of bisulfite conversion, the first nucleobase comprises one or more of unmodified cytosine, 5-formylcytosine, 5-carboxycytosine, or other bisulfite-affected cytosine forms, and the second nucleobase may comprise one or more of mC and hmC, such as mC and optionally hmC. Sequencing of the bisulfite-treated DNA identifies the position read as a cytosine as the mC position or the hmC position. Simultaneously, positions read as T are identified as T or bisulfite-susceptible forms of C, such as unmodified cytosine, 5-formylcytosine, or 5-carboxycytosine. Therefore, bisulfite conversion of the first subsample as described herein facilitates the identification of positions containing mC or hmC using sequence reads obtained from the first subsample. For an exemplary description of bisulfite conversion, see, for example, Moss et al., Nat Commun. 2018; 9: 5068.
[0090] In some embodiments, the procedures that differently affect the first and second nucleotides of the DNA in the first subsample include oxidized bisulfite (Ox-BS) conversion. In some embodiments, the procedures that differently affect the first and second nucleotides of the DNA in the first subsample include Tet-assisted substituted borane reducing agent conversion, optionally wherein the substituted borane reducing agent is 2-methylpyridineborane, pyridineborane, tert-butylamineborane, or aminoborane. In some embodiments, the procedures that differently affect the first and second nucleotides of the DNA in the first subsample include chemically assisted substituted borane reducing agent conversion, optionally wherein the substituted borane reducing agent is 2-methylpyridineborane, pyridineborane, tert-butylamineborane, or aminoborane. In some implementations, procedures that differently affect the first and second nucleobases in the DNA of the first subsample include APOBEC-coupled epigenetic (ACE) transformation.
[0091] In some implementations, procedures that differently affect the first and second nucleotides in the DNA of the first subsample include enzymatic conversion of the first nucleotide, for example, as in EM-Seq. See, for example, Vaisvila R et al. (2019) Em-seq: Detection of DNA methylation at single base resolution from picograms of DNA. bioRxiv; DOI: 10.1101 / 2019.12.20.884692, available at www.biorxiv.org / content / 10.1101 / 2019.12.20.884692v1. For example, TET2 and T4-βGT can be used to convert 5mC and 5hmC into substrates that cannot be deaminated by deaminases (e.g., APOBEC3A), and then the deaminase (e.g., APOBEC3A) can be used to deaminate unmodified cytosine, converting it to uracil.
[0092] In some implementations, the procedure that differently affects the first nucleobase in the DNA of the first subsample and the second nucleobase in the DNA includes separating the DNA that initially contains the first nucleobase from the DNA that initially does not contain the first nucleobase.
[0093] In some embodiments, the first nucleobase is a modified or unmodified adenine, and the second nucleobase is a modified or unmodified adenine. In some embodiments, the modified adenine is N6-methyladenine (mA). In some embodiments, the modified adenine is one or more of N6-methyladenine (mA), N6-hydroxymethyladenine (hmA), or N6-formyladenine (fA).
[0094] Techniques including methylated DNA immunoprecipitation (MeDIP) can be used to separate DNA containing modified bases (such as mA) from other DNA. See, for example, Kumar et al., Frontiers Genet. 2018; 9: 640; Greer et al., Cell 2015; 161: 868-878. Antibodies specific to mA are described in Sun et al., Bioessays 2015; 37:1155-62. Antibodies against various modified nucleobases (such as thymine / uracil forms, including halogenated forms such as 5-bromouracil) are commercially available. Various modified bases can also be detected based on changes in their base pairing specificity. For example, hypoxanthine is a modified form of adenine that can be produced by deamination and is read as G in sequencing. See, for example, U.S. Patent 8,486,630; Brown, Genomes, 2nd Edition, John Wiley & Sons, Inc., New York, NY, 2002, Chapter 14, “Mutation, Repair, and Recombination”.
[0095] Enrichment / capture steps; amplification; adaptor; barcode In some embodiments, the methods disclosed herein include the step of capturing one or more target regions of DNA, such as cfDNA. Capture can be performed using any suitable method known in the art. In some embodiments, capture includes contacting the DNA to be captured with a set of target-specific probes. The target-specific probe set may have any of the characteristics of the target-specific probe set described herein, including but not limited to the characteristics set forth above and in the probe-related sections below. One or more subsamples prepared during the methods disclosed herein may be captured. In some embodiments, DNA is captured from at least a first subsample or a second subsample, e.g., at least a first subsample and a second subsample. If the first subsample undergoes a separation step (e.g., separating DNA initially containing a first nucleobase (e.g., hmC) from DNA initially not containing a first nucleobase, such as hmC-seal), capture can be performed on any one, any two, or all of the DNA initially containing a first nucleobase (e.g., hmC), the DNA initially not containing a first nucleobase, and the second subsample. In some embodiments, the subsamples are differentially tagged (e.g., as described herein) and then pooled prior to capture.
[0096] The capture step can be performed using conditions suitable for a specific nucleic acid hybridization, which are typically dependent to some extent on the characteristics of the probe, such as length, base composition, etc. Given the general knowledge of nucleic acid hybridization in the art, those skilled in the art will be familiar with appropriate conditions. In some embodiments, a complex of the target-specific probe and DNA is formed.
[0097] In some embodiments, the method described herein includes capturing more than one set of target regions of cfDNA obtained from a test subject. The target regions include epigenetic target regions that may exhibit differences in methylation levels and / or fragmentation patterns, depending on whether they originate from tumor cells or healthy cells. The target regions also include sequence-variable target regions that may exhibit sequence differences, depending on whether they originate from tumor cells or healthy cells. The capture step produces a capture set of cfDNA molecules, and within the capture set of cfDNA molecules, cfDNA molecules corresponding to the sequence-variable target region set are captured with a greater capture yield than cfDNA molecules corresponding to the epigenetic target region set. For further discussion of the capture steps, capture yield, and related aspects, see WO2020 / 160414, which is incorporated herein by reference for all purposes.
[0098] In some implementations, the method described herein includes contacting cfDNA obtained from a test subject with a target-specific probe set, wherein the target-specific probe set is configured to capture cfDNA corresponding to a sequence-variable target region set with a greater capture yield than cfDNA corresponding to a set of epigenetic target regions.
[0099] Capturing cfDNA corresponding to a set of sequence-variable target regions at a higher capture yield than that corresponding to the epigenetic target region set is advantageous, because analyzing sequence-variable target regions with sufficient confidence or accuracy may require a greater sequencing depth than analyzing epigenetic target regions. The amount of data required to determine fragmentation patterns (e.g., testing for perturbations of transcription start sites or CTCF binding sites) or fragment abundance (e.g., in hypermethylated and hypomethylated regions) is generally less than the amount of data required to determine the presence or absence of cancer-related sequence mutations. Capturing the target region set at different yields can facilitate sequencing the target regions to different sequencing depths within the same sequencing run (e.g., using pooled mixtures and / or within the same sequencing pool).
[0100] In various embodiments, the method further includes sequencing the captured cfDNA to, for example, different sequencing depths for epigenetic target sets and sequence-variable target sets, consistent with those discussed herein. In some embodiments, the complex of the target-specific probe and DNA is separated from DNA not bound to the target-specific probe. For example, in cases where the target-specific probe is covalently or nonvalently bound to a solid support, washing or aspiration steps can be used to separate the unbound material. Alternatively, chromatography can be used where the complex has different chromatographic properties than the unbound material (e.g., where the probe contains ligands that bind to chromatographic resins).
[0101] As discussed in detail elsewhere herein, target-specific probe sets may include more than one set, such as probes for a sequence-variable target region set and probes for an epigenetic target region set. In some such embodiments, the capture step is performed simultaneously using probes for sequence-variable targets and probes for epigenetic targets in the same container; for example, probes for both the sequence-variable and epigenetic target regions are in the same composition. This method provides a relatively more efficient workflow. In some embodiments, the concentration of probes for the sequence-variable target region set is greater than the concentration of probes for the epigenetic target region set.
[0102] Optionally, a capture step is performed in a first container with a sequence-variable target region probe set and in a second container with an epigenetic target region probe set, or a contact step is performed in the first time and the first container with a sequence-variable target region probe set and in a second time before or after the first time with an epigenetic target region probe set. This method allows for the preparation of separate first and second compositions comprising captured DNA corresponding to a sequence-variable target region set and captured DNA corresponding to an epigenetic target region set. The compositions can be processed individually as desired (e.g., graded based on methylation, as described elsewhere herein) and recombine in appropriate proportions to provide material for further processing and analysis, such as sequencing.
[0103] In some implementations, DNA is amplified. In some implementations, amplification occurs before the capture step. In some implementations, amplification occurs after the capture step.
[0104] In some implementations, the DNA contains an adaptor. This can be performed simultaneously with the amplification process, for example, by providing the adaptor in the 5' portion of the primer, as described above. Alternatively, the adaptor can be added by other methods such as ligation.
[0105] In some implementations, the DNA contains a tag, which may be a barcode or contain a barcode. The tag can aid in identifying the origin of the nucleic acid. For example, a barcode can be used to allow identification of the origin of the DNA after pooling more than one sample for parallel sequencing, e.g., from a subject. This can be performed concurrently with the amplification procedure, e.g., by providing a barcode in the 5' portion of the primer, as described above. In some implementations, the adaptor and the tag / barcode are provided by the same primer or primer set. For example, the barcode may be located at the 3' of the adaptor and the 5' of the target hybridization portion of the primer. Alternatively, the barcode can be added by other methods, such as ligation, optionally along with the adaptor in the same ligation substrate.
[0106] Further details regarding amplification, labeling, and barcodes are discussed in the following “General Characteristics of the Method” section, and these details can be combined, to a degree of feasibility, with any of the aforementioned implementation schemes and the implementation schemes described in the “Introduction and Overview” section.
[0107] capture set In some embodiments, a capture set of DNA (e.g., cfDNA) is provided. For the disclosed methods, the capture set of DNA may be provided, for example, by performing a capture step after a partitioning step as described herein. The capture set may include DNA corresponding to a set of sequence-variable target regions, DNA corresponding to a set of epigenetic target regions, or a combination thereof. In some embodiments, the amount of captured sequence-variable target region DNA is greater than the amount of captured epigenetic target region DNA when normalized for differences in target region size (footprint size).
[0108] Optionally, a first capture set and a second capture set may be provided, comprising DNA corresponding to a sequence-variable target region set and DNA corresponding to an epigenetic target region set, respectively. The first capture set and the second capture set may be combined to provide a combined capture set.
[0109] In some embodiments, where the capture set, which includes DNA corresponding to both sequence-variable target regions and epigenetic target regions, comprises a combination of capture sets as discussed above, the DNA corresponding to the sequence-variable target regions may be present at a higher concentration than the DNA corresponding to the epigenetic target regions, for example, 1.1 to 1.2 times higher, 1.2 to 1.4 times higher, 1.4 to 1.6 times higher, 1.6 to 1.8 times higher, 1.8 to 2.0 times higher, or 2.0 to 2.2 times higher. Concentration, 2.2 to 2.4 times larger concentration, 2.4 to 2.6 times larger concentration, 2.6 to 2.8 times larger concentration, 2.8 to 3.0 times larger concentration, 3.0 to 3.5 times larger concentration, 3.5 to 4.0 times larger concentration, 4.0 to 4.5 times larger concentration, 4.5 to 5.0 times larger concentration, 5.0 to 5.5 times larger concentration, 5.5 to 6.0 times larger concentration, 6.0 to 6.5 times larger concentration, 6.5 to 7.0 times larger concentration, 7.0 to 7. 5 times larger concentration, 7.5 to 8.0 times larger concentration, 8.0 to 8.5 times larger concentration, 8.5 to 9.0 times larger concentration, 9.0 to 9.5 times larger concentration, 9.5 to 10.0 times larger concentration, 10 to 11 times larger concentration, 11 to 12 times larger concentration, 12 to 13 times larger concentration, 13 to 14 times larger concentration, 14 to 15 times larger concentration, 15 to 16 times larger concentration, 16 to 17 times larger concentration, 17 to 18 times larger concentration, 18 times larger to Concentrations ranging from 19 to 20 times, 20 to 30 times, 30 to 40 times, 40 to 50 times, 50 to 60 times, 60 to 70 times, 70 to 80 times, 80 to 90 times, 90 to 100 times, 10 to 20 times, 10 to 40 times, 10 to 50 times, 10 to 70 times, or 10 to 100 times. The degree of concentration variation is calculated based on normalization for the target footprint size, as discussed in the definitions section.
[0110] Epigenetic target region set An epigenetic target set may include one or more types of target regions that can distinguish DNA from DNA from vegetative (e.g., tumor or cancer) cells from DNA from healthy cells (e.g., non-vegetative circulating cells). Example types of such regions are discussed in detail herein. An epigenetic target set may also include one or more control regions, such as those described herein. In some embodiments, the epigenetic target set has a footprint of at least 100 kb, for example, at least 200 kb, at least 300 kb, or at least 400 kb. In some embodiments, the epigenetic target set has a footprint in the range of 100-1000 kb, for example, 100-200 kb, 200-300 kb, 300-400 kb, 400-500 kb, 500-600 kb, 600-700 kb, 700-800 kb, 800-900 kb, and 900-1000 kb.
[0111] Highly methylated variable target region In some implementations, the epigenetic target set includes one or more hypermethylated variable target regions. Typically, a hypermethylated variable target region refers to a region in which, for example in a cfDNA sample, an observed increase in methylation levels indicates an increased likelihood that the sample (e.g., cfDNA) contains DNA produced by invasive cells (such as tumor cells or cancer cells). For example, hypermethylation of tumor suppressor gene promoters has been repeatedly observed. See, for example, Kang et al., Genome Biol. 18:53 (2017) and the references cited therein. In examples, a hypermethylated variable target region may include a region in which the methylation is not necessarily different from that of DNA from the same type of healthy tissue, but is indeed different from that of typical cfDNA in healthy subjects (e.g., having more methylation). For example, such a hypermethylated variable target region can be used at least in part to detect cancer when the presence of cancer leads to increased cell death (such as apoptosis corresponding to the tissue type of cancer). In some implementations, the hypermethylated variable target region includes one or more genomic regions in which the methylation status of cfDNA molecules in these regions is not different from that in cfDNA from healthy subjects, but the presence / increased amount of hypermethylated cfDNA in these regions indicates a specific tissue type (e.g., cancer origin) and is presented as cfDNA in circulation due to increased apoptosis (e.g., tumor shedding).
[0112] Hypermethylated target regions can be obtained, for example, from the Cancer Genome Atlas. Kang et al., Genome Biology 18:53 (2017) describe a probabilistic approach to constructing a cancer locator using hypermethylated target regions from the breast, colon, kidney, liver, and lung. In some embodiments, the hypermethylated target regions can be specific to one or more types of cancer. Thus, in some embodiments, the hypermethylated target regions comprise one, two, three, four, or five subsets of hypermethylated target regions that collectively exhibit hypermethylation in one, two, three, four, or five of the following cancers: breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer.
[0113] In some embodiments, probes targeting a set of epigenetic target regions include probes specific to one or more hypermethylated variable target regions. Hypermethylated variable target regions can be any of the hypermethylated variable target regions listed above. For example, in some embodiments, probes specific to hypermethylated variable target regions include probes specific to more than one locus listed in Table 1 (e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1). In some embodiments, probes specific to hypermethylated variable target regions include probes specific to more than one locus listed in Table 2 (e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 2). In some embodiments, probes specific to hypermethylated variable target regions include probes specific to more than one locus listed in Table 1 or Table 2 (e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1 or Table 2). In some embodiments, for each locus included as a target region, there may be one or more probes having hybridization sites that bind between the transcription start site and the stop codon (or the final stop codon for alternatively spliced genes) of that gene. In some embodiments, one or more probes bind within 300 bp (e.g., within 200 bp or 100 bp) of the listed positions. In some embodiments, the probes have hybridization sites that overlap with the positions listed above. In some embodiments, probes specific to hypermethylated target regions include probes specific to one, two, three, four, or five subsets of hypermethylated target regions that collectively exhibit hypermethylation in one, two, three, four, or five of breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer.
[0114] Hypomethylated variable target region Universal hypomethylation is a phenomenon commonly observed in a variety of cancers. See, for example, Hon et al., GenomeRes. 22:246-258 (2012) (breast cancer); Ehrlich, Epigenomics 1:239-259 (2009) (a review article noting observations of hypomethylation in colon cancer, ovarian cancer, prostate cancer, leukemia, hepatocellular carcinoma, and cervical cancer). For example, regions that are normally methylated in healthy cells (such as repetitive elements (e.g., LINE1 elements, Alu elements, centromere tandem repeats, pericentromere tandem repeats, and satellite DNA) and intergenic regions) may show reduced methylation in tumor cells. Therefore, in some embodiments, the epigenetic target set includes hypomethylated variable target regions, where the observed reduction in methylation levels indicates an increased likelihood that the sample (e.g., cfDNA) contains DNA produced by cytoplasmic cells (such as tumor cells or cancer cells). In examples, a hypomethylated variable target region may include a region whose methylation state is not necessarily different in cancerous tissue relative to DNA from the same type of healthy tissue, but is indeed different in methylation (e.g., less methylated) relative to cfDNA typical of healthy subjects. For example, such a hypomethylated variable target region can be used at least in part to detect cancer when the presence of cancer leads to increased cell death (such as apoptosis corresponding to the tissue type of cancer). In some embodiments, the hypomethylated variable target region includes one or more genomic regions in which the methylation state of cfDNA molecules in these regions is not different from cfDNA from healthy subjects, but the amount of hypomethylated cfDNA present / increased in these regions indicates a specific tissue type (e.g., cancer origin) and is presented as cfDNA entering circulation accompanied by increased apoptosis (e.g., tumor shedding).
[0115] In some embodiments, the hypomethylated variable target region includes repeating elements and / or intergenic regions. In some embodiments, repeating elements include one, two, three, four, or five of the following: LINE1 elements, Alu elements, centromere tandem repeat sequences, pericentromere tandem repeat sequences, and / or satellite DNA.
[0116] Example specific genomic regions exhibiting cancer-related hypomethylation include nucleotides 8403565-8953708 and 151104701-151106035 on human chromosome 1. In some embodiments, the hypomethylated variable target region overlaps with or includes one or both of these regions.
[0117] In some implementations, probes targeting a set of epigenetic target regions include probes specific to one or more hypomethylated variable target regions. Hypomethylated variable target regions can be any of the hypomethylated target regions listed above. For example, probes specific to one or more hypomethylated variable target regions may include probes targeting regions such as repetitive elements (e.g., LINE1 elements, Alu elements, centromere tandem repeats, pericentromere tandem repeats, and satellite DNA) and intergenic regions that are normally methylated in healthy cells but may exhibit reduced methylation in tumor cells.
[0118] In some embodiments, probes specific to hypomethylated variable target regions include probes specific to repetitive elements and / or intergenic regions. In some embodiments, probes specific to repetitive elements include probes specific to one, two, three, four, or five of the following: LINE1 elements, Alu elements, centromere tandem repeat sequences, pericentromere tandem repeat sequences, and / or satellite DNA.
[0119] Example probes specific to genomic regions exhibiting cancer-related hypomethylation include probes specific to nucleotides 8403565-8953708 and / or 151104701-151106035 of human chromosome 1. In some embodiments, probes specific to hypomethylated variable target regions include probes specific to regions that overlap with or contain nucleotides 8403565-8953708 and / or 151104701-151106035 of human chromosome 1.
[0120] Probes used to detect this set of regions may include probes for detecting genomic regions of interest (hotspot regions) and nucleosome-sensing probes (e.g., KRAS codons 12 and 13), and can be designed to optimize capture based on analysis of cfDNA coverage and fragment size variation and GC sequence composition affected by nucleosome binding patterns. The regions used in this paper may also include non-hotspot regions optimized based on nucleosome location and GC models.
[0121] Subjects In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject with cancer. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject suspected of having cancer. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject with a tumor. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject suspected of having a tumor. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject with a growth. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject suspected of having a growth. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject in remission from a tumor, cancer, or growth (e.g., after chemotherapy, surgical resection, radiation, or a combination thereof). In any of the foregoing embodiments, the cancer, tumor, or growth, or suspected cancer, tumor, or growth, may be of the lung, colon, rectum, kidney, breast, prostate, or liver. In some embodiments, the cancer, tumor, or growth, or suspected cancer, tumor, or growth, is of the lung. In some embodiments, the cancer, tumor, or growth, or suspected cancer, tumor, or growth, is of the colon or rectum. In some embodiments, the cancer, tumor, or growth, or suspected cancer, tumor, or growth, is of the breast. In some embodiments, the cancer, tumor, or growth, or suspected cancer, tumor, or growth, is of the prostate. In any of the foregoing embodiments, the subject may be a human subject.
[0122] In some embodiments, the sequence-variable target probe set has a footprint of at least 0.5 kb (e.g., at least 1 kb, at least 2 kb, at least 5 kb, at least 10 kb, at least 20 kb, at least 30 kb, or at least 40 kb). In some embodiments, the epigenetic target probe set has a footprint in the range of 0.5–100 kb (e.g., 0.5–2 kb, 2–10 kb, 10–20 kb, 20–30 kb, 30–40 kb, 40–50 kb, 50–60 kb, 60–70 kb, 70–80 kb, 80–90 kb, and 90–100 kb).
[0123] In some embodiments, probes specific to a sequence-variable target set include probes specific to targets from at least 10, 20, 30, or 35 cancer-related genes, such as AKT1, ALK, BRAF, CCND1, CDK2A, CTNNB1, EGFR, ERBB2, ESR1, FGFR1, FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11, TP53, and U2AF1.
[0124] Composition containing captured DNA This document provides a combination of a first group and a second group comprising captured DNA. The first group may comprise or be derived from DNA with a greater proportion of cytosine modifications than the second group. The first group may comprise a form of a first nucleobase initially present in the DNA with altered base pairing specificity and a second nucleobase without altered base pairing specificity, wherein the form of the first nucleobase initially present in the DNA before the alteration of base pairing specificity is a modified or unmodified nucleobase, and the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the form of the first nucleobase initially present in the DNA before the alteration of base pairing specificity and the second nucleobase have the same base pairing specificity. The second group does not comprise a form of the first nucleobase initially present in the DNA with altered base pairing specificity. In some embodiments, the cytosine modification is cytosine methylation. In some embodiments, the first nucleobase is modified or unmodified cytosine, and the second nucleobase is modified or unmodified cytosine. The first and second nucleobases can be any nucleobases discussed in the overview or regarding procedures that subject the first subsample to different effects on the first and second nucleobases in the DNA of the first subsample.
[0125] In some implementations, the first group includes sequence tags selected from one or more sequence tags in the first group, and the second group includes sequence tags selected from one or more sequence tags in the second group, wherein the sequence tags in the second group are different from the sequence tags in the first group. The sequence tags may include barcodes.
[0126] In some embodiments, the first population comprises protected hmC, such as glucosylated hmC. In some embodiments, the first population undergoes any of the transformation procedures discussed herein, such as bisulfite transformation, Ox-BS transformation, TAB transformation, ACE transformation, TAP transformation, TAPSβ transformation, or CAP transformation. In some embodiments, the first population is protected by hmC followed by deamination of mC and / or C. In some combined embodiments, the first population comprises or is derived from DNA having a larger proportion of cytosine modification than the second population, and the first population comprises both first and second subpopulations, and the first nucleotide is a modified or unmodified nucleotide, the second nucleotide is a modified or unmodified nucleotide different from the first nucleotide, and the first and second nucleotides have the same base-pairing specificity. In some embodiments, the second population does not contain the first nucleotide. In some embodiments, the first nucleotide is a modified or unmodified cytosine, and the second nucleotide is a modified or unmodified cytosine, optionally wherein the modified cytosine is mC or hmC. In some embodiments, the first nucleobase is a modified or unmodified adenine, and the second nucleobase is a modified or unmodified adenine, optionally wherein the modified adenine is mA.
[0127] In some embodiments, the first nucleobase (e.g., modified cytosine) is biotinylated. In some embodiments, the first nucleobase (e.g., modified cytosine) is a product of Huisgen cycloaddition of β-6-azido-glucosyl-5-hydroxymethylcytosine, the product containing an affinity tag (e.g., biotin).
[0128] In any of the combinations described herein, the captured DNA may include cfDNA. The captured DNA may have any of the characteristics described herein with respect to the capture set, including, for example, a higher concentration of DNA corresponding to a sequence-variable target region set than the concentration of DNA corresponding to an epigenetic target region set (as normalized for footprint size as discussed above). In some embodiments, the captured DNA contains a sequence tag, which may be added to the DNA as described herein. Typically, the inclusion of a sequence tag results in DNA molecules that differ from their naturally occurring, untagged form.
[0129] This combination may also include the probe set or sequencing primers described herein, each of which may differ from naturally occurring nucleic acid molecules. For example, the probe set described herein may contain a capture portion, and the sequencing primers may contain a non-naturally occurring marker.
[0130] Computer systems and processing of real-world evidence (RWE) The methods of this disclosure can be implemented using or by means of a computer system. For example, such a method may include: partitioning a sample into more than one subsample, said more than one subsample including a first subsample and a second subsample, wherein the first subsample contains DNA with a larger proportion of cytosine modification than the second subsample; subjecting the first subsample to procedures that differently affect a first nucleobase and a second nucleobase in the DNA of the first subsample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first and second nucleobases have the same base-pairing specificity; and sequencing the DNA in the first subsample and the DNA in the second subsample in a manner that distinguishes the first and second nucleobases in the DNA of the first subsample.
[0131] In one aspect, this disclosure provides a non-transitory computer-readable medium including computer-executable instructions that, when executed by at least one electronic processor, perform at least a portion of a method comprising: collecting cfDNA from a test subject; capturing more than one set of target regions from the cfDNA, wherein the more than one set of target regions includes a sequence-variable target region set and an epigenetic target region set, thereby generating a capture set of cfDNA molecules; sequencing the captured cfDNA molecules, wherein the captured cfDNA molecules of the sequence-variable target region set are sequenced to a deeper sequencing depth than the captured cfDNA molecules of the epigenetic target region set; obtaining more than one sequence read generated by a nucleic acid sequencer through sequencing the captured cfDNA molecules; mapping the more than one sequence read to one or more reference sequences to generate mapped sequence reads; and processing the mapped sequence reads corresponding to the sequence-variable target region set and the epigenetic target region set to determine the likelihood that the subject has cancer.
[0132] The code can be pre-compiled and configured for use with machines having processors suitable for executing the code, or it can be compiled during runtime. The code can be provided in a programming language, which can be selected to enable the code to be executed in a pre-compiled or just-in-time (JIT) compiled manner.
[0133] Additional details relating to computer systems and networks, databases, and computer program products are provided, for example, in the following: Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th edition (2011); Kurose, Computer Networking: A Top-Down Approach, Pearson, 7th edition (2016); Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th edition (2010); Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11th edition (2014); Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd edition (2006); and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011), all of which are incorporated herein by reference in their entirety. Example
[0134] Example 1: Predicted Tissue of Origin (TOO) As described, the clinical utility of TOO identification guides treatment decisions and potentially improves prognosis. In addition to identifying the primary site of unknown and CUP sites, TOO identification can also decipher insights into molecular pathology, drivers and targetable alterations, as well as gene expression and methylation imprinting.
[0135] In addition, it provides identification of cancer of unknown origin (CUP) when the cancer cells or tissue found at the site are different from the expected cancer cell type, and in the case of metastatic cancer where the origin / site of the primary cancer is unknown.
[0136] TOO sites can also be ranked based on association probability reports, including an initial classification involving cancer versus non-cancer, and a two-step procedure for subsequent TOO identification using a multi-class classification model. To support TOO development, comparisons of genomic MAF and epigenomic MAF for selected samples used in computer LoD estimation were generated, such as... Figure 1 As shown in the image.
[0137] Example 2: Two-layer frame TOOs are classified in a two-step process using methylation data from the dataset. The first step is a binary classification of cancer conditions, and the second step is a multi-class classification of TOOs. An example is shown below. Figure 2 Binary classification is typically accomplished using methods such as logistic regression. In multi-class classification, there are more than two categories for a given target variable (here, cancer type).
[0138] Example 3: One-layer frame As another example, a second framework for analysis could include using a classifier built for each cancer type, where a score is determined for each cancer type and a prediction is made when the score exceeds a threshold.
[0139] To support the identification of examples, multicancer methylome data used for screening, minimal residual disease (MRD), and epigenomic assays can be utilized, taking into account various versions of groups, workflows, and cancer types / stages. Those skilled in the art will readily understand that larger group sizes or whole-genome methylation data improve accuracy.
[0140] Example 4: Classification of Early Detection of Multiple Cancers As described, the process involves a first-step cancer / non-cancer binary classifier, followed by a second-step tissue of origin (TOO) classification using multi-class classification, such as... Figure 3 As shown in the diagram, logistic regression is used for binary classification of cancer conditions. In another example, multi-class classification is typically performed using Naive Bayes, decision trees, support vector machines (SVM), random forest classifiers, k-nearest neighbors (KNN), or neural networks. In the second step, the TOO MC model is applied to each sample that passes the 98% specificity threshold.
[0141] Features of Example 5 While this paper describes methylation as an example of classifying cancer conditions, those skilled in the art will readily understand that other features, such as fragmentomics, histone modifications, chromatin state, proteomics, and microbiome, can be used. As data quality improves and new analytical methods are developed, integrating such features to train models can yield better results. In some cases, multiple sets of detections can be used together.
[0142] Example 6: Classifier Performance The performance of the TOO multi-cancer classifier (where instances range from colorectal cancer (CRC) to gastric cancer) using a multi-cancer methyl group dataset is described here. Figure 4The confusion matrix shown is used to describe the model's performance on the test data for each classification problem. The performance of the multi-class classifier is illustrated here. Those skilled in the art will readily understand that larger group sizes or whole-genome methylation data improve accuracy. Larger sample sizes will also improve accuracy, particularly for breast-gastric data.
[0143] All patent applications, websites, other publications, registration numbers, etc., cited above or below are incorporated by reference in their entirety for all purposes, to the extent that each individual item is specifically and individually indicated by reference. If different versions of a sequence are associated with a registration number at different times, it refers to the version associated with that registration number on the effective filing date of this application. The effective filing date refers to the earlier of the actual filing date using that registration number or the filing date of the priority application (if applicable). Similarly, if different versions of publications, websites, etc., are published at different times, it refers to the most recently published version on the effective filing date of the application, unless otherwise indicated. Any feature, step, element, embodiment, or aspect of this disclosure may be used in combination with any other feature, step, element, embodiment, or aspect, unless otherwise specifically indicated. Although this disclosure has been described in considerable detail by way of illustration and example for purposes of clarity and understanding, it will be apparent that certain changes and modifications may be made within the scope of the appended claims.
Claims
1. A computer-implemented method, comprising: Receiving one or more datasets in a computer system that includes one or more hardware processors and one or more computer-readable storage media. The dataset includes molecular phenotypes obtained from test samples from patients, and the computer-readable medium includes instructions that, when executed by the processor, cause the one or more hardware processors to perform one or more classifications of the test samples.
2. The computer-implemented method according to claim 1, wherein the molecular phenotype includes the methylation status of more than one genomic region, histone modifications, chromatin status, fragment length, or transcription factor occupancy.
3. The computer-implemented method according to claim 1, wherein the molecular phenotype includes epigenetic data, genome sequence data, proteomic data, microbiome data, imaging data, histological data, and / or metadata.
4. The computer-implemented method according to claim 3, wherein the epigenetic data includes methylation, histone acetylation, chromatin state, or DNA circularization interaction data.
5. The computer-implemented method of claim 1 further includes performing a first classification and a subsequent second classification, wherein the second classification is performed only if the first classification includes a predetermined category.
6. The computer-implemented method according to claim 5, wherein the first classification and the second classification are performed using a logistic regression model.
7. The computer-implemented method of claim 5, wherein the first classification is performed using a logistic regression model, and the second classification is performed using Naive Bayes, decision tree, support vector machine (SVM), random forest classifier, k-nearest neighbor (KNN) or neural network.
8. The computer-implemented method according to any one of the preceding claims, wherein the first classification is a cancer condition and the second classification is a cancer type.
9. The computer-implemented method according to any one of the preceding claims, wherein the cancer condition is a binary category including both cancer and non-cancer conditions.
10. The computer-implemented method according to any one of the preceding claims, wherein the cancer type includes breast cancer, colorectal cancer, lung cancer, bladder cancer, pancreatic cancer, ovarian cancer, liver cancer, stomach cancer, esophageal cancer, kidney cancer, melanoma, gallbladder cancer, or uterine cancer.
11. The computer-implemented method according to any one of the preceding claims, wherein the cancer type further includes the tissue from which the cancer originates in the patient.
12. The computer-implemented method according to any one of the preceding claims, wherein the cancer type further includes tissue of unknown primary cancer (CUP) originating in the patient.
13. The computer-implemented method of claim 1, wherein the classification is based on a set of cancer-specific models.
14. The computer-implemented method of claim 13, wherein each cancer-specific model outputs a score for the cancer type.
15. The computer-implemented method of claim 14, wherein when the score exceeds a threshold, a cancer type prediction is performed.
16. The computer-implemented method of claim 14, wherein cancer type prediction is not performed when the score is below a threshold.
17. The computer-implemented method of claim 15, wherein tumor score estimation is performed for the cancer type.
18. The computer-implemented method of claim 13, wherein all scoring outputs below a threshold have a tumor score of zero (TF=0) and no cancer marker.
19. The computer-implemented method according to any one of the preceding claims, wherein the sample comprises cell-free DNA (cfDNA).
20. The computer-implemented method according to any one of the preceding claims, wherein the sample comprises blood, plasma, saliva, or urine.
21. The computer-implemented method according to any one of the preceding claims, wherein the sample comprises a biological fluid, a biological solid, or a biological tissue.
22. A method comprising: Samples obtained or already obtained from the subjects; Detect one or more features in the sample; as well as The cancer status of the subjects was classified.
23. The method of claim 22, wherein the cancer condition includes identifying the tissue of origin for one or more cells in the sample.
24. The method of claim 22, further comprising determining the tissue of origin for one or more cells in the sample.
25. The method of claim 22, wherein the one or more features include the methylation status of more than one genomic region, histone modifications, chromatin status, fragment length, or transcription factor occupancy.
26. The method of claim 25, wherein one or more features include epigenetic data, genome sequence data, proteomic data, microbiome data, imaging data, histological data, and / or metadata.
27. The method of claim 26, wherein the epigenetic data includes methylation, histone acetylation, chromatin state, or DNA circularization interaction data.
28. The method of claim 22, further comprising performing a first classification and a subsequent second classification, wherein the second classification is performed only if the first classification includes a predetermined category.
29. The method of claim 28, wherein the first classification and the second classification are performed using a logistic regression model.
30. The method of claim 29, wherein the first classification is performed using a logistic regression model, and the second classification is performed using a Naive Bayes, decision tree, support vector machine (SVM), random forest classifier, k-nearest neighbor (KNN), or neural network.
31. The method according to any one of the preceding claims, wherein the first classification is a cancer condition and the second classification is a cancer type.
32. The method according to any one of the preceding claims, wherein the cancer condition is a binary category including both cancer and non-cancer conditions.
33. The method according to any one of the preceding claims, wherein the cancer type includes breast cancer, colorectal cancer, lung cancer, bladder cancer, pancreatic cancer, ovarian cancer, liver cancer, stomach cancer, esophageal cancer, kidney cancer, melanoma, gallbladder cancer, or uterine cancer.
34. The method according to any one of the preceding claims, wherein the cancer type further includes the tissue from which the cancer originates in the patient.
35. The method according to any one of the preceding claims, wherein the cancer type further comprises tissue of unknown primary cancer (CUP) originating in the patient.
36. The method of claim 22, wherein the classification is based on a cancer-specific model set.
37. A computer-implemented method, comprising: Receiving one or more datasets in a computer system that includes one or more hardware processors and one or more computer-readable storage media. The dataset includes molecular phenotypes obtained from test samples from patients, and the computer-readable medium includes instructions that, when executed by the processor, cause one or more hardware processors to perform one or more classifications of the test samples, the one or more classifications including a first classification and a subsequent second classification, the first classification including a cancer condition, the second classification including a cancer type, wherein the second classification is performed only if the first classification includes a predetermined category, wherein the first classification and / or the second classification is based on a cancer-specific model set, wherein the molecular phenotype includes epigenetic data, and wherein the cancer type also includes the tissue from which the cancer originates in the patient.
38. A system capable of performing any of the preceding claims.
Citation Information
Patent Citations
Methods for accurate sequence data and modified base position determination
US8486630B2
Methods and systems for analyzing nucleic acid molecules
WO2018119452A2
Compositions and methods for isolating cell-free DNA
WO2020160414A1