Method for layering and early detection of advanced adenoma and / or colorectal cancer using DNA methylation markers
Through DNA methylation marker-bound fragment size analysis, the shortcomings of existing tools in colorectal cancer screening were solved, and high-precision detection and diagnosis of advanced adenomas and colorectal cancer were achieved, improving the accuracy and compliance of screening.
Patent Information
- Application Number
- CN202380090622.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-12
- Filing Date
- 2023-11-15
- Publication Date
- 2025-08-08
AI Technical Summary
The lack of existing disease screening and diagnostic tools in colorectal cancer screening has led to a large number of patients not being promptly detected, and more effective markers are needed for non-invasive and highly adherent screening environments.
DNA methylation marker-bound fragment size analysis was used to identify and verify the methylation status in cell-free DNA particles through next-generation sequencing technology, and was used to detect, diagnose, predict and stage advanced adenomas and colorectal cancer.
High-precision screening of advanced adenomas and colorectal cancer is achieved, reducing negative and false positive events, and improving the accuracy of patient diagnosis, prognosis and clinical outcomes.
Smart Images

Figure BDA0005483559010000021 
Figure BDA0005483559010000031 
Figure BDA0005483559010000041
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to and the benefit of U.S. Provisional Application No. 63 / 425,882, filed on November 16, 2022, the contents of which are incorporated herein by reference in their entirety. Technical Field
[0003] The present invention generally relates to methods and systems for detecting, diagnosing, predicting, monitoring, screening, staging, and / or providing a survival prognosis for advanced adenomas and / or colorectal cancer. Background Art
[0004] Disease detection is an important part of preventing disease progression, diagnosis, and treatment. For example, early detection of colorectal cancer (CRC) has been shown to improve outcomes in CRC patients through early treatment of CRC. However, despite existing tools for screening and diagnosing CRC and other cancers, millions of individuals still die each year from diseases such as CRC, which can be treated through early intervention and detection. Existing disease screening and diagnostic tools are insufficient.
[0005] Deoxyribonucleic acid (DNA) methylation (DNAme) is an important epigenetic marker across various species. DNA methylation in vertebrates is characterized by the addition of a methyl or hydroxymethyl group to the C5 position of cytosine, primarily occurring in the context of CG dinucleotides. Alterations in DNA methylation status are a mechanism for the inactivation of cancer-associated genes, including tumor suppressor genes, in CRC and other human cancers. Figure 1 A simplified example showing how DNA methylation alters gene activation in tumor cells compared to normal cells.
[0006] DNA methylation affects a variety of cellular processes, including, for example, cell differentiation. Therefore, dysregulated methylation can lead to diseases, including cancer. Accumulation of altered DNA methylation (e.g., hypermethylation or hypomethylation) can produce cancer cells, particularly when the alterations are located in key genes. If these alterations in methylation status are detected, they can be used to predict a subject's susceptibility to cancer, as well as the occurrence or presence of cancer.
[0007] New, more effective markers that can be used in a non-invasive and high-compliance CRC screening setting are needed. Summary of the Invention
[0008] The present disclosure provides, among other things, methods and systems for detecting, diagnosing, predicting, monitoring, screening, staging, and / or providing a survival prognosis for advanced adenoma and / or colorectal cancer by using DNA methylation markers and further studying the fragment size of cell-free DNA particles carrying the methylation signal. For example, the present disclosure describes DNA methylation markers combined with fragment size values that enable high-precision classification and prognostic assessment of patients with advanced adenoma (AA) and / or colorectal cancer (CRC) from human biological samples.
[0009] In certain embodiments, the marker or a marker panel consisting of the marker provides high specificity for advanced adenoma and / or colorectal cancer screening at an early or precancerous stage, minimizing negative and false positive events when testing samples extracted from subjects. The methylation markers described herein achieve high accuracy AA and CRC stratification for patients with anonymized data obtained from The Cancer Genome Atlas (TCGA) database.
[0010] Without wishing to be bound by a particular theory, methylation events acquired in the promoter regions of tumor suppressor genes are thought to silence expression, whereas demethylation events in the promoter regions of proto-oncogenes are thought to activate these genes and thereby promote tumorigenesis. For use as a diagnostic tool, DNA methylation is considered to be a chemically and biologically more stable process than RNA or protein expression. In addition, methylation events are thought to be highly tissue specific, making these markers more informative and sensitive than DNA mutations alone, and providing high specificity. Aberrant methylation patterns of cpG islands located in the promoter regions of tumor suppressor genes are an important mechanism of gene inactivation. DNA methylation affects the readability of DNA sequences wrapped around nucleosomes, linking methylation analysis to cfDNA fragment size parameters.
[0011] In experiments conducted concurrently with the development of the biomarker panel described here, markers were first identified from tissue whole-genome bisulfite sequencing (WGBS) data and then validated in several cytoplasmic validation studies.
[0012] The resulting regions shown in Table 1 were further analyzed using comparisons of gene expression and methylation status of DNA markers in data from 455 patients obtained using the Illumina 450k methylation microarray, which were collected by the TCGA consortium.
[0013] Table 1 : List of 82 genomic regions found to have significantly altered methylation patterns in CRC patients
[0014]
[0015]
[0016]
[0017]
[0018]
[0019] Candidate markers provide the feasibility of differentially stratifying tumor-derived cells and predicting patient prognosis and clinical outcomes. Thus, this analysis provides specific markers or marker combinations for the purpose of stratifying AA and / or CRC and / or identifying tumor origins. As a result, improved accuracy in patient diagnosis, prognosis, clinical outcome, and survival prediction can be achieved.
[0020] In addition, in some embodiments, the marker combination is a combination of 1600 or fewer genomic regions, or, in some embodiments, the marker is a single region. In some embodiments, the marker is a high-density promoter region. In some embodiments, the marker and / or marker panel includes methylation and other mutations.
[0021] In one aspect, the present invention relates to a method for identifying (e.g., subtyping, stratifying, detecting, diagnosing, predicting, monitoring, screening, staging, and / or providing a survival prognosis) a condition in a human subject, the method comprising: determining the methylation state of each of one or more (e.g., at least 3, at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, or at least 82) markers identified, e.g., in deoxyribonucleic acid (DNA) fragments (DNA fragments) in a sample obtained from the subject; and identifying (e.g., subtyping, stratifying, detecting, diagnosing, predicting, monitoring, screening, staging, and / or providing a survival prognosis) the condition in the subject based at least in part on the determined methylation state of each of the one or more markers identified in the DNA fragments, wherein each of the one or more markers is a member selected from the group consisting of the 82 DMRs listed in Table 1 (i.e., SEQ ID NO: 1). 1 to 82) or a methylation locus of at least one single differentially methylated region (DMR) or a portion of a DMR (e.g., each of said portions contains at least one (1) CpG and / or the length of each of said methylation loci is equal to or less than 302 bp).
[0022] In certain embodiments, the identifying step comprises identifying a condition in the subject based at least in part on (i) the determined methylation status of each of the one or more markers identified in the DNA fragments and (ii) the determined size of each of the DNA fragments in which the one or more markers were identified.
[0023] In certain embodiments, the identifying step comprises identifying the condition in the subject based at least in part on (i) the determined methylation status of each of the one or more markers identified in the DNA fragments and (ii) the determined size and start / end nucleotide sequence of each of the DNA fragments in which the one or more markers are identified.
[0024] In certain embodiments, the disorder is colorectal cancer (CRC).
[0025] In certain embodiments, the condition is advanced adenoma (AA).
[0026] In certain embodiments, the method identifies whether a subject has colorectal cancer (CRC) or advanced adenoma (AA) as a non-differential diagnosis (i.e., the method identifies that the subject has CRC or AA, but the method does not indicate which of the two diagnoses the subject has).
[0027] In certain embodiments, the method identifies whether the subject has colorectal cancer (CRC) or advanced adenoma (AA) as a differential diagnosis (i.e., the method identifies that the subject has CRC or AA, and the method also indicates which of these two diagnoses the subject has).
[0028] In certain embodiments, the sample is a member selected from the group consisting of a tissue sample, a blood sample, a stool sample, and a blood product sample.
[0029] In certain embodiments, the sample comprises DNA isolated from blood or plasma of a human subject.
[0030] In certain embodiments, the DNA is cell-free DNA (cfDNA) from a human subject.
[0031] In certain embodiments, the method comprises detecting the methylation status of each of the one or more markers using next generation sequencing (NGS).
[0032] In certain embodiments, the method comprises using one or more capture baits enriched for a region of interest to capture one or more corresponding methylated loci.
[0033] In certain embodiments, each methylation locus is equal to or less than 302 bp in length.
[0034] In certain embodiments, the method includes: converting unmethylated cytosine in multiple DNA fragments in a sample to uracil to produce multiple converted DNA fragments, wherein the multiple DNA fragments are obtained from a biological sample; and sequencing the multiple converted DNA fragments to produce multiple sequence reads, wherein each sequence read corresponds to a converted DNA fragment.
[0035] In certain embodiments, the plurality of DNA fragments comprises (collectively) at least 1 ng, at least 5 ng, at least 10 ng, or at least 20 ng of DNA.
[0036] In certain embodiments, the plurality of DNA fragments consists of, or consists essentially of, DNA fragments each having a length of 10 bp to 800 bp (e.g., about 50 bp to about 250 bp; about 250 bp to about 480 bp; about 480 bp to about 800 bp; about 125 bp to about 200 bp; or about 140 bp to about 160 bp (e.g., for cfDNA)) (e.g., about 50 bp to about 150 bp, about 150 bp to about 350 bp; or about 200 bp to about 300 bp (e.g., for sheared DNA)).
[0037] In certain embodiments, the plurality of DNA fragments consists of, or consists essentially of, DNA fragments each having a length of 1000 bp to 200,000 bp [e.g., an average length of approximately 10,000 bp (e.g., for genomic DNA from a sample comprising tissue or a buffy coat)].
[0038] In certain embodiments, each of the plurality of sequence reads is at least 50 bp, at least 100 bp, at least 150 bp, at least 200 bp, at least 300 bp or more.
[0039] In another aspect, the present invention relates to a method for detecting hypermethylation of a marker (e.g., a biomarker, a DMR), the method comprising: detecting the methylation status of each of one or more markers identified in a deoxyribonucleic acid (DNA) fragment (DNA fragment) in a sample obtained from a human subject susceptible to (e.g., suspected of having) colorectal cancer and / or advanced adenoma, wherein each of the one or more markers is a methylation locus comprising at least one single differentially methylated region (DMR) or a portion of a DMR selected from the 82 DMRs listed in Table 1 (i.e., SEQ ID NO. 1 to 82) (e.g., each of the portions comprises at least one (1) CpG and / or the length of each of the methylation loci is equal to or less than 302 bp), wherein detecting the methylation status comprises determining whether at least one methylation site within at least one of the one or more markers is hypermethylated.
[0040] In certain embodiments, the sample is a member selected from the group consisting of a tissue sample, a blood sample, a stool sample, and a blood product sample.
[0041] In certain embodiments, the sample comprises DNA isolated from blood or plasma of a human subject.
[0042] In certain embodiments, the DNA is cell-free DNA (cfDNA) from a human subject.
[0043] In certain embodiments, the method comprises determining the methylation status of each of the one or more markers using next generation sequencing (NGS).
[0044] In certain embodiments, the method comprises using one or more capture baits enriched for a region of interest to capture one or more corresponding methylated loci.
[0045] In certain embodiments, each methylation locus is equal to or less than 302 bp in length.
[0046] In certain embodiments, at least one of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO. 3 or SEQ ID NO. 14.
[0047] In certain embodiments, at least one of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO.2.
[0048] In certain embodiments, at least one of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO. 1.
[0049] In certain embodiments, the first of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO.1, the second of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO.2, and the third of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO.3.
[0050] In certain embodiments, the first of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO. 70, and the second of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO. 72.
[0051] In another aspect, the present invention relates to a method for detecting the methylation status of a marker (e.g., a biomarker, a DMR), the method comprising: converting unmethylated cytosine of a plurality of DNA fragments in a sample into uracil to produce a plurality of converted DNA fragments, wherein the plurality of DNA fragments are obtained from a biological sample; sequencing the plurality of converted DNA fragments to produce a plurality of sequence reads, wherein each sequence read corresponds to the converted DNA fragment; and detecting the methylation status of each of one or more markers identified in the sequence reads, wherein each of the one or more markers is a methylation locus comprising at least one single differentially methylated region (DMR) or a portion of a DMR selected from the 82 DMRs listed in Table 1 (i.e., SEQ ID NO. 1 to 82) (e.g., each of the portions comprises at least one (1) CpG and / or the length of each of the methylation loci is equal to or less than 302 bp).
[0052] In certain embodiments, the plurality of DNA fragments comprises (collectively) at least 1 ng, at least 5 ng, at least 10 ng, or at least 20 ng of DNA.
[0053] In certain embodiments, the plurality of DNA fragments consists essentially of DNA fragments each having a length of 10 bp to 800 bp (e.g., about 50 bp to about 250 bp; about 250 bp to about 480 bp; about 480 bp to about 800 bp; about 125 bp to about 200 bp; or about 140 bp to about 160 bp (e.g., for cfDNA)) (e.g., about 50 bp to about 150 bp, about 150 bp to about 350 bp; or about 200 bp to about 300 bp (e.g., for sheared DNA)).
[0054] In certain embodiments, the plurality of DNA fragments consists essentially of DNA fragments each having a length of 1000 bp to 200,000 bp [e.g., an average length of approximately 10,000 bp (e.g., for genomic DNA from a sample comprising tissue or a buffy coat)].
[0055] In certain embodiments, each of the plurality of sequence reads is at least 50 bp, at least 100 bp, at least 150 bp, at least 200 bp, at least 300 bp or more.
[0056] In certain embodiments, the first of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO.1, the second of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO.2, and the third of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO.3.
[0057] In certain embodiments, the first of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO. 70, and the second of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO. 72.
[0058] In certain embodiments, the biological sample is a sample from a subject susceptible to (eg, suspected of having) colorectal cancer and / or advanced adenoma.
[0059] definition
[0060] In this application, the use of "or" means "and / or" unless otherwise indicated. The articles "a" and "an" used herein refer to one or more than one (i.e., at least one) grammatical object of the article. For example, "an element" refers to one element or more than one element. The term "comprising" and its variants (such as "including" and "containing") used in this application are not intended to exclude other additives, components, integers or steps. The terms "about" and "approximately" used in this application are used equivalently. Any number used in this application with or without "about" / "approximately" is intended to cover any normal fluctuations understood by those skilled in the relevant art. In certain embodiments, the term "about" or "approximately" refers to a numerical range that falls within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1% in either direction (above or below) of the reference value, unless otherwise specified or obvious from the context (unless the number exceeds 100% of the possible value).
[0061] Administration: The term "administration" as used herein generally refers to, for example, administering a composition to a subject or system in order to achieve the delivery of a medicament, which is contained in the composition or delivered by other means by the composition. Any appropriate route can be used for administration to an animal subject (e.g., a human being). For example, in certain embodiments, administration can be carried out by intrabronchial (including bronchial infusion), buccal, intestinal, intercutaneous, intraarterial, intradermal, intragastric, intramedullary, intramuscular, intranasal, intraperitoneal, intrathecal, intravenous, intraventricular, intracerebral, intracerebral, intracerebral organs (e.g., intrahepatic), mucosal, nasal, oral, rectal, subcutaneous, sublingual, topical, tracheal (including intratracheal infusion), transdermal, vaginal, and vitreous. In certain embodiments, administration may involve intermittent administration. In certain embodiments, administration may involve continuous administration (e.g., infusion) for at least a selected duration. As is known in the art, antibody therapy is typically administered parenterally (e.g., by intravenous or subcutaneous injection).
[0062] Advanced adenoma: As used herein, the term "advanced adenoma" refers to cells that show initial signs of relatively abnormal, uncontrolled, and / or spontaneous growth but are not yet classified as cancerous. In the context of colon tissue, "advanced adenoma" refers to a neoplastic growth that shows high-grade dysplasia and / or a size >= 10 mm and / or a villous histology and / or a serrated histology, as well as signs of any type of dysplasia.
[0063] Agent: As used herein, the term "agent" refers to an entity (e.g., a small molecule, peptide, polypeptide, nucleic acid, lipid, polysaccharide, complex, composition, mixture, system, or phenomenon such as heat, electric current, electric field, magnetism, magnetic field, etc.).
[0064] Improvement: As used herein, the term "improvement" refers to preventing, alleviating, relieving, or improving a subject's condition. Improvement includes, but does not require, complete recovery or complete prevention of a disease, disorder, or condition.
[0065] Amplicon or amplicon molecule: As used herein, the term "amplicon" or "amplicon molecule" refers to a nucleic acid molecule produced by transcription of a template nucleic acid molecule or a nucleic acid molecule having a sequence complementary thereto, or a double-stranded nucleic acid comprising any such nucleic acid molecule. Transcription can be initiated from a primer.
[0066] Amplification: As used herein, the term "amplification" refers to the use of a template nucleic acid molecule in combination with various reagents to produce additional nucleic acid molecules from the template nucleic acid molecule, which may be identical or similar (e.g., at least 70% identical, such as at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identical) to a segment of the template nucleic acid molecule and / or a sequence complementary thereto. Amplification generally refers to the production of multiple copies of a specific nucleic acid portion, typically starting from a small amount of polynucleotide. This should be distinguished from non-specific template replication (e.g., replication that is template-dependent but not specific). Template specificity is often described as "target" specificity. Target sequences are "targets" in the sense that they are to be sorted out from other nucleic acids. Amplification techniques, such as, but not limited to, polymerase chain reaction (PCR), are primarily designed for such sorting.
[0067] Amplification reaction mixture: As used herein, the term "amplification reaction mixture" or "amplification reaction" refers to a template nucleic acid molecule and reagents sufficient for amplifying the template nucleic acid molecule.
[0068] Biological pathway: As used herein, the term "biological pathway" refers to a series of interactions or changes in a cell that lead to the production of a specific product or alteration of cell fate. Each pathway can integrate with other pathways or, independently, lead to the production of a specific molecule, the switching of gene activity, or the directing of cells to a specific location. Here, this series of interactions is used to represent the relationship between aberrantly methylated regions and their biological functions, determining their clinical and diagnostic value.
[0069] Biological sample: As described herein, the term "biological sample" as used herein generally refers to a sample obtained or derived from a biological source of interest (e.g., a tissue or organism or cell culture). In some embodiments, the source of interest includes an organism, such as an animal or a human. In some embodiments, a biological sample is or includes a biological tissue or fluid. In some embodiments, a biological sample can be or include bone marrow; blood; blood cells; ascites; a tissue or fine needle biopsy sample; a body fluid containing cells; free floating nucleic acids; sputum; saliva; urine; cerebrospinal fluid, peritoneal fluid; pleural fluid; feces; lymph; gynecological fluid; skin swab; vaginal swab; oral swab; nasal swab; a wash or lavage fluid (e.g., duct lavage fluid or bronchoalveolar lavage fluid); an aspirate; a scraping; a bone marrow specimen; a tissue biopsy specimen; a surgical specimen; feces, other body fluids, secretions, and / or excretions; and / or cells therein, etc. In some embodiments, a biological sample is or includes cells obtained from an individual. In some embodiments, the cell obtained is or includes cells from an individual from which a sample is obtained. In some embodiments, the sample is a "raw sample" obtained directly from the source of interest by any appropriate means. For example, in some embodiments, the original biological sample is obtained by a method selected from the group consisting of biopsy (e.g., fine needle aspiration or tissue biopsy), surgery, body fluid collection (e.g., blood, lymph, feces, etc.). In some embodiments, as the context indicates, the term "sample" refers to a preparation obtained by processing the original sample (e.g., by removing one or more components and / or by adding one or more agents). For example, a semipermeable membrane is used for filtration. This "processed sample" may include, for example, nucleic acids or proteins extracted from the sample, or nucleic acids or proteins obtained by techniques such as mRNA amplification or reverse transcription, separation and / or purification of some components to the original sample.
[0070] Biomarker: The term "biomarker" as used herein is consistent with its usage in the art and refers to an entity whose presence, level or form is associated with a specific biological event or state of interest and is therefore considered a "marker" for that event or state. It will be understood by those skilled in the art that, for example, in the case of a DNA biomarker, a biomarker can be or include a locus (e.g., one or more methylated loci) and / or the state of a locus (e.g., the state of one or more methylated loci). To name just a few examples of biomarkers, in some embodiments, such as those described herein, a biomarker can be or include a marker for a specific disease, disorder, or condition, or a marker for, for example, a quantitative or qualitative probability that a subject will develop, present, or recur a specific disease, disorder, or condition. In some embodiments, such as those described herein, a biomarker can be or include a marker for a specific treatment outcome, or a marker for a quantitative or qualitative probability thereof. Thus, in various embodiments, such as those described herein, a biomarker can be a prediction, prognosis, and / or diagnosis of a related biological event or state of interest. A biomarker can be an entity of any chemical class. For example, in some embodiments, such as described herein, a biomarker can be or include nucleic acids, polypeptides, lipids, carbohydrates, small molecules, inorganic substances (such as metals or ions) or combinations thereof. In some embodiments, such as described herein, a biomarker is a cell surface marker. In some embodiments, such as described herein, a biomarker is an intracellular marker. In some embodiments, such as described herein, a biomarker is present in an extracellular space (such as secreted or otherwise produced or present in an extracellular space, such as in body fluids such as blood, urine, tears, saliva, cerebrospinal fluid). In some embodiments, such as described herein, a biomarker is a methylation state of a methylation locus. In some instances, such as described herein, a biomarker may be referred to as a "marker".
[0071] As just one example of a biomarker, in some embodiments, such as described herein, the term refers to the expression of a product encoded by a gene, whose expression is characteristic of a particular tumor, tumor subtype, tumor stage, etc. Alternatively, in some embodiments, such as described herein, the presence or level of a particular marker can be correlated with the activity (or activity level) of a particular signaling pathway (e.g., a signaling pathway whose activity is characteristic of a particular tumor type).
[0072] Those skilled in the art will appreciate that a biomarker can alone determine a specific biological event or state of interest, or can represent or contribute to determining the statistical probability of a specific biological event or state of interest. Those skilled in the art will appreciate that the specificity and / or sensitivity of a marker may vary depending on the specific biological event or state of interest.
[0073] Biomolecule: As used herein, "biomolecule" refers to a molecule with biological activity, diagnostic properties, and prophylactic properties. Biomolecules useful in the present invention include, but are not limited to, synthetic, recombinant, or isolated peptides and proteins, such as antibodies and antigens, receptor ligands, enzymes, and adhesion peptides; nucleotides and polynucleic acids, such as DNA and antisense nucleic acid molecules; activated sugars and polysaccharides; bacteria; viruses; and chemical drugs, such as antibiotics, anti-inflammatory drugs, and antifungal agents.
[0074] Bisulfite reagent: As used herein, the term "bisulfite reagent" refers to a reagent comprising bisulfite, disulfite, hydrogen sulfite, or a combination thereof, for distinguishing between methylated and unmethylated cytosine, such as cytosine in a CpG dinucleotide sequence.
[0075] Blood components: As used herein, the term "blood components" refers to any component of whole blood, including red blood cells, white blood cells, plasma, platelets, endothelial cells, mesothelial cells, epithelial cells, and cell-free DNA. Blood components also include components of plasma, including proteins, metabolites, lipids, nucleic acids, and carbohydrates, as well as any other cells that may be present in the blood due to pregnancy, organ transplantation, infection, injury, or disease.
[0076] Cancer: As used herein, the terms "cancer," "malignancy," "neoplasm," "tumor," and "cancer" are used interchangeably to refer to a disease, disorder, or condition in which cells exhibit or have exhibited relatively abnormal, uncontrolled, and / or spontaneous growth, such that they display or have displayed an abnormally elevated proliferation rate and / or an abnormal growth phenotype. In some embodiments, such as described herein, a cancer can include one or more tumors. In some embodiments, such as described herein, a cancer can be or include precancerous (e.g., benign), malignant, premetastatic, metastatic, and / or non-metastatic cells. In some embodiments, such as described herein, a cancer can be or include a solid tumor. In some embodiments, such as described herein, a cancer can be or include a hematological tumor. In general, examples of different types of cancer known in the art include, for example, colorectal cancer, hematopoietic cancers (including leukemias, lymphomas (Hodgkin's and non-Hodgkin's)), myelomas and myeloproliferative disorders; sarcomas, melanomas, adenomas, solid tissue cancers, squamous cell carcinoma of the oral cavity, pharynx, larynx and lung, liver cancer, genitourinary cancers (e.g., prostate cancer, cervical cancer, bladder cancer, uterine cancer and endometrial cancer) and renal cell carcinoma, bone cancer, colorectal cancer, skin cancer, cutaneous or intraocular melanoma, cancers of the endocrine system, thyroid cancer, parathyroid cancer, head and neck cancer, breast cancer, gastrointestinal cancer and nervous system cancers, benign lesions (e.g., papilloma), etc.
[0077] Chemotherapeutic agent: As used herein, the term "chemotherapeutic agent" refers to one or more agents known or having known characteristics that treat or aid in the treatment of cancer. In particular, chemotherapeutic agents include pro-apoptotic agents, cytostatic agents, and / or cytotoxic agents. In some embodiments, such as described herein, the chemotherapeutic agent can be or include an alkylating agent, an anthracycline, a cytoskeletal disruptor (e.g., a microtubule targeting moiety, such as a taxane, maytansine, and its analogs), an epothilone, a histone deacetylase inhibitor (HDACs), a topoisomerase inhibitor (e.g., a topoisomerase I and / or topoisomerase II inhibitor), a kinase inhibitor, a nucleotide analog or nucleotide precursor analog, a polypeptide antibiotic, a platinum agent, retinoic acid, a vinca alkaloid, and / or an analog thereof having a related anti-proliferative activity. In some specific embodiments, such as those described herein, the chemotherapeutic agent can be or include actinomycin, all-trans retinoic acid, auristatin, azacitidine, azathioprine, bleomycin, bortezomib, carboplatin, capecitabine, cisplatin, chlorambucil, cyclophosphamide, curcumin, cytarabine, daunorubicin, docetaxel, doxifluridine, doxorubicin, epirubicin, epothilones, etoposide, fluorouracil, gemcitabine, hydroxyurea, idarubicin, imatinib, irinotecan, maytansine and / or its analogs (e.g., DM1), nitrogen mustard, mercaptopurine, methotrexate, mitoxantrone, maytansines, oxaliplatin, paclitaxel, pemetrexed, teniposide, thioguanine, topotecan, valrubicin, vinblastine, vincristine, vindesine, vinorelbine, or a combination thereof. In some embodiments, chemotherapeutic drugs may be used in the context of antibody-drug conjugates, such as those described herein.In some embodiments, e.g., as described herein, the chemotherapeutic agent is one present in an antibody-drug conjugate selected from the group consisting of hLL1-doxorubicin, hRS7-SN-38, hMN-14-SN-38, hLL2-SN-38, hA20-SN-38, hPAM4-SN-38, hLL1-SN-38, hRS7-Pro-2-P-Dox, hMN-14-Pro-2-P-Dox, hLL2-Pro-2-P-Dox, hA20-Pro-2-P-Dox, hPAM4-Pro-2-P-Dox, hLL1-Pro-2-P-Dox, P4 / D10-doxorubicin, gemtuzumab ozogamicin, brentuximab vedotin, trastuzumab emtansine, and trastuzumab vedotin. emtansine), inotuzumab ozogamicin, glembatumomab vedotin, SAR3419, SAR566658, BIIB015, BT062, SGN-75, SGN-CD19A, AMG-172, AMG-595, BAY-94-9343, ASG-5ME, ASG-22ME, ASG-16M8F, MDX-1203, MLN-0264, anti-PSMA ADC, RG-7450, RG-7458, RG-7593, RG-7596, RG-7598, RG-7599, RG-7600, RG-7636, ABT-414, IMGN-853, IMGN-529, vorsetuzumab In some embodiments, such as those described herein, the chemotherapeutic agent can be or include farnesyl-thiosalicylic acid (FTS), 4-(4-chloro-2-methylphenoxy)-N-hydroxybutyramide (CMH), estradiol (E2), tetramethoxystilbene (TMS), delta-tocotrienol, salinomycin, or curcumin.
[0078] Combination therapy: As used herein, the term "combination therapy" refers to the administration of two or more agents or treatment regimens to a subject such that the two or more agents or treatment regimens collectively treat the subject's disease, condition, or disorder. In some embodiments, such as those described herein, the two or more therapeutic agents or treatment regimens may be administered in simultaneous, sequential, or overlapping dosage regimens. One skilled in the art will understand that combination therapy includes, but does not require, the administration of the two agents or treatment regimens together in a single composition, nor does it require simultaneous administration.
[0079] Comparability: As used herein, the term "comparable" refers to members of a group of two or more conditions, settings, agents, entities, populations, etc. that may not be identical to one another but are sufficiently similar to permit comparisons between them such that one skilled in the art will understand that conclusions can reasonably be drawn based on observed differences or similarities. In some embodiments, such as those described herein, groups of comparable conditions, settings, agents, entities, populations, etc. are generally characterized by having a plurality of substantially identical features and zero, one, or more different features. One skilled in the art will understand in the context what degree of identity is required for members of a group to be comparable. For example, one skilled in the art will understand that members of a group of conditions, settings, agents, entities, populations, etc. are comparable to one another when they are characterized by a sufficient number and type of substantially identical features to warrant a reasonable conclusion that observed differences can be attributed, in whole or in part, to non-identical features therein.
[0080] Corresponding to: As used herein, the term "corresponding to" refers to a relationship between two or more entities. For example, the term "corresponding to" can be used to specify the position / property of a structural element in a compound or composition relative to another compound or composition (e.g., relative to an appropriate reference compound or composition). For example, in some embodiments, a monomer residue in a polymer (e.g., a nucleic acid residue in a polynucleotide) can be identified as "corresponding to" a residue in an appropriate reference polymer. One skilled in the art will readily understand how to identify "corresponding" nucleic acids. For example, one skilled in the art is aware of various sequence alignment strategies, including software programs such as BLAST, CS-BLAST, CUSASW++, DIAMOND, FASTA, GGSEARCH / GLSEARCH, Genoogle, HMMER, HHpred / HHsearch, IDF, Infernal, KLAST, USEARCH, parasail, PSI-BLAST, PSI-Search, ScalaBLAST, Sequilab, SAM, SSEARCH, SWAPHI, SWAPHI-LS, SWIMM, or SWIPE, which can be used to identify "corresponding" residues in the nucleic acids of the present invention. Those skilled in the art will also understand that, in some instances, the term "corresponding to" can be used to describe an event or entity that has a relevant similarity to another event or entity (e.g., an appropriate reference event or entity). As just one example, a DNA fragment in a sample from a subject can be described as "corresponding to" a gene to indicate that, in some embodiments, it shows a particular degree of sequence identity or homology, or has particular characteristic sequence elements.
[0081] Detectable moiety: As used herein, the term "detectable moiety" refers to any detectable element, molecule, functional group, compound, fragment, or other moiety. In some embodiments, such as described herein, a detectable molecule is provided or used alone. In some embodiments, such as described herein, a detectable molecule is provided and / or used in conjunction with (e.g., linked to) another agent. Examples of detectable molecules include, but are not limited to, various ligands, radionuclides (e.g., 3 H. 14 C. 18 F. 19 F. 32 P. 35 S. 135 I. 125 I. 123 I. 64 Cu, 187 Re、 111 In, 90 Y. 99m Tc, 177 Lu, 89 Zr, etc. ), fluorescent dyes, chemiluminescent agents, bioluminescent agents, spectrally resolvable inorganic fluorescent semiconductor nanocrystals (i.e., quantum dots), metal nanoparticles, nanoclusters, paramagnetic metal ions, enzymes, colorimetric labels, biotin, dioxigenin, haptens, and proteins for which antisera or monoclonal antibodies are available.
[0082] Diagnosis: As used herein, the term "diagnosis" refers to determining whether a subject has or will have a disease, disorder, condition, or state and / or the qualitative and quantitative probability thereof. For example, in the diagnosis of cancer, diagnosis can include determining the risk, type, stage, malignancy, or other classification of the cancer. In some instances, such as those described herein, diagnosis can be or include determining the prognosis and / or likely response to treatment involving one or more conventional or specific therapeutic agents or treatment regimens.
[0083] Diagnostic Information: As used herein, the terms "diagnostic information" or "information used for diagnosis" refer to any type of information that helps determine whether a patient has a disease, disorder, or condition and / or helps classify a disease, disorder, or condition into a phenotypic category, or that is important for the prognosis of the disease, disorder, or condition or the likely response to treatment (conventional or any specific treatment) of the disease, disorder, or condition. Similarly, "diagnosis" refers to providing any type of diagnostic information, including but not limited to whether a subject is likely to have or has developed a disease, disorder, or condition, the state, stage, or characteristics of the disease, disorder, or condition manifested in a subject, information regarding the nature or classification of a tumor, information regarding prognosis, and / or information useful for selecting an appropriate treatment. Treatment selection may include selecting a specific therapeutic agent or other treatment modality, such as surgery, radiation, etc., selecting whether to withhold or provide treatment, selecting information regarding a dosage regimen (e.g., the frequency or level of one or more doses of a specific therapeutic agent or combination of therapeutic agents), etc. Diagnostic information may include, but is not limited to, biomarker status information.
[0084] Differential Methylation: As used herein, the term "differentially methylated" refers to a methylation site whose methylation status differs between a first instance and a second instance. A differentially methylated methylation site may be referred to as a differentially methylated site. In some instances, such as those described herein, a DMR is defined by an amplicon generated by amplification using oligonucleotide primers. For example, a pair of oligonucleotide primers is selected to amplify the DMR or amplify a DNA region of interest present in the amplicon. In some instances, such as those described herein, a DMR is defined as a DNA region amplified by a pair of oligonucleotide primers that includes a region having the sequence of the oligonucleotide primers or a sequence complementary to the oligonucleotide primers. In some instances, such as those described herein, a DMR is defined as a DNA region amplified by a pair of oligonucleotide primers that excludes a region having the sequence of the oligonucleotide primers or a sequence complementary to the oligonucleotide primers. As used herein, a specific DMR may be specifically identified by the name of the associated gene followed by a three-digit starting position. For example, a DMR starting at position 100785927 of ZAN may be identified as ZAN'927. As used herein, a specifically provided DMR can be unambiguously identified by the chromosome number followed by the start and end positions of the DMR.
[0085] Differentially methylated region: As used herein, the term "differentially methylated region" (DMR) refers to a region of DNA comprising one or more differentially methylated sites. Under selected conditions of interest (e.g., a cancerous state), a DMR comprising a greater number or a higher frequency of methylated sites may be referred to as a hypermethylated DMR. Under selected conditions of interest (e.g., a cancerous state), a DMR comprising a smaller number or a lower frequency of methylated sites may be referred to as a hypomethylated DMR. A DMR that is a methylation biomarker for colorectal cancer may be referred to as a colorectal cancer DMR. A DMR that is a methylation biomarker for advanced adenoma may be referred to as an advanced adenoma DMR. In some instances, such as those described herein, a DMR may be a single nucleotide that is a methylation site. In some instances, such as those described herein, a DMR may be at least 10, at least 15, at least 20, at least 30, at least 50, or at least 75 base pairs in length. In some examples, such as those described herein, the length of the DMR is equal to or less than 5000bp, 4000bp, 3000bp, 2000bp, 1000bp, 950bp, 900bp, 850bp, 800bp, 750bp, 700bp, 650bp, 600bp, 550bp, 500bp, 450bp, 400bp, 350bp, 300bp, 250bp, 200bp, 150bp, 100bp, 50bp, 40bp, 30bp, 20bp, or 10bp (e.g., wherein the methylation status is determined using quantitative polymerase chain reaction (qPCR), such as methylation-sensitive restriction enzyme quantitative polymerase chain reaction (MSRE-qPCR)) (e.g., wherein the methylation status is determined using next-generation sequencing technology, such as targeted next-generation sequencing). In some examples, such as those described herein, a DMR that is a methylation biomarker for advanced adenoma can also be used to identify colorectal cancer, and vice versa.
[0086] DNA region: As used herein, "DNA region" refers to any contiguous portion of a larger DNA molecule. Those skilled in the art are familiar with techniques for determining whether a first DNA region and a second DNA region correspond based on sequence similarity (e.g., sequence identity or homology) and / or context (e.g., sequence identity or homology of nucleic acids upstream and / or downstream of the first and second DNA regions) of the first and second DNA regions.
[0087] Unless otherwise indicated herein, sequences found in or associated with humans (e.g., sequences that hybridize to human DNA) are sequences found in, based on, and / or derived from the exemplary representative human genome sequence commonly referred to and familiar to those skilled in the art as Homo sapiens (human) genome assembly GRCh38, hg38, and / or Reference Genome Consortium Human Construct 38. Those skilled in the art will further understand that DNA regions of hg38 can be referred to by known systems that include identifying specific nucleotide positions or ranges thereof by designated numbering.
[0088] Dosage regimen: As used herein, the term "dosage regimen" may refer to a group of one or more identical or different unit doses administered to a subject, typically comprising a plurality of unit doses, the administration of each unit dose being separated by a period of time from the administration of the other unit doses. In various embodiments, such as described herein, one or more or all of the unit doses in a dosage regimen may be the same, or may be different (e.g., increased over time, decreased over time, or adjusted based on the decision of the subject and / or physician). In various embodiments, such as described herein, one or more or all of the time periods between each administration may be the same, or may be different (e.g., increased over time, decreased over time, or adjusted based on the decision of the subject and / or physician). In some embodiments, such as described herein, a given therapeutic agent has a recommended dosage regimen, which may involve one or more doses. Typically, at least one recommended dosage regimen for a marketed drug is known to those skilled in the art. In some embodiments, such as described herein, a dosage regimen is associated with a desired or beneficial outcome when administered in a relevant population (i.e., is a therapeutic dosage regimen).
[0089] Downstream: The term "downstream" as used herein means that the first DNA region is closer to the C-terminus of the nucleic acid comprising the first DNA region and the second DNA region relative to the second DNA region.
[0090] Early stage: As used herein, the term "early stage" refers to a localized stage where the cancer has not spread to nearby lymph nodes (N0) or distant sites (M0). For example, pathologically, this is a stage 0 to stage IIC cancer.
[0091] Gene: As used herein, the term "gene" refers to a single DNA region, such as a DNA region in a chromosome, which comprises a coding sequence for a coding product (e.g., an RNA product and / or a polypeptide product), and together comprises all, part, or no DNA sequence that contributes to regulating the expression of the coding sequence. In some embodiments, such as described herein, a gene comprises one or more non-coding sequences. In some specific embodiments, such as described herein, a gene comprises exon sequences and intron sequences. In some embodiments, such as described herein, a gene comprises one or more regulatory elements, such as, for example, elements that can control or influence one or more aspects of gene expression (e.g., cell type-specific expression, inducible expression, etc.). In some embodiments, such as described herein, a gene comprises a promoter. In some embodiments, such as described herein, a gene comprises one or both of (i) DNA nucleotides extending a predetermined number of nucleotides upstream of the coding sequence and (ii) DNA nucleotides extending a predetermined number of nucleotides downstream of the coding sequence. In various embodiments, such as described herein, the predetermined number of nucleotides can be 500 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 75 kb, or 100 kb.
[0092] Gene Ontology: The "Gene Ontology" (GO) is a set of concepts that describe the functions of gene products within an organism. It represents computational aspects of biological systems. GO has distinct categories, two of which were used in this paper: biological process and molecular function. Within each category, concepts related to gene products in different biological systems were analyzed. These concepts were used to demonstrate the relationship between aberrantly methylated regions and their biological functions, thereby determining their clinical and diagnostic value.
[0093] Homology: As used herein, the term "homology" refers to the overall relatedness between polymer molecules, such as between nucleic acid molecules (e.g., DNA molecules and / or RNA molecules) and / or between polypeptide molecules. One skilled in the art will appreciate that homology can be defined, for example, by percent identity or percent homology (sequence similarity). In some embodiments, such as described herein, polymer molecules are considered "homologous" to one another if their sequences are at least 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical. In some embodiments, such as described herein, polymer molecules are considered "homologous" to each other if their sequences are at least 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% similar.
[0094] Hybridization: As used herein, "hybridization" refers to the binding of a first nucleic acid to a second nucleic acid to form a double-stranded structure, which is achieved through complementary pairing of nucleotides. Those skilled in the art will understand that complementary sequences, etc. can hybridize. In various embodiments, such as described herein, hybridization can occur between nucleotide sequences having at least 70% complementarity, such as at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% complementarity. Those skilled in the art will further recognize that whether hybridization occurs between a first nucleic acid and a second nucleic acid can depend on various reaction conditions. The conditions under which hybridization occurs are known in the art.
[0095] Hypomethylation: As used herein, the term "hypomethylation" refers to a state of a methylated locus in which the methylated nucleotide in the state of interest is at least one less than that in a reference state (e.g., colorectal cancer has at least one less methylated nucleotide than a healthy control).
[0096] Hypermethylation: The term "hypermethylation" as used herein refers to a state of a methylated locus in which the state of interest has at least one more methylated nucleotide than a reference state (e.g., colorectal cancer has at least one more methylated nucleotide than a healthy control).
[0097] Identity, identical: The terms "identity" and "identical" as used herein refer to the overall correlation between polymer molecules (e.g., between nucleic acid molecules (e.g., DNA molecules and / or RNA molecules) and / or between polypeptide molecules). Methods for calculating the percent identity between two given sequences are known in the art. For example, the percent identity of two nucleic acid or polypeptide sequences can be calculated, for example, by aligning the two sequences (or the complements of one or both sequences) for optimal alignment purposes (e.g., gaps can be introduced into one or both of the first and second sequences for optimal alignment purposes, and non-identical sequences can be ignored for alignment purposes). The nucleotides or amino acids at corresponding positions are then compared. When a position in the first sequence is occupied by a residue (e.g., a nucleotide or amino acid) that is identical to the residue at the corresponding position in the second sequence, the molecules at that position are identical. The percent identity between the two sequences is related to the number of identical positions shared by the sequences, and optionally takes into account the number of gaps and the length of each gap, which may need to be introduced when the two sequences are optimally aligned. Comparing sequences and determining percent identity between two sequences can be accomplished using a computational algorithm, such as BLAST (Basic Local Alignment Search Tool).
[0098] "Improve," "increase," or "decrease": These terms, or grammatically comparable comparative terms, as used herein, refer to values relative to comparable reference measurements. For example, in some embodiments, such as described herein, an assessment achieved using an agent of interest can be "improved" relative to an assessment obtained using a comparable reference agent or without the agent. Alternatively or additionally, in some embodiments, such as described herein, an assessment of a subject or system of interest can be "improved" relative to an assessment obtained for the same subject or system under different conditions or at different time points (e.g., before or after an event such as administration of an agent of interest), or in different comparable subjects (e.g., in a comparable subject or system different from the subject or system of interest, where one or more indicators of the particular disease, disorder, or condition of interest of the subject or system differ, or where prior exposure to a condition or agent differs). In some embodiments, such as described herein, comparative terms refer to statistically relevant differences (e.g., differences that are significant and / or of sufficient magnitude to be statistically relevant). One skilled in the art will recognize or readily determine what degree and / or significance of difference is necessary or sufficient to achieve such statistical significance in a given situation.
[0099] The Kaplan-Meier estimator, or product limit estimator, is a statistical test that measures the probability of an event occurring at a specific point in time. The probability at a specific point in time is multiplied by the previous probability to produce the result. In this article, it is used to estimate the survival function from lifespan data.
[0100] A nucleic acid locus refers to a subregion of a nucleic acid, such as a CpG island, a single nucleotide, a gene on a chromosome, etc.
[0101] Marker: As used herein, a marker refers to an entity or moiety whose presence or level is characteristic of a particular state or event. In some embodiments, the presence or level of a particular marker may be characteristic of the presence or staging of a disease, disorder, or condition. By way of example only, in some embodiments, the term refers to a gene expression product that is characteristic of a particular tumor, tumor subtype, tumor stage, or the like. Alternatively or additionally, in some embodiments, the presence or level of a particular marker is associated with the activity (or level of activity) of a particular signaling pathway, which may be characteristic of a particular type of tumor, for example. The statistical significance of the presence or absence of a marker may vary depending on the particular marker. In some embodiments, detection of a marker is highly specific, reflecting that the tumor is likely to be of a particular subtype. This specificity may come at the expense of sensitivity (i.e., a negative result may occur even if the tumor is one that would be expected to express the marker).
[0102] Conversely, a marker with high sensitivity may have lower specificity than a marker with low sensitivity. According to the present invention, a useful marker does not have to distinguish a particular subtype of tumor with 100% accuracy. The marker can be a metabolite, a lipid, a fatty acid, and / or a polyunsaturated fatty acid. In some embodiments, the term marker can refer to a ratio of two entities (e.g., moieties).
[0103] Methylation: As used herein, the term "methylation" includes methylation at any of the following positions: (i) the C5 position of cytosine; (ii) the N4 position of cytosine; and (iii) the N6 position of adenine. Methylation also includes (iv) other types of nucleotide methylation. Methylated nucleotides may be referred to as "methylated nucleotides" or "methylated nucleotide bases." Thus, "methylated nucleotides" or "methylated nucleotide bases" refer to the presence of a methyl moiety on a nucleotide base that is not present in recognized typical nucleotide bases. In some embodiments, such as those described herein, methylation specifically refers to the methylation of cytosine residues. In some instances, methylation specifically refers to the methylation of cytosine residues present at CpG sites.
[0104] Methylation detection: As used herein, the term "methylation detection" refers to any technique that can be used to determine the methylation status of a methylated locus. For example, methylation detection can refer to a technique that determines the methylation status of one or more CpG dinucleotide sequences in a nucleic acid sequence.
[0105] Methylation biomarker: As used herein, the term "methylation biomarker" refers to a biomarker that is characterized by a change in the methylation state of one or more nucleic acid loci between a first state and a second state (e.g., a cancerous state and a non-cancerous state) and / or a methylation state of at least one methylated locus.
[0106] Methylation locus: As used herein, the term "methylation locus" refers to a region of DNA that contains at least one differentially methylated region. Under a selected condition of interest (e.g., a cancerous state), a methylation locus that contains a greater number or a higher frequency of methylation sites may be referred to as a hypermethylated locus. Under a selected condition of interest (e.g., a cancerous state), a methylation locus that contains a smaller number or a lower frequency of methylation sites may be referred to as a hypomethylated locus. In some examples, such as described herein, the methylation locus is at least 10, at least 15, at least 20, at least 30, at least 50, or at least 75 base pairs in length. In some examples, e.g., as described herein, the length of the methylated locus is less than 5000bp, 4000bp, 3000bp, 2000bp, 1000bp, 950bp, 900bp, 850bp, 800bp, 750bp, 700bp, 650bp, 600bp, 550bp, 500bp, 450bp, 400bp, 350bp, 300bp, 250bp, 200bp, 150bp, 100bp, 50bp, 40bp, 30bp, 20bp, or 10bp (e.g., wherein the methylation status is determined using quantitative polymerase chain reaction (qPCR), e.g., methylation-sensitive restriction enzyme quantitative polymerase chain reaction (MSRE-qPCR)).
[0107] Methylation site: As used herein, a methylation site refers to a nucleotide or nucleotide position that is methylated in at least one instance. In the methylated state, a methylation site can be referred to as a methylated site.
[0108] The term "methylation-specific restriction enzyme" or "methylation-sensitive restriction enzyme" refers to an enzyme that selectively digests nucleic acids based on the methylation status of the recognition site.
[0109] Methylation state: As used herein, "methylation profile," "methylation state," or "methylation signature" refers to the amount, frequency, or pattern of methylation at methylation sites within a methylation locus. Thus, a change in methylation state between a first state and a second state can be or include an increase in the amount, frequency, or pattern of methylation sites, or can be or include a decrease in the amount, frequency, or pattern of methylation sites. In various examples, a change in methylation state is a change in methylation value.
[0110] Methylation value: As used herein, the term "methylation value" refers to a numerical representation of the methylation state, for example, a numerical representation of the methylation frequency or ratio of a methylated locus. In some instances, such as those described herein, a methylation value can be generated by a method comprising quantifying the amount of intact nucleic acid present in a sample following restriction digestion of the sample with a methylation-dependent restriction enzyme. In some instances, such as those described herein, a methylation value can be generated by a method comprising comparing the amplification characteristics of the sample following a bisulfite reaction. In some instances, such as those described herein, a methylation value can be generated by comparing bisulfite-treated and untreated nucleic acid sequences. In some instances, such as those described herein, a methylation value comprises or is based on quantitative PCR results.
[0111] The methylation state of a specific nucleic acid sequence (e.g., a genetic marker or DNA region described herein) can represent the methylation state of each base in the sequence, with or without providing precise information about the location of methylation in the sequence, or can represent the methylation state of a subset of bases within the sequence (e.g., one or more cytosines), or can represent information related to regional methylation density within the sequence.
[0112] The terms "methylation status", "methylation state" and "methylation signature" also refer to the relative concentration, absolute concentration or pattern of methylated C or unmethylated C in any particular region of a nucleic acid in a biological sample. For example, if the cytosine (C) residues in a nucleic acid sequence are methylated, it can be referred to as "hypermethylation" or "increased methylation", while if the cytosine (C) residues in a DNA sequence are not methylated, it can be referred to as "hypomethylation" or "reduced methylation". Similarly, if the cytosine (C) residues in one nucleic acid sequence are methylated compared to another nucleic acid sequence (e.g., from a different region or a different individual, etc.), the sequence is considered to be hypermethylated, or to have increased methylation, compared to the other nucleic acid sequence. Alternatively, if the cytosine (C) residues in a DNA sequence are not methylated compared to another nucleic acid sequence (e.g., from a different region or a different individual, etc.), the sequence is considered to be hypomethylated, or to have decreased methylation, compared to the other nucleic acid sequence.
[0113] When sequences differ in degree (e.g., increased or decreased methylation of one sequence relative to another), frequency, or pattern of methylation, they are said to be "differentially methylated" or to have a "methylation difference" or to have a "different methylation state." The term "methylation difference" can refer to a difference in the level or pattern of nucleic acid methylation in a cancer-positive sample compared to the level or pattern of nucleic acid methylation in a cancer-negative sample.
[0114] Mutation: As used herein, the term "mutation" refers to a genetic variation in a biomolecule (e.g., a nucleic acid or protein) compared to a reference biomolecule. For example, in some embodiments, a mutation in a nucleic acid can include a base substitution, a deletion of one or more bases, an insertion of one or more bases, an inversion of two or more bases, or a truncation compared to a reference nucleic acid molecule. Similarly, a mutation in a protein can include an amino acid substitution, insertion, inversion, or truncation compared to a reference polypeptide. Other mutations, such as fusions and indels, are also known to those skilled in the art. In some embodiments, a mutation includes a genetic variation associated with loss of function of a gene product. Loss of function can be a complete loss of function, such as loss of enzymatic activity of an enzyme, or a partial loss of function, such as reduced enzymatic activity of an enzyme. In some embodiments, a mutant includes a genetic variation associated with a gain of function, such as a genetic variation associated with a negative or undesirable change in the characteristic or activity of a gene product. In some embodiments, a mutant is characterized by a reduction or loss of a desired level or activity compared to a reference; in some embodiments, a mutant is characterized by an increase or gain of an undesirable level or activity compared to a reference. In some embodiments, the reference biomolecule is a wild-type biomolecule.
[0115] Nucleic acid: As used herein, the term "nucleic acid" in its broadest sense refers to any compound and / or substance that is or can be incorporated into an oligonucleotide chain. In some embodiments, such as described herein, nucleic acid refers to a compound and / or substance that is or can be incorporated into an oligonucleotide chain via a phosphodiester linkage. As will be appreciated from the context, in some embodiments, such as described herein, the term nucleic acid refers to a single nucleic acid residue (e.g., a nucleotide and / or nucleoside), and in some embodiments, such as described herein, the term nucleic acid refers to a polynucleotide chain composed of a plurality of single nucleic acid residues. Nucleic acid can be or include DNA, RNA, or a combination thereof. Nucleic acid can comprise natural nucleic acid residues, nucleic acid analogs, and / or synthetic residues. In some embodiments, such as described herein, nucleic acid comprises natural nucleotides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine). In some embodiments, e.g., as described herein, a nucleic acid is or includes one or more nucleotide analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrole-pyrimidine, 3-methyladenosine, 5-methylcytidine, C-5 propynylcytidine, C-5 propynyluridine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyluridine, C5-propynylcytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanosine, 2-thiocytidine, methylated bases, intercalated bases, and combinations thereof).
[0116] In some embodiments, such as described herein, nucleic acid has a nucleotide sequence encoding a functional gene product (such as RNA or protein). In some embodiments, such as described herein, nucleic acid includes one or more introns. In some embodiments, such as described herein, nucleic acid includes one or more genes. In some embodiments, such as described herein, nucleic acid is prepared by one or more of the following methods: separation from natural sources, enzymatic polymerization synthesis based on complementary templates (in vivo or in vitro), replication in recombinant cells or systems, and chemical synthesis.
[0117] In some embodiments, such as described herein, nucleic acid analogs are different from nucleic acid in that they do not use a phosphodiester backbone. For example, in some embodiments, such as described herein, nucleic acid can include one or more peptide nucleic acids known in the art, having a peptide bond rather than a phosphodiester bond in its backbone. Alternatively or additionally, in some embodiments, such as described herein, nucleic acid has one or more phosphorothioate and / or 5'-N-phosphoamide connections rather than a phosphodiester bond. In some embodiments, such as described herein, nucleic acid comprises one or more modified sugars (such as 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose) compared to the sugar in natural nucleic acids.
[0118] In some embodiments, e.g., as described herein, a nucleic acid is or includes at least 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600 In some embodiments, the nucleic acid comprises at least 0, 180, 190, 20, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000 or more residues. In some embodiments, such as those described herein, the nucleic acid is partially or fully single-stranded or partially or fully double-stranded.
[0119] Nucleic acid detection assay: As used herein, the term "nucleic acid detection assay" refers to any method for determining the nucleotide composition of a nucleic acid of interest. Nucleic acid detection methods include, but are not limited to, DNA sequencing methods (e.g., next-generation sequencing methods), polymerase chain reaction-based methods, probe hybridization methods, ligase chain reaction, and the like.
[0120] Nucleotide: As used herein, the term "nucleotide" refers to a structural component or building block of a polynucleotide (e.g., a DNA and / or RNA polymer). A nucleotide comprises a base (e.g., adenine, thymine, uracil, guanine, or cytosine) and a molecule of sugar and at least one phosphate group. Nucleotides used herein can be methylated nucleotides or unmethylated nucleotides. It will be understood by those skilled in the art that nucleic acid terms (e.g., "locus" or "nucleotide") can refer to the locus or nucleotides of a single nucleic acid molecule and / or a cumulative population of loci or nucleotides within a plurality of nucleic acids (e.g., in a sample and / or representing a plurality of nucleic acids of a subject) representing a locus or nucleotide (e.g., having the same nucleic acid sequence and / or nucleic acid sequence context, or having substantially the same nucleic acid sequence and / or nucleic acid context).
[0121] Oligonucleotide primer: As used herein, the term oligonucleotide primer or primer refers to a nucleic acid molecule that is used, can be used, or is used to produce an amplicon from a template nucleic acid molecule. Under conditions that allow transcription (e.g., in the presence of nucleotides and a DNA polymerase, and at a suitable temperature and pH), an oligonucleotide primer can provide a point at which transcription begins from a template to which the oligonucleotide primer hybridizes. Typically, an oligonucleotide primer is a single-stranded nucleic acid having a length of 5 to 200 nucleotides. Those skilled in the art will appreciate that the optimal primer length for producing an amplicon from a template nucleic acid molecule can vary with changes in conditions including temperature parameters, primer composition, transcription or amplification methods, and the like. A pair of oligonucleotide primers used herein refers to a set of two oligonucleotide primers that are complementary to the first and second strands of a double-stranded template nucleic acid molecule, respectively. The first and second members of a pair of oligonucleotide primers can be referred to as a "forward" oligonucleotide primer and a "reverse" oligonucleotide primer, respectively, with respect to a template nucleic acid strand, wherein the forward oligonucleotide primer is capable of hybridizing to a nucleic acid strand complementary to the template nucleic acid strand and the reverse oligonucleotide primer is capable of hybridizing to the template nucleic acid strand, and the position of the forward oligonucleotide primer relative to the template nucleic acid strand is 5' of the position of the reverse oligonucleotide primer sequence relative to the template nucleic acid strand. Those skilled in the art will appreciate that identifying the first and second oligonucleotide primers as the forward oligonucleotide primer and the reverse oligonucleotide primer, respectively, is arbitrary, as these designations depend on whether a given nucleic acid strand or its complementary strand is used as a template nucleic acid molecule.
[0122] Overlap: As used herein, the term "overlap" refers to two DNA regions, each containing a subsequence that is substantially identical to a subsequence of the same length in the other region (e.g., the two DNA regions have a common subsequence). "Substantially identical" means that two subsequences of the same length differ by less than a given number of base pairs. In some examples, such as those described herein, each subsequence is at least 20 base pairs in length and differs from each other by less than 4, 3, 2, or 1 base pairs (e.g., the two subsequences have at least 80%, at least 85%, at least 90%, at least 95% similarity, at least 97% similarity, at least 98% similarity, at least 99% similarity, or at least 99.5% similarity). In some examples, such as those described herein, each subsequence is at least 24 base pairs in length and differs by less than 5, 4, 3, 2, or 1 base pairs (e.g., two subsequences are at least 80%, at least 85%, at least 90%, at least 95% similar, at least 97% similar, at least 98% similar, at least 99% similar, or at least 99.5% similar). In some examples, such as those described herein, each subsequence is at least 50 base pairs in length and differs by less than 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 base pairs (e.g., two subsequences are at least 80%, at least 85%, at least 90%, at least 95% similar, at least 97% similar, at least 98% similar, at least 99% similar, or at least 99.5% similar). In some examples, such as those described herein, each subsequence is at least 100 base pairs in length and differs by fewer than 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 base pairs (e.g., two subsequences are at least 80%, at least 85%, at least 90%, at least 95% similar, at least 97% similar, at least 98% similar, at least 99% similar, or at least 99.5% similar). In some examples, such as those described herein, each subsequence is at least 200 base pairs in length and differs by fewer than 40, 30, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 base pairs (e.g., two subsequences are at least 80%, at least 85%, at least 90%, at least 95% similar, at least 97% similar, at least 98% similar, at least 99% similar, or at least 99.5% similar).In some examples, such as those described herein, each subsequence is at least 250 base pairs in length and differs by fewer than 50, 40, 30, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 base pairs (e.g., two subsequences are at least 80%, at least 85%, at least 90%, at least 95% similar, at least 97% similar, at least 98% similar, at least 99% similar, or at least 99.5% similar). In some examples, such as those described herein, each subsequence is at least 300 base pairs in length and differs by fewer than 60, 50, 40, 30, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 base pairs (e.g., two subsequences are at least 80%, at least 85%, at least 90%, at least 95% similar, at least 97% similar, at least 98% similar, at least 99% similar, or at least 99.5% similar). In some examples, such as those described herein, each subsequence is at least 500 base pairs in length and differs by fewer than 100, 60, 50, 40, 30, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 base pairs (e.g., two subsequences are at least 80%, at least 85%, at least 90%, at least 95% similar, at least 97% similar, at least 98% similar, at least 99% similar, or at least 99.5% similar). In some examples, such as those described herein, each subsequence is at least 1000 base pairs in length and differs by less than 200, 100, 60, 50, 40, 30, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 base pair (e.g., two subsequences are at least 80%, at least 85%, at least 90%, at least 95% similar, at least 97% similar, at least 98% similar, at least 99% similar, or at least 99.5% similar). In some embodiments, such as those described herein, a subsequence of a first region of two DNA regions can include all of a second region of the two DNA regions (or vice versa) (e.g., a common subsequence can include all of either region or both regions).In certain embodiments, where a methylation locus has a sequence comprising "at least a portion" of a DMR sequence listed herein (e.g., at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the DMR sequence), the overlapping portion of the methylation locus has at least 95% similarity, at least 98% similarity, or at least 99% similarity to the overlapping portion of the DMR sequence (e.g., if the overlapping portion is 100 bp, the overlapping portion of the methylation locus differs from the DMR portion by no more than 1 bp, no more than 2 bp, or no more than 5 bp). In certain embodiments, when a methylation locus has a sequence comprising "at least a portion" of a DMR sequence listed herein, this means that the methylation locus and the DMR sequence have a common subsequence whose contiguous base sequence covers at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the DMR sequence, e.g., wherein the common subsequence differs by no more than 1 bp, no more than 2 bp, or no more than 5 bp. In certain embodiments, when a methylation locus has a sequence comprising "at least a portion" of a DMR sequence listed herein, this means that the methylation locus comprises at least a portion (e.g., at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%) of CpG dinucleotides corresponding to CpG dinucleotides within the DMR sequence.
[0123] Pharmaceutical composition: As used herein, the term "pharmaceutical composition" refers to a composition comprising an active pharmaceutical agent formulated with one or more pharmaceutically acceptable carriers. In some embodiments, such as those described herein, the active pharmaceutical agent is present in a unit dosage suitable for administration to a subject (e.g., a therapeutic regimen that demonstrates a statistically significant probability of achieving a predetermined therapeutic effect when administered to a relevant population). In some embodiments, such as those described herein, the pharmaceutical compositions can be formulated into a particular form of administration (e.g., a solid form or a liquid form) and / or can be particularly suitable for: for example, oral administration (e.g., as a drench (aqueous or non-aqueous solution or suspension), tablet, capsule, pill, powder, granule, paste, etc., which can be specially formulated for oral, sublingual or systemic absorption); parenteral administration (e.g., by subcutaneous, intramuscular, intravenous or epidural injection, such as a sterile solution or suspension, or a sustained-release formulation, etc.); topical administration (e.g., as a sterile solution or suspension or sustained-release formulation, etc., by oral, sublingual or epidural injection); local administration (e.g., as a cream, ointment, patch or spray applied to, for example, the skin, lungs or mouth); intravaginal or rectal administration (e.g., in the form of a vaginal suppository, suppository, cream or foam, etc.); ophthalmic administration; nasal or pulmonary administration, etc.
[0124] Pharmaceutically acceptable: As used herein, the term "pharmaceutically acceptable" applies to one or more or all ingredients of the formulation of the compositions disclosed herein, meaning that each ingredient must be compatible with the other ingredients of the composition and not have a deleterious effect on the recipient thereof.
[0125] Pharmaceutically acceptable carrier: The term "pharmaceutically acceptable carrier" as used herein refers to a pharmaceutically acceptable material, composition or vehicle, such as a liquid or solid filler, diluent, excipient or solvent encapsulating material, that facilitates formulation and / or altered bioavailability of a pharmaceutical agent (e.g., a therapeutic agent). Examples of materials that can serve as pharmaceutically acceptable carriers include sugars such as lactose, glucose, and sucrose; starches such as corn starch and potato starch; cellulose and its derivatives such as sodium carboxymethylcellulose, ethylcellulose, and cellulose acetate; tragacanth powder; malt; gelatin; talc; excipients such as cocoa butter and suppository waxes; oils such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil, and soybean oil; glycols such as propylene glycol; polyols such as glycerol, sorbitol, mannitol, and polyethylene glycol; esters such as ethyl oleate and ethyl laurate; agar; buffers such as magnesium hydroxide and aluminum hydroxide; alginic acid; pyrogen-free water; isotonic saline; Ringer's solution; ethanol; pH buffer solutions; polyesters, polycarbonates, and / or polycyanates; and other nontoxic, compatible substances used in pharmaceutical formulations.
[0126] Polyposis syndrome: As used herein, the terms "polyposis" and "polyposis syndrome" refer to genetic conditions including, but not limited to, familial adenomatous polyposis (FAP), hereditary nonpolyposis colorectal cancer (HNPCC) / Lynch syndrome, Gardner syndrome, Turcot syndrome, MUTYH polyposis, Peutz-Jeghers syndrome, Cowden disease, familial juvenile polyposis, and hyperplastic polyposis. In certain embodiments, the polyposis syndrome comprises serrated polyposis syndrome. Serrated polyposis syndrome is classified as: a patient having five or more serrated polyps in the proximal sigmoid colon, two or more of which are at least 10 mm in size; a patient having serrated polyps in the proximal sigmoid colon and a family history of serrated polyposis; and / or a patient having 20 or more serrated polyps throughout the colon.
[0127] When used in relation to a gene, a portion refers to a fragment of the gene. The size of a fragment can range from a few nucleotides to the entire gene sequence minus one nucleotide. Thus, "nucleotides comprising at least a portion of a gene" can include a gene fragment or the entire gene.
[0128] Prevent or Prophylaxis: As used herein, the terms "prevent" and "prevent" refer to the occurrence of a disease, disorder, or condition and refer to reducing the risk of developing the disease, disorder, or condition; delaying the onset of the disease, disorder, or condition; delaying the onset of one or more features or symptoms of the disease, disorder, or condition; and / or reducing the frequency and / or severity of one or more features or symptoms of the disease, disorder, or condition. Prevention can refer to prevention in a specific subject or to a statistical effect in a population of subjects. Prevention is considered accomplished when the onset of the disease, disorder, or condition is delayed for a predetermined period of time.
[0129] Probe: As used herein, the terms "probe," "capture probe," or "bait" refer to a single-stranded or double-stranded nucleic acid molecule capable of hybridizing to a complementary target, and in certain embodiments, include a detectable portion. In certain embodiments, such as those described herein, the probe is a restriction digest or a synthetically produced nucleic acid, such as one produced by recombination or amplification. In some instances, such as those described herein, the probe is a capture probe for detecting, identifying, and / or isolating a target sequence (e.g., a gene sequence). In various instances, such as those described herein, the detectable portion of the probe can be, for example, an enzyme (e.g., ELISA and enzyme-based histochemical assays), a fluorescent moiety, a radioactive moiety, or a moiety associated with a luminescent signal.
[0130] Prognosis: As used herein, the term "prognosis" refers to a qualitative determination of the quantitative probability of at least one possible future outcome or event. As used herein, prognosis can be a determination of the likely course of a disease, disorder, or condition (e.g., cancer) in a subject, a determination of the life expectancy of a subject, or a determination of the response to treatment (e.g., to a particular therapy).
[0131] Prognostic information: As used herein, the terms "prognostic information" and "predictive information" refer to any information that can be used to indicate any aspect of a disease or condition, either without or with treatment. Such information may include, but is not limited to, the average life expectancy of a patient, the likelihood that a patient will survive a given time (e.g., 6 months, 1 year, 5 years, etc.), the likelihood that a patient's disease will be cured, the likelihood that a patient's disease will respond to a particular therapy (wherein response can be defined in any of a variety of ways). Both prognostic information and predictive information are included within the broad category of diagnostic information. Prognostic information may include, but is not limited to, biomarker status information.
[0132] Promoter: As used herein, "promoter" may refer to a DNA regulatory region that directly or indirectly (eg, through a protein or substance that binds to the promoter) associates with RNA polymerase and participates in initiating transcription of a coding sequence.
[0133] Ratio: As used herein, the term "ratio" refers to a calculable relationship that is used to compare the amounts of two substances and indicates the relative amounts of the two substances. The substance can be a marker, such as a metabolite and / or a lipid (e.g., a fatty acid). The ratio can be directly proportional or inversely proportional (e.g., the first amount divided by the second amount, or the second amount divided by the first amount, respectively). The ratio can be weighted and / or normalized (numerator, denominator, or both). The two quantities can be physical quantities or arbitrary values corresponding to physical quantities. For example, the ratio can be calculated by measuring two intensity quantities (i.e., arbitrary units) of two substances (e.g., markers) by mass spectrometry.
[0134] Reference: As used herein, a reference describes a standard or control relative to which a comparison is made. For example, in some embodiments, as described herein, an agent, subject, animal, individual, population, sample, sequence, or value of interest is compared to a reference or control agent, subject, animal, individual, population, sample, sequence, or value. In some embodiments, as described herein, the testing and / or determination of a reference or characteristic thereof is performed substantially simultaneously with the testing or determination of the characteristic in the sample of interest. In some embodiments, as described herein, a reference is a historical reference, optionally embodied in a tangible medium. Typically, those skilled in the art understand that a reference is determined or characterized under conditions or environments comparable to the conditions or environments under which the reference is evaluated (e.g., related to the sample). Those skilled in the art will understand when there is sufficient similarity to justify reliance on and / or comparison of a particular possible reference or control.
[0135] Risk: The term "risk" as used herein with respect to a disease, disorder, or condition refers to a qualitative (expressed as a percentage or otherwise) measure of the quantitative probability that a particular individual will develop a disease, disorder, or condition. In some embodiments, such as described herein, risk is expressed as a percentage. In some embodiments, such as described herein, risk is a qualitative measure of a quantitative probability that is equal to or greater than 0%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100%. In some embodiments, such as described herein, risk is expressed as a qualitative or quantitative risk level relative to a reference risk or level, or the risk of the same outcome attributable to a reference. In some embodiments, e.g., as described herein, the relative risk is increased or decreased by more than 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, 3, 4, 5, 6, 7, 8, 9, 10 fold compared to a reference sample.
[0136] Sample: As used herein, the term "sample" generally refers to an aliquot of material obtained or derived from a source of interest. In some embodiments, such as described herein, the source of interest is a biological or environmental source. In some embodiments, such as described herein, the sample is a "raw sample" obtained directly from the source of interest. In some embodiments, such as described herein, the term "sample" refers to a preparation obtained by processing the raw sample (e.g., removing one or more components of the raw sample and / or adding one or more agents to the raw sample), as the context dictates. Such a "processed sample" can include, for example, cells, nucleic acids, or proteins extracted from the sample, or cells, nucleic acids, or proteins obtained by subjecting the raw sample to techniques such as nucleic acid amplification or reverse transcription, separation and / or purification of certain components.
[0137] In some instances, such as those described herein, the processed sample can be an amplified (e.g., pre-amplified) DNA sample. Thus, in various cases, such as those described herein, the identified sample can refer to the original form of the sample or to the processed form of the sample. In some instances, such as those described herein, the enzymatically digested DNA sample can refer to the initial enzymatically digested DNA (the direct product of the enzymatic digestion) or a further processed sample, such as an enzymatically digested DNA that has undergone an amplification step (e.g., an intermediate amplification step, such as pre-amplification) and / or a filtration step, a purification step, or a step in which the sample is modified to facilitate other steps, such as in determining the methylation state (e.g., the methylation state of the original sample of DNA and / or the DNA in the context of its original source).
[0138] Screening: As used herein, the term "screening" refers to any method, technique, process, or procedure designed to generate diagnostic and / or prognostic information. Thus, one skilled in the art will understand that the term "screening" encompasses methods, techniques, processes, or procedures for determining whether an individual has, is likely to have, or is developing, or is at risk of having or developing, a disease, disorder, or condition (e.g., colorectal cancer, advanced adenoma).
[0139] Specificity: As described herein, the "specificity" of a biomarker refers to the percentage (true negative rate) at which a biomarker measured in a sample characterized by the absence of the event or state of interest accurately indicates the absence of the event or state of interest. In various embodiments, such as those described herein, characterization of negative samples is independent of the biomarker and can be achieved by any relevant measurement method, such as any relevant measurement method known to those skilled in the art. Thus, specificity reflects the probability that a biomarker detects the absence of the event or state of interest when measured in a sample that does not characterize the event or state of interest. In specific embodiments where the event or state of interest is colorectal cancer, such as those described herein, specificity refers to the probability that a biomarker detects the absence of colorectal cancer in a subject that does not suffer from colorectal cancer. The absence of colorectal cancer can be determined, for example, by histological methods.
[0140] Sensitivity: As described herein, the "sensitivity" of a biomarker refers to the percentage (true positive rate) of biomarkers measured in samples characterized by the presence of an event or state of interest that accurately indicates the presence of the event or state of interest. In various embodiments, such as those described herein, the characterization of a positive sample is independent of the biomarker and can be achieved by any relevant measurement method, such as any relevant measurement method known to those skilled in the art. Therefore, sensitivity reflects the probability that a biomarker detects the presence of the event or state when measured in a sample characterized by the event or state of interest. In specific embodiments where the event or state of interest is colorectal cancer, such as those described herein, sensitivity refers to the probability that a biomarker detects the presence of colorectal cancer in a subject suffering from colorectal cancer. Suffering from colorectal cancer can be determined, for example, by histological methods.
[0141] Single Nucleotide Polymorphism (SNP): As used herein, the term "single nucleotide polymorphism" or "SNP" refers to a specific base position in the genome where alternative bases are known to distinguish one allele from another. In some embodiments, one or a few SNPs and / or CNPs are sufficient to distinguish complex genetic variations from one another, and thus, for analytical purposes, one or a group of SNPs and / or CNPs can be considered as characteristic of a particular variation, trait, cell type, individual, species, etc., or a collection thereof. In some embodiments, one or a group of SNPs and / or CNPs can be considered as defining a particular variation, trait, cell type, individual, species, etc., or a collection thereof.
[0142] Solid tumor: As used herein, the term "solid tumor" refers to an abnormal mass of tissue that contains cancer cells. In various embodiments, such as described herein, a solid tumor is or includes an abnormal mass of tissue that does not contain cysts or fluid areas. In some embodiments, such as described herein, a solid tumor can be benign; in some embodiments, a solid tumor can be malignant. Examples of solid tumors include carcinomas, lymphomas, and sarcomas. In some embodiments, such as described herein, a solid tumor can be or include a tumor of the adrenal gland, bile duct, bladder, bone, brain, breast, cervix, colon, endometrium, esophagus, eye, gallbladder, gastrointestinal tract, kidney, larynx, liver, lung, nasal cavity, nasopharynx, oral cavity, ovary, penis, pituitary gland, prostate, retina, salivary gland, skin, small intestine, stomach, testicle, thymus, thyroid, uterus, vagina, and / or ovary.
[0143] Cancer Stage: As used herein, the term "cancer stage" refers to a qualitative or quantitative assessment of how advanced a cancer is. In some embodiments, such as those described herein, criteria used to determine cancer stage may include, but are not limited to, one or more of the following: the location of the cancer in the body, the size of the tumor, whether the cancer has spread to lymph nodes, whether the cancer has spread to one or more different parts of the body, etc. In some embodiments, such as those described herein, cancer may be staged using the so-called TNM system, in which T refers to the size and extent of the main tumor (often referred to as the primary tumor); N refers to the number of nearby lymph nodes with cancer; and M refers to whether the cancer has metastasized. In some embodiments, such as those described herein, cancer may be referred to as Stage 0 (abnormal cells are present but have not spread to nearby tissues, also known as carcinoma in situ, or CIS; CIS is not cancer but can become cancer), Stages I-III (cancer is present; the higher the number, the larger the tumor and the greater the extent of spread to nearby tissues), or Stage IV (cancer has spread to distant parts of the body). In some embodiments, such as described herein, a cancer can be assigned a stage selected from the group consisting of: in situ (abnormal cells are present but have not spread to nearby tissues); local (cancer is confined to where it started, with no signs of spread); regional (cancer has spread to nearby lymph nodes, tissues, or organs); distant (cancer has spread to distant parts of the body); and unknown (not enough information is available to determine the cancer's stage).
[0144] Stratification: "Stratification" in this context refers to any analytical process that divides patients into distinct groups. These groups may share similar characteristics or features that make them distinct. Stratification can be applied to any study characteristics that aid in CRC diagnosis prediction, early detection, monitoring, treatment guidance, and survival prognosis.
[0145] Survival: In this context, "survival" refers to the time a patient lives from the onset of a disease (e.g., cancer) or the start of treatment. It is a term that can be used to describe how effective a new approach (including prognosis, screening, diagnosis, treatment, and monitoring) is in the face of disease progression.
[0146] Susceptible: An individual who is "susceptible" to a disease, disorder, or condition is at risk for developing the disease, disorder, or condition. In some embodiments, such as described herein, an individual who is susceptible to a disease, disorder, or condition does not exhibit any symptoms of the disease, disorder, or condition. In some embodiments, such as described herein, an individual who is susceptible to a disease, disorder, or condition has not yet been diagnosed with the disease, disorder, and / or condition. In some embodiments, such as described herein, an individual who is susceptible to a disease, disorder, or condition is an individual who has been exposed to conditions associated with developing the disease, disorder, or condition or who exhibits a biomarker state (e.g., a methylation state) associated with developing the disease, disorder, or condition. In some embodiments, such as described herein, the risk of developing a disease, disorder, and / or condition is based on the risk of a population (e.g., family members of an individual with the disease, disorder, or condition).
[0147] Subject: As used herein, the term "subject" refers to an organism, typically a mammal (e.g., a human). In some embodiments, such as described herein, a subject suffers from a disease, disorder, or condition. In some embodiments, such as described herein, a subject is susceptible to a disease, disorder, or condition. In some embodiments, such as described herein, a subject exhibits one or more symptoms or characteristics of a disease, disorder, or condition. In some embodiments, such as described herein, a subject does not suffer from a disease, disorder, or condition. In some embodiments, such as described herein, a subject does not show any symptoms or characteristics of a disease, disorder, or condition. In some embodiments, such as described herein, a subject is a subject with characteristics of susceptibility or risk for one or more diseases, disorders, or conditions. In some embodiments, such as described herein, a subject is a patient. In some embodiments, such as described herein, a subject is an individual who has been diagnosed and / or to whom therapy has been administered. In some instances, such as described herein, a human subject may be referred to interchangeably as an "individual."
[0148] Therapeutic agent: As used herein, the term "therapeutic agent" refers to any agent that produces a desired pharmacological effect when administered to a subject. In some embodiments, such as described herein, an agent is considered a therapeutic agent if it shows a statistically significant effect in an appropriate population. In some embodiments, such as described herein, an appropriate population can be a model organism population or a human population. In some embodiments, such as described herein, an appropriate population can be defined by various criteria, such as a specific age group, gender, genetic background, existing clinical symptoms, etc. In some embodiments, such as described herein, a therapeutic agent is a substance that can be used to treat a disease, disorder, or condition. In some embodiments, such as described herein, a therapeutic agent is an agent that has been or requires approval by a government agency before it can be marketed for human administration. In some embodiments, such as described herein, a therapeutic agent is an agent that requires a medical prescription for administration to humans.
[0149] Therapeutically effective amount: As used herein, the term "therapeutically effective amount" refers to an amount that produces the desired effect of administration. In some embodiments, such as described herein, the term refers to an amount that is sufficient to treat the disease, disorder, or condition when administered according to a therapeutic dosage regimen to a population suffering from or susceptible to a disease, disorder, or condition. Those skilled in the art will understand that the term therapeutically effective amount does not actually require successful treatment in a specific individual. Rather, a therapeutically effective amount can be an amount that provides a specific desired pharmacological response in a large number of subjects when administered to an individual in need of such treatment. In some embodiments, such as described herein, reference to a therapeutically effective amount can refer to an amount measured in one or more specific tissues (e.g., tissues affected by a disease, disorder, or condition) or body fluids (e.g., blood, saliva, serum, sweat, tears, urine, etc.). Those skilled in the art will understand that in some embodiments, a therapeutically effective amount of a particular agent can be formulated and / or administered in a single dose. In some embodiments, such as described herein, a therapeutically effective amount of an agent can be formulated and / or administered in multiple doses, such as as part of a multi-dose treatment regimen.
[0150] Treatment: As used herein, the term "treatment" (also referred to as "treating" or "therapy") refers to the administration of a therapy that partially or completely alleviates, ameliorates, alleviates, suppresses, delays the onset of, reduces the severity of, and / or reduces the incidence of one or more symptoms, features, and / or causes of a particular disease, disorder, or condition, or that is administered with the purpose of achieving any such result. In some embodiments, such as described herein, such treatment may be performed on subjects who do not exhibit the relevant disease, disorder, or symptoms and / or subjects who exhibit only early symptoms of the disease, disorder, or symptoms. Alternatively or additionally, treatment may also be performed on subjects who exhibit one or more known characteristics of the relevant disease, disorder, and / or condition. In some embodiments, such as described herein, treatment may be performed on subjects who have been diagnosed with the relevant disease, disorder, and / or condition. In some embodiments, such as described herein, treatment may be performed on subjects who are known to have one or more predisposing factors that are statistically correlated with an increased risk of developing the relevant disease, disorder, or condition. In various examples, treatment is performed on cancer.
[0151] Upstream: The term "upstream" as used herein refers to a first DNA region being closer to the N-terminus of a nucleic acid comprising the first and second DNA regions relative to a second DNA region.
[0152] Unit Dose: As used herein, the term "unit dose" refers to an amount administered as a physically discrete unit of a single dose and / or pharmaceutical composition. In many embodiments, such as those described herein, a unit dose comprises a predetermined amount of an active pharmaceutical agent. In some embodiments, such as those described herein, a unit dose comprises an entire single dose of a pharmaceutical agent. In some embodiments, such as those described herein, more than one unit dose is administered to achieve the entire single dose. In some embodiments, such as those described herein, administration of multiple unit doses is necessary or anticipated to achieve a desired effect. For example, a unit dose can be a volume of a liquid (e.g., an acceptable carrier) containing a predetermined amount of one or more therapeutic moieties, a predetermined amount of one or more therapeutic moieties in solid form, a sustained-release formulation or drug delivery device containing a predetermined amount of one or more therapeutic moieties, and the like. It will be understood that a unit dose can be present in a formulation comprising any of a variety of components in addition to the therapeutic agent. For example, an acceptable carrier (e.g., a pharmaceutically acceptable carrier), a diluent, a stabilizer, a buffer, a preservative, and the like can be included. One skilled in the art will appreciate that, in many embodiments, such as those described herein, an appropriate total daily dose of a particular therapeutic agent can comprise a portion or multiple unit doses and can be determined by a physician within the scope of sound medical judgment. In some embodiments, as described herein, a specific effective dosage level for any particular subject or organism may depend on a variety of factors, including the disorder being treated and the severity of the disorder; the activity of the specific active compound being used; the specific composition being used; the age, weight, general health, sex, and diet of the subject; the time of administration and the rate of excretion of the specific active compound being used; the duration of the treatment; drugs and / or additional therapies used in combination or concomitantly with the specific compound being used, and like factors well known in the medical art.
[0153] Unmethylated: As used herein, the terms "unmethylated" and "nonmethylated" are used interchangeably to mean that the identified DNA region does not contain methylated nucleotides.
[0154] Variant: As used herein, the term "variant" refers to an entity that has significant structural identity to a reference entity but differs structurally from the reference entity in the presence, absence, or level of one or more chemical moieties as compared to the reference entity. In some embodiments, such as those described herein, a variant also differs functionally from the reference entity. Generally, whether a particular entity is appropriately considered a "variant" of a reference entity is determined based on the degree of structural identity it shares with the reference entity. A variant can be a molecule that is comparable but not identical to a reference. For example, a variant nucleic acid can differ from a reference nucleic acid by one or more differences in nucleotide sequence. In some embodiments, such as those described herein, a variant nucleic acid exhibits an overall sequence identity of at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, or 99% to a reference nucleic acid. In many embodiments, such as those described herein, a nucleic acid of interest is considered a "variant" of a reference nucleic acid if it has a sequence identical to that of the reference nucleic acid but with minor sequence changes at specific positions. In some embodiments, such as those described herein, the variant has 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 substituted residue compared to a reference. In some embodiments, such as those described herein, the variant has no more than 5, 4, 3, 2, or 1 addition, substitution, or deletion of residues compared to a reference nucleic acid. In various embodiments, such as those described herein, the number of additions, substitutions, or deletions is less than about 25, about 20, about 19, about 18, about 17, about 16, about 15, about 14, about 13, about 10, about 9, about 8, about 7, about 6 residues, and typically less than about 5, about 4, about 3, or about 2 residues. BRIEF DESCRIPTION OF THE DRAWINGS
[0155] The foregoing and other objects, aspects, features and advantages of the present invention will become more apparent by referring to the following description taken in conjunction with the accompanying drawings, in which:
[0156] Figure 1 Schematic diagram illustrating the comparison of DNA methylation in normal cells and cancer cells.
[0157] Figure 2 is a bar graph showing biological pathway marker regions belonging to the Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway database.
[0158] Figure 3 is a graph showing Kaplan-Meier analysis of patient groups divided by hierarchical clustering from TCGA-COAD / READ, according to an illustrative embodiment.
[0159] Figure 4is a graph showing sensitivity and specificity values for prediction results (advanced adenoma and colorectal cancer) using an 82-member panel (genomic regions in Table 1, Seq ID No. 1-82) according to an illustrative embodiment with reference to Example 4 presented herein.
[0160] Figure 5 is a schematic diagram depicting the data details for the samples presented in Example 4 herein, according to an illustrative embodiment. DETAILED DESCRIPTION
[0161] It is contemplated that the systems, architectures, devices, methods, and processes of the present invention encompass variations and changes developed using information from the embodiments described herein. As contemplated by this specification, changes and / or modifications may be made to the systems, architectures, devices, methods, and processes described herein.
[0162] Throughout this specification, when articles, devices, systems, and architectures are described as having, including, or comprising particular components, or when processes and methods are described as having, including, or comprising particular steps, it is additionally contemplated that there are articles, devices, systems, and architectures of the invention that consist of or consist essentially of the components, and that there are processes and methods of the invention that consist essentially of or consist of the steps.
[0163] It should be understood that the order of steps or the order in which certain actions are performed is not important, as long as the present invention remains operable. In addition, two or more steps or actions can be performed simultaneously.
[0164] Reference to any publication herein (e.g., in the Background section) is not an admission that the publication is prior art with respect to any claim presented herein. The Background section is provided for clarity purposes and is not intended to be a description of prior art with respect to any claim.
[0165] Documents incorporated herein by reference are noted. Furthermore, the entire contents of all documents cited herein are incorporated herein by reference, regardless of whether they are specifically noted. In the event of any ambiguity in the meaning of a particular term, the meaning provided in the definitions above will prevail.
[0166] Headings are provided for the convenience of the reader—the presence and / or placement of headings is not intended to limit the scope of the subject matter described herein.
[0167] DNA methylation has previously been shown to have the potential to diagnose and predict colorectal cancer (CRC). Methylation changes can also be used for patient stratification and monitoring. Patient stratification can improve treatment outcomes by providing clinicians with better methods to distinguish between CRC patient subtypes.
[0168] Cell-free DNA (cfDNA) in the bloodstream is primarily a byproduct of cell death, and its fragments are relatively short, mostly corresponding to the average length of mononucleosomes. Circulating fetal DNA has been shown to be shorter than maternal DNA in plasma, and these size differences have been used to improve the sensitivity of non-invasive prenatal diagnosis. Similar findings have been noted in cancer patients, with tumor-derived DNA fragments (ctDNA) being shorter than non-tumor cell-derived fractions.
[0169] Fragment length differences in circulating DNA, combined with methylation signals, can be used to increase the sensitivity of detecting the presence of ctDNA and for non-invasive genomic analysis of cancer.
[0170] Summary of the experiments conducted
[0171] The goals of these experiments were to evaluate putative methylation markers in the context of early cancer development and diagnosis, and to further investigate the biological significance of these regions.
[0172] Biomarker discovery was performed by whole-genome bisulfite sequencing (WGBS) of 88 CRC, 48 advanced adenomas (AA), and corresponding adjacent normal tissue (NAT) samples. A short list of significantly hypermethylated regions (DMRs) was correlated with transcriptomic data from 512 CRC patients from The Cancer Genome Atlas (TCGA) consortium. Pathway enrichment for biological pathway analysis of DMRs was performed using the Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway database. Survival analyses were performed using Kaplan–Meier regression for patient subgroups distinguished by the methylation status of individual markers. Finally, the significance of individual markers in selected regions was assessed by analyzing plasma samples from 26 early-stage (stage I-IIA) CRC samples and 42 colonoscopy validation controls (CNTs) using targeted methylation sequencing assays.
[0173] 4167 putative marker regions were identified from the biomarkers found using WGBS. Differential signals can be observed between AA and NAT and between CRC and NAT, and some of these regions are also differentially methylated between AA and CRC samples, indicating that biological signals change as adenoma progresses to carcinoma. 84 DMRs from several validation studies were further evaluated based on transcriptome data from TCGA, where 69 genes were found to overlap. 19 of these genes were significantly downregulated (p < 0.05), indicating a connection between hypermethylation and gene expression. 2 genes were significantly upregulated (p < 0.05), which may indicate the presence of other epigenetic processes. KEGG pathway analysis revealed that the main pathways involved are axon guidance, ephrin receptor signaling, epithelial-mesenchymal transition, and FGF signaling, which all play an important role in the context of cancer occurrence and progression. Kaplan-Meier analysis showed that the prediction of patient 5-year survival was significantly associated with 3 genes, FGF14 (p = 0.025, HR = 1.75), DPY19L2P1 (p = 0.012, HR = 1.86), and PTPRO (p = 0.046, HR = 1.63). Targeted sequencing analysis of plasma samples from patients with early-stage (I-IIA) colorectal cancer and age- and sex-matched colonoscopy-validated controls showed high accuracy for individual markers, with AUCs of FGF14 = 0.78, DPY19L2P1 = 0.81, and PTPRO = 0.73.
[0174] In the early stages of CRC, methylation markers have distinct signals that can distinguish early-stage cancers from matched controls with high individual accuracy. These regions influence gene expression and can be linked to relevant biological pathways. Extending the early detection potential of these markers to further prognostication and stratification could lead to better outcomes and improved patient survival.
[0175] Methylation testing
[0176] The methylation status of a particular marker can be assessed using various techniques, and thus are not limited to any particular method for measuring the methylation status of a gene. For example, methylation status can be measured by genome sequencing methods, in which the entire genome is scanned at a resolution of base pairs. Another approach can involve analyzing changes in methylation patterns using PCR-based methods that involve digesting DNA with a methylation-sensitive restriction enzyme prior to PCR amplification.
[0177] In the case of MSRE-qPCR, the total amount of DNA is determined by directly measuring the extracted DNA fraction in its native form using real-time or digital PCR. This fraction is then digested with a restriction enzyme that degrades unmethylated DNA, leaving only the methylated strands intact. The resulting methylated sequence is then determined again using real-time or digital PCR.
[0178] The method based on real-time PCR includes producing a standard curve of unmethylated target by using internal standard. A standard curve is established by at least two points and the real-time Ct value of undigested and digested DNA is associated with a known quantitative standard. Next, the test sample Ct value of the digested and undigested colony is determined, and the genomic DNA equivalent is calculated by the standard curve generated. The Ct value of the undigested and digested DNA is assessed to establish markers of a Ct value that is truly digested and produces 45, as well as markers that fail to be amplified by the sample and can be regarded as failure values. Then, the value from the digested DNA after filtering and correction can be directly compared between different condition groups to establish the relative difference of methylation levels between groups. In addition, the Ct value difference between the undigested and digested DNA can be used for the same purpose.
[0179] Digital PCR-based methods distribute reactions in microfluidic devices across 96- or 384-well plates, resulting in an average initial template DNA concentration of less than one molecule per reaction chamber. Amplification of methylated DNA molecules occurs in only a few PCR wells, thus representing a digital readout of the original number of template molecules in each sample.
[0180] In addition, other techniques using bisulfite treatment of DNA as a starting point for methylation analysis can be used. These techniques include methylation-specific PCR (MSP) (Herman et al., (1992) Proc. Natl. Acad. 20 Sci. USA 93: 9821-9826), methylation-specific nuclease-assisted minor allele enrichment (MS-NaME)-PCR (Liu Y et al. Nucleic Acids Res. 2017 Apr 7; 45 (6): e39), methylation-sensitive high-resolution melting (MS-HRM) PCR (Hussmann D, Hansen LL. Methods Mol Biol. 2018; 1708: 551-571), all of which are incorporated herein by reference.
[0181] When assessing methylation status, the methylation status is typically expressed as the fraction or percentage of each DNA strand that is methylated at a specific site (e.g., at a single nucleotide, at a specific region or locus, at a longer sequence of interest (e.g., a subsequence of DNA up to ~50-bp, 100-bp, 150-bp, 200-bp, 500-bp, 1000-bp or greater)) relative to the total population of DNA in the sample that contains the specific site.
[0182] When the methylation ratio of one or more DNA methylation markers is different from the level of bisulfite-treated DNA copy number of a reference gene, it can also indicate that the biological sample has produced a tumor, wherein the one or more DNA methylation markers include bases in a differentially methylated region (DMR). The methylation ratio includes the ratio of the methylation level of the DNA methylation marker and the methylation level of the region in the reference gene determined by the same method as the methylation level of the biomarker. Typically, the methylation ratio is expressed as the ratio of the methylation level of the DNA methylation marker and the methylation level of the region in the reference gene determined by the same method as the methylation level of the DNA methylation marker.
[0183] The methylation ratio can be the ratio of the methylation level of a DNA methylation marker to the methylation level of a region in a reference gene, both of which are quantitatively measured using real-time polymerase chain reaction (PCR) or droplet digital PCR (ddPCR) methods. For example, a pair of primers and an oligonucleotide probe can be used to quantitatively measure the methylation level of a DNA methylation marker from a subject sample, wherein, for example, after the nucleic acid is modified by a modifying agent (e.g., digesting unmethylated DNA with a methylation-specific enzyme or converting unmethylated cytosine into converted nucleic acid using bisulfite), one primer, two primers, an oligonucleotide probe, or both a primer and an oligonucleotide probe can distinguish and selectively amplify methylated and unmethylated nucleic acids.
[0184] Biomarker discovery
[0185] The purpose of this example is to identify differentially methylated regions (DMRs) in the DNA of colorectal cancer and colon adenoma samples (e.g., samples from subjects with advanced adenomas). DMRs are identified by comparing DNA from subjects with colorectal cancer and / or colon adenoma with matched control samples. This comparison allows the development of methods to elucidate methylation patterns associated with colorectal cancer and advanced adenomas from cell-free (cfDNA).
[0186] Whole-genome bisulfite sequencing (WGBS) was used to identify differences in methylation status in genomic DNA (gDNA) and cfDNA samples obtained from various sources. gDNA was obtained from tissue samples with different histological backgrounds (e.g., colorectal cancer, colon adenoma, lung cancer, breast cancer, colorectal cancer, gastric cancer, and matched controls) and buffy coat samples.
[0187] Genomic DNA (gDNA) was extracted from tissue and buffy coat samples using the DNeasy blood and tissue kit (Qiagen) according to the manufacturer's protocol. The extracted gDNA was then further processed to fragment it. For example, the gDNA was fragmented using a Covaris S220 sonicator to form segments of approximately 400 bp in length.
[0188] cfDNA was extracted from plasma samples using the QIAamp circulating nucleic acid kit (Qiagen) according to the manufacturer's protocol.
[0189] The extracted and fragmented gDNA (genomic DNA) and cfDNA were subjected to bisulfite conversion using the EZ DNA Methylation-Lightning kit (ZymoResearch). Sequencing libraries were prepared from bisulfite-converted DNA fragments using the Accel-NGS Methyl-seq DNA library kit (SwiftBiosciences). The converted DNA fragments were sequenced using paired-end sequencing with an average depth of 37.5× using a NovaSeq6000 (Illumina) device. For this experiment, paired-end sequencing was performed so that each 150bp (e.g., 2×150) of the ends of the converted DNA fragments was covered. The sequencing reads were aligned with the bisulfite-converted human genome (Ensembl91 assembly) using the Bisulfite Read Mapper using Bowtie2. The following steps are used to align the sequencing reads with the bisulfite-converted human genome:
[0190] 1. Evaluate Sequencing Quality
[0191] 2. Comparison with the reference genome (hG38)
[0192] 3. Removal of repeats and elimination of adapter dimers
[0193] 4. Methylation calling (e.g., identification of methylated nucleic acids)
[0194] Differentially methylated regions were analyzed by comparing the beta (β) values of individual CpGs in colon cancer and / or colon adenoma tissue samples with matched control tissues. The β value reflects the methylation level of the CpG reads in the sample. A β value of 0 indicates that no methylated reads were found at a specific CpG position, while a β value of 1 indicates that all reads were fully methylated. The individual CpG methylation scores were combined into regions with at least 3 CpGs within 50 bp of each other. The q value of the region (i.e., the p value corrected by the inter-group label permutation test) was evaluated to select DNA regions with significantly different methylation from the same region in the DNA obtained from the subjects with colorectal cancer and / or colon adenoma. A q value < 0.05 was considered to be a differentially methylated region (DMR) showing a high statistical significance. The significant regions were further evaluated to determine whether there was a significant methylation signal compared to tissue samples from non-colorectal cancer sources, control tissue samples from non-colorectal cancer sources, buffy coat samples, and cfDNA from healthy individuals.
[0195] A total of 4167 DMRs were initially identified as significant for colorectal cancer and / or advanced adenoma. These DMRs included regions more indicative of colorectal cancer, DMRs more indicative of different histological subtypes of colon adenoma, and regions indicative of both colorectal cancer and advanced adenoma.
[0196] Further cancer signal analysis was performed on selected target regions from whole-genome sequencing data using a read-by-read signal scoring method. A threshold was calculated in paired tissue-control samples to maximize the discrimination between cancer and control reads. The calculated score was applied to each read obtained from the subject's plasma cfDNA.
[0197] Example 1: Feature evaluation of transcriptome data
[0198] The signature evaluation study was performed by evaluating data from The Cancer Genome Atlas (TCGA) research network (http: / / cancergenome.nih.gov / ) on CRC (TCGA-COAD, TCGA-READ) and normal tissue data.
[0199] From the initial list of 4167 DMRs, a short list of 82 hypermethylated DMRs (Table 1) was validated in several plasma cfDNA studies and selected for further evaluation based on transcriptome data from TCGA. This study involved several steps, first matching the DMR regions with the 450K Illumina methylation array data available in TCGA to identify regions present in the TCGA data.
[0200] The comparison was performed in a supervised mode, with the two groups to be compared designated as healthy tissue and tumor tissue. Fisher's exact test and Benjamini-Hochberg multiple hypothesis correction were used to compare the frequency of each motif flanking the positive CpG probes with the frequency of the background defined by all distal probes on the array. Those CpGs that passed the minimum threshold (significant methylation difference of 0.2, P value of 0.05) were considered to cause functional changes.
[0201] Second, we sought to identify methylation patterns potentially associated with the clonal evolution of specific tumor types. To achieve this, we performed KEGG pathway analysis to capture biological processes related to regulatory networks, morphogenesis, development, and cell differentiation.
[0202] result
[0203] We identified 69 genes that overlapped with the 409 individual CpG counts. We then correlated the methylation signals of these 69 genes with gene expression levels using an enhancer-associated methylation / expression relationship algorithm, thereby linking methylation levels to gene expression, which can indicate functional changes due to changes in methylation levels.
[0204] A total of 19 genes (22 regions) (Table 2) showed significant downregulation (p < 0.05), indicating a link between hypermethylation and gene expression. Two genes (2 regions) (Table 3) showed significant upregulation (p < 0.05), which may indicate the presence of other epigenetic processes.
[0205] Table 2 : List of 22 genomic regions (19 genes) found to be significantly downregulated in CRC patients
[0206]
[0207]
[0208] Table 3 : List of 2 genomic regions (2 genes) found to be significantly upregulated in CRC patients
[0209] 7:93889427-93891122(SEQ ID NO:70) 8:66961046-66962606(SEQ ID NO:72)
[0210] COAD / READ KEGG pathway analysis (see Figure 2 ) revealed major pathways involving axon guidance, ephrin receptor signaling, epithelial-mesenchymal transition, and FGF signaling, all of which play important roles in the context of cancer initiation and progression.
[0211] Example 2: Feature evaluation for patient stratification and 5-year survival prediction
[0212] Third, to evaluate the 5-year survival function, Kaplan-Meier analysis was used to evaluate each marker in the hypermethylated (methylated reads > 20%) and hypomethylated (methylated reads < 20%) patient groups. A total of 142 patients were assigned to the high methylation group, and a total of 142 patients were assigned to the low methylation group. The groups were balanced in terms of age, sex, cancer stage, and site.
[0213] result
[0214] The prediction of 5-year survival was closely associated with three genes (Table 4): FGF14 (p = 0.025, HR = 1.75), DPY19L2P1 (p = 0.012, HR = 1.86), and PTPRO (p = 0.046, HR = 1.63), where HR = hazard ratio ( Figure 3 ).
[0215] Table 4: List of three genomic regions (three genes) significantly associated with five-year survival prediction in CRC patients
[0216]
[0217] Example 3: Feature evaluation for early detection
[0218] Fourth, targeted hybridization capture-based methylation sequencing analysis was performed on cfDNA extracted from plasma samples of patients with early-stage (I-IIA) colorectal cancer (26 patients) and age- and sex-matched with colonoscopy validation controls (43 patients) to examine the performance of individual markers for survival stratification association markers (Table 4) for early cancer detection.
[0219] The results of the above examples
[0220] Targeted hybrid capture-based methylation sequencing analysis of plasma samples from patients with early-stage (I-IIA) colorectal cancer matched for age and sex to colonoscopy-validated controls showed high accuracy for individual markers, with AUCs of 78% for FGF14, 81% for DPY19L2P1, and 73% for PTPRO ( Figure 3 ).
[0221] Example 4: Large multi-team study shows that cell-free DNA methylation and fragmentation signals can accurately detect patients with early colorectal cancer and advanced adenomas
[0222] This study used cell-free DNA (cfDNA) methylation, fragmentation signatures of selected cancer-associated biomarker regions, tumor-derived signal derivation, and machine learning algorithms to improve a blood test for the early detection of colon cancer and advanced adenoma (AA). The aim of the study was to evaluate the diagnostic accuracy of the test for CRC.
[0223] This was a prospective, international (Spain, Ukraine, Germany, and the United States [part of the NCT04792684 study] population), observational, cohort-based study. Plasma samples were collected from 997 patients before scheduled screening colonoscopy or colon surgery for primary CRC. CFDNA samples were collected from 170 patients with early-stage (stage I-II) and 128 patients with advanced (stage III-IV) CRC (mean age 66 [44-84] years, 48% female, 60% distal cancer), 149 patients with AA (63 with high-grade dysplasia; 84 with low-grade, >1 cm), and 550 age-, sex-, and country-of-origin matched colonoscopy controls. Among the controls, 155 had a negative colonoscopy (cNEG), 337 had a benign colonoscopy with diverticular disease, hemorrhoids, previously undiagnosed gastrointestinal disease, and / or hyperplastic polyps (BEN), and 58 had a colonoscopy with a non-advanced adenoma (NAA). Samples were analyzed using a hybridization capture sequencing approach. A panel of targeted biomarkers had previously been identified through tissue- and plasma-based discovery and validation workflows. Individual cfDNA fragments belonging to each biomarker region were scored for cancer-specific methylation and fragmentation signals. Ultimately, the calculated scores were used to develop a predictive model and to confirm the panel's accuracy.
[0224] By comparing the methylation patterns of cancerous and normal adjacent tissue samples, scores were calculated for each region and each DNA fragment, which were then applied to plasma cfDNA. For each region in the cfDNA, the fragment size was also calculated for each sample and used as input to a machine learning algorithm based on a standard random forest classifier.
[0225] The prediction model used a combination of 82 methylation and fragmentation scores derived from biomarkers belonging to pathways related to cancer development and progression (such as axon guidance, ephrin receptor signaling, epithelial-mesenchymal transition, and FGF signaling), and correctly classified 93% (276 / 298) of CRC patients and 54% (81 / 149) of AA patients. A set of 82 genomic regions shown in Table 1 (SEQ ID No.1-82) was used. The sensitivity for each cancer stage was: 85% (48 / 56) for stage I, 94% (107 / 114) for stage II, 94% (90 / 96) for stage III, and 97% (31 / 32) for stage IV. The fragmentation signal contributed most to early-stage cancer (stage I-II), while the methylation signal had a greater impact on the detection of late-stage cancer (stage III-IV). The sensitivity of high-grade AA with dysplasia was 52% (33 / 63), while the sensitivity of low-grade AA >1 cm was 57% (48 / 84). The model had a specificity of 92% (504 / 550) and correctly identified 83% (48 / 58) of patients with NAA, 93% (312 / 337) of patients with BEN, and 93% (144 / 155) of patients with cNEG. Lesion location, sex, age, BMI, and country of origin were not significantly associated with prediction results (P>0.05). Differential methylation signatures can be identified at the early stage of cancer or even at the precancerous level, and the markers belong to different biological pathways that play an important role in the development and progression of cancer.
[0226] For more detailed biological relevance, see the following examples: DPY19L2P1 promotes glycosyltransferase activity. Altered glycosyltransferase levels and glycosylation patterns are associated with tumorigenesis and metastasis. FGF14 functions as a tumor suppressor by inhibiting the PI3K / AKT / mTOR pathway. PTPRO has prognostic potential and functions as a tumor suppressor in human squamous cell lung carcinoma.
[0227] This study found that using methylation and fragmentation features of cancer-associated cfDNA regions, combined with a machine learning algorithm, had high accuracy for early-stage (stage I-II) CRC (sensitivity of 91%) and AA (sensitivity of 54%) at a specificity of 92%.
[0228] Figure 4The sensitivity and specificity values for the prediction results (advanced adenoma and colorectal cancer) using the 82-member panel (genomic regions in Table 1, Seq. ID No. 1-82) are shown. "AA low" indicates low-grade advanced adenoma with low-grade dysplasia of >= 1 cm. "AA high" indicates advanced adenoma with high-grade dysplasia. "CRC" indicates colorectal cancer. "cNEG" indicates patients without colonoscopy findings. "BEN" indicates patients with benign colonoscopy findings (e.g., diverticula, hemorrhoids, previously undiagnosed gastrointestinal disease, inflammatory and / or hyperplastic polyps). "NAA" indicates non-advanced adenoma.
[0229] Figure 5 is a schematic diagram depicting the data details of the samples used in this Example 4. "CRC" means colorectal cancer. "cNEG" means patients without colonoscopic findings. "BEN" means patients with benign colonoscopic findings (e.g., diverticula, hemorrhoids, previously undiagnosed gastrointestinal diseases, inflammatory and / or hyperplastic polyps). "NAA" means non-advanced adenoma. "AA" means advanced adenoma. "CRC" means colorectal cancer. For the prediction results of each subgroup, p-value analysis was performed using the Benjamini-Hochberg method and false discovery rate (FDR) adjustment to evaluate the differences in prediction results for each demographic subgroup, such as different age ranges, BMI ranges, genders, study countries, lesion locations, advanced adenoma histology, advanced adenoma dysplasia and cancer stage. p>0.05 indicates that the results between subgroups are not significantly different.
[0230] Methods for determining the methylation status of selected markers
[0231] The target nucleic acid can be isolated from the sample by different methods, such as by direct gene capture, for example, by removing the detection inhibitor to produce a clarified sample, capturing the target nucleic acid (if present) from the clarified sample with a capture reagent to form a capture complex, separating the capture complex from the clarified sample, and recovering the target nucleic acid (if present) from the capture complex in a nucleic acid solution, thereby isolating the target nucleic acid from the sample.
[0232] The isolated DNA fragments are amplified using a primer oligonucleotide set and an amplification enzyme. Amplification of several DNA fragments can be performed simultaneously in the same reaction vessel. Amplification can be performed using polymerase chain reaction (PCR). The length of the amplicon is typically 100 to 2000 base pairs.
[0233] After DNA amplification, the DNA can be further treated with a methylation-specific enzyme, and the methylation status can be measured using qPCR or ddPCR. MSRE-qPCR is a method for analyzing the presence of 5-methylcytosine in nucleic acids based on the methylation-specific restriction enzyme method described in Beikircher et al. Methods Mol Biol. 2018; 1708: 407-424 or its variants. Restriction enzymes have been used to study DNA methylation patterns. Methylation-sensitive restriction enzymes (MSREs) containing CpG motifs in their recognition sites are generally used. Because the activity of these enzymes is blocked by 5-methylcytosine, only unmethylated sites are digested, while methylated regions are unaffected. Therefore, after successful digestion, PCR amplification is performed, and only methylated DNA produces detectable PCR products.
[0234] The design of MSRE-qPCR assays is generally less complex than bisulfite-based assays because “native” DNA can be targeted only by satisfying the requirement that primers must cover the target region where at least one MSRE cleavage site is present, but the more cleavage sites, the better the results. CpG-rich regions are often also candidate regions for differential methylation, and these regions are good targets for MSRE-qPCR assay design because they typically contain a large number of suitable MSRE cleavage sites. The use of more than one MSRE (particularly the restriction enzymes AciI, Hin6I, HpyCH4IV, and HpaII) generally provides good coverage of CpG-rich sequences. Sample availability is often a limiting factor, especially in the case of cell-free DNA, but this can be overcome by using pre-amplification during MSRE digestion. Methylation-specific restriction enzyme methods can also be combined with digital PCR.
[0235] Methylation markers can also be detected by using methylation-specific primer oligonucleotides. This technique (MSP) has been described in U.S. Pat. No. 6,265,171 to Herman, which is incorporated herein by reference in its entirety. Amplification of bisulfite-treated DNA using methylation-state-specific primers allows for the differentiation of methylated and unmethylated nucleic acids. An MSP primer pair comprises at least one primer that hybridizes to a bisulfite-treated CpG dinucleotide. Thus, the sequence of the primer comprises at least one CpG dinucleotide. The MSP primer for non-methylated DNA contains a "T" at the C position in the CpG.
[0236] Another method for analyzing nucleic acids for the presence of 5-methylcytosine is based on the bisulfite method described by Frommer et al. for the detection of 5-methylcytosine in DNA (Frommer et al., Proc Natl Acad Sci US A. 1992 Mar 1;89(5):1827-31) or its variants.
[0237] The bisulfite method for mapping 5-methylcytosine is based on the observation that cytosine, but not 5-methylcytosine, reacts with bisulfite ions (also known as bisulfite). Therefore, DNA treated with bisulfite only retains methylated cytosine. According to the prior art, various methylation detection processes can be used in combination with bisulfite treatment. These detection methods are able to determine the methylation status of one or more CpG dinucleotides (e.g., CpG islands) in a nucleic acid sequence. Among other techniques, the detection method also involves sequencing of bisulfite-treated nucleic acids, PCR (sequence-specific amplification), methylation-specific PCR, methylation-specific nuclease-assisted minor allele enrichment PCR, methylation-specific restriction enzyme qPCR, methylation-sensitive high-resolution melting, etc.
[0238] Targeted sequencing-based protocols require PCR amplification of the target region of interest from bisulfite-converted genomic DNA, followed by preparation of DNA sequencing libraries using technologies such as the standard Illumina protocol or the transposase-based NexteraXT technology. Next-generation sequencing (NGS) technology provides CpG methylation level reads with single-base resolution.
[0239] In addition, such as "MethyLight TM "(Fluorescence-based real-time PCR technology) and other detection methods (Campan M et al. Methods Mol Biol. 2018; 1708: 497-513) can be used to assess methylation status. MethyLight is a fluorescence-based quantitative real-time PCR method that can sensitively detect and quantify DNA methylation in candidate genomic regions. MethyLight combines methylation-specific primers and methylation-specific fluorescent probes, making it particularly suitable for detecting low-frequency methylated DNA regions against a high background of unmethylated DNA (Campan M et al. Methods Mol Biol. 2018; 1708: 497-513). In addition, MethyLight can be combined with digital PCR for highly sensitive detection of single methylated molecules for disease detection and screening (Campan M et al. Methods Mol Biol. 2018; 1708: 497-513)
[0240] MSP (methylation-specific PCR) allows the assessment of the methylation status of virtually any group of CpG sites within a CpG island, regardless of the use of methylation-sensitive restriction enzymes (Herman et al., Proc. Natl. Acad. Sci. USA 93:9821-9826, 1996). MSP allows for highly sensitive detection of locus-specific DNA methylation (detection levels are as low as 0.1% of the allele with complete specificity) using PCR amplification of bisulfite-converted DNA. Bisulfite-modified DNA is PCR amplified using specific primer sets, one of which binds specifically to methylated sequences and the other to unmethylated sequences only. MSP results are then obtained using gel electrophoresis without the need for further restriction enzyme digestion or sequencing analysis.
[0241] Quantitative multiplex methylation-specific PCR (QM-MSP) is another method for sensitive quantification of DNA methylation using methylation-specific primers (Fackler MJ, Sukumar S, Methods Mol Biol. 2018; 1708: 473-496). QM-MSP is a two-step PCR method in which, in the first step, a pair of gene-specific primers (forward and reverse) simultaneously amplify methylated and unmethylated copies of the same gene in a single PCR reaction. This methylation-independent amplification step can produce 10 per μL after 36 PCR cycles. 9 In the second step, the amplicons from the first reaction are quantified using real-time PCR and two independent fluorophores (e.g., 6FAM and VIC) with a standard curve to detect methylated / unmethylated DNA for each gene in the same well. One methylated copy of a reference gene can be detected among 100,000 copies.
[0242] Methylation-sensitive high-resolution melting (MS-HRM) (Hussmann D, Hansen LL. Methods Mol Biol. 2018; 1708: 551-571). Methylation-sensitive high-resolution melting (MS-HRM) is a PCR-based in vitro method for detecting methylation levels at specific loci of interest. Unique primer design contributes to high detection sensitivity, enabling the detection of methylated alleles as low as 0.1-1% in an unmethylated background.
[0243] Primers used for MS-HRM detection are designed to be complementary to the methylated allele, and a specific annealing temperature allows these primers to anneal to both methylated and unmethylated alleles, thereby increasing detection sensitivity. Bisulfite treatment of the DNA prior to MS-HRM ensures that the base composition between methylated and unmethylated DNA is distinct, which is used to separate the resulting amplicons by high-resolution melting.
[0244] The MS-NaME (methylation-specific nuclease-assisted minor allele enrichment) reaction (Liu Y et al., Nucleic Acids Res. 2017 Apr 7;45(6):e39.) can be used alone or in combination with one or more of these methods. Minor allele enrichment (MS-NaME) uses double-stranded specific DNA nuclease (DSN) to remove excess DNA with normal methylation patterns. The technique utilizes oligonucleotide probes that simultaneously direct the activity of DSN to multiple targets in bisulfite-treated DNA. Oligonucleotide probes targeting unmethylated sequences create localized double-stranded regions, resulting in digestion of unmethylated targets while leaving methylated targets intact, and vice versa. The target region is then amplified, thereby enriching for the targeted methylated or unmethylated minority epigenetic allele.
[0245] Ms-SNuPE TM Ms-SNuPE can be used to perform strand-specific PCR to generate DNA templates for quantitative methylation analysis using a methylation-sensitive single nucleotide primer extension (MSNP) reaction (Gonzalgo ML and Liang G, Nat Protoc. 2007; 2(8): 1931-6). SNuPE is then performed using oligonucleotides designed to hybridize immediately upstream of the CpG site to be detected. The reaction products are electrophoresed on polyacrylamide gels and visualized and quantified by phosphorimaging analysis.
[0246] The fragments obtained by amplification can also carry labels that can be detected directly or indirectly, such as fluorescent tags, radionuclides, or separable molecular fragments with typical masses that can be detected in a mass spectrometer. Detection can be performed and visualized by methods such as matrix-assisted laser desorption / ionization mass spectrometry (MALDI) or using electron spray mass spectrometry (ESI).
[0247] Colorectal cancer
[0248] In certain embodiments, the methods and compositions of the present invention can be used to detect, diagnose, predict, monitor, screen, stage, and / or provide a survival prognosis for colorectal cancer.
[0249] Colorectal cancer includes colorectal cancer at any of the various possible stages known in the art, including, for example, stage 0, stage I, stage II, stage III, and stage IV colorectal cancer. Colorectal cancer includes all stages of the tumor / node / metastasis (TNM) staging system. With respect to colorectal cancer, T may refer to whether the tumor is confined to the top layer of colorectal duct cells or has not invaded deeper tissues; N may refer to whether the tumor has spread to the lymph nodes, and if so, the number of lymph nodes and the location of the lymph nodes; M may refer to whether the cancer has spread to other parts of the body, and if so, the location and extent of spread. The specific stages of T, N, and M are known in the art. T staging may include TX, T0, Tis, T1, T2, T3, T4; N staging may include NX, N0, N1, N2; M staging may include M0, M1.
[0250] In some cases, the present invention includes screening for early colorectal cancer. Early colorectal cancer may include, for example, colorectal cancer that is localized in the subject's body and has not yet spread to the subject's lymph nodes (e.g., lymph nodes near the cancer) (N0 stage) or to distant sites (M0 stage). Early cancer includes colorectal cancer corresponding to stage 0 to stage II.
[0251] In some cases, the present disclosure includes techniques (e.g., methods, systems) for screening two or more conditions (i.e., a group of conditions) without clearly identifying which condition the subject suffers from. For example, as part of a non-differential diagnosis, the present disclosure includes determining whether the subject suffers from one of a group of conditions. The term "non-differential diagnosis" as used herein refers to a method of identifying that a subject suffers from one of a group of conditions, but not determining which of the conditions in the group the subject suffers from. For example, the methods disclosed herein can screen a group of conditions including CRC (e.g., early CRC) and AA by determining whether the subject suffers from any of the conditions in the group using a detection method (as described herein). Identifying a change in the methylation state (e.g., hypermethylation) of one or more markers (e.g., DMRs) indicates that the subject suffers from CRC (e.g., early CRC), suffers from AA, or suffers from CRC and AA.
[0252] In some cases, the methylation status of one or more markers (e.g., methylation loci) comprising at least one single differentially methylated region (DMR) or a portion of a DMR selected from the 82 DMRs listed in Table 1 (i.e., SEQ ID NOs. 1 to 82) is used to determine that the subject has CRC (e.g., early CRC) or AA, without identifying which disease the subject has. In some cases, the methylation status of one or more markers (e.g., methylation loci) comprising at least one single differentially methylated region (DMR) or a portion of a DMR selected from the DMRs listed in Table 2 is used to determine that the subject has CRC (e.g., early CRC) or AA, without identifying which disease the subject has. In some cases, the methylation status of one or more markers (e.g., methylation loci) comprising at least one single differentially methylated region (DMR) or a portion of a DMR selected from the DMRs listed in Table 3 (i.e., SEQ ID NOs. 70 and 72) is used to determine that the subject has CRC (e.g., early CRC) or AA, without identifying which disease the subject has. In some cases, the methylation status of one or more markers (e.g., methylation loci) comprising at least one single differentially methylated region (DMR) or a portion of a DMR selected from the DMRs listed in Table 4 (i.e., SEQ ID NOs. 1 to 3) is used to determine whether the subject has CRC (e.g., early CRC) or AA, without identifying which disease the subject has.
[0253] In some cases, the present invention includes techniques (e.g., methods, systems) for screening one or more conditions in a subject and determining which condition the subject suffers from. For example, as part of differential diagnosis, the present disclosure includes determining whether a subject suffers from a certain disease. The term "differential diagnosis" as used herein refers to a method for identifying that a subject suffers from a certain disease. For example, the methods disclosed herein are used to screen for CRC (e.g., early-stage CRC) and AA, and determine which disease, if any, the subject suffers from.
[0254] In some cases, the methylation status of one or more markers (e.g., methylation loci) comprising at least one single differentially methylated region (DMR) or a portion of a DMR selected from the 82 DMRs listed in Table 1 (i.e., SEQ ID NOs. 1 to 82) is used to determine that the subject has CRC (e.g., early CRC) or AA. In some cases, the methylation status of one or more markers (e.g., methylation loci) comprising at least one single differentially methylated region (DMR) or a portion of a DMR selected from the DMRs listed in Table 2 is used to determine that the subject has CRC (e.g., early CRC) or AA. In some cases, the methylation status of one or more markers (e.g., methylation loci) comprising at least one single differentially methylated region (DMR) or a portion of a DMR selected from the DMRs listed in Table 3 (i.e., SEQ ID NOs. 70 and 72) is used to determine that the subject has CRC (e.g., early CRC) or AA. In some cases, the methylation status of one or more markers (e.g., methylation loci) comprising at least one single differentially methylated region (DMR) or a portion of a DMR selected from the DMRs listed in Table 4 (i.e., SEQ ID NOs. 1 to 3) is used to determine whether the subject has CRC (e.g., early CRC) or AA.
[0255] The methods and compositions of the present invention can be used to stratify all forms and stages of colorectal cancer, including but not limited to colorectal cancers named herein or known in the art, and all subsets thereof. Thus, one skilled in the art will understand that colorectal cancer referred to herein includes but is not limited to all forms and stages of colorectal cancer, including but not limited to colorectal cancers named herein or known in the art, and all subsets thereof.
[0256] Subjects and samples
[0257] The sample analyzed using the method and composition provided herein can be any biological sample and / or any sample comprising nucleic acid. In various specific embodiments, the sample analyzed using the method and composition provided herein can be a sample from a mammal. In various specific embodiments, the sample analyzed using the method and composition provided herein can be a sample from a human subject. In various specific embodiments, the sample analyzed using the method and composition provided herein can be a sample from a mouse, rat, pig, horse, chicken or cattle.
[0258] In various cases, human subject refers to or seeks to be diagnosed as the subject suffering from colorectal cancer, is diagnosed as or seeks to be diagnosed as the subject with the risk of suffering from colorectal cancer and / or is diagnosed as or seeks to be diagnosed as the subject with the direct risk of suffering from colorectal cancer.In various cases, human subject is the subject being accredited as and needs to screen for colorectal cancer.In some cases, human subject is the subject being accredited as and needs to carry out colorectal cancer screening by a doctor.In various cases, human subject is accredited as and needs to carry out colorectal cancer screening due to age reasons, for example, due to age being equal to or greater than 40 years old, for example, due to age being equal to or greater than 49,45,50,55,60,65,70,75,80,85 or 90 years old, although in some cases, subject more than 18 years old may be accredited as and has risk and / or needs to carry out colorectal cancer screening.In various cases, according to, but not limited to, family medical history, previous diagnosis and / or doctor's assessment, human subject is accredited as and has high risk and / or needs to carry out colorectal cancer screening. In various instances, a human subject is a subject who has not been diagnosed with, is not at risk of, is not immediately at risk of, has not been diagnosed with, and / or is not seeking diagnosis of cancer (eg, colorectal cancer).
[0259] The sample from a subject (e.g., a human or other mammalian subject) can be a sample such as blood, a blood component (e.g., plasma, buffy coat), cfDNA (cell-free DNA), ctDNA (circulating tumor DNA), feces, or tissue. In some specific embodiments, the sample is excreta or body fluid (e.g., feces, blood, plasma, lymph, or urine) or a tissue sample of the subject. The sample of the subject can be a cell or tissue sample, such as a cancer cell of a tumor or metastatic tissue, or a cell or tissue sample of a cancer cell comprising a tumor or metastatic tissue. In various embodiments, a sample from a subject (e.g., a human or other mammalian subject) can be obtained by biopsy or surgery.
[0260] In various specific embodiments, the sample is a cell-free DNA (cfDNA) sample. cfDNA is typically present in biological fluids (e.g., plasma, serum, or urine) in the form of short double-stranded fragments. The concentration of cfDNA is typically low, but can increase significantly under certain conditions, including but not limited to pregnancy, autoimmune diseases, myocardial infarction, and cancer. Circulating tumor DNA (ctDNA) is a component of circulating DNA that specifically comes from cancer cells. ctDNA can be present in human body fluids. For example, in certain cases, ctDNA can be bound to and / or associated with white blood cells and red blood cells. In certain cases, ctDNA can be found not to be bound to and / or associated with white blood cells and red blood cells. Various detection methods for detecting tumor-derived cfDNA are based on detecting genetic or epigenetic modifications specific to cancer (e.g., related cancers). Cancer-specific genetic or epigenetic modifications may include but are not limited to oncogenic mutations or cancer-associated mutations in tumor suppressor genes, activated oncogenes, hypermethylation, and / or chromosomal disorders. Detection of genes or epigenetic modifications specific to cancer or precancer can confirm that the detected cfDNA is ctDNA.
[0261] cfDNA and ctDNA provide real-time or near-real-time measurements of the methylation status of the tissue of origin. cfDNA and ctDNA have a half-life of approximately 2 hours in the blood, so a sample collected at a given time can reflect the status of the tissue of origin relatively promptly.
[0262] Various methods for isolating nucleic acids from samples (e.g., isolating cfDNA from blood or plasma) are known in the art. Isolation of nucleic acids can be performed using standard DNA purification techniques, such as, but not limited to, direct gene capture techniques (e.g., by clarifying the sample to remove detection inhibitors, and capturing the target nucleic acid (if present) from the clarified sample with a capture agent to produce a capture complex, and separating the capture complex to recover the target nucleic acid).
[0263] In certain embodiments, a sample may have a minimum amount of DNA (e.g., cfDNA, gDNA) (e.g., DNA fragments) required for subsequent determination of methylation status. For example, in certain embodiments, a sample may be required to have at least 5 ng, 10 ng, 20 ng (or more) of DNA.
[0264] Methods for measuring methylation status
[0265] Methylation status can be measured by various methods known in the art and / or the methods provided herein. Those skilled in the art will understand that methods for measuring methylation status are generally applicable to samples of any origin and type, and will further be aware of the processing steps that can be used to modify a sample into a form suitable for measurement by a given method.
[0266] In some embodiments, the processing step involves fragmenting or shearing the DNA of the sample. For example, genomic DNA (e.g., gDNA) obtained from cells, tissues, or other sources may need to be fragmented before sequencing. In some embodiments, DNA can be fragmented using physical methods (e.g., using an ultrasonicator, atomizer technology, fluid dynamic shearing, etc.) before methylation status is measured. In some embodiments, DNA can be fragmented using enzymatic methods (e.g., using endonucleases or transposases). Some samples (e.g., cfDNA samples) may not require fragmentation. cfDNA fragments are about 200bp in length and are suitable for use with some of the methods provided herein. DNA fragments of about 100-1000bp in length are suitable for analysis with some of the NGS technologies described herein, including, for example, those based on Some technologies may require DNA fragments of approximately 100-1000 bp. In contrast, DNA fragments larger than approximately 10 kb are suitable for long-read sequencing technologies.
[0267] Methods for measuring methylation status include, but are not limited to, whole genome bisulfite sequencing, targeted bisulfite sequencing, targeted enzymatic methylation sequencing, methylation state-specific polymerase chain reaction (PCR), methods including mass spectrometry, methylation arrays, methods including methylation-specific nucleases, methods including mass-based separation, methods including target-specific capture (e.g., hybrid capture), and methods including methylation-specific oligonucleotide primers. Certain specific methylation detection methods utilize bisulfite reagents (e.g., bisulfite ions) or enzymatic conversion reagents (e.g., Tet methylcytosine dioxidase 2).
[0268] Bisulfite reagent can especially comprise bisulfite, bisulfite, sodium metabisulfite or its combination, and these reagents can be used to distinguish methylated and unmethylated nucleic acid.Bisulfite is to the interaction of cytosine and 5-methylcytosine differently.In conventional bisulfite method, DNA (such as single-stranded DNA, double-stranded DNA) is contacted with bisulfite so that unmethylated cytosine is deaminated (such as converted) to uracil, and methylated cytosine is unaffected.Methylated cytosine is selectively retained, and unmethylated cytosine does not.Therefore, in the sample that bisulfite is treated, uracil residue replaces unmethylated cytosine residue, thereby provides the identification signal of unmethylated cytosine residue, and the remaining (methylated) cytosine residue provides the identification signal of methylated cytosine residue.The sample that bisulfite is treated can be analyzed by next generation sequencing (NGS) or other methods disclosed herein.
[0269] In some embodiments, the bisulfite treated sample can be treated using a bisulfite to DNA ratio of at least 100 Å. In some embodiments, the bisulfite treated sample comprises single-stranded DNA fragments or double-stranded DNA fragments.
[0270] In some embodiments, bisulfite treatment includes performing one or more denaturation-conversion cycles on DNA fragments (e.g., double-stranded DNA) to convert unmethylated cytosine in the DNA fragments into uracil. Denaturation converts the double-stranded DNA fragments in the sample into single-stranded DNA fragments. Conversion converts unmethylated cytosine in the single-stranded DNA into uracil. In some embodiments, only one denaturation-conversion cycle is performed. In some embodiments, two, three, four, five, six, seven, eight, nine, ten, fifteen, or more than twenty denaturation-conversion cycles are performed. In some embodiments, the temperature of the denaturation step is performed at a temperature of about 80-100° C. (e.g., about 90-97° C., such as about 96° C.). In some embodiments, the denaturation step is performed for less than 10 minutes (e.g., less than 5 minutes, less than 5 minutes, less than 2 minutes or less). In certain embodiments, the conversion step is performed for less than 2.5 hours (e.g., less than 2 hours, less than 1 hour, less than 30 minutes, less than 15 minutes or less). In certain embodiments, the conversion step is performed at a temperature of 55 to 65°C. In certain embodiments, after the denaturation-conversion cycle, the converted DNA fragments can be stored at a temperature of about 4°C. In certain embodiments, bisulfite treatment can be performed before library preparation. In certain embodiments, bisulfite treatment can be performed after library preparation.
[0271] The enzymatic conversion reagent may include Tet methylcytosine dioxidase 2 (TET2). TET2 oxidizes 5-methylcytosine, thereby protecting it from the continuous deamination of APOBEC. APOBEC deaminates unmethylated cytosine to uracil, while oxidized 5-methylcytosine is unaffected. Therefore, in TET2-treated samples, uracil residues replace unmethylated cytosine residues, thereby providing an identification signal for unmethylated cytosine residues, while the remaining (methylated) cytosine residues provide an identification signal for methylated cytosine residues. The sample treated with TET2 can be analyzed by, for example, next generation sequencing (NGS). In certain embodiments, APOBEC refers to a member (or multiple members) of the apolipoprotein B mRNA editing catalytic polypeptide-like (APOBEC) family. In certain embodiments, APOBEC may refer to APOBEC-1, APOBEC-2, APOBEC-3A, APOBEC-3B, APOBEC-3C, APOBEC-3D, APOBEC-3E, APOBEC-3F, APOBEC-3G, APOBEC-3H, APOBEC-4, and / or activation-induced (cytidine) deaminase (AID).
[0272] Methods for measuring methylation status may include, but are not limited to, massively parallel sequencing (e.g., next generation sequencing) to determine methylation status, such as sequencing by synthesis, real-time (e.g., single molecule) sequencing, bead emulsion sequencing, nanopore sequencing, or other sequencing technologies known in the art. In some embodiments, methods for measuring methylation status may include whole genome sequencing, for example, measuring whole genome methylation status at base pair resolution from bisulfite or enzyme-treated material.
[0273] In some embodiments, methods of measuring methylation status include reduced representation bisulfite sequencing, for example, using restriction enzymes to measure the methylation status of high CpG content regions at base pair resolution in bisulfite or enzyme-treated material.
[0274] In some embodiments, methods of measuring methylation status can include targeted sequencing, for example, measuring the methylation status of pre-selected genomic locations at base pair resolution.
[0275] In some embodiments, the pre-selected (capture) (e.g., enrichment) of the region of interest (e.g., DMR) can be accomplished by complementary in vitro synthesized oligonucleotide sequences (e.g., capture bait / probes). Capture probes (e.g., oligonucleotide capture probes, oligonucleotide capture baits) can be used for targeted sequencing (e.g., NGS) technology of the region of particular interest in enrichment oligonucleotide (e.g., DNA) sequences. For example, when sequencing the sequence of a particular predetermined region of DNA, the enrichment target region is very useful. In some embodiments, the capture probe is about 10 to 1000bp long (e.g., about 10bp to about 200bp long) (e.g., about 120bp long). In certain embodiments, the target of one or more capture probes is to capture the region of interest (e.g., genomic markers) corresponding to one or more methylation loci (e.g., a methylation loci comprising at least a portion of one or more DMRs). In certain embodiments, capture probes target hypomethylated or hypermethylated methylation loci. For example, capture probes can target specific methylation loci. However, if the DNA fragments corresponding to the methylated loci are converted (e.g., bisulfite conversion or enzymatic conversion) before being enriched with capture probes, the sequences of the converted DNA fragments may change as described herein due to the specific cytosine residues that are not methylated. Therefore, if cytosine is hypomethylated, targeting the unconverted DNA region may result in some mismatches. Although capture probe-target sequence hybridization can tolerate some mismatches, a second probe may be needed to enrich for the hypomethylated DNA region.
[0276] In some embodiments, the ability of multiple regions of the genome that capture probe (for example, before sequencing) targeting is paid close attention to is assessed. For example, when designing the capture probe of the region (for example DMR) that targeting is particularly concerned, the ability of multiple regions of capture probe targeting genome can be considered. As discussed herein, the mispairing in pairing (for example non-Watson-Crick pairing) makes the capture probe hybridize to other unexpected regions of genome. In addition, specific target sequence may repeat somewhere else in the genome. Repetitive sequences are common in highly repeated sequences. In some embodiments, capture probe is designed to only target several similar regions in the genome. In some embodiments, capture probe can hybridize to similar regions below 500, below 100, below 50, below 10, below 5 in the genome. In some embodiments, the region similar to the target in the region paid close attention to is calculated using the 24bp window moving in the whole genome, and according to sequence order similarity, the window region is matched with the reference sequence. Window and / or technology of other sizes can be used.
[0277] For example, one or more DNA fragments (e.g., ctDNA, fragmented gDNA) can be hybridized and captured using capture probes targeting predetermined regions of interest in a genome. In certain embodiments, the capture probes target at least 2 (e.g., at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 50, 75, 100, 150 or more) predetermined regions of interest (e.g., genomic markers, such as DMRs). In certain embodiments, the capture probes overlap. In certain embodiments, the overlapping probes overlap by at least 10%, 20%, 30%, 40%, 50%, or 60%.
[0278] In some embodiments, the capture probe is a nucleic acid probe (e.g., a DNA probe, an RNA probe). In some embodiments, the method can also include identifying a mutation region (e.g., a single nucleotide base) using targeted sequencing, such as determining whether a mutation exists at one or more pre-selected genomic locations (e.g., a genomic marker, such as a mutation marker). In some embodiments, mutations can also be identified at base pair resolution from bisulfite or enzyme-treated DNA.
[0279] In some embodiments, a method for measuring methylation status can include an Illumina methylation assay, for example, which quantitatively measures over 85,000 methylation sites across the genome at single nucleotide resolution.
[0280] Various methylation detection protocols can be used in conjunction with bisulfite treatment to determine the methylation status of target sequences (e.g., DMRs). These detection methods can include, among others, methylation-specific restriction enzyme qPCR, sequencing of bisulfite-treated nucleic acids, PCR (e.g., sequence-specific amplification), methylation-specific nuclease-assisted minor allele enrichment PCR, and methylation-sensitive high-resolution melting. In some embodiments, DMRs are amplified from transformed (e.g., bisulfite or enzyme-converted) DNA fragments for library preparation.
[0281] In some embodiments, the method can be based on, for example, the Illumina protocol, Methyl-SeqDNA library kit (Swift Bioscience) scheme, Nextera XT scheme based on transposition, etc., use oligonucleotide fragments (such as cfDNA, gDNA fragments, synthetic nucleotide sequences, etc.) converted (such as bisulfite or enzyme conversion) to prepare sequencing libraries. In certain embodiments, oligonucleotide fragments are DNA fragments that have been converted (such as bisulfite conversion or enzyme conversion). In certain embodiments, the DNA fragments used to prepare sequencing libraries can be single-stranded DNA fragments or double-stranded DNA fragments. In certain embodiments, the library can be prepared by connecting a joint (adapter) to the DNA fragments. The joint comprises a short sequence (such as about 100bp to about 1000bp) (such as an oligonucleotide sequence), which allows the oligonucleotide fragments of the library (such as a DNA library) to combine and generate clusters in a flow cell for, for example, next-generation sequencing (NGS). Before NGS, the joint can be connected to the library fragment. In certain embodiments, the joint is covalently linked to the library fragment by a ligase. In certain embodiments, the joint is connected to the 5' end and 3' end of the converted DNA fragment. In certain embodiments, the attaching step is performed such that at least 40%, at least 50%, at least 60%, at least 70% of the transformed DNA fragments are attached to the adaptor. In certain embodiments, the attaching step is performed such that at least 40%, at least 50%, at least 60%, at least 70% of the transformed DNA fragments have an adaptor attached at both the 5' end and the 3' end.
[0282] In certain embodiments, the joint used herein comprises an oligonucleotide sequence that contributes to sample identification. For example, in certain embodiments, the joint comprises a sample index. The sample index is a short sequence (for example, about 8 to about 10 bases) of a nucleic acid (for example, DNA, RNA) that plays a sample identifier role and especially allows multiplexing and / or pooling multiple samples on a flow cell (for example, for NGS technology). In certain embodiments, the joint at the 5' end, 3' end, or both ends of the converted single-stranded DNA fragment comprises a sample index. In certain embodiments, the joint sequence may include a molecular barcode. The molecular barcode may serve as a unique molecular identifier for identifying a target molecule in a process such as DNA sequencing. In certain embodiments, the DNA barcode may be randomly generated. In certain embodiments, the DNA barcode may be predetermined or pre-designed. In certain embodiments, the DNA barcode is different on each DNA fragment. In certain embodiments, the DNA barcode may be identical for two single-stranded DNA fragments that are not complementary to each other (for example, Watson-Crick pairing) in a biological sample. In certain embodiments, the DNA fragment may be amplified (for example, using PCR) after the joint is connected to the DNA fragment. In certain embodiments, at least 40% (eg, at least 50%, at least 60%, at least 70%) of the transformed DNA fragments have adapters attached at both the 5' and 3' ends.
[0283] In certain embodiments, high-throughput and / or next-generation sequencing (NGS) technologies are used to achieve base pair-level resolution of oligonucleotide (e.g., DNA) sequences, allowing analysis of methylation status and / or identification of mutations. For example, in certain embodiments, NGS may include single-end or paired-end sequencing. In single-end sequencing, a technique reads sequencing fragments in one direction—from one end of the fragment to the other end of the fragment. In certain embodiments, this will produce a single DNA sequence that can then be compared with a reference sequence. In paired-end sequencing, sequencing fragments are read in a first direction from one end of the fragment to the other end. Sequencing fragments can be read until a specified read is reached. Then, sequencing fragments are read in a second direction opposite to the first direction. In certain embodiments, preserving multiple read pairs can help improve read alignment and / or identify mutations (e.g., insertions, deletions, inversions, etc.) that may not be detected by single-end reads.
[0284] Another method that can be used for methylation detection includes PCR amplification using methylation-specific oligonucleotide primers (MSP method), for example, as applied to bisulfite-treated samples (see, for example, Herman 1992 Proc. Natl. Acad. Sci. USA 93:9821-9826, which is incorporated herein by reference for methods of determining methylation status). Amplification of bisulfite-treated DNA using methylation-specific oligonucleotide primers allows for the discrimination of methylated nucleic acids from unmethylated nucleic acids. The oligonucleotide primer pairs used in the MSP method include at least one oligonucleotide primer that hybridizes to a sequence containing a methylation site (e.g., a CpG site). An oligonucleotide primer containing a T residue at a position complementary to a cytosine residue will selectively hybridize to a template in which the cytosine was unmethylated prior to bisulfite treatment, while an oligonucleotide primer containing a G residue at a position complementary to a cytosine residue will selectively hybridize to a template in which the cytosine was methylated prior to bisulfite treatment. MSP results can be obtained with or without sequencing the amplicons, for example using gel electrophoresis. MSP (methylation-specific PCR) allows for highly sensitive detection of locus-specific DNA methylation (detection levels down to 0.1% of the allele, with complete specificity) by using PCR amplification of bisulfite-converted DNA.
[0285] Another method that can be used to determine the methylation status of a sample after bisulfite treatment is methylation-sensitive high-resolution melting (MS-HRM) PCR (see, for example, Hussmann 2018 Methods Mol Biol. 1708: 551-571, which is incorporated herein by reference for methods of determining methylation status). MS-HRM is an in vitro PCR-based method that can detect the methylation level of a specific locus of interest based on hybrid melting. Bisulfite treatment of the DNA prior to MS-HRM ensures that the base composition between methylated and unmethylated DNA is different, which separates the resulting amplicons by high-resolution melting. The unique primer design contributes to the high detection sensitivity, enabling the detection of as little as 0.1-1% methylated alleles in an unmethylated background. The oligonucleotide primers used for MS-HRM detection are designed to be complementary to the methylated allele, and the specific annealing temperature enables these primers to anneal to both methylated and unmethylated alleles, thereby increasing the sensitivity of the detection.
[0286] Another method that can be used to determine the methylation status of a sample after bisulfite treatment is quantitative multiplex methylation-specific PCR (QM-MSP). QM-MSP uses methylation-specific primers to perform sensitive quantitative analysis of DNA methylation (e.g., see Fackler 2018 Methods Mol Biol. 1708: 473-496, which is incorporated herein by reference for methods of determining methylation status). QM-MSP is a two-step PCR method in which, in the first step, a pair of gene-specific primers (forward and reverse) simultaneously multiplex amplify methylated and non-methylated copies of the same gene in one PCR reaction. This methylation-independent amplification step can produce 10 per μL after 36 PCR cycles. 9 In the second step, the amplicons from the first reaction are quantified using real-time PCR and two independent fluorophores (e.g., 6FAM and VIC) with a standard curve to detect methylated / unmethylated DNA for each gene in the same well. One methylated copy can be detected among 100,000 copies of the reference gene.
[0287] Another method that can be used to determine methylation status after bisulfite treatment of a sample is methylation-specific nuclease-assisted minor allele enrichment (MS-NaME) (see, for example, Liu 2017 Nucleic Acids Res. 45(6):e39, which is incorporated herein by reference for methods of determining methylation status). Ms-NaME is based on the selective hybridization of a probe to a target sequence in the presence of a DNA nuclease specific for double-stranded (ds) DNA (DSN), such that hybridization produces a double-stranded DNA region that is subsequently digested by DSN. Thus, oligonucleotide probes targeting unmethylated sequences produce local double-stranded regions that result in digestion of the unmethylated target; oligonucleotide probes capable of hybridizing to methylated sequences produce local double-stranded regions that result in digestion of the methylated target, while the methylated target remains intact. In addition, oligonucleotide probes can simultaneously direct DSN activity to multiple targets in bisulfite-treated DNA. Subsequent amplification can enrich for undigested sequences. Ms-NaME can be used alone or in combination with other techniques provided herein.
[0288] Another method that can be used to determine the methylation status of samples after bisulfite treatment is the methylation-sensitive single nucleotide primer extension method (Ms-SNuPE TM), (see, e.g., Gonzalgo 2007 Nat Protoc. 2(8): 1931-6, which is incorporated herein by reference for methods of determining methylation status). In Ms-SNuPE, strand-specific PCR is performed using Ms-SNuPE to generate a DNA template for quantitative methylation analysis. SNuPE is then performed with oligonucleotides designed to hybridize immediately upstream of the CpG site to be detected. The reaction products can be electrophoresed on a polyacrylamide gel and visualized and quantified by phosphorimage analysis. The amplicons can also carry a label that can be detected directly or indirectly, such as a fluorescent label, a radionuclide, or a separable molecular fragment or other entity whose mass can be distinguished by mass spectrometry. Detection can be performed and / or visualized by methods such as matrix-assisted laser desorption / ionization mass spectrometry (MALDI) or using electron spray mass spectrometry (ESI).
[0289] Some methods that can be used to determine the methylation status of a sample after bisulfite treatment utilize a first oligonucleotide primer, a second oligonucleotide primer, and an oligonucleotide probe in an amplification-based method. For example, the oligonucleotide primers and probes can be used in real-time polymerase chain reaction (PCR) or droplet digital PCR (ddPCR) methods. In various cases, the first oligonucleotide primer, the second oligonucleotide primer, and / or the oligonucleotide probe selectively hybridize to methylated DNA and / or unmethylated DNA, such that the amplification or probe signal indicates the methylation status of the sample.
[0290] Other bisulfite-based methods for detecting methylation status (e.g., the presence or absence of 5-methylcytosine levels) have been disclosed, for example, in Frommer (1992 Proc Natl Acad Sci US A. 1;89(5):1827-31, which is incorporated herein by reference for methods of determining methylation status).
[0291] In some MSRE-qPCR embodiments, the total amount of DNA in an aliquot is measured in native (eg, undigested) form using, for example, real-time PCR or digital PCR.
[0292] Various amplification techniques can be used alone to detect methylation status or in combination with other techniques described herein. After reading this specification, those skilled in the art will understand how to combine various amplification techniques known in the art and / or various other methylation status determination techniques described herein. Amplification techniques include, but are not limited to, PCR, such as quantitative PCR (qPCR), real-time PCR, and / or digital PCR. Those skilled in the art will understand that polymerase amplification can perform multiple amplifications of multiple targets in a single reaction. PCR amplicons are typically 100 to 2000 base pairs in length. In various cases, amplification techniques are sufficient to determine methylation status.
[0293] The method based on digital PCR (dPCR) is to divide the sample into the wells of 96-well, 384-well or more well plates, or for example, to divide into single emulsion droplets (ddPCR) using microfluidic devices, so that some wells contain one or more template copies, while other wells do not contain template copies. Thus, before amplification, the average number of template molecules in each well is less than one. The number of wells in which the template is amplified can provide a measure of the concentration of the template. If the sample has been contacted with MSRE, the number of wells in which the template is amplified can provide a measure of the concentration of the methylated template.
[0294] In various embodiments, a MethyLight TM Fluorescence-based real-time PCR detection methods such as Campan 2018 Methods Mol Biol. 1708: 497-513, which is incorporated herein by reference for methods of determining methylation status. MethyLight is a fluorescence-based quantitative real-time PCR method that can sensitively detect and quantify DNA methylation in candidate genomic regions. MethyLight combines methylation-specific primers and methylation-specific fluorescent probes, and is particularly suitable for detecting low-frequency methylated DNA regions against a high background of unmethylated DNA. In addition, MethyLight can also be used in combination with digital PCR for highly sensitive detection of single methylated molecules for disease detection and screening.
[0295] The method based on real-time PCR for determining methylation state generally comprises the step of the standard curve according to the analysis generation of unmethylated DNA to external standard.Standard curve can be made up of at least two points, and allows the real-time Ct value of digested DNA and / or the real-time Ct value of undigested DNA to be compared with known quantitative standard.In particular cases, can determine by the sample Ct value of MSRE digestion and / or undigested sample or aliquot sample, and calculate the genome equivalent of DNA according to standard curve.Can assess through the Ct value of MSRE digestion and undigested DNA, to differentiate digested amplicon (for example effectively digestion; For example, produce 45 Ct value).Also can identify the amplicon that does not all increase under digestion or undigested condition.Then can directly compare the correction Ct value of the amplicon being paid close attention to between different conditions, to set up the relative difference of methylation state under different conditions.Alternatively, the △ difference between the Ct value of digested DNA and undigested DNA also can be used for setting up the relative difference of methylation state between different conditions.
[0296] In certain specific embodiments, targeted bisulfite sequencing (e.g., using hybridization capture), among other techniques, is particularly useful for determining the methylation state of a methylation biomarker for a disease and / or condition. For example, a methylation biomarker for colorectal neoplasms (e.g., advanced adenoma and / or colorectal cancer) is or includes a single methylation locus. In certain specific embodiments, targeted bisulfite sequencing, among other techniques, is useful for determining the methylation state of a methylation biomarker that is or includes two or more methylation loci.
[0297] Those skilled in the art will understand that in embodiments where the methylation status of multiple methylation loci (e.g., multiple DMRs) is analyzed using the colorectal cancer screening method provided herein, the methylation status of each methylation locus can be measured or represented in any of a variety of forms, and the methylation status of multiple methylation loci (preferably each methylation locus measured and / or represented in the same, similar, or comparable manner) can be analyzed or represented together or cumulatively in any of a variety of forms. In various embodiments, the methylation status of each methylation locus can be measured as a methylation portion. In various embodiments, the methylation status of each methylation locus can be expressed as a percentage value of methylation reads to total sequencing reads compared to a reference sample. In various embodiments, the methylation status of each methylation locus can be expressed as a qualitative comparison with a reference, for example, identifying each methylation locus as hypermethylated or hypomethylated.
[0298] In some embodiments of analyzing a single methylation locus, hypermethylation of the single methylation locus constitutes a diagnosis that the subject has or may have a certain disease (e.g., cancer) (e.g., advanced adenoma, colorectal cancer), while the absence of hypermethylation of the single methylation locus constitutes a diagnosis that the subject may not have the certain disease. In some embodiments, hypermethylation of a single methylation locus (e.g., a single DMR) among a plurality of analyzed methylation loci constitutes a diagnosis that the subject has or may have a disease, while the absence of hypermethylation of any methylation locus among the plurality of analyzed methylation loci constitutes a diagnosis that the subject may not have the disease. In some embodiments, hypermethylation of a determined percentage (e.g., a predetermined percentage) of methylation loci (e.g., at least 10% (e.g., at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90% or 100%) of a plurality of analyzed methylation loci constitutes a diagnosis that the subject has or is likely to have the disorder, while the absence of hypermethylation of a determined percentage (e.g., a predetermined percentage) of methylation loci (e.g., at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90% or 100%) of a plurality of analyzed methylation loci constitutes a diagnosis that the subject is likely not to have the disorder. In some embodiments, hypermethylation of a determined number (e.g., a predetermined number) of methylation loci (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 50, 100, 150 or more DMRs) out of a plurality of analyzed methylation loci (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 50, 100, 150 or more DMRs) constitutes a risk factor for a subject having or being at risk for the disorder. 150 or more DMRs) out of a plurality of analyzed methylation loci (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 50, 100, 150 or more DMRs) constitutes a diagnosis that the subject has or is likely not suffering from the disorder.
[0299] In some embodiments, the methylation state of multiple methylation loci (e.g., multiple DMRs) is qualitatively or quantitatively measured, and the measured values for each of the multiple methylation loci are combined to provide a diagnosis. In some embodiments, the quantitatively measured methylation state of each of the multiple methylation loci is individually weighted, and the weighted values are combined to provide a single value that can be compared to a reference, thereby providing a diagnosis.
[0300] In some embodiments, the methylation state may include determining methylated and / or unmethylated reads mapped to a genomic region (e.g., DMR). For example, when using a specific sequencing technology disclosed herein (e.g., NGS, whole genome bisulfite sequencing, etc.), sequence reads are generated. A sequence read is an inferred sequence (e.g., a probabilistic sequence) of base pairs corresponding to all or part of a sequenced oligonucleotide (e.g., DNA) fragment (e.g., cfDNA fragment, gDNA fragment). In certain embodiments, a reference sequence (e.g., a bisulfite-converted reference sequence) may be used to map (e.g., align) the sequence reads to a specific region of interest to determine whether there are any changes or variations in the reads. Changes may include methylation and / or mutations. The region of interest may include one or more genomic markers, including methylation markers (e.g., DMR), mutation markers, or other markers disclosed herein.
[0301] For example, in the case of DNA fragments treated with bisulfite or enzymes, the treatment converts unmethylated cytosine to uracil, while methylated cytosine is not converted to uracil. Therefore, the sequence reads generated by a DNA fragment with methylated cytosine will be different from the sequence reads generated by the same DNA fragment without methylated cytosine. Methylation at sites where a cytosine nucleotide is followed by a guanine nucleotide (e.g., CpG sites) may be of particular concern.
[0302] Quality Control Program
[0303] In certain embodiments, a quality control step may be implemented. The quality control step is used to determine whether a particular step or process is implemented within a specific parameter range. In certain embodiments, the quality control step can be used to determine the validity of a given analysis result. In addition, the quality control step can also be used to determine the quality of sequencing data. For example, the quality control step can be used to determine the read coverage of one or more DNA regions. Quantitative indicators of quality control include but are not limited to AT loss rate, GC loss rate, bisulfite conversion rate (e.g., bisulfite conversion efficiency), etc. Failure to meet a threshold quality control condition (e.g., minimum conversion rate, maximum CG loss rate, etc.) may indicate that, for example, one or more conversion steps were not performed within the appropriate parameter range.
[0304] For example, in the methods described herein, various steps of the transformation protocol can be optimized to reduce AT and / or CG loss rates. As will be appreciated by those skilled in the art, AT and GC loss metrics indicate the extent of undercoverage of a particular target region based on AT or GC content. In certain embodiments, samples with lower GC loss rates can help identify which samples have been properly processed. For example, finding a GC loss rate of less than 10%, less than 9%, less than 8%, less than 7%, less than 6%, less than 5%, less than 4% or less can help identify samples that have been properly processed.
[0305] In some embodiments, the quality control step may include determining the hit / miss-target ratio. Sequence reads aligned with the region of interest (e.g., DMR) are considered hits, while sequence reads not aligned with the region of interest (e.g., DMR) are considered miss-targets. In some embodiments, the hit ratio is expressed as the percentage of hit bases accounting for the total number of aligned bases. In some embodiments, the hit ratio is expressed as the percentage of hit bases and near-hit bases accounting for the total number of aligned bases. Near-hit bases can be bases within the scope of a certain number of bases (e.g., within 500bp, within 200bp, within 100bp) of the target region. In some embodiments, for a sequencing experiment passed through quality control, the hit ratio is at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99% or higher. In some embodiments, the miss-target ratio is expressed as the percentage of miss-target bases accounting for the total number of aligned bases. In certain embodiments, the off-target rate of a sequencing experiment that passes quality control is less than 95%, less than 90%, less than 85%, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, or less than 1%.
[0306] In some embodiments, the quality control step may include determining the quality score of the mapped sequence reads. The quality score is a value that quantifies the probability of the sequence reads being mismapped. For example, when mapping short sequences or repetitive sequences, it is feasible that a sequence will be mapped to multiple positions in the reference genome. The quality score takes into account the best alignment of the sequence reads with the reference genome compared to other possible alignments of the sequence reads with the reference genome. In some embodiments, the quality score is a mapping quality (MAPQ) score. MAPQ is the negative logarithmic probability that a mismatch occurs in a read. A high score indicates that the confidence level of the read is correctly aligned, while a low score indicates that the confidence level of the read is correctly aligned. In some embodiments, the MAPQ score can be calculated using the following equation:
[0307] MAPQ score = -10log 10 Pr{Mapping location error}.
[0308] In certain embodiments, a MAPQ score is rounded to the nearest integer. In certain embodiments, Pr is the probability that a sequence read obtained by an alignment (e.g., mapping) tool is incorrectly mapped. In certain embodiments, a scaling factor is 1 (rather than 10) or another number.
[0309] Manual spike control
[0310] Control nucleic acid (e.g., DNA) molecules (e.g., "spiking controls") can be used to assess or estimate the efficiency of conversion of unmethylated and methylated cytosine to uracil. Control nucleic acid molecules can be used in sequencing methods involving conversion of DNA samples (e.g., bisulfite or enzymatic conversion).
[0311] When DNA is carried out conversion as described herein (such as bisulfite or enzyme conversion), conversion may be incomplete. That is, some unmethylated cytosines may not be converted into uracil. If conversion is incomplete so that unmethylated cytosine is not mostly converted, then when DNA is checked for order, unconverted unmethylated cytosine may be identified as methylated cytosine. Therefore, in order to determine whether bisulfite conversion is complete, the control DNA molecule can be converted together with the DNA fragmentation from the sample. In some embodiments, the control DNA molecule converted is checked for order (for example, using NGS technology as described herein) and multiple control sequence reads will be produced. The control sequence read can be used to determine the conversion rate that unmethylated and / or methylated cytosine is converted into uracil.
[0312] The prior art does not recognize that it is useful to include a control (e.g., a control DNA molecule) in each sample. Instead, they assume that in a given run, the conversion efficiency remains relatively consistent between samples. However, studies have found that the conversion rate of unmethylated cytosine in DNA fragments to uracil may vary significantly from one sample to another. For example, in a single batch of processed samples, the conversion efficiency may be between 10% and 110%. It is noted that there may be over-conversion, so that the conversion efficiency can be greater than 100%, for example, when 10% of the methylated cytosine is converted, the conversion efficiency is 110%. In some embodiments, the conversion efficiency is 30% to 110%. In other embodiments, the conversion efficiency is 50% to 100%.
[0313] In some embodiments, after fragmentation using bisulfite or enzyme reagents and before conversion, a control DNA molecule can be added to the sample. In some embodiments, a plurality of (e.g., two, three, four or more) control DNA sequences can be added to the DNA fragments of the sample. The control DNA molecule can be a known sequence. For example, the sequence, methylated base number, and unmethylated base number of the control sequence have been determined before the control DNA molecule is added to the sample. In some embodiments, the control sequence can be a DNA sequence produced in vitro to contain artificial methylated or unmethylated nucleotides (e.g., methylated cytidine). In some embodiments, the control sequence can be a DNA sequence produced to contain fully unmethylated DNA nucleotides.
[0314] The high conversion efficiency of the spike-in control sequence can be used to infer the conversion efficiency of DNA fragments that undergo the same conversion process as the spike-in control sequence. For example, in an unmethylated spike-in control DNA sequence, at least 98% of the unmethylated cytosines are deaminated, indicating that the conversion efficiency is high and the sample can pass the quality control assessment. In certain embodiments, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the unmethylated cytosines in multiple DNA fragments of the control DNA sequence are converted to uracil. High conversion efficiency is very important because when the DNA is subjected to bisulfite or enzyme treatment, it is ideal that all (or nearly all) unmethylated cytosines are converted to uracil. As mentioned above, unconverted unmethylated cytosines can become a source of noise in the data.
[0315] In addition, when using the conversion method to process DNA, it is disadvantageous to convert methylated cytosine into uracil. The conversion of methylated cytosine in the spike-in control indicates that the methylated cytosine has been converted into uracil in the DNA sample treated in the same manner as the methylated spike-in control. The methylated cytosine in the methylated spike-in control should not be converted into uracil. Due to the same reasons as above, the conversion of methylated cytosine into uracil may lead to the misidentification of so-called unmethylated cytosine during the methylation analysis. In certain embodiments, at most 5%, at most 4%, at most 3%, at most 2% or at most 1% of the methylated cytosine in the multiple DNA fragments of the control DNA sequence is converted into uracil. For example, the deamination of up to 2% of the methylated cytosine in the methylated spike-in control DNA sequence indicates that the conversion efficiency is high and the sample can pass the quality control assessment.
[0316] application
[0317] Method and composition of the present disclosure can be used for any in various applications.For example, method and composition of the present disclosure can be used for screening or auxiliary screening certain disease (for example colorectal cancer).In various cases, the colorectal cancer of any stage can be detected using the screening of method and composition of the present disclosure, including but not limited to early colorectal cancer.In some embodiments, the screening using method and composition of the present disclosure is applied to the individuality of more than 40 years old, for example, the individuality of more than 40,45,50,55,60,65,70,75,80,85 or 90 years old.Especially, the individuality of 40 years old or more than 40 years old is that colorectal cancer screening is paid close attention to.In some embodiments, the screening using method and composition of the present disclosure is applied to the individuality of more than 18 years old, for example, the individuality of more than 18,20,25,30,35,40,45,50,55,60,65,70,75,80,85 or 90 years old.In some embodiments, the screening using method and composition of the present disclosure is applied to the individuality of 18 to 40 years old. In various embodiments, screening using the methods and compositions of the present disclosure is applied to individuals experiencing abdominal pain or discomfort, e.g., individuals experiencing undiagnosed or incompletely diagnosed abdominal pain or discomfort. In various embodiments, screening using the methods and compositions of the present disclosure is applied to individuals without symptoms that may be associated with cancer. Thus, in certain embodiments, screening using the methods and compositions of the present disclosure is completely or partially preventative or prophylactic, at least for advanced or non-early stage cancer.
[0318] In various embodiments, cancer screening using the methods and compositions of the present disclosure can be applied to asymptomatic human subjects. In particular, a subject can be referred to as "asymptomatic" if the subject does not report and / or show symptoms of a condition that are sufficient to support a reasonable medical suspicion that the subject may be suffering from the condition through non-invasive observable indicators (e.g., without one, several, or all device-based investigations, tissue sample analysis, body fluid analysis, surgery, or cancer screening).
[0319] Those skilled in the art will appreciate that regular, preventative and / or prophylactic colorectal cancer screening will improve diagnostic yields. In general, particularly in embodiments where the screening of the present disclosure is performed annually and / or the subject is asymptomatic at the time of screening, the methods and compositions of the present invention are particularly likely to detect early-stage colorectal cancer.
[0320] In various embodiments, colorectal cancer screening according to the present disclosure is performed on a given subject one or more times. In various embodiments, colorectal cancer screening according to the present disclosure is performed on a regular basis, such as once every six months, annually, every two years, every three years, every four years, every five years, or every ten years.
[0321] In various embodiments, screening using the methods and compositions disclosed herein will provide a diagnosis of a condition (e.g., a classification of colorectal cancer). Molecular subtyping of colorectal cancer helps better stratify patients before treatment and helps guide treatment plans, which in turn can lead to better outcomes. The same is true for other cancer types (e.g., breast cancer). Traditionally, molecular subtyping is accomplished through whole-genome sequencing and copy number variation identification, based on patterns of chromosomal structural variations with potential clinical utility. DNA methylation can provide another dimension to subtyping patients by providing characteristics of the cell type of origin (which can determine the different cell sources of cancer). In other cases, screening for colorectal neoplasms using the methods and compositions disclosed herein will indicate the presence of one or more conditions, but will not confirm a specific condition. For example, screening can be used to classify a subject as having one or more conditions or a combination of conditions, including, but not limited to, advanced adenoma and / or rectal cancer. In various cases, screening using the methods and compositions disclosed herein can be followed by further diagnostic confirmatory testing that can confirm, support, weaken, or refute the diagnosis made by the previous screening (e.g., the screening disclosed herein).
[0322] In various embodiments, screening with the methods and compositions of the present disclosure can reduce colorectal cancer mortality, for example, by early diagnosis of colorectal cancer. Data indicate that colorectal cancer screening reduces colorectal cancer mortality, an effect that has persisted for over 30 years (see, e.g., Shaukat 2013 N Engl J Med. 369(12): 1106-14). In addition, colorectal cancer is particularly difficult to treat, at least in part because, without timely screening, colorectal cancer may not be detected until the cancer is past its early stages. For at least this reason, treatment of colorectal cancer is often unsuccessful. To maximize colorectal cancer outcomes for the entire population, screening using the present invention can be combined with, for example, recruitment of qualified subjects to ensure extensive screening.
[0323] In various embodiments, after colorectal neoplasm screening including one or more methods and / or compositions disclosed herein, colorectal cancer treatment, such as early colorectal cancer treatment, is performed. In various embodiments, the treatment of colorectal cancer (e.g., early colorectal cancer) includes administering a treatment regimen comprising one or more of surgery, radiotherapy, and chemotherapy. In various embodiments, the treatment of colorectal cancer (e.g., early colorectal cancer) includes administering a treatment regimen comprising one or more provided herein for treating stage 0 colorectal cancer, stage I colorectal cancer, and / or stage II (e.g., including stage IIA, stage IIB, and stage IIC) colorectal cancer.
[0324] In various embodiments, treatment of colorectal cancer includes treating early-stage colorectal cancer (e.g., stage 0 colorectal cancer or stage I colorectal cancer) by one or more of the following: surgical removal of cancerous tissue (e.g., by colonoscopy), local resection, partial colectomy, or complete colectomy).
[0325] In various embodiments, treatment of colorectal cancer includes treating early-stage colorectal cancer (e.g., stage II (e.g., including stage IIA, IIB, and IIC) colorectal cancer) by one or more of the following: surgical removal of cancerous tissue (e.g., by local resection (e.g., by colonoscopy), partial colectomy, or complete colectomy), surgical removal of lymph nodes adjacent to identified colorectal cancer tissue, and chemotherapy (e.g., administration of one or more of 5-FU and leucovorin, oxaliplatin, or capecitabine).
[0326] In various embodiments, treatment of colorectal cancer comprises treating stage III colorectal cancer by one or more of surgical resection of the cancerous tissue (e.g., by local resection (e.g., by colonoscopy-based resection), segmental colectomy, or complete colectomy), surgical resection of lymph nodes proximal to the identified colorectal cancerous tissue, chemotherapy (e.g., administration of one or more of 5-FU, leucovorin, oxaliplatin, or capecitabine, such as the following combinations: (i) 5-FU and leucovorin, (ii) 5-FU, leucovorin, and oxaliplatin (e.g., FOLFOX), or (iii) capecitabine and oxaliplatin (e.g., CAPEOX)), and radiation therapy.
[0327] In various embodiments, the treatment of colorectal cancer comprises treating stage IV colorectal cancer by one or more of the following: surgical removal of cancerous tissue (e.g., by local resection (e.g., by colonoscopy), segmental colectomy, or complete colectomy), surgical removal of lymph nodes adjacent to identified colorectal cancer tissue, surgical removal of metastases, chemotherapy (e.g., administration of one or more of 5-FU, leucovorin, oxaliplatin, capecitabine, irinotecan, VEGF targeted therapeutics (e.g., bevacizumab, aflibercept, or ramucirumab), EGFR targeted therapeutics (e.g., cetuximab or panitumumab), regorafenib, trifluridine, and tipizalin, such as a combination of or including: (i) 5-FU and leucovorin; (ii) 5-FU, leucovorin, or oxaliplatin; (e.g., FOLFOX); (iii) capecitabine and oxaliplatin (e.g., CAPEOX); (iv) leucovorin, 5-FU, oxaliplatin, and irinotecan (e.g., FOLFOX); and (v) trifluridine and tipiracil (Lonsurf), radiation therapy, hepatic artery infusion (e.g., if the cancer has metastasized to the liver), tumor ablation, tumor embolization, colon stenting, colectomy, colostomy (e.g., diverting colostomy), and immunotherapy (e.g., pembrolizumab).
[0328] It will be appreciated by those skilled in the art that the colorectal cancer treatment methods provided herein may be used alone or in any combination, in any order, treatment regimen, and / or therapy, as determined by a physician. It will be further appreciated by those skilled in the art that advanced treatment options may be suitable for early-stage cancers in subjects who have previously had cancer or colorectal cancer (e.g., subjects diagnosed with recurrent colorectal cancer).
[0329] In some embodiments, the methods and compositions provided herein for colorectal cancer screening can provide information for treatment and / or payment (e.g., reimbursement or reduction of medical (e.g., screening or treatment) expenses) decisions and / or actions by, for example, individuals, medical institutions, medical practitioners, health insurance providers, government agencies, or other parties concerned with medical costs.
[0330] In some embodiments, the methods and compositions provided herein for colorectal neoplasm screening can provide information for a health insurance provider to make decisions related to whether to reimburse a payer or recipient of medical expenses, such as reimbursement for (1) the screening itself (e.g., reimbursement for screening that is not otherwise applicable, that is only applicable for regular / routine screening, or that is only applicable for temporary and / or occasional motivations) and / or reimbursement for (2) treatment, including, for example, starting, maintaining, and / or changing therapy based on the screening results. For example, in some embodiments, the methods and compositions provided herein for colorectal neoplasm screening are used as a basis, catalyst, or support for deciding whether to provide reimbursement or fee waiver to a payer or recipient of medical expenses. In some cases, the party seeking reimbursement or fee waiver can provide the results of a screening performed in accordance with the present specification at the same time as providing a request for reimbursement or fee waiver of medical expenses. In some cases, the party deciding whether to provide reimbursement or fee waiver of medical expenses will make a decision based in whole or in part on the results of the screening performed in accordance with the present specification.
[0331] For the avoidance of any doubt, those skilled in the art will appreciate from the present disclosure that the methods and compositions for colorectal cancer diagnosis described herein are for use at least in vitro. Therefore, all aspects and embodiments of the present invention may be performed and / or used at least in vitro.
[0332] Reagent test kit
[0333] Among other things, the present invention also includes a kit comprising one or more compositions for screening purposes provided herein, optionally in combination with instructions for use for screening (e.g., screening for colorectal cancer and / or other diseases or conditions associated with abnormal methylation states, e.g., neurodegenerative diseases, gastrointestinal disorders, etc.). In various embodiments, the kit for screening for a disease or condition associated with an abnormal methylation state may comprise one or more oligonucleotide probes (e.g., one or more biotinylated oligonucleotide probes). In certain embodiments, the kit for screening optionally comprises one or more bisulfite conversion reagents disclosed herein. In certain embodiments, the screening kit optionally comprises one or more enzyme conversion reagents disclosed herein. In certain embodiments, the kit for screening may comprise one or more linkers described herein. In certain embodiments, the kit may comprise one or more reagents for library preparation. In certain embodiments, the kit may comprise, for example, software for analyzing the methylation state of DMRs.
[0334] The elements of the different embodiments described herein may be combined to form other embodiments not specifically described above. The processes, computer programs, databases, etc. described herein may exclude certain elements without adversely affecting their operation. In addition, the logic flows depicted in the figures do not require the specific order or sequential order shown to achieve the desired results. Each independent element may be combined into one or more separate elements to perform the functions described herein.
[0335] Throughout the specification, when apparatus and systems are described as having, containing or comprising specific components, or when processes and methods are described as having, containing or comprising specific steps, it can be further considered that the apparatus and systems of the present invention are essentially composed of, or consist of, the elements, and that the processes and methods according to the present invention are essentially composed of, or consist of, the processing steps.
[0336] It should be understood that the order of steps or the order in which certain operations are performed is not important as long as the present invention remains operable. In addition, two or more steps or operations can be performed simultaneously.
[0337] While the present invention has been particularly shown and described with reference to certain preferred embodiments, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the invention as defined in the following claims.
[0338] Other implementations
[0339] Although many embodiments have been described, it is apparent that the underlying disclosure and examples can provide other embodiments that utilize or encompass the compositions and methods described herein. Therefore, it will be understood that the scope of the invention should be defined by the scope that can be understood from the disclosure and the appended claims, rather than by the specific embodiments presented by way of example.
[0340] All references cited herein are incorporated by reference.
Claims
1. A method for detecting hypermethylation of a marker to identify a condition in a human subject, the method comprising: Determining the methylation status of each of one or more markers identified in a deoxyribonucleic acid (DNA) fragment (DNA fragment) from a sample of a human subject susceptible to colorectal cancer and / or advanced adenoma; as well as identifying a condition in the subject based at least in part on the determined methylation status of each of the one or more markers identified in the DNA fragments, wherein each of the one or more markers is a methylation locus comprising at least one single DMR or a portion of a DMR selected from the 82 differentially methylated regions (DMRs) listed in Table 1 (i.e., SEQ ID NOs. 1 to 82), Wherein, detecting the methylation status includes: determining whether at least one methylation site in at least one marker among the one or more markers is hypermethylated.
2. The method according to claim 1, wherein The sample is a member selected from the group consisting of a tissue sample, a blood sample, a stool sample, and a blood product sample.
3. The method according to claim 1, wherein The sample comprises DNA from blood or plasma of the human subject.
4. The method according to claim 1, wherein The DNA is cell-free DNA (cfDNA) of the human subject.
5. The method according to claim 1, wherein The method comprises determining the methylation status of each of the one or more markers using next generation sequencing (NGS).
6. The method of claim 1, wherein: The method includes using one or more capture baits enriched for target regions to capture one or more corresponding methylated loci.
7. The method of claim 1, wherein: The length of each methylation locus was equal to or less than 302 bp.
8. A method for detecting the methylation status of a marker, the method comprising: converting unmethylated cytosine in a plurality of DNA fragments in a sample into uracil to generate a plurality of converted DNA fragments, wherein the plurality of DNA fragments are obtained from a biological sample; sequencing the plurality of converted DNA fragments to generate a plurality of sequence reads, wherein each sequence read corresponds to a converted DNA fragment; and detecting the methylation status of each of the one or more markers identified in the sequence reads, Wherein, each of the one or more markers is a methylation locus comprising at least one single differentially methylated region (DMR) or a portion of a DMR selected from the 82 differentially methylated regions (DMRs) listed in Table 1 (ie, SEQ ID NOs. 1 to 82).
9. The method of claim 8, wherein: The plurality of DNA fragments comprises (in total) at least 1 ng of DNA.
10. The method of claim 8, wherein: The plurality of DNA fragments essentially consist of DNA fragments each having a length of 10 bp to 800 bp.
11. The method of claim 8, wherein: The plurality of DNA fragments essentially consist of DNA fragments each having a length of 1,000 bp to 200,000 bp.
12. The method of claim 8, wherein: Each of the plurality of sequence reads is at least 50 bp.
13. The method of claim 1, wherein: At least one of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO. 3 or SEQ ID NO.
14.
14. The method of claim 1, wherein: At least one of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO.
2.
15. The method of claim 1, wherein: At least one of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO.
1.
16. The method of claim 1, wherein: The first of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO. 1, the second of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO. 2, and the third of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO.
3.
17. The method of claim 1, wherein: A first of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO. 70, and a second of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO.
72.
18. The method of claim 8, wherein: The first of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO. 1, the second of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO. 2, and the third of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO.
3.
19. The method of claim 8, wherein: A first of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO. 70, and a second of the one or more markers is a methylated locus comprising at least a portion of SEQ ID NO.
72.
20. The method of claim 8, wherein: The biological sample is a sample from a subject susceptible to colorectal cancer and / or advanced adenoma.
Citation Information
Patent Citations
Method of detection of methylated nucleic acid using agents which modify unmethylated cytosine and distinguish modified methylated and non-methylated nucleic acids
US6265171B1