Method for amplifying bisulfite treated DNA
By linking a bisulfite-protected adaptor to the DNA molecule and then treating it with bisulfite, the problem of sequencing small amounts of DNA in existing technologies has been solved, achieving high-resolution 5mC and 5hmC identification, which is suitable for rare samples and single-cell systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- UNIVERSITY OF CHICAGO
- Filing Date
- 2019-07-26
- Publication Date
- 2026-04-21
AI Technical Summary
Existing sequencing methods for 5-methylcytosine and 5-hydroxymethylcytosine in nucleic acid molecules require large amounts of DNA input, making them difficult to apply to rare samples and single-cell systems. Furthermore, existing methods have limitations in resolution and coverage.
By linking an adaptor to a DNA molecule, wherein the adaptor comprises a bisulfite-protected cytosine, the molecule is bisulfite-treated and then hybridized with primers. The extended hybridization primers produce double-stranded DNA, which is then transcribed in vitro to produce RNA, enabling the amplification and determination of the modification status of a small amount of DNA.
It enables efficient analysis of limited amounts of DNA, and can identify 5mC and 5hmC modification information at the single-cell level, improving sequencing resolution and coverage.
Smart Images

Figure CN112714796B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 711184, filed July 27, 2018, the contents of which are incorporated herein by reference in their entirety.
[0003] Government Support Statement
[0004] This invention was completed with government support under approval number HG006827 granted by the National Institutes of Health (NIH). The government holds certain rights to this invention. Background of the Invention
[0005] I. Technical Field
[0006] Embodiments of the present invention generally relate to cell biology. In some aspects, the method includes determining the presence of 5-methylcytosine and / or 5-hydroxymethylcytosine in a nucleic acid molecule.
[0007] II. Background Technology
[0008] 5-Methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC) are important epigenetic markers in mammalian cells. Current 5mC and 5hmC sequencing methods can be summarized as follows: 1) bisulfite-based methods; 2) affinity-capture-based methods, including antibody-based pull-down and selective chemical labeling-based pull-down; and 3) restriction endonuclease-based methods. All of these existing methods require micrograms of input genomic DNA. This large input limits applications in rare samples and single-cell systems, such as the behavior of single cells during differentiation. Bisulfite-based methods are considered the gold standard due to their ability to quantitatively distinguish 5mC from ordinary C at single-base resolution. However, DNA degradation is a major drawback. Affinity-capture-based methods are relatively inexpensive but have low resolution and may lose information due to low CpG density coverage (antibody-based methods). Restriction endonuclease methods have limited resolution, and coverage depends on sequence specificity and sensitivity to methylation or hydroxymethylation. In summary, current methods cannot sequence 5mC and 5hmC in small amounts of DNA (nanogels or sub-nanogels) or obtain information on these modifications at the single-cell level. Therefore, there is a need in the art for more methods to detect cytosine modifications, such as 5mC and 5hmC, in small amounts of DNA. Summary of the Invention
[0009] The methods, compositions, and kits disclosed herein provide an efficient method for whole-genome analysis, namely an unbiased DNA analysis method, which can be performed on a limited amount of DNA and can efficiently determine the modification state of DNA. Aspects of this disclosure relate to a method for amplifying bisulfite-treated deoxyribonucleic acid (DNA) molecules, the method comprising: (a) ligating an adaptor to a DNA molecule, wherein the adaptor includes an RNA polymerase promoter containing bisulfite-protected cytosine; (b) treating the ligated DNA molecule with bisulfite; (c) hybridizing the bisulfite-treated DNA molecule with primers; (d) extending the hybridized primers to produce double-stranded DNA; and (e) in vitro transcribing the double-stranded DNA to produce RNA.
[0010] This disclosure relates to a method for amplifying bisulfite-treated deoxyribonucleic acid (DNA) molecules, the method comprising: (a) contacting the nucleic acid molecule with a ligase and an adaptor under conditions suitable for ligating an adaptor to the DNA molecule; wherein the adaptor comprises an RNA polymerase promoter containing bisulfite-protected cytosine; (b) treating the ligated DNA molecule with bisulfite; (c) contacting the DNA molecule with a primer at least partially complementary to the adaptor under conditions allowing the primer to hybridize with the bisulfite-treated DNA molecule; (d) performing primer extension to produce double-stranded DNA; and (e) contacting the double-stranded DNA with an RNA polymerase in the presence of an NTP under conditions suitable for in vitro transcription of the DNA to produce RNA.
[0011] The term "bisulfite-protected cytosine" refers to a cytosine that resists deamination upon contact with a bisulfite. Exemplary bisulfite-protected cytosines are described herein, including, for example, 5-methylcytosine and 5-hydroxymethylcytosine.
[0012] Other aspects of this disclosure relate to a method for identifying 5-hydroxymethylcytosine (5hmC) in a DNA molecule, the method comprising: (1) modifying 5hmC in the DNA molecule to protect it from oxidation; (2) oxidizing the modified DNA molecule from (1) with a methylcytosine dioxygenase to convert 5-methylcytosine (5mC) to 5-carboxycytosine (5caC); and (3) performing a method comprising: (a) attaching an adaptor to the DNA molecule, wherein the adaptor comprises an RNA polymerase promoter containing bisulfite-protected cytosine; (b) treating the attached DNA molecule with bisulfite; (c) hybridizing the bisulfite-treated DNA molecule with primers; (d) extending the hybridized primers to produce double-stranded DNA; and (e) transcribing the double-stranded DNA in vitro to produce RNA. In some embodiments, step (a) is performed before step (1). In some embodiments, step (a) is performed after step (1) and before step (2). In some implementations, (a) is performed after step (2) and before step (3). In some implementations, (a) is performed after step (3) and before step (b).
[0013] Other aspects relate to methods for identifying 5mC in DNA molecules, the methods comprising: (1) oxidizing the DNA molecule with an oxidizing agent to oxidize 5hmC to 5-formylcytosine (5fC) or 5caC; (2) performing a method comprising the steps of: (a) ligating an adaptor to the DNA molecule, wherein the adaptor comprises an RNA polymerase promoter containing bisulfite-protected cytosine; (b) treating the ligated DNA molecule with bisulfite; (c) hybridizing the bisulfite-treated DNA molecule with primers; (d) extending the hybridized primers to produce double-stranded DNA; and (e) in vitro transcribing the double-stranded DNA to produce RNA. In some embodiments, step (a) is performed before step (1). In some embodiments, step (a) is performed after step (1) and before step (2). In some embodiments, (a) is performed after (2).
[0014] Other aspects of this disclosure relate to a method for identifying 5mC in a DNA molecule, the method comprising: (a) ligating an adaptor to a DNA molecule, wherein the adaptor includes an RNA polymerase promoter containing bisulfite-protected cytosine; (b) treating the ligated DNA molecule with bisulfite; (c) hybridizing the bisulfite-treated DNA molecule with primers; (d) extending the hybridization primers to produce double-stranded DNA; and (e) in vitro transcribing the double-stranded DNA to produce RNA; (f) reverse transcribing the RNA to produce DNA; and (g) sequencing the DNA and identifying 5mC in the sequenced DNA as “C” in the sequence.
[0015] Other aspects of this disclosure relate to kits that include DNA adaptors comprising RNA promoters, wherein the cytosine of the RNA promoter is protected with bisulfite.
[0016] In some embodiments, the methylcytosine dioxygenase is TET1, TET2, or TET3, or a homologue thereof. In some embodiments, 5hmC is modified with glucose or a modified glucose. In some embodiments, 5hmC is modified by a method comprising incubating a nucleic acid molecule with a β-glucosyltransferase and glucose or a modified glucose molecule. In some embodiments, the glucose molecule is a uridine diphosphate glucose molecule. In some embodiments, the modified glucose molecule is a modified uridine diphosphate glucose molecule.
[0017] In some embodiments, the oxidation selectively oxidizes 5hmC residues. In some embodiments, the oxidant is a chemical oxidant. In some embodiments, the oxidant is a perruthenate oxidant. In some embodiments, the oxidant includes KRuO4. In some embodiments, the oxidant includes the oxidants described herein.
[0018] In some embodiments, the bisulfite-protected cytosine comprises 5mC or 5hmC. In some embodiments, the bisulfite-protected cytosine comprises oxime or hydrazone-modified 5fC. In some embodiments, 5fC is modified with a compound containing a hydroxylamine group, a hydrazide group, or an acylhydrazide group. In some embodiments, said compound is hydroxylamine; hydroxylamine hydrochloride; hydroxylammonium sulfate; hydroxylamine phosphate; O-methylhydroxylamine; O-hexylhydroxylamine; O-pentylhydroxylamine; O-benzylhydroxylamine; O-ethylhydroxylamine (EtONH2), O-alkylated or O-arylated hydroxylamine, its acid, or its salt. In some embodiments, the bisulfite-protected cytosine comprises bisulfite-protected 5caC. In some embodiments, 5caC is amide-modified 5caC. In some embodiments, 5caC is amide-modified by linking to a compound containing an amino group. In some embodiments, 5caC is linked to an amino group by incubating a DNA molecule with a carbodiimide derivative. In some embodiments, the compound containing an amino group is benzylamine, a substituted benzylamine, an alkylamine, an alkyldiamine, a xylylamine, a substituted xylylamine, a cycloalkylamine, a cycloalkyldiamine, hydroxylamine, or a substituted hydroxylamine. In some embodiments, the bisulfite-protected cytosine comprises 5'-alkylcytosine. The alkyl group may be further substituted or unsubstituted. In some embodiments, the bisulfite-protected cytosine contains functional groups that can alter the electron density of the aromatic ring and / or increase steric hindrance to protect the cytosine from deamination.
[0019] In some embodiments, the promoter is a DNA-dependent RNA polymerase. The promoter can be a prokaryotic or eukaryotic RNA polymerase, such as a bacterial or bacteriophage RNA polymerase. In some embodiments, the promoter is recognized by an RNA polymerase composed of a single subunit. In further embodiments, the RNA polymerase promoter includes SP6, T7, or T3 promoters. In some embodiments, (a) includes end modification and / or end repair of the DNA. In some embodiments, end modification includes adding an A to the end. In some embodiments, (a) includes contacting the DNA with a ligase under conditions sufficient to ligate the adaptor to the DNA molecule. In some embodiments, the adaptor further includes a 3' end-blocking molecule. A 3' end-blocking molecule refers to a nucleic acid lacking the 3' phosphate necessary for the ligation reaction. In some embodiments, the 3' end-blocking molecule includes a 3' phosphate as a terminal group. In some embodiments, the 3' end-blocking molecule lacks a 3' hydroxyl group as a terminal group.
[0020] In some embodiments, the adaptor is partially double-stranded, for example, at least or at most 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% (or any range thereof) of which are double-stranded. In other embodiments, there may be at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, sixteen, seventeen, eighteen, nineteen, twenty, or more than twenty single-stranded nucleic acid residues (or any range thereof). In some embodiments, the adaptor contains one or more primer binding sites. In some implementations, the adapter contains one or more restriction sites.
[0021] In some embodiments, (d) includes contacting the DNA molecule with a DNA polymerase. In some embodiments, (d) includes incubating the DNA under denaturing conditions that denature double-stranded DNA into single-stranded DNA. In some embodiments, (d) includes incubating the DNA under conditions sufficient to anneal the primers to single-stranded DNA. In some embodiments, (d) includes incubating the DNA under conditions sufficient to extend the primers to produce double-stranded DNA.
[0022] In some implementations, (e) includes contacting the DNA molecule with RNA polymerase and nucleoside triphosphate.
[0023] In some embodiments, the method further includes one or more purification steps. In some embodiments, the purification step includes solid-phase reversible immobilization (SPRI) beads. In some embodiments, the method further includes (e) the purification of RNA molecules. In some embodiments, the method further includes isolating or purifying nucleic acid molecules by contacting nucleic acid molecules with a capture agent, wherein the capture agent binds to an affinity tag; and separating the capture agent bound to the affinity-tagged nucleic acid molecules from surrounding components.
[0024] In some embodiments, the method further includes reverse transcription of the RNA molecule in (e) to produce the corresponding DNA molecule. In some embodiments, (a) through (e) are performed sequentially. In some embodiments, the method further includes library construction of the corresponding DNA molecule. In some embodiments, the method further includes sequencing of the corresponding DNA molecule. In some embodiments, sequencing includes sequencing performed by Sanger sequencing, Maxam-Gilbert sequencing, SOLiD sequencing, sequencing by synthesis, pyrosequencing, IonTorrent semiconductor sequencing, massively parallel sequencing, polymerase cloning sequencing, 454 pyrosequencing, Illumina dye sequencing, DNA nanosphere sequencing, or single-molecule real-time sequencing. In some embodiments, the method does not include bisulfite treatment of nucleic acids.
[0025] In some embodiments, the DNA molecule is fragmented. In some embodiments, the fragment length is 100 bp to 400 bp. In some embodiments, the fragment length is 150 bp to 300 bp. In some embodiments, the fragment length is at least, at most, or exactly 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, 125 bp, 150 bp, 175 bp, 200 bp, 225 bp, 250 bp, 275 bp, 300 bp, 325 bp, 350 bp, 375 bp, 400 bp, 425 bp, 450 bp, 475 bp, 500 bp, 525 bp, 550 bp, 575 bp, 600 bp, 625 bp, 650 bp, 675 bp, 700 bp, 725 bp, 750 bp, 775 bp, 800 bp, 8 25bp, 850bp, 875bp, 900bp, 950bp, 1000bp, 1050bp, 1100bp, 1150bp, 1200bp, 1300bp, 1400bp, 1500bp, 1600bp, 1700bp, 1800bp, 1900bp, 2000bp, 2100bp, 2200bp, 2300bp, 2400bp, 2500bp, 2600bp, 2700bp, 2800bp, 2900bp, 3000bp, 3200bp, 3400bp, 3600bp, 3800bp, or 4000bp (or any range deducible from therein). In some embodiments, the DNA molecule includes biological fragments, such as DNA that is naturally present in a fragmentary state upon isolation. In some embodiments, the DNA molecule includes cell-free DNA (cfDNA). In some embodiments, cell-free DNA is isolated from serum, whole blood, plasma, or a portion thereof. In some embodiments, cell-free DNA is isolated from tissue samples. In some embodiments, the fragment comprises genomic DNA. In some embodiments, the fragment comprises genomic DNA isolated from cells. In some embodiments, the DNA molecules are fragmented and size-graded. The size grading of DNA fragments can be performed using methods known in the art, such as gel grading, size exclusion chromatography, and by using commercially available kits (e.g., EpiNext). TM The DNA size selection kit (EpiGentek) and Select-a-Size DNA Clean & Concentrator (Zymo Research) were used for this purpose.
[0026] In some embodiments, the method further includes fragmenting the nucleic acid molecule. In some embodiments, the method further includes labeling the nucleic acid molecule. In some embodiments, the nucleic acid is labeled and / or fragmented by a transposon. In some embodiments, labeling and / or fragmenting the nucleic acid includes contacting the nucleic acid molecule with a transposon and a transposon. In some embodiments, the transposon includes a transposon containing a p7 adaptor. In some embodiments, the transposon contains an affinity tag. The affinity tag may include, for example, biotin, myc, and His tags.
[0027] In some embodiments, the amount of DNA molecules in the starting material used for the experiment is from 1 pg to 1000 ng. In some embodiments, the amount of DNA molecules in the starting material used for the experiment is from 1 pg to 100 ng. In some embodiments, the amount of DNA molecules in the starting material used for the experiment is from 1 pg to 10 ng. In some embodiments, the amount of DNA molecules in the starting material used for the experiment is from 10 pg to 1000 ng. In some embodiments, the amount of DNA molecules in the starting material used for the experiment is from 10 pg to 100 ng. In some embodiments, the amount of DNA molecules in the starting material used for the experiment is from 10 pg to 10 ng. In some embodiments, the amount of DNA molecules in the starting material used for the experiment is from 10 pg to 1 ng. In some implementations, the amount of DNA molecules used as starting material for the experiment is, at least, and at most, about 1 pg, 10 pg, 50 pg, 100 pg, 200 pg, 300 pg, 400 pg, 500 pg, 600 pg, 700 pg, 800 pg, 900 pg, or 1000 pg, or 1 ng, 5 ng, 10 ng, 20 ng, 30 ng, 40 ng, 50 ng, 60 ng, or 70 ng. ng, 80ng, 90ng, 100ng, 110ng, 120ng, 130ng, 140ng, 150ng, 160ng, 170ng, 180ng, 190ng, 200ng, 210ng, 220ng, 230ng, 240ng, 250ng, 260ng, 270ng, 280ng, 290ng or 300ng (or any range that can be derived from therein).
[0028] In some embodiments, DNA molecules are isolated from a sample of the subject. In some embodiments, DNA molecules are isolated from a biopsy sample. In some embodiments, the sample is a liquid sample or a liquid biopsy sample. In certain embodiments, the sample is derived from blood, urine, cerebrospinal fluid (CSF), or aqueous humor cfDNA. In certain embodiments, a single cell or up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 50, 100, 200, 300, 400, or 500 cells (or any range deducible therebetween) is evaluated.
[0029] Some embodiments relate to methods for identifying biomarkers comprising DNA molecules in a sample, said methods including performing the methods of this disclosure. Some embodiments relate to methods for providing diagnosis or prognosis to a patient, said methods including performing the methods of this disclosure, wherein DNA molecules are provided from a biological sample of the patient. Some embodiments relate to methods for evaluating single cells, said methods including performing the methods of this disclosure, wherein DNA molecules are provided from the genomic DNA of the single cell.
[0030] In some embodiments, the method further includes separating the cell population into isolated single cells. Cells can be classified using methods known in the art, such as FACS, or by serial dilution of the cell population. In some embodiments, the method further includes labeling the nucleic acids of each single cell having a unique nucleic acid sequence. In some embodiments, the method further includes combining the labeled nucleic acids into a single composition.
[0031] In some implementations, the method further includes end repair of nucleic acids. End repair kits are known in the art and commercially available, and can be used to convert DNA containing damaged or incompatible 5' and / or 3' protruding ends into 5' phosphorylated blunt-ended DNA.
[0032] Some methods of this disclosure use samples. Biological samples can be obtained using any method known in the art, which can provide samples suitable for the analytical methods described herein. Samples can be obtained by non-invasive methods, including but not limited to: scraping the skin or cervix, wiping the cheek, collecting saliva, collecting urine, collecting feces, collecting menstrual blood, tears, or semen.
[0033] Samples can be obtained by methods known in the art. In some embodiments, samples are obtained by biopsy. In other embodiments, samples are obtained by swabbing, scraping, venipuncture, or any other method known in the art. In some cases, components of a kit based on this method can be used to obtain, store, freeze, or transport samples. In some embodiments, samples include blood samples, serum samples, plasma samples, or portions thereof. In some embodiments, samples include urine samples.
[0034] In some embodiments, biological samples can be obtained by a physician, nurse, or other medical specialist, such as a medical technician, endocrinologist, cytologist, venipuncturist, radiologist, or pulmonologist. The medical specialist can instruct on the appropriate tests or trials to be performed on the sample. In some aspects, molecular analysis can be consulted regarding which test or trial is most appropriately indicated. In other aspects of the invention, the patient or subject can obtain biological samples for testing without the assistance of a medical specialist, such as whole blood samples, urine samples, stool samples, oral samples, or saliva samples.
[0035] Unless otherwise stated, methods may involve any of the steps described herein and may be in any particular order. In particular, it is anticipated that any method, kit, or composition discussed herein may be combined with any other embodiments discussed herein. Steps and embodiments may be used in any feasible combination.
[0036] The method disclosed herein may include one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, or fifteen or more of the following steps, which may be performed in any order and repeated in all any particular method embodiment: obtaining nucleic acid molecules; obtaining nucleic acid molecules from a biological sample; obtaining a biological sample containing nucleic acids from a subject; isolating nucleic acid molecules; purifying nucleic acid molecules; obtaining an array or microarray containing the nucleic acid to be modified; denaturing nucleic acid molecules; cleaving or cutting nucleic acids; hybridizing nucleic acid molecules; fragmenting nucleic acids; incubating nucleic acid molecules with an enzyme; incubating nucleic acid molecules with an unmodified 5mC enzyme; incubating nucleic acid molecules with a restriction enzyme; linking one or more chemical groups or compounds to nucleic acids or 5mC or modified 5mC; conjugating one or more chemical groups or compounds to nucleic acids or 5mC or modified 5mC. The process involves: incubating a nucleic acid molecule with an enzyme that modifies the nucleic acid molecule or 5mC or modified 5mC by adding or removing one or more elements, chemical groups, or compounds; modifying or converting 5mC to 5-hydroxymethylcytosine (5hmC); modifying 5hmC with β-glucosyltransferase (βGT); incubating the β-glucosyltransferase with a UDP-glucose molecule and a nucleic acid substrate under conditions that promote glycosylation of nucleic acid by glucose molecules (which may or may not be modified) and produce nucleic acids glycosylated on one or more 5-hydroxymethylcytosines; ligating an adaptor to a DNA molecule, wherein the adaptor includes an RNA polymerase promoter containing bisulfite-protected cytosine; treating the ligated DNA molecule with bisulfite; hybridizing the bisulfite-treated DNA molecule with primers; extending the hybridized primers to produce double-stranded DNA; and transcribing the double-stranded DNA in vitro to produce RNA.
[0037] Some implementation schemes are expected to involve steps performed outside the body, such as by a person, or a person who controls or uses machinery, to perform one or more steps.
[0038] The methods or compositions will involve purified nucleic acids, modifying reagents or enzymes, labels, chemically modified moieties, modified UDP-Glc and / or enzymes, such as β-glucosyltransferases. Such protocols are known to those skilled in the art.
[0039] In some embodiments, the purification process can produce molecules with a purity of about or at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or greater than 99.9% relative to any contaminating component (by weight / weight or weight / volume) (or any range thereof).
[0040] Other methods include steps including, but not limited to, obtaining information (qualitative and / or quantitative) on one or more cytosine modifications in a nucleic acid sample; arranging a detection to determine, identify, and / or map cytosine modifications in the nucleic acid sample; reporting information (qualitative and / or quantitative) on one or more cytosine modifications in the nucleic acid sample; and comparing this information with information on different cytosine modifications in a control or comparison sample. Unless otherwise stated, in the context of a sample, the terms “determination,” “analysis,” “detection,” and “evaluation” refer to the chemical or physical transformation of a sample to collect qualitative and / or quantitative data about the sample. Furthermore, the term “map” indicates the identification of the position of a specific nucleotide within a nucleic acid sequence.
[0041] In some embodiments, nucleic acid molecules may be DNA, RNA, or a combination of both. Nucleic acids may be recombinant, genomic, or synthetic. In other embodiments, the method involves isolating and / or purifying nucleic acid molecules. In some embodiments, nucleic acid molecules are fragmented. In some embodiments, nucleic acid molecules are naturally fragmented. Naturally fragmented refers to nucleic acid molecules that exist as fragments in nature, such as cell-free DNA and fetal DNA. In some embodiments, nucleic acids may be isolated from cells or biological samples. Certain embodiments involve isolating nucleic acids from eukaryotic cells, mammalian cells, or human cells. In some cases, they are isolated from non-nucleic acid sources. In some embodiments, nucleic acid molecules are eukaryotic; in some cases, nucleic acids are mammalian, which may be human. This means that the nucleic acid molecules are isolated from human cells and / or have sequences that identify them as human. In certain embodiments, the nucleic acid molecules are not expected to be prokaryotic nucleic acids, such as bacterial nucleic acid molecules. In other embodiments, the isolated nucleic acid molecules are on an array. In certain cases, the array is a microarray. In some cases, nucleic acids are isolated using any technique known to those skilled in the art, including but not limited to the use of gels, columns, matrices, or filters. In some implementations, the gel is a polyacrylamide or agarose gel.
[0042] The methods and compositions may also involve one or more enzymes. In some embodiments, the enzyme is a polymerase. In some cases, the embodiments involve a restriction enzyme. The restriction enzyme may be methylation-insensitive. The steps intended to achieve the result using the enzyme involve incubating the enzyme under reaction conditions to achieve that result. Such conditions include, but are not limited to, temperature, pressure, pH, viscosity, volume, and the presence of any cofactors used in the reaction, as known to those skilled in the art. It may include one or more reaction buffers. In some embodiments, the reaction may be stopped by heat inactivation, dilution, changing the pH, or adding a compound that interferes with the reaction, or by changing the conditions for stopping the reaction.
[0043] Methods or compositions relating to the detection, characterization, and / or differentiation of cytosine modifications. Methods may involve identifying 5mC in nucleic acids by comparing the modified nucleic acid with unmodified nucleic acid or with nucleic acids whose modification state is known. Detection of modifications may involve various recombinant nucleic acid technologies. In some embodiments, the modified nucleic acid molecule is incubated with a polymerase, at least one primer, and one or more nucleotides under conditions that allow the modified nucleic acid to polymerize. In other embodiments, the method may involve sequencing the modified nucleic acid molecule. In still other embodiments, the modified nucleic acid is used for primer extension detection.
[0044] Methods and compositions may involve control nucleic acids. Controls can be used to assess whether modification or other enzymatic or chemical reactions have occurred. Alternatively, controls can be used to compare modification states. Controls can be negative controls or positive controls. They can be controls that are not incubated with one or more reagents during the modification reaction. Alternatively, control nucleic acids can be reference nucleic acids, meaning their modification state (based on qualitative and / or quantitative information related to modification at 5 mC, or its absence) is used for comparison with the nucleic acid being evaluated. In some embodiments, multiplex nucleic acids from different sources provide the basis for control nucleic acids. Furthermore, in some cases, control nucleic acids are derived from routine samples regarding specific attributes such as disease or state or other phenotypes. In some embodiments, controls include non-cancerous tissues. In some embodiments, controls include cutoff values. In some embodiments, control samples are derived from different patient populations, different cell or organ types, different disease states, different stages or severity of disease states, different prognoses, different stages of development, etc.
[0045] The embodiments also relate to kits, which can be in suitable containers and used to implement the methods. Embodiments of this disclosure relate to kits comprising DNA adaptors, said adaptors comprising RNA promoters, wherein the cytosine of the RNA promoter is protected with bisulfite. In some embodiments, the kit comprises a ligase and / or a ligase buffer. In some embodiments, the kit comprises bisulfite. In some embodiments, the kit further comprises a primer complementary to the adaptor. In some embodiments, the kit further comprises dNTPs. In some embodiments, the kit further comprises DNA polymerase. In some embodiments, the kit further comprises nuclease-free water. In some embodiments, the kit further comprises a protease. In some embodiments, the kit further comprises RNA polymerase. In some embodiments, the kit further comprises NTPs. In some embodiments, the kit further comprises an oxidizing agent. In some embodiments, the kit further comprises a dioxygenase. In some embodiments, the kit further comprises a compound containing a hydroxylamine group, a hydrazide group, or an acylhydrazide group. In some embodiments, the kit further comprises a compound containing an amino group. In some embodiments, the kit further comprises a 3' end blocking molecule. In some embodiments, the kit further comprises SPRI beads. In some embodiments, the kit further comprises reverse transcriptase. In some embodiments, the kit also includes glucose or modified glucose. In some embodiments, the kit also includes β-glucosyltransferase.
[0046] In some embodiments, the kit contents may include a methylcytosine dioxygenase or a homologue thereof and a 5-hydroxymethylcytosine modifier. In other aspects, the methylcytosine dioxygenase is TET1, TET2, or TET3. In other embodiments, the kit includes a TET1, TET2, or TET3 catalytic domain. In some aspects, the 5hmC modifier is a β-glucosyltransferase, wherein the 5hmC modifier refers to a reagent capable of modifying 5hmC.
[0047] In another embodiment, the kit further contains a 5hmC modification, such as uridine diphosphate glucose or a modified uridine diphosphate glucose molecule. In a specific embodiment, the modified uridine diphosphate glucose molecule may be a uridine diphosphate 6-N3-glucose molecule. In another embodiment, the kit further contains biotin.
[0048] Some embodiments involve kits comprising a vector and a 5-hydroxymethylcytosine modifier, said vector comprising a promoter operatively linked to a nucleic acid fragment encoding a dioxygenase or a portion thereof. In some aspects, the nucleic acid fragment encodes TET1, TET2, or TET3 or their catalytic domain. In some aspects, the 5hmC modifier is a β-glucosyltransferase. In other aspects, the kit also contains a 5hmC modifier, such as uridine diphosphate glucose or a modified uridine diphosphate glucose molecule. In a particular embodiment, the modified uridine diphosphate glucose molecule may be a uridine diphosphate 6-N3-glucose molecule. In another embodiment, the kit also contains biotin.
[0049] In some embodiments, kits exist that contain one or more modifiers (enzymatic or chemical) and one or more modification moieties. Molecules may have or involve different types of modifiers. In further embodiments, the kit may include one or more buffers, such as buffers for nucleic acids or for reactions involving nucleic acids. In addition to β-glucosyltransferases or alternative β-glucosyltransferases, other enzymes may be included in the kit. In some embodiments, the enzyme is a polymerase. The kit may also include nucleotides used with the polymerase. In some cases, restriction enzymes are included in addition to polymerases or alternative polymerases. In some embodiments, the kit includes nucleic acid probes. The nucleic acid probes may or may not have been modified. In some embodiments, the kit includes modification moieties for linking to the nucleic acid probe.
[0050] Other embodiments involve arrays or microarrays containing nucleic acid molecules that have been modified with nucleotides at 5 hmC and / or 5 mC. In some embodiments, the microarray comprises fragmented nucleic acids isolated from a sample.
[0051] The following patent applications describe embodiments useful to the method of the present invention: WO2011127136, WO2012138973 and WO2014165770, which are incorporated herein by reference.
[0052] When used with the term "comprising" in the claims and / or specification, the absence of a quantifier before an element may indicate "one," but it also means "one or more," "at least one," and "one or more than one."
[0053] It is anticipated that any embodiments discussed herein can be implemented with respect to any method or composition of the present invention, and vice versa. Furthermore, the compositions and kits of the present invention can be used to implement the methods of the present invention. Additionally, any step listed in the context of one method can be used in the context of any other method disclosed herein (as an alternative step or as an added step). One or more steps listed herein may be omitted from any method.
[0054] In this application, the term "about" is used to indicate that the value includes the standard deviation of the error of the apparatus or method used to determine the value.
[0055] The use of the term "or" in the claims means "and / or" unless it is explicitly stated that it refers only to the selection or that the selections are mutually exclusive, although the disclosure supports the definition of "and / or" referring only to the selection. It is also contemplated that anything listed using the term "or" may be specifically excluded.
[0056] As used in this specification and claims, the words “comprising,” “having,” “including,” or “containing” are inclusive or open-ended and do not exclude additional, unlisted elements or method steps. It is contemplated that any embodiment discussed herein containing the term “comprising” may be replaced by the phrases “consisting of” or “substantially consisting of”, as these terms are understood in the context of patent law.
[0057] Other objects, features, and advantages of the invention will become apparent from the following detailed description. However, it should be understood that while the detailed description and specific embodiments illustrate particular implementations of the invention, they are given by way of illustration only, as various changes and modifications within the spirit and scope of the invention will become apparent to those skilled in the art. Attached Figure Description
[0058] The following figures form part of this specification and are included to further illustrate certain aspects of the invention. A better understanding of the invention can be achieved by referring to one or more of these figures in conjunction with a detailed description of the specific embodiments provided herein.
[0059] Figures 1A to 1E (A) LABS-seq workflow: ligation of cell-free DNA or fragmented genomic DNA with a fully methylated T7 adaptor; bisulfite treatment and T7 promoter extension; IVT amplification; preparation of a total RNA library and sequencing. (B) Mapping efficiency, overall CpG methylation level, and repetition level (including technical replication) for LABS-seq libraries constructed from 10 pg, 20 pg, 50 pg, 100 pg, 1 ng, 10 ng, and 100 ng. Each point represents a library. (C) Comparison of chromosome coverage between LABS-seq, MethylC-seq, and EpiGnome. (D) Genome browser view of DNA methylation profiles. (E) Scatter plot showing the correlation between LABS-seq and batch references and LABS-seq repetitions, with Pearson correlation (r) displayed.
[0060] Figures 2A to 2C (A) Saturation plots of LABS-seq, EpiGnome, and MethylC-seq illustrate the relationship between CpG coverage and sequencing depth. The plot shows the function of the number of unique CpGs covered (y-axis) versus aligned reads (x-axis). (B) GC content bias of the three methods. (C) Amplification uniformity obtained from Lorenz curves of cumulative read fractions versus cumulative genome fractions. Completely uniform coverage corresponds to the diagonal; deviation from the diagonal represents amplification bias.
[0061] Figures 3A to 3D (A) At annotated genomic sites, normalized methylated CpG exceeds total CpG residues. (B) Methylation levels vary with the activation histone marker H3K4me3 and the repression marker H3K27me3. (C) CpG counts span different genomic backgrounds. Each bar above each genomic background represents methyl-seq–1ng, EpiGnome–1ng, and LABS-seq–50pg from left to right. (D) Comparison of normalized color annotation among the three WGBS methods. Each bar above the right side of the ChromHMM state represents methyl-seq–1ng, EpiGnome–1ng, and LABS-seq–50pg from top to bottom.
[0062] Figures 4A to 4B (A) Box percentages showing hypomethylation (MD < -3SD), hypermethylation (MD > 3SD), and differential methylation (|MD| > 3SD) in CRC and PAN patients. (C) Methylation analysis of a healthy control, a CRC patient, and a pancreatic cancer patient. Methylation z-scores for the three samples are located from the outer ring to the inner ring.
[0063] Figures 5A to 5DcfDNA DMRs used to group CRC patients and healthy individuals. (A) Heatmap of 109 DMRs in CRC patients and healthy controls. Clustering of DMRs and samples by Euclidean distance. (B) Genome-wide distribution of DMRs. (C) Hierarchical cluster analysis between the CRC group and the healthy group. (D) PCA plot of all CRC patients and healthy individuals.
[0064] Figure 6A To Figure 6C: (A) Percentage contribution of different tissues to plasma cfDNA in 6 healthy individuals and 6 CRC patients. Each bar represents the proportion for each individual. (B) Contribution of colon cells to plasma in CRC patients and healthy individuals. (C) Contribution of neutrophils to plasma in CRC patients and healthy individuals.
[0065] Figures 7A to 7B (A) Comparison of genome coverage among LABS-seq (strands 1-6), MethylC-seq (strands 10-12), and EpiGnome (strands 7-9). (B) Comparison of CpG counts among LABS-seq (strands 1-6), MethylC-seq (strands 10-12), and EpiGnome (strands 7-9).
[0066] Figure 8 : The number of CpGs covered along the increasing cutoff readings of 1X, 3X, and 5X.
[0067] Figure 9 CpGs log2 enrichment scores spanned different genomic backgrounds: 3'UTR, LINE, SINE, exons, introns, intergenic, promoters, 5'UTR, CpG islands, LTR, simple repeats, and satellites.
[0068] Figure 10 Comparison of cfDNA LABS-seq libraries between LAB-seq, MethylC-seq, and EpiGnome in terms of (A) genome coverage; (B) CpG count; and (C) repeat level. All these cfDNA libraries were constructed using technical repeats from 100 pg and 1 ng of cancer-free cfDNA. Each site is a library. Detailed Implementation
[0069] Whole-genome sequencing of 5-methylcytosine (5mC) from samples with limited DNA material, such as liquid biopsy samples (down to subnanograms), remains challenging at single-base resolution due to strategic difficulties, bisulfite degradation, and low library complexity. This disclosure relates to a method known as linear amplification-based whole-genome bisulfite sequencing (LABS-seq), in which trace amounts of cell-free DNA material can be uniformly and linearly amplified via in vitro transcription without loss or bias of methylome information. This method demonstrates high genome coverage along repeat levels and with high mapping ratios. LABS-seq provides high-quality data, particularly at the subnanogram level, enabling the exploration of DNA modification dynamics and differential methylation signals in cell-free DNA liquid biopsies for tumor marker identification and tissue origin prediction.
[0070] I. Molecular Biology Methods
[0071] The methods disclosed herein include certain molecular biology applications known and described in detail in the art. For example, the methods disclosed herein include ligating an adaptor to a DNA molecule. A typical ligase reaction may include a DNA ligase (e.g., T4 DNA ligase), a ligase buffer, DNA fragments to be ligated by the ligation reaction, and a diluent, such as water or nuclease-free water. Ligases include ligases from *E. coli*, T4 DNA ligases from bacteriophage T4, mammalian ligases, and thermostable ligases (e.g., Ampligase DNA ligase) and variants and modified forms thereof. After the components of the ligation reaction are prepared and mixed, the reaction is typically carried out by incubating the components under conditions suitable for ligase activity. In some embodiments, these conditions may include a temperature of about 16°C and a time of about two or more two hours. In some embodiments, the time may be reduced by using a high concentration of T4 DNA. In some embodiments, the reaction is then heated (e.g., 65°C) to inactivate the enzyme.
[0072] In some implementations, the method includes the construction of a library of nucleic acid molecules. The term "library" refers to a collection (e.g., multiple vectors) containing nucleic acid molecules. Vectors can be mediators, constructs, arrays, or other physical carriers. A "vector" or "construct" (sometimes referred to as a gene delivery or gene transfer "vector") refers to a macromolecule, molecular complex, or viral particle, including polynucleotides to be delivered to host cells in vitro or in vivo. Polynucleotides can be linear or circular molecules. Those skilled in the art should be able to construct vectors using standard recombination techniques (see, for example, Maniatis et al., 1988; Ausubel et al., 1994, both incorporated herein by reference). Arrays contain a solid support to which nucleic acid probes are attached. Arrays typically contain multiple different nucleic acid probes that bind to the matrix surface at different known locations. These arrays, also known as “microarrays” or commonly referred to as “chips,” have been widely described in the art, for example, in U.S. Patents 5,143,854, 5,445,934, 5,744,305, 5,677,195, 6,040,193, 5,424,186, and Fodor et al. (1991), the entire contents of which are incorporated herein by reference. Techniques for synthesizing these arrays using mechanical synthesis methods are described, for example, in U.S. Patent 5,384,261, the entire contents of which are incorporated herein by reference. Although planar array surfaces are used in some respects, arrays can be fabricated on surfaces of virtually any shape or even multiple surfaces. Arrays can be nucleic acids on beads, gels, polymer surfaces, fibers such as optical fibers, glass, or any other suitable matrix, see U.S. Patents 5,770,358, 5,789,162, 5,708,153, 6,040,193, and 5,800,992, the entire contents of which are incorporated herein by reference. Vectors may contain restriction endonuclease sites, which can be used for further cloning or other molecular biological processes. Vectors may also contain primer binding sites, which can be used for sequencing of PCR amplification and / or for quantitative PCR studies.
[0073] In some embodiments, nucleic acids are purified. Purification of nucleic acids can be performed between any of the steps described herein. In some embodiments, nucleic acid purification does not include column filtration. In some embodiments, purification includes phenol-chloroform, magnetic beads, silica-based methods (e.g., silica membranes in the presence of ionizing salts), and anion exchange. In some embodiments, the purification comprises SPRI (Solid-Phase Reversible Immobilization) beads. SPRI beads are paramagnetic (magnetic only in a magnetic field), and this prevents them from clumping and detaching from solution. Each bead is made of polystyrene and surrounded by a layer of magnetite coated with carboxyl molecules. It is these that reversibly bind DNA in the presence of the “crowding agent” polyethylene glycol (PEG) and a salt (20% PEG, 2.5M NaCl is a magic mix). PEG causes negatively charged DNA to bind to the carboxyl groups on the bead surface. Since immobilization depends on the concentrations of PEG and salt in the reaction, the bead-to-DNA volume ratio is important. SPRI is particularly suitable for the removal of low concentrations of DNA.
[0074] Some embodiments of this disclosure relate to primer extension. Primer extension refers to annealing a primer to DNA (referred to as the template because it serves as the template for generating new strands) and adding polymerase and dNTPs, and culturing the reaction under conditions that allow the generation of DNA molecules extending from the primer in a 5'-3' direction. Primer annealing is generally achieved by incubating the single-stranded DNA with the primer under conditions suitable for binding to the single-stranded DNA. The primer should be at least partially complementary to allow binding. In some embodiments, the primer may have a non-complementary region that allows the addition of a sequence not present in the template DNA. The primer may have a region that is 100%, 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, or 85% (or any deducible range therebetween) complementary to the template DNA. This region is capable of annealing to the template DNA. The annealing step can be performed at a temperature that allows the template DNA to bind to the primer. This temperature is generally about 3 to 5 degrees Celsius below the primer's Tm. Once the primers are annealed, the annealed primers and template DNA can be incubated with dNTPs (or NTPs if the reaction is for RNA production) and polymerase under conditions suitable for producing nucleic acids via primer extension. The typical polymerase used is Taq polymerase (a thermophilic aquatic bacteria polymerase). The extension temperature depends on the temperature at which the polymerase exhibits activity. For Taq, the typical extension temperature is about 70 to 80 degrees Celsius.
[0075] In some embodiments, the DNA molecules of the methods disclosed herein can be amplified by PCR. A basic PCR setup requires several components and reagents, including: a DNA template containing the DNA target region for amplification; a DNA polymerase; an enzyme for polymerizing new DNA strands; thermostable Taq polymerase is particularly common because it may remain intact during high-temperature DNA denaturation; one or more DNA primers complementary to the 3' (three primers) ends of each sense and senseless strand of the DNA target; deoxyribonucleotide triphosphates or dNTPs (sometimes called "deoxynucleotide triphosphates"; nucleotides containing triphosphate groups), from which the DNA polymerase synthesizes new DNA strands; a buffer solution providing a suitable chemical environment for the activity and stability of the DNA polymerase; and divalent cations, typically magnesium (Mg) or manganese (Mn) ions; Mg 2+ It is the most conventional, but Mn 2+ It can be used for PCR-mediated DNA mutagenesis because of its higher Mn content. 2+ Concentration increases the error rate during DNA synthesis; monovalent cations, typically potassium (K) ions.
[0076] Embodiments of this disclosure relate to treating DNA molecules with bisulfite. Treatment of DNA with bisulfite converts cytosine residues into uracil. The method involves protecting cytosine residues from the effects of bisulfite. This method is further described herein. Bisulfite treatment may provide HSO3. - Arbitrary manipulation of ions, for example, using ions containing HSO3 - Salts of ions, such as sodium bisulfite. In some embodiments, the DNA treated with sodium bisulfite is single-stranded.
[0077] In some embodiments, the DNA molecule is transcribed in vitro. In vitro transcription includes a DNA template containing an RNA polymerase promoter, such as a double-stranded RNA promoter, and ribonucleotide triphosphates (NTPs). In some embodiments, the DNA template is purified. In some embodiments, the DNA template is straight-stranded. In vitro transcription may also include incubating the DNA template, RNA polymerase, and NTPs under conditions suitable for in vitro transcription. In some embodiments, in vitro transcription includes incubation with a buffer solution. In some embodiments, the buffer solution includes a redox reagent and a cation. In some embodiments, the redox reagent includes dithiothreitol (DTT). In some embodiments, the cation includes a divalent cation. In some embodiments, the cation includes a magnesium cation. In some embodiments, the DNA molecule used for in vitro transcription is straight-stranded and contains an RNA polymerase promoter in the correct orientation relative to the target sequence to be transcribed. For example, the minimal T7 promoter contains: TAATACGACTCACTATA (SEQ ID NO:1), which can be inserted into the 5' end of the target DNA. The uncontrolled transcript will have a sequence in the 3' region following the promoter. In some embodiments, in vitro transcription includes contacting a DNA molecule with a reaction mixture containing RNA polymerase and NTPs under conditions suitable for in vitro transcription of DNA molecules to produce corresponding RNA molecules. In some embodiments, the conditions include incubating the reaction mixture at a temperature where the polymerase is active. For example, T7 polymerase may require temperatures of 30°C to 45°C or 35°C to 40°C for at least 1 hour, 2 hours, 4 hours, 12 hours, or 24 hours (or any range derived therefrom), or longer than 24 hours. In vitro transcription systems are commercially available. For example, HiScribe is sold by New England Biological Laboratories. TM This provides an in vitro transcription kit that can be used with the methods of this disclosure. It is anticipated that any RNA polymerase can be used in this system along with its corresponding promoter. In some embodiments, the RNA promoter includes the SP6 promoter, and the RNA polymerase includes SP6 polymerase. In some embodiments, the RNA promoter includes the T3 promoter, and the polymerase includes T3 polymerase. In some embodiments, the RNA promoter includes the T7 promoter, and the polymerase includes T7 polymerase.
[0078] The methods disclosed herein relate to embodiments utilizing DNA-linkable adaptors. In some embodiments, the adaptor contains one or more primer-binding sites that can be used for sequencing, PCR, or real-time quantitative PCR. The adaptor may contain one or more cloning sites that allow for other replication biology techniques. The adaptor may contain a reporter gene, such as an antibiotic resistance gene, or a marker gene, such as a fluorescence-providing gene (green fluorescent protein and its derivatives).
[0079] II. Nucleic acid modification
[0080] In some embodiments, the method involves protecting specific cytosine or cytosine variants from bisulfite treatment, particularly bisulfite-mediated deamination. This protection may include chemical modifications of these variants such that the reads in bisulfite sequencing of the modified sequence can differ from those of the unmodified control nucleic acid.
[0081] Treatment of nucleic acids with bisulfite can convert cytosine residues to uracil, while 5-methylcytosine or 5-hydroxymethylcytosine residues remain unaffected. Therefore, bisulfite treatment can induce specific changes in DNA sequencing, producing single-nucleotide resolution information about the methylation state of DNA fragments that depend on the methylation state of individual cytosine residues. This information can be retrieved through various analyses of the altered sequences. One objective of this analysis can be reduced to distinguishing single nucleotide polymorphisms (cytosine and thymidine or uracil) resulting from bisulfite conversion.
[0082] Some embodiments involve 5fC and / or 5caC modifications to protect cytosine from bisulfite treatment. Other embodiments involve 5fC and / or 5caC modifications to facilitate bisulfite deamination. Exemplary modifications that can be used in the disclosed methods are described below for various purposes, such as protecting cytosine from bisulfite deamination or for differential detection of various cytosine modifications.
[0083] A.5fC Modification
[0084] Some embodiments are directed to methods and compositions for modifying nucleic acids containing 5fC or for modifying, detecting, and / or evaluating 5fC in nucleic acids. In some aspects, nucleic acids are modified to protect the 5fC from bisulfite-mediated deamination. For example, nucleic acids can be modified to oximes using compounds comprising hydroxylamine groups (such as R-NH-OH), hydrazine groups (such as R-NH-NH2), or acylhydrazine groups (such as RC(=O)-NH-NH2).
[0085] The methods described herein can be used to bind or link functional groups (e.g., hydroxylamine) to nucleic acids. Such binding or linking of functional groups allows for further labeling or tagging of cytosine residues with biotin or other tags. Labeling or tagging of the 5fC can be performed using, for example, click chemistry or other functional groups / linking groups known to those skilled in the art. Labelled or tagged nucleic acid fragments containing the 5fC can be enriched, isolated, detected, and / or evaluated.
[0086] Hydroxylamine groups that can be used in certain applications include those having the following general formula or having functional groups including the following general formula:
[0087]
[0088] R1 and R2 are hydrogen atoms, R3 is selected from hydrogen, lower alkyl groups, and aromatic groups; and these are water-soluble salts of hydroxylamines. Lower alkyl groups can typically have 1 to 8 carbon atoms, and aromatic groups can be, for example, phenyl, benzyl, and tolyl.
[0089] Non-limiting examples of suitable compounds containing hydroxylamine include hydroxylamine; hydroxylamine hydrochloride; hydroxylamine sulfate; hydroxylamine phosphate; O-methylhydroxylamine; O-hexylhydroxylamine; O-pentylhydroxylamine; O-benzylhydroxylamine; in particular, O-ethylhydroxylamine (EtONH2), or any O-alkylated or O-aromaticated hydroxylamine may be used.
[0090] In some respects, compounds that produce hydroxylamine when added to aqueous systems are also applicable.
[0091] Compounds containing a hydroxylamine group can also include substituted hydroxylamine derivatives. If the hydroxyl hydrogen is substituted, it is called an O-hydroxylamine. Similar to common amines, primary, secondary, and tertiary hydroxylamines can be distinguished, the latter two referring to compounds in which two or three hydrogens are substituted, respectively.
[0092] "Hydrazine" can refer to a divalent group –NR 1 R 2 -NH2, where R 1 and R 2 It can be alkyl, aromatic, or benzyl. As used herein, "hydrazine" includes, but is not limited to, hydrazine, acylhydrazine, aminourea, carbazine, aminothiourea, thiocarbazine, hydrazine carboxylate, and hydrazine carbonate. Examples of hydrazine used herein include N-alkylhydrazine, N-arylhydrazine, N-benzylhydrazine, N,N-dialkylhydrazine, N,N-diarylhydrazine, N,N-dibenzylhydrazine, N,N-alkylbenzylhydrazine, N,N-arylbenzylhydrazine, and N,N-alkylarylhydrazine.
[0093] "Hydrazine" can refer to a common functional group characterized by a nitrogen-nitrogen covalent bond with four substituents, at least one of which is an acyl group. The general structure of a hydrazine group can be RC(=O)-NR. 3 -NH2 or R-(SO2)R 3 -NH2, where R can be an alkyl or aromatic group, R 3It can be hydrogen, alkyl, aromatic, or benzyl. Important members of this class are sulfonyl hydrazides, such as p-toluenesulfonyl hydrazide, which are useful reagents in organic chemistry, such as in the Shapiro reaction. This reagent can be prepared by reacting toluenesulfonyl chloride with hydrazine. Examples of acyl hydrazides used herein include p-toluenesulfonyl hydrazide, N-acyl hydrazide, N,N-alkylacyl hydrazide, N,N-benzylacyl hydrazide, N,N-arylacyl hydrazide, N-sulfonyl hydrazide, N,N-alkylsulfonyl hydrazide, N,N-benzylsulfonyl hydrazide, and N,N-arylsulfonyl hydrazide.
[0094] Modification of B.5caC
[0095] Some embodiments relate to methods and compositions for modifying nucleic acids containing 5caC. In some aspects, the target nucleic acid is modified to protect 5caC from bisulfite-mediated deamination. For example, the nucleic acid can be converted to an amide by reacting with an amine-containing compound or a compound containing an amine group.
[0096] Amine compounds can have the general formula NH2-R, where R = alkyl such as –CH2CH3 or –CH-(CH3)2; cycloalkyl, aromatic, or benzyl. For example, the amino group can be alkylamine, cycloalkylamine, benzylamine, xyleneamine, or hydroxylamine. Amine compounds can be alkylamine, cycloalkylamine, or benzylamine. The amino group can be attached to a detected label or compound such as biotin.
[0097] An amine is a functional group containing a basic nitrogen atom with a lone pair of electrons. Amines are derivatives of ammonia in which one or more hydrogen atoms are replaced by substituents such as alkyl or aromatic groups. Compounds containing an amine can be aliphatic or aromatic amines, primary amines, secondary amines, tertiary amines, or cyclic amines.
[0098] Aliphatic amines do not have an aromatic ring directly bonded to the nitrogen atom. Aromatic amines, on the other hand, have their nitrogen atom bonded to an aromatic ring, as in different anilines. The aromatic ring reduces the basicity of the amine, depending on its substituents. The presence of an amino group significantly increases the reactivity of the aromatic ring due to its electron-donating effect.
[0099] Amines can also be divided into four subclasses:
[0100] Primary amines are formed when one of the three hydrogen atoms in ammonia is replaced by an alkyl or aromatic group. Important primary alkyl amines include methylamine, ethanolamine (2-aminoethanol), and the buffer tris(hydroxymethyl)aminomethane, while primary aromatic amines include aniline.
[0101] Secondary amines have two substituents (alkyl, aromatic, or both) bonded to nitrogen and a hydrogen atom. Important examples include dimethylamine and methylethanolamine, while an example of an aromatic amine is diphenylamine.
[0102] Tertiary amines – In tertiary amines, all three hydrogen atoms are replaced by organic substituents. Examples include trimethylamine or triphenylamine; trimethylamine has a distinct fishy odor.
[0103] Cyclic amines are secondary or tertiary amines. Examples of cyclic amines include the three-membered aziridine and the six-membered piperidine. N-methylpiperidine and N-phenylpiperidine are examples of cyclic tertiary amines.
[0104] In some embodiments, a coupling agent may be used to modify or label 5caC with an amine or thiol group, such as a carbodiimide derivative, i.e., a compound having a functional group consisting of the formula R1N=C=NR2. In specific embodiments, R1 and R2 may be the same or different and may be alkyl or aromatic groups. For example, the carbodiimide derivative may be 1-ethyl-3-(3-dimethylaminopropyl)carbodiimide (EDC) or N,N'-dicyclohexylcarbodiimide (DCC).
[0105] Modification with C.5mC and / or 5hmC
[0106] In some embodiments, the 5mC and / or 5hmC in the nucleic acid can be modified, for example, oxidized to 5caC. In further embodiments, the oxidation of 5mC and / or 5hmC to 5caC can be accomplished by contacting the nucleic acid with a methylcytosine dioxygenase (e.g., TET1, TET2, and TET3) or an enzyme having activity or catalytic regions similar to methylcytosine dioxygenases. The nucleic acid can be isolated nucleic acid, nucleic acid in a sample, nucleic acid modified by the methods described above (modification of 5fC and / or 5caC), or unmodified nucleic acid.
[0107] 5-Methylcytosine (5mC) in DNA plays an important role in gene expression, genomic imprinting, and transposon repression. It is known that 5mC can be converted to 5-hydroxymethylcytosine (5hmC) by the Tet (decyl-eleventh translocation) protein. Therefore, embodiments of this disclosure include methods wherein 5mC is oxidized to 5hmC and / or 5hmC is oxidized to 5fC and / or 5fC is oxidized to 5caC. The Tet protein can convert 5mC to 5-formylcytosine (5fC) and 5-carboxycytosine (5caC) in an enzymatically dependent manner (Ito et al., 2011, incorporated herein by reference).
[0108] Modification of 5mC can be carried out using enzymes or chemical reagents to catalyze or cause the modified portion to be converted into 5mC, thereby producing modified 5mC (m5mC). This strategy is beneficial for incorporating modifications of 5mC to label or tag 5mC in eukaryotic nucleic acids.
[0109] Chemical tags can be used to determine the precise location of 5mC in a high-throughput manner. The inventors demonstrate that 5mC modification makes the tagged DNA resistant to digestion and / or polymerization by certain restriction enzymes. In some respects, modified and unmodified genomic DNA can be treated with restriction enzymes, followed by various sequencing methods to reveal the precise location of each cytosine modification that hinders digestion.
[0110] The inventors state that modified portions, such as functional groups (e.g., azide groups), can be incorporated into DNA using the methods described herein. This incorporation of functional groups allows for further labeling or tagging of cytosine residues with biotin or other tags. The labeling or tagging of 5mC can be performed using, for example, click chemistry or other functional groups / linking groups known to those skilled in the art. DNA fragments containing m5mC-labeled or tagged fragments can be isolated and / or evaluated using methods currently used to evaluate modifications containing 5mC of nucleic acids.
[0111] Furthermore, the methods and compositions disclosed herein can be used to introduce spatially large groups into 5mC. The presence of large groups on the DNA template strand will interfere with the synthesis of nucleic acid chains by DNA polymerase or RNA polymerase, or interfere with the efficient cleavage of DNA by restriction endonucleases or inhibit other enzymatic modifications of nucleic acids containing 5mC. As a result, primer extension or other assays can be performed, for example to evaluate partially extended primers of a specific length, and modification sites can be revealed by sequencing the partially extended primers. Other methods utilizing this chemical labeling approach are also considered.
[0112] Certain embodiments pertain to methods and compositions for modifying, detecting, and / or evaluating 5hmC in nucleic acids. In some aspects, 5hmC is glycosylated. In further aspects, 5hmC is linked to a labeled or diluted glucose moiety. In some aspects, the target nucleic acid is contacted with a β-glucosyltransferase and a UDP matrix containing a modified or modifiable glucose moiety. Using the methods described herein, a wide variety of detectable groups (biotin, fluorescently labeled radioactive groups, etc.) can be linked to 5hmC via glucose modification. Methods and compositions are described in PCT application PCT / US2011 / 031370, filed April 6, 2011, which is again incorporated herein by reference in its entirety.
[0113] Modification of 5hmC can be performed using the enzyme β-glucosyltransferase (βGT) or a similar enzyme that catalyzes the transfer of a glucose moiety from uridine diphosphate glucose to the hydroxyl group of 5hmC, thereby producing β-glycosyl-5-hydroxymethylcytosine (ghmC). The inventors have discovered that this enzymatic glycosylation provides a strategy for incorporating modified glucose molecules to tag or label 5hmC in eukaryotic nucleic acids. For example, a chemically modified glucose molecule containing an azide (N3) group can be covalently linked to 5hmC via enzymatically catalyzed glycosylation. Phosphine activators, including but not limited to biotin-phosphine, fluorophore-phosphine, and NHS-phosphine, or other affinity tags, can then be specifically attached to the glycosylated 5hmC through a reaction with the azide.
[0114] 5mC and / or 5hmC can be modified directly or indirectly with several functional groups or labeled molecules. One example is the oxidation of 5mC followed by labeling with a functionalized or labeled glucose molecule. In some embodiments, 5mC can be modified with a modified moiety or functional group before further modification via the linking of the glucosyl moiety.
[0115] In another embodiment, functionalized or labeled glucose molecules may be used in combination with βGT to modify 5hmC in nucleic acid polymers such as DNA or RNA. In some aspects, the βGT UDP matrix contains a functionalized or labeled glucose moiety.
[0116] In another respect, the modification portion can be modified or functionalized using click chemistry or other coupling chemistry known in the art. Click chemistry is a chemical philosophy introduced by K. Barry Sharpless in 2001 (Kolb et al., 2001; Evans, 2007) that describes the chemistry of rapidly and reliably generating substances by linking small units.
[0117] The inventors state that functional groups (such as azide groups) can be incorporated into DNA using the methods described herein. This incorporation of functional groups allows for further labeling or tagging of cytosine residues with biotin or other tags. The labeling or tagging of 5hmC can be performed using, for example, click chemistry or other functional groups / linking groups known to those skilled in the art. The labeled or tagged DNA fragments containing 5hmC can be isolated and / or evaluated using modified methods currently used to evaluate nucleic acids containing 5mC.
[0118] In some respects, it is possible to assess differential modifications of nucleic acids between two or more samples. Studies including those of the heart, liver, lungs, kidneys, muscles, testes, spleen, and brain have shown that under normal conditions, 5hmC is primarily found in normal brain cells. Assessing and comparing 5hmC levels can be used to evaluate various disease states and compare diverse nucleic acid samples.
[0119] D.TET protein
[0120] Deca-11 translocation (TET) proteins are a family of DNA hydroxylases that have been found to possess enzymatic activity at the 5-methyl cytosine (5-methylcytosine [5mC]). The TET protein family includes three members: TET1, TET2, and TET3. TET proteins are thought to convert 5mC to 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), and 5-carboxycytosine (5caC) through three consecutive oxidation reactions.
[0121] The first member of the ET family of proteins, the TET1 gene, was first detected in acute myeloid leukemia (AML) as a fusion partner for the histone H3 Lys 4 (H3K4) methyltransferase MLL (mixed lineage leukemia) (Ono et al., 2002; Lorsbach et al., 2003). The human TET1 protein was first found to possess enzymatic activity, capable of hydroxylating 5mC to 5hmC (Tahirani et al., 2009). Subsequently, all members of the mouse TET protein family (TET 1–3) have been shown to possess 5mC hydroxylase activity (Ito et al., 2010).
[0122] TET proteins typically possess several conserved domains, including a CXXC zinc finger domain with high affinity for aggregated unmethylated CpG dinucleotides, a typical catalytic domain of iron(II)- and 2-oxoglutarate (2OG)-dependent dioxygenases, and a cysteine-rich region (Wu and Zhang, 2011; Tahiliani et al., 2009).
[0123] In some implementations, TET1, TET2, or TET3 is expected to be a human or mouse protein. Human TET1 is designated NM_030625.2; human TET2 is designated NM_001127208.2 or NM_017628.4; and human TET3 is designated NM_144993.1. Mouse TET1 is designated NM_027384.1; mouse TET2 is designated NM_001040400.2; and mouse TET3 is designated NM_183138.2.
[0124] E. β-glycosyltransferase (β-GT)
[0125] Glucosyl-DNA-β-glucosyltransferase (EC 2.4.1.28, β-glycosyltransferase (β-GT)) is an enzyme that catalyzes chemical reactions in which β-D-glucosyl residues are transferred from UDP-glucose to glucosylhydroxymethylcytosine residues in nucleic acids. This enzyme is similar in this respect to DNA β-glucosyltransferase. This enzyme belongs to the glycosyltransferase family, particularly hexosyltransferases. The systematic name for this class of enzymes is UDP-glucose:D-glucosyl-DNA-β-D-glucosyltransferase. Other commonly used names include T6-glucosyl-HMC-β-glucosyltransferase, T6-β-glucosyltransferase, uridine diphosphate glucose-glucosyldeoxyribonucleic acid, and β-glucosyltransferase.
[0126] In some respects, β-glucosyltransferase is a His-tagged fusion protein with an amino acid sequence (β-GT begins at amino acid 25 (met)):
[0127]
[0128] F. Functional Groups
[0129] Nucleic acids, especially cytosine and / or modified cytosine, can be modified directly or indirectly with several functional groups or labeled molecules. One example is oxidation at 5 mC followed by labeling with a functionalized, protecting agent or labeled glucose molecule. In some embodiments, 5 mC may be modified with the modified moiety or functional group before further modification via the linker of the glucosyl moiety.
[0130] In another embodiment, functionalized or labeled glucose molecules may be used in combination with βGT to modify 5hmC in nucleic acid polymers such as DNA or RNA. In some aspects, the βGT UDP matrix contains a functionalized or labeled glucose moiety.
[0131] In another respect, the modification portion can be modified or functionalized using click chemistry or other coupling chemistry known in the art. Click chemistry is a chemical philosophy introduced by K. Barry Sharpless in 2001 (Kolb et al., 2001; Evans, 2007) that describes the chemistry of rapidly and reliably generating substances by linking small units.
[0132] Chemical reactions that cause covalent bonding include, for example, cycloaddition reactions (such as the Diels-Alder reaction, the 1,3-dipolar cycloaddition Huisgen reaction, and similar “click reactions”), condensation, nucleophilic and electrophilic addition reactions, nucleophilic and electrophilic substitution reactions, addition and elimination reactions, alkylation reactions, rearrangement reactions, and any other known organic reactions involving functional groups.
[0133] Representative examples of functional groups include, but are not limited to, acyl halides, aldehydes, alkoxy groups, alkynes, amides, amines, aryloxy groups, azides, aziridines, azo compounds, carbamates, carbonyl groups, carboxyl groups, carboxylic acid esters, cyano groups, dienes, dienophiles, epoxy groups, guanidine groups, guanine groups, halides, acyl hydrazides, hydrazines, hydroxyl groups, hydroxylamines, imino groups, isocyanates, nitro groups, phosphates, phosphonates, sulfinyl groups, sulfonamides, sulfonates, thioalkoxy groups, thioaryloxy groups, thiocarbamates, thiocarbonyl groups, thiohydroxy groups, thioureas, and ureas, as these terms are defined below.
[0134] Exemplary first and second functional groups that are chemically complementary to the other functional group described herein include, but are not limited to, hydroxyl and carboxylic acids that form ester bonds; thiols and carboxylic acids that form thioester bonds; amines and carboxylic acids that form amide bonds; aldehydes and amines, hydrazides, acylhydrazides, hydroxylamines, phenylhydrazides, aminoureas, or aminothioureas that form Schiff bases (imine bonds); alkenes and dienes that react between them via cycloaddition reactions; and functional groups that can participate in click reactions.
[0135] Further examples of functional group pairs capable of reacting with another functional group include azides and alkynes, unsaturated carbon-carbon bonds (e.g., acrylates, methacrylates, maleimides) and thiols, unsaturated carbon-carbon bonds and amines, carboxylic acids and amines, hydroxyl groups and isocyanates, carboxylic acids and isocyanates, amines and isocyanates, and thiols and isocyanates. Additional examples include amines, hydroxyl groups, thiols, or carboxylic acids, as well as nucleophilic leaving groups (e.g., hydroxysuccinimides, halogens).
[0136] In some embodiments, the functional group may be a latent group that is exposed during a chemical reaction, such that a reaction (e.g., the formation of a covalent bond) occurs once the latent group is exposed. Exemplary such groups include, but are not limited to, functional groups as described above, which are protected by protecting groups that are unstable under selected reaction conditions.
[0137] Examples of unstable protecting groups include, for example, carboxylic acid esters, which can be hydrolyzed to alcohols and carboxylic acids upon exposure to acidic or basic conditions; silyl ethers, such as trialkylsilyl ethers, which can be hydrolyzed to alcohols by acid or fluoride ions; p-methoxybenzyl ethers, which can be hydrolyzed to alcohols, for example, by oxidative or acidic conditions; tert-butyloxycarbonyl and 9-fluorenemethoxycarbonyl, which can be hydrolyzed to amines upon exposure to basic conditions; sulfonamides, which can be hydrolyzed to sulfonates and amines upon exposure to suitable reagents such as samarium iodide or tributyltin hydride; acetals and ketals, which can be hydrolyzed to aldehydes or ketones, respectively, upon exposure to acidic conditions, and alcohols or diacetals. Alcohols; acyl groups (i.e., where a carbon atom is attached to two carboxylic acid ester groups), which can be hydrolyzed to ketones, for example, by exposure to Lewis acids; orthoesters (i.e., where a carbon atom is attached to three alkoxy or aryloxy groups), which can be hydrolyzed to carboxylic acid esters by exposure to weakly acidic conditions (which can be further hydrolyzed as described above); 2-cyanoethyl phosphates, which can be converted to phosphate esters by exposure to mildly alkaline conditions; methyl phosphates, which can be hydrolyzed to phosphate esters by exposure to strong nucleophiles; phosphate esters, which can be hydrolyzed to alcohols, for example, by exposure to phosphatases; and aldehydes, which can be converted to carboxylic acids, for example, by exposure to oxidants.
[0138] According to some embodiments of the present invention, the bond-forming reaction between the two functional groups (first and second functional groups) results in the formation of a linking group.
[0139] According to some embodiments of the present invention, exemplary linking portions formed between the first and second functional groups as described herein include, but are not limited to, amides, lactones, lactams, carboxylic acid esters (esters), cycloolefins (e.g., cyclohexene), heterobicyclic compounds, heteroaryl compounds, triazines, triazoles, disulfides, imines, aldehyde imines, ketimines, hydrazones, and saccharin. Other linking portions are defined below.
[0140] For example, reactions between diene and dienophile functional groups, such as the Diels-Alde reaction, will form a cycloalkene linker, and in most cases a cyclohexene linker. In another example, upon reaction with a carboxyl functional group, an amine functional group will form an amide linker. In another example, upon reaction with a carboxyl functional group, a hydroxyl functional group will form an ester linker. In another example, when reacting with another thiol functional group under oxidizing conditions, a thiol functional group will form a disulfide (--S--S--) linker, or when reacting with a halogen functional group or another leaving functional group, a thioether (thioalkoxy) linker will form. In yet another example, when reacting with an azide functional group via "click chemistry," an alkynyl functional group will form a triazole linker.
[0141] The term "click reaction," also known as "click chemistry," is a name frequently used to describe the Huisgen 1,3-dipolar cycloaddition of azides and alkynes to generate stepwise variants of 1,2,3-triazoles. This reaction is carried out under ambient conditions or under mild microwave irradiation, typically in the presence of a Cu(I) catalyst, and exhibits exclusive regioselectivity for the 1,4-disubstituted triazole product when mediated by a catalytic amount of Cu(I) salt (V. Rostovtsev, LG Green, VV Fokin, KB Sharpless, Angew. Chem. Int. Ed. 2002, 41, 2596; HC Kolb, M. Finn, KB Sharpless, Angew. Chem. Int. 2001, 40, 2004).
[0142] The “click reaction” is particularly relevant in the context of embodiments of the present invention because it can be carried out without damaging the DNA molecule, and it can attach the label to the 5hmC of the DNA molecule in a high chemical yield using mild conditions in an aqueous medium. The selectivity of this reaction allows the reaction to proceed with minimal or no use of protecting groups, which often results in cumbersome, multi-step synthetic processes.
[0143] G. DNA transposon markers
[0144] In some cases, nucleic acids are labeled with transposons. For example, nucleic acid molecules can be contacted with transposons and transposases to allow the transposons to integrate nonspecifically into the nucleic acid molecule.
[0145] As used throughout, the term transposon refers to a double-stranded DNA containing the nucleotide sequence necessary for the formation of a complex by a transposase or integrase, which functions in an in vitro transposition reaction. The transposon forms a complex, synaptic complex, or transposome complex. The transposon can also form a transposome composition with a transposase or integrase that recognizes and binds to the transposon sequence, and the complex is capable of inserting or transposing into target DNA, incubating with the target DNA in an in vitro transposition reaction.
[0146] Labeling nucleic acid molecules with transposons can also involve fragmenting the labeled DNA. In some embodiments, transposases can be used to catalyze the integration of oligonucleotides into the target nucleic acid at a high density (e.g., about 300 base pairs per molecule). For example, transposases such as Nextera's TRANSPOSOME TM The technology can be used to generate random dsDNA disruption. TMThe complex contains a free transposon terminus and a transposase. When this complex is incubated with dsDNA, the DNA is fragmented and the transfer strand of the transposon-terminal oligonucleotide is covalently linked to the end of the DNA fragment. In some embodiments, it is linked to the 3' end. In some embodiments, it is linked to the 5' end. In some applications, the transposon terminus can be attached with a primer site. By changing the buffer and reaction conditions (e.g., TRANSPOSOME...),... TM The concentration of the complex can control the size distribution of fragmented and labeled DNA libraries.
[0147] In some embodiments, the transposon also includes a tag or affinity label, such as biotin. Other affinity labels include electronic tags, Flag tags, HA tags, His tags, Myc tags, etc. In some embodiments, the affinity label is attached to the end of the P7 interposer. In some embodiments, the affinity label is attached to the 5' end of the interposer.
[0148] III. Sequencing Methods
[0149] A. Massive parallel signal sequencing (MPSS).
[0150] The first of the next-generation sequencing technologies, massively parallel signal sequencing (or MPSS), was developed by Lynx Therapeutics in the 1990s. MPSS is a bead-based approach that uses adaptor ligation followed by adaptor decoding, a complex method for reading sequences in four-nucleotide increments. This approach makes it susceptible to sequence-specific bias or loss of specific sequences. Because the technology was so complex, MPSS could only be performed "internally" at Lynx Therapeutics, and no DNA sequencing machines were sold to independent laboratories. In 2004, Lynx Therapeutics merged with Solexa (later acquired by Illumina), leading to the development of synthetic sequencing, a simpler method derived from Manteia Predictive Medicine, which rendered MPSS obsolete. However, the fundamental characteristics of MPSS output are typical of later "next-generation" data types, consisting of thousands of short DNA sequences. In the case of MPSS, these are often used for cDNA sequencing to measure gene expression levels. In fact, powerful Illumina HiSeq2000, HiSeq2500, and MiSeq systems are all based on MPSS.
[0151] B. Polymerase cloning and sequencing.
[0152] The polymerase cloning sequencing method, developed in George M. Church's lab at Harvard, was one of the first next-generation sequencing systems and was used for whole-genome sequencing in 2005. It combines in vitro paired-tag libraries with emulsion PCR, automated microscopy, and ligation-based sequencing chemistry to sequence the *E. coli* genome with >99.9999% accuracy at approximately one-ninth the cost of Sanger sequencing. The technology was licensed to Agencourt Biosciences, subsequently spun off into Agencourt Personal Genomics, and eventually integrated into the Applied Biosystems SOLiD platform, now owned by Life Technologies.
[0153] C.454 pyrosequencing.
[0154] 454Life Sciences developed a parallel version of pyrosequencing before the company was acquired by Roche Diagnostics. This method amplifies DNA within droplets in an oil solution (emulsion PCR), where each droplet contains a DNA template linked to a single primer-covered bead, which then forms cloning colonies. The sequencing machine contains numerous microliter-volume wells, each containing a single bead and sequencing enzyme. Pyrosequencing uses luciferase to generate light for detecting individual molecules added to nascent DNA and uses combined data to generate sequence reads. Compared to Sange sequencing at one end and Solexa and SOLiD sequencing at the other, this technology offers improved intermediate read lengths and cost per base.
[0155] D. Illumina (Solexa) sequencing.
[0156] Solexa, now part of Ilumina, developed a sequencing method based on reversible dye terminator technology and engineered polymerases, which was developed internally. Termination chemistry was developed within Solexa, and the concept for the Solexa system was invented by Balasubramanian and Klennerman of the Department of Chemistry at the University of Cambridge. In 2004, Solexa acquired Manteia Predictive Medicine to gain access to massively parallel sequencing technology based on “DNA clusters,” which involves the clonal amplification of DNA on a surface. The cluster technology was acquired in conjunction with Lynx Therapeutics in California. Solexa Ltd. later merged with Lynx to form Solexa Inc.
[0157] In this method, DNA molecules and primers are first attached to a glass slide and amplified using polymerase, forming localized clonal DNA colonies, later known as "DNA clusters." To determine the sequence, four reversible terminator bases (RT-bases) are added, and uncombined nucleotides are washed away. A camera captures images of the fluorescently labeled nucleotides, and then the dye is chemically removed from the DNA along with a 3' terminator, allowing the next cycle to begin. Unlike pyrosequencing, the DNA strand extends one nucleotide at a time, and images can be acquired at delayed moments, allowing for the capture of very large arrays of DNA colonies using continuous images from a single camera.
[0158] Separating enzymatic reactions from image capture enables optimal throughput and theoretically unlimited sequencing capabilities. With optimal configuration, the final achievable instrument throughput is determined solely by the camera's analog-to-digital conversion rate, multiplied by the number of cameras, and divided by the number of pixels per DNA colony required for optimal visualization (approximately 10 pixels per colony). In 2012, with cameras operating at A / D conversion rates exceeding 10 MHz and available optical, fluidic, and enzymatic capabilities, throughput could be multiples of one million nucleotides per second, roughly equivalent to covering one human genome per hour per instrument, and re-sequencing one human genome per instrument (with a single camera) per day (approximately 30 times).
[0159] E.SOLiD sequencing
[0160] Applied Biosystems (now Life Technologies brand)'s SOLiD technology performs sequencing via ligation. Here, pools of all possible fixed-length oligonucleotides are labeled according to their sequencing positions. The oligonucleotides are annealed and ligated; preferential ligation by a DNA ligase used to match sequences creates a signal for the nucleotides at that position. DNA is amplified by emulsion PCR before sequencing. The resulting beads, each containing a single copy of the same DNA molecule, are deposited on a glass slide. The result is a sequence quantity and length comparable to Illumina sequencing. This ligation-based sequencing method has reportedly encountered some problems when sequencing palindromic sequences.
[0161] F. Ion Torrent semiconductor sequencing.
[0162] Ion Torrent Systems Inc. (now owned by Life Technologies) developed a system based on standard sequencing chemistry but employing a novel semiconductor-based detection system. Unlike the optical methods used in other sequencing systems, this approach is based on the detection of hydrogen ions released during DNA polymerization. Microwells containing the template DNA strand to be sequenced are filled with a single type of nucleotide. If the introduced nucleotide is complementary to the dominant template nucleotide, it is incorporated into the growth of the complementary strand. This results in the release of hydrogen ions, which triggers a highly sensitive ion sensor, indicating that a reaction has occurred. If a homopolymeric repeat sequence is present in the template sequence, multiple nucleotides are introduced in a single cycle. This results in a correspondingly large amount of hydrogen released and a proportionally higher electronic signal.
[0163] G. DNA nanosphere sequencing.
[0164] DNA nanosphere sequencing is a high-throughput sequencing technology used to determine the whole genome sequence of an organism. Complete Genomics uses this technology to sequence samples submitted by independent researchers. The method uses rolling circle replication to amplify small fragments of genomic DNA into DNA nanospheres. Unbound sequencing via ligation is then used to determine the nucleotide sequence. Compared to other next-generation sequencing platforms, this DNA sequencing method allows for sequencing of a large number of DNA nanospheres per run at low reagent cost. However, only short DNA sequences can be determined from the DNA nanospheres, making it difficult to map short reads to a reference genome. This technology has been used in several genome sequencing projects and is planned for use in more projects.
[0165] H.Heliscope single-molecule sequencing.
[0166] Heliscope sequencing is a single-molecule sequencing method developed by Helicos Biosciences. It uses DNA fragments with added poly-A tails as adaptors, which are attached to the surface of mobile cells. The following steps involve extension-based sequencing, using fluorescently labeled nucleotides to cyclically wash the mobile cells (one nucleotide type at a time, similar to the Sanger method). Reads are then performed using the Heliscope sequencer. These reads are short, up to 55 bases per run, but recent improvements allow for more precise readings of one type of nucleotide extension. This sequencing method and instrument were used to sequence the genome of M13 bacteriophage.
[0167] I. Single-molecule real-time (SMRT) sequencing.
[0168] SMRT sequencing is based on a synthesis-based sequencing approach. DNA is synthesized in a zero-mode waveguide (ZMW)—a small container located at the bottom of a well, similar to a trapping tool. Sequencing is performed using a modified polymerase (attached to the bottom of the ZMW) and fluorescently labeled nucleotides that flow freely in solution. Wells are constructed using a method that detects fluorescence only at the bottom of the well. The fluorescent label separates from the nucleotides as it is incorporated into the DNA strand, leaving the unmodified DNA strand. According to Pacific Biosciences, the developers of SMRT technology, this technique allows the detection of nucleotide modifications (such as cytosine methylation). This occurs by observing polymerase kinetics. This method allows for the reading of 20,000 or more nucleotides, with an average read length of 5 kilobases.
[0169] IV. Instructions for Use
[0170] A. Identification of DNA methylation variants
[0171] The field of DNA methylation analysis has recently expanded with the identification of various cytosine variants. Traditional DNA methylation involves transferring a methyl group to the carbon-5 position of cytosine to produce 5-methylcytosine (5mC). However, studies have shown that the Tet family of cytosine oxidases are involved in oxidizing 5-methylcytosine to 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), and 5-carboxycytosine (5caC).
[0172] 5-Formylcytosine (5fC) is one of the DNA variants produced when Tet enzymes act on 5-hydroxymethylcytosine. Further oxidation of 5-formylcytosine by Tet enzymes results in the conversion of 5-carboxycytosine. Oxidation of 5-methylcytosine via different DNA methylation variants is believed to represent a DNA demethylation mechanism, and this demethylation pathway is believed to be useful in development and germ cell programming. 5-Formylcytosine is present in mouse embryonic stem (ES) cells and major mouse organs. This DNA modification is also observed in post-fertilized parental pronuclei, accompanied by the disappearance of 5-methylcytosine, indicating its involvement in DNA demethylation.
[0173] 5-Carboxycytosine (5caC) has been identified as one of the DNA methylation variants produced during the oxidation of 5-hydroxymethylcytosine and subsequently 5-formylcytosine by Tet enzymes. The oxidation of 5-methylcytosine to 5-carboxycytosine is believed to represent a DNA demethylation mechanism, and this demethylation pathway is believed to be useful in development and germ cell programming. It has been proposed to isolate 5caC from genomic DNA via thymine DNA glycosylase (TDG), which returns the cytosine residues to their unmodified state. 5-Carboxycytosine has been identified in mouse embryonic stem (ES) cells. This DNA modification, accompanied by the disappearance of 5-methylcytosine in the post-fertilized parental pronucleus, further supports the view that this variant is part of a DNA demethylation pathway.
[0174] 5-Methylcytosine (5mC) is a DNA modification resulting from the transfer of a methyl group from S-adenosylmethionine (also known as AdoMet or SAM) to the carbon 5 position of a cytosine residue. This transfer is catalyzed by DNA methyltransferases (DNMTs). 5-Methylcytosine is the most common and widely studied form of DNA methylation. It typically occurs within CpG dinucleotide motifs, although non-CpG methylation has been identified in embryonic stem cells.
[0175] 5-Hydroxymethylcytosine (5hmC) is a DNA methylation modification resulting from the enzymatic oxidation of 5-methylcytosine (5mC) by iron-dependent deoxygenases of the 3Tet family. Elevated levels of 5-hydroxymethylcytosine can be found in certain mammalian tissues, such as mouse Purkinje cells and granule neurons. Alternatively, 5hmC can be generated by adding formaldehyde to DNA cytosine via the DNMT protein.
[0176] Other methods for distinguishing epigenetic modifications have been provided. It is anticipated that the current method can be applied and combined with other methods disclosed in the art. Examples of methods disclosed in the art include U.S. Provisional Patent Application No. 61 / 656924, U.S. Patent Application No. 13 / 095505, U.S. Provisional Patent Application No. 61 / 321198, PCT Application No. PCT / US2011 / 031370, PCT Application No. PCT / US2012 / 032489, U.S. Provisional Patent Application No. 61 / 472435, Provisional Patent Application No. 61 / 512334, PCT Application No. PCT / US2014 / 032997 and PCT / US2018 / 021591, and U.S. Publication No. 20140178881, each of which is incorporated herein by reference in its entirety. In some embodiments, the current method may or may not include the steps described in the patent applications cited above.
[0177] B. Clinical and Diagnostic Applications
[0178] The methods disclosed herein are advantageous for evaluating DNA for clinical and / or diagnostic purposes. Some embodiments relate to methods for evaluating samples comprising DNA molecules. This evaluation may be a detection or assay of a specific cytosine modification or a differential detection or assay of a specific modification.
[0179] Samples can be obtained from biopsies, such as needle aspiration biopsies, core needle biopsies, vacuum-assisted biopsies, excisional biopsies, excisional biopsies, drill biopsies, scraping biopsies, and skin biopsies. In some embodiments, samples can be obtained from biopsies of cancerous tissue using any of the foregoing biopsy methods. In other embodiments, samples can be obtained from any tissue described herein, including but not limited to gallbladder, skin, heart, lung, breast, pancreas, liver, muscle, kidney, smooth muscle, bladder, colon, intestine, brain, prostate, esophagus, or thyroid tissue. Alternatively, samples can be obtained from any other source, including but not limited to blood, sweat, hair follicles, oral tissue, tears, menstrual blood, feces, or saliva. In some aspects, samples are obtained from cystic fluid or fluid from a tumor or growth. In other embodiments, the cyst, tumor, or growth is from the colon or rectum. In some aspects of the present method, any medical expert, such as a physician, nurse, or medical technician, can obtain biological samples for testing. Furthermore, biological samples can be obtained without the assistance of a medical expert.
[0180] Samples may include, but are not limited to, tissues, cells, or biological material derived from or derived from cells of an object. In some embodiments, the sample includes cell-free DNA. In some embodiments, the sample includes a fertilized egg, zygote, blastocyst, or blastomeres. Biological samples can be heterogeneous or homogeneous populations of cells or tissues. Biological samples can be obtained using any method known in the art that can provide samples suitable for the analytical methods described herein. Samples can be obtained using non-invasive methods, including but not limited to: scraping the skin or cervix, wiping the cheek, collecting saliva, collecting urine, collecting feces, collecting menstrual blood, tears, or semen.
[0181] In some embodiments, the methods of this disclosure can be used to discover novel biomarkers for diseases or conditions. In some embodiments, the methods of this disclosure can be performed on samples from patients to provide prognosis for certain diseases or conditions in patients. In some embodiments, the methods of this disclosure can be performed on samples from patients to predict patient response to specific treatments. In some embodiments, diseases include cancer. For example, cancer can be pancreatic cancer, colon cancer, acute myeloid leukemia, adrenocortical carcinoma, AIDS-related cancer, AIDS-related lymphoma, anal cancer, appendiceal cancer, astrocytoma, pediatric cerebellar or basal cell carcinoma, bile duct cancer, extrahepatic bladder cancer, bone cancer, osteosarcoma / malignant fibrous histiocytoma, brainstem glioma, brain tumor, cerebellar astrocytoma brain tumor, brain astrocytoma / malignant glioma brain tumor, ependymoma brain tumor, medulloblastoma brain tumor, supratentorial primitive neuroectodermal tumor brain tumor, visual pathway and hypothalamic glioma, breast cancer, lymphoma, bronchial adenoma / carcinoid, tracheal cancer, Burkitt lymphoma, Carcinoid tumors, pediatric carcinoid tumors, primary gastrointestinal cancers of unknown origin, central nervous system lymphomas, primary cerebellar astrocytomas, pediatric astrocytomas / malignant gliomas, pediatric cervical cancer, childhood cancers, chronic lymphocytic leukemia, chronic myeloid leukemia, chronic myeloproliferative disorders, cutaneous T-cell lymphomas, desmoplastic small round cell tumors, endometrial cancer, ependymoma, esophageal cancer, Ewing's disease, pediatric extragonadal germ cell tumors, extrahepatic bile duct cancer, ocular cancer, intraocular melanoma, retinoblastoma, gallbladder cancer, gastric (stomach) cancer, gastrointestinal carcinoid tumors, gastrointestinal stromal tumors (GIST), germ cell tumors: cranial Extragonadal or ovarian tumors, gestational trophoblastic tumors, brainstem gliomas, gliomas, pediatric astrocytomas, pediatric visual pathway and hypothalamic gliomas, gastric carcinoid tumors, hairy cell leukemia, head and neck cancer, gastric cardia cancer, hepatocellular carcinoma, Hodgkin's lymphoma, hypopharyngeal cancer, hypothalamic and visual pathway gliomas, pediatric intraocular melanoma, islet cell carcinoma (endocrine pancreas), Kaposi's sarcoma, renal cell carcinoma, laryngeal cancer, leukemia, acute lymphoblastic leukemia (also known as lymphocytic leukemia), acute myeloid leukemia, chronic lymphocytic leukemia (also known as chronic lymphocytic leukemia). Leukemia, chronic myeloid leukemia (also known as chronic myeloid leukemia), pilocellular carcinoma of the lip, oral cancer, liposarcoma, primary liver cancer, non-small cell lung cancer, small cell lung cancer, lymphoma, AIDS-related lymphoma, Burkitt lymphoma, cutaneous T-cell lymphoma, Hodgkin lymphoma, non-Hodgkin lymphoma (the old classification of all lymphomas other than Hodgkin lymphoma), primary central nervous system lymphoma, Waldenstrom's macroglobulinemia, malignant fibrous histiocytoma of bone / osteosarcoma, medulloblastoma in children, melanoma, intraocular (ocular) melanoma, Merkel cell carcinoma, malignant mesothelioma in adults, mesothelioma in children.Metastatic squamous neck cancer, oral cancer, multiple endocrine neoplasia syndrome, multiple myeloma / plasma cell tumor, mycosis fungoides, myelodysplastic syndrome, myelodysplastic / myeloproliferative disorders, chronic myeloid leukemia, adult acute myeloid leukemia, childhood acute myeloid leukemia, multiple myeloma, chronic myeloproliferative disorders, nasal cavity and paranasal sinus cancer, nasopharyngeal carcinoma, neuroblastoma, oral cancer, oropharyngeal cancer, osteosarcoma / malignant, osteofibrous histiocytoma, ovarian cancer, ovarian epithelial carcinoma (surface epithelial-stromal tumor), ovarian germ cell tumor, low-potency ovarian tumor, pancreatic cancer, islet cell paranasal sinus and nasal cavity cancer, parathyroid carcinoma, penile cancer, pharyngeal cancer, pheochromocytoma, pineal astrocytoma, pineal germ cell tumor, pineal blastoma, and supratentorial primitive neuroectodermal tumor. Tumors, including: pituitary adenoma in children, plasmacytoma / multiple myeloma, pleural pulmonary blastoma, primary central nervous system lymphoma, prostate cancer, rectal cancer, renal cell carcinoma (renal carcinoma), transitional cell carcinoma of the renal pelvis and ureter, retinoblastoma, rhabdomyosarcoma, salivary gland carcinosarcoma in children, Ewing's tumor family, Kaposi's sarcoma, soft tissue sarcoma, Cezari syndrome sarcoma of the uterus, skin cancer (non-melanoma), skin cancer (melanoma), squamous neck skin cancer with occult primary and metastatic gastric cancer, supratentorial primitive neuroectodermal tumors, T-cell lymphoma in children, testicular cancer, laryngeal cancer, thymoma, thymoma in children, thymic carcinoma, thyroid cancer, urethral cancer, uterine cancer, endometrial sarcoma, vaginal cancer, visual pathway and hypothalamic glioma, vulvar cancer in children, and nephroblastoma (renal carcinoma).
[0182] In some embodiments, the cancer includes ovarian cancer, prostate cancer, colon cancer, or lung cancer. In some embodiments, the method is used to determine novel biomarkers for ovarian cancer, prostate cancer, colon cancer, or lung cancer by evaluating cell-free DNA using the methods of this disclosure. In some embodiments, the methods of this disclosure can be used to isolate fetal DNA from a pregnant woman. In some embodiments, the methods of this disclosure can use fetal DNA isolated from a pregnant woman for prenatal diagnosis. In some embodiments, the methods of this disclosure can be used to evaluate fertilized embryos such as zygotes or blastocysts to determine embryo quality or the presence or absence of specific disease markers.
[0183] V. Reagent Kit
[0184] The present invention optionally provides a kit for performing the methods of the present invention. The kit may include one or more reagents described throughout this disclosure and / or one or more reagents known in the art for performing one or more steps described throughout this disclosure. For example, a kit may include one or more of the following substances: ligase buffer, ligase, T4 ligase, amplification enzyme, nuclease-free water, one or more primers, SPRI beads, cross-linking agent, polyethylene glycol, magnetic beads, DNA polymerase, Taq polymerase, dNTPs, DNA polymerase buffer, divalent cations, monovalent cations, bisulfite, sodium bisulfite, RNA polymerase, DTT, redox reagents, Mg2+, K+, adaptor, DNA adaptor, DNA including an RNA promoter, primers complementary to and / or capable of annealing (at least partially annealing) the DNA template described herein, protease, NTPs, oxidizing agents, dioxygenase, compounds containing hydroxylamine, hydrazine, or acylhydrazine groups, compounds containing amine groups, 3' end blocking molecules, reverse transcriptase, glucose or modified glucose and / or β-glucosyltransferase compounds.
[0185] The kit may include 5mC or 5hmC modifiers or reagents such as TET, GT, and modified motifs.
[0186] One or more reagents are preferably provided in a solid form or as a liquid buffer suitable for stock storage and are subsequently added to the reaction medium when the reagents are used. Suitable packaging is provided. The kit may optionally include additional components useful in the process. These optional components include buffers, capture reagents, developing agents, labels, reaction surfaces, detection methods, control samples, instructions, and explanatory information.
[0187] Each kit may also include additional components that can be used for amplifying nucleic acids, sequencing nucleic acids, or other applications described herein. The kit may optionally include additional components useful in the process. These optional components include buffers, capture reagents, developing agents, labels, reaction surfaces, detection methods, control samples, instructions, and explanatory information.
[0188] VI. Examples
[0189] The following examples are provided for the purpose of illustrating various embodiments of the invention and are not intended to limit the invention in any way. Those skilled in the art will readily understand that the invention is well-suited to achieving the mentioned endpoints and advantages, as well as those inherent therein. These examples and the methods described herein are currently representative of certain embodiments and are provided as examples, and are not intended to be limiting of the scope of the invention. Variations and other uses will be conceived by those skilled in the art within the spirit of the invention as defined by the claims.
[0190] Example 1: Linear amplification bisulfite sequencing for cell-free DNA cancer detection
[0191] It is known that cell-free DNA (cfDNA) in peripheral blood originates from cell apoptosis and necrosis. 1 In cancer patients, in addition to hematopoietic lineage cells, it also contributes to tumor cells and specific tissue cells. These tumor cell fragments flowing into the bloodstream can be traced to their cellular origin, making them a potential biomarker for minimally invasive liquid biopsies of solid tumors. 2 Furthermore, cfDNA can be repeatedly sampled and monitored, and is unaffected by intratumoral heterogeneity in tissue biopsies. These advantages, coupled with higher patient compliance and clinical convenience, make cfDNA a potential candidate for significant advancements in cancer screening, diagnosis, and prognosis. 3,4
[0192] In addition to detecting point mutations or copy number variations (CNV), 2,5,6 Epigenetic markers such as DNA methylation can function in complementary and valuable ways. DNA methylation 5mC is a chemically stable, highly abundant epigenetic modification that has long been considered to regulate gene expression. In tumorigenesis, hypermethylation of promoters leads to the silencing of tumor suppressor genes; similarly, hypomethylation of oncogenes can promote disease development. 7-9 Abnormal methylation signatures may be the most sensitive and earliest indicator of early-stage or even precancerous lesions.
[0193] Recent studies have used WGBS, RRBS, MCTA, and MIP-seq combined with algorithms such as random forest, logistic regression, and probabilistic model building to explore potential methylation markers in cfDNA. 5,10-14However, when dealing with precious, ultra-low-input clinical samples, sensitivity barriers and widespread DNA degradation can lead to drastically limited useful reads, poor data quality, and high sequencing costs. Whole-genome amplification (WGA) can amplify minute amounts of cfDNA to large scale. MDA (multiple substitution amplification), a widely used hyperbranched strand substitution exponential amplification WGA technique, can provide near-complete genome coverage, but it exhibits unavoidable bias and error accumulation in an exponential manner. Furthermore, its low efficiency when amplifying fragmented DNA makes it unsuitable for cfDNA. 17,18 Instead, the inventors labeled T7 during in vitro transcription because linear amplification not only provides uniform coverage across the entire genome but also exhibits less amplification bias than MDA. More importantly, it allows for the efficient integration of cfDNA fragments, overcoming obstacles in liquid biopsy studies. 19
[0194] To reliably obtain robust and accurate base resolution information from the rare clinical samples presented in this paper, the inventors developed a BS-seq-based T7 linear amplification method for subnacogram liquid biopsy materials (LABS-seq). LABS-seq linearly generates multiple copies of RNA transcripts from the 5'- to 3'- ends of cfDNA fragments or up to the BS-damaged nicking base, providing greater preservation of methylome information, accurate representation of the 5mC state, and more uniform whole-genome amplification. This method demonstrates improved CpG and genome coverage for application in colon cancer cfDNA marker identification and tissue origin analysis.
[0195] A. Result
[0196] 1. Design of LABS-seq strategy
[0197] To amplify T7 after bisulfite conversion, the inventors first ligated a fully methylated T7 adaptor (replacing all C with 5mC) to both ends of the DNA fragment, and then subjected the ligation product to bisulfite (BS) conversion. Figure 1A The adaptor combines the promoter sequence of T7 RNA polymerase with a short 3'-terminal closed helper sequence to form a partially double-stranded DNA structure with helper ligation. During bisulfite (BS) treatment, C is converted to U, while 5 mC is retained, thus keeping the T7 promoter sequence intact. Complementary T7 primers are then used, and the promoter region is annealed to initiate in vitro transcription. Due to the unbiased linear amplification of in vitro transcription, minute amounts of DNA fragments can be uniformly amplified into multiple RNA replications. Finally, these effective amounts of RNA products are reverse transcribed and then used for library construction for sequencing.
[0198] 2. Evaluation of LABS-seq performance
[0199] This invention evaluates LABS-seq on genomic DNA (gDNA) of E14Tg2a mouse embryonic stem cells (ESCs). mESC libraries starting with 100 ng, 10 ng, 1 ng, and 100 pg of gDNA, respectively, were sequenced using an Illumina NextSeq 500SE75, and LABS-seq was systematically compared with two other commercially available WGBS-seq methods: MethylC-seq (pre-BS) and EpiGnome (post-BS). The inventors further performed biological replication of 50 pg, 20 pg, and 10 pg mESC libraries using the same method to investigate the technical limitations of LABS-seq.
[0200] Each library yielded approximately 20 million to 40 million single-end reads. The inventors obtained the methylated CpG% of all CpG dinucleotides, averaging 42.7% of all CpGs. The mapping efficiency of the LABS-seq library was 60.3 ± 2.4%. Figure 1B Furthermore, this ratio did not decrease with decreasing input material, even down to 10 pg. More importantly, the replication levels of 1 ng and 100 pg LABS-seq libraries were 7.7 ± 0.5% and 11.3 ± 2.8%, respectively, while the corresponding levels for EpiGnome and MethylC-seq libraries were extremely high (82.8% to 98.5%). To avoid comparison bias from different sequencing depths, the inventors downsampled 15 million reads per sample for further analysis. Comparisons were made of genome coverage, CpG coverage, and chromosome coverage with other methods (…). Figure 1C Compared to Figure 7), for LABS-seq, the reduction in coverage due to decreased input slowed significantly. Furthermore, with the increase in the number of CpG sites covered by multiple RNA transcripts (called multiple reads), IVT amplification yielded more precise methylation states for each CpG, particularly for low-input materials. In the analysis of 100 pg libraries, LABS-seq obtained at least 300-fold 5X read coverage of CpGs. Figure 8 Furthermore, deeper sequencing yields more CpG, as saturation maps do not reach stable levels even in 100pg LAB-seq libraries. Figure 2A ).
[0201] Next, the inventors tested the accuracy of LABS-seq. Overall methylation levels of CHG and CHH (0.6% to 0.9%) indicated a conversion rate greater than 99% across all LABS-seq samples. The inventors used a large pool of 100 ng and 10 ng libraries from MethyC-Seq and Epigenome as a reference and compared them with libraries generated by LABS-Seq. A high correlation coefficient (Pearson r = 0.88) was observed, indicating high accuracy of the method. Figure 1E Although the coverage of MethylC-seq and Epigenome libraries decreased significantly when the input material was reduced from 100 ng to 100 pg, a similar trend was observed in LABS-seq libraries even with an input of 50 pg. Figure 1D Using a limited input library of 10 pg, the inventors were still able to identify more CpGs than in a standard 100 pg library. All these results indicate that LABS-seq can provide sensitive methylome information down to sub-nanog. The inventors further tested the reproducibility of individual CpGs. High reproducibility between replications was observed, with Pearson r ranging from 0.92 to 0.84. The slight decrease in the coefficient was due to the input DNA material being reduced from 10 ng to 50 pg.
[0202] 3. Sequencing preference analysis and uniform amplification
[0203] Sequencing bias in WGBS can arise from incomplete BS conversion, PCR amplification, and library strategies. This bias can affect the accurate estimation of methylation levels and lead to distortion of relative differences. 20 To assess potential sequencing bias, the inventors plotted normalized coverage in genome bins with varying GC contents across all three libraries. GC profiles for various biases were observed using two commercial protocols: the MethylC-seq and Epigenome libraries showed significantly increased coverage in AT-enriched and GC-enriched regions, respectively, but lacked coverage in the opposite GC-enriched and AT-enriched regions. Conversely, the LABS-seq library exhibited less GC bias and moderate coverage in extreme regions (GC < 20% and > 70%). Figure 2B This may benefit from uniform linear amplification of the entire genome with minimal base bias. The inventors also investigated the strand composition of the mapped reads to examine whether IVT involved strand bias: LABS-seq had similar percentages of negative and positive strands as the other two methods and did not show strand bias. Figure 1B ).
[0204] Furthermore, the inventors applied the Lorenz curve to evaluate the uniformity of LABS-seq. The black diagonal line represents perfect coverage uniformity, while deviation from the ideal diagonal line represents a biased read distribution. Figure 2CThe inventors compared the curves of 0.1 ng to 1 ng WGBS libraries constructed by LABS-seq with those of two other pre-BS and post-BS methods. This demonstrated that LABS-seq exhibits the best uniformity across the entire genome.
[0205] 4. Methylation characteristics of mESCs in LABS-seq libraries
[0206] The methylome features of the libraries under the three schemes were almost identical in different aspects. The inventors first plotted the methylation density on various genomic features revealed by the three methods ( Figure 3A The inventors observed a significant reduction in 3kb from upstream to the transcription start site (TSS), followed by localized depletion at methylation in the 5'UTR region, with the density recovering to similar levels (~40%) in the upstream 3kb region, exons, introns, the 3'UTR, and downstream regions. Furthermore, consistently, by examining methylation rates across different histone modification signatures, the inventors found a significant decrease in methylation at the activating histone marker H3K4me3 and an increase in methylation at the repressing marker H3K9me3. Figure 3B The inventors then investigated CpG counts across different genomic backgrounds. Figure 3C This also demonstrates superior consistency among the three methods. However, it is noteworthy that the 50 pg LABS-seq library showed similar or even better coverage of various backgrounds compared to other libraries using 20-fold gDNA material. The Log2 enrichment score further illustrates the normalized reads described above. Figure 9 The three methods showed consistent biases in the 3'UTR, exons, promoters, 5'UTR, and CpG island regions. The inventors also used transgenomic ChromHMM chromatin states to quantify the different coverage biases of the three methods. The EpiGenome method preferentially annotated states 1, 2, and 4 (active promoters, strong enhancers, and transcription / conversion, respectively), while LABS-seq preferentially annotated states 2 and 4. When the input material was reduced to 100 pg, the LABS-seq library maintained a comparable trend across different states and provided more annotative information than Epigenome and MethylC-seq.
[0207] 5. LABS-seq was used to perform whole-genome mapping of 5mC in cfDNA.
[0208] Because LABS-seq performs exceptionally well with constructed libraries from limited inputs, the inventors tested LABS-seq on mixed cancer-free plasma cfDNA. The above compares the Epignome and MethylC-seq protocols (…). Figure 10LABS-seq demonstrated overwhelming advantages in genome coverage, CpG coverage, and replication levels. The inventors then applied the method to patient cfDNA samples. Fifteen patients were involved, including six healthy patients, six colon cancer patients, and three pancreatic cancer patients.
[0209] 6. Genome-wide hypomethylation / hypermethylation of cancer cfDNA
[0210] Because overall methylation changes are found in many cancer types, genome-wide hypo / hypermethylation was investigated. The inventors examined the methylation density (MD) in each 1-Mb bin (a total of 2734 bins, excluding sex chromosomes) of the whole genome. All healthy individuals showed relatively stable levels within each bin, while CRC and PAN showed high variability between their respective bins. In CRC and PAN cfDNA, the number of bins with MD > 3SD differing from the mean in healthy cfDNA were 177 ± 83.6 and 108 ± 19.2, respectively. The inventors determined the percentage of patients with highly variable cfDNA features ( Figure 4A The annotation of these highly variable features revealed that they were mostly located in intergenic and intronic regions, while also showing a high enrichment in promoters. Functional enrichment analysis found that paraxial mesenchymal cells were highly enriched in the CRC-derived hypo / hypermethylation box (FDR = 5.36e-6), consistent with the important role of mesenchymal cells in colorectal cancer. With more samples collected, the inventors could effectively differentiate cancer patients based on genome-wide hypo / hypermethylation using relatively low sequencing depth.
[0211] 7. DMR as a taxonomic biomarker
[0212] To examine whether LABS-Seq data could lead to the discovery of differentially methylated regions (DMRs), the inventors used 1-kb bins to compare different patient groups and found 109 significantly differentially methylated regions between CRC and healthy individuals. Figure 5A These DMRs are distributed throughout the entire genome. Figure 5B (Excluding sex chromosomes and mitochondria). Between them, 47 DMRs in the CRC group were hypomethylated, and 62 DMRs were hypermethylated. The inventors also observed that CRC-specific hypomethylated DMRs were particularly enriched in repetitive regions such as LINE and SINE, while CRC-specific hypermethylated DMRs were enriched in promoter and exon regions associated with gene expression regulation, indicating significant differences in 5mC profiles between healthy individuals and CRC patients, which can reveal cancer-related biological functions. All these results suggest that this method can provide robust 5mC genomic analysis and reveal plausible 5mC differences between patient groups.
[0213] To further evaluate the classification performance of these DMRs, the inventors used hierarchical clustering to aggregate features from different groups. Figure 5C Through unsupervised clustering, two patient groups can be directly and completely separated in both sensitive (6 / 6 CRC) and specific (6 / 6 healthy) ways. The inventors also performed PCA analysis on DMR features and noted that different patients can be easily grouped based on disease classification. Figure 5D In the DMR between CRC and healthy groups, the inventors obtained several potential CRC marker genes, such as CCR9 (chemokine receptor 9, associated with CRC invasion and metastasis), BRD4 (a domain-containing protein 4, which is frequently aberrantly methylated in colorectal cancer), LGAL9 (a tumor suppressor gene that is highly methylated in CRC patients), and IRF1 (a tumor suppressor known in a variety of cancers and capable of inducing tumor cell apoptosis).
[0214] CEA (carcinoembryonic antigen) 21 is a clinically useful blood biomarker for CRC (reference range 0 ng / mL to 3.4 ng / mL). Serum CEA test results for patients CRC5, CRC4, and CRC6 were 1.7 ng / mL, 4.5 ng / mL, and 88.2 ng / mL, respectively. In DMR clustering and PCA, all three patient groups were naturally classified as cancer. This demonstrates that the DMR biomarker discovered by the LABS-seq method can be a more sensitive indicator than CEA or other protein serum biomarkers. CRC1 is a patient with SSA / polyps, a high-risk precancerous lesion, and CRC1 was grouped into the CRC group. This demonstrates the potential for using 5mCcfDNA in early screening for high-risk precancerous lesions. If there are more precancerous / cancer patients in the process, the inventors can obtain 5mC biomarkers for precancerous diseases and cancers separately.
[0215] 8. Organizational source deconvolution
[0216] cfDNA can be considered a mixture of DNA derived from different tissues and a large number of hematopoietic cells. Although different tissue types actually have the same DNA sequence, their 5mC methylation profiles are highly specific. 14 Using the 5mC feature, which exhibits tissue-specific characteristics that are negligible or present in small proportions in other tissues, the inventors can obtain the mixing proportions of different tissues through deconvolution algorithms, trace tissue origins, and have the potential to predict specific cancers. 10,13,22
[0217] Using this method, the inventors reanalyzed public WGBS data from 14 organizations generated by the Roadmap project (found on the World Wide Web at genboree.org / epigenomeatlas / index.rhtml) and two other studies. 23,24By comparing one tissue with the rest, a total of 5823 significant DMRs were identified. These DMRs were then used as markers for deconvolution. As a proof of concept, the inventors found that neutrophils constituted the most significant portion of almost all cfDNA samples. Figure 6A (45% to 68% of those in good health; 20% to 62% of those in CRC), consistent with previous studies. 22 The inventors noted a significant decrease in the percentage of neutrophils in the CRC group, while the contribution from the colon increased (Figure 6C). The proportion of colon tissue in cfDNA in the four CRC patients (4 / 6) ranged from 3.1% to 25%, but the contribution was undetectable in the healthy group. This result supports the concept of using DNA methylation signatures for cancer type prediction. By combining genome-wide hyper / hypomethylation detection with tissue-derived deconvolution, the inventors were able to detect not only cancer but also predict the location of solid tumors.
[0218] B. Discussion
[0219] LABS-seq is the first liquid biopsy WGBS-seq method based on whole-genome amplification. Using LABS-seq technology, the inventors conducted a comprehensive study of 5mC signals in a colorectal cancer patient. Compared to healthy individuals, this patient showed significant DNA methylation abnormalities in plasma cfDNA. Although only 6 / 6 of CRC / healthy individuals participated in unsupervised clustering and biomarker identification for DMR, the inventors obtained more refined biomarkers such as CCR9 and BRD4, as well as fully confident clustering. In CRC, the inventors found that precancerous samples were grouped into the CRC group. This suggests that cfDNA DNA methylation may be useful in early screening, especially in precancerous lesions that are not large enough to be detected or are indeterminate in imaging. Since cfDNA is a balance of DNA fragments from different sources, it is able to provide a picture of the whole DNA methylation in the human body. Most importantly, different types of cancer show highly different methylation signatures, thus enabling tissue origin prediction and enhancing solid tumor detection based on methylation signatures. In future work, to make full use of LABS-seq, longitudinal monitoring can be used to assist in tumor treatment, targeting the unique cfDNA characteristics of primary solid tumors and metastatic lesions.
[0220] In addition to blood, this technology can also be applied to cfDNA in, but is not limited to, urine, cerebrospinal fluid (CSF), and aqueous humor, where DNA fragments have reportedly been smaller in size or at lower concentrations than those in blood. LAB-seq can also be used for single-cell studies. After fragmenting single-cell gDNA using transposases or nucleases, this technology can be used to amplify materials from micrograms to nanograms. It provides researchers with a powerful and accurate tool to utilize in-depth insights into cancer and embryonic development, observing epigenetic changes in circulating tumor cells, cellular heterogeneity, and the entire genome in embryos.
[0221] C. Method
[0222] 1. Patient recruitment and plasma samples
[0223] The study recruited 6 patients with colon cancer, 3 patients with pancreatic cancer, and 6 healthy individuals.
[0224] Blood samples were collected from patients with colon cancer before resection of the malignant tumor. The diagnosis was confirmed by surgical pathology after lesion resection. Blood samples were also collected from patients with pancreatic cancer. Healthy individuals were defined as those who underwent routine health checks and had no malignant or premalignant disease. Blood samples were collected from healthy individuals during surgery. Blood samples were collected in streck tubes and centrifuged twice at 1350g for 12 minutes at 4°C, followed by centrifugation twice at 1350g for 5 minutes at 4°C. cfDNA was isolated from plasma from 0.1 mL to 1 mL using a QIAamp Circulating Nucleic Acid Ki kit (Qiagen, 55114).
[0225] 2. Cell culture and genomic DNA isolation
[0226] mESC cells were grown on gelatin-coated plates in Dulbecco-modified Eagle medium (DMEM) (Invitrogen Cat, No. 11995), supplemented with 15% FBS (Gibco), 2 mM L-glutamine (Gibco), 1X non-essential amino acids (Gibco), 1% penicillin / streptavidin (Gibco), 1X β-mercaptoethanol (Sigma), 1000 u / mL leukemia inhibitory factor (Millipore Cat, No. ESG1107), 1 μM PD0325901 (Stemgent, dissolved in DMSO), and 3 μM CHIR99021 (Stemgent, dissolved in DMSO). All cells were cultured at 37°C and 5.0% CO2, and passaged every 2 days.
[0227] For genomic DNA isolation, cells were harvested by centrifugation at 1000x g for 3 minutes. DNA was then extracted from Qiagen using the AllPrep DNA / RNA Mini Kit, according to the protocol.
[0228] 3. Preparation of T7 linker
[0229] Oligonucleotides were purchased from IDT and purified by HPLC. The T7 sequence / 5Phos / iMe-dC / iMe-dC / iMe-dC / TATAGTGAGT / iMe-dC / GTATTAATTT / iMe-dC / G / iMe-dC / GGGG / iMe-dC / T (SEQ ID NO:4), where iMe-dC refers to the internal 5-methyldeoxycytidine and the short auxiliary sequence CGACTCACTATAGGGT / 3Phos / (SEQ ID NO:5) was dissolved in annealing buffer (10 mM Tris-HCl pH 8.0, 0.1 mM EDTA, 50 mM NaCl). The T7 adaptor was prepared by uniformly mixing the two oligonucleotides at 50 μM and annealing using a PCR instrument (95 °C for 5 min, cooling to 4 °C at -0.25 °C / min). The adaptor was diluted to 15 μM by annealing in buffer and stored at -20 °C.
[0230] 4. Bisulfite Conversion
[0231] Bisulfite treatment was performed using the MethylCode™ Bisulfite Conversion Kit (Invitrogen, MECOV50). In short, 20 μL of gDNA or cfDNA was mixed with 130 μL of CT conversion buffer and incubated at 98°C for 10 min, then at 64°C for 2.5 h, and finally held at 4°C. The converted DNA was purified using a spin column with on-column desulfurization provided in the kit, and eluted in 9 μL to 20 μL of preheated (55°C) nuclease-free water (Ambion, AM9937).
[0232] 5. Methyl C-seq whole genome bisulfite sequencing
[0233] The mESC gDNA is fragmented into 150bp to 400bp dsDNA fragments. This is done according to the manufacturer's protocol. The sequencing library was prepared using the Bisulfite Sequencing Kit (PerkinElmer). In short, after end repair and 3'-polyadenylation, methylated adaptors were attached to both ends of the DNA fragment. Then, the DNA was subjected to bisulfite conversion. Finally, the library was amplified using KAPA Hifi Uracil Plus Polymerase (Kapa Biosystems) and purified twice with 0.8X AMPure XP beads. The library was then subjected to NextSeq 500SR80.
[0234] 6. Epignom whole genome bisulfite sequencing
[0235] The TruSeq DNA Methylation Kit was used. First, GDNA or cfDN was transformed with bisulfite. The synthesized random primers were annealed to the transformed ssDNA following the manufacturer's protocol. DNA strands containing specific sequence markers from the random primers were synthesized. Then, the known sequence was added to the 3' end of the DNA strand. The double-labeled DNA was purified using 1.6X AMPureXP beads. The library was amplified using the Failsafe PCR enzyme system and purified with 1.0X AMPureXP beads. The library was sequenced using a NextSeq 500SR80.
[0236] 7. Whole-genome LABS-Seq
[0237] The cfDNA or fragmented gDNA was ligated to the T7 adaptor using the KAPA Hyper Prep Kit (KAPA BIOSYSTEMS, KK8502). In short, 10 μL of nuclease-free H2O containing DNA was mixed with 1.4 μL of End Repair & A-Tailling Buffer and 0.6 μL of End Repair & A-Tailling Enzyme Mix, and incubated at 20°C for 30 minutes, then at 65°C for 30 minutes. Next, 1 μL of domestically produced T7 adaptor, 1 μL of H2O, 6 μL of ligation buffer, and 2 μL of DNA ligase were added, and the mixture was incubated at 20°C for 4 hours or overnight at 4°C. After cleaning following ligation using 1.4X AMPure XP beads, the DNA was treated with bisulfite. DNA converted to bisulfite in 10 μL of H2O was mixed with 3 μL of 5X EpiMark Buffer, 0.3 μL of 10 μM T7 primer (AGCCCCGCGAAATTAATACGACTCACTATAGGG (SEQ ID NO:3), purified by HPLC using IDT), 0.3 μL of 10 mM dNTP (NEB, N0447S), 0.3 μL of EpiMark Hot Start Taq DNA polymerase (NEB, M0490S), and 1.1 μL of nuclease-free H2O. T7 primer annealing and extension were performed at 95 °C for 60 seconds, 59 °C for 60 seconds, 68 °C for 5 minutes, and held at 4 °C. Then, 1 μL of 0.85 mg / mL QIAGEN protease was added and incubated at 50 °C for 2 hours, followed by heat inactivation at 75 °C for 30 minutes. Dilute the T7-labeled dsDNA fragment with 22 μL of H2O and react it with 60 μL of T7 reaction premix (NEBHiScribe). TM The following ingredients were used in the T7 High-Yield RNA Synthesis Kit E2040S: 1X T7 reaction buffer, 10 μL RNA polymerase mixture, 10 mM ATP / GTP / UTP / CTP, and 0.4 U / μL SUPERase In RNase inhibitor (Life Technologies, AM2694). The reaction was carried out at 37°C for 12 to 16 hours.
[0238] After overnight in vitro transcription at T7, DNase I and digestion buffer were added to the abstract DNA template at room temperature for 20 minutes. The RNA transcripts were then purified using the RNAClean & Concentrator kit (Zymo Research, R1013) and eluted in 15 μL H2O. DNA yield was quantified using the Qubit 2.0 RNA HS Assay Kit (Life Technologies, Q32855). The majority of the 100 ng RNA was used for library construction (KAPA BIOSYSTEMS, KAPA RNA HyperPrep kit, KK8540). Following the manufacturer's protocol, RNA in 10 μL H2O was mixed with 10 μL of 2X fragment, primers, and elution buffer. The reaction was heated at 65 °C for 1 minute and terminated on ice. First-strand synthesis was initiated by adding 10 μL of the reaction mixture (buffer and KAPA transcriptase) and incubated at 25 °C for 10 minutes, 42 °C for 15 minutes, 70 °C for 15 minutes, and held at 4 °C. Add 30 μL of second-strand synthesis and A-tail to the reaction mixture and incubate at 16 °C for 30 min, 62 °C for 10 min, and hold at 4 °C. Then, ligate the Illumina adaptor at 20 °C for 15 min. After two rounds of ligation cleanup using a KAPAPure Bead, elute and PCR amplify (98 °C for 30 s; for the following 10 to 12 cycles: 98 °C for 15 s, 60 °C for 30 s, 72 °C for 45 s; 72 °C for 1 min), and purify using a 1X Ampure XP Bead. Finally, sequence the library using a NextSeq 500SR80.
[0239] 8. Data processing and analysis
[0240] The raw sequencing reads were first trimmed using Trim Galore to remove sequencing adaptors and low-quality nucleotides. Bismark was then used to map the trimmed reads to mm9 and hg19 reference genomes, respectively. Further duplication removal and mC calls were performed using packaging scripts within the Bismark package. Cytosine residues in the CpG background were further used for downstream analysis. Samtool and Bedtool were used for interval correlation calculations. CG bias was quantified using the Picard package. Metagenome profiling was calculated using Deeptool.
[0241] 9. Whole-genome hypomethylation / hypermethylation of cfDNA
[0242] To quantify overall DNA hypomethylation and hypermethylation in patient cfDNA samples, the inventors divided the entire genome into contiguous 1-Mb bins and calculated the methylation density for each bin. Methylation density was defined as the number of methylated cytosines divided by the total number of cytosines in the CpG background within each bin. The mean and variation of methylation density in healthy samples were calculated, and a z-score was used to determine whether certain samples were normal. Any bin with a z-score greater than 3 or less than -3 was considered to be either hypermethylated or hypomethylated and used as a potential feature of cancer. The RCircos package was used for further genome-wide visualization.
[0243] 10. DMR Analysis
[0244] Feature selection of differentially methylated regions (DMRs) was performed by analyzing unsupervised clustering downstream of the MethylKi package. Based on logistic regression and SLIM adjustment, DMRs were defined as 1-kb bins where the methylation difference between groups was greater than 0.4 and the q-value was less than 0.001. Further clustering and PCA analyses were visualized using functions within the MethylKi package. Functional annotation was performed using Homer software, and functional enrichment analysis was conducted on all intervals using the GREAT online tool.
[0245] 11. Plasma DNA Tissue Mapping
[0246] The inventors used a quadratic programming algorithm to deconvolve cfDNA methylation profiles. WGBS data from 14 tissues in the Roadmap project and two other studies used to discover tissue-specific DMRs were used as reference features. The methylation density of each feature was then calculated in both the reference tissues and patient samples. Deconvolution was then performed using R-code.
[0247] ***
[0248] According to this disclosure, all methods disclosed and claimed herein can be performed and carried out without excessive experimentation. Although the compositions and methods of the invention have been described according to preferred embodiments, it will be apparent to those skilled in the art that variations can be made to the methods and the steps or order of steps in the methods described herein without departing from the concept, spirit, and scope of the invention. More specifically, it will be apparent that specific chemically and physiologically relevant reagents can be substituted for those described herein when achieving the same or similar results. All such similar substitutions and modifications that will be apparent to those skilled in the art are considered to be within the spirit, scope, and concept of the invention as defined by the appended claims.
[0249] References
[0250] The following references provide exemplary procedures or other supplementary details to the details set forth herein, and are expressly incorporated herein by reference.
[0251] 1.Schwarzenbach,H.,Hoon,D.S.B.&Pantel,K.Cell-free nucleic acids asbiomarkers in cancer patients.Nature Reviews Cancer 11,426(2011).
[0252] 2.Bettegowda.C.et al.Detection ofcirculating tumor DNA in early-andlate-stage human malignancies.Science translational medicine 6,224ra224-224ra224(2014).
[0253] 3.Alix-Panabières,C.&Pantel,K.Clinical applications of circulatingtumor cells and circulating tumor DNA as liquid biopsy.Cancer discovery 6,479-491(2016).
[0254] 4.Murtaza,M.et al.Non-invasive analysis of acquired resistance tocancer therapy by sequencing of plasma DNA.Nature 497,108(2013).
[0255] 5.Chan,K.A.et al.Noninvasive detection of cancer-associated genome-wide hypomethylation and copy number aberrations by plasma DNA bisulfitesequencing.Proceedings of the National Academy of Sciences 110,18761-18768(2013).
[0256] 6.Shaw,J.A.et al.Genomic analysis of circulating cell-free DNA infersbreast cancer dormancy.Genome research 22,220-231(2012).
[0257] 7.Klose,R.J.&Bird,A.P.Genomic DNA methylation:the mark and itsmediators.Trends in Biochemical Sciences 31,89-97(2006).
[0258] 8.Esteller,M.Epigenetics in cancer.New England Journal of Medicine358,1148-1159(2008).
[0259] 9.Almouzni,G.&Cedar,H.Maintenance of epigenetic information.ColdSpring Harbor perspectives in biology 8,a019372(2016).
[0260] 10.Guo,S.et al.Identification of methylation haplotype blocks aids indeconvolution of heterogeneous tissue samples and tumor tissue-of-originmapping from plasma DNA.Nature genetics 49,635(2017).
[0261] 11.Wen,L.et al.Genome-scale detection of hypermethylated CpG islandsin circulating cell-free DNA of hepatocellular carcinoma patients.Cellresearch 25,1250(2015).
[0262] 12.Xu,R.-h.et al.Circulating tumour DNA methylation markers fordiagnosis and prognosis of hepatocellular carcinoma.Nature Materials 16,1155(2017).
[0263] 13.Kang,S.et al.CancerLocator:non-invasive cancer diagnosis andtissue-of-origin prediction using methylation profiles of cell-freeDNA.Genome biology 18,53(2017).
[0264] 14.Feng,H.,Jin,P.&Wu,H.Disease prediction by cell-free DNAmethylation.Briefings in bioinformatics.
[0265] 15.Dean,F.B.et al.Comprehensive human genome amplification usingmultiple displacement amplification.Proceedings of the National Academy ofSciences 99,5261-5266(2002).
[0266] 16.Spits,C.et al.Whole-genome multiple displacement amplificationfrom single cells.Nature protocols 1,1965(2006).
[0267] 17.Lage,J.M.et al.Whole genome analysis of genetic alterations insmall DNA samples using hyperbranched strand displacement amplification andarray-CGH.Genome research 13,294-307(2003).
[0268] 18.Gawad.C.,Koh,W.&Quake,S.R.Single-cell genome sequencing:currentstate of the science.Nature Reviews Genetics 17,175(2016).
[0269] 19.Chen,C.et al.Single-cell whole-genome analyses by LinearAmplification via Transposon Insertion(LIANTI).Science 356,189-194(2017).
[0270] 20.Olova,N.et al.Comparison of whole-genome bisulfite sequencinglibrary preparation strategies identifies sources of biases affecting DNAmethylation data.Genome biology 19,33(2018).
[0271] 21.Fakih,M.G.&Padmanabhan,A.CEA monitoring in colorectalcancer.ONCOLOGY-WILLISTON PARK THEN HUNTINGTON THE MELVILLE NEW YORK-20,579(2006).
[0272] 22.Sun,K.et al.Plasma DNA tissue mapping by genome-wide methylationsequencing for noninvasive prenatal,cancer,and transplantationassessments.Proceedings of the National Academy of Sciences 112,E5503-E5512(2015).
[0273] 23.Lun.F.M.et al.Noninvasive prenatal methylomic analysis bygenomewide bisulfite sequencing of maternal plasma DNA.Clinical chemistry 59,1583-1594(2013).
[0274] 24.Hodges,E.et al.Directional DNA methylation changes and complexintermediate states accompany lineage specificity in the adult hematopoieticcompartment.Molecular cell 44,17-28(2011).
[0275] 25.Smallwood,S.A.et al.Single-cell genome-wide bisulfite sequencingfor assessing epigenetic heterogeneity.Nature methods 11,817-820(2014).
Claims
1. A method for amplifying deoxyribonucleic acid (DNA) molecules treated with bisulfite, the method comprising: (a) Connecting an adaptor to a DNA molecule, wherein the adaptor comprises an RNA polymerase promoter containing 5-methylcytosine (5mC), wherein all cytosines in the RNA polymerase promoter are 5mC; (b) Treating the linked DNA molecules with bisulfite; (c) Hybridize the bisulfite-treated DNA molecules with primers; (d) Extend the hybridized primers to produce double-stranded DNA; and (e) In vitro transcription of double-stranded DNA to produce RNA.
2. A non-diagnostic method for identifying 5-hydroxymethylcytosine (5hmC) in a DNA molecule, the method comprising: (1) Modify the 5hmC in the DNA molecule to protect it from oxidation; (2) The modified DNA molecule from (1) was oxidized with methylcytosine dioxygenase to convert 5-methylcytosine (5mC) to 5-carboxycytosine (5caC); (3) Perform the following steps: (a) Connecting an adaptor to a DNA molecule, wherein the adaptor includes an RNA polymerase promoter containing 5 mC, wherein all cytosines in the RNA polymerase promoter are 5 mC; (b) Treating the linked DNA molecules with bisulfite; (c) Hybridize the bisulfite-treated DNA molecules with primers; (d) Extend the hybridized primers to produce double-stranded DNA; (e) In vitro transcription of double-stranded DNA to produce RNA; and (f) Identify 5hmC in DNA molecules.
3. A non-diagnostic method for identifying 5mC in a DNA molecule, the method comprising: (1) Oxidize DNA molecules with an oxidizing agent to oxidize 5hmC to 5-formylcytosine (5fC) or 5caC; (2) The method includes the following steps: (a) Connecting an adaptor to a DNA molecule, wherein the adaptor includes an RNA polymerase promoter containing 5 mC, wherein all cytosines in the RNA polymerase promoter are 5 mC; (b) Treating the linked DNA molecules with bisulfite; (c) Hybridize the bisulfite-treated DNA molecules with primers; (d) Extend the hybridized primers to produce double-stranded DNA; (e) In vitro transcription of double-stranded DNA to produce RNA; and (f) Identify 5hmC in DNA molecules.
4. The method according to any one of claims 1 to 3, wherein 5fC is modified with a compound comprising a hydroxylamine group, a hydrazine group, or an acylhydrazine group.
5. The method according to claim 4, wherein the compound is hydroxylamine; hydroxylamine hydrochloride; hydroxylamine sulfate; hydroxylamine phosphate; O-methylhydroxylamine; O-hexylhydroxylamine; O-pentylhydroxylamine; O-benzylhydroxylamine; O-ethylhydroxylamine (EtONH2), O-alkylated or O-arylated hydroxylamine, its acid or salt.
6. The method according to any one of claims 1 to 3, wherein 5caC is amide-modified 5caC.
7. The method of claim 6, wherein 5caC is amide modified by linking to a compound containing an amino group.
8. The method of claim 7, wherein 5caC is linked to the amino group by incubating the DNA molecule together with the carbodiimide derivative.
9. The method of claim 7, wherein the compound comprising an amino group is benzylamine, substituted benzylamine, alkylamine, alkyldiamine, xylylamine, substituted xylylamine, cycloalkylamine, cycloalkyldiamine, hydroxylamine, or substituted hydroxylamine.
10. The method according to any one of claims 1 to 3, wherein the RNA polymerase promoter comprises an SP6, T7, or T3 promoter.
11. The method according to any one of claims 1 to 3, wherein (a) comprises DNA end modification and / or DNA end repair.
12. The method of claim 11, wherein the end modification comprises adding an A to the end.
13. The method according to any one of claims 1 to 3, wherein (a) comprises contacting the DNA with a ligase under conditions sufficient to link the adaptor to the DNA molecule.
14. The method according to any one of claims 1 to 3, wherein the adapter further comprises a 3' end-closed molecule.
15. The method according to any one of claims 1 to 3, wherein the connector is partially double-stranded.
16. The method according to any one of claims 1 to 3, wherein (d) comprises contacting the DNA molecule with DNA polymerase.
17. The method according to any one of claims 1 to 3, wherein (d) comprises incubating the DNA under denaturing conditions that allow double-stranded DNA to denature into single-stranded DNA.
18. The method of claim 17, wherein (d) comprises incubating the DNA under conditions sufficient to anneal the primers to single-stranded DNA.
19. The method of claim 18, wherein (d) comprises incubating the DNA under conditions sufficient to extend the primers to produce double-stranded DNA.
20. The method according to any one of claims 1 to 3, wherein (e) comprises contacting the DNA molecule with RNA polymerase and nucleoside triphosphate.
21. The method according to any one of claims 1 to 3, wherein the method further comprises one or more purification steps.
22. The method of claim 21, wherein the purification step comprises solid-phase reversible immobilization (SPRI) beads.
23. The method according to any one of claims 1 to 3, wherein the method further comprises (e) the purification of the RNA molecule.
24. The method according to any one of claims 1 to 3, wherein the method further comprises reverse transcription of the RNA molecule of (e) to produce the corresponding DNA molecule.
25. The method of claim 1, wherein (a) through (e) are performed in sequence.
26. The method of claim 23, wherein the method further comprises library construction of the corresponding DNA molecule.
27. The method of claim 23, wherein the method further comprises sequencing the corresponding DNA molecule.
28. The method according to any one of claims 1 to 3, wherein the DNA molecule is fragmented.
29. The method of claim 28, wherein the length of the segment is from 100 bp to 300 bp.
30. The method of claim 28, wherein the DNA molecule is a biological fragment.
31. The method according to any one of claims 1 to 3, wherein the method does not include an enrichment step.
32. The method according to any one of claims 1 to 3, wherein the DNA molecule comprises free DNA (cfDNA).
33. The method according to any one of claims 1 to 3, wherein the DNA molecule comprises genomic DNA.
34. The method according to any one of claims 1 to 3, wherein the amount of said DNA molecule is from 1 pg to 100 ng.
35. The method according to any one of claims 1 to 3, wherein the DNA molecule is isolated from a sample of the object.
36. The method according to any one of claims 1 to 3, wherein the DNA molecule is isolated from the biopsy sample.
37. The method of claim 36, wherein the sample is a liquid sample.
38. The method according to claim 2, wherein the methylcytosine dioxygenase is TET1, TET2 or TET3, or a homologue thereof.
39. The method of claim 2, wherein 5hmC is modified with glucose or modified glucose.
40. The method of claim 39, wherein 5hmC is modified by means of incubating nucleic acid molecules with β-glucosyltransferase and glucose or modified glucose molecules.
41. The method of claim 40, wherein the glucose molecule is a uridine diphosphate glucose molecule.
42. The method according to claim 41, wherein the modified glucose molecule is a modified uridine diphosphate glucose molecule.
43. The method of claim 3, wherein the oxidation selectively oxidizes 5hmC residues.
44. The method according to claim 3, wherein the oxidant is a chemical oxidant.
45. The method of claim 3, wherein the oxidant is a perruthenate oxidant.
46. The method of claim 3, wherein the oxidant comprises KRuO4.
47. A non-diagnostic method for identifying 5-methylcytosine (5mC) in a DNA molecule, the method comprising: (a) Connecting an adaptor to a DNA molecule, wherein the adaptor includes an RNA polymerase promoter containing 5 mC, wherein all cytosines in the RNA polymerase promoter are 5 mC; (b) Treating the linked DNA molecules with bisulfite; (c) Hybridize the bisulfite-treated DNA molecules with primers; (d) Extend the hybridized primers to produce double-stranded DNA; and (e) In vitro transcription of double-stranded DNA to produce RNA; (f) Reverse transcription of RNA to produce DNA; (g) Sequencing the DNA and identifying the 5mC in the sequenced DNA as "C" in the sequence.
48. A method for identifying methylcytosine in a sample containing DNA molecules, the method comprising performing the method of any one of claims 1 to 3.
49. Use of an adapter comprising an RNA polymerase promoter in the preparation of a kit for identifying 5mC in a DNA molecule derived from a patient and / or for assisting in patient diagnosis or prognosis, wherein the RNA polymerase promoter comprises 5mC, wherein all cytosines in the RNA polymerase promoter are 5mC, wherein the method as defined in any one of claims 1 to 3 is performed, wherein the DNA molecule is provided from a biological sample of the patient.
50. A non-diagnostic method for evaluating single cells, the method comprising performing the method of any one of claims 1 to 3, wherein DNA molecules are provided from the genomic DNA of the single cell.
51. The method according to any one of claims 1 to 3, wherein the RNA polymerase promoter is a prokaryotic RNA polymerase promoter.
Citation Information
Patent Citations
Methods for Detection of Nucleotide Modification
US20140178881A1
Large scale photolithographic solid phase synthesis of polypeptides and receptor binding screening thereof
US5143854A
Very large scale immobilized polymer synthesis using mechanically directed flow paths
US5384261A
Very large scale immobilized polymer synthesis
US5424186A
Array of oligonucleotides on a solid substrate
US5445934A