Method for sequential preparation of nucleic acid libraries
The method addresses the challenge of invasive and error-prone nucleic acid detection by using sequential amplification and targeted enrichment, improving the accuracy and sensitivity of detecting genetic and epigenetic alterations in nucleic acids for early cancer detection.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NATERA INC
- Filing Date
- 2025-10-15
- Publication Date
- 2026-04-23
AI Technical Summary
Current methods for detecting cancer, particularly early-stage cancer or cancer relapse, are invasive and prone to errors due to chemical or enzymatic treatments of nucleic acid molecules, leading to inaccurate detection of genetic variants or mutations, and there is a need for improved methods to analyze small amounts of nucleic acids like cell-free DNA or RNA for early detection and error reduction.
A method involving sequential amplification and targeted enrichment of nucleic acids, followed by sequencing, using capture moieties like biotin or click chemistry to enrich specific loci, allowing for error correction and improved sensitivity in detecting genetic or epigenetic alterations.
The method enhances the accuracy and sensitivity of detecting genetic variants and epigenetic changes associated with cancer, enabling early detection and minimizing false negatives by reusing original nucleic acid molecules for multiple library preparations.
Smart Images

Figure IMGF000080_0001 
Figure IMGF000080_0002 
Figure IMGF000083_0001
Abstract
Description
Attorney Docket No. N.057.W0.01METHOD FOR SEQUENTIAL PREPARATION OF NUCLEIC ACID LIBRARIESCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefits of U.S. Provisional Application No. 63 / 708,153 filed October 16, 2024, the contents of which are hereby incorporated by reference in their entirety.TECHNICAL FIELD
[0002] The present disclosure relates generally to methods for preparing and analyzing nucleic acid molecules, and, more particularly, to methods for preparing and analyzing nucleic acid molecules that include one or more genetic variants or mutations or epigenetic changes. Such methods may be used alone or in combination with other methods for preparing and analyzing nucleic acid molecules or other analytes or biomarkers indicative of cancer or other disease or biological state.BACKGROUND
[0003] Detection of cancer, cancer relapse, or cancer metastasis has traditionally relied on imaging and tissue biopsy. The biopsy of tumor tissue is invasive and carries risk of potentially contributing to metastasis or surgical complications, while imaging-based detection is not sufficiently sensitive. Moreover, neither method is suited for detecting asymptomatic cancer, or detecting relapse or metastasis at an early stage. Improved and less invasive methods are needed for early detection and for detecting relapse or metastasis of cancers. There is a need for improved methods for preparing, quantifying, and otherwise evaluating small amounts of nucleic acid molecules in samples, such as cell-free DNA (cfDNA) or cell-free RNA (cfRNA), carrying alterations such as genetic variants or mutations or epigenetic changes (e.g., SNVs, SNPs, indels, and changes in methylation patterns or status), such as circulating tumor DNA (ctDNA), circular RNA, miRNA, especially where it is desirable to perform multiple assays (e.g. replicate assays or assays involving different signals) on the small amounts of nucleic acid molecules. Chemical or enzymatic treatment of the nucleic acid molecules in certain assays (e.g. chemical or enzymatic conversion often used for methylation status sensing) not only can damage the nucleic acid molecules but is also prone to errors caused by, for example, incomplete or nonspecific reactions14904-2842-5073.2Attorney Docket No. N.057.W0.01(such as conversion, digestion, or protection errors). Furthermore. PCR and sequencing artifacts are seen in methods of detecting alterations when a large number of targets are analyzed. Accordingly, there is a need for error and noise reduction for accurate cancer or other disease or biological state detection. Methods described herein address these needs.SUMMARY
[0004] One aspect of the invention described herein relates to a method for amplifying and sequencing nucleic acids, comprising: (a) extracting nucleic acids from a sample of a subject to obtain extracted nucleic acids; (b) appending an adaptor having a capture moiety to the extracted nucleic acids to obtain adapted nucleic acids; (c) performing a first amplification on the adapted nucleic acids and generating first amplified DNA, and separating the adapted nucleic acids from the first amplified DNA; (d) performing a second amplification on the adapted nucleic acids and generating second amplified DNA, and separating the adapted nucleic acids from the second amplified DNA; (e) performing a first targeted enrichment on the first amplified DNA or its derivative to enrich a plurality of target loci each encompassing at least one alteration associated with a phenotype and generating first enriched DNA; (f) performing a second targeted enrichment on the second amplified DNA or its derivative to enrich the plurality of target loci and generating second enriched DNA; and (g) sequencing the first enriched DNA and the second enriched DNA or their respective derivative and producing a first set of sequence reads and a second set of sequence reads, respectively.
[0005] In some embodiments, the method further comprises detecting at least one alteration associated with the phenotype present in both the first enriched DNA and the second enriched DNA based on the first set of sequence reads and the second set of sequence reads.
[0006] In some embodiments, the extracted nucleic acids comprise DNA or RNA. In some embodiments, the nucleic acids comprise total or cell-free DNA. In some embodiments, the nucleic acids comprise total or cell-free RNA.
[0007] In some embodiments, the capture moiety comprises biotin, biotin dT, biotin-TEG, photocleavable (PC) biotin, or desthiobiotin-TEG, or biotin azide, wherein the separating is performed using streptavidin beads. In some embodiments, the capture moiety comprises a click chemistry capture moiety. In some embodiments, the capture moiety comprises a thiol group, and the separating is performed using gold beads. In some embodiments, the capture moiety comprises 24904-2842-5073.2Attorney Docket No. N.057.W0.01 digoxigenin and the separating is performed using anti-digoxigenin antibody-modified beads. In some embodiments, the capture moiety comprises acrydite, and the separating is performed using thiol-modified beads.
[0008] In some embodiments, the method further comprises (dl) performing a third amplification on the adapted nucleic acids and generating third amplified DNA, and separating the adapted nucleic acids from the third amplified DNA, and optionally repeating step (dl) to generate additional amplified DNA.
[0009] In some embodiments, the adaptor comprises a universal priming sequence, wherein the adapted nucleic acids are amplified using a primer that binds to the universal priming sequence.
[0010] In some embodiments, an error in amplification, enrichment, or sequencing is identified from the first set of sequence reads in combination with the second sets of sequence reads without using any molecular barcode.
[0011] In some embodiments, the method does not comprise tagging the extracted nucleic acids with a molecular barcode.
[0012] In some embodiments, the plurality of target loci comprises 25-20,000 target loci.
[0013] In some embodiments, performing targeted enrichment comprises performing targeted probe capture to enrich the target loci. In some embodiments, performing targeted probe capture comprises performing linked target capture (LTC) using probe-dependent primers. In some embodiments, performing targeted probe capture comprises performing capture by circularization using molecular inversion probes (MIP). In some embodiments, performing targeted enrichment comprises performing targeted multiplex amplification of 25-20,000 target loci together in the same reaction volume.
[0014] In some embodiments, the method comprises whole genome sequencing or whole exome sequencing of a tumor tissue or its matched normal sample of the subject. In some embodiments, the method does not comprise whole genome sequencing or whole exome sequencing of a tumor tissue sample or its matched normal of the subject.
[0015] In some embodiments, the identification of at least one alteration associated with the disease that is present in both the first enriched DNA and the second enriched DNA is indicative of minimal residual disease.
[0016] In some embodiments, at least one alteration associated with cancer is identified in both the first enriched DNA and the second enriched DNA, and the method further comprises:34904-2842-5073.2Attorney Docket No. N.057.W0.01 extracting cellular DNA from a whole blood or buffy coat fraction of the blood sample; performing targeted enrichment on the extracted cellular DNA or their derivative to enrich a subset of the plurality of target loci and generating leukocyte-derived enriched DNA, wherein the subset comprises the at least one alteration; and sequencing the leukocyte-derived enriched DNA or their derivative to determine whether the at least one alteration is also present in the leukocyte- derived enriched DNA. In some embodiments, the method further comprises identifying at least one alteration that is not present in the leukocyte-derived enriched DNA, thereby identifying at least one alteration associated with the phenotype. In some embodiments, the method further comprises identifying at least one alteration that is also present in the leukocyte-derived enriched DNA, thereby identifying at least one clonal hematopoiesis (CH) alteration. In some embodiments, the method further comprises identifying at least one germline mutation present in the leukocyte-derived enriched DNA.
[0017] In some embodiments, the alteration associated with the phenotype comprises a genetic variant. In some embodiments, the alteration associated with the phenotype comprises a single nucleotide variant (SNV). a multi-nucleotide variant (MNV), an indel, a gene fusion, a structural variant, or a combination thereof.
[0018] In some embodiments, the alteration associated with the phenotype comprises an epigenetic alteration. In some embodiments, the alteration associated with the phenotype comprises a change in methylation status.
[0019] In some embodiments, the cancer is a solid tumor. In some embodiments, the solid tumor is breast cancer, advanced adenoma, colorectal cancer, kidney cancer, liver cancer, lung cancer, ovarian cancer, pancreatic cancer, bladder cancer, stomach cancer, esophagus cancer, small intestine cancer, prostate cancer, endometrial cancer, cervical cancer, melanoma, or head and neck cancers. In some embodiments, the cancer is a non-solid tumor. In some embodiments, the nonsolid tumor is a blood cancer. In some embodiments, the blood cancer is leukemia, lymphoma, or multiple myeloma.
[0020] In some embodiments, the subject has been treated with surgery, first-line chemotherapy, and / or adjuvant therapy. In some embodiments, the method further comprises longitudinally collecting a plurality of blood samples from the subject and repeating steps (a) to (g) for each of the plurality of blood samples.44904-2842-5073.2Attorney Docket No. N.057.W0.01
[0021] Another aspect of the invention described herein relates to a method for amplifying and sequencing nucleic acids, comprising: (a) extracting nucleic acids from a sample of a subject to obtain extracted nucleic acids; (b) ligating an adaptor having a capture moiety to the extracted nucleic acids to obtain adapted nucleic acids; (c) preparing a first library of amplified DNA comprising first amplified DNA by performing a first amplification on the adapted nucleic acids and separating the adapted nucleic acids from the first amplified DNA; (d) sequencing at least a portion of the first library of amplified DNA or its derivative and generating a first set of sequence reads; (e) preparing a second library of amplified DNA comprising second amplified DNA by performing a second amplification on the adapted nucleic acids and separating the adapted nucleic acids from the second amplified DNA; and (f) sequencing at least a portion of the second library of amplified DNA or its derivative and generating a second set of sequence reads.
[0022] Yet another aspect of the invention described herein relates to a method of amplifying and sequencing DNA, comprising; (a) extracting cell-free DNA from a liquid sample from a subject to generate extracted DNA; (b) ligating an adaptor having a capture moiety to the extracted DNA to obtain adapted DNA; (c) preparing a first library of methylation-preserved amplified DNA having one or more methylation-transferred DNA molecules, wherein said first library of methylation- preserved amplified DNA is prepared by (i) performing a first copying of one or more template DNA molecules in the adapted DNA and generating one or more copied DNA molecules, (ii) transferring methylation status from the one or more template DNA molecules to the one or more copied DNA molecules and generating one or more methylation-transferred copied DNA molecules, (iii) optionally repeating the copying and the transferring one or more rounds, and (iv) separating the methylation-transferred copied DNA molecules from the adapted DNA; (d) treating the first library of methylation-preserved amplified DNA or a portion thereof with an agent or a combination of agents and generating treated first methylation-preserved amplified DNA, wherein the agent or the combination of agents discriminates between methylated and unmethylated cytosines; (e) performing a first targeted enrichment on the treated first methylation-preserved amplified DNA or its derivative to enrich a first plurality of differentially methylated regions that are differentially methylated in a phenotype and generating first enriched DNA; and (f) sequencing the first enriched DNA or its derivative and generating a first set of sequence reads.
[0023] In some embodiments, the method further comprises (g) preparing a second library of methylation-preserved amplified DNA having one or more methylation-transferred DNA54904-2842-5073.2Atorney Docket No. N.057.W0.01 molecules, wherein said second library of methylation-preserved amplified DNA is prepared by (i) performing a second copying of one or more template DNA molecules in the adapted DNA and generating one or more copied DNA molecules, (ii) transferring methylation status from the one or more template DNA molecules to the one or more copied DNA molecules and generating one or more methylation-transferred copied DNA molecules, (iii) optionally repeating the copying and the transferring one or more rounds, and (iv) separating the methylation-transferred copied DNA molecules from the adapted DNA; (h) treating the second library of methylation-preserved amplified DNA or a portion thereof with an agent or a combination of agents and generating treated second methylation-preserved amplified DNA, wherein the agent or the combination of agents discriminates between methylated and unmethylated cytosines; (i) performing a second targeted enrichment on the treated second methylation-preserved amplified DNA or its derivative to enrich the first plurality of differentially methylated regions that are differentially methylated in a phenotype and generating second enriched DNA; (j) sequencing the second enriched DNA or its derivative and generating a second set of sequence reads; and (k) identifying the presence of or the likelihood of the presence of the phenotype based on the first set of sequence reads and the second set of sequence reads.
[0024] In some embodiments, the method further comprises (g) preparing a second library of methylation-preserved amplified DNA having one or more methylation-transferred DNA molecules, wherein said second library of methylation-preserved amplified DNA is prepared by (i) performing a second copying of one or more template DNA molecules in the adapted DNA and generating one or more copied DNA molecules, (ii) transferring methylation status from the one or more template DNA molecules to the one or more copied DNA molecules and generating one or more methylation-transferred copied DNA molecules, (iii) optionally repeating the copying and the transferring one or more rounds, and (iv) separating the methylation-transferred copied DNA molecules from the adapted DNA; (h) treating the second library of methylation-preserved amplified DNA or a portion thereof with an agent or a combination of agents and generating treated second methylation-preserved amplified DNA, wherein the agent or the combination of agents discriminates between methylated and unmethylated cytosines; (i) performing a second targeted enrichment on the treated second methylation-preserved amplified DNA or its derivative to enrich a second plurality of target loci each encompassing at least one alteration associated with the phenotype and generating second enriched DNA, wherein the second plurality of target loci64904-2842-5073.2Attorney Docket No. N.057.W0.01 optionally comprise at least a subset of the first plurality of differentially methylated regions that are differentially methylated in the phenotype; (j) sequencing the second enriched DNA or its derivative and generating a second set of sequence reads; and (k) identifying the presence of or the likelihood of the presence of the phenotype based on the first set of sequence reads and the second set of sequence reads.
[0025] In some embodiments, the capture moiety comprises biotin, biotin dT, biotin-TEG, photocleavable (PC) biotin, or desthiobiotin-TEG, or biotin azide, wherein the separating is performed using streptavidin beads. In some embodiments, the capture moiety comprises click chemistry capture moiety.
[0026] In some embodiments, the method further comprises (gl) preparing a third library of methylation-preserved amplified DNA having one or more methylation-transferred DNA molecules, wherein said third library of methylation-preserved amplified DNA is prepared by (i) performing a third copying of one or more template DNA molecules in the adapted DNA and generating one or more copied DNA molecules, (ii) transferring methylation status from the one or more template DNA molecules to the one or more copied DNA molecules and generating one or more methylation-transferred copied DNA molecules, (iii) optionally repeating the copying and the transferring one or more rounds, and (iv) separating the methylation-transferred copied DNA molecules from the adapted DNA; and optionally repeating step (gl) to generate additional library of methylation-preserved amplified DNA having one or more methylation-transferred DNA molecules.
[0027] In some embodiments, the copying comprises isothermal extension. In some embodiments, the adaptor comprises a universal priming sequence, and wherein the isothermal extension is performed using a primer that binds to the universal priming sequence. In some embodiments, the primer comprises a locked nucleic acid (LNA) oligonucleotide or a 5’ linked primer.
[0028] In some embodiments, the adaptor comprises a molecular barcode. In some embodiments, the adaptor comprises at least one methylated cytosine and at least one unmethylated cytosine. In some embodiments, the methylation status is transferred using a methyltransferase.
[0029] In some embodiments, the agent or the combination of agents that discriminates between methylated and unmethylated cytosines comprises a deaminating agent or a combination of an oxidizing agent and a deaminating agent, wherein the deaminating agent is sodium bisulfite or a deaminase.74904-2842-5073.2Attorney Docket No. N.057.W0.01
[0030] In some embodiments, an error in isothermal amplification, methyltransferase treatment, chemical or enzymatic methylation status detection (e.g. bisulfite or enzymatic conversion, MSRE or MDRE treatment), enrichment (e.g. DNA or RNA polymerase-based amplification) or sequencing is identified from the first set of sequence reads in combination with the second set of sequence reads without using any molecular barcode.
[0031] In some embodiments, the panel of differentially methylated regions comprises 50-50.000 differentially methylated regions.
[0032] In some embodiments, performing targeted enrichment comprises performing targeted probe capture to enrich the panel of differentially methylated regions. In some embodiments, performing targeted probe capture comprises performing linked target capture (LTC) using probedependent primers. In some embodiments, performing targeted enrichment comprises performing targeted multiplex amplification to enrich the panel of differentially methylated regions. In some embodiments, the targeted multiplex amplification comprises amplification of 50-20,000 target loci together in the same reaction volume.
[0033] In some embodiments, the liquid sample is a blood, plasma, serum, or urine sample. In some embodiments, the sample comprises cell-free nucleic acid from a subject having or carrying a tumor, a transplanted organ, or a fetus. In some embodiments, the sample comprises cell-free DNA from a tumor, a transplant, or a fetus. In some embodiments, the sample comprises cell-free RNA from a tumor, a transplant, or a fetus.
[0034] In some embodiments, the disease is a cancer. In some embodiments, the cancer is breast cancer, advanced adenoma, colorectal cancer, kidney cancer, liver cancer, lung cancer, multiple myeloma, ovarian cancer, or pancreatic cancer. In some embodiments, the subject has been treated with surgery, first- line chemotherapy, and / or adjuvant therapy.
[0035] In some embodiments, the method further comprises longitudinally collecting a plurality of liquid samples from the subject and repeating steps (a) to (f) for each of the plurality of blood samples.
[0036] In some embodiments, the identification of one or more differentially methylated regions associated with cancer that is present in both the first enriched DNA and the second enriched DNA is indicative of minimal residual disease.
[0037] In some embodiments, the present disclosure provides methods and compositions for preparing and analyzing DNA molecules to identify at least one somatic variant or mutation or84904-2842-5073.2Attorney Docket No. N.057.W0.01 epigenetic change associated with cancer. Cell-free DNA (cfDNA) is extracted from at least one plasma fraction of a blood sample of a subject, and optionally cellular DNA is extracted from a fraction of the blood sample containing a plurality of leukocytes. cfDNA is captured and the same cfDNA is amplified sequentially at least twice, creating at least two sequentially amplified DNA libraries. The sequentially amplified DNA libraries are analyzed by sequencing to identify at least one consensus alteration present in the plasma-derived cfDNA that is not a clonal hematopoiesis (CH) mutation present in the leukocyte-derived amplified DNA.
[0038] Further details regarding aspects and embodiments of the present disclosure are provided throughout this patent application. Sections and section headers are for ease of reading and are not intended to limit combinations of disclosure, such as methods, compositions, and kits or functional elements therein across sections. Further details regarding aspects and embodiments of the present disclosure are provided throughout this patent application. Sections and section headers are for ease of reading and are not intended to limit combinations of disclosure, such as methods, compositions, or other functional elements therein across sections.BRIEF DESCRIPTION OF THE DRAWINGS
[0039] FIG. 1 is a workflow diagram of one embodiment of the method described herein. Stars denote a biotin capture moiety.
[0040] FIG. 2 shows the average depth of read (DOR) of the two target replicate libraries for each test sample.
[0041] FIG. 3 shows VAF distribution across twelve base substitution types for two replicate libraries separately amplified from the biotinylated original library (biotinylated library rep 1 and biotinylated library rep 2).
[0042] FIG. 4 shows total SNVs called for the two sequentially amplified libraries (replicate 1 and replicate 2) and consensus calls between the libraries across eight different healthy cfDNA samples (M1-M8).
[0043] FIG. 5 shows that the presently disclosed methods detected two spiked-in mutations (i.e., in samples M7 and M8) with high confidence in both amplified libraries.
[0044] Unless otherwise explained, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The singular terms “a,” “an,” and “the” include plural referents unless the context clearly94904-2842-5073.2Attorney Docket No. N.057.W0.01 indicates otherwise. “Comprising A or B” means including A, or B, or A and B. It is further to be understood that all base sizes or amino acid sizes, and all molecular weight or molecular mass values, given for DNA molecules or polypeptides are approximate, and are provided for description.
[0045] Further, ranges provided herein are understood to be shorthand for all of the values within the range. For example, a range of 1 to 50 is understood to include any number, combination of numbers, or sub-range from the group consisting, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 1 to 49, 1 to 25, 1.7 to 31.9, and so forth (as well as fractions thereof unless the context clearly dictates otherwise). Any concentration range, percentage range, ratio range, or integer range is to be understood to include the value of any integer within the recited range and, when appropriate, fractions thereof (such as one tenth and one hundredth of an integer), unless otherwise indicated. Also, any number range recited herein relating to any physical feature, such as polymer subunits, size or thickness, are to be understood to include any integer within the recited range, unless otherwise indicated. When multiple low and multiple high values for ranges are given that overlap, a skilled artisan will recognize that a selected range will include a low value that is less than the high value.
[0046] As used herein, “about” or “consisting essentially of’ mean ± 10% of the indicated range, value, or structure, unless otherwise indicated. As used herein, the terms “include” and “comprise” are open ended and are used synonymously. As used herein, “comprising” is synonymous with “including,” “containing,” or “characterized by,” and is inclusive or open-ended and does not exclude additional, unrecited elements or method steps. As used herein, “consisting of’ excludes any element, step, or ingredient not specified in the claim element. As used herein, “consisting essentially of’ does not exclude materials or steps that do not materially affect the basic and novel characteristics of the claim. In each instance herein any of the terms “comprising”, “consisting essentially of’ and “consisting of’ may be replaced with either of the other two terms. The invention illustratively described herein suitably may be practiced in the absence of any element or elements, limitation or limitations which is not specifically disclosed herein.
[0047] Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, suitable methods and materials are described below. All publications, patent applications, patents, and other references mentioned104904-2842-5073.2Attorney Docket No. N.057.W0.01 herein are incorporated by reference in their entireties. In case of conflict, the present specification, including explanations of terms, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting.
[0048] It is appreciated that certain features of aspects and embodiments herein, which are, for clarity, discussed in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various aspects and embodiments, which are, for brevity, discussed in the context of a single aspect or embodiment, may also be provided separately or in any suitable sub-combination. All combinations of aspects and embodiments are specifically embraced herein and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all sub-combinations of the various aspects and embodiments and elements thereof are also specifically disclosed herein even if each and every such subcombination is not individually and explicitly disclosed herein.DETAILED DESCRIPTION
[0049] The present disclosure addresses many long-felt needs and long-standing problems in the art, such as. but not limited to, those mentioned in the Background section herein. For example, methods are provided herein for preparing nucleic acid molecules useful for identifying at least one genetic variant or mutation or epigenetic change (i.e., alteration) associated with a phenotype, for example, cancer. In some embodiments described herein, identification of at least one alteration associated with a phenotype is useful for early detection of a disease, such as asymptomatic cancers, and for detecting relapse or metastasis of cancers, for example in minimal residual disease (MRD).
[0050] Previously, workflows for preparing nucleic acid molecules sometimes rely on error correction methods based on dividing an original nucleic acid sample and performing parallel library preparation for each aliquot. Since circulating tumor nucleic acid, such as circulating tumor DNA (ctDNA), is present at a low relative abundance within a cell-free nucleic acid, such as cell- free DNA (cfDNA), sample, dividing the original nucleic acid sample may result in artifactually low or high ctDNA in the two or more aliquots, resulting in potential false negative calls. The presently disclosed methods solve this problem by reusing the original nucleic acid molecules in library amplification. As described herein, reusing original nucleic acid molecules allows for robust error rate correction and improved sensitivity for target nucleic acid detection. In addition,114904-2842-5073.2Attorney Docket No. N.057.W0.01 having the original nucleic acid molecules available for two or more library preparations suitable for two or more downstream analyses (e.g. library preparations for detecting genetic or epigenetic alterations) allows for an integrated approach for disease or cancer detection or prediction and therapy selection based on multi-omics, such as the presence or absence of polymorphisms or mutation (e.g. SNVs, indels, copy number variations), altered (increased or decreased) levels of total or particular cfDNA, cfRNA, microRNA (miRNA); altered (increased or decreased) tumor fraction; altered (increased or decreased) methylation levels, altered (increased or decreased) DNA integrity, altered (increased or decreased) or alternative mRNA splicing, or altered fragmentomics, etc.
[0051] Accordingly, as illustrated in FIG. 1, provided herein in some aspects are methods for preparing nucleic acid molecules (e.g. deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) molecules) that have been extracted from a blood sample of a subject. In some embodiments, the subject is at a risk or suspected of having a disease, such as cancer. In some embodiments, the subject has previously been diagnosed of and / or treated for a disease, such as cancer. In some embodiments, the subject is a pregnant woman. In some embodiments, the subject is an organ transplant recipient or a candidate for receiving an organ transplant.
[0052] In general, the invention described herein encompasses sequential preparation of a plurality of DNA libraries from the same original nucleic acids. Such methods may comprise the steps of: extracting nucleic acids from a sample of a subject; ligating an adaptor having a capture moiety to the extracted nucleic acids; preparing a first library of amplified DNA comprising first amplified DNA by performing a first amplification on the adapted nucleic acids and separating the adapted nucleic acids from the first amplified DNA; sequencing at least a portion of the first library of amplified DNA or its derivative and generating a first set of sequence reads; preparing a second library of amplified DNA comprising second amplified DNA by performing a second amplification on the adapted nucleic acids and separating the adapted nucleic acids from the second amplified DNA; and sequencing at least a portion of the second library of amplified DNA or its derivative and generating a second set of sequence reads, wherein the sequencing of the first and second libraries can be separate or together.
[0053] The invention described herein is particularly advantageous in improving methods for detecting alterations associated with a phenotype (e.g., somatic variants or mutations associated with cancer). Such methods may comprise the steps of: extracting nucleic acids from a sample of124904-2842-5073.2Attorney Docket No. N.057.W0.01 a subject to obtain extracted nucleic acids; ligating an adaptor having a capture moiety to the extracted nucleic acids to obtain adapted nucleic acids; performing a first amplification on the adapted nucleic acids and generating first amplified DNA, and separating the adapted nucleic acids from the first amplified DNA; performing a second amplification on the adapted nucleic acids and generating second amplified DNA, and separating the adapted nucleic acids from the second amplified DNA; performing a first targeted enrichment on the first amplified DNA or its derivative to enrich a plurality of target loci each encompassing at least one alteration associated with a phenotype and generating first enriched DNA; performing a second targeted enrichment on the second amplified DNA or its derivative to enrich the plurality of target loci and generating second enriched DNA; sequencing the first enriched DNA and the second enriched DNA or their respective derivative and producing a first set of sequence reads and a second set of sequence reads, respectively, wherein the sequencing of the first and second enriched DNA can be separate or together in the same sequencing run, and wherein at least one alteration associated with the phenotype present in both the first enriched DNA and the second enriched DNA can be identified based on the first set of sequence reads and the second set of sequence reads.
[0054] The invention described herein is also advantageous in improving methods for detecting an epigenetic change associated with a phenotype (e.g., differentially methylated regions that are differentially methylated in cancer). Such methods may comprise the steps of: extracting cell-free DNA from a liquid sample from a subject to generate extracted DNA; ligating an adaptor having a capture moiety to the extracted DNA to obtain adapted DNA; preparing a first library of methylation-preserved amplified DNA having one or more methylation-transferred DNA molecules, wherein said first library of methylation-preserved amplified DNA is prepared by (i) performing a first copying of one or more template DNA molecules in the adapted DNA and generating one or more copied DNA molecules, (ii) transferring methylation status from the one or more template DNA molecules to the one or more copied DNA molecules and generating one or more methylation-transferred copied DNA molecules, (iii) optionally repeating the copying and the transferring one or more rounds, and (iv) separating the methylation-transferred copied DNA molecules from the adapted DNA; treating the first library of methylation-preserved amplified DNA or a portion thereof with an agent or a combination of agents and generating treated first methylation-preserved amplified DNA, wherein the agent or the combination of agents discriminates between methylated and unmethylated cytosines; performing a first targeted134904-2842-5073.2Attorney Docket No. N.057.W0.01 enrichment on the treated first methylation-preserved amplified DNA or its derivative to enrich a plurality of differentially methylated regions that are differentially methylated in a phenotype and generating first enriched DNA; preparing a second library of methylation-preserved amplified DNA having one or more methylation-transferred DNA molecules, wherein said second library of methylation-preserved amplified DNA is prepared by (i) performing a second copying of one or more template DNA molecules in the adapted DNA and generating one or more copied DNA molecules, (ii) transferring methylation status from the one or more template DNA molecules to the one or more copied DNA molecules and generating one or more methylation-transferred copied DNA molecules, (iii) optionally repeating the copying and the transferring one or more rounds, and (iv) separating the methylation-transferred copied DNA molecules from the adapted DNA; treating the second library of methylation-preserved amplified DNA or a portion thereof with an agent or a combination of agents and generating treated second methylation-preserved amplified DNA, wherein the agent or the combination of agents discriminates between methylated and unmethylated cytosines; performing a second targeted enrichment on the treated second methylation-preserved amplified DNA or its derivative to enrich the plurality of differentially methylated regions that are differentially methylated in a phenotype and generating second enriched DNA; sequencing the first enriched DNA and the second enriched DNA or their respective derivative and producing a first set of sequence reads and a second set of sequence reads, respectively, wherein the sequencing of the first and second enriched DNA can be separate or together in the same sequencing run, and wherein at least one differentially methylated region associated with the phenotype can be identified based on the first set of sequence reads and the second set of sequence reads.
[0055] The invention described herein is particularly advantageous in improving methods for integrating multiple omics and signals associated with a phenotype (e.g., cancer). The methods described herein allow preparation of multiple types of libraries, each suitable for a different type of biomarker, from the same original nucleic acid molecules.
[0056] In some embodiments, the methods described herein comprise steps of extracting cell-free DNA from a plasma fraction of a blood sample of a subject, and optionally extracting cellular DNA from leukocytes from whole blood or from a buffy coat fraction of the blood sample; performing targeted enrichment (for example, by targeted probe capture or targeted multiplex amplification) on the extracted cell-free DNA or DNA derived therefrom to enrich a panel of a144904-2842-5073.2Attorney Docket No. N.057.W0.01 plurality of target loci each encompassing at least one somatic variant or mutation or epigenetic change (i.e., alteration) associated with the cancer to obtain a non-naturally occurring composition of plasma-derived enriched DNA, and optionally performing targeted enrichment (for example, by targeted probe capture or targeted multiplex amplification) on the extracted cellular DNA or DNA derived therefrom to enrich one or more of the target loci from the panel of target loci to obtain a non-naturally occurring composition of leukocyte-derived enriched DNA; and analyzing the non- naturally occurring composition of plasma-derived enriched DNA and optionally the non-naturally occurring composition of leukocyte-derived enriched DNA by sequencing to identify the at least one alteration that is present in the plasma-derived DNA and that is not an alteration, such as a CH mutation, present in the leukocyte-derived DNA. In some embodiments, the panel of target loci comprises at least 25, at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 750, at least 1000, at least 1500. at least 2500, at least 4500, or at least 4999 target loci. In some embodiments, the panel of target loci comprises no more than 5000, no more than 4500, no more than 2500, no more than 1500, no more than 1000, no more than 750, no more than 500, no more than 400. no more than 300, no more than 200, no more than 100. no more than 50, or no more than 25 target loci. In some embodiments, the panel of target loci comprises between 25-200, 200-400, 25-5000, 50-4000, 100-3000, 200-2000, or 400-1000 target loci. In some embodiments, the panel of target loci has a combined amplicon size of at least Ikb, at least 5kb, at least lOkb, at least 25kb, at least 50kb, at least lOOkb, at least 200kb, at least 300kb, at least 400kb, or at least 500kb. In some embodiments, the panel of target loci has a combined amplicon size of no more than 500kb, no more than 400kb, no more than 300kb, no more than 200kb, no more than lOOkb, no more than 50kb, no more than 25kb, no more than 10 kb, no more than 5kb, or no more than Ikb. In some embodiments, the panel of target loci has a combined amplicon size between Ikb- 500kb, 25kb-400kb, 50kb-300kb, 100kb-2000kb, or 250kb-500kb. In some embodiments, at least one of the target loci in the panel of target loci is selected based on sequencing of tumor samples from one or more patients known to have cancer. In some embodiments, at least 10%, at least 25%, at least 50%, at least 75%, at least 80%, or at least 90% of the target loci in the panel of target loci are selected based on sequencing of tumor samples from one or more patients known to have cancer. In some embodiments, at least one of the target loci in the panel of target loci is a hotspot mutation. In some embodiments, at least 10%, at least 25%, at least 50%, at least 75%, at least 80%, or at least 90% of the target loci in the panel of target loci are hotspot mutations.154904-2842-5073.2Attorney Docket No. N.057.W0.01
[0057] In some embodiments, targeted enrichment is by targeted probe capture. In some embodiments, targeted enrichment is by linked target capture with probe-dependent primers to enrich the target loci. In some embodiments, targeted enrichment is by targeted multiplex amplification using target-specific primers. In some embodiments, all the target loci in the panel of target loci are enriched together in the same reaction volume. In some embodiments, the target loci in the panel of target loci are enriched in 2, 3, 4, 5, 10, 20, or more reaction volumes or pools, each enriching an about equal number of different target loci from the panel of target loci. For example, each reaction volume or pool may enrich 20, 50, 100, 200, 500, or 1000 target loci.
[0058] In some embodiments, extraction, amplification and targeted enrichment of leukocyte- derived cellular DNA or derivatives therefrom are performed in parallel of extraction, amplification and targeted enrichment of plasma-derived cell-free DNA or derivatives therefrom, wherein the same panel of target loci are enriched in both processes. In some embodiments of the methods described herein, targeted enrichment of leukocyte-derived cellular DNA or derivatives therefrom is only performed if and after one or more alterations are identified in plasma-derived DNA and / or the plasma sample from the subject is identified as having circulating tumor DNA (ctDNA). In some embodiments, only a sub-pool of target loci comprising the one or more target loci encompassing the one or more somatic variants or mutations or epigenetic changes identified in the plasma-derived DNA are enriched from the leukocyte-derived cellular DNA or derivatives therefrom. In some embodiments, only the one or more target loci encompassing the one or more somatic variants or mutations or epigenetic changes identified in the plasma-derived DNA are enriched from the leukocyte-derived cellular DNA or derivatives therefrom.
[0059] In some embodiments, if a somatic variant or mutation or epigenetic change is identified in the cell-free DNA and the same somatic variant or mutation or epigenetic change is also identified in the cellular DNA from matched whole blood / buffy coat, the alteration would be considered a non-tumor specific alteration, such as a CH mutation, and filtered. In some embodiments, if a somatic variant or mutation or epigenetic change is identified in the cell-free DNA and the same alteration is not identified in the cellular DNA from matched whole blood / buffy coat, the alteration can be considered a true tumor specific alteration.
[0060] In some embodiments, at least one adaptor comprises a capture moiety. In some embodiments the capture moiety is biotin or a biotin derivative, which binds to streptavidin or a streptavidin derivative immobilized on a substrate. In some embodiments, the substrate is a164904-2842-5073.2Attorney Docket No. N.057.W0.01 magnetic bead. In some embodiments, the adapted nucleic acid is amplified before it binds to the substrate. In some embodiments, the adapted nucleic acid is amplified after it binds to the substrate. In some embodiments, (i) the adapted nucleic acid is amplified producing a first set of amplicons lacking capture moieties, (ii) the substrate is added to the amplified mixture thereby binding and immobilizing the adapted nucleic acid, (iii) the first set of amplicons is removed to a first pool, (iv) the adapted nucleic acid is reamplified producing a second set of amplicons, and (v) the second set of amplicons is removed to a second pool.
[0061] A person skilled in the art would appreciate that a variety of techniques are suitable for immobilizing an adapted nucleic acid to a substrate, provided there is sufficient affinity between the capture moiety and the substrate.
[0062] In some embodiments, at least one adaptor comprises a biotin or biotin derivative as the capture moiety. In some embodiments, the capture moiety comprises biotin, biotin dT, biotin- TEG, photocleavable (PC) biotin, desthiobiotin-TEG, or biotin azide, and the separating is performed using streptavidin beads. In some embodiments, the adaptor is a 5’ biotinylated adaptor.
[0063] In some embodiments, after a first amplification is performed on the adapted nucleic acids with biotinylated adaptors, the adapted nucleic acids are captured on streptavidin beads, and the first amplified DNA in the supernatant are not biotinylated and can be separate and subject to targeted enrichment to produce first enriched DNA. Beads containing the bound adapted nucleic acids with biotinylated adaptors are then used for a second amplification (e.g., on-bead PCR), and again the second amplified DNA in the supernatant are not biotinylated and can be separate and subject to targeted enrichment to produce second enriched DNA.
[0064] In some embodiments, at least one adaptor is chemically modified with a thiol group, and the separating can be performed using a gold substrate (e.g., gold nanoparticles) via Au-S bonding. See Cao et al., J. Am. Chem. Soc., 123:7961-7962 (2001); Kim et al., Angew. Chem., Int. Ed. Engl., 50:9185-9190 (2011), which are incorporated herein by reference in their entirety.
[0065] In some embodiments, at least one adaptor is chemically modified with an amine group, and the separating can be performed using an aldehyde group-modified substrate via formation of a Schiff base (e.g. EDC (1 -(3 -dimethyl aminopropyl)-3-ethyl carbodiimide hydrochloride)-NHS (N-hydroxy succinimide) reaction and Schiff base reaction). See Fuentes et al., Biosens.174904-2842-5073.2Attorney Docket No. N.057.W0.01Bioelectron., 21:1574-1580 (2006); Nguyen et al. Sci. Rep., 8:337. (2018), which are incorporated herein by reference in their entirety.
[0066] In some embodiments, the capture moiety is a click chemistry capture moiety. In some embodiments, at least one adaptor is chemically modified with an azide group, and the separating can be performed using an alkyne-functionalized substrate. See McKenna et al., Sensors and Actuators B: Chemical, 236: 286-293 (2016); El-Sagheer et al., Chem. Soc. Rev., 39:1388-1405 (2010), which are incorporated herein by reference in their entirety.
[0067] In some embodiments, sample barcodes are added to the extracted nucleic acid molecules before or after targeted enrichment. In some embodiments, sequencing reads from the at least two library amplifications are combined to identify at least one somatic variant or mutation or epigenetic change (i.e., alteration) that is present in the plasma-derived nucleic acid molecules.
[0068] In some embodiments, the method described herein comprises sequencing a plurality of sequentially amplified DNA libraries (e.g., sequentially amplified from the same adapted nucleic acids) or DNA derived therefrom to obtain a plurality of sets of sequence reads, and combining and analyzing the plurality of sets of sequence reads to identify the presence or absence of at least one alteration associated with cancer. In some embodiments, the methods described herein comprise performing targeted enrichment on the adapted- amplified DNA or DNA derived therefrom to enrich a plurality of target loci each encompassing a potential alteration associated with cancer, prior to sequencing.
[0069] In some embodiments, at least one processing or sequencing error, if present, is identified from the combined reads from sequencing of the DNA derived from each of the two or more sequentially amplified libraries. In some embodiments, the method does not comprise tagging the extracted cell-free nucleic acid molecules with a molecular barcode; instead, sequence reads derived from the two or more sequentially amplified libraries can cross-check each other to eliminate processing / sequencing errors or false positives.
[0070] In some embodiments, the method may further comprise treating the adapted nucleic acids separated from the amplified DNA used to identify somatic variants or mutations, or amplicons derived from the adapted original nucleic acids through one or more methylation-preserving amplifications, with an agent or a combination of agents and generating treated nucleic acids or treated DNA, wherein the agent or the combination of agents discriminates between methylated and unmethylated cytosines. In some embodiments, the method may further comprise subjecting184904-2842-5073.2Atorney Docket No. N.057.W0.01 the treated nucleic acids or treated DNA to targeted enrichment and sequencing of a plurality of differentially methylated regions that are differentially methylated in the cancer.
[0071] In some embodiments, the method may further comprise targeted enrichment and sequencing of a plurality of sequences that encode proteins or miRNA that are differentially expressed in the cancer. In some embodiments, targeted enrichment and sequencing of a plurality of differentially expressed miRNA or protein sequences is performed on a library of amplified DNA amplified from the same adapted nucleic acids as another library of amplified DNA used to identify somatic variants or mutations. In some embodiments, the method may further be combined with fragmentomics.
[0072] In some embodiments, the extracted cell-free DNA and / or extracted cellular DNA is end repaired and A-tailed, prior to adaptor ligation. In some embodiments, at least one adaptor comprises a universal priming sequence. In some embodiments, prior to performing targeted enrichment, the adapted DNA is amplified using the universal priming sequence. In some embodiments, a universal amplification is performed after the targeted enrichment to obtain a non- naturally occurring composition of plasma-derived amplified DNA. In some embodiments, a universal amplification is performed after the targeted enrichment to obtain a non-naturally occurring composition of leukocyte-derived amplified DNA. In some embodiments, at least one target loci is amplified using two or more target-specific primers having overlapping sequences.
[0073] In some embodiments, analyzing the sequencing reads comprises identifying an observed genomic (e.g. SNV) or epigenomic (e.g. methylation) variant as a true or false variant or having a percentage likelihood of being a true or false variant (sometimes referred to as a variant caller). In some embodiments, analyzing the sequencing reads comprises classifying the sample as containing cancer-derived DNA (positive) or not containing cancer-derived DNA (negative) (sometimes referred to as a sample caller). In some embodiments of any of the methods described herein, analyzing the sequencing reads comprises using an error model. In some embodiments, analyzing the sequencing reads comprises using a position-based background error model. In some embodiments, analyzing the sequence reads comprises using position-specific priors. In some embodiments, analyzing the sequencing reads comprises using VAF priors. In some embodiments, analyzing the sequencing reads comprises using consequence priors. In some embodiments, analyzing the sequencing reads comprises using gene-level priors. In some embodiments, analyzing the sequencing reads comprises using sample priors. In some embodiments, analyzing194904-2842-5073.2Atorney Docket No. N.057.W0.01 the sequencing reads comprises using one or more priors to make a variant call. In some embodiments, analyzing the sequencing reads comprises using one or more priors to make a sample call. In some embodiments, analyzing the sequencing reads comprises adjusting a confidence level based on concordance between two plasma replicates. In some embodiments, analyzing the sequencing reads comprises filtering variants called in both plasma-derived and whole blood / buffy-derived samples. In some embodiments, analyzing the sequencing reads comprises using methylation-preserving amplification priors.
[0074] In some embodiments, the somatic variant or mutation associated with cancer comprises a single nucleotide variant (SNV), a multi-nucleotide variant (MNV), an indel, a gene fusion, a structural variant, or a combination thereof. In some embodiments, the somatic variant or mutation associated with cancer comprises an SNV or an indel, or a combination thereof. In some embodiments, the methods described herein may further comprise identifying and / or filtering one or more germline variants present in the leukocyte-derived cellular DNA. In some embodiments, the epigenetic change associated with the cancer comprises an altered DNA methylation pattern.
[0075] In one aspect of the methods described herein, samples are called by using a plasma variant caller to first identify a variant or mutation or epigenetic change, or the likelihood of having a variant or mutation or epigenetic change, from each of the two or more replicate sequencing libraries sequentially amplified from the same original nucleic acid molecules and using a consensus variant caller to identify one or more consensus variants. In some embodiments, 2 or more replicate sequencing libraries sequentially amplified from the same original nucleic acid molecules are used for the consensus variant caller. In some embodiments, 3 or more replicate sequencing libraries sequentially amplified from the same original nucleic acid molecules are used for the consensus variant caller. In some embodiments, 4 or more replicate sequencing libraries sequentially amplified from the same original nucleic acid molecules are used for the consensus variant caller. In some embodiments, 5 or more replicate sequencing libraries sequentially amplified from the same original nucleic acid molecules are used for the consensus variant caller. In some embodiments, 10 or more replicate sequencing libraries sequentially amplified from the same original nucleic acid molecules are used for the consensus variant caller. A consensus variant will only be called if it is present or likely to be present in each of the replicate sequencing libraries or in the majority of the replicate sequencing libraries. The data is then filtered for nontumor specific alterations such as CH mutations using a matched whole blood / buffy sample. If an204904-2842-5073.2Atorney Docket No. N.057.W0.01 alteration is called in both the whole blood / buffy coat and plasma samples, the alteration is considered a non-tumor specific alteration and is filtered out. If an alteration is called in a plasma sample but absent from matched whole blood / buffy coat, the alteration is considered a tumor specific alteration. The sample is then called as cancer positive or cancer negative using a sample caller based on data and likelihoods collected from priors, such as sample priors and target WES / WGS priors.
[0076] The methods disclosed herein can be used in virtually any applications involving a small amount of starting nucleic acid molecules from a subject (e.g. from a liquid biopsy), including for example cancer or cancer-recurrence detection / diagnosis, other disease detection / diagnosis, therapy selection, non-invasive prenatal testing (NIPT), organ health or organ transplant evaluation, prediction and monitoring, and veterinary applications (e.g. aging or health status evaluation, prediction and monitoring of domesticated animals).Subjects
[0077] Subjects of the methods described herein can be virtually any animal, in illustrative embodiments a mammal, and in further illustrative embodiments, a human. In some embodiments, the subject has, had, or is suspected or at risk of having a disease, in some embodiments, cancer or precancerous lesions. In some embodiments, the subject has received or is receiving a treatment for cancer and is being tested for minimal residual disease (MRD). In some embodiments, the subject is a pregnant female. In some embodiments, the subject is a subject comprising an organ from another individual.
[0078] In embodiments where the subject has. is suspected to have, or is at risk of having cancer, the cancer can be any type of cancer. In some embodiments, the genome of cancerous cells of the subject have portions of their genome that are differentially methylated compared to non- cancerous cells of the subject. Typically, some, most, almost all or all cancer cells of the subject have region(s) of their genome that are methylated that are not methylated in non-cancerous cells of the subject, or that are more methylated than non-cancerous Thus, in some embodiments, the subject has, had, or is suspected or at risk of having one or more cancers (e.g., one cancer) including ovarian cancer, soft tissue sarcoma, peripheral T cell cancer, colorectal cancer, intrahepatic cholangiocarcinoma, glioblastoma, esophageal cancer, cutaneous T cell lymphoma, non-Hodgkin lymphoma, urothelial cancer, basal cell carcinoma, epithelioid sarcoma, pancreatic214904-2842-5073.2Atorney Docket No. N.057.W0.01 cancer, non-small cell lung carcinoma, Hodgkin lymphoma, renal cell carcinoma, mesothelioma, metastatic uveal melanoma, kidney cancer, blood cancer, HER2-expressing cancers, nonmelanoma skin cancer, liposarcoma, hepatocellular carcinoma, small lymphocytic lymphoma, prostate cancer, breast cancer, anal cancer, marginal zone lymphoma, cutaneous squamous cell carcinoma, thyroid cancer, medullary thyroid cancer, triple-negative breast cancer, neuroendocrine prostate cancer, bladder cancer, paraganglioma, medulloblastoma, superficial basal cell carcinoma, head and neck squamous cell carcinoma, hematologic malignancies, melanoma, B-cell lymphoma, relapsed / refractory acute myeloid leukemia, angiosarcoma, bone sarcoma, refractory cervical cancer, cholangiocarcinoma, osteosarcoma, biliary tract cancer, castration-resistant prostate cancer, gastroesophageal adenocarcinomas, rhabdomyosarcoma, carcinoma, non-muscle invasive bladder cancer, uveal melanoma, small cell lung cancer, cervical cancer, primary open angle glaucoma, follicular lymphoma, synovial sarcoma, liver cancer, carcinosarcoma, leptomeningeal brain tumors, T-cell lymphoma, lymphoma, small cell lung cancer, mantle cell lymphoma, B-cell malignancies, endometrial cancer, myxoid / round cell liposarcoma, metastatic Merkel cell carcinoma, neuroblastoma, chronic lymphocytic leukemia, tenosynovial giant cell tumors, sarcoma, acute myeloid leukemia, skin cancer, nasopharyngeal carcinoma, relapsed / refractory Ewing sarcoma, bone cancer, glioma, salivary gland carcinoma, gastric cancer, benign tumor, low- grade serous ovarian cancer, metastatic breast cancer, multiple myeloma, diffuse large B cell lymphoma, relapsed / refractory lymphoma, metastatic colorectal cancer, advanced malignancies, and acute lymphoblastic leukemia.
[0079] In certain embodiments of methods herein, the subject has, had, or is suspected or is at risk of having a cancer of the abdomen or abdominal wall, adrenal gland, anus, appendix, bladder, blood, bone, bone marrow, brain, breast, cervix, chest wall, colon, diaphragm, duodenum, ear, endometrium, esophagus, fallopian tube, gallbladder, gastro-esophageal junction, head and neck, kidney, larynx, liver, lung, lymph node, malignant effusions, mediastinum, nasal cavity, omentum, ovaries, pancreas, pancreatobiliary, parotid gland, pelvis, penis, pericardium, peritoneum, pleura, prostate, rectum, salivary gland, skin, small intestine, soft tissue, spleen, stomach, thyroid, tongue, trachea, ureter, uterus, vagina, vulva, and Whipple resection. In illustrative embodiments, the cancer is selected from nasopharyngeal carcinoma, hepatocellular carcinoma, breast cancer, ovarian cancer, pancreatic cancer, colorectal cancer, lung cancer, gastroesophageal cancer, prostate cancer, bladder cancer, melanoma, and acute leukemia. In illustrative embodiments, the224904-2842-5073.2Atorney Docket No. N.057.W0.01 cancer is selected from colorectal cancer. In some embodiments, the cancer is a solid tumor. In some embodiments, the solid tumor is colorectal cancer.Sample Collection
[0080] The methods disclosed herein are contemplated to be used to monitor or detect a wide variety of phenotypes, including cancers, in a patient. Different types of cancer may require collection of different types of samples as described herein. In some embodiments, the cancer is a solid tumor, and the biological sample is a tumor biopsy sample. Performing a biopsy generally involves using a sharp tool to remove a small amount of tissue from the area suspected to contain diseased cells or tissue such as a tumor. There are many different types of biopsies such as needle biopsy, CT-guided biopsy, ultrasound guided biopsy, bone biopsy, bone marrow biopsy, liver biopsy, kidney biopsy, aspiration biopsy, prostate biopsy, skin biopsy, surgical biopsy such as laparoscopic biopsy. In some embodiments, the biological sample is obtained by liquid biopsy. In some embodiments, the biological sample is a blood, serum, plasma, or urine sample. Further, biological liquid samples may be extracted from variety of animal fluids containing cell-free DNA, including but not limited to blood, serum, plasma, bone marrow, urine, vitreous, sputum, tears, perspiration, saliva, semen, mucosa excretions, feces, bile, ascites fluid, pleural fluid, peritoneal fluid, cerebrospinal fluid, amniotic fluid, lymph fluid, exhaled breath condensate, bronchoalveolar lavage fluid, and so on. In some embodiments, cell-free DNA may be transplant donor or fetal in origin (via fluid taken from a pregnant subject).
[0081] For many embodiments, the sample is a mixed sample with DNA or RNA from one or more target cells and one or more non-target cells. In some embodiments, the target cells are cells that have a CNV, such as a deletion or duplication of interest, and the non-target cells are cells that do not have the copy number variation of interest (such as a mixture of cells with the deletion or duplication of interest and cells without any of the deletions or duplications being tested). In some embodiments, the target cells are cells that are associated with a disease or disorder or an increased risk for disease or disorder (such as cancer cells), and the non-target cells are cells that are not associated with a disease or disorder or an increased risk for disease or disorder (such as noncancerous cells). In some embodiments, the target cells all have the same genetic or epigenetic alteration. In some embodiments, two or more target cells have different genetic or epigenetic alterations. In some embodiments, one or more of the target cells has a CNV, polymorphism,234904-2842-5073.2Attorney Docket No. N.057.W0.01 mutation, or epigenetic alteration associated with the disease or disorder or an increased risk for disease or disorder that is not found in at least one other target cell. In some such embodiments, the fraction of the cells that are associated with the disease or disorder or an increased risk for disease or disorder out of the total cells from a sample is assumed to be greater than or equal to the fraction of the most frequent of these CNVs, polymorphisms, mutations, or epigenetic alterations in the sample. For example if 6% of the cells have a K-ras mutation, and 8% of the cells have a BRAF mutation, at least 8% of the cells are assumed to be cancerous.
[0082] In some embodiments, the biological sample is a liquid sample. In some embodiments, the biological sample is blood, serum, plasma, buffy coat, or bone marrow sample. In some embodiments, cell-free DNA / RNA and cellular DNA / RNA are both obtained from a blood sample of the subject by isolating and separating the plasma (which contains cell-free DNA / RNA) and buffy coat (which contains most or all of the white blood cells) fractions. The cellular DNA / RNA obtained from the white blood cells from whole blood or in the buffy coat may serve as a matched normal DNA / RNA sample to the cell-free DNA / RNA obtained from the plasma fraction, which may include circulating tumor DNA / RNA.
[0083] In some embodiments, the methods of the present disclosure further comprise longitudinally collecting a plurality of liquid biopsy samples from the patient. In some embodiments, the liquid biopsy sample is obtained from the patient after the patient has been treated for the cancer. In some embodiments, the liquid biopsy sample is a blood, serum, plasma, buffy coat, or urine sample. In some embodiments, the methods described herein are repeated for each of the longitudinally collected samples. In some embodiments, the subject has received treatment for the cancer. In some embodiments, the longitudinally collected samples are collected after the subject has been treated with surgery, first-line chemotherapy, and / or adjuvant therapy. In some embodiments, the identification of the at least one somatic mutation in the longitudinally collected samples using the methods described herein is indicative of minimal residual disease. In further embodiments, the subject is suspected or at risk of having cancer and has not received treatment for the cancer, and the identification of the at least one alteration is indicative of the presence of the cancer in the subject.
[0084] In some embodiments, the methods described herein can be used at a very early period of time following cancer treatment, for example as early as the day of treatment, one day after treatment, two days after treatment, three days after treatment, four days after treatment, five days244904-2842-5073.2Attorney Docket No. N.057.W0.01 after treatment, six days after treatment, a week after treatment, two weeks after treatment, three weeks after treatment, four weeks after treatment, one month after treatment, two months after treatment, three months after treatment, four months after treatment, five months after treatment, six months after treatment, seven months after treatment, eight months after treatment, nine months after treatment, ten months after treatment, eleven months after treatment, or a year or more after treatment. In some embodiments, the methods described herein can be used prior to being diagnosed or and / or receiving treatment for cancer.Extraction of Nucleic Acid Molecules
[0085] In some embodiments, the method includes isolating or purifying the DNA and / or RNA. In some embodiments, the sample may be centrifuged to separate various layers. In some embodiments, the DNA or RNA may be isolated using filtration. In some embodiments, the preparation of the DNA or RNA may involve amplification, separation, purification by chromatography, liquid separation, isolation, preferential enrichment, preferential amplification, targeted amplification, or any of a number of other techniques either known in the art or described herein. In some embodiments for the isolation of DNA, RNase is used to degrade RNA. In some embodiments for the isolation of RNA, DNase (such as DNase I from Invitrogen, Carlsbad, Calif., USA) is used to degrade DNA. In some embodiments, an RNeasy mini kit (Qiagen), is used to isolate RNA according to the manufacturer's protocol, hi some embodiments, small RNA molecules are isolated using the mirVana PARIS kit (Ambion, Austin, Tex., USA) according to the manufacturer’s protocol (Gu et al., J. Neurochem. 122:641-649, 2012, which is hereby incorporated by reference in its entirety). The concentration and purity of RNA may optionally be determined using Nanovue (GE Healthcare, Piscataway, N.I., USA), and RNA integrity may optionally be measured by use of the 2100 Bioanalyzer (Agilent Technologies, Santa Clara, Calif., USA) (Gu et al., J. Neurochem. 122:641-649, 2012, which is hereby incorporated by reference in its entirety). In some embodiments, TRIZOL or RNAlater (Ambion) is used to stabilize RNA during storage.
[0086] Methods provided herein, in certain embodiments, are specially adapted for extracting and amplifying nucleic acid molecules, especially circulating tumor DNA (ctDNA) fragments. Such fragments are typically about 160 nucleotides in length.254904-2842-5073.2Attorney Docket No. N.057.W0.01
[0087] Cell-free nucleic acid (cfNA), e.g., cell-free DNA (cfDNA) and cell-free RNA (cfRNA), can be released into the circulation via various forms of cell death such as apoptosis, necrosis, autophagy and necroptosis. cfDNA is fragmented and the size distribution of the fragments varies from 150- 350 bp to> 10000 bp. By contrast, for example, the size distributions of ctDNA fragments from hepatocellular carcinoma (HCC) patients spanned a range of 100-220 bp in length with a peak in count frequency at about 166bp and the highest tumor DNA concentration in fragments of 150-180 bp in length.
[0088] In an illustrative embodiment, cell-free DNA (cfDNA) containing circulating tumor DNA (ctDNA) is extracted from the plasma fraction of blood after removal of cellular debris and platelets by centrifugation. The plasma samples can be stored at -80°C until the DNA is extracted using, for example, QIAamp DNA Mini Kit (Qiagen, Hilden, Germany), (e.g., Hamakawa et al., Br J Cancer. 2015; 112:352-356). Hamakava et al. reported median concentration of extracted cfDNA of all samples 43.1 ng per ml plasma (range 9.5-1338 ng / ml) and a mutant fraction range of 0.001-77.8%, with a median of 0.90%.
[0089] Samples that are useful for methods herein can be virtually any nucleic acid sample. In illustrative embodiments, the nucleic acid sample is extracted or isolated from a subject. In certain illustrative embodiments, DNA molecules are extracted from a sample such as a tissue sample, and for illustrative embodiments herein, a liquid sample. Methods that are particularly useful in exemplary embodiments include methods for isolating cfDNA from a liquid sample, and in illustrative embodiments from a blood, serum, bone marrow, urine, vitreous, sputum, saliva, tears, perspiration, mucosa excretions, feces, bile, lymph fluid, cervical mucus, ascites fluid, pleural fluid, exhaled breath condensate, peritoneal fluid, cerebrospinal fluid, amniotic fluid, bronchoalveolar lavage fluid, or semen sample, and in further illustrative embodiments, a plasma sample.
[0090] In certain illustrative embodiments, isolation of cfDNA or cellular DNA from a liquid (e.g., blood or blood derivative sample such as a serum, plasma, or buffy coat sample) can involve binding DNA molecules from a sample to a matrix and isolating the DNA molecules in the presence of a solvent. In some embodiments, the method further comprises incubating the biological sample comprising DNA molecules with a protease, prior to contacting the DNA molecules to the matrix. In some embodiments, the method can further include the steps of264904-2842-5073.2Atorney Docket No. N.057.W0.01 washing the matrix with a wash buffer to remove impurities, and optionally, drying the matrix. Enriched nucleic acid samples can be eluted from the matrix with an elution buffer.
[0091] Other methods for nucleic acid isolation, for example cfDNA or cellular DNA isolation, and optional enrichment of certain cfDNA or cellular DNA can include ion exchange columns, or microfluidic devices, such as solid phase isolation, based on DNA capture by immobilized beads or functionalized surface. Additional methods include liquid phase isolation, utilizing an electric field, or chemical reagents, instead of a functionalized surface. In some embodiments, isolation of cfDNA from a patient sample is performed using a DNA isolation kit (e.g., QIAamp Circulating Nucleic Acid kit (Qiagen)). In some embodiments, capture by hybridization with hybrid capture probes is used to preferentially enrich the DNA, for example using probes that bind to specific nucleic acid sequences at or near, target regions.
[0092] In some embodiments, cfDNA or their derivatives of certain sizes can be enriched before or after subjecting the cfDNA to the sequential amplification methods described herein. In some embodiments, size selection can be performed before the sequencing library preparation. In some embodiments, size selection can be performed after the sequencing library preparation and before sequencing. In some embodiments, size selection is performed on a sequencing-ready pool. Enriched cfDNA molecules can be, for example 50 to 1200 base pairs in length, 70 to 500 base pairs in length, 100 to 200 base pairs in length, or 130 to 170 base pairs in length. In some embodiments, the enriched cfDNA molecules are from 50 to 200 bp in length. In some embodiments, the enriched cfDNA molecules are between 60 and 200 bp in length, between 60 and 150 bp in length, or between 60 and 100 bp in length before the enriched cfDNA molecules, or derivatives thereof, are ligated to adaptors in methods herein. In some embodiments, the enriched cfDNA molecules are less than 150, 100, 90, 75, or 50 bp in length before they are ligated to adaptors. Such enrichment methods can be performed for example using the methods described in WO2018 / 156418 and WO2019161244, each of which is incorporated herein by reference in its entirety.
[0093] In some embodiments, the sample is enriched for tumor DNA molecules, which are typically less than 160 bp and have a peak at about 145 bp in length. In illustrative embodiments, the enriched nucleic acid sample is circulating tumor DNA (ctDNA), or amplicons thereof. In such embodiments, the enriched nucleic acid is less than 160, 150, 145, 120, 100, 90, 75, or 50 bp in length. In some embodiments, size selection is used to filter out cfDNA molecules carrying274904-2842-5073.2Attorney Docket No. N.057.W0.01 mutations derived from clonal hematopoiesis of indeterminate potential (CHIP) but not tumor- derived mutations, which are typically longer than ctDNA molecules, at about 165bp.
[0094] In some embodiments, the sample is enriched for fetal DNA molecules. In illustrative embodiments of such embodiments, the enriched nucleic acid sample is fetal cell-free DNA or amplicons thereof. In such embodiments, the enriched nucleic acid is between 100 bp and 220 bp in length. In some embodiments, the length of the nucleic acid sample of fetal cfDNA ranges from 100 bp to 200 bp, from 120 bp to 180 bp, from 140 bp to 160 bp, from 150 bp to 170 bp, from 160 to 190 bp, or from 170 bp to 220 bp.
[0095] In some embodiments, the sample is enriched for transplant donor DNA molecules. In illustrative embodiments of such embodiments, the enriched nucleic acid sample is transplant donor cell-free DNA or amplicons thereof. In such embodiments, the enriched nucleic acid is between 100 bp and 220 bp in length. In some embodiments, the length of the nucleic acid sample of transplant donor cfDNA ranges from 100 bp to 200 bp, from 120 bp to 180 bp, from 140 bp to 160 bp, from 150 bp to 170bp, from 160 bp to 190 bp, or from 170 bp to 220 bp.Nucleic Acid Processing
[0096] In some embodiments, the method described herein comprises ligating an adaptor having a capture moiety to the extracted nucleic acids to obtain adapted nucleic acids. In some embodiments, the method described herein comprises performing a first amplification on the adapted nucleic acids and generating first amplified DNA, and separating the adapted nucleic acids from the first amplified DNA. In some embodiments, the method described herein comprises performing a second amplification on the adapted nucleic acids and generating second amplified DNA, and separating the adapted nucleic acids from the second amplified DNA.
[0097] Additional embodiments of nucleic acid processing relating to analysis of epigenetic changes are described in “Template Copying” and “Methylation Transfer” sections. In some embodiments, the method described herein comprises preparing a first library of methylation- preserved amplified DNA having one or more methylation-transferred DNA molecules, wherein said first library of methylation-preserved amplified DNA is prepared by (i) performing a first copying of one or more template DNA molecules in the adapted DNA and generating one or more copied DNA molecules, (ii) transferring methylation status from the one or more template DNA molecules to the one or more copied DNA molecules and generating one or more methylation-284904-2842-5073.2Atorney Docket No. N.057.W0.01 transferred copied DNA molecules, (iii) optionally repeating the copying and the transferring one or more rounds, and (iv) separating the methylation-transferred copied DNA molecules from the adapted DNA. In some embodiments, the method described herein comprises preparing a second library of methylation-preserved amplified DNA having one or more methylation-transferred DNA molecules, wherein said second library of methylation-preserved amplified DNA is prepared by (i) performing a second copying of one or more template DNA molecules in the adapted DNA and generating one or more copied DNA molecules, (ii) transferring methylation status from the one or more template DNA molecules to the one or more copied DNA molecules and generating one or more methylation-transferred copied DNA molecules, (iii) optionally repeating the copying and the transferring one or more rounds, and (iv) separating the methylation-transferred copied DNA molecules from the adapted DNA.
[0098] Typically, methods herein include a step of appending, in some embodiments ligating, nucleic acid adaptors to sample DNA / RNA molecules, or in illustrative embodiments to nucleic acid derivatives generated therefrom. For example, adaptors may be appended on to the DNA / RNA molecules by ligation or PCR. The sample nucleic acid molecules in illustrative embodiments are cell-free DNA or cellular DNA extracted from a sample of a subject. In some embodiments methods include exposing sample DNA molecules to one or more polymerases or kinases, such as Klenow Large Fragment Polymerase and T4 polynucleotide kinase (PNK), as well as a ligase, such as T4 DNA ligase. In some embodiments, sample nucleic acid molecules or fragmented nucleic acid molecules are exposed to one or more polymerases and / or kinases to generate the nucleic acid derivatives generated therefrom. In some embodiments, the method further comprises ligating adaptors to the nucleic acid derivatives generated therefrom. In some embodiments, sample DNA molecules are not fragmented prior to appending nucleic acid adaptors thereto. In some embodiments, sample nucleic acid molecules are cfDNA molecules or cfRNA molecules.
[0099] In some embodiments, before adaptor ligation, extracted or isolated nucleic acid molecules can be modified to form sample nucleic acid derivatives, for example to make them more amenable to adaptor ligation. For example, sample DNA molecules can be blunt ended, nucleotides can be added to sample DNA molecules or blunted-ended derivative therefrom, and / or phosphate moieties can be added or removed from the ends of sample DNA molecules or derivatives thereof. In some embodiments, prior to ligation, sample DNA molecules may be blunt294904-2842-5073.2Attorney Docket No. N.057.W0.01 ended, and then a single adenosine base can be added to the 3’ end. Prior to ligation the DNA may be cleaved using a restriction enzyme or some other cleavage method. During ligation the 3’ adenosine of the sample fragments and the complementary 3’ thymidine overhang of an adaptor can enhance ligation efficiency. In some embodiments, adaptor ligation is performed using a T4 DNA ligase.
[0100] In some embodiments, adaptors are appended to sample nucleic acid molecules by primer extension. For example, in some embodiments, adaptors are appended to sample nucleic acid molecules using target specific primers with tails comprising universal priming sequences. The tails may also include other useful sequences such as molecular or sample barcodes.
[0101] In some embodiments, adaptors containing one or more universal priming sequences are utilized in methods herein. In some embodiments, the adaptors are Y adaptors, for example in illustrative methods in which target region amplicons are sequenced using NGS. In some embodiments, the adaptors each comprise a universal priming site. The adaptors may or may not include methylated cytosine residues. In some embodiments, the adaptor comprises at least one methylated cytosine and at least one unmethylated cytosine. In some embodiments, all cytosines in the adaptor are methylated. In some embodiments, the adaptor comprises at least one MSRE recognition site in which all cytosines are methylated. In some embodiments, the adaptor comprises at least one MDRE recognition site in which all cytosines are unmethylated.
[0102] In some embodiments, the method further comprises amplifying the adapted nucleic acids using a primer that binds to the universal priming site and generating adapted-amplified nucleic acids before performing targeted enrichment. In some embodiments, the adapted-amplified nucleic acids further comprises a sequencing adaptor sequence or sequencing primer binding site for high-throughput sequence. In some embodiments, the adapted-amplified nucleic acids further comprise a sample barcode or index sequence, which allows multiplexed sequencing of pooled sequencing libraries (e.g., multiplexed sequencing of sequencing libraries generated from multiple samples). Thus, multiple samples can be analyzed in the same sequencing run. The sample barcode or index sequence can be used to process data according to the sample from which the data was generated.
[0103] In some embodiments, the adaptors do not comprise a molecular barcode or index sequence. In some embodiments, sequence reads generated from the high-throughput sequencing can be grouped together using the fragment-end sequence of the extracted cell-free nucleic acids304904-2842-5073.2Attorney Docket No. N.057.W0.01 or DNA derived therefrom. In some embodiments, the sequence reads that are grouped together using the molecular barcode or index sequence and / or the fragment-end sequence can be subject to error correction to correct sequencing errors and generate a consensus sequence.
[0104] In some embodiments, at least one adaptor comprises a molecular barcode or index sequence. In some embodiments, sequence reads generated from high-throughput sequencing can be grouped together using the molecular barcode or index sequence. In some embodiments, the number of adaptors having different molecular barcodes or index sequences is between 10 to 1,000, and wherein the ratio of the total number of sample nucleic acid or cfDNA molecules to the number of different molecular barcodes molecular barcodes in the ligation reaction is at least 1,000:1. The number of different molecular barcodes or index sequences in the ligation reaction, in certain embodiments, ranges from 10 to 50, 10 to 100, 50 to 200, 100 to 300, 200 to 500, 300 to 600. 500 to 700, 600 to 800 or 700 to 1,000. In some embodiments, there are at least 1, 10, 20, 30, 40, 50, or at least 100; 200, 300, 400, 500, 600, 700, 800, 900, or 1000 different molecular barcodes or index sequences in the ligation reaction. In some embodiments, the ratio of the total number of sample nucleic acid or cfDNA molecules to the number of different molecular barcodes or index sequences in the ligation reaction is at least 10,000:1. In some embodiments, the ratio of the total number of sample nucleic acid or cfDNA molecules to the number of different molecular barcodes or index sequences in the ligation reaction ranges from 50,000: 1 to 50:1, from 25,000: 1 to 100:1, from 10,000:1 to 100:1, from 10:000:1 to 8,000:1 to 500:1, from 5,000:1 to 200:1, from 10,000:1 to 50:1. In some embodiments, the methods disclosed herein result in at least 100; 200; 500; 750; 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000 different molecular barcodes molecular barcodes to each one sample nucleic acid or cfDNA molecules.
[0105] In some embodiments, at least one adaptor comprises a capture moiety. In some embodiments, the capture moiety comprises biotin, biotin dT, biotin-TEG, photocleavable (PC) biotin, desthiobiotin-TEG, or biotin azide, and the separating is performed using streptavidin beads. In some embodiments, the adaptor is a 5’ biotinylated adaptor.
[0106] In some embodiments, after a first amplification is performed on the adapted nucleic acids with biotinylated adaptors, the adapted nucleic acids are captured on streptavidin beads, and the first amplified DNA in the supernatant are not biotinylated and can be separate and subject to targeted enrichment to produce first enriched DNA. Beads containing the bound adapted nucleic acids with biotinylated adaptors are then used for a second amplification (e.g., on-bead PCR). and314904-2842-5073.2Attorney Docket No. N.057.W0.01 again the second amplified DNA in the supernatant are not biotinylated and can be separate and subject to targeted enrichment to produce second enriched DNA.
[0107] In some embodiments, at least one adaptor is chemically modified with a thiol group, and the separating can be performed using a gold substrate (e.g„ gold nanoparticles) via Au-S bonding. See Cao et al., J. Am. Chem. Soc., 123:7961-7962 (2001); Kim et al., Angew. Chem., Int. Ed. Engl., 50:9185-9190 (2011), which are incorporated herein by reference in their entirety.
[0108] In some embodiments, at least one adaptor is chemically modified with an amine group, and the separating can be performed using an aldehyde group-modified substrate via formation of a Schiff base (e.g. EDC (l-(3 -dimethyl aminopropyl)-3-ethyl carbodiimide hydrochloride)-NHS (N-hydroxy succinimide) reaction and Schiff base reaction). See Fuentes et al., Biosens. Bioelectron., 21:1574-1580 (2006); Nguyen et al. Sci. Rep., 8:337. (2018), which are incorporated herein by reference in their entirety.
[0109] In some embodiments, the capture moiety is a click chemistry capture moiety. In some embodiments, at least one adaptor is chemically modified with an azide group, and the separating can be performed using an alkyne-functionalized substrate. See McKenna et al.. Sensors and Actuators B: Chemical, 236: 286-293 (2016); El-Sagheer et al., Chem. Soc. Rev., 39:1388-1405 (2010), which are incorporated herein by reference in their entirety.Target Loci
[0110] The methods described herein comprise targeted enrichment of a panel of target loci. In some embodiments, the panel of target loci may comprise one or more somatic variants including single nucleotide polymorphism (SNPs), single nucleotide variants (SNVs). multi-nucleotide variants (MNVs), indels, CNVs, gene fusions, structural variant, or a combination thereof (i.e., alterations). In other embodiments, the panel of target loci may comprise one or more epigenetic changes, such as changes in DNA methylation. In some embodiments, the panel of target loci may include somatic or epigenetic alterations that have been previously determined to be present in cancer cells and not present in non-cancerous cells, or present in one type of cancer cells and not present in another type of cancer cells or non-cancerous cells. In some embodiments, the panel of target loci may include somatic or epigenetic alterations that have been previously identified from tumor samples. In some embodiments, the panel of target loci comprises 5-50,000 target loci, 10- 10,000 target loci, 25-5,000 target loci, 50-1,000 target loci, or 100-500 loci. In some324904-2842-5073.2Attorney Docket No. N.057.W0.01 embodiments, the panel of target loci comprises at least 5 target loci, at least 10 target loci, at least 25 target loci, at least 50 target loci, at least 100 target loci, at least 200 target loci, at least 300 target loci, at least 400 target loci, at least 1000 target loci, at least 1500, at least 2500, a least 4500, or at least 5000, at least 10.000 target loci, at least 50,000 target loci, or at least 100,000 target loci. In some embodiments, the panel of target loci comprises no more than 5000, no more than 4500, no more than 2500, no more than 1500, no more than 1000, no more than 750, no more than 500, no more than 400, no more than 300, no more than 200, no more than 100, no more than 50, or no more than 25 target loci.
[0111] In some embodiments, the panel of target loci is determined based on historical sequencing data, such as whole genome sequencing (WGS), whole exome sequencing (WES), or targeted panel sequencing data, generated from cancer patients. In some embodiments, the data is generated from population-wide SNV patterns and prior knowledge of cancer specific somatic variants or mutations. In some embodiments, the data is generated from synonymous and non- synonymous mutations from a plurality of cancer patient samples. In some embodiments, the data is generated from synonymous and non-synonymous mutations from 10 or more, from 25 or more, from 50 or more, from 100 or more, from 250 or more, from 500 or more, from 1000 or more, from 1500 or more, from 2000 or more, from 2500 or more, from 5000 or more, or from 10,000 or more cancer patient samples. In some embodiments, the data is generated based on tumor SNVs detected in a plurality of tumor samples. In some embodiments, the data is generated based on tumor SNVs detected in 10 or more, 25 or more, 50 or more, 100 or more, 250 or more, from 500 or more, 1000 or more, 1500 or more, 2000 or more, 2500 or more, 5000 or more, 10,000 or more, or 20,000 or more tumor samples. In some embodiments, the panel of target loci is determined to ensure optimal mutation coverage and representation across subpopulations, including by cancer stage, cancer type, cancer subtype, high-risk mutation carriers, and ancestry. In some embodiments, the panel of target loci comprises clustered mutations. In some embodiments, the panel of target loci comprises methylated DNA. In some embodiments, the panel of target loci comprises mutations with high prevalence in cancer patients. In some embodiments, the target loci comprise one or more unstable microsatellite loci. In some embodiments, the target loci comprise one or more cancer hotspot mutations. In some embodiments, the target loci comprise one or more COSMIC driver or tumor suppressor gene loci.334904-2842-5073.2Attorney Docket No. N.057.W0.01
[0112] In some embodiments, the panel of target loci is based on previously determined cancerspecific somatic or epigenetic alterations. In some embodiments, the previously determined cancer-specific somatic variants or mutations are ranked by prevalence and the panel is limited to the top 5 somatic or epigenetic alterations, the top 10 somatic or epigenetic alterations, the top 50 somatic or epigenetic alterations, the top 100 somatic or epigenetic alterations, the top 200 somatic or epigenetic alterations, the top 300 somatic or epigenetic alterations, the top 400 somatic or epigenetic alterations, the top 500 somatic or epigenetic alterations, the top 1000 somatic or epigenetic alterations, the top 10,000 somatic or epigenetic alterations, the top 50,000 somatic or epigenetic alterations, or top 100,000 somatic or epigenetic alterations. In some embodiments, the previously determined cancer-specific somatic or epigenetic alterations are ranked by prevalence and the panel is limited to the top 10-100,000 somatic or epigenetic alterations, the top 25-5,000 somatic or epigenetic alterations, the top 50-1.000 somatic or epigenetic alterations, or the top 100-500 somatic or epigenetic alterations.Targeted Enrichment
[0113] In some embodiments, the method described herein comprises performing a first targeted enrichment on the first amplified DNA or its derivative to enrich a plurality of target loci each encompassing at least one alteration associated with a phenotype and generating first enriched DNA. In some embodiments, the method described herein comprises performing a second targeted enrichment on the second amplified DNA or its derivative to enrich the plurality of target loci and generating second enriched DNA. In some embodiments, the method described herein comprises performing targeted enrichment on the extracted cellular DNA or their derivative to enrich a subset of the plurality of target loci and generating leukocyte-derived enriched DNA.
[0114] In some embodiments, the method described herein comprises performing a first targeted enrichment on the treated first methylation-preserved amplified DNA or its derivative to enrich a first plurality of differentially methylated regions that are differentially methylated in a phenotype and generating first enriched DNA. In some embodiments, the method described herein comprises performing a second targeted enrichment on the treated second methylation-preserved amplified DNA or its derivative to enrich the first plurality of differentially methylated regions that are differentially methylated in a phenotype and generating second enriched DNA. In some embodiments, the method described herein comprises performing a second targeted enrichment on344904-2842-5073.2Atorney Docket No. N.057.W0.01 the treated second methylation-preserved amplified DNA or its derivative to enrich a second plurality of target loci each encompassing at least one alteration associated with the phenotype and generating second enriched DNA, wherein the second plurality of target loci optionally comprise at least a subset of the first plurality of differentially methylated regions that are differentially methylated in the phenotype.
[0115] In some embodiments, performing targeted enrichment comprises performing targeted probe capture to enrich the target loci. In some embodiments, performing targeted enrichment comprises performing linked target capture with probe-dependent primers to enrich the target loci. In some embodiments, performing targeted enrichment comprises performing targeted multiplex amplification to enrich the target loci.Hybrid Capture
[0116] In some embodiments, the targeted enrichment technique can involve fragment capture by hybridization (i.e., hybrid capture). Although any hybrid capture method can be used to perform methods herein that include a targeted enrichment step, in some embodiments, a method of the present disclosure may involve using any of the hybrid capture methods disclosed herein to selectively enrich DNA. In some embodiments described herein, the targeted enrichment steps can be performed after cfDNA molecules are extracted. In some embodiments described herein, the targeted enrichment steps can be performed after appending adaptors to the extracted cfDNA molecules. In some embodiments described herein, the targeted enrichment steps can be performed after library amplification of the adapted DNA. In some embodiments described herein, the targeted enrichment steps can be performed after preparation of the methylation-preserved amplified DNA. In some embodiments described herein, the targeted enrichment steps can be performed after treatment of the methylation-preserved amplified DNA with an agent that discriminates between methylated and unmethylated cytosines.
[0117] In capture by hybridization, hybrid capture oligonucleotide probes complementary to one or both strands of specific target DNA sequences, or DNA derived therefrom in a sample, are utilized, i.e., the probes may be strand specific. In some embodiments, the probes may be complementary to the target sequence in either a 100% or 0% methylated state. Accordingly, each strand specific probe may have two designs, assuming both a completely methylated and a completely unmethylated status of each target DNA sequence, yielding a total of four unique probes per target. The specific target DNA sequence in illustrative embodiments overlaps with or 354904-2842-5073.2Attorney Docket No. N.057.W0.01 is found within a target region of a sample DNA molecule such as a cfDNA. Thus, hybrid capture probes when used in methods herein can be designed to bind to a DNA molecule that contains at least one target region or a portion thereof. In some embodiments, the hybrid capture probes can be designed to bind to a target DNA sequence within or overlapping a target region,. In other examples, the hybrid capture probes can be designed to bind to a common region that is flanking but not overlapping the target region and that can be a common region that was added to some, most, almost all or all of the DNA in a sample, or added to all amplicons using a common sequence on at least one primer of a primer pair. In illustrative embodiments, a hybrid capture probe or set thereof, are designed to bind to a target DNA sequence within target region, or set of target regions, respectively.
[0118] Hybrid capture probes may be added to a prepared sample and hybridized through a denature -reannealing process to form duplexes of exogenous-endogenous fragments (e.g„ hybrid capture probes bound to sample DNA molecules, or DNA derived therefrom). These duplexes may then be physically separated from the sample by various means. In some embodiments, once the hybrid capture probes are removed, the sample DNA molecules, or DNA derived therefrom can be amplified. Some ways to physically remove the hybrid capture probes are by covalently bonding the hybrid capture probes to a solid support, for example a magnetic bead, or a chip. Another way to physically remove the hybrid capture probes is by covalently bonding them to a molecular moiety with a strong affinity for another molecular moiety. An example of such a molecular pair is biotin and streptavidin, such as is used in SURESELECT (Agilent). Thus, hybrid capture probes, for example that bind to a target DNA sequence within or overlapping a target region of a DNA molecule obtained or derived from a sample, can be covalently attached to a biotin molecule, and after hybridization with sample DNA or DNA derived therefrom, a solid support with streptavidin affixed can be used to pull down the biotinylated hybrid capture probes, which are hybridized to DNA molecules obtained or derived from a sample that include a target region that includes the target DNA sequence recognized by the hybrid capture probes. Thus, in some embodiments, the hybrid capture probes are immobilized, directly or indirectly to a solid support. In some embodiments, the hybrid capture probes include a binding partner, for example biotin.
[0119] In some embodiments of any of the aspects herein, the hybrid capture probes can be a part of a set of at least two hybrid capture probes. In some embodiments, the set includes at least one364904-2842-5073.2Atorney Docket No. N.057.W0.01 hybrid capture probe for each target region. In some embodiments, the set includes two or more hybrid capture probes for each target region. In some embodiments, the set includes three or more hybrid capture probes for each target region. In some embodiments, the set includes four or more hybrid capture probes for each target region.
[0120] In some embodiments of any of the aspects herein, the hybrid capture probes can have a length in the range of 30 bases to 170 bases, 30 bases to 160 bases, 30 bases to 150 bases, 30 bases to 140 bases, 30 bases to 130 bases, 30 bases to 120 bases, 30 bases to 110 bases, 30 bases to 100 bases, 30 bases to 90 bases, 30 bases to 80 bases, 30 bases to 70 bases, 30 bases to 60 bases, 30 bases to 50 bases, 40 bases to 160 bases, 40 bases to 150 bases, 40 bases to 140 bases, 40 bases to 130 bases, 40 bases to 120 bases, 40 bases to 110 bases, 40 bases to 100 bases, 40 bases to 90 bases, 40 bases to 80 bases, 40 bases to 70 bases, 40 bases to 60 bases, 50 bases to 150 bases, 50 bases to 140 bases. 50 bases to 130 bases, 50 bases to 120 bases. 50 bases to 110 bases, 50 bases to 100 bases, 50 bases to 90 bases, 50 bases to 80 bases, 50 bases to 70 bases, 60 bases to 140 bases, 60 bases to 130 bases, 60 bases to 120 bases, 60 bases to 110 bases, 60 bases to 100 bases, 60 bases to 90 bases. 60 bases to 80 bases. 70 bases to 130 bases, 70 bases to 120 bases. 70 bases to 110 bases, 70 bases to 100 bases, 70 bases to 90 bases, 80 bases to 120 bases, 80 bases to 110 bases, 80 bases to 100 bases, 90 bases to 120 bases, 90 bases to 110 bases, 100 bases to 165 bases, 100 bases to 150 bases, 100 bases to 140 bases, 100 bases to 130 bases, 100 bases to 120 bases, 110 bases to 150 bases, 110 bases to 140 bases, 110 bases to 130 bases, 120 bases to 150 bases, or 130 bases to 160 bases.Linked Target Capture
[0121] In some embodiments, the targeted enrichment technique can involve probe-dependent primers. Probe-dependent primers (PDPs) have been disclosed (Pel, et al. “Rapid and highly- specific generation of targeted DNA sequencing libraries enabled by linking capture probes with universal primers” PLoS ONE 13(12):e0208283 (2018); WO2017 / 168332A1 “Linked duplex target capture”; W02020 / 039261 “Linked target capture and ligation”, which are hereby incorporated by reference in their entirety). Such embodiments can be considered linked target capture (LTC) methods. Briefly, in an LTC method, a target-specific probe is linked to a universal primer. The target specific probe is designed to hybridize to a target of interest such as one of the target variant loci. The universal primer linked to the probe is designed to hybridize to a universal priming site in the adaptors that have been ligated to the cell-free DNA. The binding of the probe 374904-2842-5073.2Atorney Docket No. N.057.W0.01 to the target brings the linked universal primer into proximity with the universal priming site and in fact, the ability of the universal primer to bind to the universal priming site and be extended depends on the probe binding to its target. Results to-date have shown that the universal primers do not hybridize to adaptors attached to fragments that do not include the target. The linked target capture is highly target specific and primer extension depends on successful probe binding. For that reason, the universal primers linked to the probes are "probe-dependent primers". LTC may be performed using only one PDP, e.g„ for a linear amplification, but preferably uses paired forward and reverse PDPs (as shown in Fig. 1(b) of Pel 2018 as target-capture PCR1) to amplify the fragment exponentially. The bound probe does not interfere with primer extension to copy the entire fragment (including the distal adaptor) when a strand-displacing polymerase is used. It is noted that the cell-free DNA fragment is copied by primers extended from within the ligated adaptors, so the entirety of the fragment is copied into amplicons. Due to the probe, LTC gives the target specificity of conventional PCR with gene-specific primers while due to the priming sites in the adaptors, the entire fragment is amplified. The resulting amplification products will include a copy of the entire target-containing fragment with adaptors at both ends. Preferably, PDPs are designed to incorporate non-extendable capture probes linked 5’ to 5’ with a primer. Multiple linker types are possible as discussed below. Typically, probes of PDPs can be between 30 to 70 nucleotides in length, and include or comprise a 3’ inverted dT base or 3’ C3 spacer to inhibit polymerase extension. In some embodiments, probes are designed to cover the desired region with overlap between forward and reverse probes. In some embodiments, the probes are between 20 and 100 nucleotides in length. In some embodiments, the size of the probe can be between 20 and 40 nucleotides, between 30 and 50 nucleotides, between 40 and 60, between 50 and 70 between 60 and 80, between 70 and 90, 80 and 100, 90 and 110, 100 and 120 nucleotides in length. In some embodiments, at least one of the primers of a PDP pair comprises a sample index.
[0122] In PDPs, forward and reverse probes can be designed to bind to nucleic acid sequences within or near a genomic region of interest on a sample DNA molecule to enrich nucleic acid molecules comprising the genomic region of interest or copies thereof. In some embodiments, at least one of the probe binding regions can include one or more somatic variants or mutations. In some embodiments, one of the probe binding regions can include one or more somatic variants or mutations. In some embodiments, both of the probe binding sites of the probe binding regions can include one or more somatic variants or mutations. In some embodiments, neither of the probe384904-2842-5073.2Attorney Docket No. N.057.W0.01 binding sites of the probe binding region comprises a somatic mutation. In some embodiments, at least one of the probe binding regions can include one or more CpG sites. In some embodiments, both of the probe binding sites of the probe binding regions can include one or more CpG sites.
[0123] In some embodiments, the primer portion of a PDP is a universal primer designed to bind to a universal primer site on the ligated adaptor. In some embodiments, the PDP is designed with a sequencer binding sequence, such as an Illumina flow cell binding sequence, incorporated therein. In some embodiments, the sequencer flow cell binding sequence is between the probe and universal primer, and adjacent to the primer. Linked primers of the invention may also include sequencing tags to ensure that all cluster reads originate from the same linked template molecule. The lengths of the primers can be extended or shortened at the 5' end or the 3' end to produce primers with desired melting temperatures. Also, the annealing position of each primer pair can be designed such that the sequence and length of the primer pairs yield the desired melting temperature. In illustrative embodiments, the primer is a low melting temperature universal primer complementary to a portion of the ligated adaptor.
[0124] The primer can be tailed or untailed on either side of the forward and reverse side depending on the specific requirements. In some embodiments, the universal primer comprises an A tail. In some embodiments, the universal primer is blunt ended. The length of the primers of the PDP can range from 5 to 40 nucleotides in length. In certain embodiments, the PDP primers are between 10 and 25 nucleotides long. In embodiments, the primers of the PDP can range from 5 to 15 nucleotides, from 10 to 25 nucleotides, from 15 to 35 nucleotides, or from 25 to 40 nucleotides in length.
[0125] In some embodiments, probe dependent primers comprise a linker between the probe and the primer. In some embodiments, probe and primer portions of the PDP are linked by a polyethylene glycol derivative, an oligosaccharide, a lipid, a hydrocarbon, a polymer, or a protein. In some embodiments, the linker is a PEG molecule, or derivative thereof. In some embodiments, the linker is an oligosaccharide. In some embodiments, the linker is a lipid. In some embodiments, the linker is a hydrocarbon. In some embodiments, the linker is a polymer. In some embodiments, the linker is a protein, or portion thereof. Linkers based on click chemistry is described in WO2017 / 168332A1, which is incorporated herein by reference in its entirety.394904-2842-5073.2Attorney Docket No. N.057.W0.01Targeted Multiplex Amplifications
[0126] In some embodiments, the targeted enrichment technique can involve targeted multiplex amplification (e.g., PCR or isothermal amplification). Methods in some aspects herein include performing one or, in some embodiments, two or more amplifications. Such amplifications in certain illustrative embodiments include at least one targeted amplification wherein at least one primer and in certain embodiments both primers of a primer pair, one or more primer pairs, or a set of primer pairs used for the amplification are each designed to bind to a specific nucleic acid sequence at or near a genomic region of interest (i.e., are target-specific primers) to generate target region amplicons. In some embodiments, methods herein include one or more universal amplifications. Exemplary target enrichment protocols based on targeted multiplex amplification are provided in Zimmermann et al., Prenat. Diagn. 32:1233-1241 (2012); and Sigdel et al., J. Clin. Med. 8(1): 19 (2019), each of which is incorporated herein by reference in its entirety.
[0127] A number of amplification technologies can be used with methods herein. For example, such amplification can be an isothermal amplification (e.g., recombinase polymerase amplification (RPA), a ligase-based amplification, PCR. or a combination thereof (e.g.. ligation-mediated PCR). In some illustrative embodiments, the targeted amplification is a targeted PCR(s) that is performed using a PCR reaction mixture that includes one primer pair, or in illustrative embodiments a set of primer pairs, and at least a portion of the library of DNA molecules comprising the extracted cfDNA or DNA derived therefrom (e.g., adapted DNA, adapted-amplified DNA, methylation- preserved amplified DNA, or treated methylation-preserved amplified DNA).
[0128] In some illustrative embodiments, the targeted enrichment is a targeted PCR(s) that is performed using a PCR reaction mixture that includes one primer pair, or in illustrative embodiments a set of primer pairs, and DNA molecules (e.g., cfDNA or cellular DNA), or universally amplified DNA molecules generated therefrom.
[0129] In some embodiments, at least one primer of a primer pair used for targeted amplification is a target-specific primer designed to bind to a specific nucleic acid sequence at or near a genomic region of interest, which in illustrative examples can be genomic regions where one or more somatic variants or mutations or epigenetic changes (i.e., alterations) are associated with or indicative of formation or presence of a phenotype, such as cancer, and can, for example, include mutations in promoter regions of tumor suppressor genes or changes in DNA methylation associated with or indicative of formation or presence of cancer. A target-specific primer can be404904-2842-5073.2Attorney Docket No. N.057.W0.01 designed to bind to any sequence at or near a target region for amplification of the target region encompassing one or more cancer-specific variants. In some embodiments, a target- specific primer can be designed to bind to a sequence that includes one or more cancer- specific variant loci. In other embodiments, a target- specific primer can be designed to bind to a sequence that is upstream or downstream to one or more SNV or indel loci. In some embodiments, a target- specific primer can be designed to bind to a sequence that includes a CpG site. In other embodiments, a target- specific primer can be designed to bind to a sequence that is upstream or downstream to one or more CpG sites.
[0130] In some methods herein, a universal amplification(s) can be performed before the targeted amplification(s). In some embodiments, a universal amplification(s) can be performed after the targeted amplification(s). A universal amplification can be performed for example using a universal primer pair. In some embodiments, the universal primer pair binds primer binding sites in an adaptor added by ligation or during a previous amplification reaction, step, or cycle, and in some embodiments, during a previous targeted amplification reaction, step, or cycle. In some embodiments, the methods herein include performing a universal PCR using the plurality of adapted methylation-transferred DNA molecules, and a universal PCR primer pair comprising primers designed to bind universal primer binding sequences on the adaptors, before performing one or more targeted PCRs.
[0131] In some embodiments, at least one of the primer pairs comprises a universal primer and a target- specific primer. In some embodiments, at least one of the primer pairs comprises two targetspecific primers. In some embodiments, at least one of the primers comprises a sequencing tag. In some embodiments, at least one of the primers comprises a sample index. In some embodiments, performing a PCR further comprises using primers comprising a sequencing primer binding site. In some embodiments, performing a PCR further comprises using primers comprising a sample index. In some embodiments, the primers of the primer pairs are probe-dependent primers and the amplification is a target capture polymerase chain reaction.
[0132] In some embodiments, the target regions each comprises one or more CpG sites differentially methylated in one or more cancers, although the nucleic acids comprising the target region may have been treated with an agent that discriminates between methylated and unmethylated cytosines (e.g., conversion of unmethylated C to U). In some embodiments, the set414904-2842-5073.2Attorney Docket No. N.057.W0.01 of loci comprises two or more loci that are differentially methylated in a different type of cancer from each other.
[0133] Methods as described herein, in some embodiments, can include multiple amplification cycles (e.g. multiple PCR temperature cycles), and in some embodiments can include several sequential PCR reactions performed during the same set of temperature cycles. In some embodiments, amplification cycles can include at least 2, 3, 4, 5, 6, 7, 8. 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 cycles. In some embodiments, amplification cycles can include at least 7, 8, 9, or 10 cycles. In illustrative embodiments, amplification cycles can include at least 11, 12, 13, 14, 15, 16, or 17 cycles.
[0134] In some embodiments described herein, PCR amplification is performed by adding a PCR reaction mixture to the DNA template (e.g., cfDNA or cellular DNA) followed by addition of a polymerase enzyme, and then amplified through multiple amplification cycles. In some embodiments, the PCR reaction mixture contains one or more primer pairs, deoxynucleotides (dNTPs), PCR reaction buffer, and deionized water. In some embodiments, the dNTPs can comprise a mixture of dATP, dCTP, dGTP and dTTP. In some embodiments, the final concentration of each dNTP in the reaction mixture can range from 0.05 mM to 0.5 mM dNTPs, for example, from 0.05 mM to 0.5 mM, 0.05 mM to 0.1 mM, 0.05 mM to 0.15 mM, 0.05 to 0.2 mM, 0.05 mM to 0.25 mM, 0.05 mM to 0.3 mM, 0.05 mM to 0.35 mM, 0.05 to 0.4 mM 0.05 mM to 0.45 mM, from .1 mM to 0.5 mM, 0.15 mM to 0.5 mM, or 0.2 mM to 0.5 mM, 0.25 to 0.5 mM, 0.3 mM to 0.5 mM, 0.35 mM to 0.5 mM, 0.4 mM 0.5 mM, or from 0.45 mM to 0.5 mM. In illustrative embodiments, the final concentration of each dNTP in the reaction mixture ranges from 0.15 mM to 0.25 mM. In some embodiments, the final concentration of each dNTP in the reaction mixture is 0.2 mM.
[0135] PCR buffer solution creates a suitable environment for the polymerase chain reaction and can contain many different components, including magnesium chloride (MgC12), potassium chloride (KC1), dimethyl sulfoxide (DMSO), and glycerin or bovine serum albumin (BSA). In some embodiments, the concentration of KC1 can be between 25 and 50 mM, between 25 and 75 mM, between 25 and 100 mM, between 30 and 100 mM, between 50 and 100 mM, or between 70 and 100 mM. In some embodiments, the MgC12 concentration can be in the range of 0.5 mM to 5 mM, between 0.5 mM to 4.5 mM, 0.5 to 4.0 mM, 0.5 to 3.5 mM, 0.5 to 3.0 mM, 0.5 to 2.5 mM, 0.5 to 2.0 mM, 0.5 to 1.5 mM, 1.0 to 5 mM, 1.5 to 5mM, 2.0 to 5 mM, 2.5 to 5 mM, 3.0 to 5 mM,424904-2842-5073.2Attorney Docket No. N.057.W0.013.5 to 5 mM, 4.0 to 5 mM, or 4.5 to 5 mM. In some embodiments, the concentration of MgC12 is 2.0 mM.
[0136] In some embodiments, the buffer solution is a Q5® Reaction Buffer (B9027S, New England Biolabs. Inc.). In some embodiments, the reaction buffer is Standard Taq Reaction Buffer (B9014S, New England Biolabs, Inc). In some embodiments, the reaction buffer is a Standard Taq (Mg-free) Reaction Buffer (B9015S, New England Biolabs, Inc.).
[0137] In some embodiments, a DNA polymerase is used to produce DNA amplicons using DNA as a template. In some embodiments, the polymerase is a Q5® DNA Polymerase, such as Q5® High-Fidelity DNA Polymerase (M0491S, New England BioLabs. Inc.) or Q5® Hot Start High- Fidelity DNA Polymerase (M0493S, New England BioLabs, Inc.). Q5® High-Fidelity DNA polymerase is a high-fidelity, thermostable, DNA polymerase with 3" > 5" exonuclease activity, fused to a processivity-enhancing Sso7d domain. Q5® High-Fidelity DNA polymerase lacks 5"— > 3 "exonuclease activity and strand displacement activity.
[0138] In some embodiments, the polymerase is a T4 DNA polymerase (M0203S, New England BioLabs, Inc.). T4 DNA Polymerase catalyzes the synthesis of DNA in the 5'— ► 3' direction and requires the presence of template and primer. This enzyme has a 3'— > 5' exonuclease activity which is much more active than that found in DNA Polymerase I. T4 DNA polymerase lacks 5'— 3' exonuclease activity and strand displacement activity.
[0139] In some embodiments of any of the aspects herein, the length of the primers can be between 10 to 100 nucleotides, such as between 10 to 75 nucleotides, 10 to 40 nucleotides, 10 to 35 nucleotides, 10 to 30 nucleotides, 10 to 20 nucleotides, 15 to 100 nucleotides, 20 to 100 nucleotides, from 25 to 100 nucleotides, from 30 to 100 nucleotides from 35 to 100 nucleotides, from 40 to 100 nucleotides, from 45 to 100 nucleotides, from 50 to 100 nucleotides, from 55 to 100 nucleotides, from 60 to 100 nucleotides, from 65 to 100 nucleotides, from 70 to 100 nucleotides, or from 75 to 100 nucleotides. In some embodiments, the range of the length of the primers is between 5 to 50 nucleotides, such as 5 to 40 nucleotides, 5 to 20 nucleotides, or 5 to 10 nucleotides. In some embodiments, the primers are between 5 and 50 bp in length, between 10 and 40 bp in length, between 15 and 30 bp in length, between 15 and 25 bp in length, between 20 and 40 bp in length, between 25 and 50 bp in length, or between 30 and 50 bp in length. In some embodiments, the primers are between 25 and 100 bp in length, between 35 and 100 bp in length,434904-2842-5073.2Attorney Docket No. N.057.W0.01 between 45 and 100 bp in length, between 55 and 100 bp in length, between 65 and 100 bp in length, or between 75 and 100 bp in length.
[0140] In some embodiments of any of the aspects or embodiments herein, the number of primer pairs can range from 1 to 100.000 primer pairs that each bind to one or more primer binding sequences. In some embodiments, the primer pairs are a part of a set of primer pairs. In some embodiments, the set of primers range from 2 to 100,000, from 2 to 10,000, from 2 to 1,000, from 2 to 100, from 2 to 50, from 10 to 100, from 50 to 100, from 100 to 200, from 100 to 500, from 100 to 1,000, from 100 to 10,000, from 100 to 100,000, from 1,000 to 100,00, or from 10,000 to 100,000 primer pairs. In some embodiments, the number of primer pairs can range from 10 to 10,000, 10 to 1,000, 10 to 100, 10 to 50, 10 to 40, 10 to 30, 15 to 30, or 15 to 25 primer pairs.
[0141] In some embodiments, PCR is used to generate very short amplicons. cfDNA (such as fetal cfDNA in maternal serum or necroptically- or apoptotically-released cancer cfDNA) is highly fragmented. For fetal cfDNA, the fragment sizes are distributed in approximately a Gaussian fashion with a mean of 160 bp, a standard deviation of 15 bp, a minimum size of about 100 bp, and a maximum size of about 220 bp. Because cfDNA fragments are short, the likelihood of both primer sites being present the likelihood of a fragment of length L comprising both the forward and reverse primers sites is the ratio of the length of the amplicon to the length of the fragment. Under ideal conditions, assays in which the amplicon is 45, 50, 55, 60, 65, or 70 bp will successfully amplify from 72%, 69%, 66%, 63%, 59%, or 56%, respectively, of available template fragment molecules. Thus, in some embodiments target amplicons generated in method herein are between 40 and 100, 40 and 75, or 45 and 70 bp in length. In certain embodiments that relate most preferably to cfDNA from samples of individuals suspected of having cancer, the cfDNA is amplified using primers that yield a maximum amplicon length of 85, 80, 75 or 70 bp, and in certain preferred embodiments 75 bp, and that have a melting temperature between 50 and 65°C, and in certain preferred embodiments, between 54-60.5°C. The amplicon length is the distance between the 5 -prime ends of the forward and reverse priming sites. Amplicon length that is shorter than typically used by those known in the art may result in more efficient measurements of the desired methylation sites by only requiring short sequence reads. In an embodiment, a substantial fraction of the amplicons are between 25 on the low end of the range, and 100 bp, 90 bp, 80 bp, 70 bp, 65 bp, 60 bp, 55 bp, 50 bp, or 45 bp on the high end of the range.444904-2842-5073.2Attorney Docket No. N.057.W0.01
[0142] In some embodiments, the PCR is multiplexed. In any of the methods for detecting somatic variants or mutations herein, improved amplification parameters for multiplex PCR can be employed. For example, wherein the amplification reaction is a PCR reaction and the annealing temperature is between 1, 2, 3, 4, 5, 6, 7. 8, 9, or 10°C greater than the melting temperature on the low end of the range, and 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15°C on the high end the range for at least 10, 20, 25, 30, 40, 50, 06, 70, 75, 80, 90, 95 or 100% the primers of the set of primers.
[0143] In certain embodiments, wherein the amplification reaction is a multiplex PCR reaction the length of the annealing step in the PCR reaction is between 10, 15, 20, 30, 45, and 60 minutes on the low end of the range, and 15, 20, 30, 45, 60, 120, 180, or 240 minutes on the high end of the range. In certain embodiments, the primer concentration in the amplification, such as the multiplex PCR reaction is between 1 and 10 nM. Furthermore, in exemplary embodiments, the primers in the set of primers, are designed to minimize primer dimer formation.
[0144] Multiplex PCR may involve a single round of PCR in which all targets are amplified or it may involve one round of PCR followed by one or more rounds of nested PCR or some variant of nested PCR. Nested PCR consists of a subsequent round or rounds of PCR amplification using one or more new primers that bind internally, by at least one base pair, to the primers used in a previous round. Nested PCR reduces the number of spurious amplification targets by amplifying, in subsequent reactions, only those amplification products from the previous one that have the correct internal sequence. Reducing spurious amplification targets improves the number of useful measurements that can be obtained, especially in sequencing. Nested PCR typically entails designing primers completely internal to the previous primer binding sites, necessarily increasing the minimum DNA segment size required for amplification. For samples such as cancer patient plasma cfDNA, in which the DNA is highly fragmented, the larger assay size reduces the number of distinct cfDNA molecules from which a measurement can be obtained. In an embodiment, to offset this effect, one may use a partial nesting approach where one or both of the second round primers overlap the first binding sites extending internally some number of bases to achieve additional specificity while minimally increasing in the total assay size.
[0145] In some embodiments, a multiplex pool of PCR assays are designed to amplify potentially heterozygous SNVs or other polymorphic or non-polymorphic loci on one or more chromosomes and these assays are used in a single reaction to amplify DNA. The number of PCR assays may be more than 10, more than 25, more than 50. more than 100, more than 200, more than 300, more454904-2842-5073.2Atorney Docket No. N.057.W0.01 than 400, more than 500, or more than 1000 PCR assays in a single reaction. The SNV frequencies of each locus may be determined by clonal or some other method of sequencing of the amplicons. Note that this method is equally well applicable to detecting translocations, deletions, duplications, and other chromosomal abnormalities.
[0146] In some embodiments, tails with no homology to the target genome may also be added to the 3-prime or 5-prime end of any of the primers. These tails facilitate subsequent manipulations, procedures, or measurements. In an embodiment, the tail sequence can be the same for the forward and reverse target specific primers. In an embodiment, different tails may be used for the forward and reverse target specific primers. In an embodiment, a plurality of different tails may be used for different loci or sets of loci. Certain tails may be shared among all loci or among subsets of loci. For example, using forward and reverse tails corresponding to forward and reverse sequences required by any of the current sequencing platforms can enable direct sequencing following amplification. In an illustrative embodiment, the tails can be used as common priming sites among all amplified targets that can be used to add other useful sequences. In some embodiments, the inner primers may contain a region that is designed to hybridize either upstream or downstream of the targeted polymorphic locus. In some embodiments, the primers may contain a molecular barcode. In some embodiments, the primer may contain a universal priming sequence designed to allow PCR amplification.
[0147] In an illustrative embodiment, a PCR assay pool is created such that forward and reverse primers have tails corresponding to the required forward and reverse sequences required by a high throughput sequencing instrument such as the HIS EQ, GAIIX, or MYSEQ available from ILLUMINA. In addition, included 5-prime to the sequencing tails is an additional sequence that can be used as a priming site in a subsequent PCR to add nucleotide barcode sequences to the amplicons, enabling multiplex sequencing of multiple samples in a single lane of the high throughput sequencing instrument.
[0148] Multiplexed PCR can often result in the production of a very high proportion of product DNA that results from unproductive side reactions such as primer dimer formation. In an embodiment, the particular primers that are most likely to cause unproductive side reactions may be removed from the primer library to give a primer library that will result in a greater proportion of amplified DNA that maps to the genome. The step of removing problematic primers, that is, those primers that are particularly likely to form dimers has unexpectedly enabled extremely high464904-2842-5073.2Attorney Docket No. N.057.W0.01PCR multiplexing levels for subsequent analysis by sequencing. In systems such as sequencing, where performance significantly degrades by primer dimers and / or other mischief products, greater than 10, greater than 50, and greater than 100 times higher multiplexing than other described multiplexing has been achieved. Note this is opposed to probe-based detection methods, e.g., microarrays, TAQMAN, PCR etc. where an excess of primer dimers will not affect the outcome appreciably.
[0149] There are a number of ways to choose primers for a library where the amount of nonmapping primer-dimer or other primer mischief products are minimized. Empirical data indicate that a small number of ‘bad’ primers are responsible for a large amount of non-mapping primer dimer side reactions. Removing these ‘bad’ primers can increase the percent of sequence reads that map to targeted loci. One way to identify the ‘bad’ primers is to look at the sequencing data of DNA that was amplified by targeted amplification; those primer dimers that are seen with greatest frequency can be removed to give a primer library that is significantly less likely to result in side product DNA that does not map to the genome. There are also publicly available programs that can calculate the binding energy of various primer combinations, and removing those with the highest binding energy will also give a primer library that is significantly less likely to result in side product DNA that does not map to the genome.
[0150] When primers can be designed it is possible to attempt to identify primer pairs likely to form spurious products by evaluating the likelihood of spurious primer duplex formation between all possible pairs of primers using published thermodynamic parameters for DNA duplex formation. Primer interactions may be ranked by a scoring function related to the interaction and primers with the worst interaction scores are eliminated until the number of primers desired is met. In cases where SNVs / SNPs likely to be heterozygous are most useful, it is possible to also rank the list of assays and select the most heterozygous compatible assays. Experiments have validated that primers with high interaction scores are most likely to form primer dimers.
[0151] Note that there are other methods for determining which PCR probes are likely to form dimers. In an embodiment, analysis of a pool of DNA that has been amplified using a nonoptimized set of primers may be sufficient to determine problematic primers. For example, analysis may be done using sequencing, and those dimers which are present in the greatest number are determined to be those most likely to form dimers and may be removed.474904-2842-5073.2Attorney Docket No. N.057.W0.01
[0152] The use of tags on the primers may reduce amplification and sequencing of primer dimer products. Tag-primers can be used to shorten necessary target- specific sequences to below 20, below 15, below 12, and even below 10 base pairs. This can be serendipitous with standard primer design when the target sequence is fragmented within the primer binding site or, or it can be designed into the primer design. Advantages of this method include at least the following: it increases the number of assays that can be designed for a certain maximal amplicon length, and it shortens the “non-informative” sequencing of primer sequence. It may also be used in combination with internal tagging (see elsewhere in this document).
[0153] In an embodiment, the relative amount of nonproductive products in the multiplexed targeted PCR amplification can be reduced by raising the annealing temperature, lowering primer concentrations and using longer annealing times. For example, methods of designing and optimizing primers can be found in U.S. Patent No. 11,312,996, incorporated herein by reference.
[0154] To select target locations, one may start with a pool of candidate primer pair designs and create a thermodynamic model of potentially adverse interactions between primer pairs, and then use the model to eliminate designs that are incompatible with other designs in the pool.
[0155] One issue with fragmented DNA is that since it is short in length, the chance that a variant is close to the end of a DNA strand is higher than for a long strand. Since PCR capture of a variant requires a primer binding site of suitable length on both sides of the variant, a significant number of strands of DNA with the targeted variant will be missed due to insufficient overlap between the primer and the targeted binding site. In cases where the binding region is shorter than the 18 bp typically required for hybridization, the region (cr) on the primer that is complementary to the library tag is able to increase the binding energy to a point where the PCR can proceed. Note that any specificity that is lost due to a shorter binding region can be made up for by other PCR primers with suitably long target binding regions. Note that this embodiment can be used in combination with direct PCR, or any of the other methods described herein, such as nested PCR, semi nested PCR, hemi nested PCR, one sided nested or semi or hemi nested PCR, or other PCR protocols.
[0156] In some embodiments, ligation mediated universal-PCR amplification of fragmented DNA may be used. The ligation mediated universal-PCR amplification can be used to amplify plasma- derrived DNA, which can then be divided into multiple parallel reactions. It may also be used to preferentially amplify short fragments, thereby enriching fetal fraction. In some embodiments the484904-2842-5073.2Attorney Docket No. N.057.W0.01 addition of tags to the fragments by ligation can enable detection of shorter fragments, use of shorter target sequence specific portions of the primers and / or annealing at higher temperatures which reduces unspecific reactions.RNA Amplification, Quantification, and Analysis
[0157] Any of the following exemplary methods may be used to amplify and optionally quantify RNA, such as such as cfRNA, cellular RNA, cytoplasmic RNA, coding cytoplasmic RNA, noncoding cytoplasmic RNA, mRNA, miRNA, mitochondrial RNA, rRNA, or tRNA. In some embodiments, the miRNA is any of the miRNA molecules listed in the miRBase database available at the world wide web at mirbase.org, which is hereby incorporated by reference in its entirety. Exemplary miRNA molecules include miR-509; miR-21, and miR-146a.
[0158] In some embodiments, reverse-transcriptase multiplex ligation-dependent probe amplification (RT-MLPA) is used to amplify RNA. In some embodiments, each set of hybridizing probes consists of two short synthetic oligonucleotides spanning the SNP and one long oligonucleotide (Li et al., Arch Gynecol Obstet. “Development of noninvasive prenatal diagnosis of trisomy 21 by RT-MLPA with a new set of SNP markers,” Jul. 5, 2013, DOI 10.1007 / s00404- 013-2926-5; Schouten et al. “Relative quantification of 40 nucleic acid sequences by multiplex ligation-dependent probe amplification.” Nucleic Acids Res 30:e57, 2002; Deng et al. (2011) “Non-invasive prenatal diagnosis of trisomy 21 by reverse transcriptase multiplex ligationdependent probe amplification,” Clin, Chem. Lab Med. 49:641-646, 2011 , which are each hereby incorporated by reference in its entirety).
[0159] In some embodiments, RNA is amplified with reverse-transcriptase PCR. In some embodiments, RNA is amplified with real-time reverse-transcriptase PCR, such as one-step realtime reverse-transcriptase PCR with SYBR GREEN I as previously described (Li et al., Arch Gynecol Obstet. “Development of noninvasive prenatal diagnosis of trisomy 21 by RT-MLPA with a new set of SNP markers,” Jul. 5, 2013, DOI 10.1007 / s00404-013-2926-5; Lo et al., “Plasma placental RNA allelic ratio permits noninvasive prenatal chromosomal aneuploidy detection,” Nat Med 13:218-223, 2007; Tsui et al„ Systematic micro-array based identification of placental mRNA in maternal plasma: towards non-invasive prenatal gene expression profiling. J Med Genet 41:461-467, 2004; Gu et al., J. Neurochem. 122:641-649, 2012, which are each hereby incorporated by reference in its entirety).494904-2842-5073.2Attorney Docket No. N.057.W0.01
[0160] In some embodiments, a microarray is used to detect RNA. For example, a human miRNA microarray from Agilent Technologies can be used according to the manufacturer's protocol. Briefly, isolated RNA is dephosphorylated and ligated with pCp-Cy3. Labeled RNA is purified and hybridized to miRNA arrays containing probes for human mature miRNAs on the basis of Sanger miRBase release 14.0. The arrays is washed and scanned with use of a microarray scanner (G2565BA, Agilent Technologies). The intensity of each hybridization signal is evaluated by Agilent extraction software v9.5.3. The labeling, hybridization, and scanning may be performed according to the protocols in the Agilent miRNA microarray system (Gu et al., J. Neurochem. 122:641-649, 2012, which is hereby incorporated by reference in its entirety).
[0161] In some embodiments, a TaqMan assay is used to detect RNA. An exemplary assay is the TaqMan Array Human MicroRNA Panel vl.O (Early Access) (Applied Biosystems), which contains 157 TaqMan MicroRNA Assays, including the respective reverse-transcription primers, PCR primers, and TaqMan probe (Chim et al., “Detection and characterization of placental microRNAs in maternal plasma,” Clin Chem. 54(3):482-90, 2008, which is hereby incorporated by reference in its entirety).
[0162] If desired, the mRNA splicing pattern of one or more mRNAs can be determined using standard methods (Fackenthall and Godley, Disease Models & Mechanisms 1: 37-42, 2008, doi:10.1242 / dmm.000331, which is hereby incorporated by reference in its entirety). For example, high-density microarrays and / or high-throughput DNA sequencing can be used to detect mRNA splice variants.
[0163] In some embodiments, whole transcriptome shotgun sequencing or an array is used to measure the transcriptome.Methylation-Preserving Library PreparationTemplate Copying
[0164] In some embodiments, the method described herein comprises preparing a first library of methylation-preserved amplified DNA having one or more methylation-transferred DNA molecules, wherein said first library of methylation-preserved amplified DNA is prepared by (i) performing a first copying of one or more template DNA molecules in the adapted DNA and generating one or more copied DNA molecules, (ii) transferring methylation status from the one or more template DNA molecules to the one or more copied DNA molecules and generating one or504904-2842-5073.2Atorney Docket No. N.057.W0.01 more methylation-transferred copied DNA molecules, (iii) optionally repeating the copying and the transferring one or more rounds, and (iv) separating the methylation-transferred copied DNA molecules from the adapted DNA. In some embodiments, the method described herein comprises preparing a second library of methylation-preserved amplified DNA having one or more methylation-transferred DNA molecules, wherein said second library of methylation-preserved amplified DNA is prepared by (i) performing a second copying of one or more template DNA molecules in the adapted DNA and generating one or more copied DNA molecules, (ii) transferring methylation status from the one or more template DNA molecules to the one or more copied DNA molecules and generating one or more methylation-transferred copied DNA molecules, (iii) optionally repeating the copying and the transferring one or more rounds, and (iv) separating the methylation-transferred copied DNA molecules from the adapted DNA.
[0165] Methods described herein include copying of the template DNA molecules for example by primer extension. In some embodiments, adaptors may be appended to the 5’ and / or 3’ ends of the template DNA molecule, and primers specific for an adaptor sequence, such as a universal priming sequence, hybridize to the adaptor sequence and are extended using a polymerase enzyme, for example, to generate a copied DNA molecule.
[0166] In some embodiments of the methods described herein, primer extension is carried out under isothermal conditions. In some embodiments, thermal cycling is employed. In some embodiments, primer extension is performed at between about 37 °C and 64 °C, between about 40 °C and 64 °C. between about 50 °C and 64 °C, or between about 58 °C and 64 °C. In some embodiments, primer extension is performed at about 60 °C. Methods that can be carried out under such isothermal conditions include adaptor priming, strand invasion, and enhanced strand invasion, rolling circle amplification, and variations thereof. i) Adaptor priming
[0167] In some embodiments, adaptors comprising primer binding sequences may be appended to the ends of the template DNA molecules. In some embodiments, Y-adaptors having unique tail sequences and comprising primer binding sequences may be appended to the ends of the template DNA molecules. In some embodiments, primers hybridize to a single stranded “Y” portion of the adaptor. This allows for single strand extension and avoids the need for heat denaturation of the template DNA molecule. In some embodiments, the DNA is ssDNA and the adaptor comprising514904-2842-5073.2Atorney Docket No. N.057.W0.01 the primer binding sequence is appended to each single stranded template. Using this method and variations thereof, a maximum of 2-fold amplification can be achieved. ii) Strand invasion
[0168] In some embodiments, adaptors can be designed to be more susceptible to fraying at specific temperatures. For example, adaptors having higher AT content will be prone to fraying at around 60 °C. In such methods, primer hybridization can take advantage of fraying DNA ends, allowing strand invasion to occur. In some embodiments, primers are designed to improve strand invasion capability. For example, primers comprising locked nucleic acids (LNAs) may be used. In some embodiments described herein, recombinase polymerase amplification (RPA) may be used to enhance primer binding. Using these methods and variations thereof, 10-fold or more amplification can be achieved. iii) Enhanced strand invasion
[0169] In some embodiments, the primers used in the methods described herein are physically linked (5 ’-5’) to increase primer concentration and on-rate. One challenge with isothermal amplification is creating single-stranded regions where primers can bind. This is particularly challenging since only the 5’ end of a primer extension product is ‘controlled’ i.e., can be determined by a primer sequence (and priming occurs on the 3’ end of a template strand). By utilizing 5’ linked primers, the local concentration of primers can be dramatically increased for the subsequent cycle, which leads to a higher rate of strand invasion.(iv) Extension reaction mixture
[0170] Typically, in embodiments described herein, primer extension is performed by adding an extension reaction mixture to the template DNA (e.g., adapted sample DNA, such as adapted sample cfDNA) followed by addition of a polymerase enzyme. In some embodiments, the extension reaction mixture contains one or more primers, deoxynucleotides (dNTPs), reaction buffer, and deionized water. In some embodiments, the dNTPs can comprise a mixture of dATP, dCTP, dGTP and dTTP. In some embodiments, the final concentration of each dNTP in the reaction mixture can range from 0.05 mM to 0.5 mM dNTPs, for example, from 0.05 mM to 0.5 mM, 0.05 mM to 0.1 mM, 0.05 mM to 0.15 mM, 0.05 to 0.2 mM, 0.05 mM to 0.25 mM, 0.05 mM to 0.3 mM, 0.05 mM to 0.35 mM, 0.05 to 0.4 mM 0.05 mM to 0.45 mM, from 0.1 mM to 0.5 mM, 0.15 mM to 0.5 mM, or 0.2 mM to 0.5 mM, 0.25 to 0.5 mM, 0.3 mM to 0.5 mM, 0.35 mM to 0.5524904-2842-5073.2Attorney Docket No. N.057.W0.01 mM, 0.4 mM 0.5 mM, or from 0.45 mM to 0.5 mM. In some embodiments, the final concentration of each dNTP in the reaction mixture is between 0.15 mM and 0.25 mM. In some embodiments, the final concentration of each dNTP in the reaction mixture is 0.2 mM.
[0171] In illustrative embodiments, the reaction buffer is a BST polymerase buffer such as the ThermoPol® Reaction Buffer (B9004S, New England BioLabs, Inc.). In some embodiments, the reaction buffer is an isothermal amplification buffer (e.g., B0537S, New England BioLabs, Inc.). In some embodiments, the reaction buffer is Q5® Reaction Buffer (B9027S, New England BioLabs, Inc.). In some embodiments, the reaction buffer is Standard Taq Reaction Buffer (B9014S, New England BioLabs, Inc). In some embodiments, the reaction buffer is a Standard Taq (Mg-free) Reaction Buffer (B9015S, New England BioLabs, Inc.).
[0172] In some embodiments, the polymerase is a BST DNA polymerase, such as BST DNA Polymerase (M0328S, New England BioLabs. Inc.) or BST 2.0 DNA Polymerase (M0537S, New England BioLabs, Inc.). BST DNA polymerases have 5' —> 3' polymerase and double-strand specific 5' — »■ 3' exonuclease activity, but lack 3' 5' exonuclease activity. In some embodiments, a warm start BST DNA polymerase may be used to help reduce non-specific activity. For example, in some embodiments, the polymerase is Bst 2.0 WarmStart® DNA Polymerase (M0538S, New England BioLabs, Inc.).
[0173] In some embodiments, the polymerase is selected based on the particular activity levels at specific temperatures. This is important when using the polymerase for primer extension prior to methyl transfer, as explained in more detail below. For example. BST DNA polymerase provides 100% activity at 60-65 °C, but low activity (10-15%) at 37 °C. Accordingly, there will be little amplification occurring when the methyl transfer step is being performed. In some embodiments, isothermal conditions are used for primer extension, such that the reaction is performed at around 60 °C for example. In some embodiments, primer extension may be performed for a short amount of time. For example, in some embodiments, primer extension is performed for between 1 and 5 mins, between 1 and 10 mins, between 1 and 20 mins, or between 1 and 30 mins.
[0174] Buffer solution creates a suitable environment for the polymerase enzyme and can contain many different components, including magnesium chloride (MgC12), potassium chloride (KC1), dimethyl sulfoxide (DMSO), and glycerin or bovine serum albumin (BSA). In some embodiments, the concentration of KC1 can be between 25 and 50 mM, between 25 and 75 mM, between 25 and 100 mM, between 30 and 100 mM, between 50 and 100 mM, or between 70 and 100 mM. In some534904-2842-5073.2Attorney Docket No. N.057.W0.01 embodiments, the MgC12 concentration can be in the range of 0.5 mM to 5 mM, between 0.5 mM to 4.5 mM, 0.5 to 4.0 mM, 0.5 to 3.5 mM, 0.5 to 3.0 mM, 0.5 to 2.5 mM, 0.5 to 2.0 mM, 0.5 to 1.5 mM, 1.0 to 5 mM, 1.5 to 5mM, 2.0 to 5 mM, 2.5 to 5 mM, 3.0 to 5 mM, 3.5 to 5 mM, 4.0 to 5 mM, or 4.5 to 5 mM. In some embodiments, the concentration of MgC12 is 2.0 mM.
[0175] In some embodiments, the primers are universal primers designed to hybridize to sequences on the appended adaptors. In some embodiments, the length of the primers can be between 10 to 100 nucleotides, such as between 10 to 75 nucleotides, 10 to 40 nucleotides, 10 to 35 nucleotides, 10 to 30 nucleotides, 10 to 20 nucleotides, 15 to 100 nucleotides, 20 to 100 nucleotides, from 25 to 100 nucleotides, from 30 to 100 nucleotides from 35 to 100 nucleotides, from 40 to 100 nucleotides, from 45 to 100 nucleotides, from 50 to 100 nucleotides, from 55 to 100 nucleotides, from 60 to 100 nucleotides, from 65 to 100 nucleotides, from 70 to 100 nucleotides, or from 75 to 100 nucleotides. In some embodiments, the range of the length of the primers is between 5 to 50 nucleotides, such as 5 to 40 nucleotides, 5 to 20 nucleotides, or 5 to 10 nucleotides. In some embodiments, the primers are between 5 and 50 bp in length, between 10 and 40 bp in length, between 15 and 30 bp in length, between 15 and 25 bp in length, between 20 and 40 bp in length, between 25 and 50 bp in length, or between 30 and 50 bp in length. In some embodiments, the primers are between 25 and 100 bp in length, between 35 and 100 bp in length, between 45 and 100 bp in length, between 55 and 100 bp in length, between 65 and 100 bp in length, or between 75 and 100 bp in length.(v) Single-stranded DNA (ssDNA) molecules
[0176] In some embodiments, the template DNA molecules are single- stranded DNA (ssDNA). In some embodiments, at least one adapter having a primer binding sequence is appended to one or more ssDNA template molecules. In some embodiments, only a single primer is needed for primer extension, which does not require heat-denaturing. The primer binds to and extends from the primer binding sequence in the adapter. In some embodiments, primers and adapters in the methods disclosed herein may include nucleotides with methylation modifications (e.g. methylated cytosines) for use in downstream methylation detection assays. ssDNA can be prepared using single stranded library preparation methods such as IDT xGen™ ssDNA & Low- Input DNA Library Preparation Kit or ClaretBIO SRSLY kit. In some embodiments, a second adapter is appended before or after the primer extension and methylation transfer steps. In some544904-2842-5073.2Attorney Docket No. N.057.W0.01 embodiments, after the first primer extension, other amplification methods disclosed herein (e.g., strand invasion, enhanced strand invasion) can be used for additional amplification.Methylation Transfer
[0177] Template copying or amplification may not preserve the methylation status of the template DNA molecule on the copied DNA molecule. Accordingly, after the template DNA molecule is copied, the methods described herein include a step of methylation transfer to generate methylation-transferred DNA molecules.
[0178] In the methods described herein, one or more methylating agents (for example, methyltransferase) are used to preserve the methylation pattern of the template DNA molecule on the copied DNA molecule. In some illustrative embodiments, the methylating agent is DNMT1. DNMT1 is a maintenance DNA methyltransferase that propagates the CpG DNA methylation pattern in dividing cells; DNMT1 acts on hemi-methylated DNA, where one strand (the original, template strand) is methylated and the other strand (the newly synthesized strand) is not methylated, and copies the methylation pattern from the old strand to the new strand with high efficiency and specificity. DNMT1 is not thermostable, and accordingly in some embodiments the method described herein comprises adding DNMT1 fresh after each cycle of template amplification or copying. In some embodiments, other methylating agents may be used, such as mammalian methyltransferases DNMT3a and DNMT3b, plant methyltransferases DRM2, MET1, and CMT3, and bacterial methyltransferase Dam. In some embodiments, methyltransferases such as DNMT1 or other suitable methyl transferases are used with one or more sources of methyl groups, such as S-adenosylmethionine (SAM), and may be used with or without cofactors such as NP95 (Uhrfl). In some embodiments, a DNMT1 reaction buffer may be used with the methyltransferase that includes DNMT1 reaction buffer. For example, the DNMT1 reaction buffer may comprise 200 mM NaCl. 50 mM Tris-HCl. 1 mM EDTA, 1 mM DTT, and 50% glycerol. In some embodiments, the DNMT1 reaction buffer may further comprise BSA and S- adenosylmethionine (SAM).
[0179] In some embodiments, a methyltransferase, such as DNMT1, may require conditions different to those present in the template copying reactions. For example, methyl transfer reactions may require buffer conditions that do not include ions (such as cations), such as magnesium ions or manganese ions which may be a component of a primer extension reaction. Accordingly, in554904-2842-5073.2Attorney Docket No. N.057.W0.01 some embodiments, a chelating agent such as EDTA is required after the primer extension step to chelate ions, such as magnesium ions in order for the methylation step to be carried out. In some embodiments, magnesium is replenished back into the primer extension reaction mixture for the next round of primer extension after the completion of methyl transfer reaction.
[0180] As explained above, the primer extension step is generally performed at about 60 °C, when the activity of the polymerase, such as BST, is highest. Methyl transferase on the other hand, such as DNMT1 has a recommended temperature of 37 °C and is deactivated at 65 °C. In some embodiments, the methyl transfer step is performed for between 5 and 120 mins, for between 10 and 90 mins, for between 20 and 60 mins, or for between 30 and 45 mins. Accordingly, in illustrative examples, a two-step process of primer extension at approximately 60 °C (or in some embodiments, between about 37 °C and 64 °C, between about 40 °C and 64 °C, between about 50 °C and 64 °C. or between about 58 °C and 64 °C) to copy template DNA molecules, and DNA methyl transfer at about 37 °C (in some embodiments, between about 30 °C and 45 °C, between about 33 °C and 42 °C, or between about 35 °C and 40 °C) to copy the methylation pattern on to the copied DNA molecule can be repeated multiple times, to generate a library of amplified DNA molecules comprising one or more methylation-transferred DNA molecules, during which the two enzymes will not be active at the same time. In some embodiments, the primer extension step is performed for a shorter amount of time than the methyl transfer step.
[0181] In some embodiments, the two-step process of primer extension and methylation transfer may be repeated for at least 1 cycle, at least two cycles, at least three cycles, at least four cycles, at least five cycles, at least 10 cycles, at least 15 cycles, at least 20 cycles, at least 30 cycles, at least 40 cycles, or at least 50 cycles. In some embodiments, the two-step process of primer extension and methylation transfer may be repeated for more than 50 cycles. In some embodiments, the two- step process of primer extension and methylation transfer may be repeated for between 1 and 50 cycles, between 1 and 40 cycles, between 1 and 30 cycles, between 1 and 20 cycles, between 1 and 15 cycles, between 1 and 10 cycles, or between 1 and 5 cycles.Methylation Assays
[0182] DNA methylation biomarkers are increasingly being utilized in the development of novel assays for use in, for example, cancer and other disease detection, women’s health, organ health, and veterinary health. For example, methylation profiling is being used in non-invasive prenatal564904-2842-5073.2Atorney Docket No. N.057.W0.01 testing (NIPT) for monitoring of placental and fetal epigenomic changes, pre- symptomatic detection of preterm birth, preeclampsia, placental insufficiency, and fetal growth restriction. Methylation markers can also be indicative of congenital diseases of the fetus. Furthermore, methylation biomarkers are also used for monitoring of organ health such as predicting and monitoring organ rejection in transplant patients, monitoring of immune changes in transplant rejection, and for monitoring of organ health in high risk or predisposed individuals. In addition, methylation biomarkers can be used to determine biological age or detect age-related diseases. Accordingly, the sample described herein may be a maternal sample, a fetal sample, a sample from a transplant patient, or a sample from a subject having, had, suspected of having, or at risk of having a certain phenotype.
[0183] In some embodiments, the method described herein comprises treating the first library of methylation-preserved amplified DNA or a portion thereof with an agent or a combination of agents and generating treated first methylation-preserved amplified DNA, wherein the agent or the combination of agents discriminates between methylated and unmethylated cytosines. In some embodiments, the method described herein comprises treating the second library of methylation- preserved amplified DNA or a portion thereof with an agent or a combination of agents and generating treated second methylation-preserved amplified DNA, wherein the agent or the combination of agents discriminates between methylated and unmethylated cytosines.
[0184] In some embodiments, a library of methylation-preserved amplified DNA or a portion thereof is treated with a deaminating agent that converts (directly or indirectly through one or more other agents) unmethylated but not methylated cytosines to uracils before performing targeted enrichment In some embodiments, a library of methylation-preserved amplified DNA or a portion thereof is treated with a deaminating agent (directly or indirectly through one or more other agents) that converts unmethylated but not methylated cytosines to uracils after performing targeted enrichment. In some embodiments, the deaminating agent (e.g., a chemical reagent such as sodium bisulfite) directly converts unmethylated cytosines to uracils. In some embodiments, the deaminating agent converts unmethylated cytosines to uracils in combination with one or more other agents - for example, in order to convert unmethylated cytosines to uracils, a library of methylation-preserved amplified DNA or a portion thereof may be first treated with an oxidizing agent (e.g., TET or its catalytic domain) that oxidizes one or more forms of methylated cytosines574904-2842-5073.2Attorney Docket No. N.057.W0.01(e.g., 5hmC or 5mC), followed by treatment with a deaminase (e.g., APOB EC or its catalytic domain).
[0185] In some embodiments, the deaminating agent comprises a chemical reagent (e.g., sodium bisulfite). In some embodiments, the deaminating agent comprises a deaminase (e.g., APOBEC) or its catalytic domain. In some embodiments, the deaminating agent comprises APOBEC3A (A3A) or its catalytic domain.
[0186] In some embodiments, a library of methylation-preserved amplified DNA or a portion thereof is treated with a reducing agent that converts certain forms of modified cytosines, but not unmodified cytosines, into dihydrouridine (DHU) before the contacting with the panel of oligonucleotide probes. In some embodiments, the selected DNA or derivative thereof is treated with a reducing agent that converts certain forms of modified cytosines, but not unmodified cytosines, into DHU after the contacting with the panel of oligonucleotide probes. In some embodiments, the reducing agent comprises pyridine borane.
[0187] In some embodiments, a library of methylation-preserved amplified DNA or a portion thereof is treated with an oxidizing agent that oxidizes one or more forms of methylated cytosines before being treated with a deaminating agent or a reducing agent. In some embodiments, the oxidizing agent oxidizes 5hmC and 5mC to 5caC (5-carboxylcytosine). In some embodiments, the oxidizing agent selectively oxidizes 5hmC, but not 5mC, to 5fC (5-formylcytosine). In some embodiments, the oxidizing agent comprises an enzyme of the ten-eleven translocation (TET) families or its catalytic domain. In some embodiments, the oxidizing agent comprises TET1 or its catalytic domain. In some embodiments, the oxidizing agent comprises TET2 or its catalytic domain. In some embodiments, the oxidizing agent comprises TET3 or its catalytic domain. In some embodiments, the oxidizing agent comprises potassium perruthenate (KRuO4). In some embodiments, the oxidizing agent comprises potassium ruthenate (K2RuO4).
[0188] In some embodiments, a library of methylation-preserved amplified DNA or a portion thereof is treated with a glycosylation agent that glycosylates 5hmC to glucosyl-5hmC (5ghmC) to protect it from oxidation, deamination, or reduction before being treated with an oxidizing agent, a deaminating agent, or a reducing agent. In some embodiments, the glycosylation agent is a glycosyltransferase. In some embodiments, the glycosyltransferase is a glucosyltransferase. In some embodiments, the glucosyltransferase is P-glucosyltransferase (|3-GT). In some embodiments, the glucosyltransferase is (3-GT of T4 phase.584904-2842-5073.2Attorney Docket No. N.057.W0.01
[0189] In some embodiments, a library of methylation-preserved amplified DNA or a portion thereof is treated with a methylation sensitive restriction enzyme (MSRE), methylation dependent restriction enzymes (MDRE), a 5-methylcytosine (5mC) antibody, a 5-hydroxymethylcytosine (5hmC) antibody, a methyl-CpG-binding domain (MBD) protein, or a DNA methyltransferase, before selective enrichment with a panel of oligonucleotide probes. In some embodiments, the selected DNA or derivative thereof is treated with a methylation sensitive restriction enzyme (MSRE), methylation dependent restriction enzymes (MDRE), a 5-methylcytosine (5mC) antibody, a 5-hydroxymethylcytosine (5hmC) antibody, a methyl-CpG-binding domain (MBD) protein, or a DNA methyltransferase, after selective enrichment with a panel of oligonucleotide probes. In some embodiments, other methylation detection methods such as methylated DNA immunoprecipitation (MeDIP) or direct detection by long read sequencers may also be used.
[0190] Libraries of methylation-transferred DNA molecules (e.g., methylation-preserved amplified DNA) generated using the methods described herein or portions thereof can be used for many downstream applications. One of such applications is for the analysis of methylation biomarkers, including for non-invasive prenatal testing (NIPT), cancer detection or diagnosis, and transplant health diagnosis. For purposes of illustration only, some of the methods that can be used for the analysis of methylated DNA are described below.Conversion-based methylation detection
[0191] In some embodiments, the methods described herein include treating a library of methylation-preserved amplified DNA or a portion thereof with one or more chemical or enzymatic agents or a combination of chemical and enzymatic agents (e.g. deaminating agents, reducing agents, oxidizing agents, glycosylation agents) that allow discrimination between methylated and unmethylated cytosines, or between 5mC and 5hmC. In some embodiments, a portion of a library of DNA molecules or their derivatives is treated with one or more such chemical and / or enzymatic agents. For example, in some embodiments, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the library is treated with one or more such chemical and / or enzymatic agents. In some embodiments, the entire library of DNA molecules or their derivatives is treated with one or more such chemical or enzymatic agents.
[0192] In some embodiments, the agent comprises a deaminating agent. In some embodiments, the deaminating agent comprises an enzyme such as a deaminase or its catalytic domain. In some594904-2842-5073.2Attorney Docket No. N.057.W0.01 embodiments, the deaminase may be APOBEC or its catalytic domain. In some embodiments, the deaminase may be APOBEC3A (A3A) or its catalytic domain, which deaminates unmethylated C, 5mC, and 5hmC, but not 5ghmC, 5fC, or 5caC. In some embodiments, the method described herein involves treating a library of methylation-preserved amplified DNA or a portion thereof, with TET2 / oxidation enhancer, which converts 5mC and 5hmC to 5caC and protects them from deamination by APOBEC. followed by treatment with APOBEC to deaminate the unmethylated cytosines to uracils. In some embodiments, the DNA sample is treated with a glycosylation agent (e.g. T4 P-glucosyltransferase) to glycosylates 5hmC to 5ghmC and protects it from deamination by APOBEC.
[0193] In some embodiments, the deaminating agent comprises a bisulfite reagent. Bisulfite reagents, such as sodium bisulfite, convert unmethylated cytosine to uracil and leave methylated cytosine (including 5mC and 5hmC) unchanged. Therefore, after bisulfite treatment, 5mC and 5hmC in the DNA remains as cytosine and unmodified or unmethylated cytosine will be changed to uracil. In some embodiments, the bisulfite reagent may be sodium bisulfite, potassium bisulfite, ammonium bisulfite, magnesium bisulfite, sodium metabisulfite, potassium metabisulfite, ammonium metabisulfite, magnesium metabisulfite, and equivalents. In some embodiments, a combination of one or more chemical reagents and / or one or more enzymes can be used to discriminate between various forms of methylated cytosines (e.g., mC or hmC) and unmethylated cytosines. For example, in some embodiments, the DNA sample is treated with an oxidizing agent (e.g. KRuO4 or K2RuO4) that selectively oxidizes 5hmC, but not 5mC. to 5fC before bisulfite treatment, which converts unmethylated cytosine (including 5fC converted from 5hmC) to uracil but leaving 5mC unchanged. In some embodiments, a glycosylation agent (e.g. T4 p- glucosyltransferase or its catalytic domain) is used to glycosylate 5hmC to 5ghmC and an oxidizing agent (e.g. a TET enzyme or its catalytic domain) is used to convert 5mC to 5caC before bisulfite treatment, which converts unmethylated cytosine (including 5caC converted from 5mC) to uracil but leaving methylated cytosine (including 5ghmC converted from 5hmC) unchanged.
[0194] The bisulfite treatment can be performed by commercial kits such as the Imprint DNA Modification Kit (Sigma), EZ DNA Methylation- DirectTM Kit ( ZYMO), and the EZ DNA Methylation-Gold Kit (ZYMO). After DNA bisulfite conversion, single stranded DNA is captured, desulphonated and cleaned. The bisulfite-treated DNA can be captured by purification columns or magnetic beads and eluted. Bisulfite-treated single stranded DNA can be converted604904-2842-5073.2Attorney Docket No. N.057.W0.01 into dsDNA through an enzyme-catalyzed DNA strand synthesis with appropriate primers and polymerase. The polymerase will recognize the uracil in the ssDNA template as thymine and add an adenine to the complementary strand. Further polymerase extension on the complementary strand will result in replication of the original bisulfite treated ssDNA template, substituting uracil with thymine. Identification of cytosine to thymine conversion and guanine to adenine conversions (complementary strand) through comparing to the reference genome, will determine all unmodified cytosines, while the remaining cytosines are considered to be methylated.
[0195] In some embodiments, the agent comprises a reducing agent (e.g. pyridine borane) that selectively reduces 5fC and 5caC to dihydrouridine (DHU), which, like uracil, is converted to T base following PCR amplification. In some embodiments, the method described herein involves treating a library of methylation-preserved amplified DNA or a portion thereof, with TET2 / oxidation enhancer, which converts 5mC and 5hmC to 5caC, followed by treatment with borane, which reduces 5fC and 5caC (including 5caC converted from 5mC and 5hmC) to DHU but leaves unmodified cytosine (dC) unchanged. In some embodiments, the DNA sample is treated with an oxidizing agent (e.g. KRuO4 or K2RuO4) that selectively oxidizes 5hmC, but not 5mC, to 5fC, followed by treatment with borane, which reduces 5fC (including 5fC converted from 5hmC) and 5caC to DHU but leaves dC and 5mC unchanged. In some embodiments, the DNA sample is first treated with a glycosylation agent (e.g. T4 0-glucosyltransferase) and an oxidizing agent (e.g. TET) such that 5hmC is glycosylated to 5ghmC and only 5mC is oxidized and converted to 5caC. The subsequent treatment with borane reduces 5fC and 5caC (including 5caC converted from 5mC) to DHU but leaves dC and 5ghmC unchanged.
[0196] Other conversion-based methyl detection methods, including EM-seq, oxidative bisulfite sequencing (oxBS-seq), TET-assisted bisulfite sequencing (TAB-seq), TET-assisted pyridine borane sequencing (TAPS), TAPS with T4-0GT protection sequencing (TAPS0), chemical- assisted pyridine borane sequencing (CAPS), and many variants of these that can discriminate between C, mC, and / or hmC. These methods can be performed with commercial kits such as NEBNext® Enzymatic Methyl-seq (EM-seq™) (New England Biolabs, Inc.), the 5hmC TAB-Seq Kit (WiseGene), and the EpiTect® Bisulfite Kit (Qiagen), for example.Restriction enzyme-based methylation detection methods614904-2842-5073.2Attorney Docket No. N.057.W0.01
[0197] In some embodiments, the methods described herein include treating the methylation- transferred DNA molecules, or their derivatives, with one or more restriction enzymes. In some embodiments, the one or more restriction enzymes are one or more methylation sensitive restriction enzymes (MSREs). In some embodiments, a portion of a library of amplified DNA molecules comprising the one or more methylation-transferred DNA molecules is treated with one or more MSREs. For example, in some embodiments, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the library is treated with one or more MSREs. In some embodiments, the entire library of amplified DNA molecules comprising the one or more methylation-transferred DNA molecules is treated with one or more MSREs.
[0198] MSREs are useful in analyzing the methylation status of cytosine residues in CpG sites. As the name implies, these enzymes are not able to cleave their palindromic target sites when the cytosine residues are methylated. The size of the MSRE targets sites range from 4 bp up to 76bp, but typically are in the 4 to 8 bp range.
[0199] In one aspect, methods herein typically include contacting methylation- transferred DNA molecules, or their derivatives, with one or more MSREs. MSREs selectively cleave a sample nucleic acid when the MSRE recognition site is unmethylated, but not when the MSRE recognition site is methylated. An “isoschizomer” of an MSRE is a restriction enzyme that recognizes the same recognition site as a methylation sensitive restriction enzyme but cleaves both methylated CGs and unmethylated CGs. Isoschizomer of the selected MSREs may be used in control reactions. Non-limiting examples of methylation sensitive restriction enzyme include, and thus in some embodiments, the one or more MSREs can include, Aatll, Acc65I, AccI, Acil, Acll, Afel, Agel, Agel-HF®, AhdI, Alel-v2, Apal, ApaLI ApeKI, Asci, AsiSI, Aval, Avail, Bael, BanI, BbvCI, BceAI,, Bcgl, BcoDI, BfuAI, Bgll, BmgBI, BsaAI, BsaBI, BsaHI, BsaI-HF®v2, BseYI, BsiE, BsiWI, BsiWI-HF®, BslI, BsmAI, BsmBI-v2, BsmFI, BspDI, BspEI, BsrBI, BsrFI-v2, BssHII, BstAPI, BstBI, BstUI, BstZ17I-HF®, BtgZI, Cac8I, Clal, Dpnl, Dralll-HF®, DrdI, Eael, Eagl-HF®, Earl, Ecil, Eco53kl, EcoRI, EcoRI-HF®,EcoRV, EcoRV-HF®, Esp31, Faul, Fnu4HI, FokI, Fsel, FspI, Haell, Hgal, Hhal, HinPlI, HincII, Hinfl, Hpal, Hpall, Hpyl66II, Hpyl88III, Hpy99I, HpyAV, HpyCH4IV, KasI, Mbol, Mid, MluI-HF®, Mmel, MspAlI, Mwol, Nael. Narl, Neil, NgoMIV, Nhel-HF®, NlalV, Notl, Notl-HF®, Nrul, NruI-HF®, Nt.BbvCI, Nt.BsmAI, Nt.CviPII, PaeR7I, PaqCI, Piel, PluTI, Pmel, Pmll, PshAI, PspOMI, PspXI, Pvul, PvuI-HF®, Rsal, RsrII. SacI-HF®, SacII, Sall, Sall-HF®. Sau3AI, Sau96I, ScrFI, SfaNI. Sfil, Sfol, SgrAI.624904-2842-5073.2Attorney Docket No. N.057.W0.01Smal, SnaBI, Srfl, StyD4I, Tfil, Tsel, TspMI, Xhol, Xmal. and / or Zral. In illustrative embodiments, the one or more MSREs can include Hhal, Hpall, BstUI, and / or HpyCH4IV.
[0200] In some embodiments, the one or more restriction enzymes are one or more methylation dependent restriction enzymes (MDREs). MDREs selectively cleave a sample nucleic acid when one or more nucleotides in the MDRE recognition site is methylated. In some embodiments, one or more MDREs can be used in combination with one or more MSREs. In some embodiments, the one or more MDREs can be AbaSI, AoxI, BisI, BlsI, Dpnl, FspEI, Glal, Glul, Krol, LpnPI, Mall, MspJI, Mtel, Pcsl, PkrI, or Sgel.
[0201] The MSRE or MDRE can be selected based on differentially methylated CpG sites in target DNA molecules, such as tumors or ctDNA from specific cancer targets, or a diverse spectrum of tumors. Further criteria for selection may include low background methylation in normal tissues, size and number of cleavage fragments, number of base pairs of recognition sequence, whether the cleavage results in blunt vs. tailed end fragments, and whether the enzymes have the same or similar reaction conditions such that the contacting step can be done under the same set of conditions and / or in a single reaction.
[0202] In some embodiments, more than one, a plurality, or a set of MSREs (and / or MDREs) can be used to contact methylation-transferred DNA molecules, or their derivatives, comprising one or more CpG sites of interest. The number of selected MSREs in certain embodiments is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 20, 25, 50 or 100. In some embodiments, the number of MSREs range from 1 to 2, 3, 4, 5, 6, 7, 8, 9, 10, or 20, 25, 50 or 100; from 2 to 3, 4, 5, 6, 7, 8. 9, 10, or 20. 25. 50 or 100; or from 5 to 6, 7, 8, 9, 10, or 20, 25, 50 or 100. In some embodiments, the number of MSREs range from 1 to 10, 1 to 5, from 2 to 5, from 3 to 5, from 4 to 5, 3 to 7, from 5 to 10, from 6 to 10, or from 7 to 10. Thus, criteria such as target site, target site methylation status, number of base pairs in tailed end fragments, reaction conditions including but not limited to buffer conditions, incubation time and temperature for optimal activity, as well as deactivation time and temperature for each MSRE can be used for selecting the plurality or set of MSREs and / or MDREs to include in the method. As activity measured by units is specific to each enzyme, the target DNA molecules can be contacted with from 1 to 5 Units (U) of each MSREs in the reaction sample. In some embodiments, the methylation-transferred DNA molecules, or their derivatives, are contacted with from 1 to 5 U, from 1.5 to 4 U, from 2 to 3 U, from 2.5 to 4 U, or from 3 to 5 U of each MSRE. In some embodiments, one or more of the MSRE is selected from Hpall, Sall,634904-2842-5073.2Attorney Docket No. N.057.W0.01Bbel, Notl, Smal, Xmal, Mbol, BstUI, BstBI, Clal, Mini, Nael, Narl, Pvul. SacII, HpyCH41V. Hhal, and combinations thereof. In exemplary embodiments, the one or more MSREs is selected from one or more of Hpall, Hhal, HpyCH41V, and BstUI. In some embodiments, the one or more MSREs are selected from one or more of Hpall. Hhal, HpyCH41V, or BstUI. In some embodiments of the methods as described herein, the contacting comprises contacting with two or more MSREs. In some embodiments, the contacting comprises contacting with three or more MSREs. In some embodiments, the contacting comprises contacting with four or more MSREs. In some embodiments, the contacting comprises contacting with the two or more MSREs in a single reaction.
[0203] In some embodiments, MSREs and / or MDREs are included that are able to cleave the MSRE sites and / or MDRE sites in the same cleavage buffer conditions. In some embodiments, the cleavage buffer conditions include 5-500 mM potassium acetate, for example 25-100 mM potassium acetate. In some embodiments, the cleavage buffer conditions include 2-200 mM Trisacetate, for example 10-40 mM Tris-acetate. In some embodiments, the cleavage buffer conditions include 1-100 mM magnesium acetate, for example 5-20 mM magnesium acetate. In some embodiments, the cleavage buffer conditions include 10-1000 pg / ml recombinant albumin, for example 50-200 pg / ml recombinant albumin. In some embodiments, the pH is between 6.9 and 8.9 at 25 °C, for example, between 7.4 and 8.4, 7.5 and 8.3, 7.6 and 8.2, 7.7 and 8.1, or 7.8 and 8, or about 7.9 at 25 °C. In some embodiments, the cleavage buffer conditions include 1-100 mM bis- tris-propane-HCl, for example 5-20 mM bis-tris-propane-HCl. In some embodiments, the cleavage buffer conditions include 1-100 mM MgC12, for example 5-20 mM MgC12. In some embodiments, the pH is between 6 and 8 at 25 °C, for example, between 6.5 and 7.5, 6.6 and 7.4, 6.7 and 7.3. 6.8 and 7.2, or 6.9 and 7.1. or about 7.0 at 25 °C. In some embodiments, the cleavage buffer conditions include 5-500 mM NaCl, for example 25-100 mM NaCl. In some embodiments, the cleavage buffer conditions include 1-100 mM Tris-HCl, for example 5-20 mM Tris-HCl. In some embodiments, the cleavage buffer conditions include 10-1000 mM NaCl, for example 50- 200 mM NaCl. In some embodiments, the cleavage buffer conditions include 5-500 mM Tris-HCl, for example 25-100 mM Tris-HCl. In some embodiments, the cleavage buffer conditions include 25-100 mM potassium acetate, 10-40 mM Tris-acetate, 5-20 mM magnesium acetate, and 50-200 pg / ml recombinant albumin, and the pH is between 7.6 and 8.2 at 25 °C. In some embodiments, the cleavage buffer conditions include 5-20 mM bis-tris-propane-HCl, 5-20 mM MgC12, and 50-644904-2842-5073.2Attorney Docket No. N.057.W0.01200 pg / ml recombinant albumin, and the pH is between 6.7 and 7.3at 25 °C. In some embodiments, the cleavage buffer conditions include 25-100 mM NaCl, 5-20 mM Tris-HCl, 5-20 mM MgC12, and 50-200 pg / ml recombinant albumin, and the pH is between 7.6 and 8.2 at 25 °C. In some embodiments, the cleavage buffer conditions include 50-200 mM NaCl 25-100 mM Tris- HCl, 5-20 mM MgC12, and 50-200 pg / ml recombinant albumin, and the pH is between 7.6 and 8.2 at 25 °C.
[0204] In some embodiments, the method described herein comprises amplifying the treated methylation-preserved amplified DNA or its derivative after treatment with an agent or a combination of agents that discriminates between methylated and unmethylated cytosine. In some embodiments, the method described herein comprises amplifying the treated methylation- preserved amplified DNA or its derivative before performing targeted enrichment. In some embodiments, the method described herein comprises amplifying the treated methylation- preserved amplified DNA or its derivative after performing targeted enrichment. In some embodiments, the amplifying adds a sample barcode, and a plurality of libraries of the enriched DNA or derivative thereof that are produced sequentially from the same sample are sequenced together in one sequencing lane. In some embodiments, the amplifying adds a sample barcode, and a plurality of libraries of the enriched DNA or derivative thereof that are produced from a plurality of different samples are sequenced together in one sequencing lane. In some embodiments, the method further comprises performing size selection before or after the amplifying step.Error correction
[0205] Some embodiments of the methods described herein use molecular barcode or index sequences for identifying and removing errors produced at different stages of the methods.
[0206] In previous methods of conversion based detection of a methylation pattern of a target region, amplification of a template is not performed prior to the conversion, for example chemical and / or enzymatic conversion. Accordingly, a molecular barcode or index sequence would only represent a single conversion event since there are no molecular copies to be converted. If a mistake is made during conversion, this mistake will be propagated to any molecular copies made post-conversion, resulting in an incorrect consensus sequence. Likewise, in previous restrictionenzyme-based methylation detection methods, a mistake during the restriction enzyme digestion would be propagated to any molecular copies made post-restriction digest.654904-2842-5073.2Atorney Docket No. N.057.W0.01
[0207] Because molecular barcode or index sequences are appended to the template DNA molecules before template copying and methylation transfer, which in turn are performed prior to bisulfite or enzymatic conversion, molecular barcode or index sequences represent multiple conversion events, and thus can be used to correct errors in conversion and / or amplification. Consensus can be made which preserve the initial state of the template DNA molecule (both sequence and methylation). Unlike consensus sequences often used for determining genetic variants, the consensus required to detect a methylated base does not need to be over 50% of the bases at a given position for a given MIT. For example, as long as the methylated base signal is above the background methylation non-specific activity, library preparation, PCR, sequencing errors and other error and non-specific activity at that position, a methylated base may be called. Background errors can be modeled and / or measured experimentally. The use of methylation amplification in conjugation with molecular barcode or index sequences to create a methylation consensus from multiple conversion events was previously unknown and represents a significant improvement over existing technology.Sequencing to Generate Sequence Reads
[0208] In some embodiments, the method described herein comprises sequencing the first enriched DNA and the second enriched DNA or their respective derivative and producing a first set of sequence reads and a second set of sequence reads, respectively. In some embodiments, the method described herein comprises detecting at least one somatic variant or mutation associated with a phenotype (e.g., a disease such as cancer) present in both the first enriched DNA and the second enriched DNA based on the first set of sequence reads and the second set of sequence reads. In some embodiments, the method described herein comprises identifying an error in amplification, enrichment, or sequencing from the first set of sequence reads in combination with the second sets of sequence reads, optionally without using molecular barcode or index sequences as the first and second sets of sequence reads are capable of cross-checking each other to eliminate processing errors and false-positives. In some embodiments, the identification of one or more somatic variants or mutations associated with cancer that are present in both the first enriched DNA and the second enriched DNA is indicative of cancer or minimal residual disease.
[0209] In some embodiments, the method described herein comprises sequencing the first enriched DNA or its derivative and generating a first set of sequence reads. In some664904-2842-5073.2Attorney Docket No. N.057.W0.01 embodiments, the method described herein comprises sequencing the second enriched DNA or its derivative and generating a second set of sequence reads. In some embodiments, the method described herein comprises detecting at least one differentially methylated region associated with a phenotype (e.g., a disease such as cancer) based on the first set of sequence reads and the second set of sequence reads. In some embodiments, the method described herein comprises identifying an error in isothermal amplification, methyltransferase treatment, bisulfite or enzymatic conversion, enrichment or sequencing from the first set of sequence reads in combination with the second sets of sequence reads, optionally without using molecular barcode or index sequences as the first and second sets of sequence reads are capable of cross-checking each other to eliminate processing errors and false-positives. In some embodiments, the identification of one or more differentially methylated regions associated with cancer that are present in both the first enriched DNA and the second enriched DNA is indicative of cancer or minimal residual disease.
[0210] Methods as described herein include detecting and optionally quantifying nucleic acids, including DNA, cfDNA. enriched subsets of DNA having target regions, and in illustrative embodiments, target region amplicons or amplicons derived therefrom. In some embodiments, cfDNA from a blood sample from the individual is analyzed. Not to be limited by theory, cfDNA is believed to be released from certain cells, such as cancer cells, for example when they undergo necrosis or apoptosis. In some embodiments, methods herein can be used to detect somatic variants or mutations in target regions or nucleic acid sequence of interest that is present in a small percentage of DNA in a sample, such as cfDNA, for example from a fetus, a cell from a donated organ, or in illustrative embodiments, a cancer cell. In some embodiments, methylation-transferred DNA molecules prepared from the cfDNA from a blood sample from the individual is analyzed.
[0211] In some embodiments, cellular DNA from normal tissue, such as the huffy coat of the blood samples from the individual not suspected of having blood cancer, is analyzed. Non-tumor specific mutations, such as CH mutations, potentially contribute to a significant number of false positive tumor calls. Accordingly, in some embodiments, matched normal tissue samples, in some embodiments buffy coat samples and in some embodiments whole blood samples, from the individual may be used to enable identification and bioinformatic elimination of CH and other non-tumor specific mutations to reduce false positive variants and enhance the specificity of the methods described herein.674904-2842-5073.2Attorney Docket No. N.057.W0.01
[0212] In some embodiments, the sequencing is a next-generation sequencing or high-throughput sequencing. In methods herein, detecting or quantifying comprises counting sequence reads generated from target regions. Quantifying can also comprise determining a depth of read (DOR) per target region for at least some of the target regions. DOR for each of the target regions can be normalized relative to a DOR for a normalization sequence. DNA sequences used for normalization can be derived from genomic DNA or from control plasmids and will depend on the experimental conditions. The normalization sequence can be derived from a lambda control plasmid or synthetic sequence. The normalization sequence can be a spike-in control DNA sample. A spike-in control DNA sample can be a genomic, plasmid, or synthetic DNA sample. Methods herein, can include more than one, for example 2, 3, 4, 5, or more control or spike-in control samples. DNA sequencing techniques, particularly high throughput next- generation sequencing techniques (often referred to as massively parallel sequencing techniques) such as those employed in MYSEQ (ILLUMINA), HISEQ (ILLUMINA), ION TORRENT (LIFE TECHNOLOGIES), GENOME ANALYZER ILX (ILLUMINA), NOVASEQ (ILLUMINA), GS FLEX+(ROCHE 454). T7 (BGI). AVITI (ELEMENT BIOSCIENCES), ONSO (PACBIO). UG100 (ULTIMA GENOMICS), G4 (SINGULAR GENOMICS) etc., can be used for quantitative measurements of the number of copies of a target region present, for example, but not limiting to, target region amplicons or enriched subsets of amplified DNA, and thus provide quantitative information regarding the number and / or amount of target regions in sample DNA molecules. High throughput genetic sequencers are amenable to the use of barcoding (i.e., sample tagging with distinctive nucleic acid sequences) so as to identify specific samples from individuals thereby permitting the simultaneous analysis of multiple samples in a single run of the DNA sequencer. The number of times a given region of the genome in a library preparation (or other nucleic preparation of interest) is sequenced (number of reads) will be proportional to the number of copies of that sequence. Methods as described herein that utilize NGS detection, in some embodiments can have an average DOR of at least 200, 500, 1000, 2000, 2900, 3000, 3500, 4000, 5000, 10,000, 50,000, 75,000, 100,000, 120,000, 130,000, 150,000, 175,000, or 200,000, per target locus. In some embodiments, described herein, the sequencing has a depth of read (DOR) of between 50,000 to 200,000 per target locus.
[0213] Methods herein can include analyzing data obtained from next- generation sequencing techniques. In some embodiments of methods herein, target region amplicons or enriched subsets684904-2842-5073.2Attorney Docket No. N.057.W0.01 of amplified DNA can be subjected to sequencing using next- generation sequencing techniques. In some embodiments of methods herein, deaminated, or MSRE treated, methylation-transferred DNA molecules can be subjected to sequencing using next- generation sequencing techniques. Nucleic acid sequencing data can be generated for amplicons created by PCR, for example a multiplex targeted PCR. In some embodiments, the multiplex PCR can be a tiled multiplex PCR. For a skilled artisan, algorithm design tools are available that can be used and / or adapted to analyze the sequencing data. In addition, those skilled in the art can determine appropriate parameters for measuring alignment to a consensus sequence and / or to a known target region sequence, including any algorithms needed to achieve maximal alignment over the length of the sequences being compared.
[0214] Sequencing reads can be demultiplexed using an in-house tool and mapped using the Burrows-Wheeler alignment software, Bwa mem function (BWA, Burrows-Wheeler Alignment Software (see Li H. and Durbin R. (2010) Fast and accurate long-read alignment with Burrows- Wheeler Transform. Bioinformatics.) on single end mode using pear merged reads to the hgl9 genome. Amplification statistics QC can be performed by analyzing one or more of, but not limiting to, total reads, number of mapped reads, number of mapped reads on target, and number of reads counted.
[0215] Methods herein can include a background error model that can be constructed using normal, or healthy liquid samples, in illustrative embodiments, normal, or healthy plasma samples, which are sequenced on the same sequencing run to account for run-specific artifacts. In some embodiments, 5, 10, 15, 20, 25, 30, 40, 50, 100, 150, 200, 250, or more than 250 normal, or healthy liquid samples, in illustrative embodiments, plasma samples can be analyzed on the same sequencing run. The number of samples that can be sequenced on the same sequencing run can be in the range of 5 to 500, 5 to 400, 5 to 300, 5 to 250, 20 to 250, 30 to 250, 50 to 250, 75 to 250, 100 to 250, 50 to 500, or 100 to 500. Sample barcodes are used in illustrative embodiments. In some illustrative embodiments, 20, 25, 40, or 50 normal samples (e.g., plasma samples) can be analyzed on the same sequencing run. Outlier samples can be iteratively removed from the model to account for noise and contamination. In some embodiments, samples with a Z score of greater than 5, 6, 7, 8, 9, or 10 are removed from the data analysis. For each base substitution of every genomic loci, the DOR weighted mean and standard deviation of the error can be calculated.694904-2842-5073.2Atorney Docket No. N.057.W0.01
[0216] Methods herein can include calculating percent identity that can be calculated by determining the number of matched positions in aligned DNA sequences, dividing the number of matched positions by the total number of aligned DNA sequences, and multiplying by 100. A matched position refers to a position in which identical nucleotides occur at the same position in aligned DNA sequences. The percent identity over a particular length can be determined by counting the number of matched positions over that length and dividing that number by the length followed by multiplying the resulting value by 100. A non-limiting example for calculating the percent identity, can be, if (i) a 500-nucleotide DNA target sequence is compared to a subject DNA sequence, (ii) an alignment program presents 200 nucleotides from the target DNA sequence aligned with a region of the subject DNA sequence where the first and last nucleotides of that 200- nucleotide region are matches, and (iii) the number of matches over those 200 aligned nucleotides is 180, then the 500-nucleotide nucleic acid target sequence contains a length of 200 and a sequence identity over that length of 90 percent (i.e., 180, 200x100=90).
[0217] In some embodiments, the uniformity in DOR can be measured using standard methods such as. but not limiting to, DOR slope, normalized median depth of read (nmDOR), or breadth of read (BOR). DOR slope represents the slope of the line in the linear portion of a list of loci sorted in descending DOR order. Closer to zero is better, as it represents a flat line. In some embodiments, the uniformity in DOR can be measured using the percent of reads in the 90th-95thpercentile. For this measurement, the loci are sorted in descending DOR order. In illustrative embodiments, a DOR distribution using the 90th-95thpercentile contains 5 percent of reads. The reads of all loci between the 90thpercentile and 95thpercentile can be counted and divided by the total reads for all loci.
[0218] In some embodiments, the magnitude of the DOR slope can be less than 0.005. 0.001, 0.0005, 0.0001, 0.00005, 0.00001, 0.000005, or 0.000001. The magnitude of the DOR slope can be between 0 and 0.005, such as 0.000001 to 0.005, such as between 0.000005 to 0.00001, 0.00001 to 0.00005, 0.00005 to 0.0001, 0.0001 to 0.0005, 0.0005 to 0.001, or 0.001 to 0.005. The percent of reads in the 90th-95thpercentile can be between 0.2 and 9 percent, such as between 0.2 to 8 percent, 0.2 to 7 percent, 0.2 to 6 percent, 0.4 to 9 percent. 0.4 to 8 percent, 0.4 to 7 percent, 0.4 to 6 percent, 1 to 9 percent, 1 to 8 percent, 1 to 7 percent, 1 to 6 percent, 2 to 9 percent, 2 to 8 percent, 2 to 7 percent, 2 to 6 percent, 3 to 9 percent, 3 to 8 percent, 3 to 7 percent, 3 to 6 percent, 0.2 to 1.0 percent, 1 to 2 percent, 2 to 3 percent, 2 to 4 percent. 3 to 4 percent, 4 to 5 percent, 5 to 6704904-2842-5073.2Attorney Docket No. N.057.W0.01 percent, or 6 to 8 percent, or 7 to 9 percent. In some embodiments of methods herein, the method or the amplification steps in the method can produce a composition comprising at least 100 different amplicons (e.g., at least 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000. 28,000, 30,000. 40.000, 50,000. 75.000, or 100,000 non-identical amplicons) with the magnitude of the DOR slope in any of the ranges herein, or with a percent of reads in the 90th-95thpercentile in any of the ranges herein. In some embodiments, different amplicons can range in between 100 to 500,000, 100 to 400,000, 100 to 300,000, 100 to 200,000, 100 to 100,000, 100 to 75,000, 100 to 50,000, 100 to 40,000, 100 to 30,000, 100 to 25,000, 100 to 20,000, or 100 to 15.000 non-identical amplicons.
[0219] In some embodiments of methods herein, in addition, or, in some embodiments, as an alternative to analyzing an altered (increased or decreased) methylation levels in a sample, one or more other factors can be analyzed if desired. These factors can be used to increase the accuracy of the diagnosis (such as determining the presence or absence of cancer or an increased risk for cancer, classifying the cancer, or staging the cancer) or prognosis. These factors can also be used to select a particular therapy or treatment regimen that is likely to be effective in the subject.Limit of Detection
[0220] Exemplary methods herein are to detect target region amplicons generated by a targeted amplification and / or selectively enriched, and used to determine the status of one or more, in illustrative embodiments a plurality of, somatic variants or mutations or epigenetic changes on a DNA / RNA molecule in a nucleic acid sample. A target region can include 1, 2, 3, 4, 5, 6 or more sites that are predominantly mutated in cells of a certain origin (e.g. a cancer). In illustrative embodiments, sample nucleic acid molecules comprising the target regions are fragments of genomic nucleic acid or cell-free nucleic acid (cfNA) of a subject. Target regions can be selectively amplified and / or selectively enriched using methods herein. Exemplary methods herein, in some embodiments, have a mean or median limit of detection of as low as 1.0%, 0.5%, 0.1%, 0.05%. 0.01%, 0.005%, 0.002%, or 0.001%, wherein the method is capable of detecting somatic variants or mutations or epigenetic changes present at 1.0%, 0.5%, 0.1%, 0.05%, 0.01%, 0.005%, 0.002%, or 0.001% or more in a mixture of nucleic acid molecules.
[0221] In some embodiments, measurements can be adjusted for bias, such as bias due to differences in amplification efficiency or adjusted for sequencing errors. In some embodiments,714904-2842-5073.2Attorney Docket No. N.057.W0.01 differentiation between mutated and non-mutated samples can be analyzed using a control sample that is non-mutated, for example a healthy sample for normalization of quantitative results. In some embodiments, the differentiation between the mutated and non-mutated samples recited in this paragraph can be achieved after normalization of a detected and quantified signal using one or more (e.g. 2, 3, 4, 5, or 6) controls that do not have the mutation present.
[0222] In certain embodiments, the method is capable of detecting circulating tumor nucleic acids (ctNA) when the somatic variant or mutation or epigenetic change is present in 10%, 5%, 2.5%, 2%, 1%, 0.5%, 0.1%, 0.05%, 0.01%, or 0.001% of the cfNA in a sample. In some embodiments, methods herein detect or are capable of detecting ctNA from a sample when it is present at a range between 0.01% on the low end and 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5%, 2.5%, 1%, or 0.5% on the high end of the range of total cfNA in the sample. Methods herein, in some embodiments, detect or are capable of detecting 5%, 4%, 3%, 2%, 1%, 0.5%. 0.1%, 0.05%, 0.01% or less ctNA, as a percentage of total cfNA from a sample. The percentage of ctNA in a cfNA sample can be determined or approximated by VAF of one or more DNA / RNA variants, such as single nucleotide variants (“SNVs”), present in tumor cells but not in normal cells. (See e.g., W02019200228, incorporated by reference herein in its entirety). VAF can be used as a surrogate measure of the percentage of ctNA in the cfNA sample.Bioinformatics
[0223] Any of the embodiments disclosed herein may be implemented in digital electronic circuitry, integrated circuitry, specially designed ASICs (application- specific integrated circuits), computer hardware, firmware, software, or in combinations thereof. Apparatus of the presently disclosed embodiments can be implemented in a computer program product tangibly embodied in a machine-readable storage device for execution by a programmable processor; and method steps of the presently disclosed embodiments can be performed by a programmable processor executing a program of instructions to perform functions of the presently disclosed embodiments by operating on input data and generating output. The presently disclosed embodiments can be implemented advantageously in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.724904-2842-5073.2Atorney Docket No. N.057.W0.01Each computer program can be implemented in a high-level procedural or object-oriented programming language or in assembly or machine language if desired; and in any case, the language can be a compiled or interpreted language. A computer program may be deployed in any form, including as a stand-alone program, or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program may be deployed to be executed or interpreted on one computer or on multiple computers at one site, or distributed across multiple sites and interconnected by a communication network.
[0224] Computer readable storage media, as used herein, refers to physical or tangible storage (as opposed to signals) and includes without limitation volatile and non-volatile, removable and nonremovable media implemented in any method or technology for the tangible storage of information such as computer-readable instructions, data structures, program modules or other data. Computer readable storage media includes, but is not limited to, RAM, ROM, EPROM, EEPROM, flash memory or other solid state memory technology, CD-ROM, DVD, or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other physical or material medium which can be used to tangibly store the desired information or data or instructions and which can be accessed by a computer or processor.
[0225] Any of the methods described herein may include the output of data in a physical format, such as on a computer screen, or on a paper printout. In explanations of any embodiments elsewhere in this document, it should be understood that the described methods may be combined with the output of the actionable data in a format that can be acted upon by a physician. In addition, the described methods may be combined with the actual execution of a clinical decision that results in a clinical treatment, or the execution of a clinical decision to make no action. Some of the embodiments described in the document for determining genetic data pertaining to a target individual may be combined with a clinical decision or action. Some of the embodiments described in the document for determining genetic data pertaining to a target individual may be combined with the notification of a potential cancer status, or lack thereof, with a medical professional. Some of the embodiments described herein may be combined with the output of the actionable data, and the execution of a clinical decision that results in a clinical treatment, or the execution of a clinical decision to make no action.
[0226] For example, in some embodiments of the methods described herein, a system may be used that includes a computing system (which may be or may include one or more computing devices,734904-2842-5073.2Attorney Docket No. N.057.W0.01 co-located or remote to each other), a sequencing system (capable of, e.g., generating DNA sequencing data), an information system (such as a management or clinical information system that may record and / or provide information regarding samples or patients), vessels (e.g., for holding samples), sample processing system (e.g.. a system that may include robotic components for preparing, moving, and / or processing samples), and imaging and / or therapy systems (which may include e.g., treatment unit and imager I sensors for imaging patients using various imaging modalities, providing treatments such as radiotherapy, and / or performing other evaluation or therapeutic procedures based on what is learned from patient samples). The computing system (e.g., one or more computing devices) may be used to control and / or exchange signals and / or data with sequencing system, sample processing system, information system, and / or imaging and / or therapy systems, directly (e.g., through wireless and / or wired communication) or indirectly via another component of system (e.g., via any combination of wireless and / or wired communication). In certain embodiments, a computing system may be used to control and / or exchange data or other signals with a sequencing system, information system, sample processing system, and / or imaging and / or therapy systems. The computing system may include one or more processors and one or more volatile and / or non-volatile memories for storing computing code and data that are captured, acquired, recorded, and / or generated.
[0227] The computing system may include a controller that is configured to exchange control signals with sequencing system, information system, sample processing system, and / or imaging and / or therapy systems, and / or any components thereof, allowing the computing system to be used to control, for example, processing of samples, capture of images, acquisition of signals by sensors, positioning or repositioning of samples being tested or imaged, recording or obtaining other information, and / or running assays or tests.
[0228] A transceiver allows the computing system to exchange readings, control commands, and / or other data or signals, wirelessly or via wires, directly or indirectly via networking protocols, with, for example, other components of system. One or more user interfaces allow the computing system to receive user inputs (e.g., via a keyboard, touchscreen, microphone, camera, biometric scanners, etc.) and provide outputs (e.g., via a display screen, audio speakers, light emitters, etc.) with users. The computing system may additionally include one or more databases for storing, for example, data acquired from one or more systems or devices, signals acquired via one or more sensors, sequencing data, biomarkers, images, etc. In some implementations, database744904-2842-5073.2Attorney Docket No. N.057.W0.01(or portions thereof) may alternatively or additionally be part of another computing device that is co-located or remote (e.g., via “cloud computing”) and in communication with computing system, sequencing system, information system, sample processing system, imaging and / or therapy systems, and / or components thereof.
[0229] Imaging and / or therapy systems may include a treatment unit, which may include any systems and / or devices employed to administer a treatment to a patient, such as a radiation therapy system. Imager and / or sensors may be or may include, for example, any system or device that is involved in, for example, capturing images of patients and / or samples. Imagers may include detectors for visible light and / or light in any frequencies of interest, such as (but not limited to) the spectrum from infrared to ultraviolet. Imagers may employ any suitable optical components (e.g., lenses, mirrors, filters, beam splitters, prisms, diffusers, diffraction gratings, etc.), digital components (e.g., charge-coupled devices (CCDs), as well as an area for placement of samples, computing components (e.g., one or more processors, such as digital signal processors) to process and / or pre-process images, etc. Imagers may be, or may employ, any microscopes or microscopy systems (e.g.. confocal microscopy), tools, and / or techniques that will provide the desired imaging data. The imager and / or sensors may have the capability of receiving control signals from a computing device or system to, for example, initiate or cease image capture, and / or return images or other imaging data or status signals to the computing device or system (e.g., computing system). Imaging and / or therapy systems may include any tools for obtaining additional data about samples and / or patients (such as spectrometers, chromatographs, etc.). Sensors may be used to detect, for example, other aspects of samples and / or patients, such as temperature, humidity, location, etc.
[0230] Sample processing systems may include any components used to automate preparation of samples and running tests on samples. Robotics, for example, may include any combination of actuators, stepper motors, servomotors, control system, robot arms, end effectors like grippers and manipulators, etc. Robotics may be, or may comprise, for example, an automated robotics system with process automation software. In various embodiments, robotics may employ vessels, such as vials or plates, that can be moved and manipulated, for example, for positioning of samples to be evaluated using a sequencing system or components thereof. Sample processing systems may include components that may be used to maintain the integrity of samples during storage, transport, and testing, such as heaters, passive or active coolers, etc. Sensors may be used by754904-2842-5073.2Attorney Docket No. N.057.W0.01 sample processing systems to, for example, monitor, guide, and evaluate the progress of tests (e.g., by detecting temperature, fill level of vessels, or other states or conditions). Sequencing systems may include sequencing platforms or devices as well as various analytical tools used to process data from the sequencing platforms. Analytical tools may include software and / or hardware used to generate computer files with raw and / or processed sequencing data.
[0231] In various implementations, components of a system may be rearranged or integrated in other configurations. For example, a computing system (or components thereof) may be integrated with one or more of the sequencing system, sample processing system, and / or components thereof. The sequencing system, sample processing system, and / or components thereof may be directed to a vessel in which a sample can be situated (e.g., so as to test or evaluate biological samples). In various embodiments, the vessel may be movable (e.g., using any combination of motors, magnets, etc.) to allow for positioning and repositioning of samples (such as micro-adjustments for positioning of samples to be sequenced). It is also noted that not all components of a system are required to implement the disclosed approach, and in various embodiments, only a subset of the components of a system may be employed. For example, in various embodiments, a computing system may obtain and process data that was obtained via a sequencing system or another system that is or is not in direct communication with the computing system. In certain embodiments, data may be obtained for samples that were not processed in an automated fashion but rather manually by one or more users.
[0232] A data acquisition unit may retrieve, acquire, or otherwise obtain various data, such as sequencing data. The data acquisition unit may, for example, obtain data stored in a database, sequencing system, information system, sample processing system, and / or imaging and / or therapy systems. An interaction unit may interact (e.g., via user interfaces) with users (e.g.. laboratory technicians, data scientists, clinicians, etc.) to obtain information or commands needed for system or components thereof to function, or to start operations, repeat processes, and / or stop operations. In certain embodiments, the data acquisition unit may obtain data from users via interaction unit.
[0233] A preprocessor may perform any number of steps disclosed herein in preprocessing sequencing data. Machine learning platforms may be configured to train and update machine learning models, as further discussed herein. Machine learning platforms may, for example, employ certain machine learning techniques and algorithms to train and update predictive models. Machine learning platforms may include a training engine which may, for example, fit models to764904-2842-5073.2Attorney Docket No. N.057.W0.01 sequencing data. The training engine may, for example, obtain sequencing data from or via data acquisition unit, sequencing system or components thereof, and / or information system. Inference engines may apply models (e.g., models obtained from machine learning platform or components thereof) to generate outputs as disclosed herein. Biomarker generators may perform analyses on data from machine learning platforms to, for example, generate biomarkers to be used for various conditions or sub-populations.
[0234] A reporter may generate reports that include, for example, information on sample sequencing (e.g., via sequencing system or components thereof), imaging or administration of treatments (e.g., via imaging and / or therapy systems or components thereof and / or via information system), model training (e.g., via machine learning platform), and biomarker generation (e.g., via biomarker generator). For example, reporters may provide or identify trained and / or updated models, and / or biomarkers. In certain embodiments, reporters may obtain information to be reported from, for example, a database and / or information system. Data that may be reported may be, for example, transmitted to another system or device (which may or may not be part of system), saved in non-transitory computer- readable storage media (e.g.. database, information system, and / or elsewhere), and / or presented or otherwise provided to users (e.g., via a display device that is part of user interfaces).Methylation Amplification Error Model
[0235] In order to improve methylation detection from sequencing data, an error model may be incorporated into the methylation calling procedure that incorporates element specific priors from one or more steps of the workflow, for example from methylation preserving amplification and / or methylation conversion. In some embodiments, the element specific priors described herein may be determined based on published data. In some embodiments, the element specific priors described herein may be determined based on experimental analysis.
[0236] In some embodiments, the error model may incorporate priors based on particular motifs where DNMT1 may transfer methylation with a lower or higher likelihood. In some embodiments, the error model may incorporate priors based on motifs where DNMT1 may add a methyl group where a methyl group did not exist on the other strand. In some embodiments, the error model may incorporate priors based on the overall efficiency of DNMT1 (e.g., the fraction of mC that774904-2842-5073.2Atorney Docket No. N.057.W0.01 remain as mC after DNMT1 treatment, per cycle, regardless of motif). In some embodiments, the error model may incorporate priors based on bias of DNMT1 with respect to fragment length and number of CpG sites. In some embodiments, the error model may incorporate priors based on bias of DNMT1 with respect to mismatches (e.g., due to polymerase errors). In some embodiments, the error model may incorporate priors based on the overall non-specific activity of DNMT1 (e.g., the fraction of C that DNMT1 converts from C to mC inadvertently). In some embodiments, the error model may incorporate priors based on motifs where BST has lower copying efficiency, including methylation motifs (e.g., in instances where CmCGA is challenging for BST to amplify with a methyl group present). In some embodiments, the error model may incorporate priors based on motifs where BST make incorrect base insertions. In some embodiments, the error model may incorporate priors based on performance as a function of the methyl amplification time and temperature (e.g., long time and higher temperature may result in higher / different efficiency and motifs). In some embodiments, motifs may be 1 bp, 2, bp, 3 bp, 4 bp, 5 bp, or longer.
[0237] In some embodiments, the error model may incorporate any combination of the priors described herein depending on the workflow implemented. For example, EM-Seq and BS-Seq error models may each have bias towards specific sequences or template length.
[0238] In some embodiments, the methods described herein may include a position-based methylation error model. A cohort of healthy plasma samples can be used for error model construction. In some embodiments, 10 healthy samples, 50 healthy samples, 100 healthy samples, 200 healthy samples, 500 healthy samples, or 1000 healthy samples may be used for error model construction. In some embodiments, methylation transfer efficiency and error rates (spontaneous methylation) are estimated for each target in the panel of differentially methylated regions. Using a set of normal samples that are not expected to have any phenotype-related differential methylation, the per position efficiency and error rate per cycle can be estimated. The error model accounts for DNA input amount by adjusting the number of library prep cycles based on input amount.784904-2842-5073.2Attorney Docket No. N.057.W0.01Methylation Calling
[0239] In some embodiments, a methylation caller algorithm calculates a confidence score using the likelihoods generated from the error model and combining them with uniform priors across a grid of ctDNA amounts. In some embodiments, the methylation caller calculates a likelihood for each of the targets in a beta binomial model. The confidence score is then calculated by getting the maximum likelihood across that grid, and dividing it by that (max likelihood+negative likelihood).
[0240] Training Data: Di k=denotes
[0241] Test Data:
[0242] Result: Mutation call confidence scores for non-reference alleles in the test set for all bases 1,2,.. ,,B.
[0243] For i = 1,2, ..., B do1. Estimate efficiency and error from training data for base z, using the data Di k. Estimate methylation transfer and error rate parameters using training data: Transfer rate (p); Error rate (pe).2. Estimate starting copy for base i for test data at base i. Xo, the total number of starting fragments at a given base.
[0244] In some embodiments, an improved methylation caller algorithm as disclosed herein can be used. This algorithm provides an improved way of calculating confidence scores, rescaling it to significantly differentiate between positive and negative confidences, and allows the addition of non-uniform priors. This would result in more stable and more accurate confidences.
[0245] In some embodiments, in the improved variant caller algorithm, likelihoods are combined with priors, derived from training data ahead of time as described above, for smoother, more realistic and easily adjusted confidences.
[0246] Using the same error model as previously described, for each differentially methylated allele, at each position p, the likelihood of data at allele a, position p, MAF v LIK D(a, p) | v).794904-2842-5073.2Attorney Docket No. N.057.W0.01
[0247] Smoothed allele likelihood ratio: Computed by summing over MAF likelihoods scaled by MRD MAF priors.
[0248] Position likelihood ratio: Compute the weighted position likelihood ratio rat(p) assuming only one of the alleles at a position is the true differentially methylated allele.
[0249] Sample positive probability: Recursively compute the sample positive probability, as the probability of at least one target being positive, and find combined probability of sample being positive, taking into account methylation hotspots, via methylation priors, and control sample outcome via sample prior.
[0250] In some embodiments, sample prior is input to the algorithm and will be adjusted to achieve desired false positive, false negative and no call rate. Allele and position priors are calculated from MRD data positive prevalence and can always be adjusted given additional information. MAF priors are calculated from MAF rates of positive MRD samples.Mean Sample Methylated Allele Frequency (MAF)
[0251] Methods herein can include quantification of ctDNA in a sample, such as an MRD sample, by calculating a sample MAF. In some embodiments, only highly methylated CpG sites are used for calculating a mean sample MAF. In some embodiments, statistical weighting is further applied. In some embodiments, mean sample MAF may be determined by computing the weighted average of MAFs, where each differentially methylated site’s contribution is weighted by its corresponding prior probability, incorporating prior knowledge or assumptions about the likelihood of differentially methylated site’s presence in specific genes. In some embodiments, mean sample MAF may be determined by computing the weighted average of MAFs, where each differentially methylated site’s contribution is weighted by its corresponding gene posterior probability, reflecting the confidence in the variant calls based on posterior distribution. In some embodiments, the observed MAF for highly methylated CpG sites may undergo adjustment to account for the background error rate. In some embodiments, the background error rate may be derived from a position-based error model. In some embodiments, the average observed error is subtracted from the observed MAF to obtain an adjusted value. In some embodiments, the mean sample MAF is only calculated if a positive call is made for the sample.804904-2842-5073.2Attorney Docket No. N.057.W0.01SNV / Indel Error Model
[0252] In some embodiments, the methods described herein may include a position-based somatic variant error model, e.g., an SNV / indel error model. A cohort of healthy plasma samples can be used for error model construction. In some embodiments. 10 healthy samples, 50 healthy samples, 100 healthy samples, 200 healthy samples. 500 healthy samples, or 1000 healthy samples may be used for error model construction.
[0253] In some embodiments, per cycle PCR efficiency and error rates are estimated for each target in the panel. Using a set of normal samples that are not expected to have any cancer related mutation, the per position efficiency and error rate per cycle can be estimated. The error model accounts for DNA input amount by adjusting the number of library prep cycles based on input amount.Mean Sample VAF
[0254] Methods herein can include quantification of ctDNA in a sample, such as an MRD sample, by calculating a sample VAF. In some embodiments, only consensus variants are used for calculating a mean sample VAF. In some embodiments, statistical weighting is further applied. In some embodiments, mean sample VAF may be determined by computing the weighted average of variant VAFs, where each variant’s contribution is weighted by its corresponding gene prior probability, incorporating prior knowledge or assumptions about the likelihood of variant presence in specific genes. In some embodiments, mean sample VAF may be determined by computing the weighted average of variant VAFs, where each variant's contribution is weighted by its corresponding gene posterior probability, reflecting the confidence in the variant calls based on posterior distribution. In some embodiments, the observed VAF for consensus SNVs may undergo adjustment to account for the background error rate. In some embodiments, the background error rate may be derived from a position-based error model. In some embodiments, the average observed error is subtracted from the observed VAF to obtain an adjusted value. In some embodiments, the mean sample VAF is only calculated if a positive call is made for the sample.SNV Calling
[0255] In some embodiments, a variant caller algorithm calculates a confidence score using the likelihoods generated from the error model and combining them with uniform priors across a grid of ctDNA amounts. In some embodiments, the variant caller calculates a likelihood for each of the814904-2842-5073.2Atorney Docket No. N.057.W0.01 targets in a beta binomial model. The confidence score is then calculated by getting the maximum likelihood across that grid, and dividing it by that (max likelihood+negative likelihood).
[0256] Training Data: Di k= ^Ri k,RefAllelei,Ai>k, Ci i,Gi>k, ti k) where i E {1,2,denotes
[0257] Test Data:
[0258] Result: Mutation call confidence scores for non-reference alleles in the test set for all bases 1,2,.. ,,B.
[0259] For i = 1,2, ..., B do1. Estimate efficiency and error from training data for base i, using the data Di k. Estimate PCR parameters using training data: Replication rate (p); Error rate (pe2. Estimate starting copy for base i for test data at base i. Xo, the total number of starting fragments at a given base.
[0260] In some embodiments, an improved variant caller algorithm as disclosed herein can be used. This algorithm provides an improved way of calculating confidence scores, rescaling it to significantly differentiate between positive and negative confidences, and allows the addition of non-uniform priors. This would result in more stable and more accurate confidences.
[0261] In some embodiments, in the improved variant caller algorithm, likelihoods are combined with priors, derived from training data ahead of time as described above, for smoother, more realistic and easily adjusted confidences.
[0262] Using the same error model as previously described, for each mutation allele, at each position p, the likelihood of data at allele a, position p, VAF v LIK D(a, p)|v).
[0263] Smoothed allele likelihood ratio: Computed by summing over VAF likelihoods scaled by MRD VAF priors.
[0264] Position likelihood ratio: Compute the weighted position likelihood ratio rat(p) assuming only one of the alleles at a position is the true mutation allele.
[0265] Sample positive probability: Recursively compute the sample positive probability, as the probability of at least one target being positive, and find combined probability of sample being positive, taking into account hotspots, via position priors, and control sample outcome via sample prior.
[0266] In some embodiments, sample prior is input to the algorithm and will be adjusted to achieve desired false positive, false negative and no call rate. Allele and position priors are824904-2842-5073.2Atorney Docket No. N.057.W0.01 calculated from MRD data positive prevalence and can always be adjusted given additional information. VAF priors are calculated from VAF rates of positive MRD samples.InDei Calling
[0267] In some embodiments, an InDel-based sample caller combines position-based InDei calls with context-based InDei calls to generate InDel-based sample calls. In some embodiments, the sample calls generated by this caller are combined with the SNV calls to generate a combined sample-level call.
[0268] Position-based InDei caller
[0269] In some embodiments, position-based calling involves training a target specific error model for each InDei target that is encountered in the test sample. In one example, the error model is trained using a cohort of 200 healthy donor samples.
[0270] In some embodiments, the position-based caller comprises (1) target consolidation, (2) outlier detection using a beta binominal distribution, and (3) likelihood from tail probability.
[0271] In some embodiments, targets are consolidated to account for alignment differences that may result in incorrect characterization of target error rates, and outliers are detected using a beta binomial distribution.
[0272] The probability mass function (PMF) for the binomial distribution is:for k E {0,1, ... , n], n > 0, a > 0, b > 0, where B a, b)is the beta function.
[0273] where the beta-binomial distribution is a binomial distribution, whose probability of success p follows a beta distribution B(a, 0). A Beta distribution is fitted to the error model target VAFs, by computing the maximum likelihood estimates for a and 0.
[0274] Log likelihood ratio for the target is then computed using the following formula: LogLikelihoodRatio = -log(test tail probability / train tail probability)
[0275] Target posterior probability can then be calculated as:PosteriorProbability = eLLR / (1 + eLLR)
[0276] Finally, in some embodiments, only replicate consensus targets that are found to have sufficient support in two plasma replicates are called.Context-based InDei caller834904-2842-5073.2Attorney Docket No. N.057.W0.01
[0277] Insertions and deletions that do not occur in a repeat context are likely to have very low rates of background error. The relatively small number of training samples used to fit the error model is unable to characterize the background error profile for such targets, resulting in poor fit, or fit failure. For such targets a context-based calling strategy is used.
[0278] In some embodiments, the context-based caller comprises features to fit a Beta-Binomial regression model, which is used to compute target confidence.
[0279] The context-based caller is primarily used to call less noisy targets where position-based calling is not feasible.Ensemble InDei caller or combined position and context InDei caller
[0280] In some embodiments, the ensemble caller combines position- and context-based InDei calls to generate a single InDei target call.
[0281] In some embodiments, the ensemble caller uses the position-based call for InDei targets when the Beta-Binomial background error fit is successful and meets goodness of fit (GoF) criteria.
[0282] In some embodiments, the ensemble caller uses the context-based call for InDei targets when Beta-Binomial background error fit fails, Beta-Binomial fit does not meet GoF criteria, or the number of non-zero mutant DOR error models samples is below threshold, i.e., only use context-based calling for low noise targets.Sample InDei calling
[0283] In some embodiments, InDei sample calling uses an approach very similar to SNV sample caller, with primary difference being InDei specific WES / WGS priors, which are generated using filtered InDei targets for the same CRC sample cohort as was used for SNVs.
[0284] The sample caller algorithm recursively computes the sample likelihood ratio incorporating target likelihoods and target WES / WGS priors.Consensus Calling
[0285] PCR and sequencing artifacts are a result of random process noise and are expected to be discordant between library pools. Accordingly, consensus calling helps filter a significant fraction of such discordant targets.
[0286] In some embodiments, the methods described herein are used to filter targets that are likely a result of process noise and generate a candidate target set that will be input to the sample caller844904-2842-5073.2Atorney Docket No. N.057.W0.01 to compute the sample posterior positive probability. Because the sample level call is made by the sample caller by integrating target WES / WGS prior data, this version of the target caller is intentionally tuned to be more permissive than previous versions. The primary objective here is to filter targets lacking sufficient support or likely to be a result of random process noise.
[0287] In some embodiments, the variant caller described herein uses data likelihoods, target VAFs and MRD VAF priors, to first call targets separately for each library pool, followed by a combined call using both library pools.
[0288]
[0289] CH Mutation Filtering
[0290] A significant number of false positive calls have the potential to be CH (clonal hematopoiesis) mutations, including CHIP (clonal hematopoiesis of indeterminate potential) mutations. It is expected that mutations that have CH origin will be detected in matched whole blood or buffy coat data. Accordingly, in some embodiments, the methods described herein implement a reflex strategy to filter CH mutations to provide a significant improvement in specificity of the caller.
[0291] In some embodiments, the reflex strategy involves filtering targets called in plasma that are likely to be CH targets. In some embodiments, evidence of variants or potential variants identified at this step are the input to a sample caller.
[0292] All patents and publications mentioned in the specification are indicative of the levels of skill of those skilled in the art to which the invention pertains. References cited herein are incorporated by reference herein in their entirety to indicate the state of the art as of their publication or filing date and it is intended that this information can be employed herein, if needed, to exclude specific aspects that are in the prior art. For example, when composition of matter are claimed, it should be understood that compounds known and available in the art prior to Applicant's invention, including compounds for which an enabling disclosure is provided in the references cited herein, are not intended to be included in the composition of matter claims herein.
[0293] One of ordinary skill in the art will appreciate that starting materials, biological materials, reagents, synthetic methods, purification methods, analytical methods, assay methods, and854904-2842-5073.2Attorney Docket No. N.057.W0.01 biological methods other than those specifically exemplified can be employed in the practice of the invention without resort to undue experimentation. All art-known functional equivalents of any such materials and methods are intended to be included in this invention. The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention has been specifically disclosed by illustrative aspects and optional features, modification and variation of the concepts herein disclosed may be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention as defined by the appended claims.
[0294] The disclosed embodiments, examples and experiments are not intended to limit the scope of the disclosure or to represent that the experiments below are all or the only experiments performed. Efforts have been made to ensure accuracy with respect to numbers used (e.g., amounts, temperature, etc.) but some experimental errors and deviations should be accounted for. It should be understood that variations in the methods as described may be made without changing the fundamental aspects that the experiments are meant to illustrate.
[0295] Those skilled in the art can devise many modifications and other embodiments within the scope and spirit of the present disclosure. Indeed, variations in the materials, methods, drawings, experiments, examples, and embodiments described may be made by skilled artisans without changing the fundamental aspects of the present disclosure. Any of the disclosed embodiments can be used in combination with any other disclosed embodiment.
[0296] In some instances, some concepts have been described with reference to specific embodiments. However, one of ordinary skill in the art appreciates that various modifications and changes can be made without departing from the scope of the invention as set forth in the claims below. Accordingly, the specification and figures are to be regarded in an illustrative rather than a restrictive sense, and all such modifications are intended to be included within the scope of invention.The methods disclosed herein can be used in virtually any application involving a small amount of starting nucleic acid molecules (of either a single origin or mixed origins) from a subject (e.g. in a liquid biopsy). The methods disclosed herein can be used with multiple downstream applications864904-2842-5073.2Attorney Docket No. N.057.W0.01 and are especially advantageous in detecting a small amount of target nucleic acid molecules mixed with contaminating nucleic acid molecules. In some embodiments, the target nucleic acids and the contaminating nucleic acids may be from individuals who are genetically related. For example, genetic abnormalities in a fetus (target) may be detected from maternal plasma which contains fetal (target) nucleic acids and also maternal (contaminating) nucleic acids; the abnormalities include whole chromosome abnormalities (e.g. aneuploidy) partial chromosome abnormalities (e.g. deletions, duplications, inversions, translocations), polynucleotide polymorphisms (e.g. STRs), single nucleotide polymorphisms, and / or other genetic abnormalities or differences. In some embodiments, the target nucleic acids and the contaminating nucleic acids may be from individuals who may not be genetically related, such as an organ donor and a transplant recipient. For example, organ rejection may be detected from recipient plasma which contains donor (target) nucleic acids and also recipient (contaminating) nucleic acids. In some embodiments, the target and contaminating nucleic acids may be from the same individual, but where the target and contaminating DNA or RNA are different by one or more mutations or epigenetic alterations, for example in the case of cancer. The methods herein can be used to diagnose, detect, or monitor a disease (whether symptomatic, pre-symptomatic, or asymptomatic), phenotype, or condition, including for example cancer or cancer-recurrence detection / diagnosis, other disease detection / diagnosis, therapy selection, non-invasive prenatal testing (NIPT), organ health or organ transplant evaluation, prediction and monitoring, and veterinary applications (e.g. aging or health status evaluation, prediction and monitoring of domesticated animals).
[0297] The following non-limiting examples are provided purely by way of illustration of exemplary embodiments, and in no way limit the scope and spirit of the present disclosure. Furthermore, it is to be understood that any inventions disclosed or claimed herein encompass all variations, combinations, and permutations of any one or more features described herein. Any one or more features may be explicitly excluded from the claims even if the specific exclusion is not set forth explicitly herein. It should also be understood that disclosure of a reagent for use in a method is intended to be synonymous with (and provide support for) that method involving the use of that reagent, according either to the specific methods disclosed herein, or other methods known in the art unless one of ordinary skill in the art would understand otherwise. In addition, where the specification and / or claims disclose a method, any one or more of the reagents disclosed874904-2842-5073.2Attorney Docket No. N.057.W0.01 herein may be used in the method, unless one of ordinary skill in the art would understand otherwise.EXAMPLESExample 1. Sequential Library AmplificationsOverview
[0298] Library preparation was performed on a set of 8 cfDNA samples using 5' biotinylated adaptors. After library amplification with standard non-biotinylated primers, the original biotinylated library molecules were captured on streptavidin beads. The amplified library in the supernatant was subject to targeted multiplex PCR using a pool of target- specific primers targeting certain SNV loci to obtain a first library. Beads with captured original molecules were then used to re-amplify the library using the same amplification protocol, and the second amplified library in the supernatant was again subject to targeted multiplex PCR using the same pool of target- specific primers targeting the same SNV loci to obtain a second library. The separately barcoded replicate libraries were pooled and sequenced and the data analyzed to detect expected SNVs in positive samples, and to determine the background error rates at positions of interest. SNV detection and background error rates were compared between the two replicate libraries.Experiment
[0299] 8 healthy cfDNA samples and one MNase-treated DNA sample from cell line HCT-116 were used in the experiment. The healthy cfDNA samples were not expected to contain SNVs, and the HCT-116 cell line contains 2 mutations (KRAS G13D, a C>T mutation present at 54% VAF and PIK3CA H1047R, an A>G mutation present at 37% VAF in the cell line).
[0300] The MNased HCT-116 cell-line DNA was spiked into cfDNA samples 7 and 8 at -2% by mass, targeting a VAF of -1% for the KRAS mutation and -0.75% for the PIK3CA mutation.
[0301] 33 ng of each cfDNA sample was used in library preparation using biotinylated adaptors. The Y-adaptors are biotinylated on the 5' end of the adaptor arm in the single-stranded region, and in this way each original cfDNA molecule that is converted to library will be 5' biotinylated.
[0302] The library was amplified using 12 amplification cycles. Following library amplification, the original cfDNA molecules remained biotinylated, and cfDNA copy molecules were not884904-2842-5073.2Attorney Docket No. N.057.W0.01 biotinylated. Streptavidin beads were used after library amplification to retrieve the biotinylated molecules, and amplicons in the supernatant were removed for use in a targeted multiplex PCR reaction as discussed below.
[0303] The beads containing the bound biotinylated library molecules were then used for a second library amplification with on-bead PCR, with otherwise the same protocol as the first library amplification except using 13 amplification cycles. The supernatant (containing an independently amplified library from the same starting material) was removed and used for a second targeted multiplex PCR reaction.
[0304] The first and second library amplification materials were used in parallel targeted multiplex PCR reactions using the same pool of target-specific primers and separate barcoding primers. Following purification, the two barcoded target-enriched replicate libraries were combined at equimolar amounts, and sequenced on one NovaSeq 6000 SP flow cell.Discussion
[0305] As shown in FIG. 2, DOR averaged >100,000 and was similar for most of the samples between the two replicate libraries.
[0306] The sequence data was further analyzed to determine error rates after removing spiked-in targets from samples 7-8. As shown in FIG. 3, error rates were observed to be similar between the two replicate libraries across each base substitution.
[0307] Target SNV calls were made for each replicate using a relaxed variant-calling threshold. Candidate SNV calls contributed to the sample caller only if identified in both replicates (i.e., consensus calls, see FIG. 4). Additional filters were applied within the sample caller before reporting positive SNV calls.
[0308] As shown in FIG. 5, the two spiked-in targets were both called positive in both spiked-in samples (M7. M8) but not in healthy, non-spiked-in samples (Ml - M6) for each of the two replicate libraries.
[0309] An exemplary novel SNV-detection workflow has been demonstrated, comprising biotinylated library prep for tagging of the original DNA molecules in a sample, followed by sequential library amplification steps directly from the original molecules captured on streptavidin beads, and independent targeted multiplex PCR reactions using the separate library amplified material for SNV detection and error correction. The workflow generates similar DORs for the894904-2842-5073.2Atorney Docket No. N.057.W0.01 first and second amplification samples, it demonstrates SNV error correction by relying on consensus target calls between the two amplification samples, and it successfully detects spiked-in SNVs in the VAF range of 1% and below. The workflow improved SNV detection by reducing the number of library prep reactions per sample and allowing a higher cfDNA input amount per library prep (entire cfDNA amount available), compared to splitting a sample into multiple aliquots.
[0310] This exemplary workflow shows that a method for tagging and retrieving the original cfDNA molecules from a sample can be used to perform sequential library amplification reactions from the same original starting material, which can be used for error correction in SNV analyses. The original library molecules can also be used for additional analyses, such as methylation analysis.* *904904-2842-5073.2
Claims
Attorney Docket No. N.057.W0.01What is claimed is:
1. A method for amplifying and sequencing nucleic acids, comprising:(a) extracting nucleic acids from a sample of a subject to obtain extracted nucleic acids;(b) ligating an adaptor having a capture moiety to the extracted nucleic acids to obtain adapted nucleic acids;(c) performing a first amplification on the adapted nucleic acids and generating first amplified DNA, and separating the adapted nucleic acids from the first amplified DNA;(d) performing a second amplification on the adapted nucleic acids and generating second amplified DNA, and separating the adapted nucleic acids from the second amplified DNA;(e) performing a first targeted enrichment on the first amplified DNA or its derivative to enrich a plurality of target loci each encompassing at least one alteration associated with a phenotype and generating first enriched DNA;(f) performing a second targeted enrichment on the second amplified DNA or its derivative to enrich the plurality of target loci and generating second enriched DNA; and(g) sequencing the first enriched DNA and the second enriched DNA or their respective derivative and producing a first set of sequence reads and a second set of sequence reads, respectively.
2. The method of claim 1, further comprising detecting at least one alteration associated with the phenotype present in both the first enriched DNA and the second enriched DNA based on the first set of sequence reads and the second set of sequence reads.
3. The method of any of claims 1-2, wherein the extracted nucleic acids comprise DNA or RNA.
4. The method of any of claims 1-3, wherein the extracted nucleic acids comprise cell-free DNA.
5. The method of any of claims 1-4, wherein the capture moiety comprises biotin, biotin dT, biotin-TEG, photocleavable (PC) biotin, or desthiobiotin-TEG, or biotin azide; and wherein the separating is performed using streptavidin beads.914904-2842-5073.2Attorney Docket No. N.057.W0.
016. The method of any of claims 1-4, wherein the capture moiety comprises click chemistry capture moiety.
7. The method of any of claims 1-6, further comprising (dl) performing a third amplification on the adapted nucleic acids and generating third amplified DNA, and separating the adapted nucleic acids from the third amplified DNA, and optionally repeating step (dl) to generate additional amplified DNA.
8. The method of any of claims 1-7, wherein the adaptor comprises a universal priming sequence, and wherein the adapted nucleic acids is amplified using a primer that binds to the universal priming sequence.
9. The method of any of claims 1-8, wherein an error in amplification, enrichment, or sequencing is identified from the first set of sequence reads in combination with the second sets of sequence reads without using any molecular barcode.
10. The method of any of claims 1-9. wherein the method does not comprise tagging the extracted nucleic acids with a molecular barcode.
11. The method of any of claims 1-10, wherein the plurality of target loci comprises 25-5,000 target loci.
12. The method of any of claims 1-11, wherein performing targeted enrichment comprises performing targeted probe capture to enrich the target loci.
13. The method of any of claims 1-11, wherein performing targeted enrichment comprises performing targeted multiplex amplification to enrich the target loci.
14. The method of claim 13, wherein the targeted multiplex amplification comprises amplification of 50-500 target loci together in the same reaction volume.
15. The method of any of claims 1-14, wherein the method does not comprise whole genome sequencing or whole exome sequencing of a tumor tissue sample of the subject.
16. The method of any of claims 1-15, wherein the identification of the at least one alteration associated with cancer that is present in both the first enriched DNA and the second enriched DNA is indicative of minimal residual disease.924904-2842-5073.2Attorney Docket No. N.057.W0.0117. The method of any of claims 1-16, wherein at least one alteration associated with cancer is identified in both the first enriched DNA and the second enriched DNA, and the method further comprises: extracting cellular DNA from a whole blood or buffy coat fraction of the blood sample; performing targeted enrichment on the extracted cellular DNA or their derivative to enrich a subset of the plurality of target loci and generating leukocyte-derived enriched DNA, wherein the subset comprises the at least one alteration; and sequencing the leukocyte-derived enriched DNA or their derivative to determine whether the at least one alteration is also present in the leukocyte-derived enriched DNA.
18. The method of claim 17, further comprises identifying at least one alteration that is not present in the leukocyte-derived enriched DNA, thereby identifying at least one alteration associated with the phenotype.
19. The method of claim 17, further comprises identifying at least one alteration that is also present in the leukocyte-derived enriched DNA, thereby identifying at least one clonal hematopoiesis (CH) alteration.
20. The method of claim 17, further comprises identifying at least one germline mutation present in the leukocyte-derived enriched DNA.
21. The method of any of claims 1-20, wherein the alteration associated with the phenotype comprises a genetic variant.
22. The method of claim 21, wherein the alteration associated with the phenotype comprises a single nucleotide variant (SNV), a multi-nucleotide variant (MNV), an indel, a gene fusion, a structural variant, or a combination thereof.
23. The method of any of claims 1-20, wherein the alteration associated with the phenotype comprises an epigenetic alteration.
24. The method of claim 23, wherein the alteration associated with the phenotype comprises a change in methylation status.
25. The method of any of claims 1-24, wherein the cancer is a solid tumor.934904-2842-5073.2Attorney Docket No. N.057.W0.0126. The method of claim 25, wherein the solid tumor is breast cancer, advanced adenoma, colorectal cancer, kidney cancer, liver cancer, lung cancer, multiple myeloma, ovarian cancer, or pancreatic cancer.
27. The method of any of claims 1-26, wherein the subject has been treated with surgery, first- line chemotherapy, and / or adjuvant therapy.
29. The method of claim 27, further comprises longitudinally collecting a plurality of blood samples from the subject and repeating steps (a) to (g) for each of the plurality of blood samples.
30. A method for amplifying and sequencing nucleic acids, comprising:(a) extracting nucleic acids from a sample of a subject to obtain extracted nucleic acids;(b) ligating an adaptor having a capture moiety to the extracted nucleic acids to obtain adapted nucleic acids;(c) preparing a first library of amplified DNA comprising first amplified DNA by performing a first amplification on the adapted nucleic acids and separating the adapted nucleic acids from the first amplified DNA;(d) sequencing at least a portion of the first library of amplified DNA or its derivative and generating a first set of sequence reads;(e) preparing a second library of amplified DNA comprising second amplified DNA by performing a second amplification on the adapted nucleic acids and separating the adapted nucleic acids from the second amplified DNA; and(f) sequencing at least a portion of the second library of amplified DNA or its derivative and generating a second set of sequence reads.
31. A method of amplifying and sequencing DNA, comprising:(a) extracting cell-free DNA from a liquid sample from a subject to generate extracted DNA;(b) ligating an adaptor having a capture moiety to the extracted DNA to obtain adapted DNA;(c) preparing a first library of methylation-preserved amplified DNA having one or more methylation-transferred DNA molecules, wherein said first library of methylation-preserved amplified DNA is prepared by (i) performing a first copying of one or more template DNA molecules in the adapted DNA and generating one or more copied DNA molecules, (ii)944904-2842-5073.2Atorney Docket No. N.057.W0.01 transferring methylation status from the one or more template DNA molecules to the one or more copied DNA molecules and generating one or more methylation-transferred copied DNA molecules, (iii) optionally repeating the copying and the transferring one or more rounds, and (iv) separating the methylation-transferred copied DNA molecules from the adapted DNA;(d) treating the first library of methylation-preserved amplified DNA or a portion thereof with an agent or a combination of agents and generating treated first methylation-preserved amplified DNA, wherein the agent or the combination of agents discriminates between methylated and unmethylated cytosines;(e) performing a first targeted enrichment on the treated first methylation-preserved amplified DNA or its derivative to enrich a first plurality of differentially methylated regions that are differentially methylated in a phenotype and generating first enriched DNA;(f) sequencing the first enriched DNA or its derivative and generating a first set of sequence reads.
32. The method of claim 31, further comprising;(g) preparing a second library of methylation-preserved amplified DNA having one or more methylation-transferred DNA molecules, wherein said second library of methylation- preserved amplified DNA is prepared by (i) performing a second copying of one or more template DNA molecules in the adapted DNA and generating one or more copied DNA molecules, (ii) transferring methylation status from the one or more template DNA molecules to the one or more copied DNA molecules and generating one or more methylation-transferred copied DNA molecules, (iii) optionally repeating the copying and the transferring one or more rounds, and (iv) separating the methylation-transferred copied DNA molecules from the adapted DNA;(h) treating the second library of methylation-preserved amplified DNA or a portion thereof with an agent or a combination of agents and generating treated second methylation- preserved amplified DNA, wherein the agent or the combination of agents discriminates between methylated and unmethylated cytosines;(i) performing a second targeted enrichment on the treated second methylation-preserved amplified DNA or its derivative to enrich the first plurality of differentially methylated regions that are differentially methylated in a phenotype and generating second enriched DNA;(j) sequencing the second enriched DNA or its derivative and generating a second set of954904-2842-5073.2Attorney Docket No. N.057.W0.01 sequence reads;(k) identifying the presence of or the likelihood of the presence of the phenotype based on the first set of sequence reads and the second set of sequence reads.
33. The method of claim 31, further comprising:(g) preparing a second library of methylation-preserved amplified DNA having one or more methylation-transferred DNA molecules, wherein said second library of methylation- preserved amplified DNA is prepared by (i) performing a second copying of one or more template DNA molecules in the adapted DNA and generating one or more copied DNA molecules, (ii) transferring methylation status from the one or more template DNA molecules to the one or more copied DNA molecules and generating one or more methylation-transferred copied DNA molecules, (iii) optionally repeating the copying and the transferring one or more rounds, and (iv) separating the methylation-transferred copied DNA molecules from the adapted DNA;(h) treating the second library of methylation-preserved amplified DNA or a portion thereof with an agent or a combination of agents and generating treated second methylation- preserved amplified DNA, wherein the agent or the combination of agents discriminates between methylated and unmethylated cytosines;(i) performing a second targeted enrichment on the treated second methylation-preserved amplified DNA or its derivative to enrich a second plurality of target loci each encompassing at least one alteration associated with the phenotype and generating second enriched DNA, wherein the second plurality of target loci optionally comprise at least a subset of the first plurality of differentially methylated regions that are differentially methylated in the phenotype;(j) sequencing the second enriched DNA or its derivative and generating a second set of sequence reads;(k) identifying the presence of or the likelihood of the presence of the phenotype based on the first set of sequence reads and the second set of sequence reads.
34. The method of any of claims 31-33, wherein the capture moiety comprises biotin, biotin dT, biotin-TEG, photocleavable (PC) biotin, or desthiobiotin-TEG, or biotin azide; and wherein the separating is performed using streptavidin beads.
35. The method of any of claims 31-33, wherein the capture moiety comprises click chemistry capture moiety.964904-2842-5073.2Attorney Docket No. N.057.W0.0136. The method of any of claims 31-35, further comprising (gl) preparing a third library of methylation-preserved amplified DNA having one or more methylation-transferred DNA molecules, wherein said third library of methylation-preserved amplified DNA is prepared by (i) performing a third copying of one or more template DNA molecules in the adapted DNA and generating one or more copied DNA molecules, (ii) transferring methylation status from the one or more template DNA molecules to the one or more copied DNA molecules and generating one or more methylation-transferred copied DNA molecules, (iii) optionally repeating the copying and the transferring one or more rounds, and (iv) separating the methylation-transferred copied DNA molecules from the adapted DNA; and optionally repeating step (gl) to generate an additional library of methylation-preserved amplified DNA having one or more methylation-transferred DNA molecules.
37. The method of any of claims 31-36, wherein the copying comprises isothermal extension.
38. The method of claim 37, wherein the adaptor comprises a universal priming sequence, and wherein the isothermal extension is performed using a primer that binds to the universal priming sequence.
39. The method of claim 38, wherein the primer comprises a locked nucleic acid (LNA) oligonucleotide or a 5’ linked primer.
40. The method of any of claims 31-39, wherein the adaptor comprises a molecular barcode.
41. The method of any of claims 31-40, wherein the adaptor comprises at least one methylated cytosine and at least one unmethylated cytosine.
42. The method of any of claims 31-41, wherein the methylation status is transferred using a methyltransferase.
43. The method of any of claims 31-42, wherein the agent or the combination of agents that discriminates between methylated and unmethylated cytosines comprises an deaminating agent or a combination of an oxidizing agent and a deaminating agent, wherein the deaminating agent is sodium bisulfite or a deaminase.
44. The method of any of claims 31-43, wherein an error in isothermal amplification, methyltransferase treatment, bisulfite or enzymatic conversion, enrichment or sequencing is974904-2842-5073.2Attorney Docket No. N.057.W0.01 identified from the first set of sequence reads in combination with the second sets of sequence reads without using any molecular barcode.
45. The method of any of claims 31-44, wherein the panel of differentially methylated regions comprises 50-50,000 differentially methylated regions.
46. The method of any of claims 31-45, wherein performing targeted enrichment comprises performing targeted probe capture to enrich the panel of differentially methylated regions.
47. The method of any of claims 31-45, wherein performing targeted enrichment comprises performing targeted multiplex amplification to enrich the panel of differentially methylated regions.
48. The method of claim 47, wherein the targeted multiplex amplification comprises amplification of 50-20,000 target loci together in the same reaction volume.
49. The method of any of claims 31-48, wherein the liquid sample is a blood, plasma, serum, or urine sample.
50. The method of any of claims 31-49, wherein the sample comprises cell-free DNA from a tumor, a transplant, or a fetus.51 . The method of any of claims 31-50, wherein the disease is a cancer.
52. The method of claim 51 , wherein the cancer is breast cancer, advanced adenoma, colorectal cancer, kidney cancer, liver cancer, lung cancer, multiple myeloma, ovarian cancer, or pancreatic cancer.
53. The method of claim 51 or 52, wherein the subject has been treated with surgery, first-line chemotherapy, and / or adjuvant therapy.
54. The method of claim 53, further comprises longitudinally collecting a plurality of liquid samples from the subject and repeating steps (a) to (f) for each of the plurality of blood samples.
55. The method of claim 54, wherein the identification of one or more differentially methylated regions associated with cancer that is present in both the first enriched DNA and the second enriched DNA is indicative of minimal residual disease.984904-2842-5073.2
Citation Information
Patent Citations
Methods for simultaneous amplification of target loci
US11312996B2
Linked duplex target capture
WO2017168332A1
Compositions, methods, and kits for isolating nucleic acids
WO2018156418A1
Methods for isolating nucleic acids with size selection
WO2019161244A1
Linked target capture and ligation
WO2020039261A1