Methods of nucleic acid preparation and analysis
The method of extracting and enriching cell-free DNA from plasma to identify somatic mutations addresses the invasiveness and sensitivity issues of current cancer detection methods, providing accurate and early cancer relapse monitoring.
Patent Information
- Application Number
- PCT/US2025/038160
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-19
- Filing Date
- 2025-07-17
- Publication Date
- 2026-01-22
AI Technical Summary
Current methods for detecting cancer, particularly cancer relapse or metastasis, are invasive, lack sensitivity, and produce PCR and sequencing artifacts, making early detection of asymptomatic cancer challenging and prone to errors.
A method involving the extraction of cell-free DNA from plasma, targeted enrichment of specific loci associated with somatic mutations, and sequencing to identify cancer-specific variants, optionally combined with leukocyte DNA analysis to confirm tumor-specific mutations, while avoiding molecular barcodes to reduce errors.
Enables accurate, less invasive detection of somatic mutations associated with cancer, facilitating early detection of asymptomatic cancers and monitoring relapse or metastasis, reducing noise and improving sensitivity.
Smart Images

Figure US2025038160_22012026_PF_FP_ABST
Abstract
Description
Attorney Docket No. N.053.WO.01 METHODS OF NUCLEIC ACID PREPARATION AND ANALYSIS CROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefits of U.S. Provisional Application No. 63 / 673,519 filed July 19, 2024, the contents of which are hereby incorporated by reference in their entirety. TECHNICAL FIELD
[0002] The present disclosure relates generally to methods for preparing and analyzing DNA molecules, and, more particularly, to methods for preparing and analyzing DNA molecules that include one or more somatic mutations. Such methods may be used alone or in combination with other methods for preparing and analyzing nucleic acid molecules that include one or more other genetic or epigenetic mutations or other analytes or biomarkers indicative of cancer or other disease or biological state. BACKGROUND
[0003] The present disclosure relates generally to methods for preparing and analyzing DNA molecules, and, more particularly, to methods for preparing and analyzing DNA molecules that include one or more somatic mutations. Detection of cancer, cancer relapse, or cancer metastasis has traditionally relied on imaging and tissue biopsy. The biopsy of tumor tissue is invasive and carries risk of potentially contributing to metastasis or surgical complications, while imaging- based detection is not sufficiently sensitive. Moreover, neither method is suited for detecting asymptomatic cancer, or detecting relapse or metastasis at an early stage. Improved and less invasive methods are needed for early detection and for detecting relapse or metastasis of cancers. Furthermore, a large number of PCR and sequencing artifacts are seen in methods of detecting somatic mutations when a large number of targets are analyzed. Accordingly, there is a need for error and noise reduction for accurate cancer or other disease or biological state detection. Methods described herein address these needs. 1 4934-6677-8453.1Attorney Docket No. N.053.WO.01 SUMMARY
[0004] One aspect of the invention described herein relates to a method for preparing a plurality of non-naturally occurring compositions of amplified DNA from a blood sample of a subject, comprising: (a) extracting cell-free DNA from a plasma fraction of a blood sample of a subject; (b) performing targeted enrichment on the extracted cell-free DNA or DNA derived therefrom to enrich a panel of a plurality of target loci each encompassing at least one somatic mutation associated with the cancer to obtain a non-naturally occurring composition of plasma-derived amplified DNA; (c) sequencing the non-naturally occurring composition of plasma-derived amplified DNA to identify at least one somatic variant that is present in the plasma-derived amplified DNA; and optionally, if at least one somatic variant present in the plasma-derived amplified DNA is identified in step (c), the method further comprises: (d) extracting cellular DNA from a fraction of the blood sample containing a plurality of leukocytes; (e) performing targeted enrichment on the extracted cellular DNA or DNA derived therefrom to enrich a subset of the panel of the plurality of target loci to obtain a non-naturally occurring composition of leukocyte- derived amplified DNA, wherein the subset comprises the at least one somatic variant identified in the plasma-derived amplified DNA; and (f) sequencing the non-naturally occurring composition of leukocyte-derived enriched DNA to determine whether the at least one somatic variant identified in the plasma-derived enriched DNA is also present in the leukocyte-derived enriched DNA, thereby identifying at least one somatic variant that is present in the plasma-derived enriched DNA and that is not present in the leukocyte-derived enriched DNA.
[0005] In some embodiments, the at least one somatic variant that is present in the plasma-derived enriched DNA and that is not present in the leukocyte-derived enriched DNA is a somatic mutation associated with the cancer. In some embodiments, performing targeted enrichment comprises targeted multiplex amplification or targeted probe capture on the extracted cell-free DNA or DNA derived therefrom to enrich the panel of the plurality of target loci.
[0006] In some embodiments, the plurality of target loci comprises 25-5,000 target loci. In some embodiments, the method further comprises identifying at least one somatic variant present in the plasma-derived enriched DNA that is also present in the leukocyte-derived enriched DNA, thereby identifying a clonal hematopoiesis (CH) mutation.
[0007] In some embodiments, the blood sample is split into a plurality of sub-samples prior to extracting cell-free DNA from the plasma fraction and extracting cellular DNA from the fraction 2 4934-6677-8453.1Attorney Docket No. N.053.WO.01 of the sample containing the plurality of leukocytes. In some embodiments, steps (b) and (c) are performed on at least two of the sub-samples, wherein sequencing reads from the at least two sub- samples are combined to identify at least one somatic variant that is present in the plasma-derived enriched DNA of each of the at least two sub-samples.
[0008] In some embodiments, the plasma fraction of the blood sample is split into a plurality of plasma sub-samples prior to extracting cell-free DNA from the plasma fraction. In some embodiments, steps (b) and (c) are performed on at least two of the plasma sub-samples, wherein sequencing reads from the at least two plasma sub-samples are combined to identify at least one somatic variant that is present in the plasma-derived enriched DNA of each of the at least two plasma sub-samples.
[0009] In some embodiments, the cell-free DNA extracted from the plasma fraction is split into a plurality of cell-free DNA sub-samples and steps (b) and (c) are performed on at least two of the cell-free DNA sub-samples, wherein sequencing reads from the at least two cell-free DNA sub- samples are combined to identify at least one somatic variant that is present in the plasma-derived enriched DNA of each of the at least two cell-free DNA sub-samples.
[0010] One aspect of the invention described herein relates to a method for preparing a plurality of non-naturally occurring compositions of amplified DNA from a blood sample of a subject, comprising: (a) extracting cell-free DNA from a plasma fraction of a blood sample of a subject; (b) adding adaptors to a first portion and a second portion of the extracted cell-free DNA and generating a first portion and a second portion of adapted DNA, and wherein the adaptors are added to the first portion and the second portion of the extracted cell-free DNA under the same conditions; (c) performing a first targeted enrichment on the first portion of the adapted DNA or DNA derived therefrom to enrich a panel of a plurality of target loci each encompassing at least one somatic variant associated with a cancer to obtain a first non-naturally occurring composition of plasma-derived amplified DNA; (d) performing a second targeted enrichment on a second portion of the adapted DNA or DNA derived therefrom to enrich the panel of the plurality of target loci to obtain a second non-naturally occurring composition of plasma-derived amplified DNA, wherein the second targeted enrichment is performed under the same conditions as the first targeted enrichment; (e) sequencing DNA derived from the first non-naturally occurring composition to obtain a first set of reads, and sequencing DNA derived from the second non- naturally occurring composition to obtain a second set of reads, wherein the first and second sequencing are performed under the same conditions; and (f) combining and analyzing the first 3 4934-6677-8453.1Attorney Docket No. N.053.WO.01 and second sets of reads to identify the presence or absence of the at least one somatic variant associated with cancer in the extracted cell-free DNA. In some embodiments, the first and second portions of the extracted cell-free DNA have an about equal amount of DNA.
[0011] In some embodiments, at least one somatic variant associated with cancer is identified in step (f), and the method further comprises: (g) extracting cellular DNA from a fraction of the blood sample containing a plurality of leukocytes; (h) performing targeted enrichment on the extracted cellular DNA or DNA derived therefrom to enrich a subset of the panel of the plurality of target loci to obtain a non-naturally occurring composition of leukocyte-derived amplified DNA, wherein the subset comprises the at least one somatic variant identified in step (f); and (i) sequencing the non-naturally occurring composition of leukocyte-derived enriched DNA to determine whether the at least one somatic variant identified in step (f) is also present in the leukocyte-derived enriched DNA, thereby identifying at least one somatic variant that is present in the plasma-derived enriched DNA and that is not present in the leukocyte-derived enriched DNA.
[0012] In some embodiments, performing targeted enrichment comprises targeted multiplex amplification or targeted probe capture on the extracted cell-free DNA or DNA derived therefrom to enrich the panel of the plurality of target loci. In some embodiments, the plurality of target loci comprises 25-5,000 target loci.
[0013] In some embodiments, at least one processing or sequencing error, if present, is identified from the combined reads without using any molecular barcode. In some embodiments, the method does not comprise tagging the extracted cell-free DNA with a molecular barcode. In some embodiments, prior to performing targeted enrichment, the extracted cell-free DNA is end repaired, A-tailed, and ligated to at least one adaptor comprising a universal priming sequence to obtain adaptor-ligated DNA. In some embodiments, prior to performing targeted enrichment, the adaptor-ligated DNA is amplified using the universal priming sequence. In some embodiments, prior to performing targeted enrichment, the extracted cellular DNA is not ligated to an adaptor and is not amplified.
[0014] In some embodiments, the somatic variant associated with the cancer comprises a single nucleotide variant (SNV), a multi-nucleotide variant (MNV), an indel, a gene fusion, a structural variant, or a combination thereof. In some embodiments, the somatic variant associated with the cancer comprises a single nucleotide variant (SNV) or an InDel, or a combination thereof. In some embodiments, the methods further comprise identifying one or more germline mutations present in the leukocyte-derived amplified DNA. 4 4934-6677-8453.1Attorney Docket No. N.053.WO.01
[0015] In some embodiments, the cancer is a solid tumor. In some embodiments, the solid tumor is a cancer or tumor of abdomen or abdominal wall, adrenal gland, anus, appendix, bladder, bone, brain, breast, cervix, chest wall, colon, diaphragm, duodenum, ear, endometrium, esophagus, fallopian tube, gallbladder, gastroesophageal junction, head and neck, kidney, larynx, liver, lung, lymph node, malignant effusions, mediastinum, nasal cavity, omentum, ovarian, pancreas, pancreatobiliary, parotid gland, pelvis, penis, pericardium, peritoneum, pleura, prostate, rectum, salivary gland, skin, small intestine, soft tissue, spleen, stomach, thyroid, tongue, trachea, ureter, uterus, vagina, vulva, or Whipple resection. In some embodiments, the solid tumor is breast cancer, advanced adenoma, colorectal cancer, gastrointestinal cancer, kidney cancer, lung cancer, multiple myeloma, ovarian cancer, or pancreatic cancer. In some embodiments, the solid tumor is colorectal cancer.
[0016] In some embodiments, the method further comprises longitudinally collecting a plurality of blood samples from the subject and repeating each of the steps for each of the plurality of blood samples. In some embodiments, the subject has received treatment for the cancer. In some embodiments, the plurality of blood samples are collected after the subject has been treated with surgery, first-line chemotherapy, and / or adjuvant therapy. In some embodiments, the identification of the at least one somatic variant is indicative of minimal residual disease. In some embodiments, the subject has not received treatment for the cancer. In some embodiments, the identification of the at least one somatic mutation is indicative of the presence of the cancer in the subject.
[0017] In some embodiments, the target loci comprise between 100 and 500 SNV and / or InDel loci. In some embodiments, the target loci are selected based on population-wide SNV and / or InDel patterns. In some embodiments, the target loci are cross-validated across sub-populations including by cancer stage, CRC subtype, high-risk mutation carriers, and ancestry. In some embodiments, the target loci comprise unstable microsatellite loci. In some embodiments, the target loci comprise one or more COSMIC driver or tumor suppressor gene loci.
[0018] In some embodiments, the method further comprises targeted enrichment and sequencing of a plurality of differentially methylated regions that are differentially methylated in the cancer, from cell-free DNA extracted from a different blood sample of the subject or a different fraction of the same blood sample, or DNA derived therefrom. In some embodiments, the method does not comprise whole genome sequencing or whole exosome sequencing of a tumor tissue sample of the subject. In some embodiments, the targeted enrichment comprises amplifying at least 50 of the 25- 5,000 target loci in a single reaction volume. 4 In some embodiments, the targeted enrichment 5 4934-6677-8453.1Attorney Docket No. N.053.WO.01 comprises amplifying 50-200 of the 25-5,000 target loci in a single reaction volume. In some embodiments, the target loci cover 1-5000 kilobases (kb) of genomic sequences. In some embodiments, the sequencing has a depth of read (DOR) of between 50,000 to 200,000 per target locus.
[0019] In some embodiments, one or more universal amplifications are performed after the targeted enrichment to obtain the non-naturally occurring composition(s) of plasma-derived amplified DNA. In some embodiments, one or more universal amplifications are performed after the targeted enrichment to obtain the non-naturally occurring composition of leukocyte-derived amplified DNA. In some embodiments, at least one target loci is enriched using two or more target-specific primers having overlapping sequences.
[0020] In some embodiments, the analyzing further comprises using an error model based on variant allele frequency (VAF) priors. In some embodiments, the analyzing further comprises using position-specific priors.
[0021] In some embodiments, the methods further comprise classifying the sample as containing cancer-derived DNA (positive) or not containing cancer-derived DNA (negative). In some embodiments, the classifying further comprises using gene-specific priors. In some embodiments, the classifying further comprises using position-specific, gene-specific, and / or cancer-type- specific priors.
[0022] In some embodiments, the classifying further comprises adjusting a confidence level based on concordance between two plasma replicates. In some embodiments, the methods further comprise generating a confidence score for the identified mutation based on a smoothed allele likelihood ratio, wherein the smoothed allele likelihood ratio is determined by summing over VAF likelihoods scaled by VAF likelihood priors. In some embodiments, the methods further comprise generating a confidence score for the identified variant based on a position likelihood ratio, wherein the position likelihood ratio is determined by computing a weighted position likelihood ratio ^^^(^) assuming only one of the alleles at a position is the true mutation allele. In some embodiments, the methods further comprise generating a confidence score for the identified variant based on a sample likelihood ratio, wherein the sample likelihood ratio is a likelihood for at least one target being positive. In some embodiments, a sample posterior probability of position being positive is determined based on the sample likelihood ratio, wherein the sample posterior probability is used to classify the sample. 6 4934-6677-8453.1Attorney Docket No. N.053.WO.01
[0023] In some embodiments, identification of an InDel is based on a position-based caller, wherein the position-based caller comprises training a target specific error model for each InDel target that is encountered in the sample. In some embodiments, the position-based caller comprises target consolidation, outlier detection using a beta binominal distribution, and a likelihood from tail probability. In some embodiments, the target consolidation comprises consolidating targets having the same start coordinate or share the same InDel repeat unit. In some embodiments, the InDel identification is based on a context-based caller. In some embodiments, the context-based caller comprises features to fit a Beta-Binomial regression model. In some embodiments, a positive identification by the context-based caller requires calling the InDel in at least two replicates of the sample. In some embodiments, identification of the InDel comprises an ensemble caller combining both the position-based caller and the context-based caller. In some embodiments, the method further comprises generating a confidence score for the sample being positive for a mutation using a sample caller. In some embodiments, the sample caller is an SNV sample caller, an InDel sample caller, or a combination thereof. In some embodiments, the sample caller computes a sample posterior probability.
[0024] Further details regarding aspects and embodiments of the present disclosure are provided throughout this patent application. Sections and section headers are for ease of reading and are not intended to limit combinations of disclosure, such as methods, compositions, and kits or functional elements therein across sections. Further details regarding aspects and embodiments of the present disclosure are provided throughout this patent application. Sections and section headers are for ease of reading and are not intended to limit combinations of disclosure, such as methods, compositions, or other functional elements therein across sections. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] FIG. 1 is a workflow diagram of one embodiment of the method described herein.
[0026] FIG. 2 is a workflow diagram of one embodiment of the method described herein.
[0027] FIG. 3 is a workflow diagram of one embodiment of the method described herein.
[0028] FIGs. 4 A-E shows distribution of the CRC samples used across different subpopulations. (A) Gender distribution; (B) Cancer stage distribution; (c) Age of onset distribution; (D) Age of onset by cancer stage; and (E) Ethnicity distribution. N = 26,162. Ancestry: EUR - European; AFR - African; SAS - South Asian; EAS - East Asian. 7 4934-6677-8453.1Attorney Docket No. N.053.WO.01
[0029] FIG. 5 shows overall patient coverage of a CRC v1 panel. Results based on the standard balanced test set of 3544 samples.
[0030] FIGs. 6 A-F shows overall patient coverage of a CRC v1 panel by subgroups. A-C shows cumulative distribution of mutation number in the v1 panel. D-F shows box plots by subgroups in each category. MSS: Microsatellite-stable patients; MSI - Microsatellite-unstable patients; Ancestry: EUR - European; AFR - African; SAS - South Asian; EAS - East Asian.
[0031] FIG. 7 shows an example of the mutation weighting function in KRAS G12 (toy panel constructed over 4 genes). Two moderately recurrent mutation clusters surround the main G12D peak to the left and to the right.
[0032] FIG. 8 shows overall patient coverage of a CRC v2 panel. Results based on the standard balanced test set of 3544 samples.
[0033] FIGs. 9A-F shows overall patient coverage of a CRC v2 panel by subgroups. A-C shows cumulative distribution of mutation number in the v2 panel. D-F shows box plots by subgroups in each category. MSS: Microsatellite-stable patients; MSI - Microsatellite-unstable patients; Ancestry: EUR - European; AFR - African; SAS - South Asian; EAS - East Asian.
[0034] FIG. 10 shows correlation between coverage and approximate amplicon number for the v2 panel.
[0035] Unless otherwise explained, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The singular terms “a,” “an,” and “the” include plural referents unless the context clearly indicates otherwise. “Comprising A or B” means including A, or B, or A and B. It is further to be understood that all base sizes or amino acid sizes, and all molecular weight or molecular mass values, given for DNA molecules or polypeptides are approximate, and are provided for description.
[0036] Further, ranges provided herein are understood to be shorthand for all of the values within the range. For example, a range of 1 to 50 is understood to include any number, combination of numbers, or sub-range from the group consisting, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 1 to 49, 1 to 25, 1.7 to 31.9, and so forth (as well as fractions thereof unless the context clearly dictates otherwise). Any concentration range, percentage range, ratio range, or integer range is to be understood to include the value of any integer within the recited range and, when appropriate, fractions thereof (such as one tenth and one hundredth of an integer), 8 4934-6677-8453.1Attorney Docket No. N.053.WO.01 unless otherwise indicated. Also, any number range recited herein relating to any physical feature, such as polymer subunits, size or thickness, are to be understood to include any integer within the recited range, unless otherwise indicated. When multiple low and multiple high values for ranges are given that overlap, a skilled artisan will recognize that a selected range will include a low value that is less than the high value.
[0037] As used herein, “about” or “consisting essentially of” mean ± 10% of the indicated range, value, or structure, unless otherwise indicated. As used herein, the terms “include” and “comprise” are open ended and are used synonymously. As used herein, “comprising” is synonymous with “including,” “containing,” or “characterized by,” and is inclusive or open-ended and does not exclude additional, unrecited elements or method steps. As used herein, “consisting of” excludes any element, step, or ingredient not specified in the claim element. As used herein, “consisting essentially of” does not exclude materials or steps that do not materially affect the basic and novel characteristics of the claim. In each instance herein any of the terms “comprising”, “consisting essentially of” and “consisting of” may be replaced with either of the other two terms. The invention illustratively described herein suitably may be practiced in the absence of any element or elements, limitation or limitations which is not specifically disclosed herein.
[0038] Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, suitable methods and materials are described below. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entireties. In case of conflict, the present specification, including explanations of terms, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting.
[0039] It is appreciated that certain features of aspects and embodiments herein, which are, for clarity, discussed in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various aspects and embodiments, which are, for brevity, discussed in the context of a single aspect or embodiment, may also be provided separately or in any suitable sub-combination. All combinations of aspects and embodiments are specifically embraced herein and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all sub-combinations of the various aspects and embodiments and elements thereof are also specifically disclosed herein even if each and every such sub- combination is not individually and explicitly disclosed herein. 9 4934-6677-8453.1Attorney Docket No. N.053.WO.01 DETAILED DESCRIPTION
[0040] The present disclosure addresses many long-felt needs and long-standing problems in the art, such as, but not limited to, those mentioned in the Background section herein. For example, methods are provided herein for preparing deoxyribonucleic acid (DNA) molecules useful for identifying at least one somatic mutation associated with cancer. In some embodiments described herein, identification of at least one somatic mutation associated with cancer is useful for early detection of asymptomatic cancers and for detecting relapse or metastasis of cancers, for example in minimal residual disease (MRD).
[0041] Accordingly, as illustrated in FIGs. 1 and 2, provided herein in some aspects are methods for preparing deoxyribonucleic acid (DNA) molecules that have been extracted from a blood sample of a subject. In some embodiments, the subject is at a risk or suspected of having cancer. In some embodiments, the subject has previously been diagnosed of and / or treated for cancer.
[0042] As illustrated in FIG. 1, some aspects of the methods described herein comprise steps of extracting cell-free DNA from a plasma fraction of a blood sample of a subject, and optionally extracting cellular DNA from leukocytes from whole blood or from a buffy coat fraction of the blood sample; performing targeted enrichment (for example, by targeted probe capture or targeted multiplex amplification) on the extracted cell-free DNA or DNA derived therefrom to enrich a panel of a plurality of target loci each encompassing at least one somatic mutation associated with the cancer to obtain a non-naturally occurring composition of plasma-derived enriched DNA, and optionally performing targeted enrichment (for example, by targeted probe capture or targeted multiplex amplification) on the extracted cellular DNA or DNA derived therefrom to enrich one or more of the target loci from the panel of target loci to obtain a non-naturally occurring composition of leukocyte-derived enriched DNA; and analyzing the non-naturally occurring composition of plasma-derived enriched DNA and optionally the non-naturally occurring composition of leukocyte-derived enriched DNA by sequencing to identify the at least one somatic mutation that is present in the plasma-derived DNA and that is not a mutation, such as a CH mutation, present in the leukocyte-derived DNA. In some embodiments, the panel of target loci comprises at least 25, at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 750, at least 1000, at least 1500, at least 2500, at least 4500, or at least 4999 target loci. In some embodiments, the panel of target loci comprises no more than 5000, no more than 4500, no more than 2500, no more than 1500, no more than 1000, no more than 750, no more than 500, no 10 4934-6677-8453.1Attorney Docket No. N.053.WO.01 more than 400, no more than 300, no more than 200, no more than 100, no more than 50, or no more than 25 target loci. In some embodiments, the panel of target loci comprises between 25- 200, 200-400, 25-5000, 50-4000, 100-3000, 200-2000, or 400-1000 target loci. In some embodiments, the panel of target loci has a combined amplicon size of at least 1kb, at least 5kb, at least 10kb, at least 25kb, at least 50kb, at least 100kb, at least 200kb, at least 300kb, at least 400kb, or at least 500kb. In some embodiments, the panel of target loci has a combined amplicon size of no more than 500kb, no more than 400kb, no more than 300kb, no more than 200kb, no more than 100kb, no more than 50kb, no more than 25kb, no more than 10 kb, no more than 5kb, or no more than 1kb. In some embodiments, the panel of target loci has a combined amplicon size comprising between 1kb-500kb, 25kb-400kb, 50kb-300kb, 100kb-2000kb, or 250kb-500kb. In some embodiments, at least one of the target loci in the panel of target loci is selected based on sequencing of tumor samples from one or more patients known to have cancer. In some embodiments, at least 10%, at least 25%, at least 50%, at least 75%, at least 80%, or at least 90% of the target loci in the panel of target loci are selected based on sequencing of tumor samples from one or more patients known to have cancer. In some embodiments, at least one of the target loci in the panel of target loci is a hotspot mutation. In some embodiments, at least 10%, at least 25%, at least 50%, at least 75%, at least 80%, or at least 90% of the target loci in the panel of target loci are hotspot mutations.
[0043] In some embodiments, targeted enrichment is by targeted multiplex amplification using target-specific primers. In some embodiments, all the target loci in the panel of target loci are amplified in a single reaction volume. In some embodiments, the target loci in the panel of target loci are amplified in 2, 3, 4, 5, 10, 20, or more reaction volumes or pools, each amplifying an about equal number of different target loci from the panel of target loci. For example, each reaction volume or pool may amplify 20, 50, 100, 200, 500, or 1000 target loci.
[0044] In some embodiments, extraction and targeted enrichment of leukocyte-derived cellular DNA or derivatives therefrom is performed in parallel of extraction and targeted enrichment of plasma-derived cell free DNA or derivatives therefrom, wherein the same panel of target loci are enriched in both processes. In some embodiments of the methods described herein, targeted enrichment of leukocyte-derived cellular DNA or derivatives therefrom is only performed if and after one or more somatic mutations are identified in plasma-derived DNA and / or the plasma sample from the subject is identified as having circulating tumor DNA (ctDNA). In some embodiments, only the one or more sub-pools of target loci comprising one or more target loci 11 4934-6677-8453.1Attorney Docket No. N.053.WO.01 encompassing the one or more somatic mutations identified in the plasma-derived DNA are enriched from the leukocyte-derived cellular DNA or derivatives therefrom in one or more multiplex enrichment reactions. In some embodiments, only the one or more target loci encompassing the one or more somatic mutations identified in the plasma-derived DNA are enriched from the leukocyte-derived cellular DNA or derivatives therefrom in one or more multiplex enrichment reactions.
[0045] In some embodiments, if a somatic variant or potential somatic variant is identified in extracted cell-free DNA and the same variant or potential variant is identified in the cellular DNA from the matched whole blood / buffy coat, the variant or potential variant would be considered a non-tumor specific mutation, such as a CH mutation, and filtered. In some embodiments, if a somatic variant is identified in extracted cell-free DNA and the same variant is not identified in the cellular DNA from the matched whole blood / buffy coat, the variant would be considered a true tumor specific somatic mutation.
[0046] In some embodiments of the methods described herein, the blood sample from the subject is split into a plurality of sub-samples prior to extracting cell-free DNA from the plasma fraction and extracting cellular DNA from the whole blood or buffy coat fraction of the sample. In some embodiments, the blood sample of the subject is split into a plurality of sub-samples and the extraction, enrichment (e.g., amplification) and sequencing steps of the methods described herein are performed on at least two of the sub-samples. In some embodiments, adaptors may be added to the extracted DNA before or after enrichment. In some embodiments, sample barcodes are added to the extracted DNA before or after enrichment. In some embodiments, sequencing reads from the at least two sub-samples are combined to identify at least one somatic mutation that is present in the plasma-derived DNA of each of the at least two sub-samples.
[0047] In some embodiments of the methods described herein, the plasma fraction of the blood sample is split into a plurality of plasma sub-samples prior to extracting cell-free DNA from the plasma fraction. In some embodiments, the plasma fraction of the blood sample is split into a plurality of plasma sub-samples and the amplification and sequencing steps of the methods described herein are performed on at least two of the sub-samples. In some embodiments, adaptors may be added to the extracted DNA before or after amplification. Sequencing reads from the at least two sub-samples are combined to identify at least one somatic mutation that is present in the plasma-derived amplified DNA of each of the at least two sub-samples. 12 4934-6677-8453.1Attorney Docket No. N.053.WO.01
[0048] In some embodiments of the methods described herein, the cell-free DNA extracted from the plasma fraction is split into a plurality of cell-free DNA sub-samples and the amplification and sequencing steps of the methods described herein are performed on at least two of the cell-free DNA sub-samples. In some embodiments, adaptors may be added to the extracted DNA before or after amplification. Sequencing reads from the at least two cell-free DNA sub-samples are combined to identify at least one somatic mutation that is present in the plasma-derived amplified DNA of each of the at least two cell-free DNA sub-samples.
[0049] As illustrated in FIG. 2, some aspects of the methods described herein comprise steps of extracting cell-free DNA from a plasma fraction of a blood sample of a subject, which is then further processed in two or more replicates, including adding adaptors to two or more replicates of the extracted cell-free DNA to generate two or more replicates of adapted cell-free DNA, wherein the adaptors are added to each replicate of the extracted cell-free DNA under the same conditions. In some embodiments the methods described herein comprise performing targeted amplification on two or more replicates of the adapted cell-free DNA or DNA derived therefrom to amplify a plurality of target loci each encompassing at least one somatic mutation associated with a cancer to obtain a non-naturally occurring composition of plasma-derived amplified DNA for each replicate, wherein each targeted enrichment is performed under the same conditions. In some embodiments, the methods described herein comprise sequencing DNA derived from each replicate of non-naturally occurring compositions to obtain a set of reads for each replicate, wherein the sequencing of each replicate is performed under the same conditions; and combining and analyzing the sets of reads for each replicate to identify the presence or absence of the at least one somatic mutation associated with cancer in the extracted cell-free DNA.
[0050] In some embodiments, targeted enrichment is by targeted multiplex amplification using target-specific primers. In some embodiments, the target loci in the panel of target loci are amplified in a single reaction volume for each of the replicate targeted multiplex amplifications. In some embodiments, the target loci in the panel of target loci are amplified in 2, 3, 4, 5 or more separate reaction volumes or pools for each of the replicate targeted multiplex amplifications, each amplifying an about equal number of different target loci from the panel of target loci. In some embodiments, each of the separate reaction volumes or pools for each of the replicate targeted multiplex amplifications, amplifies at least 50 of the plurality of target loci. In some embodiments, each of the separate reaction volumes or pools for each of the replicate targeted multiplex amplifications, amplifies 50-200 of the plurality of target loci. In some embodiments, the method 13 4934-6677-8453.1Attorney Docket No. N.053.WO.01 comprises amplifying the target loci in 4 separate reaction volumes or pools for each of the replicate targeted multiplex amplifications, wherein each of the 4 separate reaction volumes or pools amplifies 100 different target loci.
[0051] In some embodiments, each replicate of the two or more replicates of extracted cell-free DNA has an about equal amount of DNA. In some embodiments, each replicate of the two or more replicates of extracted cell-free DNA has an amount of DNA differing by up to 1%. In some embodiments, each replicate of the two or more replicates of extracted cell-free DNA has an amount of DNA differing by up to 5%. In some embodiments, each replicate of the two or more replicates of extracted cell-free DNA has an amount of DNA differing by up to 10%, up to 15%, up to 20%, or up to 25%.
[0052] In some embodiments, at least one processing or sequencing error, if present, is identified from the combined reads from sequencing of the DNA derived from each of the two or more replicates of the non-naturally occurring compositions. In some embodiments, the method does not comprise tagging the extracted cell-free DNA with a molecular barcode.
[0053] In some embodiments described herein, the method may further comprise targeted enrichment and sequencing of a plurality of differentially methylated regions that are differentially methylated in the cancer. In some embodiments, targeted enrichment and sequencing of a plurality of differentially methylated regions that are differentially methylated in the cancer is performed in DNA extracted from the same sample used to identify somatic mutations. In some embodiments, targeted enrichment and sequencing of a plurality of differentially methylated regions that are differentially methylated in the cancer is performed in a second cell-free DNA sample extracted from the blood of the subject, or DNA derived therefrom, which is different to the sample used for somatic mutation identification.
[0054] In some embodiments described herein, the method may further comprise targeted enrichment and sequencing of a plurality of sequences that encode proteins or miRNA that are differentially expressed in the cancer. In some embodiments, targeted enrichment and sequencing of a plurality of differentially expressed miRNA or protein sequences is performed in DNA extracted from the same sample used to identify somatic mutations. In some embodiments, targeted enrichment and sequencing of a plurality of differentially expressed miRNA or protein sequences is performed in a second cell-free DNA sample extracted from the blood of the subject, or DNA derived therefrom, which is different to the sample used for somatic mutation identification. In some embodiments, the method may further be combined with fragmentomics. 14 4934-6677-8453.1Attorney Docket No. N.053.WO.01
[0055] In some embodiments of any of the methods described herein, prior to performing targeted enrichment, the extracted cell-free DNA and / or extracted cellular DNA is end repaired, A-tailed, and at least one adaptor comprising a universal priming sequence is added to obtain adapted DNA. In some embodiments, the at least one adaptor comprising a universal priming sequence is added by ligation. In some embodiments, prior to performing targeted enrichment, the adaptor-ligated DNA is amplified using the universal priming sequence. In some embodiments, prior to performing targeted enrichment, the extracted DNA is not ligated to an adaptor and is not amplified. In some embodiments, at least one adaptor comprising a universal priming sequence is added to the extracted DNA or derivatives therefrom by primer extension using one or more pairs of target-specific primers with tails including a universal priming sequence. In some embodiments, one or more universal amplifications are performed after the targeted enrichment to obtain the non-naturally occurring composition(s) of plasma-derived amplified DNA. In some embodiments, one or more universal amplifications are performed after the targeted enrichment to obtain the non-naturally occurring composition of leukocyte-derived amplified DNA. In some embodiments, at least one target loci is amplified using two or more target-specific primers having overlapping sequences.
[0056] In some embodiments, analyzing the sequencing reads comprises identifying an observed variant as a true or false variant (sometimes referred to as a variant caller). In some embodiments, analyzing the sequencing reads comprises classifying the sample as containing cancer-derived DNA (positive) or not containing cancer-derived DNA (negative) (sometimes referred to as a sample caller). In some embodiments of any of the methods described herein, analyzing the sequencing reads comprises using an error model. In some embodiments, analyzing the sequencing reads comprises using a position-based background error model. In some embodiments, analyzing the sequence reads comprises using position-specific priors. In some embodiments, analyzing the sequencing reads comprises using VAF priors. In some embodiments, analyzing the sequencing reads comprises using consequence priors. In some embodiments, analyzing the sequencing reads comprises using gene-level priors. In some embodiments, analyzing the sequencing reads comprises using sample priors. In some embodiments, analyzing the sequencing reads comprises using one or more priors to make a variant call. In some embodiments, analyzing the sequencing reads comprises using one or more priors to make a sample call. In some embodiments, analyzing the sequencing reads comprises adjusting a confidence level based on concordance between two plasma replicates. In some embodiments, 15 4934-6677-8453.1Attorney Docket No. N.053.WO.01 analyzing the sequencing reads comprises filtering variants called in both plasma-derived and whole blood / buffy-derived samples.
[0057] In some embodiments, the somatic mutation associated with the cancer comprises a single nucleotide variant (SNV), a multi-nucleotide variant (MNV), an InDel, a gene fusion, a structural variant, or a combination thereof. In some embodiments, the somatic mutation associated with the cancer comprises an SNV or an InDel, or a combination thereof. In some embodiments, the methods described herein may further comprise identifying and / or filtering one or more germline variants present in the leukocyte-derived cellular DNA.
[0058] In one aspect of the methods described herein, as illustrated in FIG. 3, samples are called by using a plasma variant caller to first identify variants from each of the two or more plasma replicates and using a plasma consensus variant caller to identify one or more consensus variants. In some embodiments, 2 or more replicates are used. In some embodiments, 3 or more replicates are used. In some embodiments, 4 or more replicates are used. In some embodiments, 5 or more replicates are used. In some embodiments, 10 or more replicates are used. A consensus variant will only be called if it is present in each of the replicates or in the majority of the replicates. The data is then filtered for non-tumor specific mutations such as CH mutations using matched whole blood / buffy samples. If a variant is called in both the whole blood / buffy coat and plasma samples, the variant is considered a non-tumor specific mutation and is filtered out. The sample is then called as cancer positive or cancer negative using a sample caller based on data and likelihoods collected from priors, such as sample priors and target WES / WGS priors. Sample Collection
[0059] The methods disclosed herein are contemplated to be used to monitor or detect a wide variety of diseases or phenotypes, including cancers, in a patient. Different types of cancer may require collection of different types of samples as described herein. In some embodiments, the cancer is a solid tumor, and the biological sample is a tumor biopsy sample. Performing a biopsy generally involves using a sharp tool to remove a small amount of tissue from the are suspected to containing diseased cells or tissue such as a tumor. There are many different types of biopsies such as needle biopsy, CT-guided biopsy, ultrasound guided biopsy, bone biopsy, bone marrow biopsy, liver biopsy, kidney biopsy, aspiration biopsy, prostate biopsy, skin biopsy, surgical biopsy such as laparoscopic biopsy. In some embodiments, the biological sample is obtained by liquid biopsy. In some embodiments, the biological sample is a blood, serum, plasma, or urine sample. Further, 16 4934-6677-8453.1Attorney Docket No. N.053.WO.01 biological liquid samples may be extracted from variety of animal fluids containing cell-free DNA, including but not limited to blood, serum, plasma, bone marrow, urine, vitreous, sputum, tears, perspiration, saliva, semen, mucosa excretions, mucus, spinal fluid, amniotic fluid, lymph fluid, and so on. In some embodiments, cell-free DNA may be transplant donor or fetal in origin (via fluid taken from a pregnant subject) or may be derived from tissue of the subject itself.
[0060] In some embodiments, the biological sample is a liquid sample. In some embodiments, the biological sample is blood, serum, plasma, buffy coat, or bone marrow sample. In some embodiments, cell-free DNA and cellular DNA are both obtained from a blood sample of the subject by isolating and separating the plasma (which contains cell-free DNA) and buffy coat (which contains most or all of the white blood cells) fractions. The cellular DNA obtained from the white blood cells from whole blood or in the buffy coat may serve as a matched normal DNA sample to the cell-free DNA obtained from the plasma fraction, which may include circulating tumor DNA.
[0061] In some embodiments, the methods of the present disclosure further comprise longitudinally collecting a plurality of liquid biopsy samples from the patient. In some embodiments, the liquid biopsy sample is obtained from the patient after the patient has been treated for the cancer. In some embodiments, the liquid biopsy sample is a blood, serum, plasma, buffy coat, or urine sample. In some embodiments, the methods described herein are repeated for each of the longitudinally collected samples. In some embodiments, the subject has received treatment for the cancer. In some embodiments, the longitudinally collected samples are collected after the subject has been treated with surgery, first-line chemotherapy, and / or adjuvant therapy. In some embodiments, the identification of the at least one somatic mutation in the longitudinally collected samples using the methods described herein is indicative of minimal residual disease. In further embodiments, the subject is suspected or at risk of having cancer and has not received treatment for the cancer, and the identification of the at least one somatic mutation is indicative of the presence of the cancer in the subject.
[0062] In some embodiments, the methods described herein can be used at a very early period of time following cancer treatment, for example as early as the day of treatment, one day after treatment, two days after treatment, three days after treatment, four days after treatment, five days after treatment, six days after treatment, a week after treatment, two weeks after treatment, three weeks after treatment, four weeks after treatment, one month after treatment, two months after treatment, three months after treatment, four months after treatment, five months after treatment, 17 4934-6677-8453.1Attorney Docket No. N.053.WO.01 six months after treatment, seven months after treatment, eight months after treatment, nine months after treatment, ten months after treatment, eleven months after treatment, or a year or more after treatment. In some embodiments, the methods described herein can be used prior to being diagnosed or and / or receiving treatment for cancer. Sample Extraction and Enrichment of DNA Molecules
[0063] Methods provided herein, in certain embodiments, are specially adapted for extracting and amplifying DNA fragments, especially tumor DNA fragments that are found in circulating tumor DNA (ctDNA). Such fragments are typically about 160 nucleotides in length.
[0064] Cell-free nucleic acid (cfNA), e.g., cell-free DNA (cfDNA), can be released into the circulation via various forms of cell death such as apoptosis, necrosis, autophagy and necroptosis. cfDNA is fragmented and the size distribution of the fragments varies from 150- 350 bp to> 10000 bp. By contrast, for example, the size distributions of ctDNA fragments from hepatocellular carcinoma (HCC) patients spanned a range of 100-220 bp in length with a peak in count frequency at about 166bp and the highest tumor DNA concentration in fragments of 150-180 bp in length.
[0065] In an illustrative embodiment, cell-free DNA (cfDNA) containing circulating tumor DNA (ctDNA) is extracted from the plasma fraction of blood after removal of cellular debris and platelets by centrifugation. The plasma samples can be stored at -80ºC until the DNA is extracted using, for example, QIAamp DNA Mini Kit (Qiagen, Hilden, Germany), (e.g., Hamakawa et al., Br J Cancer. 2015; 112:352-356). Hamakava et al. reported median concentration of extracted cfDNA of all samples 43.1 ng per ml plasma (range 9.5-1338 ng / ml) and a mutant fraction range of 0.001-77.8%, with a median of 0.90%.
[0066] In some embodiments, the plasma fraction of the blood sample is split into a plurality of plasma sub-samples prior to extracting cfDNA from the plasma fraction. For example, the plasma fraction of a blood sample may be split into 2, 3, 4, 5, 6, 7, 8, 9, 10, or more plasma sub-samples. In some embodiments, each of the plurality of plasma sub-samples contain substantially equal amounts of cfDNA. In some embodiments, each of the plurality of plasma sub-samples do not contain equal amounts of cfDNA.
[0067] In some embodiments, the cfDNA extracted from the plasma fraction is split into a plurality of cfDNA sub-samples. For example, the cfDNA may be split into 2, 3, 4, 5, 6, 7, 8, 9, 10, or more cfDNA sub-samples. In some embodiments, each of the plurality of cfDNA sub- 18 4934-6677-8453.1Attorney Docket No. N.053.WO.01 samples contain substantially equal amounts of cfDNA. In some embodiments, each of the plurality of plasma sub-samples do not contain equal amounts of cfDNA.
[0068] Samples that are useful for methods herein can be virtually any nucleic acid sample. In illustrative embodiments, the nucleic acid sample is extracted or isolated from a subject. In certain illustrative embodiments, DNA molecules are extracted from a sample such as a tissue sample, and for illustrative embodiments herein, a liquid sample. Methods that are particularly useful in exemplary embodiments include methods for isolating cfDNA from a liquid sample, and in illustrative embodiments from a blood, serum, urine, vitreous, sputum, saliva, tears, perspiration, feces, bile, lymph, cervical mucus, or semen sample, and in further illustrative embodiments, a plasma sample.
[0069] In certain illustrative embodiments, extracted or isolated DNA molecules can be enriched. Reagents, kits and associated methods for nucleic acid isolation and enrichment from biological samples, including for isolation of cfDNA from liquid samples, are generally available commercially. For example, isolation of cfDNA or cellular DNA from a liquid (e.g., blood or blood derivative sample such as a serum, plasma, or buffy coat sample) can involve binding DNA molecules from a sample to a matrix and isolating the DNA molecules in the presence of a solvent. In some embodiments, the method further comprises incubating the biological sample comprising DNA molecules with a protease, prior to contacting the DNA molecules to the matrix. In some embodiments, the method can further include the steps of washing the matrix with a wash buffer to remove impurities, and optionally, drying the matrix. Enriched nucleic acid samples can be eluted from the matrix with an elution buffer.
[0070] Other methods for nucleic acid isolation, for example cfDNA or cellular DNA isolation, and optional enrichment of certain cfDNA or cellular DNA can include ion exchange columns, or microfluidic devices, such as solid phase isolation, based on DNA capture by immobilized beads or functionalized surface. Additional methods include liquid phase isolation, utilizing an electric field, or chemical reagents, instead of a functionalized surface. In some embodiments, isolation of cfDNA from a patient sample is performed using a DNA isolation kit (e.g., QIAamp Circulating Nucleic Acid kit (Qiagen)). In some embodiments, capture by hybridization with hybrid capture probes is used to preferentially enrich the DNA, for example using probes that bind to specific nucleic acid sequences at or near, target regions.
[0071] In some embodiments, cfDNA of certain sizes can be enriched before subjecting the enriched cfDNA to methods herein. Enriched cfDNA molecules can be, for example 50 to 1200 19 4934-6677-8453.1Attorney Docket No. N.053.WO.01 base pairs in length, or 70 to 500 base pairs in length, or 100 to 200 base pairs in length, or 130 to 170 base pairs in length. In some embodiments, the enriched cfDNA molecules are from 50 to 200 bp in length. In some embodiments, the enriched cfDNA molecules are between 60 and 200 bp in length, between 60 and 150 bp in length, or between 60 and 100 bp in length before the enriched cfDNA molecules, or derivatives thereof, are ligated to adapters in methods herein. In some embodiments, the enriched cfDNA molecules are less than 150, 100, 90, 75, or 50 bp in length before they are ligated to adapters. Such enrichment methods can be performed for example using the methods of WO2018156418 A1, Stray, et al. (incorporated herein, in its entirety).
[0072] In some embodiments, the sample is enriched for cancer DNA molecules. In illustrative embodiments of such embodiments, the enriched nucleic acid sample is circulating tumor DNA (ctDNA). In such embodiments, the enriched nucleic acid the enriched nucleic acid is less than 160, 150, 145, 120, 100, 90, 75, or 50 bp in length. In some embodiments, size selection is used to filter out cfDNA molecules carrying mutations derived from clonal hematopoiesis but not tumor- derived mutations, which are typically longer than ctDNA molecules, at about 165bp.
[0073] DNA methylation biomarkers are increasingly being utilized in the development of novel assays for use in women’s health and / or organ health. For example, methylation profiling is being used in non-invasive prenatal testing (NIPT) for monitoring of placental and fetal epigenomic changes, pre-symptomatic detection of preterm birth, preeclampsia, placental insufficiency, and fetal growth restriction. Methylation markers can also be indicative of congenital diseases of the fetus. Furthermore, methylation biomarkers are also used for monitoring of organ health such as predicting and monitoring organ rejection in transplant patients, monitoring of immune changes in transplant rejection, and for monitoring of organ health in high risk or predisposed individuals. Accordingly, the sample described herein may be a maternal sample, a fetal sample, or a sample from a transplant patient.
[0074] In some embodiments, the sample is enriched for fetal DNA molecules. In illustrative embodiments of such embodiments, the enriched nucleic acid sample is fetal circulating free DNA or amplicons thereof. In such embodiments, the enriched nucleic acid is between 100 bp and 220 bp in length. In some embodiments, the length of the nucleic acid sample of fetal cfDNA ranges from 100 bp to 200 bp, from 120 bp to 180 bp, from 140 bp to 160 bp, from 150 bp to 170 bp, from 160 to 190 bp, or from 170 bp to 220 bp.
[0075] In some embodiments, the sample is enriched for transplant donor DNA molecules. In illustrative embodiments of such embodiments, the enriched nucleic acid sample is transplant 20 4934-6677-8453.1Attorney Docket No. N.053.WO.01 donor circulating free DNA or amplicons thereof. In such embodiments, the enriched nucleic acid is between 100 bp and 220 bp in length. In some embodiments, the length of the nucleic acid sample of transplant donor cfDNA ranges from 100 bp to 200 bp, from 120 bp to 180 bp, from 140 bp to 160 bp, from 150 bp to 170bp, from 160 bp to 190 bp, or from 170 bp to 220 bp. Subjects
[0076] Subjects used in methods herein can be virtually any animal, in illustrative embodiments a mammal, and in further illustrative embodiments, a human. In some embodiments, the subject has or is suspected or at risk of having a disease, in some embodiments, cancer. In some embodiments, the subject is a pregnant female. In some embodiments, the subject is a subject comprising an organ from another individual. In embodiments where the subject has, is suspected to have, or is at risk of having cancer, the cancer can be any type of cancer. In some embodiments, the subject has, is suspected to have, or is at risk of having one or more cancers (e.g., one cancer) including ovarian cancer, soft tissue sarcoma, peripheral T cell cancer, colorectal cancer, intrahepatic cholangiocarcinoma, glioblastoma, esophageal cancer, cutaneous T cell lymphoma, non-Hodgkin lymphoma, urothelial cancer, basal cell carcinoma, epithelioid sarcoma, pancreatic cancer, non-small cell lung carcinoma, Hodgkin lymphoma, renal cell carcinoma, mesothelioma, metastatic uveal melanoma, kidney cancer, blood cancer, HER2-expressing cancers, non-melanoma skin cancer, liposarcoma, hepatocellular carcinoma, small lymphocytic lymphoma, prostate cancer, breast cancer, anal cancer, marginal zone lymphoma, cutaneous squamous cell carcinoma, thyroid cancer, medullary thyroid cancer, triple-negative breast cancer, neuroendocrine prostate cancer, bladder cancer, paraganglioma, medulloblastoma, superficial basal cell carcinoma, head and neck squamous cell carcinoma, hematologic malignancies, melanoma, B-cell lymphoma, relapsed / refractory acute myeloid leukemia, angiosarcoma, bone sarcoma, refractory cervical cancer, cholangiocarcinoma, osteosarcoma, biliary tract cancer, castration-resistant prostate cancer, gastroesophageal adenocarcinomas, rhabdomyosarcoma, carcinoma, non-muscle invasive bladder cancer, uveal melanoma, small cell lung cancer, cervical cancer, primary open angle glaucoma, follicular lymphoma, synovial sarcoma, liver cancer, carcinosarcoma, leptomeningeal brain tumors, T-cell lymphoma, lymphoma, small cell lung cancer, mantle cell lymphoma, B-cell malignancies, endometrial cancer, myxoid / round cell liposarcoma, metastatic Merkel cell carcinoma, neuroblastoma, chronic lymphocytic leukemia, tenosynovial giant cell tumors, sarcoma, acute 21 4934-6677-8453.1Attorney Docket No. N.053.WO.01 myeloid leukemia, skin cancer, nasopharyngeal carcinoma, relapsed / refractory Ewing sarcoma, bone cancer, glioma, salivary gland carcinoma, gastric cancer, benign tumor, low-grade serous ovarian cancer, metastatic breast cancer, multiple myeloma, diffuse large B cell lymphoma, relapsed / refractory lymphoma, metastatic colorectal cancer, advanced malignancies, and acute lymphoblastic leukemia.
[0077] In certain embodiments of methods herein, the subject has, is suspected to have, or is at risk of having a cancer of the abdomen or abdominal wall, adrenal gland, anus, appendix, bladder, bone, brain, breast, cervix, chest wall, colon, diaphragm, duodenum, ear, endometrium, esophagus, fallopian tube, gallbladder, gastro-esophageal junction, head and neck, kidney, larynx, liver, lung, lymph node, malignant effusions, mediastinum, nasal cavity, omentum, ovaries, pancreas, pancreatobiliary, parotid gland, pelvis, penis, pericardium, peritoneum, pleura, prostate, rectum, salivary gland, skin, small intestine, soft tissue, spleen, stomach, thyroid, tongue, trachea, ureter, uterus, vagina, vulva, and Whipple resection. In illustrative embodiments, the cancer is selected from nasopharyngeal carcinoma, hepatocellular carcinoma, breast cancer, ovarian cancer, pancreatic cancer, colorectal cancer, lung cancer, gastroesophageal cancer, prostate cancer, bladder cancer, melanoma, and acute leukemia. In illustrative embodiments, the cancer is selected from colorectal cancer. In some embodiments, the cancer is a solid tumor. In some embodiments, the solid tumor is colorectal cancer. Nucleic Acid Processing
[0078] Typically, methods herein include a step of adding, in some embodiments ligating, nucleic acid adapters to sample DNA molecules, or in illustrative embodiments to nucleic acid derivatives generated therefrom. The sample DNA molecules in illustrative embodiments are from a subject. In some embodiments, nucleic acid adapters are added after sample cellular DNA molecules are fragmented to form fragmented cellular DNA molecules. In some embodiments methods include exposing sample DNA molecules to one or more polymerases or kinases, such as Klenow Large Fragment Polymerase and T4 polynucleotide kinas (PNK), as well as a ligase, such as T4 ligase. In some embodiments, sample DNA molecules or the fragmented DNA molecules, are exposed to one or more polymerases and / or kinases to generate the nucleic acid derivatives generated therefrom. In some embodiments, the method further comprises ligating adapters to the nucleic acid derivatives generated therefrom. 22 4934-6677-8453.1Attorney Docket No. N.053.WO.01
[0079] In some embodiments, adapters are ligated to sample DNA molecules in. In some embodiments, before such ligation, extracted or isolated DNA molecules can be modified to form sample nucleic acid derivatives, for example to make them more amenable to adapter ligation. For example, sample DNA molecules can be blunt ended, nucleotides can be added to sample DNA molecules or blunted-ended derivative therefrom, and / or phosphate moieties can be added or removed from the ends of sample DNA molecules or derivatives thereof. In some embodiments, prior to ligation, sample DNA molecules may be blunt ended, and then a single adenosine base can be added to the 3’ end. Prior to ligation the DNA may be cleaved using a restriction enzyme or some other cleavage method. During ligation the 3’ adenosine of the sample fragments and the complementary 3’ tyrosine overhang of an adapter can enhance ligation efficiency. In some embodiments, adapter ligation is performed using a T4 ligase.
[0080] In some embodiments, adapters are appended to sample DNA molecules by primer extension. For example, in some embodiments, adapters are appended to sample DNA molecules using target specific primers with tails comprising universal priming sequences. The tails may also include other useful sequences such as molecular or sample barcodes.
[0081] In some embodiments, adapters containing one or more universal priming sequences are utilized in methods herein. In some embodiments, the adapters are Y adapters, for example in illustrative methods in which target region amplicons are sequenced using NGS. In some embodiments, the adapters each comprises a universal priming site.
[0082] In some embodiments, the adapters each further comprises a sample barcode. Thus, multiple samples can be analyzed in the same sequencing reaction. The sample barcode can be used to process data according to the sample from which the data was generated.
[0083] In some embodiments, the adapters each further comprises a molecular barcode. In some embodiments, the number of adapters having different molecular barcodes is between 10 to 1,000, and wherein the ratio of the total number of sample nucleic acid or cfDNA molecules to the number of different molecular barcodes in the ligation reaction is at least 1,000:1. The number of different molecular barcodes in the ligation reaction, in certain embodiments, ranges from 10 to 50, 10 to 100, 50 to 200, 100 to 300, 200 to 500, 300 to 600, 500 to 700, 600 to 800 or 700 to 1,000. In some embodiments, there are at least 1, 10, 20, 30, 40, 50, or at least 100; 200, 300, 400, 500, 600, 700, 800, 900, or 1000 different molecular barcodes in the ligation reaction. In some embodiments, the ratio of the total number of sample nucleic acid or cfDNA molecules to the number of different molecular barcodes in the ligation reaction is at least 10,000:1. In some 23 4934-6677-8453.1Attorney Docket No. N.053.WO.01 embodiments, the ratio of the total number of sample nucleic acid or cfDNA molecules to the number of different molecular barcodes in the ligation reaction ranges from 50,000: 1 to 50:1, from 25,000: 1 to 100:1, from 10,000:1 to 100:1, from 10:000:1 to 8,000:1 to 500:1, from 5,000:1 to 200:1, from 10,000:1 to 50:1. In some embodiments, the methods disclosed herein result in at least 100; 200; 500; 750; 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000 different molecular barcodes to each one sample nucleic acid or cfDNA molecules. Targets
[0084] The methods described herein comprise targeted enrichment of a panel of target loci. In some embodiments, the panel of target loci may comprise one or more somatic variants including single nucleotide variants (SNVs), multi-nucleotide variants (MNVs), InDels, gene fusions, structural variant, or a combination thereof. In some embodiments, the panel of target loci may include somatic variants that have been previously determined to be present in cancer cells and not present in non-cancerous cells, or present in one type of cancer cells and not present in another type of cancer cells or non-cancerous cells. In some embodiments, the panel of target loci may include somatic variants that have been previously identified from tumor samples. In some embodiments, the panel of target loci comprises 5-50,000 target loci, 10-10,000 target loci, 25- 5,000 target loci, 50-1,000 target loci, or 100-500 loci. In some embodiments, the panel of target loci comprises at least 5 target loci, at least 10 target loci, at least 25 target loci, at least 50 target loci, at least 100 target loci, at least 200 target loci, at least 300 target loci, at least 400 target loci, at least 1000 target loci, at least 1500, at least 2500, a least 4500, or at least 5000, at least 10,000 target loci, at least 50,000 target loci, or at least 100,000 target loci. In some embodiments, the panel of target loci comprises no more than 5000, no more than 4500, no more than 2500, no more than 1500, no more than 1000, no more than 750, no more than 500, no more than 400, no more than 300, no more than 200, no more than 100, no more than 50, or no more than 25 target loci.
[0085] In some embodiments, the panel of target loci is determined based on historical sequencing data, such as whole genome sequencing (WGS), whole exome sequencing (WES), or targeted panel sequencing data, generated from cancer patients. In some embodiments, the data is generated from population-wide SNV patterns and prior knowledge of cancer specific somatic mutations. In some embodiments, the data is generated from synonymous and non-synonymous mutations from a plurality of cancer patient samples. In some embodiments, the data is generated from synonymous and non-synonymous mutations from 10 or more, from 25 or more, from 50 or 24 4934-6677-8453.1Attorney Docket No. N.053.WO.01 more, from 100 or more, from 250 or more, from 500 or more, from 1000 or more, from 1500 or more, from 2000 or more, from 2500 or more, from 5000 or more, or from 10,000 or more cancer patient samples. In some embodiments, the data is generated based on tumor SNVs detected in a plurality of tumor samples. In some embodiments, the data is generated based on tumor SNVs detected in 10 or more, 25 or more, 50 or more, 100 or more, 250 or more, from 500 or more, 1000 or more, 1500 or more, 2000 or more, 2500 or more, 5000 or more, 10,000 or more, or 20,000 or more tumor samples. In some embodiments, the panel of target loci is determined to ensure optimal mutation coverage and representation across subpopulations, including by cancer stage, cancer type, cancer subtype, high-risk mutation carriers, and ancestry. In some embodiments, the panel of target loci comprises clustered mutations. In some embodiments, the panel of target loci comprises mutations with high prevalence in cancer patients. In some embodiments, the target loci comprise one or more unstable microsatellite loci. In some embodiments, the target loci comprise one or more cancer hotspot mutations. In some embodiments, the target loci comprise one or more COSMIC driver or tumor suppressor gene loci.
[0086] In some embodiments, the panel of target loci is based on previously determined cancer- specific somatic mutations. In some embodiments, the previously determined cancer-specific somatic mutations are ranked by prevalence and the panel is limited to the top 5 somatic mutations, the top 10 somatic mutations, the top 50 somatic mutations, the top 100 somatic mutations, the top 200 somatic mutations, the top 300 somatic mutations, the top 400 somatic mutations, the top 500 somatic mutations, the top 1000 somatic mutations, the top 10,000 somatic mutations, the top 50,000 somatic mutations, or top 100,000 somatic mutations. In some embodiments, the previously determined cancer-specific somatic mutations are ranked by prevalence and the panel is limited to the top 10-100,000 somatic mutations, the top 25-5,000 somatic mutations, the top 50-1,000 somatic mutations, or the top 100-500 somatic mutations. Amplifications
[0087] Methods in some aspects herein include performing one or, in some embodiments, two or more amplifications. Such amplifications in certain illustrative embodiments include at least one targeted amplification wherein at least one primer and in certain embodiments both primers of a primer pair, one or more primer pairs, or a set of primer pairs used for the amplification are each designed to bind to a specific nucleic acid sequence at or near a genomic region of interest (i.e., 25 4934-6677-8453.1Attorney Docket No. N.053.WO.01 are target-specific primers) to generate target region amplicons. In some embodiments, methods herein include one or more universal amplifications.
[0088] A number of amplification technologies can be used with methods herein. For example, such amplification can be an isothermal amplification (e.g., recombinase polymerase amplification (RPA), a ligase-based amplification, PCR, or a combination thereof (e.g., ligation-mediated PCR).
[0089] In some illustrative embodiments, the targeted enrichment is a targeted PCR(s) that is performed using a PCR reaction mixture that includes one primer pair, or in illustrative embodiments a set of primer pairs, and DNA molecules (e.g., cfDNA or cellular DNA), or universally amplified DNA molecules generated therefrom.
[0090] In some embodiments, at least one primer of a primer pair used for targeted amplification herein, is a target-specific primer designed to bind to a specific nucleic acid sequence at or near a genomic region of interest, which in illustrative examples can be genomic regions where one or more somatic mutations are associated with or indicative of formation or presence of cancer, and can, for example, include mutations in promoter regions of tumor suppressor genes. A target- specific primer can be designed to bind to any sequence at or near a target region for amplification of the target region encompassing one or more cancer-specific variants. In some embodiments, a target-specific primer can be designed to bind to a sequence that includes one or more cancer- specific variant loci. In other embodiments, a target-specific primer can be designed to bind to a sequence that is upstream or downstream to one or more SNV or indel loci.
[0091] In some methods herein, a universal amplification(s) can be performed before the targeted amplification(s). In some embodiments, a universal amplification(s) can be performed after the targeted amplification(s). A universal amplification can be performed for example using a universal primer pair. In some embodiments, the universal primer pair binds primer binding sites in an adapter added by ligation or during a previous amplification reaction, step, or cycle, and in some embodiments, during a previous targeted amplification reaction, step, or cycle.
[0092] In some embodiments, at least one of the primer pairs comprises a universal primer and a target-specific primer. In some embodiments, at least one of the primer pairs comprises two target- specific primers. In some embodiments, at least one of the primers comprises a sequencing tag. In some embodiments, at least one of the primers comprises a sample index. In some embodiments, performing a PCR further comprising using primers comprising a sequencing primer binding site. In some embodiments, performing a PCR further comprising using primers comprising a sample 26 4934-6677-8453.1Attorney Docket No. N.053.WO.01 index. In some embodiments, the primers of the primer pairs are probe-dependent primers and the amplification is a target capture polymerase chain reaction.
[0093] Methods as described herein, in some embodiments, can include multiple amplification cycles (e.g. multiple PCR temperature cycles), and in some embodiments can include several sequential PCR reactions performed during the same set of temperature cycles. In some embodiments, amplification cycles can include at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 cycles. In some embodiments, amplification cycles can include at least 7, 8, 9, or 10 cycles. In illustrative embodiments, amplification cycles can include at least 11, 12, 13, 14, 15, 16, or 17 cycles.
[0094] In some embodiments described herein, PCR amplification is performed by adding a PCR reaction mixture to the DNA template (e.g., cfDNA or cellular DNA) followed by addition of a polymerase enzyme, and then amplified through multiple amplification cycles. In some embodiments, the PCR reaction mixture contains one or more primer pairs, deoxynucleotides (dNTPs), PCR reaction buffer, and deionized water. In some embodiments, the dNTPs can comprise a mixture of dATP, dCTP, dGTP and dTTP. In some embodiments, the final concentration of each dNTP in the reaction mixture can range from 0.05 mM to 0.5 mM dNTPs, for example, from 0.05 mM to 0.5 mM, 0.05 mM to 0.1 mM, 0.05 mM to 0.15 mM, 0.05 to 0.2 mM, 0.05 mM to 0.25 mM, 0.05 mM to 0.3 mM, 0.05 mM to 0.35 mM, 0.05 to 0.4 mM 0.05 mM to 0.45 mM, from .1 mM to 0.5 mM, 0.15 mM to 0.5 mM, or 0.2 mM to 0.5 mM, 0.25 to 0.5 mM, 0.3 mM to 0.5 mM, 0.35 mM to 0.5 mM, 0.4 mM 0.5 mM, or from 0.45 mM to 0.5 mM. In illustrative embodiments, the final concentration of each dNTP in the reaction mixture is between from 0.15 mM and 0.25 mM. In some embodiments, the final concentration of each dNTP in the reaction mixture is 0.2 mM.
[0095] PCR buffer solution creates a suitable environment for the polymerase chain reaction and can contain many different components, including magnesium chloride (MgCl2), potassium chloride (KCl), dimethyl sulfoxide (DMSO), and glycerin or bovine serum albumin (BSA). In some embodiments, the concentration of KCl can be between of 25 and 50 mM, between 25 and 75 mM, between 25 and 100 mM, between 30 and 100 mM, between 50 and 100 mM, or between 70 and 100 mM. In some embodiments, the MgCl2 concentration can be in the range of 0.5 mM to 5 mM, between 0.5 mM to 4.5 mM, 0.5 to 4.0 mM, 0.5 to 3.5 mM, 0.5 to 3.0 mM, 0.5 to 2.5 mM, 0.5 to 2.0 mM, 0.5 to 1.5 mM, 1.0 to 5 mM, 1.5 to 5mM, 2.0 to 5 mM, 2.5 to 5 mM, 3.0 to 5 mM, 27 4934-6677-8453.1Attorney Docket No. N.053.WO.01 3.5 to 5 mM, 4.0 to 5 mM, or 4.5 to 5 mM. In some embodiments, the concentration of MgCl2 is 2.0 mM.
[0096] In some embodiments, the buffer solution is a Q5® Reaction Buffer (B9027S, New England Biolabs, Inc.). In some embodiments, the reaction buffer is Standard Taq Reaction Buffer (B9014S, New England Biolabs, Inc). In some embodiments, the reaction buffer is a Standard Taq (Mg-free) Reaction Buffer (B9015S, New England Biolabs, Inc.).
[0097] In some embodiments, a DNA polymerase is used to produce DNA amplicons using DNA as a template. In some embodiments, the polymerase is a Q5® DNA Polymerase, such as Q5® High-Fidelity DNA Polymerase (M0491S, New England BioLabs, Inc.) or Q5® Hot Start High- Fidelity DNA Polymerase (M0493S, New England BioLabs, Inc.). Q5® High-Fidelity DNA polymerase is a high-fidelity, thermostable, DNA polymerase with 3´→ 5´ exonuclease activity, fused to a processivity-enhancing Sso7d domain. Q5® High-Fidelity DNA polymerase lacks 5´→ 3´exonuclease activity and strand displacement activity.
[0098] In some embodiments, the polymerase is a T4 DNA polymerase (M0203S, New England BioLabs, Inc.). T4 DNA Polymerase catalyzes the synthesis of DNA in the 5´→ 3´ direction and requires the presence of template and primer. This enzyme has a 3´→ 5´ exonuclease activity which is much more active than that found in DNA PolymeraseDNA polymerase lacks 5´→ 3´ exonuclease activity and strand displacement activity.
[0099] In some embodiments of any of the aspects herein, the length of the primers can be between 10 to 100 nucleotides, such as between 10 to 75 nucleotides, 10 to 40 nucleotides, 10 to 35 nucleotides, 10 to 30 nucleotides, 10 to 20 nucleotides, 15 to 100 nucleotides, 20 to 100 nucleotides, from 25 to 100 nucleotides, from 30 to 100 nucleotides from 35 to 100 nucleotides, from 40 to 100 nucleotides, from 45 to 100 nucleotides, from 50 to 100 nucleotides, from 55 to 100 nucleotides, from 60 to 100 nucleotides, from 65 to 100 nucleotides, from 70 to 100 nucleotides, or from 75 to 100 nucleotides. In some embodiments, the range of the length of the primers is between 5 to 50 nucleotides, such as 5 to 40 nucleotides, 5 to 20 nucleotides, or 5 to 10 nucleotides. In some embodiments, the primers are between 5 and 50 bp in length, between 10 and 40 bp in length, between 15 and 30 bp in length, between 15 and 25 bp in length, between 20 and 40 bp in length, between 25 and 50 bp in length, or between 30 and 50 bp in length. In some embodiments, the primers are between 25 and 100 bp in length, between 35 and 100 bp in length, between 45 and 100 bp in length, between 55 and 100 bp in length, between 65 and 100 bp in length, or between 75 and 100 bp in length. 28 4934-6677-8453.1Attorney Docket No. N.053.WO.01
[0100] In some embodiments of any of the aspects or embodiments herein, the number of primer pairs can range from 1 to 100,000 primer pairs that each bind to one or more primer binding sequences. In some embodiments, the primer pairs are a part of a set of primer pairs. In some embodiments, the set of primers range from 2 to 100,000, from 2 to 10,000, from 2 to 1,000, from 2 to 100, from 2 to 50, from 10 to 100, from 50 to 100, from 100 to 200, from 100 to 500, from 100 to 1,000, from 100 to 10,000, from 100 to 100,000, from 1,000 to 100,00, or from 10,000 to 100,000 primer pairs. In some embodiments, the number of primer pairs can range from 10 to 10,000, 10 to 1,000, 10 to 100, 10 to 50, 10 to 40, 10 to 30, 15 to 30, or 15 to 25 primer pairs.
[0101] In some embodiments, PCR is used to generate very short amplicons. cfDNA (such as fetal cfDNA in maternal serum or necroptically- or apoptotically-released cancer cfDNA) is highly fragmented. For fetal cfDNA, the fragment sizes are distributed in approximately a Gaussian fashion with a mean of 160 bp, a standard deviation of 15 bp, a minimum size of about 100 bp, and a maximum size of about 220 bp. Because cfDNA fragments are short, the likelihood of both primer sites being present the likelihood of a fragment of length L comprising both the forward and reverse primers sites is the ratio of the length of the amplicon to the length of the fragment. Under ideal conditions, assays in which the amplicon is 45, 50, 55, 60, 65, or 70 bp will successfully amplify from 72%, 69%, 66%, 63%, 59%, or 56%, respectively, of available template fragment molecules. Thus, in some embodiments target amplicons generated in method herein are between 40 and 100, 40 and 75, or 45 and 70 bp in length. In certain embodiments that relate most preferably to cfDNA from samples of individuals suspected of having cancer, the cfDNA is amplified using primers that yield a maximum amplicon length of 85, 80, 75 or 70 bp, and in certain preferred embodiments 75 bp, and that have a melting temperature between 50 and 65°C, and in certain preferred embodiments, between 54–60.5°C. The amplicon length is the distance between the 5-prime ends of the forward and reverse priming sites. Amplicon length that is shorter than typically used by those known in the art may result in more efficient measurements of the desired methylation sites by only requiring short sequence reads. In an embodiment, a substantial fraction of the amplicons are between 25 on the low end of the range, and 100 bp, 90 bp, 80 bp, 70 bp, 65 bp, 60 bp, 55 bp, 50 bp, or 45 bp on the high end of the range.
[0102] In some embodiments, the PCR is multiplexed. In any of the methods for detecting somatic mutations herein, improved amplification parameters for multiplex PCR can be employed. For example, wherein the amplification reaction is a PCR reaction and the annealing temperature is between 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10°C greater than the melting temperature on the low end of the 29 4934-6677-8453.1Attorney Docket No. N.053.WO.01 range, and 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15°C on the high end the range for at least 10, 20, 25, 30, 40, 50, 06, 70, 75, 80, 90, 95 or 100% the primers of the set of primers.
[0103] In certain embodiments, wherein the amplification reaction is a multiplex PCR reaction the length of the annealing step in the PCR reaction is between 10, 15, 20, 30, 45, and 60 minutes on the low end of the range, and 15, 20, 30, 45, 60, 120, 180, or 240 minutes on the high end of the range. In certain embodiments, the primer concentration in the amplification, such as the multiplex PCR reaction is between 1 and 10 nM. Furthermore, in exemplary embodiments, the primers in the set of primers, are designed to minimize primer dimer formation.
[0104] Multiplex PCR may involve a single round of PCR in which all targets are amplified or it may involve one round of PCR followed by one or more rounds of nested PCR or some variant of nested PCR. Nested PCR consists of a subsequent round or rounds of PCR amplification using one or more new primers that bind internally, by at least one base pair, to the primers used in a previous round. Nested PCR reduces the number of spurious amplification targets by amplifying, in subsequent reactions, only those amplification products from the previous one that have the correct internal sequence. Reducing spurious amplification targets improves the number of useful measurements that can be obtained, especially in sequencing. Nested PCR typically entails designing primers completely internal to the previous primer binding sites, necessarily increasing the minimum DNA segment size required for amplification. For samples such as cancer patient plasma cfDNA, in which the DNA is highly fragmented, the larger assay size reduces the number of distinct cfDNA molecules from which a measurement can be obtained. In an embodiment, to offset this effect, one may use a partial nesting approach where one or both of the second round primers overlap the first binding sites extending internally some number of bases to achieve additional specificity while minimally increasing in the total assay size.
[0105] In an embodiment, a multiplex pool of PCR assays are designed to amplify potentially heterozygous SNVs or other polymorphic or non-polymorphic loci on one or more chromosomes and these assays are used in a single reaction to amplify DNA. The number of PCR assays may be more than 10, more than 25, more than 50, more than 100, more than 200, more than 300, more than 400, more than 500, or more than 1000 PCR assays in a single reaction. The SNV frequencies of each locus may be determined by clonal or some other method of sequencing of the amplicons. Note that this method is equally well applicable to detecting translocations, deletions, duplications, and other chromosomal abnormalities. 30 4934-6677-8453.1Attorney Docket No. N.053.WO.01
[0106] In an embodiment, tails with no homology to the target genome may also be added to the 3-prime or 5-prime end of any of the primers. These tails facilitate subsequent manipulations, procedures, or measurements. In an embodiment, the tail sequence can be the same for the forward and reverse target specific primers. In an embodiment, different tails may used for the forward and reverse target specific primers. In an embodiment, a plurality of different tails may be used for different loci or sets of loci. Certain tails may be shared among all loci or among subsets of loci. For example, using forward and reverse tails corresponding to forward and reverse sequences required by any of the current sequencing platforms can enable direct sequencing following amplification. In an embodiment, the tails can be used as common priming sites among all amplified targets that can be used to add other useful sequences. In some embodiments, the inner primers may contain a region that is designed to hybridize either upstream or downstream of the targeted polymorphic locus. In some embodiments, the primers may contain a molecular barcode. In some embodiments, the primer may contain a universal priming sequence designed to allow PCR amplification.
[0107] In an embodiment, a PCR assay pool is created such that forward and reverse primers have tails corresponding to the required forward and reverse sequences required by a high throughput sequencing instrument such as the HISEQ, GAIIX, or MYSEQ available from ILLUMINA. In addition, included 5-prime to the sequencing tails is an additional sequence that can be used as a priming site in a subsequent PCR to add nucleotide barcode sequences to the amplicons, enabling multiplex sequencing of multiple samples in a single lane of the high throughput sequencing instrument.
[0108] In some embodiments, ligation mediated universal-PCR amplification of fragmented DNA may be used. The ligation mediated universal-PCR amplification can be used to amplify plasma DNA, which can then be divided into multiple parallel reactions. It may also be used to preferentially amplify short fragments, thereby enriching fetal fraction. In some embodiments the addition of tags to the fragments by ligation can enable detection of shorter fragments, use of shorter target sequence specific portions of the primers and / or annealing at higher temperatures which reduces unspecific reactions. Circularizing Probes
[0109] Some embodiments of the present disclosure involve the use of “Linked Inverted Probes” (LIPs), which have been previously described in the literature. LIPs is a generic term meant to 31 4934-6677-8453.1Attorney Docket No. N.053.WO.01 encompass technologies that involve the creation of a circular molecule of DNA, where the probes are designed to hybridize to targeted region of DNA on either side of a targeted allele, such that addition of appropriate polymerases and / or ligases, and the appropriate conditions, buffers and other reagents, will complete the complementary, inverted region of DNA across the targeted allele to create a circular loop of DNA that captures the information found in the targeted allele. LIPs may also be called pre-circularized probes, pre-circularizing probes, or circularizing probes. The LIPs probe may be a linear DNA molecule between 50 and 500 nucleotides in length, and in an embodiment between 70 and 100 nucleotides in length; in some embodiments, it may be longer or shorter than described herein. Others embodiments of the present disclosure involve different incarnations, of the LIPs technology, such as Padlock Probes and MOLECULAR INVERSION PROBES (MIPs).
[0110] One method to target specific locations for sequencing is to synthesize probes in which the 3’ and 5’ ends of the probes anneal to target DNA at locations adjacent to and on either side of the targeted region, in an inverted manner, such that the addition of DNA polymerase and DNA ligase results in extension from the 3’ end, adding bases to single stranded probe that are complementary to the target molecule (gap-fill), followed by ligation of the new 3’ end to the 5’ end of the original probe resulting in a circular DNA molecule that can be subsequently isolated from background DNA. The probe ends are designed to flank the targeted region of interest. One aspect of this approach is commonly called MIPS and has been used in conjunction with array technologies to determine the nature of the sequence filled in. One drawback to the use of MIPs in the context of measuring allele ratios is that the hybridization, circularization and amplification steps do not happed at equal rates for different alleles at the same loci. This results in measured allele ratios that are not representative of the actual allele ratios present in the original mixture.
[0111] In an embodiment, the circularizing probes are constructed such that the region of the probe that is designed to hybridize upstream of the targeted polymorphic locus and the region of the probe that is designed to hybridize downstream of the targeted polymorphic locus are covalently connected through a non-nucleic acid backbone. This backbone can be any biocompatible molecule or combination of biocompatible molecules. Some examples of possible biocompatible molecules are poly(ethylene glycol), polycarbonates, polyurethanes, polyethylenes, polypropylenes, sulfone polymers, silicone, cellulose, fluoropolymers, acrylic compounds, styrene block copolymers, and other block copolymers. 32 4934-6677-8453.1Attorney Docket No. N.053.WO.01
[0112] In an embodiment of the present disclosure, this approach has been modified to be easily amenable to sequencing as a means of interrogating the filled in sequence. In order to retain the original allelic proportions of the original sample at least one key consideration must be taken into account. The variable positions among different alleles in the gap-fill region must not be too close to the probe binding sites as there can be initiation bias by the DNA polymerase resulting in differential of the variants. Another consideration is that additional variations may be present in the probe binding sites that are correlated to the variants in the gap-fill region which can result unequal amplification from different alleles. In an embodiment of the present disclosure, the 3’ ends and 5’ ends of the pre-circularized probe are designed to hybridize to bases that are one or a few positions away from the variant positions (polymorphic sites) of the targeted allele. The number of bases between the polymorphic site (SNV or otherwise) and the base to which the 3’ end and / or 5’ of the pre-circularized probe is designed to hybridize may be one base, it may be two bases, it may be three bases, it may be four bases, it may be five bases, it may be six bases, it may be seven to ten bases, it may be eleven to fifteen bases, or it may be sixteen to twenty bases, twenty to thirty bases, or thirty to sixty bases. The forward and reverse primers may be designed to hybridize a different number of bases away from the polymorphic site. Circularizing probes can be generated in large numbers with current DNA synthesis technology allowing very large numbers of probes to be generated and potentially pooled, enabling interrogation of many loci simultaneously. It has been reported to work with more than 300,000 probes.
[0113] In some embodiments of the methods disclosed herein, the genetic material of the target individual is optionally amplified, followed by hybridization of the pre-circularized probes, performing a gap fill to fill in the bases between the two ends of the hybridized probes, ligating the two ends to form a circularized probe, and amplifying the circularized probe, using, for example, rolling circle amplification. Once the desired target allelic genetic information is captured by circularizing appropriately designed oligonucleic probes, such as in the LIPs system, the genetic sequence of the circularized probes may be being measured to give the desired sequence data. In an embodiment, the appropriately designed oligonucleotides probes may be circularized directly on unamplified genetic material of the target individual, and amplified afterwards. Note that a number of amplification procedures may be used to amplify the original genetic material, or the circularized LIPs, including rolling circle amplification, MDA, or other amplification protocols. Different methods may be used to measure the genetic information on the target genome, for example using high throughput sequencing, Sanger sequencing, other sequencing methods, 33 4934-6677-8453.1Attorney Docket No. N.053.WO.01 capture-by-hybridization, capture-by-circularization, multiplex PCR, other hybridization methods, and combinations thereof.
[0114] Once the genetic material of the individual has been measured using one or a combination of the above methods, an informatics based method, along with the appropriate genetic measurements, can then be used to determination the cancer status of a subject.
[0115] It is important to note that LIPs may be used as a method for targeting specific loci in a sample of DNA for genotyping by methods other than sequencing. For example, LIPs may be used to target DNA for genotyping using SNP arrays or other DNA or RNA based microarrays. Ligation-mediated PCR
[0116] Ligation-mediated PCR is method of PCR used to preferentially enrich a sample of DNA by amplifying one or a plurality of loci in a mixture of DNA, the method comprising: obtaining a set of primer pairs, where each primer in the pair contains a target specific sequence and a non- target sequence, where the target specific sequence is designed to anneal to a target region, one upstream and one downstream from the polymorphic site, and which can be separated from the polymorphic site by 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11-20, 21-30, 31-40, 41-50, 51-100, or more than 100; polymerization of the DNA from the 3-prime end of upstream primer to the fill the single strand region between it and the 5-prime end of the downstream primer with nucleotides complementary to the target molecule; ligation of the last polymerized base of the upstream primer to the adjacent 5-prime base of the downstream primer; and amplification of only polymerized and ligated molecules using the non-target sequences contained at the 5-prime end of the upstream primer and the 3-prime end of the downstream primer. Pairs of primers to distinct targets may be mixed in the same reaction. The non-target sequences serve as universal sequences such that of all pairs of primers that have been successfully polymerized and ligated may be amplified with a single pair of amplification primers. Probe-Dependent Primers
[0117] In some embodiments, one or more primers herein are probe-dependent primers. Probe- dependent primers (PDPs) have been disclosed (Pel, et al. “Rapid and highly-specific generation of targeted DNA sequencing libraries enabled by linking capture probes with universal primers” PLoS ONE 13(12):e0208283 (2018); WO 2020 / 039261 “Linked target capture and ligation”, 34 4934-6677-8453.1Attorney Docket No. N.053.WO.01 which are hereby incorporated by reference in their entirety). Such embodiments can be considered LTC methods. Briefly, in an LTC method, PDPs are designed to incorporate non- extendable capture probes linked 5’ to 5’ with a primer. Multiple linker types are possible as discussed below. In some embodiments, probes of PDPs can be between 30 to 70 nucleotides in length and include or comprise a 3’ inverted dT base to inhibit polymerase extension. In some embodiments, probes are designed to cover the desired region with zero gap between forward and reverse probes. In some embodiments, the probes are between 20 and 100 nucleotides in length. In some embodiments, the size of the probe can be between 20 and 40 nucleotides, between 30 and 50 nucleotides, between 40 and 60, between 50 and 70 between 60 and 80, between 70 and 90, 80 and 100, 90 and 110, 100 and 120 nucleotides in length. In some embodiments, at least one of the probes of a PDP pair comprises a sample index.
[0118] In PDPs, forward and reverse probes can be designed to bind to nucleic acid sequences within or near a genomic region of interest on a sample DNA molecule to enrich nucleic acid molecules comprising the genomic region of interest or copies thereof. In some embodiments, at least one of the probe binding regions can include one or more somatic mutations. In some embodiments, one of the probe binding regions can include one or more somatic mutations. In some embodiments, both of the probe binding sites of the probe binding regions can include one or more somatic mutations. In some embodiments, neither of the probe binding sites of the probe binding region comprises a somatic mutation.
[0119] In some embodiments, the primer portion of a PDP is a universal primer designed to bind to a universal primer site on the ligated adapter. In some embodiments, the PDP is designed with a sequencer binding sequence, such as an Illumina flow cell binding sequence, incorporated therein. In some embodiments, the sequencer flow cell binding sequence is between the probe and universal primer, and adjacent to the primer. Linked primers of the invention may also include sequencing tags to ensure that all cluster reads originate from the same linked template molecule. The lengths of the primers can be extended or shortened at the 5' end or the 3' end to produce primers with desired melting temperatures. Also, the annealing position of each primer pair can be designed such that the sequence and length of the primer pairs yield the desired melting temperature. In illustrative embodiments, the primer is a low melting temperature universal primer complimentary to a portion of the ligated adapter.
[0120] The primer can be tailed or untailed one either side of the forward and reverse side depending on the specific requirements. In some embodiments, the universal primer comprises an 35 4934-6677-8453.1Attorney Docket No. N.053.WO.01 A tail. In some embodiments, the universal primer is blunt ended. The length of the primers of the PDP can range from 5 to 40 nucleotides in length. In certain embodiments, the PDP primers are between 10 and 25 nucleotides long. In embodiments, the primers of the PDP can range from 5 to 15 nucleotides, from 10 to 25 nucleotides, from 15 to 35 nucleotides, or from 25 to 40 nucleotides in length.
[0121] In some embodiments, probe dependent primers comprise a linker between the probe and the primer. In some embodiments, probe and primer portions of the PDP are linked by a polyethylene glycol derivative, an oligosaccharide, a lipid, a hydrocarbon, a polymer, or a protein. In some embodiments, the linker is a PEG molecule, or derivative thereof. In some embodiments, the linker is an oligosaccharide. In some embodiments, the linker is a lipid. In some embodiments, the linker is a hydrocarbon. In some embodiments, the linker is a polymer. In some embodiments, the linker is a protein, or portion thereof. Capture by Hybridization
[0122] Preferential enrichment of a specific set of sequences in a target genome can be accomplished in a number of ways. Elsewhere in this document is a description of how LIPs can be used to target a specific set of sequences, but in all of those applications, other targeting and / or preferential enrichment methods can be used equally well for the same ends. One example of another targeting method is the capture by hybridization approach. Some examples of commercial capture by hybridization technologies include AGILENT’s SURE SELECT and ILLUMINA’s TRUSEQ. In capture by hybridization, a set of oligonucleotides that is complimentary or mostly complimentary to the desired targeted sequences is allowed to hybridize to a mixture of DNA, and then physically separated from the mixture. Once the desired sequences have hybridized to the targeting oligonucleotides, the effect of physically removing the targeting oligonucleotides is to also remove the targeted sequences. Once the hybridized oligos are removed, they can be heated to above their melting temperature and they can be amplified. Some ways to physically remove the targeting oligonucleotides is by covalently bonding the targeting oligos to a solid support, for example a magnetic bead, or a chip. Another way to physically remove the targeting oligonucleotides is by covalently bonding them to a molecular moiety with a strong affinity for another molecular moiety. An example of such a molecular pair is biotin and streptavidin, such as is used in SURE SELECT. Thus, targeted sequences could be covalently attached to a biotin 36 4934-6677-8453.1Attorney Docket No. N.053.WO.01 molecule, and after hybridization, a solid support with streptavidin affixed can be used to pull down the biotinylated oligonucleotides, to which are hybridized to the targeted sequences.
[0123] Hybrid capture involves hybridizing probes that are complementary to the targets of interest to the target molecules. Hybrid capture probes were originally developed to target and enrich large fractions of the genome with relative uniformity between targets. In that application, it was important that all targets be amplified with enough uniformity that all regions could be detected by sequencing, however, no regard was paid to retaining the proportion of alleles in original sample. Following capture, the alleles present in the sample can be determined by direct sequencing of the captured molecules. These sequencing reads can be analyzed and counted according to the allele type. However, using the current technology, the measured allele distributions of the captured sequences are typically not representative of the original allele distributions.
[0124] In an embodiment, detection of the alleles is performed by sequencing. In order to capture the allele identity at the polymorphic site, it is essential that the sequencing read span the allele in question in order to evaluate the allelic composition of that captured molecule. Since the capture molecules are often of variable lengths upon sequencing cannot be guaranteed to overlap the variant positions unless the entire molecule is sequenced. However, cost considerations as well as technical limitations as to the maximum possible length and accuracy of sequencing reads make sequencing the entire molecule unfeasible. In an embodiment, the read length can be increased from about 30 to about 50 or about 70 bases can greatly increase the number of reads that overlap the variant positions within the targeted sequences.
[0125] Another way to increase the number of reads that interrogate the position of interest is to decrease the length of the probe, as long as it does not result in bias in the underlying enriched alleles. The length of the synthesized probe should be long enough such that two probes designed to hybridize to two different alleles found at one locus will hybridize with near equal affinity to the various alleles in the original sample. Currently, methods known in the art describe probes that are typically longer than 120 bases. In a current embodiment, if the allele is one or a few bases then the capture probes may be less than about 110 bases, less than about 100 bases, less than about 90 bases, less than about 80 bases, less than about 70 bases, less than about 60 bases, less than about 50 bases, less than about 40 bases, less than about 30 bases, and less than about 25 bases, and this is sufficient to ensure equal enrichment from all alleles. When the mixture of DNA that is to be enriched using the hybrid capture technology is a mixture comprising free floating 37 4934-6677-8453.1Attorney Docket No. N.053.WO.01 DNA isolated from blood, for example blood from a subject having or suspected of having cancer, the average length of DNA is quite short, typically less than 200 bases. The use of shorter probes results in a greater chance that the hybrid capture probes will capture desired DNA fragments. Larger variations may require longer probes. In an embodiment, the variations of interest are one (a SNP) to a few bases in length. In an embodiment, targeted regions in the genome can be preferentially enriched using hybrid capture probes wherein the hybrid capture probes are of a length below 90 bases, and can be less than 80 bases, less than 70 bases, less than 60 bases, less than 50 bases, less than 40 bases, less than 30 bases, or less than 25 bases. In an embodiment, to increase the chance that the desired allele is sequenced, the length of the probe that is designed to hybridize to the regions flanking the polymorphic allele location can be decreased from above 90 bases, to about 80 bases, or to about 70 bases, or to about 60 bases, or to about 50 bases, or to about 40 bases, or to about 30 bases, or to about 25 bases.
[0126] There is a minimum overlap between the synthesized probe and the target molecule in order to enable capture. This synthesized probe can be made as short as possible while still being larger than this minimum required overlap. The effect of using a shorter probe length to target a polymorphic region is that there will be more molecules that overlap the target allele region. The state of fragmentation of the original DNA molecules also affects the number of reads that will overlap the targeted alleles. Some DNA samples such as plasma samples are already fragmented due to biological processes that take place in vivo. However, samples with longer fragments by benefit from fragmentation prior to sequencing library preparation and enrichment. When both probes and fragments are short (~60-80 bp) maximum specificity may be achieved relatively few sequence reads failing to overlap the critical region of interest.
[0127] In an embodiment, the hybridization conditions can be adjusted to maximize uniformity in the capture of different alleles present in the original sample. In an embodiment, hybridization temperatures are decreased to minimize differences in hybridization bias between alleles. Methods known in the art avoid using lower temperatures for hybridization because lowering the temperature has the effect of increasing hybridization of probes to unintended targets. However, when the goal is to preserve allele ratios with maximum fidelity, the approach of using lower hybridization temperatures provides optimally accurate allele ratios, despite the fact that the current art teaches away from this approach. Hybridization temperature can also be increased to require greater overlap between the target and the synthesized probe so that only targets with substantial overlap of the targeted region are captured. In some embodiments of the present 38 4934-6677-8453.1Attorney Docket No. N.053.WO.01 disclosure, the hybridization temperature is lowered from the normal hybridization temperature to about 40°C, to about 45°C, to about 50°C, to about 55°C, to about 60°C, to about 65°C, or to about 70°C.
[0128] In an embodiment, the hybrid capture probes can be designed such that the region of the capture probe with DNA that is complementary to the DNA found in regions flanking the polymorphic allele is not immediately adjacent to the polymorphic site. Instead, the capture probe can be designed such that the region of the capture probe that is designed to hybridize to the DNA flanking the polymorphic site of the target is separated from the portion of the capture probe that will be in van der Waals contact with the polymorphic site by a small distance that is equivalent in length to one or a small number of bases. In an embodiment, the hybrid capture probe is designed to hybridize to a region that is flanking the polymorphic allele but does not cross it; this may be termed a flanking capture probe. The length of the flanking capture probe may be less than about 120 bases, less than about 110 bases, less than about 100 bases, less than about 90 bases, and can be less than about 80 bases, less than about 70 bases, less than about 60 bases, less than about 50 bases, less than about 40 bases, less than about 30 bases, or less than about 25 bases. The region of the genome that is targeted by the flanking capture probe may be separated by the polymorphic locus by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11-20, or more than 20 base pairs.
[0129] For small insertions or deletions, one or more probes that overlap the mutation may be sufficient to capture and sequence fragments comprising the mutation. Hybridization may be less efficient between the probe-limiting capture efficiency, typically designed to the reference genome sequence. To ensure capture of fragments comprising the mutation one could design two probes, one matching the normal allele and one matching the mutant allele. A longer probe may enhance hybridization. Multiple overlapping probes may enhance capture. Finally, placing a probe immediately adjacent to, but not overlapping, the mutation may permit relatively similar capture efficiency of the normal and mutant alleles.
[0130] For Simple Tandem Repeats (STRs), a probe overlapping these highly variable sites is unlikely to capture the fragment well. To enhance capture a probe could be placed adjacent to, but not overlapping the variable site. The fragment could then be sequenced as normal to reveal the length and composition of the STR.
[0131] For large deletions, a series of overlapping probes, a common approach currently used in exome capture systems may work. However, with this approach it may be difficult to determine whether or not an individual is heterozygous. Targeting and evaluating SNVs / SNPs within the 39 4934-6677-8453.1Attorney Docket No. N.053.WO.01 captured region could potentially reveal loss of heterozygosity across the region indicating that an individual is a carrier. In an embodiment, it is possible to place non-overlapping or singleton probes across the potentially deleted region and use the number of fragments captured as a measure of heterozygosity. In the case where an individual carries a large deletion, one-half the number of fragments are expected to be available for capture relative to a non-deleted (diploid) reference locus. Consequently, the number of reads obtained from the deleted regions should be roughly half that obtained from a normal diploid locus. Aggregating and averaging the sequencing read depth from multiple singleton probes across the potentially deleted region may enhance the signal and improve confidence of the diagnosis. The two approaches, targeting SNVs / SNPs to identify loss of heterozygosity and using multiple singleton probes to obtain a quantitative measure of the quantity of underlying fragments from that locus can also be combined. Either or both of these strategies may be combined with other strategies to better obtain the same end.
[0132] There are a number of ways to decrease depth of read (DOR) variability: for example, one could increase primer concentrations, one could use longer targeted amplification probes, or one could run more STA cycles (such as more than 25, more than 30, more than 35, or even more than 40). Primer Design
[0133] Multiplexed PCR can often result in the production of a very high proportion of product DNA that results from unproductive side reactions such as primer dimer formation. In an embodiment, the particular primers that are most likely to cause unproductive side reactions may be removed from the primer library to give a primer library that will result in a greater proportion of amplified DNA that maps to the genome. The step of removing problematic primers, that is, those primers that are particularly likely to firm dimers has unexpectedly enabled extremely high PCR multiplexing levels for subsequent analysis by sequencing. In systems such as sequencing, where performance significantly degrades by primer dimers and / or other mischief products, greater than 10, greater than 50, and greater than 100 times higher multiplexing than other described multiplexing has been achieved. Note this is opposed to probe-based detection methods, e.g., microarrays, TAQMAN, PCR etc. where an excess of primer dimers will not affect the outcome appreciably.
[0134] There are a number of ways to choose primers for a library where the amount of non- mapping primer-dimer or other primer mischief products are minimized. Empirical data indicate 40 4934-6677-8453.1Attorney Docket No. N.053.WO.01 that a small number of ‘bad’ primers are responsible for a large amount of non-mapping primer dimer side reactions. Removing these ‘bad’ primers can increase the percent of sequence reads that map to targeted loci. One way to identify the ‘bad’ primers is to look at the sequencing data of DNA that was amplified by targeted amplification; those primer dimers that are seen with greatest frequency can be removed to give a primer library that is significantly less likely to result in side product DNA that does not map to the genome. There are also publicly available programs that can calculate the binding energy of various primer combinations, and removing those with the highest binding energy will also give a primer library that is significantly less likely to result in side product DNA that does not map to the genome.
[0135] When primers can be designed it is possible to attempt to identify primer pairs likely to form spurious products by evaluating the likelihood of spurious primer duplex formation between all possible pairs of primers using published thermodynamic parameters for DNA duplex formation. Primer interactions may be ranked by a scoring function related to the interaction and primers with the worst interaction scores are eliminated until the number of primers desired is met. In cases where SNVs / SNPs likely to be heterozygous are most useful, it is possible to also rank the list of assays and select the most heterozygous compatible assays. Experiments have validated that primers with high interaction scores are most likely to form primer dimers.
[0136] Note that there are other methods for determining which PCR probes are likely to form dimers. In an embodiment, analysis of a pool of DNA that has been amplified using a non- optimized set of primers may be sufficient to determine problematic primers. For example, analysis may be done using sequencing, and those dimers which are present in the greatest number are determined to be those most likely to form dimers and may be removed.
[0137] The use of tags on the primers may reduce amplification and sequencing of primer dimer products. Tag-primers can be used to shorten necessary target-specific sequences to below 20, below 15, below 12, and even below 10 base pairs. This can be serendipitous with standard primer design when the target sequence is fragmented within the primer binding site or, or it can be designed into the primer design. Advantages of this method include at least the following: it increases the number of assays that can be designed for a certain maximal amplicon length, and it shortens the “non-informative” sequencing of primer sequence. It may also be used in combination with internal tagging (see elsewhere in this document).
[0138] In an embodiment, the relative amount of nonproductive products in the multiplexed targeted PCR amplification can be reduced by raising the annealing temperature, lowering primer 41 4934-6677-8453.1Attorney Docket No. N.053.WO.01 concentrations and using longer annealing times. For example, methods of designing and optimizing primers can be found in U.S. Patent No. 11,312,996, incorporated herein by reference.
[0139] To select target locations, one may start with a pool of candidate primer pair designs and create a thermodynamic model of potentially adverse interactions between primer pairs, and then use the model to eliminate designs that are incompatible with other the designs in the pool. Primers with adaptor binding region
[0140] One issue with fragmented DNA is that since it is short in length, the chance that a variant is close to the end of a DNA strand is higher than for a long strand. Since PCR capture of a variant requires a primer binding site of suitable length on both sides of the variant, a significant number of strands of DNA with the targeted variant will be missed due to insufficient overlap between the primer and the targeted binding site. In cases where the binding region is shorter than the 18 bp typically required for hybridization, the region (cr) on the primer than is complementary to the library tag is able to increase the binding energy to a point where the PCR can proceed. Note that any specificity that is lost due to a shorter binding region can be made up for by other PCR primers with suitably long target binding regions. Note that this embodiment can be used in combination with direct PCR, or any of the other methods described herein, such as nested PCR, semi nested PCR, hemi nested PCR, one sided nested or semi or hemi nested PCR, or other PCR protocols. Diagnostic Box
[0141] In an embodiment, the present disclosure comprises a diagnostic box that is capable of partly or completely carrying out any of the methods described in this disclosure. In some embodiments, a diagnostic box refers to one or a combination of machines designed to perform one or a plurality of aspects of the methods disclosed herein. In an embodiment, the diagnostic box may perform targeted amplification followed by sequencing. In an embodiment, the diagnostic box may be placed at a point of patient care. For example, the diagnostic box may be located at a physician’s office, a hospital laboratory, or any suitable location reasonably proximal to the point of patient care. The box may be able to run the entire method in a wholly automated fashion, or the box may require one or a number of steps to be completed manually by a technician. In an embodiment, the box may be able to analyze at least the genotypic data measured 42 4934-6677-8453.1Attorney Docket No. N.053.WO.01 on the samples from the subject. In an embodiment, the box may be linked to means to transmit the genotypic data measured on the diagnostic box to an external computation facility which may then analyze the genotypic data, and possibly also generate a report. The diagnostic box may include a robotic unit that is capable of transferring aqueous or liquid samples from one container to another. It may comprise a number of reagents, both solid and liquid. It may comprise a high throughput sequencer. It may comprise a computer. Kits
[0142] In some embodiments, a kit may be formulated that comprises a plurality of primers designed to achieve the methods described in this disclosure. The primers may be outer forward and reverse primers, inner forward and reverse primers as disclosed herein, they could be primers that have been designed to have low binding affinity to other primers in the kit as disclosed in the section on primer design, they could be hybrid capture probes or pre-circularized probes as described in the relevant sections, or some combination thereof. In an embodiment, a kit may be formulated for determining the cancer status of a cancer patient or a subject suspected of having cancer and designed to be used with the methods disclosed herein, the kit comprising a plurality of inner forward primers and optionally the plurality of inner reverse primers, and optionally outer forward primers and outer reverse primers, where each of the primers is designed to hybridize to the region of DNA immediately upstream and / or downstream from one of the polymorphic sites on the target chromosome, and optionally additional chromosomes. In an embodiment, the primer kit may be used in combination with the diagnostic box described elsewhere in this document. Variant and Sample Calling
[0143] Methods as described herein include detecting and optionally quantifying nucleic acids, including DNA, cfDNA, enriched subsets of DNA having target regions, and in illustrative embodiments, target region amplicons or amplicons derived therefrom. In some embodiments, cfDNA from a blood sample from the individual is analyzed. Not to be limited by theory, cfDNA is believed to be released from certain cells, such as cancer cells, for example when they undergo necrosis or apoptosis. In some embodiments, methods herein can be used to detect somatic mutations in target regions or nucleic acid sequence of interest that is present in a small percentage 43 4934-6677-8453.1Attorney Docket No. N.053.WO.01 of DNA in a sample, such as cfDNA, for example from a fetus, a cell from a donated organ, or in illustrative embodiments, a cancer cell.
[0144] In some embodiments, cellular DNA from normal tissue, such as the buffy coat of the blood samples from the individual not suspected of having blood cancer, is analyzed. Non-tumor specific mutations, such as CHmutations, potentially contribute to a significant number of false positive tumor calls. Accordingly, in some embodiments, matched normal tissue samples, in some embodiments buffy coat samples and in some embodiments whole blood samples, from the individual may be used to enable identification and bioinformatic elimination of CH and other non-tumor specific mutations to reduce false positive variants and enhance the specificity of the methods described herein.
[0145] In some embodiments, at least some of the regions of the cfDNA comprising somatic mutations can be enriched using a set of hybrid capture probes to form at least some of the enriched subsets of target loci before detecting or quantifying an amount for at least some of the enriched subsets. In methods that include detecting or quantifying the target region amplicons, such target region amplicons can be enriched using a set of hybrid capture probes before detecting or quantifying. In some embodiments, methods here in can include detecting or quantifying the target region amplicons without a selective enrichment step.
[0146] In some embodiments, the method further comprises an additional amplification reaction that amplifies at least some of the amplified target region amplicons or amplifies at least some of the enriched subsets of amplified DNA, for detecting or quantifying. The additional amplification reaction can be, for example, a quantitative PCR (qPCR reaction) such as TAQMAN assay (LIFE TECHNOLOGIES), or an INVADER assay (THIRD WAVE TECHNOLOGIES), a digital PCR, or any other method for detecting and / or quantifying a target DNA, which in methods described herein comprise a target region that includes one or more SNVs / InDels.
[0147] In some embodiments, the additional amplification reaction is a clonal amplification reaction to form clonally amplified target region amplicons, or to form clonally amplified enriched target loci. For detecting or quantifying at least some of the amplified and / or selectively enriched target region amplicons, methods herein can comprise i) performing an additional amplification reaction that amplifies at least some of the amplified target region amplicons, wherein the additional amplification reaction is a clonal amplification reaction to form clonally amplified target region amplicons, and ii) performing a next-generation sequencing reaction on the clonally amplified target region amplicons. For detecting or quantifying at least some of the enriched 44 4934-6677-8453.1Attorney Docket No. N.053.WO.01 subsets of amplified DNA, methods herein can comprise i) performing an additional amplification reaction that amplifies at least some of the enriched subsets of amplified DNA, wherein the additional amplification reaction is a clonal amplification reaction to form clonally amplified enriched subsets of amplified DNA, and ii) performing a next-generation sequencing reaction on the clonally amplified enriched subsets of amplified DNA. The detecting or quantifying included in methods herein, can comprise performing a sequencing reaction on the clonally amplified target region amplicons.
[0148] In some embodiments, the sequencing reaction is a next-generation sequencing reaction. In methods herein, detecting or quantifying comprises counting sequence reads generated from target region amplicons. Quantifying can also comprise determining a depth of read (DOR) per target region for at least some of the target regions. DOR for each of the target region can be normalized relative to a DOR for a normalization sequence. DNA sequences used for normalization can be derived from genomic DNA or from control plasmids and will depend on the experimental conditions. The normalization sequence can be derived from a lambda control plasmid or synthetic sequence. The normalization sequence can be a spike-in control DNA sample. A spike-in control DNA sample can be a genomic, plasmid, or synthetic DNA sample. Methods herein, can include more than one, for example 2, 3, 4, 5, or more control or spike-in control sample. DNA sequencing techniques, particularly high throughput next-generation sequencing techniques (often referred to as massively parallel sequencing techniques) such as those employed in MYSEQ (ILLUMINA), HISEQ (ILLUMINA), ION TORRENT (LIFE TECHNOLOGIES), GENOME ANALYZER ILX (ILLUMINA), GS FLEX+(ROCHE 454) etc., can be used for quantitative measurements of the number of copies of a target region present, for example, but not limiting to, target region amplicons or enriched subsets of amplified DNA, and thus provide quantitative information regarding the number and / or amount of target regions in sample DNA molecules. High throughput genetic sequencers are amenable to the use of barcoding (i.e., sample tagging with distinctive nucleic acid sequences) so as to identify specific samples from individuals thereby permitting the simultaneous analysis of multiple samples in a single run of the DNA sequencer. The number of times a given region of the genome in a library preparation (or other nucleic preparation of interest) is sequenced (number of reads) will be proportional to the number of copies of that sequence. Methods as described herein that utilize NGS detection, in some embodiments can have an average DOR of at least 200, 500, 1000, 2000, 2900, 3000, 3500, 4000, 5000, 10,000, 50,000, 75,000, 100,000, 120,000, 130,000, 150,000, 175,000, or 200,000, per 45 4934-6677-8453.1Attorney Docket No. N.053.WO.01 target locus. In some embodiments, described herein, the sequencing has a depth of read (DOR) of between 50,000 to 200,000 per target locus.
[0149] Methods herein can include analyzing data obtained from next-generation sequencing techniques. In some embodiments of methods herein, target region amplicons or enriched subsets of amplified DNA can be subjected to sequencing using next-generation sequencing techniques. Nucleic acid sequencing data can be generated for amplicons created by PCR, for example a multiplex targeted PCR. In some embodiments, the multiplex PCR can be a tiled multiplex PCR. For a skilled artisan, algorithm design tools are available that can be used and / or adapted to analyze the sequencing data. In addition, those skilled in the art can determine appropriate parameters for measuring alignment to a consensus sequence and / or to a known target region sequence, including any algorithms needed to achieve maximal alignment over the length of the sequences being compared.
[0150] Sequencing reads can be demultiplexed using an in-house tool and mapped using the Burrows-Wheeler alignment software, Bwa mem function (BWA, Burrows-Wheeler Alignment Software (see Li H. and Durbin R. (2010) Fast and accurate long-read alignment with Burrows- Wheeler Transform. Bioinformatics.) on single end mode using pear merged reads to the hg19 genome. Amplification statistics QC can be performed by analyzing one or more of, but not limiting to, total reads, number of mapped reads, number of mapped reads on target, and number of reads counted.
[0151] Methods herein can include a background error model that can be constructed using normal, or healthy liquid samples, in illustrative embodiments, normal, or healthy plasma samples, which are sequenced on the same sequencing run to account for run-specific artifacts. In some embodiments, 5, 10, 15, 20, 25, 30, 40, 50, 100, 150, 200, 250, or more than 250 normal, or healthy liquid samples, in illustrative embodiments, plasma samples can be analyzed on the same sequencing run. The number of samples that can be sequenced on the same sequencing run can be in the range of 5 to 500, 5 to 400, 5 to 300, 5 to 250, 20 to 250, 30 to 250, 50 to 250, 75 to 250, 100 to 250, 50 to 500, or 100 to 500. Sample barcodes are used in illustrative embodiments. In some illustrative embodiments, 20, 25, 40, or 50 normal samples (e.g., plasma samples) can be analyzed on the same sequencing run. Outlier samples can be iteratively removed from the model to account for noise and contamination. In some embodiments, samples with a Z score of greater than 5, 6, 7, 8, 9, or 10 are removed from the data analysis. For each base substitution of every genomic loci, the DOR weighted mean and standard deviation of the error can be calculated. 46 4934-6677-8453.1Attorney Docket No. N.053.WO.01
[0152] Methods herein can include calculating percent identity that can be calculated by determining the number of matched positions in aligned DNA sequences, dividing the number of matched positions by the total number of aligned DNA sequences, and multiplying by 100. A matched position refers to a position in which identical nucleotides occur at the same position in aligned DNA sequences. The percent identity over a particular length can be determined by counting the number of matched positions over that length and dividing that number by the length followed by multiplying the resulting value by 100. A non-limiting example for calculating the percent identity, can be, if (i) a 500-nucleotide DNA target sequence is compared to a subject DNA sequence, (ii) an alignment program presents 200 nucleotides from the target DNA sequence aligned with a region of the subject DNA sequence where the first and last nucleotides of that 200- nucleotide region are matches, and (iii) the number of matches over those 200 aligned nucleotides is 180, then the 500-nucleotide nucleic acid target sequence contains a length of 200 and a sequence identity over that length of 90 percent (i.e., 180, 200x100=90).
[0153] In some embodiments, the uniformity in DOR can be measured using standard methods such as, but not limiting to, DOR slope, normalized median depth of read (nmDOR), or breadth of read (BOR). DOR slope represents the slope of the line in the linear portion of a list of loci sorted in descending DOR order. Closer to zero is better, as it represents a flat line. In some embodiments, the uniformity in DOR can be measured using the percent of reads in the 90th-95thpercentile. For this measurement, the loci are sorted in descending DOR order. In illustrative embodiments, a DOR distribution using the 90th-95thpercentile contains 5 percent of reads. The reads of all loci between the 90thpercentile and 95thpercentile can be counted and divided by the total reads for all loci.
[0154] In some embodiments, the magnitude of the DOR slope can be less than 0.005, 0.001, 0.0005, 0.0001, 0.00005, 0.00001, 0.000005, or 0.000001. The magnitude of the DOR slope can be between 0 and 0.005, such as 0.000001 to 0.005, such as between 0.000005 to 0.00001, 0.00001 to 0.00005, 0.00005 to 0.0001, 0.0001 to 0.0005, 0.0005 to 0.001, or 0.001 to 0.005. The percent of reads in the 90th-95thpercentile can be between 0.2 and 9 percent, such as between 0.2 to 8 percent, 0.2 to 7 percent, 0.2 to 6 percent, 0.4 to 9 percent, 0.4 to 8 percent, 0.4 to 7 percent, 0.4 to 6 percent, 1 to 9 percent, 1 to 8 percent, 1 to 7 percent, 1 to 6 percent, 2 to 9 percent, 2 to 8 percent, 2 to 7 percent, 2 to 6 percent, 3 to 9 percent, 3 to 8 percent, 3 to 7 percent, 3 to 6 percent, 0.2 to 1.0 percent, 1 to 2 percent, 2 to 3 percent, 2 to 4 percent, 3 to 4 percent, 4 to 5 percent, 5 to 6 percent, or 6 to 8 percent, or 7 to 9 percent. In some embodiments of methods herein, the method 47 4934-6677-8453.1Attorney Docket No. N.053.WO.01 or the amplification steps in the method can produce a composition comprising at least 100 different amplicons (e.g., at least 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000, 50,000, 75,000, or 100,000 non-identical amplicons) with the magnitude of the DOR slope in any of the ranges herein, or with a percent of reads in the 90th-95thpercentile in any of the ranges herein. In some embodiments, different amplicons can range in between 100 to 500,000, 100 to 400,000, 100 to 300,000, 100 to 200,000, 100 to 100,000, 100 to 75,000, 100 to 50,000, 100 to 40,000, 100 to 30,000, 100 to 25,000, 100 to 20,000, or 100 to 15,000 non-identical amplicons.
[0155] In some embodiments of methods herein, in addition, or, in some embodiments, as an alternative to analyzing an altered (increased or decreased) methylation levels in a sample, one or more other factors can be analyzed if desired. These factors can be used to increase the accuracy of the diagnosis (such as determining the presence or absence of cancer or an increased risk for cancer, classifying the cancer, or staging the cancer) or prognosis. These factors can also be used to select a particular therapy or treatment regimen that is likely to be effective in the subject. Limits of Detection
[0156] Exemplary methods herein are to detect target region amplicons generated by a targeted amplification and / or selectively enriched, and used to determine the status of one or more, in illustrative embodiments a plurality of, somatic mutations on a DNA molecule in a DNA sample. A target region can include 1, 2, 3, 4, 5, 6 or more sites that are predominantly mutated in cells of a certain origin (e.g. a cancer). In illustrative embodiments, sample DNA molecules comprising the target regions are fragments of genomic DNA or circulating free DNA (cfDNA) of a subject. Target regions can be selectively amplified and / or selectively enriched using methods herein. Exemplary methods herein, in some embodiments, have a mean or median limit of detection of as low as 1.0%, 0.5%, 0.1%, 0.05%, 0.01%, 0.005%, 0.002%, or 0.001%, wherein the method is capable of detecting somatic mutations present at 1.0%, 0.5%, 0.1%, 0.05%, 0.01%, 0.005%, 0.002%, or 0.001% or more in a mixture of DNA molecules.
[0157] In some embodiments, measurements can be adjusted for bias, such as bias due to differences in amplification efficiency or adjusted for sequencing errors. In some embodiments, differentiation between mutated and non-mutated samples can be analyzed using a control sample that is non-mutated, for example a healthy sample for normalization of quantitative results. In some embodiments, the differentiation between the mutated and non-mutated samples recited in 48 4934-6677-8453.1Attorney Docket No. N.053.WO.01 this paragraph can be achieved after normalization of a detected and quantified signal using one or more (e.g. 2, 3, 4, 5, or 6) controls that do not have the mutation present.
[0158] In certain embodiments, the method is capable of detecting ctDNA when the somatic mutation is present in 10%, 5%, 2.5%, 2%, 1%, 0.5%, 0.1%, 0.05%, 0.01%, or 0.001% of the circulating free DNA in a sample. In some embodiments, methods herein detect or are capable of detecting ctDNA from a sample when it is present at a range between 0.01% on the low end and 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5%, 2.5%, 1%, or 0.5% on the high end of the range of total cfDNA in the sample. Methods herein, in some embodiments, detect or are capable of detecting 5%, 4%, 3%, 2%, 1%, 0.5%, 0.1%, 0.05%, 0.01% or less ctDNA, as a percentage of total cfDNA from a sample. The percentage of ctDNA in a cfDNA sample can be determined or approximated by VAF of one or more DNA variants, such as single nucleotide variants (“SNVs”), present in tumor cells but not in normal cells. (See e.g., WO2019200228, incorporated by reference herein in its entirety). VAF can be used as a surrogate measure of the percentage of ctDNA in the cfDNA sample. Bioinformatics
[0159] Any of the embodiments disclosed herein may be implemented in digital electronic circuitry, integrated circuitry, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, or in combinations thereof. Apparatus of the presently disclosed embodiments can be implemented in a computer program product tangibly embodied in a machine-readable storage device for execution by a programmable processor; and method steps of the presently disclosed embodiments can be performed by a programmable processor executing a program of instructions to perform functions of the presently disclosed embodiments by operating on input data and generating output. The presently disclosed embodiments can be implemented advantageously in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device. Each computer program can be implemented in a high-level procedural or object-oriented programming language or in assembly or machine language if desired; and in any case, the language can be a compiled or interpreted language. A computer program may be deployed in any form, including as a stand-alone program, or as a module, component, subroutine, or other unit 49 4934-6677-8453.1Attorney Docket No. N.053.WO.01 suitable for use in a computing environment. A computer program may be deployed to be executed or interpreted on one computer or on multiple computers at one site, or distributed across multiple sites and interconnected by a communication network.
[0160] Computer readable storage media, as used herein, refers to physical or tangible storage (as opposed to signals) and includes without limitation volatile and non-volatile, removable and non- removable media implemented in any method or technology for the tangible storage of information such as computer-readable instructions, data structures, program modules or other data. Computer readable storage media includes, but is not limited to, RAM, ROM, EPROM, EEPROM, flash memory or other solid state memory technology, CD-ROM, DVD, or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other physical or material medium which can be used to tangibly store the desired information or data or instructions and which can be accessed by a computer or processor.
[0161] Any of the methods described herein may include the output of data in a physical format, such as on a computer screen, or on a paper printout. In explanations of any embodiments elsewhere in this document, it should be understood that the described methods may be combined with the output of the actionable data in a format that can be acted upon by a physician. In addition, the described methods may be combined with the actual execution of a clinical decision that results in a clinical treatment, or the execution of a clinical decision to make no action. Some of the embodiments described in the document for determining genetic data pertaining to a target individual may be combined with a clinical decision or action. Some of the embodiments described in the document for determining genetic data pertaining to a target individual may be combined with the notification of a potential cancer status, or lack thereof, with a medical professional. Some of the embodiments described herein may be combined with the output of the actionable data, and the execution of a clinical decision that results in a clinical treatment, or the execution of a clinical decision to make no action.
[0162] For example, in some embodiments of the methods described herein, a system may be used that includes a computing system (which may be or may include one or more computing devices, co-located or remote to each other), a sequencing system (capable of, e.g., generating DNA sequencing data), an information system (such as a management or clinical information system that may record and / or provide information regarding samples or patients), vessels (e.g., for holding samples), sample processing system (e.g., a system that may include robotic components for preparing, moving, and / or processing samples), and imaging and / or therapy systems (which 50 4934-6677-8453.1Attorney Docket No. N.053.WO.01 may include e.g., treatment unit and imager / sensors for imaging patients using various imaging modalities, providing treatments such as radiotherapy, and / or performing other evaluation or therapeutic procedures based on what is learned from patient samples). The computing system (e.g., one or more computing devices) may be used to control and / or exchange signals and / or data with sequencing system, sample processing system, information system, and / or imaging and / or therapy systems, directly (e.g., through wireless and / or wired communication) or indirectly via another component of system (e.g., via any combination of wireless and / or wired communication). In certain embodiments, computing system may be used to control and / or exchange data or other signals with sequencing system, information system, sample processing system, and / or imaging and / or therapy systems. The computing system may include one or more processors and one or more volatile and / or non-volatile memories for storing computing code and data that are captured, acquired, recorded, and / or generated.
[0163] The computing system may include a controller that is configured to exchange control signals with sequencing system, information system, sample processing system, and / or imaging and / or therapy systems, and / or any components thereof, allowing the computing system to be used to control, for example, processing of samples, capture of images, acquisition of signals by sensors, positioning or repositioning of samples being tested or imaged, recording or obtaining other information, and / or running assays or tests.
[0164] A transceiver allows the computing system to exchange readings, control commands, and / or other data or signals, wirelessly or via wires, directly or indirectly via networking protocols, with, for example, other components of system. One or more user interfaces allow the computing system to receive user inputs (e.g., via a keyboard, touchscreen, microphone, camera, biometric scanners, etc.) and provide outputs (e.g., via a display screen, audio speakers, light emitters, etc.) with users. The computing system may additionally include one or more databases for storing, for example, data acquired from one or more systems or devices, signals acquired via one or more sensors, sequencing data, biomarkers, images, etc. In some implementations, database (or portions thereof) may alternatively or additionally be part of another computing device that is co-located or remote (e.g., via “cloud computing”) and in communication with computing system, sequencing system, information system, sample processing system, imaging and / or therapy systems, and / or components thereof.
[0165] Imaging and / or therapy systems may include a treatment unit, which may include any systems and / or devices employed to administer a treatment to a patient, such as a radiation therapy 51 4934-6677-8453.1Attorney Docket No. N.053.WO.01 system. Imager and / or sensors may be or may include, for example, any system or device that is involved in, for example, capturing images of patients and / or samples. Imagers may include detectors for visible light and / or light in any frequencies of interest, such as (but not limited to) the spectrum from infrared to ultraviolet. Imagers may employ any suitable optical components (e.g., lenses, mirrors, filters, beam splitters, prisms, diffusers, diffraction gratings, etc.), digital components (e.g., charge-coupled devices (CCDs), as well as an area for placement of samples, computing components (e.g., one or more processors, such as digital signal processors) to process and / or pre-process images, etc. Imagers may be, or may employ, any microscopes or microscopy systems (e.g., confocal microscopy), tools, and / or techniques that will provide the desired imaging data. The imager and / or sensors may have the capability of receiving control signals from a computing device or system to, for example, initiate or cease image capture, and / or return images or other imaging data or status signals to the computing device or system (e.g., computing system). Imaging and / or therapy systems may include any tools for obtaining additional data about samples and / or patients (such as spectrometers, chromatographs, etc.). Sensors may be used to detect, for example, other aspects of samples and / or patients, such as temperature, humidity, location, etc.
[0166] Sample processing system may include any components used to automate preparation of samples and running tests on samples. Robotics, for example, may include any combination of actuators, stepper motors, servomotors, control system, robot arms, end effectors like grippers and manipulators, etc. Robotics may be, or may comprise, for example, an automated robotics system with process automation software. In various embodiments, robotics may employ vessels, such as vials or plates, that can be moved and manipulated, for example, for positioning of samples to be evaluated using sequencing system or components thereof. Sample processing system may include components that may be used to maintain the integrity of samples during storage, transport, and testing, such as heaters, passive or active coolers, etc. Sensors may be used by sample processing system to, for example, monitor, guide, and evaluate the progress of tests (e.g., by detecting temperature, fill level of vessels, or other states or conditions). Sequencing system may include sequencing platforms or devices as well as various analytical tools used to process data from the sequencing platforms. Analytical tools may include software and / or hardware used to generate computer files with raw and / or processed sequencing data.
[0167] In various implementations, components of system may be rearranged or integrated in other configurations. For example, computing system (or components thereof) may be integrated 52 4934-6677-8453.1Attorney Docket No. N.053.WO.01 with one or more of the sequencing system, sample processing system, and / or components thereof. The sequencing system, sample processing system, and / or components thereof may be directed to a vessel in which a sample can be situated (e.g., so as to test or evaluate biological samples). In various embodiments, the vessel may be movable (e.g., using any combination of motors, magnets, etc.) to allow for positioning and repositioning of samples (such as micro-adjustments for positioning of samples to be sequenced). It is also noted that not all components of system are required to implement the disclosed approach, and in various embodiments, only a subset of the components of system may be employed. For example, in various embodiments, computing system may obtain and process data that was obtained via sequencing system or another system that is or is not in direct communication with the computing system. In certain embodiments, data may be obtained for samples that were not processed in an automated fashion but rather manually by one or more users.
[0168] A data acquisition unit may retrieve, acquire, or otherwise obtain various data, such as sequencing data. The data acquisition unit may, for example, obtain data stored in database, sequencing system, information system, sample processing system, and / or imaging and / or therapy systems. An interaction unit may interact (e.g., via user interfaces) with users (e.g., laboratory technicians, data scientists, clinicians, etc.) to obtain information or commands needed for system or components thereof to function, or to start operations, repeat processes, and / or stop operations. In certain embodiments, the data acquisition unit may obtain data from users via interaction unit.
[0169] A preprocessor may perform any number of steps disclosed herein in preprocessing sequencing data. Machine learning platform may be configured to train and update machine learning models, as further discussed herein. Machine learning platform may, for example, employ certain machine learning techniques and algorithms to train and update predictive models. Machine learning platform may include a training engine which may, for example, fit models to sequencing data. The training engine may, for example, obtain sequencing data from or via data acquisition unit, sequencing system or components thereof, and / or information system. Inference engine may apply models (e.g., models obtained from machine learning platform or components thereof) to generate outputs as disclosed herein. Biomarker generator may perform analyses on data from machine learning platform to, for example, generate biomarkers to be used for various conditions or sub-populations.
[0170] A reporter may generate reports that include, for example, information on sample sequencing (e.g., via sequencing system or components thereof), imaging or administration of 53 4934-6677-8453.1Attorney Docket No. N.053.WO.01 treatments (e.g., via imaging and / or therapy systems or components thereof and / or via information system), model training (e.g., via machine learning platform), and biomarker generation (e.g., via biomarker generator). For example, reporter may provide or identify trained and / or updated models, and / or biomarkers. In certain embodiments, reporter may obtain information to be reported from, for example, database and / or information system. Data that may be reported may be, for example, transmitted to another system or device (which may or may not be part of system), saved in non-transitory computer-readable storage media (e.g., database, information system, and / or elsewhere), and / or presented or otherwise provided to users (e.g., via a display device that is part of user interfaces). Panel Design
[0171] Only a relatively small subset of somatic mutations, such as SNVs and InDels, accumulated before and during tumorigenesis – including driver and passenger, clonal and subclonal mutations – contribute to tumor inception, and are therefore recurrent hotspot mutations that can be leveraged for early detection of asymptomatic cancers or recurrence, relapse or metastasis of cancers. It is crucial to identify the most prevalent of these mutations to develop a cost-effective and accurate somatic mutation panel for early detection of asymptomatic cancers or recurrence, relapse or metastasis of cancers. Further, tumor-related mutation prevalence may vary by a number of factors, such as ancestral population, patient age, cancer subtype, anatomical location (left vs right colon), cancer stage, smoking, high-risk germline mutation carriers, etc. Sub- population level assessment of a putative panel can inform panel design strategy.
[0172] For cancer detection, a tumor-informed, patient-bespoke approach may not be viable for detecting asymptomatic cancers or when tissue biopsies are unavailable from a patient. In some embodiments, a cancer panel that relies on population-wide somatic mutation patterns and prior knowledge of their prevalence, distribution, and relevance is created. In some embodiments, a plurality of targets are identified based on population-wide SNV and / or InDel patterns and prior knowledge. In some embodiments, a panel of targets may be designed using synonymous and non- synonymous mutations from cancer patients and additional computational approaches described herein to encode prior knowledge and tumor SNVs from cancer samples and trials with comprehensive clinical annotations to ensure optimal mutation coverage across subpopulations. In some embodiments, a clustering algorithm (such as DB-SCAN) is used to find areas of high target density. In some embodiments, improved algorithms disclosed herein are used to identify optimal 54 4934-6677-8453.1Attorney Docket No. N.053.WO.01 cancer-driving mutations with a balanced performance across various subpopulations. For example, computational approaches disclosed herein were used to encode prior knowledge and tumor mutations from more than 20,000 CRC samples from clinical studies including the Circulate and the US Bespoke trials with comprehensive clinical annotations to ensure optimal mutation coverage across subpopulations. Exemplary targets from chromosomes 9 and 22 that may be included in a panel for CRC screening or detection according to various embodiments are shown below: chrom start end gene chr9 127765840 127765875 SCAI chr9 990862 990878 DMRT3 chr9 35091448 35091468 PIGO chr22 36877667 36877720 TXN2 chr22 24176331 24176370 SMARCB1 chr22 30823252 30823270 MTFP1 chr22 29440867 29440878 ZNRF3
[0173] The methods described herein are further applicable for identifying tumor-specific mutations for other cancers or precancerous adenoma. Error Model
[0174] In some embodiments, the methods described herein may include a position-based somatic variant error model, e.g., an SNV / InDel error model. A cohort of healthy plasma samples can be used for error model construction. In some embodiments, 10 healthy samples, 50 healthy samples, 100 healthy samples, 200 healthy samples, 500 healthy samples, or 1000 healthy samples may be used for error model construction.
[0175] In some embodiments, per cycle PCR efficiency and error rates are estimated for each target in the panel.
[0176] Using a set of normal samples that are not expected to have any cancer related mutation, the per position efficiency and error rate per cycle can be estimated. The error model accounts for DNA input amount by adjusting the number of library prep cycles based on input amount. Calculating Mean Sample VAF 55 4934-6677-8453.1Attorney Docket No. N.053.WO.01
[0177] Methods herein can include quantification of ctDNA in a sample, such as an MRD sample, by calculating a sample VAF. In some embodiments, only consensus variants identified in both replicates are used for calculating a mean sample VAF. In some embodiments, statistical weighting is further applied. In some embodiments, mean sample VAF may be determined by computing the weighted average of variant VAFs, where each variant’s contribution is weighted by its corresponding gene prior probability, incorporating prior knowledge or assumptions about the likelihood of variant presence in specific genes. In some embodiments, mean sample VAF may be determined by computing the weighted average of variant VAFs, where each variant's contribution is weighted by its corresponding gene posterior probability, reflecting the confidence in the variant calls based on posterior distribution. In some embodiments, the observed VAF for consensus SNVs may undergo adjustment to account for the background error rate. In some embodiments, the background error rate may be derived from a position-based error model. In some embodiments, the average observed error is subtracted from the observed VAF to obtain an adjusted value. In some embodiments, the mean sample VAF is only calculated if a positive call is made for the sample. SNV Calling
[0178] In some embodiments, a variant caller algorithm calculates a confidence score using the likelihoods generated from the error model and combining them with uniform priors across a grid of ctDNA amounts. In some embodiments, the variant caller calculates a likelihood for each of the targets in a beta binomial model. The confidence score is then calculated by getting the maximum likelihood across that grid, and dividing it by that (max likelihood+negative likelihood).
[0179] Training Data: ^^,^ = ^^^,^, ^^^^^^^^^^, ^^,^, ^^,^, ^^,^, ^^,^^ where ^ ∈ {1,2, … , ^} denotes
[0180] Test Data: ^^,^ = ^^^
[0181] Result:for all bases 1,2,...,B.
[0182] For ^ = 1,2, ... , ^ do1. Estimate efficiency and error from training data for base i, using the data ^^,^. Estimate PCR parameters using training data: Replication rate (p); Error rate (pe). 2. Estimate starting copy for base i for test data at base i. X0, the total number of starting fragments at a given base. 56 4934-6677-8453.1Attorney Docket No. N.053.WO.01
[0183] In some embodiments, an improved variant caller algorithm as disclosed herein can be used. This algorithm provides an improved way of calculating confidence scores, rescaling it to significantly differentiate between positive and negative confidences, and allows the addition of non-uniform priors. This would result in more stable and more accurate confidences.
[0184] In some embodiments, in the improved variant caller algorithm, likelihoods are combined with priors, derived from training data ahead of time as described above, for smoother, more realistic and easily adjusted confidences.
[0185] Using the same error model as previously described, for each mutation allele, at eachposition p, the likelihood of data at allele a, position p, VAF % &'((^(^, ^)|%).
[0186] Smoothed allele likelihood ratio: Computed by summing over VAF likelihoods scaled by MRD VAF priors.
[0187] Position likelihood ratio: Compute the weighted position likelihood ratio ^^^(^) assuming only one of the alleles at a position is the true mutation allele.
[0188] Sample positive probability: Recursively compute the sample positive probability, as the probability of at least one target being positive, and find combined probability of sample being positive, taking into account hotspots, via position priors, and control sample outcome via sample prior.
[0189] In some embodiments, sample prior is input to the algorithm and will be adjusted to achieve desired false positive, false negative and no call rate. Allele and position priors are calculated from MRD data positive prevalence and can always be adjusted given additional information. VAF priors are calculated from VAF rates of positive MRD samples. InDel Calling
[0190] In some embodiments, an InDel-based sample caller combines position-based InDel calls with context-based InDel calls to generate InDel-based sample calls. In some embodiments, the sample calls generated by this caller are combined with the SNV calls to generate a combined sample-level call.
[0191] Position-based InDel caller
[0192] In some embodiments, position-based calling involves training a target specific error model for each InDel target that is encountered in the test sample. In one example, the error model is trained using a cohort of 200 healthy donor samples. 57 4934-6677-8453.1Attorney Docket No. N.053.WO.01
[0193] In some embodiments, the position-based caller comprises (1) target consolidation, (2) outlier detection using a beta binominal distribution, and (3) likelihood from tail probability.
[0194] In some embodiments, targets are consolidated to account for alignment differences that may result in incorrect characterization of target error rates, and outliers are detected using a beta binominal distribution.
[0195] The probability mass function (PMF) for the binomial distribution is: +^(* + ^, - − * + / )^(*) = (^) ^(^, / )^"^ * ∈ {0,1, … , -}, / )^5 ^ℎ^ / ^^^ ^6-7^^"-.
[0196] where the beta-binomial distribution is a binomial distribution, whose probability of success p follows a beta distribution B(α, β). A Beta distribution is fitted to the error model target VAFs, by computing the maximum likelihood estimates for α and β.
[0197] Log likelihood ratio for the target is then computed using the following formula: LogLikelihoodRatio = -log(test tail probability / train tail probability)
[0198] Target posterior probability can then be calculated as: PosteriorProbability = eLLR / (1 + eLLR)
[0199] Finally, in some embodiments, only replicate consensus targets that are found to have sufficient support in two plasma replicates are called. Context-based InDel caller
[0200] Insertions and deletions that do not occur in a repeat context are likely to have very low rates of background error. The relatively small number of training samples used to fit the error model is unable to characterize the background error profile for such targets, resulting in poor fit, or fit failure. For such targets a context-based calling strategy is used.
[0201] In some embodiments, the context-based caller comprises features to fit a Beta-Binomial regression model, which is used to compute target confidence.
[0202] The context-based caller is primarily used to call less noisy targets where position-based calling is not feasible. Ensemble InDel caller or combined position and context InDel caller
[0203] In some embodiments, the ensemble caller combines position- and context-based InDel calls to generate a single InDel target call. 58 4934-6677-8453.1Attorney Docket No. N.053.WO.01
[0204] In some embodiments, the ensemble caller uses the position-based call for InDel targets when the Beta-Binomial background error fit is successful and meets goodness of fit (GoF) criteria.
[0205] In some embodiments, the ensemble caller uses the context-based call for InDel targets when Beta-Binomial background error fit fails, Beta-Binomial fit does not meet GoF criteria, or the number of non-zero mutant DOR error models samples is below threshold, i.e., only use context-based calling for low noise targets. Sample InDel calling
[0206] In some embodiments, InDel sample calling uses an approach very similar to SNV sample caller, with primary difference being InDel specific WES / WGS priors, which are generated using filtered InDel targets for the same CRC sample cohort as was used for SNVs.
[0207] The sample caller algorithm recursively computes the sample likelihood ratio incorporating target likelihoods and target WES / WGS priors. Replicate Consensus Calling
[0208] A large number of PCR and sequencing artifacts are seen due to the large panel size. Such artifacts are a result of random process noise and are expected to be discordant between replicates. Accordingly, replicate consensus calling helps filter a significant fraction of such discordant targets.
[0209] In some embodiments, the methods described herein are used to filter targets that are likely a result of process noise and generate a candidate target set that will be input to the sample caller to compute the sample posterior positive probability. Because the sample level call is made by the sample caller by integrating target WES / WGS prior data, this version of the target caller is intentionally tuned to be more permissive than previous versions. The primary objective here is to filter targets lacking sufficient support or likely to be a result of random process noise. In some embodiments, the variant caller described herein uses data likelihoods, target VAFs and MRD VAF priors, to first call targets separately for each replicate, followed by a combined call using both replicates. 2 LDOR replicates were used per sample with 100K target amplicon median DOR.
[0210] Both replicates (combined) target calling: During the final target calling step information from both replicates is combined to call targets that are likely to be either CH mutations or tumor 59 4934-6677-8453.1Attorney Docket No. N.053.WO.01 derived. In some embodiments, evidence of variants or potential variants identified at this step are the input to the v2 sample caller. CH Mutation Filtering
[0211] A significant number of false positive calls have the potential to be CH (clonal hematopoiesis) mutations, including CHIP (clonal hematopoiesis of indeterminate potential) mutations. It is expected that mutations that have CH origin will be detected in matched whole blood or buffy coat data. Accordingly, in some embodiments, the methods described herein implement a reflex strategy to filter CH mutations to provide a significant improvement in specificity of the caller.
[0212] In some embodiments, the reflex strategy involves filtering targets called in plasma that are likely to be CH targets. In some embodiments, evidence of variants or potential variants identified at this step are the input to a sample caller.
[0213] All patents and publications mentioned in the specification are indicative of the levels of skill of those skilled in the art to which the invention pertains. References cited herein are incorporated by reference herein in their entirety to indicate the state of the art as of their publication or filing date and it is intended that this information can be employed herein, if needed, to exclude specific aspects that are in the prior art. For example, when composition of matter are claimed, it should be understood that compounds known and available in the art prior to Applicant's invention, including compounds for which an enabling disclosure is provided in the references cited herein, are not intended to be included in the composition of matter claims herein.
[0214] One of ordinary skill in the art will appreciate that starting materials, biological materials, reagents, synthetic methods, purification methods, analytical methods, assay methods, and biological methods other than those specifically exemplified can be employed in the practice of the invention without resort to undue experimentation. All art-known functional equivalents of any such materials and methods are intended to be included in this invention. The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention has been specifically disclosed by illustrative aspects and optional features, 60 4934-6677-8453.1Attorney Docket No. N.053.WO.01 modification and variation of the concepts herein disclosed may be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention as defined by the appended claims.
[0215] The disclosed embodiments, examples and experiments are not intended to limit the scope of the disclosure or to represent that the experiments below are all or the only experiments performed. Efforts have been made to ensure accuracy with respect to numbers used (e.g., amounts, temperature, etc.) but some experimental errors and deviations should be accounted for. It should be understood that variations in the methods as described may be made without changing the fundamental aspects that the experiments are meant to illustrate.
[0216] Those skilled in the art can devise many modifications and other embodiments within the scope and spirit of the present disclosure. Indeed, variations in the materials, methods, drawings, experiments, examples, and embodiments described may be made by skilled artisans without changing the fundamental aspects of the present disclosure. Any of the disclosed embodiments can be used in combination with any other disclosed embodiment.
[0217] In some instances, some concepts have been described with reference to specific embodiments. However, one of ordinary skill in the art appreciates that various modifications and changes can be made without departing from the scope of the invention as set forth in the claims below. Accordingly, the specification and figures are to be regarded in an illustrative rather than a restrictive sense, and all such modifications are intended to be included within the scope of invention.
[0218] The following non-limiting examples are provided purely by way of illustration of exemplary embodiments, and in no way limit the scope and spirit of the present disclosure. Furthermore, it is to be understood that any inventions disclosed or claimed herein encompass all variations, combinations, and permutations of any one or more features described herein. Any one or more features may be explicitly excluded from the claims even if the specific exclusion is not set forth explicitly herein. It should also be understood that disclosure of a reagent for use in a method is intended to be synonymous with (and provide support for) that method involving the use of that reagent, according either to the specific methods disclosed herein, or other methods known in the art unless one of ordinary skill in the art would understand otherwise. In addition, where the specification and / or claims disclose a method, any one or more of the reagents disclosed herein may be used in the method, unless one of ordinary skill in the art would understand otherwise. 61 4934-6677-8453.1Attorney Docket No. N.053.WO.01 EXAMPLES Example 1. CRC v1 panel differences due to ancestry, subtype, stage
[0219] SNV prevalence in cancer may vary by ancestral population. For example, differences in SNV landscapes between ascending and descending colon tumors are known to exist. Further, FDA approval requires demonstration of comparable performance across ancestral groups.
[0220] DB-SCAN was applied to create a v1 CRC panel based on WES data from CRC tissue samples and prior knowledge. In this example, ancestral annotations previously derived from germline variants in Circulate and US Bespoke CRC MRD trials data were used to assess sub- population performance. A reusable module was developed to assess the predicted performance of a given targeted panel and WES data. Cross-validation stratified by cancer stage, CRC subtype, high-risk mutation carriers, and ancestry was used to output expected rate of detection.
[0221] The subpopulation distribution of the samples used is shown in FIG. 4A-F, including sex, age, cancer stage and ethnicity of 26,162 samples. The microsatellite instability (MSI) status for 12,677 samples was also known (10.9% MSI) as well as germline risk variant information for ~2,000 of the samples.
[0222] Patient coverage of the v1 CRC panel was analyzed. Results showed that 96% of samples were covered with 1 mutation and 82% of samples were covered with 2 mutations (FIG. 5). The results were based on the standard balanced test set of 3544 samples.
[0223] Subpopulations were analyzed to determine differences by subgroups. As can be seen in FIGs. 6A-F, patient coverage of the v1 panel does not appear to vary significantly different due to stage, MSI status or inferred ancestry. Example 2. Algorithm development to identify and improve SNV / InDel groupings
[0224] Driver gene point mutations often cluster in small genomic regions. It is thus desirable to select variants in a manner that balances individual SNV prevalence and genomic proximity to highly prevalent tumor-related variants to maximize mutation coverage while minimizing panel size and costs. Improvements on the DB-SCAN pipeline may include (1) continuous scoring function / mutation weighting; (2) probabilistic generative model of mutations in CRC & panel cost function derived therefrom; (3) flexible contiguity (“clustering factor”) and sparsity priors to improve the balance of cost and coverage. This expresses prior knowledge about spatial correlation in tumor-related mutations. 62 4934-6677-8453.1Attorney Docket No. N.053.WO.01
[0225] The approach using DB-SCAN clustering of commonly occurring SNVs does not balance these three goals, but optimizes them sequentially. The balancing can be controlled and achieved instead by the choice of meta-parameters in Total Variation - L1 (TV-L1). Panels were trained at several levels of sparsity to assess the cost-performance trade-off.
[0226] Sparse, spatially contiguous regularization over any smooth objective can be achieved, for example, with TV-L1, as used in medical imaging. Here, the algorithm was adapted to 1D genetic array data, over the objective of identifying optimal CRC-driving SNVs. Parameters of the regularization were optimized to maximize CRC detection. To ensure optimal subpopulation performance, the algorithm was trained in a subpopulation-“balanced” manner. For this, balanced loss function was used, using subpopulation data as additional grouping variables.
[0227] Accordingly, the improved algorithmic approach naturally allows for subject weighting, useful for subgroup stratification, and for expanding a single-cancer panel to pan-cancer panels. For example, each subject’s contribution to the coverage quality score can be inversely related to the fraction of the subject’s subpopulation in the whole dataset. In this way, performance in smaller groups will be treated equally with performance in larger groups in identifying a target set of variants.
[0228] In some embodiments, the approach in DB-SCAN was followed, focusing only on mutation frequency and location, but not type. One approach to model coverage is to (a) maximize the number of variants each sample has in the panel, and (b) maximize “fairness,” i.e., ensure that all samples are covered as similarly as possible.
[0229] FIG. 7 shows an example of the mutation weighting function in KRAS G12 (toy panel constructed over 4 genes). Two moderately recurrent mutation clusters surround the main G12D peak to the left and to the right. Although the cluster to the left is more recurrent, the algorithm ranks the cluster to the right higher. This is likely because the cluster to the right improves incremental coverage more as these mutations are less concurrent with the main peak than the mutations to the left.
[0230] In some embodiments, filters can be implemented for data pre-processing, including for example a homopolymer SNV filter, a homopolymer InDel filter, and an InDel consensus filter. Example 3. CRC v2 panel differences due to ancestry, subtype, stage
[0231] Improved algorithms described above were used to create a v2 CRC panel based on WES data from CRC tissue samples and prior knowledge. Patient coverage of the v2 CRC panel was 63 4934-6677-8453.1Attorney Docket No. N.053.WO.01 analyzed as described above. As shown in FIG. 8, 97% of samples were covered with 1 mutation and 90.2% of samples were covered with 2 mutations. The results were based on the standard balanced test set of 3544 samples.
[0232] Subpopulations were then analyzed to determine differences by subgroups. As can be seen in FIGs. 9A-F, Like the v1 panel, the v2 panel patient coverage does not appear to vary significantly different due to stage or inferred ancestry. However, there is significantly higher coverage for MSI patients (defined as having an MSI Sensor score of 5 or greater), compared to MS-Stable patients: Mean MSS mutations in panel: 3.18; Mean MSI mutations in panel: 4.64.
[0233] Coverage vs. size and approximate amplicon number was then measured as a function of number of regions. A series of theoretical panels were constructed using regions initially identified by the algorithm. This was done as follows: 1. a bag of regions of top N-weighted genomic positions was derived; as N increases, the average size of the top K regions in a bag also increases; 2. several bags of regions for several values of N were derived, with each subsequent bag having larger sets of regions; 3. For each bag of regions, a greedy optimization was performed to sort the regions for maximum marginal coverage; 4. out-of-fold (cross-validation) coverage was assessed for theoretical panels of top K regions, for each region bag. FIG. 10 shows correlation between coverage and approximate amplicon number for the v2 panel. Example 4. Performance results for replicate consensus calling
[0234] The primary objective of this example was to filter targets lacking sufficient support or likely to be a result of random process noise, 2 LDOR replicates were used per sample with 100K target amplicon median DOR.
[0235] During the final target calling step, information from both replicates were combined to call targets that were likely to be either CH mutations or tumor derived (Table 1). Targets called at this step are then input to the v2 sample caller. Table 1. Calling results of MRD positive samples Sig_VAF_bins Signatera Positive Percentage Agreement (PPA) (0.0, 0.02] 0.277 64 4934-6677-8453.1Attorney Docket No. N.053.WO.01 (0.02, 0.04] 0.500 (0.04, 0.1] 0.857 (0.1, 0.2] 1.000 (0.2, 100.0] 0.933 Example 5. Performance evaluation for a reflex strategy
[0236] As mentioned above, a significant number of false positive calls have the potential to be CH mutations. It was expected that mutations that have CH origin will be detected in matched whole blood or buffy coat data. Accordingly, this example demonstrates results from the reflex strategy described herein to filter CH mutations.
[0237] As shown in Table 2 below, 106 MRD samples previously tested using Signatera™ were analyzed using the v2 CRC panel, of which 47 samples were initially called as positive and subject to reflex with buffy coat data. The 47 samples subject to buffy reflex included 10 Signatera negative samples and 37 Signatera positive samples. Table 2. Buffy reflex samples Sample type Samples Buffy reflex rate Weighted reflex rate Overall 106 47 0.432 Signatera Negatives 59 10 0.12 Signatera Positives 47 37 0.228
[0238] As seen in Table 3 below, 9 / 10 false positive samples (Signatera negatives) were filtered using the reflex strategy at 98.3% target specificity. The data demonstrate that a reflex-based strategy can be successfully used to filter putative CH mutations and significantly reduce the false positive rate. Table 3. Buffy reflex results Sample Buffy CH Samples Sample Specificity PPA Sample Specificity PPA type target targets calls w / o w / o CH w / o calls with CH with calls CH filter filter CH with CH filter CH filter filer filter 65 4934-6677-8453.1Attorney Docket No. N.053.WO.01 Signatera 292 24 47 37 - 79% 34 - 72.3% Positives Signatera 408 38 59 10 83% - 1 98.3 % - Negatives * * * *66 4934-6677-8453.1
Claims
Attorney Docket No. N.053.WO.01 What is claimed is:
1. A method for preparing a plurality of non-naturally occurring compositions of amplified DNA from a blood sample of a subject, comprising: (a) extracting cell-free DNA from a plasma fraction of a blood sample of a subject; (b) performing targeted enrichment on the extracted cell-free DNA or DNA derived therefrom to enrich a panel of a plurality of target loci each encompassing at least one somatic mutation associated with the cancer to obtain a non-naturally occurring composition of plasma- derived amplified DNA; (c) sequencing the non-naturally occurring composition of plasma-derived amplified DNA to identify at least one somatic variant that is present in the plasma-derived amplified DNA; and optionally, if at least one somatic variant present in the plasma-derived amplified DNA is identified in step (c), the method further comprises: (d) extracting cellular DNA from a fraction of the blood sample containing a plurality of leukocytes; (e) performing targeted enrichment on the extracted cellular DNA or DNA derived therefrom to enrich a subset of the panel of the plurality of target loci to obtain a non-naturally occurring composition of leukocyte-derived amplified DNA, wherein the subset comprises the at least one somatic variant identified in the plasma-derived amplified DNA; and (f) sequencing the non-naturally occurring composition of leukocyte-derived enriched DNA to determine whether the at least one somatic variant identified in the plasma-derived enriched DNA is also present in the leukocyte-derived enriched DNA, thereby identifying at least one somatic variant that is present in the plasma-derived enriched DNA and that is not present in the leukocyte-derived enriched DNA.
2. The method of claim 1, wherein the at least one somatic variant that is present in the plasma-derived enriched DNA and that is not present in the leukocyte-derived enriched DNA is a somatic mutation associated with the cancer.
3. The method of claim 1, wherein performing targeted enrichment comprises targeted multiplex amplification or targeted probe capture on the extracted cell-free DNA or DNA derived therefrom to enrich the panel of the plurality of target loci. 67 4934-6677-8453.1Attorney Docket No. N.053.WO.01 4. The method of any of claims 1-3, wherein the plurality of target loci comprises 25-5,000 target loci.
5. The method of any of claims 1-4, further comprising identifying at least one somatic variant present in the plasma-derived enriched DNA that is also present in the leukocyte-derived enriched DNA, thereby identifying a clonal hematopoiesis (CH) mutation.
6. The method of claim 5, wherein the CH mutation is a clonal hematopoiesis of indeterminate potential (CHIP) mutation.
7. The method of any of claims 1-6, wherein the blood sample is split into a plurality of sub- samples prior to extracting cell-free DNA from the plasma fraction and extracting cellular DNA from the fraction of the sample containing the plurality of leukocytes.
8. The method of claim 7, wherein steps (b) and (c) are performed on at least two of the sub- samples, wherein sequencing reads from the at least two sub-samples are combined to identify at least one somatic variant that is present in the plasma-derived enriched DNA of each of the at least two sub-samples.
9. The method of claim 1, wherein the plasma fraction of the blood sample is split into a plurality of plasma sub-samples prior to extracting cell-free DNA from the plasma fraction.
10. The method of claim 9, wherein steps (b) and (c) are performed on at least two of the plasma sub-samples, wherein sequencing reads from the at least two plasma sub-samples are combined to identify at least one somatic variant that is present in the plasma-derived enriched DNA of each of the at least two plasma sub-samples.
11. The method of claim 1, wherein the cell-free DNA extracted from the plasma fraction is split into a plurality of cell-free DNA sub-samples and steps (b) and (c) are performed on at least two of the cell-free DNA sub-samples, wherein sequencing reads from the at least two cell-free DNA sub-samples are combined to identify at least one somatic variant that is present in the plasma-derived enriched DNA of each of the at least two cell-free DNA sub-samples.
12. A method for preparing a plurality of non-naturally occurring compositions of amplified DNA from a blood sample of a subject, comprising: 68 4934-6677-8453.1Attorney Docket No. N.053.WO.01 (a) extracting cell-free DNA from a plasma fraction of a blood sample of a subject; (b) adding adaptors to a first portion and a second portion of the extracted cell-free DNA and generating a first portion and a second portion of adapted DNA, and wherein the adaptors are added to the first portion and the second portion of the extracted cell-free DNA under the same conditions; (c) performing a first targeted enrichment on the first portion of the adapted DNA or DNA derived therefrom to enrich a panel of a plurality of target loci each encompassing at least one somatic variant associated with a cancer to obtain a first non-naturally occurring composition of plasma-derived amplified DNA; (d) performing a second targeted enrichment on a second portion of the adapted DNA or DNA derived therefrom to enrich the panel of the plurality of target loci to obtain a second non- naturally occurring composition of plasma-derived amplified DNA, wherein the second targeted enrichment is performed under the same conditions as the first targeted enrichment; (e) sequencing DNA derived from the first non-naturally occurring composition to obtain a first set of reads, and sequencing DNA derived from the second non-naturally occurring composition to obtain a second set of reads, wherein the first and second sequencing are performed under the same conditions; and (f) combining and analyzing the first and second sets of reads to identify the presence or absence of the at least one somatic variant associated with cancer in the extracted cell-free DNA.
13. The method of claim 12, wherein the first and second portions of the extracted cell-free DNA have an about equal amount of DNA.
14. The method of claim 12 or 13, wherein at least one somatic variant associated with cancer is identified in step (f), and the method further comprises: (g) extracting cellular DNA from a fraction of the blood sample containing a plurality of leukocytes; (h) performing targeted enrichment on the extracted cellular DNA or DNA derived therefrom to enrich a subset of the panel of the plurality of target loci to obtain a non-naturally occurring composition of leukocyte-derived amplified DNA, wherein the subset comprises the at least one somatic variant identified in step (f); and (i) sequencing the non-naturally occurring composition of leukocyte-derived enriched DNA to determine whether the at least one somatic variant identified in step (f) is also present in 69 4934-6677-8453.1Attorney Docket No. N.053.WO.01 the leukocyte-derived enriched DNA, thereby identifying at least one somatic variant that is present in the plasma-derived enriched DNA and that is not present in the leukocyte-derived enriched DNA.
15. The method of any of claims 12-14, wherein performing targeted enrichment comprises targeted multiplex amplification or targeted probe capture on the extracted cell-free DNA or DNA derived therefrom to enrich the panel of the plurality of target loci.
16. The method of any of claims 12-15, wherein the plurality of target loci comprises 25-5,000 target loci.
17. The method of any of claims 8 or 10-16, wherein at least one processing or sequencing error, if present, is identified from the combined reads without using any molecular barcode.
18. The method of claim 17, wherein the method does not comprise tagging the extracted cell- free DNA with a molecular barcode.
19. The method of any of claims 1-18, wherein prior to performing targeted enrichment, the extracted cell-free DNA is end repaired, A-tailed, and ligated to at least one adaptor comprising a universal priming sequence to obtain adaptor-ligated DNA.
20. The method of claim 19, wherein prior to performing targeted enrichment, the adaptor- ligated DNA is amplified using the universal priming sequence.
21. The method of any of claims 1-18, wherein prior to performing targeted enrichment, the extracted cellular DNA is not ligated to an adaptor and is not amplified.
22. The method of any of claims 1-21, wherein the somatic variant associated with the cancer comprises a single nucleotide variant (SNV), a multi-nucleotide variant (MNV), an indel, a gene fusion, a structural variant, or a combination thereof.
23. The method of any of claims 1-22, wherein the somatic variant associated with the cancer comprises a single nucleotide variant (SNV) or an InDel, or a combination thereof.
24. The method of any of claims 1-23, further comprising identifying one or more germline mutations present in the leukocyte-derived amplified DNA. 70 4934-6677-8453.1Attorney Docket No. N.053.WO.01 25. The method of any of claims 1-24, wherein the cancer is a solid tumor.
26. The method of claim 25, wherein the solid tumor is a cancer or tumor of abdomen or abdominal wall, adrenal gland, anus, appendix, bladder, bone, brain, breast, cervix, chest wall, colon, diaphragm, duodenum, ear, endometrium, esophagus, fallopian tube, gallbladder, gastroesophageal junction, head and neck, kidney, larynx, liver, lung, lymph node, malignant effusions, mediastinum, nasal cavity, omentum, ovarian, pancreas, pancreatobiliary, parotid gland, pelvis, penis, pericardium, peritoneum, pleura, prostate, rectum, salivary gland, skin, small intestine, soft tissue, spleen, stomach, thyroid, tongue, trachea, ureter, uterus, vagina, vulva, or whipple resection.
27. The method of claim 25 or 22, wherein the solid tumor is breast cancer, advanced adenoma, colorectal cancer, gastrointestinal cancer, kidney cancer, lung cancer, multiple myeloma, ovarian cancer, or pancreatic cancer.
28. The method of any of claims 25-27, wherein the solid tumor is colorectal cancer.
29. The method of any of claims 1-28, further comprising longitudinally collecting a plurality of blood samples from the subject and repeating each of the steps for each of the plurality of blood samples.
30. The method of any of claims 1-29, wherein the subject has received treatment for the cancer.
31. The method of claim 29, wherein the plurality of blood samples are collected after the subject has been treated with surgery, first-line chemotherapy, and / or adjuvant therapy.
32. The method of claim 30 or 31, wherein the identification of the at least one somatic variant is indicative of minimal residual disease.
33. The method of any of claims 1-32, wherein the subject has not received treatment for the cancer.
34. The method of claim 33, wherein the identification of the at least one somatic mutation is indicative of the presence of the cancer in the subject. 71 4934-6677-8453.1Attorney Docket No. N.053.WO.01 35. The method of any of claims 1-34, wherein the target loci comprise between 100 and 500 SNV and / or InDel loci.
36. The method of any of claims 1-35, wherein the target loci are selected based on population-wide SNV and / or InDel patterns.
37. The method of any of claims 1-36, wherein the target loci are cross-validated across sub- populations including by cancer stage, CRC subtype, high-risk mutation carriers, and ancestry.
38. The method of any of claims 1-37, wherein the target loci comprise unstable microsatellite loci.
39. The method of any of claims 1-38, wherein the target loci comprise one or more COSMIC driver or tumor suppressor gene loci.
40. The method of any of claims 1-39, wherein the method further comprises targeted enrichment and sequencing of a plurality of differentially methylated regions that are differentially methylated in the cancer, from cell-free DNA extracted from a different blood sample of the subject or a different fraction of the same blood sample, or DNA derived therefrom.
41. The method of any of claims 1-40, wherein the method does not comprise whole genome sequencing or whole exosome sequencing of a tumor tissue sample of the subject.
42. The method of any of claims 1-41, wherein the targeted enrichment comprises amplifying at least 50 of the 25-5,000 target loci in a single reaction volume.
43. The method of any of claims 1-42, wherein the targeted enrichment comprises amplifying 50-200 of the 25-5,000 target loci in a single reaction volume.
44. The method of any of claims 1-43, wherein the target loci cover 1-5000 kilobases (kb) of genomic sequences.
45. The method of any of claims 1-44, wherein the sequencing has a depth of read (DOR) of between 50,000 to 200,000 per target locus. 72 4934-6677-8453.1Attorney Docket No. N.053.WO.01 46. The method of any of claims 1-45, wherein one or more universal amplifications are performed after the targeted enrichment to obtain the non-naturally occurring composition(s) of plasma-derived amplified DNA.
47. The method of any of claims 1-46, wherein one or more universal amplifications are performed after the targeted enrichment to obtain the non-naturally occurring composition of leukocyte-derived amplified DNA.
48. The method of any of claims 1-47, wherein at least one target loci is enriched using two or more target-specific primers having overlapping sequences.
49. The method of any of claims 1-48, wherein the analyzing further comprises using an error model based on variant allele frequency (VAF) priors.
50. The method of any of claims 1-49, wherein the analyzing further comprises using position- specific priors.
51. The method of any of claims 1-50, further comprising classifying the sample as containing cancer-derived DNA (positive) or not containing cancer-derived DNA (negative).
52. The method of claim 51, wherein the classifying further comprises using gene-specific priors.
53. The method of claim 51 or 52, wherein the classifying further comprises using position- specific, gene-specific, and / or cancer-type-specific priors.
54. The method of any of claims 51-53, wherein the classifying further comprises adjusting a confidence level based on concordance between two plasma replicates.
55. The method of any of claims 49-54, further comprising generating a confidence score for the identified mutation based on a smoothed allele likelihood ratio, wherein the smoothed allele likelihood ratio is determined by summing over VAF likelihoods scaled by VAF likelihood priors.
56. The method of any of claims 49-54, further comprising generating a confidence score for the identified variant based on a position likelihood ratio, wherein the position likelihood ratio is 73 4934-6677-8453.1Attorney Docket No. N.053.WO.01 determined by computing a weighted position likelihood ratio ^^^(^) assuming only one of the alleles at a position is the true mutation allele.
57. The method of any of claims 49-54, further comprising generating a confidence score for the identified variant based on a sample likelihood ratio, wherein the sample likelihood ratio is a likelihood for at least one target being positive.
58. The method of claim 57, wherein a sample posterior probability of position being positive is determined based on the sample likelihood ratio, wherein the sample posterior probability is used to classify the sample.
59. The method of any of claims 1-58, wherein identification of an InDel is based on a position-based caller, wherein the position-based caller comprises training a target specific error model for each InDel target that is encountered in the sample.
60. The method of claim 59, wherein the position-based caller comprises target consolidation, outlier detection using a beta binominal distribution, and a likelihood from tail probability.
61. The method of claim 60, wherein the target consolidation comprises consolidating targets having the same start coordinate or share the same InDel repeat unit.
62. The method of any of claims 1-58, wherein the InDel identification is based on a context- based caller.
63. The method of claim 62, wherein the context-based caller comprises features to fit a Beta- Binomial regression model.
64. The method of claim 62 or 63, wherein a positive identification by the context-based caller requires calling the InDel in at least two replicates of the sample.
65. The method of and of claims 63-64, wherein identification of the InDel comprises an ensemble caller combining both the position-based caller and the context based caller.
66. The method of any of claims 1-65, further comprising generating a confidence score for the sample being positive for a mutation using a sample caller. 74 4934-6677-8453.1Attorney Docket No. N.053.WO.01 67. The method of claim 66, wherein the sample caller is an SNV sample caller, an InDel sample caller, or a combination thereof.
68. The method of claim 66, wherein the sample caller computes a sample posterior probability. 75 4934-6677-8453.1
Citation Information
Patent Citations
Methods for simultaneous amplification of target loci
US11312996B2
Compositions, methods, and kits for isolating nucleic acids
WO2018156418A1
Methods for cancer detection and monitoring by means of personalized detection of circulating tumor DNA
WO2019200228A1
Linked target capture and ligation
WO2020039261A1
Methods and compositions for analyses of cancer
US20230002831A1