Methods for cancer detection and monitoring
Patent Information
- Application Number
- JP2024539961
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-01-04
- Filing Date
- 2023-01-04
- Publication Date
- 2025-12-10
AI Technical Summary
Current methods for detecting early recurrence or metastasis of cancer rely heavily on invasive biopsies and image diagnosis, which are not sufficient and carry risks, while non-invasive methods for analyzing blood cells or bone marrow mutations are lacking.
A method for preparing amplified DNA from biological samples using isolated DNA from tumor biopsies or blood/bone marrow samples to identify patient-specific mutations, including clone hematopoietic (CHIP) mutations, which serve as indicators of cancer recurrence or metastasis through targeted gene amplification and analysis.
Enables early detection of cancer recurrence or metastasis with high sensitivity and specificity, reducing the need for invasive procedures and providing reliable indicators for clinical management.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] Detecting early cancer recurrence or metastasis has traditionally relied on diagnostic imaging and tissue biopsy. While tumor tissue biopsy is invasive and carries risks that potentially contribute to metastasis or surgical complications, diagnostic imaging-based detection is not sensitive enough to detect early recurrence or metastasis. Better, less invasive methods for detecting cancer recurrence or metastasis are needed, particularly methods that incorporate analysis of somatic mutations in blood cells or bone marrow known as clonal hematopoiesis with undetermined potential (CHIP). Summary of the Invention
[0002] In one aspect, the disclosure provides a method for preparing a preparation of amplified DNA from a biological sample of a patient diagnosed with cancer, useful for determining cancer recurrence or metastasis, comprising: (a) sequencing DNA isolated from hematopoietic cells in the patient's blood or bone marrow sample, or a fraction thereof, to determine the presence or absence of one or more clonal hematopoietic with undetermined potential (CHIP) mutations; (b) sequencing DNA isolated from (i) a tumor biopsy sample of the patient or (ii) cell-free DNA isolated from the blood or bone marrow sample, or a fraction thereof, to identify multiple patient-specific somatic mutations associated with the cancer; and (c) sequencing the cell-free DNA isolated from the patient's longitudinally collected biological sample, or a fraction thereof, to identify multiple patient-specific somatic mutations associated with the cancer. (d) analyzing the amplified DNA preparation by sequencing the amplified DNA to determine the presence or absence of the patient-specific somatic mutations, wherein the presence of two or more patient-specific somatic mutations associated with cancer and the presence of one or more CHIP mutations is indicative of cancer recurrence or metastasis.
[0003] In some embodiments, step (a) comprises performing whole-exome or whole-genome sequencing on DNA isolated from the buffy coat fraction of the blood or bone marrow sample to determine the presence or absence of one or more CHIP mutations.
[0004] In some embodiments, step (a) comprises enriching a panel of genomic loci associated with a bone marrow disorder from DNA isolated from a buffy coat fraction of a blood or bone marrow sample to obtain enriched genomic loci, and subsequently sequencing the enriched genomic loci to determine the presence or absence of one or more CHIP mutations.
[0005] In some embodiments, step (b) comprises performing whole-exome or whole-genome sequencing on cell-free DNA isolated from the plasma fraction of the blood or bone marrow sample to identify multiple patient-specific somatic mutations associated with the cancer.
[0006] In some embodiments, step (b) comprises performing whole-exome or whole-genome sequencing on DNA isolated from a tumor biopsy sample from the patient to identify a plurality of patient-specific somatic mutations associated with the cancer.
[0007] In some embodiments, step (b) comprises enriching a panel of genomic loci associated with cancer from cell-free DNA isolated from the plasma fraction of the blood or bone marrow sample to obtain enriched genomic loci, and subsequently sequencing the enriched genomic loci to identify a plurality of patient-specific somatic mutations associated with cancer.
[0008] In some embodiments, step (b) comprises enriching a panel of genomic loci associated with cancer from DNA isolated from a tumor biopsy sample of the patient to obtain enriched genomic loci, and subsequently sequencing the enriched genomic loci to identify a plurality of patient-specific somatic mutations associated with the cancer.
[0009] In some embodiments, the panel of genomic loci associated with myeloid disorders is enriched by hybrid capture and / or targeted amplification. In some embodiments, the panel of genomic loci associated with myeloid disorders is enriched by hybrid capture and / or multiplexed targeted amplification. In some embodiments, the panel of genomic loci associated with myeloid disorders is enriched by multiplexed targeted PCR.
[0010] In some embodiments, the panel of genomic loci associated with cancer is enriched by hybrid capture and / or targeted amplification. In some embodiments, the panel of genomic loci associated with cancer is enriched by multiplexed targeted amplification. In some embodiments, the panel of genomic loci associated with cancer is enriched by multiplexed targeted PCR.
[0011] In some embodiments, the panel of genomic loci associated with myeloid disorders and / or the panel of genomic loci associated with cancer comprise one or more genomic loci in exons, introns, gene regulatory regions, non-coding RNAs, rearranged genes, or combinations thereof.
[0012] In some embodiments, the patient-specific somatic mutations associated with cancer comprise single nucleotide variants (SNVs), multi-nucleotide variants (MNVs), indels, gene fusions, structural variants, or combinations thereof.
[0013] In some embodiments, step (c) comprises the targeted multiplex amplification of at least 8 target loci, each spanning at least one patient-specific cancer mutation associated with cancer, within one reaction volume. In some embodiments, step (c) comprises the targeted multiplex amplification of at least 16 target loci, each spanning at least one patient-specific cancer mutation associated with cancer, within one reaction volume. In some embodiments, step (c) comprises the targeted multiplex amplification of at least 32 target loci, each spanning at least one patient-specific cancer mutation associated with cancer, within one reaction volume. In some embodiments, step (c) comprises the targeted multiplex amplification of at least 64 target loci, each spanning at least one patient-specific cancer mutation associated with cancer, within one reaction volume. In some embodiments, step (c) comprises the targeted multiplex amplification of at least 128 target loci, each spanning at least one patient-specific cancer mutation associated with cancer, within one reaction volume.
[0014] In some embodiments, the method further comprises identifying one or more germline mutations in the patient, wherein the target locus amplified in step (c) does not span the one or more germline mutations. In some embodiments, the one or more germline mutations are identified by sequencing DNA isolated from hematopoietic cells in a blood or bone marrow sample or a fraction thereof.
[0015] In some embodiments, the cancer is a cancer or tumor of the abdomen or abdominal wall, adrenal gland, anus, appendix, bladder, bone, brain, breast, cervix, chest wall, colon, diaphragm, duodenum, ear, endometrium, esophagus, fallopian tube, gallbladder, gastroesophageal junction, head and neck, kidney, larynx, liver, lung, lymph node, malignant effusion, mediastinum, nasal cavity, omentum, ovary, pancreas, pancreaticobiliary duct, parotid gland, pelvis, penis, pericardium, peritoneum, pleura, prostate, rectum, salivary gland, skin, small intestine, soft tissue, spleen, stomach, thyroid, tongue, trachea, ureter, uterus, vagina, vulva, or Whipple resection.
[0016] In some embodiments, the cancer is breast cancer, colon cancer, gastrointestinal cancer, renal cancer, lung cancer, multiple myeloma, ovarian cancer, or pancreatic cancer.
[0017] In some embodiments, the method further comprises collecting multiple biological samples longitudinally from the patient and repeating steps (c) and (d) for each of the biological samples.
[0018] In some embodiments, one or more biological samples are collected after the patient has been treated with surgery, initial chemotherapy, and / or adjuvant therapy. In some embodiments, the patient has been treated with surgery prior to collection of the liquid biopsy sample. In some embodiments, the patient has been treated with chemotherapy prior to collection of the liquid biopsy sample. In some embodiments, the patient has been treated with adjuvant or neoadjuvant therapy prior to collection of the liquid biopsy sample. In some embodiments, the patient has been treated with radiation therapy prior to collection of the liquid biopsy sample. In some embodiments, the liquid biopsy sample is collected from the patient about 2-12 weeks after surgery, initial chemotherapy, adjuvant therapy, and / or neoadjuvant therapy. In some embodiments, the liquid biopsy sample is collected from the patient about 4-8 weeks after surgery, initial chemotherapy, adjuvant therapy, and / or neoadjuvant therapy. In some embodiments, the liquid biopsy sample is collected from the patient about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 weeks after surgery. In some embodiments, the liquid biopsy sample is collected from the patient about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 weeks after initial chemotherapy. In some embodiments, the liquid biopsy sample is collected from the patient about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 weeks after adjuvant or neoadjuvant therapy. In some embodiments, the liquid biopsy sample is collected from the patient about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 weeks after adjuvant chemotherapy (ACT).
[0019] In some embodiments, the presence of two or more patient-specific somatic mutations associated with cancer and the presence of two or more CHIP mutations is indicative of cancer recurrence or metastasis.
[0020] In another aspect, the disclosure provides a method for preparing a preparation of amplified DNA from a biological sample of a patient diagnosed with cancer, useful for determining cancer recurrence or metastasis, comprising: (a) sequencing (i) DNA isolated from a tumor biopsy sample of the patient, or (ii) cell-free DNA isolated from a blood or bone marrow sample, or a fraction thereof, of the patient, to identify a plurality of patient-specific somatic mutations associated with the cancer; and (b) performing targeted multiplex amplification on cell-free DNA isolated from a longitudinally collected biological sample of the patient, or a fraction thereof, to amplify a plurality of target loci to obtain amplified DNA, thereby preparing the preparation of amplified DNA; (c) analyzing the amplified DNA preparation by sequencing the amplified DNA to determine the presence or absence of the patient-specific somatic mutations; and (d) sequencing DNA isolated from hematopoietic cells in the patient's biological sample or a fraction thereof to determine the presence or absence of one or more CHIP mutations, wherein the presence of two or more patient-specific somatic mutations associated with cancer and the presence of one or more CHIP mutations is indicative of cancer recurrence or metastasis.
[0021] In some embodiments, step (a) comprises performing whole-exome or whole-genome sequencing on cell-free DNA isolated from the plasma fraction of the blood or bone marrow sample to identify multiple patient-specific somatic mutations associated with the cancer.
[0022] In some embodiments, step (a) comprises performing whole-exome or whole-genome sequencing on DNA isolated from a tumor biopsy sample from the patient to identify a plurality of patient-specific somatic mutations associated with the cancer.
[0023] In some embodiments, step (a) comprises enriching a panel of genomic loci associated with cancer from cell-free DNA isolated from the plasma fraction of a blood or bone marrow sample to obtain enriched genomic loci, and subsequently sequencing the enriched genomic loci to identify a plurality of patient-specific somatic mutations associated with cancer.
[0024] In some embodiments, step (a) comprises enriching a panel of genomic loci associated with cancer from DNA isolated from a tumor biopsy sample of the patient to obtain enriched genomic loci, and subsequently sequencing the enriched genomic loci to identify a plurality of patient-specific somatic mutations associated with the cancer.
[0025] In some embodiments, step (d) comprises performing whole-exome or whole-genome sequencing on DNA isolated from the buffy coat fraction of the blood or bone marrow sample to determine the presence or absence of one or more CHIP mutations.
[0026] In some embodiments, step (d) comprises enriching a panel of genomic loci associated with a bone marrow disorder from DNA isolated from the buffy coat fraction of a blood or bone marrow sample to obtain enriched genomic loci, and subsequently sequencing the enriched genomic loci to determine the presence or absence of one or more CHIP mutations.
[0027] In some embodiments, the panel of genomic loci associated with myeloid disorders is enriched by hybrid capture and / or targeted amplification. In some embodiments, the panel of genomic loci associated with myeloid disorders is enriched by hybrid capture and / or multiplexed targeted amplification. In some embodiments, the panel of genomic loci associated with myeloid disorders is enriched by multiplexed targeted PCR.
[0028] In some embodiments, the panel of genomic loci associated with cancer is enriched by hybrid capture and / or targeted amplification. In some embodiments, the panel of genomic loci associated with cancer is enriched by multiplexed targeted amplification. In some embodiments, the panel of genomic loci associated with cancer is enriched by multiplexed targeted PCR.
[0029] In some embodiments, the panel of genomic loci associated with myeloid disorders and / or the panel of genomic loci associated with cancer comprise one or more genomic loci in exons, introns, gene regulatory regions, non-coding RNAs, rearranged genes, or combinations thereof.
[0030] In some embodiments, the patient-specific somatic mutations associated with cancer comprise single nucleotide variants (SNVs), multi-nucleotide variants (MNVs), indels, gene fusions, structural variants, or combinations thereof.
[0031] In some embodiments, step (b) comprises the targeted multiplex amplification of at least 8 target loci, each spanning at least one patient-specific cancer mutation associated with cancer, within one reaction volume. In some embodiments, step (b) comprises the targeted multiplex amplification of at least 16 target loci, each spanning at least one patient-specific cancer mutation associated with cancer, within one reaction volume. In some embodiments, step (b) comprises the targeted multiplex amplification of at least 32 target loci, each spanning at least one patient-specific cancer mutation associated with cancer, within one reaction volume. In some embodiments, step (b) comprises the targeted multiplex amplification of at least 64 target loci, each spanning at least one patient-specific cancer mutation associated with cancer, within one reaction volume. In some embodiments, step (b) comprises the targeted multiplex amplification of at least 128 target loci, each spanning at least one patient-specific cancer mutation associated with cancer, within one reaction volume.
[0032] In some embodiments, the method further comprises identifying one or more germline mutations in the patient, wherein the target locus amplified in step (b) does not span the one or more germline mutations. In some embodiments, the one or more germline mutations are identified by sequencing DNA isolated from hematopoietic cells in a blood or bone marrow sample or a fraction thereof.
[0033] In some embodiments, the cancer is a cancer or tumor of the abdomen or abdominal wall, adrenal gland, anus, appendix, bladder, bone, brain, breast, cervix, chest wall, colon, diaphragm, duodenum, ear, endometrium, esophagus, fallopian tube, gallbladder, gastroesophageal junction, head and neck, kidney, larynx, liver, lung, lymph node, malignant effusion, mediastinum, nasal cavity, omentum, ovary, pancreas, pancreaticobiliary duct, parotid gland, pelvis, penis, pericardium, peritoneum, pleura, prostate, rectum, salivary gland, skin, small intestine, soft tissue, spleen, stomach, thyroid, tongue, trachea, ureter, uterus, vagina, vulva, or Whipple resection.
[0034] In some embodiments, the cancer is breast cancer, colon cancer, gastrointestinal cancer, renal cancer, lung cancer, multiple myeloma, ovarian cancer, or pancreatic cancer.
[0035] In some embodiments, the method further comprises collecting multiple biological samples longitudinally from the patient and repeating steps (b) and (c) for each of the biological samples.
[0036] In some embodiments, one or more biological samples are collected after the patient has been treated with surgery, initial chemotherapy, and / or adjuvant therapy. In some embodiments, the patient has been treated with surgery prior to collection of the liquid biopsy sample. In some embodiments, the patient has been treated with chemotherapy prior to collection of the liquid biopsy sample. In some embodiments, the patient has been treated with adjuvant or neoadjuvant therapy prior to collection of the liquid biopsy sample. In some embodiments, the patient has been treated with radiation therapy prior to collection of the liquid biopsy sample. In some embodiments, the liquid biopsy sample is collected from the patient about 2-12 weeks after surgery, initial chemotherapy, adjuvant therapy, and / or neoadjuvant therapy. In some embodiments, the liquid biopsy sample is collected from the patient about 4-8 weeks after surgery, initial chemotherapy, adjuvant therapy, and / or neoadjuvant therapy. In some embodiments, the liquid biopsy sample is collected from the patient about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 weeks after surgery. In some embodiments, the liquid biopsy sample is collected from the patient about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 weeks after initial chemotherapy. In some embodiments, the liquid biopsy sample is collected from the patient about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 weeks after adjuvant or neoadjuvant therapy. In some embodiments, the liquid biopsy sample is collected from the patient about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 weeks after adjuvant chemotherapy (ACT).
[0037] In some embodiments, the presence of two or more patient-specific somatic mutations associated with cancer and the presence of two or more CHIP mutations is indicative of cancer recurrence or metastasis.
[0038] In a further aspect, the disclosure relates to a method for sequencing DNA from a biological sample of a patient diagnosed with cancer, comprising performing whole-exome or whole-genome sequencing on DNA isolated from hematopoietic cells in the patient's blood or bone marrow sample or a fraction thereof to determine the presence or absence of one or more CHIP mutations, and identifying the patient as having an increased risk of disease progression due to the presence of the one or more CHIP mutations.
[0039] The presently disclosed embodiments will be further described with reference to the accompanying drawings, in which like structure is referred to by like numerals throughout the several views. The drawings shown are not necessarily to scale, emphasis instead generally being placed upon illustrating the principles of the presently disclosed embodiments. [Brief explanation of the drawings]
[0040] [Figure 1] Characteristics of the cohort and identified CHIP mutations (A-D). Analysis revealed CHIP mutations present in 16% (392 / 2484) of patients. The majority of patients with CHIP (82%, 320) had a single mutation, and 18% (72) of patients had two to four mutations detected. The most commonly affected genes in patients with CHIP within this cohort were DNMT3A (46%), TET2 (16%), TP53 (13%), NOTCH1 and EZH2 (6% each), and CDKN2A and ASXL1 (5% each). [Figure 2] Association of CHIP incidence with age and cancer type (A-B). The incidence of CHIP increased exponentially from 7% in patients under 40 years of age to 23% in patients over 60 years of age. Patients with renal cell carcinoma (32%), multiple myeloma (27%), lung cancer (23%), and pancreatic cancer (20%) had a higher prevalence of CHIP compared with patients with breast cancer (15%) and colorectal cancer (14%). [Figure 3]Disease progression and CHIP status. (A) Kaplan-Meier curves showing the proportion of patients with progression-free survival over time, stratified by CHIP status. (B) Time to disease progression for each patient by CHIP status. CHIP-positive patients showed significantly shorter time to progression (p=0.02*). DETAILED DESCRIPTION OF THE INVENTION
[0041] I. Overview The methods and compositions provided herein improve the detection, diagnosis, staging, screening, treatment, and management of cancer. In one aspect, the disclosure provides a method for preparing a preparation of amplified DNA from a biological sample of a patient diagnosed with cancer, useful for determining cancer recurrence or metastasis, comprising: (a) sequencing DNA isolated from hematopoietic cells in the patient's blood or bone marrow sample, or a fraction thereof, to determine the presence or absence of one or more clonal hematopoietic with undetermined potential (CHIP) mutations; (b) sequencing DNA isolated from (i) a tumor biopsy sample of the patient, or (ii) cell-free DNA isolated from the blood or bone marrow sample, or a fraction thereof, to identify multiple patient-specific somatic mutations associated with the cancer; and (c) sequencing the cell-free DNA isolated from the patient's longitudinally collected biological sample, or a fraction thereof, to identify multiple patient-specific somatic mutations associated with the cancer. (d) analyzing the amplified DNA preparation by sequencing the amplified DNA to determine the presence or absence of the patient-specific somatic mutations, wherein the presence of two or more patient-specific somatic mutations associated with cancer and the presence of one or more CHIP mutations is indicative of cancer recurrence or metastasis.
[0042] In another aspect, the disclosure provides a method for preparing a preparation of amplified DNA from a biological sample of a patient diagnosed with cancer, useful for determining cancer recurrence or metastasis, comprising: (a) sequencing (i) DNA isolated from a tumor biopsy sample of the patient, or (ii) cell-free DNA isolated from a blood or bone marrow sample, or a fraction thereof, of the patient, to identify a plurality of patient-specific somatic mutations associated with the cancer; and (b) performing targeted multiplex amplification on cell-free DNA isolated from a longitudinally collected biological sample of the patient, or a fraction thereof, to amplify a plurality of target loci to obtain amplified DNA, thereby preparing the preparation of amplified DNA; (c) analyzing the amplified DNA preparation by sequencing the amplified DNA to determine the presence or absence of the patient-specific somatic mutations; and (d) sequencing DNA isolated from hematopoietic cells in the patient's biological sample or a fraction thereof to determine the presence or absence of one or more CHIP mutations, wherein the presence of two or more patient-specific somatic mutations associated with cancer and the presence of one or more CHIP mutations is indicative of cancer recurrence or metastasis.
[0043] In a further aspect, the disclosure relates to a method for sequencing DNA from a biological sample of a patient diagnosed with cancer, comprising performing whole-exome or whole-genome sequencing on DNA isolated from hematopoietic cells in the patient's blood or bone marrow sample or a fraction thereof to determine the presence or absence of one or more CHIP mutations, and identifying the patient as having an increased risk of disease progression due to the presence of the one or more CHIP mutations.
[0044] In some embodiments, the multiplex amplification reactions each target 1-500 target loci, 1-20 target loci, or 20-50 target loci, or 50-100 target loci, or 100-200 target loci, or 200-500 target loci spanning at least one cancer-specific mutation within one reaction volume.
[0045] The methods provided herein, in exemplary embodiments, analyze single nucleotide variant mutations (SNVs) in circulating fluids, particularly cell-free and / or circulating tumor DNA. The methods offer the advantage of identifying many of the mutations found in tumors and clones, as well as subclonal mutations, in a single test, rather than multiple tests utilizing tumor samples that would be required for any effectiveness. The methods and compositions may be useful by themselves, or they may be useful when used in conjunction with other methods for cancer detection, diagnosis, staging, screening, treatment, and management, e.g., to help confirm the results of these other methods and provide more reliable and / or conclusive results.
[0046] Thus, in one embodiment, provided herein is a method for determining cancer-specific mutations (e.g., SNVs, MNVs, indels, or gene fusions) present in a cancer by determining the cancer-specific mutations present in a ctDNA sample from an individual, such as an individual having or suspected of having cancer (e.g., lung cancer, breast cancer, bladder cancer, or colorectal cancer), using the ctDNA amplification / sequencing workflow provided herein. In some embodiments, the method detects at least one cancer-specific mutation in at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% of patients with early recurrence or metastasis of the cancer.
[0047] In some embodiments, the methods described herein can detect patient-specific cancer-associated mutations in patients with early cancer recurrence or metastasis at least 30 days, at least 60 days, at least 100 days, at least 150 days, at least 200 days, at least 250 days, or at least 300 days before clinical determination of cancer recurrence or metastasis detectable by imaging and / or well-established biomarkers. Exemplary imaging methods include X-ray, magnetic resonance imaging (MRI), positron emission tomography (PET), nuclear medicine scan, computed tomography (CT) imaging, mammogram, or ultrasound. Imaging methods for diagnosing cancer may include microscopic examination and histological staining of biological samples. In some embodiments, the methods described herein are capable of detecting patient-specific breast cancer-associated mutations in patients with early recurrence or metastasis of breast cancer at least 30 days, at least 60 days, at least 100 days, at least 150 days, at least 200 days, at least 250 days, or at least 300 days before elevated CA15-3 levels.
[0048] In some embodiments, the methods described herein have a specificity of at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or at least 99.9% in detecting early cancer recurrence or metastasis when one or more patient-specific cancer-associated mutations are detected above a predetermined confidence threshold (e.g., 0.95, 0.96, 0.97, 0.98, or 0.99). In some embodiments, the methods detect at least one cancer-specific mutation in at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% of patients with early cancer recurrence or metastasis.
[0049] II. Sample Collection It is contemplated that the methods disclosed herein may be used to monitor or detect a wide variety of cancers in patients, and one of skill in the art will understand that different types of cancers will require the collection of different types of samples, as described herein.
[0050] In some embodiments, the cancer is a solid tumor and the biological sample is a tumor biopsy sample. Performing a biopsy generally involves using a sharp tool to remove a small amount of tissue from something suspected of containing diseased cells or tissue, such as a tumor. There are many different types of biopsies, such as needle biopsy, CT-guided biopsy, ultrasound-guided biopsy, bone biopsy, bone marrow biopsy, liver biopsy, kidney biopsy, aspiration biopsy, prostate biopsy, skin biopsy, and surgical biopsies such as laparoscopic biopsy. In some embodiments, the biological sample is obtained by liquid biopsy. In some embodiments, the biological sample is a blood, serum, plasma, or urine sample. Furthermore, biological liquid samples can be extracted from various animal fluids containing cell-free DNA, including, but not limited to, blood, serum, plasma, bone marrow, urine vitreous, sputum, tears, sweat, saliva, semen, mucosal excretions, mucus, spinal fluid, amniotic fluid, lymphatic fluid, etc. Cell-free DNA can be derived from a fetus (via fluid collected from a pregnant subject) or from the subject's own tissues.
[0051] In some embodiments, the cancer is a blood cancer and the biological sample is a liquid sample. In some embodiments, the cancer is a blood cancer and the biological sample is a blood, serum, plasma, or bone marrow sample. In some embodiments, both the cancer-derived DNA and the matched normal DNA are obtained from a blood sample by isolating and separating the plasma and buffy coat. The DNA obtained from the buffy coat can serve as the matched normal DNA to the circulating tumor DNA obtained from the plasma fraction.
[0052] In some embodiments, the methods of the present disclosure further include longitudinally collecting multiple liquid biopsy samples from the patient. In some embodiments, the liquid biopsy samples are obtained from the patient after the patient has undergone treatment for cancer. In some embodiments, the liquid biopsy samples are blood, serum, plasma, or urine samples.
[0053] In certain embodiments, the method provided herein is particularly adapted to amplify DNA fragments, particularly tumor DNA fragments found in circulating tumor DNA (ctDNA).Such fragments are typically about 160 nucleotides in length.
[0054] It is known in the art that cell-free nucleic acids (cfNAs), such as cfDNA, can be released into the circulation through various forms of cell death, such as apoptosis, necrosis, autophagy, and necroptosis. cfDNA is fragmented, and the size distribution of the fragments varies from 150-350 bp to over 10,000 bp (see Kalnina et al. World J Gastroenterol. 2015 Nov 7;21(41):11636-11653). For example, the size distribution of plasma DNA fragments in hepatocellular carcinoma (HCC) patients ranged from 100-220 bp in length, with a peak frequency of approximately 166 bp, and the highest tumor DNA concentration among the fragments was 150-180 bp in length (see Jiang et al. Proc Natl Acad Sci USA 112:E1317-E1325).
[0055] In an exemplary embodiment, circulating tumor DNA (ctDNA) is isolated from blood using EDTA-2Na tubes after centrifugation to remove cellular debris and platelets. Plasma samples can be stored at -80°C until DNA is extracted, for example, using a QIAamp DNA Mini Kit (Qiagen, Hilden, Germany) (see, e.g., Hamakawa et al., Br J Cancer. 2015;112:352-356). reported that the median concentration of extracted cell-free DNA for all samples was 43.1 ng / ml of plasma (range 9.5-1338 ng / ml), with a mutant fraction range of 0.001-77.8%, with a median of 0.90%.
[0056] In certain exemplary embodiments, the sample is tumor.In view of the teachings of this specification, the method for isolating nucleic acid from tumor and the method for making nucleic acid library from such DNA sample are known in the art.In addition, in view of the teachings of this specification, those skilled in the art will know how to make nucleic acid library suitable for the method of this specification from other samples, such as other liquid samples in which DNA is suspended in a free state, in addition to ctDNA sample.
[0057] III. Identification of cancer-specific mutations After sample collection, targeted sequencing or whole exome sequencing (WES) can be performed on circulating tumor DNA, cell-free DNA, or cellular DNA obtained from solid tumor or liquid biopsy samples, and matched normal tissues or cells, as described above, depending on the type of cancer being analyzed. Comparing sequences from tumor or cancer cells with sequences from normal tissues or cells allows for the identification of cancer-specific mutations. Following the identification of personalized cancer-specific mutations for a patient, cancer in the patient can be detected or monitored using the personalized cancer-specific mutations. Detection of personalized cancer-specific mutations before, during, and after cancer treatment can be an indicator of cancer relapse, recurrence, or metastasis.
[0058] In some embodiments, cancer-specific mutations include one or more somatic mutations. Somatic mutations can be distinguished from germline mutations by, for example, sequencing nucleic acids isolated from a patient's non-cancer cells to identify one or more non-cancer-specific germline mutations, and the nucleic acids are enriched in a panel of cancer-related genomic loci. In some embodiments, the non-cancer cells are obtained from the buffy coat of a patient's blood sample. Germline mutations can be filtered by first performing a number of targets selected for a first patient-specific assay on the non-cancer DNA obtained from the buffy coat, and then selecting cancer-specific variants for a second patient-specific assay.
[0059] In some embodiments, the disclosed methods further include comparing the sequences of amplified DNA prepared from two longitudinally collected liquid biopsy samples to identify one or more non-cancer-specific germline mutations. Germline mutations have a variant allele frequency (VAF) of about 50% in the consecutive biological samples. In some embodiments, where the level of ctDNA is very high, the copy number of the variant region may need to be considered to determine and filter germline mutations.
[0060] In some embodiments, germline mutations can be determined by separating cell-free DNA from plasma samples into long and short DNA fractions, and analyzing both fractions using a custom (individualized or patient-specific) assay. Tumor-specific variants are expected to have higher variant allele frequencies in samples with shorter DNA fractions. Alternatively, in some embodiments, shorter fragments can be enriched, and germline mutations can be identified by comparing the variant allele frequencies for mutations in the enriched sample with those in the original sample.
[0061] In some embodiments, the methods of the present disclosure further comprise comparing the sequence of the nucleic acid isolated from the biological sample to a germline mutation database to identify one or more non-cancer-specific germline mutations.
[0062] Upon identifying a patient's cancer-specific mutations, multiplex PCR is performed to amplify multiple target loci from cell-free DNA isolated from the patient's liquid biopsy sample to obtain amplified DNA. In some embodiments, the multiplex amplification targets 1 to 100 target loci, or 1 to 20 target loci, or 1 to 10 target loci, or 10 to 20 target loci, or 20 to 50 target loci, each spanning at least one cancer-specific mutation. In some embodiments, the multiplex amplification targets 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 target loci, each spanning at least one cancer-specific mutation.
[0063] In one aspect, cancer-specific mutations are identified by performing whole exome sequencing (WES) on DNA obtained from a liquid sample or solid tumor sample and compared with whole exome sequencing of normal tissue. In some embodiments, whole exome sequencing is performed on cellular DNA obtained from the solid tumor and matched normal tissue. In some embodiments, whole exome sequencing is performed on cell-free DNA from a liquid biopsy sample, such as blood or plasma. In some embodiments, WES is performed on cell-free or cellular DNA obtained from a blood sample from a patient suffering from a blood cancer to identify cancer-specific blood cancer mutations. By comparing sequencing data from DNA obtained from a blood cancer or solid tumor with DNA obtained from a normal matched tissue, cancer-specific mutations are identified and can be used to monitor or detect cancer during the clinical progression of a patient's cancer.
[0064] As used herein, "whole exome sequencing" refers to the sequencing of all protein-coding regions of genes in the genome, also known as the exome. Thus, whole exome sequencing can first involve isolating a subset of DNA-encoded proteins, known as the exome, prior to sequencing. This initial step can be performed by isolated exon capture techniques, i.e., array-based capture or in-solution capture, as described elsewhere herein.
[0065] In another embodiment, cancer-specific mutations are identified by targeted sequencing of nucleic acids from biological samples obtained from patients.Biological samples can be obtained by solid tumor biopsy or liquid biopsy as described above.Cancer nucleic acids can be cellular DNA obtained from solid tumors, cell-free DNA or circulating DNA obtained from any liquid sample as described above, or cancer DNA can be cell-free DNA or cellular DNA obtained from a blood sample of a patient suffering from blood cancer.Normal matched DNA can be cellular DNA obtained from non-cancerous cells or tissues from the patient.
[0066] In some embodiments of the present disclosure, targeted sequencing is performed by enriching nucleic acids obtained from a patient in a panel of cancer-associated genes or genomic loci to reduce the number of target loci or nucleic acid bases required to identify patient-specific tumor or cancer cell mutations. In some embodiments, targeted sequencing involves enriching nucleic acids (e.g., cellular DNA) obtained from a patient's solid tumor biopsy sample in a panel of cancer-associated genes. In some embodiments, targeted sequencing is performed by enriching nucleic acids (e.g., cfDNA) obtained from a patient's blood, plasma, serum, or urine sample in a panel of cancer-associated genes.
[0067] In some embodiments, the panel includes 2,000 or fewer cancer-associated genes or genomic loci, or 1,000 or fewer cancer-associated genes or genomic loci, or 500 or fewer cancer-associated genes or genomic loci, or 100 to 1,000 cancer-associated genes or genomic loci, or 200 to 500 cancer-associated genes or genomic loci. In some embodiments, the panel includes about 100 to about 300 cancer-associated genes or genomic loci, about 300 to about 450 cancer-associated genes or genomic loci, about 200 to about 350 cancer-associated genes or genomic loci, about 500 to about 1,000 cancer-associated genes or genomic loci, about 1,000 to about 1,500 cancer-associated genes or genomic loci, about 1,500 to about 2,000 cancer-associated genes or genomic loci, or about 1,650 to about 2,000 cancer-associated genes or genomic loci. In some embodiments, the panel comprises from about 100, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1000, 1500, 1850, or 2000 cancer-associated genes or genomic loci.
[0068] In some embodiments, sequencing of nucleic acid isolated from a first biological sample obtained from a patient produces up to 5,000,000 bases of DNA sequence, or up to 4,000,000 bases of DNA sequence, or up to 3,000,000 bases of DNA sequence, or up to 2,000,000 bases of DNA sequence, or between 500,000 and 2,000,000 bases of DNA sequence, or between 1,000,000 and 1,500,000 bases of DNA sequence. As used herein, the term "cancer-associated genomic locus" refers to any genomic locus determined to be useful for monitoring or detecting cancer in a patient. Cancer-associated genomic loci may be associated with (i) the likelihood of cancer metastasis, the likelihood of metastasis to specific organs, the risk of recurrence, and / or the course of the tumor; (ii) the stage of the tumor; (iii) the patient's prognosis in the absence of treatment for the cancer; (iv) a prognosis of the patient's response (e.g., tumor shrinkage or progression-free survival) to treatment (e.g., chemotherapy, radiation therapy, surgery to remove the tumor, etc.); (v) a diagnosis of the patient's actual response to current and / or past treatments; (vi) a determination of a preferred course of treatment for the patient; (vii) a prognosis for the patient's recurrence after treatment (either in general or some specific treatment); (viii) a prognosis for the patient's life expectancy (e.g., a prognosis for overall survival);
[0069] Thus, in some embodiments, cancer-associated genomic loci are associated with rapidly proliferating (and therefore more aggressive) cancer cells. Such cancers in patients often mean that the patient has an increased likelihood of recurrence after treatment (e.g., cancer cells that are not killed or eliminated by treatment grow quickly). Such cancers may also mean that the patient has an increased likelihood of cancer progression due to more rapid progression (e.g., rapidly proliferating cells cause any tumor to grow rapidly, become more toxic, and / or metastasize). Such cancers may also mean that the patient may require relatively more aggressive treatment. Thus, in some embodiments, the invention provides methods of classifying cancers, comprising determining the status of a gene panel comprising at least two or more cancer-associated genomic loci, wherein an abnormal status indicates an increased likelihood of recurrence or progression.
[0070] In some embodiments, the panel of cancer-associated genomic loci comprises exons, introns, gene regulatory regions, non-coding RNA, and rearranged genes. In some embodiments, the cancer-specific mutations comprise one or more single nucleotide variants (SNVs), one or more multi-nucleotide variants (MNVs), one or more copy number variants (CNVs), one or more indels, one or more gene fusions, one or more structural variants, or a combination thereof.
[0071] In some embodiments, the panel of cancer-associated genomic loci includes any genomic alteration of any size, from a single nucleotide change to an alteration of a genomic region of more than 1 kilobase (kb). The term "indel" refers to both insertions and deletions of nucleic acids within a genome. As used herein, the term "structural variant" refers to a genomic alteration, such as a deletion or insertion, that involves a DNA segment of more than 1 kilobase (kb) and can be either microscopic or submicroscopic. The term "gene fusion" refers to any genomic alteration that results in the fusion of two different genomic loci caused by the insertion and / or deletion of DNA within the genome. The resulting genomic alteration caused by a gene fusion can involve a DNA segment of any size.
[0072] Noncoding RNAs (ncRNAs) are functional RNA molecules that are transcribed from DNA but not translated into proteins. Epigenetically relevant ncRNAs include miRNAs, siRNAs, piRNAs, and lncRNAs. Generally, ncRNAs function to regulate gene expression at the transcriptional and posttranscriptional levels. Those ncRNAs that appear to be involved in epigenetic processes can be divided into two major groups: short ncRNAs (<30 nt) and long ncRNAs (>200 nt). The three major classes of short noncoding RNAs are microRNAs (miRNAs), small interfering RNAs (siRNAs), and piwi-interacting RNAs (piRNAs). Both major groups have been shown to play roles in heterochromatin formation, histone modification, DNA methylation targeting, and gene silencing.
[0073] In some embodiments, the panel of cancer-associated genomic loci includes a list or set of known cancer genes, oncogenes, or any genes reported to be altered in cancer cells or tumor tissue. Cancer-associated genes refer to genes associated with an altered risk for cancer (e.g., breast cancer, bladder cancer, or colon cancer) or an altered prognosis for cancer. Exemplary cancer-associated genes that promote cancer include oncogenes, genes that promote cell proliferation, invasion, or metastasis, genes that inhibit apoptosis, and pro-angiogenic genes. Cancer-associated genes that inhibit cancer include, but are not limited to, tumor suppressor genes, genes that inhibit cell proliferation, invasion, or metastasis, genes that promote apoptosis, and anti-angiogenic genes.
[0074] In some embodiments, the cancer-associated genomic loci of the panel are AKT1 (14q32.33), ALK (2p23.2-23.1), APC (5q22.2), AR (Xq12), ARAF (Xp11.3), ARID1A (1p36.11), ATM (11q22.3), BRAF (7q34), BRCA1 (17q21.31), BRCA2 (13q13.1), CCND1 (11q13.3), CCND2 (12p13.32), CCNE1 (19q12), CDH1 (16q22.1), CDK4 (12q14.1), CDK6 (7q21. 2), CDKN2A(9p21.3), CTNNB1(3p22.1), DDR2(1q23.3), EGFR(7p11.2), ERBB2(17q12), ESR1(6q25.1-25.2), EZH2(7q36.1), FBXW7(4q31.3), FGFR1( 8p11.23), FGFR2(10q26.13), FGFR3(4p16.3), GATA3(10p14), GNA11(19p13.3), GNAQ(9q21.2), GNAS(20q13.32), HNF1A(12q24.31), HRAS(11p15.5) , IDH1(2q34), IDH2(15q26.1), JAK2(9p24.1), JAK3(19p13.11), KIT(4q12), KRAS(12p12.1), MAP2K1(15q22.31), MAP2K2(19p13.3), MAPK1(22q11. 22), MAPK3 (16p11.2), MET (7q31.2), MLH1 (3p22.2), MPL (1p34.2), MTOR (1p36.22), MYC (8q24.21), NF1 (17q11.2), NFE2L2 (2q31.2), NOTCH1 (9q34.3) ), NPM1(5q35.1), NRAS(1p13.2), NTRK1(1q23.1), NTRK3(15q25.3), PDGFRA(4q12), PIK3CA(3q26.32), PTEN(10q23.31), PTPN11(12q24.13), RAF1(3 p25.2), RB1(13q14.2), RET(10q11.21), RHEB(7q36.1), RHOA(3p21.31), RIT1(1q22), ROS1(6q22.1), SMAD4(18q21.2), SMO(7q32.1), STK11(19p13.3), TERT (5p15.33), TP53 (17p13.1), TSC1 (9q34.13), and / or VHL (3p25.3). An embodiment of the mutation detection method begins with selecting a region of the gene to target. The region with the known mutation is used to develop primers for mPCR-NGS to amplify and detect the mutation.
[0075] The methods provided herein can be used to detect virtually any type of mutation, particularly mutations known to be associated with cancer, and most particularly, the methods provided herein are directed to mutations, particularly single nucleotide variants (SNVs), copy number variations (CNVs), indels, or gene fusions or rearrangements associated with cancer. Exemplary SNVs may be in one or more of the following genes: EGFR, FGFR1, FGFR2, ALK, MET, ROS1, NTRK1, RET, HER2, DDR2, PDGFRA, KRAS, NF1, BRAF, PIK3CA, MEK1, NOTCH1, MLL2, EZH2, TET2, DNMT3A, SOX2, MYC, KEAP1, CDKN2A, NRG1, TP53, LKB1, and PTEN, which have been identified as mutated, or in copy number gain, or fused to other genes, and combinations thereof, in a variety of lung cancer samples (Non-small-cell lung cancers: a heterogeneous set of diseases. Chen et al. Nat. Rev. Cancer. 2014 Aug 14(8):535-551). In another example, the list of genes is as listed above and the SNVs are reported, such as in the Chen et al. reference.
[0076] Exemplary embodiments of potential cancer-associated genomic loci include the exonic regions of the following genes (e.g., for detection of SNVs, CNVs, and indels): ABL1 ACVR1B AKT1 AKT2 AKT3 ALK ALOX12B AMER1 (FAM123B) APC AR ARAF ARFRP1 ARID1A ASXL1 ATM ATR ATRX AURKA AURKB AXIN1 AXL BAP1 BARD1 BCL2 BCL2L1 BCL2L2 BCL6 BCOR BCORL1 BRAF BRCA1 BRCA2 BRD4 BRIP1 BTG1 BTG2 BTK C11orf30 (EMSY) CALR CARD11 CASP8 CBFB CBL CCND1 CCND2 CCND3 CCNE1 CD22 CD274 (PD-L1) CD70 CD79A CD79B CDC73 CDH1 CDK12 CDK4 CDK6 CDK8 CDKN1A CDKN1B CDKN2A CDKN2B CDKN2C CEBPA CHEK1 CHEK2 CIC CREBBP CRKL CSF1R CSF3R CTCF CTNNA1 CTNNB1 CUL3 CUL4A CXCR4 CYP17A1 DAXX DDR1 DDR2 DIS3 DNMT3A DOT1L EED EGFR EP300 EPHA3 EPHB1 EPHB4 ERBB2 ERBB3 ERBB4 ERCC4 ERG ERRFI1 ESR1 EZH2 FAM46C FANCA FANCC FANCG FANCL FAS FBXW7 FGF10 FGF12 FGF14 FGF19 FGF23 FGF3 FGF4 FGF6 FGFR1 FGFR2 FGFR3 FGFR4 FH FLCN FLT1 FLT3 FOXL2 FUBP1 GABRA6 GATA3 GATA4 GATA6 GID4 (C17orf39) GNA11 GNA13 GNAQ GNAS GRM3 GSK3B H3F3A HDAC1 HGF HNF1A HRAS HSD3B1 ID3 IDH1 IDH2 IGF1R IKBKE IKZF1 INPP4B IRF2 IRF4 IRS2 JAK1 JAK2 JAK3 JUN KDM5A KDM5C KDM6A KDRKEAP1 KEL KIT KLHL6 KMT2A (MLL) KMT2D (MLL2) KRAS LTK LYN MAF MAP2K1 (MEK1) MAP2K2 (MEK2) MAP2K4 MAP3K1 MAP3K13 MAPK1 MCL1 MDM2 MDM4 MED12 MEF2B MEN1 MERTK MET MITF MKNK1 MLH1 MPL MRE11A MSH2 MSH3 MSH6 MST1R MTAP MTOR MUTYH MYC MYCL (MYCL1) MYCN MYD88 NBN NF1 NF2 NFE2L2 NFKBIA NKX2-1 NOTCH1 NOTCH2 NOTCH3 NPM1 NRAS NT5C2 NTRK1 NTRK2 NTRK3 P2RY8 PALB2 PARK2 PARP1 PARP2 PARP3 PAX5 PBRM1 PDCD1 (PD-1) PDCD1LG2 (PD-L2) PDGFRA PDGFRB PDK1 PIK3C2B PIK3C2G PIK3CA PIK3CB PIK3R1 PIM1 PMS2 POLD1 POLE PPARG PPP2R1A PPP2R2A PRDM1 PRKAR1A PRKCI PTCH1 PTEN PTPN11 PTPRO QKI RAC1 RAD21 RAD51 RAD51B RAD51C RAD51D RAD52 RAD54L RAF1 RARA RB1 RBM10 REL RET RICTOR RNF43 ROS1 RPTOR SDHA SDHB SDHC SDHD SETD2 SF3B1 SGK1 SMAD2 SMAD4 SMARCA4 SMARCB1 SMO SNCAIP SOCS1 SOX2 SOX9 SPEN SPOP SRC STAG2 STAT3 STK11 SUFU SYK TBX3 TEK TET2 TGFBR2 TIPARP TNFAIP3 TNFRSF14 TP53 TSC1 TSC2 TYRO3 U2AF1 VEGFA VHL WHSC1 (MMSET) WHSC1L1 WT1 XPO1 XRCC2 ZNF217Also, exemplary embodiments of potential cancer-associated genomic loci include intronic regions, promoter regions, and non-coding RNA sequences of the following genes (e.g., for detection of gene fusions or rearrangements): ALK BCL2 BCR BRAF BRCA1 BRCA2 CD74 EGFR ETV4 ETV5 ETV6 EWSR1 EZR FGFR1 FGFR2 FGFR3 KIT KMT2A (MLL) MSH2 MYB MYC NOTCH2 NTRK1 NTRK2 NUTM1 PDGFRA RAF1 RARA RET ROS1 RSPO2 SDC4 SLC34A2 TERC TERT TMPRSS2.
[0077] IV. Methods for Enriching for Nucleic Acids in a Panel of Cancer-Related Genes or Isolating Exonic Genomic DNA for Whole Exome Sequencing Target enrichment methods allow for selective capture of genomic regions of interest from DNA samples prior to sequencing by enrichment methods such as hybrid capture or targeted PCR. The genomic regions of interest can be any subset of genomic loci, such as the cancer-associated genomic loci described above, or all exonic regions of the genome for preparing samples for whole exome sequencing (WES).
[0078] Generally, hybrid capture involves designing an oligonucleotide sequence that can bind to a genomic DNA sequence of interest through complementarity. The oligonucleotide is bound to a solid surface or bead, allowing the genomic sequence bound to the oligonucleotide to be separated from the unbound genomic sequence. The unbound genomic DNA sequence can then be washed away, leaving the genomic sequence of interest bound to the solid surface or bead for further processing and / or amplification. In some embodiments, a panel of cancer-associated genomic loci is enriched by hybrid capture, such as array-based hybrid capture or in-solution hybrid capture.
[0079] In some embodiments, target enrichment can be an array-based hybrid capture method. In some embodiments, array-based hybrid capture methods can involve designing a microarray by immobilizing single-stranded oligonucleotide sequences from the human genome and tiling regions of interest immobilized on the surface of a microarray chip or surface. Genomic DNA is sheared to form double-stranded fragments. The fragments undergo end repair to generate blunt ends, and adapters with universal priming sequences are added. These fragments are hybridized to oligos on the microarray chip or surface. Unhybridized fragments are washed away, and the desired fragments are eluted. The fragments are then amplified using polymerase chain reaction. The microarray used for array-based hybrid capture can be a Roche Nimblegen™ array, an Agilent™ capture array, or a similar comparative genomic hybridization array that can be used for hybrid capture of target sequences. In some embodiments, a panel of cancer-associated genomic loci is enriched by hybrid capture. In other embodiments, the target enrichment strategy can be an in-solution capture strategy. To capture genomic regions of interest using solution capture, a pool of custom oligonucleotides (probes) is synthesized and hybridized to a fragmented genomic DNA sample in solution. The probes (labeled on beads) selectively hybridize to the genomic regions of interest, after which the beads (containing the DNA fragments of interest) can be pulled down and washed to remove excess material. The beads can then be removed and the genomic fragments sequenced, allowing for selective DNA sequencing of the genomic regions of interest (e.g., exons, introns, promoter regions or other gene regulatory regions, or non-coding RNA sequences).
[0080] In contrast to hybrid capture, in solution capture, there is an excess of probes targeting the region of interest that exceeds the amount of template required. The optimal target size is about 3.5 megabases, which provides excellent sequence coverage of the target region. The preferred method depends on several factors, including the number of base pairs in the region of interest, the demand for reads in the target, in-house equipment, etc.
[0081] Alternatively, cancer-related genomic loci can be enriched by targeted amplification. Targeted amplification of genomic loci can be achieved by multiplex PCR using primers designed to target specific regions. Protocols for multiplex PCR of multiple desired targets are described in detail elsewhere herein.
[0082] V. Cancer The terms "cancer" and "cancerous" refer to or describe a physiological condition in animals that is typically characterized by uncontrolled cell growth. A "tumor" contains one or more cancerous cells. There are several major types of cancer. Carcinoma is cancer that begins in the skin or in tissues that outline or cover internal organs. Sarcoma is cancer that begins in bone, cartilage, fat, muscle, blood vessels, or other connective or supportive tissue. Leukemia is cancer that begins in blood-forming tissues, such as the bone marrow, and causes large numbers of abnormal blood cells to be produced and enter the blood. Lymphoma and multiple myeloma are cancers that begin in cells of the immune system. Central nervous system cancer is cancer that begins in the tissues of the brain and spinal cord.
[0083] In some embodiments, the cancer is a cancer or tumor of the abdomen or abdominal wall, adrenal gland, anus, appendix, bladder, bone, brain, breast, cervix, chest wall, colon, diaphragm, duodenum, ear, endometrium, esophagus, fallopian tube, gallbladder, gastroesophageal junction, head and neck, kidney, larynx, liver, lung, lymph node, malignant effusion, mediastinum, nasal cavity, omentum, ovary, pancreas, pancreaticobiliary duct, parotid gland, pelvis, penis, pericardium, peritoneum, pleura, prostate, rectum, salivary gland, skin, small intestine, soft tissue, spleen, stomach, thyroid, tongue, trachea, ureter, uterus, vagina, vulva, or Whipple resection.
[0084] In some embodiments, the cancer is lung cancer, breast cancer, bladder cancer, or colon cancer.
[0085] In some embodiments, the cancer is acute lymphoblastic leukemia, acute myeloid leukemia, adrenocortical carcinoma, AIDS-related cancer, AIDS-related lymphoma, anal cancer, appendix cancer, astrocytoma, atypical teratoid / rhabdoid tumor, basal cell carcinoma, bladder cancer, brain stem glioma, brain tumor (including brain stem glioma, central nervous system atypical teratoid / rhabdoid tumor, central nervous system embryonal tumor, astrocytoma, craniopharyngioma, ependymoblastoma, ependymoma, medulloblastoma, medulloepithelioma, intermediate pineal parenchymal tumor, supratentorial primitive neuroectodermal tumor, and pineoblastoma), breast cancer, bronchial tumor, Burkitt's lymphoma, primary tumor, Cancer of unknown location, carcinoid tumor, carcinoma of unknown primary site, central nervous system atypical teratoid / rhabdoid tumor, central nervous system embryonal tumor, cervical cancer, childhood cancer, chordoma, chronic lymphocytic leukemia, chronic myeloid leukemia, chronic myeloproliferative disorder, colon cancer, colorectal cancer, craniopharyngioma, cutaneous T-cell lymphoma, endocrine pancreatic islet cell tumor, endometrial cancer, ependymoblastoma, ependymoma, esophageal cancer, nasal neuroblastoma, Ewing's sarcoma, extracranial germ cell tumor, extragonadal germ cell tumor, extrahepatic bile duct cancer, gallbladder cancer, gastric (stomach) cancer, gastrointestinal carcinoid tumor, gastrointestinal stromal cell tumor, gastrointestinal stromal tumor (GIST), pregnancy Chronic trophoblastic tumor, glioma, hairy cell leukemia, head and neck cancer, heart cancer, Hodgkin's lymphoma, hypopharyngeal cancer, intraocular melanoma, pancreatic islet tumor, Kaposi's sarcoma, kidney cancer, Langerhans cell histiocytosis, laryngeal cancer, lip cancer, liver cancer, malignant fibrous histiocytoma / bone cancer, medulloblastoma, medulloepithelioma, melanoma, Merkel cell carcinoma, Merkel cell skin carcinoma, mesothelioma, metastatic squamous cell neck cancer of unknown primary, oral cancer, multiple endocrine neoplasia syndrome, multiple myeloma, multiple myeloma / plasma cell neoplasm, mycosis fungoides, myelodysplastic syndrome, myeloproliferative neoplasm, nasal cavity cancer, nasopharyngeal cancer, neuroblastoma, non- Hodgkin's lymphoma, non-melanoma skin cancer, non-small cell lung cancer, mouth cancer, oral cavity cancer, oropharyngeal cancer, osteosarcoma, other brain and spinal cord tumors, ovarian cancer, ovarian epithelial cancer, ovarian germ cell tumor, ovarian low malignant potential tumor, pancreatic cancer, papillomatosis, sinus cancer, parathyroid cancer, pelvic cancer, penile cancer, pharyngeal cancer, intermediate pineal parenchymal tumor, pineoblastoma, pituitary tumor, plasma cell neoplasm / multiple myeloma, pleuropulmonary blastoma, primary central nervous system (CNS) lymphoma, primary hepatocellular carcinoma, prostate cancer, rectal cancer, kidney cancer, renal cell (kidney) cancer, renal cell carcinoma, respiratory tract cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer,Sézary syndrome, small cell lung cancer, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma, squamous cell neck cancer, gastric (stomach) cancer, supratentorial primitive neuroectodermal tumor, T-cell lymphoma, testicular cancer, throat cancer, thymic cancer, thymoma, thyroid cancer, transitional cell carcinoma, transitional cell carcinoma of the renal pelvis and ureter, trophoblastic tumor, ureteral cancer, urethral cancer, uterine cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenstrom's macroglobulinemia, or Wilms' tumor.
[0086] In another embodiment, provided herein is a method for detecting cancer in a blood sample or fraction thereof from an individual, such as an individual suspected of having cancer, comprising determining single nucleotide variants present in the sample by determining single nucleotide variants present in the ctDNA sample using the ctDNA SNV amplification / sequencing workflow provided herein. The presence of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 SNVs at the lower end of the range and 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, or 50 SNVs at the upper end of the range in the sample at a plurality of single nucleotide loci is indicative of the presence of cancer.
[0087] In another embodiment, a method for detecting clonal single nucleotide variants (SNVs) in an individual's tumor is provided herein. The method includes, for example, performing a ctDNA amplification / sequencing workflow as provided herein in the Examples, and determining the variant allele frequency of each SNV locus based on the sequence of multiple copies of a series of amplicons. A relatively high allele frequency of multiple single nucleotide variant loci compared to other single nucleotide variants is indicative of clonal single nucleotide variants in the tumor. Variant allele frequencies are well known in the art of sequencing.
[0088] In certain embodiments, the method further includes determining a treatment plan, a therapy, and / or administering to the individual a compound that targets one or more clonal single nucleotide variants. In certain instances, subclones and / or other clonal SNVs are not targeted by the therapy. Specific therapies and associated mutations are provided in other sections of this specification and are known in the art. Thus, in certain instances, the method further includes administering to the individual a compound that is known to be specifically effective in treating cancers with one or more of the determined single nucleotide variants.
[0089] In certain aspects of this embodiment, a variant allele frequency of greater than 0.25%, 0.5%, 0.75%, 1.0%, 5% or 10% is indicative of a clonal single nucleotide variant.
[0090] In certain examples of this embodiment, the cancer is stage 1a, 1b, or 2a breast cancer, bladder cancer, or colon cancer. In certain examples of this embodiment, the cancer is stage 1a or 1b breast cancer, bladder cancer, or colon cancer. In certain examples of this embodiment, the individual does not undergo surgery. In certain examples of this embodiment, the individual does not undergo a biopsy.
[0091] In some examples of this embodiment, a clonal SNV is identified or further identified when other tests, such as direct tumor tests, suggest that the SNV under test is a clonal SNV for SNVs under test that have a test variable allele frequency that is at least one-quarter, one-third, one-half, or more than three-quarters of other determined single nucleotide variants.
[0092] In some embodiments, the methods herein for detecting SNVs in ctDNA can be used in place of direct analysis of DNA from the tumor.
[0093] In certain examples of any of the embodiments of the methods provided herein, data is provided about SNVs found in tumors from an individual before targeted amplification is performed on ctDNA from the individual. Thus, in these embodiments, SNV amplification / sequencing reactions are performed on one or more tumor samples from the individual. In this method, the ctDNA SNV amplification / sequencing reactions provided herein are still advantageous because they provide a liquid biopsy of clonal and subclonal mutations. Furthermore, as provided herein, clonal mutations can be more clearly identified in individuals with cancer when a high VAF rate, for example, a VAF greater than 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10%, is determined for SNVs in ctDNA samples from the individual.
[0094] In certain embodiments, the method provided herein can be used to determine whether to isolate and analyze ctDNA from circulating free nucleic acid from an individual with cancer.First, determine whether the cancer is breast cancer, bladder cancer, or colon cancer.If the cancer is breast cancer, bladder cancer, or colon cancer, isolate circulating free nucleic acid from the individual.In some examples, the method further includes determining the stage of cancer.
[0095] In some methods, provided herein are compositions and / or solid supports of the invention: A composition comprising circulating tumor nucleic acid fragments comprising universal adaptors, wherein the circulating tumor nucleic acid was derived from breast cancer, bladder cancer, or colon cancer.
[0096] In some embodiments, provided herein are compositions of the invention comprising circulating tumor nucleic acid fragments containing universal adaptors, where the circulating tumor nucleic acid was derived from a blood sample or fraction thereof of an individual with cancer. These methods typically involve forming ctDNA fragments containing universal adaptors. Furthermore, such methods typically involve forming a solid support, particularly a solid support for high-throughput screening, that comprises a clonal population of a plurality of nucleic acids, where the clonal population comprises amplicons generated from a sample of circulating free nucleic acid, and is ctDNA. In exemplary embodiments based on the surprising results provided herein, the ctDNA was derived from a cancer.
[0097] Similarly, provided herein as one embodiment of the present invention is a solid support comprising a clonal population of a plurality of nucleic acids, the clonal population comprising nucleic acid fragments generated from a sample of circulating free nucleic acids from a sample of blood or a fraction thereof from an individual with cancer.
[0098] In certain embodiments, nucleic acid fragments in different clonal populations contain the same universal adaptor. Such compositions are typically formed during high-throughput sequencing reactions in the methods of the invention.
[0099] A clonal collection of nucleic acids can be derived from nucleic acid fragments from a set of samples from two or more individuals, in these embodiments, the nucleic acid fragments comprise one of a set of molecular barcodes corresponding to a sample in the set of samples.
[0100] VI. Analysis Methods SNV1 and 2 Detailed analysis methods are provided herein as SNV Method 1 and SNV Method 2 in the Analysis section of this specification. Any of the methods provided herein can further include the analysis steps provided herein. Thus, in certain examples, a method for determining whether a single nucleotide variant is present in a sample includes identifying a confidence value for each allele determination at each of a set of single nucleotide variant loci, which may be based at least in part on the read depth for the locus. Confidence limits can be set at at least 75%, 80%, 85%, 90%, 95%, 96%, 96%, 98%, or 99%. Confidence limits can be set at different levels for different types of mutations.
[0101] The method can be performed at a read depth for a set of at least 5, 10, 15, 20, 25, 50, 100, 150, 200, 250, 500, 1,000, 10,000, 25,000, 50,000, 100,000, 250,000, 500,000, or 1 million single nucleotide variant loci.
[0102] In certain embodiments, any of the methods described herein may include determining the efficiency and / or error rate per cycle for each amplification reaction of a multiplex amplification reaction of single nucleotide variant loci. The efficiency and error rate can then be used to determine whether a single nucleotide variant at a set of single variant loci is present in a sample. Further detailed analysis steps provided in SNV Method 2 may also be included in certain embodiments.
[0103] In an exemplary embodiment of any of the methods herein, the set of single nucleotide variant loci includes all of the single nucleotide variant loci identified in the TCGA and COSMIC datasets for cancer.
[0104] In certain embodiments of any of the methods herein, the set of single nucleotide variant loci includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 75, 100, 250, 500, 1000, 2500, 5000, or 10,000 single nucleotide variant loci known to be associated with cancer at the lower end of the range, and 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 75, 100, 250, 500, 1000, 2500, 5000, 10,000, 20,000, and 25,000 single nucleotide variant loci known to be associated with cancer at the higher end of the range.
[0105] VII. PCR method In any of the methods for detecting SNVs herein, including the ctDNA SNV amplification / sequencing workflow, improved amplification parameters for multiplex PCR can be employed. For example, when the amplification reaction is a PCR reaction, the annealing temperature is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10°C higher than the melting temperature of at least 10, 20, 25, 30, 40, 50, 06, 70, 75, 80, 90, 95, or 100% of the primers in the set of primers at the lower end of the range, and 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15°C higher at the upper end of the range.
[0106] In certain embodiments, when the amplification reaction is a PCR reaction, the length of the annealing step in the PCR reaction is 10, 15, 20, 30, 45, and 60 minutes at the lower end of the range, and 15, 20, 30, 45, 60, 120, 180, or 240 minutes at the higher end of the range. In certain embodiments, the primer concentration in the amplification, such as a PCR reaction, is 1 to 10 nM. Furthermore, in exemplary embodiments, the primers in the primer set are designed to minimize primer-dimer formation.
[0107] Thus, in one example of any of the methods herein that include an amplification step, the amplification reaction is a PCR reaction, the annealing temperature is 1-10°C higher than the melting temperature of at least 90% of the primers in the primer set, the length of the annealing step in the PCR reaction is 15-60 minutes, the primer concentration in the amplification reaction is 1-10 nM, and the primers in the primer set are designed to minimize primer-dimer formation. In a further aspect of this example, the multiplex amplification reaction is performed under limiting primer conditions.
[0108] VIII. Use in Cancer Diagnostics In another embodiment, provided herein is a method for confirming a diagnosis of cancer for an individual, such as an individual suspected of having cancer, from a sample of blood from the individual or a fraction thereof, comprising performing a DNA amplification / sequencing workflow provided herein to determine whether one or more single nucleotide variants are present at a plurality of single nucleotide variant loci. In this embodiment, the following elements, statements, guidelines, or rules apply: the absence of a single nucleotide variant supports a diagnosis of Stage 1a, 1b, or 2a adenocarcinoma; the presence of a single nucleotide variant supports a diagnosis of squamous cell carcinoma or Stage 2b or 3a adenocarcinoma; and / or the presence of 10 or more single nucleotide variants supports a diagnosis of squamous cell carcinoma or Stage 2b or 3 adenocarcinoma.
[0109] These results identify analysis of lung ADC and SCC samples from individuals using the ctDNA SNV amplification / sequencing workflow as a valuable method for identifying SNVs found in ADC tumors, particularly for stage 2b and 3a ADC tumors, and particularly for SCC tumors at any stage.
[0110] IX. Use in Prescribing Treatment Regimen In certain embodiments, the methods herein for detecting SNVs can be used to prescribe a treatment regimen. Therapies targeting specific mutations associated with ADC and SCC are available and under development (Nature Review Cancer. 14:535-551 (2014)). For example, detection of EGFR mutations at L858R or T790M can be beneficial for selecting a therapy. Erlotinib, gefitinib, afatinib, AZK9291, CO-1686, and HM61713 are currently approved therapies in the United States and in clinical trials that target specific EGFR mutations. In another example, a G12D, G12C, or G12V mutation in KRAS can be used to prescribe a combination of selumetinib and docetaxel to an individual. In another example, a V600E mutation in BRAF can be used to prescribe a treatment with vemurafenib, dabrafenib, or trametinib to a subject.
[0111] X. Library Preparation In certain embodiments, the methods of the present invention typically include creating and amplifying a nucleic acid library from a sample (i.e., library preparation). Nucleic acids from a sample during the library preparation process can have attached ligation adapters, often referred to as library tags or ligation adapter tags (LT), which contain universal priming sequences followed by universal amplification. In one embodiment, this can be done using standard protocols designed to create sequencing libraries after fragmentation. In one embodiment, the DNA sample can be blunt-ended, and then an A can be added to its 3' end. A Y adapter with a T overhang can be added and ligated. In some embodiments, other sticky ends other than A or T overhangs can be used. In some embodiments, other adapters can be added, such as looped ligation adapters. In some embodiments, the adapters can have tags designed for PCR amplification.
[0112] XI. DNA amplification / sequencing workflow for monitoring or detecting cancer in patients. Some embodiments provided herein include detecting cancer-specific mutations in ctDNA, cfDNA, or cellular DNA samples. Such methods in illustrative embodiments include amplification and sequencing steps (sometimes referred to herein as "ctDNA amplification / sequencing workflows"). In illustrative examples, the DNA amplification / sequencing workflow can include: generating a set of amplicons by performing a multiplex amplification reaction on nucleic acids isolated from a sample of blood or a fraction thereof from an individual, such as an individual suspected of having cancer, e.g., breast cancer, bladder cancer, or colorectal cancer, wherein each amplicon in the set of amplicons spans at least one cancer-associated genomic locus in a set of cancer-associated genomic loci, such as an SNV locus, known to be associated with cancer; and determining the sequence of at least a segment of each amplicon in the set of amplicons, wherein the segment includes the cancer-associated genomic locus. In some embodiments, the cancer-associated genomic locus is identified by a single nucleotide variant. The amplification / sequencing workflow may include single nucleotide variants (SNVs), copy number variations (CNVs), indels, rearranged genes, or variations in exons, introns, gene regulatory sequences, or non-coding RNA sequences. An exemplary DNA amplification / sequencing workflow, more specifically, may include forming an amplification reaction mixture by combining a polymerase, nucleotide triphosphates, and nucleic acid fragments from a nucleic acid library created from the sample with a set of primers that each bind a valid distance from the single nucleotide variant locus or a set of primer pairs that each span a valid region that includes the cancer-associated genomic locus. The amplification reaction mixture may then be subjected to amplification conditions to produce a set of amplicons that include at least one cancer-associated genomic locus from the set of cancer-associated genomic loci, and determining the sequence of at least a segment of each amplicon in the set of amplicons, where the segment includes the cancer-associated genomic locus.
[0113] The effective distance of primer binding can be within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 75, 100, 125, or 150 base pairs of the cancer-associated genomic locus. The effective range spanned by a pair of primers typically includes the cancer-associated genomic locus and is typically 160 base pairs or less, and can be 150, 140, 130, 125, 100, 75, 50, or 25 base pairs or less. In other embodiments, the working range spanned by a pair of primers is 20, 25, 30, 40, 50, 60, 70, 75, 100, 110, 120, 125, 130, 140, or 150 nucleotides at the lower end of the range and 25, 30, 40, 50, 60, 70, 75, 100, 110, 120, 125, 130, 140, or 150, 160, 170, 175, or 200 nucleotides at the higher end of the range from the cancer-associated genomic locus.
[0114] Further details regarding amplification methods that can be used in ctDNA amplification / sequencing workflows to detect cancer-associated genomic loci for use in the methods of the present invention are provided in other sections herein.
[0115] XII. SNV Call Analysis During the methods provided herein, nucleic acid sequencing data is generated for the amplicons generated by tiled multiplex PCR. Algorithm design tools are available that can be used and / or adapted to analyze this data to determine, within certain confidence limits, whether a cancer-associated genomic locus, such as a single nucleotide variant (SNV), is present in a target gene known to be associated with cancer onset, recurrence, metastasis, treatment response, or prognosis.
[0116] Sequencing reads can be demultiplexed using in-house tools and mapped in single-end mode using paired merged reads against the hg19 genome using the BWA mem function of the Burrows-Wheeler alignment software (BWA, Burrows-Wheeler Alignment Software (see Li H. and Durbin R. (2010) Fast and accurate long-read alignment with Burrows-Wheeler Transform. Bioinformatics, Epub. [PMID: 20080505]). Amplification statistics QC can be performed by analyzing total reads, number of mapped reads, number of on-target mapped reads, and number of counted reads.
[0117] In certain embodiments, any analytical method for detecting SNVs from nucleic acid sequencing data detection can be used with the methods of the invention that include detecting an SNV or determining whether an SNV is present. In certain exemplary embodiments, methods of the invention that utilize SNV Method 1 below are used. In yet other further exemplary embodiments, methods of the invention that include detecting an SNV or determining whether an SNV is present at an SNV locus utilize SNV Method 2 below.
[0118] SNV Method 1: For this embodiment, a background error model is constructed using normal plasma samples sequenced in the same sequencing run to account for run-specific artifacts. In certain embodiments, 5, 10, 15, 20, 25, 30, 40, 50, 100, 150, 200, 250, or more than 250 normal plasma samples are analyzed in the same sequencing run. In certain exemplary embodiments, 20, 25, 40, or 50 normal plasma samples are analyzed in the same sequencing run. Noise positions with a median common variant allele frequency greater than a cutoff are removed. For example, in certain embodiments, this cutoff is greater than 0.1%, 0.2%, 0.25%, 0.5%, 1%, 2%, 5%, or 10%. In certain exemplary embodiments, noise positions with a median common variant allele frequency greater than 0.5% are removed. To account for noise and contamination, outlier samples are repeatedly removed from this model. In certain embodiments, samples with a Z-score greater than 5, 6, 7, 8, 9, or 10 are removed from data analysis. For each base substitution at all genomic loci, the read-depth-weighted mean and standard deviation of the errors are calculated. Positions in tumor or cell-free plasma samples with at least five variant reads and a Z-score of 10 relative to the background error model can be called, for example, as candidate mutations.
[0119] SNV Method 2: For this embodiment, single nucleotide variants (SNVs) are determined using plasma ctDNA data. The PCR process is modeled as a stochastic process, and a training set is used to estimate parameters and make final SNV calls for a separate test set. Error propagation over multiple PCR cycles is determined, and the mean and variance of background error are calculated; in an exemplary embodiment, background error is distinguished from actual mutations.
[0120] For each base, the following parameters are estimated: p = efficiency (probability that each read is replicated during each cycle) p e = error rate per cycle for variant e (probability of an error of type e occurring) X0 = initial number of molecules
[0121] As a read is replicated over the course of the PCR process, more errors are introduced. Thus, the error profile of a read is determined by its degree of separation from the original read. A read is said to be of generation k if it has undergone k replications before it is created.
[0122] For each base, let's define the following variables: X ij = number of i-th generation reads produced in PCR cycle j Y ij = total number of reads in generation i at the end of cycle j X ij e = number of i-th generation reads with mutation e generated in PCR cycle j
[0123] Furthermore, in addition to the normal molecule X0, an additional f with mutation e at the start of the PCR process e If X0 molecules are present (thus fe / (1+fe) will be the fraction of mutated molecules in the initial mixture).
[0124] Considering the total number of i-1 generation reads in cycle j-1, the number of i-th generation reads produced in cycle j is determined by the sample size Y i-1,j-1 and has a binomial distribution with probability parameter p. Therefore, E(X ij, |Y i-1,j-1 ,p)=pY i-1,j-1 and Var(X ij, |Y i-1,j-1 ,p)=p(1-p)Y i-1,j-1 is.
[0125] The present inventors
number
[0126] Finally, E(X ij e |Y i-1,j-1 ,p e )=p e Y i-1,j-1 and Var(X ij e |Y i-1,j-1 ,p)=p e (1-p e )Y i-1,j-1 Using these, E(X ij e ) and Var(X ij e ) can be calculated.
[0127] In certain embodiments, SNV method 2 is performed as follows. a) The training dataset is used to estimate PCR efficiency and error rate per cycle. b) Using the distribution of efficiencies estimated in step (a), estimate the number of starting molecules for the test data set at each base. c) If necessary, update the efficiency estimate for the test dataset using the number of starting molecules estimated in step (b). d) Using the test set data and parameters estimated in steps (a), (b), and (c), estimate the mean and variance for the total number of molecules, background error molecules, and actual mutant molecules (for a search space consisting of the initial proportion of actual mutant molecules). e) Fit the distribution to the number of total error molecules (background error and actual mutations) in the total molecules and calculate the likelihood of the proportion of each actual mutation in the search space. f) determining the proportion of most likely actual mutations and calculating the confidence using the data from step (e);
[0128] A confidence cutoff can be used to identify SNVs at SNV loci, for example, SNVs can be called using a confidence cutoff of 90%, 95%, 96%, 97%, 98%, or 99%.
[0129] Exemplary SNV Method 2 Algorithm The algorithm begins by using a training set to estimate the efficiency and error rate per cycle, where n denotes the total number of PCR cycles.
[0130] Read R at each base b b The number of b ) n X0, where p b is the efficiency with base b. Then, (R b / X0) 1 / n Using 1+p b Then, over all training samples, p b The mean and standard deviation of can be determined to estimate the parameters of the probability distribution for each base (e.g., normal, beta, or similar distribution).
[0131] Similarly, the read R of error e at each base b b e Using the number of e After determining the mean and standard deviation of the error rate over all training samples, its probability distribution (e.g., normal, beta, or similar) is estimated, and this mean and standard deviation value is used to estimate its parameters.
[0132] Next, for the test data, the initial starting copies at each base are
number
[0133]
number
[0134] Therefore, we estimated the parameters used in the stochastic process, and then used these estimates to estimate the mean and variance of the molecules created in each cycle (separately for normal, error, and mutant molecules).
[0135] Finally, by using a probabilistic method (e.g., maximum likelihood or similar), the best f that best fits the distribution of errors, mutations, and normal molecules is determined. e More specifically, the inventors have determined the value of various f e For values, estimate the expected ratio of error molecules to total molecules, determine the likelihood of the data for each of these values, and then select the value with the greatest likelihood.
[0136] XIII. Primer Design / Library Preparation Primer tails can improve detection of fragmented DNA from universally tagged libraries. When the library tag and primer tail contain homologous sequences, hybridization can be improved (e.g., lowering the melting temperature (Tm)) and primer extension can occur when only a portion of the primer target sequence is present in the sample DNA primer fragment. In some embodiments, 13 or more target-specific base pairs can be used. In some embodiments, 10-12 target-specific base pairs can be used. In some embodiments, 8-9 target-specific base pairs can be used. In some embodiments, 6-7 target-specific base pairs can be used.
[0137] In one embodiment, a library is generated from a sample by ligating adapters to the ends of DNA fragments in the sample or to the ends of DNA fragments generated from DNA isolated from the sample. The fragments can then be amplified using PCR, for example, according to the following exemplary protocol:
[0138] 95°C for 2 min; 15× [95°C for 20 s, 55°C for 20 s, 68°C for 20 s], 68°C for 2 min, hold at 4°C.
[0139] Many kits and methods are known in the art for creating nucleic acid libraries containing universal primer binding sites for subsequent amplification, e.g., clonal amplification, and subsequent sequencing. To facilitate adapter ligation, library preparation and amplification can include end repair and adenylation (i.e., A-tailing). Kits specifically adapted for preparing libraries from small nucleic acid fragments, particularly circulating free DNA, can be useful for performing the methods provided herein. For example, the NEXTflex Cell Free Kit available from Bioo Scientific or the Natera Library Prep Kit (available from Natera, Inc., San Carlos, CA). However, such kits are typically modified to include customized adapters for the amplification and sequencing steps of the methods provided herein. Adapter ligation can be performed using commercially available kits, such as the ligation kit found in the AGILENT SURESELECT kit (Agilent, CA).
[0140] The target region of a nucleic acid library made from DNA isolated from a sample, particularly a circulating free DNA sample for the method of the present invention, is then amplified. For this amplification, the set of primers or primer pairs can include 5, 10, 15, 20, 25, 50, 100, 125, 150, 250, 500, 1000, 2500, 5000, 10,000, 20,000, 25,000, or 50,000 primers at the lower end of the range, and 15, 20, 25, 50, 100, 125, 150, 250, 500, 1000, 2500, 5000, 10,000, 20,000, 25,000, 50,000, 60,000, 75,000, or 100,000 primers at the upper end of the range, each binding to one of the set of primer binding sites.
[0141] Primer design can be generated using Primer3 (Untergrasser A, Cutcutache I, Koressaar T, Ye J, Faircloth BC, Remm M, Rozen SG (2012) "Primer3 - new capabilities and interfaces." Nucleic Acids Research 40(15):e115 and Koressaar T, Remm M (2007) "Enhancements and modifications of primer design program Primer3." Bioinformatics 23(10):1289-91. Source code is available at primer3.sourceforge.net). Primer specificity can be assessed by BLAST and added to existing primer design pipeline criteria.
[0142] Primer specificity can be determined using the BLASTn program from the ncbi-blast-2.2.29+ package. The task option "blastn-short" can be used to map primers to the hg19 human genome. A primer design can be determined to be "specific" if the primer has fewer than 100 hits to the genome, and the top hit is the target-complementary primer-binding region of that genome and scores at least 2 points higher than the other hits (scores are defined by the BLASTn program). This can be done to have a unique hit to the genome and not many other hits throughout the genome.
[0143] The final selected primers can be visualized in IGV (James T. Robinson, Helga Thorvaldsdottir, Wendy Winckler, Mitchell Guttman, Eric S. Lander, Gad Getz, Jill P. Mesirov. Integrative Genomics Viewer. Nature Biotechnology 29, 24-26 (2011)) and the UCSC browser (Kent WJ, Sugnet CW, Furey TS, Roskin KM, Pringle TH, Zahler AM, Haussler D. The human genome browser at UCSC. Genome Res. 2002 Jun;12(6):996-1006) using bed files and coverage maps for validation.
[0144] XIV.PCR reaction mixture In certain embodiments, the methods of the present invention include forming an amplification reaction mixture. The reaction mixture is typically generated by combining a polymerase, nucleotide triphosphates, nucleic acid fragments from a nucleic acid library generated from a sample, and a set of forward and reverse primers specific to a target region containing an SNV. The reaction mixtures provided herein, in exemplary embodiments, themselves form a separate aspect of the present invention.
[0145] Amplification reaction mixtures useful in the present invention contain components known in the art for nucleic acid amplification, particularly PCR amplification. For example, reaction mixtures typically contain nucleotide triphosphates, polymerase, and magnesium. Polymerases useful in the present invention can include any polymerase that can be used in amplification reactions, particularly those useful in PCR reactions. In certain embodiments, hot-start Taq polymerase is particularly useful. Amplification reaction mixtures useful for performing the methods provided herein, such as AmpliTaq Gold Master Mix (Life Technologies, Carlsbad, CA), are commercially available.
[0146] PCR amplification (e.g., temperature cycling) conditions are well known in the art. The methods provided herein can include any PCR cycling conditions that result in the amplification of target nucleic acids, such as target nucleic acids from a library. Non-limiting exemplary cycling conditions are provided in the Examples section of this specification.
[0147] There are many possible workflows when performing PCR, and several workflows typical of the methods disclosed herein are provided herein. The steps outlined herein are not meant to exclude other possible steps, nor do they imply that any of the steps described herein are necessary for the method to function properly. Variations of many parameters or other modifications are known in the literature and can be made without affecting the essence of the invention.
[0148] In certain embodiments of the methods provided herein, at least a portion, and in illustrative examples, the entire sequence of an amplicon, such as an outer primer-targeted amplicon, is determined. Methods for determining the sequence of an amplicon are known in the art. Any of the sequencing methods known in the art, such as Sanger sequencing, can be used to determine such a sequence. In illustrative embodiments, high-throughput next-generation sequencing technologies (also referred to herein as massively parallel sequencing technologies), such as, but not limited to, those employed in MYSEQ (ILLUMINA), HISEQ (ILLUMINA), ION TORRENT (LIFE TECHNOLOGIES), GENOME ANALYZER ILX (ILLUMINA), and GS FLEX+ (ROCHE 454), can be used to sequence the amplicons generated by the methods provided herein.
[0149] High-throughput gene sequencers can be modified to accommodate the use of barcoding (i.e., sample tagging with a distinctive nucleic acid sequence) to identify specific samples from individuals, thereby enabling simultaneous analysis of multiple samples in a single DNA sequencer run. The number of times a given region of the genome is sequenced (number of reads) in a library preparation (or other nucleic acid preparation of interest) will be proportional to the copy number (or expression level, in the case of a cDNA-containing preparation) of that sequence in the genome of interest. Bias in amplification efficiency can be taken into account in such quantitative determinations.
[0150] In certain embodiments, the methods of the present invention include forming an amplification reaction mixture. This reaction mixture is typically formed by combining a polymerase, nucleotide triphosphates, nucleic acid fragments from a nucleic acid library created from a sample with a set of forward target-specific outer primers and a first-strand reverse outer universal primer. Another illustrative embodiment is a reaction mixture that includes a forward target-specific inner primer instead of the forward target-specific outer primer, and amplicons from a first PCR reaction using the outer primer instead of the nucleic acid fragments from the nucleic acid library. In illustrative embodiments, the reaction mixture provided herein itself forms a separate aspect of the present invention. In an illustrative embodiment, the reaction mixture is a PCR reaction mixture. The PCR reaction mixture typically includes magnesium.
[0151] In some embodiments, the reaction mixture contains ethylenediaminetetraacetic acid (EDTA), magnesium, tetramethylammonium chloride (TMAC), or any combination thereof. In some embodiments, the concentration of TMAC is 20-70 mM, inclusive. Without being bound by any particular theory, it is believed that TMAC binds to DNA, stabilizes the duplex, increases primer specificity, and / or equalizes the melting temperatures of different primers. In some embodiments, TMAC increases the uniformity of the amount of amplified product for different targets. In some embodiments, the concentration of magnesium (e.g., from magnesium chloride) is 1-8 mM.
[0152] The large number of primers used in multiplex PCR of multiple targets can chelate a lot of magnesium (two phosphate groups in a primer chelate one magnesium). For example, if enough primers are used so that the concentration of primer-derived phosphate groups is about 9 mM, the primers can reduce the effective magnesium concentration to about 4.5 mM. In some embodiments, because high concentrations of magnesium can lead to PCR errors, such as amplification of non-target loci, EDTA is used to reduce the amount of magnesium available as a polymerase cofactor. In some embodiments, the concentration of EDTA reduces the amount of available magnesium to 1-5 mM (e.g., 3-5 mM).
[0153] In some embodiments, the pH is 7.5-8.5, e.g., 7.5-8, 8-8.3, or 8.3-8.5 (inclusive). In some embodiments, Tris is used at a concentration of, e.g., 10-100 mM, e.g., 10-25 mM, 25-50 mM, 50-75 mM, or 25-75 mM (inclusive). In some embodiments, any of these concentrations of Tris is used at a pH of 7.5-8.5. In some embodiments, a combination of KCl and (NH)SO is used, such as 50-150 mM KCl and 10-90 mM (NH)SO (inclusive). In some embodiments, the concentration of KCl is 0-30 mM, 50-100 mM, or 100-150 mM (inclusive). In some embodiments, the concentration of (NH4)2SO4 is 10-50 mM, 50-90 mM, 10-20 mM, 20-40 mM, 40-60 mM, or 60-80 mM (NH4)2SO4, inclusive. + ] concentration is 0 to 160 mM, for example, 0 to 50, 50 to 100, or 100 to 160 mM (boundaries included). In some embodiments, the sum of the potassium concentration and the ammonium concentration ([K + ]+[NH4 +]) is 0 to 160 mM, for example, 0 to 25, 25 to 50, 50 to 150, 50 to 75, 75 to 100, 100 to 125, or 125 to 160 mM (boundary values included). + ]+[NH4 + An exemplary buffer with a pH of 0.05 is 20 mM KCl and 50 mM (NH4)2SO4. In some embodiments, the buffer comprises 25-75 mM Tris (pH 7.2-8), 0-50 mM KCl, 10-80 mM ammonium sulfate, and 3-6 mM magnesium (inclusive). In some embodiments, the buffer comprises 25-75 mM Tris (pH 7-8.5), 3-6 mM MgCl2, 10-50 mM KCl, and 20-80 mM (NH4)2SO4 (inclusive). In some embodiments, 100-200 units / mL of polymerase are used. In some embodiments, 7 ul of DNA template in a final volume of 20 ul in 100 mM KCl, 50 mM (NH4)2SO4, 3 mM MgCl2, 7.5 nM of each primer in the library, 50 mM TMAC, and pH 8.1 is used.
[0154] In some embodiments, a crowding agent such as polyethylene glycol (PEG, such as PEG 8,000) or glycerol is used. In some embodiments, the amount of PEG (e.g., PEG 8,000) is 0.1-20%, e.g., 0.5-15%, 1-10%, 2-8%, or 4-8%, inclusive. In some embodiments, the amount of glycerol is 0.1-20%, e.g., 0.5-15%, 1-10%, 2-8%, or 4-8%, inclusive. In some embodiments, the crowding agent allows for the use of either a lower polymerase concentration and / or a shorter annealing time. In some embodiments, the crowding agent improves the uniformity of DOR and / or reduces dropouts (undetected alleles). Polymerases: In some embodiments, polymerases with proofreading activity, polymerases without (or with negligible) proofreading activity, or mixtures of polymerases with and without (or with negligible) proofreading activity are used. In some embodiments, hot-start polymerases, non-hot-start polymerases, or mixtures of hot-start and non-hot-start polymerases are used. In some embodiments, HotStarTaq DNA polymerase is used (see, e.g., QIAGEN catalog number 203203). In some embodiments, AmpliTaq Gold® DNA polymerase is used. In some embodiments, PrimeSTAR GXL DNA polymerase is used (Takara Clontech, Mountain View, CA), a high-fidelity polymerase that provides efficient PCR amplification when excess template is present in the reaction mixture and when amplifying long products. In some embodiments, KAPA Taq DNA polymerase or KAPA Taq HotStart DNA polymerase is used. These are based on the single subunit wild-type Taq DNA polymerase of the thermophilic bacterium Thermus aquaticus.KAPA Taq and KAPA Taq HotStart DNA polymerases have 5'-3' polymerase activity and 5'-3' exonuclease activity, but not 3'-5' exonuclease (proofreading) activity (see, e.g., KAPA BIOSYSTEMS catalog number BK1000). In some embodiments, Pfu DNA polymerase is used. This polymerase is a thermostable DNA polymerase derived from the hyperthermophilic archaeon Pyrococcus furiosus. This enzyme catalyzes the template-dependent polymerization of nucleotides into double-stranded DNA in the 5'->3' direction. Pfu DNA polymerase also exhibits 3'->5' exonuclease (proofreading) activity, allowing the polymerase to correct nucleotide incorporation errors. This polymerase does not have 5'->3' exonuclease activity (see, e.g., Thermo Scientific catalog number EP0501). In some embodiments, Klentaq1 is used. It is a Klenow fragment analog of Taq DNA polymerase and does not have exonuclease or endonuclease activity (see, e.g., DNA POLYMERASE TECHNOLOGY, Inc., St. Louis, Missouri, catalog number 100). In some embodiments, the polymerase is a PHUSION DNA polymerase, such as PHUSION High Fidelity DNA Polymerase (M0530S, New England BioLabs, Inc.) or PHUSION Hot Start Flex DNA Polymerase (M0535S, New England BioLabs, Inc.). In some embodiments, the polymerase is a Q5® DNA polymerase, such as Q5® High-Fidelity DNA Polymerase (M0491S, New England BioLabs, Inc.) or Q5® Hot Start High-Fidelity DNA Polymerase (M0493S, New England BioLabs, Inc.).In some embodiments, the polymerase is T4 DNA polymerase (M0203S, New England BioLabs, Inc.).
[0155] In some embodiments, 5 to 600 units / mL (units per mL of reaction volume), for example, 5 to 100, 100 to 200, 200 to 300, 300 to 400, 400 to 500, or 500 to 600 units / mL (boundary values included), of polymerase is used.
[0156] XV. PCR method In some embodiments, hot-start PCR is used to reduce or prevent polymerization prior to PCR thermal cycling. Exemplary hot-start PCR methods include initial inhibition of the DNA polymerase or physical separation of reaction components until the reaction mixture reaches a higher temperature. In some embodiments, delayed release of magnesium is used. Because DNA polymerase requires magnesium ions for activity, magnesium is chemically separated from the reaction by binding to a chemical compound and released into solution only at elevated temperatures. In some embodiments, non-covalent binding of an inhibitor is used. In this method, a peptide, antibody, or aptamer non-covalently binds to the enzyme at low temperatures and inhibits its activity. After incubation at elevated temperatures, the inhibitor is released and the reaction begins. In some embodiments, cold-sensitive Taq polymerase, such as a modified DNA polymerase that has little activity at low temperatures, is used. In some embodiments, chemical modification is used. In this method, a molecule is covalently attached to the side chain of an amino acid in the active site of the DNA polymerase. The molecule is released from the enzyme by incubating the reaction mixture at elevated temperatures. Once the molecule is released, the enzyme is activated.
[0157] In some embodiments, the amount of nucleic acid (e.g., RNA or DNA sample) for template assembly is between 20 and 5,000 ng, e.g., between 20 and 200, 200 and 400, 400 and 600, 600 and 1,000, 1,000 and 1,500, or 2,000 and 3,000 ng, inclusive.
[0158] In some embodiments, the QIAGEN Multiplex PCR Kit is used (QIAGEN catalog number 206143). For a 100 x 50 μl multiplex PCR reaction, the kit contains 2 x QIAGEN Multiplex PCR Master Mix (3 x 0.85 ml, providing a final concentration of 3 mM MgCl), 5 x Q-Solution (1 x 2.0 ml), and RNase-free water (2 x 1.7 ml). QIAGEN Multiplex PCR Master Mix (MM) contains a combination of KCl and (NH4)2SO4, as well as the PCR additive Factor MP, which increases the local concentration of primers at the template. Factor MP stabilizes specifically bound primers, allowing efficient primer extension by HotStarTaq DNA Polymerase. HotStarTaq DNA Polymerase is a modified form of Taq DNA polymerase that has no polymerase activity at ambient temperature. In some embodiments, HotStarTaq DNA Polymerase is activated by incubation at 95° C. for 15 minutes, which can be incorporated into any existing thermal cycler program.
[0159] In some embodiments, a final concentration of 1×QIAGEN MM (recommended concentration), 7.5 nM of each primer in the library, 50 mM TMAC, and 7 ul of DNA template in a final volume of 20 ul are used. In some embodiments, PCR thermocycling conditions include 95° C. for 10 minutes (hot start), 20 cycles of 96° C. for 30 seconds, 65° C. for 15 minutes, 72° C. for 30 seconds, followed by 72° C. for 2 minutes (final extension), then a 4° C. hold.
[0160] In some embodiments, a final concentration of 2x QIAGEN MM (twice the recommended concentration), 2 nM of each primer in the library, 70 mM TMAC, and 7 ul of DNA template in a total volume of 20 ul are used. In some embodiments, up to 4 mM EDTA is also included. In some embodiments, PCR thermocycling conditions include 95°C for 10 minutes (hot start), 96°C for 30 seconds, 65°C for 20, 25, 30, 45, 60, 120, or 180 minutes, optionally 25 cycles of 72°C for 30 seconds, followed by 72°C for 2 minutes (final extension), then a 4°C hold.
[0161] Another exemplary set of conditions involves a semi-nested PCR procedure. The first PCR reaction uses a 20 μl reaction volume containing a 2× QIAGEN MM final concentration, 1.875 nM of each primer (outer forward and reverse primers) in the library, and DNA template. Thermal cycling parameters include 95°C for 10 minutes, 96°C for 30 seconds, 65°C for 1 minute, 58°C for 6 minutes, 60°C for 8 minutes, 25 cycles of 65°C for 4 minutes and 72°C for 30 seconds, followed by 72°C for 2 minutes, followed by a 4°C hold. 2 μl of the resulting product, diluted 1:200, is then used as input for the second PCR reaction. This reaction uses a 10 μl reaction volume containing a 1× QIAGEN MM final concentration, 20 nM of each inner forward primer, and 1 μM of the reverse primer tag. Thermal cycling parameters include 95° C. for 10 minutes, 95° C. for 30 seconds, 65° C. for 1 minute, 60° C. for 5 minutes, 15 cycles of 65° C. for 5 minutes and 72° C. for 30 seconds, then 72° C. for 2 minutes, then a hold at 4° C. The annealing temperature can optionally be higher than the melting temperature of some or all of the primers, as discussed herein (see U.S. Patent Application No. 14 / 918,544, filed October 20, 2015, which is incorporated by reference in its entirety).
[0162] Melting point (T mThe annealing temperature (T) is the temperature at which half (50%) of the DNA duplex of an oligonucleotide (e.g., a primer) and its perfect complement dissociates into single-stranded DNA. A ) is the temperature at which the PCR protocol is run. For conventional methods, this temperature is usually the lowest T of the primers used. m At temperatures 5°C below 100°C, nearly all possible duplexes are formed (so that virtually all primer molecules bind to the template nucleic acid). Although this is highly efficient, lower temperatures are sure to result in more nonspecific reactions. A One consequence of having T that is too low is that internal single-base mismatches or partial annealing may be tolerated, allowing the primer to anneal to sequences other than the true target. A T m At higher temperatures, at a given moment, only a small fraction of the target has primers annealed (e.g., only about 1-5%). As these are extended, they are removed from the equilibrium of primer and target annealing and dissociation (T) until extension exceeds 70°C. m By allowing the reaction to anneal for a long time, approximately 1-5% of the targets will be copied per cycle (to rapidly increase the number of copies).
[0163] In various embodiments, the annealing temperature is at least 25, 50, 60, 70, 75, 80, 90, 95, or 100% of the melting temperatures (empirically measured or calculated T mIn various embodiments, the annealing temperature is at least 25, 50, 75, 100, 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000, 50,000, 75,000, 100,000, or all of the melting temperatures (empirically measured or calculated T m In various embodiments, the annealing temperature is 1 to 15°C (e.g., 1 to 10, 1 to 5, 1 to 3, 3 to 5, 5 to 10, 5 to 8, 8 to 10, 10 to 12, or 12 to 15°C, inclusive) higher than the melting points (empirically measured or calculated T ) of at least 25%, 50%, 60%, 70%, 75%, 80%, 90%, 95% or all of the non-identical primers. m The annealing step is performed at a temperature of 1 to 15°C (for example, 1 to 10, 1 to 5, 1 to 3, 3 to 5, 3 to 8, 5 to 10, 5 to 8, 8 to 10, 10 to 12, or 12 to 15°C (inclusive)) higher than the initial temperature, and the length of the annealing step (per PCR cycle) is 5 to 180 minutes, for example, 15 to 120 minutes, 15 to 60 minutes, 15 to 45 minutes, or 20 to 60 minutes (inclusive).
[0164] XVI. Exemplary Multiplex PCR Methods In various embodiments, long annealing times and / or low primer concentrations are used. Indeed, in certain embodiments, limited primer concentrations and / or conditions are used. In various embodiments, the length of the annealing step ranges from 15, 20, 25, 30, 35, 40, 45, or 60 minutes at the lower end of the range to 20, 25, 30, 35, 40, 45, 60, 120, or 180 minutes at the higher end of the range. In various embodiments, the length of the annealing step (per PCR cycle) is 30 to 180 minutes. For example, the annealing step can be 30 to 60 minutes, and the concentration of each primer can be less than 20, 15, 10, or 5 nM. In other embodiments, primer concentrations are from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 25 nM at the lower end of the range to 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, and 50 at the higher end of the range.
[0165] At high levels of multiplexing, the solution may become viscous due to the large amount of primers in the solution. If the solution is too viscous, the primer concentration can be reduced to a level that is still sufficient for the primers to bind to the template DNA. In various embodiments, 1,000 to 100,000 different primers are used, and the concentration of each primer is less than 20 nM, for example, less than 10 nM or 1 to 10 nM (boundary values included).
[0166] XVII. Copy Number Variation (CNV) Detection In addition to SNVs and indels, the methods for monitoring and detecting early recurrence and metastasis described herein can also benefit from the detection of CNVs.
[0167] In one aspect, the present invention generally relates, at least in part, to improved methods for determining the presence or absence of copy number variations, such as deletions or duplications of chromosomal segments or entire chromosomes. The methods are particularly useful for detecting small deletions or duplications that can be difficult to detect with high specificity and sensitivity using conventional methods due to the limited data available from the relevant chromosomal segments. The methods include improved analytical methods, improved bioassay methods, and combinations of improved analytical and bioassay methods. The methods of the present invention can also be used to detect deletions or duplications that are present in only a small percentage of cells or nucleic acid molecules tested. This allows for detection of deletions or duplications before the onset of disease (e.g., precancerous conditions) or early in the disease, e.g., before a large number of diseased cells (e.g., cancer cells) with deletions or duplications have accumulated. More accurate detection of deletions or duplications associated with a disease or disorder allows for improved methods for diagnosing, predicting, preventing, delaying, stabilizing, or treating the disease or disorder. Some deletions or duplications are known to be associated with cancer or severe intellectual or physical disabilities.
[0168] XVIII. SNV detection In another aspect, the present invention generally relates, at least in part, to improved methods for detecting single nucleotide variations (SNVs). These improved methods include improved analytical methods, improved bioassay methods, and improved methods using a combination of improved analytical and bioassay methods. In certain exemplary embodiments, the methods are used to detect, diagnose, monitor, or stage cancer in samples, e.g., circulating free DNA samples, in which SNVs are present at very low concentrations, e.g., less than 10%, 5%, 4%, 3%, 2.5%, 2%, 1%, 0.5%, 0.25%, or 0.1% of the total number of normal copies of the SNV locus. That is, in certain exemplary embodiments, these methods are particularly well suited to samples in which a relatively low proportion of mutations or variants are present relative to the normal polymorphic alleles present for that locus. Finally, methods are provided herein that combine improved methods for detecting copy number variations with improved methods for detecting single nucleotide variations.
[0169] Successful treatment of diseases such as cancer often depends on early diagnosis, accurate disease stage determination, selection of effective treatment regimens, and close monitoring to prevent or detect recurrence. For cancer diagnosis, histological evaluation of tumor material obtained from tissue biopsy is often considered the most reliable method. However, the invasive nature of biopsy-based sampling makes it impractical for mass screening and regular follow-up. Therefore, the present method has the advantage of being relatively low-cost and can be performed non-invasively when a fast turnaround time is desired. Targeted sequencing, which can be used by the method of the present invention, requires fewer reads than shotgun sequencing, such as a few hundred million reads instead of 40 million reads, thereby reducing costs. Multiplex PCR and next-generation sequencing, which can be used, increase throughput and reduce costs.
[0170] In some exemplary embodiments, analysis of AAI patterns in ctDNA provides more detailed insight into the clonal architecture of tumors, helping to predict their therapeutic response and optimize treatment strategies. Thus, in certain embodiments, mmPCR-NGS panels targeting clinically causative CNVs and SNVs are selected. In certain exemplary embodiments, such panels are particularly useful for patients with cancers in which CNVs represent a substantial proportion of the mutation burden, as is common in breast, ovarian, and lung cancers.
[0171] In some embodiments, the method is used to detect deletions, duplications, or single nucleotide variants in an individual. A sample from an individual containing cells or nucleic acids suspected of having a deletion, duplication, or single nucleotide variant can be analyzed. In some embodiments, the sample is derived from a tissue or organ suspected of having a deletion, duplication, or single nucleotide variant, such as a cell or mass suspected of being cancerous. The method of the present invention can be used to detect deletions, duplications, or single nucleotide variants present in only one cell or a small number of cells in a mixture containing cells with a deletion, duplication, or single nucleotide variant and cells lacking the deletion, duplication, or single nucleotide variant. In some embodiments, cfDNA or cfRNA from a blood sample from an individual is analyzed. In some embodiments, cfDNA or cfRNA is secreted by cells, such as cancer cells. In some embodiments, cfDNA or cfRNA is released by cells undergoing necrosis or apoptosis, such as cancer cells. The method of the present invention can be used to detect deletions, duplications, or single nucleotide variants present in only a small percentage of cfDNA or cfRNA. In some embodiments, one or more cells from an embryo are tested.
[0172] In addition to determining the presence or absence of copy number variations, one or more other factors can be analyzed if desired. These factors can be used to improve the accuracy of diagnosis (such as determining the presence or absence of cancer or an increased risk of cancer, classifying cancer, or determining the stage of cancer) or prognosis. These factors can also be used to select a specific therapy or treatment regimen that is likely to be effective in a subject. Exemplary factors include the presence or absence of polymorphisms or mutations, changes (increases or decreases) in the levels of total or specific cfDNA, cfRNA, or microRNA (miRNA), changes (increases or decreases) in tumor fraction, changes (increases or decreases) in methylation levels, changes (increases or decreases) in DNA integrity, and altered (increases or decreases) or alternative mRNA splicing.
[0173] The following sections describe methods for detecting deletions or duplications using phasing data (such as inferred or measured phasing data) or non-phasing data, samples that can be tested, methods for sample preparation, amplification, and quantification, methods for phasing genetic data, polymorphisms, mutations, nucleic acid changes, mRNA splicing, and changes in nucleic acid levels that can be detected, databases with results from the methods, other risk factors, and screening methods, cancers that can be diagnosed or treated, cancer treatments, cancer models for testing treatments, and methods for prescribing and administering treatments.
[0174] XIX. ILLUSTRATIVE EMBODIMENTS A. Exemplary Method for Determining Ploidy Using Phasing Data Some of the methods of the present invention are based, in part, on the discovery that using phased data to detect CNVs reduces the false-negative and false-positive rates compared to using non-phased data. This improvement is greatest for samples with CNVs present at low levels. Thus, phased data improves the accuracy of CNV detection compared to using non-phased data (methods that calculate allele ratios at one or more loci or aggregate allele ratios to determine an aggregate value (e.g., average) across a chromosome or chromosome segment without considering whether the allele ratios at different loci indicate that the same or different haplotypes appear to be present in unusual amounts). The use of phased data allows for more accurate determination of whether differences between measured and predicted allele ratios are due to noise or the presence of CNVs. For example, if differences between measured and predicted allele ratios at most or all loci within a region indicate that the same sample haplotype is overrepresented, then a CNV is likely present. Using the association between alleles in haplotypes allows for determination of whether the measured genetic data corresponds to the same overrepresented haplotype (rather than random noise). In contrast, if the difference between the measured and predicted allele ratios is due solely to noise (such as experimental error), in some embodiments, the first haplotype will appear over-represented about half the time and the second haplotype will appear over-represented about half the time.
[0175] In some embodiments, the phasing gene data is used to determine whether there is an overrepresentation of copies of a first homologous chromosomal segment compared to a second homologous chromosomal segment in an individual's genome (e.g., in the genome of one or more cells, or in cfDNA or cfRNA). Exemplary overrepresentation includes a duplication of the first homologous chromosomal segment or a deletion of the second homologous chromosomal segment. In some embodiments, there is no overrepresentation because the first chromosomal segment and the homologous chromosomal segment are present in equal proportions (e.g., one copy of each segment in a diploid sample). In some embodiments, the calculated allele ratio in the nucleic acid sample is compared to the predicted allele ratio to determine whether there is overrepresentation, as described further below. As used herein, the phrase "first homologous chromosomal segment compared to a second homologous chromosomal segment" refers to the first homolog of the chromosomal segment and the second homolog of the chromosomal segment.
[0176] In some embodiments, the method includes: obtaining phased genetic data for a first homologous chromosome segment, the phased genetic data including, for each locus in a set of polymorphic loci on a first homologous chromosome segment, the identity of the allele present at the locus on the first homologous chromosome segment; obtaining phased genetic data for a second homologous chromosome segment, the identity of the allele present at the locus on the second homologous chromosome segment, the identity of the allele present at the locus on the second homologous chromosome segment; and, for each allele at each of the loci in the set of polymorphic loci, obtaining measured genetic allele data including the amount of each allele present in a sample of DNA or RNA from one or more target cells and one or more non-target cells from the individual. In some embodiments, the method includes enumerating a set of one or more hypotheses indicating the degree of overrepresentation of the first homologous chromosomal segment; for each of the hypotheses, calculating predicted genetic data for a plurality of loci in the sample from phased genetic data obtained for one or more possible ratios of DNA or RNA from one or more target cells to total DNA or RNA in the sample; for each possible DNA or RNA ratio and for each hypothesis, calculating (e.g., computing on a computer) a data fitting between the obtained genetic data of the sample and the predicted genetic data for the sample for that possible DNA or RNA ratio and for that hypothesis; ranking one or more of the hypotheses according to the data fitting; and selecting the highest ranked hypothesis, thereby determining the degree of overrepresentation of the copy number of the first homologous chromosomal segment in the genome of one or more cells from the individual.
[0177] In some embodiments, the method involves obtaining phasing genetic data using any of the methods described herein or any known method. In some embodiments, the method involves simultaneously, or sequentially in any order, (i) obtaining phasing genetic data for a first homologous chromosome segment, the data including, for each locus in a set of polymorphic loci on the first homologous chromosome segment, the identity of the allele present at the locus on the first homologous chromosome segment; (ii) obtaining phasing genetic data for a second homologous chromosome segment, the identity of the allele present at the locus on the second homologous chromosome segment, the identity of the allele present at the locus on the second homologous chromosome segment, and (iii) obtaining measured gene allele data including the amount of each allele at each locus in the set of polymorphic loci in a sample of DNA from one or more cells from an individual.
[0178] In some embodiments, the method involves calculating an allele ratio for one or more loci in a set of polymorphic loci that are heterozygous in at least one cell from which the sample was derived. In some embodiments, the calculated allele ratio for a particular locus is the measured amount of one of the alleles divided by the total measured amount of all alleles for that locus. In some embodiments, the calculated allele ratio for a particular locus is the measured amount of one of the alleles (such as an allele on a first homologous chromosome segment) divided by the measured amount of one or more other alleles for that locus (such as an allele on a second homologous chromosome segment). The calculated allele ratio can be calculated using any of the methods described herein or any standard method (such as any mathematical transformation of the calculated allele ratios described herein).
[0179] In some embodiments, the method involves determining whether there is an overrepresentation of a copy number of a first homologous chromosomal segment by comparing a calculated allele ratio for one or more loci to an expected allele ratio for the locus when the first and second homologous chromosomal segments are present in equal proportions. In some embodiments, the expected allele ratio assumes that possible alleles for a locus are equally likely to be present. When the calculated allele ratio for a particular locus is the measured amount of one allele divided by the total measured amount of all alleles for that locus, in some embodiments, the corresponding expected allele ratio is 0.5 for a biallelic locus or 1 / 3 for a triallelic locus. In some embodiments, the expected allele ratio is the same for all loci, e.g., 0.5 for all loci. In some embodiments, the expected allele ratio assumes that possible alleles for a locus may have different likelihoods of being present, e.g., based on the frequency of each allele within a particular set to which the subject belongs, such as a set based on the subject's ancestry. Such allele frequencies are publicly available (see, e.g., HapMap Project; Perlegen Human Haplotype Project; web ncbi.nlm.nih.gov / projects / SNP / ; Sherry ST, Ward MH, Kholodov M, et al. dbSNP: the NCBI database of genetic variation. Nucleic Acids Res. 2001 Jan 1;29(1):308-11, each of which is incorporated herein by reference in its entirety). In some embodiments, the predicted allele ratio is the allele ratio predicted for a particular individual tested for a particular hypothesis indicating the degree of overrepresentation of the first homologous chromosomal segment.For example, a predicted allele ratio for a particular individual can be determined based on phased or non-phased genetic data from that individual (e.g., from a sample from an individual unlikely to have the deletion or duplication, such as a non-cancerous sample), or data from one or more relatives of that individual.
[0180] In some embodiments, a calculated allele ratio is indicative of overrepresentation of copies of a first homologous chromosomal segment if either (i) the allele ratio for the measured amount of an allele present at a locus on a first homologous chromosomal segment divided by the total measured amount of all alleles for that locus is greater than the predicted allele ratio for that locus, or (ii) the allele ratio for the measured amount of an allele present at a locus on a second homologous chromosomal segment divided by the total measured amount of all alleles for that locus is less than the predicted allele ratio for that locus. In some embodiments, a calculated allele ratio is considered to be indicative of overrepresentation only if it is significantly greater or less than the predicted ratio for that locus. In some embodiments, a calculated allele ratio is indicative of the absence of overrepresentation of copies of a first homologous chromosomal segment if (i) the allele ratio for the measured amount of an allele present at a locus on a first homologous chromosomal segment, divided by the total measured amount of all alleles for that locus, is less than or equal to the predicted allele ratio for that locus, or (ii) the allele ratio for the measured amount of an allele present at a locus on a second homologous chromosomal segment, divided by the total measured amount of all alleles for that locus, is greater than or equal to the predicted allele ratio for that locus. In some embodiments, calculated ratios that are equal to the predicted corresponding ratios are ignored (as they are indicative of the absence of overrepresentation).
[0181] In various embodiments, one or more of the calculated allele ratios are compared to the corresponding predicted allele ratio(s) using one or more of the following methods: In some embodiments, it is determined whether the calculated allele ratio is above or below the predicted allele ratio for a particular locus, regardless of the magnitude of the difference; In some embodiments, it is determined the magnitude of the difference between the calculated allele ratio and the predicted allele ratio for a particular locus, regardless of whether the calculated allele ratio is above or below the predicted allele ratio; In some embodiments, it is determined whether the calculated allele ratio is above or below the predicted allele ratio and the magnitude of the difference for a particular locus; In some embodiments, it is determined whether the average or weighted average of the calculated allele ratios is above or below the average or weighted average of the predicted allele ratios, regardless of the magnitude of the difference. In some embodiments, the magnitude of the difference between the average or weighted average calculated allele rations and the average or weighted average predicted allele rations is determined, regardless of whether the average or weighted average calculated allele rations is above or below the average or weighted average predicted allele rations. In some embodiments, the magnitude of the difference between the average or weighted average calculated allele rations and the average or weighted average predicted allele rations is determined, regardless of whether the average or weighted average calculated allele rations is above or below the average or weighted average predicted allele rations. In some embodiments, the magnitude of the difference between the calculated allele rations and the predicted allele rations is determined.
[0182] In some embodiments, the magnitude of the difference between the calculated allele ratio and the expected allele ratio for one or more loci is used to determine whether an overrepresentation of a copy number of a first homologous chromosomal segment is due to a duplication of the first homologous chromosomal segment or a deletion of a second homologous chromosomal segment in the genome of one or more cells.
[0183] In some embodiments, copy number overrepresentation of the first homologous chromosomal segment is determined to be present if one or more of the following conditions are met: In some embodiments, the numerical value of the calculated allele ratio indicative of copy number overrepresentation of the first homologous chromosomal segment is above a threshold; In some embodiments, the numerical value of the calculated allele ratio indicative of no copy number overrepresentation of the first homologous chromosomal segment is below a threshold; In some embodiments, the magnitude of the difference between the calculated allele ratio indicative of copy number overrepresentation of the first homologous chromosomal segment and the corresponding predicted allele ratio is above a threshold; In some embodiments, the sum of the magnitudes of the differences between the calculated allele ratio and the corresponding predicted allele ratio for all calculated allele ratios indicative of overrepresentation is above a threshold; In some embodiments, the magnitude of the difference between the calculated allele ratio indicative of no copy number overrepresentation of the first homologous chromosomal segment and the corresponding predicted allele ratio is below a threshold. In some embodiments, the average or weighted average of the calculated allele ratios for the measured amounts of alleles present on the first homologous chromosomal segment, divided by the total measured amounts of all alleles for that locus, is greater than the average or weighted average of the predicted allele ratios by at least one threshold. In some embodiments, the average or weighted average of the calculated allele ratios for the measured amounts of alleles present on the second homologous chromosomal segment, divided by the total measured amounts of all alleles for that locus, is less than the average or weighted average of the predicted allele ratios by at least one threshold. In some embodiments, the data fit between the calculated allele ratios and the allele ratios predicted for copy number overrepresentation of the first homologous chromosomal segment is below a threshold (indicating good data fit). In some embodiments, the data fit between the calculated allele ratios and the allele ratios predicted for no copy number overrepresentation of the first homologous chromosomal segment is above a threshold (indicating poor data fit).
[0184] In some embodiments, copy number overrepresentation of the first homologous chromosomal segment is determined to be absent if one or more of the following conditions are met: In some embodiments, the numerical value of the calculated allele ratio indicative of copy number overrepresentation of the first homologous chromosomal segment is below a threshold; In some embodiments, the numerical value of the calculated allele ratio indicative of copy number overrepresentation of the first homologous chromosomal segment is above a threshold; In some embodiments, the magnitude of the difference between the calculated allele ratio indicative of copy number overrepresentation of the first homologous chromosomal segment and the predicted corresponding allele ratio is below a threshold; In some embodiments, the magnitude of the difference between the calculated allele ratio indicative of copy number overrepresentation of the first homologous chromosomal segment and the predicted corresponding allele ratio is above a threshold. In some embodiments, the average or weighted average of the calculated allele ratios for the measured amounts of alleles present on the first homologous chromosomal segment, divided by the total measured amounts of all alleles for that locus, minus the average or weighted average of the predicted allele ratios, is below a threshold. In some embodiments, the average or weighted average of the predicted allele ratios, minus the average or weighted average of the calculated allele ratios for the measured amounts of alleles present on the second homologous chromosomal segment, divided by the total measured amounts of all alleles for that locus, is below a threshold. In some embodiments, the data fit between the calculated allele ratios and the predicted allele ratios for copy number overrepresentation of the first homologous chromosomal segment exceeds the threshold. In some embodiments, the data fit between the calculated allele ratios and the predicted allele ratios for copy number overrepresentation of the first homologous chromosomal segment is below the threshold. In some embodiments, the threshold is determined from empirical testing of samples known to have the CNV of interest and / or samples known to lack the CNV.
[0185] In some embodiments, determining whether copy number overrepresentation of the first homologous chromosomal segment exists includes enumerating a set of one or more hypotheses indicating the degree of overrepresentation of the first homologous chromosomal segment. An exemplary hypothesis is the absence of overrepresentation because the first chromosomal segment and the homologous chromosomal segment are present in equal proportions (e.g., one copy of each segment in a diploid sample). Another exemplary hypothesis includes the first homologous chromosomal segment being replicated one or more times (e.g., one, two, three, four, five, or more excess copies of the first homologous chromosomal segment compared to the copy number of the second homologous chromosomal segment). Another exemplary hypothesis includes a deletion of the second homologous chromosomal segment. Yet another exemplary hypothesis is a deletion of both the first and second homologous chromosomal segments. In some embodiments, the predicted allele ratio for a locus that is heterozygous in at least one cell is estimated for each hypothesis, taking into account the degree of overrepresentation indicated by that hypothesis. In some embodiments, the likelihood that the hypothesis is correct is calculated by comparing the calculated allele ratios with the predicted allele ratios, and the hypothesis with the greatest likelihood is selected.
[0186] In some embodiments, an expected distribution of test statistics is calculated using the predicted values of allele ratios for each hypothesis, hi some embodiments, the likelihood that the hypothesis is correct is calculated by comparing the test statistic calculated using the calculated values of allele ratios with the expected distribution of test statistics calculated using the predicted values of allele ratios, and the hypothesis with the greatest likelihood is selected.
[0187] In some embodiments, a predicted allele ratio for a locus that is heterozygous in at least one cell is estimated taking into account the phasing genetic data for the first homologous chromosome segment, the phasing genetic data for the second homologous chromosome segment, and the degree of overrepresentation indicated by the hypothesis. In some embodiments, the likelihood that the hypothesis is correct is calculated by comparing the calculated allele ratio with the predicted allele ratio, and the hypothesis with the greatest likelihood is selected.
[0188] B. Use of mixed samples It will be understood that for many embodiments, the sample is a mixed sample containing DNA or RNA from one or more target cells and one or more non-target cells. In some embodiments, the target cells are cells that have a CNV, such as a deletion or duplication of interest, and the non-target cells are cells that do not have the copy number variation of interest (e.g., a mixture of cells with the deletion or duplication of interest and cells that do not contain any of the deletions or duplications being tested). In some embodiments, the target cells are cells associated with a disease or disorder or an elevated risk of a disease or disorder (e.g., cancer cells), and the non-target cells are cells that are not associated with a disease or disorder or an elevated risk of a disease or disorder (e.g., non-cancerous cells). In some embodiments, the target cells all have the same CNV. In some embodiments, two or more target cells have different CNVs. In some embodiments, one or more of the target cells have a CNV, polymorphism, or mutation associated with a disease or disorder or an elevated risk of a disease or disorder that is not found in at least one other target cell. In some such embodiments, the fraction of cells associated with a disease or disorder or an elevated risk of a disease or disorder among all cells from a sample is assumed to be equal to or greater than the fraction of the most frequent of these CNVs, polymorphisms, or mutations in the sample. For example, if 6% of cells have a K-ras mutation and 8% of cells have a BRAF mutation, then it is assumed that at least 8% of the cells are cancerous.
[0189] In some embodiments, the ratio of DNA (or RNA) from one or more target cells to the total DNA (or RNA) in the sample is calculated. In some embodiments, a set of one or more hypotheses indicating the degree of overrepresentation of a first homologous chromosomal segment is enumerated. In some embodiments, predicted allele ratios for loci that are heterozygous in at least one cell are estimated given the calculated DNA or RNA ratios, and the degree of overrepresentation indicated by that hypothesis is estimated for each hypothesis. In some embodiments, the likelihood that the hypothesis is correct is calculated by comparing the calculated allele ratios with the predicted allele ratios, and the hypothesis with the greatest likelihood is selected.
[0190] In some embodiments, an expected distribution of a test statistic calculated using the predicted allele ratios and calculated DNA or RNA ratios is estimated for each hypothesis. In some embodiments, the likelihood that the hypothesis is correct is determined by comparing the test statistic calculated using the calculated allele ratios and calculated DNA or RNA ratios with the expected distribution of a test statistic calculated using the predicted allele ratios and calculated DNA or RNA ratios, and the hypothesis with the greatest likelihood is selected.
[0191] In some embodiments, the method includes enumerating a set of one or more hypotheses that indicate the degree of overrepresentation of the first homologous chromosomal segment. In some embodiments, the method includes, for each hypothesis, estimating either (i) a predicted allele ratio for a locus that is heterozygous in at least one cell, taking into account the degree of overrepresentation indicated by that hypothesis, or (ii) for one or more possible ratios of DNA or RNA, an expected distribution of a test statistic calculated using the predicted allele ratio and the possible ratios of DNA or RNA from one or more target cells to the total DNA or RNA in the sample. In some embodiments, the data fitting is calculated by either (i) taking the calculated allele ratio as the predicted allele ratio, or (ii) comparing any of the test statistics calculated using the calculated allele ratio and the possible ratios of DNA or RNA to the expected distribution of a test statistic calculated using the predicted allele ratio and the possible ratios of DNA or RNA. In some embodiments, one or more of the hypotheses are ranked according to the data fitting, and the highest-ranked hypothesis is selected. In some embodiments, a technique or algorithm, such as a search algorithm, is used for one or more of computing a data fitting, ranking the hypotheses, or selecting the highest ranked hypothesis. In some embodiments, the data fitting is a fit to a beta-binomial distribution or a fit to a binomial distribution. In some embodiments, the technique or algorithm is selected from the group consisting of maximum likelihood estimation, empirical maximum estimation, Bayesian estimation, dynamic estimation (such as dynamic Bayesian estimation), and expectation maximization estimation. In some embodiments, the method includes applying the technique or algorithm to the obtained genetic data and predicted values of the genetic data.
[0192] In some embodiments, the method includes generating a distribution of possible ratios of DNA or RNA from one or more target cells to total DNA or RNA in the sample, ranging from a lower limit to an upper limit. In some embodiments, a set of one or more hypotheses is enumerated, each of which indicates the degree of overrepresentation of a first homologous chromosomal segment. In some embodiments, the method includes, for each possible ratio of DNA or RNA in the distribution and for each hypothesis, estimating either (i) a predicted allele ratio for a locus that is heterozygous in at least one cell, given the possible ratios of DNA or RNA and the degree of overrepresentation indicated by that hypothesis, or (ii) an expected distribution of test probabilities calculated using the predicted allele ratios and the possible ratios of DNA or RNA. In some embodiments, the method calculates, for each possible ratio of DNA or RNA in the distribution and for each hypothesis, the likelihood that the hypothesis is correct by either (i) comparing the calculated allele ratio to the predicted allele ratio, or (ii) comparing a test statistic calculated using the calculated allele ratio and the possible ratio of DNA or RNA to the expected distribution of test statistics calculated using the predicted allele ratio and the possible ratio of DNA or RNA. In some embodiments, a composite probability for each hypothesis is determined by adding up the probabilities of that hypothesis for each possible ratio in the distribution, and the hypothesis with the largest composite probability is selected. In some embodiments, the composite probability for each hypothesis is determined by weighting the probability of a hypothesis based on the likelihood that the possible ratio is the correct ratio for a particular possible ratio.
[0193] In some embodiments, a technique selected from the group consisting of maximum likelihood estimation, empirical maximum estimation, Bayesian estimation, dynamic estimation (such as dynamic Bayesian estimation), and expectation maximization estimation is used to estimate the ratio of DNA or RNA from one or more target cells to total DNA or RNA in the sample. In some embodiments, the ratio of DNA or RNA from one or more target cells to total DNA or RNA in the sample is assumed to be the same for two or more (or all) of the CNVs of interest. In some embodiments, the ratio of DNA or RNA from one or more target cells to total DNA or RNA in the sample is calculated for each CNV of interest.
[0194] C. Exemplary Methods for Using Incomplete Fading Data It will be understood that for many embodiments, incomplete phasing data is used. For example, it may not be known with 100% certainty which alleles are present for one or more of the loci on the first and / or second homologous chromosome segments. In some embodiments, prior probabilities for the individual's possible haplotypes (such as haplotypes based on population-based haplotype frequencies) are used in calculating the probability of each hypothesis. In some embodiments, the prior probabilities for the possible haplotypes are adjusted either by using another method for phasing genetic data or by using phasing data from other subjects (such as previous subjects) to refine the population data used for the individual's informatics-based phasing.
[0195] In some embodiments, the phasing genetic data includes probability data for two or more possible sets of phasing genetic data, each possible set of phasing data including the possible identity of an allele present at each locus in a set of polymorphic loci on a first homologous chromosomal segment and the possible identity of an allele present at each locus in a set of polymorphic loci on a second homologous chromosomal segment. In some embodiments, a probability for at least one of the hypotheses is determined for each possible set of phasing genetic data. In some embodiments, a composite probability for a hypothesis is determined by combining the probabilities of that hypothesis for each possible set of phasing genetic data, and the hypothesis with the largest composite probability is selected.
[0196] Any of the methods disclosed herein or any known method can be used to generate incomplete phasing data (such as using ensemble-based haplotype frequencies to infer the most likely phase) for use in the claimed methods. In some embodiments, the phasing data is obtained by probabilistically combining haplotypes of smaller segments. For example, possible haplotypes can be determined based on the possible combinations of one haplotype from a first region with another haplotype from another region on the same chromosome. The probability that particular haplotypes from different regions are part of the same, larger haplotype block on the same chromosome can be determined, for example, using ensemble-based haplotype frequencies and / or known recombination rates between different regions.
[0197] In some embodiments, a single-hypothesis rejection test is used for the null hypothesis of disomy. In some embodiments, the probability of the disomy hypothesis is calculated, and the disomy hypothesis is rejected if the probability is below a given threshold (e.g., less than 1 in 1,000). If the null hypothesis is rejected, this may be due to an error in incomplete phasing data or the presence of CNVs. In some embodiments, more accurate phasing data is obtained (e.g., phasing data from any of the molecular phasing methods disclosed herein to obtain actual phasing data, rather than phasing data estimated based on bioinformatics). In some embodiments, the probability of the disomy hypothesis is recalculated using this more accurate phasing data to determine whether the disomy hypothesis should still be rejected. Rejection of this hypothesis indicates the presence of a duplication or deletion of a chromosomal segment. If desired, the false positive rate can be changed by adjusting the threshold.
[0198] D. Further Exemplary Embodiments for Determining Ploidy Using Phasing Data In an illustrative embodiment, provided herein is a method for determining the ploidy of a chromosome segment in a sample from an individual. The method includes receiving allele frequency data at each locus in a set of polymorphic loci on the chromosome segment, the allele frequency data including the amount of each allele present in the sample, generating phased allele information for the set of polymorphic loci by estimating the phase of the allele frequency data, using the allele frequency data to generate individual probabilities of allele frequencies for the polymorphic loci for different ploidy states, using the individual probabilities and the phased allele information to generate a composite probability for the set of polymorphic loci, and selecting a best-fitting model that is indicative of chromosome ploidy based on the composite probability, thereby determining the ploidy of the chromosome segment.
[0199] As disclosed herein, allele frequency data (also referred to herein as measured gene allele data) can be generated by methods known in the art. For example, the data can be generated using qPCR or microarrays. In one exemplary embodiment, the data is generated using nucleic acid sequence data, particularly high-throughput nucleic acid sequence data.
[0200] In certain illustrative examples, the allele frequency data is corrected for errors before being used to generate the individual probabilities. In specific illustrative embodiments, the errors corrected include allele amplification efficiency bias. In other embodiments, the errors corrected include ambient contamination and genotype contamination. In some embodiments, the errors corrected include allele amplification bias, sequencing errors, ambient contamination, and genotype contamination.
[0201] In certain embodiments, the individual probabilities are generated using a set of models of different ploidy states and allelic imbalance fractions for the set of polymorphic loci. In these and other embodiments, the composite probability is generated by considering the linkages between polymorphic loci on a chromosomal segment.
[0202] Thus, in one illustrative embodiment that combines several of these embodiments, provided herein is a method for detecting chromosomal ploidy in a sample from an individual, the method including: receiving nucleic acid sequence data for alleles at a set of polymorphic loci on a chromosomal segment in the individual; using the nucleic acid sequence data to detect allele frequencies at the set of loci; correcting allele amplification efficiency bias in the detected allele frequencies to create corrected allele frequencies for the set of polymorphic loci; creating phasing allele information for the set of polymorphic loci by estimating the phase of the nucleic acid sequence data; comparing the corrected allele frequencies to a set of models of different ploidy states and allele imbalance fractions for the set of polymorphic loci to create individual probabilities of allele frequencies for the polymorphic loci for different ploidy states; combining the individual probabilities that take into account the linkages between polymorphic loci on the chromosomal segment to create a composite probability for the set of polymorphic loci; and selecting a best-fitting model that is indicative of chromosomal aneuploidy based on the composite probability.
[0203] As disclosed herein, individual probabilities can be generated using a set of models or hypotheses of different ploidy states and average allelic imbalance fractions for a set of polymorphic loci. For example, in a particularly illustrative example, individual probabilities are generated by modeling the ploidy states of a first homolog of a chromosomal segment and a second homolog of the chromosomal segment. The modeled ploidy states include: (1) none of the cells have a deletion or amplification of the first homolog or the second homolog of the chromosomal segment; (2) at least some of the cells have a deletion of the first homolog or an amplification of the second homolog of the chromosomal segment; and (3) at least some of the cells have a deletion of the second homolog of the chromosomal segment or an amplification of the first homolog.
[0204] It will be understood that the above model may also be referred to as hypotheses used to constrain the model. Thus, the above demonstrates three hypotheses that may be used.
[0205] The modeled average allelic imbalance fraction can include any range of average allelic imbalance, including the actual average allelic imbalance of the chromosomal segment. For example, in certain illustrative embodiments, the modeled average allelic imbalance range can be 0, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.75, 1, 2, 2.5, 3, 4, and 5% at the lower end and 1, 2, 2.5, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 95, and 99% at the upper end. The interval for modeling with this range can be any interval, depending on the computing power used and the time allowed for analysis. For example, intervals of 0.01, 0.05, 0.02, or 0.1 can be modeled.
[0206] In certain exemplary embodiments, the samples have an average allelic imbalance for chromosomal segments of 0.4% to 5%. In certain embodiments, the average allelic imbalance is low. In these embodiments, the average allelic imbalance is typically less than 10%. In certain exemplary embodiments, the allelic imbalance has lower limits of 0.25, 0.3, 0.4, 0.5, 0.6, 0.75, 1, 2, 2.5, 3, 4, and 5% and upper limits of 1, 2, 2.5, 3, 4, and 5%. In other exemplary embodiments, the average allelic imbalance has lower limits of 0.4, 0.45, 0.5, 0.6, 0.7, 0.8, 0.9, or 1.0% and upper limits of 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.5, 2.0, 3.0, 4.0, or 5.0%. For example, the average allelic imbalance of the sample is, in illustrative examples, 0.45-2.5%. In other examples, the average allelic imbalance is detected with a sensitivity of 0.45, 0.5, 0.6, 0.8, 0.8, 0.9, or 1.0%. That is, the test method can detect chromosomal aneuploidies down to an AAI of 0.45, 0.5, 0.6, 0.8, 0.8, 0.9, or 1.0%. Exemplary samples with low allelic imbalance for use in the methods of the present invention include plasma samples from individuals with cancer who have circulating tumor DNA or plasma samples from pregnant women who have circulating fetal DNA.
[0207] It will be understood that for SNV, the proportion of abnormal DNA is typically measured using mutant allele frequency (the number of mutant alleles at a locus / the total number of alleles at that locus).Because the difference between the amount of two homologs in tumors is similar, the proportion of abnormal DNA for CNV is measured by average allele imbalance (AAI), which is defined as |(H1-H2)| / (H1+H2), where Hi is the average copy number of homolog i in the sample, and Hi / (H1+H2) is the abundance fraction of homolog i, i.e., the homolog ratio.The maximum homolog ratio is the homolog ratio of the more abundant homolog.
[0208] The assay dropout rate is the proportion of SNPs with no reads, estimated using all SNPs. The single allele dropout (ADO) rate is the proportion of SNPs with only one allele present, estimated using only heterozygous SNPs. Genotype confidence can be determined by fitting a binomial distribution to the number of reads at each SNP that were B allele reads and estimating the probability of each genotype using the ploidy state of the SNP's focal region.
[0209] For tumor tissue samples, chromosomal aneuploidies (exemplified in this paragraph by CNVs) can be represented by transitions between allele frequency distributions. In plasma samples from cancer patients, individuals suspected of having cancer, individuals previously diagnosed with cancer, or as cancer screening for at-risk individuals or the general population, CNVs can be identified by a maximum likelihood algorithm that searches for plasma CNVs in regions known to exhibit aneuploidy in cancer and / or when tumor samples from the same individuals also have CNVs. In an exemplary embodiment, the algorithm uses haplotype phase information of the individual whose sample is analyzed for the presence of circulating tumor DNA and fits the measured and corrected allele counts of the test sample to predicted allele counts, for example, using a joint distribution model. Such haplotype phase information can be deduced from any sample from an individual that contains most, or at least 60, 70, 80, 90, 95, 96, 97, 98, 99%, or all normal cellular DNA, such as, but not limited to, a buffy coat sample, saliva sample, or skin sample, from parental genetic information, or by de novo haplotype phasing, which can be performed using a variety of methods (e.g., Snyder, M., et al., Haplotype-resolved genome sequencing: experimental methods and applications. Nat Rev Genet 16, 344-358 (2015)), such as whole-genome haplotyping by dilution (Kaper, F., et al., Whole-genome haplotyping by dilution, amplification, and sequencing. Proc Natl Acad Sci USA 110, 5552-5557 (2013)), or long-read sequencing (Kuleshov, V., et al., Haplotype-resolved genome sequencing: experimental methods and applications. Nat Rev Genet 16, 344-358 (2015)). This can be achieved by (Whole-genome haplotyping using long reads and statistical methods. Nat Biotech 32, 261-266 (2014)).The algorithm can model predicted allele frequencies across all allelic imbalance rates at 0.025% intervals for three sets of hypotheses: (1) all cells are normal (no allelic imbalance); (2) some / all cells have a deletion of homolog 1 or an amplification of homolog 2; or (3) some / all cells have a deletion of homolog 2 or an amplification of homolog 1. The likelihood of each hypothesis can be determined at each SNP using a Bayesian classifier based on a beta-binomial model of predicted and observed allele frequencies at all heterozygous SNPs. The joint likelihood across multiple SNPs can then be calculated, in certain illustrative embodiments, taking into account the joint SNP loci, as illustrated herein. Indeed, in illustrative embodiments, the haplotype phase information of normal cells, obtained as disclosed above, is used by the algorithm to fit measured and typically corrected test sample allele counts to predicted allele counts using a joint distribution model. The maximum likelihood hypothesis can then be selected.
[0210] Consider a chromosomal region with an average of N copies in the tumor, and let c denote the fraction of DNA in plasma that comes from a mixture of normal and tumor cells in the disomic region. The AAI is calculated as follows:
number
[0211] In certain illustrative examples, the allele frequency data is corrected for errors before being used to generate individual probabilities. Correction of different types of errors and / or biases is disclosed herein. In specific illustrative embodiments, the error corrected is allele amplification efficiency bias. In other embodiments, the error corrected includes sequencing errors, ambient contamination, and genotypic contamination. In some embodiments, the errors corrected include allele amplification bias, sequencing errors, ambient contamination, and genotypic contamination.
[0212] It will be appreciated that allele amplification efficiency bias can be determined for a given allele as part of an experimental or laboratory determination that includes the sample under test, or can be determined at a different time using a set of samples that include the allele for which the efficiency is being calculated. Ambient contamination and genotypic contamination are typically determined in the same run as the sample under test analysis.
[0213] In certain embodiments, ambient contamination and genotypic contamination are determined for homozygous alleles in the sample. For any given sample from an individual, even if a locus is selected for analysis because it has a relatively high heterozygosity in the population, it will be understood that some loci in the sample will be heterozygous and others will be homozygous. In some embodiments, it is advantageous to determine the ploidy of a chromosome segment using heterozygous loci for an individual, while ambient contamination and genotypic contamination can be calculated using homozygous loci.
[0214] In one particular illustrative example, the selection is made by analyzing the magnitude of the difference between the phasing allele information and the estimated allele frequencies generated for the model.
[0215] In an illustrative example, individual probabilities of allele frequencies are generated based on a beta-binomial model of the predicted and observed allele frequencies at a set of polymorphic loci. In an illustrative example, the individual probabilities are generated using a Bayesian classifier.
[0216] In certain exemplary embodiments, nucleic acid sequence data is generated by performing high-throughput DNA sequencing of multiple copies of a series of amplicons generated using a multiplex amplification reaction, wherein each amplicon in the series spans at least one polymorphic locus of a set of polymorphic loci, and each of the polymorphic loci in the set is amplified. In certain embodiments, the multiplex amplification reaction is performed under limiting primer conditions for at least half of the reactions. In some embodiments, limiting primer concentrations are used in 1 / 10, 1 / 5, 1 / 4, 1 / 3, 1 / 2, or all of the reactions in the multiplex reaction. Factors to consider for achieving limiting primer conditions in amplification reactions such as PCR are provided herein.
[0217] In certain embodiments, the methods provided herein detect the ploidy of multiple chromosome segments across multiple chromosomes. Thus, in these embodiments, the chromosome ploidy is determined for a set of chromosome segments in a sample. For these embodiments, a larger number of multiplex amplification reactions are required. Thus, for these embodiments, the multiplex amplification reaction can include, for example, 2,500 to 50,000 multiplex reactions. In certain embodiments, multiplex reactions are performed in the following ranges: from 100, 200, 250, 500, 1000, 2500, 5000, 10,000, 20,000, 25,000, and 50,000 at the lower end of the range to 200, 250, 500, 1000, 2500, 5000, 10,000, 20,000, 25,000, 50,000, and 100,000 at the upper end of the range.
[0218] In an exemplary embodiment, the set of polymorphic loci is a set of loci known to exhibit high heterozygosity. However, for any given individual, some of these loci are expected to be homozygous. In certain exemplary embodiments, the methods of the present invention utilize nucleic acid sequence information for both the individual's homozygous and heterozygous loci. The individual's homozygous loci are used, for example, for error correction, while the heterozygous loci are used to determine the allelic imbalance of the sample. In certain embodiments, at least 10% of the polymorphic loci are heterozygous loci for the individual.
[0219] As disclosed herein, preference is given to analyzing target SNP loci that are known to be heterozygous in the population. Thus, in certain embodiments, polymorphic loci are selected where 10, 20, 25, 50, 75, 80, 90, 95, 99, or 100% of the polymorphic loci are known to be heterozygous in the population.
[0220] As disclosed herein, in certain embodiments, the sample is a plasma sample from a pregnant woman.
[0221] In some examples, the method further includes performing the method on a control sample with a known average allelic imbalance ratio. The control may have an average allelic imbalance ratio for a particular allelic state that is indicative of aneuploidy of a chromosomal segment of 0.4-10%, for example, to mimic the average allelic imbalance of alleles in a sample that are present at low concentrations, as would be expected for circulating free DNA from a tumor.
[0222] In some embodiments, PlasmArt control is used as control as disclosed herein.Therefore, in certain aspects, this is the sample that comprises fragmenting the nucleic acid sample that is known to show chromosomal aneuploidy into fragments that mimic the size of the fragments of DNA that circulate in individual's plasma, and is prepared by a method.In certain aspects, the control that does not have aneuploidy for chromosomal segment is used.
[0223] In exemplary embodiments, data from one or more controls can be analyzed with the method along with the test sample. The controls can include, for example, a different sample from an individual not suspected of containing chromosomal aneuploidy, or a sample suspected of containing CNV or chromosomal aneuploidy. For example, if the test sample is a tumor sample suspected of containing circulating free tumor DNA, the method can be performed on a control sample derived from a tumor from the subject as well as a plasma sample. As disclosed herein, the control sample can be prepared by fragmenting a DNA sample known to exhibit chromosomal aneuploidy. Such fragmentation can obtain a DNA sample that mimics the DNA composition of apoptotic cells, particularly when the sample is from an individual suffering from cancer. Data from the control sample will increase the reliability of the detection of chromosomal aneuploidy.
[0224] In certain embodiments of the methods of determining ploidy, the sample is a plasma sample from an individual suspected of having cancer. In these embodiments, the method further comprises determining whether copy number variations are present in tumor cells of the individual based on the selecting. For these embodiments, the sample may be a plasma sample from the individual. In these embodiments, the method may further comprise determining whether cancer is present in the individual based on the selecting.
[0225] These embodiments for determining the ploidy of a chromosomal segment can further include detecting a single nucleotide variant at a single nucleotide dispersed position in the set of single nucleotide dispersed positions, wherein detecting either a chromosomal aneuploidy or a single nucleotide variant, or both, indicates the presence of circulating tumor nucleic acid in the sample.
[0226] These embodiments may further include receiving haplotype information of chromosomal segments for the individual's tumor and using this haplotype information to generate a set of models of different ploidy states and allelic imbalance fractions for the set of polymorphic loci.
[0227] As disclosed herein, certain embodiments of methods for determining ploidy can further include removing outliers from the initial or corrected allele frequency data before comparing the initial or corrected allele frequencies to the set of models. For example, in certain embodiments, locus allele frequencies that are at least two or three standard deviations below or above the mean values for other loci on the chromosome segment are removed from the data before being used for modeling.
[0228] As mentioned herein, it will be understood that for many of the embodiments provided herein, including those for determining the ploidy of a chromosome segment, incomplete or complete phasing data is preferably used. It will also be understood that several features are provided herein that provide improvements over conventional methods for detecting ploidy, and that many different combinations of these features may be used.
[0229] In certain embodiments, computer systems and computer-readable media are provided herein for performing any of the methods of the present invention. These include systems and computer-readable media for performing methods for determining ploidy. Thus, as a non-limiting example of a system embodiment, and to demonstrate that any of the methods provided herein can be performed using a system and computer-readable medium using the disclosures herein, in another aspect, a system for detecting chromosomal ploidy in a sample from an individual is provided herein, the system comprising: an input processor configured to receive allele frequency data including the amount of each allele present in the sample at each locus in a set of polymorphic loci on a chromosomal segment; a modeler configured to create phasing allele information for the set of polymorphic loci by estimating the phase of the allele frequency data, use the allele frequency data to create individual probabilities of allele frequencies for the polymorphic loci for different ploidy states, and use the individual probabilities and the phasing allele information to create a composite probability for the set of polymorphic loci; and a hypothesis manager configured to select a best-fitting model that is indicative of chromosomal ploidy based on the composite probability, thereby determining the ploidy of the chromosomal segment.
[0230] In certain embodiments of this system, the allele frequency data is data generated by a nucleic acid sequencing system. In certain embodiments, the system further comprises an error correction unit configured to correct errors in the allele frequency data, and the corrected allele frequency data is used by the modeler to generate the individual probabilities. In certain embodiments, the error correction unit corrects for allele amplification efficiency bias. In certain embodiments, the modeler generates the individual probabilities using a set of models of different ploidy states and allele imbalance fractions for the set of polymorphic loci. In certain exemplary embodiments, the modeler generates the composite probabilities by considering linkages between polymorphic loci on chromosomal segments.
[0231] In one illustrative embodiment, provided herein is a system for detecting chromosomal ploidy in a sample from an individual, the system comprising: an input processor configured to receive nucleic acid sequence data for alleles at a set of polymorphic loci on a chromosomal segment in the individual and to use the nucleic acid sequence data to detect allele frequencies at the set of loci; an error correction unit configured to correct errors in the detected allele frequencies and create corrected allele frequencies for the set of polymorphic loci; a modeler configured to create phased allele information for the set of polymorphic loci by estimating the phase of the nucleic acid sequence data, create individual probabilities of allele frequencies for the polymorphic loci for different ploidy states by comparing the phased allele information to a set of models of different ploidy states and allele imbalance fractions of the set of polymorphic loci, and create a composite probability for the set of polymorphic loci by combining the individual probabilities taking into account the relative distance between the polymorphic loci on the chromosomal segment; and a hypothesis manager configured to select a best-fitting model that is indicative of chromosomal aneuploidy based on the composite probability.
[0232] In certain exemplary system embodiments provided herein, the set of polymorphic loci includes 1,000 to 50,000 polymorphic loci. In certain exemplary system embodiments provided herein, the set of polymorphic loci includes 100 known heterozygosity hotspot loci. In certain exemplary system embodiments provided herein, the set of polymorphic loci includes 100 loci that are within 0.5 kb of or within a recombination hotspot.
[0233] In certain exemplary system embodiments provided herein, the best fitting model analyzes the following ploidy states of a first homolog of a chromosomal segment and a second homolog of a chromosomal segment: (1) none of the cells have a deletion or amplification of the first homolog or the second homolog of the chromosomal segment; (2) some or all of the cells have a deletion of the first homolog of the chromosomal segment or an amplification of the second homolog; or (3) some or all of the cells have a deletion of the second homolog of the chromosomal segment or an amplification of the first homolog.
[0234] In certain exemplary system embodiments provided herein, the errors corrected include allele amplification efficiency bias, contamination, and / or sequencing errors. In certain exemplary system embodiments provided herein, the contamination includes ambient contamination and genotypic contamination. In certain exemplary system embodiments provided herein, ambient contamination and genotypic contamination are determined for homozygous alleles.
[0235] In certain exemplary system embodiments provided herein, the hypothesis manager is configured to analyze the magnitude of the difference between the phasing allele information generated for the model and the estimated allele frequencies. In certain exemplary system embodiments provided herein, the modeler generates individual probabilities of allele frequencies based on a beta-binomial model of predicted and observed allele frequencies at the set of polymorphic loci. In certain exemplary system embodiments provided herein, the modeler generates the individual probabilities using a Bayesian classifier.
[0236] In certain exemplary system embodiments provided herein, nucleic acid sequence data is generated by performing high-throughput DNA sequencing of multiple copies of a series of amplicons generated using a multiplex amplification reaction, where each amplicon in the series spans at least one polymorphic locus of a set of polymorphic loci, and each of the polymorphic loci in the set is amplified. In certain exemplary system embodiments provided herein, the multiplex amplification reaction is performed under restrictive primer conditions for at least half of the reactions. In certain exemplary system embodiments provided herein, the samples have an average allelic imbalance of 0.4% to 5%.
[0237] In certain exemplary system embodiments provided herein, the sample is a plasma sample from an individual suspected of having cancer, and the hypothesis manager is further configured to determine whether copy number variations are present in tumor cells of the individual based on the best fitting model.
[0238] In certain exemplary system embodiments provided herein, the sample is a plasma sample from the individual, and the hypothesis manager is further configured to determine whether cancer is present in the individual based on the best fitting model. In these embodiments, the hypothesis manager can be further configured to detect single nucleotide variants at the single nucleotide variance positions in the set of single nucleotide variance positions, wherein detecting either a chromosomal aneuploidy or a single nucleotide variant, or both, indicates the presence of circulating tumor nucleic acid in the sample.
[0239] In certain exemplary system embodiments provided herein, the input processor is further configured to receive haplotype information of chromosomal segments for the individual's tumor, and the modeler is configured to use the haplotype information to create a set of models of different ploidy states and allelic imbalance fractions for the set of polymorphic loci.
[0240] In certain exemplary system embodiments provided herein, the modeler creates models over a range of allelic imbalance fractions from 0% to 25%.
[0241] It will be understood that any of the methods provided herein may be performed by computer-readable code stored on a non-transitory computer-readable medium. Thus, in one embodiment, provided herein is a non-transitory computer-readable medium for detecting chromosomal ploidy in a sample from an individual, the non-transitory computer-readable medium including computer-readable code that, when executed by a processing device, causes the processing device to receive allele frequency data including the amount of each allele present in the sample at each locus in a set of polymorphic loci on a chromosomal segment, generate phasing allele information for the set of polymorphic loci by estimating the phase of the allele frequency data, use the allele frequency data to generate individual probabilities of allele frequencies for the polymorphic loci for different ploidy states, use the individual probabilities and the phasing allele information to generate a composite probability for the set of polymorphic loci, and select a best-fitting model that is indicative of chromosomal ploidy based on the composite probability, thereby determining the ploidy of the chromosomal segment.
[0242] In certain computer-readable medium embodiments, the allele frequency data is generated from nucleic acid sequence data. Certain computer-readable medium embodiments further include correcting errors in the allele frequency data and using the corrected allele frequency data to generate the individual probabilities. In certain computer-readable medium embodiments, the error corrected is allele amplification efficiency bias. In certain computer-readable medium embodiments, the individual probabilities are generated using a set of models of different ploidy states and allele imbalance fractions for the set of polymorphic loci. In certain computer-readable medium embodiments, the composite probability is generated by considering linkages between polymorphic loci on a chromosomal segment.
[0243] In one particular embodiment, provided herein is a non-transitory computer-readable medium for detecting chromosomal ploidy in a sample from an individual, the non-transitory computer-readable medium comprising computer-readable code that, when executed by a processing device, causes the processing device to receive nucleic acid sequence data for alleles at a set of polymorphic loci on a chromosomal segment in the individual, use the nucleic acid sequence data to detect allele frequencies at the set of loci, correct for allele amplification efficiency bias in the detected allele frequencies to create corrected allele frequencies for the set of polymorphic loci, estimate the phase of the nucleic acid sequence data to create phasing allele information for the set of polymorphic loci, compare the corrected allele frequencies to a set of models of different ploidy states and allele imbalance fractions for the set of polymorphic loci to create individual probabilities of allele frequencies for the polymorphic loci for different ploidy states, combine the individual probabilities that take into account the linkages between the polymorphic loci on the chromosomal segment to create a composite probability for the set of polymorphic loci, and select a best-fitting model that is indicative of chromosomal aneuploidy based on the composite probability.
[0244] In certain illustrative computer-readable medium embodiments, the selecting is performed by analyzing the magnitude of the difference between the phasing allele information and the estimated allele frequencies generated for the model.
[0245] In certain illustrative computer-readable medium embodiments, individual probabilities of allele frequencies are generated based on a beta-binomial model of predicted and observed allele frequencies at a set of polymorphic loci.
[0246] It will be appreciated that any of the method embodiments provided herein may be performed by executing code stored on a non-transitory computer-readable medium.
[0247] E. Exemplary Embodiments for Detecting Cancer In certain aspects, the present invention provides a method for detecting cancer. It will be understood that the sample can be a tumor sample or a liquid sample, such as plasma, from an individual suspected of having cancer. The method is particularly effective for detecting genetic mutations, such as single nucleotide changes such as SNVs, or copy number changes such as CNVs in samples containing low levels of these genetic changes as part of the total DNA in the sample. Thus, the sensitivity for detecting DNA or RNA from cancer in a sample is exceptional. To achieve this exceptional sensitivity, the method can combine any or all of the improvements provided herein for detecting CNVs and SNVs.
[0248] Thus, in certain embodiments, provided herein is a method for determining whether circulating tumor nucleic acid is present in a sample from an individual, and a non-transitory computer-readable medium comprising computer-readable code that, when executed by a processing device, causes the processing device to perform the method. The method comprises: analyzing the sample to determine the ploidy at a set of polymorphic loci on a chromosomal segment in the individual; and determining a level of average allelic imbalance present at the polymorphic loci based on the ploidy determination, wherein an average allelic imbalance of 0.4%, 0.45%, 0.5%, 0.6%, 0.7%, 0.75%, 0.8%, 0.9%, or 1% or greater is indicative of the presence of circulating tumor nucleic acid, such as ctDNA, in the sample.
[0249] In certain illustrative examples, an average allelic imbalance of more than 0.4, 0.45, or 0.5% is indicative of the presence of ctDNA. In certain embodiments, the method for determining whether circulating tumor nucleic acid is present further comprises detecting a single nucleotide variant at the single nucleotide variance site in the set of single nucleotide variance positions, and detecting an allelic imbalance of 0.5% or more, or detecting a single nucleotide variant, or both, is indicative of the presence of circulating tumor nucleic acid in the sample. It will be understood that any of the methods provided for detecting chromosomal ploidy or CNV can be used to determine the level of allelic imbalance, typically expressed as average allelic imbalance. It will be understood that any of the methods provided herein for detecting SNV can be used to detect a single nucleotide for this aspect of the present invention.
[0250] In certain embodiments, the method for determining whether circulating tumor nucleic acid is present further comprises performing the method on a control sample having a known average allelic imbalance ratio. The control can be, for example, a sample from an individual's tumor. In some embodiments, the control has the average allelic imbalance expected for the sample under analysis, for example, an AAI of 0.5% to 5%, or an average allelic imbalance ratio of 0.5%.
[0251] In certain embodiments, the analyzing step in the methods for determining whether circulating tumor nucleic acid is present comprises analyzing a set of chromosomal segments known to exhibit aneuploidy in cancer. In certain embodiments, the analyzing step in the methods for determining whether circulating tumor nucleic acid is present comprises analyzing 1,000 to 50,000 or 100 to 1,000 polymorphic loci for ploidy. In certain embodiments, the analyzing step in the methods for determining whether circulating tumor nucleic acid is present comprises analyzing 100 to 1,000 single nucleotide variant sites. For example, in these embodiments, the analyzing step can comprise performing multiplex PCR to amplify amplicons across 1,000 to 50,000 polymorphic loci and 100 to 1,000 single nucleotide variant sites. This multiplex reaction can be configured as a single reaction or as a pool of multiplex reactions of different subsets. The multiplex reaction methods provided herein, such as the massively multiplex PCR disclosed herein, provide exemplary processes for performing amplification reactions to help achieve improved multiplexing and, therefore, sensitivity levels.
[0252] In certain embodiments, multiplex PCR reactions are performed under limiting primer conditions for at least 10%, 20%, 25%, 50%, 75%, 90%, 95%, 98%, 99%, or 100% of the reactions. Improved conditions for performing large-scale multiplex reactions provided herein can be used.
[0253] In certain aspects, the above-described methods for determining whether circulating tumor nucleic acid is present in a sample from an individual, and all embodiments thereof, can be performed using a system. The present disclosure provides teachings regarding specific functional and structural features for performing the above-described methods. As non-limiting examples, the system may include:
[0254] an input processor configured to analyze data from the sample to determine ploidy at a set of polymorphic loci on a chromosomal segment in the individual; and
[0255] A modeler configured to determine a level of allelic imbalance present at a polymorphic locus based on the ploidy determination, wherein an allelic imbalance of 0.5% or greater is indicative of the presence of cycling.
[0256] F. Exemplary Embodiments for Detecting Single Nucleotide Variants In certain aspects, provided herein are methods for detecting single nucleotide variants in a sample. The improved methods provided herein can achieve a detection limit of 0.015, 0.017, 0.02, 0.05, 0.1, 0.2, 0.3, 0.4, or 0.5% SNVs in a sample. All embodiments for detecting SNVs can be implemented using a system. The present disclosure provides teachings regarding specific functional and structural features for implementing the methods. Further provided herein are embodiments including a non-transitory computer-readable medium containing computer-readable code that, when executed by a processing device, causes the processing device to implement the methods for detecting SNVs provided herein.
[0257] Thus, in one embodiment, provided herein is a method for determining whether a single nucleotide variant is present at a set of genomic locations in a sample from an individual, the method comprising: for each genomic location, using a training dataset to generate estimates of the efficiency and error rate per cycle for the amplicon spanning that genomic location; receiving independently observed nucleotide identity information for each genomic location in the sample; using the estimated amplification efficiency and error rate per cycle for each genomic location to determine a set of probabilities for the fraction of single nucleotide variants resulting from one or more actual mutations at each genomic location by comparing the observed nucleotide identity information at each genomic location to models of different variant fractions; and determining from the set of probabilities for each genomic location the fraction and confidence of the most likely actual variant.
[0258] In an exemplary embodiment of a method for determining whether a single nucleotide variant is present, estimates of efficiency and error rate per cycle are generated for a set of amplicons spanning a genomic location. For example, the set may include 2, 3, 4, 5, 10, 15, 20, 25, 50, 100, or more amplicons spanning the genomic location.
[0259] In an exemplary embodiment of a method for determining whether a single nucleotide variant is present, the observed nucleotide identity information includes the observed number of total reads for each genomic location and the observed number of variant allele reads for each genomic location.
[0260] In an illustrative embodiment of the method for determining whether a single nucleotide variant is present, the sample is a plasma sample and the single nucleotide variant is present in circulating tumor DNA of the sample.
[0261] In another embodiment, provided herein is a method for estimating the proportion of single nucleotide variants present in a sample from an individual, comprising: at a set of genomic locations, using a training dataset to generate estimates of efficiency and error rate per cycle for one or more amplicons spanning those genomic locations; receiving observed nucleotide identity information for each genomic location in the sample; using the amplicon amplification efficiency and error rate per cycle to generate mean and variance estimates for the total number of molecules, background error molecules, and actual mutant molecules for a search space containing an initial proportion of actual mutant molecules; and determining the proportion of single nucleotide variants present in the sample that result from actual mutations by fitting a distribution to the observed nucleotide identity information in the sample using the mean and variance estimates.
[0262] In an illustrative example of this method for estimating the proportion of a single nucleotide variant present in a sample, the sample is a plasma sample and the single nucleotide variant is present in circulating tumor DNA of the sample.
[0263] The training dataset in this embodiment of the present invention typically includes samples from one healthy individual or, preferably, a group of healthy individuals. In certain exemplary embodiments, the training dataset is analyzed on the same day, or even in the same run, as one or more samples under test. For example, samples from a group of 2, 3, 4, 5, 10, 15, 20, 25, 30, 36, 48, 96, 100, 192, 200, 250, 500, 1000, or more healthy individuals can be used to create the training dataset. If data is available for a larger number of healthy individuals, e.g., 96 or more, the confidence in the estimate of amplification efficiency increases, even if a run is performed before performing the method on the samples under test. Because PCR error rates are per amplicon, nucleic acid sequence information generated for the entire amplified region surrounding the SNV can be used, not just for the SNV base position. For example, if samples from 50 individuals are used to sequence a 20-base pair amplicon surrounding the SNV, error frequency data from 1000-base reads can be used to determine the error frequency rate.
[0264] Typically, amplification efficiency is estimated by estimating the mean and standard deviation of the amplification efficiency for the amplified segments and then fitting this to a distribution model such as a binomial or beta-binomial distribution. The error rate is determined for a PCR with a known number of cycles, and then the error rate per cycle is estimated.
[0265] In certain exemplary embodiments, estimating the starting molecules for the test dataset further includes updating the efficiency estimate for the test dataset using the starting molecule number estimated in step (b) if the observed number of reads significantly differs from the estimate of the number of reads. This estimate can then be updated for the new efficiency and / or starting molecules.
[0266] The search space used to estimate the total number of molecules, background error molecules, and actual mutant molecules can include a search space with a lower limit of 0.1%, 0.2%, 0.25%, 0.5%, 1%, 2.5%, 5%, 10%, 15%, 20%, or 25% of copies of the base at the SNV position, which is the SNV base, and an upper limit of 1%, 2%, 2.5%, 5%, 10%, 12.5%, 15%, 20%, 25%, 50%, 75%, 90%, or 95%. Lower ranges, such as a lower limit of 0.1%, 0.2%, 0.25%, 0.5%, or 1%, and an upper limit of 1%, 2%, 2.5%, 5%, 10%, 12.5%, or 15%, can be used in an illustrative example for plasma samples, where the method detects circulating tumor DNA. Higher ranges are used for tumor samples.
[0267] A distribution is fitted to the number of total error molecules (background errors and actual mutations) in the total molecules to calculate the likelihood or probability for each possible actual mutation in the search space. This distribution can be a binomial or beta-binomial distribution.
[0268] The most likely actual mutation is determined by determining the proportion of the most likely actual mutation and calculating the confidence using the data from the distribution fitting. As an illustrative example, without intending to limit the clinical interpretation provided herein, when the average mutation rate is high, the confidence rate required to make a positive SNV determination is lower. For example, if the average mutation rate for SNVs in samples using the most likely hypothesis is 5% and the confidence rate is 99%, a positive SNV call will be made. On the other hand, for this illustrative example, if the average mutation rate for SNVs in samples using the most likely hypothesis is 1% and the confidence rate is 50%, in certain circumstances, a positive SNV call will not be made. It will be understood that the clinical interpretation of data may be a function of sensitivity, specificity, prevalence, and the availability of alternative products.
[0269] In one exemplary embodiment, the sample is a circulating DNA sample, such as a circulating tumor DNA sample.
[0270] In another embodiment, provided herein is a method for detecting one or more single nucleotide variants in a test sample from an individual, the method comprising the steps of:
[0271] determining a median variant allele frequency for a plurality of control samples from each of a plurality of normal individuals for each single nucleotide variant position in the set of single nucleotide variant positions based on the results produced in the sequencing run to identify selected single nucleotide variant positions having a median variant allele frequency in the normal samples that is below a threshold, and determining a background error for each of the single nucleotide variant positions after removing outlier samples for each of the single nucleotide variant positions; determining a weighted mean and variance of the read depth observed for the selected single nucleotide variant positions for the test samples based on the data produced in the sequencing run for the test samples; and using a computer to identify one or more single nucleotide variant positions having a statistically significant weighted mean of read depth compared to the background error for that position, thereby detecting the one or more single nucleotide variants.
[0272] In certain embodiments of this method for detecting one or more SNVs, the sample is a plasma sample, the control sample is a plasma sample, and the one or more detected single nucleotide variants detected are present in circulating tumor DNA of the sample. In certain embodiments of this method for detecting one or more SNVs, the plurality of control samples comprises at least 25 samples. In certain exemplary embodiments, the plurality of control samples is at least 5, 10, 15, 20, 25, 50, 75, 100, 200, or 250 samples at the lower end, and at least 10, 15, 20, 25, 50, 75, 100, 200, 250, 500, and 1000 samples at the upper end.
[0273] In certain embodiments of this method for detecting one or more SNVs, outliers are removed from the data generated by high-throughput sequencing run, and the weighted average of observed read depth is calculated, and observed variance is determined.In certain embodiments of this method for detecting one or more SNVs, the read depth for each single nucleotide variant position for test sample is at least 100 reads.
[0274] In certain embodiments of this method for detecting one or more SNVs, the sequencing run includes a multiplex amplification reaction performed under limited primer reaction conditions. Improved methods for performing multiplex amplification reactions provided herein are used to perform these embodiments in illustrative examples.
[0275] Without being limited by theory, the method of the present embodiment utilizes a background error model using normal plasma samples sequenced in the same sequencing run as the sample under test to account for run-specific artifacts. Noisy positions with median common variant allele frequencies above thresholds, e.g., 0.1%, 0.2%, 0.25%, 0.5%, 0.75%, and 1.0%, are removed.
[0276] Outlier samples are repeatedly removed from this model to account for noise and contamination. For each base substitution at all genomic loci, the read-depth-weighted mean and standard deviation of the error are calculated. In certain exemplary embodiments, samples such as tumor or cell-free plasma samples that have at least a threshold number of reads, for example, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 250, 500, or 1000 variant reads, and in certain embodiments, have a single nucleotide variant position with an a1 Z-score greater than 2.5, 5, 7.5, or 10 against the background error model, are counted as candidate mutations.
[0277] In certain embodiments, a read depth of more than 100, 250, 500, 1,000, 2000, 2500, 5000, 10,000, 20,000, 25,0000, 50,000, or 100,000 at the lower end of the range, and 2000, 2500, 5,000, 7,500, 10,000, 25,000, 50,000, 100,000, 250,000, or 500,000 at the upper end of the range, is achieved in a sequencing run for each single nucleotide variant position in a set of single nucleotide variant positions. Typically, the sequencing run is a high-throughput sequencing run. In exemplary embodiments, the average or median value generated for the samples under test is weighted by read depth. Thus, the likelihood that a variant allele determination is actual in a sample in which one variant allele is detected in 1000 reads is weighted more heavily than in a sample in which one variant allele is detected in 10,000 reads. Because variant allele (i.e., mutation) determinations are not made with 100% confidence, identified single nucleotide variants may be considered candidate variants or mutations.
[0278] G. Exemplary Test Statistics for Analysis of Fading Data Exemplary test statistics are described below for the analysis of phasing data from a sample known or suspected to be a mixed sample containing DNA or RNA derived from two or more genetically non-identical cells. f denotes the fraction of DNA or RNA of interest, e.g., the fraction of DNA or RNA containing a CNV of interest, or the fraction of DNA or RNA from a cell of interest, such as a cancer cell. In some embodiments of cancer testing, f denotes the fraction of DNA or RNA from a cancer cell in a mixture of cancer cells and normal cells, or f denotes the fraction of cancer cells in a mixture of cancer cells and normal cells. Note that this refers to the fraction of DNA from a cell of interest, assuming two copies of DNA are provided by each cell of interest. This is different from the fraction of DNA from a cell of interest with deleted or duplicated segments.
[0279] The possible allele values for each SNP are designated A and B. AA, AB, BA, and BB are used to represent all possible ordered allele pairs. In some embodiments, SNPs containing the ordered alleles AB or BA are analyzed. i indicates the number of sequence reads for the i-th SNP, and A i and B i Let denote the number of reads of the i-th SNP representing alleles A and B, respectively. Assume the following: N i =A i +B i .
[0280] Allele ratio R i is defined as follows:
number
[0281] Let T denote the number of targeted SNPs.
[0282] Without loss of generality, some embodiments focus on a single chromosome segment.For further clarity, in this specification, the phrase "a first homologous chromosome segment compared to a second homologous chromosome segment" refers to the first homolog of the chromosome segment and the second homolog of the chromosome segment.In some such embodiments, all of the target SNPs are contained in the target segment chromosome.In other embodiments, multiple chromosome segments are analyzed for possible copy number variations.
[0283] MAP estimation This method exploits knowledge of phasing through ordered allele pairs to detect deletions or duplications of target segments. For each SNP i, define
number
[0284] Then, the following definitions are made:
number
[0285] X under various copy number hypotheses (such as disomy, deletion of the first or second homolog, or duplication of the first or second homolog) i The distribution of and S is given below.
[0286] Disomy hypothesis Under the hypothesis that the target segment is not deleted or duplicated,
number
number
[0287] Assuming a constant read depth N, we have a binomial distribution S with the following parameters:
number
[0288] Deletion hypothesis Under the hypothesis that the first homolog is deleted (i.e., the AB SNP becomes B and the BA SNP becomes A), R i has a binomial distribution and parameters for AB SNPs
number
number
number
number
[0289] Under the hypothesis that the second homolog is deleted (i.e., the AB SNP becomes A and the BA SNP becomes B), R i has a binomial distribution and parameters for AB SNPs
number
number
number
number
[0290] overlap hypothesis Under the hypothesis that the first homolog is duplicated (i.e., the AB SNP becomes AAB and the BA SNP becomes BBA), R i has a binomial distribution and parameters for AB SNPs
number
number
number
number
[0291] Under the hypothesis that the second homolog is duplicated (i.e., the AB SNP becomes ABB and the BA SNP becomes BAA), R i has a binomial distribution and parameters for AB SNPs
number
number
number
number
[0292] classification As shown in the above chapter, X i is a binary random variable with
number
[0293] This allows the probability of the test statistic S to be calculated under each hypothesis. The probability of each hypothesis given the measured data can be calculated. In some embodiments, the hypothesis with the greatest probability is selected. If desired, a distribution for S can be calculated for each N i can be simplified either by approximating with a constant penetration depth N or by truncating the read depth to a constant value N. This simplification results in:
number
[0294] The value of f can be estimated by using an algorithm (e.g., a search algorithm) such as maximum likelihood estimation, empirical maximum estimation, or Bayesian estimation that takes into account the measured data to select the most likely value of f, such as the value of f that produces the best data fit. In some embodiments, multiple chromosomal segments are analyzed, and a value of f is estimated based on the data for each segment. If all target cells have these duplications or deletions, the estimates of f based on the data for these different segments will be similar. In some embodiments, f is measured experimentally, such as by determining the fraction of DNA or RNA from cancer cells based on the methylation difference (hypomethylated or hypermethylated) between cancer and non-cancerous DNA or RNA.
[0295] Single hypothesis rejection The distribution of S for the disomy hypothesis does not depend on f. Therefore, the probability of the measured data can be calculated for the disomy hypothesis without calculating f. A single-hypothesis rejection test can be used for the null hypothesis of disomy. In some embodiments, the probability of S under the disomy hypothesis is calculated, and the disomy hypothesis is rejected if the probability is below a given threshold (such as less than 1 in 1,000). This indicates the presence of a duplication or deletion of a chromosomal segment. If desired, the false positive rate can be changed by adjusting the threshold.
[0296] H. Exemplary Methods for Analysis of Fading Data Exemplary methods are described below for analyzing data from samples known or suspected to be mixed samples containing DNA or RNA from two or more cells that are not genetically identical. In some embodiments, phasing data is used. In some embodiments, the method involves determining, for each calculated allele ratio, whether the calculated allele ratio is above or below the expected allele ratio and the magnitude of the difference for a particular locus. In some embodiments, a likelihood distribution is determined for the allele ratios at loci for a particular hypothesis; the closer the calculated allele ratio is to the center of the likelihood distribution, the more likely the hypothesis is correct. In some embodiments, the method involves determining the likelihood that a hypothesis is correct for each locus. In some embodiments, the method involves determining the likelihood that a hypothesis is correct for each locus and combining the probabilities of the hypotheses for each locus; the hypothesis with the greatest combined probability is selected. In some embodiments, the method involves determining the likelihood that a hypothesis is correct for each locus and for each possible ratio of DNA or RNA from one or more target cells to total DNA or RNA in the sample. In some embodiments, a composite probability for each hypothesis is determined by combining the probabilities of the hypotheses for each locus and each possible ratio, and the hypothesis with the largest composite probability is selected.
[0297] In one embodiment, the following hypotheses are considered: 11 (All cells are normal), H 10 (presence of cells with only homolog 1, and therefore lack of homolog 2), H 01 (presence of cells with only homolog 2, and therefore lack of homolog 1), H 21 (presence of cells with homolog 1 duplication), H 12 (Presence of cells with a homolog 2 duplication). For a fraction f of target cells (or a fraction of DNA or RNA from target cells), such as cancer cells or mosaic cells, the predicted allele ratio for a heterozygous (AB or BA) SNP can be found as follows: Equation (1):
number
[0298] Correction of bias, contamination, and sequencing errors: Observation D in SNP s is the number of originally mapped reads in which each allele is present, n A 0 and n B 0 Then, using the predicted bias in the amplification of the A and B alleles, the corrected read n A and n B can be found.
[0299] c a indicates ambient contamination (e.g., contamination from DNA in the air or polluted environment), and r(c a ) denotes the allele ratio for the ambient pollutant (initially taken as 0.5). Furthermore, c g indicates the genotype contamination rate (e.g., contamination from another sample), and r(c g ) is the allele ratio for that contaminant. e (A, B) and se (B,A) indicates a sequencing error that calls one allele a different allele (such as by falsely detecting the A allele when the B allele is present).
[0300] By correcting for ambient contamination, genotypic contamination, and sequencing errors, for a given predicted allele ratio r, the observed allele ratio q(r,c a ,r(c a ),c g ,r(c g ),s e (A,B),s e (B,A)) can be found.
[0301] Since the genotype of the contaminant is unknown, the ensemble frequency is used to calculate P(r(c g ) can be found. More specifically, let p be the population frequency for one of the alleles (which may be referred to as the reference allele). Then, P(r(c g )=0)=(1-p) 2 , P(r(c g )=0)=2p(1-p), and P(r(cg)=0)=p 2 r(c g ) to find the conditional expectation over E[q(r,c a ,r(c a ),c g ,r(c g ),s e (A,B),s e (B, A)) can be determined. Note that ambient contamination and genotypic contamination are determined using homozygous SNPs and are therefore not affected by the absence or presence of deletions or duplications. Furthermore, if desired, ambient contamination and genotypic contamination can be measured using a reference chromosome.
[0302] Likelihood at each SNP: The following equation takes into account the allele ratio r and A and n BFind the probability of observing Equation (2):
number
[0303] D s Let denote the data for SNP s. Each hypothesis h∈{H 11 ,H 01 ,H 10 ,H 21 ,H 12}, in equation (1), r = r(AB,h) or r = r(BA,h), and then r(c g ) and find the conditional expectation over the observed allele ratios E[q(r,c a ,r(c a ),c g ,r(c g ))] can then be determined by using equation (2): r = E[q(r,c a ,r(c a ),c g ,r(c g ),s e (A,B),s e (B,A))], and P(D s |h,f) can be determined.
[0304] Search algorithm: In some embodiments, SNPs with allele ratios that appear to be outliers are ignored (such as by ignoring or excluding SNPs with allele ratios at least two or three standard deviations above or below the mean), with an identified advantage of this approach being that it ensures that SNPs are not trimmed due to mosaicism, as allele ratio variability may be high in the presence of higher rates of mosaicism.
[0305] F={f1,….,f N} denotes the search space for mosaic proportion (e.g., tumor fraction). P(D s|h,f) can be determined and the likelihoods across all SNPs can be combined.
[0306] The algorithm runs through each f for each hypothesis. Using a search method, it concludes that mosaic exists if there exists a range F* of f where the confidence of the deletion or duplication hypothesis is higher than the confidence of the no deletion and no duplication hypotheses. In some embodiments, P(D s The maximum likelihood estimate of |h,f) is determined. If desired, the conditional expectation over f∈F* can be determined. If desired, the confidence for each hypothesis can be determined.
[0307] In some embodiments, a beta-binomial distribution is used instead of a binomial distribution, hi some embodiments, a reference chromosome or chromosome segment is used to determine sample-specific parameters of the beta-binomial.
[0308] Theoretical performance using simulation: If desired, the theoretical performance of the algorithm can be evaluated by randomly assigning the number of reference reads to SNPs at a given depth of read (DOR). Typically, p = 0.5 is used for the binomial probability parameter, and p is adjusted accordingly for deletions or duplications. Exemplary input parameters for each simulation are: (1) the number of SNPs S, (2) a constant DOR D per SNP, (3) p, and (4) the number of experiments.
[0309] First simulation experiment: This experiment focused on S ∈ {500, 1000}, D ∈ {500, 1000}, and p ∈ {0%, 1%, 2%, 3%, 4%, 5%}. For each setting, 1,000 simulation experiments were performed (thus, 24,000 experiments with phases and 24,000 experiments without phases). The number of reads was simulated from a binomial distribution (other distributions can be used if desired). The false positive rate (for p = 0%) and false negative rate (for p > 0%) were determined both with and without phase information. Note that phase information is very useful, especially for S = 1000, D = 1000. However, for S = 500, D = 500, the algorithm has the highest false positive rate, regardless of the presence or absence of phase-out from the tested condition.
[0310] Phase information is especially useful at low mosaic rates (≤3%). Without phase information, the confidence in the deletion is limited to H 10 and H 01 Because the probability of a hypothesis being determined by assigning equal chances to the probability of a hypothesis, a high level of false negatives is observed for p=1%, and small deviations in favor of one hypothesis are not enough to compensate for the low likelihood of the other hypothesis. This is also true for overlaps. The algorithm also appears to be more sensitive to read depth compared to the number of SNPs. For results using phase information, we assume that complete phase information is available for a large number of consecutive heterozygous SNPs. If desired, haplotype information can be obtained by probabilistically combining haplotypes for smaller segments.
[0311] Second simulation experiment: This experiment focused on S ∈ {100, 200, 300, 400, 500}, D ∈ {1000, 2000, 3000, 4000, 5000}, and p ∈ {0%, 1%, 1.5%, 2%, 2.5%, 3%}, and 10,000 random experiments for each setting. The false positive rate (when p = 0%) and false negative rate (when p > 0%) were determined both with and without phase information. Using haplotype information, the false negative rate is less than 10% for D ≥ 3000 and N ≥ 200, while achieving the same performance for D = 5000 and N ≥ 400. For small mosaic fractions, the difference between the false negative rates was particularly striking. For example, when p = 1%, without haplotype data, the false negative rate never reached less than 20%, while it was close to 0% for N ≥ 300 and D ≥ 3000. For p=3%, a 0% false negative rate is observed with haplotype data, while without haplotype data, N≧300 and D≧3000 are required to reach the same performance.
[0312] I. Exemplary Methods for Detecting Deletions and Duplications Without Phasing Data In some embodiments, unphased genetic data is used to determine whether there is an overrepresentation of the copy number of a first homologous chromosomal segment compared to a second homologous chromosomal segment in the individual's genome (such as in the genome of one or more cells, or in cfDNA or cfRNA). In some embodiments, phased genetic data is used, but phasing is ignored. In some embodiments, the DNA or RNA sample is a mixed cfDNA or cfRNA sample from an individual containing cfDNA or cfRNA from two or more genetically distinct cells. In some embodiments, the method utilizes the magnitude of the difference between the calculated allele ratio and the predicted allele ratio for each locus.
[0313] In some embodiments, the methods involve obtaining genetic data at a set of polymorphic loci on a chromosome or chromosome segment in a sample of DNA or RNA from one or more cells from an individual by measuring the amount of each allele at each locus. In some embodiments, allele ratios are calculated for loci that are heterozygous in at least one cell from which the sample was derived. In some embodiments, the calculated allele ratio for a particular locus is the measured amount of one of the alleles divided by the total measured amount of all alleles for that locus. In some embodiments, the calculated allele ratio for a particular locus is the measured amount of one of the alleles (e.g., an allele on a first homologous chromosome segment) divided by the measured amount of one or more other alleles for that locus (e.g., an allele on a second homologous chromosome segment). Calculated allele ratios and predicted allele ratios can be calculated using any of the methods described herein or any standard method (e.g., any mathematical transformation of the calculated allele ratios or predicted allele ratios described herein).
[0314] In some embodiments, a test statistic is calculated based on the magnitude of the difference between the calculated allele ratio and the expected allele ratio for each of the loci. In some embodiments, the test statistic Δ is calculated using the following formula:
number
[0315] For example, if the predicted allele ratio is 0.5, then δi can be defined as follows:
number
[0316] μ i and σ i The value for R i is a binomial random variable. In some embodiments, the standard deviation is assumed to be the same for all loci. In some embodiments, the mean or weighted mean of the standard deviations, or an estimate of the standard deviation, is σ i 2 In some embodiments, the test statistic is assumed to have a normal distribution. For example, the central limit theorem suggests that as the number of loci (such as the number of SNPs, T) increases, the distribution of Δ converges to a normal distribution.
[0317] In some embodiments, a set of one or more hypotheses indicating the copy number of a chromosome or chromosome segment in one or more genomes of the cell is enumerated. In some embodiments, the most likely hypothesis is selected based on a test statistic, thereby determining the copy number of a chromosome or chromosome segment in one or more genomes of the cell. In some embodiments, a hypothesis is selected if the probability that the test statistic belongs to the distribution of the test statistic for a hypothesis exceeds an upper threshold. One or more of the hypotheses is rejected if the probability that the test statistic belongs to the distribution of the test statistic for a hypothesis is below a lower threshold, or a hypothesis is neither selected nor rejected if the probability that the test statistic belongs to the distribution of the test statistic for a hypothesis is between the lower and upper thresholds, or if the probability cannot be determined with sufficiently high confidence. In some embodiments, the upper and / or lower thresholds are determined from an empirical distribution, such as a distribution from training data (e.g., diploid samples or samples with known copy numbers, such as samples known to have a particular deletion or duplication). Such an empirical distribution can be used to select a threshold for single-hypothesis rejection testing. Note that the test statistic Δ is independent of S, so both can be used independently if desired.
[0318] J. Exemplary Methods for Detecting Deletions or Duplications Using Allele Distributions or Patterns This section includes methods for determining whether there is an overrepresentation of a copy number of a first homologous chromosomal segment relative to a second homologous chromosomal segment. In some embodiments, the methods involve enumerating (i) a plurality of hypotheses indicating the copy number of a chromosome or chromosomal segment present in the genome of one or more cells (e.g., cancer cells) of an individual, or (ii) a plurality of hypotheses indicating the degree of overrepresentation of a copy number of a first homologous chromosomal segment relative to a second homologous chromosomal segment in the genome of one or more cells of the individual. In some embodiments, the methods involve obtaining genetic data from an individual at a plurality of polymorphic loci (e.g., SNP loci) on a chromosome or chromosomal segment. In some embodiments, a probability distribution of the individual's predicted genotype for each of the hypotheses is created. In some embodiments, a data fit between the individual's obtained genetic data and the probability distribution of the individual's predicted genotype is calculated. In some embodiments, one or more hypotheses are ranked according to the data fit, and the highest ranked hypothesis is selected. In some embodiments, a technique or algorithm, such as a search algorithm, is used for one or more of calculating the data fit, ranking the hypotheses, or selecting the highest ranked hypothesis. In some embodiments, the data fitting is fitting to a beta-binomial distribution or fitting to a binomial distribution. In some embodiments, the technique or algorithm is selected from the group consisting of maximum likelihood estimation, empirical maximum estimation, Bayesian estimation, dynamic estimation (such as dynamic Bayesian estimation), and expectation maximization estimation. In some embodiments, the method includes applying the technique or algorithm to the obtained genetic data and predicted values of the genetic data.
[0319] In some embodiments, the methods involve enumerating (i) a plurality of hypotheses indicating the copy number of a chromosome or chromosome segment present in the genome of one or more cells (e.g., cancer cells) of the individual, or (ii) a plurality of hypotheses indicating the degree of copy number overrepresentation of a first homologous chromosome segment compared to a second homologous chromosome segment in the genome of one or more cells of the individual. In some embodiments, the methods involve obtaining genetic data from the individual at a plurality of polymorphic loci (e.g., SNP loci) on the chromosome or chromosome segment. In some embodiments, the genetic data includes allele counts for the plurality of polymorphic loci. In some embodiments, a joint distribution model is created for predicted values of allele counts at the plurality of polymorphic loci on the chromosome or chromosome segment for each hypothesis. In some embodiments, the relative probabilities of one or more of the hypotheses are determined using the joint distribution model and the allele counts measured for the sample, and the hypothesis with the greatest probability is selected.
[0320] In some embodiments, the distribution or pattern of alleles (such as the pattern of calculated allele ratios) is used to determine the presence or absence of a CNV, such as a deletion or duplication. If desired, the parental origin of the CNV can be determined based on this pattern.
[0321] K. Exemplary Counting / Quantitation Methods In some embodiments, one or more counting methods (also referred to as quantification methods) are used to detect one or more CNS, such as deletions or duplications of chromosome segments or entire chromosomes. In some embodiments, one or more counting methods are used to determine whether the overrepresentation of the copy number of a first homologous chromosome segment is due to duplication of the first homologous chromosome segment or deletion of a second homologous chromosome segment. In some embodiments, one or more counting methods are used to determine the excess copy number of a duplicated chromosome segment or chromosome (e.g., whether there are 1, 2, 3, 4, or more excess copies). In some embodiments, one or more counting methods are used to distinguish samples with many duplications and a smaller tumor fraction from samples with fewer duplications and a larger tumor fraction. For example, one or more counting methods can be used to distinguish a sample with four excess chromosome copies and a 10% tumor fraction from a sample with two excess chromosome copies and a 20% tumor fraction. Exemplary methods are disclosed, for example, in U.S. Publication Nos. 2007 / 0184467, 2013 / 0172211, and 2012 / 0003637, U.S. Patent Nos. 8,467,976, 7,888,017, 8,008,018, 8,296,076, and 8,195,415, U.S. Application No. 62 / 008,235 filed June 5, 2014, and U.S. Application No. 62 / 032,785 filed August 4, 2014, each of which is incorporated by reference herein in its entirety.
[0322] In some embodiments, the counting method involves counting the number of DNA sequence-based reads that map to one or more given chromosomes or chromosomal segments. Some such methods involve creating a reference value (cutoff value) for the number of DNA sequence reads that map to a particular chromosome or chromosomal segment, with excess read counts being indicative of a particular genetic abnormality.
[0323] In some embodiments, the total measured amount of all alleles for one or more loci (such as the total number of polymorphic or non-polymorphic loci) is compared to a reference value. In some embodiments, the reference amount is (i) a threshold or (ii) a predicted amount for a particular copy number hypothesis. In some embodiments, the reference amount (for the absence of CNV) is the total measured amount of all alleles for one or more loci for one or more chromosomes or chromosome segments known or predicted to not have a deletion or duplication. In some embodiments, the reference amount (for the presence of CNV) is the total measured amount of all alleles for one or more loci for one or more chromosomes or chromosome segments known or predicted to have a deletion or duplication. In some embodiments, the reference amount is the total measured amount of all alleles for one or more loci for one or more reference chromosomes or chromosome segments. In some embodiments, the reference amount is the mean or median of values determined for two or more different chromosomes, chromosome segments, or different samples. In some embodiments, random (e.g., massively parallel shotgun sequencing) or targeted sequencing is used to determine the amount of one or more polymorphic or non-polymorphic loci.
[0324] In some embodiments utilizing a reference amount, the method includes: (a) measuring the amount of genetic material on a chromosome or chromosome segment of interest; (b) comparing the amount from step (a) to the reference amount; and (c) identifying the presence or absence of a deletion or duplication based on the comparison.
[0325] In some embodiments utilizing a reference chromosome or chromosome segment, the method includes sequencing DNA or RNA from the sample to obtain multiple sequence tags that align to target loci. In some embodiments, the sequence tags are sufficiently long (e.g., 15-100 nucleotides long) to be assigned to specific target loci, and the target loci are derived from multiple different chromosomes or chromosome segments, including at least one first chromosome or chromosome segment suspected of having an abnormal distribution in the sample and at least one second chromosome or chromosome segment presumed to be normally distributed in the sample. In some embodiments, the multiple sequence tags are assigned to their corresponding target loci. In some embodiments, the number of sequence tags that align to the target loci of the first chromosome or chromosome segment and the number of sequence tags that align to the target loci of the second chromosome or chromosome segment are determined. In some embodiments, these numbers are compared to determine the presence or absence of an abnormal distribution (e.g., deletion or duplication) of the first chromosome or chromosome segment.
[0326] In some embodiments, a value of f (e.g., tumor fraction) is used in CNV determination, such as to compare the observed difference between the amounts of two chromosomes or chromosome segments to the difference expected for a particular type of CNV given the value of f (see, e.g., U.S. Publication Nos. 2012 / 0190020, 2012 / 0190021, 2012 / 0190557, and 2012 / 0191358, each of which is incorporated by reference herein in its entirety). For example, the difference in the amount of a duplicated chromosome segment in a tumor compared to a disomic reference chromosome segment increases as tumor fraction increases. In some embodiments, the method includes comparing the relative frequency of a chromosome or chromosome segment of interest to a reference chromosome or chromosome segment (a chromosome or chromosome segment predicted or known to be disomic) to the value of f to determine the likelihood of a CNV. For example, the difference in dosage between a first chromosome or chromosome segment and a reference chromosome or chromosome segment can be compared to what would be expected given the values of f for various possible CNVs (such as one or two extra copies of the chromosome segment of interest).
[0327] The following hypothetical example illustrates the use of counting / quantification methods to distinguish between duplications of a first homologous chromosomal segment and deletions of a second homologous chromosomal segment. If the host's normal disomic genome is considered the baseline, analysis of a mixture of normal and cancer cells yields the average difference between the baseline and cancer DNA in the mixture. For example, imagine that 10% of the DNA in a sample is derived from cells with a deletion across the chromosomal region targeted by the assay. In some embodiments, the quantification approach indicates that the amount of reads corresponding to this region is predicted to be 95% of the amount expected for a normal sample. This means that because one of the two target chromosomal regions in each tumor cell with a deletion of the target region is missing, the total amount of DNA mapping to this region is 90% (for normal cells) + ½ × 10% (for tumor cells) = 95%. Alternatively, in some embodiments, the allele approach indicates that the allele ratio at heterozygous loci is, on average, 19:20. Next, imagine that 10% of the DNA in a sample is derived from cells with a 5-fold focal amplification of the chromosomal region targeted by the assay. In some embodiments, a quantitative approach indicates that the amount of reads corresponding to this region is predicted to be 125% of the amount expected for a normal sample. This means that one of the two target chromosomal regions in each of the tumor cells with a 5-fold focal amplification is overcopied 5 times across the target region, so the total amount of DNA mapping to this region is 90% (for normal cells) + (2 + 5) × 10% (for tumor cells) / 2 = 125%. Alternatively, in some embodiments, an allele approach indicates that the allele ratio at heterozygous loci is, on average, 25:20. Note that using only an allele approach, a 5-fold focal amplification across a chromosomal region in a sample containing 10% cfDNA may appear to be the same as a deletion across the same region in a sample containing 40% cfDNA.In these two cases, the under-represented haplotype in the deletion case appears to be the haplotype without a CNV in the focal duplication case, and the haplotype without a CNV in the deletion case appears to be the over-represented haplotype in the focal duplication case. The likelihoods generated by the allele approach and the quantitative approach are combined to distinguish between these two probabilities.
[0328] L. Exemplary Counting / Quantification Methods Using Reference Samples Exemplary quantification methods using one or more reference samples are described in U.S. Application No. 62 / 008,235, filed June 5, 2014, and U.S. Application No. 62 / 032,785, filed August 4, 2014, which are incorporated by reference in their entireties. In some embodiments, one or more reference samples (e.g., normal samples) that are most likely to be free of CNVs on one or more chromosomes or chromosomes of interest are identified by selecting the sample with the highest tumor DNA fraction, selecting the sample with a z-score closest to 0, selecting a sample whose data fit a hypothesis corresponding to the absence of CNVs with the highest confidence or likelihood, selecting a sample known to be normal, selecting a sample from an individual with the lowest likelihood of having cancer (e.g., young age, male when screening for breast cancer, no family history, etc.), selecting the sample with the highest DNA input, selecting the sample with the highest signal-to-noise ratio, selecting samples based on other criteria believed to be correlated with the likelihood of having cancer, or selecting samples using some combination of criteria. Once a reference set is selected, these cases can be assumed to be disomy, and bias per SNP, i.e., experiment-specific amplification and other processing biases for each locus, can be estimated. This experiment-specific bias estimate can then be used to correct bias in measurements of the chromosome of interest, such as the locus of chromosome 21, and, if appropriate, for other chromosomal loci, for samples that are not part of the subset in which disomy is not assumed for chromosome 21. Once bias has been corrected for these samples with unknown ploidy, the data for these samples can be analyzed twice using the same or different methods to determine whether the individual has trisomy 21. For example, a quantification method can be used on the remaining samples with unknown ploidy, and z-scores can be calculated using the corrected genetic data measurements for chromosome 21.Alternatively, the tumor fraction of a sample from an individual suspected of having cancer can be calculated as part of a preliminary estimation of the ploidy state of chromosome 21. The expected corrected read fraction for disomy (disomy hypothesis) and the expected corrected read fraction for trisomy (trisomy hypothesis) can be calculated for a case with that tumor fraction. Alternatively, if the tumor fraction has not been previously measured, a set of disomy and trisomy hypotheses can be created for different tumor fractions. For each case, an expected distribution of corrected read fractions can be calculated, given expected statistical variation in the selection and measurement of various DNA loci. The observed corrected read fractions can be compared to the expected distribution of corrected read fractions, and likelihood ratios can be calculated for the disomy and trisomy hypotheses for each sample with unknown ploidy. The ploidy state associated with the hypothesis with the highest calculated likelihood can be selected as the correct ploidy state.
[0329] In some embodiments, a subset of samples with a sufficiently low likelihood of having cancer can be selected to serve as a control set of samples. This subset can be a fixed number or a variable number based on selecting only samples below a threshold. Quantitative data from a subset of samples can be combined, averaged, or combined using a weighted average, with the weighting based on the likelihood of the sample being normal. The quantitative data can be used to determine per-locus bias for amplification sequencing of samples in an immediate batch of control samples. Per-locus bias can also include data from other batches of samples. Per-locus bias can indicate the relative over- or under-amplification observed for that locus compared to other loci, assuming that the subset of samples does not contain CNVs and that any observations of over- or under-amplification are due to amplification and / or sequencing or other biases. Per-locus bias can take into account the GC content of the amplicon. Loci can be grouped into locus groups for the purpose of calculating per-locus bias. Once the bias per locus has been calculated for each locus in the plurality of loci, the sequencing data for one or more of the samples not in the sample subset, and optionally one or more of the samples in the sample subset, can be corrected by adjusting the quantitative measurement for each locus to remove the effect of bias at that locus. For example, if SNP1 is observed to have a read depth twice as large as the average in a subset of patients, the adjustment can involve replacing the corresponding number of reads from SNP1 with a number half that size. If the locus in question is a SNP, the adjustment can involve reducing the number of reads corresponding to each allele at that locus by half. Once the sequencing data for each locus in one or more samples has been adjusted, it can be analyzed using a method to detect the presence of CNVs in one or more chromosomal regions.
[0330] In one example, sample A is a mixture of amplified DNA derived from a mixture of normal and cancerous cells analyzed using a quantitative method. The following shows exemplary possible data: A region of the q arm on chromosome 22 is found to have only 90% of the expected value of DNA mapping to that region, a focal region corresponding to the HER2 gene is found to have 150% of the expected value of DNA mapping to that region, and the p arm of chromosome 5 is found to have 105% of the expected value of DNA mapping to that region. A clinician can infer that the sample has a deletion of a region on the q arm of chromosome 22 and a duplication of the HER2 gene. Because 22q deletions are common in breast cancer and because cells with deletions of the 22q regions on both chromosomes do not usually survive, a clinician can infer that approximately 20% of the DNA in the sample was derived from cells with a 22q deletion on one of the two chromosomes. Clinicians can also infer that the cells contained a five-fold duplication of the HER2 region if the DNA from a mixed sample derived from tumor cells was derived from a set of genetic tumor cells that were homogeneous for the HER2 and 22q regions.
[0331] In one example, sample A is also analyzed using the allele method. The following shows exemplary possible data: Two haplotypes for the same region on the q arm of chromosome 22 are present in a ratio of 4:5, two haplotypes in the focal region corresponding to the HER2 gene are present in a ratio of 1:2, and two haplotypes in the p arm of chromosome 5 are present in a ratio of 20:21. All other assayed regions of the genome do not contain any haplotypes in statistically significant excess. A clinician can infer that the sample contains DNA from a tumor with CNVs in the 22q region, the HER2 region, and the 5p arm. Based on the knowledge that 22q deletions are very common in breast cancer and / or a quantitative analysis showing an underrepresentation of DNA mapping to the 22q region of the genome, a clinician can infer the presence of a tumor with 22q deletion. Based on the knowledge that HER2 amplification is very common in breast cancer and / or quantitative analysis showing an overrepresentation of the amount of DNA mapping to the HER2 region of the genome, clinicians can infer the presence of a tumor with HER2 amplification.
[0332] M. Exemplary Reference Chromosomes or Chromosome Segments In some embodiments, any of the methods described herein are also performed on one or more reference chromosomes or chromosome segments, and the results are compared to the results for one or more chromosomes or chromosome segments of interest.
[0333] In some embodiments, a reference chromosome or chromosome segment is used as a control predicted to be free of CNVs. In some embodiments, the reference is the same chromosome or chromosome segment from one or more different samples that is known or predicted to have no deletions or duplications in the chromosome or chromosome segment. In some embodiments, the reference is a different chromosome or chromosome segment from the sample being tested that is predicted to be disomic. In some embodiments, the reference is a different segment from one of the chromosomes of interest in the same sample being tested. For example, the reference can be one or more segments outside the region of the potential deletion or duplication. Having a reference for the same chromosome being tested avoids variability between different chromosomes, such as differences in metabolism, apoptosis, histone, inactivation, and / or amplification between chromosomes. Analysis of segments that do not contain CNVs on the same chromosome being tested can also be used to determine differences in metabolism, apoptosis, histone, inactivation, and / or amplification between homologs, allowing the level of variability between homologs in the absence of CNV to be determined for comparison with results from potential CNVs. In some embodiments, the magnitude of the difference between the calculated and expected allele ratios for the potential CNV is greater than the corresponding magnitude for the reference, thereby confirming the presence of the CNV.
[0334] In some embodiments, a reference chromosome or chromosome segment is used as a control where a CNV, such as a specific deletion or duplication of interest, is expected to exist. In some embodiments, the reference is the same chromosome or chromosome segment from one or more different samples known or predicted to have a deletion or duplication in the chromosome or chromosome segment. In some embodiments, the reference is a different chromosome or chromosome segment from the test sample known or predicted to have a CNV. In some embodiments, the magnitude of the difference between the calculated and predicted allele ratios for the potential CNV is similar (e.g., not significantly different) to the corresponding magnitude for the reference CNV, thereby confirming the presence of the CNV. In some embodiments, the magnitude of the difference between the calculated and predicted allele ratios for the potential CNV is smaller (e.g., significantly smaller) than the corresponding magnitude for the reference CNV, thereby confirming the absence of the CNV. In some embodiments, one or more loci where the genotype of cancer cells (or DNA or RNA from cancer cells, such as cfDNA or cfRNA) differs from the genotype of non-cancerous cells (or DNA or RNA from non-cancerous cells, such as cfDNA or cfRNA) are used to determine tumor fraction. The tumor fraction can be used to determine whether the overrepresentation of a copy number of a first homologous chromosomal segment is due to a duplication of the first homologous chromosomal segment or a deletion of the second homologous chromosomal segment. The tumor fraction can also be used to determine the excess copy number of a duplicated chromosomal segment or chromosome (whether there are 1, 2, 3, 4, or more excess copies), such as to distinguish a sample with 4 excess chromosomal copies and a tumor fraction of 10% from a sample with 2 excess chromosomal copies and a tumor fraction of 20%. The tumor fraction can also be used to determine how well the observed data matches the predicted data for possible CNVs. In some embodiments, the degree of overrepresentation of CNVs is used to select a particular therapy or treatment regimen for an individual.For example, some therapeutic agents are only effective against at least four, six, or more copies of a chromosomal segment.
[0335] In some embodiments, the one or more loci used to determine tumor fraction are located on a reference chromosome or chromosome segment, e.g., a chromosome or chromosome segment known or predicted to be disomic, a chromosome or chromosome segment rarely duplicated or deleted in cancer cells generally or in a particular type of cancer in individuals known to have or at elevated risk of having, or a chromosome or chromosome segment with a low likelihood of aneuploidy (e.g., such a segment predicted to cause cell death if deleted or duplicated). In some embodiments, any of the methods of the invention are used to confirm that the reference chromosome or chromosome segment is disomic in both cancer cells and non-cancerous cells. In some embodiments, one or more chromosomes or chromosome segments with high confidence in calling disomy are used.
[0336] Exemplary loci that can be used to determine tumor fraction include polymorphisms or mutations (e.g., SNPs) in cancer cells (or DNA or RNA, such as cfDNA or cfRNA, from cancer cells) that are not present in non-cancerous cells (or DNA or RNA from non-cancerous cells) in an individual. In some embodiments, tumor fraction is determined by identifying polymorphic loci at which cancer cells (or DNA or RNA from cancer cells) have alleles that are not present in non-cancerous cells (or DNA or RNA from non-cancerous cells) in a sample (e.g., a plasma sample or tumor specimen) from the individual, and using the amount of alleles unique to the cancer cells at one or more of the identified polymorphic loci to determine the tumor fraction in the sample. In some embodiments, the non-cancerous cells are homozygous for a first allele at the polymorphic locus, and the cancer cells are (i) heterozygous for the first allele and a second allele, or (ii) homozygous for the second allele at the polymorphic locus. In some embodiments, non-cancerous cells are heterozygous for a first and a second allele at a polymorphic locus, and cancer cells have (i) one or two copies of a third allele at the polymorphic locus. In some embodiments, cancer cells are assumed or known to have only one copy of an allele not present in non-cancerous cells. For example, if the genotype of non-cancerous cells is AA and the genotype of cancer cells is AB, and 5% of the signals at that locus in a sample are from the B allele and 95% are from the A allele, the tumor fraction of the sample is 10%. In some embodiments, cancer cells are assumed or known to have two copies of an allele not present in non-cancerous cells. For example, if the genotype of non-cancerous cells is AA and the genotype of cancer cells is BB, and 5% of the signals at that locus in a sample are from the B allele and 95% are from the A allele, the tumor fraction of the sample is 5%.In some embodiments, multiple loci at which cancer cells have alleles that are not present in non-cancerous cells are analyzed to determine which loci in cancer cells are heterozygous and which are homozygous. For example, for loci at which non-cancerous cells have AA, if the signal from the B allele is about 5% at some loci and about 10% at some loci, then the cancer cells are assumed to be heterozygous at loci with about 5% B alleles and homozygous at loci with about 10% B alleles (indicating a tumor fraction of about 10%).
[0337] Exemplary loci that can be used to determine tumor fraction include loci where cancer cells and non-cancerous cells share one allele in common (e.g., a locus where cancer cells are AB and non-cancerous cells are BB, or a locus where cancer cells are BB and non-cancerous cells are AB). The amount of A signal, the amount of B signal, or the ratio of A to B signals in a mixed sample (containing DNA or RNA from cancer cells and non-cancerous cells) is compared with the corresponding value for (i) a sample containing DNA or RNA from cancer cells only, or (ii) a sample containing DNA or RNA from non-cancerous cells only. The difference in values is used to determine the tumor fraction of the mixed sample.
[0338] In some embodiments, loci that can be used to determine tumor fraction are selected based on the genotype of (i) a sample containing DNA or RNA from cancer cells only and / or (ii) a sample containing DNA or RNA from non-cancerous cells only. In some embodiments, loci are selected based on analysis of mixed samples, such as loci where the absolute or relative amount of each allele differs from the amount expected if both cancer cells and cancerous cells have the same genotype at a particular locus. For example, if cancer cells and non-cancerous cells have the same genotype, a locus is expected to produce a 0% B signal if all cells are AA, a 50% B signal if all cells are AB, or a 100% B signal if all cells are BB. Other values of the B signal indicate that the genotypes of cancer cells and non-cancerous cells differ at that locus, and therefore the locus can be used to determine tumor fraction.
[0339] In some embodiments, the tumor fraction calculated based on alleles at one or more loci is compared to the tumor fraction calculated using one or more of the counting methods disclosed herein.
[0340] N. Exemplary Methods for Detecting Phenotypes or Analyzing Multiple Mutations In some embodiments, the method involves analyzing a sample for a set of mutations associated with a disease or disorder (e.g., cancer) or an elevated risk of a disease or disorder. There is a strong correlation between events within a class (e.g., cancer class M or C) that can be used to improve the signal-to-noise ratio of a method and classify tumors into distinct clinical subsets. For example, a borderline result for several mutations (e.g., several CNVs) on one or more chromosomes or chromosomal segments considered together can be a very strong signal. In some embodiments, determining the presence or absence of multiple polymorphisms or mutations of interest (e.g., 2, 3, 4, 5, 8, 10, 12, 15, or more) increases the sensitivity and / or specificity of determining the presence or absence of a disease or disorder, such as cancer, or an elevated risk of a disease or disorder, such as cancer. In some embodiments, correlations between events across multiple chromosomes are used to see a stronger signal compared to looking at each of them individually. The design of the method itself can be optimized to optimally classify tumors. This can be very useful for early detection and screening for recurrence, where sensitivity to one particular mutation / CNV may be most important. In some embodiments, events are not always correlated, but have a probability of being correlated. In some embodiments, a matrix estimation composition is used that has a noise covariance matrix with off-diagonal terms.
[0341] In some embodiments, the invention features a method for detecting a phenotype (such as a cancer phenotype) in an individual, where the phenotype is defined by the presence of at least one of a set of mutations. In some embodiments, the method includes obtaining DNA or RNA measurements on a sample of DNA or RNA from one or more cells from the individual, where one or more cells are suspected of having the phenotype, and analyzing the DNA or RNA measurements to determine, for each mutation in the set of mutations, the likelihood that at least one of the cells has that mutation. In some embodiments, the method includes determining that the individual has the phenotype if either (i) for at least one of the mutations, the likelihood that at least one of the cells has that mutation is greater than a threshold, or (ii) for at least one of the mutations, the likelihood that at least one of the cells has that mutation is less than a threshold, and for multiple mutations, the combined likelihood that at least one of the cells has at least one of the mutations is greater than a threshold. In some embodiments, one or more cells have a subset or all of the mutations in the set of mutations. In some embodiments, the subset of mutations is associated with cancer or an elevated risk of cancer. In some embodiments, the set of mutations includes a subset or all of the mutations in the M class of cancer mutations (Ciriello, Nat Genet. 45(10):1127-1133, 2013, doi:10.1038 / ng.2762, incorporated herein by reference in its entirety). In some embodiments, the set of mutations includes a subset or all of the mutations in the C class of cancer mutations (Ciriello, supra). In some embodiments, the sample comprises cell-free DNA or RNA. In some embodiments, the DNA or RNA measurements include measurements at a set of polymorphic loci (e.g., the dosage of each allele at each locus) on one or more chromosomes or chromosomal segments of interest.
[0342] O. Exemplary Combinations of Methods To increase the accuracy of the results, more than one method for detecting the presence or absence of a CNV (such as any of the methods of the present invention or any known method) is performed. In some embodiments, more than one method for analyzing factors that are indicative of the presence or absence of a disease or disorder, or an increased risk of a disease or disorder (such as any of the methods of the present invention or any known method) is performed.
[0343] In some embodiments, standard mathematical techniques are used to calculate the covariance and / or correlation between two or more methods. Standard mathematical techniques may also be used to determine the combined probability of a particular hypothesis based on two or more tests. Exemplary techniques include meta-analysis, Fisher's combined probability test for independent tests, Brown's method for combining dependent p-values with known covariances, and cost methods for combining dependent p-values with unknown covariances. In cases where likelihoods are determined by a first method that is orthogonal or independent to the way likelihoods are determined for a second method, combining the likelihoods is straightforward and can be done by multiplication and normalization, or by using a formula such as R comb =R1R2 / [R1R2+(1-R1)(1-R2)]
[0344] R comb is the joint likelihood, and R1 and R2 are the individual likelihoods. For example, if the likelihood of trisomy from Method 1 is 90% and the likelihood of trisomy from Method 2 is 95%, by combining the outputs from the two methods, a clinician can conclude that the fetus is trisomic with a likelihood of (0.90)(0.95) / [(0.90)(0.95)+(1-0.90)(1-0.95)]=99.42%. The likelihoods can still be combined if the first and second methods are not orthogonal, i.e., if there is a correlation between the two methods.
[0345] Exemplary methods for analyzing multiple factors or variables are disclosed in U.S. Pat. No. 8,024,128, issued Sep. 20, 2011, U.S. Publication No. 2007 / 0027636, filed July 31, 2006, and U.S. Publication No. 2007 / 0178501, filed December 6, 2006, each of which is incorporated herein by reference in its entirety.
[0346] In various embodiments, the combined probability of a particular hypothesis or diagnosis is greater than 80, 85, 90, 92, 94, 96, 98, 99, or 99.9%, or greater than some other threshold.
[0347] P. Detection Limit As demonstrated by the experiments provided in the Examples section, the methods provided herein can detect an average allelic imbalance in a sample with a limit of detection or sensitivity of 0.45% AAI (which is the limit of detection for aneuploidy in an exemplary method of the present invention). Similarly, in certain embodiments, the methods provided herein can detect an average allelic imbalance in a sample of 0.45, 0.5, 0.6, 0.8, 0.8, 0.9, or 1.0%. That is, the test method can detect chromosomal aneuploidy in a sample down to an AAI of 0.45, 0.5, 0.6, 0.8, 0.8, 0.9, or 1.0%. As demonstrated by the experiments provided in the Examples section, the methods provided herein can detect the presence of SNVs in a sample with a limit of detection or sensitivity of 0.2%, which is the limit of detection for at least some SNVs in one exemplary embodiment. Similarly, in certain embodiments, the methods are capable of detecting SNVs at a frequency or SNV AAI of 0.2, 0.3, 0.4, 0.5, 0.6, 0.8, 0.8, 0.9, or 1.0%, i.e., the test methods are capable of detecting SNVs in a sample down to a detection limit of 0.2, 0.3, 0.4, 0.5, 0.6, 0.8, 0.8, 0.9, or 1.0% of the total number of alleles at the chromosomal locus of the SNV.
[0348] In some embodiments, the detection limit of mutations (such as SNVs or CNVs) of the methods of the present invention is 10, 5, 2, 1, 0.5, 0.1, 0.05, 0.01, or 0.005% or less. In some embodiments, the detection limit of mutations (such as SNVs or CNVs) of the methods of the present invention is 15 to 0.005%, for example, 10 to 0.005%, 10 to 0.01%, 10 to 0.1%, 5 to 0.005%, 5 to 0.01%, 5 to 0.1%, 1 to 0.005%, 1 to 0.01%, 1 to 0.1%, 0.5 to 0.005%, 0.5 to 0.01%, 0.5 to 0.1%, or 0.1 to 0.01 (inclusive).
[0349] In some embodiments, the detection limit is a value that allows (or can allow) detection of a mutation (such as an SNV or CNV) present in 10, 5, 2, 1, 0.5, 0.1, 0.05, 0.01, or 0.005% or less of the DNA or RNA comprising a locus in a sample (such as a cfDNA or cfRNA sample). For example, a mutation can be detected when 10, 5, 2, 1, 0.5, 0.1, 0.05, 0.01, or 0.005% or less of the DNA or RNA molecules comprising the locus (e.g., instead of a wild-type or non-mutated version of the locus or a different mutation at the locus) have a mutation in the locus. In some embodiments, the detection limit is a value that allows (or can allow) detection of a mutation (such as an SNV or CNV) present in 10, 5, 2, 1, 0.5, 0.1, 0.05, 0.01, or 0.005% or less of the DNA or RNA molecules in a sample (such as a cfDNA or cfRNA sample). In some embodiments where the CNV is a deletion, the deletion can be detected even if it is present in no more than 10, 5, 2, 1, 0.5, 0.1, 0.05, 0.01, or 0.005% of the DNA or RNA molecules in the sample that have the region of interest, which may or may not contain the deletion. In some embodiments where the CNV is a deletion, the deletion can be detected even if it is present in no more than 10, 5, 2, 1, 0.5, 0.1, 0.05, 0.01, or 0.005% of the DNA or RNA molecules in the sample. In some embodiments where the CNV is a duplication, the duplication can be detected even if the excess duplicated RNA or DNA present in the sample is present in no more than 10, 5, 2, 1, 0.5, 0.1, 0.05, 0.01, or 0.005% of the DNA or RNA molecules in the sample that have the region of interest, which may or may not overlap. In some embodiments where the CNV is a duplication, the duplication can be detected even if the excess duplicated RNA or DNA present is only present in an amount less than or equal to 10, 5, 2, 1, 0.5, 0.1, 0.05, 0.01, or 0.005% of the DNA or RNA molecules in the sample.
[0350] Q. Illustrative Samples In some embodiments of any of the aspects of the invention, the sample contains intracellular and / or extracellular genetic material from cells suspected of having a deletion or duplication, such as cells suspected of being cancerous. In some embodiments, the sample includes any tissue or bodily fluid suspected of containing cells, DNA, or RNA with a deletion or duplication, such as a tumor or other sample containing cancer cells, DNA, or RNA. The genetic measurements used as part of these methods can be performed on any sample containing DNA or RNA, including, but not limited to, tissue, blood, serum, plasma, urine, hair, tears, saliva, skin, fingernails, feces, bile, lymph, cervical mucus, semen, tumor, or other cells or substances containing nucleic acids. The sample can include any cell type, or DNA or RNA from any cell type can be used (such as cells from any organ or tissue suspected of being cancerous, or neurons). In some embodiments, the sample contains nuclear and / or mitochondrial DNA. In some embodiments, the sample is derived from any of the target individuals disclosed herein. In some embodiments, the target individual is a cancer patient.
[0351] Exemplary samples include those containing cfDNA or cfRNA. In some embodiments, cfDNA is available for analysis without the need for cell lysis. Cell-free DNA can be obtained from various tissues, such as tissues in liquid form, for example, blood, plasma, lymph, ascites, or cerebrospinal fluid. In some cases, cfDNA consists of DNA derived from fetal cells. In some cases, cfDNA is isolated from plasma separated from whole blood by centrifugation to remove cellular material. cfDNA can be a mixture of DNA derived from target cells (such as cancer cells) and non-target cells (such as non-cancer cells).
[0352] In some embodiments, the sample contains or is suspected of containing a mixture of DNA (or RNA), such as a mixture of DNA (or RNA) from cancer cells and DNA (or RNA) from non-cancerous (i.e., normal) cells. In some embodiments, at least 0.5, 1, 3, 5, 7, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 92, 94, 95, 96, 98, 99, or 100% of the cells in the sample are cancer cells. In some embodiments, at least 0.5, 1, 3, 5, 7, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 92, 94, 95, 96, 98, 99, or 100% of the DNA (e.g., cfDNA) or RNA (e.g., cfRNA) in the sample is derived from cancer cell(s). In various embodiments, the percentage of cells in a sample that are cancerous cells is between 0.5 and 99%, e.g., between 1 and 95%, 5 and 95%, 10 and 90%, 5 and 70%, 10 and 70%, 20 and 90%, or 20 and 70%, inclusive. In some embodiments, the sample is enriched for cancer cells or enriched for DNA or RNA from cancer cells. In some embodiments of a sample enriched for cancer cells, at least 0.5, 1, 2, 3, 4, 5, 6, 7, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 92, 94, 95, 96, 98, 99, or 100% of the cells in the enriched sample are cancer cells. In some embodiments of a sample enriched for DNA or RNA from cancer cells, at least 0.5, 1, 2, 3, 4, 5, 6, 7, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 92, 94, 95, 96, 98, 99, or 100% of the DNA or RNA in the enriched sample is derived from cancer cell(s).In some embodiments, cell sorting (fluorescence-activated cell sorting (FACS)) is used to enrich for cancer cells (Barteneva et al., Biochim Biophys Acta., 1836(1):105-22, Aug 2013. doi:10.1016 / j.bbcan.2013.02.004. Epub 2013 Feb 24, and Ibrahim et al., Adv Biochem Eng Biotechnol. 106:19-39, 2007, each of which is incorporated by reference in its entirety).
[0353] In some embodiments, the sample is enriched for fetal cells. In some embodiments of samples enriched for fetal cells, at least 0.5, 1, 2, 3, 4, 5, 6, 7%, or more of the cells in the enriched sample are fetal cells. In some embodiments, the percentage of cells in the sample that are fetal cells is between 0.5 and 100%, e.g., between 1 and 99%, 5 and 95%, 10 and 95%, 10 and 95%, 20 and 90%, or 30 and 70%, inclusive. In some embodiments, the sample is enriched for fetal DNA. In some embodiments of samples enriched for fetal DNA, at least 0.5, 1, 2, 3, 4, 5, 6, 7%, or more of the DNA in the enriched sample is fetal DNA. In some embodiments, the percentage of DNA in the sample that is fetal DNA is between 0.5 and 100%, e.g., between 1 and 99%, between 5 and 95%, between 10 and 95%, between 10 and 95%, between 20 and 90%, or between 30 and 70% (inclusive).
[0354] In some embodiments, a sample comprises a single cell or comprises DNA and / or RNA from a single cell. In some embodiments, multiple individual cells (e.g., at least 5, 10, 20, 30, 40, or 50 cells from the same or different subjects) are analyzed in parallel. In some embodiments, cells from multiple samples from the same individual are combined, reducing the amount of work compared to analyzing the samples separately. Combining multiple samples can also allow multiple tissues to be tested for cancer simultaneously (which can be used to provide a more thorough screening for cancer or to determine whether the cancer may have spread to other tissues).
[0355] In some embodiments, the sample contains a single cell or a small number of cells, such as 2, 3, 5, 6, 7, 8, 9, or 10 cells. In some embodiments, the sample has 1 to 100, 100 to 500, or 500 to 1,000 cells (inclusive). In some embodiments, the sample contains 1 to 10 picograms, 10 to 100 picograms, 100 picograms to 1 nanogram, 1 to 10 nanograms, 10 to 100 nanograms, or 100 nanograms to 1 microgram of RNA and / or DNA (inclusive).
[0356] In some embodiments, the sample is embedded in parafilm. In some embodiments, the sample is preserved with a preservative such as formaldehyde and optionally embedded in paraffin, which may cause cross-linking of DNA so that a small amount is available for PCR. In some embodiments, the sample is a formaldehyde-fixed, paraffin-embedded (FFPE) sample. In some embodiments, the sample is a fresh sample (such as a sample obtained within one or two days of analysis). In some embodiments, the sample is frozen prior to analysis. In some embodiments, the sample is a historical sample.
[0357] These samples can be used in any of the methods of the present invention.
[0358] R. Exemplary Sample Preparation Methods In some embodiments, the method includes isolating or purifying DNA and / or RNA. There are several standard procedures known in the art to achieve this goal. In some embodiments, the sample may be centrifuged to separate the various layers. In some embodiments, DNA or RNA may be isolated using filtration. In some embodiments, DNA or RNA preparation may involve amplification, separation, chromatographic purification, liquid separation, isolation, preferential enrichment, preferential amplification, targeted amplification, or any of several other techniques known in the art or described herein. In some embodiments for DNA isolation, RNase is used to degrade RNA. In some embodiments for RNA isolation, DNase (such as DNase I from Invitrogen, Carlsbad, CA, USA) is used to degrade DNA. In some embodiments, the RNeasy Mini Kit (Qiagen) is used to isolate RNA according to the manufacturer's protocol. In some embodiments, small RNAs are isolated using the mirVana PARIS kit (Ambion, Austin, TX, USA) according to the manufacturer's protocol (Gu et al., J. Neurochem. 122:641-649, 2012, incorporated herein by reference in its entirety). RNA concentration and purity can optionally be determined using Nanovue (GE Healthcare, Piscataway, NJ, USA), and RNA integrity can optionally be measured by using a 2100 Bioanalyzer (Agilent Technologies, Santa Clara, CA, USA) (Gu et al., J. Neurochem. 122:641-649, 2012, incorporated herein by reference in its entirety). In some embodiments, TRIZOL or RNAlater (Ambion) is used to stabilize RNA during storage.
[0359] In some embodiments, universal tagged adapters are added to create the library. Prior to ligation, the sample DNA may be blunt-ended, and then a single adenosine base is added to the 3' end. Prior to ligation, the DNA may be cleaved using a restriction enzyme or some other cleavage method. During ligation, the 3' adenosine of the sample fragment and the complementary 3' tyrosine overhang of the adapter can increase ligation efficiency. In some embodiments, adapter ligation is performed using a ligation kit such as found in an AGILENT SURESELECT kit. In some embodiments, the library is amplified using universal primers. In one embodiment, the amplified library is fractionated by size separation or by using products such as AGENCORT AMPURE beads or other similar methods. In some embodiments, PCR amplification is used to amplify the target loci. In some embodiments, the amplified DNA is sequenced (such as using an ILLUMINA IIGAX or HiSeq sequencer). In some embodiments, the amplified DNA is sequenced from each end of the amplified DNA to reduce sequencing errors: if there is a sequence error at a particular base when sequencing from one end of the amplified DNA, there is less likely to be a sequence error in the complementary base when sequencing from the other end of the amplified DNA (compared to multiple sequencing from the same end of the amplified DNA).
[0360] In some embodiments, whole genome amplification (WGA) is used to amplify nucleic acid samples. Several methods are available for WGA, including ligation-mediated PCR (LM-PCR), degenerate oligonucleotide primer PCR (DOP-PCR), and multiple displacement amplification (MDA). In LM-PCR, short DNA sequences called adapters are ligated to blunt ends of DNA. These adapters contain universal amplification sequences that are used to amplify DNA by PCR. In DOP-PCR, random primers that also contain universal amplification sequences are used in the first round of annealing and PCR. A second round of PCR is then used to further amplify the sequences using the universal primer sequences. MDA uses phi-29 polymerase, a highly processive, non-specific enzyme that replicates DNA and has been used in single-cell analysis. In some embodiments, WGA is not performed.
[0361] In some embodiments, selective amplification or enrichment is used to amplify or enrich target loci. In some embodiments, amplification and / or selective enrichment techniques may involve PCR, such as ligation-mediated PCR, fraction capture by hybridization, molecular inversion probes, or other circularization probes. In some embodiments, real-time quantitative PCR (RT-qPCR), digital PCR, or emulsion PCR, single-allele base extension reactions followed by mass spectrometry are used (Hung et al., J Clin Pathol 62:308-313, 2009, incorporated herein by reference in its entirety). In some embodiments, hybridization capture using hybrid capture probes is used to preferentially enrich DNA. In some embodiments, methods for amplification or selective enrichment may involve using probes such that, upon correct hybridization to the target sequence, the 3' or 5' end of the nucleotide probe is separated from the polymorphic site of the polymorphic allele by a small number of nucleotides. This separation reduces preferential amplification of one allele, referred to as allelic bias. This is an improvement over methods involving the use of probes in which the 3' or 5' end of a correctly hybridized probe is immediately adjacent to or very close to the polymorphic site of an allele. In one embodiment, probes whose hybridizing region may or certainly does contain a polymorphic site are excluded. A polymorphic site at the hybridization site may cause unequal hybridization at some alleles or may inhibit hybridization entirely, resulting in preferential amplification of certain alleles. These embodiments are an improvement over other methods involving targeted amplification and / or selective enrichment, where the sample is a pure genomic sample from a single individual or a mixture of individuals, in that they better preserve the original allele frequency of the sample at each polymorphic locus.
[0362] In some embodiments, PCR (referred to as mini-PCR) is used to generate very short amplicons (see U.S. Application No. 13 / 683,604, filed November 21, 2012; U.S. Publication No. 2013 / 0123120, filed November 18, 2011; U.S. Publication No. 2012 / 0270212, filed November 18, 2011; and U.S. Application No. 61 / 994,791, filed May 16, 2014, each of which is incorporated by reference in its entirety). cfDNA (such as cancer cfDNA released by necrosis or apoptosis) is highly fragmented. For fetal cfDNA, fragment sizes are distributed in an approximately Gaussian manner, with a mean of 160 bp, a standard deviation of 15 bp, a minimum size of approximately 100 bp, and a maximum size of approximately 220 bp. A polymorphic site at a particular target locus can occupy any position, from first to last, among the various fragments derived from that locus. Because cfDNA fragments are short, the likelihood that both primer sites are present—that a fragment of length L contains both forward and reverse primer sites—is the ratio of amplicon length to fragment length. Under ideal conditions, assays with amplicons of 45, 50, 55, 60, 65, or 70 bp will successfully amplify from 72%, 69%, 66%, 63%, 59%, or 56% of the available template fragment molecules, respectively. In certain embodiments, most preferably related to cfDNA from samples of individuals suspected of having cancer, cfDNA is amplified using primers that produce a maximum amplicon length of 85, 80, 75, or 70 bp, or in certain preferred embodiments, 75 bp, and have a melting temperature of 50-65°C, or in certain preferred embodiments, 54-60.5°C. The amplicon length is the distance between the 5' ends of the forward and reverse priming sites. Amplicon lengths shorter than those typically used by those known in the art can result in more efficient measurement of the desired polymorphic locus by requiring only short sequence reads.In one embodiment, a substantial fraction of the amplicons are less than 100 bp, less than 90 bp, less than 80 bp, less than 70 bp, less than 65 bp, less than 60 bp, less than 55 bp, less than 50 bp, or less than 45 bp.
[0363] In some embodiments, the amplification is performed using direct multiplex PCR, sequential PCR, nested PCR, double nested PCR, one-and-a-half PCR, or any combination thereof. Mini-PCR can be performed using one-sided nested PCR, fully nested PCR, one-sided fully nested PCR, one-sided nested PCR, hemi-nested PCR, hemi-nested PCR, triplex hemi-nested PCR, semi-nested PCR, one-sided semi-nested PCR, reverse semi-nested PCR, or one-sided PCR, as described in U.S. Application No. 13 / 683,604, filed November 21, 2012, U.S. Publication No. 2013 / 0123120, U.S. Application No. 13 / 300,235, filed November 18, 2011, U.S. Publication No. 2012 / 0270212, and U.S. Application No. 61 / 994,791, filed May 16, 2014, which are incorporated by reference in their entireties. Any of these methods can be used for mini-PCR, if desired.
[0364] If desired, the extension step of the PCR amplification can be limited in time to reduce amplification from fragments longer than 200, 300, 400, 500, or 1,000 nucleotides, which can result in enrichment of fragmented or shorter DNA (such as fetal DNA or DNA from cancer cells undergoing apoptosis or necrosis) and improved test performance.
[0365] In some embodiments, multiplex PCR is used. In some embodiments, a method for amplifying target loci in a nucleic acid sample involves (i) contacting the nucleic acid sample with a library of primers that simultaneously hybridize to at least 100, 200, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 20,000, 25,000, 30,000, 40,000, 50,000, 75,000, or 100,000 different target loci to generate a reaction mixture, and (ii) subjecting the reaction mixture to primer extension reaction conditions (such as PCR conditions) to generate amplified products containing target amplicons. In some embodiments, at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the target loci are amplified. In various embodiments, less than 60, 50, 40, 30, 20, 10, 5, 4, 3, 2, 1, 0.5, 0.25, 0.1, or 0.05% of the amplified products are primer dimers. In some embodiments, the primers are in solution (e.g., dissolved in a liquid phase rather than a solid phase). In some embodiments, the primers are in solution and not immobilized on a solid support. In some embodiments, the primers are not part of a microarray. In some embodiments, the primers do not comprise a molecular inversion probe (MIP).
[0366] In some embodiments, two or more (e.g., three or four) target amplicons (e.g., amplicons from the mini-PCR methods disclosed herein) are ligated together, and the ligated product is then sequenced. Combining multiple amplicons into a single ligation product increases the efficiency of the subsequent sequencing step. In some embodiments, the target amplicons are less than 150, 100, 90, 75, or 50 base pairs in length before they are ligated. Selective enrichment and / or amplification may involve tagging each individual molecule with a different tag, molecular barcode, tag for amplification, and / or tag for sequencing. In some embodiments, the amplified products are analyzed by sequencing (e.g., by high-throughput sequencing) or by hybridization to arrays such as SNP arrays, ILLUMINA INFINIUM arrays, or AFFYMETRIX gene chips. In some embodiments, nanopore sequencing is used, such as the nanopore sequencing technology developed by Genia (see, e.g., the World Wide Web at geniachip.com / technology, incorporated herein by reference in its entirety). In some embodiments, double-stranded sequencing is used (Schmitt et al., "Detection of ultra-rare mutations by next-generation sequencing," Proc Natl Acad Sci U S A. 109(36):14508-14513, 2012, incorporated herein by reference in its entirety). This approach significantly reduces errors by independently tagging and sequencing each of the two strands of a DNA duplex. Because the two strands are complementary, true mutations are found at the same position in both strands. In contrast, PCR or sequencing errors can be discounted as technical errors because they only result in mutations in one strand. In some embodiments, the method involves tagging both strands of double-stranded DNA with random but complementary double-stranded nucleotide sequences, referred to as double-stranded tags.The double-stranded tag sequence is incorporated into a standard sequencing adapter by first introducing a single-stranded randomized nucleotide sequence into one adapter strand, then using DNA polymerase to extend the opposite strand, generating a complementary double-stranded tag. After ligating the tagged adapter to the sheared DNA, the individually labeled strands are PCR amplified from the asymmetric primer sites on the adapter tail and subjected to paired-end sequencing. In some embodiments, a sample (such as a DNA or RNA sample) is divided into multiple fractions, such as different wells (e.g., wells of a WaferGen SmartChip). By dividing the sample into different fractions (e.g., at least 5, 10, 20, 50, 75, 100, 150, 200, or 300 fractions), the proportion of molecules with mutations is higher in some wells than in the overall sample, thereby increasing the sensitivity of the analysis. In some embodiments, each fraction contains less than 500, 400, 200, 100, 50, 20, 10, 5, 2, or 1 DNA or RNA molecule. In some embodiments, the molecules in each fraction are sequenced separately. In some embodiments, the same barcode (e.g., random or non-human sequence) is added to all molecules in the same fraction (such as by amplification with primers containing the barcode or by ligation of the barcode), and different barcodes are added to molecules in different fractions. The barcoded molecules can be pooled and sequenced together. In some embodiments, the molecules are amplified before pooling and sequencing, such as by using nested PCR. In some embodiments, one forward primer and two reverse primers, or two forward primers and one reverse primer, are used.
[0367] S. Detection Limit In some embodiments, mutations (such as SNVs or CNVs) present in less than 10, 5, 2, 1, 0.5, 0.1, 0.05, 0.01, or 0.005% of the DNA or RNA molecules in a sample (such as a cfDNA or cfRNA sample) are detected (or can be detected). In some embodiments, mutations (such as SNVs or CNVs) present in less than 1,000, 500, 100, 50, 20, 10, 5, 4, 3, or 2 original DNA or RNA molecules (before amplification) in a sample (such as a cfDNA or cfRNA sample from a blood sample) are detected (or can be detected). In some embodiments, mutations (such as SNVs or CNVs) present in only one original DNA or RNA molecule (before amplification) in a sample (such as a cfDNA or cfRNA sample from a blood sample) are detected (or can be detected).
[0368] For example, if the detection limit for a mutation (such as a single nucleotide variant (SNV)) is 0.1%, dividing the fraction into multiple fractions, such as 100 wells, allows for the detection of a mutation present at 0.01%. The majority of wells will contain no copies of the mutation. For the few wells that do contain a mutation, the mutation will be present at a significantly higher percentage of reads. In one example, there are 20,000 initial copies of DNA from the target locus, and two of these copies contain the SNV of interest. If a sample is divided into 100 wells, 98 wells will contain the SNV, and 2 wells will have the SNV at 0.5%. The DNA in each well can be barcoded, amplified, pooled with DNA from other wells, and sequenced. The wells without SNVs can be used to measure the background amplification / sequencing error rate and determine whether the signal from the outlier wells exceeds the background level of noise.
[0369] T. Detection Methods In some embodiments, the amplified products are detected using an array, such as an array, particularly a microarray, that employs probes for one or more chromosomes of interest (e.g., chromosomes 13, 18, 21, X, Y, or any combination thereof. It will be appreciated that commercially available SNP detection microarrays can be used, such as, for example, Illumina (San Diego, CA) GoldenGate, DASL, Infinium, or CytoSNP-12 genotyping assays, or SNP detection microarray products from Affymetrix, e.g., OncoScan microarrays.
[0370] In some embodiments involving sequencing, read depth is the number of sequencing reads that map to a given locus. Read depth may be normalized across the total number of reads. In some embodiments of sample read depth, read depth is the average read depth across the target locus. In some embodiments of locus read depth, read depth is the number of reads measured by the sequencer that map to that locus. Generally, the greater the read depth at a locus, the closer the allele ratio at that locus is to the allele ratio in the original DNA sample. Read depth can be expressed in a variety of different ways, including but not limited to, a percentage or ratio. Thus, for example, on a highly parallel DNA sequencer, such as an Illumina HISEQ that generates 1 million clonal sequences, sequencing a locus 3,000 times would result in a read depth at that locus of 3,000 reads. The percentage of reads at that locus is 3,000 divided by 1 million total reads, or 0.3% of the total reads.
[0371] In some embodiments, allele data is obtained, and the allele data includes quantitative measurements that are indicative of the copy number of a particular allele at a polymorphic locus. In some embodiments, the allele data includes quantitative measurements that are indicative of the copy number of each of the alleles observed at the polymorphic locus. Typically, quantitative measurements are obtained for all possible alleles at a polymorphic locus of interest. For example, any of the methods described in the preceding paragraphs for determining alleles at SNP or SNV loci, such as microarrays, qPCR, DNA sequencing, e.g., high-throughput DNA sequencing, can be used to generate quantitative measurements of the copy number of a particular allele at a polymorphic locus. This quantitative measurement is referred to herein as allele frequency data or measurements of gene allele data. Methods that use allele data are sometimes referred to as quantitative allele methods, in contrast to quantitative methods that exclusively use quantitative data from non-polymorphic loci or from polymorphic loci but not regarding allele identity. When allelic data is measured using high-throughput sequencing, the allelic data typically includes the number of reads for each allele that maps to the locus of interest.
[0372] In some embodiments, non-allelic data are obtained, and the non-allelic data include quantitative measurements that are indicative of the copy number of a particular locus. A locus can be polymorphic or non-polymorphic. In some embodiments when a locus is non-polymorphic, the non-allelic data does not contain information about the relative or absolute amounts of individual alleles that may be present at the locus. Methods that use only non-allelic data (i.e., quantitative data from non-polymorphic alleles, or quantitative data from polymorphic alleles but not regarding the allelic identity of each fragment) are referred to as quantitative methods. Typically, quantitative measurements are obtained for all possible alleles at a polymorphic locus of interest, and a single value is associated with the measured amounts for all alleles at the locus in total. Non-allelic data for a polymorphic locus can be obtained by summing the quantitative alleles for each allele at the locus. When allelic data are measured using high-throughput sequencing, the non-allelic data typically include the number of reads that map to the locus of interest. The sequencing measurements may indicate the relative and / or absolute number of each allele present at that locus, and the non-allelic data includes the sum of reads mapping to that locus, regardless of allelic identity. In some embodiments, the same set of sequencing measurements can be used to generate both allelic and non-allelic data. In some embodiments, the allelic data is used as part of a method to determine copy number at a chromosome of interest, and the generated non-allelic data can be used as part of a different method to determine copy number at the chromosome of interest. In some embodiments, the two methods are statistically orthogonal and, combined, provide a more accurate determination of copy number at a chromosome of interest.
[0373] In some embodiments, obtaining genetic data includes (i) obtaining DNA sequence information by laboratory techniques, e.g., by use of an automated high-throughput DNA sequencer, or (ii) obtaining information previously obtained by laboratory techniques, which is transmitted electronically, e.g., by a computer via the internet or by electronic transmission from a sequencing device.
[0374] Further exemplary sample preparation, amplification and quantification methods are described in U.S. Application No. 13 / 683,604, filed November 21, 2012 (U.S. Publication No. 2013 / 0123120 and U.S. Application No. 61 / 994,791, filed May 16, 2014, which are incorporated by reference in their entireties). These methods can be used in analyzing any of the samples disclosed herein.
[0375] U. Exemplary Quantification Methods for Cell-Free DNA If desired, the amount or concentration of cfDNA or cfRNA can be measured using standard methods.In some embodiments, the amount or concentration of cell-free mitochondrial DNA (cf mDNA) is determined.In some embodiments, the amount or concentration of cell-free DNA (cf nDNA) derived from nuclear DNA is determined.In some embodiments, the amount or concentration of cf mDNA and cf nDNA is determined simultaneously.
[0376] In some embodiments, qPCR is used to measure cf nDNA and / or cfm DNA (Kohler et al., "Levels of plasma circulating cell free nuclear and mitochondrial DNA as potential biomarkers for breast tumors." Mol Cancer 8:105, 2009, 8:doi:10.1186 / 1476-4598-8-105, which is incorporated by reference in its entirety). For example, one or more loci from cf nDNA (such as glyceraldehyde-3-phosphate-dehydrogenase, GAPDH) and one or more loci from cf mDNA (ATPase 8, MTATP 8) can be measured using multiplex qPCR. In some embodiments, fluorescently labeled PCR is used to measure cf nDNA and / or cf mDNA (Schwarzenbach et al., "Evaluation of cell-free tumor DNA and RNA in patients with breast cancer and benign breast disease." Mol Biosys 7:2848-2854, 2011, incorporated herein by reference in its entirety). If desired, normal distribution of the data can be determined using standard methods such as a Shapiro-Wilk test. If desired, cf nDNA and mDNA levels can be compared using standard methods such as a Mann-Whitney U test. In some embodiments, cf nDNA and / or mDNA levels are compared to other established prognostic factors using standard methods such as a Mann-Whitney U test or a Kruskal-Wallis test.
[0377] V. Exemplary RNA Amplification, Quantification, and Analysis Methods Any of the following exemplary methods can be used to amplify and optionally quantify RNA, such as cfRNA, cellular RNA, cytoplasmic RNA, coding cytoplasmic RNA, non-coding cytoplasmic RNA, mRNA, miRNA, mitochondrial RNA, rRNA, or tRNA. In some embodiments, the miRNA is any of the miRNA molecules listed in miRBase, available on the World Wide Web at mirbase.org, the entire contents of which are incorporated herein by reference. Exemplary miRNA molecules include miR-509, miR-21, and miR-146a.
[0378] In some embodiments, reverse transcriptase multiplex ligation-dependent probe amplification (RT-MLPA) is used to amplify RNA. In some embodiments, each set of hybridizing probes consists of two short synthetic oligonucleotides spanning the SNP and one long oligonucleotide (see Li et al., Arch Gynecol Obstet. "Development of noninvasive prenatal diagnosis of trisomy 21 by RT-MLPA with a new set of SNP markers," July 5, 2013, DOI 10.1007 / s00404-013-2926-5; Schouten et al., "Relative quantification of 40 nucleic acid sequences by multiplex ligation-dependent probe amplification," Nucleic Acids Res 30:e57, 2002; Deng et al. (2011) "Non-invasive prenatal diagnosis of trisomy 21 by reverse transcriptase multiplex ligation-dependent probe amplification," Clin. Chem. Lab. Med.49:641-646,2011).
[0379] In some embodiments, RNA is amplified by reverse transcriptase PCR. In some embodiments, RNA is amplified using real-time reverse transcriptase PCR, such as one-step real-time reverse transcriptase PCR using SYBR GREEN I, as previously described (see Li et al., Arch Gynecol Obstet. "Development of noninvasive prenatal diagnosis of trisomy 21 by RT-MLPA with a new set of SNP markers," July 5, 2013, DOI 10.1007 / s00404-013-2926-5; Lo et al., "Plasma placental RNA allelic ratio permits noninvasive prenatal chromosomal aneuploidy detection," Nat Med 13:218-223, 2007; Tsui et al., Systematic micro-array based identification of placental mRNA in maternal plasma: toward non-invasive prenatal gene expression profiling. J Med Genet, 2007, each of which is incorporated by reference in its entirety). 41:461-467, 2004, Gu et al., J. Neurochem. 122:641-649, 2012).
[0380] In some embodiments, a microarray is used to detect RNA. For example, a human miRNA microarray from Agilent Technologies can be used according to the manufacturer's protocol. Briefly, isolated RNA is dephosphorylated and ligated with pCp-Cy3. The labeled RNA is purified and hybridized to an miRNA array containing probes for human mature miRNAs based on Sanger miRBase release 14.0. The array is washed and scanned using a microarray scanner (G2565BA, Agilent Technologies). The intensity of each hybridization signal is evaluated using Agilent Extraction Software v9.5.3. Labeling, hybridization, and scanning can be performed according to the protocol of the Agilent miRNA microarray system (Gu et al., J. Neurochem. 122:641-649, 2012, the entire contents of which are incorporated herein by reference).
[0381] In some embodiments, a TaqMan assay is used to detect RNA. An exemplary assay is the TaqMan Array Human MicroRNA Panel v1.0 (Early Access) (Applied Biosystems), which contains 157 TaqMan MicroRNA assays, each containing a reverse transcription primer, PCR primer, and TaqMan probe (Chim et al., "Detection and characterization of placental microRNAs in maternal plasma," ClinChem. 54(3):482-90, 2008, which is incorporated herein by reference in its entirety).
[0382] If desired, the mRNA splicing pattern of one or more mRNAs can be determined using standard methods (Fackenthal and Godley, Disease Models & Mechanisms 1:37-42, 2008, doi:10.1242 / dmm.000331, which is incorporated herein by reference in its entirety). For example, high-density microarrays and / or high-throughput DNA sequencing can be used to detect mRNA splice variants.
[0383] In some embodiments, whole transcriptome shotgun sequencing or arrays are used to measure the transcriptome.
[0384] W. Exemplary Amplification Methods Improved PCR amplification methods have also been developed that minimize or prevent interference due to amplification of nearby or adjacent target loci in the same reaction volume (such as part of a sample multiplex PCR in which all target loci are amplified simultaneously). These methods can be used to simultaneously amplify nearby or adjacent target loci, which is faster and cheaper than methods that require dividing nearby target loci into different reaction volumes so that the target loci can be amplified separately and avoid interference.
[0385] In some embodiments, amplification of target loci is performed using a polymerase (e.g., a DNA polymerase, an RNA polymerase, or a reverse transcriptase) with low 5' to 3' exonuclease activity and / or low strand displacement activity. In some embodiments, low levels of 5' to 3' exonuclease reduce or prevent degradation of nearby primers (e.g., primers that are not extended or primers that have one or more nucleotides added during primer extension). In some embodiments, low levels of strand displacement activity reduce or prevent displacement of nearby primers (e.g., primers that are not extended or primers that have one or more nucleotides added during primer extension). In some embodiments, target loci that are adjacent to each other (e.g., no bases between the target loci) or nearby (e.g., loci within 50, 40, 30, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 base apart) are amplified. In some embodiments, the 3' end of one locus is within 50, 40, 30, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 base of the 5' end of the next downstream locus.
[0386] In some embodiments, at least 100, 200, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 20,000, 25,000, 30,000, 40,000, 50,000, 75,000, or 100,000 different target loci are amplified, such as by simultaneous amplification in one reaction volume. In some embodiments, at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the amplified products are target amplicons. In various embodiments, the amount of amplified products that are target amplicons is between 50 and 99.5%, e.g., between 60 and 99%, between 70 and 98%, between 80 and 98%, between 90 and 99.5%, or between 95 and 99.5%, inclusive. In some embodiments, at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the target loci are amplified (e.g., at least 5-, 10-, 20-, 30-, 50-, or 100-fold compared to the amount before amplification), e.g., by co-amplification in one reaction volume. In various embodiments, the amount of target loci amplified (e.g., at least 5-, 10-, 20-, 30-, 50-, or 100-fold compared to the amount before amplification) is 50-99.5%, e.g., 60-99%, 70-98%, 80-99%, 90-99.5%, 95-99.9%, or 98-99.99%, inclusive. In some embodiments, fewer non-target amplicons are produced, e.g., fewer amplicons generated from the forward primer from the first primer pair and the reverse primer from the second primer pair. Such unwanted non-target amplicons can be generated using conventional amplification methods, for example, when the reverse primer from a first primer pair and / or the forward primer from a second primer pair are degraded and / or displaced.
[0387] In some embodiments, these methods allow for the use of longer extension times because the polymerase bound to the extended primer is less likely to degrade and / or displace nearby primers (such as the next downstream primer) given the polymerase's low 5' to 3' exonuclease activity and / or low strand displacement activity. In various embodiments, reaction conditions (such as extension time and temperature) are used that allow the extension rate of the polymerase to add nucleotides to the extended primer to be 80, 90, 95, 100, 110, 120, 130, 140, 150, 175, or 200% or more of the number of nucleotides between the 3' end of its primer binding site and the 5' end of the next downstream primer binding site on the same strand.
[0388] In some embodiments, a DNA polymerase is used to generate DNA amplicons using DNA as a template. In some embodiments, an RNA polymerase is used to generate RNA amplicons using DNA as a template. In some embodiments, a reverse transcriptase is used to generate cDNA amplicons using RNA as a template.
[0389] In some embodiments, the low level of 5' to 3' exonuclease activity of the polymerase is less than 80, 70, 60, 50, 40, 30, 20, 10, 5, 1, or 0.1% of the activity of the same amount of Thermus aquaticus polymerase under the same conditions ("Taq" polymerase, a commonly used DNA polymerase from a thermophilic bacterium, PDB 1BGX, EC 2.7.7.7, incorporated herein by reference in its entirety; Murali et al., "Crystal structure of Taq DNA polymerase in complex with an inhibitory Fab: the Fab is directed against an intermediate in the helix-coil dynamics of the enzyme," Proc. Natl. Acad. Sci. USA 95:12562-12567, 1998). In some embodiments, the low level strand displacement activity of the polymerase is less than 80, 70, 60, 50, 40, 30, 20, 10, 5, 1, or 0.1% of the activity of the same amount of Taq polymerase under the same conditions.
[0390] In some embodiments, the polymerase is a PUSHION DNA polymerase, such as PHUSION High Fidelity DNA polymerase (M0530S, New England BioLabs, Inc.) or PHUSION Hot Start Flex DNA polymerase (M0535S, New England BioLabs, Inc.; Frey and Suppman BioChemica. 2:34-35, 1995; Chester and Marshak Analytical Biochemistry. 209:284-290, 1993, each of which is incorporated herein by reference in its entirety). PHUSION DNA polymerase is a Pyrococcus-like enzyme fused to a processivity-enhancing domain. PHUSION DNA polymerase possesses 5' to 3' polymerase activity and 3' to 5' endonuclease activity, generating blunt-ended products. PHUSION DNA polymerase lacks 5' to 3' exonuclease activity and strand displacement activity.
[0391] In some embodiments, the polymerase is a Q5® DNA polymerase, such as Q5® High-Fidelity DNA Polymerase (M0491S, New England BioLabs, Inc.) or Q5® Hot Start High-Fidelity DNA Polymerase (M0493S, New England BioLabs, Inc.). Q5® High-Fidelity DNA polymerase is a high-fidelity, thermostable DNA polymerase with 3' to 5' exonuclease activity fused to a processivity-enhancing Sso7d domain. Q5® High-Fidelity DNA polymerase lacks 5' to 3' exonuclease activity and strand displacement activity.
[0392] In some embodiments, the polymerase is T4 DNA polymerase (M0203S, New England BioLabs, Inc.; Tabor and Struh. (1989). "DNA-Dependent DNA Polymerases," Ausebel et al. (Ed.), Current Protocols in Molecular Biology, 3.5.10-3.5.12. New York: John Wiley & Sons, Inc., 1989; Sambrook et al. Molecular Cloning: A Laboratory Manual. (2nd ed.), 5.44-5.47. Cold Spring Harbor: Cold Spring Harbor Laboratory Press, 1989, each of which is incorporated by reference in its entirety). T4 DNA polymerase catalyzes the synthesis of DNA in the 5' to 3' direction and requires the presence of a template and a primer. This enzyme possesses a 3' to 5' exonuclease activity that is significantly greater than that found in DNA Polymerase I. T4 DNA polymerase lacks 5' to 3' exonuclease activity and strand displacement activity.
[0393] In some embodiments, the polymerase is Sulfolobus DNA Polymerase IV (M0327S, New England BioLabs, Inc.; Boudsocq et al. (2001). Nucleic Acids Res., 29:4607-4616, 2001; McDonald et al. (2006). Nucleic Acids Res., 34:1102-1111, 2006, each of which is incorporated by reference in its entirety). Sulfolobus DNA Polymerase IV is a thermostable Y-family lesion bypass DNA polymerase that efficiently synthesizes DNA across a variety of DNA template lesions (McDonald, JP et al. (2006). Nucleic Acids Res., 34:1102-1111, each of which is incorporated by reference in its entirety). Sulfolobus DNA Polymerase IV lacks 5' to 3' exonuclease activity and strand displacement activity.
[0394] In some embodiments, when primers bind to a region containing a SNP, the primers may bind to and amplify different alleles with different efficiencies, or may bind to and amplify only one allele. For heterozygous subjects, one of the alleles may not be amplified by the primers. In some embodiments, primers are designed for each allele. For example, if there are two alleles (e.g., a biallelic SNP), two primers can be used to bind to the same position of the target locus (e.g., a forward primer to bind to the "A" allele and a forward primer to bind to the "B" allele). Standard methods, such as the dbSNP database, can be used to determine the locations of known SNPs, such as SNP hotspots with high heterozygosity rates.
[0395] In some embodiments, the amplicons are similarly sized. In some embodiments, the length of the target amplicon ranges from less than 100, 75, 50, 25, 15, 10, or 5 nucleotides. In some embodiments (e.g., amplification of target loci in fragmented DNA or RNA), the length of the target amplicon is 50-100 nucleotides, e.g., 60-80 nucleotides or 60-75 nucleotides, inclusive. In some embodiments (e.g., amplification of multiple target loci of exons or entire genes), the length of the target amplicon is 100-500 nucleotides, e.g., 150-450 nucleotides, 200-400 nucleotides, 200-300 nucleotides, or 300-400 nucleotides, inclusive.
[0396] In some embodiments, multiple target loci are simultaneously amplified using a primer pair comprising a forward and reverse primer for each target locus being amplified in the reaction volume. In some embodiments, one round of PCR is performed using one primer per target locus, and then a second round of PCR is performed using one primer pair per target locus. For example, the first round of PCR can be performed using one primer per target locus, such that all primers bind to the same strand (e.g., using a forward primer for each target locus). This allows PCR to amplify in a linear manner, reducing or eliminating amplification bias between amplicons due to sequence or length differences. In some embodiments, amplicons are then amplified using a forward and reverse primer for each target locus.
[0397] X. Exemplary Primer Design Methods If desired, multiplex PCR can be performed using primers with reduced likelihood of forming primer dimers. In particular, highly multiplexed PCR can often result in the generation of a very high percentage of product DNA resulting from unproductive side reactions such as primer dimer formation. In one embodiment, specific primers that are most likely to cause unproductive side reactions can be removed from the primer library, resulting in a primer library that increases the percentage of amplified DNA that maps to the genome. The process of removing problematic primers, i.e., primers that are particularly likely to stabilize dimers, unexpectedly enabled a very high PCR multiplexing level for subsequent analysis by sequencing.
[0398] There are several methods for selecting primers for libraries that minimize the amount of non-mapping primer dimers or other primer interference products. Empirical data indicates that a small number of "bad" primers are responsible for a large amount of non-mapping primer dimer side reactions. Removing these "bad" primers can increase the percentage of sequence reads that map to the target locus. One method for identifying "bad" primers is to look at the sequencing data of DNA amplified by targeted amplification, and removing the most frequently occurring primer dimers can yield a primer library that is significantly less likely to produce by-product DNA that does not map to the genome. Publicly available programs also exist that can calculate the binding energy of various primer combinations; removing those with the highest binding energy will also yield a primer library that is significantly less likely to produce by-product DNA that does not map to the genome.
[0399] In some embodiments for selecting primers, an initial library of candidate primers is created by designing one or more primers or primer pairs for candidate target loci. A set of candidate target loci (e.g., SNPs) can be selected based on publicly available information regarding desirable parameters for target loci, such as the frequency of the SNP or the heterozygosity rate of the SNP within the target set. In one embodiment, PCR primers can be designed using the Primer3 program (available on the World Wide Web at primer3.sourceforge.net:libprimer3 release 2.2.3, incorporated herein by reference in its entirety). If desired, primers can be designed to anneal within a specific annealing temperature range, have a specific range of GC content, have a specific size range, generate target amplicons in a specific size range, and / or have other parameter characteristics. Starting with multiple primers or primer pairs per candidate target locus increases the likelihood that a primer or primer pair will remain in the library for most or all of the target loci. In one embodiment, the selection criteria can require that at least one primer per target gene remain in the library. Then, when the final primer library is used, most or all of the target loci will be amplified. This is desirable for applications such as s...
Claims
1. 1. A method for preparing a preparation of amplified DNA from a biological sample of a patient diagnosed with cancer, useful for determining cancer recurrence or metastasis, comprising: (a) sequencing DNA isolated from hematopoietic cells in a blood or bone marrow sample or fraction thereof of said patient to determine the presence or absence of one or more clonal hematopoietic with undetermined potential (CHIP) mutations; (b) sequencing (i) DNA isolated from a tumor biopsy sample of the patient, or (ii) cell-free DNA isolated from the blood or bone marrow sample, or a fraction thereof, to identify a plurality of patient-specific somatic mutations associated with the cancer; (c) preparing a preparation of amplified DNA by performing targeted multiplex amplification on cell-free DNA isolated from the patient's longitudinally collected biological sample or a fraction thereof to amplify multiple target loci to obtain amplified DNA, wherein each of the target loci spans the patient-specific somatic mutation identified in step (b) and does not span any of the CHIP mutations identified in step (a), and the biological sample is a blood, urine, or bone marrow sample; (d) analyzing the preparation of amplified DNA by sequencing the amplified DNA to determine the presence or absence of the patient-specific somatic mutations, wherein the presence of two or more patient-specific somatic mutations associated with the cancer and the presence of one or more CHIP mutations is indicative of recurrence or metastasis of the cancer.
2. 2. The method of claim 1, wherein step (a) comprises performing whole-exome or whole-genome sequencing on DNA isolated from the buffy coat fraction of the blood or bone marrow sample to determine the presence or absence of one or more CHIP mutations.
3. 2. The method of claim 1, wherein step (a) comprises enriching a panel of genomic loci associated with a bone marrow disorder from DNA isolated from a buffy coat fraction of the blood or bone marrow sample to obtain enriched genomic loci, and subsequently sequencing the enriched genomic loci to determine the presence or absence of one or more CHIP mutations.
4. 2. The method of claim 1, wherein step (b) comprises performing whole-exome or whole-genome sequencing on cell-free DNA isolated from the plasma fraction of the blood or bone marrow sample to identify a plurality of patient-specific somatic mutations associated with the cancer.
5. 2. The method of claim 1, wherein step (b) comprises performing whole-exome or whole-genome sequencing on DNA isolated from the tumor biopsy sample of the patient to identify a plurality of patient-specific somatic mutations associated with the cancer.
6. 2. The method of claim 1, wherein step (b) comprises enriching a panel of genomic loci associated with cancer from cell-free DNA isolated from the plasma fraction of the blood or bone marrow sample to obtain enriched genomic loci, and subsequently sequencing the enriched genomic loci to identify a plurality of patient-specific somatic mutations associated with the cancer.
7. 2. The method of claim 1, wherein step (b) comprises enriching a panel of genomic loci associated with cancer from DNA isolated from the tumor biopsy sample of the patient to obtain enriched genomic loci, and subsequently sequencing the enriched genomic loci to identify a plurality of patient-specific somatic mutations associated with the cancer.
8. 4. The method of claim 3, wherein the panel of genomic loci associated with myeloid disorders is enriched by hybrid capture and / or targeted amplification.
9. 7. The method of claim 6, wherein the panel of cancer-associated genomic loci is enriched by hybrid capture and / or targeted amplification.
10. 9. The method of claim 8, wherein the panel of genomic loci associated with myeloid disorders and / or the panel of genomic loci associated with cancer comprise one or more genomic loci in exons, introns, gene regulatory regions, non-coding RNA, rearranged genes, or combinations thereof.
11. 2. The method of claim 1, wherein the patient-specific somatic mutation associated with the cancer comprises a single nucleotide variant (SNV), a multi-nucleotide variant (MNV), an indel, a gene fusion, a structural variant, or a combination thereof.
12. 10. The method of claim 1, further comprising identifying one or more germline mutations in the patient, wherein the target loci amplified in step (c) do not span the one or more germline mutations.
13. 13. The method of claim 12, wherein the one or more germline mutations are identified by sequencing DNA isolated from the hematopoietic cells in the blood or bone marrow sample or a fraction thereof.
14. 10. The method of claim 1, wherein the cancer is a cancer or tumor of the abdomen or abdominal wall, adrenal gland, anus, appendix, bladder, bone, brain, breast, cervix, chest wall, colon, diaphragm, duodenum, ear, endometrium, esophagus, fallopian tube, gallbladder, gastroesophageal junction, head and neck, kidney, larynx, liver, lung, lymph node, malignant effusion, mediastinum, nasal cavity, omentum, ovary, pancreas, pancreaticobiliary duct, parotid gland, pelvis, penis, pericardium, peritoneum, pleura, prostate, rectum, salivary gland, skin, small intestine, soft tissue, spleen, stomach, thyroid, tongue, trachea, ureter, uterus, vagina, vulva, or Whipple resection.
15. 10. The method of claim 1, wherein the cancer is breast cancer, colon cancer, gastrointestinal cancer, renal cancer, lung cancer, multiple myeloma, ovarian cancer, or pancreatic cancer.
16. 10. The method of claim 1, further comprising longitudinally collecting a plurality of biological samples from the patient and repeating steps (c) and (d) for each of the biological samples.
17. 17. The method of claim 16, wherein the plurality of biological samples are collected after the patient has been treated with surgery, primary chemotherapy, and / or adjuvant therapy.
18. 2. The method of claim 1, wherein the presence of two or more patient-specific somatic mutations associated with the cancer and the presence of two or more CHIP mutations is indicative of recurrence or metastasis of the cancer.
19. 1. A method for preparing a preparation of amplified DNA from a biological sample of a patient diagnosed with cancer, useful for determining cancer recurrence or metastasis, comprising: (a) sequencing (i) DNA isolated from a tumor biopsy sample of the patient, or (ii) cell-free DNA isolated from a blood or bone marrow sample, or a fraction thereof, of the patient, to identify a plurality of patient-specific somatic mutations associated with the cancer; (b) preparing a preparation of amplified DNA by performing targeted multiplex amplification on cell-free DNA isolated from the patient's longitudinally collected biological sample or a fraction thereof to amplify multiple target loci to obtain amplified DNA, wherein each of the target loci spans the patient-specific somatic mutation identified in step (a), and the biological sample is a blood, urine, or bone marrow sample; (c) analyzing the preparation of amplified DNA by sequencing the amplified DNA to determine the presence or absence of the patient-specific somatic mutation; (d) sequencing DNA isolated from hematopoietic cells in the patient's biological sample or a fraction thereof to determine the presence or absence of one or more CHIP mutations, wherein the presence of two or more patient-specific somatic mutations associated with the cancer and the presence of one or more CHIP mutations is indicative of recurrence or metastasis of the cancer.
20. 1. A method for sequencing DNA from a biological sample of a patient diagnosed with cancer, comprising performing whole exome or whole genome sequencing on DNA isolated from hematopoietic cells in a blood or bone marrow sample or fraction thereof of the patient to determine the presence or absence of one or more CHIP mutations, and identifying the patient as having an increased risk of disease progression due to the presence of one or more CHIP mutations.