Cancer detection and monitoring methods using personalized detection of circulating tumor DNA

The method of creating amplicons from patient-specific nucleotide variants in blood or urine samples addresses the invasiveness and sensitivity issues of current detection methods, enabling early and accurate cancer recurrence or metastasis detection.

JP7855659B2Active Publication Date: 2026-05-08NATERA INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NATERA INC
Filing Date
2024-10-15
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Current methods for detecting early cancer recurrence or metastasis, such as imaging and tissue biopsy, are invasive and not sufficiently sensitive.

Method used

A method involving the creation of amplicons from patient-specific single nucleotide variant loci in nucleic acids from blood or urine samples, followed by sequencing to detect specific single nucleotide variants as indicators of early recurrence or metastasis.

Benefits of technology

Enables early detection of cancer recurrence or metastasis with high sensitivity and specificity, potentially before clinical recurrence is detectable by imaging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007855659000077
    Figure 0007855659000077
  • Figure 0007855659000078
    Figure 0007855659000078
  • Figure 0007855659000079
    Figure 0007855659000079
Patent Text Reader

Abstract

To provide methods for detecting single nucleotide variants in breast cancer, bladder cancer, or colorectal cancer; and to provide additional methods and compositions, such as reaction mixtures and solid supports comprising clonal populations of nucleic acids.SOLUTION: For example, provided here is a method for monitoring and detection of early relapse or metastasis of breast cancer, bladder cancer, or colorectal cancer, comprising :generating a set of amplicons by performing a multiplex amplification reaction on nucleic acids isolated from a sample of blood or urine or a fraction thereof from a patient who has been treated for a breast cancer, bladder cancer, or colorectal cancer, each amplicon of the set of amplicons spanning at least one single nucleotide variant locus of a set of patient-specific single nucleotide variant loci associated with the breast cancer, bladder cancer, or colorectal cancer; and determining the sequence of at least a segment of each amplicon of the set of amplicons that comprises a patient-specific single nucleotide variant locus, detection of one or more patient-specific single nucleotide variants being indicative of early relapse or metastasis of breast cancer, bladder cancer, or colorectal cancer.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference of related applications This application claims priority to U.S. Provisional Application No. 62 / 657,727 filed on 14 April 2018, U.S. Provisional Application No. 62 / 669,330 filed on 9 May 2018, U.S. Provisional Application No. 62 / 693,843 filed on 3 July 2018, U.S. Provisional Application No. 62 / 715,143 filed on 6 August 2018, U.S. Provisional Application No. 62 / 746,210 filed on 16 October 2018, U.S. Provisional Application No. 62 / 777,973 filed on 11 December 2018, and U.S. Provisional Application No. 62 / 804,566 filed on 12 February 2019. Each of these applications cited above is incorporated herein by reference in its entirety. [Background technology]

[0002] The detection of early cancer recurrence or metastasis has traditionally relied on imaging and tissue biopsy. While tumor tissue biopsy is invasive and carries a risk of potentially contributing to metastasis or surgical complications, imaging-based detection is not sufficiently sensitive to detect early recurrence or metastasis. Better and less invasive methods are needed to detect cancer recurrence or metastasis. [Overview of the project] [Problems that the invention aims to solve]

[0003] One aspect of the present invention described herein is a method for monitoring and detecting early recurrence or metastasis of cancer (e.g., breast cancer, bladder cancer, or colorectal cancer), comprising creating a set of amplicons by performing multiple amplification reactions on nucleic acids isolated from a blood or urine sample or fraction thereof from a patient treated for cancer (e.g., breast cancer, bladder cancer, or colorectal cancer), wherein each amplicon in the set of amplicons is at least one single nucleotide variant locus from a set of patient-specific single nucleotide variant loci associated with cancer (e.g., breast cancer, bladder cancer, or colorectal cancer). A method comprising creating a mononucleotide variant locus and determining the sequence of at least one segment of each amplicon in a set of amplicons containing a patient-specific mononucleotide variant locus, wherein the detection of one or more (or two or three or four or five or six or seven or eight or nine or ten or more) patient-specific mononucleotide variants is an indicator of early recurrence or metastasis of cancer (e.g., breast cancer, bladder cancer, or colorectal cancer). [Means for solving the problem]

[0004] In addition to breast cancer, bladder cancer, and colorectal cancer, the methods described herein are also applicable to other types of cancer, such as acute lymphoblastic leukemia, acute myeloid leukemia, adrenocortical carcinoma, AIDS-related cancer, AIDS-related lymphoma, anal cancer, appendiceal cancer, astrocytoma, atypical teratoma / rhabdoid tumor, basal cell carcinoma, brainstem glioma, brain tumor (brainstem glioma, atypical teratoma / rhabdoid tumor of the central nervous system, germ cell tumor of the central nervous system, astrocytoma, craniopharyngioma, ependymoblastoma, ependymoldoma, medulloblastoma, medullary epithelioma, moderately differentiated pineal parenchymal tumor, supratentorial primitive neuroectoderm tumor, and pineal bud tumor. (including cell tumors), bronchial tumors, Burkitt lymphoma, cancer of unknown primary site, carcinoid tumors, carcinomas of unknown primary site, atypical teratomas of the central nervous system / rhabdoid tumors, germ blastomas of the central nervous system, cervical cancer, childhood cancer, chordoma, chronic lymphocytic leukemia, chronic myeloid leukemia, chronic myeloproliferative disorders, colon cancer, craniopharyngioma, cutaneous T-cell lymphoma, pancreatic islet cell tumors, endometrial cancer, ependymoblastoma, ependymoma, esophageal cancer, nasal neuroblastoma, Ewing's sarcoma, extracranial germ cell tumors, extragonadal germ cell tumors, extrahepatic cholangiocarcinoma, gallbladder cancer, gastric cancer (gastric (stomach)) Cancer, gastrointestinal carcinoid tumor, gastrointestinal stromal cell tumor, gastrointestinal stromal tumor (GIST), choriocarcinoma of pregnancy, glioma, hairy cell leukemia, head and neck cancer, cardiac cancer, Hodgkin lymphoma, hypopharyngeal cancer, intraocular melanoma, islet cell tumor, Kaposi's sarcoma, kidney cancer, Langerhans cell histiocytosis, laryngeal cancer, lip cancer, liver cancer, malignant fibrous histiocytoma, bone cancer, medulloblastoma, medullary epithelioma, melanoma, Merkel cell carcinoma, Merkel cell carcinoma, mesothelioma, metastatic cervical squamous cell carcinoma of unknown primary origin, mouth cancer, multiple endocrine neoplasia syndrome, multiple myeloma, multiple myeloma / plasmacytoma, mycosis fungoides, myelodysplastic syndrome, myeloproliferative neoplasm, nasal cavity cancer, nasopharyngeal cancer, neuroblastoma, non-Hodgkin lymphoma, non-melanoma, skin cancer, non-small cell lung cancer, mouth cancerCancer, oral cancer, oropharyngeal cancer, osteosarcoma, other brain and spinal cord tumors, ovarian cancer, epithelial ovarian cancer, ovarian germ cell tumor, low-grade ovarian tumor, pancreatic cancer, papillomatosis, paranasal sinus cancer, parathyroid cancer, pelvic cancer, penile cancer, pharyngeal cancer, moderately differentiated pineal parenchymal tumor, pineoblastoma, pituitary tumor, plasmacytoma / multiple myeloma, pleuropulmonary blastoma, primary central nervous system (CNS) lymphoma, primary hepatocellular carcinoma, prostate cancer, rectal cancer, kidney cancer, renal cell carcinoma, renal cell carcinoma, airway cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, Sézary syndrome, small cell lung cancer, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma, cervical squamous cell carcinoma, stomach cancer It can also be used to monitor and detect early recurrence or metastasis of cancer, supratentorial primordial neuroectodermal tumors, T-cell lymphomas, testicular cancers, pharyngeal cancers, thymic carcinomas, thymomas, thyroid cancers, transitional cell carcinomas, transitional cell carcinomas of the renal pelvis and ureters, trophoblastic tumors, ureteral cancers, urinary tract cancers, uterine cancers, uterine sarcomas, vaginal cancers, vulvar cancers, waldenström macroglobulinemia, or Wilms' tumors.

[0005] In some embodiments, nucleic acids are isolated from the patient's tumor before determining the sequence of at least one segment of each amplicon in a set of amplicons for a blood or urine sample or fraction thereof, and somatic mutations are identified in the tumor for a set of patient-specific single nucleotide variant loci, and single nucleotide variants.

[0006] In some embodiments, the method involves collecting blood or urine samples from patients over a long period and sequencing them.

[0007] In some embodiments, at least two or at least five SNVs were detected, and the presence of at least two or at least five SNVs is an indicator of early recurrence or metastasis of breast cancer, bladder cancer, or colorectal cancer.

[0008] In some embodiments, breast cancer, bladder cancer, or colorectal cancer is stage 1 or stage 2 breast cancer, bladder cancer, or colorectal cancer. In some embodiments, breast cancer, bladder cancer, or colorectal cancer is stage 3 or stage 4 breast cancer, bladder cancer, or colorectal cancer.

[0009] In some embodiments, the individual is treated surgically before the isolation of a blood or urine sample.

[0010] In some embodiments, the individuals are treated with chemotherapy before the isolation of blood or urine samples.

[0011] In some embodiments, the individual is treated with an adjuvant or neoadjuvant before isolation of a blood or urine sample.

[0012] In some embodiments, the individual is treated with radiotherapy before the isolation of blood or urine samples.

[0013] In some embodiments, the method further comprises administering a compound to an individual, the compound being known to be specifically effective in treating breast cancer, bladder cancer, or colorectal cancer having one or more of the determined single nucleotide variants.

[0014] In some embodiments, the method further includes determining the variant allele frequency for each single nucleotide variant from sequencing.

[0015] In some embodiments, the treatment plan for breast cancer, bladder cancer, or colorectal cancer is determined based on the determination of variant allele frequencies.

[0016] In some embodiments, the method further comprises administering a compound to an individual, the compound being known to be specifically effective in treating breast, bladder, or colorectal cancer in which one of the single nucleotide variants has a variable allele frequency greater than at least half of the other single nucleotide variants determined.

[0017] In some embodiments, the sequence is determined by high-throughput DNA sequencing of multiple single-nucleotide variant loci.

[0018] In some embodiments, the method further comprises detecting clonal single-nucleotide variants in breast cancer, bladder cancer, or colorectal cancer by determining the variant allele frequency for each SNV locus based on the sequences of multiple copies of a series of amplicons, wherein a relatively high allele frequency compared to other single-nucleotide variants at multiple single-nucleotide variant loci is an indicator of a clonal single-nucleotide variant in breast cancer, bladder cancer, or colorectal cancer.

[0019] In some embodiments, the method further comprises administering the compound to an individual that targets one or more clonal single nucleotide variants but does not target other single nucleotide variants.

[0020] In some embodiments, a variant allele frequency greater than 1.0% is an indicator of a clonal single-nucleotide variant.

[0021] In some embodiments, the method further comprises forming an amplification reaction mixture by combining a polymerase, nucleotide triphosphates, nucleic acid fragments from a nucleic acid library prepared from a sample, and a set of primers that each bind within 150 base pairs of a single nucleotide variant locus or a set of primer pairs that each extend to a region of 160 base pairs or less containing a single nucleotide variant locus; and subjecting the amplification reaction mixture to amplification conditions to create a set of amplicons.

[0022] In some embodiments, determining whether a single nucleotide variant is present in a sample involves, at least in part, identifying a confidence value for each allele determination for each set of single nucleotide variant loci, based on read depth for each locus.

[0023] In some embodiments, a single nucleotide variant call is made if the confidence level for the presence of a single nucleotide variant is greater than 90%.

[0024] In some embodiments, a single nucleotide variant call is made if the confidence level for the presence of a single nucleotide variant is greater than 95%.

[0025] In some embodiments, the set of single-nucleotide variant loci includes all single-nucleotide variant loci identified in the TCGA and COSMIC datasets for breast cancer, bladder cancer, or colorectal cancer.

[0026] In some embodiments, the set of single-nucleotide dispersion sites includes all of the single-nucleotide dispersion sites identified in the TCGA and COSMIC datasets for breast cancer, bladder cancer, or colorectal cancer.

[0027] In some embodiments, the method is performed with a read depth of at least 1,000 single nucleotide variant loci.

[0028] In some embodiments, a set of single nucleotide variant loci includes 25 to 1,000 single nucleotide variant loci known to be associated with breast cancer, bladder cancer, or colorectal cancer.

[0029] In some embodiments, the efficiency and error rate per cycle are determined for each amplification reaction of the multiple amplification reaction of single-nucleotide variant loci, and the efficiency and error rate are used to determine whether single-nucleotide variants are present in the sample at a set of single-variant loci.

[0030] In some embodiments, the amplification reaction is a PCR reaction, and the annealing temperature is 1–15°C higher than the melting point of at least 50% of the primers in the set of primers.

[0031] In some embodiments, the amplification reaction is a PCR reaction, and the length of the annealing step during the PCR reaction is 15 to 120 minutes.

[0032] In some embodiments, the amplification reaction is a PCR reaction, and the length of the annealing step during the PCR reaction is 15 to 120 minutes.

[0033] In some embodiments, the primer concentration in the amplification reaction is 1 to 10 nM.

[0034] In some embodiments, the primers in a set of primers are designed to minimize primer dimerization.

[0035] In some embodiments, the amplification reaction is a PCR reaction, the annealing temperature is 1–15°C higher than the melting point of at least 50% of the primers in the primer set, the length of the annealing step during the PCR reaction is 15–120 minutes, the primer concentration in the amplification reaction is 1–10 nM, and the primers in the primer set are designed to minimize primer dimerization.

[0036] In some embodiments, the multiple amplification reaction is carried out under restricting primer conditions.

[0037] Another aspect of the present invention described herein relates to a composition comprising a circulating tumor nucleic acid fragment including a universal adapter, wherein the circulating tumor nucleic acid is derived from breast cancer, bladder cancer, or colorectal cancer.

[0038] In some embodiments, the circulating tumor nucleic acid was derived from a sample or fragment of blood or urine from an individual with breast cancer, bladder cancer, or colorectal cancer.

[0039] Another aspect of the present invention described herein relates to a composition comprising a solid support containing a clone assembly of a plurality of nucleic acids, wherein the clone assembly comprises an amplicon prepared from a sample of circulating free nucleic acid, and the circulating tumor nucleic acid was derived from breast cancer, bladder cancer, or colorectal cancer.

[0040] In some embodiments, the circulating free nucleic acids were derived from blood or urine samples or fragments thereof from individuals with breast cancer, bladder cancer, or colorectal cancer.

[0041] In some embodiments, nucleic acid fragments in different clone assemblies include the same universal adapter.

[0042] In some embodiments, the nucleic acid clone assembly is derived from nucleic acid fragments from a set of samples from two or more individuals.

[0043] In some embodiments, the nucleic acid fragment includes one of a series of molecular barcodes corresponding to a sample in a set of samples.

[0044] Further aspects of the present invention as described herein are methods for monitoring and detecting early recurrence or metastasis of breast cancer, bladder cancer, or colorectal cancer, comprising: selecting a set of at least eight or sixteen patient-specific single nucleotide variant loci based on somatic mutations identified in tumor samples from patients diagnosed with breast cancer, bladder cancer, or colorectal cancer; collecting one or more blood or urine samples from patients over a long period after the patients have been treated with surgery, first-line chemotherapy and / or adjuvant therapy; and creating a set of amplicons by performing multiple amplification reactions on nucleic acids isolated from each blood or urine sample or a fraction thereof, wherein each amplicon of the set of amplicons The method comprises creating a single nucleotide variant that extends to at least one single nucleotide variant locus of a set of patient-specific single nucleotide variant loci associated with breast cancer, bladder cancer, or colorectal cancer, and determining the sequence of at least one segment of each amplicon in the set of amplicons containing the patient-specific single nucleotide variant locus, and determining that the detection of one or more (or two or three or four or five or six or seven or eight or nine or ten or more) patient-specific single nucleotide variants from a blood or urine sample is an indicator of early recurrence or metastasis of breast cancer, bladder cancer, or colorectal cancer.

[0045] Further aspects of the present invention as described herein are methods for treating breast cancer, bladder cancer, or colorectal cancer, comprising: treating a patient diagnosed with breast cancer, bladder cancer, or colorectal cancer with surgery, first-line chemotherapy and / or adjuvant therapy; collecting one or more blood or urine samples from the patient over a long period; and creating a set of amplicons by performing multiple amplification reactions on nucleic acids isolated from each blood or urine sample or a fraction thereof, wherein each amplicon of the set of amplicons is associated with breast cancer, bladder cancer, or colorectal cancer and comprises at least one single nucleotide from a set of at least eight or sixteen patient-specific single nucleotide variant loci selected based on somatic mutations identified in the patient's tumor sample. A method comprising: creating an otidovariant locus and determining the sequence of at least one segment of each amplicon of a set of amplicons containing a patient-specific single nucleotide variant locus, determining that the detection of one or more (or two or three or four or five or six or seven or eight or nine or ten or more) patient-specific single nucleotide variants from a blood or urine sample is an indicator of early recurrence or metastasis of breast cancer, bladder cancer, or colorectal cancer; and administering to an individual a compound known to be effective in treating breast cancer, bladder cancer, or colorectal cancer having one or more of the single nucleotide variants detected from a blood or urine sample.

[0046] Further aspects of the present invention as described herein are methods for monitoring or predicting a response to treatment for breast cancer, bladder cancer, or colorectal cancer, comprising collecting one or more blood or urine samples over a long period from a patient being treated for breast cancer, bladder cancer, or colorectal cancer, and creating a set of amplicons by performing multiple amplification reactions on nucleic acids isolated from each blood or urine sample or a fraction thereof, wherein each amplicon in the set of amplicons is associated with breast cancer, bladder cancer, or colorectal cancer and is selected based on somatic mutations identified in the patient's tumor sample, comprising at least eight or sixteen patient-specific single nucleic acids A method comprising creating a single nucleotide variant locus that extends to at least one single nucleotide variant locus of a set of rheotide variant loci, and determining the sequence of at least one segment of each amplicon of a set of amplicons containing a patient-specific single nucleotide variant locus, and determining that the detection of one or more (or two or three or four or five or six or seven or eight or nine or ten or more) patient-specific single nucleotide variants from a blood or urine sample is an indicator of poor response to treatment for breast cancer, bladder cancer, or colorectal cancer.

[0047] In some embodiments, the methods described herein involve detecting ctDNA in the plasma of breast cancer patients before and / or during neoadjuvant therapy (e.g., after cycle 1, cycle 2, cycle 3, cycle 4, etc.). In some embodiments, the treatment plan is defined based on ctDNA concentration determination (e.g., presence / absence) and the rate of decline during neoadjuvant therapy.

[0048] In some embodiments, the methods described herein include evaluating the presence and level of ctDNA in each cancer patient (i.e., targeting mutations actually present in the tumor). In some embodiments, the methods described herein include detecting two or more, four or more, ten or more, sixteen or more, thirty-two or more, fifty or more, sixty-four or more, or one hundred or more mutations actually present in the patient's tumor(s).

[0049] According to some embodiments of the present invention, at least 50%, or at least 60%, or at least 70%, or at least 80%, or at least 90%, or about 100% of patients who may have metastatic recurrence (e.g., after neoadjuvant therapy and surgery) have detectable ctDNA at baseline.

[0050] According to some embodiments of the present invention, at least 50%, or at least 60%, or at least 70%, or at least 80%, or at least 90%, or about 100% of patients who may have metastatic recurrence (e.g., after neoadjuvant therapy and surgery) have detectable ctDNA after cycle 1 of neoadjuvant therapy.

[0051] According to some embodiments of the present invention, at least 50%, or at least 60%, or at least 70%, or at least 80%, or at least 90%, or about 100% of patients who may have metastatic recurrence (e.g., after neoadjuvant therapy and surgery) have detectable ctDNA after cycle 2 of neoadjuvant therapy.

[0052] According to some embodiments of the present invention, at least 50%, or at least 60%, or at least 70%, or at least 80%, or at least 90%, or about 100% of patients who may have metastatic recurrence (e.g., after neoadjuvant therapy and surgery) have detectable ctDNA after neoadjuvant therapy and before surgery.

[0053] According to some embodiments of the present invention, at least 50%, or at least 60%, or at least 70%, or at least 80%, or at least 90%, or about 100% of patients who may have metastatic recurrence (e.g., after neoadjuvant therapy and surgery) have detectable ctDNA after surgery.

[0054] According to some embodiments of the present invention, at least 50%, or at least 60%, or at least 70%, or at least 80%, or at least 90%, or about 100% of patients with detectable ctDNA (e.g., after surgery) may have metastasis without further treatment recurrence (e.g., after neoadjuvant therapy and surgery).

[0055] According to some embodiments of the present invention, at least 50%, or at least 60%, or at least 70%, or at least 80%, or at least 90%, or about 100% of patients whose ctDNA is increasing between baseline and cycle 1 or cycle 2, etc., may have metastatic recurrence after surgery if no further treatment is given.

[0056] In some embodiments, the methods described herein include detecting the onset, recurrence, or metastasis of a specific subtype of cancer, including a specific subtype of breast cancer. In some embodiments, the methods described herein include detecting the onset, recurrence, or metastasis of HR+ / HER2- tumors, including HR+ / HER2- breast cancer (e.g., hormone receptor-positive ERα+ and / or PR+). HR+ tumors are typically less aggressive and have a good prognosis with a 5-year survival rate of over 90%.

[0057] In some embodiments, the methods described herein include detecting the development, recurrence, or metastasis of HER2+ tumors, including HER2+ breast cancer (human epidermal growth factor receptor 2 positive). HER2+ tumors are generally more invasive, have a worse prognosis, and are more likely to recur and metastasize than HR+ / HER2- breast cancers.

[0058] In some embodiments, the methods described herein include detecting the onset, recurrence, or metastasis of HR- / HER2 breast cancer, including HR- / HER2 breast cancer (TNBC, i.e., triple-negative BC). Triple-negative breast cancer (TNBC) does not express ERα, PR, or HER2. These tumors are the most aggressive and tend to have the worst prognosis of all breast cancer subtypes.

[0059] In some embodiments, the methods described herein can detect patient-specific single nucleotide variants in at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% of patients with early recurrence or metastasis of cancer.

[0060] In some embodiments, the methods described herein can detect patient-specific single nucleotide variants in at least 80%, at least 85%, at least 90%, at least 95%, or at least 98% of patients with early recurrence or metastasis of HER2+ breast cancer.

[0061] In some embodiments, the methods described herein can detect patient-specific single nucleotide variants in at least 80%, at least 85%, at least 90%, at least 95%, or at least 98% of patients with early recurrence or metastasis of triple-negative breast cancer.

[0062] In some embodiments, the methods described herein can detect patient-specific single nucleotide variants in at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% of patients with early recurrence or metastasis of HR+ / HER2- breast cancer.

[0063] In some embodiments, the methods described herein make it possible to detect patient-specific single nucleotide variants in patients with early recurrence or metastasis of cancer at least 100 days, at least 150 days, at least 200 days, at least 250 days, or at least 300 days prior to clinical recurrence or metastasis of cancer detectable by imaging, and / or at least 100 days, at least 150 days, at least 200 days, at least 250 days, or at least 300 days prior to elevation of CA15-3 levels.

[0064] In some embodiments, the methods described herein make it possible to detect patient-specific single nucleotide variants in patients with early recurrence or metastasis of HER2+ breast cancer at least 100 days, at least 150 days, at least 200 days, at least 250 days, or at least 300 days prior to clinical recurrence or metastasis of HER2+ breast cancer detectable by imaging, and / or at least 100 days, at least 150 days, at least 200 days, at least 250 days, or at least 300 days prior to elevation of CA15-3 levels.

[0065] In some embodiments, the methods described herein make it possible to detect patient-specific single nucleotide variants in patients with early recurrence or metastasis of triple-negative breast cancer at least 100 days, at least 150 days, at least 200 days, at least 250 days, or at least 300 days prior to clinical recurrence or metastasis of triple-negative breast cancer detectable by imaging, and / or at least 100 days, at least 150 days, at least 200 days, at least 250 days, or at least 300 days prior to elevation of CA15-3 levels.

[0066] In some embodiments, the methods described herein make it possible to detect patient-specific single nucleotide variants in patients with early recurrence or metastasis of HR+ / HER2- breast cancer at least 100 days, at least 150 days, at least 200 days, at least 250 days, or at least 300 days prior to clinical recurrence or metastasis of HR+ / HER2- breast cancer detectable by imaging, and / or at least 100 days, at least 150 days, at least 200 days, at least 250 days, or at least 300 days prior to elevation of CA15-3 levels.

[0067] In some embodiments, the methods described herein do not detect patient-specific single nucleotide variants in at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or at least 99.9% of patients without early recurrence or metastasis of cancer.

[0068] In some embodiments, the methods described herein do not detect patient-specific single nucleotide variants in at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or at least 99.9% of patients with HER2+ breast cancer that do not have early recurrence or metastasis.

[0069] In some embodiments, the methods described herein do not detect patient-specific single nucleotide variants in at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or at least 99.9% of patients with triple-negative breast cancer who do not have early recurrence or metastasis.

[0070] In some embodiments, the methods described herein do not detect patient-specific single nucleotide variants in at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or at least 99.9% of patients with HR+ / HER2- breast cancer that do not have early recurrence or metastasis.

[0071] In some embodiments, the methods described herein have a specificity of at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or at least 99.9% in detecting early recurrence or metastasis of cancer when two or more patient-specific single nucleotide variants are detected above a predetermined confidence threshold (e.g., 0.95, 0.96, 0.97, 0.98, or 0.99).

[0072] In some embodiments, the methods described herein have at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or at least 99.9% specificity in detecting early recurrence or metastasis of HER2+ breast cancer when two or more patient-specific single nucleotide variants are detected above a predetermined confidence threshold (e.g., 0.95, 0.96, 0.97, 0.98, or 0.99).

[0073] In some embodiments, the methods described herein have a specificity of at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or at least 99.9% in detecting early recurrence or metastasis of triple-negative breast cancer when two or more patient-specific single nucleotide variants are detected above a predetermined confidence threshold (e.g., 0.95, 0.96, 0.97, 0.98, or 0.99).

[0074] In some embodiments, the methods described herein have a specificity of at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or at least 99.9% in detecting early recurrence or metastasis of HR+ / HER2- breast cancer when two or more patient-specific single nucleotide variants are detected above a predetermined confidence threshold (e.g., 0.95, 0.96, 0.97, 0.98, or 0.99).

[0075] In some embodiments, the methods described herein detect patient-specific single nucleotide variants in at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% of patients with early recurrence or metastasis of muscle-invasive bladder cancer (MIBC).

[0076] In some embodiments, the methods described herein detect patient-specific single nucleotide variants in patients with early recurrence or metastasis of cancer at least 100 days, at least 150 days, at least 200 days, or at least 250 days prior to clinical recurrence or metastasis of MIBC detectable by imaging.

[0077] In some embodiments, the methods described herein do not detect patient-specific single nucleotide variants in at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or at least 99.9% of patients without early recurrence or metastasis of MIBC.

[0078] In some embodiments, the methods described herein have a specificity of at least 95%, at least 98%, at least 99%, at least 99%, at least 99.5%, at least 99.8%, or at least 99.9% in detecting early recurrence or metastasis of MIBC when two or more patient-specific single nucleotide variants are detected above a predetermined confidence threshold (e.g., 0.95, 0.96, 0.97, 0.98, or 0.99).

[0079] In addition to, or instead of, single nucleotide variant detection, the methods described herein may also be based on the detection of other genomic variants (e.g., indels, multinucleotide variants, and / or gene fusions).

[0080] Accordingly, further aspects of the present invention as described herein are methods for monitoring and detecting early recurrence or metastasis of breast cancer, bladder cancer or colorectal cancer, comprising: selecting multiple genomic variant loci (e.g., SNVs, indels, multinucleotide variants and gene fusions) based on somatic mutations identified in tumor samples from patients diagnosed with breast cancer, bladder cancer or colorectal cancer; collecting one or more blood or urine samples from patients over a long period after the patients have been treated with surgery, first-line chemotherapy and / or adjuvant therapy; and creating a set of amplicons by performing multiple amplification reactions on nucleic acids isolated from each blood or urine sample or a fraction thereof. A method comprising: creating a set of amplicons in which each amplicon extends to at least one genomic variant locus of a set of patient-specific genomic variant loci associated with breast cancer, bladder cancer, or colorectal cancer; and sequencing at least one segment of each amplicon in the set of amplicons containing the patient-specific genomic variant locus, and determining that the detection of one or more (or two or three or four or five or six or seven or eight or nine or ten or more) patient-specific genomic variants from a blood or urine sample is an indicator of early recurrence or metastasis of breast cancer, bladder cancer, or colorectal cancer.

[0081] Further aspects of the present invention as described herein are methods for treating breast cancer, bladder cancer, or colorectal cancer, comprising treating a patient diagnosed with breast cancer, bladder cancer, or colorectal cancer with surgery, first-line chemotherapy and / or adjuvant therapy, collecting one or more blood or urine samples from the patient over a long period, and creating a set of amplicons by performing multiple amplification reactions on nucleic acids isolated from each blood or urine sample or a fraction thereof, wherein each amplicon of the set of amplicons is associated with breast cancer, bladder cancer, or colorectal cancer and is associated with at least one genomic variant locus from a set of at least eight or sixteen patient-specific genomic variant loci selected based on somatic mutations identified in the patient's tumor sample (e.g., The present invention relates to a method comprising: creating a set of amplicons containing patient-specific genomic variant loci (including SNVs, indels, multinucleotide variants, and gene fusions); determining that the detection of one or more patient-specific genomic variants (or two or three or four or five or six or seven or eight or nine or ten or more) from a blood or urine sample is an indicator of early recurrence or metastasis of breast cancer, bladder cancer, or colorectal cancer; and administering to an individual a compound known to be effective in treating breast cancer, bladder cancer, or colorectal cancer having one or more of the genomic variants detected from a blood or urine sample.

[0082] Further aspects of the present invention as described herein are methods for monitoring or predicting a response to treatment for breast cancer, bladder cancer, or colorectal cancer, comprising: collecting one or more blood or urine samples over a long period from a patient being treated for breast cancer, bladder cancer, or colorectal cancer; and creating a set of amplicons by performing multiple amplification reactions on nucleic acids isolated from each blood or urine sample or a fraction thereof, wherein each amplicon in the set of amplicons is associated with breast cancer, bladder cancer, or colorectal cancer, and comprises at least eight or sixteen patient-specific genomic variant loci selected based on somatic mutations identified in the patient's tumor sample. A method comprising: creating a set of amplicons containing patient-specific genomic variant loci (e.g., SNVs, indels, multinucleotide variants, and gene fusions) that extends to at least one genomic variant locus (e.g., SNVs, indels, multinucleotide variants, and gene fusions); and determining the sequence of at least one segment of each amplicon in a set of amplicons containing patient-specific genomic variant loci; and determining that the detection of one or more (or two or more, or three or more, or four or more, or five or more, or six or more, or seven or more, or eight or more, or nine or more, or ten or more) patient-specific genomic variants from a blood or urine sample is an indicator of poor response to treatment for breast cancer, bladder cancer, or colorectal cancer.

[0083] In addition to, or instead of, patient-specific genomic variants, the methods described herein may also be based on the detection of recurrent cancer-associated mutations (e.g., hotspot cancer mutations, drug resistance markers, cancer panel mutations) that are common in many cancer patients.

[0084] Accordingly, further aspects of the present invention as described herein are methods for monitoring and detecting early recurrence or metastasis of breast cancer, bladder cancer or colorectal cancer, comprising: selecting multiple mutations associated with recurrent cancer; collecting one or more blood or urine samples from a patient over a long period after the patient has been treated with surgery, first-line chemotherapy and / or adjuvant therapy; and creating a set of amplicons by performing multiple amplification reactions on nucleic acids isolated from each blood or urine sample or fraction thereof, wherein each amplicon of the set of amplicons is associated with breast cancer, A method comprising: creating a set of amplicons containing recurrent cancer-associated mutations that are spread across at least one of a set of recurrent mutations associated with bladder cancer or colorectal cancer; and determining the sequence of at least one segment of each amplicon in the set of amplicons containing the recurrent cancer-associated mutations, wherein the detection of one or more (or two or three or four or five or six or seven or eight or nine or ten or more) recurrent cancer-associated mutations from a blood or urine sample is an indicator of early recurrence or metastasis of breast cancer, bladder cancer, or colorectal cancer.

[0085] Further aspects of the present invention as described herein are methods for treating breast cancer, bladder cancer, or colorectal cancer, comprising: treating a patient diagnosed with breast cancer, bladder cancer, or colorectal cancer with surgery, first-line chemotherapy and / or adjuvant therapy; collecting one or more blood or urine samples from the patient over a long period; and creating a set of amplicons by performing multiple amplification reactions on nucleic acids isolated from each blood or urine sample or a fraction thereof, wherein each amplicon of the set of amplicons is at least one recurrent cancer-related mutation (e.g., hotspot cancer mutation, drug resistance marker) from a set of at least eight or sixteen recurrent mutations associated with breast cancer, bladder cancer, or colorectal cancer. The present invention relates to a method comprising: creating a set of amplicons containing mutations associated with recurrent cancer (cancer panel mutations), determining that the detection of one or more (or two or three or four or five or six or seven or eight or nine or ten or more) recurrent cancer-associated mutations from a blood or urine sample is an indicator of early recurrence or metastasis of breast cancer, bladder cancer, or colorectal cancer; and administering to an individual a compound known to be effective in treating breast cancer, bladder cancer, or colorectal cancer having one or more of the recurrent cancer-associated mutations detected from a blood or urine sample.

[0086] Further aspects of the present invention as described herein are methods for monitoring or predicting a response to treatment for breast cancer, bladder cancer, or colorectal cancer, comprising collecting one or more blood or urine samples over a long period from patients being treated for breast cancer, bladder cancer, or colorectal cancer, and creating a set of amplicons by performing multiple amplification reactions on nucleic acids isolated from each blood or urine sample or a fraction thereof, wherein each amplicon in the set of amplicons is at least one of a set of at least eight or sixteen recurrent mutations associated with breast cancer, bladder cancer, or colorectal cancer. A method comprising: creating a set of amplicons containing mutations associated with oncogenesis (e.g., hotspot cancer mutations, drug resistance markers, cancer panel mutations); determining the sequence of at least one segment of each amplicon in a set of amplicons containing mutations associated with recurrent cancer; and determining that the detection of one or more (or two or three or four or five or six or seven or eight or nine or ten or more) mutations associated with recurrent cancer from a blood or urine sample is an indicator of poor response to treatment for breast cancer, bladder cancer, or colorectal cancer.

[0087] In addition to first identifying somatic mutations from tumor samples of patients diagnosed with breast cancer, bladder cancer, or colorectal cancer, the methods described herein may also be based on identifying somatic mutations from other biological samples of the patient, such as blood, serum, plasma, urine, hair, tears, saliva, skin, fingernails, feces, bile, lymph, cervical mucus, or semen.

[0088] Accordingly, further aspects of the present invention as described herein are methods for monitoring and detecting early recurrence or metastasis of breast cancer, bladder cancer, or colorectal cancer, comprising: selecting multiple genomic variant loci (e.g., SNVs, indels, multinucleotide variants, and gene fusions) based on somatic mutations identified in a biological sample (e.g., blood, serum, plasma, urine, hair, tears, saliva, skin, fingernails, feces, bile, lymph, cervical mucus, or semen) containing cancer-related mutations from a patient diagnosed with breast cancer, bladder cancer, or colorectal cancer; collecting one or more blood or urine samples from the patient over a long period after the patient has been treated with surgery, first-line chemotherapy, and / or adjuvant therapy; and performing multiple amplification reactions on nucleic acids isolated from each blood or urine sample or a fraction thereof. A method comprising: creating a set of amplicons such that each amplicon in the set of amplicons extends to at least one of the patient-specific genomic variant loci associated with breast cancer, bladder cancer, or colorectal cancer; and determining the sequence of at least one segment of each amplicon in the set of amplicons containing the patient-specific genomic variant loci, thereby determining that the detection of one or more (or two or three or four or five or six or seven or eight or nine or ten or more) patient-specific genomic variants from a blood or urine sample is an indicator of early recurrence or metastasis of breast cancer, bladder cancer, or colorectal cancer.

[0089] Further aspects of the present invention as described herein are methods for treating breast cancer, bladder cancer, or colorectal cancer, comprising: treating a patient diagnosed with breast cancer, bladder cancer, or colorectal cancer with surgery, first-line chemotherapy and / or adjuvant therapy; collecting one or more blood or urine samples from the patient over a long period; and creating a set of amplicons by performing multiple amplification reactions on nucleic acids isolated from each blood or urine sample or a fraction thereof, wherein each amplicon of the set of amplicons is associated with breast cancer, bladder cancer, or colorectal cancer, and comprises at least eight or sixteen patient-specific genomic variant loci selected based on somatic mutations identified in the patient's biological samples (e.g., blood, serum, plasma, urine, hair, tears, saliva, skin, fingernails, feces, bile, lymph, cervical mucus, or semen) containing cancer-related mutations. A method comprising: creating a set of amplicons containing patient-specific genomic variant loci (e.g., SNVs, indels, multinucleotide variants, and gene fusions) that extend to at least one genomic variant locus (e.g., SNVs, indels, multinucleotide variants, and gene fusions); determining the sequence of at least one segment of each amplicon in a set of amplicons containing patient-specific genomic variant loci; determining that the detection of one or more (or two or three or four or five or six or seven or eight or nine or ten or more) patient-specific genomic variants from a blood or urine sample is an indicator of early recurrence or metastasis of breast cancer, bladder cancer, or colorectal cancer; and administering to an individual a compound known to be effective in treating breast cancer, bladder cancer, or colorectal cancer having one or more of the genomic variants detected from a blood or urine sample.

[0090] Further aspects of the present invention as described herein are methods for monitoring or predicting a response to treatment for breast cancer, bladder cancer, or colorectal cancer, comprising collecting one or more blood or urine samples over a long period from patients being treated for breast cancer, bladder cancer, or colorectal cancer, and creating a set of amplicons by performing multiple amplification reactions on nucleic acids isolated from each blood or urine sample or a fraction thereof, wherein each amplicon in the set of amplicons is associated with breast cancer, bladder cancer, or colorectal cancer and is selected based on somatic mutations identified in a patient's biological sample (e.g., blood, serum, plasma, urine, hair, tears, saliva, skin, fingernails, feces, bile, lymph, cervical mucus, or semen) containing cancer-related mutations. The method comprises creating, and determining the sequence of at least one segment of each amplicon in a set of amplicons containing a patient-specific genomic variant locus (e.g., SNVs, indels, multinucleotide variants, and gene fusions), which extends to at least one genomic variant locus (e.g., SNVs, indels, multinucleotide variants, and gene fusions) from a set of at least eight or sixteen patient-specific genomic variant loci, and determining that the detection of one or more (or two or three or four or five or six or seven or eight or nine or ten or more) patient-specific genomic variants from a blood or urine sample is an indicator of poor response to treatment for breast cancer, bladder cancer, or colorectal cancer.

[0091] Other embodiments, features, and advantages of the disclosed invention will be apparent from the following detailed description and claims. [Brief explanation of the drawing]

[0092] The patent or application file must include at least one drawing in color. A copy of the published patent or patent application containing the color drawing will be provided by the Office upon request and payment of the necessary fees.

[0093] The embodiments disclosed herein will be further described with reference to the accompanying drawings, and similar structures will be referred to by similar figures throughout some of the drawings. The drawings shown are not necessarily to scale, but rather are generally emphasized when illustrating the principles of the embodiments disclosed herein.

[0094] [Figure 1] This is a workflow diagram. [Figure 2] Upper panel: Number of SNVs per sample; Lower panel: Working assay sorted by driver category. [Figure 3] Measured cfDNA concentration. Each data point refers to a plasma sample. [Figure 4] This sample shows a good correlation between previously determined tissue VAF measurements (x-axis) and the current measurements using mPCR-NGS (y-axis). Each sample is represented by a separate rectangle, and the VAF data points are color-coded according to the tissue sub-section. [Figure 5] Samples showing a weak correlation between previously determined tissue VAF measurements (x-axis) and the current measurements using mPCR-NGS (y-axis). Each sample is shown as a separate rectangle, and the VAF data points are color-coded according to the tissue sub-section. [Figure 6A-B] Depth of the read histogram as a function of the obtained calls. Top: This assay did not detect expected plasma SNVs. Bottom: This assay detected expected plasma SNVs. [Figure 7] The number of SNVs detected in plasma by histological type. [Figure 8] SNV detection in plasma (left) and sample (right) according to tumor stage. [Figure 9] Plasma VAF as a function of tumor stage and SNV clonality. [Figure 10] The number of SNVs detected in plasma from each sample as a function of the amount of cfDNA input. [Figure 11]Plasma VAF as a function of mean tumor VAF. Mean tumor VAF was calculated across all tumor sub-sub [Figure 12] The clonal ratio (red vs. blue) and mutant variant allele frequency (MutVAF) for each detected SNV are shown. All SNVs detected from each sample are placed in a single column, and the sample is classified by tumor stage (pTNM stage). Samples in which no SNVs were detected are included. The clonal ratio is defined as the ratio between the number of tumor segments in which SNVs were observed and the total number of segments analyzed from that tumor. [Figure 13] The clonal status (blue for clones, red for subclones) and mutant variant allele frequency (MutVAF) of each detected SNV are shown. All SNVs detected from each sample are placed in a single column, and the sample is classified by tumor stage (pTNM stage). Samples in which no SNVs were detected are included. Clonal status was determined by PyCloneCluster using whole exome sequencing data from tumor tissue. [Figure 14] The clonal status (blue for clones, red for subclones) and mutant variant allele frequency (MutVAF) of each detected SNV are shown. The upper panel shows only clonal SNVs, and the lower panel shows only subclonal SNVs. All SNVs detected from each sample are placed in a single column, and the sample is classified by tumor stage (pTNM stage). Samples in which no SNVs were detected are included. Clonal status was determined by PyCloneCluster using whole exome sequencing data from tumor tissue. [Figure 15] This shows the number of SNVs detected in plasma as a function of histological type and tumor size. Histological type and tumor stage were determined by pathology reports. Each data point is color-coded by size, with red indicating the largest tumor size and blue indicating the smallest tumor size. [Figure 16]This is a table of cfDNA analysis results showing DNA concentration, genomic copy equivalents in library preparations, plasma hemolysis grade, and cDNA profile for all samples. [Figure 17] This is a table of SNVs detected in plasma for each sample. [Figure 18] This is a table of further SNVs detected in plasma. [Figure 19] This is an example of a detection assay for plasma samples at the time of relapse and its background allele fraction (LTX103). [Figure 20A-B] Schematic diagrams of clinical and molecular protocols. [Figure 21] Exam overview. [Figure 22] A summary of patient surveillance and plasma collection data over a 36-month period. [Figure 23A-B] Recurrence risk is stratified by postoperative ctDNA status. [Figure 24A-B] Post-therapy recurrence risk is stratified by postoperative ctDNA status. [Figure 25] The effectiveness of adjuvant therapy in preventing recurrence. [Figure 26A-B] Time to release based on radiology and ctDNA. [Figure 27A-D] Early detection of recurrence and prediction of treatment response. [Figure 28] A schematic diagram of clinical sample collection. [Figure 29] Plasma sequencing QC. [Figure 30A-F] Early detection of recurrence. [Figure 31] Recurrence-free survival rates and ctDNA status at diagnosis and after cystectomy. [Figure 32A-B] Neoadjuvant therapy response. [Figure 33] Signatera (RUO) process. [Figure 34] Plasma sequencing QC. [Figure 35] Sensitivity for single SNV detection. [Figure 36]VAF observed using predicted input versus Signatera (RUO). [Figure 37] Summary of patients for the breast cancer trial in Example 6. [Figure 38A-H] This is a table containing information about the samples analyzed in the test of Example 6. Figure 38A is the first part of the table. Figure 38B is a continuation of the table. Figure 38C is a continuation of the table. Figure 38D is a continuation of the table. Figure 38E is a continuation of the table. Figure 38F is a continuation of the table. Figure 38G is a continuation of the table. Figure 38H is a continuation of the table. [Figure 39] Patient composition in the breast cancer trial of Example 6. WES raw data from 50 patients (including driver variants from 35 patients) was received. 218 plasma samples were received at various time points (between 1 and 8). 108 additional extracted DNA samples were received. Recurrence status was also collected. Blood samples were taken after adjuvant therapy with a 6-month time interval. [Figure 40] Summary of WES analysis and pool design for the breast cancer trial in Example 6. Pool A is based on the Signatera method. Pool B includes 25 patients and is indicated with an asterisk in the box plot. 19 patients in Pool B had low tumor purity. 6 patients had very early-stage HER2- tumors. Pool B includes driver variants. [Figure 41] Plasma samples for the breast cancer test in Example 6. The median plasma volume was 4 mL. The median DNA input was 26 ng. The median DNA input was lower than that of the CRC and MIBC samples (45 ng and 66 ng, respectively). [Figure 42] Sequence quality controls showing the median process error rate per type and median read assay depth for the breast cancer test in Example 6. A total of 326 plasma sequencing samples were processed. The mutation call FP rate is estimated to be 0.28%. [Figure 43]Re-execution of plasma samples for breast cancer testing in Example 6. 319 sequencing samples, including 214 unique plasma samples, are represented for 49 patients. [Figure 44] Results from Pool A for breast cancer testing in Example 6. Of 49 patients, 11 were baseline positive. Three had only one time point. The remaining 8 patients remained positive at all time points. Pool B and the driver yielded similar results. Driver information: 16 recurrent samples had the driver mutation. Eleven recurrent samples underwent at least one assay using the driver. [Figure 45] Summary of 16 patients in whom ctDNA was detected. [Figure 46] Figure 38 provides a graphical explanation of the data corresponding to patient CD047 (TNBC). [Figure 47] Figure 38 provides a graphical explanation of the data corresponding to patient CD033 (TNBC). [Figure 48] Figure 38 illustrates the data corresponding to patient CD037 (HER2+). [Figure 49] Figure 38 illustrates the data corresponding to patient CD040 (HER2+). [Figure 50] Figure 38 illustrates the data corresponding to patient CD048 (HER2-). [Figure 51] Figure 38 illustrates the data corresponding to patient CD005 (HER2-). [Figure 52] Figure 38 illustrates the data corresponding to patient CD036 (HER2-). [Figure 53] Figure 38 illustrates the data corresponding to patient CD044 (HER2-). [Figure 54] Figure 38 provides a diagrammatic explanation of the data corresponding to patient CD049. [Figure 55] Figure 38 provides a diagrammatic explanation of the data corresponding to patient CD029. [Figure 56]Figure 38 provides a diagrammatic explanation of the data corresponding to patient CD026. [Figure 57] Figure 38 provides a diagrammatic explanation of the data corresponding to patient CD017. [Figure 58] Figure 38 illustrates the data corresponding to patient CD031. HW: SHC2, PKD1, COLEC12 [Figure 59] Figure 38 illustrates the data corresponding to patient CD025. In this patient, ctDNA was observed for FGF9 mutations at two consecutive time points. This patient is likely to experience a relapse in the near future. [Figure 60] Patient recruitment and clinical sample collection. For 49 BC women monitored in this study, collected tumor tissue and serial plasma samples were analyzed in a blinded manner using the Signatera® RUO workflow. Exome alterations were determined by paired-end sequencing of FFPE tumor tissue specimens and matched normal DNA. Patient-specific panels were designed, including 16 somatic mutations identified from WES. Plasma samples were processed using these corresponding custom panels. 208 samples were analyzed for ctDNA detection. [Figure 61A-C] Summary of ctDNA analysis results. (A) Summary of treatment regimens for each patient (n=49) along with the results of the analyzed serial plasma samples (n=208). (B) Summary table showing the total number of patients, number of recurrences, percentage detected by ctDNA analysis, and median lead time in days for each breast cancer subtype. (C) Comparison of molecules colored by breast cancer subtypes HR+, HER2+, and TNBC and clinical recurrence using paired Wilcoxon signed-rank test (p<0.001). [Figure 62A-B]CtDNA detection in serial plasma samples predicts recurrence-free survival. (A) Recurrence-free survival rate based on ctDNA detection in any follow-up plasma sample after surgery [HR: 35.84 (7.9626~161.32)] p < 0.001. (B) Recurrence-free survival rate based on ctDNA detection in the first postoperative plasma sample [HR: 11.784 (4.2784~32.457)]. Data are from n=49 patients, and p < 0.001. [Figure 63] (A-E) Plasma levels of ctDNA across multiple plasma time points for five breast cancer patients (one patient per panel). Primary tumor and matched normal whole exome sequencing identified patient-specific somatic mutations. Using an analytically validated Signatera® RUO workflow, each patient-specific assay was designed to target 16 somatic SNVs and indelvariants using large-scale parallel sequencing (median depth per target over 100,000x). Mean VAF is shown as a dark blue circle, and the solid line represents the mean VAF profile over time. Lead time is calculated as the difference between clinical and molecular relapse. CA15-3 levels are graphed over time, with baseline levels shown as light blue shading. (F) Summary of VAF and target counts detected at molecular and clinical relapse for all ctDNA-positive samples, except for patients with only one time point. [Figure 64A-C] Signatera variant selection strategy for a patient-specific panel of 49 patients. (Top) VAF distribution of tumor tissue in patient-specific panels. Different colors represent different subtypes: HER2- (dark blue), triple-negative (orange), and HER2+ (green). (Center) Number of estimated clonal and subclonal variants in patient-specific panels. The median number of clonal variants in the 49 patient-specific panels is 13 out of 16. (Bottom) Number of estimated clonal and subclonal variants in patient WES data. [Figure 65](A-L) Plasma levels of ctDNA across multiple plasma time points in 12 breast cancer patients (11 with recurrence, 1 without recurrence). Primary tumor and matched normal whole exome sequencing identified patient-specific somatic mutations. Using an analytically validated Signatera® workflow, each patient-specific assay was designed to target 16 somatic SNVs and indelvariants using large-scale parallel sequencing (median depth per target over 100,000x). Mean VAF is shown as a dark blue circle, and the solid line represents the mean VAF profile over time. Lead time is calculated as the difference between clinical recurrence and molecular recurrence. CA15-3 levels are graphed over time, with baseline levels shown as light blue shading. [Figure 66] Distribution of VAFs and variant counts. In total, 251 targets were detected in ctDNA-positive plasma samples. The VAFs of the detected targets ranged from 0.01% to 64%, with a median of 0.82%. The number of tumor molecules present in patient plasma samples was calculated using the observed variant VAFs and total DNA molecule counts for each sample. The number of variant molecules detected for the 251 positive targets ranged from 1 to 6500 variant molecules, with a median of 39 molecules. [Figure 67A-D]Signatera Quality Control Process: Quality control was performed at every stage of the workflow. Of a total of 215 plasma samples, 208 passed the inventors' sample QC process, and of the 784 unique assays designed, 767 passed the inventors' QC (corresponding to a total of 3237 assays that passed out of 3328 across all samples). A) cfDNA extracted per 1 mL. cfDNA extracted from each plasma sample was quantified using the Quant-iT High Sensitivity dsDNA Assay Kit. Samples with a quantified cfDNA amount of less than 5 ng were flagged as "WARNING". The cfDNA extracted per 1 mL ranged from 1 to 21.4 ng, with a median of 4.7 ng. B) DNA input amount for library preparation. Up to 66 ng of cfDNA from each plasma sample was used as input to the library preparation protocol. The library DNA input amount ranged from 1 to 66 ng, with a median of 25.02. The purified libraries underwent QC before proceeding to the next step. C) Sequencing coverage. Assays with coverage less than 5000x were excluded from analysis. Subsequently, samples with fewer than 8 successful assays failed sequencing coverage QC. The median read depth for assays that passed coverage QC was 110,000x. D) Sample consistency. To track sample integrity, SNP tracers were used to measure consistency between patient samples. For each plasma sample, a genotyping concordance score was calculated by comparing it to its corresponding matched normal genotyping data. Samples were considered to be from the same patient if at least 85% of their SNPs had the same genotype. Six plasma samples identified as swapped were excluded from ctDNA analysis. [Figure 68A-B]Analysis and validation results. (A) Single-target detection sensitivity. Approximately 60% analytical sensitivity was achieved using Signatera for mutation detection in approximately 0.03% of spike-in tumor DNA. (B) Estimated sample-level sensitivity of Signatera when at least two mutations are detected from a set of 16 target variants. [Figure 69] Following screening and recruitment, patients were followed up with six monthly blood samples. HER2 status was determined by immunohistochemical assay and fluorescence in in situ hybridization assay. If either assay was positive, the patient was considered to have HER2-positive cancer. NACT: Neoadjuvant chemotherapy, ACT: Adjuvant chemotherapy. [Figure 70] Workflow diagram of the muscle-invasive bladder cancer test in Example 9. [Figure 71A-G] Summary of patients with muscle-invasive bladder cancer in Example 9. Figure 71A shows the rates of synonymous and non-synonymous mutations called from WES. One patient's tumor was highly mutated, with a mutational burden of 126 mutations / Mb, and showed the POLD1 mutation, which has been previously associated with a high-frequency mutagenesis (Campbell, BB et al., Comprehensive Analysis of Hypermutation in Human Cancer. Cell 171, 1042-1056. e10 (2017)). Figure 71B shows the relative contribution of mutational signatures associated with bladder cancer. Figure 71C shows mutations in genes that frequently mutate in bladder cancer (TCGA) (Robertson, AG et al., Comprehensive Molecular Characterization of Muscle-Invasive Bladder Cancer. Cell 171, 540-556.e25(2017). Figure 71D shows harmful mutations in DNA damage response (DDR) related genes that were mutated in more than 5% of the 68 samples. Figure 71E shows the total number of harmful DDR mutations. Figure 71F shows the clinical and histopathological features. Figure 71G shows the summarized ctDNA status. [Figure 72]A diagram outlining the clinical protocol and sampling plan for the muscle-invasive bladder cancer test in Example 9. [Figure 73] A diagram outlining the Signatera (trademark) workflow. [Figure 74] Long-term display of ctDNA results for all analytical samples corresponding to the muscle-invasive bladder cancer test in Example 9. Patients are divided into three groups based on their ctDNA status. The upper panel shows patients who were ctDNA positive before and after cystectomy (CX), the middle panel shows patients who were ctDNA positive only before CX, and the lower panel shows patients who were ctDNA negative. Horizontal lines represent the disease course of each patient, circles represent ctDNA status, and red circles indicate samples with at least two positive assays. Treatment and imaging information is shown for each patient. [Figure 75A-E] Graph showing prognostic values ​​of ctDNA detection for muscle-invasive bladder cancer testing in Example 9. Kaplan-Meier survival analysis showing the probability of recurrence-free survival (RFS) and overall survival (OS) stratified by ctDNA status before chemotherapy (Figure 75A), before cystectomy (CX) (Figure 75B), and after cystectomy (CX) (Figure 75C). Figure 75D shows the relationship between disease recurrence and ctDNA status before chemotherapy, before cystectomy, and after cystectomy, and disease recurrence and lymph node status before cystectomy. Figure 75E shows the relationship between ctDNA status before cystectomy (CX) and pathological status at cystectomy (CX). Statistical significance was assessed using Wilcoxon rank-sum tests for continuous variables and Fisher's exact tests for categorical variables. [Figure 76] A graph showing ctDNA changes in the individual disease course for muscle-invasive bladder cancer testing in Example 9. Figure 76 shows a detailed display of the disease course, applied treatment, and associated long-term ctDNA analysis from selected patients. ctDNA status, applied treatment, and imaging results are presented according to legend. A good lead time for recurrence detection based on ctDNA is shown. [Figure 77]Graph showing the time difference between molecular recurrence (ctDNA positive) and clinical recurrence (positive radiographic imaging) for muscle-invasive bladder cancer in Example 9. The p-value was calculated using the paired Wilcoxon signed-rank test. [Figure 78A-H] Graph showing predictive markers for chemotherapy response for muscle-invasive bladder cancer testing in Example 9. Figure 78A shows the relationship between disease recurrence and response to chemotherapy. Figure 78B shows the relative contribution of Signature 5 for all patients stratified by response to chemotherapy and ERCC2 mutation status, respectively. Figure 78C shows the fraction of patients responding to treatment in relation to ERCC2 mutation status. Figure 78D. RNA subtype graph_new graph. Figure 78E shows the association between ctDNA and response to chemotherapy for patients who are ctDNA negative throughout the entire disease course, patients whose ctDNA levels drop to zero, and patients whose ctDNA levels remain positive. Figure 78F shows ctDNA levels for all patients with detectable ctDNA before, during, and after chemotherapy. Patients are grouped by response to chemotherapy and their recurrence status is indicated. [Figure 79A-D] A graph showing the total number of identified mutations per patient in relation to the number of ERCC2 status or disruptive DNA damage response (DDR) mutations for muscle-invasive bladder cancer testing in Example 9. [Figure 80] Graph showing genomic heterogeneity between primary tumor and metastatic recurrence for muscle-invasive bladder cancer testing in Example 9. Whole exome sequencing (WES) data from the primary tumor were compared with ctDNA. WES data from plasma samples with high ctDNA variant allele frequencies (VAFs) were detected at the time of metastatic recurrence. Genomic locations with mutations identified in either plasma or tumor exome data were examined in terms of base pairs. Allele frequencies identified and obtained in plasma or tumor exome data are shown. Individual mutations are color-coded according to the statistical probability (intensity) of mutation calls. The Venn diagram represents the number of mutations exclusively identified in tumor, plasma, or both. [Figure 81] Graphs showing variant allele frequencies (VAF%) at various days relative to cystectomy (CX), derived from eight patients in the muscle-invasive bladder cancer study in Example 9. [Figure 82] A graph showing ctDNA levels in plasma from 10 patients previously analyzed by ddPCR, compared to ultra-deep sequencing for muscle-invasive bladder cancer testing in Example 9. [Figure 83A-E] Graphs showing clinical, histopathological, and molecular parameters for all 125 patients. Figure 83A shows the relative contributions of the five most relevant colorectal cancer-associated mutation signatures. Figure 83B shows the rates of synonymous and non-synonymous mutations called from WES. Figure 83C shows a graph of mutations in genes that frequently mutate in colorectal cancer (TCGA) {Cancer Genome Atlas, 2012 #52}. Figure 83D shows clinical and histopathological features. Figure 83E shows a graph summarizing pre- and post-operative ctDNA status. [Figure 84] A diagram illustrating patient registration, sample collection, and patient subgroup definitions used to address defined clinical problems. Abbreviations: ctDNA (circulating tumor DNA); CT scan (computed tomography scan); post-op (post-surgery); TTR (time to recurrence). [Figure 85A-C]A graph showing the quality control (QC) testing of the workflow for whole exome sequencing of patient samples. Of 795 plasma samples, 793 (99%) passed the sample QC process. 194 samples (from 70 patients) were run with SNP tracers to test for agreement between the plasma samples and their corresponding tissue biopsies. All 194 plasma samples passed agreement QC. Figure 85A shows the amount of DNA input for library preparation. Up to 66 ng of cell-free DNA (cfDNA) from each plasma sample was used as input to the library preparation protocol. Library DNA input amounts ranged from 1 to 66 ng, with a median of 45.66. The purified libraries underwent quality control before proceeding to the next step. One sample failed library preparation QC. Figure 85B shows the sequencing coverage. Assays with coverage less than 5000x were excluded from analysis. Subsequently, samples with fewer than eight successful assays failed sequencing coverage QC. One sample failed the sequencing coverage requirements. The median read depth for assays that passed coverage QC was 105,000x. Figure 85C shows the sequencing error rates measured for all plasma samples. The mean transition error rate was 5e-5, and the mean transversion error rate was 8e-6. [Figure 86] The results and dynamics of circulating tumor DNA (ctDNA) for each individual patient are shown. [Figure 87A-F]Figure 87A shows ctDNA status pre-op, 30 days post-operatively, and during adjuvant chemotherapy (ACT). Figure 87A shows pre-operative detection of ctDNA. Figure 87B shows recurrence rates. Figure 87C shows Kaplan-Meier estimates of TTR for 94 stage I-III patients stratified by ctDNA status at 30 days post-operatively. Figure 87D shows the effect of ACT on ctDNA-positive patients, assessed by recurrence rates and long-term ctDNA status. Figure 87E shows recurrence rates stratified by ctDNA status at the first follow-up visit after ACT. Figure 87F shows Kaplan-Meier estimates of TTR for 58 ACT-treated patients stratified by ctDNA status and first follow-up visit after ACT. [Figure 88] This shows the postoperative detection of carcinoembryonic antigen (CEA) in 125 patients with stage i-III chronic carcinoma (CRC). [Figure 89] A schematic diagram of ctDNA profiling results for plasma samples included in ctDNA analysis at day 30, ordered by relapse status and disease stage, is shown. Patients marked with a(a) have synchronous CRC. Plasma marked with ** is positive only in the second pool (n=1). [Figure 90A-B] A schematic diagram of ctDNA profiling results for a subset of plasma samples undergoing ACT and included in 30-day ctDNA analysis is shown, ordered by relapse status and disease stage. Patients marked with a (or more) have synchronous CRC. [Figure 91] A schematic diagram of ctDNA profiling results for plasma samples included in long-term post-ACT ctDNA analysis, ordered by recurrence status, postoperative ctDNA status, and length of follow-up. Patients marked with a(a) have synchronous CRC (n=2). Plasma samples marked with ** are positive only in the second pool (n=1). [Figure 92]A schematic diagram of CEA profiling results for plasma samples included in long-term ACT-postoperative ctDNA analysis is shown, ordered by recurrence status, postoperative ctDNA status, and length of follow-up. Patients marked with a (or more) have synchronous CRC (n=2). Plasma marked with ** is positive only in the second pool (n=1). [Figure 93A-D] (Figures 93-D) Graphs showing the association between ctDNA status and relapse after curative treatment. Figure 93A shows relapse rates stratified by long-term ctDNA status. Figure 93B shows Kaplan-Meier estimates of TTR for 75 patients with long-term samples, stratified by long-term ctDNA status. Figure 93C shows a graph comparing time to radiological relapse and time to ctDNA relapse. Figure 93D shows that variant allele frequencies (VAF) of ctDNA in plasma increased towards radiological relapse. Early time points before and during ACT are omitted. [Figure 94] Schematic diagram of long-term ctDNA profiling results from plasma samples of patients with relapse and those without relapse. Patients with only one positive plasma sample during surveillance are considered positive. [Figure 95] Schematic diagram of long-term CEA profiling results from relapsed and non-relapsed patients. Patients with only one positive plasma sample during surveillance are considered positive. [Figure 96] A graph comparing the time to radiological recurrence and the time to CEA recurrence. [Figure 97A-C]Detection of causative mutations in relapsed patients. Figure 97A shows the percentage of ctDNA+ relapsed patients in whom causative mutations were detected during surveillance. The first ctDNA+ sample (left column) and all ctDNA+ plasma samples (right column). Figure 97B shows the causative variants called in the blood. Correlation between mean blood VAF calculated using the Signatera ctDNA+ assay and variant allele frequencies (VAF) of the causative mutation, plotted using logarithmic scales on both the horizontal and vertical axes. Figure 97C shows sequential ctDNA profiling of two representative relapsed patients with causative mutations. [Figure 98] A schematic comparison of current standard treatments and postoperative patient management guided by potential ctDNA. [Figure 99] A graph showing the reduction in ctDNA due to adjuvant chemotherapy (ACT). [Figure 100A-B] Graphs showing the association between ctDNA status and relapse after curative treatment. Figure 100A shows relapse rates stratified by long-term ctDNA status and Kaplan-Meier estimates of TTR for 58 patients with long-term samples stratified by long-term ctDNA status. Figure 100B shows relapse rates stratified by CEA analysis and Kaplan-Meier estimates of TTR for 58 patients with long-term samples stratified by CEA analysis.

[0095] The figures identified above are provided for illustrative purposes only and are not limited to them. [Modes for carrying out the invention]

[0096] The methods and compositions provided herein improve the detection, diagnosis, staging, screening, treatment, and management of cancer (e.g., breast cancer, bladder cancer, or colorectal cancer). The methods provided herein, in exemplary embodiments, analyze single nucleotide variant mutations (SNVs) in circulating fluids, particularly in circulating tumor DNA. The methods offer the advantage of identifying many of the mutations found in tumors and clones, as well as subclonal mutations, in a single test, compared to multiple tests that require the use of tumor samples if even slightly effective. The methods and compositions may be useful on their own, or they may be useful when used in conjunction with other methods for the detection, diagnosis, staging, screening, treatment, and management of cancer (e.g., breast cancer, bladder cancer, or colorectal cancer), for example, to support the results of these other methods and provide more reliable and / or definitive results.

[0097] Accordingly, in one embodiment, a method for determining single nucleotide variants present in cancer (e.g., breast cancer, bladder cancer, or colorectal cancer) is provided herein by determining single nucleotide variants present in ctDNA samples from an individual having or suspected of having cancer (e.g., breast cancer, bladder cancer, or colorectal cancer) using the ctDNA SNV amplification / sequencing workflow provided herein.

[0098] The terms “cancer” and “malignant” refer to or describe a physiological condition in animals typically characterized by uncontrolled cell proliferation. A “tumor” contains one or more cancerous cells. There are several major types of cancer. Carcinomas are cancers that begin in the skin or in the tissues that form the contours of or cover internal organs. Sarcomas are cancers that begin in bone, cartilage, fat, muscle, blood vessels, or other connective or supporting tissues. Leukemia is a cancer that begins in hematopoietic tissues such as bone marrow, where a large number of abnormal blood cells are produced and enter the bloodstream. Lymphoma and multiple myeloma are cancers that begin in the cells of the immune system. Central nervous system cancers are cancers that begin in the tissues of the brain and spinal cord.

[0099] In some embodiments, cancer includes acute lymphoblastic leukemia, acute myeloid leukemia, adrenocortical carcinoma, AIDS-related cancer, AIDS-related lymphoma, anal cancer, appendiceal cancer, astrocytoma, atypical teratoma / rhabdoid tumor, basal cell carcinoma, bladder cancer, brainstem glioma, brain tumor (including brainstem glioma, atypical teratoma / rhabdoid tumor of the central nervous system, germ cell tumor of the central nervous system, astrocytoma, craniopharyngioma, ependymoblastoma, ependymodium, medulloblastoma, medullary epithelioma, moderately differentiated pineal parenchymal tumor, supratentorial primitive neuroectoderm tumor, and pineoblastoma), breast cancer, and bronchial cancer. Tumors, Burkitt lymphoma, cancer of unknown primary site, carcinoid tumor, carcinoma of unknown primary site, atypical teratoma / rhabdoid tumor of the central nervous system, germ cell tumor of the central nervous system, cervical cancer, childhood cancer, chordoma, chronic lymphocytic leukemia, chronic myeloid leukemia, chronic myeloproliferative disorder, colon cancer, colorectal cancer, craniopharyngioma, cutaneous T-cell lymphoma, pancreatic islet cell tumor, endometrial cancer, ependymoblastoma, ependymoma, esophageal cancer, nasal neuroblastoma, Ewing's sarcoma, extracranial germ cell tumor, extragonadal germ cell tumor, extrahepatic cholangiocarcinoma, gallbladder cancer, gastric cancer (stomach) Cancer, gastrointestinal carcinoid tumor, gastrointestinal stromal cell tumor, gastrointestinal stromal tumor (GIST), choriocarcinoma of pregnancy, glioma, hairy cell leukemia, head and neck cancer, cardiac cancer, Hodgkin lymphoma, hypopharyngeal cancer, intraocular melanoma, islet cell tumor, Kaposi's sarcoma, kidney cancer, Langerhans cell histiocytosis, laryngeal cancer, lip cancer, liver cancer, malignant fibrous histiocytoma, bone cancer, medulloblastoma, medullary epithelioma, melanoma, Merkel cell carcinoma, Merkel cell carcinoma, mesothelioma, metastatic cervical squamous cell carcinoma of unknown primary origin, mouth cancer, multiple endocrine neoplasia syndrome, multiple myeloma, multiple myeloma / plasmacytoma, mycosis fungoides, myelodysplastic syndrome, myeloproliferative neoplasm, nasal cavity cancer, nasopharyngeal cancer, neuroblastoma, non-Hodgkin lymphoma, non-melanoma, skin cancer, non-small cell lung cancer, mouth cancerCancer, oral cancer, oropharyngeal cancer, osteosarcoma, other brain and spinal cord tumors, ovarian cancer, epithelial ovarian cancer, ovarian germ cell tumor, low-grade ovarian tumor, pancreatic cancer, papillomatosis, paranasal sinus cancer, parathyroid cancer, pelvic cancer, penile cancer, pharyngeal cancer, moderately differentiated pineal parenchymal tumor, pineoblastoma, pituitary tumor, plasmacytoma / multiple myeloma, pleuropulmonary blastoma, primary central nervous system (CNS) lymphoma, primary hepatocellular carcinoma, prostate cancer, rectal cancer, kidney cancer, renal cell carcinoma, renal cell carcinoma, airway cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, Sézary syndrome, small cell lung cancer, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma, cervical squamous cell carcinoma, stomach cancer This includes cancer, supratentorial primitive neuroectoderm tumors, T-cell lymphomas, testicular cancers, pharyngeal cancers, thymic carcinomas, thymomas, thyroid cancers, transitional cell carcinomas, transitional cell carcinomas of the renal pelvis and ureters, trophoblastic tumors, ureteral cancers, urinary tract cancers, uterine cancers, uterine sarcomas, vaginal cancers, vulvar cancers, Valdenström macroglobulinemia, or Wilms' tumors.

[0100] In another embodiment, a method is provided herein for detecting cancer (e.g., breast cancer, bladder cancer, or colorectal cancer) in a blood sample or fraction thereof from an individual, e.g., an individual suspected of having cancer, the method comprising determining single nucleotide variants present in the sample by determining the single nucleotide variants present in the ctDNA sample using the ctDNA SNV amplification / sequencing workflow provided herein. The presence of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 SNVs at the lower limit of the range and 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, or 50 SNVs at multiple single nucleotide loci in the sample is an indicator of the presence of cancer (e.g., breast cancer, bladder cancer, or colorectal cancer).

[0101] In another embodiment, a method for detecting clonal single-nucleotide variants in tumors of an individual (e.g., breast cancer, bladder cancer, or colorectal cancer) is provided herein. This method comprises performing a ctDNA SNV amplification / sequencing workflow as provided herein and determining the variant allele frequency of each SNV locus based on the sequences of multiple copies of a set of amplicons. A relatively high allele frequency of multiple single-nucleotide variant loci compared to other single-nucleotide variants is an indicator of a clonal single-nucleotide variant in a tumor. Variant allele frequencies are well known in the art of sequencing. Supporting this embodiment is provided, for example, in Figures 12–14.

[0102] In certain embodiments, the method further comprises determining a treatment plan, a therapy, and / or administering a compound targeting one or more clonal single nucleotide variants to an individual. In certain examples, subclones and / or other clonal SNVs are not targeted by the therapy. Specific therapies and associated mutations are provided in other chapters of this specification and are known in the art. Thus, in certain examples, the method further comprises administering a compound to an individual, which is known to be specifically effective in treating cancers (e.g., breast cancer, bladder cancer, or colorectal cancer) having one or more of the determined single nucleotide variants.

[0103] In certain embodiments of this model, variant allele frequencies greater than 0.25%, 0.5%, 0.75%, 1.0%, 5%, or 10% are indicators of clonal single-nucleotide variants. These cutoffs are supported by the tabular data in Figures 20A–20B.

[0104] In a particular example of this embodiment, the cancer is breast cancer, bladder cancer, or colorectal cancer of stage 1a, 1b, or 2a. In a particular example of this embodiment, the cancer is breast cancer, bladder cancer, or colorectal cancer of stage 1a or 1b. In a particular example of this embodiment, the individual does not undergo surgery. In this particular embodiment, the individual does not undergo a biopsy.

[0105] In some examples of this embodiment, clonal SNVs are identified or further identified if other tests, such as direct tumor tests, suggest that the SNV under test is a clonal SNV, in any given SNV, the variable allele frequency is greater than at least one-quarter, one-third, half, or three-quarters of the other single nucleotide variants determined.

[0106] In some embodiments, the methods herein for detecting SNVs in ctDNA may be used instead of direct analysis of DNA from tumors. The results provided herein show that SNVs that are considerably more likely to be clonal SNVs have a higher VAF (see Figures 12–14).

[0107] In certain examples of any embodiment of the method provided herein, data are provided for SNVs found in tumors from an individual before targeted amplification is performed on ctDNA from that individual. Thus, in these embodiments, the SNV amplification / sequencing reaction is performed on one or more tumor samples from the individual. In this method, the ctDNA SNV amplification / sequencing reaction provided herein is still advantageous because it provides a liquid biopsy of clonal and subclonal mutations. Furthermore, as provided herein, clonal mutations can be more clearly identified in individuals with cancer (e.g., breast cancer, bladder cancer, or colorectal cancer) if a high VAF percentage (e.g., greater than 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10% VAF in a ctDNA sample from the individual) is determined for a given SNV.

[0108] In certain embodiments, the methods provided herein can be used to determine how to isolate and analyze ctDNA from circulating free nucleic acids from an individual having cancer (e.g., breast cancer, bladder cancer, or colorectal cancer). First, it is determined whether the cancer is breast cancer, bladder cancer, or colorectal cancer. If the cancer is breast cancer, bladder cancer, or colorectal cancer, circulating free nucleic acids are isolated from the individual. In some examples, the methods further include determining the stage of the cancer.

[0109] In some ways, compositions and / or solid supports of the present invention are provided herein. A composition comprising a circulating tumor nucleic acid fragment comprising a universal adapter, wherein the circulating tumor nucleic acid was derived from breast cancer, bladder cancer, or colorectal cancer.

[0110] Compositions of the present invention are provided herein, in some embodiments, comprising a circulating tumor nucleic acid fragment containing a universal adapter, wherein the circulating tumor nucleic acid is derived from a sample of blood or a fraction thereof from an individual having cancer (e.g., breast cancer, bladder cancer, or colorectal cancer). These methods typically involve the formation of a ctDNA fragment containing a universal adapter. Furthermore, such methods typically involve the formation of a solid support, particularly for high-throughput screening, comprising a clonal assembly of multiple nucleic acids, wherein the clonal assembly comprises an amplicon created from a sample of circulating free nucleic acid and is ctDNA. In exemplary embodiments based on remarkable results provided herein, the ctDNA was derived from cancer (e.g., breast cancer, bladder cancer, or colorectal cancer).

[0111] Similarly, a solid support is provided herein as one embodiment of the present invention, which is a solid indicator comprising a clonal assembly of a plurality of nucleic acids, wherein the clonal assembly comprises nucleic acid fragments prepared from a sample of circulating free nucleic acids from a sample of blood or a fraction thereof from a solid having cancer (e.g., breast cancer, bladder cancer, or colorectal cancer).

[0112] In certain embodiments, nucleic acid fragments in different clone assemblies contain the same universal adapter. Such compositions are typically formed during high-throughput sequencing reactions in the methods of the present invention.

[0113] The nucleic acid clone assembly may originate from nucleic acid fragments from a set of samples from two or more individuals. In these embodiments, the nucleic acid fragment includes one of a series of molecular barcodes corresponding to a sample in the set of samples.

[0114] Detailed analytical methods are provided herein as SNV Method 1 and SNV Method 2 in the Analysis chapter of this specification. Either method provided herein may further include the analytical steps provided herein. Thus, in a particular example, a method for determining whether a single nucleotide variant is present in a sample includes identifying a confidence value for each allele determination for each set of single nucleotide variant loci, which may be based at least in part on read depth for the loci. The confidence limits can be set at least 75%, 80%, 85%, 90%, 95%, 96%, 96%, 98%, or 99%. The confidence limits can be set at different levels for different types of mutations.

[0115] This method can be performed with read depths of at least 5, 10, 15, 20, 25, 50, 100, 150, 200, 250, 500, 1,000, 10,000, 25,000, 50,000, 100,000, 250,000, 500,000, or 1 million single nucleotide variant loci.

[0116] In certain embodiments, the methods of any embodiment herein include determining the efficiency and / or error rate per cycle, determined for each amplification reaction of a multiple amplification reaction of a single nucleotide variant locus. The efficiency and error rate may then be used to determine whether a single nucleotide variant is present in the sample at a set of single variant loci. Further detailed analytical steps provided in SNV Method 2 provided in the analytical methods may also be included in certain embodiments.

[0117] In exemplary embodiments, in any of the methods herein, the set of single nucleotide variant loci includes all single nucleotide variant loci identified in the TCGA and COSMIC datasets for cancer (e.g., breast cancer, bladder cancer, or colorectal cancer).

[0118] In certain embodiments of any method according to this specification, the set of single nucleotide variant loci includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 75, 100, 250, 500, 1000, 2500, 5000 or 10,000 single nucleotide variant loci known to be associated with cancer (e.g., breast cancer, bladder cancer, or colorectal cancer) at the lower end of the range, and 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 75, 100, 250, 500, 1000, 2500, 5000, 10,000, 20,000 and 25,000 single nucleotide variant loci known to be associated with cancer (e.g., breast cancer, bladder cancer, or colorectal cancer) at the upper end of the range.

[0119] In any method for detecting SNVs according to this specification, including ctDNA SNV amplification / sequencing workflows, improved amplification parameters for multiplex PCR may be used. For example, if the amplification reaction is a PCR reaction, the annealing temperature may be 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10°C higher than the melting point of at least 10, 20, 25, 30, 40, 50, 06, 70, 75, 80, 90, 95, or 100% of the primers in the primer set at the lower end of the range, and 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15°C higher at the upper end of the range.

[0120] In certain embodiments, when the amplification reaction is a PCR reaction, the length of the annealing step during the PCR reaction is 10, 15, 20, 30, 45, and 60 minutes at the lower end of the range, and 15, 20, 30, 45, 60, 120, 180, or 240 minutes at the upper end of the range. In certain embodiments, the primer concentration in amplification (e.g., PCR reaction) is 1 to 10 nM. Furthermore, in exemplary embodiments, the primers in the primer set are designed to minimize primer dimerization.

[0121] Accordingly, in an example of any method herein that includes an amplification step, the amplification reaction is a PCR reaction, the annealing temperature is 1 to 10°C higher than the melting point of at least 90% of the primers in the primer set, the length of the annealing step during the PCR reaction is 15 to 60 minutes, the primer concentration in the amplification reaction is 1 to 10 nM, and the primers in the primer set are designed to minimize primer dimerization. In a further embodiment of this example, the multiple amplification reaction is carried out under restriction primer conditions.

[0122] In another embodiment, a method is provided herein for supporting a diagnosis of cancer (e.g., breast cancer, bladder cancer, or colorectal cancer) in an individual suspected of having cancer (e.g., breast cancer, bladder cancer, or colorectal cancer) from a blood sample or fraction thereof from an individual, comprising performing a ctDNA SNV amplification / sequencing workflow provided herein to determine whether one or more single nucleotide variants are present at multiple single nucleotide variant loci. In this embodiment, the following elements, descriptions, guidelines, or rules apply: The absence of a single nucleotide variant supports a diagnosis of stage 1a, 1b, or 2a adenocarcinoma; the presence of a single nucleotide variant supports a diagnosis of squamous cell carcinoma or stage 2b or 3a adenocarcinoma; and / or the presence of 10 or more single nucleotide variants supports a diagnosis of squamous cell carcinoma or stage 2b or 3 adenocarcinoma.

[0123] These results identify the analysis of lung ADC and SCC samples from individuals using a ctDNA SNV amplification / sequencing workflow as a valuable method for identifying SNVs found in ADC tumors, particularly in stage 2b and 3a ADC tumors, and especially in SCC tumors at any stage (see Figures 15 and 20A-B).

[0124] In certain embodiments, the methods described herein for detecting SNVs may be used to direct treatment regimens. Therapies targeting specific mutations associated with ADCs and SCCs are available and under development (Nature Review Cancer. 14:535-551 (2014). For example, detection of EGFR mutations at L858R or T790M may be beneficial in selecting a therapy. Erlotinib, gefitinib, afatinib, AZK9291, CO-1686, and HM61713 are current therapies approved in the U.S. and in clinical trials that target specific EGFR mutations. In another example, G12D, G12C, or G12V mutations in KRAS may be used to direct an individual to a combination therapy of selumetinib and docetaxel. In yet another example, the V600E mutation in BRAF may be used to direct a subject to treatment with vemurafenib, dabrafenib, and trametinib.

[0125] The sample analyzed by the method of the present invention is, in certain exemplary embodiments, a blood sample or a fraction thereof. The method provided herein is adapted, in certain embodiments in particular, to amplify DNA fragments, especially tumor DNA fragments found in circulating tumor DNA (ctDNA). Such fragments are typically about 160 nucleotides long.

[0126] Cell-free nucleic acids (cfNA), such as cfDNA, are known in the art to be released into circulation through various forms of cell death, including apoptosis, necrosis, autophagy, and necroptosis. cfDNA is fragmented, and the fragment size distribution varies from 150–350 bp to over 10,000 bp (see Kalnina et al., World J Gastroenterol. 2015 Nov. 7;21(41):11636-11653). For example, the size distribution of plasma DNA fragments in hepatocellular carcinoma (HCC) patients ranges from 100–220 bp in length, with a frequency peak at approximately 166 bp, and the highest tumor DNA concentration in fragments is at 150–180 bp in length (see Jiang et al., Proc Natl Acad Sci USA 112:E1317-E1325).

[0127] In an exemplary embodiment, after removing cell fragments and platelets by centrifugation, circulating tumor DNA (ctDNA) is isolated from the blood using an EDTA-2Na tube. The plasma sample may be stored at -80°C until the DNA is extracted, for example, using the QIAamp DNA Mini Kit (Qiagen, Hilden, Germany) (e.g., Hamakawa et al., Br J Cancer. 2015;112:352-356). Hamakava et al. reported that the median concentration of extracted cell-free DNA for all samples was 43.1 ng per ml of plasma (range 9.5–1338 ng / ml), with a variant frequency range of 0.001–77.8% and a median of 0.90%.

[0128] In certain exemplary embodiments, the sample is a tumor. Given the teachings herein, methods for isolating nucleic acids from tumors and for preparing nucleic acid libraries from such DNA samples are known in the art. Furthermore, given the teachings herein, those skilled in the art will recognize how to prepare nucleic acid libraries suitable for the methods herein from other samples, such as other liquid samples in which DNA is suspended in a free state, in addition to ctDNA samples.

[0129] The method of the present invention, in certain embodiments, typically includes the step of preparing and amplifying a nucleic acid library from a sample (i.e., library preparation). During the library preparation step, nucleic acids from the sample may have an accompanying ligation adapter (often called a library tag or ligation adapter tag (LT)), the ligation adapter containing a universal priming sequence, followed by universal amplification. In one embodiment, this may be done using a standard protocol designed to prepare a sequencing library after fragmentation. In one embodiment, the DNA sample may be blunt-ended, and then an A may be added to its 3' end. A Y adapter with a T overhang may be added and ligated. In some embodiments, other sticky ends other than A or T overhangs may be used. In some embodiments, other adapters, such as loop-shaped ligation adapters, may be added. In some embodiments, the adapter may have a tag designed for PCR amplification.

[0130] Some embodiments provided herein involve detecting SNVs in a ctDNA sample. Such methods in exemplary embodiments include amplification and sequencing steps (sometimes referred to herein as the “ctDNA SNV amplification / sequencing workflow”). In exemplary examples, the ctDNA amplification / sequencing workflow may include creating a set of amplicons by performing multiple amplification reactions on nucleic acids isolated from a blood sample or fraction thereof from an individual, e.g., an individual suspected of having cancer (e.g., breast cancer, bladder cancer, or colorectal cancer), wherein each amplicon in the set of amplicons extends to at least one single nucleotide variant locus of a set of single nucleotide variant loci, e.g., an SNV locus known to be associated with cancer (e.g., breast cancer, bladder cancer, or colorectal cancer); and sequencing at least one segment of each amplicon in the set of amplicons, wherein the segment contains a single nucleotide variant locus. In this method, this exemplary method determines single nucleotide variants present in the sample.

[0131] An exemplary ctDNA SNV amplification / sequencing workflow may more specifically involve forming an amplification reaction mixture by combining polymerase, nucleotide triphosphates, and nucleic acid fragments from a nucleic acid library prepared from a sample with a set of primers that bind at an effective distance from a single nucleotide variant locus, or a set of primer pairs that extend to an effective region containing a single nucleotide variant locus. In exemplary embodiments, the single nucleotide variant locus is one known to be associated with cancer (e.g., breast cancer, bladder cancer, or colorectal cancer). The amplification reaction mixture is then subjected to amplification conditions to preferably create a set of amplicons containing at least one single nucleotide variant locus from a set of single nucleotide variant loci known to be associated with cancer (e.g., breast cancer, bladder cancer, or colorectal cancer), and sequencing at least one segment of each amplicon in the set of amplicons, wherein the segment contains a single nucleotide variant locus.

[0132] The effective distance for primer binding may be within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 75, 100, 125, or 150 base pairs of the SNV locus. The effective range that a pair of primers extend to typically includes the SNV and is typically 160 base pairs or less, and may be 150, 140, 130, 125, 100, 75, 50, or 25 base pairs or less. In another embodiment, the effective range to which the primer pair extends is 20, 25, 30, 40, 50, 60, 70, 75, 100, 110, 120, 125, 130, 140, or 150 nucleotides at the lower limit of the range from the SNV locus, and 25, 30, 40, 50, 60, 70, 75, 100, 110, 120, 125, 130, 140, or 150, 160, 170, 175, or 200 at the upper limit of the range.

[0133] Further details regarding amplification methods available in ctDNA SNV amplification / sequencing workflows for detecting SNVs for use in the methods described herein are provided in other chapters herein.

[0134] SNV Call Analysis While performing the methods provided herein, nucleic acid sequencing data is generated for amplicons created by tiled multiplex PCR. Algorithmic design tools are available that can be used and / or adapted to analyze this data to determine, within a specific confidence limit, whether a mutation (e.g., SNV) is present in the target gene.

[0135] Sequencing reads may be demultiplexed using in-house tools and mapped in single-ended mode using paired merged reads for the hg19 genome with the BWA mem function (BWA) of Burrows-Wheeler Alignment Software (see Li H. and Durbin R. (2010) Fast and accurate long-read alignment with Burrows-Wheeler Transform. Bioinformatics, Epub. [PMID:20080505]). Amplification statistics QC can be performed by analyzing the total number of reads, the number of mapped reads, the number of mapped reads on the target, and the number of measured reads.

[0136] In certain embodiments, any analytical method for detecting SNVs from nucleic acid sequencing data may be used in conjunction with the method of the present invention according to the present invention, which includes a step of detecting SNVs or determining whether SNVs are present. In certain exemplary embodiments, the method of the present invention utilizing the following SNV method 1 is used. In other further exemplary embodiments, the method of the present invention, which includes a step of detecting SNVs or determining whether SNVs are present at an SNV locus, utilizes the following SNV method 2.

[0137] SNV Method 1: With respect to this embodiment, the background error model was constructed using normal plasma samples and sequenced in the same sequencing run to account for run-specific artifacts. In certain embodiments, 5, 10, 15, 20, 25, 30, 40, 50, 100, 150, 200, 250, or more normal plasma samples are analyzed in the same sequencing run. In certain exemplary embodiments, 20, 25, 40, or 50 normal plasma samples are analyzed in the same sequencing run. Noise locations with a median normal variant allele frequency above the cutoff are removed. For example, this cutoff is greater than 0.1%, 0.2%, 0.25%, 0.5%, 1%, 2%, 5%, or 10% in certain embodiments. In certain exemplary embodiments, noise locations with a median normal variant allele frequency above 0.5% are removed. Outlier samples were repeatedly removed from this model to account for noise and contamination. In certain embodiments, samples with a Z-score greater than 5, 6, 7, 8, 9, or 10 are removed from data analysis. For each base substitution at all genomic loci, the mean and standard deviation of the error are calculated, weighted by read depth. Locations in plasma samples that do not contain tumors or cells and have at least five variant reads and a Z-score of 10 against the background error model can be called, for example, candidate mutations.

[0138] SNV Method 2: In this embodiment, single nucleotide variants (SNVs) are determined using plasma ctDNA data. The PCR process is modeled as a stochastic process, parameters are estimated using a training set, and a final SNV call is made for a separate test set. Error propagation across multiple PCR cycles is determined, the mean and variance of background errors are calculated, and in exemplary embodiments, background errors are distinguished from actual mutations.

[0139] For each base, the following parameters are estimated. p = efficiency (probability that each read is replicated during each cycle) p e = error rate per cycle for mutant e (probability that an e - type error occurs) X0 = initial number of molecules

[0140] As reads are replicated over a series of PCR processes, more errors occur. Thus, the error profile of a read is determined by the degree of separation from the original read. If a read has undergone k replications before being created, the read is called the k - th generation.

[0141] Let's define the following variables for each base. X ij = number of reads of the i - th generation created in PCR cycle j Y ij = total number of reads of the i - th generation at the end of cycle j X ij e = number of reads of the i - th generation having mutation e created in PCR cycle j

[0142] Furthermore, in addition to the normal molecules X0, there are an additional f e X0 molecules with mutation e at the start of the PCR process (thus, fe / (1 + fe) would be the fraction of mutated molecules in the initial mixture).

[0143] Considering the total number of reads of the i - 1 - th generation in cycle j - 1, the number of reads of the i - th generation created in cycle j has a binomial distribution with sample size Y i-1,j-1 and probability parameter p. Thus, E(X ij ,|Y i-1,j-1 ,p)=pY i-1,j-1 and Var(X ij ,|Y i-1,j-1 ,p)=p(1 - p)Y i-1,j-1 That is.

[0144] The inventors of the present application

Number

[0145] Finally, E(Xi j e |Y i-1,j-1 ,p e )=p e Y i-1,j-1 and Var(X ij e |Y i-1,j-1 ,p)=p e (1-p e )Y i-1,j-1 And using these, E(X ij e ) and Var(X ij e ) can be calculated.

[0146] In a particular embodiment, SNV method 2 is carried out as follows.

[0147] a) Use the training dataset to estimate PCR efficiency and error rate per cycle.

[0148] b) Using the efficiency distribution estimated in step (a), estimate the initial number of molecules for the test dataset at each base.

[0149] c) If necessary, update the efficiency estimate for the test dataset using the initial number of molecules estimated in step (b).

[0150] d) Using the test set data and parameters estimated in steps (a), (b), and (c), estimate the total number of molecules, the mean and variance of background error molecules and actual mutant molecules (for the search space consisting of the initial proportion of actual mutant molecules).

[0151] e) Fit the distribution of the total number of error molecules (background errors and actual mutations) in the total molecules and calculate the likelihood of the proportion of each actual mutation in the search space.

[0152] f) Determine the most likely proportion of actual mutations and calculate the reliability using data from process (e).

[0153] SNVs can be identified at an SNV locus using confidence cutoffs. For example, SNVs can be called using confidence cutoffs of 90%, 95%, 96%, 97%, 98%, or 99%.

[0154] Exemplary SNV Method 2 Algorithm This algorithm begins by estimating efficiency and error rate per cycle using a training set. n represents the total number of PCR cycles.

[0155] Read R at each base b b The number is (1+p b ) n It can be approximated by X0, where p b This is the efficiency at base b. (R b / X0) 1 / n Using 1+p b This can be estimated. Then, across all training samples, p b By determining the mean and standard deviation, the parameters of the probability distribution for each base (e.g., a normal distribution, a beta distribution, or a similar distribution) can be estimated.

[0156] Similarly, read R of error e at each base b. be Using the number of, p e can be estimated. After determining the mean and standard deviation of the error rate across all training samples, the probability distribution (e.g., normal distribution, beta distribution or similar distribution) is approximated, and using the values of this mean and standard deviation, its parameters are estimated.

[0157] Next, for the test data, the initial starting copy at each base is

Number

Number

[0158] Therefore, this parameter is estimated and used in the probability process. Next, by using these estimated values, the mean and variance of the molecules created in each cycle can be estimated (this is done separately for normal molecules, error molecules and mutant molecules).

[0159] Finally, by using a probability method (e.g., maximum likelihood or similar method), the best f that best fits the distributions of error, mutant and normal molecules e value can be determined. More specifically, for the various f e values in the final lead, the predicted ratio of error molecules to all molecules is estimated, the likelihood of the inventors' data for each of these values is determined, and then the value with the maximum likelihood is selected.

[0160] Primer tails can improve the detection of fragmented DNA from universally tagged libraries. If the library tag and primer tail contain homologous sequences, hybridization can be improved (e.g., by lowering the melting point (Tm)), and if only a portion of the primer target sequence is present in the sample DNA primer fragment, the primer can be extended. In some embodiments, 13 or more target-specific base pairs may be used. In some embodiments, 10 to 12 target-specific base pairs may be used. In some embodiments, 8 to 9 target-specific base pairs may be used. In some embodiments, 6 to 7 target-specific base pairs may be used.

[0161] In one embodiment, the library is created from the sample by ligating adapters to the ends of DNA fragments in the sample, or to the ends of DNA fragments created from DNA isolated from the sample. The fragments can then be amplified using PCR, for example, according to the following exemplary protocol.

[0162] 95°C for 2 minutes; 15 × [95°C for 20 seconds, 55°C for 20 seconds, 68°C for 20 seconds], 68°C for 2 minutes, hold at 4°C.

[0163] Many kits and methods are known in the art for the preparation of nucleic acid libraries containing universal primer binding sites for subsequent amplification (e.g., clonal amplification) and subsequent sequencing. To facilitate adapter ligation, library preparation and amplification may include end repair and adenylation (i.e., A-tailing). Kits specifically adapted for preparing libraries from small nucleic acid fragments (particularly circulating free DNA) may be useful for carrying out the methods provided herein. For example, the NEXTflex Cell Free Kit available from Bioo Scientific() or the Natera Library Prep Kit (available from Natera, Inc., San Carlos, CA). However, such kits are typically modified to include adapters customized for the amplification and sequencing steps of the methods provided herein. Adapter ligation can be carried out using commercially available kits, such as the ligation kit found in the AGILENT SURESELECT kit (Agilent, CA).

[0164] Next, a target region of a nucleic acid library prepared from a sample, in particular from DNA isolated from a circulating free DNA sample for the method of the present invention, is amplified. For this amplification, a series of primers or primer pairs may include 5, 10, 15, 20, 25, 50, 100, 125, 150, 250, 500, 1000, 2500, 5000, 10,000, 20,000, 25,000 or 50,000 primers at the lower end of the range, and 15, 20, 25, 50, 100, 125, 150, 250, 500, 1000, 2500, 5000, 10,000, 20,000, 25,000, 50,000, 60,000, 75,000 or 100,000 primers at the upper end of the range, each binding to one of the series of primer binding sites.

[0165] Primer designs may be created in conjunction with Primer3 (Untergrasser A, Cutcutache I, Koressaar T, Ye J, Faircloth BC, Remm M, Rozen SG (2012) "Primer3 - new capabilities and interfaces." Nucleic Acids Research 40(15):e115 and Koressaar T, Remm M (2007) "Enhancements and modifications of primer design program Primer3." Bioinformatics 23(10):1289-91). The source code is available at primer3.sourceforge.net. Primer specificity may be evaluated by BLAST and added to existing primer design pipeline criteria.

[0166] Primer specificity can be determined using the BLASTn program from the ncbi-blast-2.2.29+ package. The task option "blastn-short" may be used to map primers to the hg19 human genome. A primer design can be determined to be "specific" if it has fewer than 100 hits to the genome, and the top hit is the target complementary primer binding region of that genome, with a score at least 2 higher than other hits (the score is defined by the BLASTn program). This can be done to have hits unique to that genome and not have many other hits throughout the genome.

[0167] The final selected primers can be visualized using bed files and coverage maps for validation with IGV (James T. Robinson, Helga Thorvaldsdottir, Wendy Winckler, Mitchell Guttman, Eric S. Lander, Gad Getz, Jill P. Mesirov. Integrative Genomics Viewer. Nature Biotechnology 29, 24-26 (2011)) and the UCSC browser (Kent WJ, Sugnet CW, Furey TS, Roskin KM, Pringle TH, Zahler AM, Haussler D. The human genome browser at UCSC. Genome Res. June 2002;12(6):996-1006).

[0168] In certain embodiments, the method of the present invention includes forming an amplification reaction mixture. This reaction mixture is typically prepared by combining polymerase, nucleotide triphosphates, and nucleic acid fragments from a nucleic acid library prepared from a sample with a set of forward and reverse primers specific to a target region containing an SNV. In exemplary embodiments, the reaction mixture provided herein itself forms a distinct aspect of the present invention.

[0169] The amplification reaction mixture useful for the present invention contains components known in the art of nucleic acid amplification, particularly PCR amplification. For example, the reaction mixture typically contains nucleotide triphosphates, polymerase, and magnesium. The polymerase useful for the present invention may include any polymerase available for amplification reactions, particularly those useful for PCR reactions. In certain embodiments, hot-start Taq polymerase is particularly useful. Amplification reaction mixtures useful for carrying out the methods provided herein, such as AmpliTaq Gold Master Mix (Life Technologies, Carlsbad, CA), are commercially available.

[0170] PCR amplification (e.g., temperature cycling) conditions are well known in the art. The methods provided herein may include any PCR cycling conditions for amplifying a target nucleic acid (e.g., a target nucleic acid from a library). Non-limiting exemplary cycling conditions are provided in the Examples section of this specification.

[0171] Many workflows exist for performing PCR, and several workflows typical of the method disclosed herein are provided herein. The steps outlined herein are not intended to exclude other possible steps, nor are they implied to be necessary for any of the steps described herein to function properly. Numerous parameter variations or other modifications are known in the literature and can be made without affecting the essence of the invention.

[0172] In certain embodiments of the methods provided herein, at least a portion of an amplicon (e.g., an outer primer target amplicon), and in exemplary examples, the entire sequence, is determined. Methods for determining the sequence of amplicons are known in the art. Any sequencing method known in the art, such as Sanger sequencing, can be used for such sequence determination. In exemplary embodiments, high-throughput next-generation sequencing technologies (also referred herein as massively parallel sequencing technologies), such as, but not limited to, those used in MYSEQ (ILLUMINA), HISEQ (ILLUMINA), ION TORRENT (LIFE TECHNOLOGIES), GENOME ANALYZER ILX (ILLUMINA), and GS FLEX+ (ROCHE 454), can be used to sequence amplicons produced by the methods provided herein.

[0173] High-throughput gene sequencers can be modified to accommodate the use of barcoding (i.e., sample tagging using characteristic nucleic acid sequences) to identify unique samples from individuals, thereby enabling simultaneous analysis of multiple samples in a single run of the DNA sequencer. The number of times a given region of the genome is sequenced (reads) in library preparation (or other nucleic acid preparation of interest) will be proportional to the number of copies of that sequence in the genome of interest (or expression level in the case of preparations containing cDNA). Bias in amplification efficiency may be taken into consideration in such quantitative determinations.

[0174] target genes In exemplary embodiments of the present invention, the target genes are cancer-related genes, and in many exemplary embodiments, they are cancer-related genes. Cancer-related genes (e.g., cancer-related genes, bladder cancer-related genes, or colorectal cancer-related genes) refer to genes associated with changes in the risk of cancer (e.g., breast cancer, bladder cancer, or colorectal cancer) or changes in the prognosis of cancer. Exemplary cancer-related genes that promote cancer include oncogenes, genes that promote cell proliferation, invasion, or metastasis, genes that inhibit apoptosis, and pro-angiogenic genes. Cancer-related genes that inhibit cancer include, but are not limited to, tumor suppressor genes, genes that inhibit cell proliferation, invasion, or metastasis, genes that promote apoptosis, and anti-angiogenic genes.

[0175] One embodiment of the mutation detection method begins with the selection of a region of the target gene. Using a region with known mutations, primers for mPCR-NGS are developed to amplify the mutations and detect them.

[0176] The methods provided herein can be used to detect substantially any type of mutation, in particular mutations known to be associated with cancer, and most specifically, the methods provided herein target mutations (in particular SNVs) associated with cancer, in particular breast cancer, bladder cancer, or colorectal cancer. Exemplary SNVs may be one or more of the following genes: EGFR, FGFR1, FGFR2, ALK, MET, ROS1, NTRK1, RET, HER2, DDR2, PDGFRA, KRAS, NF1, BRAF, PIK3CA, MEK1, NOTCH1, MLL2, EZH2, TET2, DNMT3A, SOX2, MYC, KEAP1, CDKN2A, NRG1, TP53, LKB1, and PTEN have been identified as either mutated, having increased copy numbers, fused to other genes, or a combination thereof in various lung cancer samples (Non-small-cell lung cancers: a heterogeneous set of diseases. Chen et al., Nat. Rev. Cancer. 2014 August 14(8):535-551). In another example, the list of genes is as enumerated above, and SNVs are reported, for example, in the references of Chen et al.

[0177] Amplification (e.g., PCR) reaction mixture: The method of the present invention, in certain embodiments, includes forming an amplification reaction mixture. This reaction mixture is typically formed by combining polymerase, nucleotide triphosphates, and nucleic acid fragments from a nucleic acid library prepared from a sample with a series of forward-direction target-specific outer primers and first-chain reverse-direction outer universal primers. Another exemplary embodiment is a reaction mixture comprising forward-direction target-specific inner primers instead of forward-direction target-specific outer primers, and an amplicon from a first PCR reaction using outer primers instead of nucleic acid fragments from a nucleic acid library. The reaction mixtures provided herein, in exemplary embodiments, themselves form distinct aspects of the present invention. In exemplary embodiments, the reaction mixture is a PCR reaction mixture. The PCR reaction mixture typically contains magnesium.

[0178] In some embodiments, the reaction mixture includes ethylenediaminetetraacetic acid (EDTA), magnesium, tetramethylammonium chloride (TMAC), or any combination thereof. In some embodiments, the concentration of TMAC is 20–70 mM (including boundary values). While not intended to be bound by any particular theory, TMAC is thought to bind to DNA, stabilize the double helix, increase primer specificity, and / or equalize the melting points of different primers. In some embodiments, TMAC enhances the uniformity of the amount of amplification product for different targets. In some embodiments, the concentration of magnesium (e.g., magnesium derived from magnesium chloride) is 1–8 mM.

[0179] Multiple primers used in multiplex PCR for multiple targets can chelate a large amount of magnesium (two phosphate groups in the primer chelate one magnesium). For example, when using sufficient primers such that the concentration of phosphate groups from the primers is about 9 mM, the primers can reduce the effective magnesium concentration to about 4.5 mM. In some embodiments, since high concentrations of magnesium can cause PCR errors (e.g., amplification of non-target loci), EDTA is used to reduce the amount of magnesium available as a cofactor for the polymerase. In some embodiments, the concentration of EDTA reduces the amount of available magnesium to 1 - 5 mM (e.g., 3 - 5 mM).

[0180] In some embodiments, the pH is 7.5 - 8.5, e.g., 7.5 - 8, 8 - 8.3 or 8.3 - 8.5 (including the boundary values). In some embodiments, Tris is used at a concentration of, e.g., 10 - 100 mM, e.g., 10 - 25 mM, 25 - 50 mM, 50 - 75 mM or 25 - 75 mM (including the boundary values). In some embodiments, Tris at any of these concentrations is used at a pH of 7.5 - 8.5. In some embodiments, a combination of KCl and (NH4)2SO4 is used, e.g., 50 - 150 mM KCl and 10 - 90 mM (NH4)2SO4 (including the boundary values). In some embodiments, the concentration of KCl is 0 - 30 mM, 50 - 100 mM or 100 - 150 mM (including the boundary values). In some embodiments, the concentration of (NH4)2SO4 is 10 - 50 mM, 50 - 90 mM, 10 - 20 mM, 20 - 40 mM, 40 - 60 mM or 60 - 80 mM (NH4)2SO4 (including the boundary values). In some embodiments, the ammonium [NH4 + concentration is 0 - 160 mM, e.g., 0 - 50, 50 - 100 or 100 - 160 mM (including the boundary values). In some embodiments, the total of the potassium concentration and the ammonium concentration ([K + +[NH4 +]) is 0 to 160 mM, for example, 0 to 25, 25 to 50, 50 to 150, 50 to 75, 75 to 100, 100 to 125 or 125 to 160 mM (including boundary values). [K + ]+[NH4 + An exemplary buffer with a pH of 120 mM is 20 mM KCl and 50 mM (NH4)2SO4. In some embodiments, the buffer contains 25–75 mM Tris (pH 7.2–8), 0–50 mM KCl, 10–80 mM ammonium sulfate, and 3–6 mM magnesium (including boundary values). In some embodiments, the buffer contains 25–75 mM Tris (pH 7–8.5), 3–6 mM MgCl2, 10–50 mM KCl, and 20–80 mM (NH4)2SO4 (including boundary values). In some embodiments, 100–200 units / mL of polymerase is used. In some embodiments, 100 mM KCl, 50 mM (NH4)2SO4, 3 mM MgCl2, 7.5 nM of each primer in the library, and 7 µl of DNA template in a final volume of 20 µl at pH 8.1 are used.

[0181] In some embodiments, a crowding agent, such as polyethylene glycol (PEG, e.g., PEG8,000) or glycerol, is used. In some embodiments, the amount of PEG (e.g., PEG8,000) is 0.1–20%, e.g., 0.5–15%, 1–10%, 2–8%, or 4–8% (including boundary values). In some embodiments, the amount of glycerol is 0.1–20%, e.g., 0.5–15%, 1–10%, 2–8%, or 4–8% (including boundary values). In some embodiments, the crowding agent allows the use of either a low polymerase concentration and / or a shorter annealing time. In some embodiments, the crowding agent improves the uniformity of the DOR and / or reduces dropout (undetectable alleles) of polymerase. In some embodiments, proofreading polymerases, polymerases without (or with negligible) proofreading activity, or mixtures of proofreading polymerases and polymerases without (or with negligible) proofreading activity are used. In some embodiments, hot-start polymerases, non-hot-start polymerases, or mixtures of hot-start polymerases and non-hot-start polymerases are used. In some embodiments, HotStarTaq DNA polymerase is used (see, for example, QIAGEN catalog number 203203). In some embodiments, AmpliTaq Gold® DNA polymerase is used. In some embodiments, PrimeSTAR GXL DNA polymerase (Takara Clontech, Mountain View, CA), a high-fidelity polymerase that provides efficient PCR amplification when an excess template is present in the reaction mixture and when amplifying a long product, is used. In some embodiments, KAPA Taq DNA polymerase or KAPA Taq HotStart DNA polymerase is used. These are derived from single-subunit wild-type Taq DNA polymerase from the thermophilic bacterium Thermus aquaticus.KAPA Taq and KAPA Taq HotStart DNA Polymerase possess 5'-3' polymerase activity and 5'-3' exonuclease activity, but lack 3'-5' exonuclease (proofreading) activity (see, e.g., KAPA BIOSYSTEMS catalog number BK1000). In some embodiments, Pfu DNA polymerase is used. This polymerase is a high-temperature stable DNA polymerase derived from the hyperthermophilic archaeon Pyrococcus furiosus. This enzyme catalyzes template-dependent polymerization from nucleotides to double-stranded DNA in the 5'→3' direction. Pfu DNA Polymerase also exhibits 3'→5' exonuclease (proofreading) activity, allowing this polymerase to correct nucleotide integration errors. This polymerase does not possess 5'→3' exonuclease activity (see, e.g., Thermo Scientific catalog number EP0501). In some embodiments, Klentaq1 is used. This is a Klenow fragment analog of Taq DNA polymerase and does not possess exonuclease or endonuclease activity (see, e.g., DNA POLYMERASE TECHNOLOGY, Inc., St. Louis, Missouri, catalog number 100). In some embodiments, the polymerase is PHUSION DNA polymerase, e.g., PHUSION High Fidelity DNA polymerase (M0530S, New England BioLabs, Inc.) or PHUSION Hot Start Flex DNA polymerase (M0535S, New England BioLabs, Inc.). In some embodiments, the polymerase is Q5® DNA polymerase, for example, Q5® High-Fidelity DNA Polymerase (M0491S, New England BioLabs, Inc.) or Q5® Hot Start High-Fidelity DNA Polymerase (M0493S, New England BioLabs, Inc.). In some embodiments, the polymerase is T4 DNA polymerase (M0203S, New England BioLabs, Inc.).

[0182] In some embodiments, polymerases are used with concentrations of 5-600 units / mL (number of units per 1 mL of reaction volume), for example, 5-100, 100-200, 200-300, 300-400, 400-500, or 500-600 units / mL (including boundary values).

[0183] PCR method In some embodiments, hot-start PCR is used to reduce or prevent polymerization before the PCR thermal cycle. Exemplary hot-start PCR methods include initial inhibition of DNA polymerase, or physical separation of reaction components until the reaction mixture reaches a higher temperature. In some embodiments, delayed release of magnesium is used. Because DNA polymerase requires magnesium ions for activity, magnesium is chemically separated from the reaction by binding to a chemical compound and released into solution only at high temperatures. In some embodiments, non-covalent bonding of an inhibitor is used. In this method, a peptide, antibody, or aptamer non-covalently bonds to the enzyme at low temperatures, inhibiting its activity. After incubation at high temperatures, the inhibitor is released and the reaction begins. In some embodiments, cold-sensitive Taq polymerase, e.g., a modified DNA polymerase that has little activity at low temperatures, is used. In some embodiments, chemical modification is used. In this method, a molecule covalently bonds to the side chain of an amino acid at the active site of DNA polymerase. This molecule is released from the enzyme by incubating the reaction mixture at high temperatures. Once the molecule is released, the enzyme is activated.

[0184] In some embodiments, the amount of nucleic acid (e.g., RNA or DNA sample) to assemble with the template is 20 to 5,000 ng, for example, 20 to 200, 200 to 400, 400 to 600, 600 to 1,000, 1,000 to 1,500, or 2,000 to 3,000 ng (including boundary values).

[0185] In some embodiments, the QIAGEN Multiplex PCR Kit is used (QIAGEN catalog number 206143). For a 100 × 50 μl multiplex PCR reaction, the kit contains 2 × QIAGEN Multiplex PCR Master Mix (3 × 0.85 ml, providing a final concentration of 3 mM MgCl2), 5 × Q-Solution (1 × 2.0 ml), and RNase-Free Water (2 × 1.7 ml). The QIAGEN Multiplex PCR Master Mix (MM) contains a combination of KCl and (NH4)2SO4, as well as the PCR additive Factor MP, which increases the local concentration of primers in the template. Factor MP stabilizes specifically bound primers, enabling efficient primer extension by HotStarTaq DNA Polymerase. HotStarTaq DNA Polymerase is a modified form of Taq DNA polymerase that does not exhibit polymerase activity at ambient temperature. In some embodiments, HotStarTaq DNA Polymerase is activated by incubation at 95°C for 15 minutes, which can be incorporated into any existing thermal cycler program.

[0186] In some embodiments, 1×QIAGEN MM at final concentration (recommended concentration), 7.5 nM of each primer in the library, 50 mM TMAC, and 7 µl of DNA template in a final volume of 20 µl are used. In some embodiments, PCR thermal cycling conditions include 20 cycles of 10 minutes at 95°C (hot start), 30 seconds at 96°C, 15 minutes at 65°C, and 30 seconds at 72°C, followed by 2 minutes at 72°C (final extension), and then holding at 4°C.

[0187] In some embodiments, 2×QIAGEN MM final concentration (twice the recommended concentration), 2 nM of each primer in the library, 70 mM TMAC, and 7 µl of DNA template in a total volume of 20 µl are used. In some embodiments, up to 4 mM EDTA is also included. In some embodiments, PCR thermal cycling conditions include 10 minutes at 95°C (hot start), 30 seconds at 96°C, 20, 25, 30, 45, 60, 120 or 180 minutes at 65°C (25 cycles of 30 seconds at 72°C if applicable), then 2 minutes at 72°C (final extension), followed by holding at 4°C.

[0188] Another exemplary set of conditions involves a semi-nested PCR technique. The first PCR reaction uses a reaction volume of 20 μl containing each primer (forward and reverse outer primers) and DNA template in a library of 2 × QIAGEN MM at a final concentration of 1.875 nM. The thermal cycling parameters include 25 cycles of 10 minutes at 95°C, 30 seconds at 96°C, 1 minute at 65°C, 6 minutes at 58°C, 8 minutes at 60°C, 4 minutes at 65°C, and 30 seconds at 72°C, followed by 2 minutes at 72°C, followed by a hold at 4°C. Next, 2 μl of the obtained product, diluted to 1:200, is used as the input for the second PCR reaction. This reaction uses a reaction volume of 10 μl containing 1 × QIAGEN MM final concentration, 20 nM each of the inner forward primers and 1 μM reverse primer tags. The thermal cycling parameters include 15 cycles of 10 minutes at 95°C, 30 seconds at 95°C, 1 minute at 65°C, 5 minutes at 60°C, 5 minutes at 65°C, and 30 seconds at 72°C, followed by 2 minutes at 72°C, followed by a hold at 4°C. The annealing temperature may, in some cases, be higher than the melting points of some or all of the primers, as discussed herein (see U.S. Patent Application No. 14 / 918,544, filed October 20, 2015, which is incorporated herein by reference in its entirety).

[0189] Melting point (T mThe annealing temperature (T) is the temperature at which half (50%) of the double-stranded DNA of an oligonucleotide (e.g., a primer) and its complete complement dissociates, resulting in single-stranded DNA. A ) is the temperature at which the PCR protocol is performed. For conventional methods, this temperature is usually the lowest T of the primers used. m Because the temperature is 5°C lower, all possible near-double strands are formed (resulting in virtually all primer molecules binding to the template nucleic acid). While this is highly efficient, it is certain that lower temperatures will result in more nonspecific reactions. A One consequence of having too low a factor is that internal single-base mismatches or partial annealing may be tolerated, allowing the primer to anneal to a sequence other than the true target. In some embodiments of the present invention, T A is T m At a higher level, only a small portion of the target has annealed primer at a given moment (e.g., only about 1-5%). As these are extended, they are removed from the equilibrium of annealing and dissociation of the primer and target (extension is T m (To rapidly increase the temperature above 70°C), approximately 1-5% of the target will have primers. Therefore, by allowing the reaction to run for a longer time for annealing, approximately 100% of the target will be copied with each cycle.

[0190] In various embodiments, the annealing temperature is the melting point of at least 25, 50, 60, 70, 75, 80, 90, 95, or 100% of the non-identical primers (e.g., empirically measured or calculated T). m) is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13°C or 15°C higher than the upper limit of the range. In various embodiments, the annealing temperature is at least 25, 50, 75, 100, 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000, 50,000, 75,000, 100,000, or all melting points (e.g., empirically measured or calculated T) of non-identical primers. m ) is 1 to 15°C higher than (e.g., 1 to 10, 1 to 5, 1 to 3, 3 to 5, 5 to 10, 5 to 8, 8 to 10, 10 to 12 or 12 to 15°C (including boundary values)). In various embodiments, the annealing temperature is at least 25%, 50%, 60%, 70%, 75%, 80%, 90%, 95%, or all of the melting points of the non-identical primers (e.g., empirically measured or calculated T). m The temperature is 1–15°C higher than (e.g., 1–10, 1–5, 1–3, 3–5, 3–8, 5–10, 5–8, 8–10, 10–12, or 12–15°C (including boundary values)), and the length of the annealing process (per PCR cycle) is 5–180 minutes, e.g., 15–120 minutes, 15–60 minutes, 15–45 minutes, or 20–60 minutes (including boundary values).

[0191] Exemplary Multiplex PCR Method In various embodiments, long annealing times (as discussed herein and illustrated in Example 10) and / or low primer concentrations are used. In fact, in certain embodiments, limited primer concentrations and / or conditions are used. In various embodiments, the length of the annealing step ranges from 15, 20, 25, 30, 35, 40, 45 or 60 minutes at the lower end of the range to 20, 25, 30, 35, 40, 45, 60, 120 or 180 minutes at the upper end of the range. In various embodiments, the length of the annealing step (per PCR cycle) ranges from 30 to 180 minutes. For example, the annealing step may be 30 to 60 minutes, and the concentration of each primer may be less than 20, 15, 10 or 5 nM. In other embodiments, the primer concentration ranges from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 25 nM at the lower end of the range to 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, and 50 nM at the upper end of the range.

[0192] At high levels of multiplexing, the solution may become viscous due to the large amount of primer present. If the solution is too viscous, the primer concentration may be reduced to an amount that is still sufficient for the primer to bind to the template DNA. In various embodiments, 1,000 to 100,000 different primers are used, with the concentration of each primer being less than 20 nM, for example, less than 10 nM or 1 to 10 nM (including boundary values).

[0193] Detection of copy number variations (CNVs) In addition to SNVs and indels, the monitoring and detection methods for early recurrence and metastasis described herein can also benefit from the detection of CNVs.

[0194] In one embodiment, the present invention generally relates, at least in part, to an improved method for determining the presence or absence of copy number variations (e.g., deletions or duplications of chromosomal segments or entire chromosomes). This method is particularly useful for detecting small deletions or duplications that may be difficult to detect with high specificity and sensitivity using conventional methods due to the limited data available from the relevant chromosomal segments. This method includes improved analytical methods, improved bioassay methods, and combinations of improved analytical and bioassay methods. The methods of the present invention can also be used to detect deletions or duplications that are present in only a small percentage of the cells or nucleic acid molecules being tested. This makes it possible to detect deletions or duplications before the onset of a disease (e.g., in a precancerous state) or early in the disease, for example, before a large number of diseased cells (e.g., cancer cells) with deletions or duplications accumulate. More accurate detection of deletions or duplications associated with a disease or disorder enables improved methods for diagnosing, predicting, preventing, delaying, stabilizing, or treating that disease or disorder. Some deletions or duplications are known to be associated with cancer or severe intellectual or physical disability.

[0195] In another embodiment, the present invention generally relates, at least in part, to improved methods for detecting single nucleotide variants (SNVs). These improved methods include improved analytical methods, improved bioassay methods, and improved methods using a combination of improved analytical and bioassay methods. In certain exemplary embodiments, these methods are used to detect, diagnose, monitor, or stage cancer in samples (e.g., circulating free DNA samples) in which SNVs are present at very low concentrations (e.g., less than 10%, 5%, 4%, 3%, 2.5%, 2%, 1%, 0.5%, 0.25%, or 0.1% of the total number of normal copies of the SNV locus). That is, these methods are particularly well suited in samples in certain exemplary embodiments to which relatively low proportions of mutations or variants are present relative to the normal polymorphic alleles present for a locus. Finally, methods combining improved methods for detecting copy number variations with improved methods for detecting single nucleotide variants are provided herein.

[0196] The success of treating diseases such as cancer largely depends on early diagnosis, correct staging of the disease, selection of an effective treatment regimen, and close monitoring to prevent or detect recurrence. For cancer diagnosis, histological evaluation of tumor material obtained from tissue biopsy is often considered the most reliable method. However, the invasive nature of biopsy-based sampling makes it impractical for large-scale screening and regular follow-up. Therefore, this method has the advantage of being non-invasive and relatively low-cost, and suitable when a fast turnaround time is desired. Targeted sequencing usable with the method of the present invention requires fewer reads than shotgun sequencing (e.g., several hundred reads instead of 40 million reads), thereby reducing costs. Multiplex PCR and usable next-generation sequencing increase throughput and reduce costs.

[0197] In some exemplary embodiments, analysis of AAI patterns in ctDNA provides more detailed insights into the tumor's clonal architecture, helping to predict its therapeutic response and optimize treatment strategies. Therefore, in certain embodiments, mmPCR-NGS panels targeting clinically pathogenic CNVs and SNVs are selected. Such panels are particularly useful in patients with cancers in which CNVs represent a substantial proportion of the mutational burden, as is common in breast, ovarian, and lung cancers in certain exemplary embodiments.

[0198] In some embodiments, the method is used to detect deletions, duplications, or single nucleotide variants in an individual. Samples from an individual containing cells or nucleic acids suspected of having deletions, duplications, or single nucleotide variants may be analyzed. In some embodiments, the sample is derived from tissue or organs suspected of having deletions, duplications, or single nucleotide variants, e.g., cells or masses suspected of being cancerous. Using the method of the present invention, deletions, duplications, or single nucleotide variants present in only one or a few cells can be detected in a mixture containing cells with and without deletions, duplications, or single nucleotide variants. In some embodiments, cfDNA or cfRNA from blood samples derived from an individual is analyzed. In some embodiments, cfDNA or cfRNA is secreted by cells (e.g., cancer cells). In some embodiments, cfDNA or cfRNA is released by cells undergoing necrosis or apoptosis (e.g., cancer cells). Using the methods of the present invention, deletions, duplications, or single-nucleotide variants present in only a small percentage of cfDNA or cfRNA can be detected. In some embodiments, one or more cells derived from embryos are tested.

[0199] In addition to determining the presence or absence of copy number variations, one or more other factors may be analyzed as desired. These factors can be used to improve the accuracy of the diagnosis (e.g., determining the presence or absence of cancer or an increased risk of cancer, classifying cancer, or determining the stage of cancer) or the accuracy of the prognosis. These factors can also be used to select a particular therapy or treatment regimen that is likely to be effective in the subject. Exemplary factors include the presence or absence of polymorphisms or mutations, changes (increases or decreases) in the levels of overall or specific cfDNA, cfRNA, or microRNA (miRNA), changes (increases or decreases) in the tumor fraction, changes (increases or decreases) in methylation levels, changes (increases or decreases) in DNA integrity, changes (increases or decreases) or alternative mRNA splicing.

[0200] The following chapters describe methods for detecting deletions or duplications using phasing data (e.g., inferred or measured phasing data) or non-phasing data, testable samples, sample preparation, amplification and quantification methods, methods for phasing genetic data, detectable polymorphisms, mutations, nucleic acid changes, mRNA splicing changes and changes at the nucleic acid level, databases derived from this method, other risk factors and screening methods, cancers that can be diagnosed or treated, cancer treatments, cancer models for testing treatments, and methods for prescribing and administering treatments.

[0201] An exemplary method for determining ploidy using fading data Some of the methods of the present invention are based in part on the finding that using phasing data to detect CNVs reduces false negative and false positive rates compared to using non-phasing data. This improvement is greatest for samples with CNVs present at low levels. Therefore, phasing data improves the accuracy of CNV detection compared to using non-phasing data (for example, a method that aggregates allele ratios to give aggregate values ​​(e.g., mean values) across a chromosome or chromosomal segment without calculating allele ratios at one or more loci or considering whether allele ratios at different loci appear to indicate the presence of the same or different haplotypes in abnormal amounts). By using phasing data, it becomes possible to make a more accurate determination as to whether the difference between measured and predicted allele ratios is due to noise or to the presence of CNVs. For example, if the difference between measured and predicted allele ratios at most or all loci within a region indicates an overpopulation of the sample haplotype, then CNVs are likely to be present. By using the allele junctions in haplotypes, it is possible to determine whether the measured genetic data corresponds to the same haplotype that is overpopulating (rather than random noise). In contrast, if the difference between the measured allele ratio and the predicted allele ratio is due solely to noise (e.g., experimental error), in some embodiments, the first haplotype may appear to be overpopulating for about half the time, and the second haplotype for about the other half of the time.

[0202] In some embodiments, phasing gene data is used to determine whether there is copy number overpopulation of the first homologous chromosome segment compared to a second homologous chromosome segment in the genome of an individual (e.g., in the genome of one or more cells, or in cfDNA or cfRNA). Exemplary overpopulations include duplication of the first homologous chromosome segment or deletion of the second homologous chromosome segment. In some embodiments, there is no overpopulation because the first chromosome segment and homologous chromosome segments are present in equal proportions (e.g., one copy of each segment in a diploid sample). In some embodiments, calculated allele ratios in a nucleic acid sample are compared to predicted allele ratios to determine whether there is overpopulation, as further described below. In this specification, the phrase "first homologous chromosome segment compared to a second homologous chromosome segment" means the first homolog of the chromosome segment and the second homolog of the chromosome segment.

[0203] In some embodiments, the method includes obtaining phasing gene data for a first homologous chromosome segment, including the identity of the allele present at each locus in the set of polymorphic loci on the first homologous chromosome segment; obtaining phasing gene data for a second homologous chromosome segment, including the identity of the allele present at each locus in the set of polymorphic loci on the second homologous chromosome segment; and obtaining measured gene allele data for each allele at each locus in the set of polymorphic loci, including the amount of each allele present in DNA or RNA samples from one or more target cells and one or more non-target cells from an individual. In some embodiments, the method includes: listing a set of one or more hypotheses indicating the degree of overpopulation of a first homologous chromosome segment; for each of the above hypotheses, calculating predicted gene data for multiple loci in the sample from phasing gene data obtained for one or more possible ratios of DNA or RNA from one or more target cells to total DNA or RNA in the sample; calculating (e.g., by computer) a data fitting between the obtained gene data of the sample and the predicted gene data for the sample for each possible ratio of DNA or RNA and for each hypothesis; ranking the one or more above hypotheses according to this data fitting; and determining the degree of copy number overpopulation of a first homologous chromosome segment in the genome of one or more cells from an individual by selecting the highest-ranked hypothesis.

[0204] In some embodiments, the method involves obtaining phasing gene data using any of the methods described herein or any known method. In some embodiments, the method involves simultaneously or sequentially in any order obtaining phasing gene data for a first homologous chromosome segment, including the identity of the alleles present at each locus in the set of polymorphic loci on the first homologous chromosome segment; obtaining phasing gene data for a second homologous chromosome segment, including the identity of the alleles present at each locus in the set of polymorphic loci on the second homologous chromosome segment; and obtaining measured gene allele data, including the amount of each allele for each locus in the set of polymorphic loci in a DNA sample from one or more cells from an individual.

[0205] In some embodiments, the method involves calculating allele ratios for one or more loci in a set of heterozygous polymorphic loci from at least one cell from which the sample originates. In some embodiments, the calculated allele ratio for a particular locus is obtained by dividing one allele measurement by the total allele measurements for all alleles at that locus. In some embodiments, the calculated allele ratio for a particular locus is obtained by dividing one allele measurement (e.g., an allele on a first homologous chromosome segment) by the allele measurements for one or more other alleles at that locus (e.g., alleles on a second homologous chromosome segment). The calculated allele ratios may be calculated using any of the methods described herein or any standard method (e.g., any mathematical transformation of the calculated allele ratios described herein).

[0206] In some embodiments, the method involves determining whether there is an over-recurrence of the first homologous chromosome segment by comparing a calculated allele ratio for one or more alleles for a given locus with a predicted allele ratio for that locus, given that the first homologous chromosome segment and the second homologous chromosome segment are present in equal proportions. In some embodiments, the predicted allele ratio assumes that the likelihood of multiple possible alleles existing for a given locus is equal. In some embodiments, where the calculated allele ratio for a particular locus is obtained by dividing a single allele measurement by the total allele measurements for all alleles at that locus, the corresponding predicted allele ratio is 0.5 for two allele loci or 1 / 3 for three allele loci. In some embodiments, the predicted allele ratio is the same for all loci, for example, 0.5 for all loci. In some embodiments, the predicted allele ratio assumes that the likelihood of possible alleles existing for a given locus may differ, for example, based on the frequency of each allele in a particular set to which the subject belongs (e.g., a set based on the subject's ancestors). Such allele frequencies are publicly available (e.g., HapMap Project; Perlegen Human Haplotype Project; at ncbi.nlm.nih.gov / projects / SNP / on the web; see Sherry ST, Ward MH, Kholodov M et al., dbSNP: the NCBI database of genetic variation. Nucleic Acids Res. January 1, 2001; 29(1):308-11, each incorporated herein in its entirety by reference). In some embodiments, the predicted allele ratio is the predicted allele ratio for a particular individual being tested for a specific hypothesis indicating the degree of overpopulation of a first homologous chromosome segment.For example, the predicted allele ratio for a particular individual may be determined based on phasing gene data or non-phasing gene data from that individual (e.g., samples from that individual that are unlikely to have deletions or duplications, such as non-cancerous samples), or data from one or more relatives of that individual.

[0207] In some embodiments, the calculated allele ratio is an indicator of copy number overpopulation in the first homologous chromosome segment if (i) the allele ratio for the measured amount of alleles present at a locus on the first homologous chromosome segment is greater than the predicted allele ratio for that locus when divided by the total measured amount of all alleles at that locus, or (ii) the allele ratio for the measured amount of alleles present at a locus on the second homologous chromosome segment is less than the predicted allele ratio for that locus when divided by the total measured amount of all alleles at that locus. In some embodiments, the calculated allele ratio is considered an indicator of overpopulation only if it is significantly larger or smaller than the predicted ratio for that locus. In some embodiments, the calculated allele ratio is an indicator of no copy number overpopulation of the first homologous chromosome segment if (i) the allele ratio for the measured amount of alleles present at a locus on the first homologous chromosome segment divided by the total measured amount of all alleles at that locus is less than or equal to the predicted allele ratio for that locus, or (ii) the allele ratio for the measured amount of alleles present at a locus on the second homologous chromosome segment divided by the total measured amount of all alleles at that locus is greater than or equal to the predicted allele ratio for that locus. In some embodiments, calculated ratios that are equal to the predicted values ​​of the corresponding ratios are ignored (because these are indicators of no overpopulation).

[0208] In various embodiments, one or more of the following methods are used to compare one or more calculated allele ratios with the corresponding predicted allele ratios. In some embodiments, it is determined whether the calculated allele ratio is above or below the predicted allele ratio for a particular locus, regardless of the magnitude of the difference. In some embodiments, the magnitude of the difference between the calculated allele ratio and the predicted allele ratio for a particular locus is determined, regardless of whether the calculated allele ratio is above or below the predicted allele ratio. In some embodiments, it is determined whether the calculated allele ratio is above or below the predicted allele ratio, and the magnitude of the difference for a particular locus. In some embodiments, it is determined whether the mean or weighted mean of the calculated allele ratios is above or below the mean or weighted mean of the predicted allele ratios, regardless of the magnitude of the difference. In some embodiments, the magnitude of the difference between the average or weighted average of the calculated allele ratios and the average or weighted average of the predicted allele ratios is determined, regardless of whether the average or weighted average of the calculated allele ratios is above or below the average or weighted average of the predicted allele ratios. In some embodiments, the magnitude of the difference between the average or weighted average of the calculated allele ratios and the predicted allele ratios is determined, and whether the average or weighted average of the calculated allele ratios is above or below the average or weighted average of the predicted allele ratios. In some embodiments, the average or weighted average of the magnitude of the difference between the calculated allele ratios and the predicted allele ratios is determined.

[0209] In some embodiments, the magnitude of the difference between the calculated allele ratio and the predicted allele ratio for one or more loci is used to determine whether the over-recurrence of the first homologous chromosome segment is due to duplication of the first homologous chromosome segment or deletion of the second homologous chromosome segment in the genome of one or more cells.

[0210] In some embodiments, over-reproduction of the first homologous chromosome segment is determined to exist if one or more of the following conditions are met: In some embodiments, the calculated value of the allele ratio, which is an indicator of over-reproduction of the first homologous chromosome segment, is above a threshold. In some embodiments, the calculated value of the allele ratio, which is an indicator of the absence of over-reproduction of the first homologous chromosome segment, is below a threshold. In some embodiments, the magnitude of the difference between the calculated value of the allele ratio, which is an indicator of over-reproduction of the first homologous chromosome segment, and the predicted value of the corresponding allele ratio is above a threshold. In some embodiments, for all calculated allele ratios that are indicators of over-reproduction, the sum of the magnitudes of the differences between the calculated allele ratio and the predicted value of the corresponding allele ratio is above a threshold. In some embodiments, the magnitude of the difference between the calculated value of the allele ratio, which is an indicator of the absence of over-reproduction of the first homologous chromosome segment, and the predicted value of the corresponding allele ratio is below a threshold. In some embodiments, the average or weighted mean of the calculated allele ratios for alleles present on a first homologous chromosome segment, divided by the total alleles for that locus, is greater than the average or weighted mean of the predicted allele ratios by at least one threshold. In some embodiments, the average or weighted mean of the calculated allele ratios for alleles present on a second homologous chromosome segment, divided by the total alleles for that locus, is less than the average or weighted mean of the predicted allele ratios by at least one threshold. In some embodiments, the data fitting between the calculated allele ratios and the predicted allele ratios for the copy number overpopulation of the first homologous chromosome segment is below a threshold (an indicator of good data fitting).In some embodiments, the data fitting between the calculated allele ratio and the allele ratio predicted by the absence of copy number overpopulation of the first homologous chromosome segment exceeds a threshold (an indicator of poor data fitting).

[0211] In some embodiments, the over-recurrence of the first homologous chromosome segment is determined to be absent if one or more of the following conditions are met: In some embodiments, the calculated value of the allele ratio, which is an indicator of the over-recurrence of the first homologous chromosome segment, is below a threshold. In some embodiments, the calculated value of the allele ratio, which is an indicator of the absence of the over-recurrence of the first homologous chromosome segment, is above a threshold. In some embodiments, the magnitude of the difference between the calculated value of the allele ratio, which is an indicator of the over-recurrence of the first homologous chromosome segment, and the predicted value of the corresponding allele ratio is below a threshold. In some embodiments, the magnitude of the difference between the calculated value of the allele ratio, which is an indicator of the absence of the over-recurrence of the first homologous chromosome segment, and the predicted value of the corresponding allele ratio is above a threshold. In some embodiments, the calculation or weighted mean of the allele ratio for alleles present on a first homologous chromosome segment, divided by the total alleles present at that locus, minus the calculation or weighted mean of the predicted allele ratio, falls below a threshold. In some embodiments, the calculation or weighted mean of the allele ratio for all alleles present on a second homologous chromosome segment, subtracted from the calculation or weighted mean of the predicted allele ratio, divided by the total alleles present at that locus, falls below a threshold. In some embodiments, the data fitting between the calculated allele ratio and the predicted allele ratio for copy number overpopulation of the first homologous chromosome segment exceeds a threshold. In some embodiments, the data fitting between the calculated allele ratio and the predicted allele ratio for the absence of copy number overpopulation of the first homologous chromosome segment falls below a threshold. In some embodiments, the threshold is determined from empirical testing of samples known to have the CNV of the target and / or samples known to lack the CNV.

[0212] In some embodiments, determining whether there is copy number overpopulation of a first homologous chromosome segment involves enumerating a set of one or more hypotheses indicating the degree of overpopulation of the first homologous chromosome segment. An exemplary hypothesis is that there is no overpopulation because homologous chromosome segments to the first chromosome segment are present in equal proportions (e.g., one copy of each segment in a diploid sample). Another exemplary hypothesis involves the first homologous chromosome segment being replicated one or more times (e.g., one, two, three, four, five or more excess copies of the first homologous chromosome segment compared to the copy number of the second homologous chromosome segment). Yet another exemplary hypothesis involves deletion of the second homologous chromosome segment. Yet another exemplary hypothesis is deletion of both the first and second homologous chromosome segments. In some embodiments, predictive allele ratios for a locus that is heterozygous in at least one cell are estimated for each hypothesis, taking into account the degree of overpopulation indicated by that hypothesis. In some embodiments, the likelihood that a hypothesis is correct is calculated by comparing the calculated allele ratios with the predicted allele ratios, and the hypothesis with the highest likelihood is selected.

[0213] In some embodiments, the expected distribution of the test statistics is calculated using the predicted allele ratios for each hypothesis. In some embodiments, the likelihood that a hypothesis is correct is calculated by comparing the test statistics calculated using the calculated allele ratios with the expected distribution of the test statistics calculated using the predicted allele ratios, and the hypothesis with the highest likelihood is selected.

[0214] In some embodiments, the predicted allele ratio for a locus that is heterozygous in at least one cell is estimated considering phasing gene data for a first homologous chromosome segment, phasing gene data for a second homologous chromosome segment, and the degree of overpopulation indicated by the hypothesis. In some embodiments, the likelihood that the hypothesis is correct is calculated by comparing the calculated allele ratio with the predicted allele ratio, and the hypothesis with the highest likelihood is selected.

[0215] Use of mixed samples In many embodiments, the sample will be understood to be a mixed sample containing DNA or RNA from one or more target cells and one or more non-target cells. In some embodiments, the target cells are cells having a CNV (e.g., a deletion or duplication of interest), and the non-target cells are cells that do not have the copy number variation of interest (e.g., a mixture of cells having the deletion or duplication of interest and cells that do not contain any of the deletions or duplications to be tested). In some embodiments, the target cells are cells associated with a certain disease or disorder or an increased risk of disease or disorder (e.g., cancer cells), and the non-target cells are cells not associated with a certain disease or disorder or an increased risk of disease or disorder (e.g., non-cancerous cells). In some embodiments, all target cells have the same CNV. In some embodiments, two or more target cells have different CNVs. In some embodiments, one or more of the target cells have a CNV, polymorphism, or mutation associated with the disease or disorder or an increased risk of the disease or disorder that is not found in at least one of the other target cells. In some such embodiments, it is assumed that among all cells from a sample, the portion of cells associated with the disease or disorder or an increased risk of the disease or disorder is greater than or equal to the portion of the sample with the highest frequency of these CNVs, polymorphisms, or mutations. For example, if 6% of the cells have a K-ras mutation and 8% of the cells have a BRAF mutation, it is assumed that at least 8% of the cells are cancerous.

[0216] In some embodiments, the ratio of DNA (or RNA) from one or more target cells to the total DNA (or RNA) in the sample is calculated. In some embodiments, a set of one or more hypotheses indicating the degree of overpopulation of a first homologous chromosome segment is enumerated. In some embodiments, predicted allele ratios for loci that are heterozygous in at least one cell are estimated considering the calculated DNA or RNA ratios, and the degree of overpopulation indicated by each hypothesis is estimated for each hypothesis. In some embodiments, the likelihood that a hypothesis is correct is calculated by comparing the calculated allele ratios with the predicted allele ratios, and the hypothesis with the highest likelihood is selected.

[0217] In some embodiments, a predicted distribution of test statistics calculated using predicted allele ratios and calculated DNA or RNA ratios is estimated for each hypothesis. In some embodiments, the likelihood that the hypothesis is correct is determined by comparing the test statistics calculated using the calculated allele ratios and DNA or RNA ratios with the predicted distribution of test statistics calculated using predicted allele ratios and DNA or RNA ratios, and the hypothesis with the highest likelihood is selected.

[0218] In some embodiments, the method includes listing a set of one or more hypotheses indicating the degree of overpopulation of a first homologous chromosome segment. In some embodiments, for each hypothesis, the method includes estimating either (i) a predicted allele ratio for a locus that is heterozygous in at least one cell, taking into account the degree of overpopulation indicated by the hypothesis, or (ii) a predictive distribution of test statistics calculated using the predicted allele ratio and the possible ratio of DNA or RNA from one or more target cells to the total DNA or RNA in the sample for one or more possible ratios of DNA or RNA. In some embodiments, data fitting is performed by comparing (i) the calculated allele ratio with the predicted allele ratio, or (ii) a test statistics calculated using the calculated allele ratio and the possible ratios of DNA or RNA with the predictive distribution of test statistics calculated using the predicted allele ratio and the possible ratios of DNA or RNA. In some embodiments, one or more hypotheses are ranked according to the data fitting, and the highest-ranked hypothesis is selected. In some embodiments, a technique or algorithm, such as a search algorithm, is used in one or more of the steps of calculating data fitting, ranking hypotheses, or selecting the highest-ranked hypothesis. In some embodiments, the data fitting is fitting to a beta-binomial distribution or fitting to a binomial distribution. In some embodiments, the technique or algorithm is selected from the group consisting of maximum likelihood estimation, empirical maximum estimation, Bayesian estimation, dynamic estimation (e.g., dynamic Bayesian estimation), and expectation maximization estimation. In some embodiments, the method includes applying the techniques or algorithms described above to the obtained gene data and the predicted values ​​of the gene data.

[0219] In some embodiments, the method involves creating a distribution of possible ratios ranging from a lower limit to an upper limit for the ratio of DNA or RNA from one or more target cells to the total DNA or RNA in the sample. In some embodiments, a set of one or more hypotheses indicating the degree of overpopulation of a first homologous chromosome segment is enumerated. In some embodiments, the method involves, for each possible ratio of DNA or RNA in the distribution and for each hypothesis, estimating either (i) a predicted allele ratio for a locus that is heterozygous in at least one cell, taking into account the possible ratio of DNA or RNA and the degree of overpopulation indicated by the hypothesis, or (ii) a predicted distribution of test probabilities calculated using the predicted allele ratio and the possible ratio of DNA or RNA. In some embodiments, the method calculates the likelihood that a hypothesis is correct for each possible ratio of DNA or RNA in the distribution and for each hypothesis by comparing (i) the calculated allele ratio with the predicted allele ratio, or (ii) the test statistics calculated using the calculated allele ratio and the possible ratio of DNA or RNA, with the predicted distribution of the test statistics calculated using the predicted allele ratio and the possible ratio of DNA or RNA. In some embodiments, the binding probability for each hypothesis is determined by combining the probabilities of the hypothesis for each possible ratio in the distribution, and the hypothesis with the highest binding probability is selected. In some embodiments, the binding probability for each hypothesis is determined by weighting the probability of a given hypothesis for a particular possible ratio based on the likelihood that the possible ratio is the correct ratio.

[0220] In some embodiments, the ratio of DNA or RNA from one or more target cells to total DNA or RNA in a sample is estimated using a technique selected from the group consisting of maximum likelihood estimation, empirical maximum estimation, Bayesian estimation, dynamic estimation (e.g., dynamic Bayesian estimation), and expectation maximization estimation. In some embodiments, it is assumed that the ratio of DNA or RNA from one or more target cells to total DNA or RNA in a sample is the same for two or more (or all) of the CNVs of interest. In some embodiments, for each CNV of interest, the ratio of DNA or RNA from one or more target cells to total DNA or RNA in a sample is calculated.

[0221] An example of using incomplete fading data It should be understood that in many embodiments, incomplete phasing data is used. For example, it may not be 100% certain which alleles are present at one or more loci on the first and / or second homologous chromosome segments. In some embodiments, prior probabilities for possible haplotypes of an individual (e.g., haplotypes based on set-based haplotype frequencies) are used when calculating the probability of each hypothesis. In some embodiments, prior probabilities for possible haplotypes are refined by using a different method for phasing the genetic data, or by using phasing data from other subjects (e.g., previous subjects) to refine the set-based data used for phasing based on individual informatics.

[0222] In some embodiments, the phasing gene data includes probability data for two or more possible sets of phasing gene data, each possible set of phasing gene data including possible allele identities present at each locus in a set of polymorphic loci on a first homologous chromosome segment and possible allele identities present at each locus in a set of polymorphic loci on a second homologous chromosome segment. In some embodiments, the probability for at least one of the hypotheses is determined for each possible set of phasing gene data. In some embodiments, the binding probability for a hypothesis is determined by combining the probabilities of that hypothesis for each possible set of phasing gene data, and the hypothesis with the highest binding probability is selected.

[0223] Incomplete phasing data for use in the claimed method may be prepared using any of the methods disclosed herein or any known method (e.g., using set-based haplotype frequencies to infer the most likely phase). In some embodiments, phasing data is obtained by probabilistically combining haplotypes from smaller segments. For example, possible haplotypes may be determined based on possible combinations of one haplotype from a first region and another haplotype from another region on the same chromosome. The probability that particular haplotypes from different regions are part of the same, larger haplotype block on the same chromosome may be determined, for example, using set-based haplotype frequencies and / or known recombination rates between different regions.

[0224] In some embodiments, a single-hypothesis rejection test is used for the null hypothesis of disomy. In some embodiments, the probability of the disomy hypothesis is calculated, and the disomy hypothesis is rejected if its probability falls below a given threshold (e.g., less than 1 in 1,000). If the null hypothesis is rejected, this may be due to an error in incomplete phasing data or to the presence of a CNV. In some embodiments, more accurate phasing data is obtained (e.g., phasing data from one of the molecular phasing methods disclosed herein for obtaining actual phasing data, rather than phasing data inferred based on bioinformatics). In some embodiments, the probability of the disomy hypothesis is recalculated using this more accurate phasing data to determine whether the disomy hypothesis should still be rejected. Rejection of this hypothesis indicates the presence of a chromosomal segment duplication or deletion. If desired, the false positive rate can be altered by adjusting the threshold.

[0225] Further exemplary embodiments for determining ploidy using fading data In an exemplary embodiment, a method for determining the ploidy of a chromosomal segment in a sample of an individual is provided herein. The method includes: receiving allele frequency data, including the amount of each allele present in the sample, at each locus in a set of polymorphic loci on the chromosomal segment; creating phasing allele information for the set of polymorphic loci by estimating the phase of the allele frequency data; creating individual probabilities of allele frequencies for different ploidy states for the polymorphic locus using the allele frequency data; creating binding probabilities for the set of polymorphic loci using the individual probabilities and phasing allele information; and determining the ploidy of the chromosomal segment by selecting a best-fitting model, which is an index of chromosomal ploidy, based on the binding probabilities.

[0226] As disclosed herein, allele frequency data (also referred herein as allele data of the gene to be measured) may be prepared by methods known in the art. For example, the data may be prepared using qPCR or microarrays. In one exemplary embodiment, the data is generated using nucleic acid sequence data, in particular high-throughput nucleic acid sequence data.

[0227] In certain exemplary embodiments, allele frequency data are corrected for errors before being used to create individual probabilities. In specific exemplary embodiments, the errors corrected include allele amplification efficiency bias. In other embodiments, the errors corrected include ambient contamination and genotype contamination. In some embodiments, the errors corrected include allele amplification bias, sequencing errors, ambient contamination, and genotype contamination.

[0228] In certain embodiments, individual probabilities are constructed using a set of models for different ploidy states and allelic imbalance fractions for a set of polymorphic loci. In these embodiments and other embodiments, binding probabilities are constructed by considering binding between polymorphic loci on chromosomal segments.

[0229] Accordingly, in an exemplary embodiment combining some of these embodiments, a method for detecting chromosomal ploidy in a sample of an individual is provided herein, comprising the steps of: receiving nucleic acid sequence data for alleles in a set of polymorphic loci on a chromosomal segment in the individual; detecting allele frequencies in the set of loci using the nucleic acid sequence data; correcting for allele amplification efficiency bias in the detected allele frequencies to create corrected allele frequencies for the set of polymorphic loci; creating phasing allele information for the set of polymorphic loci by estimating the phase of the nucleic acid sequence data; creating individual probabilities of allele frequencies for different ploidy states in the polymorphic loci for different ploidy states by comparing the corrected allele frequencies with a set of models of different ploidy states and allele imbalance fractions for the set of polymorphic loci; creating a binding probability for the set of polymorphic loci by combining individual probabilities that take into account binding between polymorphic loci on a chromosomal segment; and selecting a best-fitting model, which is an indicator of chromosomal aneuploidy, based on the binding probability.

[0230] As disclosed herein, individual probabilities may be constructed using a set of different ploidy states and mean allele imbalance fraction models or hypotheses for a set of polymorphic loci. For example, in particularly exemplary cases, individual probabilities are constructed by modeling the ploidy states of the first homolog and the second homolog of the chromosomal segment. The ploidy states to be modeled include: (1) all cells have no deletion or amplification of the first or second homolog of the chromosomal segment; (2) at least some cells have a deletion or amplification of the first homolog of the chromosomal segment; and (3) at least some cells have a deletion or amplification of the first homolog of the chromosomal segment.

[0231] It will be understood that the model above may also be referred to as the hypothesis used to constrain the model. Therefore, the three available hypotheses are shown above.

[0232] The mean allele disequilibrium fraction to be modeled may include any range of mean allele disequilibrium, including the actual mean allele disequilibrium of the chromosomal segment. For example, in a particular exemplary embodiment, the range of mean allele disequilibrium to be modeled may be 0, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.75, 1, 2, 2.5, 3, 4, and 5% at the lower limit, and 1, 2, 2.5, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 95, and 99% at the upper limit. The interval for modeling with this range may be any interval, depending on the computing power used and the time allowed for analysis. For example, intervals of 0.01, 0.05, 0.02, or 0.1 may be modeled.

[0233] In certain exemplary embodiments, the sample has a mean allele mismatch of 0.4% to 5% for chromosomal segments. In certain embodiments, the mean allele mismatch is low. In these embodiments, the mean allele mismatch is typically less than 10%. In certain exemplary embodiments, the allele mismatch is 0.25, 0.3, 0.4, 0.5, 0.6, 0.75, 1, 2, 2.5, 3, 4, and 5% at the lower limit and 1, 2, 2.5, 3, 4, and 5% at the upper limit. In other exemplary embodiments, the mean allele mismatch is 0.4, 0.45, 0.5, 0.6, 0.7, 0.8, 0.9, or 1.0% at the lower limit and 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.5, 2.0, 3.0, 4.0, or 5.0% at the upper limit. For example, the mean allele imbalance in a sample is 0.45–2.5% in an exemplary example. In another example, the mean allele imbalance is detected with sensitivities of 0.45, 0.5, 0.6, 0.8, 0.8, 0.9, or 1.0%. In other words, this test method can detect chromosomal aneuploidy where AAI drops to 0.45, 0.5, 0.6, 0.8, 0.8, 0.9, or 1.0%. Exemplary samples with low allele imbalance in the method of the present invention include plasma samples from individuals with cancer containing circulating tumor DNA or plasma samples from pregnant women containing circulating fetal DNA.

[0234] For SNVs, it will be understood that the proportion of abnormal DNA is typically measured using the mutant allele frequency (number of mutant alleles at a given locus / total number of alleles at that locus). Since the difference in the amounts of the two homologs in the tumor is similar, the mean allele imbalance (AAI) is used to measure the proportion of abnormal DNA for CNVs (defined as |(H1-H2)| / (H1+H2)), where Hi is the mean copy number of homolog i in the sample, and Hi / (H1+H2) is the fraction of homolog i present, i.e., the homolog ratio. The maximum homolog ratio is the homolog ratio of the more abundant homolog.

[0235] The assay dropout rate is the proportion of SNPs that do not have reads, estimated using all SNPs. The single allele dropout (ADO) rate is the proportion of SNPs that have only one allele, estimated using only heterozygous SNPs. Genotype reliability is determined by fitting the binomial distribution to the number of reads at each SNP that was a B allele read, using the ploidy state of the focal region of the SNP, and the probability of each genotype can be estimated.

[0236] For tumor tissue samples, chromosomal aneuploidy (exemplified in this paragraph by CNV) can be represented by transitions between allele frequency distributions. In plasma samples from cancer patients, individuals suspected of having cancer, individuals previously diagnosed with cancer, or as part of cancer screening for individuals at risk or general populations, CNVs can be identified by a maximum likelihood algorithm that searches for plasma CNVs in regions known to exhibit aneuploidy in cancer, and / or if tumor samples from the same individual also have CNVs. In an exemplary embodiment, this algorithm uses haplotype phase information of the individual from which the sample is analyzed for the presence of circulating tumor DNA to fit the measured and corrected allele counts of the test sample to predicted allele counts, for example, using a binding distribution mode. Such haplotype phase information can be deduced from parental genetic information or by de novo haplotype phasing from any sample from an individual, e.g., buffy coat samples, saliva samples, or skin samples, containing the majority, or at least 60, 70, 80, 90, 95, 96, 97, 98, 99%, or all normal cellular DNA, e.g., buffy coat samples, saliva samples, or skin samples, e.g., from parental genetic information, or by de novo haplotype phasing. These can be deduced by various methods (e.g., Snyder, M. et al., Haplotype-resolved genome sequencing: experimental methods and applications. Nat Rev Genet 16, 344-358 (2015)), e.g., haplotyping by dilution (Kaper, F. et al., Whole-genome haplotyping by dilution, amplification, and sequencing. Proc Natl Acad Sci USA 110, 5552-5557 (2013)) or long-read sequencing (Kuleshov, V. et al., Whole-genome haplotyping using long This can be achieved by readings and statistical methods. (Nat Biotech 32, 261-266 (2014)).This algorithm can model predicted allele frequencies across all allele imbalance ratios at 0.025% intervals for three sets of hypotheses: (1) all cells are normal (no allele imbalance), (2) some / all cells have homolog 1 deletion or homolog 2 amplification, or (3) some / all cells have homolog 2 deletion or homolog 1 amplification. The likelihood of each hypothesis can be determined for each SNP using a Bayesian classifier based on a beta-binomial model of predicted and observed allele frequencies at all heterozygous SNPs, and then the binding likelihood across multiple SNPs can be calculated, taking into account the binding of SNP loci, as illustrated herein in certain exemplary embodiments. In fact, in exemplary embodiments, the haplotype phase information of normal cells obtained as disclosed above is measured and typically used by the algorithm to fit the predicted allele numbers of the test sample to the predicted allele numbers using a binding distribution model. The maximum likelihood hypothesis can then be selected.

[0237] Considering the chromosomal regions in the tumor that have an average of N copies, c represents the fraction of DNA in the plasma derived from a mixture of normal and tumor cells in the disomy region. AAI is calculated as follows:

number

[0238] In certain exemplary embodiments, allele frequency data are corrected for errors before being used to create individual probabilities. Different types of error and / or bias corrections are disclosed herein. In specific exemplary embodiments, the error corrected is allele amplification efficiency bias. In other embodiments, errors corrected include sequencing errors, ambient contamination, and genotype contamination. In some embodiments, errors corrected include allele amplification bias, sequencing errors, ambient contamination, and genotype contamination.

[0239] It will be understood that allele amplification efficiency bias can be determined for a particular allele as part of an experimental or laboratory decision involving the sample under test, or it can be determined at different times using a set of samples containing the alleles for which efficiency is calculated. Ambient contamination and genotype contamination are typically determined in the same run as the sample analysis under test.

[0240] In certain embodiments, surrounding contamination and genotype contamination are determined for homozygous alleles in the sample. It will be understood that some loci in the sample are heterozygous and others homozygous, even if, for any given sample from an individual, a certain locus is selected for analysis because it has relatively high heterozygosity within the set. In some embodiments, it is advantageous to determine the ploidy of a chromosomal segment using heterozygous loci for a given individual, while surrounding contamination and genotype contamination can be calculated using homozygous loci.

[0241] In certain exemplary cases, the selection described above is performed by analyzing the magnitude of the difference between the phasing allele information and the estimated allele frequencies created for the model.

[0242] In the illustrative example, the individual probabilities of allele frequencies are constructed based on a beta-binomial model of predicted and observed allele frequencies for a set of polymorphic loci. In the illustrative example, the individual probabilities are constructed using a Bayesian classifier.

[0243] In certain exemplary embodiments, nucleic acid sequence data are generated by high-throughput DNA sequencing of multiple copies of a series of amplicons created using a multiple amplification reaction, where each amplicon in the series spreads to at least one polymorphic locus of a set of polymorphic loci, and each of these polymorphic loci is amplified. In certain embodiments, the multiple amplification reaction is carried out under restricted primer conditions for at least half of the reaction. In some embodiments, the restricted primer concentration is used for 1 / 10, 1 / 5, 1 / 4, 1 / 3, 1 / 2, or all of the reactions in the multiple amplification reaction. Factors to consider for achieving restricted primer conditions in amplification reactions such as PCR are provided herein.

[0244] In certain embodiments, the methods provided herein detect ploidy for multiple chromosomal segments across multiple chromosomes. Thus, chromosomal ploidy in these embodiments is determined for a set of chromosomal segments in a sample. For these embodiments, more multiple amplification reactions are required. Therefore, for these embodiments, the multiple amplification reactions may include, for example, 2,500 to 50,000 multiple reactions. In certain embodiments, multiple reactions are performed within the following ranges: from 100, 200, 250, 500, 1000, 2500, 5000, 10,000, 20,000, 25000, and 50000 at the lower end of the range, to 200, 250, 500, 1000, 2500, 5000, 10,000, 20,000, 25000, 50000, and 100,000 at the upper end of the range.

[0245] In exemplary embodiments, a set of polymorphic loci is a set of loci known to exhibit high heterozygosity. However, for any given individual, some of these loci are expected to be homozygous. In certain exemplary embodiments, the method of the present invention utilizes nucleic acid sequence information for both homozygous and heterozygous loci of an individual. The homozygous loci of an individual are used, for example, for error correction, while the heterozygous loci are used to determine allele imbalances in a sample. In certain embodiments, at least 10% of the polymorphic loci are heterozygous loci of the individual.

[0246] As disclosed herein, it is preferable to analyze target SNP loci known to be heterozygous in the aggregate. Thus, in certain embodiments, 10, 20, 25, 50, 75, 80, 90, 95, 99, or 100% of the polymorphic loci are selected from polymorphic loci known to be heterozygous in the aggregate.

[0247] As disclosed herein, in certain embodiments, the sample is a plasma sample derived from a pregnant woman.

[0248] In some cases, the method further includes performing the method on a control sample having a known mean allele disequilibrium ratio. The control may have a mean allele disequilibrium ratio for a specific allele state, which is an indicator of aneuploidy in 0.4–10% of chromosomal segments, to mimic the mean allele disequilibrium in the sample that is present at low concentrations, for example, as expected for circulating free DNA from a tumor.

[0249] In some embodiments, as disclosed herein, PlasmArt controls are used as controls. Thus, in certain embodiments, this is a sample prepared by a method comprising fragmenting a nucleic acid sample known to exhibit chromosomal aneuploidy into fragments that mimic the size of DNA fragments circulating in the plasma of an individual. In certain embodiments, controls that do not exhibit aneuploidy with respect to chromosomal segments are used.

[0250] In exemplary embodiments, data from one or more controls may be analyzed together with the test sample using the method. Controls may include, for example, different samples from individuals not suspected of containing chromosomal aneuploidy, or samples suspected of containing CNV or chromosomal aneuploidy. For example, if the test sample is a tumor sample suspected of containing circulating free tumor DNA, the method can be performed on a tumor-derived control sample from that subject, along with its plasma sample. As disclosed herein, the control sample may be prepared by fragmenting a DNA sample known to exhibit chromosomal aneuploidy. Such fragmentation can yield a DNA sample that mimics the DNA composition of apoptotic cells, particularly if the sample is from an individual with cancer. Data from the control sample will enhance the reliability of the detection of chromosomal aneuploidy.

[0251] In certain embodiments of the method for determining ploidy, the sample is a plasma sample from an individual suspected of having cancer. In these embodiments, the method further includes determining whether copy number variations are present in the tumor cells of the individual, based on the selection described above. For these embodiments, the sample may be a plasma sample from the individual. In these embodiments, the method may further include determining whether cancer is present in the individual, based on the selection described above.

[0252] These embodiments for determining the ploidy of a chromosomal segment may further include detecting a single nucleotide variant at a single nucleotide variant location in a set of single nucleotide variant locations, and detecting either chromosomal aneuploidy or a single nucleotide variant, or both, indicates the presence of circulating tumor nucleic acids in the sample.

[0253] These embodiments may further include receiving haplotype information of a chromosomal segment for a tumor in an individual, and using this haplotype information to create a set of models for different ploidy states and allelic imbalance fractions for a set of polymorphic loci.

[0254] As disclosed herein, certain embodiments of methods for determining ploidy may further include removing outliers from the initial or modified allele frequency data before comparing the initial or modified allele frequencies to a set of models. For example, in certain embodiments, locus allele frequencies that are at least two or three standard deviations above or below the mean for other loci on a chromosomal segment are removed from the data before being used for modeling.

[0255] As referred to herein, it will be understood that for many of the embodiments provided herein, including those for determining the ploidy of chromosome segments, incomplete or complete phasing data is preferably used. Several features that provide improvements over conventional methods for detecting ploidy are provided herein, and it will also be understood that many different combinations of these features may be used.

[0256] In certain embodiments, computer systems and computer-readable media for performing any method of the present invention are provided herein. These include systems and computer-readable media for performing methods for determining ploidy. Accordingly, as a non-limiting example of embodiments of the system, in another embodiment to demonstrate that any of the methods provided herein can be performed using the disclosure herein and the system and computer-readable media, a system for detecting chromosomal ploidy in a sample of an individual is provided herein, comprising: an input processor configured to receive allele frequency data, including the amount of each allele present in the sample at each locus in a set of polymorphic loci on a chromosomal segment; a modeler configured to create phasing allele information for the set of polymorphic loci by estimating the phase of the allele frequency data, create individual probabilities of allele frequencies for the polymorphic loci for different ploidy states using the allele frequency data, and create binding probabilities for the set of polymorphic loci using the individual probabilities and the phasing allele information; and a hypothesis manager that determines the ploidy of a chromosomal segment by selecting a best-fitting model, which is an index of chromosomal ploidy, based on the binding probabilities.

[0257] In a particular embodiment of this system, the allele frequency data is data generated by a nucleic acid sequencing system. In a particular embodiment, the system further includes an error correction unit configured to correct errors in the allele frequency data, and the corrected allele frequency data is used by a modeler to create individual probabilities. In a particular embodiment, the error correction unit corrects allele amplification efficiency bias. In a particular embodiment, the modeler uses a set of models for different ploidy states and allele imbalance fractions for a set of polymorphic loci to create individual probabilities. In a particular exemplary embodiment, the modeler creates binding probabilities by considering binding between polymorphic loci on chromosomal segments.

[0258] In one exemplary embodiment, a system for detecting chromosomal ploidy in a sample of an individual is provided herein, comprising: an input processor configured to receive nucleic acid sequence data for alleles at a set of polymorphic loci on a chromosomal segment of the individual, and to use the nucleic acid sequence data to detect allele frequencies at the set of loci; an error correction unit configured to correct errors in the detected allele frequencies and to create corrected allele frequencies for the set of polymorphic loci; a modeler configured to create phasing allele information for the set of polymorphic loci by estimating the phase of the nucleic acid sequence data, to create individual probabilities of allele frequencies for different polymorphic states at the polymorphic loci by comparing the phasing allele information with a set of models of different ploidy states and allele imbalance fractions of the set of polymorphic loci, and to create a binding probability for the set of polymorphic loci by combining the individual probabilities, taking into account the relative distances between polymorphic loci on the chromosomal segment; and a hypothesis manager configured to select a best-fitting model, which is an index of chromosomal aneuploidy, based on the binding probability.

[0259] In certain exemplary system embodiments provided herein, the set of polymorphic loci includes 1,000 to 50,000 polymorphic loci. In certain exemplary system embodiments provided herein, the set of polymorphic loci includes 100 known heterozygous hotspot loci. In certain exemplary system embodiments provided herein, the set of polymorphic loci includes 100 loci located within or inside a 0.5kb recombinant hotspot.

[0260] In a particular exemplary system embodiment provided herein, the best-fitting model analyzes the following ploidy states of the first homolog and the second homolog of the chromosomal segment: (1) all cells have no deletion or amplification of the first or second homolog of the chromosomal segment; (2) some or all cells have a deletion of the first homolog or amplification of the second homolog of the chromosomal segment; (3) some or all cells have a deletion of the second homolog or amplification of the first homolog of the chromosomal segment.

[0261] In the embodiment of the particular exemplary system provided herein, the errors corrected include allele amplification efficiency bias, contamination, and / or sequencing errors. In the embodiment of the particular exemplary system provided herein, contamination includes ambient contamination and genotype contamination. In the embodiment of the particular exemplary system provided herein, ambient contamination and genotype contamination are determined for homozygous alleles.

[0262] In a particular exemplary system embodiment provided herein, the hypothesis manager is configured to analyze the magnitude of the difference between the phasing allele information created for the model and the estimated allele frequencies. In a particular exemplary system embodiment provided herein, the modeler creates individual probabilities of allele frequencies based on a beta-binomial model of predicted and observed allele frequencies for a set of polymorphic loci. In a particular exemplary system embodiment provided herein, the modeler creates individual probabilities using a Bayesian classifier.

[0263] In certain exemplary system embodiments provided herein, nucleic acid sequence data are generated by high-throughput DNA sequencing of multiple copies of a series of amplicons created using a multiple amplification reaction, where each amplicon of the series spreads to at least one polymorphic locus of a set of polymorphic loci, and each of the polymorphic loci of this set is amplified. In certain exemplary system embodiments provided herein, the multiple amplification reaction is carried out under restricted primer conditions for at least half of the reaction. In certain exemplary system embodiments provided herein, the sample has an average allele imbalance of 0.4% to 5%.

[0264] In a particular exemplary embodiment of the system provided herein, the sample is a plasma sample from an individual suspected of having cancer, and the hypothesis manager is configured to further determine, based on a best-fit model, whether copy number variations are present in the individual's tumor cells.

[0265] In certain exemplary systems provided herein, the sample is a plasma sample from an individual, and the hypothesis manager is further configured to determine whether cancer is present in the individual based on a best-fitting model. In these embodiments, the hypothesis manager may also be configured to detect single nucleotide variants at single nucleotide variant locations in a set of single nucleotide variant locations, and the detection of either chromosomal aneuploidy or single nucleotide variants, or both, indicates the presence of circulating tumor nucleic acids in the sample.

[0266] In a particular exemplary system embodiment provided herein, the input processor is configured to further receive haplotype information of a chromosomal segment for a tumor in an individual, and the modeler is configured to use this haplotype information to create a set of models of different ploidy states and allelic imbalance fractions for a set of polymorphic loci.

[0267] In the embodiment of the particular exemplary system provided herein, the modeler creates a model over allele imbalance fractions ranging from 0% to 25%.

[0268] It will be understood that any method provided herein may be performed by computer-readable code stored on a non-temporary computer-readable medium. Accordingly, in one embodiment, a non-temporary computer-readable medium for detecting chromosomal ploidy in an individual sample is provided herein, comprising computer-readable code and, when performed by a processing device, causing the processing device to receive allele frequency data, including the amount of each allele present in the sample at each locus in a set of polymorphic loci on a chromosomal segment; to create phasing allele information for the set of polymorphic loci by estimating the phase of the allele frequency data; to create individual probabilities of allele frequencies for the polymorphic loci for different ploidy states using the allele frequency data; to create binding probabilities for the set of polymorphic loci using the individual probabilities and the phasing allele information; and to determine the ploidy of a chromosomal segment by selecting a best-fitting model, which is an index of chromosomal ploidy, based on the binding probabilities.

[0269] In a particular embodiment of the computer-readable medium, allele frequency data is created from nucleic acid sequence data. The particular embodiment of the computer-readable medium further includes correcting errors in the allele frequency data and using the corrected allele frequency data in the process of creating individual probabilities. In a particular embodiment of the computer-readable medium, the error corrected is an allele amplification efficiency bias. In a particular embodiment of the computer-readable medium, individual probabilities are created using a set of models for different ploidy states and allele imbalance fractions for a set of polymorphic loci. In a particular embodiment of the computer-readable medium, binding probabilities are created by considering bindings between polymorphic loci on chromosomal segments.

[0270] In a particular embodiment, a non-temporary computer-readable medium for detecting chromosomal ploidy in a sample of an individual is provided herein, comprising computer-readable code, which, when executed by a processing device, causes the processing device to receive nucleic acid sequence data for alleles at a set of polymorphic loci on a chromosomal segment of the individual; uses the nucleic acid sequence data to detect allele frequencies at the set of loci; corrects for allele amplification efficiency bias in the detected allele frequencies to create corrected allele frequencies for the set of polymorphic loci; estimates the phase of the nucleic acid sequence data to create phasing allele information for the set of polymorphic loci; compares the corrected allele frequencies with a set of models of different ploidy states and allele imbalance fractions for the set of polymorphic loci to create individual probabilities of allele frequencies for different ploidy states for the polymorphic loci; combines individual probabilities considering the binding between polymorphic loci on the chromosomal segment to create binding probabilities for the set of polymorphic loci; and selects a best-fitting model, which is an index of chromosomal aneuploidy, based on the binding probabilities.

[0271] In a particular exemplary embodiment of a computer-readable medium, the above selection is made by analyzing the magnitude of the difference between the phasing allele information and the estimated allele frequencies created for the model.

[0272] In a particular exemplary embodiment of a computer-readable medium, the individual probabilities of allele frequencies are constructed based on a beta-binomial model of predicted and observed allele frequencies in a set of polymorphic loci.

[0273] It will be understood that any embodiment of the method provided herein may be carried out by executing code stored in a non-temporary computer-readable medium.

[0274] Exemplary Embodiments for Cancer Detection In certain embodiments, the present invention provides a method for detecting cancer. It will be understood that the sample may be a tumor sample or a liquid sample, such as plasma, from an individual suspected of having cancer. The method is particularly effective for detecting gene mutations, such as single nucleotide changes (SNVs), or copy number changes, such as CNVs in a sample containing low levels of these gene changes as part of the total DNA in the sample. Thus, the sensitivity for detecting DNA or RNA from cancer in a sample is exceptional. To achieve this exceptional sensitivity, the method may combine any or all of the improvements provided herein for detecting CNVs and SNVs.

[0275] Accordingly, in certain embodiments provided herein, a method for determining whether circulating tumor nucleic acids are present in a sample of an individual, and a non-temporary computer-readable medium comprising computer-readable code, which, when executed by a processing device, causes the processing device to carry out the method. The method comprises the steps of analyzing a sample to determine ploidy at a set of polymorphic loci on a chromosomal segment in the individual, and determining, based on the ploidy determination, the level of mean allele imbalance present at the polymorphic loci, where mean allele imbalances equal to or greater than 0.4%, 0.45%, 0.5%, 0.6%, 0.7%, 0.75%, 0.8%, 0.9%, or 1% are indicators of the presence of circulating tumor nucleic acids (e.g., ctDNA) in the sample.

[0276] In certain exemplary embodiments, mean allele imbalances greater than 0.4, 0.45, or 0.5% are indicators of the presence of ctDNA. In certain embodiments, a method for determining the presence of circulating tumor nucleic acids further comprises detecting a single nucleotide variant at a single nucleotide dispersion site in a set of single nucleotide dispersion sites, detecting an allele imbalance equal to or greater than 0.5%, or detecting a single nucleotide variant, or both, which are indicators of the presence of circulating tumor nucleic acids in a sample. It will be understood that the level of allele imbalance (typically expressed as mean allele imbalance) can be determined using any of the methods provided herein for detecting chromosomal ploidy or CNVs. It will be understood that a single nucleotide for this embodiment of the invention can be detected using any of the methods provided herein for detecting SNVs.

[0277] In certain embodiments, a method for determining the presence of circulating tumor nucleic acids further comprises performing the method on a control sample having a known mean allele imbalance ratio. The control may be, for example, a sample from an individual's tumor. In some embodiments, the control has a predicted mean allele imbalance relative to the sample under analysis. For example, AAI is 0.5% to 5%, or the mean allele imbalance ratio is 0.5%.

[0278] In certain embodiments, the analytical step of a method for determining the presence of circulating tumor nucleic acids includes analyzing a set of chromosomal segments known to exhibit aneuploidy in cancer. In certain embodiments, the analytical step of a method for determining the presence of circulating tumor nucleic acids includes analyzing 1,000 to 50,000 or 100 to 1,000 polymorphic loci for ploidy. In certain embodiments, the analytical step of a method for determining the presence of circulating tumor nucleic acids includes analyzing 100 to 1,000 single nucleotide variant sites. For example, in these embodiments, the analytical step may include performing multiplex PCR to amplify amplicons across 1,000 to 50,000 polymorphic loci and 100 to 1,000 single nucleotide variant sites. This multiplex reaction can be set up as a single reaction or as a pool of multiplex reactions of different subsets. The multiplex reaction methods provided herein (e.g., large-scale multiplex PCR disclosed herein) provide exemplary processes for performing amplification reactions to help achieve improved multiplexing and thus sensitivity levels.

[0279] In certain embodiments, the multiplex PCR reaction is carried out under restricted primer conditions for at least 10%, 20%, 25%, 50%, 75%, 90%, 95%, 98%, 99%, or 100% of the reaction. Improved conditions for carrying out large-scale multiplex reactions provided herein can be used.

[0280] In certain embodiments, the above-described method for determining whether circulating tumor nucleic acids are present in a sample of an individual, and all embodiments thereof, can be carried out using a system. This disclosure provides teachings relating to specific functional and structural features for performing the methods described above. In non-limiting examples, the system includes:

[0281] An input processor configured to analyze data from a sample and determine ploidy in a set of polymorphic loci on a chromosomal segment in an individual,

[0282] A modeler that determines the level of allele imbalance present at polymorphic loci based on the determination of ploidy, where an allele imbalance equal to or greater than 0.5% is an indicator of the presence of a cycle.

[0283] Exemplary Embodiment for Detecting a Single nucleotide Variant In certain embodiments, methods for detecting single nucleotide variants in a sample are provided herein. The improved methods provided herein can achieve detection limits of 0.015, 0.017, 0.02, 0.05, 0.1, 0.2, 0.3, 0.4, or 0.5% of SNVs in a sample. All embodiments for detecting SNVs can be performed using a system. This disclosure provides teachings relating to specific functional and structural features for performing the methods described above. Furthermore, embodiments are provided herein that include a non-transient computer-readable medium, which, when executed by a processing device, causes the processing device to perform the methods for detecting SNVs provided herein.

[0284] Accordingly, in one embodiment, a method is provided herein for determining whether a single nucleotide variant is present in a set of genomic locations in a sample from an individual, comprising: creating estimates of the efficiency and error rate per cycle for each genomic location using a training dataset; receiving observed nucleotide identity information for each genomic location in the sample; independently using the estimates of amplification efficiency and error rate per cycle for each genomic location, determining a set of probabilities of the proportion of single nucleotide variants obtained from one or more actual mutations at each genomic location by comparing the observed nucleotide identity information at each genomic location with a model of different variant proportions; and determining the proportion and reliability of the most likely actual variant from the set of probabilities for each genomic location.

[0285] In an exemplary embodiment of a method for determining the presence of a single nucleotide variant, estimates of efficiency and error rate per cycle are made for a set of amplicons spread across a genomic location. For example, this may include 2, 3, 4, 5, 10, 15, 20, 25, 50, 100, or more amplicons spread across a genomic location.

[0286] In an exemplary embodiment of a method for determining the presence of a single nucleotide variant, the observed nucleotide identity information includes the total number of reads observed for each genomic location and the number of variant allele reads observed for each genomic location.

[0287] In an exemplary embodiment of a method for determining the presence of a single nucleotide variant, the sample is a plasma sample, and the single nucleotide variant is present in the circulating tumor DNA of the sample.

[0288] In another embodiment, a method for estimating the proportion of single nucleotide variants present in a sample from an individual is provided herein. The method includes: creating estimates of efficiency and error per cycle for one or more amplicons spread across a set of genomic locations using a training dataset; receiving observed nucleotide identity information for each genomic location in the sample; using the amplification efficiency and error per cycle of the amplicons, creating mean and variance estimates for the total number of molecules, background error molecules, and actual mutant molecules in a search space containing an initial proportion of actual mutant molecules; and determining the proportion of single nucleotide variants present in the sample from actual mutations by using the mean and variance estimates and fitting the distribution to the observed nucleotide identity information in the sample to determine the proportion of the most likely actual single nucleotide variants.

[0289] In an exemplary example of this method for estimating the proportion of single nucleotide variants present in a sample, the sample is a plasma sample, and the single nucleotide variants are present in the circulating tumor DNA of the sample.

[0290] The training dataset for this embodiment of the present invention typically includes a sample from one healthy individual or preferably a group of healthy individuals. In certain exemplary embodiments, the training dataset is analyzed on the same day or in the same run for one or more samples under test. For example, a training dataset may be created using samples from a group of 2, 3, 4, 5, 10, 15, 20, 25, 30, 36, 48, 96, 100, 192, 200, 250, 500, 1000, or more healthy individuals. If data is available for an even larger number of healthy individuals (e.g., 96 or more), the reliability of the amplification efficiency estimate increases, even if runs are performed before the method is run on the samples under test. Since the error rate of PCR is per amplicon, nucleic acid sequence information created for the entire amplified region around the SNV, not just for the SNV base position, may be used. For example, if samples from 50 individuals are used and the 20-base pair amplicons around the SNV are sequenced, the error frequency can be determined using error frequency data from 1000 base reads.

[0291] Typically, amplification efficiency is estimated by estimating the mean and standard deviation of the amplification efficiency for the segment being amplified, and then fitting this to a distribution model (e.g., a binomial distribution or a beta-binomial distribution). For PCR with a known number of cycles, the error rate is determined, and then the error rate per cycle is estimated.

[0292] In certain exemplary embodiments, estimating the initial molecule in a test dataset further includes updating the efficiency estimate for the test dataset using the initial molecule number estimated in step (b) if the number of read observations differs significantly from the estimated number of reads. This estimate can then be updated for new efficiencies and / or initial molecules.

[0293] The search space used to estimate the total number of molecules, background error molecules, and actual mutant molecules may include a search space with a lower limit of 0.1%, 0.2%, 0.25%, 0.5%, 1%, 2.5%, 5%, 10%, 15%, 20%, or 25% of copies of the base at the SNV position, which is the SNV base, and an upper limit of 1%, 2%, 2.5%, 5%, 10%, 12.5%, 15%, 20%, 25%, 50%, 75%, 90%, or 95%. A lower range, such as 0.1%, 0.2%, 0.25%, 0.5%, or 1% at the lower limit and 1%, 2%, 2.5%, 5%, 10%, 12.5%, or 15% at the upper limit, may be used in exemplary cases for plasma samples, where the method detects circulating tumor DNA. A higher range is used for tumor samples.

[0294] The distribution is fitted to the total number of error molecules in the total molecules (background errors and actual mutations), and the likelihood or probability is calculated for each possible actual mutation in the search space. This distribution may be a binomial distribution or a beta-binomial distribution.

[0295] The most likely actual mutation is determined by determining the proportion of the most likely actual mutation and calculating the reliability using data from the distribution fitting. Exemplary examples, without intending to limit the clinical interpretations provided herein, show that a higher mean mutation rate results in a lower confidence rate required to make a positive determination for SNVs. For example, if the mean mutation rate for SNVs in a sample using the most likely hypothesis is 5% and the confidence rate is 99%, a positive SNV call would be made. On the other hand, for this example, if the mean mutation rate for SNVs in a sample using the most likely hypothesis is 1% and the confidence rate is 50%, a positive SNV call would not be made in certain circumstances. It will be understood that the clinical interpretation of the data may be a function of sensitivity, specificity, prevalence, and the availability of alternative products.

[0296] In one exemplary embodiment, the sample is a circulating DNA sample, for example, a circulating tumor DNA sample.

[0297] In another embodiment, a method for detecting one or more single nucleotide variants in a test sample from an organism is provided herein. The method according to this embodiment includes the following steps.

[0298] The process involves: determining the median variant allele frequency for each single nucleotide variant position in a set of single nucleotide variant positions based on the results generated by a sequencing run, using multiple control samples from multiple normal individuals to identify selected single nucleotide variant positions that have a median variant allele frequency below a threshold in normal samples, removing outlier samples for each single nucleotide variant position, and then determining the background error for each single nucleotide variant position; determining the weighted mean and variance of the observed read depth for the selected single nucleotide variant positions in the test sample based on the data generated by a sequencing run for the test sample; and detecting one or more single nucleotide variants by using a computer to identify one or more single nucleotide variant positions with a statistically significant weighted mean of read depth by comparing them with the background error for those positions.

[0299] In a particular embodiment of this method for detecting one or more SNVs, the sample is a plasma sample, the control sample is a plasma sample, and the one or more detected single nucleotide variants are present in the circulating tumor DNA of the sample. In a particular embodiment of this method for detecting one or more SNVs, the multiple control samples include at least 25 samples. In a particular exemplary embodiment, the multiple control samples include at least 5, 10, 15, 20, 25, 50, 75, 100, 200 or 250 samples at the lower limit, and 10, 15, 20, 25, 50, 75, 100, 200, 250, 500 and 1000 samples at the upper limit.

[0300] In a particular embodiment of this method for detecting one or more SNVs, outliers are removed from the data produced by a high-throughput sequencing run, a weighted mean of the observed read depths is calculated, and the observed variance is determined. In a particular embodiment of this method for detecting one or more SNVs, the read depth for each single-nucleotide variant position for the test sample is at least 100 reads.

[0301] In certain embodiments of this method for detecting one or more SNVs, the sequencing run includes a multiple amplification reaction carried out under restricted primer reaction conditions. These embodiments are carried out in exemplary examples using an improved method for carrying out the multiple amplification reaction provided herein.

[0302] While not limited to theory, the method of this embodiment utilizes a background error model using a normal plasma sample, sequencing it in the same sequencing run as the sample under test, and taking run-specific artifacts into account. Noise locations with median normal variant allele frequencies above thresholds, e.g., 0.1%, 0.2%, 0.25%, 0.5%, 0.75%, and 1.0%, are removed.

[0303] To account for noise and contamination, outlier samples are iteratively removed from this model. For each base substitution at all genomic loci, the read-depth-weighted mean and standard deviation of the error are calculated. In certain exemplary embodiments, samples (e.g., plasma samples that do not contain tumors or cells) that have at least a threshold number of reads (e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 250, 500, or 1000 variant reads) and have single nucleotide variant locations with an a1 Z score greater than 2.5, 5, 7.5, or 10 against the background error model in certain embodiments are counted as candidate mutations.

[0304] In a particular embodiment, read depths of 100, 250, 500, 1,000, 2,000, 2,500, 5,000, 10,000, 20,000, 25,0000, 50,000, or 100,000 at the lower limit of the range and 2,000, 2,500, 5,000, 7,500, 10,000, 25,000, 50,000, 100,000, 250,000, or 500,000 reads at the upper limit are achieved in a sequencing run for each single nucleotide variant position in a set of single nucleotide variant positions. Typically, the sequencing run is a high-throughput sequencing run. The mean or median value created for the samples under test is weighted by read depth in exemplary embodiments. Thus, the likelihood that a variant allele determination is real in a sample where one variant allele is detected in 1,000 reads is weighted more heavily than in a sample where one variant allele is detected in 10,000 reads. Because the determination of variant alleles (i.e., mutations) is not performed with 100% confidence, identified single nucleotide variants may be considered candidate variants or candidate mutations.

[0305] Exemplary test statistics for the analysis of fading data Exemplary test statistics are described below for the analysis of phasing data from samples known or suspected to be mixed samples containing DNA or RNA from two or more genetically non-identical cells. f represents the fraction of the DNA or RNA of interest, e.g., the fraction of DNA or RNA containing the CNV of interest, or the fraction of DNA or RNA from the cells of interest, e.g., cancer cells. In some embodiments of the cancer test, f represents the fraction of DNA or RNA from cancer cells in a mixture of cancer and normal cells, or f represents the fraction of cancer cells in a mixture of cancer and normal cells. Note that this refers to the fraction of DNA from the cells of interest, assuming that two copies of DNA are provided by each cell of interest. This is different from the fraction of DNA from the cells of interest in deleted or duplicated segments.

[0306] The possible allele values ​​for each SNP are indicated by A and B. AA, AB, BA, and BB are used to represent all possible ordered allele pairs. In some embodiments, SNPs containing ordered alleles AB or BA are analyzed. i This indicates the number of sequence reads for the i-th SNP, and A i and B i These represent the number of reads for the i-th SNP, which represents alleles A and B, respectively. The following assumptions are made: N i =A i +B i

[0307] Allele ratio R i It is defined as follows:

number

[0308] T represents the number of target SNPs.

[0309] Without loss of generality, some embodiments focus on a single chromosomal segment. For further clarity, the phrase “first homologous chromosomal segment compared to a second homologous chromosomal segment” means the first homolog of the chromosomal segment and the second homolog of the chromosomal segment. In some such embodiments, all target SNPs are contained within the target segment chromosome. In other embodiments, multiple chromosomal segments are analyzed for possible copy number variations.

[0310] MAP estimation This method leverages knowledge of phasing mediated by ordered allele pairs to detect deletions or duplications of target segments. For each SNPi, it is defined as follows:

number

[0311] Next, we define it as follows:

number

[0312] X under various copy number hypotheses (e.g., disomy hypothesis, deletion of the first or second homolog, or duplication of the first or second homolog) i The distribution of S is described below.

[0313] Disomy hypothesis Under the assumption that the target segments are not missing or overlapping,

number

number

[0314] Assuming a constant read depth N, we obtain a binomial distribution S with the following parameters.

number

[0315] Negative hypothesis Under the hypothesis that the first homolog is deleted (i.e., AB SNP becomes B and BA SNP becomes A), R i It has a binomial distribution, and for AB SNPs, the parameter

number

number

number

[0316] Assuming a constant read depth N, we obtain a binomial distribution S with the following parameters.

number

[0317] Under the hypothesis that the second homolog is deleted (i.e., AB SNP becomes A and BA SNP becomes B), i It has a binomial distribution, and for AB SNPs, the parameter

number

number

number

[0318] Assuming a constant read depth N, we obtain a binomial distribution S with the following parameters.

number

[0319] Overlap Hypothesis Under the hypothesis that the first homologs overlap (i.e., AB SNPs become AAB and BA SNPs become BBA), R i It has a binomial distribution, and for AB SNPs, the parameter

number

number

number

[0320] Assuming a constant read depth N, we obtain a binomial distribution S with the following parameters.

number

[0321] Under the hypothesis that the second homolog overlaps (i.e., AB SNPs become ABB and BA SNPs become BAA), R i It has a binomial distribution, and for AB SNPs, the parameter

number

number

number

[0322] Assuming a constant read depth N, we obtain a binomial distribution S with the following parameters.

number

[0323] classification As shown in the previous chapter, X i This is a binary random variable having the following characteristics:

number

[0324] This allows us to calculate the probability of the test statistic S under each hypothesis. We can calculate the probability of each hypothesis considering the measured data. In some embodiments, the hypothesis with the highest probability is selected. If desired, the distribution for S is given for each N i This can be simplified by estimating it with a constant reachable depth N, or by rounding down the read depth to a constant value N. This simplification gives the following:

number

[0325] The value of f can be estimated by selecting the most likely value of f that takes the measured data into account, for example, by selecting the value of f that produces the best data fitting using an algorithm (e.g., a search algorithm), for example, maximum likelihood estimation, empirical maximum estimation, or Bayesian estimation. In some embodiments, multiple chromosomal segments are analyzed, and the value of f is estimated based on the data for each segment. If all target cells have these overlaps or deletions, the estimates of f based on the data for these different segments will be similar. In some embodiments, f is measured experimentally, for example, by determining the fraction of DNA or RNA from cancer cells based on the difference in methylation (hypomethylation or hypermethylation) of cancer and non-cancerous DNA or RNA.

[0326] Single Hypothesis Rejection The distribution of S for the disomy hypothesis is independent of f. Therefore, the probability of the measured data can be calculated for the disomy hypothesis without calculating f. A single-hypothesis rejection test can be used for the null hypothesis of disomy. In some embodiments, the probability of S for the disomy hypothesis is calculated, and the disomy hypothesis is rejected if its probability falls below a given threshold (e.g., less than 1 in 1000). This indicates the presence of a chromosomal segment duplication or deletion. If desired, the false positive rate can be altered by adjusting the threshold.

[0327] Exemplary methods for analyzing fading data Exemplary methods are described below for the analysis of data from a sample known or suspected to be a mixed sample containing DNA or RNA from two or more genetically non-identical cells. In some embodiments, phasing data is used. In some embodiments, the method involves determining, for each calculated allele ratio, whether the calculated allele ratio for a particular locus is above or below the predicted allele ratio, and the magnitude of the difference. In some embodiments, a likelihood distribution is determined for the allele ratios at a locus for a particular hypothesis, and the closer the calculated allele ratio is to the center of the likelihood distribution, the more likely the hypothesis is correct. In some embodiments, the method involves determining the correct likelihood for a hypothesis for each locus. In some embodiments, the method involves determining the correct likelihood for a hypothesis for each locus and combining this with the probability of that hypothesis for each locus, and selecting the hypothesis with the highest binding probability. In some embodiments, the method involves determining the correct likelihood for a hypothesis for each locus and for each possible ratio of DNA or RNA from one or more target cells to the total DNA or RNA in the sample. In some embodiments, the binding probability for each hypothesis is determined by combining the probabilities of the hypotheses for each locus and each possible ratio, and the hypothesis with the highest binding probability is selected.

[0328] In one embodiment, the following hypothesis is considered: H 11 (All cells are normal), H 10 (Presence of cells containing only homolog 1, and therefore deletion of homolog 2), H 01 (Presence of cells containing only homolog 2, and therefore deletion of homolog 1), H 21 (Presence of cells with homolog 1 duplication), H 12 (Presence of cells with homolog 2 duplication). For the fraction f of target cells such as cancer cells or mosaic cells (or the fraction of DNA or RNA from target cells), the predicted value of the allele ratio for a heterozygous (AB or BA) SNP can be found as follows. Formula (1):

Number

[0329] Correction of bias, contamination and sequencing errors: Observed D at SNP s is the number n of the originally mapped reads in which each allele is present A 0 and n B 0 Consists of. Then, using the predicted values of the bias in the amplification of alleles A and B, the corrected reads n A and n B Can be found.

[0330] c a Indicates ambient contamination (e.g., contamination from DNA in air or the environment), and r(c a ) indicates the allele ratio for the ambient contaminant (initially considered 0.5). Further, c g Indicates the genotype contamination rate (e.g., contamination from another sample), and r(c g ) is the allele ratio for that contaminant. s e (A,B) and s e (B,A) indicate sequencing errors in which one allele is called as a different allele (e.g., by erroneously detecting the A allele when the B allele is present).

[0331] By correcting for ambient contamination, genotype contamination and sequencing errors, for the predicted value r of a given allele ratio, the observed value q of the allele ratio (r,c a ,r(c a ),cg , r(c g ), s e (A, B), s e (B, A)) can be found.

[0332] Since the genotype of the contaminant is unknown, the population frequency is used to find P(r(c g )). More specifically, p is the population frequency for one of the alleles (which may be called the reference allele). Then, P(r(c g ) = 0) = (1 - p) 2 , P(r(c g ) = 0) = 2p(1 - p) and P(r(cg) = 0) = p 2 . Using the conditional expectation over r(c g ), E[q(r, c a , r(c a ), c g , r(c g ), s e (A, B), s e (B, A))] can be determined. Note that the ambient contamination and genotype contamination are determined using homozygous SNPs and are thus not affected by the presence or absence of deletions or duplications. Further, if desired, the reference chromosome can be used to measure the ambient contamination and genotype contamination.

[0333] Likelihood at each SNP: The following equation gives the probability of observing n A and n B considering the allele ratio r. Equation (2): [Number]

[0334] D s represents the data for the SNP. For each hypothesis h ε {H 11 , H 01 , H 10 , H 21 , H 12Regarding}, in formula (1), set r=r(AB,h) or r=r(BA,h), and r(c g We found the conditional expected value over ) and the observed allele ratio E[q(r,c a ,r(c a ), c g ,r(c g ))] can be determined. Next, in equation (2), r = E[q(r,c a ,r(c a ), c g ,r(c g ),s e (A,B),s e (B,A))], P(D s |h,f) can be determined.

[0335] Search algorithm: In some embodiments, SNPs with allele ratios that appear to be outliers are ignored (for example, by ignoring or excluding SNPs with allele ratios that are at least 2 or 3 standard deviations above or below the mean). The advantage identified for this method is that it allows for greater variability of allele ratios in the presence of a higher proportion of mosaicism, thus ensuring that SNPs are not trimmed due to mosaicism.

[0336] F={f1,···,f N} represents the search space for the proportion of mosaicism (e.g., tumor fraction). P(Ds|h,f) can be determined for each SNP and fε F, and the likelihood can be matched for all SNPs.

[0337] This algorithm is performed for each hypothesis over each f. Using a search method, if the reliability of a missing or overlapping hypothesis is higher than the reliability of a hypothesis that has no missing or overlapping hypotheses, then the range of f is F. * When F is present, we conclude that a mosaic exists. In some embodiments, F * P(D) s The maximum likelihood estimate of |h,f) is determined. If desired, fε F *Conditional expectations across the range may be determined. If desired, the reliability of each hypothesis can be determined.

[0338] In some embodiments, a beta-binomial distribution is used instead of a binomial distribution. In some embodiments, a reference chromosome or chromosomal segment is used to determine sample-specific parameters of the beta-binomial formula.

[0339] Theoretical performance using simulations: If desired, the theoretical performance of the algorithm can be evaluated by randomly assigning the number of reference reads to SNPs at a given read depth (DOR). Typically, a binomial probability parameter of p=0.5 is used, and p is adjusted accordingly for deletions or duplicates. The exemplary input parameters for each simulation are as follows: (1) the number of SNPs S, (2) a constant DOR per SNP D, (3) p, and (4) the number of experiments.

[0340] First simulation experiment: This experiment focused on Sε{500,1000}, Dε{500,1000} and pε{0%,1%,2%,3%,4%,5%}. For each setting, 1,000 simulation experiments were performed (thus, 24,000 experiments with and 24,000 without phases). The number of reads from a binomial distribution was simulated (other distributions may be used if desired). False positive rates (for p=0%) and false negative rates (for p>0%) were determined with or without phase information. Phase information is particularly useful for S=1000 and D=1000. However, for S=500 and D=500, this algorithm exhibits the highest false positive rate, regardless of whether phase-out from the tested condition is performed.

[0341] Phase information is particularly useful for low mosaicism rates (≤3%). Without phase information, the reliability of deletions is low. 10 and H 01Because this is determined by assigning equal chances to each, a high level of false negatives is observed for p=1%, and a small deviation favoring one hypothesis is not sufficient to compensate for the low likelihood from other hypotheses. This is also true for overlaps. Furthermore, this algorithm appears to be more sensitive to read depth than to the number of SNPs. For results using phase information, we assume that complete phase information is available for a large number of consecutive heterozygous SNPs. If desired, haplotype information can be obtained by probabilistically matching haplotypes for smaller segments.

[0342] Second simulation experiment: This experiment focused on randomized experiments with Sε{100,200,300,400,500}, Dε{1000,2000,3000,4000,5000} and pε{0%,1%,1.5%,2%,2.5%,3%} and 10000 for each setting. False positive rates (when p=0%) and false negative rates (when p>0%) were determined with or without phase information. False negative rates, using haplotype information, were less than 10% for D≧3000 and N≧200, while achieving the same performance for D=5000 and N≧400. The difference in false negative rates was particularly noticeable for small mosaic proportions. For example, at p=1%, without haplotype data, false negative rates of less than 20% were never achieved, while for N≧300 and D≧3000, they were close to 0%. When p=3%, a false negative rate of 0% is observed when haplotype data is used, whereas without haplotype data, N≧300 and D≧3000 are required to achieve the same performance.

[0343] An exemplary method for detecting deletions and duplicates without using fading data. In some embodiments, non-phasing gene data is used to determine whether there is a copy number overexpression of a first homologous chromosome segment compared to a second homologous chromosome segment in the genome of an individual (e.g., in the genome of one or more cells, or in cfDNA or cfRNA). In some embodiments, phasing gene data is used, but phasing is ignored. In some embodiments, the DNA or RNA sample is a mixed sample of cfDNA or cfRNA from an individual containing cfDNA or cfRNA from two or more genetically distinct cells. In some embodiments, the method utilizes the magnitude of the difference between the calculated allele ratio and the predicted allele ratio for each locus.

[0344] In some embodiments, the method involves obtaining genetic data for a set of polymorphic loci on a chromosome or chromosomal segment in a DNA or RNA sample from one or more cells from an individual by measuring the amount of each allele at each locus. In some embodiments, the allele ratio is calculated for loci that are heterozygous in at least one cell from which the sample originates. In some embodiments, the calculated allele ratio for a particular locus is obtained by dividing one measured amount of the allele by the total measured amount of all alleles for that locus. In some embodiments, the calculated allele ratio for a particular locus is obtained by dividing one measured amount of the allele (e.g., an allele on a first homologous chromosomal segment) by the measured amounts of one or more other alleles for that locus (e.g., alleles on a second homologous chromosomal segment). Calculated and predicted allele ratios may be calculated using any of the methods described herein or any standard method (e.g., any mathematical transformation of the calculated or predicted allele ratios described herein).

[0345] In some embodiments, the test statistic is calculated for each locus based on the magnitude of the difference between the calculated allele ratio and the predicted allele ratio. In some embodiments, the test statistic Δ is calculated using the following formula.

number

[0346] In the formula, δ i This is the magnitude of the difference between the calculated allele ratio and the predicted allele ratio for the i-th locus. In the formula, μ i is, δ i It is the average value, During the ceremony,

number

[0347] For example, if the predicted allele ratio is 0.5, δ i This can be defined as follows:

number

[0348] μ i and δ i The value of R i It can be calculated using the fact that is a binomial random variable. In some embodiments, it is assumed that the standard deviation is the same for all loci. In some embodiments, the mean or weighted mean of the standard deviation, or an estimate of the standard deviation,

number

[0349] In some embodiments, a set of one or more hypotheses indicating the copy number of chromosomes or chromosomal segments in one or more genomes of a cell is enumerated. In some embodiments, the most likely hypothesis is selected based on test statistics, thereby determining the copy number of chromosomes or chromosomal segments in one or more genomes of a cell. In some embodiments, a hypothesis is selected if the probability that the test statistics belong to the distribution of test statistics for a hypothesis exceeds an upper threshold. If the probability that the test statistics belong to the distribution of test statistics for a hypothesis falls below a lower threshold, one or more hypotheses are rejected, or if the probability that the test statistics belong to the distribution of test statistics for a hypothesis is between the lower and upper thresholds, or if the probability cannot be determined with sufficiently high confidence, the hypothesis is neither selected nor rejected. In some embodiments, the upper and / or lower thresholds are determined, for example, from an empirical distribution from training data (e.g., samples with known copy numbers, e.g., diploid samples or samples known to have specific deletions or duplications). Such empirical distributions can be used to select thresholds for single-hypothesis rejection tests. Furthermore, since the test statistic Δ is independent of S, both can be used independently if desired.

[0350] Exemplary Methods for Detecting Deletions or Duplications Using Allelic Distributions or Patterns This chapter includes a method for determining whether there is an over-occurrence of copy numbers in a first homologous chromosome segment compared to a second homologous chromosome segment. In some embodiments, the method involves (i) listing several hypotheses indicating the copy numbers of chromosomes or chromosomal segments present in the genome of one or more cells of an individual (e.g., cancer cells), or (ii) listing several hypotheses indicating the degree of over-occurrence of copy numbers in a first homologous chromosome segment compared to a second homologous chromosome segment in the genome of one or more cells of an individual. In some embodiments, the method involves obtaining genetic data from an individual at multiple polymorphic loci (e.g., SNP loci) on a chromosome or chromosomal segment. In some embodiments, a probability distribution of the individual's predicted genotype is constructed for each hypothesis. In some embodiments, a data fitting is calculated between the obtained genetic data of the individual and the probability distribution of the individual's predicted genotype. In some embodiments, one or more hypotheses are ranked according to the data fitting, and the highest-ranked hypothesis is selected. In some embodiments, a technique or algorithm, such as a search algorithm, is used in one or more of the steps of calculating data fitting, ranking hypotheses, or selecting the highest-ranked hypothesis. In some embodiments, the data fitting is fitting to a beta-binomial distribution or a binomial distribution. In some embodiments, the technique or algorithm is selected from the group consisting of maximum likelihood estimation, empirical maximum estimation, Bayesian estimation, dynamic estimation (e.g., dynamic Bayesian estimation), and expectation maximization estimation. In some embodiments, the method includes applying the techniques or algorithms described above to the obtained gene data and the predicted values ​​of the gene data.

[0351] In some embodiments, the method involves (i) listing several hypotheses indicating the copy number of a chromosome or chromosomal segment present in the genome of one or more cells of an individual (e.g., cancer cells), or (ii) listing several hypotheses indicating the degree of copy number overpopulation of a first homologous chromosomal segment compared to a second homologous chromosomal segment in the genome of one or more cells of an individual. In some embodiments, the method involves obtaining genetic data from an individual at multiple polymorphic loci (e.g., SNP loci) on a chromosome or chromosomal segment. In some embodiments, the genetic data includes the allele count for multiple polymorphic loci. In some embodiments, a bound distribution model is constructed for the predicted allele counts at multiple polymorphic loci on a chromosome or chromosomal segment for each hypothesis. In some embodiments, the relative probabilities of one or more of the hypotheses are determined using the bound distribution model and the allele counts measured for the sample, and the hypothesis with the highest probability is selected.

[0352] In some embodiments, the distribution or pattern of alleles (e.g., the pattern of calculated allele ratios) is used to determine the presence or absence of CNVs (e.g., deletions or duplications). If desired, the parental origin of the CNV can be determined based on this pattern.

[0353] Exemplary counting / quantification methods In some embodiments, one or more counting methods (also called quantitative methods) are used to detect one or more CNSs (e.g., deletions or duplications of chromosomal segments or entire chromosomes). In some embodiments, one or more counting methods are used to determine whether an over-occurrence of a copy number in a first homologous chromosomal segment is due to duplication of the first homologous chromosomal segment or deletion of a second homologous chromosomal segment. In some embodiments, one or more counting methods are used to determine whether there is an over-copy number in a duplicated chromosomal segment or chromosome (e.g., whether there are 1, 2, 3, 4, or more over-copys). In some embodiments, one or more counting methods are used to distinguish a sample with many over-copys and a low tumor fraction from a sample with few over-copys and a high tumor fraction. For example, one or more counting methods may be used to distinguish a sample with 4 over-copys and a tumor fraction of 10% from a sample with 2 over-copys and a tumor fraction of 20%. Exemplary methods are disclosed, for example, in U.S. Publications 2007 / 0184467, 2013 / 0172211 and 2012 / 0003637, U.S. Patents 8,467,976, 7,888,017, 8,008,018, 8,296,076 and 8,195,415, U.S. Application 62 / 008,235 filed on 5 June 2014 and U.S. Application 62 / 032,785 filed on 4 August 2014, each incorporated herein by reference as a whole.

[0354] In some embodiments, the counting method includes counting the number of reads based on DNA sequences that map to one or more given chromosomes or chromosomal segments. Some such methods involve creating a reference value (cutoff value) for the number of DNA sequence reads that map to a particular chromosome or chromosomal segment, where an excess number of reads is an indicator of a particular genetic abnormality.

[0355] In some embodiments, the total allele measurement for one or more loci (e.g., the total number of polymorphic or non-polymorphic loci) is compared to a reference value. In some embodiments, the reference value is (i) a threshold or (ii) a predictor for a particular copy number hypothesis. In some embodiments, the reference value (if no CNVs exist) is the total allele measurement for one or more loci on one or more chromosomes or chromosomal segments that are known or predicted to be non-deleting or non-duplication. In some embodiments, the reference value (if CNVs exist) is the total allele measurement for one or more loci on one or more chromosomes or chromosomal segments that are known or predicted to be deleting or non-duplication. In some embodiments, the reference value is the total allele measurement for one or more loci on one or more reference chromosomes or chromosomal segments. In some embodiments, the reference value is the mean or median of values ​​determined for two or more different chromosomes, chromosomal segments, or different samples. In some embodiments, random (e.g., massively parallel shotgun sequencing) or targeted sequencing is used to determine the amount of one or more polymorphic or non-polymorphic loci.

[0356] In some embodiments utilizing a reference quantity, the method includes (a) measuring the amount of genetic material in a chromosome or chromosomal segment of interest, (b) comparing the amount from step (a) with a reference quantity, and (c) identifying the presence or absence of deletions or duplications based on this comparison.

[0357] In some embodiments utilizing a reference chromosome or chromosomal segment, the method involves sequencing DNA or RNA from a sample to obtain a plurality of sequence tags aligned to target loci. In some embodiments, the sequence tags are long enough to be assigned to a specific target locus (e.g., 15–100 nucleotides long), and the target loci are derived from a plurality of different chromosomes or chromosomal segments, including at least one first chromosome or chromosomal segment suspected of having an abnormal distribution in the sample and at least one second chromosome or chromosomal segment presumed to be normally distributed in the sample. In some embodiments, the plurality of sequence tags are assigned to their corresponding target loci. In some embodiments, the number of sequence tags to assign to the target loci of the first chromosome or chromosomal segment and the number of sequence tags to assign to the target loci of the second chromosome or chromosomal segment are determined. In some embodiments, these numbers are compared to determine whether there is an abnormal distribution (e.g., deletion or duplication) in the first chromosome or chromosomal segment.

[0358] In some embodiments, the value of f (e.g., tumor fraction) is used in CNV determination to compare, for example, an observed difference in the amount of two chromosomes or chromosomal segments with a difference predicted for a particular type of CNV, taking the value of f into account (see, for example, U.S. Publications 2012 / 0190020, 2012 / 0190021, 2012 / 0190557, and 2012 / 0191358, respectively, which are incorporated herein by reference as a whole). For example, the difference in the amount of overlapping chromosomal segments in the tumor compared to a disomy reference chromosomal segment increases as the tumor fraction increases. In some embodiments, the method includes determining the likelihood of a CNV by comparing the relative frequency of the chromosome or chromosomal segment of interest with a value of f and a reference chromosome or chromosomal segment (e.g., a chromosome or chromosomal segment predicted or known to be disomy). For example, the difference in quantity between the first chromosome or chromosomal segment and the reference chromosome or chromosomal segment may be compared to what is predicted considering the f value for various possible CNVs (e.g., one or two extra copies of the chromosomal segment of interest).

[0359] The following hypothetical examples illustrate the use of counting / quantification methods to distinguish between duplication of a first homologous chromosome segment and deletion of a second homologous chromosome segment. Assuming a normal disomy genome of the host is the baseline, analysis of a mixture of normal and cancer cells will give the mean difference between the baseline and cancer DNA in the mixture. For example, imagine that 10% of the DNA in a sample comes from cells with deletions across the chromosomal region targeted by the assay. In some embodiments, the quantification method shows that the amount of reads corresponding to this region is predicted to be 95% of the amount predicted for a normal sample. This is because, since one of the two target chromosomal regions is missing in each tumor cell with a deletion in the target region, the total amount of DNA mapping to this region is 90% (for normal cells) + 1 / 2 × 10% (for tumor cells) = 95%. Alternatively, in some embodiments, the allelic method shows that the allele ratio at heterozygous loci is 19:20 on average. Next, consider a case where 10% of the DNA in a sample originates from cells with a 5x focal amplification of a chromosomal region targeted by the assay. In some embodiments, the quantitative method shows that the amount of reads corresponding to this region is predicted to be 125% of the amount predicted for a normal sample. This is because one of the two target chromosomal regions in each tumor cell with 5x focal amplification is over-copied 5 times across the target region, so the total amount of DNA mapping to this region is 90% (for normal cells) + (2 + 5) × 10% (for tumor cells) / 2 = 125%. Alternatively, in some embodiments, the allele method shows that the allele ratio at the heterozygous locus is 25:20 on average. Note that when using the allele method alone, a 5x focal amplification across a chromosomal region in a sample containing 10% cfDNA may appear equivalent to a deletion across the same region in a sample containing 40% cfDNA.In these two cases, the haplotype that is under-occurring in the case of deletion appears to be the haplotype that does not contain CNV in the case of focal duplication, and the haplotype that does not contain CNV in the case of deletion appears to be the haplotype that is over-occurring in the case of focal duplication. By combining the likelihoods created by this allelic approach with the likelihoods created by the quantitative approach, we can distinguish between these two probabilities.

[0360] Exemplary counting / quantification methods using a reference sample Exemplary quantification methods using one or more reference samples are described in U.S. Application No. 62 / 008,235 filed June 5, 2014 and U.S. Application No. 62 / 032,785 filed August 4, 2014, which are incorporated herein by reference in their entirety. In some embodiments, one or more reference samples (e.g., normal samples) that are most likely to not have a CNV on one or more chromosomes or on the chromosome of interest are identified by selecting the sample with the highest tumor DNA fraction, selecting the sample with the closest z-score to 0, selecting the sample whose data fits the hypothesis corresponding to the absence of a CNV with the highest reliability or likelihood, selecting a sample known to be normal, selecting a sample from an individual with the lowest likelihood of having cancer (e.g., young age, being male when screened for breast cancer, no family history), selecting the sample with the highest amount of DNA input, selecting the sample with the highest signal-to-noise ratio, selecting a sample based on other criteria considered to correlate with the likelihood of having cancer, or selecting a sample using a combination of some of the criteria. Once a reference set is selected, it can be assumed that these cases are disomy, and a bias per SNP, i.e., amplification specific to the experiment and other processing biases for each locus, can be estimated. This estimate of experiment-specific bias can then be used to correct for bias in the measurement of loci on the chromosome of interest, e.g., chromosome 21, and, where appropriate, for other chromosomal loci, for samples that are not part of a subset where disomy is not assumed for chromosome 21. Once the bias has been corrected for these samples with unknown ploidy, the data for these samples can be analyzed twice using the same or different methods to determine whether the individual has trisomy 21. For example, a quantitative method may be used for the remaining samples with unknown ploidy, and the z-score can be calculated using the measured genetic data corrected for chromosome 21. Alternatively, as part of a preliminary estimate of the ploidy status of chromosome 21, the tumor fraction of samples from individuals suspected of having cancer can be calculated.The predicted proportion of corrected reads for disomy (disomy hypothesis) and trisomy (trisomy hypothesis) can be calculated for the tumor fraction. Alternatively, if the tumor fraction is not measured beforehand, sets of disomy and trisomy hypotheses may be created for different tumor fractions. For each case, the predicted distribution of the proportion of corrected reads may be calculated considering a given predictive statistical variation, with respect to the selection and measurement of various DNA loci. Observed values ​​of the proportion of corrected reads may be compared to the predicted distribution of the proportion of corrected reads, and likelihood ratios can be calculated for the disomy and trisomy hypotheses for each sample with unknown ploidy. The ploidy state associated with the hypothesis with the highest calculated likelihood can be selected as the correct ploidy state.

[0361] In some embodiments, a subset of samples with a sufficiently low likelihood of having cancer may be selected to serve as a control set of samples. This subset may be a fixed number or a variable number based on selecting only samples below a threshold. Quantitative data from the subset of samples may be combined, averaged, or combined using a weighted average, the weighting of which is based on the likelihood of normal samples. The quantitative data may be used to determine locus bias for amplification in sequencing the samples in an immediate batch of control samples. The locus bias may also include data from other batches of samples. The locus bias may indicate relative overamplification or relative underamplification observed for that locus compared to other loci, and assuming the subset of samples does not contain CNVs, any observation of overamplification or underamplification may indicate that it is due to amplification and / or sequencing or other biases. The locus bias may take into account the GC content of the amplicons. Loci may be grouped into locus clusters for the purpose of calculating the locus bias. Once a bias per locus is calculated for each locus among multiple loci, sequencing data for one or more samples that are not in the sample subset, and optionally for one or more samples that are in the sample subset, may be adjusted by adjusting the quantitative measurements for each locus to remove the effect of bias at that locus. For example, if in a subset of patients SNP1 is observed to have a read depth twice the average size, the adjustment may involve replacing the corresponding number of reads from SNP1 with half the number of reads of that size. If the locus in question is an SNP, the adjustment may involve halving the number of reads corresponding to each allele at that locus. Once the sequencing data for each locus in one or more samples has been adjusted, it may be analyzed using a method for the purpose of detecting the presence of CNVs in one or more chromosomal regions.

[0362] In one example, sample A is a mixture of amplified DNA derived from a mixture of normal and cancerous cells analyzed using a quantitative method. The following are illustrative possible data: The region of the q arm on chromosome 22 is found to have only 90% of the predicted value of DNA mapping to that region; the focal region corresponding to the HER2 gene is found to have 150% of the predicted value of DNA mapping to that region; and the p arm on chromosome 5 is found to have 105% of the predicted value of DNA mapping to that region. A physician may infer that the sample has a deletion in the region on the q arm on chromosome 22 and a duplication of the HER2 gene. A physician may infer that approximately 20% of the DNA in the sample comes from a cell with a 22q deletion on one of the two chromosomes, given that 22q deletions are common in breast cancer and that cells with deletions in the 22q region on both chromosomes do not usually survive. A physician could also infer that if DNA from a mixed sample derived from tumor cells originates from a set of genetically derived tumor cells in which the HER2 and 22q regions are homogeneous, then those cells contain a 5-fold duplication of the HER2 region.

[0363] For example, sample A is also analyzed using an alleleological method. The following are exemplary possible data: Two haplotypes for the same region on the q arm of chromosome 22 are present in a ratio of 4:5, two haplotypes in the focal region corresponding to the HER2 gene are present in a ratio of 1:2, and two haplotypes in the p arm of chromosome 5 are present in a ratio of 20:21. All other assayed regions of the genome do not contain any haplotype in a statistically significant excess. A physician may infer that the sample contains DNA from a tumor with CNVs in the 22q region, the HER2 region, and the 5p arm. Based on the knowledge that 22q deletions are very common in breast cancer and / or quantitative analysis showing an underprevalence of the amount of DNA mapping to the 22q region of the genome, a physician may infer the presence of a tumor with a 22q deletion. Based on the knowledge that HER2 amplification is very common in breast cancer and / or quantitative analysis showing an overabundance of DNA that maps to the HER2 region of the genome, physicians may infer the presence of a tumor with HER2 amplification.

[0364] Exemplary reference chromosome or chromosomal segment In some embodiments, one of the methods described herein is also performed on one or more reference chromosomes or chromosomal segments, and the results are compared with the results for one or more chromosomes or chromosomal segments of interest.

[0365] In some embodiments, a reference chromosome or chromosomal segment is used as a control in which CNVs are expected to be absent. In some embodiments, the reference is the same chromosome or chromosomal segment from one or more different samples that are known or expected to have no deletions or duplications in the chromosome or chromosomal segment. In some embodiments, the reference is a different chromosome or chromosomal segment from the sample being tested that is expected to be disomy. In some embodiments, the reference is a different segment from one of the chromosomes of interest in the same sample being tested. For example, the reference may be one or more segments outside the region of potential deletion or duplication. Having a reference on the same chromosome being tested avoids differences between different chromosomes, e.g., differences in metabolism, apoptosis, histones, inactivation, and / or amplification between chromosomes. Analyzing a segment without CNVs on the same chromosome being tested can also be used to determine differences in metabolism, apoptosis, histones, inactivation, and / or amplification between chromosomes, allowing for the determination of the level of variability between CNV-free homologs to be compared to results from potential CNVs. In some embodiments, the magnitude of the difference between the calculated and predicted allele ratios for potential CNVs is greater than the corresponding magnitude for a reference, thereby confirming the presence of CNVs.

[0366] In some embodiments, a reference chromosome or chromosomal segment is used as a control in which a CNV (e.g., a specific deletion or duplication of interest) is expected to be present. In some embodiments, the reference is the same chromosome or chromosomal segment from one or more different samples known or expected to have a deletion or duplication in the chromosome or chromosomal segment. In some embodiments, the reference is a different chromosome or chromosomal segment from the sample being tested that is known or expected to have a CNV. In some embodiments, the magnitude of the difference between the calculated and predicted allele ratios for a potential CNV is similar to the corresponding magnitude for the reference for the CNV (e.g., not significantly different), thereby confirming the presence of the CNV. In some embodiments, the magnitude of the difference between the calculated and predicted allele ratios for a potential CNV is smaller than the corresponding magnitude for the reference for the CNV (e.g., significantly smaller), thereby confirming the absence of the CNV. In some embodiments, the tumor fraction is determined using one or more loci (or DNA or RNA from cancer cells, such as cfDNA or cfRNA) for the genotype of cancer cells, which is different from the genotype of non-cancerous cells (or DNA or RNA from non-cancerous cells, e.g., cfDNA or cfRNA). Tumor fractions can be used to determine whether the overpopulation of a first homologous chromosome segment is due to duplication of the first homologous chromosome segment or deletion of a second homologous chromosome segment. Tumor fractions can also be used to determine the number of duplicated chromosome segments or the overpopulation of chromosomes (e.g., whether there are 1, 2, 3, 4, or more overpopulations), allowing for differentiation between a sample with 2 overpopulations and a tumor fraction of 20% and a sample with 4 overpopulations and a tumor fraction of 10%. Tumor fractions can also be used to determine how well the observed data fit with predictive data for possible CNVs. In some embodiments, the degree of CNV overpopulation is used to select a specific therapy or treatment regimen for an individual.For example, some therapeutic drugs are effective against only four, six, or even more copies of a chromosomal segment.

[0367] In some embodiments, one or more loci used to determine the tumor fraction are to a reference chromosome or chromosomal segment (e.g., a chromosome or chromosomal segment known or predicted to be disomy, a chromosome or chromosomal segment that is generally not duplicated or deleted in cancer cells or in certain types of cancer in individuals known to have or at increased risk of having, or a chromosome or chromosomal segment with a low probability of aneuploidy (e.g., such a segment that, if deleted or duplicated, is predicted to cause cell death). In some embodiments, one or more chromosomes or chromosomal segments are used to confirm that the reference chromosome or chromosomal segment is disomy in both cancer cells and non-cancerous cells using any of the methods of the present invention. In some embodiments, one or more chromosomes or chromosomal segments with high reliability in calling for disomy are used.

[0368] Exemplary loci that can be used to determine tumor fraction include polymorphisms or mutations (e.g., SNPs) in cancer cells (or DNA or RNA such as cfDNA or cfRNA from cancer cells) that are not present in non-cancerous cells (or DNA or RNA from non-cancerous cells) in an individual. In some embodiments, tumor fraction is determined by identifying these polymorphic loci where cancer cells (or DNA or RNA from cancer cells) in a sample from an individual (e.g., a plasma sample or tumor specimen) have alleles not present in non-cancerous cells (or DNA or RNA from non-cancerous cells), and by using the amount of alleles specific to cancer cells at one or more of the identified polymorphic loci to determine the tumor fraction in the sample. In some embodiments, non-cancerous cells are homozygous for a first allele at a polymorphic locus, and cancer cells are (i) heterozygous for a first allele and a second allele, or (ii) homozygous for a second allele at a polymorphic locus. In some embodiments, non-cancerous cells are heterozygous for a first and second allele at a polymorphic locus, and cancer cells have (i) one or two copies of a third allele at the polymorphic locus. In some embodiments, it is assumed or known that cancer cells have only one copy of an allele that is not present in non-cancerous cells. For example, if the genotype of non-cancerous cells is AA and the cancer cells are AB, and 5% of the signal at that locus in a sample is from the B allele and 95% is from the A allele, then the tumor fraction of the sample is 10%. In some embodiments, it is assumed or known that cancer cells have two copies of an allele that is not present in non-cancerous cells. For example, if the genotype of non-cancerous cells is AA and the cancer cells are BB, and 5% of the signal at that locus in a sample is from the B allele and 95% is from the A allele, then the tumor fraction of the sample is 5%. In some embodiments, cancer cells analyze multiple loci that have alleles not present in non-cancerous cells to determine which loci in the cancer cells are heterozygous and which are homozygous. For example, if, for a locus where non-cancerous cells are AA, the signal from the B allele is approximately 5% at some loci and approximately 10% at others, it is assumed that cancer cells are heterozygous at loci with approximately 5% B alleles and homozygous at loci with approximately 10% B alleles (indicating a tumor fraction of approximately 10%).

[0369] Exemplary loci that can be used to determine the tumor fraction include loci where cancer cells and non-cancerous cells share a single allele (e.g., loci where cancer cells are AB and non-cancerous cells are BB, or where cancer cells are BB and non-cancerous cells are AB). The amount of A signal, the amount of B signal, or the ratio of A signal to B signal in a mixed sample (containing DNA or RNA from cancer cells and non-cancerous cells) is compared to the corresponding values ​​for (i) a sample containing DNA or RNA from cancer cells only or (ii) a sample containing DNA or RNA from non-cancerous cells only. The difference in these values ​​is used to determine the tumor fraction of the mixed sample.

[0370] In some embodiments, loci available for determining tumor fraction are selected based on the genotypes of (i) samples containing DNA or RNA from cancer cells only and / or (ii) samples containing DNA or RNA from non-cancerous cells only. In some embodiments, loci are selected based on an analysis of a mixed sample, for example, loci where the absolute or relative amount of each allele differs from the amount predicted if both cancer cells and cancerous cells had the same genotype at a particular locus. For example, if cancer cells and non-cancerous cells have the same genotype, a locus is predicted to produce a 0% B signal if all cells are AA, a 50% B signal if all cells are AB, or a 100% B signal if all cells are BB. Other values ​​of the B signal indicate that the genotypes of cancer cells and non-cancerous cells differ at that locus, and therefore that the locus can be used to determine tumor fraction.

[0371] In some embodiments, tumor fractions calculated based on alleles at one or more gene loci are compared with tumor fractions calculated using one or more of the counting methods disclosed herein.

[0372] Exemplary methods for detecting phenotypes or analyzing multiple mutations In some embodiments, the method involves analyzing a sample for a set of mutations associated with a disease or disorder (e.g., cancer) or an increased risk of a disease or disorder. Strong correlations exist between events within a class (e.g., cancer classes M or C), which can be used to improve the signal-to-noise ratio of the method and classify tumors into distinct clinical subsets. For example, boundary results for several mutations (e.g., several CNVs) on one or more chromosomes or chromosomal segments considered together can be a very strong signal. In some embodiments, determining the presence or absence of multiple polymorphisms or mutations of interest (e.g., 2, 3, 4, 5, 8, 10, 12, 15, or more) increases the sensitivity and / or specificity of determining the presence or absence of a disease or disorder (e.g., cancer) or an increased risk of a disease or disorder (e.g., cancer). In some embodiments, using correlations between events across multiple chromosomes yields a stronger signal compared to viewing each of them individually. The design of the method itself can be optimized to best classify tumors. This would be very useful for early detection and screening for recurrence, where sensitivity to one particular mutation / CNV may be most important. In some embodiments, events are not always correlated, but have a probability of being correlated. In some embodiments, a matrix estimation composition is used that has a noise covariance matrix with off-diagonal terms.

[0373] In some embodiments, the present invention features a method for detecting a phenotype (e.g., a cancer phenotype) in an individual, where the phenotype is defined by the presence of at least one of a set of mutations. In some embodiments, the method includes obtaining a DNA or RNA measurement of a sample of DNA or RNA from one or more cells of an individual, such that one or more cells are suspected to have the phenotype, and analyzing the DNA or RNA measurement to determine the likelihood that, for each mutation in the set of mutations, at least one cell has that mutation. In some embodiments, the method includes determining that an individual has the phenotype if (i) for at least one of the mutations, the likelihood that at least one cell has that mutation is greater than a threshold, or (ii) for at least one of the mutations, the likelihood that at least one cell has that mutation is less than a threshold, and for multiple mutations, the bound likelihood that at least one cell has at least one of the mutations is greater than a threshold. In some embodiments, one or more cells have a subset or all of the mutations in the set of mutations. In some embodiments, the subset of mutations is associated with cancer or an increased risk of cancer. In some embodiments, the set of mutations includes a subset or all of the mutations in the M class of cancer mutations (Ciriello, Nat Genet. 45(10):1127-1133, 2013, doi:10.1038 / ng.2762, the whole of which is incorporated herein by reference). In some embodiments, the set of mutations includes a subset or all of the mutations in the C class of cancer mutations (Ciriello, ibid.). In some embodiments, the sample includes cell-free DNA or RNA. In some embodiments, the DNA or RNA measurement includes measurements at a set of polymorphic loci on one or more chromosomes or chromosomal segments of interest (e.g., the amount of each allele at each locus).

[0374] Exemplary combinations of methods To improve the accuracy of the results, two or more methods (e.g., any of the methods of the present invention or any known method) are used to detect the presence or absence of CNV. In some embodiments, one or more methods (e.g., any of the methods of the present invention or any known method) are used to analyze factors that are indicators of the presence or absence of a disease or disorder or an increased risk of a disease or disorder.

[0375] In some embodiments, standard mathematical techniques are used to calculate the covariance and / or correlation between two or more methods. Standard mathematical techniques may also be used to determine the combined probability of a particular hypothesis based on two or more trials. Exemplary techniques include meta-analysis, Fisher's combined probability test for independent trials, Brown's method for combining dependent p-values ​​and known covariances, and the cost method for combining dependent p-values ​​and unknown covariances. Combining likelihoods is straightforward when the likelihood is orthogonal to the method determined by the second method, or determined by the first method in an unrelated way, and can be done by multiplication and normalization, or by using formulas such as the following:

[0376] R 結合 =R1R2 / [R1R2+(1-R1)(1-R2)]

[0377] R 結合 R1 is the combined likelihood, and R2 are the individual likelihoods. For example, if the likelihood of trisomy from Method 1 is 90% and the likelihood of trisomy from Method 2 is 95%, by combining the outputs from these two methods, the physician can conclude that the fetus has trisomy with a likelihood of (0.90)(0.95) / [(0.90)(0.95)+(1-0.90)(1-0.95)] = 99.42%. Likelihoods can also be combined if the first and second methods are not orthogonal, i.e., if there is a correlation between the two methods.

[0378] Exemplary methods for analyzing multiple factors or variables are disclosed in U.S. Patent No. 8,024,128, registered September 20, 2011, U.S. Publication No. 2007 / 0027636, filed July 31, 2006, and U.S. Publication No. 2007 / 0178501, filed December 6, 2006, which are incorporated herein by reference, respectively.

[0379] In various embodiments, the binding probability of a particular hypothesis or diagnosis is greater than 80, 85, 90, 92, 94, 96, 98, 99, or 99.9%, or greater than some other threshold.

[0380] Detection limit As demonstrated by the experiments provided in the Examples section, the methods provided herein can detect mean allele imbalance in a sample with a detection or sensitivity limit of 0.45% AAI (which is the detection limit for aneuploidy in the exemplary methods of the present invention). Similarly, in certain embodiments, the methods provided herein can detect mean allele imbalance in samples at 0.45%, 0.5%, 0.6%, 0.8%, 0.8%, 0.9%, or 1.0%. That is, the test methods can detect chromosomal aneuploidy in a given sample where the AAI drops to 0.45%, 0.5%, 0.6%, 0.8%, 0.8%, 0.9%, or 1.0%. As demonstrated by the experiments provided in the Examples section, the methods provided herein can detect the presence of at least some SNVs in a given sample with a detection or sensitivity limit of 0.2%, which is the detection limit for at least some SNVs in one exemplary embodiment. Similarly, in certain embodiments, the method can detect SNVs at frequencies of 0.2, 0.3, 0.4, 0.5, 0.6, 0.8, 0.8, 0.9, or 1.0% or by SNV AAI. That is, the test method can detect SNVs in samples where the detection limit is reduced to 0.2, 0.3, 0.4, 0.5, 0.6, 0.8, 0.8, 0.9, or 1.0% of the total alleles at the chromosomal locus of the SNV.

[0381] In some embodiments, the detection limits for the variations (e.g., SNVs or CNVs) of the methods of the present invention are less than or equal to 10, 5, 2, 1, 0.5, 0.1, 0.05, 0.01, or 0.005%. In some embodiments, the detection limits for the variations (e.g., SNVs or CNVs) of the methods of the present invention are 15 to 0.005%, for example, 10 to 0.005%, 10 to 0.01%, 10 to 0.1%, 5 to 0.005%, 5 to 0.01%, 5 to 0.1%, 1 to 0.005%, 1 to 0.01%, 1 to 0.1%, 0.5 to 0.005%, 0.5 to 0.01%, 0.5 to 0.1%, or 0.1 to 0.01 (including boundary values).

[0382] In some embodiments, the detection limit is the value at which a mutation (e.g., SNV or CNV) present in a sample (e.g., a sample of cfDNA or cfRNA) in an amount less than or equal to 10, 5, 2, 1, 0.5, 0.1, 0.05, 0.01, or 0.005% of the DNA or RNA molecules containing the locus can be detected (or is detectable). For example, a mutation can be detected if it is less than or equal to 10, 5, 2, 1, 0.5, 0.1, 0.05, 0.01, or 0.005% of the DNA or RNA molecules containing the locus with the mutation (e.g., instead of the wild-type or non-mutant form of the locus or a different mutation present in that locus). In some embodiments, the detection limit is the value at which mutations (e.g., SNVs or CNVs) present in amounts less than or equal to 10, 5, 2, 1, 0.5, 0.1, 0.05, 0.01, or 0.005% of DNA or RNA molecules in the sample (e.g., a sample of cfDNA or cfRNA) are detected (or can be detected). In some embodiments where the CNV is a deletion, the deletion can be detected even if it is present in amounts less than or equal to 10, 5, 2, 1, 0.5, 0.1, 0.05, 0.01, or 0.005% of DNA or RNA molecules having the region of interest, which may or may not contain the deletion. In some embodiments where CNV is duplication, duplication can be detected even if the present excess duplicated RNA or DNA is present in amounts less than or equal to 10, 5, 2, 1, 0.5, 0.1, 0.05, 0.01, or 0.005% of the DNA or RNA molecule having the region of interest which may or may not be duplicated in the sample.In some embodiments where CNV is duplication, duplication can be detected even if the present excess duplicated RNA or DNA is present in amounts less than or equal to 10, 5, 2, 1, 0.5, 0.1, 0.05, 0.01, or 0.005% of the DNA or RNA molecules in the sample.

[0383] Exemplary sample In some embodiments of any aspect of the present invention, the sample comprises intracellular and / or extracellular genetic material from cells suspected of having deletions or duplications, e.g., cells suspected of being cancerous. In some embodiments, the sample comprises any tissue or bodily fluid (e.g., tumor) suspected of containing cells, DNA or RNA, or other samples containing cancer cells, DNA or RNA. Genetic measurements used as part of these methods may be performed on any sample containing DNA or RNA, e.g., but not limited to tissue, blood, serum, plasma, urine, hair, tears, saliva, skin, fingernails, feces, bile, lymph, cervical mucus, semen, tumor, or other cells or substances containing nucleic acids. The sample may contain any cell type, or use DNA or RNA from any cell type (e.g., cells or neurons from any organ or tissue suspected of being cancerous). In some embodiments, the sample comprises nuclear and / or mitochondrial DNA. In some embodiments, the sample is derived from any of the target individuals disclosed herein. In some embodiments, the target individual is a cancer patient.

[0384] Exemplary samples include those containing cfDNA or cfRNA. In some embodiments, cfDNA is available for analysis without requiring a cell lysis step. Cell-free DNA may be obtained from various tissues, for example, tissues in liquid form, such as blood, plasma, lymph, ascites, or cerebrospinal fluid. In some cases, cfDNA consists of DNA derived from fetal cells. In other cases, cfDNA is isolated from plasma isolated from whole blood after centrifugation to remove cellular material. cfDNA may be a mixture of DNA derived from target cells (e.g., cancer cells) and non-target cells (e.g., non-cancer cells).

[0385] In some embodiments, the sample contains, or is suspected to contain, a mixture of DNA (or RNA), for example, a mixture of DNA (or RNA) derived from cancer cells and DNA (or RNA) derived from non-cancerous (i.e., normal) cells. In some embodiments, at least 0.5, 1, 3, 5, 7, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 92, 94, 95, 96, 98, 99, or 100% of the cells in the sample are cancer cells. In some embodiments, at least 0.5, 1, 3, 5, 7, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 92, 94, 95, 96, 98, 99, or 100% of the DNA (e.g., cfDNA) or RNA (e.g., cfRNA) in the sample is derived from cancer cells. In various embodiments, the percentage of cells that are cancerous in the sample is 0.5–99%, for example, 1–95%, 5–95%, 10–90%, 5–70%, 10–70%, 20–90%, or 20–70% (including boundary values). In some embodiments, the sample is enriched with cancer cells or with DNA or RNA from cancer cells. In some embodiments of the cancer cell-enriched sample, at least 0.5%, 1, 2, 3, 4, 5, 6, 7, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 92, 94, 95, 96, 98, 99%, or 100% of the cells in the enriched sample are cancer cells. In some embodiments of a sample enriched with DNA or RNA from cancer cells, at least 0.5, 1, 2, 3, 4, 5, 6, 7, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 92, 94, 95, 96, 98, 99, or 100% of the DNA or RNA in the enriched sample is derived from cancer cells.In some embodiments, cancer cells are enriched using cell sorting (e.g., fluorescence-activated cell sorting (FACS)) (Barteneva et al., Biochim Biophys Acta., 1836(1):105-22, August 2013 doi:10.1016 / j.bbcan.2013.02.004.Epub February 24, 2013, and Ibrahim et al., Adv Biochem Eng Biotechnol. 106:19-39, 2007, both incorporated herein by reference).

[0386] In some embodiments, the sample is enriched with fetal cells. In some embodiments of the sample enriched with fetal cells, at least 0.5, 1, 2, 3, 4, 5, 6, 7%, or more of the cells in the enriched sample are fetal cells. In some embodiments, the percentage of cells that are fetal cells in the sample is 0.5–100%, for example, 1–99%, 5–95%, 10–95%, 10–95%, 20–90%, or 30–70% (including boundary values). In some embodiments, the sample is enriched with fetal DNA. In some embodiments of the sample enriched with fetal DNA, at least 0.5, 1, 2, 3, 4, 5, 6, 7%, or more of the DNA in the enriched sample is fetal DNA. In some embodiments, the percentage of DNA that is fetal DNA in the sample is 0.5–100%, for example, 1–99%, 5–95%, 10–95%, 10–95%, 20–90%, or 30–70% (including boundary values).

[0387] In some embodiments, the sample contains a single cell or DNA and / or RNA from a single cell. In some embodiments, multiple individual cells (e.g., at least 5, 10, 20, 30, 40, or 50 cells from the same or different subjects) are analyzed in parallel. In some embodiments, cells from multiple samples from the same individual are combined, reducing the workload compared to analyzing these samples separately. Combining multiple samples also makes it possible to test multiple tissues for cancer simultaneously (which can be used to provide a more thorough screening for cancer or to determine if cancer may have metastasized to other tissues).

[0388] In some embodiments, the sample comprises a single cell or a small number of cells, e.g., 2, 3, 5, 6, 7, 8, 9, or 10 cells. In some embodiments, the sample comprises 1 to 100, 100 to 500, or 500 to 1,000 cells (including boundary values). In some embodiments, the sample comprises 1 to 10 picograms, 10 to 100 picograms, 100 picograms to 1 nanogram, 1 to 10 nanograms, 10 to 100 nanograms, or 100 nanograms to 1 microgram of RNA and / or DNA (including boundary values).

[0389] In some embodiments, the sample is embedded in Parafilm. In some embodiments, the sample is preserved with a preservative such as formaldehyde and, optionally, encapsulated in paraffin, which may cause DNA crosslinking so that a small amount is available for PCR. In some embodiments, the sample is a formaldehyde-fixed paraffin-embedded (FFPE) sample. In some embodiments, the sample is a fresh sample (e.g., a sample obtained from analysis on day 1 or 2). In some embodiments, the sample is frozen before analysis. In some embodiments, the sample is a historical sample.

[0390] These samples can be used in any of the methods of the present invention.

[0391] Exemplary Sample Preparation Method In some embodiments, the method includes isolating or purifying DNA and / or RNA. Several standard procedures known in the art exist to achieve such objectives. In some embodiments, the sample may be centrifuged to separate various layers. In some embodiments, DNA or RNA may be isolated by filtration. In some embodiments, the preparation of DNA or RNA may involve amplification, separation, purification by chromatography, liquid-liquid separation, isolation, preferential concentration, preferential amplification, targeted amplification, or any of several other techniques known in the art or described herein. In some embodiments for DNA isolation, RNA is degraded using RNase. In some embodiments for RNA isolation, DNA is degraded using DNase (e.g., DNase I from Invitrogen, Carlsbad, CA, USA). In some embodiments, RNA is isolated using an RNeasy mini-kit (Qiagen) according to the manufacturer's protocol. In some embodiments, small RNAs are isolated using the mirVana PARIS kit (Ambion, Austin, TX, USA) according to the manufacturer's protocol (Gu et al., J. Neurochem. 122:641-649, 2012, the whole of which is incorporated herein by reference). The concentration and purity of the RNA may optionally be determined using Nanovue (GE Healthcare, Piscataway, NJ, USA), and the integrity of the RNA may optionally be measured using the 2100 Bioanalyzer (Agilent Technologies, Santa Clara, CA, USA) (Gu et al., J. Neurochem. 122:641-649, 2012, the whole of which is incorporated herein by reference). In some embodiments, the RNA is stabilized during storage using TRIZOL or RNAlater (Ambion).

[0392] In some embodiments, a universal tagging adapter is added to create a library. Prior to ligation, the sample DNA may be blunt-ended and then have a single adenosine base added to the 3' end. Prior to ligation, the DNA may be cleaved using restriction enzymes or some other cleavage method. During ligation, the 3'-adenosine of the sample fragment and the complementary 3'-tyrosine overhang of the adapter can enhance ligation efficiency. In some embodiments, adapter ligation is performed using a ligation kit found in the AGILENT SURESELECT kit. In some embodiments, the library is amplified using universal primers. In one embodiment, the amplified library is fractionated by size separation or by using a product such as AGENCORT AMPURE beads or other similar methods. In some embodiments, the target locus is amplified using PCR amplification. In some embodiments, the amplified DNA is sequenced (e.g., ILLUMINA IIGAX or HiSeq sequencer). In some embodiments, the amplified DNA is sequenced from each end of the amplified DNA to reduce sequencing errors. When sequencing from one end of the amplified DNA, if a sequence error exists at a specific base, the likelihood of a sequence error in the complementary base when sequencing from the other end of the amplified DNA is lower (compared to multiple sequencing attempts from the same end of the amplified DNA).

[0393] In some embodiments, whole-genome applications (WGA) are used to amplify nucleic acid samples. Several methods are available for WGA, including ligation-mediated PCR (LM-PCR), denatured oligonucleotide-primer PCR (DOP-PCR), and multiple substitution amplification (MDA). In LM-PCR, short DNA sequences called adapters are ligated to the blunt ends of the DNA. These adapters contain universal amplification sequences used to amplify the DNA by PCR. In DOP-PCR, random primers, which also contain universal amplification sequences, are used in the first round of annealing and PCR. A second round of PCR is then used to further amplify the sequence using the universal primer sequences. MDA uses phi-29 polymerase, a highly processive nonspecific enzyme that has been used to replicate DNA and for single-cell analysis. In some embodiments, WGA is not performed.

[0394] In some embodiments, selective amplification or enrichment is used to amplify or enrich a target gene locus. In some embodiments, amplification and / or selective enrichment techniques may involve PCR (e.g., ligation-mediated PCR), fraction capture by hybridization, molecular inversion probes or other cyclized probes. In some embodiments, real-time quantitative PCR (RT-qPCR), digital PCR, or emulsion PCR, or mass spectrometry after a single-allele base extension reaction is used (Hung et al., J Clin Pathol 62:308-313, 2009, the whole of which is incorporated herein by reference). In some embodiments, DNA is preferentially enriched using capture by hybridization with a hybrid capture probe. In some embodiments, the method for amplification or selective enrichment may involve using a probe in which, upon correct hybridization to the target sequence, the 3' or 5' end of the nucleotide probe is separated from the polymorphic site of the polymorphic allele by a small number of nucleotides. This separation reduces the preferential amplification of one allele, known as allele bias. This is an improvement to the method, involving the use of probes in which the 3' or 5' end of the correctly hybridized probe is directly adjacent to, or very close to, the polymorphic site of the allele. In one embodiment, probes in which the hybridizing region may contain, or certainly contains, a polymorphic site are excluded. Polymorphic sites in the hybridization site may cause uneven hybridization or complete inhibition of hybridization in some alleles, which may result in preferential amplification of certain alleles. These embodiments are improvements to other methods involving targeted amplification and / or selective enrichment in that they well preserve the original allele frequencies of the sample at each polymorphic locus, where the sample is a pure genomic sample from a single individual or a mixture of individuals.

[0395] In some embodiments, PCR (referred to as mini-PCR) is used to create very short amplicons (U.S. Application No. 13 / 683,604 filed November 21, 2012, U.S. Publication No. 2013 / 0123120, U.S. Application No. 13 / 300,235 filed November 18, 2011, U.S. Publication No. 2012 / 0270212 filed November 18, 2011, and U.S. Application No. 61 / 994,791 filed May 16, 2014, each incorporated herein by reference as a whole). cfDNA (e.g., cancer cfDNA released by necrosis or apoptosis) is highly fragmented. In the case of fetal cfDNA, fragment sizes are distributed in a nearly Gaussian manner, with a mean of 160 bp, a standard deviation of 15 bp, a minimum size of approximately 100 bp, and a maximum size of approximately 220 bp. Polymorphic sites of a particular target locus may occupy any position from beginning to end of the various fragments derived from that locus. Because cfDNA fragments are short, the likelihood of both primer sites being present, and the likelihood of a fragment of length L containing both forward and reverse primer sites, is the ratio of the amplicon length to the fragment length. Under ideal conditions, assays with amplicons of 45, 50, 55, 60, 65, or 70 bp successfully amplify 72%, 69%, 66%, 63%, 59%, or 56% of the available template fragment molecules, respectively. In a particular embodiment most preferably relevant to cfDNA from a sample of an individual suspected of having cancer, the cfDNA is amplified using primers having a melting point of 50–65°C, 54–60.5°C, giving a maximum amplicon length of 85, 80, 75, or 70 bp, or 75 bp in a particular preferred embodiment. The amplicon length is the distance between the 5' ends of the forward and reverse priming sites. Shorter amplicon lengths than those typically used by those known in the art may result in more efficient measurement of desired polymorphic loci by requiring only short sequence reads. In one embodiment, the substantial fraction of the amplicon is less than 100 bp, less than 90 bp, less than 80 bp, less than 70 bp, less than 65 bp, less than 60 bp, less than 55 bp, less than 50 bp, or less than 45 bp.

[0396] In some embodiments, amplification is performed using direct multiplexed PCR, serial PCR, nested PCR, double nested PCR, one-and-a-half This is performed using sided nested PCR, fully nested PCR, one-sided fully nested PCR, one-sided nested PCR, heminested PCR, heminested PCR, triple heminested PCR, seminested PCR, one-sided seminested PCR, reverse seminested PCR, or one-sided PCR, which are described in U.S. Application No. 13 / 683,604, U.S. Publication No. 2013 / 0123120, filed November 21, 2012; U.S. Application No. 13 / 300,235, U.S. Publication No. 2012 / 0270212, filed November 18, 2011; and U.S. Application No. 61 / 994,791, filed May 16, 2014 (which are incorporated herein by reference in their entirety). If desired, any of these methods may be used for miniPCR.

[0397] If desired, the extension step of PCR amplification may be limited in terms of time to reduce amplification from fragments longer than 200, 300, 400, 500, or 1,000 nucleotides. This may result in enrichment of fragmented DNA or shorter DNA (e.g., fetal DNA or DNA from apoptotic or necrotic cancer cells), which may improve test performance.

[0398] In some embodiments, multiplex PCR is used. In some embodiments, a method for amplifying target loci in a nucleic acid sample involves (i) contacting the nucleic acid sample with a library of primers that simultaneously hybridize at least 100, 200, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 20,000, 25,000, 30,000, 40,000, 50,000, 75,000, or 100,000 different target loci to generate a reaction mixture, and (ii) subjecting this reaction mixture to primer extension reaction conditions (e.g., PCR conditions) to generate an amplification product containing the target amplicon. In some embodiments, at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the target loci are amplified. In various embodiments, less than 60, 50, 40, 30, 20, 10, 5, 4, 3, 2, 1, 0.5, 0.25, 0.1, or 0.05% of the amplification product is the primer dimer. In some embodiments, the primer is in solution (e.g., dissolved in a liquid phase rather than a solid phase). In some embodiments, the primer is in solution and not immobilized on a solid support. In some embodiments, the primer is not part of a microarray. In some embodiments, the primer does not contain a molecular inversion probe (MIP).

[0399] In some embodiments, two or more (e.g., three or four) target amplicons (e.g., amplicons from the miniPCR methods disclosed herein) are ligated together, and the ligated product is then sequenced. Combining multiple amplicons to form a single ligation product increases the efficiency of the subsequent sequencing step. In some embodiments, the target amplicons are less than 150, 100, 90, 75, or 50 base pairs in length before they are ligated. Selective enrichment and / or amplification may involve tagging each individual molecule with different tags, molecular barcodes, tags for amplification, and / or tags for sequencing. In some embodiments, the amplified product is analyzed by sequencing (e.g., high-throughput sequencing) or by hybridization to an array, e.g., an SNP array, an ILLUMINA INFINIUM array, or an AFFYMETRIX gene chip. In some embodiments, nanopore sequencing is used, for example, the nanopore sequencing technology developed by Genia (see, for example, the entire article incorporated herein by reference, on the World Wide Web at geniachip.com / technology). In some embodiments, dual sequencing is used (Schmitt et al., "Detection of ultra-rare mutations by next-generation sequencing," Proc Natl Acad Sci US A.109(36):14508-14513, 2012, the entire article incorporated herein by reference). This method significantly reduces errors by independently tagging and sequencing each of the two strands of a double-stranded DNA. Because these two strands are complementary, true mutations are found at the same location on both strands. In contrast, errors in PCR or sequencing can be discounted as technical errors because they result in mutations on only one strand. In some embodiments, the method involves tagging both strands of double-stranded DNA with random but complementary double-stranded nucleotide sequences (called double-stranded tags).First, a single-stranded randomized nucleotide sequence is introduced into one adapter strand, and then the opposite strand is extended using DNA polymer...

Claims

1. A method for monitoring and detecting minimal residual disease in colorectal cancer, Whole exome sequencing is performed on tumor samples from patients diagnosed with colorectal cancer to identify somatic mutations associated with the colorectal cancer, and a set of at least eight or sixteen patient-specific single nucleotide variant (SNV) loci is selected based on the somatic mutations identified in the tumor samples. A set of amplicons is created by performing a multiple amplification reaction on nucleic acids isolated from the patient's blood sample or a fraction thereof, wherein each amplicon in the set of amplicons extends to the patient-specific SNV locus of the patient-specific set of SNV loci associated with colorectal cancer. High-throughput sequencing is performed to determine the sequence of at least one segment of each amplicon in the set of amplicons spread across patient-specific SNV loci, wherein the high-throughput sequencing is performed with a read depth of at least 100,000 per locus, and the detection of two or more patient-specific SNVs exceeding a confidence threshold of 0.97 from the set of at least eight or sixteen patient-specific SNV loci is an indicator of the presence of circulating tumor DNA (ctDNA) in the blood sample and minimal residual disease in the patient. Methods that include...

2. A method for monitoring and detecting minimal residual disease in colorectal cancer, Creating a set of amplicons by performing multiple amplification reactions on nucleic acids isolated from blood samples or fractions thereof from patients diagnosed with colorectal cancer, wherein each amplicon in the set of amplicons extends to at least one SNV locus from a set of at least eight or sixteen patient-specific single nucleotide variant (SNV) loci associated with colorectal cancer, selected based on somatic mutations associated with colorectal cancer identified in the patient's tumor sample. High-throughput sequencing is performed to determine the sequence of at least one segment of each amplicon in the set of amplicons spread across patient-specific SNV loci, wherein the high-throughput sequencing is performed with a read depth of at least 100,000 per locus, and the detection of two or more patient-specific SNVs exceeding a confidence threshold of 0.97 from the set of at least eight or sixteen patient-specific SNV loci is an indicator of the presence of circulating tumor DNA (ctDNA) in the blood sample and minimal residual disease in the patient. Methods that include...

3. The method according to claim 1 or 2, wherein the method detects two or more patient-specific SNVs in a patient having early recurrence or metastasis of cancer at least 10.2 months prior to clinical recurrence or metastasis of cancer detectable by CT imaging.

4. The method according to any one of claims 1 to 3, wherein the method detects two or more patient-specific SNVs in at least 85% of patients with early recurrence or metastasis of cancer.

5. The method according to any one of claims 1 to 4, wherein the method does not detect two or more patient-specific SNVs in at least 95% of patients who do not have early recurrence or metastasis of cancer.

6. A method for monitoring and detecting minimal residual disease in colorectal cancer, Creating a set of amplicons by performing multiple amplification reactions on nucleic acids isolated from blood samples or fractions thereof from patients treated for colorectal cancer, wherein each amplicon in the set of amplicons extends to at least one SNV locus from a set of at least eight or sixteen patient-specific single nucleotide variant (SNV) loci associated with colorectal cancer, selected based on somatic mutations associated with colorectal cancer identified in the patient's tumor sample. High-throughput sequencing is performed to determine the sequence of at least one segment of each amplicon in the set of amplicons spread across patient-specific SNV loci, wherein the high-throughput sequencing is performed with a read depth of at least 100,000 per locus, and the detection of two or more patient-specific SNVs exceeding a confidence threshold of 0.97 from the set of at least eight or sixteen patient-specific SNV loci is an indicator of the presence of circulating tumor DNA (ctDNA) in the blood sample and minimal residual disease in the patient. Methods that include...

7. The method according to claim 6, wherein the method has at least 95% specificity in detecting minimal residual disease in the patient.

8. The method according to claim 6 or 7, wherein at least five patient-specific SNVs are detected from the blood sample.

9. The method according to any one of claims 6 to 8, wherein the treatment is neoadjuvant therapy or adjuvant therapy.

10. The method according to any one of claims 6 to 9, wherein the patient is receiving chemotherapy, radiotherapy, or adjuvant therapy before the blood sample is taken from the patient.

11. The method according to any one of claims 6 to 10, further comprising determining the variant allele frequency for each of the patient-specific SNVs detected from the blood sample.

12. The method according to claim 11, further comprising selecting a treatment that targets the clonal SNV, wherein a variant allele frequency of more than 1% or more than 5% is an indicator of clonal SNV in colorectal cancer.

13. The method according to any one of claims 1 to 12, further comprising: forming an amplification reaction mixture by combining polymerase, nucleotide triphosphates, nucleic acid fragments from a nucleic acid library prepared from the blood sample, and a set of target-specific primers that each bind within 150 base pairs of the SNV gene locus; and subjecting the amplification reaction mixture to amplification conditions to create a set of amplicons.

14. The method according to any one of claims 1 to 13, wherein the efficiency and error rate per cycle are determined for each amplification of the multiple amplification reaction of the set of SNV loci, and the efficiency and error rate are used to determine whether patient-specific SNVs are present in the sample.

15. The aforementioned multiple amplification reaction is a PCR reaction performed using a set of primers. (i) The annealing temperature is 1 to 15°C higher than the melting point of at least 50% of the set of primers. (ii) The length of the annealing step during the PCR reaction is 15 to 120 minutes. (iii) The concentration of each primer in the set of primers in the amplification reaction mixture is 1 to 10 nM, and / or (iv) The method according to any one of claims 1 to 14, wherein the set of primers is designed to minimize primer dimer formation.

16. The method according to any one of claims 1 to 15, wherein the patient-specific set of SNV loci includes 64 or more somatic mutations identified in the tumor sample.