Tumor detection and cancer monitoring

The method improves cancer detection and monitoring by employing multiplex DNA amplification and dual machine learning models with diverse target loci, addressing sensitivity and comprehensiveness issues in existing sequencing methods, thereby enhancing detection accuracy and reducing false positives.

WO2026161515A1PCT designated stage Publication Date: 2026-07-30NATERA INC +3
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
NATERA INC
Filing Date
2026-01-22
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing cancer detection and monitoring methods, particularly whole exome sequencing, suffer from sensitivity and comprehensiveness issues, leading to incomplete analysis and high false positive rates due to focusing on limited high-risk targets.

Method used

A method involving multiplex amplification of DNA from cell-free samples, followed by high-throughput sequencing and analysis using two machine learning models with distinct feature sets, to generate a ctDNA positivity classification, incorporating both tumor-specific and additional target loci for improved detection and monitoring.

Benefits of technology

Enhances the sensitivity and reliability of cancer detection by providing a comprehensive analysis of genetic variants, reducing false positives through the use of multiple machine learning models and diverse target loci, enabling precise monitoring of cancer progression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2026012089_30072026_PF_FP_ABST
    Figure US2026012089_30072026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are techniques for monitoring and detecting a tumor status in a subject. Cell-free DNA may be extracted from a sample or a fraction thereof and sequence reads may be obtained for certain loci. A circulating tumor DNA positivity classification may be generated based on an analysis of a first result from a first machine learning model and a second result from a second ML model that is different from the first ML model. The first ML model may receive a first feature set and the second ML model may receive a second feature set different from the first feature set. On-target loci data, in combination with off-target loci data, may be used to obtain an accurate ctDNA positivity classification.
Need to check novelty before this filing date? Find Prior Art

Description

N.062.WO.01TUMOR DETECTION AND CANCER MONITORINGCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] The present application claims priority to U. S. Provisional Patent Application No.63 / 748,333, titled TUMOR DETECTION AND CANCER MONITORING, filed January 22, 2025, which is hereby incorporated by reference herein in its entirety.BACKGROUND

[0002] Previous methods for detecting and monitoring cancer in patients have primarily relied on whole exome sequencing (WES) with few target regions. This approach, while useful, has inherent limitations in terms of sensitivity and comprehensiveness. Typically, these methods focus on the top 3–4 high-risk targets, which can result in incomplete analysis and potential oversight of significant genetic variants. Additionally, maintaining a high specificity requirement has been challenging, leading to a higher likelihood of false positives. Therefore, systems and methods are needed to improve the performance and reliability of cancer detection and monitoring.SUMMARY

[0003] In some aspects, the techniques described herein relate to a method for preparing a composition of amplified DNA from a sample useful for monitoring and detecting a tumor status in a cancer patient, wherein the method includes the steps of: (a) extracting cell-free DNA from the sample or a fraction thereof; (b) preparing a composition of amplified DNA by performing a multiplex amplification reaction on cell-free DNA extracted in (a) or DNA derived therefrom to obtain a set of amplicons encompassing a first set of target loci; (c) performing high-throughput sequencing of at least some of the amplicons to obtain sequence reads of at least a fraction of the set of amplicons encompassing the first set of target loci; and (d) generating a ctDNA positivity classification based on an analysis of a first result from a first machine learning (ML) model and a second result from a second ML model that is different from the first ML model, wherein the first ML model receives a first feature set and the second ML model receives a second feature set different from the first feature set, and wherein the first and second feature sets are based at least in part on a fraction or all of the first set of target loci.N.062.WO.01

[0004] In some aspects, the techniques described herein relate to a method, wherein the analysis aggregates the first result from the first model and the second result from the second model to generate the ctDNA positivity classification.

[0005] In some aspects, the techniques described herein relate to a method, wherein the first feature set includes a first feature that is not in the second feature set, and / or the second feature set includes a second feature that is not included in the first feature set.

[0006] In some aspects, the techniques described herein relate to a method, wherein the ctDNA positivity classification is based on sequence reads for the first set of target loci and sequence reads for one or more additional target loci from a second set of target loci different from the first set of target loci, wherein the second set of target loci are encompassed by the set of amplicons.

[0007] In some aspects, the techniques described herein relate to a method, wherein the sequence reads from the second set of target loci are used to at least one of calibrate, predict, or validate the ctDNA positivity classification. Validation may thus be internal to the analytical workflow without requiring independent experimental confirmation.

[0008] In some aspects, the techniques described herein relate to a method, wherein the first set of target loci are tumor-specific target loci and the second set of target loci is different from the first set of target loci.

[0009] In some aspects, the techniques described herein relate to a method, wherein training of the first ML model included generation of a first training feature set, and wherein training of the second ML model included generation of a second training feature set which differs from the first training feature set.

[0010] In some aspects, the techniques described herein relate to a method, wherein the first set of target loci includes at least 50 loci.

[0011] In some aspects, the techniques described herein relate to a method, wherein the first ML model employs a first algorithm and the second ML model employs a second algorithm different from the first algorithm.

[0012] In some aspects, the techniques described herein relate to a method, wherein at least one of the first algorithm, the second algorithm, or the analysis includes at least one of Platt scaling, geometric averaging, voting scheme, or gradient boosting.

[0013] In some aspects, the techniques described herein relate to a method for preparing a composition of amplified DNA from a sample useful for monitoring and detecting a tumor statusN.062.WO.01in a cancer patient, wherein the method includes the steps of: (a) extracting cell-free DNA from the sample or a fraction thereof; (b) preparing a composition of amplified DNA by performing a multiplex amplification reaction on cell-free DNA extracted in (a) or DNA derived therefrom to obtain a set of amplicons encompassing a first set of target loci and a second set of target loci, wherein the first set of target loci are tumor- specific target loci and the second set of target loci is different from the first set of target loci; (c) performing high-throughput sequencing of at least some of the amplicons to obtain sequence reads of at least a fraction of the set of amplicons encompassing the first set of target loci and at least a fraction of the set of amplicons encompassing the second set of target loci; and (d) generating a ctDNA positivity classification based on an analysis of a first result from a first machine learning (ML) model and a second result from a second ML model, wherein the first ML model receives a first feature set and the second ML model receives a second feature set, wherein the first and second feature sets are based at least in part on a fraction or all of the first set of target loci, and wherein the second set of target loci is used as validation of the ctDNA positivity classification.

[0014] In some aspects, the techniques described herein relate to a method, wherein the first feature set is based on a first feature set corresponding to the first ML model, and the second feature set is based on a second feature set generated for the second ML model.

[0015] In some aspects, the techniques described herein relate to a method, wherein the first feature set includes a first feature that is not in the second feature set, and the second feature set includes a second feature that is not included in the first feature set.

[0016] In some aspects, the techniques described herein relate to a method, wherein the analysis aggregates the first result from the first model and the second result from the second model to generate the ctDNA positivity classification.

[0017] In some aspects, the techniques described herein relate to a method, wherein the first set of target loci are tumor- specific target loci and the second set of target loci is different from the first set of target loci.

[0018] In some aspects, the techniques described herein relate to a method, wherein training of the first ML model included generation of a first feature set, and wherein training of the second ML model included generation of a second feature set different from the first feature set.N.062.WO.01

[0019] In some aspects, the techniques described herein relate to a method, wherein the first ML model employs a first algorithm and the second ML model employs a second algorithm different from the first algorithm.

[0020] In some aspects, the techniques described herein relate to a method, wherein at least one of the first algorithm, the second algorithm, or the analysis includes at least one of Platt scaling, geometric averaging, voting scheme, or gradient boosting.

[0021] In some aspects, the techniques described herein relate to a method, wherein the first set of target loci encompasses a first number of loci, and the second set of target loci encompasses a second number of target loci that is at least one order of magnitude greater than the first number of target loci.

[0022] In some aspects, the techniques described herein relate to a method, wherein the ctDNA positivity classification is indicative of the subject developing or having cancer.

[0023] In some aspects, the techniques described herein relate to the method of any of the above claims, wherein the analysis includes ensembling the first result with the second result to generate a score for the sample, wherein the first result is determined using a first threshold, the second result is determined using a second threshold, and the score for the sample is determined using a third threshold, and wherein the ctDNA positivity classification is based on the score.

[0024] In some aspects, the techniques described herein relate to the method of any of the above claims, wherein the first and second ML models were trained at least in part by generating a first training feature set for the first ML model and generating a second training feature set that is different from the first training feature set for the second ML model.

[0025] In some aspects, the techniques described herein relate to the method of any of the above claims, wherein the first and second ML models were trained at least in part by generating genetic data for a set of samples from a cohort of subjects.

[0026] In some aspects, the techniques described herein relate to a method, wherein generating the genetic data included: obtaining the set of samples; stratifying the set of samples into DNA input groups based on corresponding genetic data; and generating augmented samples corresponding to each DNA input group by repeatedly selecting and combining, for each DNA input group, a plurality of samples from the set of samples to obtain an augmented sample.

[0027] In some aspects, the techniques described herein relate to a method, wherein training the first ML model included a first feature generation algorithm, wherein the first feature generationN.062.WO.01algorithm included: obtaining a probability function by determining, for each target in the first set of target loci and the second set of target loci, a probability that a mutation is present; deriving a first set of target features from the genetic data based on the probability function; and deriving a second set of target features different from the first set of target features.

[0028] In some aspects, the techniques described herein relate to a method, wherein training the second ML model included a second feature generation algorithm, wherein the second feature generation algorithm includes: obtaining a set of posterior probabilities by determining, for each target in the first set of target loci and the second set of target loci, a probability that a mutation is present; deriving a first set of target features from the genetic data based on the set of posterior probabilities; deriving a second set of target features from the genetic data different from the first set of target features.

[0029] In some aspects, the techniques described herein relate to the method of any of the above claims, wherein at least one of the first ML model, the second ML model, and / or the analysis includes at least one Platt scaling, geometric averaging, voting scheme, or gradient boosting.

[0030] In some aspects, the techniques described herein relate to the method of any of the above claims, wherein the first ML model or the second ML model was trained as a binary classification-based algorithm.

[0031] In some aspects, the techniques described herein relate to the method of any of the above claims, wherein at least one of the first ML model or the second ML model was trained by optimizing sensitivity given a specificity constraint.

[0032] In some aspects, the techniques described herein relate to the method of claim any of the above claims, further including training the first and second ML models.

[0033] In some aspects, the techniques described herein relate to a method implemented by a first computing system including one or more processors, the method including: generating genetic data from a biological sample of a subject, the genetic data including information on a first set of target loci corresponding to cell-free DNA (cfDNA) in the sample; generating a ctDNA positivity classification based on the genetic data, wherein generating the ctDNA positivity classification includes: providing a first feature set based on the genetic data to a first machine learning (ML) model to obtain a first result, the first ML model employing a first feature set; providing a second feature set based on the genetic data to a second ML model to obtain a second result, the second ML model employing a second feature set; and aggregating the first result and the second result toN.062.WO.01obtain an aggregated result; and outputting the ctDNA positivity classification, wherein outputting the ctDNA positivity classification includes at least one of storing the ctDNA positivity classification in a non-transitory computer-readable storage medium, displaying the ctDNA positivity classification on one or more display screens of the first computing system, or transmitting the ctDNA positivity classification to a second computing system via a network.

[0034] In some aspects, the techniques described herein relate to a method, wherein generating the ctDNA positivity classification further includes validating the aggregated result based on a second set of target loci. In some aspects, generating the ctDNA positivity classification further includes calibrating or predicting the aggregated result based on the second set of target loci.

[0035] In some aspects, the techniques described herein relate to a method implemented by a first computing system including one or more processors, the method including: generating genetic data from a biological sample of a subject, the genetic data including information on a first set of target loci corresponding to cell-free DNA (cfDNA) in the sample and a second set of target loci corresponding to the cfDNA in the sample, wherein the first set of target loci includes tumorspecific target loci and the second set of target loci is different from the first set of target loci; generating a ctDNA positivity classification based on the genetic data, wherein generating the ctDNA positivity classification includes: providing a first feature set based on the genetic data to a first machine learning (ML) model to obtain a first result, the first feature set being based on the first set of target loci; providing a second feature set based on the genetic data to a second ML model to obtain a second result, the second feature set being based on the first set of target loci; aggregating the first result and the second result to obtain an aggregated result; and validating the aggregated result based on the second set of target loci; and outputting the ctDNA positivity classification, wherein outputting the ctDNA positivity classification includes at least one of storing the ctDNA positivity classification in a non-transitory computer-readable storage medium, displaying the ctDNA positivity classification on one or more display screens of the first computing system, or transmitting the ctDNA positivity classification to a second computing system via a network.

[0036] In some aspects, the techniques described herein relate to a method, wherein the first ML model is based on a first set of features and the second ML model is based on a second set of features that is different from the first set of features.N.062.WO.01BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The presently disclosed embodiments will be further explained with reference to the attached drawings, wherein like structures are referred to by like numerals throughout the several views. The drawings shown are not necessarily to scale, with emphasis instead generally being placed upon illustrating the principles of the presently disclosed embodiments.

[0038] FIG. 1 depicts a block diagram of an example computing system suitable for use in the various arrangements described herein, in accordance with one or more example implementations.

[0039] FIG. 2 depicts an example workflow of embodiments of the methods described herein.

[0040] FIG. 3 depicts an example workflow of embodiments of the methods described herein.

[0041] FIG. 4 depicts an example workflow of embodiments of the methods described herein.

[0042] FIG. 5 depicts an example workflow of embodiments of the methods described herein.

[0043] FIG. 6 is a block diagram of an example of a system, in accordance with implementing some embodiments of the present disclosure.

[0044] FIG. 7 is a block diagram of an example of a system, in accordance with implementing some embodiments of the present disclosure.

[0045] FIG. 8 depicts a block diagram of an example computing system suitable for use in the various arrangements described herein, in accordance with one or more example implementations.

[0046] FIGS. 9A-B depict graphs showing augmented samples DNA input distribution, in accordance with one or more example implementations. FIG. 9A represents augmented training samples DNA input distribution without DNA input stratification (low DNA inputs underrepresented), and FIG. 9B represents augmented training samples DNA input distribution with DNA input stratification (used in the final models; low DNA inputs well-represented).

[0047] FIG. 10 depicts a clustering representation of augmented samples and evidence generation samples in training, in accordance with one or more example implementations. FIG. 10 depicts training sample representation using the subset of available features (unsupervised analysis).

[0048] FIG. 11 depicts calculated probabilities of positive mutation at individual on-target positions, in accordance with one or more example implementations.

[0049] FIGS. 12A-B depict graphs showing a probability distribution and CDF within the first 16, 32, and 64 targets, in accordance with one or more example implementations.

[0050] FIG. 13 depicts a block diagram of an example computing system suitable for use in the various arrangements described herein, in accordance with one or more example implementations.N.062.WO.01

[0051] FIG. 14 depicts a block diagram of an example computing system suitable for use in the various arrangements described herein, in accordance with one or more example implementations.DETAILED DESCRIPTIONGeneral Overview

[0052] Methods and compositions provided herein improve the detection, diagnosis, staging, screening, treatment, and management of cancer. Example workflows of the methods disclosed herein are shown in Figures 2-5.

[0053] FIG. 2 depicts an example workflow of a method 200 for preparing a composition of amplified DNA from a sample useful for monitoring and detecting a tumor status in a cancer patient, comprising steps described herein.

[0054] Step 205 may involve acquiring DNA sampling data. Acquiring DNA sampling data may involve isolating nucleic acids from a first biological sample of a subject. This sample can be obtained through a solid tumor biopsy or a liquid biopsy, such as blood or plasma. The isolated nucleic acids may be sequenced and compared to a germline mutation database to identify non-cancer-specific germline mutations. This comparison helps distinguish between inherited genetic variations and mutations specific to the patient's cancer. In some embodiments, whole genome sequencing (WGS) may be used to sequence the entire genome of a subject, providing a comprehensive view of the genetic alterations present in the sample. Acquired samples may be G64 samples, which are acquired using WGS target information by aggregating all 64-target information, as opposed to the top 3-4 risk targets only. In other embodiments. WES may be used to sequence the entire exome of a subject, as described in greater detail below.

[0055] In some embodiments, the method 200 can include obtaining multiple different target loci, such as on-target and off-target loci. For example, step 210 may involve obtaining a first set of target loci and a second set of target loci. The first set of target loci may correspond to the cfDNA in the sample. The first set of target loci may be, for example “on-target” positions or loci (“positions” and “loci” may be used herein interchangeably). As used herein, on-target loci may refer to tumor- specific target loci. The second set of target loci may be, for example, “off-target” positions or loci, referring to target loci other than the tumor-specific target loci. Once the subject’s cancer-specific mutations are identified, multiplex PCR can be performed to amplify a plurality of target loci from cell-free DNA isolated from a second biological sample. ThisN.062.WO.01process targets specific regions of the genome that span at least one cancer-specific mutation. The number of target loci can vary, typically ranging from 1 to 100, depending on the specific requirements of the analysis. The amplified DNA may then be used for further sequencing or analysis to monitor the presence and progression of cancer-specific mutations. This targeted approach allows for precise detection and monitoring of cancer-related genetic changes in the patient's DNA.

[0056] In some embodiments, the method 200 can include generating one or more feature sets for based on the target loci. For example, step 215 may involve generating a first feature set and a second feature set. Generating a first feature set may involve calculating a probability of positive mutations using an augmented variant caller. The variant caller provides a probability value for each target position for both the first set of target loci and the second set of target loci (e.g., both on-target and off-target positions). The probability value may be used to determine the likelihood of different numbers of true mutations. The resulting first feature set may include probabilities of mutations at various thresholds and mean mutation counts across different subsets of amplicons. Generating a second feature set may involve performing posterior probability feature transformations, such as top position posterior probabilities, number of position posterior probabilities per sample above certain thresholds, and the differences between the first and second sets of target loci transformed probabilities.

[0057] In various embodiments, “on-target” loci refer to tumor-specific target loci selected for a subject (e.g., patient- specific loci). “Off-target” or “additional” loci refer to other sequenced loci distinct from the tumor- specific target loci that are analyzed within the same workflow. Unless context indicates otherwise, “first set of target loci” corresponds to on-target loci and “second set of target loci” corresponds to additional loci away from the tumor- specific loci.

[0058] In various embodiments, references to validation of a ctDNA positivity classification with respect to a second set of target loci do not require independent or orthogonal biological confirmation. Rather, such references may encompass use of sequencing information from the second set of target loci as internal calibration or normalization inputs, background or error-related modeling features, predictor features, or other confidence-related inputs within the same analytical or machine-learning workflow used to generate the ctDNA positivity classification. This usage is distinct from validation of machine-learning models using training, validation, or test datasets, which is also described herein.N.062.WO.01

[0059] In some embodiments, sequence reads corresponding to the second set of target loci may be further processed to compute one or more quantitative error parameters, such as estimated background mutation rates, amplification noise metrics, read-level uncertainty values, and / or other numerical measures characterizing sequencing artifacts. One or more of these computed error parameters may be provided as numerical inputs to at least one of the machine-learning models or may be applied to adjust one or more classification thresholds used in generating the ctDNA positivity classification.

[0060] In some embodiments, one or more features included in the first feature set or the second feature set may be computed as functions of physical amplification or sequencing characteristics, such as read depth distributions, amplification efficiency estimates, locus-specific error profiles, and / or fragment-length-dependent effects observed in the sequence reads. Such features may reflect measured properties of the sequencing process itself and vary in response to assay conditions present in the biological sample.

[0061] In some embodiments, the method 200 can include feeding the one or more feature sets into one or more machine learning (ML) models. For example, step 220 may involve feeding the first feature set into a first ML model and a second feature set into a second ML model. Step 220 may further involve training and validating the first ML model using the first feature set and training and validating the second ML model using the second feature set.

[0062] In some embodiments, the method 200 can include ensembling the results from the one or more ML models. For example, step 225 may involve ensembling results from the first ML model and the second ML model to obtain an ensembled result. Ensembling results from the first and second ML models may involve combining the models by applying a geometric average to their probability-like outputs. In various embodiments, Platt scaling, geometric averaging, voting scheme, or gradient boosting may be used to adjust the probability-like outputs to better represent probabilities. The quality of this adjustment may be estimated using a Brier Score. The ensembled result may be represented as votes from each of the models to classify the sample as positive, wherein a positive classification may represent a cancerous sample.

[0063] Step 230 may involve generating a ctDNA positivity classification based on the ensembled result. A sample may be classified as positive if it receives more than a predetermined number of votes out of 10, based on specificity / sensitivity trade-offs. In various embodiments, a “ctDNA positivity classification” may refer to a binary or categorical determination for a sampleN.062.WO.01(e.g., “positive” or “negative” for ctDNA) indicating whether circulating tumor DNA is detected at or above a defined decision threshold. The classification may be generated from one or more quantitative outputs (e.g., probability-like scores or risk scores) of the described machine learning models and / or variant callers by applying one or more thresholds as described herein.

[0064] In certain embodiments, the first and second sets of target loci are captured and amplified together in a single multiplex reaction, such that the off-target loci provide internal calibration and background modeling information specific to the exact assay conditions and sample, rather than relying on external control runs.

[0065] FIG. 2 depicts an example workflow of a method for preparing a composition of amplified DNA from a sample useful for monitoring and detecting a tumor status in a cancer patient, comprising steps described herein.

[0066] FIG. 3 depicts an example workflow of a method 300 for obtaining a first set of target loci and a second set of target loci, comprising steps described herein.

[0067] Step 305 may involve acquiring sampling data. Sampling data may be acquired from one or more databases. Sampling data may be obtained from a subject through a solid tumor biopsy or a liquid biopsy, such as blood or plasma.

[0068] Step 310 may involve extracting DNA from the sampling data, or a fraction thereof. Step 310 may further involve preparing a composition of amplified DNA by performing a multiplex amplification reaction on the extracted cell-free DNA or DNA derived from it. This process results in a set of amplicons encompassing a first set of target loci.

[0069] Step 315 may involve performing high-throughput sequencing on at least some of the amplicons to obtain sequence reads of at least a fraction of the set of amplicons encompassing the first set of target loci. Additionally, a second set of target loci, different from the first set, may also be sequenced. The sequence reads from the second set of target loci may be used to calibrate, predict, and / or validate the ctDNA positivity classification.

[0070] FIG. 4 depicts an example workflow of a method 400 for training a first ML model and a second ML model, comprising steps described herein. Step 405 may involve generating a first feature set and a second feature set using the steps described in 215. Step 410 may involve training and validating a first model on the first feature set. Training and validating the first ML model may involve defining the classification problem as a binary call on the proposition " Sample is Positive," and implementing the classification with a plurality of models trained forN.062.WO.01each of a plurality of training and validation splits. Hyperparameters, such as max_depth, eta, and gamma, may be set and kept constant across models to ensure consistency. To address uncertainty in sample labeling, models may be trained with sample weights, and monotonicity constraints may be applied to enhance model robustness and prevent overfitting. Step 410 may further involve training and validating the second model on the second feature set. Training and validating the second ML model may involve refining processes to optimize cross-validation performance. The second ML model may be trained on an augmented training set with hyperparameters selected through randomized search, focusing on sensitivity while maintaining a high level of specificity (e.g., at least 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, or another suitable level). Monotonicity constraints and early stopping (e.g., stopping after 200 rounds) may be implemented to prevent overfitting, along with balanced trees for regularization. The validation set performance may be evaluated to set the classification threshold, and the train / validation split procedure may be repeated multiple times to ensure robustness.

[0071] FIG. 5 depicts an example workflow of a method 500 for generating training data.Generating a balanced augmented training data set may comprise the steps described herein.

[0072] Step 505 may involve obtaining a set of samples. Step 510 may involve stratifying the samples into DNA input groups. This may involve, for example, stratifying the underlying 16-plex samples by DNA input into a number of bins. This stratification helps ensure that the augmented groups corresponding to these DNA input bins are obtained, maintaining diversity in the DNA input groups. Step 515 may involve repeatedly selecting and combining a plurality of samples from the set to generate augmented samples for each DNA input group. The augmented dataset may comprise, for example, 5000 negative samples and 5000 positive samples.Sample Collection

[0073] The methods disclosed herein are contemplated to be used to monitor or detect a wide variety of cancers in a patient. A person of ordinary skill in the art would understand that different types of cancer will require collection of different type of samples as described herein.

[0074] In some embodiments, the cancer is a solid tumor, and the first biological sample is a tumor biopsy sample. Performing a biopsy generally involves using a sharp tool to remove a small amount of tissue from the area suspected to contain diseased cells or tissue, such as a tumor. There are many different types of biopsies such as needle biopsy, CT-guided biopsy, ultrasound guided biopsy, bone biopsy, bone marrow biopsy, liver biopsy, kidney biopsy, aspiration biopsy, prostateN.062.WO.01biopsy, skin biopsy, surgical biopsy such as laparoscopic biopsy. In some embodiments, the first biological sample is obtained by liquid biopsy. In some embodiments, the first biological sample is a blood, serum, plasma, or urine sample. Further, biological liquid samples may be extracted from variety of animal fluids containing cell free DNA, including but not limited to blood, serum, plasma, bone marrow, urine vitreous, sputum, tears, perspiration, saliva, semen, mucosal excretions, mucus, spinal fluid, amniotic fluid, lymph fluid and so on. Cell free DNA may be fetal in origin (via fluid taken from a pregnant subject), or may be derived from tissue of the subject itself.

[0075] In some embodiments, the cancer is a blood cancer, and the first biological sample is a liquid sample. In some embodiments, the cancer is a blood cancer, and the first biological sample is blood, serum, plasma, or bone marrow sample. In some embodiments, the DNA from the cancer and the matched normal DNA are both obtained from the blood sample by isolating and separating plasma and buffy coat. The DNA obtained from the buffy coat may serve as the matched normal DNA to the circulating tumor DNA obtained from the plasma fraction.

[0076] In some embodiments, the methods of the present disclosure further comprise longitudinally collecting a plurality of second biological samples from the patient and repeating steps (b)-(c) for each of the second biological samples.

[0077] In some embodiments, the second biological sample is obtained from the patient after the patient has been treated for the cancer. In some embodiments, the second biological sample is a liquid sample. In some embodiments, the second biological sample is blood, serum, plasma, or urine sample.

[0078] Methods provided herein, in certain embodiments, are specially adapted for amplifying DNA fragments, especially tumor DNA fragments that are found in circulating tumor DNA (ctDNA). Such fragments may be, for example, about 160 nucleotides in length.

[0079] It is known in the art that cell-free nucleic acid (cfNA), e.g cfDNA, can be released into the circulation via various forms of cell death such as apoptosis, necrosis, autophagy and necroptosis. The cfDNA, is fragmented and the size distribution of the fragments varies from 150-350 bp to > 10000 bp. (see Kalnina et al. World J Gastroenterol. 2015 Nov 7; 21(41): 11636— 11653). For example the size distributions of plasma DNA fragments in hepatocellular carcinoma (HCC) patients spanned a range of 100-220 bp in length with a peak in count frequency at aboutN.062.WO.01166bp and the highest tumor DNA concentration in fragments of 150-180 bp in length (see: Jiang et al. Proc Natl Acad Sci USA 112: E1317-E1325).

[0080] In an illustrative embodiment the ctDNA is isolated from blood using EDTA-2Na tube after removal of cellular debris and platelets by centrifugation. The plasma samples can be stored at -80oC until the DNA is extracted using, for example, QIAamp DNA Mini Kit (Qiagen, Hilden, Germany), (e.g. Hamakawa et al.. Br J Cancer. 2015; 112:352-356). Hamakava et al. reported median concentration of extracted cell free DNA of all samples 43.1 ng per ml plasma (range 9.5-1338 ng ml / ) and a mutant fraction range of 0.001-77.8%, with a median of 0.90%.

[0081] In certain illustrative embodiments the sample is a tumor. Methods are known in the art for isolating nucleic acid from a tumor and for creating a nucleic acid library from such a DNA sample given the teachings here. Furthermore, given the teachings herein, a skilled artisan will recognize how to create a nucleic acid library appropriate for the methods herein from other samples such as other liquid samples where the DNA is free floating in addition to ctDNA samples.Identification of Cancer-Specific Mutations

[0082] After collecting the samples, targeted sequencing or WES may be performed on the circulating tumor DNA, cell free DNA or cellular DNA obtained from the solid tumor or the liquid biopsy samples, and the matched normal tissue or cells as described above according to the type of cancer being analyzed. Comparing sequences from tumor or cancer cells with the sequences from normal tissue or cells allows identification of cancer-specific mutations. Following identification of cancer-specific mutations personalized for a patient, the cancer in the patient may be detected or monitored by using the personalized cancer-specific mutations. The detection of the personalized cancer- specific mutations before, during, and after cancer treatment may be indicative of relapse, recurrence, or metastasis of the cancer.

[0083] In some embodiments, the cancer-specific mutations comprise one or more somatic mutations. Somatic mutations may be distinguished from germline mutations for example by sequencing nucleic acids isolated from non-cancer cells of the patient to identify one or more non-cancer-specific germline mutations, wherein the nucleic acids have been enriched at the panel of cancer-associated genomic loci. In some embodiments, the non-cancer cells are obtained from buffy coat in a blood sample of the patient. Germline mutations may be filtered out by first running a large number of targets selected for a first patient specific assay on the non-cancer DNA obtained from the buffy coat, and then select cancer specific variants for a second patient specific assay.N.062.WO.01

[0084] In some embodiments, the methods of the present disclosure further comprise comparing the sequences of the amplified DNA prepared from two longitudinally collected second biological samples to identify one or more non-cancer-specific germline mutations. Germline mutations will have variant allele frequency (VAF) of about 50% in sequential biological samples. In some embodiments, wherein the levels of ctDNA are very high, the copy number of the regions of the variants may have to be considered for determining germline mutations and filter them out.

[0085] In some embodiments, germline mutations may be determined by separating cell free DNA from plasma samples into long and short DNA fractions and analyze both fractions with the bespoke (personalized or patient- specific) assay. Tumor specific variant are expected to have higher variant allele frequency in the sample with shorter DNA fractions. Alternatively, in some embodiments, the shorter fragments may be enriched and the germline mutations can be identified by comparing variant allele frequency for the mutations in the enriched sample with the original sample.

[0086] In some embodiments, the methods of the present disclosure further comprise comparing the sequences of the nucleic acids isolated from the first biological sample to a germline mutation database to identify one or more non-cancer-specific germline mutations.

[0087] Upon identification of the patient’s cancer specific mutations, multiplex PCR is performed to amplify a plurality of target loci form cell-free DNA isolated from a second biological sample of the patient to obtain amplified DNA, In some embodiments, the multiplex amplification targets 1-100 target loci, or 1-20 target loci, or 1-10 target loci, or 10-20 target loci, or 20-50 target loci, each spanning at least one cancer- specific mutation. In some embodiments, the multiplex amplification targets 1, 2, 3, 4, 5. 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16. 17. 18, 19. or 20 target loci spanning at least one cancer-specific mutation.

[0088] In one aspect, the cancer- specific mutations are identified by performing WES on the DNA obtained from liquid samples or solid tumor samples and compared to WES of normal tissue. In some embodiments, WES is performed on cellular DNA obtained from a solid tumor and from matched normal tissue. In some embodiments, WES is performed on cell free DNA from a liquid biopsy sample such as blood or plasma. In some embodiments, WES is performed on cell free or cellular DNA obtained from a blood sample from a patient suffering from a blood cancer to identify cancer specific blood cancer mutations. By comparing sequencing data of DNA obtained from blood cancer or solid tumors with DNA obtained from normal matched tissue, the cancer specificN.062.WO.01mutations may be identified and used to monitor or detect the cancer during the clinical progression of the patient’ s cancer.

[0089] “Whole exome sequencing,” as used herein, refers to sequencing of all protein coding regions of genes in a genome, also known as exomes. Accordingly, WES may first involve a step of isolating a subset of DNA encoding protein that are known as exons before sequencing. This first step may be performed by capture techniques to isolated exons, i.e. array based capture or insolution capture as described elsewhere herein.

[0090] In another aspect, the cancer specific mutations are identified by targeted sequencing of nucleic acids derived from biological samples obtained from the patient. The biological samples may be obtained by solid tumor biopsy or by liquid biopsy as described above. The cancerous nucleic acids may be cellular DNA obtained from the solid tumor, cell free or circulating DNA obtained from any liquid sample as described above, or the cancerous DNA may be cell-free DNA or cellular DNA obtained from a blood sample of a patient suffering from blood cancer. The normal matched DNA may be cellular DNA obtained from non cancerous cells or tissue from the patient.

[0091] In some embodiments of the present disclosure, the targeted sequencing is performed by enriching the nucleic acids obtained from the patient at a panel of cancer associated genes or genomic loci to reduce the number of target loci or nucleic acid bases necessary for identification of patient specific tumor or cancer cell mutations. In some embodiments, the targeted sequencing comprises enriching the nucleic acids (e.g., cellular DNA) obtained from a solid tumor biopsy sample of the patient at a panel of cancer associated genes (e.g., FoundationOne™ panel from Foundation Medicine). In some embodiments, the targeted sequencing is performed by enriching the nucleic acids (e.g., cfDNA) obtained from a blood, plasma, serum, or urine sample of the patient at a panel of cancer associated genes (e.g., Guardant360™ panel from Guardant Health).

[0092] In some embodiments, the panel comprises 2,000 or less cancer-associated genes or genomic loci, or 1,000 or less cancer-associated genes or genomic loci, or 500 or less cancer-associated genes or genomic loci, or 100-1,000 cancer- associated genes or genomic loci, or 200-500 cancer-associated genes or genomic loci. In some embodiments, the panel comprises from about 100 to about 300 cancer- associated genes or genomic loci, from about 300 to about 450 cancer-associated genes or genomic loci from about 200 to about 350 cancer-associated genes or genomic loci from about 500 to about 1000 genes or cancer- associated genes or genomic loci fromN.062.WO.01about 1000 to about 1500 cancer-associated genes or genomic loci from about 1500 to about 2000 cancer-associated genes or genomic loci from about 1650 to about 2000 cancer-associated genes or genomic loci. In some embodiments, the panel comprises from about 100, 150. 200, 250, 300, 350, 400, 450, 500, 750, 1000, 1500, 1850, or 2000 cancer-associated genes or genomic loci.

[0093] In some embodiments, the sequencing of the nucleic acids isolated from the first biological sample obtained from the patient produces 5,000.000 bases or less of DNA sequences, or 4,000,000 bases or less of DNA sequences, or 3,000,000 bases or less of DNA sequences, or 2,000,000 bases or less of DNA sequences, or 500,000-2,000,000 bases of DNA sequences, or 1,000,000-1,500,000 bases of DNA sequences. As used herein, the term “cancer associated genomic loci” refers to any genomic loci determined to be useful for monitoring or detecting a cancer in a patient. The cancer associated genomic loci may be associated with (i) the metastatic potential of the cancer, potential to metastasize to specific organs, risk of recurrence, and / or course of the tumor; (ii) the tumor stage; (iii) the patient prognosis in the absence of treatment of the cancer; (iv) the prognosis of patient response (e.g., tumor shrinkage or progression- free survival) to treatment (e.g., chemotherapy, radiation therapy, surgery to excise tumor, etc.); (v) diagnosis of actual patient response to current and / or past treatment; (vi) determining a preferred course of treatment for the patient; (vii) prognosis for patient relapse after treatment (either treatment in general or some particular treatment); (viii) prognosis of patient life expectancy (e.g., prognosis for overall survival), etc.

[0094] Accordingly, in some embodiments, cancer associated genomic loci accompanies rapidly proliferating (and thus more aggressive) cancer cells. Such a cancer in a patient will often mean the patient has an increased likelihood of recurrence after treatment (e.g., the cancer cells not killed or removed by the treatment will quickly grow back). Such a cancer can also mean the patient has an increased likelihood of cancer progression for more rapid progression (e.g., the rapidly proliferating cells will cause any tumor to grow quickly, gain in virulence, and / or metastasize). Such a cancer can also mean the patient may require a relatively more aggressive treatment. Thus, in some embodiments the disclosure provides a method of classifying cancer comprising determining the status of a panel of genes comprising at least two or more cancer associated genomic loci, wherein an abnormal status indicates an increased likelihood of recurrence or progression.N.062.WO.01

[0095] In some embodiments, the panel of cancer-associated genomic loci comprises exons, introns, gene regulatory regions, non-coding RNA, rearranged genes. In some embodiments, the cancer-specific mutations comprise one or more single nucleotide variants (SNVs). one or more multi-nucleotide variants (MNVs), one or more copy number variants (CNVs), one or more indels, one or more gene fusions, one or more structural variants, or a combination thereof.

[0096] In some embodiments, the panel of cancer-associated genomic loci comprises any genomic alterations of any size from changes in single nucleotides to changes in genomic regions larger than 1 kilo base (kb). The term “indel” refers to both insertion and deletion of nucleic acids in the genome. As used herein, the term “structural variant” refers to a genomic alteration such as deletions or insertions that involve DNA segments larger than 1 kilo base (kb), and could be either microscopic or submicroscopic. The term “gene fusions” refers to any genomic alteration resulting in the fusion of two different genomic loci caused by insertions and / or deletions of DNA in the genome. The resulting genomic alteration caused by gene fusion may involve a DNA segment of any size.

[0097] A non-coding RNA (ncRNA) is a functional RNA molecule that is transcribed from DNA but not translated into proteins. Epigenetically related ncRNAs include miRNA, siRNA, piRNA and IncRNA. In general, ncRNAs function to regulate gene expression at the transcriptional and post-transcriptional level. Those ncRNAs that appear to be involved in epigenetic processes can be divided into two main groups; the short ncRNAs (<30 nts) and the long ncRNAs (>200 nts). The three major classes of short non-coding RNAs are microRNAs (miRNAs), short interfering RNAs (siRNAs), and piwi-interacting RNAs (piRNAs). Both major groups are shown to play a role in heterochromatin formation, histone modification, DNA methylation targeting, and gene silencing.

[0098] In some embodiments, the panel of cancer associated genomic loci comprises a list or set of well-known cancer genes, oncogenes, or any genes reported altered in cancerous cells or tumor tissue. A cancer-associated gene refers to a gene associated with an altered risk for a cancer (e.g. breast cancer, bladder cancer, or colorectal cancer) or an altered prognosis for a cancer. Example cancer-related genes that promote cancer include oncogenes; genes that enhance cell proliferation, invasion, or metastasis; genes that inhibit apoptosis; and pro-angiogenesis genes. Cancer-related genes that inhibit cancer include, but are not limited to, tumor suppressor genes; genes that inhibitN.062.WO.01cell proliferation, invasion, or metastasis; genes that promote apoptosis; and anti-angiogenesis genes.

[0099] In some embodiments, cancer-associated genomic loci of the panel may comprise AKT1 (14q32.33, ALK (2p23.2-23.1), APC (5q22.2), AR (Xql2), ARAF (Xpll.3), ARID1A (lp36.11), ATM (llq22.3), BRAF (7q34), BRCA1 (17q21.31), BRCA2 (13ql3.1), CCND1 (llql3.3), CCND2 (12pl3.32), CCNE1 (19ql2). CDH1 (16q22.1), CDK4 (12ql4.1), CDK6 (7q21.2), CDKN2A (9p21.3), CTNNB1 (3p22.1), DDR2 (lq23.3), EGFR (7pl 1.2), ERBB2 (17ql2), ESRI (6q25.1-25.2), EZH2 (7q36.1), FBXW7 (4q31.3). FGFR1 (8pl 1.23), FGFR2 (10q26.13), FGFR3 (4pl6.3), GATA3 (10pl4), GNA11 (19pl3.3), GNAQ (9q21.2), GNAS (20ql3.32), HNF1A (12q24.31), HRAS (1 lpl5.5), IDH1 (2q34), IDH2 (15q26.1), JAK2 (9p24.1), JAK3 (19pl3.11), KIT (4ql2), KRAS (12pl2.1), MAP2K1 (15q22.31), MAP2K2 (19pl3.3), MAPK1 (22qll.22), MAPK3 (16pll.2), MET (7q31.2), MLH1 (3p22.2), MPL (lp34.2), MTOR (lp36.22), MYC (8q24.21), NF1 (17qll.2), NFE2L2 (2q31.2), NOTCH1 (9q34.3), NPM1 (5q35.1), NRAS (lpl3.2), NTRK1 (lq23.1), NTRK3 (15q25.3). PDGFRA (4ql2), PIK3CA (3q26.32), PTEN (10q23.31), PTPN11 (12q24.13), RAF1 (3p25.2), RBI (13ql4.2), RET (10qll.21), RHEB (7q36.1), RHOA (3p21.31), RIT1 (lq22), ROS1 (6q22.1), SMAD4 (18q21.2), SMO (7q32.1), STK11 (19pl3.3), TERT (5pl5.33), TP53 (17pl3.1), TSC1 (9q34.13), and / or VHL (3p25.3). An embodiment of the mutation detection method begins with the selection of the region of the gene that becomes the target. The region with known mutations is used to develop primers for mPCR-NGS to amplify and detect the mutation.

[0100] Methods provided herein can be used to detect virtually any type of mutation, especially mutations known to be associated with cancer and most particularly the methods provided herein are directed to mutations, especially SNVs, CNVs, indels, or gene fusions or rearrangement, associated with cancer. Example SNVs can be in one or more of the following genes: EGFR, FGFR1, FGFR2, ALK. MET. ROS1 NTRK1, RET. HER2, DDR2. PDGFRA, KRAS. NF1, BRAF, PIK3CA, MEK1, NOTCH1, MLL2, EZH2, TET2, DNMT3A, SOX2, MYC, KEAP1, CDKN2A, NRG1, TP53, LKB1, and PTEN, which have been identified in various lung cancer samples as being mutated, having increased copy numbers, or being fused to other genes and combinations thereof (Non-small-cell lung cancers: a heterogeneous set of diseases. Chen et al. Nat. Rev. Cancer. 2014 Aug 14(8):535-551). In another example, the list of genes are those listed above, where SNVs have been reported, such as in the cited Chen et al. reference.N.062.WO.01

[0101] Example embodiments of potential cancer associated genomic loci include exonic regions of the following genes (e.g., for the detection of SNVs, CNVs, and indels): ABL1 ACVR1B AKT1 AKT2 AKT3 ALK ALOX12B AMER1 (FAM123B) APC AR ARAF ARFRP1 ARID1A ASXL1 ATM ATR ATRX AURKA AURKB AXIN1 AXL BAP1 BARD1 BCL2 BCL2L1 BCL2L2 BCL6 BCOR BCORL1 BRAF BRCA1 BRCA2 BRD4 BRIP1 BTG1 BTG2 BTK Cllorf30 (EMSY) CALR CARD11 CASP8 CBFB CBL CCND1 CCND2 CCND3 CCNE1 CD22 CD274 (PD-L1) CD70 CD79A CD79B CDC73 CDH1 CDK12 CDK4 CDK6 CDK8 CDKN1A CDKN1B CDKN2A CDKN2B CDKN2C CEBPA CHEK1 CHEK2 CIC CREBBP CRKL CSF1R CSF3R CTCF CTNNA1 CTNNB1 CUL3 CUL4A CXCR4 CYP17A1 DAXX DDR1 DDR2 DIS3 DNMT3A DOT1L EED EGFR EP300 EPHA3 EPHB1 EPHB4 ERBB2 ERBB3 ERBB4 ERCC4 ERG ERRFI1 ESRI EZH2 FAM46C FANCA FANCC FANCG FANCL FAS FBXW7 FGF10 FGF12 FGF14 FGF19 FGF23 FGF3 FGF4 FGF6 FGFR1 FGFR2 FGFR3 FGFR4 FH FLCN FLT1 FLT3 FOXL2 FUBP1 GABRA6 GATA3 GATA4 GATA6 GID4 (C17orf39) GNA11 GNA13 GNAQ GNAS GRM3 GSK3B H3F3A HDAC1 HGF HNF1A HRAS HSD3B1 ID3 IDH1 IDH2 IGF1R IKBKE IKZF1 INPP4B IRF2 IRF4 IRS2 JAK1 JAK2 JAK3 JUN KDM5A KDM5C KDM6A KDR KEAP1 KEL KIT KLHL6 KMT2A (MLL) KMT2D (MLL2) KRAS LTK LYN MAF MAP2K1 (MEK1) MAP2K2 (MEK2) MAP2K4 MAP3K1 MAP3K13 MAPK1 MCL1 MDM2 MDM4 MED 12 MEF2B MEN1 MERTK MET MITF MKNK1 MLH1 MPL MRE11A MSH2 MSH3 MSH6 MST1R MTAP MTOR MUTYH MYC MYCL (MYCL1) MYCN MYD88 NBN NF1 NF2 NFE2L2 NFKBIA NKX2-1 NOTCH1 NOTCH2 NOTCH3 NPM1 NRAS NT5C2 NTRK1 NTRK2 NTRK3 P2RY8 PALB2 PARK2 PARP1 PARP2 PARP3 PAX5 PBRM1 PDCD1 (PD-1) PDCD1LG2 (PD-L2) PDGFRA PDGFRB PDK1 PIK3C2B PIK3C2G PIK3CA PIK3CB PIK3R1 PIM1 PMS2 POLD1 POLE PPARG PPP2R1A PPP2R2A PRDM1 PRKAR1A PRKCI PTCHI PTEN PTPN11 PTPRO QKI RAC1 RAD21 RAD51 RAD51B RAD51C RAD51D RAD52 RAD54L RAF1 RARA RBI RBM10 REL RET RICTOR RNF43 ROS1 RPTOR SDHA SDHB SDHC SDHD SETD2 SF3B1 SGK1 SMAD2 SMAD4 SMARCA4 SMARCB1 SMO SNCAIP SOCS1 SOX2 SOX9 SPEN SPOP SRC STAG2 STAT3 STK11 SUFU SYK TBX3 TEK TET2 TGFBR2 TIPARP TNFAIP3 TNFRSF14 TP53 TSC1 TSC2 TYRO3 U2AF1 VEGFA VHL WHSCI (MMSET) WHSC1L1 WT1 XPO1 XRCC2 ZNF217 ZNF703. Example embodiments of potential cancer associated genomic loci also include intronic regions, promoter regions, and noncoding RNA sequences of the following genes (e.g., for the detection of gene fusion orN.062.WO.01rearrangement): ALK BCL2 BCR BRAF BRCA1 BRCA2 CD74 EGFR ETV4 ETV5 ETV6 EWSR1 EZR FGFR1 FGFR2 FGFR3 KIT KMT2A (MLL) MSH2 MYB MYC NOTCH2 NTRK1 NTRK2 NUTM1 PDGFRA RAF1 RARA RET ROS1 RSPO2 SDC4 SLC34A2 TERC TERT TMPRSS2.Methods of Enriching for Nucleic Acids Using a Panel of Cancer-Associated Genes or Isolating Exonic Genomic DNA for WES

[0102] Target-enrichment methods allow one to selectively capture genomic regions of interest from a DNA sample prior to sequencing by enrichment methods such as hybrid capture or targeted PCR. The genomic regions of interests may be any subset of genomic loci such as cancer associated genomi loci described above, or all the exonic regions of the genome to prepare samples for WES.

[0103] In general, hybrid capture involves designing oligonucleotide sequences capable of binding by complementarity to genomic DNA sequences of interest. The oligonucleotides are bound to a solid surface or beads that will allow separating genomic sequences bound to the oligonucleotides from the unbound genomic sequences. The unbound genomic DNA sequences may then be washed away, and the genomic sequences of interest remain bound to solid surface or bead for further processing and / or amplification. In some embodiments, the panel of cancer-associated genomic loci are enriched by hybrid capture such as an array-based hybrid capture method or an in-solution hybrid capture methods.

[0104] In some embodiments, target enrichment may be an array-based hybrid capture method. In some embodiments, an array-based hybrid capture method may involve designing microarrays by fixing single-stranded oligonucleotide sequences from the human genome to tile the region of interest fixed to the surface of a microarray chip or surface. Genomic DNA is sheared to form double-stranded fragments. The fragments undergo end-repair to produce blunt ends and adaptors with universal priming sequences are added. These fragments are hybridized to oligos on the microarray chip or surface. Unhybridized fragments are washed away and the desired fragments are eluted. The fragments are then amplified using polymerase chain reaction. Microarrays to be used for array-based hybrid capture may be the Roche Nimblegen™ arrays, or the Agilent™ Capture Array, or similar comparative genomic hybridization array that can by used for hybrid capture of target sequences. In some embodiments, the panel of cancer-associated genomic loci are enriched by hybrid capture. In other embodiments, the target enrichment strategy may be anN.062.WO.01in-solution capture strategy. To capture genomic regions of interest using in-solution capture, a pool of custom oligonucleotides (probes) is synthesized and hybridized in solution to a fragmented genomic DNA sample. The probes (labeled with beads) selectively hybridize to the genomic regions of interest after which the beads (now including the DNA fragments of interest) can be pulled down and washed to clear excess material. The beads are then removed and the genomic fragments can be sequenced allowing for selective DNA sequencing of genomic regions (e.g., exons, introns, promoter regions or other gene regulatory regions, or non-coding RNA sequences) of interest.

[0105] In solution capture as opposed to hybrid capture, there is an excess of probes to target regions of interest over the amount of template required. The optimal target size is about 3.5 megabases and yields excellent sequence coverage of the target regions. The preferred method is dependent on several factors including: number of base pairs in the region of interest, demands for reads on target, equipment in house, etc.

[0106] Alternatively, the cancer-associated genomic loci can be enriched by targeted amplification. Targeted amplification of genomic loci may be achieved with multiplex PCR performed with primers designed to target specific regions. Protocols for performing multiplex PCR of a plurality of desired targets are described in detail elsewhere herein.Cancers

[0107] The terms "cancer" and "cancerous" refer to or describe the physiological condition in animals that is typically characterized by unregulated cell growth. A "tumor" comprises one or more cancerous cells. There are several main types of cancer. Carcinoma is a cancer that begins in the skin or in tissues that line or cover internal organs. Sarcoma is a cancer that begins in bone, cartilage, fat, muscle, blood vessels, or other connective or supportive tissue. Leukemia is a cancer that starts in blood-forming tissue, such as the bone marrow, and causes large numbers of abnormal blood cells to be produced and enter the blood. Lymphoma and multiple myeloma are cancers that begin in the cells of the immune system. Central nervous system cancers are cancers that begin in the tissues of the brain and spinal cord.

[0108] In some embodiments, the cancer is a cancer or tumor of abdomen or abdominal wall, adrenal gland, anus, appendix, bladder, bone, brain, breast, cervix, chest wall, colon, diaphragm, duodenum, ear, endometrium, esophagus, fallopian tube, gallbladder, gastro-esophageal junction, head and neck, kidney, larynx, liver, lung, lymph node, malignant effusions, mediastinum, nasalN.062.WO.01cavity, omentum, ovarian, pancreas, pancreatobiliary, parotid gland, pelvis, penis, pericardium, peritoneum, pleura, prostate, rectum, salivary gland, skin, small intestine, soft tissue, spleen, stomach, thyroid, tongue, trachea, ureter, uterus, vagina, vulva, or whippie resection.

[0109] In some embodiments, the cancer is lung cancer, breast cancer, bladder cancer, or colorectal cancer.

[0110] In some embodiments, the cancer comprises an acute lymphoblastic leukemia; acute myeloid leukemia; adrenocortical carcinoma; AIDS-related cancers; AIDS-related lymphoma; anal cancer; appendix cancer; astrocytomas; atypical teratoid / rhabdoid tumor; basal cell carcinoma; bladder cancer; brain stem glioma; brain tumor (including brain stem glioma, central nervous system atypical teratoid / rhabdoid tumor, central nervous system embryonal tumors, astrocytomas, craniopharyngioma, ependymoblastoma, ependymoma, medulloblastoma, medulloepithelioma, pineal parenchymal tumors of intermediate differentiation, supratentorial primitive neuroectodermal tumors and pineoblastoma); breast cancer; bronchial tumors; Burkitt lymphoma; cancer of unknown primary site; carcinoid tumor; carcinoma of unknown primary site; central nervous system atypical teratoid / rhabdoid tumor; central nervous system embryonal tumors; cervical cancer; childhood cancers; chordoma; chronic lymphocytic leukemia; chronic myelogenous leukemia; chronic myeloproliferative disorders; colon cancer; colorectal cancer; craniopharyngioma; cutaneous T-cell lymphoma; endocrine pancreas islet cell tumors; endometrial cancer; ependymoblastoma; ependymoma; esophageal cancer; esthesioneuroblastoma; Ewing sarcoma; extracranial germ cell tumor; extragonadal germ cell tumor; extrahepatic bile duct cancer; gallbladder cancer; gastric (stomach) cancer; gastrointestinal carcinoid tumor; gastrointestinal stromal cell tumor; gastrointestinal stromal tumor (GIST); gestational trophoblastic tumor; glioma; hairy cell leukemia; head and neck cancer; heart cancer; Hodgkin lymphoma; hypopharyngeal cancer; intraocular melanoma; islet cell tumors; Kaposi sarcoma; kidney cancer; Langerhans cell histiocytosis; laryngeal cancer; lip cancer; liver cancer; malignant fibrous histiocytoma bone cancer; medulloblastoma; medulloepithelioma; melanoma; Merkel cell carcinoma; Merkel cell skin carcinoma; mesothelioma; metastatic squamous neck cancer with occult primary; mouth cancer; multiple endocrine neoplasia syndromes; multiple myeloma; multiple myeloma / plasma cell neoplasm; mycosis fungoides; myelodysplastic syndromes; myeloproliferative neoplasms; nasal cavity cancer; nasopharyngeal cancer; neuroblastoma; Non-Hodgkin lymphoma; nonmelanoma skin cancer; non- small cell lung cancer;N.062.WO.01oral cancer; oral cavity cancer; oropharyngeal cancer; osteosarcoma; other brain and spinal cord tumors; ovarian cancer; ovarian epithelial cancer; ovarian germ cell tumor; ovarian low malignant potential tumor; pancreatic cancer; papillomatosis; paranasal sinus cancer; parathyroid cancer; pelvic cancer; penile cancer; pharyngeal cancer; pineal parenchymal tumors of intermediate differentiation; pineoblastoma; pituitary tumor; plasma cell neoplasm / multiple myeloma; pleuropulmonary blastoma; primary central nervous system (CNS) lymphoma; primary hepatocellular liver cancer; prostate cancer; rectal cancer; renal cancer; renal cell (kidney) cancer; renal cell cancer; respiratory tract cancer; retinoblastoma; rhabdomyosarcoma; salivary gland cancer; Sezary syndrome; small cell lung cancer; small intestine cancer; soft tissue sarcoma; squamous cell carcinoma; squamous neck cancer; stomach (gastric) cancer; supratentorial primitive neuroectodermal tumors; T-cell lymphoma; testicular cancer; throat cancer; thymic carcinoma; thymoma; thyroid cancer; transitional cell cancer; transitional cell cancer of the renal pelvis and ureter; trophoblastic tumor; ureter cancer; urethral cancer; uterine cancer; uterine sarcoma; vaginal cancer; vulvar cancer; Waldenstrom macroglobulinemia; or Wilm's tumor.

[0111] In another embodiment, provided herein is a method for detecting cancer in a sample of blood or a fraction thereof from an individual, such as an individual suspected of having a cancer, that includes determining the single nucleotide variants present in a sample by determining the single nucleotide variants present in a ctDNA sample using a ctDNA SNV amplification / sequencing workflow provided herein. The presence of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 SNVs on the low end of the range, and 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, or 50 SNVs on the high end of the range, in the sample at the plurality of single nucleotide loci is indicative of the presence of cancer.

[0112] In another embodiment, provided herein is a method for detecting a clonal SNV in a tumor of an individual. The method includes performing for example a ctDNA amplification / sequencing workflow as provided herein in the working examples, and determining the variant allele frequency for each of the SNV loci based on the sequence of the plurality of copies of the series of amplicons. A higher relative allele frequency compared to the other single nucleotide variants of the plurality of single nucleotide variant loci is indicative of a clonal single nucleotide variant in the tumor. Variant allele frequencies are well known in the sequencing art.

[0113] In certain embodiments, the method further includes determining a treatment plan, therapy and / or administering a compound to the individual that targets the one or more clonal singleN.062.WO.01nucleotide variants. In certain examples, subclonal and / or other clonal SNVs are not targeted by therapy. Specific therapies and associated mutations are provided in other sections of this specification and are known in the art. Accordingly, in certain examples, the method further includes administering a compound to the individual, where the compound is known to be specifically effective in treating cancer having one or more of the determined single nucleotide variants.

[0114] In certain aspects of this embodiment, a variant allele frequency of greater than 0.25%, 0.5%, 0.75%, 1.0%, 5% or 10% is indicative a clonal single nucleotide variant.

[0115] In certain examples of this embodiment, the cancer is a stage la, lb, or 2a breast cancer, bladder cancer, or colorectal cancer. In certain examples of this embodiment, the cancer is a stage la or lb breast cancer, bladder cancer, or colorectal cancer. In certain examples of the embodiment, the individual is not subjected to surgery. In certain examples of the embodiment, the individual is not subjected to a biopsy.

[0116] In some examples of this embodiment, a clonal SNV is identified or further identified if other testing such as direct tumor testing suggest an on-test SNV is a clonal SNV, for any SNV on test that has a variable allele frequency greater than at least one quarter, one third, one half, or three quarters of the other single nucleotide variants that were determined.

[0117] In some embodiments, methods herein for detecting SNVs in ctDNA can be used instead of direct analysis of DNA from a tumor.

[0118] In certain examples of any of the method embodiments provided herein herein, before a targeted amplification is performed on ctDNA from an individual, data is provided on SNVs that are found in a tumor from the individual. Accordingly, in these embodiments, a SNV amplification / sequencing reaction is performed on one or more tumor samples from the individual. In this methods, the ctDNA SNV amplification / sequencing reaction provided herein is still advantageous because it provides a liquid biopsy of clonal and subclonal mutations. Furthermore, as provided herein, clonal mutations can be more unambiguously identified in an individual that has cancer, if a high VAF percentage, for example, more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10% VAF in a ctDNA sample from the individual is determined for an SNV.

[0119] In certain embodiment, method provided herein can be used to determine whether to isolate and analyze ctDNA from circulating free nucleic acids from an individual with cancer. First, it is determined whether the cancer is breast cancer, bladder cancer, or colorectal cancer. If the cancerN.062.WO.01is a breast cancer, bladder cancer, or colorectal cancer, circulating free nucleic acids are isolated from individual. The method in some examples, further includes determining the stage of the cancer.

[0120] In some methods, provided herein are inventive compositions and / or solid supports. A composition comprising circulating tumor nucleic acid fragments comprising a universal adapter, wherein the circulating tumor nucleic acids originated from breast cancer, bladder cancer, or colorectal cancer.

[0121] In some embodiments, provided herein is an inventive composition that includes circulating tumor nucleic acid fragments comprising a universal adapter, wherein the circulating tumor nucleic acids originated from a sample of blood or a fraction thereof, of an individual with cancer. These methods typically include formation of ctDNA fragment that include a universal adapter. Furthermore, such methods typically include the formation of a solid support especially a solid support for high throughput sequencing, that includes a plurality of clonal populations of nucleic acids, wherein the clonal populations comprise amplicons generated from a sample of circulating free nucleic acids, wherein the ctDNA. In illustrative embodiments based on the surprising results provided herein, the ctDNA originated from cancer.

[0122] Similarly, provided herein as an embodiment of the disclosure is a solid support comprising a plurality of clonal populations of nucleic acids, wherein the clonal populations comprise nucleic acid fragments generated from a sample of circulating free nucleic acids from a sample of blood or a fraction thereof, from an individual with cancer.

[0123] In certain embodiments, the nucleic acid fragments in different clonal populations comprise the same universal adapter. Such a composition is typically formed during a high throughput sequencing reaction in embodiments of methods of the present disclosure.

[0124] The clonal populations of nucleic acids can be derived from nucleic acid fragments from a set of samples from two or more individuals. In these embodiments, the nucleic acid fragments comprise one of a series of molecular barcodes corresponding to a sample in the set of samples.Analytical Methods SNV 1 and 2

[0125] Detailed analytical methods are provided herein as SNV Methods 1 and SNV Method 2 in the analytical section herein. Any of the methods provided herein can further include analytical steps provided herein. Accordingly, in certain examples, the methods for determining whether a single nucleotide variant is present in the sample, includes identifying a confidence value for eachN.062.WO.01allele determination at each of the set of single nucleotide variance loci, which can be based at least in part on a depth of read for the loci. The confidence limit can be set at least 75%, 80%, 85%, 90%, 95%. 96%, 96%, 98%, or 99%. The confidence limit can be set at different levels for different types of mutations.

[0126] The method can performed with a depth of read for the set of single nucleotide variance loci of at least 5, 10. 15. 20, 25, 50. 100, 150, 200, 250, 500, 1.000, 10,000. 25,000, 50.000, 100,000, 250,000, 500,000, or 1 million.

[0127] In certain embodiments, a method of any of the embodiments herein includes determining an efficiency and / or an error rate per cycle are determined for each amplification reaction of the multiplex amplification reaction of the single nucleotide variance loci. The efficiency and the error rate can then be used to determine whether a single nucleotide variant at the set of single variant loci is present in the sample. More detailed analytical steps provided in SNV Method 2 provided in the analytical method can be included as well, in certain embodiments.

[0128] In illustrative embodiments, of any of the methods herein the set of single nucleotide variance loci includes all of the single nucleotide variance loci identified in the TCGA and COSMIC data sets for cancer.

[0129] In certain embodiments of any of the methods herein the set of single nucleotide variant loci include 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 75, 100, 250, 500, 1000, 2500, 5000, or 10,000 single nucleotide variance loci known to be associated with cancer on the low end of the range, and, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 75, 100, 250, 500, 1000, 2500, 5000, 10,000, 20,000 and 25,000 on the high end of the range.PCR Methods

[0130] In any of the methods for detecting SNVs herein that include a ctDNA SNV amplification / sequencing workflow, improved amplification parameters for multiplex PCR can be employed. For example, wherein the amplification reaction is a PCR reaction and the annealing temperature is between 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10°C greater than the melting temperature on the low end of the range, and 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15° on the high end the range for at least 10, 20, 25, 30, 40, 50, 60, 70, 75, 80, 90, 95 or 100% the primers of the set of primers.

[0131] In certain embodiments, wherein the amplification reaction is a PCR reaction the length of the annealing step in the PCR reaction is between 10, 15, 20, 30, 45, and 60 minutes on the low end of the range, and 15, 20, 30, 45, 60, 120, 180, or 240 minutes on the high end of the range. InN.062.WO.01certain embodiments, the primer concentration in the amplification, such as the PCR reaction is between 1 and 10 nM. Furthermore, in example embodiments, the primers in the set of primers, are designed to minimize primer dimer formation.

[0132] Accordingly, in an example of any of the methods herein that include an amplification step, the amplification reaction is a PCR reaction, the annealing temperature is between 1 and 10 °C greater than the melting temperature of at least 90% of the primers of the set of primers, the length of the annealing step in the PCR reaction is between 15 and 60 minutes, the primer concentration in the amplification reaction is between 1 and 10 nM, and the primers in the set of primers, are designed to minimize primer dimer formation. In a further aspect of this example, the multiplex amplification reaction is performed under limiting primer conditions.Use in Diagnosing Cancer

[0133] In another embodiment, provided herein is a method for supporting a cancer diagnosis for an individual, such as an individual suspected of having cancer, from a sample of blood or a fraction thereof from the individual, that includes performing a DNA amplification / sequencing workflow as provided herein, to determine whether one or more single nucleotide variants are present in the plurality of single nucleotide variant loci. In this embodiment, the following elements, statements, guidelines or rules apply: the absence of a single nucleotide variant supports a diagnosis of stage la, lb, or 2a adenocarcinoma, the presence of a single nucleotide variant supports a diagnosis of squamous cell carcinoma or a stage 2b or 3a adenocarcinoma, and / or the presence of ten or more single nucleotide variants supports a diagnosis of squamous cell carcinoma or a stage 2b or 3 adenocarcinoma.

[0134] These results identify analysis using a ctDNA SNV amplification / sequencing workflow of lung ADC and SCC samples from an individual as a valuable method for identifying SNVs found in an ADC tumor, especially for stage 2b and 3a ADC tumors, and especially an SCC tumor at any stage.Use in Directing a Therapeutic Regimen

[0135] In certain embodiments, methods herein for detecting SNVs can be used to direct a therapeutic regimen. Therapies are available and under development that target specific mutations associated with ADC and SCC (Nature Review Cancer. 14:535-551 (2014). For example, detection of an EGFR mutation at L858R or T790M can be informative for selecting a therapy. Erlotinib,N.062.WO.01gefitinib, afatinib, AZD9291, CO- 1686, and HM61713 are current therapies approved in the U. S. or in clinical trials, that target specific EGFR mutations. In another example, a G12D, G12C, or G12V mutation in KRAS can be used to direct an individual to a therapy of a combination of Selumetinib plus docetaxel. As another example, a mutation of V600E in BRAF can be used to direct a subject to a treatment of Vemurafenib, dabrafenib, and trametinib.Library Preparation

[0136] Methods of the present disclosure in certain embodiments, typically include a step of generating and amplifying a nucleic acid library from the sample (i.e. library preparation). The nucleic acids from the sample during the library preparation step can have ligation adapters, often referred to as library tags or ligation adaptor tags (LTs), appended, where the ligation adapters contain a universal priming sequence, followed by a universal amplification. In an embodiment, this may be done using a standard protocol designed to create sequencing libraries after fragmentation. In an embodiment, the DNA sample can be blunt ended, and then an A can be added at the 3’ end. A Y-adaptor with a T-overhang can be added and ligated. In some embodiments, other sticky ends can be used other than an A or T overhang. In some embodiments, other adaptors can be added, for example looped ligation adaptors. In some embodiments, the adaptors may have tag designed for PCR amplification.The DNA Amplification / Sequencing Workflow for Monitoring or Detecting Cancer in a Patien t

[0137] A number of the embodiments provided herein, include detecting the cancer-specific mutations in a ctDNA, cfDNA, or cellular DNA sample. Such methods in illustrative embodiments, include an amplification step and a sequencing step (Sometimes referred to herein as a “ctDNA amplification / sequencing workflow). In an illustrative example, a DNA amplification / sequencing workflow can include generating a set of amplicons by performing a multiplex amplification reaction on nucleic acids isolated from a sample of blood or a fraction thereof from an individual, such as an individual suspected of having cancer, for example breast cancer, bladder cancer, or colorectal cancer, wherein each amplicon of the set of amplicons spans at least one cancer-associated genomic loci of a set of cancer-associated genomic loci, such as an SNV loci known to be associated with cancer; and determining the sequence of at least a segment of at each amplicon of the set of amplicons, wherein the segment comprises a cancer-associated genomic loci. In some embodiments, the cancer-associated genomic loci comprise a SNV, a CNV,N.062.WO.01an indel, a rearranged gene, or a variation in exon, intron, gene regulatory sequences, or non coding RNA sequences. Example DNA amplification / sequencing workflows in more detail can include forming an amplification reaction mixture by combining a polymerase, nucleotide triphosphates, nucleic acid fragments from a nucleic acid library generated from the sample, and a set of primers that each binds an effective distance from a single nucleotide variant loci, or a set of primer pairs that each span an effective region that includes a cancer-associated genomic locus. Then, subjecting the amplification reaction mixture to amplification conditions to generate a set of amplicons comprising at least one cancer-associated genomic locus of a set of cancer-associated genomic loci,; and determining the sequence of at least a segment of each amplicon of the set of amplicons, wherein the segment comprises a cancer-associated genomic locus.

[0138] The effective distance of binding of the primers can be within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 75, 100, 125, or 150 base pairs of a cancer-associated genomic locus. The effective range that a pair of primers spans typically includes a cancer-associated genomic locus and is typically 160 base pairs or less, and can be 150, 140, 130, 125, 100, 75, 50 or 25 base pairs or less. In other embodiments, the effective range that a pair of primers spans is 20, 25, 30, 40, 50, 60, 70, 75, 100, 110, 120, 125, 130, 140, or 150 nucleotides from a cancer-associated genomic locus on the low end of the range, and 25, 30, 40, 50, 60, 70, 75, 100, 110, 120, 125, 130, 140, or 150, 160, 170, 175, or 200 on the high end of the range.

[0139] Further details regarding methods of amplification that can be used in a ctDNA amplification / sequencing workflow to detect cancer-associated genomic loci for use in example methods of the disclosure are provided in other sections of this specification.SNV Calling Analytics

[0140] During performance of the methods provided herein, nucleic acid sequencing data is generated for amplicons created by the tiled multiplex PCR. Algorithm design tools are available that can be used and / or adapted to analyze this data to determine within certain confidence limits, whether a cancer-associated genomic locus, such as a SNV is present in a target gene known to be associated with cancer development, recurrence, metastasis, treatment response, or prognosis.

[0141] Sequencing reads can be demultiplexed using an in-house tool or a commercially available tool and mapped using a suitable aligner, such as the Burrows- Wheeler Aligner (BWA) software operating in the Maximal Exact Match (MEM) alignment mode, (BWA-MEM), (see Li H. and Durbin R. (2010) Fast and accurate long-read alignment with Burrows-Wheeler Transform.N.062.WO.01Bioinformatics, Epub. [PMID: 20080505]). In some embodiments, reads are aligned in single end mode, using, for example, PEAR-merged reads to a reference genome such as hgl9. Amplification statistics quality control (QC) can be performed by analyzing total reads, number of mapped reads, number of mapped reads on target, and number of reads counted.

[0142] In certain embodiments, any analytical method for detecting an SNV from nucleic acid sequencing data detection can be used with methods of the disclosure that include a step of detecting an SNV or determining whether an SNV is present. In certain illustrative embodiments, methods of the disclosure that utilize SNV METHOD 1 below are used. In other illustrative embodiments, methods of the disclosure that include a step of detecting an SNV or determining whether an SNV is present at an SNV loci, utilize SNV METHOD 2 below.Example SNV Method 1

[0143] For this embodiment, a background error model is constructed using normal plasma samples, which were sequenced on the same sequencing run to account for run-specific artifacts. In certain embodiments, 5, 10, 15, 20, 25, 30, 40, 50, 100, 150, 200, 250, or more than 250 normal plasma samples are analyzed on the same sequencing run. In certain illustrative embodiments, 20, 25, 40, or 50 normal plasma samples are analyzed on the same sequencing run. Noisy positions with normal median variant allele frequency greater than a cutoff are removed. For example this cutoff in certain embodiments is > 0.1%, 0.2%, 0.25%, 0.5%, 1%, 2%, 5%, or 10%. In certain illustrative embodiments noisy positions with normal medial variant allele frequency greater than 0.5% are removed. Outlier samples were iteratively removed from the model to account for noise and contamination. In certain embodiments, samples with a Z score of greater than 5, 6, 7. 8, 9, or 10 are removed from the data analysis. For each base substitution of every genomic loci, the depth of read weighted mean and standard deviation of the error are calculated. Tumor or cell-free plasma samples’ positions with at least 5 variant reads and a Z-score of 10 against the background error model for example, can be called as a candidate mutation.

[0144] SNV METHOD 2: For this embodiment SNVs are determined using plasma ctDNA data. The PCR process is modeled as a stochastic process, estimating the parameters using a training set and making the final SNV calls for a separate testing set. The propagation of the error across multiple PCR cycles is determined, and the mean and the variance of the background error are calculated, and in illustrative embodiments, background error is differentiated from real mutations.

[0145] The following parameters are estimated for each base:N.062.WO.01

[0146] p = efficiency (probability that each read is replicated in each cycle)

[0147] pe= error rate per cycle for mutation type e (probability that an error of type e occurs)

[0148] Xo = initial number of molecules

[0149] As a read is replicated over the course of PCR process, more errors occur. Hence, the error profile of the reads is determined by the degrees of separation from the original read. We refer to a read as kthgeneration if it has gone through k replications until it has been generated.

[0150] Let us define the following variables for each base:

[0151] Xij = number of generation i reads generated in the PCR cycle j

[0152] Yy = total number of generation i reads at the end of cycle j

[0153] Xif = number of generation i reads with mutation e generated in the PCR cycle j

[0154] Moreover, in addition to normal molecules Xo, if there are additional feXo molecules with the mutation e at the beginning of the PCR process (hence fe / (l+fe) will be the fraction of mutated molecules in the initial mixture).

[0155] Given the total number of generation i-1 reads at cycle j-1, the number of generation i reads generated at cycle j has a binomial distribution with a sample size of Yi-ij.i and probability parameter of p. Hence, E(Xy, \Yi-ij-i, p) - p Yi-ij-i and Var(Xy, VYi-ij-i, p)= p(l-p) Yi.ij.i.

[0156] We also have Y^ =Jk=i Xtk- Hence, by recursion, simulation or similar methods, we can determine E(Xy_). Similarly, we can determine Var(Xy) = E(Var(Xy, I p)) + Var(E(Xy, I / ?)) using the distribution of p.

[0157] finally, E(XijelYi.jj.j, pe) - peYi-ij-i and Var(X \Yi-ij-i, p)= pe(l-pe) Yi-ij-i, and we can use these to compute E(Xtf) and Var(X ).

[0158] In certain embodiments, SNV Method 2 is performed as follows:

[0159] a) Estimate a PCR efficiency and a per cycle error rate using a training data set;

[0160] b) Estimate a number of starting molecules for the testing data set at each base using the distribution of the efficiency estimated in step (a);

[0161] c) If needed, update the estimate of the efficiency for the testing data set using the starting number of molecules estimated in step (b);

[0162] d) Estimate the mean and variance for the total number of molecules, background error molecules and real mutation molecules (for a search space consisting of an initial percentage of real mutation molecules) using testing set data and parameters estimated in steps (a), (b) and (c);N.062.WO.01

[0163] e) Fit a distribution to the number of total error molecules (background error and real mutation) in the total molecules, and calculate the likelihood for each real mutation percentage in the search space; and

[0164] f) Determine the most likely real mutation percentage and calculate the confidence using the data from in step (e).

[0165] A confidence cutoff can be used to identify an SNV at an SNV loci. For example, a 90%, 95%, 96%, 97%, 98%, or 99% confidence cutoff can be used to call an SNV.Example SNV Method 2

[0166] The algorithm starts by estimating the efficiency and error rate per cycle using the training set. Let n denote the total number of PCR cycles.

[0167] The number of reads Rb at each base b can be approximated by ( 1 +pb) ” Xo, where pb is the efficiency at base b. Then (Rb / Xo)1 / ncan be used to approximate 1+pb. Then, we can determine the mean and the standard variation of pb across all training samples, to estimate the parameters of the probability distribution (such as normal, beta, or similar distributions) for each base.

[0168] Similarly the number of error e reads Rbeat each base b can be used to estimate pe. After determining the mean and the standard deviation of the error rate across all training samples, we approximate its probability distribution (such as normal, beta, or similar distributions) whose parameters are estimated using this mean and standard deviation values.

[0169] Next, for the testing data, we estimate the initial starting copy at each base as f1 R\nf(Pb)dpbwhere f(.) is an estimated distribution from the training set.

[0170] fiRb~.nf(Pb)dpbwhere f(.) is an estimated distribution from the training set.

[0171] Hence, we have estimated the parameters that will be used in the stochastic process. Then, by using these estimates, we can estimate the mean and the variance of the molecules created at each cycle (note that we do this separately for normal molecules, error molecules, and mutation molecules).

[0172] Finally, by using a probabilistic method (such as maximum likelihood or similar methods), we can determine the best fevalue that fits the distribution of the error, mutation, and normal molecules the best. More specifically, we estimate the expected ratio of the error molecules to total molecules for various Rvalues in the final reads, and determine the likelihood of our data for each of these values, and then select the value with the highest likelihood.N.062.WO.01Primer Design / Library Preparation

[0173] Primer tails can improve the detection of fragmented DNA from universally tagged libraries. If the library tag and the primer- tails contain a homologous sequence, hybridization can be improved (for example, melting temperature (Tm) is lowered) and primers can be extended if only a portion of the primer target sequence is in the sample DNA fragment. In some embodiments, 13 or more target specific base pairs may be used. In some embodiments, 10 to 12 target specific base pairs may be used. In some embodiments, 8 to 9 target specific base pairs may be used. In some embodiments, 6 to 7 target specific base pairs may be used.

[0174] In one embodiment, Libraries are generated from the samples above by ligating adaptors to the ends of DNA fragments in the samples, or to the ends of DNA fragments generated from DNA isolated from the samples. The fragments can then be amplified using PCR, for example, according to the following example protocol: 95°C, 2 min; 15 x [95°C, 20 sec, 55°C, 20 sec, 68°C, 20 sec], 68°C 2 min, 4°C hold.

[0175] Many kits and methods are known in the art for generation of libraries of nucleic acids that include universal primer binding sites for subsequent amplification, for example clonal amplification, and for subsequence sequencing. To help facilitate ligation of adapters library preparation and amplification can include end repair and adenylation (i.e. A-tailing). Kits especially adapted for preparing libraries from small nucleic acid fragments, especially circulating free DNA, can be useful for practicing methods provided herein. For example, the NEXTflex Cell Free kits available from Bioo Scientific or the Natera Library Prep Kit (available from Natera, Inc., San Carlos, CA). However, such kits would typically be modified to include adaptors that are customized for the amplification and sequencing steps of the methods provided herein. Adaptor ligation can be performed using commercially available kits such as the ligation kit found in the AGILENT SURESELECT kit (Agilent, CA).

[0176] Target regions of the nucleic acid library generated from DNA isolated from the sample, especially a circulating free DNA sample for the methods of the present disclosure, are then amplified. For this amplification, a series of primers or primer pairs, which can include between 5, 10, 15, 20, 25, 50, 100, 125, 150, 250, 500, 1000, 2500, 5000, 10,000, 20,000, 25,000, or 50,000 on the low end of the range and 15, 20, 25, 50, 100, 125, 150, 250, 500, 1000, 2500, 5000, 10,000, 20,000, 25.000, 50,000. 60.000, 75,000, or 100,000 primers on the upper end of the range, that each bind to one of a series of primer binding sites.N.062.WO.01

[0177] Primer designs can be generated with Primer3 (Untergrasser A, Cutcutache T, Koressaar T, Ye J, Faircloth BC, Remm M, Rozen SG (2012) “Primer3 - new capabilities and interfaces.” Nucleic Acids Research 40(15):ell5 and Koressaar T, Remm M (2007) “Enhancements and modifications of primer design program Primer3.” Bioinformatics 23(10): 1289-91) source code available at primer3.sourceforge.net). Primer specificity can be evaluated by BLAST and added to existing primer design pipeline criteria:

[0178] Primer specificities can be determined using the BLASTn program from the ncbi-blast-2.2.29+ package. The task option “blastn-short” can be used to map the primers against hgl9 human genome. Primer designs can be determined as “specific” if the primer has less than 100 hits to the genome and the top hit is the target complementary primer binding region of the genome and is at least two scores higher than other hits (score is defined by BLASTn program). This can be done in order to have a unique hit to the genome and to not have many other hits throughout the genome.

[0179] The final selected primers can be visualized in IGV (James T. Robinson, Helga Thorvaldsdottir, Wendy Winckler, Mitchell Guttman, Eric S. Lander, Gad Getz, Jill P. Mesirov. Integrative Genomics Viewer. Nature Biotechnology 29, 24-26 (2011)) and UCSC browser (Kent WJ, Sugnet CW, Furey TS, Roskin KM, Pringle TH, Zahler AM, Haussler D. The human genome browser at UCSC. Genome Res. 2002 Jun;12(6):996-1006 ) using bed files and coverage maps for validation.PCR Reaction Mixtures

[0180] Methods of the present disclosure, in certain embodiments, include forming an amplification reaction mixture. The reaction mixture typically is formed by combining a polymerase, nucleotide triphosphates, nucleic acid fragments from a nucleic acid library generated from the sample, a set of forward and reverse primers specific for target regions that contain SNVs. The reaction mixtures provided herein, themselves forming in illustrative embodiments, a separate aspect of the disclosure.

[0181] An amplification reaction mixture useful for embodiments of the present disclosure includes components known in the art for nucleic acid amplification, especially for PCR amplification. For example, the reaction mixture typically includes nucleotide triphosphates, a polymerase, and magnesium. Polymerases that are useful for embodiments of the present disclosure can include any polymerase that can be used in an amplification reaction especiallyN.062.WO.01those that are useful in PCR reactions. Tn certain embodiments, hot start Taq polymerases are especially useful. Amplification reaction mixtures useful for practicing the methods provided herein, such as AmpliTaq Gold master mix (Life Technologies, Carlsbad, CA), are available commercially.

[0182] Amplification (e.g. temperature cycling) conditions for PCR are well known in the art. The methods provided herein can include any PCR cycling conditions that result in amplification of target nucleic acids such as target nucleic acids from a library. Non-limiting example cycling conditions are provided in the Examples section herein.

[0183] There are many workflows that are possible when conducting PCR; some workflows typical to the methods disclosed herein are provided herein. The steps outlined herein are not meant to exclude other possible steps nor does it imply that any of the steps described herein are required for the method to work properly. A large number of parameter variations or other modifications are known in the literature, and may be made without affecting the essence of the embodiments of the disclosure.

[0184] In certain embodiments of the method provided herein, at least a portion and in illustrative examples the entire sequence of an amplicon, such as an outer primer target amplicon, is determined. Methods for determining the sequence of an amplicon are known in the art. Any of the sequencing methods known in the art, e.g. Sanger sequencing, can be used for such sequence determination. In illustrative embodiments high throughput next- generation sequencing techniques (also referred to herein as massively parallel sequencing techniques) such as, but not limited to, those employed in MYSEQ (ILLUMINA), HISEQ (ILLUMINA), ION TORRENT (LITE TECHNOLOGIES), GENOME ANALYZER ILX (ILLUMINA), GS FLEX+ (ROCHE 454), can be used for sequencing the amplicons produced by the methods provided herein.

[0185] High throughput genetic sequencers are amenable to the use of barcoding (i.e., sample tagging with distinctive nucleic acid sequences) so as to identify specific samples from individuals thereby permitting the simultaneous analysis of multiple samples in a single run of the DNA sequencer. The number of times a given region of the genome in a library preparation (or other nucleic preparation of interest) is sequenced (number of reads) will be proportional to the number of copies of that sequence in the genome of interest (or expression level in the case of cDNA containing preparations). Biases in amplification efficiency can be taken into account in such quantitative determination.N.062.WO.01

[0186] Methods of the present disclosure, in certain embodiments, include forming an amplification reaction mixture. The reaction mixture typically is formed by combining a polymerase, nucleotide triphosphates, nucleic acid fragments from a nucleic acid library generated from the sample, a series of forward target- specific outer primers and a first strand reverse outer universal primer. Another illustrative embodiment is a reaction mixture that includes forward target- specific inner primers instead of the forward target- specific outer primers and amplicons from a first PCR reaction using the outer primers, instead of nucleic acid fragments from the nucleic acid library. The reaction mixtures provided herein, themselves forming in illustrative embodiments, a separate aspect of the invention. In illustrative embodiments, the reaction mixtures are PCR reaction mixtures. PCR reaction mixtures typically include magnesium.

[0187] In some embodiments, the reaction mixture includes ethylenediaminetetraacetic acid (EDTA), magnesium, tetramethyl ammonium chloride (TMAC), or any combination thereof. In some embodiments, the concentration of TMAC is between 20 and 70 mM, inclusive. While not meant to be bound to any particular theory, it is believed that TMAC binds to DNA, stabilizes duplexes, increases primer specificity, and / or equalizes the melting temperatures of different primers. In some embodiments, TMAC increases the uniformity in the amount of amplified products for the different targets. In some embodiments, the concentration of magnesium (such as magnesium from magnesium chloride) is between 1 and 8 mM.

[0188] The large number of primers used for multiplex PCR of a large number of targets may chelate a lot of the magnesium (2 phosphates in the primers chelate 1 magnesium). For example, if enough primers are used such that the concentration of phosphate from the primers is ~9 mM, then the primers may reduce the effective magnesium concentration by ~4.5 mM. In some embodiments, EDTA is used to decrease the amount of magnesium available as a cofactor for the polymerase since high concentrations of magnesium can result in PCR errors, such as amplification of off-target loci. In some embodiments, the concentration of EDTA reduces the amount of available magnesium to between 1 and 5 mM (such as between 3 and 5 mM).

[0189] In some embodiments, the pH is between 7.5 and 8.5, such as between 7.5 and 8, 8 and 8.3, or 8.3 and 8.5, inclusive. In some embodiments, Tris is used at, for example, a concentration of between 10 and 100 mM, such as between 10 and 25 mM, 25 and 50 mM, 50 and 75 mM, or 25 and 75 mM, inclusive. In some embodiments, any of these concentrations of Tris are used at a pH between 7.5 and 8.5. In some embodiments, a combination of KC1 and (NHz^SCriis used, such asN.062.WO.01between 50 and 150 mM KC1 and between 10 and 90 mM (NFL^SC, inclusive. Tn some embodiments, the concentration of KC1 is between 0 and 30 mM, between 50 and 100 mM, or between 100 and 150 mM, inclusive. In some embodiments, the concentration of (NH^SOHs between 10 and 50 mM, 50 and 90 mM, 10 and 20 mM, 20 and 40 mM, 40 and 60 mM, or 60 and 80 mM (Nt SC, inclusive. In some embodiments, the ammonium [NH4+] concentration is between 0 and 160 mM, such as between 0 to 50, 50 to 100, or 100 to 160 mM, inclusive. In some embodiments, the sum of the potassium and ammonium concentration ([K+] + [NH4+]) is between 0 and 160 mM, such as between 0 to 25, 25 to 50, 50 to 150, 50 to 75, 75 to 100, 100 to 125, or 125 to 160 mM, inclusive. An example buffer with [K+] + [NH4+] = 120 mM is 20 mM KC1 and 50 mM (NH4)2SO4. In some embodiments, the buffer includes 25 to 75 mM Tris, pH 7.2 to 8, 0 to 50 mM KC1. 10 to 80 mM ammonium sulfate, and 3 to 6 mM magnesium, inclusive. In some embodiments, the buffer includes 25 to 75 mM Tris pH 7 to 8.5, 3 to 6 mM MgCb, 10 to 50 mM KC1, and 20 to 80 mM (NH zSC, inclusive. In some embodiments, 100 to 200 Units / mL of polymerase are used. In some embodiments, 100 mM KC1, 50 mM (NFU SC, 3 mM MgCh, 7.5 nM of each primer in the library, 50 mM TMAC, and 7 ul DNA template in a 20 ul final volume at pH 8.1 is used.

[0190] In some embodiments, a crowding agent is used, such as polyethylene glycol (PEG, such as PEG 8,000) or glycerol. In some embodiments, the amount of PEG (such as PEG 8,000) is between 0.1 to 20%, such as between 0.5 to 15%, 1 to 10%, 2 to 8%, or 4 to 8%. inclusive. In some embodiments, the amount of glycerol is between 0.1 to 20%, such as between 0.5 to 15%, 1 to 10%, 2 to 8%, or 4 to 8%, inclusive. In some embodiments, a crowding agent allows either a low polymerase concentration and / or a shorter annealing time to be used. In some embodiments, a crowding agent improves the uniformity of the Dinucleotide Odds Ratio (DOR) and / or reduces dropouts (undetected alleles).Polymerases

[0191] In some embodiments, a polymerase with proof-reading activity, a polymerase without (or with negligible) proof-reading activity, or a mixture of a polymerase with proof-reading activity and a polymerase without (or with negligible) proof-reading activity is used. In some embodiments, a hot start polymerase, a non-hot start polymerase, or a mixture of a hot start polymerase and a non-hot start polymerase is used. In some embodiments, a HotStarTaq DNA polymerase is used (see, for example, QIAGEN catalog No. 203203). In some embodiments,N.062.WO.01AmpliTaq Gold® DNA Polymerase is used. Tn some embodiments a PrimeSTAR GXL DNA polymerase, a high fidelity polymerase that provides efficient PCR amplification when there is excess template in the reaction mixture, and when amplifying long products, is used (Takara Clontech, Mountain View, CA). In some embodiments, KAPA Taq DNA Polymerase or KAPA Taq HotStart DNA Polymerase is used; they are based on the single- subunit, wild-type Taq DNA polymerase of the thermophilic bacterium Thermits aquaticus. KAPA Taq and KAPA Taq HotStart DNA Polymerase have 5'-3' polymerase and 5'-3' exonuclease activities, but no 3' to 5' exonuclease (proofreading) activity (see, for example, KAPA BIOSYSTEMS catalog No. BK1000). In some embodiments, Pfu DNA polymerase is used; it is a highly thermostable DNA polymerase from the hyperthermophilic archaeum Pyrococcus furiosus. The enzyme catalyzes the template-dependent polymerization of nucleotides into duplex DNA in the 5’— >3’ direction. Pfu DNA Polymerase also exhibits 3’— >5’ exonuclease (proofreading) activity that enables the polymerase to correct nucleotide incorporation errors. It has no 5’^-3’ exonuclease activity (see, for example, Thermo Scientific catalog No. EP0501). In some embodiments Klentaql is used; it is a Klenow-fragment analog of Taq DNA polymerase, it has no exonuclease or endonuclease activity (see, for example, DNA POLYMERASE TECHNOLOGY, Inc, St. Louis, Missouri, catalog No. 100). In some embodiments, the polymerase is a PHUSION DNA polymerase, such as PHUSION High Fidelity DNA polymerase (M0530S, New England BioLabs, Inc.) or PHUSION Hot Start Flex DNA polymerase (M0535S, New England BioLabs, Inc.). In some embodiments, the polymerase is a Q5® DNA Polymerase, such as Q5® High-Fidelity DNA Polymerase (M0491S, New England BioLabs, Inc.) or Q5® Hot Start High-Fidelity DNA Polymerase (M0493S, New England BioLabs, Inc.). In some embodiments, the polymerase is a T4 DNA polymerase (M0203S, New England BioLabs, Inc.).

[0192] In some embodiment, between 5 and 600 Units / mL (Units per 1 mL of reaction volume) of polymerase is used, such as between 5 to 100, 100 to 200, 200 to 300. 300 to 400, 400 to 500, or 500 to 600 Units / mL, inclusive.PCR Methods

[0193] In some embodiments, hot-start PCR is used to reduce or prevent polymerization prior to PCR thermocycling. Example hot-start PCR methods include initial inhibition of the DNA polymerase, or physical separation of reaction components reaction until the reaction mixture reaches the higher temperatures. In some embodiments, slow release of magnesium is used. DNAN.062.WO.01polymerase requires magnesium ions for activity, so the magnesium is chemically separated from the reaction by binding to a chemical compound, and is released into the solution only at high temperature. In some embodiments, non-covalent binding of an inhibitor is used. In this method a peptide, antibody, or aptamer are non-covalently bound to the enzyme at low temperature and inhibit its activity. After incubation at elevated temperature, the inhibitor is released and the reaction starts. In some embodiments, a cold-sensitive Taq polymerase is used, such as a modified DNA polymerase with almost no activity at low temperature. In some embodiments, chemical modification is used. In this method, a molecule is covalently bound to the side chain of an amino acid in the active site of the DNA polymerase. The molecule is released from the enzyme by incubation of the reaction mixture at elevated temperature. Once the molecule is released, the enzyme is activated.

[0194] In some embodiments, the amount to template nucleic acids (such as an RNA or DNA sample) is between 20 and 5,000 ng, such as between 20 to 200, 200 to 400, 400 to 600, 600 to 1,000; 1,000 to 1,500; or 2,000 to 3,000 ng, inclusive.

[0195] In some embodiments a QIAGEN Multiplex PCR Kit is used (QIAGEN catalog No.206143). For 100 x 50 pl multiplex PCR reactions, the kit includes 2x QIAGEN Multiplex PCR Master Mix (providing a final concentration of 3 mM MgCh, 3 x 0.85 ml), 5x Q-Solution (1 x 2.0 ml), and RNase-Free Water (2 x 1.7 ml). The QIAGEN Multiplex PCR Master Mix (MM) contains a combination of KC1 and (NH^SCE as well as the PCR additive, Factor MP, which increases the local concentration of primers at the template. Factor MP stabilizes specifically bound primers, allowing efficient primer extension by HotStarTaq DNA Polymerase. HotStarTaq DNA Polymerase is a modified form of Taq DNA polymerase and has no polymerase activity at ambient temperatures. In some embodiments, HotStarTaq DNA Polymerase is activated by a 15-minute incubation at 95 °C which can be incorporated into any existing thermal-cycler program.

[0196] In some embodiments, lx QIAGEN MM final concentration (the recommended concentration), 7.5 nM of each primer in the library, 50 mM TMAC, and 7 ul DNA template in a 20 ul final volume is used. In some embodiments, the PCR thermocycling conditions include 95 °C for 10 minutes (hot start); 20 cycles of 96°C for 30 seconds; 65°C for 15 minutes; and 72°C for 30 seconds; followed by 72°C for 2 minutes (final extension); and then a 4°C hold.

[0197] In some embodiments, 2x QIAGEN MM final concentration (twice the recommended concentration), 2 nM of each primer in the library, 70 mM TMAC, and 7 ul DNA template in a 20N.062.WO.01ul total volume is used. In some embodiments, up to 4 mM EDTA is also included. In some embodiments, the PCR thermocycling conditions include 95 °C for 10 minutes (hot start); 25 cycles of 96°C for 30 seconds; 65°C for 20, 25, 30, 45, 60, 120, or 180 minutes; and optionally 72°C for 30 seconds); followed by 72°C for 2 minutes (final extension); and then a 4°C hold.

[0198] Another example set of conditions includes a semi-nested PCR approach. The first PCR reaction uses 20 ul a reaction volume with 2x QIAGEN MM final concentration, 1.875 nM of each primer in the library (outer forward and reverse primers), and DNA template. Thermocycling parameters include 95°C for 10 minutes; 25 cycles of 96°C for 30 seconds, 65°C for 1 minute, 58°C for 6 minutes, 60°C for 8 minutes, 65°C for 4 minutes, and 72°C for 30 seconds; and then 72°C for 2 minutes, and then a 4°C hold. Next, 2 ul of the resulting product, diluted 1:200, is used as input in a second PCR reaction. This reaction uses a 10 ul reaction volume with lx QIAGEN MM final concentration, 20 nM of each inner forward primer, and 1 uM of reverse primer tag. Thermocycling parameters include 95°C for 10 minutes; 15 cycles of 95°C for 30 seconds, 65°C for 1 minute, 60°C for 5 minutes, 65°C for 5 minutes, and 72°C for 30 seconds; and then 72°C for 2 minutes, and then a 4°C hold. The annealing temperature can optionally be higher than the melting temperatures of some or all of the primers, as discussed herein (see U. S. Patent Application No. 14 / 918,544, filed Oct. 20, 2015, which is herein incorporated by reference in its entirety).

[0199] The melting temperature (Tm) is the temperature at which one-half (50%) of a DNA duplex of an oligonucleotide (such as a primer) and its perfect complement dissociates and becomes single strand DNA. The annealing temperature (TA) is the temperature one runs the PCR protocol at. For prior methods, it is usually 5 °C below the lowest Tmof the primers used, thus close to all possible duplexes are formed (such that essentially all the primer molecules bind the template nucleic acid). While this is highly efficient, at lower temperatures there are more unspecific reactions bound to occur. One consequence of having too low a TA is that primers may anneal to sequences other than the true target, as internal single-base mismatches or partial annealing may be tolerated. In some embodiments of the present inventions, the TA is higher than Tm, where at a given moment only a small fraction of the targets have a primer annealed (such as only -1-5%). If these get extended, they are removed from the equilibrium of annealing and dissociating primers and target (as extension increases Tmquickly to above 70°C), and a new -1-5% of targets has primers. Thus, by giving the reaction a long time for annealing, one can get -100% of the targets copied per cycle.N.062.WO.01

[0200] In various embodiments, the annealing temperature is between 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 °C and 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 15 °C on the high end of the range, greater than the melting temperature (such as the empirically measured or calculated Tm) of at least 25, 50, 60, 70, 75, 80, 90, 95, or 100% of the non-identical primers. In various embodiments, the annealing temperature is between 1 and 15 °C (such as between 1 to 10, 1 to 5, 1 to 3, 3 to 5, 5 to 10, 5 to 8, 8 to 10. 10 to 12, or 12 to 15 °C, inclusive) greater than the melting temperature (such as the empirically measured or calculated Tm) of at least 25; 50; 75; 100; 300; 500; 750; 1,000; 2,000; 5,000; 7,500; 10,000; 15,000; 19,000; 20,000; 25,000; 27,000; 28,000; 30,000; 40,000; 50,000; 75,000; 100,000; or all of the non-identical primers. In various embodiments, the annealing temperature is between 1 and 15 °C (such as between 1 to 10, 1 to 5, 1 to 3, 3 to 5, 3 to 8, 5 to 10, 5 to 8, 8 to 10, 10 to 12, or 12 to 15 °C, inclusive) greater than the melting temperature (such as the empirically measured or calculated Tm) of at least 25%, 50%, 60%, 70%, 75%, 80%, 90%, 95%, or all of the non-identical primers, and the length of the annealing step (per PCR cycle) is between 5 and 180 minutes, such as 15 and 120 minutes, 15 and 60 minutes, 15 and 45 minutes, or 20 and 60 minutes, inclusive.Example Multiplex PCR Methods

[0201] In various embodiments, long annealing times (as discussed herein and illustrated in Example 10) and / or low primer concentrations are used. In fact, in certain embodiments, limiting primer concentrations and / or conditions are used. In various embodiments, the length of the annealing step is between 15, 20, 25, 30, 35, 40, 45, or 60 minutes on the low end of the range and 20, 25, 30, 35, 40, 45, 60, 120, or 180 minutes on the high end of the range. In various embodiments, the length of the annealing step (per PCR cycle) is between 30 and 180 minutes. For example, the annealing step can be between 30 and 60 minutes and the concentration of each primer can be less than 20, 15, 10, or 5 nM. In other embodiments the primer concentration is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 25 nM on the low end of the range, and 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, and 50 on the high end of the range.

[0202] At high level of multiplexing, the solution may become viscous due to the large amount of primers in solution. If the solution is too viscous, one can reduce the primer concentration to an amount that is still sufficient for the primers to bind the template DNA. In various embodiments, between 1,000 and 100,000 different primers are used and the concentration of each primer is less than 20 nM, such as less than 10 nM or between 1 and 10 nM, inclusive.N.062.WO.01CNV Detection

[0203] In addition to SNVs and indels, methods for monitoring and detection of early relapse and metastasis described herein can also benefit from detection of CNVs.

[0204] In one aspect, embodiments of the present disclosure generally relate, at least in part, to improved methods of determining the presence or absence of copy number variations, such as deletions or duplications of chromosome segments or entire chromosomes. The methods are particularly useful for detecting small deletions or duplications, which can be difficult to detect with high specificity and sensitivity using prior methods due to the small amount of data available from the relevant chromosome segment. The methods include improved analytical methods, improved bioassay methods, and combinations of improved analytical and bioassay methods. Methods of the disclosure, in certain embodiments, can also be used to detect deletions or duplications that are only present in a small percentage of the cells or nucleic acid molecules that are tested. This allows deletions or duplications to be detected prior to the occurrence of disease (such as at a precancerous stage) or in the early stages of disease, such as before a large number of diseased cells (such as cancer cells) with the deletion or duplication accumulate. The more accurate detection of deletions or duplications associated with a disease or disorder enable improved methods for diagnosing, prognosticating, preventing, delaying, stabilizing, or treating the disease or disorder. Several deletions or duplications are known to be associated with cancer or with severe mental or physical handicaps.SNV Detection

[0205] In another aspect, the present disclosure generally relates, at least in part, to improved methods of detecting SNVs. These improved methods include improved analytical methods, improved bioassay methods, and improved methods that use a combination of improved analytical and bioassay methods. The methods in certain illustrative embodiments are used to detect, diagnose, monitor, or stage cancer, for example in samples where the SNV is present at very low concentrations, for example less than 10%, 5%, 4%, 3%, 2.5%, 2%, 1%, 0.5%, 0.25%, or 0.1% relative to the total number of normal copies of the SNV locus, such as circulating free DNA samples. That is, these methods in certain illustrative embodiments are particularly well suited for samples where there is a relatively low percentage of a mutation or variant relative to the normal polymorphic alleles present for that genetic loci. Finally, provided herein are methods that combineN.062.WO.01the improved methods for detecting copy number variations with the improved methods for detecting single nucleotide variations.

[0206] Successful treatment of a disease such as cancer often relies on early diagnosis, correct staging of the disease, selection of an effective therapeutic regimen, and close monitoring to prevent or detect relapse. For cancer diagnosis, histological evaluation of tumor material obtained from tissue biopsy is often considered the most reliable method. However, the invasive nature of biopsy-based sampling has rendered it impractical for mass screening and regular follow up. Therefore, the present methods have the advantage of being able to be performed non-invasively if desired for relatively low cost with fast turnaround time. The targeted sequencing that may be used by the methods of the disclosure, in certain embodiments, requires less reads than shotgun sequencing, such as a few million reads instead of 40 million reads, thereby decreasing cost. The multiplex PCR and next generation sequencing that may be used increase throughput and reduces costs.

[0207] In some example embodiments, analysis of average allelic imbalance (A Al) patterns in ctDNA provide more detailed insights into the clonal architecture of tumors to help predict their therapeutic responses and optimize treatment strategies. Therefore, in certain embodiments, mmPCR-NGS panels are selected that target clinically actionable CNVs and SNVs. Such panels in certain illustrative embodiments, are particularly useful for patients with cancers where CNVs represent a substantial proportion of the mutation load, as is common in breast, ovarian, and lung cancer. In various embodiments, “AAI” represents a measure of deviation from expected allelic balance across a plurality of loci or genomic regions within a sample.

[0208] In some embodiments, the methods are used to detect a deletion, duplication, or single nucleotide variant in an individual. A sample from the individual that contains cells or nucleic acids suspected of having a deletion, duplication, or single nucleotide variant may be analyzed. In some embodiments, the sample is from a tissue or organ suspected of having a deletion, duplication, or single nucleotide variant, such as cells or a mass suspected of being cancerous. The methods of the disclosure, in certain embodiments, can be used to detect deletion, duplication, or single nucleotide variant that are only present in one cell or a small number of cells in a mixture containing cells with the deletion, duplication, or single nucleotide variant and cells without the deletion, duplication, or single nucleotide variant. In some embodiments. cfDNA or cfRNA from a blood sample from the individual is analyzed. In some embodiments, cfDNA or cfRNA isN.062.WO.01secreted by cells, such as cancer cells. In some embodiments, cfDNA or cfRNA is released by cells undergoing necrosis or apoptosis, such as cancer cells. The methods of the disclosure, in certain embodiments, can be used to detect deletion, duplication, or single nucleotide variant that are only present in a small percentage of the cfDNA or cfRNA. In some embodiments, one or more cells from an embryo are tested.

[0209] In addition to determining the presence or absence of copy number variation, one or more other factors can be analyzed if desired. These factors can be used to increase the accuracy of the diagnosis (such as determining the presence or absence of cancer or an increased risk for cancer, classifying the cancer, or staging the cancer) or prognosis. These factors can also be used to select a particular therapy or treatment regimen that is likely to be effective in the subject. Example factors include the presence or absence of polymorphisms or mutation; altered (increased or decreased) levels of total or particular cfDNA, cfRNA, microRNA (miRNA); altered (increased or decreased) tumor fraction; altered (increased or decreased) methylation levels, altered (increased or decreased) DNA integrity, altered (increased or decreased) or alternative mRNA splicing.

[0210] The following sections describe methods for detecting deletions or duplications using phased data (such as inferred or measured phased data) or unphased data; samples that can be tested; methods for sample preparation, amplification, and quantification; methods for phasing genetic data; polymorphisms, mutations, nucleic acid alterations, mRNA splicing alterations, and changes in nucleic acid levels that can be detected; databases with results from the methods, other risk factors and screening methods; cancers that can be diagnosed or treated; cancer treatments; cancer models for testing treatments; and methods for formulating and administering treatments.Example EmbodimentsA. Example Embodiments for Data Preparation

[0211] In certain aspects, provided herein are methods for using new variant callers, including revised error models with VAF priors and new outlier detection methods, to obtain variant and sample level data.

[0212] In certain illustrative examples, training and validation data is based on augmented samples and 64-plex samples from the evidence generation set. The underlying 16-plex data for the augmented samples comprises various cohorts, including approximately 1500 negative samples and approximately 1500 Signatera positive samples.N.062.WO.01

[0213] In some examples, a balanced augmented training set and an augmented negatives test set are generated. In these examples, a balanced augmented training set is generated with, for example, 5000 negative samples and 5000 positive samples. The underlying 16-plex samples are stratified by DNA input into three bins. This stratification helps ensure that three augmented groups corresponding to these DNA input bins are obtained. The augmented data is ordered by the underlying sample amplicon, and ties are resolved randomly to avoid any unpredicted design biases that could artificially enhance performance.

[0214] Samples with Expected Positivity = 1 and 0 variants are excluded to avoid additional false positives during inference. Samples with fewer than 50 positions are excluded to ensure the quality and reliability of the training data. Samples with an expected positivity equal to 0.5 are excluded. Samples that are missing truth data which have more than 4 detected targets are assigned an expected positivity value equal to 1. Samples that are missing truth data which have less than 1 detected target are assigned an expected positivity value equal to 0.

[0215] In some embodiments, data cleaning involves filtering out bad targets and known germline mutations. In some embodiments, data cleaning involves zeroing out values of low error counts (e.g., less than 2, 4, 6, 8, 10, 25, 20, 25, 50, or 100). values of a low total of DOR positions’ posterior probabilities (e.g., less than 1,000; 2,000; 3,000; 4,000; 5,000; 6,000; 7,000; 8,000; 9,000; or 10,000), and values of a high total of DOR positions’ posterior probabilities (e.g., greater than 25,000; 50,000; 100,000, 200,000; 350,000; or 500,000). In some embodiments, data cleaning involves filtering out samples with DNA input less than 5 ng and samples with fewer than 50 remaining quality positions, particularly in machine learning model training. Hence for samples with less than 50-plex, only a variant threshold caller may be used in certain embodiments. In additional embodiments, a clustering approach is implemented to filter out unknown germline mutations (e.g., outliers).B. Example Embodiments for Training ML Sample Caller 1

[0216] In certain examples, feature generation for the classification algorithm is derived from the variant caller output for all positions, including both the first set of target loci (e.g., tumorspecific “on-target” positions) and the second set of target loci (e.g., off-target positions). The first set of target loci may be identified by performing multiplex PCR to amplify a plurality of target loci from cell-free DNA isolated from a biological sample. The process can target specific regions of the genome that span at least one cancer- specific mutation. The number of target lociN.062.WO.01can vary, typically ranging from 1 to 100, depending on the specific requirements of the analysis. The second set of loci are positions that are not specifically targeted for amplification but are still analyzed during the sequencing process. The number of positions in the second set of target loci can be one or two orders of magnitude greater than the number of on-target positions. For the on-target features, the calculations of probability of positive mutation generated by the augmented variant caller are utilized, which take into consideration the contribution of multiple callers at a variant level.

[0217] In some embodiments, the variant caller is used to obtain a probability value of a mutation being true for every target position of the sample. At a sample level, the goal is to use the individual target probabilities to calculate the probability that exactly k targets are true mutations, where k=0. 1, 2, 3,..., N. This can be calculated as the probability that k mutations are true given pl, p2, p3,..., pN, where pl, p2, p3,..., pN are the individual position probabilities.

[0218] Features based on the second set of target loci include posterior probability calculations from the revised error model. In some embodiments, information regarding the second set of target positions for the amplicons for the same G64 target positions is included as predictor features in the sample caller. The objective is to provide a baseline or internal calibration that captures noise across the sample and affects both the first set of target positions and the second set of target positions.

[0219] The classification problem is defined as a binary call on the proposition " Sample is Positive". In some embodiments, the classification implementation may be based on, for example, XGBoost, which may be chosen for its ease of use. scalability, and accuracy.Individual XGBoost models may be trained for each training / validation split in which the overall training set was divided.

[0220] Although gradient boosting decision tree models (e.g., XGBoost, CatBoost) are listed, in other embodiments the first and / or second ML models can include other classification algorithms such as logistic regression, random forests, support vector machines, neural networks, or combinations thereof. Likewise, in addition to the probability and posterior-probability features described herein, features can include, for example, per-locus allele depths, base quality metrics, fragment length distributions, and other ctDNA-related statistics derived from the first and / orN.062.WO.01second sets of target loci. Substitution among such algorithms and features, along with routine hyperparameter tuning and cross-validation, can be performed in various embodiments.C. Example Embodiments for Training ML Sample Caller 2

[0221] Accordingly, provided herein in one embodiment, is a method for feature generation with information from both first and second sets of target loci.

[0222] In certain aspects, provided herein are methods for performing several posterior probability feature transformations to form the underlying features for the Machine Learning algorithm being used. The feature list includes features such as top position posterior probabilities, number of position posterior probabilities per sample above certain thresholds, and IQR-like features (differences between different posterior probabilities’ quantiles).

[0223] In certain aspects, provided herein are methods for training and validating the ML Sample Caller 2 algorithm. The training and validation processes are refined to optimize cross-validation performance. In certain aspects, methods for training an algorithm on an augmented training set may utilize, for example, CatBoost, an efficient tree-based boosting algorithm that has demonstrated superior performance compared to other gradient boosting decision tree (GBDT) libraries on various datasets. The metric optimized through cross-validation is sensitivity, given a sufficiently high specificity. Hyperparameters are selected through randomized search and include the number of iterations, tree depth, regularization parameters, and the class weighting. Monotonicity constraints are introduced on several features, and early stopping is implemented after 200 rounds to prevent overfitting. Additionally, balanced trees are used as a form of regularization to further mitigate overfitting.

[0224] Once the optimal hyperparameters are determined, the validation set performance is evaluated. The classification threshold is set to achieve a sufficiently high specificity on the extended validation set. The test set performance for ML Sample Caller 2 is then reported. To ensure robustness, the train / validation split procedure is repeated (e.g., ten times), and sample predictions are generated (e.g., ten times).

[0225] In various embodiments, a “first feature set” and a “second feature set” are “different” when at least one feature is present in one set and absent from the other. The feature sets may partially overlap (e.g., share some features in common) or may be disjoint. Features may differ, for example, in the underlying statistics used (e.g., per-target mutation probabilities versusN.062.WO.01posterior probabilities), the aggregation across loci (e.g., per-sample counts versus quantiles), or the particular loci, coverage metrics, or error-model outputs included.

[0226] In various embodiments, when it is stated that a first machine learning model “employs a first algorithm” and a second machine learning model “employs a second algorithm different from the first algorithm,” the algorithms are considered “different” when they use distinct model families or training procedures (e.g., XGBoost and CatBoost, or a gradient boosting decision tree model and a logistic regression model). In some embodiments, the first and second algorithms may both be gradient boosting methods implemented in different libraries or with materially different training procedures (e.g., different handling of categorical features, monotonic constraints, or tree-building strategies) and may still be considered “different algorithms.”D. Example Embodiments for Model Ensembling

[0227] Provided herein in various embodiments are methods for combining multiple machine learning models to improve the overall performance and robustness of the G64 sample caller. In certain aspects, the method combines two machine learning (ML) models by applying a geometric average to the probability-like results produced by each model. Given that gradient boosting risk scores do not accurately represent probabilities, Platt scaling may be applied. This monotonic transformation adjusts the predictions to be closer to probability estimates by fitting a logistic regression of risk scores against the true labels on the validation set. The quality of this re-fit may be estimated using the Brier Score. A threshold for calling a sample positive may then be defined as any reasonably high percentile of the negative validation samples' transformed risk scores.

[0228] In this embodiment, each ML sample caller is trained and validated ten times. Consequently, Platt transformations, geometric averaging, and validation set threshold settings are performed ten times, and the final votes for each test sample are reported out of ten. A test sample is classified as positive by the ML Callers if it receives more than a predetermined number of votes out of ten.

[0229] In this embodiment further, the ensembled ML Callers are combined with an improved variant caller, which incorporates off-target false positive rate (FPR) based thresholding. FPR is controlled by choosing a conservative threshold based on the validation set’s negative predicted outcomes. A sample is classified as positive if any of the combined callers predict a positive result, using the 'OR' condition.N.062.WO.01

[0230] In various embodiments, the use of multiple machine-learning models trained on different feature sets, combined with internal calibration based on additional off-target loci, provides improved detection performance relative to prior ctDNA detection methods. For example, the described feature-engineering and ensembling approaches can increase sensitivity at a fixed high specificity for samples with low DNA input and low variant allele frequency. This can represent an improvement in the functioning of the ctDNA analysis pipeline.Example Computing Systems

[0231] Referring to FIG. 1, in potential implementations, a system 100 may include a computing system 110 (which may be or may include one or more computing devices, co-located or remote to each other), a condition detection system 160, a Clinical Information System (CIS) 170 (used interchangeably with electronic medical record (EMR) system or electronic health record (EHR) system), a platform 175, and a therapeutic system 180. The computing system 110 (e.g., one or more computing devices) may be used to control and / or exchange signals and / or data with condition detection system 160, CIS 170, platform 175, and / or therapeutic system 180, directly or via another component of system 100. In certain embodiments, computing system 110 may be used to control and / or exchange data or other signals with condition detection system 160, CIS 170, platform 175, and / or therapeutic system 180. The computing system 110 may include one or more processors and one or more volatile and non-volatile memories for storing computing code and data that are captured, acquired, recorded, and / or generated.

[0232] The computing system 110 may include a controller 112 that is configured to exchange control signals with condition detection system 160. CIS 170, platform 175, therapeutic system 180, and / or any components thereof, allowing the computing system 110 to be used to control, for example, acquisition of patient data such as test results, capture of images, acquisition of signals by sensors, positioning or repositioning of subjects and patients, recording or obtaining other subject or patient information, and applying therapies.

[0233] A transceiver 114 allows the computing system 110 to exchange readings, control commands, and / or other data, wirelessly or via wires, directly or via networking protocols, with condition detection system 160, CIS 170, platform 175, and / or therapeutic system 180, or components thereof. One or more user interfaces 116 allow the computing device 110 to receive user inputs (e.g., via a keyboard, touchscreen, microphone, camera, etc.) and provide outputs (e.g., via a display screen, audio speakers, etc.) with users. The computing device 110 may additionallyN.062.WO.01include one or more databases 118 for storing, for example, data acquired from one or more systems or devices, signals acquired via one or more sensors, biomarker signatures, etc. In some implementations, database 118 (or portions thereof) may alternatively or additionally be part of another computing device that is co-located or remote (e.g., via “cloud computing”) and in communication with computing device 110, condition detection system 160, CIS 170, platform 175, and / or therapeutic system 180 or components thereof.

[0234] Condition detection system 160 may include a testing system 162, which may be or may include, for example, any system or device that is involved in, for example, analyzing samples (e.g., blood, plasma, or urine), recording patient data, and / or obtaining laboratory or other test results (e.g., DNA data). An imaging system 164 may be any system or device used to capture imaging data, such as a magnetic resonance imaging (MRI) scanner, a positron emission tomography (PET) scanner, a single photon emission computed tomography (SPECT) scanner, a computed tomography (CT) scanner, a fluoroscopy scanner, and / or other imaging devices and / or sensors. Sensors 166 may detect, for example, a position or motion of a patient, organs, tissues, physiological readings such as lung capacity or heart activity / signals, or other states and / or conditions of the patient.

[0235] Therapeutic system 180 may include a treatment unit 180, which may be or may include, for example, a radiation source for external beam therapy (e.g., orthovoltage x-ray machines, Cobalt-60 machines, linear accelerators, proton beam machines, neutron beam machines, etc.) and / or one or more other treatment devices. Sensors 184 may be used by therapeutic system 180 to evaluate and guide a treatment (e.g., by detecting level of emitted radiation, a condition or state of the patient, or other states or conditions). In various implementations, components of system 100 may be rearranged or integrated in other configurations. For example, computing system 110 (or components thereof) may be integrated with one or more of the condition detection system 160, therapeutic system 180. and / or components thereof. The condition detection system 160, therapeutic system 180, and / or components thereof may be directed to a platform 175 on which a patient, subject, or sample can be situated (so as to test a sample, image a subject, apply a treatment or therapy to the subject, detect activity and / or motion of the subject in a stress test, etc.). In various embodiments, the platform 175 may be movable (e.g., using any combination of motors, magnets, etc.) to allow for positioning and repositioning of samples and / or subjects (such as microadjustments to compensate for motion of a subject or patient or to position the patient for scans ofN.062.WO.01different regions of interest). The platform 175 may include its own sensors to detect a condition or state of the sample, patient, and / or subject.

[0236] The computing system 110 may include a tester and imager 120 configured to direct, for example, laboratory tests, image capture, and acquisition of test and / or imaging data. Tester and imager 120 may include an image generator that may convert or transform raw imaging data from condition detection system 160 into usable medical images or into another form to be analyzed. Computing system 110 may include a test and image analyzer 122 configured to, for example, analyze raw test results to generate relevant metrics, identify features in images or imaging data, or otherwise make use of tests or testing data, and / or images or imaging data.

[0237] A data acquisition unit 124 may retrieve, acquire, or otherwise obtain various data to be used to. for example, train models using data on subjects, or apply trained models using data on patients. The data acquisition unit 124 may, for example, obtain test data stored in CIS 170. An interaction unit 126 may interact (e.g., via user interfaces 116) with users (e.g., patients and / or healthcare providers) to obtain information needed for a model or to provide information to users. In certain embodiments, the data acquisition unit 124 may obtain data from users via interaction unit 126.

[0238] A machine learning model trainer 130 (“ML model trainer”) may be configured to train machine learning models used herein, as further discussed below. The ML model trainer may include a training dataset generator that processes various data to generate one or more training datasets to be used, for example, to train the models discussed herein. In certain embodiments, as more data and outcomes become available, a model retrainer 134 may further train a previously-trained model to improve or otherwise update or revise models. The ML model trainer may apply various machine learning techniques to the training datasets. A machine learning modeler 140 (“ML modeler”) may be configured to apply trained machine learning models to particular patient data. A feature selector 142 may be configured to select which features are to be fed to a model to obtain a prediction (e.g., presence of ctDNA). The feature selector 142 may select features based on, for example, which features were previously employed to train the ML model(s) to be used and / or the parameters of the model(s). A feature extractor 144 may obtain values for selected features from various data sources. An outputting and reporting unit 150 may transmit models or predictions thereof (to, e.g., healthcare professionals, patients, etc.), or generate and provide reports based on the models or predictions thereof. The reports may include, for example, modelN.062.WO.01outputs (e.g., a prediction) along with information on the basis for a prediction, and guidelines and / or proposed recommendations and next steps based on the predictions, so as to identify for clinicians the factors most impactful on low or high likelihoods or other prognoses and guide clinicians and patients.

[0239] Referring now to FIG. 6, FIG. 6 is a block diagram of an example of a system 600, in accordance with implementing some embodiments of the present disclosure. System 600 may include a first feature set 602 and a second feature set 604, wherein both feature sets are based on input sampling data. System 600 may further include a first ML model such as a first ML Sample Caller 606 and a second ML model such as a second ML Sample Caller 608. The first feature set 602 may be used as input to the first ML Sample Caller 606. The first ML Sample Caller 606 may produce an output contributing to a classification 620 which may be a ctDNA positivity classification. The resulting output of the first ML Sample Caller 606 may indicate a likelihood of cancer presence based on the first feature set. The second feature set 604 may be used as input to the second ML Sample Caller 608. The second ML Sample Caller 606 may produce an output contributing to the classification 620. The resulting output of the second ML Sample Caller 608 may indicate a likelihood of cancer presence based on the second feature set. System 600 may further include an ensembler 610 for combining the results from the first and second ML Callers 606 and 608, respectively. The ensembler 610 may output an ensembled result used for the final ctDNA positivity classification 620. The classification 620 may indicate whether a sample from the input sampling data is positive for cancer.

[0240] Referring now to FIG. 7, FIG. 7 is a block diagram of an example of a system 700 for training and testing machine learning models, in accordance with implementing some embodiments of the present disclosure. The system 700 may include a train set 702 to be used as input for training a first ML model, which may be ML Caller 704, and for training a second ML model, which may be ML Caller 706. Validation set 708 may be used in the validation process after training the ML models. ML Caller predictions 710 may be obtained and recorded based on validation set 708. The system 700 may include Platt calibrators 712. Platt calibrators 712 may be fit based on the validation set 708 and may transform the ML caller predictions 710. Geometric transformation 714 may be included for combining the ML caller predictions by calculating the geometric average of the transformed predictions. The system 700 may include threshold 716, which may be defined based on the validation set 708. The system 700 may include test set 718N.062.WO.01for testing the machine learning models. Samples 720 with transformed combined predictions above the threshold may be recorded.

[0241] Referring now to FIG. 8, Referring to FIG. 8, an example process 800 is illustrated, according to various potential embodiments. Various elements of process 800 may be implemented by or via system 100 or components thereof. Process 800 may begin (205) with model training (on the left side of FIG. 8), which may be implemented by or via computing system 110, if a model is not already available (e.g., in database 118), or if additional models are to be generated or updated through training or retraining with new training data. Alternatively, process 800 may begin with use (application) of a model (on the right side of FIG. 8) for a patient if a trained model is already available. Predictions may be implemented by or via computing system 110 if a suitable trained model is available. In various embodiments, process 800 may comprise both model training (e.g., steps 810 - 825) followed by prognosis prediction (e.g., steps 850 - 870).

[0242] At 810, health data pertaining to subjects may be obtained. This may be obtained by or via, for example, condition detection system 160 and / or CIS 170 for a cohort of, for example, 10,000, 11,000, 12,000, 13,000, 14,000, 15,000, or more subjects. The health data should have enough data to be able to capture and account for a variety of features and build a generalizable model. Step 815 involves extracting, from (or based on) the health data, feature values corresponding to the subjects in the cohort.

[0243] At 820, one or more datasets (e.g., training datasets) may be generated using the extracted feature values corresponding health data of the subjects. This may be performed, for example, by or via ML model trainer 130, or more specifically, training and test dataset generator 132. At 825, the one or more datasets may be used to train and test a model for making predictions. The model may be stored (e.g., in database 118) for subsequent use. Process 800 may end (890), or may proceed to step 850 for use in prognosis prediction (as represented by the dotted line from step 825 to step 865, the model may subsequently be used to generate and use predictions.)

[0244] At 850, health data of a patient may be obtained and analyzed (e.g., by or via condition detection system 160 and / or CIS 170). Test data may correspond to various laboratory tests performed on samples of the patient. At 855, a set of features may be selected (e.g., by or via feature selector 142), and at 860, values for the features may be extracted from the health data of the patient (e.g., by or via feature extractor 144).N.062.WO.01

[0245] At 865, feature values extracted from patient health data may be input to a predictive model (e.g., a machine learning classifier) to predict patient prognosis. At 870, the predicted prognosis and the factors underlying the prognosis may be used in caring for the patient. For example, predicted prognosis may be used for planning for a potential outcome and / or identifying potentially preventative care. Additionally or alternatively, predicted prognosis can be used in evaluation of a patient’s health and changes or trends therein, such as whether the patient is deteriorating or improving as indicated by changes in prognoses over time from the ML modeler as new tests are run or otherwise as new health data becomes available.

[0246] Process 800 may end (890), or return to step 850 (e.g., after running another test or administering a treatment) for subsequent planning based on a change in a condition of the patient.

[0247] FIG. 13 shows a simplified block diagram of a representative server system 1300, client computing system 1314, and network 1326 usable to implement certain embodiments of the present disclosure. In various embodiments, server system 1300 or similar systems can implement services or servers described herein or portions thereof. Client computing system 1314 or similar systems can implement clients described herein. Server system 1300 can have a modular design that incorporates a number of modules 1302 (e.g., blades in a blade server embodiment); while two modules 1302 are shown, any number can be provided. Each module 1302 can include processing unit(s) 1304 and local storage 1306.

[0248] Processing unit(s) 1304 can include a single processor, which can have one or more cores, or multiple processors. In some embodiments, processing unit(s) 1304 can include a general-purpose primary processor as well as one or more special-purpose co-processors such as graphics processors, digital signal processors, or the like. In some embodiments, some or all processing units 1304 can be implemented using customized circuits, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). In some embodiments, such integrated circuits execute instructions that are stored on the circuit itself. In other embodiments, processing unit(s) 1304 can execute instructions stored in local storage 1306. Any type of processors in any combination can be included in processing unit(s) 1304.

[0249] Local storage 1306 can include volatile storage media (e.g., DRAM, SRAM, SDRAM, or the like) and / or non-volatile storage media (e.g., magnetic or optical disk, flash memory, or the like). Storage media incorporated in local storage 1306 can be fixed, removable or upgradeable as desired. Local storage 1306 can be physically or logically divided into various subunits such as aN.062.WO.01system memory, a read-only memory (ROM), and a permanent storage device. The system memory can be a read- and- write memory device or a volatile read-and-write memory, such as dynamic random-access memory. The system memory can store some or all of the instructions and data that processing unit(s) 1304 need at runtime. The ROM can store static data and instructions that are needed by processing unit(s) 1304. The permanent storage device can be a non-volatile read-and-write memory device that can store instructions and data even when module 1302 is powered down. The term “storage medium” as used herein includes any medium in which data can be stored indefinitely (subject to overwriting, electrical disturbance, power loss, or the like) and does not include carrier waves and transitory electronic signals propagating wirelessly or over wired connections.

[0250] In some embodiments, local storage 1306 can store one or more software programs to be executed by processing unit(s) 1304, such as an operating system and / or programs implementing various methods or steps thereof such as methods 300 and 500 of FIGS. 3 and 5 or any other methods or processes described herein.

[0251] “Software” refers generally to sequences of instructions that, when executed by processing unit(s) 1304 cause server system 1300 (or portions thereof) to perform various operations, thus defining one or more specific machine embodiments that execute and perform the operations of the software programs. The instructions can be stored as firmware residing in read-only memory and / or program code stored in non-volatile storage media that can be read into volatile working memory for execution by processing unit(s) 1304. Software can be implemented as a single program or a collection of separate programs or program modules that interact as desired. From local storage 1306 (or non-local storage described below), processing unit(s) 1304 can retrieve program instructions to execute and data to process in order to execute various operations described above.

[0252] In some server systems 1300, multiple modules 1302 can be interconnected via a bus or other interconnect 1308, forming a local area network that supports communication between modules 1302 and other components of server system 1300. Interconnect 1308 can be implemented using various technologies including server racks, hubs, routers, etc.

[0253] A wide area network (WAN) interface 1310 can provide data communication capability between the local area network (interconnect 1308) and the network 1326, such as the Internet.N.062.WO.01Technologies can be used, including wired (e.g., Ethernet, IEEE 802.3 standards) and / or wireless technologies (e.g., Wi-Fi, IEEE 802.11 standards).

[0254] In some embodiments, local storage 1306 is intended to provide working memory for processing unit(s) 1304, providing fast access to programs and / or data to be processed while reducing traffic on interconnect 1308. Storage for larger quantities of data can be provided on the local area network by one or more mass storage subsystems 1312 that can be connected to interconnect 1308. Mass storage subsystem 1312 can be based on magnetic, optical, semiconductor, or other data storage media. Direct attached storage, storage area networks, network-attached storage, and the like can be used. Any data stores or other collections of data described herein as being produced, consumed, or maintained by a service or server can be stored in mass storage subsystem 1312. In some embodiments, additional data storage resources may be accessible via WAN interface 1310 (potentially with increased latency).

[0255] Server system 1300 can operate in response to requests received via WAN interface 1310. For example, one of modules 1302 can implement a supervisory function and assign discrete tasks to other modules 1302 in response to received requests. Work allocation techniques can be used. As requests are processed, results can be returned to the requester via WAN interface 1310. Such operation can generally be automated. Further, in some embodiments, WAN interface 1310 can connect multiple server systems 1300 to each other, providing scalable systems capable of managing high volumes of activity. Other techniques for managing server systems and server farms (collections of server systems that cooperate) can be used, including dynamic resource allocation and reallocation.

[0256] Server system 1300 can interact with various user-owned or user-operated devices via a wide-area network such as the Internet. An example of a user-operated device is shown in FIG. 13 as client computing system 1314. Client computing system 1314 can be implemented, for example, as a consumer device such as a smartphone, other mobile phone, tablet computer, wearable computing device (e.g., smart watch, eyeglasses), desktop computer, laptop computer, and so on.

[0257] For example, client computing system 1314 can communicate via WAN interface 1310. Client computing system 1314 can include computer components such as processing unit(s) 1316, storage device 1318, network interface 1320, user input device 1322, and user output device 1324. Client computing system 1314 can be a computing device implemented in a variety of form factors,N.062.WO.01such as a desktop computer, laptop computer, tablet computer, smartphone, other mobile computing device, wearable computing device, or the like.

[0258] Processing unit(s) 1316 and storage device 1318 can be similar to processing unit(s) 1304 and local storage 1306 described above. Suitable devices can be selected based on the demands to be placed on client computing system 1314; for example, client computing system 1314 can be implemented as a “thin” client with limited processing capability or as a high-powered computing device. Client computing system 1314 can be provisioned with program code executable by processing unit(s) 1316 to enable various interactions with server system 1300.

[0259] Network interface 1320 can provide a connection to the network 1326, such as a wide area network (e.g., the Internet) to which WAN interface 1310 of server system 1300 is also connected. In various embodiments, network interface 1320 can include a wired interface (e.g., Ethernet) and / or a wireless interface implementing various RF data communication standards such as WiFi, Bluetooth, or cellular data network standards (e.g., 3G, 4G, LTE, etc.).

[0260] User input device 1322 can include any device (or devices) via which a user can provide signals to client computing system 1314; client computing system 1314 can interpret the signals as indicative of particular user requests or information. In various embodiments, user input device 1322 can include any or all of a keyboard, touch pad, touch screen, mouse or other pointing device, scroll wheel, click wheel, dial, button, switch, keypad, microphone, and so on.

[0261] User output device 1324 can include any device via which client computing system 1314 can provide information to a user. For example, user output device 1324 can include a display to display images generated by or delivered to client computing system 1314. The display can incorporate various image generation technologies, e.g., a liquid crystal display (LCD), lightemitting diode (LED) including organic light-emitting diodes (OLED), projection system, cathode ray tube (CRT), or the like, together with supporting electronics (e.g., digital-to-analog or analog-to-digital converters, signal processors, or the like). Some embodiments can include a device such as a touchscreen that function as both input and output device. In some embodiments, other user output devices 1324 can be provided in addition to or instead of a display. Examples include indicator lights, speakers, tactile “display” devices, printers, and so on.

[0262] Some embodiments include electronic components, such as microprocessors, storage and memory that store computer program instructions in a computer-readable storage medium. Many of the features described in this specification can be implemented as processes that are specifiedN.062.WO.01as a set of program instructions encoded on a computer-readable storage medium. When these program instructions are executed by one or more processing units, they cause the processing unit(s) to perform various operation indicated in the program instructions. Examples of program instructions or computer code include machine code, such as is produced by a compiler, and files including higher-level code that are executed by a computer, an electronic component, or a microprocessor using an interpreter. Through suitable programming, processing unit(s) 1304 and 1316 can provide various functionality for server system 1300 and client computing system 1314, including any of the functionality described herein as being performed by a server or client, or other functionality.

[0263] It will be appreciated that server system 1300 and client computing system 1314 are illustrative and that variations and modifications are possible. Computer systems used in connection with embodiments of the present disclosure can have other capabilities not specifically described here. Further, while server system 1300 and client computing system 1314 are described with reference to particular blocks, it is to be understood that these blocks are defined for convenience of description and are not intended to imply a particular physical arrangement of component parts. For instance, different blocks can be but need not be located in the same facility, in the same server rack, or on the same motherboard. Further, the blocks need not correspond to physically distinct components. Blocks can be configured to perform various operations, e.g., by programming a processor or providing appropriate control circuitry, and various blocks might or might not be reconfigurable depending on how the initial configuration is obtained. Embodiments of the present disclosure can be realized in a variety of apparatus including electronic devices implemented using any combination of circuitry and software.

[0264] Referring now to FIG. 14, FIG. 14 is a component diagram of an example computing system suitable for use in the various implementations described herein, according to an example implementation. For example, the computing system 1400 may implement a computing device 110 or controller 112, or various other example systems and devices described in the present disclosure.

[0265] The computing system 1400 includes a bus 1402 or other communication component for communicating information and a processor 1404 coupled to the bus 1402 for processing information. The computing system 1400 also includes main memory 1406, such as a RAM or other dynamic storage device, coupled to the bus 1402 for storing information, and instructions toN.062.WO.01be executed by the processor 1404. Main memory 1406 can also be used for storing position information, temporary variables, or other intermediate information during execution of instructions by the processor 1404. The computing system 1400 may further include a ROM 1408 or other static storage device coupled to the bus 1402 for storing static information and instructions for the processor 1404. A storage device 1410, such as a solid-state device, magnetic disk, or optical disk, is coupled to the bus 1402 for persistently storing information and instructions.

[0266] The computing system 1400 may be coupled via the bus 1402 to a display 1414, such as a liquid crystal display, or active matrix display, for displaying information to a user. An input device 1412, such as a keyboard including alphanumeric and other keys, may be coupled to the bus 1402 for communicating information, and command selections to the processor 1404. In another implementation, the input device 1412 has a touch screen display. The input device 1412 can include any type of biometric sensor, or a cursor control, such as a mouse, a trackball, or cursor direction keys, for communicating direction information and command selections to the processor 1404 and for controlling cursor movement on the display 1414.

[0267] In some implementations, the computing system 1400 may include a communications adapter 1416, such as a networking adapter. Communications adapter 1416 may be coupled to bus 1402 and may be configured to enable communications with a computing or communications network or other computing systems. In various illustrative implementations, any type of networking configuration may be achieved using communications adapter 1416. such as wired (e.g., via Ethernet), wireless (e.g., via Wi-Fi, Bluetooth), satellite (e.g., via GPS) pre-configured, ad-hoc, LAN, WAN, and the like.

[0268] According to various implementations, the processes of the illustrative implementations that are described herein can be achieved by the computing system 1400 in response to the processor 1404 executing an implementation of instructions contained in main memory 1406. Such instructions can be read into main memory 1406 from another computer-readable medium, such as the storage device 1410. Execution of the implementation of instructions contained in main memory 1406 causes the computing system 1400 to perform the illustrative processes described herein. One or more processors in a multi-processing implementation may also be employed to execute the instructions contained in main memory 1406. In alternative implementations, hard-wired circuitry may be used in place of or in combination with software instructions to implementN.062.WO.01illustrative implementations. Thus, implementations are not limited to any specific combination of hardware circuitry and software.

[0269] Example embodiments of the disclosed approach include, without limitation, the following:

[0270] Embodiment Al. A method for preparing a composition of amplified DNA from a sample useful for monitoring and detecting a tumor status in a cancer patient, wherein the method comprises the steps of: (a) extracting cell-free DNA from the sample or a fraction thereof; (b) preparing a composition of amplified DNA by performing a multiplex amplification reaction on cell-free DNA extracted in (a) or DNA derived therefrom to obtain a set of amplicons encompassing a first set of target loci; (c) performing high-throughput sequencing of at least some of the amplicons to obtain sequence reads of at least a fraction of the set of amplicons encompassing the first set of target loci; and (d) generating a ctDNA positivity classification based on an analysis of a first result from a first machine learning (ML) model and a second result from a second ML model that is different from the first ML model, wherein the first ML model receives a first feature set and the second ML model receives a second feature set different from the first feature set, and wherein the first and second feature sets are based at least in part on a fraction or all of the first set of target loci.

[0271] Embodiment A2: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein the analysis aggregates the first result from the first model and the second result from the second model to generate the ctDNA positivity classification.

[0272] Embodiment A3: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein the first feature set comprises a first feature that is not in the second feature set, and / or the second feature set comprises a second feature that is not included in the first feature set.

[0273] Embodiment A4: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein the ctDNA positivity classification is based on sequence reads for the first set of target loci and sequence reads for one or more additional target loci from a second set of target loci different from the first set of target loci, wherein the second set of target loci are encompassed by the set of amplicons.N.062.WO.01

[0274] Embodiment A5: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein the sequence reads from the second set of target loci are used to at least one of calibrate, predict, or validate the ctDNA positivity classification.

[0275] Embodiment A6: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein the first set of target loci are tumorspecific target loci and the second set of target loci is different from the first set of target loci.

[0276] Embodiment A7: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein training of the first ML model comprised generation of a first training feature set, and wherein training of the second ML model comprised generation of a second training feature set which differs from the first training feature set.

[0277] Embodiment A8: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein the first set of target loci comprises at least 50 loci.

[0278] Embodiment A9: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein the first ML model employs a first algorithm and the second ML model employs a second algorithm different from the first algorithm.

[0279] Embodiment A10: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein at least one of the first algorithm, the second algorithm, or the analysis comprises at least one of Platt scaling, geometric averaging, voting scheme, or gradient boosting.

[0280] Embodiment B 1: A method for preparing a composition of amplified DNA from a sample useful for monitoring and detecting a tumor status in a cancer patient, wherein the method comprises the steps of: (a) extracting cell-free DNA from the sample or a fraction thereof: (b) preparing a composition of amplified DNA by performing a multiplex amplification reaction on cell-free DNA extracted in (a) or DNA derived therefrom to obtain a set of amplicons encompassing a first set of target loci and a second set of target loci, wherein the first set of target loci are tumor-specific target loci and the second set of target loci is different from the first set of target loci; (c) performing high-throughput sequencing of at least some of the amplicons to obtain sequence reads of at least a fraction of the set of amplicons encompassing the first set of target lociN.062.WO.01and at least a fraction of the set of amplicons encompassing the second set of target loci; and (d) generating a ctDNA positivity classification based on an analysis of a first result from a first machine learning (ML) model and a second result from a second ML model, wherein the first ML model receives a first feature set and the second ML model receives a second feature set, wherein the first and second feature sets are based at least in part on a fraction or all of the first set of target loci, and wherein the second set of target loci is used in validation of the ctDNA positivity classification.

[0281] Embodiment B2: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein the first feature set is based on a first feature set corresponding to the first ML model, and the second feature set is based on a second feature set generated for the second ML model.

[0282] Embodiment B3: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein the first feature set comprises a first feature that is not in the second feature set, and the second feature set comprises a second feature that is not included in the first feature set.

[0283] Embodiment B4: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein the analysis aggregates the first result from the first model and the second result from the second model to generate the ctDNA positivity classification.

[0284] Embodiment B5: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein the first set of target loci are tumorspecific target loci and the second set of target loci is different from the first set of target loci.

[0285] Embodiment B6: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein training of the first ML model comprised generation of a first feature set, and wherein training of the second ML model comprised generation of a second feature set different from the first feature set.

[0286] Embodiment B7: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein the first ML model employs a first algorithm and the second ML model employs a second algorithm different from the first algorithm.

[0287] Embodiment B8: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein at least one of the first algorithm,N.062.WO.01the second algorithm, or the analysis comprises at least one of Platt scaling, geometric averaging, voting scheme, or gradient boosting.

[0288] Embodiment B9: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein the first set of target loci encompasses a first number of loci, and the second set of target loci encompasses a second number of target loci that is at least one order of magnitude greater than the first number of target loci.

[0289] Embodiment B10: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein the ctDNA positivity classification is indicative of the subject developing or having cancer.

[0290] Embodiment B11: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein the analysis comprises ensembling the first result with the second result to generate a score for the sample, the score for the sample is determined using a threshold, and the ctDNA positivity classification is based on the score.

[0291] Embodiment B12: The method of any “A” Embodiment. “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein the first and second ML models were trained at least in part by generating a first training feature set for the first ML model and generating a second training feature set that is different from the first training feature set for the second ML model.

[0292] Embodiment B13: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein the first and second ML models were trained at least in part by generating genetic data for a set of samples from a cohort of subjects.

[0293] Embodiment B14: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein generating the genetic data comprised: obtaining the set of samples; stratifying the set of samples into DNA input groups based on corresponding genetic data; and generating augmented samples corresponding to each DNA input group by repeatedly selecting and combining, for each DNA input group, a plurality of samples from the set of samples to obtain an augmented sample.

[0294] Embodiment B15: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein training the first ML model comprised a first feature generation algorithm, wherein the first feature generation algorithm comprised: obtaining a probability function by determining, for each target in the first set of targetN.062.WO.01loci and the second set of target loci, a probability that a mutation is present; deriving a first set of target features from the genetic data based on the probability function; and deriving a second set of target features different from the first set of target features.

[0295] Embodiment B16: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, or “D” Embodiment, or “E” Embodiment, wherein training the second ML model comprised a second feature generation algorithm, wherein the second feature generation algorithm comprises: obtaining a set of posterior probabilities by determining, for each target in the first set of target loci and the second set of target loci, a probability that a mutation is present; deriving a first set of target features from the genetic data based on the set of posterior probabilities; deriving a second set of target features from the genetic data different from the first set of target features.

[0296] Embodiment B17: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein at least one of the first ML model, the second ML model, and / or the analysis comprises at least one Platt scaling, geometric averaging, voting scheme, or gradient boosting.

[0297] Embodiment B18: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein the first ML model or the second ML model was trained as a binary classification-based algorithm.

[0298] Embodiment B19: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein at least one of the first ML model or the second ML model was trained by optimizing sensitivity given a specificity constraint.

[0299] Embodiment B20: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, further comprising training the first and second ML models.

[0300] Embodiment B21: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein: (i) the first ML model and the second ML model are gradient boosting decision tree models; (ii) the first feature set comprises features derived from a probability function that, for each target in the first set of target loci, represents a probability that a mutation is present and comprises sample-level probabilities that exactly k targets are true mutations for a plurality of values of k; (iii) the second feature set comprises features derived from posterior probabilities that, for each target in the first set of target loci, represent a probability that a mutation is present and comprises at least one of: (a) a set of topN.062.WO.01position posterior probabilities, (b) counts of positions per sample with posterior probabilities above one or more thresholds, or (c) inter-quantile range features of the posterior probability distribution; and (iv) generating the ctDNA positivity classification comprises applying Platt scaling to outputs of the first ML model and the second ML model, combining the Platt-scaled outputs using a geometric average to obtain an ensembled score, and comparing the ensembled score to a threshold selected based on a specificity constraint to classify the sample as ctDNA-positive or ctDNA-negative.

[0301] Embodiment Cl: A method implemented by a first computing system comprising one or more processors, the method comprising: generating genetic data from a biological sample of a subject, the genetic data comprising information on a first set of target loci corresponding to cell-free DNA (cfDNA) in the sample; generating a ctDNA positivity classification based on the genetic data, wherein generating the ctDNA positivity classification comprises: providing a first feature set based on the genetic data to a first machine learning (ML) model to obtain a first result, the first ML model employing a first feature set; providing a second feature set based on the genetic data to a second ML model to obtain a second result, the second ML model employing a second feature set; and aggregating the first result and the second result to obtain an aggregated result; and outputting the ctDNA positivity classification, wherein outputting the ctDNA positivity classification comprises at least one of storing the ctDNA positivity classification in a non-transitory computer-readable storage medium, displaying the ctDNA positivity classification on one or more display screens of the first computing system, or transmitting the ctDNA positivity classification to a second computing system via a network.

[0302] Embodiment C2: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein generating the ctDNA positivity classification further comprises validating the aggregated result based on a second set of target loci.

[0303] Embodiment DI: A method implemented by a first computing system comprising one or more processors, the method comprising: generating genetic data from a biological sample of a subject, the genetic data comprising information on a first set of target loci corresponding to cell-free DNA (cfDNA) in the sample and a second set of target loci corresponding to the cfDNA in the sample, wherein the first set of target loci comprises tumor-specific target loci and the second set of target loci is different from the first set of target loci; generating a ctDNA positivityN.062.WO.01classification based on the genetic data, wherein generating the ctDNA positivity classification comprises: providing a first feature set based on the genetic data to a first machine learning (ML) model to obtain a first result, the first feature set being based on the first set of target loci; providing a second feature set based on the genetic data to a second ML model to obtain a second result, the second feature set being based on the first set of target loci; aggregating the first result and the second result to obtain an aggregated result; and validating the aggregated result based on the second set of target loci; and outputting the ctDNA positivity classification, wherein outputting the ctDNA positivity classification comprises at least one of storing the ctDNA positivity classification in a non-transitory computer-readable storage medium, displaying the ctDNA positivity classification on one or more display screens of the first computing system, or transmitting the ctDNA positivity classification to a second computing system via a network.

[0304] Embodiment D2: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein the first ML model is based on a first set of features and the second ML model is based on a second set of features that is different from the first set of features.

[0305] Embodiment D3: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein aggregating the first result and the second result comprises applying Platt scaling to outputs of the first ML model and the second ML model, combining the Platt scaled outputs using a geometric average to obtain an ensembled score, and comparing the ensembled score to a threshold selected based on a specificity constraint to generate the ctDNA positivity classification.

[0306] Embodiment D4: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein validating the aggregated result based on the second set of target loci comprises using sequence reads from the second set of target loci as predictor features and / or error model inputs within the same analytical workflow to adjust or confirm the ctDNA positivity classification at a predetermined specificity level.

[0307] Embodiment El: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein the analysis comprises an ensemble of two or more machine-learning models.N.062.WO.01

[0308] Embodiment E2: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein a first set of target loci comprises at least 64 loci, at least 128 loci, or at least 256 loci.

[0309] Embodiment E3: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein a second set of target loci provides one or more of internal calibration inputs, background or error modeling features, predictor features, or confidence-related inputs.

[0310] Embodiment E4: The method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment, wherein sequence reads corresponding to the second set of target loci are processed to compute a sequencing-noise parameter, and wherein the sequencing-noise parameter is applied to modify a classification threshold used to generate the ctDNA positivity classification.

[0311] Embodiment Fl: A computing system comprising one or more processors and one or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the system to perform any method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment.

[0312] Embodiment F2: A system comprising one or more processors and one or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the system to: generate genetic data from a biological sample of a subject, the genetic data comprising information on a first set of target loci and a second set of target loci corresponding to cell-free DNA (cfDNA); generate a ctDNA positivity classification by providing features based on the genetic data to one or more machine-learning models and aggregating model outputs; and output the ctDNA positivity classification.

[0313] Embodiment Gl: A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause a computing system to perform operations for any method of any “A” Embodiment, “B” Embodiment, “C” Embodiment, “D” Embodiment, or “E” Embodiment.

[0314] Embodiment G2: A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause a computing system to perform operations comprising: receiving genetic data for a sample comprising information on a first set of target loci and a second set of target loci; generating features based on the genetic data; providing the featuresN.062.WO.01to one or more machine-learning models to obtain model outputs; aggregating the model outputs to generate a ctDNA positivity classification; and providing the ctDNA positivity classification for display, storage, or transmission.

[0315] While the disclosure has been described with respect to specific embodiments, one skilled in the art will recognize that numerous modifications are possible. Embodiments of the disclosure can be realized using a variety of computer systems and communication technologies including but not limited to the specific examples described herein. Embodiments of the present disclosure can be realized using any combination of dedicated components and / or programmable processors and / or other programmable devices. The various processes described herein can be implemented on the same processor or different processors in any combination. Where components are described as being configured to perform certain operations, such configuration can be accomplished, e.g., by designing electronic circuits to perform the operation, by programming programmable electronic circuits (such as microprocessors) to perform the operation, or any combination thereof. Further, while the embodiments described above may make reference to specific hardware and software components, those skilled in the art will appreciate that different combinations of hardware and / or software components may also be used and that particular operations described as being implemented in hardware might also be implemented in software or vice versa.

[0316] Computer programs incorporating various features of the present disclosure may be encoded and stored on various computer-readable storage media; suitable media include magnetic disk or tape, optical storage media such as compact disk (CD) or DVD (digital versatile disk), flash memory, and other non-transitory media. Computer-readable media encoded with the program code may be packaged with a compatible electronic device, or the program code may be provided separately from electronic devices (e.g., via Internet download or as a separately packaged computer-readable storage medium).

[0317] Thus, although the disclosure has been described with respect to specific embodiments, it will be appreciated that the disclosure is intended to cover all modifications and equivalents within the scope of the following claims. In particular, while various illustrative embodiments incorporating the principles of the present teachings have been disclosed, the present teachings are not limited to the disclosed embodiments. Instead, this application is intended to cover any variations, uses, or adaptations of the present teachings and use its general principles. Further, thisN.062.WO.01application is intended to cover such departures from the present disclosure as come within known or customary practice in the art to which these teachings pertain.

[0318] In the above detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, similar symbols typically identify similar components, unless context dictates otherwise. The illustrative embodiments described in the present disclosure are not meant to be limiting. Other embodiments may be used, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that various features of the present disclosure, as generally described herein, and illustrated in the Figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein.

[0319] The present disclosure is not to be limited in terms of the particular embodiments described in this application, which are intended as illustrations of various features. Many modifications and variations can be made without departing from its spirit and scope, as will be apparent to those skilled in the art. Functionally equivalent methods and apparatuses within the scope of the disclosure, in addition to those enumerated herein, will be apparent to those skilled in the art from the foregoing descriptions. It is to be understood that this disclosure is not limited to particular methods, reagents, compounds, compositions or biological systems, which can, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting.

[0320] With respect to the use of substantially any plural and / or singular terms herein, those having skill in the art can translate from the plural to the singular and / or from the singular to the plural as is appropriate to the context and / or application. The various singular / plural permutations may be expressly set forth herein for sake of clarity.

[0321] It will be understood by those within the art that, in general, terms used herein are generally intended as “open” terms (for example, the term “including” should be interpreted as “including but not limited to,” the term “having” should be interpreted as “having at least,” the term “includes” should be interpreted as “includes but is not limited to,” et cetera). While various compositions, methods, and devices are described in terms of “comprising” various components or steps (interpreted as meaning “including, but not limited to”), the compositions, methods, and devices can also “consist essentially of’ or “consist of’ the various components and steps, and such terminology should be interpreted as defining essentially closed-member groups.N.062.WO.01

[0322] In addition, even if a specific number is explicitly recited, those skilled in the art will recognize that such recitation should be interpreted to mean at least the recited number (for example, the bare recitation of “two recitations,” without other modifiers, means at least two recitations, or two or more recitations). Furthermore, in those instances where a convention analogous to “at least one of A, B, and C, et cetera” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (for example, “a system having at least one of A, B, and C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, et cetera). In those instances where a convention analogous to “at least one of A, B, or C, et cetera” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (for example, “a system having at least one of A, B, or C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, et cetera). It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, sample embodiments, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase “A or B” will be understood to include the possibilities of “A” or “B” or “A and B.”

[0323] In addition, where features of the disclosure are described in terms of Markush groups, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual member or subgroup of members of the Markush group.

[0324] As will be understood by one skilled in the art. for any and all purposes, such as in terms of providing a written description, all ranges disclosed herein also encompass any and all possible subranges and combinations of subranges thereof. Any listed range can be easily recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, tenths, et cetera. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, et cetera. As will also be understood by one skilled in the art all language such as “up to,” “at least,” and the like include the number recited and refer to ranges that can be subsequently broken down into subranges as discussed above. Finally, as will be understood by one skilled in the art, a range includes each individual member. Thus, for example, a group having 1–3 values refers to groups having 1, 2, orN.062.WO.013 values. Similarly, a group having 1–5 values refers to groups having 1, 2, 3, 4, or 5 values, and so forth.

[0325] The term “about.” as used herein, refers to variations in a numerical quantity that can occur, for example, through measuring or handling procedures in the real world; through inadvertent error in these procedures; through differences in the manufacture, source, or purity of compositions or reagents; and the like.

[0326] Various of the above-disclosed and other features and functions, or alternatives thereof, may be combined into many other different systems or applications. Various presently unforeseen or unanticipated alternatives, modifications, variations or improvements therein may be subsequently made by those skilled in the art, each of which is also intended to be encompassed by the disclosed embodiments.

[0327] The functions and process steps herein may be performed automatically or wholly or partially in response to user command. An activity (including a step) performed automatically is performed in response to one or more executable instructions or device operation without user direct initiation of the activity.

Claims

N.062.WO.01CLAIMS1. A method for preparing a composition of amplified DNA from a sample useful for monitoring and detecting a tumor status in a cancer patient, wherein the method comprises the steps of:(a) extracting cell-free DNA from the sample or a fraction thereof;(b) preparing a composition of amplified DNA by performing a multiplex amplification reaction on cell-free DNA extracted in (a) or DNA derived therefrom to obtain a set of amplicons encompassing a first set of target loci;(c) performing high-throughput sequencing of at least some of the amplicons to obtain sequence reads of at least a fraction of the set of amplicons encompassing the first set of target loci; and(d) generating a ctDNA positivity classification based on an analysis of a first result from a first machine learning (ML) model and a second result from a second ML model that is different from the first ML model, wherein the first ML model receives a first feature set and the second ML model receives a second feature set different from the first feature set, and wherein the first and second feature sets are based at least in part on a fraction or all of the first set of target loci.

2. The method of claim 1, wherein the analysis aggregates the first result from the first model and the second result from the second model to generate the ctDNA positivity classification.

3. The method of claim 1 or claim 2, wherein the first feature set comprises a first feature that is not in the second feature set, and / or the second feature set comprises a second feature that is not included in the first feature set.

4. The method of any one of claims 1–3, wherein the ctDNA positivity classification is based on sequence reads for the first set of target loci and sequence reads for one or more additional target loci from a second set of target loci different from the first set of target loci, wherein the second set of target loci are encompassed by the set of amplicons.N.062.WO.

015. The method of claim 4, wherein the sequence reads from the second set of target loci are used to at least one of calibrate, predict, or validate the ctDNA positivity classification.

6. The method of claim 4 or claim 5, wherein the first set of target loci are tumor-specific target loci and the second set of target loci is different from the first set of target loci.

7. The method of any one of claims 1–6, wherein training of the first ML model comprised generation of a first training feature set, and wherein training of the second ML model comprised generation of a second training feature set which differs from the first training feature set.

8. The method of any one of claims 1–7, wherein the first set of target loci comprises at least 50 loci.

9. The method of any one of claims 1–8, wherein the first ML model employs a first algorithm and the second ML model employs a second algorithm different from the first algorithm.

10. The method of claim 9, wherein at least one of the first algorithm, the second algorithm, or the analysis comprises at least one of Platt scaling, geometric averaging, voting scheme, or gradient boosting.

11. A method for preparing a composition of amplified DNA from a sample useful for monitoring and detecting a tumor status in a cancer patient, wherein the method comprises the steps of:(a) extracting cell-free DNA from the sample or a fraction thereof;(b) preparing a composition of amplified DNA by performing a multiplex amplification reaction on cell-free DNA extracted in (a) or DNA derived therefrom to obtain a set of amplicons encompassing a first set of target loci and a second set of target loci, wherein the first set of target loci are tumor- specific target loci and the second set of target loci is different from the first set of target loci;N.062.WO.01(c) performing high-throughput sequencing of at least some of the amplicons to obtain sequence reads of at least a fraction of the set of amplicons encompassing the first set of target loci and at least a fraction of the set of amplicons encompassing the second set of target loci; and (d) generating a ctDNA positivity classification based on an analysis of a first result from a first machine learning (ML) model and a second result from a second ML model, wherein the first ML model receives a first feature set and the second ML model receives a second feature set, wherein the first and second feature sets are based at least in part on a fraction or all of the first set of target loci, and wherein the second set of target loci is used in validation of the ctDNA positivity classification.

12. The method of claim 11, wherein the first feature set is based on a first feature set corresponding to the first ML model, and the second feature set is based on a second feature set generated for the second ML model.

13. The method of claim 11 or claim 12, wherein the first feature set comprises a first feature that is not in the second feature set, and the second feature set comprises a second feature that is not included in the first feature set.

14. The method of any one of claims 11–13, wherein the analysis aggregates the first result from the first model and the second result from the second model to generate the ctDNA positivity classification.

15. The method of any one of claims 11–14, wherein the first set of target loci are tumorspecific target loci and the second set of target loci is different from the first set of target loci.

16. The method of any one of claims 11–15, wherein training of the first ML model comprised generation of a first feature set, and wherein training of the second ML model comprised generation of a second feature set different from the first feature set.N.062.WO.0117. The method of any one of claims 11–16, wherein the first ML model employs a first algorithm and the second ML model employs a second algorithm different from the first algorithm.

18. The method of claim 17, wherein at least one of the first algorithm, the second algorithm, or the analysis comprises at least one of Platt scaling, geometric averaging, voting scheme, or gradient boosting.

19. The method of any one of claims 11–18, wherein the first set of target loci encompasses a first number of loci, and the second set of target loci encompasses a second number of target loci that is at least one order of magnitude greater than the first number of target loci.

20. The method of any one of claims 11–19, wherein the ctDNA positivity classification is indicative of the subject developing or having cancer.

21. The method of any one of claims 1–20, wherein the analysis comprises ensembling the first result with the second result to generate a score for the sample, the score for the sample is determined using a threshold, and the ctDNA positivity classification is based on the score.

22. The method of any one of claims 1–21, wherein the first and second ML models were trained at least in part by generating a first training feature set for the first ML model and generating a second training feature set that is different from the first training feature set for the second ML model.

23. The method of any one of claims 1–22, wherein the first and second ML models were trained at least in part by generating genetic data for a set of samples from a cohort of subjects.

24. The method of claim 23, wherein generating the genetic data comprised:obtaining the set of samples;stratifying the set of samples into DNA input groups based on corresponding genetic data; andN.062.WO.01generating augmented samples corresponding to each DNA input group by repeatedly selecting and combining, for each DNA input group, a plurality of samples from the set of samples to obtain an augmented sample.

25. The method of claim 23 or claim 24, wherein training the first ML model comprised a first feature generation algorithm, wherein the first feature generation algorithm comprised: obtaining a probability function by determining, for each target in the first set of target loci and the second set of target loci, a probability that a mutation is present;deriving a first set of target features from the genetic data based on the probability function; andderiving a second set of target features different from the first set of target features.

26. The method of any one of claims 23–25, wherein training the second ML model comprises a second feature generation algorithm, wherein the second feature generation algorithm comprises:obtaining a set of posterior probabilities by determining, for each target in the first set of target loci and the second set of target loci, a probability that a mutation is present;deriving a first set of target features from the genetic data based on the set of posterior probabilities;deriving a second set of target features from the genetic data different from the first set of target features.

27. The method of any one of claims 1–26, wherein at least one of the first ML model, the second ML model, and / or the analysis comprises at least one Platt scaling, geometric averaging, voting scheme, or gradient boosting.

28. The method of any one of claims 1–27, wherein the first ML model or the second ML model was trained as a binary classification-based algorithm.

29. The method of any one of claims 1–28, wherein at least one of the first ML model or the second ML model was trained by optimizing sensitivity given a specificity constraint.N.062.WO.0130. The method of any one of claims 1–29, further comprising training the first and second ML models.

31. A method implemented by a first computing system comprising one or more processors, the method comprising:generating genetic data from a biological sample of a subject, the genetic data comprising information on a first set of target loci corresponding to cell-free DNA (cfDNA) in the sample;generating a ctDNA positivity classification based on the genetic data, wherein generating the ctDNA positivity classification comprises:providing a first feature set based on the genetic data to a first machine learning (ML) model to obtain a first result, the first ML model employing a first feature set;providing a second feature set based on the genetic data to a second ML model to obtain a second result, the second ML model employing a second feature set; and aggregating the first result and the second result to obtain an aggregated result; andoutputting the ctDNA positivity classification, wherein outputting the ctDNA positivity classification comprises at least one of storing the ctDNA positivity classification in a non-transitory computer-readable storage medium, displaying the ctDNA positivity classification on one or more display screens of the first computing system, or transmitting the ctDNA positivity classification to a second computing system via a network.

32. The method of claim 31, wherein generating the ctDNA positivity classification further comprises validating the aggregated result based on a second set of target loci.

33. A method implemented by a first computing system comprising one or more processors, the method comprising:generating genetic data from a biological sample of a subject, the genetic data comprising information on a first set of target loci corresponding to cell-free DNA (cfDNA) in the sample and a second set of target loci corresponding to the cfDNA in the sample, wherein the first set of target loci comprises tumor- specific target loci and the second set of target loci is different from the first set of target loci;N.062.WO.01generating a ctDNA positivity classification based on the genetic data, wherein generating the ctDNA positivity classification comprises:providing a first feature set based on the genetic data to a first machine learning (ML) model to obtain a first result, the first feature set being based on the first set of target loci;providing a second feature set based on the genetic data to a second ML model to obtain a second result, the second feature set being based on the first set of target loci;aggregating the first result and the second result to obtain an aggregated result; andvalidating the aggregated result based on the second set of target loci; and outputting the ctDNA positivity classification, wherein outputting the ctDNA positivity classification comprises at least one of storing the ctDNA positivity classification in a non-transitory computer-readable storage medium, displaying the ctDNA positivity classification on one or more display screens of the first computing system, or transmitting the ctDNA positivity classification to a second computing system via a network.

34. The method of claim 33, wherein the first ML model is based on a first set of features and the second ML model is based on a second set of features that is different from the first set of features.