Methods and compositions for analyzing free biomarkers
By combining peptide detection with cfDNA methylation pattern assessment and using machine learning algorithms to summarize probability scores, the problem of insufficient sensitivity of DNA methylation sequencing in cancer detection in existing technologies has been solved, achieving early and highly sensitive cancer identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GRAIL INC
- Filing Date
- 2024-08-23
- Publication Date
- 2026-05-05
AI Technical Summary
Existing DNA methylation sequencing methods, such as WGBS, lack sufficient sensitivity in cancer detection and cannot effectively identify tumors in the early stages. Furthermore, most genomes lack differential methylation signals in cancer, resulting in unsatisfactory detection results.
By combining peptide detection with methylation pattern assessment in cfDNA, a classifier is trained to analyze the methylation and peptide levels of target molecules, thereby improving the detection sensitivity of cancer biomarkers. A machine learning algorithm is then used to aggregate probability scores to identify cancer.
It enables highly sensitive cancer detection at an early stage, improves the accuracy of cancer sample identification, and enhances the specificity and sensitivity of individual analytes.
Smart Images

Figure CN121986175A_ABST
Abstract
Description
[0001] Cross-referencing This application claims the benefit of U.S. Provisional Application No. 63 / 578,347, filed August 23, 2023, and U.S. Provisional Application No. 63 / 549,406, filed February 2, 2024, which are incorporated herein by reference in their entirety for all purposes.
[0002] Technical Field and Background Technology Cancer is a prominent global public health problem. Screening programs and early diagnosis have a significant impact on improving disease-free survival and reducing mortality in cancer patients. Because non-invasive methods for early diagnosis promote patient adherence, they can be incorporated into screening programs.
[0003] DNA methylation plays a crucial role in regulating gene expression. Aberrant DNA methylation is involved in many disease processes, including cancer. DNA methylation profiling using methylation sequencing (e.g., whole-genome bisulfite sequencing (WGBS)) is increasingly recognized as an important diagnostic tool for detecting, diagnosing, and / or monitoring cancer. For example, specific patterns in differentially methylated regions can serve as molecular biomarkers for various diseases. However, because most genomes lack differential methylation in cancer, or the local CpG density is too low to provide a robust signal, only a few percent of the genome may be usable for classification, making WGBS less than ideal for product assays.
[0004] Cancer remains a leading cause of death worldwide. While treatment options have improved over the past few decades, survival rates remain low. Successful treatments through surgical resection and drug-based approaches heavily rely on the identification of early-stage tumors. However, many current detection methods often fail to identify tumors before they reach later stages of the disease. Summary of the Invention
[0005] Given the above, a non-invasive diagnostic approach that can identify disease at its earliest stages remains necessary even when therapeutic interventions have a greater chance of success. The aspects disclosed herein address this need and offer additional advantages. For example, some aspects provided herein relate to a method for non-invasively and cost-effectively combining peptide detection with assessment of methylation patterns in cfDNA for the detection of cancer-related biomarkers in samples from subjects. By analyzing data relating both methylation patterns and peptide levels in cfDNA, higher sensitivity for detecting cancer biomarkers can be achieved with a specificity equal to or greater than that of using either analyte alone. This improvement in detection allows for the identification of true positive cancer samples that might be missed by analyzing either analyte alone.
[0006] In one aspect, this disclosure provides a method for detecting cancer in a subject. In some embodiments, the method for detecting cancer in a subject includes: (a) measuring the level of a first target molecule from a first sample from the subject; (b) measuring the level of a second target molecule from a second sample from the subject; (c) applying a trained classifier to the measured levels of the first and second target molecules to assign an overall probability score to the cancer; and (d) detecting the cancer by identifying that the overall probability score is above a threshold indicating the presence of the cancer. In some embodiments, the first target molecule comprises cell-free DNA (cfDNA) from a plurality of different target genomic regions that are differentially methylated in at least one of a plurality of cancer types. In some embodiments, the second target molecule comprises a plurality of different polypeptides differentially expressed in at least one of the plurality of cancer types. In some embodiments, applying the trained classifier includes: (i) applying a first trained model to the measured levels of the first target molecules to assign a first probability score to the cancer; (ii) applying a second trained model to the measured levels of the second target molecules to assign a second probability score to the cancer; and (iii) summing the first and second probability scores. In some embodiments, the first sample and the second sample are identical.
[0007] In some embodiments, for reference samples from (1) reference subjects with known cancer and (2) reference subjects without cancer, the trained classifier is trained using a reference first probability score from the first trained model, a reference second probability score from the second trained model, and a reference overall probability score that sums these reference first and reference second probability scores. In some embodiments, the trained classifier assigns an overall probability score to each of a plurality of different cancer types, and detecting the cancer includes identifying the cancer type as the cancer type with the highest overall probability score. In some embodiments, summing the first and second probability scores includes calculating the product of the first and second probability scores for the cancer. In some embodiments, summing the first and second probability scores includes combining the first and second probability scores for the cancer in a linear model.
[0008] In some embodiments, the plurality of distinct target genomic regions comprises at least 1,000, 5,000, 10,000, 20,000, or 30,000 target genomic regions. In some embodiments, the plurality of target genomic regions has a total collective length of at least 50 kb, 100 kb, 500 kb, or 1,000 kb. In some embodiments, each of the plurality of distinct target genomic regions contains at least five methylation sites. In some embodiments, measuring these first target molecules comprises sequencing transformed cfDNA or its amplified products from the plurality of distinct target genomic regions, wherein the transformed cfDNA comprises cfDNA treated with a deamination agent. In some embodiments, the method further comprises treating the cfDNA with the deamination agent, optionally wherein the deamination agent is a cytosine deaminase or bisulfite. In some embodiments, the sequencing yields at least 100,000 sequencing reads.
[0009] In some embodiments, measuring these first target molecules includes enriching the transformed cfDNA or its amplification product to produce an enriched polynucleotide sample. In some embodiments, the enrichment includes capturing the transformed cfDNA or its amplification product with a plurality of corresponding bait oligonucleotides. In some embodiments, the plurality of different target genomic regions used for enrichment by these bait oligonucleotides are genomic regions identified by the first trained model as differentially methylated in at least one of a plurality of cancer types relative to non-cancer tissue or relative to different types of cancer.
[0010] In some embodiments, the plurality of different polypeptides includes at least 5, 10, 25, 50, 100, 200, 500, 1000, 2000, 3000, 5000, or 7500 different polypeptides. In some embodiments, the plurality of different polypeptides includes: (a) a polypeptide identifying proteins selected from List 1; (b) a polypeptide identifying proteins selected from any of Lists 2-19; (c) a polypeptide identifying proteins selected from List 20; or (d) a polypeptide identifying one or more of CHAD, KRT19, MMP12, PTN, SERPINA3, and SPP1.
[0011] In some embodiments, the trained classifier distinguishes between subjects with cancer and subjects without cancer with specificity defined for each of the plurality of cancer types. In some embodiments, the trained classifier has a higher cancer detection sensitivity than each of the first trained model and the second trained model; optionally, the trained classifier has a cancer detection specificity equal to or greater than each of the first trained model and the second trained model. In some embodiments, the trained classifier is a binary classifier, a hybrid model classifier, a multilayer perceptron model classifier, or a logistic regression classifier.
[0012] In some embodiments, the first trained model and / or the second trained model is a binary classifier, a hybrid model classifier, a multilayer perceptron model classifier, or a logistic regression classifier. In some embodiments, the first trained model binarizes the measurement levels of these first target molecules by assigning a first value when a target genomic region is detected and a second value when a target genomic region is not detected. In some embodiments, the second trained model performs a logarithmic transformation on the measurement levels of these second target molecules normalized to control proteins present in known amounts. In some embodiments, (a) the first trained model is trained using the measurement levels of these first target molecules for a first reference sample, (b) the second trained model is trained using the measurement levels of these second target molecules for a second reference sample, and (c) the first reference sample and the second reference sample include samples from reference subjects with known cancer and reference subjects without cancer.
[0013] In some embodiments, the first sample and / or the second sample comprises a biological fluid. In some embodiments, the biological fluid comprises blood, plasma, serum, urine, saliva, pleural fluid, pericardial fluid, cerebrospinal fluid (CSF), peritoneal fluid, or any combination thereof. In some embodiments, the biological fluid comprises blood, blood fractions, plasma, or serum. In some embodiments, the first sample and / or the second sample is a plasma sample.
[0014] In some embodiments, the plurality of cancer types includes at least 10 cancer types. In some embodiments, the plurality of cancer types includes one or more of the following: anorectal cancer, bladder cancer, colorectal cancer, esophageal cancer, head and neck cancer, liver cancer, bile duct cancer, lung cancer, ovarian cancer, pancreatic cancer, stomach cancer, breast cancer, prostate cancer, kidney cancer, cervical cancer, endometrial cancer, and blood cancer.
[0015] In some embodiments of any of the methods provided herein, the method further includes treating the subject for that type of cancer. In some embodiments, the treatment includes surgical resection, radiation therapy, chemotherapy, and / or immunotherapy.
[0016] In one aspect, this document also provides a method for training a classifier for detecting target molecules from cancer. In some embodiments, the method for training a classifier for detecting target molecules from cancer includes (a) receiving a first measurement level of a first target molecule for a first sample of a reference subject; (b) training a first model by applying a first machine learning algorithm to these first measurement levels to generate a first probability score of the presence of cancer in the subject; (c) receiving a second measurement level of a second target molecule for a second sample of a reference subject; (d) training a second model by applying a second machine learning algorithm to these second measurement levels to generate a second probability score of the presence of cancer in the subject; (e) generating a reference first cancer probability score for these first samples using the trained first model; (f) generating a reference second cancer probability score for these second samples using the trained second model; (g) generating a reference overall cancer probability score for a plurality of reference subjects by summing the reference first cancer probability score and the reference second cancer probability score for each corresponding reference subject; and (h) training a classifier by applying a third machine learning algorithm to these reference first cancer probability scores, these reference second cancer probability scores, and these reference overall cancer probability scores to generate an overall cancer probability score for the subject. In some embodiments, the first target molecule comprises cell-free DNA (cfDNA) from multiple different target genomic regions that are differentially methylated in at least one of multiple cancer types. In some embodiments, the reference subjects include a first subject with a known cancer type and a second subject without cancer. In some embodiments, the second target molecule comprises multiple different peptides that are differentially expressed in at least one of the multiple cancer types.
[0017] In some embodiments, summing the first cancer probability score and the second cancer probability score includes calculating the product of the first probability score and the second probability score for the cancer. In some embodiments, summing the first probability score and the second probability score includes combining the first probability score and the second probability score for the cancer in a linear model. In some embodiments, the first machine learning algorithm, the second machine learning algorithm, and / or the third machine learning algorithm are L1 regularized logistic regression, L2 regularized logistic regression, a generalized linear model (GLM), a random forest, multinomial logistic regression, a multilayer perceptron, a support vector machine, or a neural network. In some embodiments, the first trained model binarizes the measurement levels of these first target molecules by assigning a first value when a target genomic region is detected and a second value when a target genomic region is not detected. In some embodiments, the second trained model performs a logarithmic transformation on the measurement levels of these second target molecules normalized to control proteins present in known amounts. In some embodiments, the third machine learning algorithm is logistic regression. In some embodiments, (a) the first trained model is trained using measurement levels of the first target molecules for a first reference sample, (b) the second trained model is trained using measurement levels of the second target molecules for a second reference sample, and (c) the first and second reference samples include samples from reference subjects with known cancer and reference subjects without cancer. In some embodiments, at least some of the first and second reference samples are from the same reference subject.
[0018] In some embodiments, the plurality of distinct target genomic regions comprises at least 1,000, 5,000, 10,000, 20,000, or 30,000 target genomic regions. In some embodiments, the plurality of target genomic regions has a total collective length of at least 50 kb, 100 kb, 500 kb, or 1,000 kb. In some embodiments, each of the plurality of distinct target genomic regions contains at least five methylation sites. In some embodiments, the first measurement level comprises sequencing results of cfDNA or its amplicon. In some embodiments, for each of the first samples, the sequencing results comprise at least 100,000 reads.
[0019] In some embodiments, the plurality of different polypeptides includes at least 5, 10, 25, 50, 100, 200, 500, 1000, 2000, 3000, 5000, or 7500 different polypeptides. In some embodiments, the plurality of different polypeptides includes: (a) a polypeptide identifying proteins selected from List 1; (b) a polypeptide identifying proteins selected from any of Lists 2-19; (c) a polypeptide identifying proteins selected from List 20; or (d) a polypeptide identifying one or more of CHAD, KRT19, MMP12, PTN, SERPINA3, and SPP1.
[0020] In some embodiments, the plurality of cancer types includes at least 10 cancer types. In some embodiments, the plurality of cancer types includes one or more of the following: anorectal cancer, bladder cancer, colorectal cancer, esophageal cancer, head and neck cancer, liver cancer, bile duct cancer, lung cancer, ovarian cancer, pancreatic cancer, stomach cancer, breast cancer, prostate cancer, kidney cancer, cervical cancer, endometrial cancer, and blood cancer.
[0021] In one aspect, this document provides a method for screening a subject for cancer, the method comprising (a) measuring the level of a first target molecule from a first sample from the subject, wherein the first target molecule comprises a plurality of different peptides differentially expressed in at least one of a plurality of cancer types; (b) applying a first trained model to the measured level of the first target molecule to assign a first probability score to each of the plurality of cancer types; wherein (i) the first trained model has a first specificity for cancer detection, and (ii) the first probability score for at least one of the cancer types is above a first threshold for the presence of cancer; (c) measuring the level of a second target molecule from a second sample from the subject, wherein the second target molecule comprises cell-free DNA (cfDNA) from a plurality of different target genomic regions that are differentially methylated in at least one of the plurality of cancer types; (d) applying a second trained model to the measured level of the second target molecule to assign a second probability score to the cancer; wherein (ii) the second trained model has a second specificity for cancer detection, and (ii) the second specificity is higher than the first specificity; and (e) detecting the cancer by identifying that the second probability score is higher than the threshold for the presence of cancer. In some embodiments, (a) the first trained model is trained using measurement levels of the first target molecules against a first reference sample, (b) the second trained model is trained using measurement levels of the second target molecules against a second reference sample, and (c) the first and second reference samples include samples from reference subjects with known cancer and reference subjects without cancer. In some embodiments, the first and second samples are identical. In some embodiments, the method further includes treating the subject for the type of cancer (e.g., by surgical resection, radiation therapy, chemotherapy, and / or immunotherapy).
[0022] In some embodiments, the plurality of distinct target genomic regions comprises at least 1,000, 5,000, 10,000, 20,000, or 30,000 target genomic regions. In some embodiments, the plurality of target genomic regions has a total collective length of at least 50 kb, 100 kb, 500 kb, or 1,000 kb. In some embodiments, each of the plurality of distinct target genomic regions contains at least five methylation sites. In some embodiments, measuring these second target molecules comprises sequencing transformed cfDNA or its amplified products from the plurality of distinct target genomic regions, wherein the transformed cfDNA comprises cfDNA treated with a deamination agent. In some embodiments, the sequencing yields at least 100,000 sequencing reads. In some embodiments, measuring these first target molecules comprises enriching the transformed cfDNA or its amplified products to produce an enriched polynucleotide sample. In some embodiments, the enrichment comprises capturing the transformed cfDNA or its amplified products with a plurality of corresponding decoy oligonucleotides. In some embodiments, the plurality of distinct target genomic regions for enrichment by these decoy oligonucleotides are genomic regions identified by the second trained model as differentially methylated in at least one of a plurality of cancer types relative to non-cancer tissue or relative to different types of cancer. In some embodiments, the plurality of distinct polypeptides includes at least 5, 10, 25, 50, 100, 200, 500, 1000, 2000, 3000, 5000, or 7500 distinct polypeptides. In some embodiments, the plurality of distinct polypeptides includes: (a) polypeptides identifying proteins selected from List 1; (b) polypeptides identifying proteins selected from any of Lists 2-19; (c) polypeptides identifying proteins selected from List 20; or (d) polypeptides identifying one or more of CHAD, KRT19, MMP12, PTN, SERPINA3, and SPP1. In some embodiments, the first trained model and / or the second trained model is a neural network classifier, a binary classifier, a hybrid model classifier, a multilayer perceptron model classifier, or a logistic regression classifier. In some embodiments, the second trained model binarizes the measurement levels of these second target molecules by assigning a first value when the target genomic region is detected and a second value when the target genomic region is not detected. In some embodiments, the first trained model performs a logarithmic transformation on the measurement levels of these first target molecules normalized to control peptides present in known amounts. In some embodiments, (a) the first sample and / or the second sample comprises a biological fluid; optionally, the biological fluid comprises blood, plasma, serum, urine, saliva, pleural fluid, pericardial fluid, cerebrospinal fluid (CSF), peritoneal fluid, or any combination thereof; (b) the first sample and / or the second sample is a plasma sample.In some embodiments, (a) the plurality of cancer types includes at least 10 cancer types; and / or (b) the plurality of cancer types includes one or more of anorectal cancer, bladder cancer, colorectal cancer, esophageal cancer, head and neck cancer, liver cancer, bile duct cancer, lung cancer, ovarian cancer, pancreatic cancer, stomach cancer, breast cancer, prostate cancer, kidney cancer, cervical cancer, endometrial cancer, and blood cancer.
[0023] In one aspect, this document also provides a method for treating a subject's cancer. In some embodiments, the method for treating a subject's cancer includes selecting the subject based on the results of a detection or screening assay and treating the subject's cancer, wherein: (a) the detection or screening assay includes a method for detecting or screening a subject's cancer according to any of the various aspects or embodiments described herein; and (b) the treatment includes surgical resection, radiotherapy, chemotherapy, and / or immunotherapy.
[0024] In one aspect, this document provides a non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, perform one or more steps of a method for detecting cancer in a subject according to any of the aspects or embodiments described herein.
[0025] In one aspect, this paper provides a non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, perform a method for training a classifier to detect target molecules from cancer.
[0026] By incorporating via reference All publications, patents and patent applications mentioned in this specification are incorporated herein by reference to the extent that each individual publication, patent or patent application is specifically and individually cited and incorporated herein by reference. Attached Figure Description
[0027] The novel features of this disclosure are set forth in the appended claims. A better understanding of the features and advantages of this disclosure will be obtained by referring to the following detailed description, which illustrates illustrative embodiments utilizing the principles of this disclosure and is accompanied by accompanying drawings: Figure 1A This is a flowchart illustrating the process of sequencing cell-free DNA (cfDNA) fragments according to the description of the embodiments.
[0028] Figure 1B According to the embodiments Figure 1A This diagram illustrates the process of sequencing cfDNA fragments to obtain methylation state vectors.
[0029] Figure 2The process of developing single-analyte models for cancer detection using cfDNA methylation or protein levels in a sample is illustrated.
[0030] Figure 3 Illustrative plots are provided for one-dimensional analysis based on cfDNA methylation (left plot), one-dimensional analysis at the protein level (bottom plot), and a two-dimensional analysis (middle plot) with cancer probability scores based on protein-based cancer probability as the x-axis and cfDNA methylation-based cancer probability as the y-axis. Shaded areas represent samples below a selected threshold, which are designated as non-cancer by the corresponding model. Points corresponding to samples in the white area between the two dashed lines in the middle plot represent additional cancer detections obtained by analyzing both cfDNA and protein (compared to analyzing cfDNA alone with the same specificity). All these detections correspond to samples from cancer subjects. In the left plot, samples from non-cancer subjects and samples from cancer subjects are plotted on the left and right sides, respectively. In the bottom plot, samples from non-cancer subjects and samples from cancer subjects are plotted at the bottom and top, respectively.
[0031] Figure 4 This paper demonstrates a process for developing a multi-omics classifier for cancer detection by combining assessments of cfDNA methylation and protein levels in samples.
[0032] Figure 5A This is an illustrative receiver operating characteristic (ROC) plot, which compares cancer detection by analyzing cfDNA methylation alone or in combination with protein levels in the analyzed sample. The bottom line in each pair of lines corresponds to cfDNA methylation analysis alone, while the top line corresponds to the combination of cfDNA methylation analysis and protein biomarker analysis.
[0033] Figure 5B This graph compares the sensitivity of cancer detection at 99.4% specificity by analyzing cfDNA methylation alone or in combination with protein levels in the analyzed sample. As shown in the figure, the average sensitivity at 99.4% specificity is 0.534 for the combined analysis of cfDNA methylation and protein biomarkers, while the average sensitivity at 99.4% specificity is 0.478 for cfDNA methylation analysis alone.
[0034] Figure 6 This graph compares the sensitivity of cancer detection at 99.4% specificity by analyzing cfDNA methylation in combination with the number of protein markers shown in the analysis. Detailed Implementation
[0035] Before describing the invention in more detail, it should be understood that the invention is not limited to the specific embodiments described, and variations are possible with respect to the embodiments themselves. It should also be understood that the terminology used herein is for the purpose of describing specific embodiments only and is not intended to be limiting, as the scope of the invention will be limited only by the appended claims.
[0036] Where numerical ranges are provided, it should be understood that, unless the context explicitly specifies otherwise, every intermediate value (accurate to one-tenth of the lower limit unit) between the upper and lower limits of the range, as well as any other stated value or intermediate value within the stated range, and each provided endpoint of the range, is covered within the scope of this invention. The upper and lower limits of these smaller ranges may be independently included within the smaller ranges covered by this invention, but are subject to any specific exclusions within the stated range.
[0037] Unless otherwise defined, the technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0038] As used herein, the term "about" means a range of values that includes the specified value and that would be reasonably considered by one of ordinary skill in the art to be similar to the specified value. In embodiments, "about" means within the range of standard deviations using measurements generally accepted in the art. In embodiments, "about" means a range extending to + / - 10% of the specified value. In embodiments, "about" includes the specified value.
[0039] As used herein, the term "methylation" refers to the process of adding a methyl group to a DNA molecule. For example, a hydrogen atom on the pyrimidine ring of a cytosine base can be converted into a methyl group, forming 5-methylcytosine. The term also refers to the process of adding a hydroxymethyl group to a DNA molecule, for example, by oxidizing the methyl group on the pyrimidine ring of a cytosine base. Methylation and hydroxymethylation tend to occur at dinucleotides composed of cytosine and guanine, which are referred to herein as "CpG sites." The principles described herein also apply to the detection of methylation in a non-CpG background, including non-cytosine methylation. In such embodiments, the laboratory wet assay used to detect methylation may differ from any assay described herein. Furthermore, the methylation state vector may contain elements that are typically a vector of sites that have undergone or have not undergone methylation (even if those sites are not specifically CpG sites).
[0040] The term "methylation" can also refer to the methylation state of a CpG site. CpG sites with a 5-methylcytosine moiety are methylated. CpG sites with a hydrogen atom on the pyrimidine ring of the cytosine base are not methylated.
[0041] As used herein, the term "methylation site" refers to a region of a DNA molecule where a methyl group can be added. "CpG" sites are the most common methylation sites, but methylation sites are not limited to CpG sites. For example, DNA methylation can occur on cytosine in CHG and CHH, where H is adenine, cytosine, or thymine. The methods and procedures disclosed herein can also be used to evaluate cytosine methylation in the form of 5-hydroxymethylcytosine (see, for example, US 20110236894 A1 and US 20110301045 A1, which are incorporated herein by reference) and their characteristics.
[0042] As used herein, the term "CpG site" refers to a region in a DNA molecule in which a cytosine nucleotide is followed by a guanine nucleotide in a basic linear sequence along its 5' to 3' direction. "CpG" is short for 5'-C-phosphate-G-3', meaning that cytosine and guanine are separated by only one phosphate group. The cytosine in the CpG dinucleotide can be methylated to form 5-methylcytosine.
[0043] In some embodiments, the oligonucleotide probes described herein comprise one or more CpG detection sites. As used herein, the term "CpG detection site" refers to a region of the probe configured to hybridize with a CpG site on a target DNA molecule. A CpG site on a target DNA molecule may comprise cytosine and guanine separated by one phosphate group, wherein the cytosine is methylated or unmethylated. A CpG site on a target DNA molecule may comprise uracil and guanine separated by one phosphate group, wherein the uracil is generated from unmethylated cytosine.
[0044] The term "UpG" is short for 5'-U-phosphate-G-3', meaning that uracil and guanine are separated by only one phosphate group. UpG can be generated by treating DNA with bisulfite, which converts unmethylated cytosine into uracil. Cytosine can be converted into uracil by other methods, such as chemical modification, synthesis, or enzymatic conversion.
[0045] As used herein, the terms “hypomethylated” or “hypermethylated” refer to the methylation state of a DNA molecule containing multiple CpG sites (e.g., more than 2, 3, 4, 5, 6, 7, 8, 9, 10, etc.), wherein a high percentage of CpG sites (e.g., more than 80%, 85%, 90%, or 95%, or any other percentage in the range of 50%–100%, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97.5% or more, 98% or more, 99% or more, 99.9% or more, or any other numerical percentage in the range of 50%–100% or more, wherein the ranges provided include the 50% and 100% range boundaries) are either unmethylated or methylated. For example, a "hypomethylated" nucleic acid (e.g., cfDNA) fragment can be a fragment having a certain number (e.g., 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 9 or more, 10 or more) of CpG sites, of which a certain proportion (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, or 95% or more, or 97.5% or more, 98% or more, 99% or more, 99.9% or more) of the CpG sites are unmethylated. Similarly, a “hypermethylated” nucleic acid (e.g., cfDNA) fragment can be a fragment having a certain number (e.g., 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 9 or more, 10 or more) of CpG sites, wherein a certain proportion (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, or 95% or more, or 97.5% or more, 98% or more, 99% or more, 99.9% or more) of the CpG sites are methylated. In some embodiments, a hypomethylated DNA molecule contains multiple CpG sites, wherein at least 80% are unmethylated. In some embodiments, a hypermethylated DNA molecule contains multiple CpG sites, wherein at least 80% are methylated. In some embodiments, the level of the first target molecule is the presence, number, or proportion of cfDNA molecules from a target genomic region having a specific methylation level, or the presence, number, or proportion of its sequencing reads.
[0046] As used herein, the term "methylation state vector" or "methylation status vector" refers to a vector containing multiple elements, each element indicating the methylation status of methylation sites in a DNA molecule containing multiple methylation sites arranged in order of their occurrence from 5' to 3' in the DNA molecule. For example,<Mx, Mx+1, Mx+2> ,<Mx, Mx+1, Ux+2> ......<Ux,Ux+1, Ux+2> It can be a methylation vector of a DNA molecule containing three methylation sites, where M represents a methylated methylation site and U represents an unmethylated methylation site. In some embodiments, the level of the first target molecule is the presence, number, or proportion of cfDNA molecules from a target genomic region having a specific methylation state vector, or the presence, number, or proportion of its sequencing reads.
[0047] As used herein, the terms “abnormal methylation pattern” and “anomalous methylation pattern” refer to a methylation pattern of a nucleic acid molecule (e.g., cfDNA molecule) or its methylation state vector that is found more frequently in a sample than is expected in a healthy (e.g., non-cancer) sample. In various embodiments, the probability of finding such a methylation pattern in a healthy (e.g., non-cancer) sample is below a threshold. Thus, for example, as used herein, the terms “abnormally methylated” and “anomalously methylated” describe nucleic acids (e.g., DNA, such as cfDNA), molecules, or methylation state vectors that exhibit abnormal methylation patterns. When abnormal methylation patterns distinguish between cancer and non-cancer and / or one cancer type from another, the corresponding target genomic region may be referred to as “differential methylation.” When referring to the health status of the subject from whom the subject sample was derived, whether a target genomic region is differentially methylated can be used as an indicator to determine a comparison between healthy (e.g., non-cancer) and diseased (e.g., cancer). In some embodiments provided herein, the finding and / or expected finding of a particular methylation state vector in a healthy control group including healthy individuals is represented by a p-value. In some embodiments, a low p-value score corresponds to a relatively unexpected methylation state vector compared to other methylation state vectors in samples from healthy individuals (such as individuals in a healthy control group). In some versions, a high p-value score corresponds to a relatively more expected methylation state vector compared to other methylation state vectors found in samples from healthy individuals (such as individuals in a healthy control group). In various embodiments, a methylation state vector with an anomalous / abnormal methylation pattern is a methylation state vector having a p-value equal to and / or below a threshold (e.g., 0.1, 0.01, 0.001, 0.0001, etc.), such a threshold is a threshold corresponding to a healthy (e.g., non-cancer) sample. In various embodiments, the method includes associating a methylation state vector from a sample having a p-value equal to and / or below a threshold (e.g., 0.1 or less, 0.01 or less, 0.001 or less, 0.0001 or less, etc.) with a determination that the sample is not a healthy sample (e.g., a sample from a subject with cancer). In various embodiments, thresholds are applied as filters because the application of smaller thresholds (e.g., 0.001, 0.0001, etc.) is associated with a higher expectation that the methylation state vector originates from an unhealthy sample (e.g., from an individual with cancer). Various methods can be used to calculate the p-value or expectation of the methylation pattern or methylation state vector. The exemplary method provided herein involves using Markov chain probabilities, which assume that the methylation state of a CpG site depends on the methylation state of neighboring CpG sites.The alternative methods presented herein compute the expected value of a specific methylation state vector observed in a healthy individual by utilizing a hybrid model comprising multiple hybrid components, each of which is an independent site model, wherein it is assumed that the methylation state at each CpG site is independent of the methylation state at other CpG sites. In some versions, the methods of the present invention include determining whether a nucleic acid (e.g., DNA), molecule, or methylation state vector is aberrantly methylated. In various embodiments of these methods, a generated p-value (e.g., by an analysis system) is compared to a threshold to identify vectors (e.g., nucleic acids, such as cfDNA fragments) that are aberrantly methylated relative to a control group (e.g., a group associated with one or more healthy (e.g., non-cancer) samples). Furthermore, aberrant methylation (e.g., cfDNA methylation) can be hypermethylated and / or hypomethylated, both of which can indicate an unhealthy (e.g., cancerous) state. Therefore, the methods include determining a healthy or diseased (e.g., non-cancer or cancerous) state at least in part based on a p-value (e.g., a relatively low p-value, such as a p-value below a threshold), wherein the p-value indicates aberrant methylation, such as hypermethylation and / or hypomethylation, in various respects. Low p-values (e.g., p-values equal to or below a threshold (e.g., 0.1, 0.01, 0.001, 0.0001, etc.)) can indicate anomalous methylation in a sample, such as hypermethylation and / or hypomethylation. In various embodiments, the method includes determining the health or disease (e.g., non-cancer or cancerous) status of a sample based on nucleic acid (e.g., nucleic acid fragments) or methylation vectors from samples having low p-values (e.g., equal to or less than 0.1, 0.01, or 0.001) and simultaneously being hypermethylated and hypomethylated or hypermethylated or hypomethylated. In various aspects, the method includes determining the health or disease (e.g., non-cancer or cancerous) status of a sample based at least in part on whether nucleic acid (e.g., nucleic acid fragments) or methylation vectors from the sample are simultaneously hypermethylated and hypomethylated. In some variations, determining whether a vector (e.g., a sample fragment) is anomalously methylated based on a generated p-value score includes determining whether the generated score of the vector is below a threshold score, where the threshold score is a confidence level that the vector is anomalously methylated.
[0048] As used herein, the term "amplifier" refers to the product of a polynucleotide amplification reaction; that is, a clonal population of polynucleotides replicated from one or more starter sequences, which may be single-stranded or double-stranded. The one or more starter sequences may be one or more copies of the same sequence, or they may be a mixture of different sequences. Preferably, an amplifier is formed by amplifying a single starter sequence. Amplifiers can be produced by a variety of amplification reactions, the products of which contain copies of one or more starter (or target) nucleic acids. In one aspect, the amplification reaction that produces an amplifier is "template-driven" because the base pairings of the reactants (nucleotides or oligonucleotides) have complementary sequences in the template polynucleotide required to produce the reaction product. In another aspect, the template-driven reaction is primer extension using a nucleic acid polymerase, or oligonucleotide ligation using a nucleic acid ligase. Such reactions include, but are not limited to, polymerase chain reaction (PCR), linear polymerase reaction, nucleic acid sequence-based amplification (NASBA), rolling circle amplification, etc., disclosed in the following references, each of which is incorporated herein by reference in its entirety: Mullis et al., U.S. Patent Nos. 4,683,195; 4,965,188; 4,683,202; 4,800,159 (PCR); Gelfand et al., U.S. Patent No. 5,210,015 (Real-time PCR with “TaqMan” probe); Wittwer et al., U.S. Patent No. 6,174,670; Kacian et al., U.S. Patent No. 5,399,491 (“NASBA”); Lizardi, U.S. Patent No. 5,854,033; Aono et al., Japanese Patent Publication JP 4-262799 (Rolling circle amplification); etc. In one aspect, the amplicon of the present invention is generated by PCR. If the detection chemistry for measuring the reaction products during the amplification reaction is available, the amplification reaction can be “real-time” amplification, such as “real-time PCR” or “real-time NASBA” as described in Leone et al., Nucleic Acids Research, 26: 2150-2155 (1998) and similar references.
[0049] As used herein, the term “enrichment” means increasing the proportion of one or more target molecules (e.g., target nucleic acids or target peptides) in a sample. For example, an “enriched” sample or sequencing library is therefore a sample or sequencing library in which the proportion of one or more target nucleic acids is increased relative to the proportion of non-target nucleic acids in the sample. In some embodiments, enrichment includes the physical separation of target molecules from non-target molecules.
[0050] As used herein, the term "cancer sample" refers to a sample containing genomic DNA and / or peptides from an individual with cancer. Genomic DNA can be (but is not limited to) fragments of cfDNA or chromosomal DNA from a subject with cancer. Genomic DNA can be sequenced and its methylation status can be assessed by various methods, such as bisulfite sequencing. When the genomic sequence is obtained from a public database (e.g., The Cancer Genome Atlas (TCGA)) or through sequencing experiments on the genome of an individual diagnosed with cancer, a cancer sample can refer to genomic DNA or fragments of cfDNA with a genomic sequence. Peptides can be detected by any suitable method known in the art. The term "cancer sample" in the plural form refers to a sample containing genomic DNA and / or peptides from multiple individuals, each of whom is an individual with cancer. In various embodiments, cancerous samples from more than 100, 300, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 40,000, 50,000 or more individuals diagnosed with cancer were used.
[0051] As used herein, the terms “non-cancerous sample” or “healthy sample” refer to a sample containing genomic DNA and / or peptides from a healthy individual or an individual not diagnosed with cancer. Genomic DNA can be (but is not limited to) fragments of cfDNA or chromosomal DNA from a subject who does not have cancer. Genomic DNA can be sequenced and its methylation status can be assessed by various methods, such as bisulfite sequencing. When the genomic sequence is obtained from a public database (e.g., The Cancer Genome Atlas (TCGA)) or by sequencing experiments on the genome of an individual who does not have cancer, a non-cancerous sample can refer to genomic DNA or fragments of cfDNA with a genomic sequence. Peptides can be detected by any suitable method known in the art. The term “non-cancerous sample” in the plural form refers to a sample containing genomic DNA and / or peptides from multiple individuals, each of whom is not diagnosed with cancer. In various embodiments, cancerous samples were used from more than 100, 300, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 40,000, 50,000 or more individuals who did not have cancer.
[0052] As used herein, the term "training sample" refers to a sample used to train the model or classifier described herein, select one or more target genomic regions for capture / analysis, and / or determine a probability score for cancer detection. Training samples may contain genomic DNA or modifications thereof, and / or peptides, from one or more healthy subjects and from one or more subjects with a disease condition (e.g., cancer, a specific type of cancer, cancer at a specific stage, etc.). Genomic DNA may be, but is not limited to, cfDNA fragments or chromosomal DNA. Genomic DNA may be sequenced, and its methylation status may be assessed by various methods (e.g., bisulfite sequencing). When the genomic sequence is obtained from a public database (e.g., the Cancer Genome Atlas (TCGA)) or through sequencing experiments on an individual's genome, a training sample may refer to genomic DNA or cfDNA fragments with a genomic sequence. Peptides may be detected by any suitable method known in the art.
[0053] As used herein, the term "test sample" refers to a sample from a subject whose health condition has been, is, or will be tested using the classifiers and / or assay sets described herein. Test samples may contain genomic DNA or modifications thereof, and / or peptides. Genomic DNA may be, but is not limited to, fragments of cfDNA or chromosomal DNA.
[0054] As used herein, the term "target genomic region" refers to a region in the genome selected for analysis in a test sample. An assay set is generated using oligonucleotide probes designed to hybridize (and optionally pull down) with nucleic acid fragments derived from or fragmented from the target genomic region. Oligonucleotide probes targeting the target region are also referred to herein as "decoy oligonucleotides." Nucleic acid fragments derived from the target genomic region refer to nucleic acid fragments generated by degradation, cleavage, bisulfite conversion, or other treatment of DNA from the target genomic region. In some embodiments, multiple different decoy oligonucleotides are designed to hybridize across a single target genomic region (e.g., overlapping probes tiled across the target genomic region). Typically, when referring to multiple target genomic regions, no single target genomic region is completely contained within another target genomic region. Different target genomic regions within multiple target genomic regions may overlap but will have at least different ends. In one embodiment, each target genomic region within multiple target genomic regions is separate from and does not overlap with any other target genomic region within the multiple target genomic regions.
[0055] Various target genomic regions are described according to their chromosomal locations in the sequence listing submitted herein. Chromosomal DNA is double-stranded, therefore target genomic regions consist of two DNA strands: one with the sequence provided in the listing, and the second with the reverse complementary sequence of the listed sequence. Probes can be designed to hybridize with one or both sequences. Optionally, probes hybridize with transformed sequences produced, for example, by treatment with sodium bisulfite.
[0056] As used herein, the term "off-target genomic region" refers to a region of the genome that was not selected for analysis in the test sample but has sufficient homology to a target genomic region to be likely to be bound and captured by a probe designed to target that target genomic region. In one embodiment, an off-target genomic region is a genomic region that aligns with the probe along at least 45 bp and has at least a 90% match.
[0057] The terms “transformed DNA molecule,” “transformed cfDNA molecule,” and “modified fragment obtained from processing these cfDNA molecules” refer to DNA molecules obtained by processing DNA or cfDNA molecules in a sample with the aim of distinguishing between methylated and unmethylated nucleotides in the DNA or cfDNA molecule. For example, in some embodiments, the sample may be treated with bisulfite ions (e.g., using sodium bisulfite) to convert unmethylated cytosine (“C”) to uracil (“U”). In another embodiment, the conversion of unmethylated cytosine to uracil is achieved using an enzymatic conversion reaction (e.g., using a cytidine deaminase such as APOBEC). After treatment, the transformed DNA molecule or cfDNA molecule includes additional uracil not present in the original cfDNA sample. Replication of the DNA strand containing uracil by DNA polymerase results in the addition of adenine to the nascent complementary strand, rather than guanine, which is typically added as a complement to cytosine or methylcytosine.
[0058] Generally, the terms “free,” “circulating,” and “extracellular” are used interchangeably when referring to polynucleotides (e.g., “cell-free DNA” or “cfDNA”), referring to polynucleotides or portions thereof present in a sample from a subject that can be isolated or otherwise manipulated without applying a lysis step (e.g., as in lysis for extraction from cells or viruses) to the original collected sample. Thus, even prior to collecting the subject's sample, the free polynucleotides are not encapsulated or “free” from the cells or viruses of their origin. Free polynucleotides may be produced as a byproduct of cell death (e.g., apoptosis or necrosis) or cell shedding, releasing the polynucleotides into surrounding bodily fluids or circulation. Therefore, free nucleic acids can be isolated from non-cellular fractions of blood (e.g., serum or plasma), from other bodily fluids (e.g., urine), or from non-cellular fractions of other types of samples. In some embodiments, cfDNA refers to deoxyribonucleic acid molecules circulating within the subject's body (e.g., in the bloodstream) and may originate from one or more healthy cells and / or from one or more cancer cells.
[0059] The term "circulating tumor DNA" or "ctDNA" refers to nucleic acid fragments derived from tumor cells that may be released into an individual's bloodstream due to biological processes such as apoptosis or necrosis of dying cells or active release from surviving tumor cells.
[0060] As used herein, the term "fragment" can refer to a fragment of a nucleic acid molecule. For example, in one embodiment, a fragment can refer to a cfDNA molecule in a blood or plasma sample, or a cfDNA molecule that has been extracted from a blood or plasma sample. The amplification product of a cfDNA molecule can also be referred to as a "fragment." In another embodiment, the term "fragment" refers to a sequence read or set of sequence reads that has been processed for subsequent analysis (e.g., for machine learning-based classification), as described herein. For example, a raw sequence read can be aligned to a reference genome, and matched paired-end sequence reads can be assembled into a longer fragment for subsequent analysis.
[0061] The terms “polypeptide,” “peptide,” and “protein” are used interchangeably herein to refer to polymers of amino acids of any length. These terms also cover polymers of amino acids that have been modified; for example, by disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation, such as conjugation with a labeled component. As used herein, the term “amino acid” includes natural and / or non-natural or synthetic amino acids, including glycine and D or L optical isomers, as well as amino acid analogs and peptide mimics. In some embodiments, a polypeptide is encoded by a target polynucleotide or a portion thereof. In some embodiments, a polypeptide is a polypeptide fragment.
[0062] The terms “individual” and “subject” refer to an individual person. The term “healthy individual” refers to an individual who is presumed not to have cancer or disease. In some embodiments, a subject is an individual whose DNA is being analyzed. For example, a subject may be a test subject whose DNA will be evaluated using a set of targets as described herein to assess whether the person has cancer or other disease. In some embodiments, a subject is part of a control group (also referred to as a “reference subject”) known to have (or not have) cancer or other disease. Control groups and cancer / disease groups can be used to assist in the design or validation of target sets.
[0063] As used herein, the term "sequence read" refers to a partial or complete nucleotide string that is identified as a nucleic acid molecule by a nucleic acid sequencing process. A sequence read can be a short nucleotide string (e.g., 20-150) sequenced from a nucleic acid fragment, a short nucleotide string at one or both ends of a nucleic acid fragment, or the sequencing of an entire nucleic acid fragment present in a biological sample. Sequence reads can be obtained by the various methods provided herein or by other methods known in the art.
[0064] As used herein, the term "sequencing depth" refers to the count of the number of times a given target nucleic acid in a sample has been sequenced (e.g., the count of sequence reads at a given target region), or the average number of times that the nucleic acid has actually or is expected to be sequenced based on the amount of nucleic acid being sequenced and the total read length generated by a given sequencing process (e.g., the average read depth of all sequencing regions in a given sequencing run). Increasing sequencing depth can reduce the amount of nucleic acid required to assess disease states (e.g., cancer or cancer-derived tissue).
[0065] As used herein, “treating” or “treatment” includes any means of obtaining a beneficial or desired outcome in a subject’s condition, including clinical outcomes. Beneficial or desired clinical outcomes may include, but are not limited to: reduction or improvement of one or more symptoms or conditions; reduction of disease severity; stabilization (i.e., non-exacerbation) of the disease state; prevention of the spread or diffusion of the disease; delay or slowing of disease progression; improvement or mitigation of the disease state; reduction of disease relapse; and remission (whether partial or complete, and whether detectable or undetectable). In other words, as used herein, “treatment” includes any cure, improvement, or prevention of disease. Treatment may prevent the onset of disease; inhibit the spread of disease; alleviate the symptoms of disease; completely or partially eliminate the root cause of disease; shorten the duration of disease; or achieve a combination of these.
[0066] As used herein, “treating” or “treatment” includes preventative treatment. Treatment methods include administering a therapeutically effective amount of an active agent to a subject. Administration may consist of a single administration or may include a series of administrations. The length of treatment depends on a variety of factors, such as the severity of the condition, the patient’s age, the concentration of the active agent, the activity of the composition used in the treatment, or a combination thereof. It should also be understood that the effective dose of the agent used for treatment or prevention may increase or decrease during the course of a particular treatment or prevention regimen. Changes in dose can be caused and become apparent by standard diagnostic assays known in the art. In some cases, prolonged administration may be required. For example, administering the composition to the subject in an amount and for a duration sufficient to treat the patient. In the examples, treatment or treatment is not preventative treatment.
[0067] The term "prevention," when referring to a subject's disease or condition, means reducing the occurrence of one or more corresponding symptoms in the subject. As shown above, prevention can be complete (no detectable symptoms) or partial, resulting in fewer and / or lower incidence of observed symptoms compared to the absence of treatment.
[0068] The terms "anti-cancer agent" and "anticancer agent" are used in their common and general sense and refer to compositions (e.g., compounds, drugs, antagonists, inhibitors, modulators) that have antitumor properties or the ability to inhibit cell growth or proliferation. In some embodiments, an anticancer agent is a chemotherapeutic agent. In some embodiments, an anticancer agent is a pharmaceutical agent identified herein that is useful in a method of treating cancer. In some embodiments, an anticancer agent is a pharmaceutical agent approved by the FDA or a similar regulatory agency in a country other than the United States for the treatment of cancer. Examples of anticancer agents include, but are not limited to, MEK inhibitors, alkylating agents, antimetabolites, plant alkaloids, topoisomerase inhibitors, antitumor antibiotics, platinum compounds, inhibitors of mitogen-activated protein kinase signaling, and other substances known to those skilled in the art.
[0069] In some embodiments, the anticancer agent is an epigenetic inhibitor. As used herein, “epigenetic inhibitor” means an inhibitor of an epigenetic process such as DNA methylation (DNA methylation inhibitor) or histone modification (histone modification inhibitor). An epigenetic inhibitor may be a histone deacetylase (HDAC) inhibitor, a DNA methyltransferase (DNMT) inhibitor, a histone methyltransferase (HMT) inhibitor, a histone demethylase (HDM) inhibitor, or a histone acetyltransferase (HAT). Examples of HDAC inhibitors include vorinostat, romidesin, CI-994, belistat, pabistal, givinostat, entinostat, mocetinostat, SRT501, CUDC-101, JNJ-26481585, or PCI24781. Examples of DNMT inhibitors include azacitidine and decitabine. Examples of HMT inhibitors include EPZ-5676. Examples of HDM inhibitors include pargyline and tranylcyclopropionamide. Examples of HAT inhibitors include CCT077791 and mangosteen.
[0070] In some embodiments, the anticancer agent is a multi-kinase inhibitor. A "multi-kinase inhibitor" is a small molecule inhibitor of at least one protein kinase, including tyrosine protein kinases and serine / threonine kinases. Multi-kinase inhibitors may include single-kinase inhibitors. Multi-kinase inhibitors can block phosphorylation. Multi-kinase inhibitors can act as covalent modifiers of protein kinases. Multi-kinase inhibitors can bind to the active site of the kinase or to secondary or tertiary sites that inhibit the activity of the protein kinase. Multi-kinase inhibitors can be anticancer multi-kinase inhibitors. Exemplary anticancer multikinase inhibitors include dasatinib, sunitinib, erlotinib, bevacizumab, vastarabine, vemurafenib, vandetanib, cabozantinib, poatinib, axitinib, ruxotinib, regorafenib, crizotinib, besutinib, cetuximab, gefitinib, imatinib, lapatinib, lenvatinib, mulitinib, nilotinib, panitumab, pazopanib, trastuzumab, or sorafenib.
[0071] Methods for detecting cancer In one aspect, this disclosure provides a method for detecting cancer in a subject, the method comprising: (a) measuring the level of a first target molecule from a first sample from the subject; (b) measuring the level of a second target molecule from a second sample from the subject; (c) applying a trained classifier to the measured levels of the first and second target molecules to assign an overall probability score to the cancer; and (d) detecting the cancer by identifying that the overall probability score is above a threshold for the presence of the cancer. In some embodiments, the first target molecule comprises cell-free DNA (cfDNA) from a plurality of different target genomic regions that are differentially methylated in at least one of a plurality of cancer types. In some embodiments, the second target molecule comprises a plurality of different polypeptides differentially expressed in at least one of the plurality of cancer types. In some embodiments, applying the trained classifier comprises: (i) applying a first trained model to the measured levels of the first target molecules to assign a first probability score to the cancer; (ii) applying a second trained model to the measured levels of the second target molecules to assign a second probability score to the cancer; and (iii) summing the first probability score and the second probability score.
[0072] Trained classifier This disclosure relates to trained classifiers. For example, machine learning or deep learning models (e.g., trained classifiers) can be used to determine a disease state based on the levels of a first target molecule and a second target molecule. In various embodiments, the output of the trained classifier is a probability score for cancer. For example, the trained classifier can determine a first probability score based on the level of a first target molecule and a second probability score based on the level of a second target molecule. In some embodiments, the first target molecule is a cfDNA molecule from multiple different target genomic regions that are differentially methylated in at least one of multiple cancer types. In some embodiments, the second target molecule comprises multiple different polypeptides differentially expressed in at least one of the multiple cancer types. In some embodiments, the trained classifier aggregates the first probability score and the second probability score to generate an overall probability score for cancer. For example, in some embodiments, the trained classifier computes the product of the first probability score and the second probability score for cancer to aggregate the first probability score and the second probability score, wherein the first probability score and the second probability score are determined by a first trained model and a second trained model, respectively. In some embodiments, aggregating the first probability score and the second probability score includes combining the first probability score and the second probability score for cancer in a linear model. Furthermore, the trained classifier can be combined with or otherwise used with a threshold setting to determine whether a sample is classified as cancer or non-cancer based on whether the overall probability score is above the threshold.
[0073] To determine each of the first and second probability scores, a trained classifier can apply a trained model to the measurement levels of the first and second target molecules. For example, in some embodiments, the trained classifier applies a first trained model to the measurement levels of the first target molecule to assign a first probability score to cancer, and applies a second trained model to the measurement levels of the second target molecule to assign a second probability score to cancer. In some embodiments, the first trained model binarizes the measurement levels of these first target molecules by assigning a first value when a target genomic region (e.g., a target genomic region with a specific methylation level, methylation pattern, or methylation state vector) is detected and a second value when no target genomic region is detected. In some embodiments, the second trained model performs a logarithmic transformation on the measurement levels of these second target molecules normalized to control proteins present in known amounts.
[0074] In some embodiments, a trained classifier assigns an overall probability score to each of a plurality of different cancer types. In some embodiments, the cancer type with the highest overall probability score is identified as the cancer detected in the sample. In some embodiments, the trained classifier distinguishes subjects with cancer from subjects without cancer with a specificity defined for each of the plurality of cancer types. In some embodiments, the plurality of cancer types includes at least 10 cancer types. In some embodiments, the plurality of cancer types includes one or more of anorectal cancer, bladder cancer, colorectal cancer, esophageal cancer, head and neck cancer, liver cancer, bile duct cancer, lung cancer, ovarian cancer, pancreatic cancer, stomach cancer, breast cancer, prostate cancer, kidney cancer, cervical cancer, endometrial cancer, and hematologic malignancies.
[0075] The trained classifier and trained model can be applied to any of a variety of modeling types. In some embodiments, the trained classifier is a binary classifier, a hybrid model classifier, a multilayer perceptron model classifier, or a logistic regression classifier. In some embodiments, the first trained model and / or the second trained model is a binary classifier, a hybrid model classifier, a multilayer perceptron model classifier, or a logistic regression classifier. In some embodiments, the first trained model and the second trained model are the same. In some embodiments, the first trained model and the second trained model are different. In some embodiments, the trained classifier has a higher cancer detection sensitivity than each of the first trained model and the second trained model. In some embodiments, the trained classifier has a cancer detection specificity equal to or greater than each of the first trained model and the second trained model. In some embodiments, the first and / or second trained model has a defined cancer detection specificity of 0.900 or higher (e.g., at least 0.950, 0.975, 0.980, 0.985, 0.990, 0.995 or higher). In some embodiments, the trained classifier has a defined cancer detection specificity of 0.900 or higher (e.g., at least 0.950, 0.975, 0.980, 0.985, 0.990, 0.995 or higher). In some embodiments, the trained classifier has a defined cancer detection specificity of at least 0.99 or higher. In some embodiments, the application of the trained classifier includes a sensitivity of at least 30% (e.g., at least 40%, 50%, 60%, 70%, 80% or higher).
[0076] To train a cancer type classifier, the analysis system can obtain data from a training sample set. In some embodiments, the data includes measurements of the levels of a first target molecule and a second target molecule, along with labels corresponding to the type of cancer or non-cancer state of the sample. The analysis system can utilize the training set to train the classifier to predict the cancer state of test samples. For example, a first model can be trained by applying a first machine learning algorithm to a first measurement level to generate a first probability score of the presence of cancer in the subject, and a second model can be trained by applying a second machine learning algorithm to a second measurement level to generate a second probability score of the presence of cancer in the subject. The classifier can then be trained based on an overall cancer probability score obtained by summing the scores from the first and second models.
[0077] Furthermore, in some embodiments, the analysis system can divide the training set into K subsets or folds for K-fold cross-validation. In some instances, the folds can be balanced against cancer / non-cancer status, cancer type, cancer stage, age (e.g., grouped by 10-year intervals), and / or smoking status. In some instances, the training set is divided into 5 folds, thereby training 5 independent classifiers, training on 4 / 5 of the training samples in each case, and using the remaining 1 / 5 for validation.
[0078] During training with a training set, the analysis system can fit a probabilistic model to the level of a first target molecule in the sample or the level of a second target molecule in the sample. As used herein, a “probabilistic model” is any mathematical model capable of assigning probabilities to measured levels of target molecules. During training, the analysis system receives levels of a first target analyte and a second target analyte from one or more samples from subjects with a known cancer state and can use this information to determine the probability indicative of the cancer state. A trained classifier can then be trained based on the trained model. For example, in some embodiments, for a reference sample, the trained classifier can be trained using a reference first probability score from a first trained model, a reference second probability score from a second trained model, and a reference overall probability score summing the reference first and reference second probability scores. In some embodiments, the reference samples are from reference subjects with known cancer and from reference subjects without cancer.
[0079] In some instances, the probabilistic model is a “hybrid model” fitted using a mixture of components from a base model. In some instances, the analysis system performs fitting separately for each of multiple cancer types. According to various aspects of this disclosure, other methods may be used to fit the probabilistic model or identify parameters that maximize the log-likelihood of a given level (e.g., sequence read or protein level) derived from a reference sample. For example, in some instances, Bayesian fitting (using, for example, Markov chain Monte Carlo) is used, where each parameter is not assigned a single value but is associated with a distribution. In some instances, gradient-based optimization is used, where the gradient of the likelihood (or log-likelihood) relative to the parameter values is used to iteratively approach the optimal solution through the parameter space. In some embodiments, expectation maximization is performed, where a set of latent parameters (such as the identity of the mixture from which each of the multiple cfDNA fragments is derived) is set to their expected values under the previous model parameters, and then the model parameters are assigned to maximize the likelihood under the assumed values of these latent variables. These two steps are then repeated until convergence.
[0080] In some instances, the analytics system trains a multinomial logistic regression classifier on the training data of the K folds and generates predictions on the retained data. For example, for each of the K folds, a logistic regression can be trained for each combination of hyperparameters. Such hyperparameters may include L2 penalty and / or topK (e.g., the number of highly ranked regions that should be retained for each tissue type pair (including non-cancer), as ranked according to the mutual information procedure described above). For each set of hyperparameters, performance is evaluated against cross-validation predictions on the full training set, and the set of hyperparameters with the best performance is selected for retraining on that full training set. In some instances, the analytics system uses log loss as a performance metric, which is calculated by taking the negative logarithm of the correctly labeled prediction for each sample and then summing it over the samples (i.e., a perfect prediction for a correctly labeled sample is 1.0, resulting in a log loss of 0).
[0081] To generate predictions for new samples, feature values are calculated using the same method described above, but only for features selected at the chosen top K values (region / positive class combination). The generated features are then used to create predictions using a trained logistic regression model.
[0082] In some instances, the analytics system trains a two-level classifier. For example, the analytics system trains a binary cancer classifier based on the feature vectors of training samples to distinguish between labels (cancer and non-cancer). In this case, the binary classifier outputs a probability score that indicates the likelihood of cancer being present or absent. In another instance, the analytics system trains a multi-class cancer classifier to distinguish between multiple cancer types. In this multi-class cancer classifier, the cancer classifier is trained to determine a cancer prediction that includes predicted values for each of the multiple cancer types being classified. The predicted values may correspond to the likelihood that a given sample has each of the cancer types. For example, the cancer classifier returns a cancer prediction that includes predicted values for breast cancer, lung cancer, and non-cancer. Alternatively, the cancer classifier may return a cancer prediction for a test sample that includes predicted scores for breast cancer, lung cancer, and / or no cancer.
[0083] The analysis system may train the cancer classifier using any of many methods. For example, a binary cancer classifier could be an L2-regularized logistic regression classifier trained using a log loss function. Alternatively, the classifier could be a multinomial logistic regression classifier. In practice, other techniques can be used to train any type of cancer classifier, including but not limited to L1-regularized logistic regression, generalized linear models (GLM), random forests, multilayer perceptrons, support vector machines, and neural networks. These techniques are numerous and include kernel methods, machine learning algorithms such as multilayer neural networks, etc. In particular, methods described in PCT / US2019 / 022122 and US 20190287652 A1 (which are incorporated herein by reference in their entirety) can be used in various embodiments, especially those related to training models based on methylated cfDNA molecules.
[0084] In one particular embodiment, training the classifier includes (a) receiving a first measurement level of a first target molecule for a first sample of a reference subject; (b) training a first model by applying a first machine learning algorithm to these first measurement levels to generate a first probability score of the presence of cancer in the subject; (c) receiving a second measurement level of a second target molecule for a second sample of the reference subject; (d) training a second model by applying a second machine learning algorithm to these second measurement levels to generate a second probability score of the presence of cancer in the subject; (e) generating a reference first cancer probability score for the first samples using the trained first model; (f) generating a reference second cancer probability score for the second samples using the trained second model; (g) generating a reference overall cancer probability score for a plurality of reference subjects by aggregating the reference first cancer probability score and the reference second cancer probability score for each corresponding reference subject; and (h) training the classifier by applying a third machine learning algorithm to these reference first cancer probability scores, these reference second cancer probability scores, and these reference overall cancer probability scores to generate an overall cancer probability score for the subject. In some embodiments, the reference subjects include a first subject with a known type of cancer and a second subject without cancer. In some embodiments, summing the first cancer probability score and the second cancer probability score includes calculating the product of the first probability score and the second probability score for the cancer.
[0085] In some embodiments, summing the first probability score and the second probability score includes combining the first probability score and the second probability score of the cancer in a linear model. In some embodiments, the overall probability of the cancer is determined using a linear model according to the following formula:
[0086] In the above formula: =The final probability of cancer output given DNA and protein p_DNA = Cancer prediction from DNA models p_prot = Cancer prediction from protein model β0 = Bias term, constant / intercept β1 = Learning coefficient of DNA probability β2 = Learning coefficient of protein probability β3 = Learning coefficient of the product of DNA probability and protein probability A first trained model and a second trained model can apply machine learning algorithms differently to the levels of each target molecule being analyzed. For example, the first trained model can binarize the measurement levels of a first target molecule, while the second trained model can perform a logarithmic transformation on the measurement levels of a second target molecule. In some embodiments, the first trained model binarizes the measurement levels of the first target molecule by assigning a first value if a target genomic region is detected and a second value if no target genomic region is detected. In some embodiments, the first target molecule comprises cell-free DNA (cfDNA) from multiple different target genomic regions that are differentially methylated in at least one of multiple cancer types. In some embodiments, the second trained model performs a logarithmic transformation on the measurement levels of these second target molecules normalized to control proteins present in known amounts. In some embodiments, the second target molecule comprises multiple different peptides differentially expressed in at least one of the multiple cancer types.
[0087] The machine learning algorithm used to train the classifier can be any suitable machine learning algorithm known in the art. In some embodiments, the first machine learning algorithm, the second machine learning algorithm, and / or the third machine learning algorithm is L1 regularized logistic regression, L2 regularized logistic regression, generalized linear model (GLM), random forest, multinomial logistic regression, multilayer perceptron, support vector machine, or neural network. In some embodiments, the third machine learning algorithm is logistic regression.
[0088] The trained model This disclosure provides a model for assigning probability scores based on the level of target molecules, methods for training the model, and methods for combining the model with a trained classifier. In some embodiments, the model is trained using the cfDNA levels (e.g., the presence, number, or proportion of cfDNA from a target genomic region having a specific methylation level or methylation state vector, or its sequencing reads) of a reference sample from a subject with a known cancer diagnosis (e.g., any type of cancer, a specific type of cancer, or healthy / non-cancer). In some embodiments, the model is trained using a measurement level of peptides from a reference sample from a subject with a known cancer diagnosis (e.g., any type of cancer, a specific type of cancer, or healthy / non-cancer). Further details regarding embodiments of training the model based on cfDNA methylation levels and selecting target genomic regions accordingly are provided below. Similar considerations can be applied to the analysis of peptide levels and the selection of target peptides accordingly.
[0089] Data structure generation To create a healthy control group data structure, the analysis system obtains information related to the methylation status of multiple CpG sites on sequence reads derived from multiple DNA molecules or fragments from multiple healthy subjects. The method presented herein for creating a healthy control group data structure can be similarly performed on subjects with cancer, subjects with tissue-of-origin (TOO) cancer, subjects with a known type of cancer, or subjects with other known disease states. For example, via... Figure 1B The process shown in the image generates a methylation state vector for each DNA molecule or fragment.
[0090] In some embodiments, the analysis system subdivides the methylation state vector of each DNA fragment into strings of CpG sites. In one embodiment, the analysis system subdivides the methylation state vector such that the resulting strings are all shorter than a given length. For example, a methylation state vector of length 11 may be subdivided into strings of length 3 or less, resulting in 9 strings of length 3, 10 strings of length 2, and 11 strings of length 1. In another instance, a methylation state vector of length 7 may be subdivided into strings of length 4 or less, resulting in 4 strings of length 4, 5 strings of length 3, 6 strings of length 2, and 7 strings of length 1. If the methylation state vector generated from the DNA fragment is shorter than or the same length as a specific string, the methylation state vector may be converted into a single string containing all CpG sites of that vector.
[0091] In some embodiments, the analysis system counts the number of strings in the control group that have the specified CpG site as the first CpG site in the string and possess the methylation state probability for each possible CpG site and vector. For a string of length three at a given CpG site, there are 2^3 or 8 possible string configurations. For each CpG site, the analysis system counts the number of times each possible methylation state vector appears in the control group. This may involve counting the following quantities: for each starting CpG site in the reference genome,<Mx, Mx+1, Mx+2> ,<Mx, Mx+1, Ux+2> ......<Ux, Ux+1, Ux+2> The analysis system creates a data structure that stores the statistical count of the probability of a string at each starting CpG site.
[0092] Setting an upper limit on string length has several benefits. First, the size of the data structure created by the analysis system increases dramatically with the maximum string length. For example, a maximum string length of 4 means that at most 2^4 numbers need to be counted at each CpG. Increasing the maximum string length to 5 would double the possible number of methylation states that need to be counted. Reducing the string size helps alleviate the computational and data storage burden on the data structure. In some embodiments, the string size is 3. In some embodiments, the string size is 4. A second reason for limiting the maximum string length is to avoid overfitting downstream models. If long CpG strings do not have a strong biological effect on the outcome (e.g., an anomalous prediction of the presence of cancer), calculating probabilities based on long CpG site strings can be problematic because it requires a large amount of data that may not be available, making it too sparse for the model to perform properly. For example, calculating the probability of anomalousness / cancer conditioned on the first 100 CpG sites would require counting strings of length 100 in the data structure, some of which ideally match exactly the first 100 methylation states. If only sparse counts of strings of length 100 are available, the data may be insufficient to determine whether a given string of length 100 in the test sample is anomalous.
[0093] Data Structure Validation Once a data structure has been created, the analytics system may attempt to validate that data structure and / or any downstream models that utilize it.
[0094] The first type of validation ensures the removal of potentially cancerous samples from the healthy control group to avoid affecting the purity of that control group. This validation type checks the consistency within the control group data structure. For example, the healthy control group might contain samples from individuals with undiagnosed cancer that contain multiple anomalously methylated fragments. The analysis system can perform various calculations to determine whether data from subjects with clearly undiagnosed cancer should be excluded.
[0095] The second type of validation uses counts from the data structure itself (e.g., from the healthy control group) to examine the probabilistic model used to calculate p-values. Once the analysis system generates p-values for the methylation state vectors in the validation group, it constructs a cumulative density function (CDF) using these p-values. Using the CDF, the analysis system can perform various calculations on the CDF to validate the control group's data structure. One test uses the fact that, ideally, the CDF should be equal to or lower than the identity function such that CDF(x) ≤ x. Conversely, above the identity function, some flaws are revealed within the probabilistic model used for the control group's data structure. For example, if a fragment of 1 / 100 has a p-value score of 1 / 1000, meaning CDF(1 / 1000) = 1 / 100 > 1 / 1000, then the second type of validation fails, indicating a problem with the probabilistic model. See, for example, U.S. Publication No. 2019 / 0287652, the entire contents of which are hereby incorporated by reference.
[0096] The third type of validation uses a healthy validation sample set, separate from the samples used to build the data structure. This tests whether the data structure is constructed correctly and whether the model functions properly. The third type of validation quantifies how well the healthy control group generalizes to the distribution of healthy samples. If the third type of validation fails, the healthy control group fails to generalize well to the healthy distribution.
[0097] The fourth validation type uses samples from the non-healthy validation group for testing. The analysis system calculates p-values and constructs CDFs for the non-healthy validation group. For the non-healthy validation group, the analysis system expects to see CDF(x) > x for at least some samples, or in other words, the opposite of what was expected in the second and third validation types for the healthy control and healthy validation groups. If the fourth validation type fails, it indicates that the model cannot properly identify the anomaly it was designed to identify.
[0098] In embodiments that include a step of validating the data structure, the analysis system performs a fourth type of validation test as described above, which utilizes a validation group consisting of subjects, samples, and / or fragments assumed to be similar to the control group. For example, if the analysis system selects healthy subjects without cancer for the control group, the analysis system will also use healthy subjects without cancer in the validation group.
[0099] In some embodiments, the analysis system employs a validation set and generates a set of methylation state vectors. The analysis system performs p-value calculations for each methylation state vector from the validation set. For each possible methylation state vector, the analysis system calculates a probability from the control set's data structure. Once the probability of a methylation state vector is calculated, the analysis system calculates a p-value score for that vector based on the calculated probability. The p-value score represents the likelihood of finding that particular methylation state vector in the control set, as well as other methylation state vectors that might have even lower probabilities. Therefore, a low p-value score typically corresponds to a relatively unexpected methylation state vector compared to other methylation state vectors in the control set, while a high p-value score typically corresponds to a relatively more expected methylation state vector compared to other methylation state vectors found in that control set. Once the analysis system generates p-value scores for the methylation state vectors in the validation set, it constructs a cumulative density function (CDF) using the p-value scores from that validation set. The analysis system verifies the consistency of the CDF as described above in a fourth validation type test.
[0100] Fragments with anomalous methylation According to embodiments, anomalously methylated fragments with abnormal methylation patterns are selected as target genomic regions in cancer patient samples, subjects with TOO cancer, subjects with known cancer types, or subjects with other known disease states. In some embodiments, the analysis system generates a methylation state vector from the cfDNA fragments of the sample. The analysis system may process each methylation state vector as follows.
[0101] For a given methylation state vector, the analysis system enumerates all possibilities of methylation state vectors with the same starting CpG site and the same length (set of CpG sites) as that given methylation state vector. Since each methylation state can be methylated or unmethylated, there are only two possible states at each CpG site. Therefore, the number of different possibilities for a methylation state vector depends on a power of 2, such that a methylation state vector of length n will be associated with 2^n methylation state vector possibilities.
[0102] The analysis system accesses the healthy control group data structure to calculate the probability of observing each methylation state vector for the identified initiation CpG site / methylation state vector length. In one embodiment, Markov chain probabilities are used to model the joint probability calculation when calculating the probability of observing a given possibility. In some embodiments, calculation methods other than Markov chain probabilities are used to determine the probability of observing each methylation state vector.
[0103] The analysis system uses the calculated probability of each possibility to compute a p-value score for the methylation state vector. In one embodiment, this includes identifying the calculated probabilities corresponding to the possibilities that match the methylation state vector in question. Specifically, this is the possibility of having the same set of CpG sites as the methylation state vector, or similarly having the same starting CpG sites and length as the methylation state vector. The analysis system sums the calculated probabilities of any possibilities whose probabilities are less than or equal to the identified probabilities to generate the p-value score.
[0104] The p-value represents the probability of observing a fragment's methylation state vector or other less likely methylation state vectors in a healthy control group. Therefore, a low p-value typically corresponds to a methylation state vector that is rare in healthy subjects and causes the fragment to be labeled as aberrantly methylated relative to healthy controls. A high p-value typically relates to a methylation state vector expected to be present in healthy subjects in a relative sense. For example, if the healthy control group is non-cancerous, a low p-value indicates that the fragment is aberrantly methylated relative to that non-cancerous group, and thus may indicate the presence of cancer in the test subject.
[0105] As described above, the analysis system calculates a p-value score for each of multiple methylation state vectors, each representing a fragment of cfDNA in the test sample. To identify which fragments are aberrantly methylated, the analysis system may filter the set of methylation state vectors based on their p-value scores. In one embodiment, filtering is performed by comparing the p-value scores to a threshold and retaining only those fragments below the threshold. This threshold p-value score can be on the order of 0.1, 0.01, 0.001, 0.0001, or a similar order of magnitude.
[0106] P-value score calculation To calculate the p-value score for a given test methylation state vector, the analysis system uses that test methylation state vector and enumerates the probabilities of using methylation state vectors. For example,<M23, M24, M25, U26> The length of the test methylation state vector is 4, where the 2^4 possibilities of the methylation state vector cover the indicated CpG sites 23-26. In a general instance, the number of possibilities for the methylation state vector is 2^n, where n is the length of the test methylation state vector, or alternatively, the length of the sliding window (described further below).
[0107] The analysis system calculates the probability of observing a given methylation state vector. Since methylation conditionally depends on the methylation state of neighboring CpG sites, one approach to calculating the probability of observing a given methylation state vector is to use a Markov chain model. Typically, methylation state vectors are such as...<S1, S2, …, Sn> , where S represents the methylation state, whether methylated (denoted as M), unmethylated (denoted as U), or uncertain (denoted as I), and has a joint probability that can be expanded using the probability chain rule as follows:
[0108] Markov chain models can be used to make the calculation of the conditional probability of each possibility more efficient. In one embodiment, the analysis system selects the order of the Markov chain. k This corresponds to the number of CpG sites in the vector (or window) to be considered in the conditional probability calculation, such that the conditional probability is modeled as P(S n | S1, …, S n-1 ) ~ P(S n | S n-k-2 , …,S n-1 ).
[0109] To compute the per-Markov modeling probability of the methylation state vector possibilities, the analysis system accesses the control group's data structure, specifically the counts of various CpG sites and state strings. To compute P(Mn | Sn-k-2, …, Sn-1), the analysis system uses the following ratio: stored data from the matching...<Sn-k-2, …, Sn-1, Mn> The number of strings in the data structure divided by the number of strings stored from the match<Sn-k-2, …, Sn-1, Mn> and<Sn-k-2, …, Sn-1, Un> The sum of the number of strings in the data structure. Therefore, P(Mn | Sn-k-2, …, Sn-1) is a calculated ratio of the following form:
[0110] The calculation can also be smoothed by applying a prior distribution. In one embodiment, the prior distribution is a uniform prior, as in Laplace smoothing. As an example, a constant is added to the numerator of the above equation, and another constant (e.g., twice the constant in the numerator) is added to the denominator of the above equation. In other embodiments, algorithmic techniques such as Kensier-Ney smoothing are used.
[0111] The above formula is applied to the test methylation state vector. Once the probabilities are calculated, the analysis system calculates a p-value score, which is the sum of the probabilities that the methylation state vectors that match the test methylation state vector.
[0112] In one embodiment, the computational load of calculating probabilities and / or p-value scores can be further reduced by caching at least some computations. For example, the analysis system can cache the probability calculations of methylation state vectors (or windows thereof) in transient or persistent memory. If other fragments have the same CpG sites, caching the probabilities of these possibilities allows for efficient calculation of p-value scores without recalculating the probabilities of potential possibilities. Equivalently, the analysis system can calculate a p-value score for each possibility of a methylation state vector associated with a set of CpG sites in the vector (or window thereof). The analysis system can cache these p-value scores to determine the p-value scores of other fragments that include the same CpG sites. Overall, the p-value scores of the possibilities of methylation state vectors with the same CpG sites can be used to determine the p-value scores of different possibilities under the same set of CpG sites.
[0113] Sliding window In one embodiment, the analysis system uses a sliding window to determine the probabilities of methylation state vectors and calculate p-values. The analysis system only needs to enumerate probabilities and calculate p-values for windows containing consecutive CpG sites, where the length of the window (CpG sites) is shorter than at least some segments (otherwise, the window would be meaningless), without needing to enumerate probabilities and calculate p-values for all methylation state vectors. The window length may be static, user-defined, dynamic, or otherwise selected.
[0114] When calculating the p-value for a methylation state vector larger than the window, the window is identified starting from the first CpG site in the vector, identifying a set of consecutive CpG sites within the window. The analysis system calculates a p-value score for the window including the first CpG site. Then, the analysis system "slides" the window to the second CpG site in the vector and calculates another p-value score for the second window. Therefore, for a window size *l* and a methylation vector length *m*, each methylation state vector will generate *m-l+1* p-value scores. After completing the p-value calculation for each part of the vector, the lowest p-value score across all sliding windows is used as the overall p-value score for that methylation state vector. In another embodiment, the analysis system aggregates the p-value scores of the methylation state vectors to generate an overall p-value score.
[0115] Using a sliding window helps reduce the number of possible methylation state vectors that need to be enumerated and the corresponding probability calculations that would otherwise be required. Typically, the number of possible methylation state vectors increases exponentially with the size of the methylation state vector by a factor of 2. For a practical example, a fragment may have up to 54 CpG sites. The analysis system can use, for example, a window of size 5 to perform 50 p-value calculations for each of the 50 windows representing the fragment's methylation state vectors, instead of calculating the probabilities of 2^54 (approximately 1.8 × 10^16) possibilities to generate a single p-value. Each of these 50 calculations enumerates 2^5 (32) possible methylation state vectors, totaling 50 × 2^5 (1.6 × 10^3) probability calculations. This significantly reduces the computations required without significantly affecting the accurate identification of anomalous fragments. This additional step can also be applied when validating the control group using the methylation state vectors of the validation group.
[0116] Identify fragments that indicate cancer In some embodiments, the analysis system identifies cancer-indicating DNA fragments from a filtered set of anomalously methylated fragments. In some embodiments, the fragments identified as cancer-indicating are used to select target genomic regions for analysis of subsequent samples. For example, data used in training may include information about a set of genomic regions, and a subset of these regions is selected for analysis in test samples based on the degree to which they are determined to be informative. The selection of target genomic regions for analysis can be performed by sample processing steps (e.g., selectively capturing cfDNA fragments from target genomic regions using decoy oligonucleotides) or computationally (e.g., by ignoring sequencing reads of cfDNA fragments originating from outside the target genomic region set).
[0117] hypomethylated and hypermethylated fragments In some embodiments, the analysis system can identify DNA fragments considered hypomethylated or hypermethylated from a filtered set of anomalously methylated fragments as cancer-indicating fragments. Hypomethylated and hypermethylated fragments can be defined as fragments having a certain length of CpG sites (e.g., more than 3, 4, 5, 6, 7, 8, 9, 10, etc.) where the percentage of methylated CpG sites is high (e.g., more than 80%, 85%, 90%, or 95%, or any other percentage above 50%), or the percentage of unmethylated CpG sites is high (e.g., more than 80%, 85%, 90%, or 95%, or any other percentage above 50%).
[0118] Probabilistic Model In some embodiments, the analysis system identifies cancer-indicating fragments using a probabilistic model of methylation patterns suitable for each cancer type and non-cancer type. The analysis system calculates the log-likelihood ratio of the sample using DNA fragments in the genomic region, considering various cancer types with a probabilistic model fitted for each cancer type and non-cancer type. The analysis system may determine cancer-indicating DNA fragments based on whether at least one of the log-likelihood ratios considered for various cancer types is above a threshold.
[0119] In some embodiments, the analysis system divides the genome into multiple regions through multiple stages. In the first stage, the analysis system divides the genome into CpG site blocks. Each block is defined when the interval between two adjacent CpG sites reaches and / or exceeds a certain threshold (e.g., greater than 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, or 1,000 bp). Starting from each block, in the second stage, the analysis system further subdivides each block into regions of a certain length (e.g., 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1,100 bp, 1,200 bp, 1,300 bp, 1,400 bp, or 1,500 bp). The analysis system can further make adjacent regions overlap by length percentage (e.g., 10%, 20%, 30%, 40%, 50% or 60%, or 10% or more, 20% or more, 30% or more, 40% or more, 50% or more or 60% or more).
[0120] The analysis system can analyze sequence reads of DNA fragments derived from each region. It can process samples from tissues and / or high-signal cfDNA. High-signal cfDNA samples can be identified using binary classification models, cancer staging, or other indicators.
[0121] For each cancer and non-cancer type, the analysis system fits a separate probability model to the fragment. In one instance, each probability model is a mixture model, which contains a combination of multiple mixture components, each of which is an independent site model, where it is assumed that methylation at each CpG site is independent of the methylation state at other CpG sites.
[0122] In some embodiments, the calculation is performed relative to each CpG site. Specifically, a first count is determined, namely the number of cancerous samples that include anomalously methylated DNA fragments overlapping with the CpG (cancer_count), and a second count is determined, namely the total number of samples in the set containing fragments overlapping with the CpG (total). Genomic regions can be selected based on these numbers, for example, based on a criterion that is positively correlated with the number of cancerous samples that include DNA fragments overlapping with the CpG (cancer_count) and a criterion that is negatively correlated with the total number of samples in the set containing fragments overlapping with the CpG (total).
[0123] In some embodiments, various types of cancer with different TOOs are selected from anorectal cancer, bladder cancer, colorectal cancer, esophageal cancer, head and neck cancer, liver cancer, bile duct cancer, lung cancer, ovarian cancer, pancreatic cancer, stomach cancer, breast cancer, prostate cancer, kidney cancer, cervical cancer, endometrial cancer, and blood cancer.
[0124] The analysis system can further calculate the log-likelihood ratio (“R”) of fragments for various cancer types using a probabilistic model fitted for each cancer type and non-cancer type, or for cancer TOO. This log-likelihood ratio indicates the likelihood that the fragment indicates cancer. Both probabilities can be derived from probabilistic models fitted for each of the cancer and non-cancer types, defined as calculating the likelihood of observing methylation patterns on fragments in each of a given cancer and non-cancer type. For example, probabilistic models suitable for each of the cancer and non-cancer types can be defined.
[0125] Selection of genomic regions that indicate cancer In some embodiments, the analysis system can identify genomic regions that indicate cancer. To identify these informational regions, the analysis system calculates the information gain for each genomic region, or more specifically, each CpG site, which describes the ability to distinguish between various outcomes.
[0126] Methods for identifying genomic regions capable of distinguishing between cancer and non-cancer types utilize trained classification models that can be applied to sets of anomalously methylated DNA molecules or fragments corresponding to or derived from cancerous or non-cancer groups. The trained classification models can be trained to identify any target condition that can be identified from methylation state vectors.
[0127] In one embodiment, the trained classification model is a binary model (trained based on the methylation status of cfDNA fragments or genomic sequences obtained from a cohort of subjects with cancer or cancer TOO and a cohort of healthy subjects without cancer), and is then used to classify the probability that a test subject has cancer, cancer TOO, or does not have cancer based on an anomalous methylation state vector. In other embodiments, different models can be trained using subject cohorts that are known to have a specific cancer (e.g., breast cancer, lung cancer, prostate cancer, etc.); known to have cancer with a specific TOO that is considered the origin of cancer; or known to have different stages of a specific cancer (e.g., breast cancer, lung cancer, prostate cancer, etc.). In these embodiments, different models can be trained using sequence reads obtained from samples enriched with tumor cells from a cohort of subjects known to have a specific cancer (e.g., breast cancer, lung cancer, prostate cancer, etc.). The ability of each genomic region in the classification model to distinguish between cancer and non-cancer types is used to rank the genomic regions from most informative to least informative in terms of classification performance. The analysis system can identify genomic regions from the ranking based on the information gain in the classification between non-cancer and cancer types.
[0128] Calculate the information gain of hypomethylated and hypermethylated fragments that indicate cancer. In some embodiments, the analysis system uses fragments indicating cancer to train the model. The process accesses two sample training groups: a non-cancer group and a cancer group, and obtains a non-cancer methylation state vector set and a cancer methylation state vector set (containing fragments with anomalous methylation).
[0129] For each methylation state vector, the analysis system can determine whether the methylation state vector indicates cancer. Here, a fragment indicating cancer can be defined as a hypermethylated or hypomethylated fragment determined by whether at least a certain number of CpG sites have a specific state (methylated or unmethylated, respectively) and / or has a threshold percentage of sites in the specific state (again, methylated or unmethylated, respectively). In one instance, a cfDNA fragment is identified as hypomethylated or hypermethylated, respectively, if it overlaps with at least 5 CpG sites and at least 80%, 90%, or 100% of its CpG sites are methylated, or at least 80%, 90%, or 100% are unmethylated.
[0130] In some embodiments, the analysis system considers portions of the methylation state vector and determines whether those portions are hypomethylated or hypermethylated, and can distinguish between them. This alternative addresses missing methylation state vectors that, while large in size, contain at least one dense region of hypomethylation or hypermethylation. In some embodiments, segments indicating cancer can be defined based on the likelihood output of a trained probability model.
[0131] In one embodiment, the analysis system generates a score for low methylation (P-low) and a score for high methylation (P-high) at each CpG site in the genome. To generate any score at a given CpG site, the model performs four counts at that CpG site: (1) a vector count of the (methylation status) of the cancer set labeled as low methylation overlapping with that CpG site; (2) a vector count of the cancer set labeled as high methylation overlapping with that CpG site; (3) a vector count of the non-cancer set labeled as low methylation overlapping with that CpG site; and (4) a vector count of the non-cancer set labeled as high methylation overlapping with that CpG site. Additionally, this process can normalize these counts for each group to account for differences in group size between the non-cancer and cancer groups. In some embodiments, where segments indicating cancer are more commonly used, the score can be more broadly defined as the count of segments indicating cancer at each genomic region and / or CpG site.
[0132] In some embodiments, to generate a score for hypomethylation at a given CpG site, the process takes the ratio of (1) to the sum of (1) and (3). Similarly, a score for hypermethylation can be calculated by taking the ratio of (2) to the sum of (2) and (4). Additionally, these ratios can be calculated using other smoothing techniques discussed above. The hypomethylation and hypermethylation scores are correlated with an estimate of cancer probability, assuming that fragments from the cancer set are either hypomethylated or hypermethylated.
[0133] In some embodiments, the analysis system generates an overall score for hypomethylation and an overall score for hypermethylation for each anomalous methylation state vector. The overall scores for hypermethylation and hypomethylation are determined based on the scores for hypermethylation and hypomethylation at CpG sites in the methylation state vector. In one embodiment, the overall scores for hypermethylation and hypomethylation are assigned, respectively, to the maximum hypermethylation score and the maximum hypomethylation score for each site in the state vector. In some embodiments, the overall score may be based on the mean, median, or other calculated values using the scores for hyper / hypomethylation at sites in each vector.
[0134] In some embodiments, the analysis system sorts all of the subjects' methylation state vectors based on the subjects' overall hypomethylation score and overall hypermethylation score, thereby generating two rankings for each subject. The process selects the overall hypomethylation score from the hypomethylation ranking and the overall hypermethylation score from the hypermethylation ranking. Using the selected scores, the model generates a single feature vector for each subject. In one embodiment, the scores selected from either ranking are chosen in a fixed order that is identical for each generated feature vector for each subject in each of the training groups. For example, in one embodiment, the model selects the first, second, fourth, and eighth overall hypermethylation scores from each ranking, and similarly selects each overall hypomethylation score, writing these scores into the subject's feature vector.
[0135] In some embodiments, the analysis system trains a binary model to distinguish feature vectors between cancer and non-cancer training groups. Typically, any of a variety of classification techniques can be used. In one embodiment, the model is a non-linear classifier. In a particular embodiment, the classifier is a non-linear classifier employing L2-regularized kernel logistic regression with a Gaussian radial basis function (RBF) kernel.
[0136] Specifically, in one embodiment, the number of non-cancer samples or one or more different cancer types (nothers) and the number of cancer samples or one or more cancer types (ncancers) with anomalous methylation fragments overlapping with CpG sites are counted. The probability that a sample is cancerous is then estimated by a score (“S”) that is positively correlated with ncancers and negatively correlated with nothers. The score can be calculated using the following equations: (ncancer + 1) / (ncancer + nothers + 2) or (ncancer) / (ncancer + nothers). The analysis system calculates the information gain for each cancer type and each genomic region or CpG site to determine whether that genomic region or CpG site indicates cancer. The information gain is calculated for training samples, which are given a cancer type, compared to all other training samples. For example, two random variables, “anomalous fragment” (“AF”) and “cancer type” (“CT”), are used. In one embodiment, AF is a binary variable indicating the presence of an anomalous fragment overlapping a given CpG site in a given sample (as determined by the anomalous score / eigenvector above). CT is a random variable indicating whether the cancer belongs to a particular type. The analysis system calculates the mutual information relative to CT for a given AF. That is, if the presence of anomalous fragments overlapping with a specific CpG site is known, the number of bits of information about that cancer type can be obtained.
[0137] For a given cancer type, the analysis system can use this information to rank these sites based on their cancer-specificity. This procedure is repeated for all cancer types considered. If a particular region is typically anomalously methylated in training samples for a given cancer type but not in training samples for other cancer types or in healthy training samples, then for that given cancer type, the CpG sites overlapping those anomalous fragments will tend to have high information gain. The ranked CpG sites for each cancer type are then greedily added (selected) to a set of selected CpG sites according to their ranking in the trained model.
[0138] Calculate the pairwise information gain from fragments identified as cancer indicators from a probabilistic model. Once a fragment indicative of cancer is identified, analysis can then be used to identify genomic regions. The analysis system defines a feature vector for each sample, each region, and each cancer type by counting DNA fragments with a calculated log-likelihood ratio higher than a plurality of thresholds indicating that the fragment indicates cancer, where each count is a value in the feature vector. In some embodiments, the analysis system counts the number of fragments present in regions of each cancer type in the sample that have a log-likelihood ratio higher than one or more possible thresholds. In some embodiments, the analysis system defines a feature vector for each sample by counting the number of DNA fragments in each genomic region of each cancer type, providing calculated log-likelihood ratios for fragments above a plurality of thresholds, where each count is a value in the feature vector. The analysis system can use the defined feature vectors to calculate an informatics score for each genomic region, which describes the ability of the genomic region to distinguish each pair of cancer types. For each pair of cancer types, the analysis system ranks the regions based on the informatics score. The analysis system can select regions based on the ranking according to the informatics score.
[0139] In some embodiments, the analysis system calculates an information score for each region, which describes the region's ability to distinguish each pair of cancer types. For each distinct pair of cancer types, the analysis system may designate one type as a positive type and the other as a negative type. In some embodiments, the ability of a region to distinguish between positive and negative types is based on mutual information, which is calculated using estimated scores of positive and negative cfDNA samples for which the feature is expected to be non-zero in the final assay (i.e., at least one fragment will be sequenced to that level in a targeted methylation assay). These scores are estimated using the ratio of the feature appearing in healthy cfDNA as well as in high-signal cfDNA and / or tumor samples for each cancer type. For example, if a feature appears frequently in healthy cfDNA, it is estimated that it will also appear frequently in cfDNA of any cancer type, and this may result in a low information score. The analysis system may select a number of regions from the sequence for each pair of cancer types.
[0140] In some embodiments, the analysis system further identifies predominantly hypermethylated or hypomethylated regions from the region sorting. The analysis system may load a set of one or more positive-type fragments identified as informative regions. The analysis system evaluates whether the loaded fragments are predominantly hypermethylated or hypomethylated based on the loaded fragments. If the loaded fragments are predominantly hypermethylated or hypomethylated, the analysis system can select probes corresponding to the predominant methylation pattern. If the loaded fragments are not predominantly hypermethylated or hypomethylated, the analysis system can use a hybrid probe to target both hypermethylation and hypomethylation. The analysis system may further identify a minimal set of CpG sites that overlap by more than a certain percentage of the fragments.
[0141] In some embodiments, after ranking regions based on information content scores, the analysis system labels each region with the lowest information content ranking across all cancer type pairs. For example, if a region is the 10th most informative region distinguishing between breast and lung and the 5th most informative region distinguishing between breast and colorectal, then that region will be given an overall label of "5". The analysis system can design probes starting with the lowest-labeled regions while adding regions to the set, for example, until the set's size budget has been exhausted.
[0142] The trained model and its features In some instances, the assay set is used in conjunction with a trained model that predicts the disease state of a sample, such as cancer or non-cancer prediction, tissue of origin prediction, and / or indeterminate state prediction. In some instances, cancer type models can generate features based on sequence reads by considering methylated or unmethylated DNA fragments from certain target genomic regions. For example, if a cancer type model determines that the methylation pattern of a fragment is similar to that of a certain cancer type, the model can set the feature of that fragment to 1; otherwise, if such a fragment is not present, the feature can be set to 0. In this way, a cancer type model can generate a binary feature set for each sample (30,000 features, for example). Further, in some instances, all or part of the binary feature set of a sample can be input into the cancer type model to provide a set of probability scores, such as a probability score for each cancer type category and non-cancer type category. Furthermore, in some instances, cancer type models may be used in conjunction with or otherwise combined with threshold settings to determine whether a sample is classified as cancerous or non-cancer, and / or in conjunction with or otherwise combined with uncertainty threshold settings to reflect the confidence level of a particular TOO determination. Such methods will be described further below.
[0143] To train a cancer type model, the analysis system can obtain a training sample set. In some instances, each training sample includes one or more fragment files (e.g., files containing sequence read data), a label corresponding to the sample's cancer type (TOO) or non-cancer state, and / or the sample individual's sex. The analysis system can use the training set to train a cancer type classifier to predict the disease state of the sample.
[0144] In some instances, for training, the analysis system divides the genome (e.g., the whole genome) or a subset of the genome (e.g., targeted methylation regions) into multiple regions. For example, portions of the genome can be divided into CpG “blocks,” with a new block arising whenever the interval between nearest-neighbor CpGs is at least a minimum interval distance (e.g., at least 500 bp). Further, in some instances, each block can be divided into 1000 bp regions and positioned such that neighboring regions have a certain amount of overlap (e.g., 50% or 500 bp).
[0145] Furthermore, in some instances, the analysis system can split the training set into K subsets or folds for K-fold cross-validation. In some instances, the folds can be balanced against cancer / non-cancer status, tissue of origin, cancer stage, and / or age (e.g., grouped by 10-year intervals). In some instances, the training set is split into 5 folds, thereby training 5 independent classifiers, training on 4 / 5 of the training samples in each case, and using the remaining 1 / 5 for validation.
[0146] During training using a training set, the analysis system can fit a probabilistic model to fragments derived from samples of each cancer type (and for healthy cfDNA). During training, the analysis system fits sequence reads derived from one or more samples from subjects with known diseases and can be used to determine the probability of sequence reads indicative of disease states using methylation information or methylation state vectors. Specifically, in some cases, the analysis system determines the methylation rate of each CpG site observed in the sequence read. The methylation rate represents the fraction or percentage of methylated base pairs within a CpG site. The trained probabilistic model can be parameterized by the product of methylation rates. Typically, any known probabilistic model used to assign probabilities to sequence reads from samples can be used. For example, the probabilistic model could be a binomial model, where each site on the nucleic acid fragment (e.g., a CpG site) is assigned a methylation probability; or it could be an independent site model, where the methylation of each CpG is specified by a different methylation probability, where the methylation of one site is considered independent of the methylation of one or more other sites on the nucleic acid fragment.
[0147] In some instances, the probabilistic model is a Markov model, where the methylation probability at each CpG site depends on a certain number of methylation states at pre-CpG sites in the sequence read or the nucleic acid molecule from which the sequence read is derived. See, for example, US 20190287652 A1, which is incorporated herein by reference in its entirety and may be used in various embodiments.
[0148] In some instances, the probabilistic model is a “mixture model” fitted using a mixture of components from a base model. For example, in some embodiments, multiple independent site models can be used to determine the mixture components, where the methylation (e.g., methylation rate) of each CpG site is considered independent of the methylation of other CpG sites. Using an independent site model, the probability assigned to a sequence read or the nucleic acid molecule from which it is derived is the product of the methylation probability of each CpG site where the sequence read is methylated and the methylation probability of each CpG site where the sequence read is not methylated. According to this example, the analysis system determines the methylation rate of each of the mixture components. The mixture model is parameterized by the sum of the mixture components, with each mixture component associated with a product of methylation rates. The probabilistic model Pr of n mixture components can be expressed as:
[0149] For the input segment, This indicates the location of the fragment in the reference genome. i The observed methylation states, where 0 represents unmethylated and 1 represents methylated. For each mixture component k The fraction assignment is f k ,in, f k ≥0 and Hybrid components k Location of CpG sites i The methylation probability at is β ki Therefore, the probability of unmethylation is 1- β ki Number of hybrid components n It can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, etc.
[0150] In some embodiments, the analysis system uses maximum likelihood estimation to fit a probability model to identify a set of parameters that maximizes the log-likelihood of all fragments derived from the disease state, and imposes a regularization penalty on each methylation probability with a regularization strength r. For a total of N fragments, the maximized quantity can be expressed as:
[0151] In some embodiments, the analysis system performs fitting separately for each cancer type and healthy cfDNA. According to various aspects of this disclosure, other methods may be used to fit a probabilistic model or identify parameters that maximize the log-likelihood of all sequence reads derived from a reference sample. For example, in some instances, Bayesian fitting (using, for example, Markov chain Monte Carlo) is used, where each parameter is not assigned a single value but is associated with a distribution. In some instances, gradient-based optimization is used, where the gradient of the likelihood (or log-likelihood) relative to the parameter values is used to progressively approach the optimal solution through the parameter space. In still other instances, expectation maximization is used, where the latent parameter set (such as the identity of the mixed component from which each fragment is derived) is set to expected values under the previous model parameters, and then the model parameters are assigned to maximize the likelihood conditional on the assumptions of these latent variables. These two steps are then repeated until convergence.
[0152] Furthermore, in some instances, the analysis system can generate features for each sample in the training set. For example, for each sample (regardless of labeling), for each region, for each of multiple cancer types, and for each fragment, the analysis system can evaluate the log-likelihood ratio R based on the fitted probability model as follows:
[0153] Next, for each sample, for each region, for each cancer type, and for each in the "hierarchy" value set, the analysis system can count the number of fragments where R_cancer_type > hierarchy and assign these counts as non-negative integer value features. For example, the hierarchy includes thresholds of 1, 2, 3, 4, 5, 6, 7, 8, and 9, resulting in each region containing 9 features for each cancer type.
[0154] In some instances, the analysis system may select certain features to include in the feature vector of each sample. For example, for each pair of distinct cancer types, the analysis system may designate one type as "positive" and the other as "negative," and rank the features according to their ability to distinguish these types. In some cases, the ranking is based on mutual information calculated by the analysis system. For example, mutual information can be calculated using estimated scores of positive and negative type (e.g., cancer types A and B) samples for which the feature is expected to be non-zero in the outcome assay. For example, if a feature is frequent in healthy cfDNA, the analysis system may determine that the feature is unlikely to be frequent in cfDNA associated with various types of cancer. Therefore, the feature may be a weak measure in distinguishing disease states. When calculating mutual information I, variable X is a feature (e.g., binary), and variable Y represents the disease state, such as type A or type B cancer:
[0155] X and Y The joint probability mass function is And the marginal probability mass function is and The analysis system can assume that feature loss is not informative and that any disease state is equally likely a priori, for example... The probability of observing (e.g., in cfDNA) a given binary trait of type A cancer is expressed as: ,in, f A It is the probability of observing this feature in tumor ctDNA samples (or high-signal cfDNA samples) associated with type A cancer, and f H It is the probability of observing this feature in healthy or non-cancer cfDNA samples.
[0156] In some embodiments, only features corresponding to positive types are included in the ranking, and only if the predicted incidence of these features in the positive type is higher than that in the negative type. For example, if "liver" is a positive type and "breast" is a negative type, only the "liver_x" feature is considered, and only if its estimated incidence in liver cfDNA is higher than its estimated incidence in breast cfDNA. Further, in some instances, for each region, for each cancer type pair (including non-cancer as a negative type), the analysis system retains only the best-performing hierarchy. Further, in some instances, the analysis system binarizes feature values, thereby setting any feature value greater than 0 to 1, such that all features are either 0 or 1.
[0157] In some instances, the analytics system trains a multinomial logistic regression model on the training data of the K folds and generates predictions on the retained data. For example, for each of the K folds, a logistic regression can be trained for each combination of hyperparameters. Such hyperparameters may include L2 penalty and / or topK (e.g., the number of highly ranked regions that should be retained for each tissue type pair (including non-cancer), as ranked according to the mutual information procedure described above). For each set of hyperparameters, performance is evaluated against cross-validation predictions on the full training set, and the set of hyperparameters with the best performance is selected for retraining on that full training set. In some instances, the analytics system uses log loss as a performance metric, which is calculated by taking the negative logarithm of the correctly labeled prediction for each sample and then summing it over the samples (i.e., a perfect prediction for a correctly labeled sample is 1.0, resulting in a log loss of 0).
[0158] To generate predictions for new samples, feature values are calculated using the same method described above, but only for features selected at the chosen top K values (region / positive class combination). The generated features are then used to create predictions using the logistic regression model trained above.
[0159] In some instances, the analytics system trains a two-level model. For example, the analytics system trains a binary cancer model based on the feature vectors of training samples to distinguish between labels (cancer and non-cancer). In this case, the binary model outputs a prediction score indicating the likelihood of cancer presence or absence. In another instance, the analytics system trains a multi-class cancer model to distinguish between multiple cancer types. In this multi-class cancer model, the cancer model is trained to determine a cancer prediction that includes a predicted value for each of the cancer types being classified. The predicted value may correspond to the likelihood that a given sample has each of the cancer types. For example, the cancer model returns a cancer prediction or probability score that includes predicted values for breast cancer, lung cancer, and non-cancer. For example, the cancer model may return a cancer prediction or probability score for a test sample that includes predicted scores for breast cancer, lung cancer, and / or no cancer.
[0160] The analysis system may train the model according to any of many methods. For example, a binary cancer model could be an L2-regularized logistic regression classifier trained using a log loss function. Similarly, a multi-cancer (TOO) model could be a multinomial logistic regression classifier. In practice, other techniques can be used to train both types of cancer models. These techniques are numerous, including kernel methods, machine learning algorithms such as multilayer neural networks, etc. In particular, the methods described in US20190287652A1 (which are incorporated herein by reference in their entirety) can be used in various embodiments. Furthermore, in some embodiments, the TOO model is trained only on cancer samples that the binary model successfully identifies as cancer, thereby ensuring sufficient cancer signal in those cancer samples. In some instances, the binary model is trained on training samples regardless of the TOO.
[0161] Methylation nucleic acid detection In some embodiments, the methods described herein include detecting methylation patterns in nucleic acid molecules (e.g., cfDNA molecules). Methylation patterns in nucleic acids can be detected and analyzed using any suitable method known in the art.
[0162] Figure 1AThis is a flowchart of an exemplary process 100 for processing nucleic acid samples and generating methylation state vectors of DNA fragments according to some embodiments. The method may include, but is not limited to, one or more of the following steps. For example, any step of the method may include a quantitative sub-step for quality control or other laboratory assay procedures known to those skilled in the art.
[0163] In step 105, a sample containing nucleic acids (e.g., cfDNA) is collected from the subject. The sample can be any subset of the human genome, including the whole genome. The sample can include blood, plasma, serum, urine, feces, saliva, other types of bodily fluids, or any combination thereof. In some embodiments, the method of obtaining a blood sample (e.g., using a syringe or pricking a finger) can be less invasive than a procedure to obtain a tissue biopsy (which may require surgery). The extracted sample may contain cfDNA and / or ctDNA. In healthy individuals, the body can naturally clear cfDNA and other cellular debris. If the subject has cancer or a disease, the cfDNA and / or ctDNA in the sample may be present at detectable levels for the detection of that cancer or disease.
[0164] In step 110, nucleic acids are processed to distinguish between methylated and unmethylated nucleotides, thereby generating transformed cfDNA molecules. In some embodiments, the processing includes deamination, such as treatment of the cfDNA molecule with cytidine deaminase or with bisulfite. In one embodiment, the method uses bisulfite treatment of DNA (e.g., cfDNA) that converts unmethylated cytosine to uracil without converting methylated cytosine. For example, commercially available kits are used for bisulfite conversion, such as the EZ DNA Methylation™-Gold, EZ DNA Methylation™-Direct, or EZ DNA Methylation™-Lightning kits (available from Zymo Research Corp (Irvine, CA)). In another embodiment, an enzymatic reaction is used to complete the conversion of unmethylated cytosine to uracil. For example, this conversion can be accomplished using commercially available kits, such as APOBEC-Seq (NEBiolabs, Ipswich, MA).
[0165] In step 115, a sequencing library is prepared. In some embodiments, an ssDNA ligation reaction is used to add an ssDNA adaptor to the 3'-OH end of a bisulfite-converted ssDNA molecule. In one embodiment, the ssDNA ligation reaction uses CircLigase II (Epicentre) to ligate the ssDNA adaptor to the 3'-OH end of the bisulfite-converted ssDNA molecule, wherein the 5' end of the adaptor is phosphorylated and the bisulfite-converted ssDNA has been dephosphorylated (i.e., the 5' phosphate is removed). In another embodiment, the ssDNA ligation reaction uses a thermostable 5' AppDNA / RNA ligase (available from New England BioLabs, Ipswich, Massachusetts, USA) to ligate the ssDNA adapter to the 3'-OH end of the bisulfite-converted ssDNA molecule. In this example, the first adaptor is adenylated at the 5' end and blocked at the 3' end. In another embodiment, the ssDNA ligation reaction uses T4 RNA ligase (available from New England Biolabs, Inc.) to ligate the ssDNA adapter to the 3'-OH end of the bisulfite-converted ssDNA molecule. In the second step, a second-strand DNA is synthesized in an extension reaction. For example, extension primers that hybridize to primer sequences included in the ssDNA adapter are used in a primer extension reaction to form a double-stranded bisulfite-converted DNA molecule. Optionally, in one embodiment, the extension reaction uses an enzyme capable of reading through uracil residues in the bisulfite-converted template strand. Optionally, in a third step, the dsDNA adapter is added to the double-stranded bisulfite-converted DNA molecule. Finally, the double-stranded bisulfite-converted DNA is amplified to add a sequencing adaptor. For example, PCR amplification using a forward primer including the P5 sequence and a reverse primer including the P7 sequence is used to add the P5 and P7 sequences to the bisulfite-converted DNA. Optionally, during library preparation, a unique molecular identifier (UMI) can be added to a nucleic acid molecule (e.g., a DNA molecule), such as through adaptor ligation or primer extension. Typically, a UMI is a short nucleic acid sequence (e.g., 4-10 base pairs) that acts as a tag to facilitate the identification of sequence reads derived from a specific DNA fragment. In some embodiments, the UMI contains degenerate base pair positions. During PCR amplification following adaptor ligation, the UMI is replicated along with the attached DNA fragment, providing a method for identifying sequence reads from the same original fragment in downstream analysis (either by using the UMI alone or in combination with a portion of the terminal sequence of a sample nucleic acid fragment, such as the first 2-10 nucleotides).
[0166] In step 120, the target genomic region can be enriched from the library. This method is used, for example, in the case of a targeted panel assay of the sample. During enrichment, hybridization probes (also referred to herein as “probes” or “bait oligonucleotides”) are used to target and optionally pull down nucleic acid fragments that provide information about the presence or absence of cancer (or disease), cancer status, or cancer classification (e.g., cancer type or tissue of origin). In some embodiments, the probes have features specific to this document, such as features associated with various other aspects described herein. For a given workflow, the probes can be designed to anneal (or hybridize) with a target (complementary) strand of DNA (e.g., a transformed DNA molecule). The target strand can be a “positive” strand (e.g., the strand transcribed into mRNA and subsequently translated into protein) or a complementary “negative” strand. The length of the probes can range from tens, hundreds, or thousands of base pairs. Furthermore, these probes can cover overlapping portions of the target genomic region.
[0167] In some embodiments, the decoy oligonucleotide is designed to enrich target genomic regions comprising at least 1,000, 5,000, 10,000, 20,000, or 30,000 or more target genomic regions. In some embodiments, the target genomic regions comprise 1,000 to 30,000, 5,000 to 25,000, 10,000 to 20,000, or 12,500 to 15,000 target genomic regions. In some embodiments, the target genomic regions comprise at least 5,000 target genomic regions. In some embodiments, the target genomic regions comprise at least 20,000 target genomic regions. In some embodiments, the decoy oligonucleotide is designed to target genomic regions having a collective total size of at least 50 kb, 100 kb, 500 kb, or 1,000 kb. In some embodiments, the total size of the target genomic regions (e.g., at least 10,000 target genomic regions) is 500 kb to 1000 kb, 100 kb to 500 kb, 50 kb to 100 kb, or 10 kb to 50 kb. In some embodiments, the total size of the target genomic regions is at least 100 kb. In some embodiments, the total size of the target genomic regions is at least 500 kb. In some embodiments, the total size of the target genomic regions is less than the combined length of all different bait oligonucleotides in the set used to enrich the target genomic regions, such as when the bait oligonucleotides contain overlapping sequences. In some embodiments, when overlapping nucleotide positions are counted only once, the total size of the multiple target genomic regions is given by the total length of the different oligonucleotide probes in the set used to enrich the target genomic regions.
[0168] In some embodiments, the decoy oligonucleotide is configured to hybridize with transformed DNA molecules (e.g., transformed cfDNA molecules) corresponding to or derived from one or more genomic regions. Therefore, the decoy oligonucleotide may have a sequence different from the target genomic region. For example, DNA containing unmethylated CpG sites can be converted to include UpG instead of CpG by deamination (e.g., by treatment with cytidine deaminase or bisulfite). Therefore, probes targeting such targets may be configured to hybridize with sequences including UpG instead of naturally occurring unmethylated CpG. Thus, sites complementary to unmethylated sites in the probe may contain CpA instead of CpG, and some probes targeting hypomethylated sites where all methylation sites are unmethylated may not contain guanine (G) bases. In some embodiments, at least 3%, 5%, 10%, 15%, or 20% of the probes do not contain CpG sequences. In some embodiments, at least 5% of the probes do not contain CpG sequences. In some embodiments, at least 10% of the probes do not contain CpG sequences.
[0169] In some embodiments, the probe length ranges from tens, hundreds, more than 200, or more than 300 base pairs. The probe may contain at least 50, 75, 100, or 120 nucleotides. The probe may contain less than 300, 250, 200, or 150 nucleotides. In some embodiments, the probe contains 100-150 nucleotides. In one particular embodiment, the probe contains 120 nucleotides.
[0170] In some embodiments, probes are designed in a “2× tiling” manner to cover overlapping portions of the target region. The coverage of each probe optionally overlaps at least partially with another probe in the library. In such embodiments, the collection contains multiple probe pairs, where each probe in a pair overlaps with the other by at least 25, 30, 35, 40, 45, 50, 60, 70, 75, or 100 nucleotides. In some embodiments, the overlapping sequence may be designed to be complementary to the target genomic region (or cfDNA derived therefrom) or to a sequence homologous to the target region or cfDNA. Thus, in some embodiments, at least two probes contain sequences complementary to the same sequences within the target genomic region, and nucleotide fragments corresponding to or derived from that target genomic region can be bound and captured by at least one of the probes. For a given pair of probes containing overlapping sequences, the pair may contain non-overlapping sequences complementary to the target genomic region extending from different ends of the overlapping sequences. Other tiling levels are also possible, such as 3× tiling, 4× tiling, etc., where each nucleotide in the target region can bind to more than two probes.
[0171] In some embodiments, a single base in the target genomic region is exactly overlapped by two probes. Probes extending bidirectionally beyond the target genomic region can be used to pull down cfDNA fragments containing a portion of the target genomic region and DNA sequences adjacent to it. In some cases, even relatively small target regions can be targeted with three probes. Probe sets containing three or more probes are optionally used to capture larger genomic regions. In some embodiments, a subset of probes will extend collectively across the entire target genomic region (e.g., it may be complementary to unconverted or converted fragments within that entire genomic region). Tiled probe sets optionally include probes that collectively include at least two probes overlapping each nucleotide in the target genomic region. This is done to ensure that cfDNA containing a small portion of the target genomic region at one end will have substantial overlap with at least one probe (this overlap extends into adjacent non-target genomic regions) to provide effective capture.
[0172] In some embodiments, each target genomic region is targeted by a probe set. The probe set may be designed in a tiling manner such that adjacent probes have overlapping sequences that hybridize to the same portion of the genomic region. Since DNA has two strands, the probe set may also include overlapping probes that hybridize to the other strand, for a total of four probes hybridizing to the same portion of the genomic region. In some embodiments, the probe set configured to hybridize to the target genomic region does not span the entire region; that is, at least some sequences within the target genomic region do not have corresponding probes. For example, sequences within the target genomic region may be similar to or identical to many other sequences in the genome, and no probe is designed to target that sequence because such a probe would hybridize to more than a threshold number of off-target regions.
[0173] For example, a 100 bp cfDNA fragment containing a 30 nt target genomic region will have at least 65 bp overlap with at least one of the overlapping probes. Other tiling levels are also possible. For example, to increase the target size and add more probes to the ensemble, probes can be designed to extend the 30 bp target region by at least 70 bp, 65 bp, 60 bp, 55 bp, or 50 bp. To completely capture any fragment overlapping with the target region (even if it overlaps by only 1 bp), probes can be designed to extend beyond the ends of either side of the target region, such as extending by at least 50 bp, 55 bp, 60 bp, 65 bp, 70 bp, 75 bp, 80 bp, or 85 bp. Probes can be designed to extend 75 bp beyond the ends of either side of the target region. In some embodiments, the presence of probes designed to extend beyond the ends of the target genomic region does not increase the size of that target genomic region (e.g., it is not included in determining the size of the individual target genomic region or the common size of multiple target genomic regions).
[0174] In some embodiments, the probe targets differentially methylated genomic regions between general cancerous (pan-cancer) samples and non-cancer samples, or only in cancerous samples with a specific cancer type (e.g., a lung cancer-specific target). For example, in some embodiments, the cancer assay set is designed to include differentially methylated genomic regions based on transformation (e.g., bisulfite) sequencing data generated from cfDNA and / or whole-genome DNA from cancerous and non-cancer individuals.
[0175] In some embodiments, each of the target genomic regions is differentially methylated in at least one of a plurality of cancer types. In some embodiments, the plurality of cancer types includes at least 10 cancer types. In some embodiments, the plurality of cancer types includes one or more of anorectal cancer, bladder cancer, colorectal cancer, esophageal cancer, head and neck cancer, liver cancer, bile duct cancer, lung cancer, ovarian cancer, pancreatic cancer, gastric cancer, breast cancer, prostate cancer, kidney cancer, cervical cancer, endometrial cancer, and hematologic malignancies.
[0176] In some embodiments, target genomic regions may be selected to have at least 3, 5, or 7 methylation sites. In some embodiments, each target genomic region contains at least five methylation sites. In some embodiments, the number of methylation sites (e.g., at least 5 methylation sites) is the number of methylation sites that are differentially methylated in at least one of multiple cancers.
[0177] Each of the probes (or probe pairs) can be designed to target one or more target genomic regions. Target genomic regions can be selected based on several criteria designed to improve selective enrichment of informative cfDNA fragments while reducing noise and nonspecific binding. Various filtering or modeling procedures are described herein for determining whether to include a target genomic region. In some embodiments, the target genomic regions for enrichment by decoy oligonucleotides are identified by a trained model (e.g., the trained model described herein) as genomic regions that are differentially methylated in at least one of several cancer types relative to non-cancer tissue or relative to different types of cancer. In some embodiments, two or more of the filtering or modeling procedures described herein are used in combination.
[0178] Decoy oligonucleotides can be engineered to enrich target genomic regions located at various locations within the genome, including but not limited to promoters, enhancers, exons, introns, intergenic regions, and other parts. In some embodiments, decoy oligonucleotides targeting non-human genomic regions, such as probes targeting viral genomic regions, can be added.
[0179] In some embodiments, each decoy oligonucleotide is conjugated to a solid surface (e.g., a chip or bead, such as a magnetic or paramagnetic bead) or to a non-nucleotide affinity moiety (e.g., a member of a binding pair). In some embodiments, this conjugation is used to facilitate the separation of DNA molecules bound to the decoy oligonucleotide from unbound DNA molecules. Generally, a “binding pair” refers to a first part and a second part, wherein the first part and the second part have specific binding affinity for each other. Non-limiting examples of binding pairs include antigen / antibody; biotin / avidin (or biotin / streptavidin); calmodulin-binding protein (CBP) / calmodulin; hormone / hormone receptor; lectin / carbohydrate; peptide / cell membrane receptor; enzyme / cofactor; and enzyme / substrate. In some embodiments, the affinity moiety is biotin.
[0180] Non-restricted examples of target genomic regions and decoy oligonucleotides used to enrich target genomic regions are described in US20210025011A1, US20210238693A1, US20220119890A1, US20220064737A1 and US20220098672A1, which are incorporated herein by reference.
[0181] Following hybridization step 120, the hybridized target nucleic acid is enriched (e.g., separated from unbound nucleic acids by capture or otherwise) and can also be amplified using PCR (enrichment 125). For example, the target nucleic acid can be enriched to obtain enriched sequences that can be sequenced subsequently. Typically, a variety of methods can be used to isolate and enrich the target nucleic acid from probe hybridization. For example, a biotin moiety can be added to the 5' end of the probe (i.e., biotinylation) to facilitate the isolation of the target nucleic acid from probe hybridization with a surface coated with streptavidin (e.g., streptavidin-coated beads).
[0182] In step 130, sequence reads are generated from the enriched nucleic acid fragments. Sequencing data can be obtained from the enriched DNA sequences using methods known in the art. For example, methods may include next-generation sequencing (NGS) technologies, including synthesis techniques (Illumina), pyrosequencing (454 Life Sciences), ion semiconductor technology (Ion Torrent sequencing), single-molecule real-time sequencing (Pacific Biosciences), ligation sequencing (SOLiD sequencing), nanopore sequencing (Oxford Nanopore Technologies), or paired-end sequencing. In some embodiments, massively parallel sequencing is performed using sequencing-by-synthesis with reversible dye terminators.
[0183] In some embodiments, various methods can be used to align a sequence read with a reference genome to determine alignment location information. This alignment location information may indicate the start and end positions of a region in the reference genome, corresponding to the start and end nucleotide bases of a given sequence read. The alignment location information may also include the sequence read length, which can be determined by the start and end positions. Regions in the reference genome may be associated with genes or segments of genes.
[0184] In various embodiments, sequence reads consist of read pairs denoted as R1 and R2. For example, the first read R1 may be sequenced from the first end of the nucleic acid fragment, while the second read R2 may be sequenced from the second end of the nucleic acid fragment. Therefore, the nucleotide base pairs of the first read R1 and the second read R2 can be aligned in a manner consistent with the nucleotide bases of the reference genome (e.g., in opposite directions). The alignment position information derived from read pairs R1 and R2 may include a start position in the reference genome corresponding to one end of the first read (e.g., R1) and an end position in the reference genome corresponding to one end of the second read (e.g., R2). In other words, the start and end positions in the reference genome represent the possible positions of the nucleic acid fragment within the reference genome. Output files in SAM (Sequence Alignment Map) or BAM (Binary Alignment Map) format can be generated and output for further analysis.
[0185] Based on sequence reads, the location and methylation status of each CpG site can be determined based on alignment with a reference genome. Further, a methylation status vector can be generated for each fragment, specifying the fragment's location in the reference genome (e.g., as specified by the location of the first CpG site in each fragment or another similar indicator), the number of CpG sites in the fragment, and the methylation status of each CpG site in the fragment (methylated (e.g., denoted as M), unmethylated (e.g., denoted as U), or indeterminate (e.g., denoted as I)). The methylation status vector can be stored in transient or persistent computer memory for later use and processing. Further, duplicate reads or duplicate methylation status vectors from individual subjects can be removed. In another embodiment, a fragment can be identified as having one or more CpG sites with indeterminate methylation states. Such fragments can be excluded from subsequent processing or selectively included if downstream data models interpret such indeterminate methylation states.
[0186] In step 140, a methylation state vector is generated from the sequence read. For this, the sequence read is aligned to a reference genome. The reference genome helps provide context, indicating where the DNA fragment (e.g., cfDNA) originates in the human genome. In a simplified example, the sequence read is aligned such that three CpG sites correspond to CpG sites 23, 24, and 25 (these are arbitrary reference identifiers used for ease of description). After alignment, information is available about the methylation state of all CpG sites on the cfDNA fragment, and information about which locations in the human genome these CpG sites map to. Using the methylation state and location information, a methylation state vector of the DNA fragment can be generated.
[0187] Figure 1B According to the embodiments Figure 1A This describes an exemplary procedure 100 for sequencing a cfDNA fragment to obtain a methylation state vector. For example, the analysis system uses cfDNA fragment 112. In this example, cfDNA fragment 112 contains three CpG sites. As shown, the first and third CpG sites of cfDNA fragment 112 are methylated 114. During processing step 120, cfDNA fragment 112 is transformed to generate a transformed cfDNA fragment 122. During processing 120, the cytosine at the unmethylated second CpG site is converted to uracil. However, the first and third CpG sites are not converted.
[0188] After transformation, a sequencing library 130 is prepared and sequenced 140, generating sequence reads 142. The analysis system aligns sequence read 142 with a reference genome 144 150. The reference genome 144 provides context indicating the location of the cfDNA fragment within the human genome. In this simplified example, the analysis system aligns the sequence read 150 so that the three CpG sites correspond to CpG sites 23, 24, and 25 (these are arbitrary reference identifiers used for ease of description). The analysis system then generates information about the methylation status of all CpG sites on the cfDNA fragment 112 and the locations of these CpG sites within the human genome. As shown in the figure, methylated CpG sites on sequence read 142 are read as cytosine. In this example, cytosine appears only at the first and third CpG sites of sequence read 142, suggesting that the first and third CpG sites in the original cfDNA fragment are methylated. The second CpG site is read as thymine (U is converted to T during sequencing), indicating that the second CpG site in the original cfDNA fragment is unmethylated. With this information of methylation state and location, the analysis system generates a methylation state vector 152 for cfDNA fragment 112. In this example, the resulting methylation state vector 152 is...<M23, U24, M25> Where M corresponds to a methylated CpG site, U corresponds to an unmethylated CpG site, and the subscript number corresponds to the position of each CpG site in the reference genome.
[0189] Peptide detection In some embodiments, the methods described herein include detecting one or more peptides (e.g., measuring the level of one or more peptides).
[0190] In some embodiments, the polypeptide comprises at least 5, 10, 25, 50, 100, 200, 500, 1000, 2000, 3000, 5000, or 7500 different polypeptides. In some embodiments, the polypeptide comprises at least 5 to 7500, 10 to 5000, 25 to 3000, 50 to 1000, or 100 to 500 different polypeptides. In some embodiments, the polypeptide comprises at least 20 different polypeptides. In some embodiments, the polypeptide comprises at least 100 different polypeptides. In some embodiments, the polypeptide comprises at least 500 different polypeptides. In some embodiments, the polypeptide comprises at least 3000 different polypeptides.
[0191] In some embodiments, each of the polypeptides is differentially expressed in at least one of a plurality of cancer types. In some embodiments, the plurality of cancer types includes at least 10 cancer types. In some embodiments, the plurality of cancer types includes one or more of anorectal cancer, bladder cancer, colorectal cancer, esophageal cancer, head and neck cancer, liver cancer, bile duct cancer, lung cancer, ovarian cancer, pancreatic cancer, gastric cancer, breast cancer, prostate cancer, kidney cancer, cervical cancer, endometrial cancer, and hematologic malignancies.
[0192] Various suitable methods are available for detecting one or more target peptides. Non-limiting examples include competitive and non-competitive immunoassays, enzyme immunoassays (EIA), radioimmunoassays (RIA), antigen capture assays, double-antibody sandwich assays, Western blot analysis, enzyme-linked immunosorbent assays (ELISA), colorimetric assays, chemiluminescence assays, fluorescence assays, immunohistochemistry, chromatography, liquid chromatography, size exclusion chromatography, high-performance liquid chromatography (HPLC), gas chromatography, mass spectrometry, tandem mass spectrometry, and matrix-assisted laser desorption / ionization-time-of-flight (MALDI-TOF) assays. F) Mass spectrometry, electrospray ionization (ESI) mass spectrometry, surface-enhanced laser desorption / ionization-time-of-flight (SELDI-TOF) mass spectrometry, quadrupole-time-of-flight (Q-TOF) mass spectrometry, atmospheric pressure photoionization mass spectrometry (APPI-MS), Fourier transform mass spectrometry (FTMS), matrix-assisted laser desorption / ionization-Fourier transform-ion cyclotron resonance (MALDI-FT-ICR) mass spectrometry, secondary ion mass spectrometry (SIMS), microscopy, microfluidic chip-based measurements, and surface plasmon resonance.
[0193] In some embodiments, proximity extension assay (PEA) is used to detect one or more peptides (and optionally, to determine relative levels). In embodiments, PEA comprises the simultaneous binding of a pair of proximity probes to a neighboring biomarker. When the pair of proximity probes bind to the biomarker, nucleic acid domains are able to interact and form a nucleic acid duplex, which allows at least one of the nucleic acid domains to extend from its 3' end. This extension product forms a detectable nucleic acid detection product, optionally after amplification, for example by PCR. Exemplary PEA methods are described in more detail in WO 2012 / 104261 and US2015 / 0044674, which are incorporated herein by reference. Target peptides can be detected individually, or more preferably, multiple target peptides can be detected simultaneously in a multiplex assay configuration.
[0194] In some embodiments, multiple reaction monitoring (MRM) assays are used to detect one or more peptides (and optionally, to determine relative levels). Various MRM methods are available. In examples, MRM assays use a triple quadrupole mass spectrometer coupled to liquid chromatography to detect or quantify target peptides. In the first quadrupole (Q1), a peptide corresponding to the target protein is selected. The peptide is then fragmented in the second quadrupole (Q2), and a filter is applied to allow the specific fragment to enter the third quadrupole (Q3), where the intensity of that specific fragment is measured. Target peptides can be detected individually, or more preferably, multiple target peptides can be detected simultaneously in a multiplex detection configuration. Further non-limiting examples of MRM are described in US20190277846 and US20180024108, which are incorporated herein by reference.
[0195] In some embodiments, a quantitative platform integrating nanoparticle (NP) protein crowns with liquid chromatography-mass spectrometry is used to detect one or more peptides (and optionally, determine relative levels). In an embodiment, the platform is the Proteograph platform. In an embodiment, the protein crown is a layer of proteins adsorbed onto the NP upon contact with a biofluid. Changing the physicochemical properties of engineered NPs results in different protein crown patterns, thereby enabling differentiated and reproducible probing of biological samples. In an embodiment, the Proteograph platform uses a multi-NP protein crown method and mass spectrometry. In an embodiment, this method comprises four steps: (1) NP-biological sample incubation and protein crown formation; (2) NP protein crown purification by magnet; (3) digestion of the crown protein; and (4) LC-MS / MS analysis. In this case, each biological sample-NP well is a sample, with a total of 96 samples per plate. Target peptides can be detected individually, or more preferably, multiple target peptides can be detected simultaneously in a multiplex detection manner. A non-limiting example of NP-based protein crown detection is described in WO 2020096631A2, which is incorporated herein by reference.
[0196] In some embodiments, aptamer-based assays are used to detect one or more peptides (and optionally, determine relative levels). Generally, an "aptamer" refers to a nucleic acid that has a specific binding affinity for a target molecule. In this context, the "specific binding affinity" of an aptamer for its target means that the aptamer's binding affinity to its target is generally much higher than its binding affinity to other components in the test sample. Methods for generating aptamers are known in the art and include, for example, the SELEX method (see, for example, U.S. Patent No. 5,475,096). Aptamer assays that allow aptamers to capture their targets in solution followed by separation steps designed to remove specific components of the aptamer-target mixture prior to detection are also described (see, for example, US2009 / 0042206). Exemplary solution-based aptamer assays that can be used to detect (and optionally quantify) proteins in biological samples include those described in US20210215711A1.
[0197] In some embodiments, one or more polypeptides are polypeptides that identify one or more specific proteins. Proteins can be identified by any of a variety of characteristics that are identifiable to the protein. For example, a specific protein can be identified by a specific epitope or by a specific amino acid sequence (or subsequence) that distinguishes the protein from other proteins. Thus, a polypeptide that identifies a specific protein can be a portion of that protein sufficient to identify the specific protein as the source of that portion. In some embodiments, a plurality of different polypeptides include polypeptides that identify proteins selected from List 1 of Table 1 herein (e.g., 2, 5, 10, 20, 30, 40 or more proteins). In some embodiments, a plurality of different polypeptides include polypeptides that identify each protein from List 1 of Table 1 herein. In some embodiments, a plurality of different polypeptides include polypeptides that identify proteins selected from any of Lists 2-19 of Table 1 herein (e.g., 2, 5, 10, 15 or more proteins). In some embodiments, a plurality of different polypeptides include polypeptides that identify each protein from any of Lists 2-19 of Table 1 herein. In some embodiments, a plurality of different polypeptides include polypeptides that identify proteins selected from List 20 of Table 1 herein (e.g., 2, 5, 10, 15 or more proteins). In some embodiments, the plurality of different polypeptides include a polypeptide that identifies each of the proteins from List 20 in Table 1 herein. In some embodiments, the plurality of different polypeptides include a polypeptide that identifies one or more (e.g., 2, 3, 4, 5 or 6) of CHAD, KRT19, MMP12, PTN, SERPINA3 and SPP1.
[0198] In some embodiments, the plurality of different polypeptides include polypeptides that identify proteins selected from the following list of proteins (identified by UniProt reference numbers): P31483, P21964, Q9NRD8, P16860, O60635, O96017, Q9UKL0, Q8NHS0, P58546, O43854, P40225, Q99549, P08319, P25815, Q8TE57, Q04760, Q9BYF1, O14793, Q9NWQ8, Q13444, P34913, P09496, P34947, P55259, P01375, P52789, P09668, O75354, Q9 BWV1, P22004, P05231, P46379, P40818, P62736, P51161, P09237, Q15165, Q 92558, O43186, P08670, P07585, Q15831, P19429, Q9UKP3, O95988, P36952, Q 16619, P61978, P17676, Q96N03, Q13105, O95684, P21246, P34998, Q6UWL2, Q969D9, P35218, P55082, P17516, O15354, Q12912, P31997, Q9NRV9, Q9Y2B0, O95183, P13807, P20718, Q9H5Y7, Q8NC01, O75356, Q96A56, Q9GZM7, P27352 , P12104, Q9NQX5, P12724, Q9UBU3, P35754, P41159, P09382, P40189, Q92692 , Q15067, Q16620, P21583, P31431, P09417, Q8WVQ1, Q15846, Q9UKJ0, O0016 1. Q6WN34, Q92823, P00568, Q13043, P09525, Q05315, Q9UHL4, Q03154, P1064 4. O94903, P16234, Q9H773, O14917, Q9H7M9, NT-proBNP, P31949, Q9Y4X3, P 01222, P21980, P21549, Q9UMF0, Q6GTS8, Q9NY25, Q9HBB8, P16112, P55285, O 60664, P08263, P52888, Q969P0, O75340, Q9ULL4, P41218, P48357, Q9Y286, P51693, O95502, O75791, Q06418, Q12864, Q9Y5X1, Q13541, Q9UHD0, Q8WX77,Q8WTU2、P78380、Q99674、Q8NI22、P23526、Q8IW75、P09601、Q9BQR3、Q6PJW8、Q8IZP9、P06858、Q13158、Q9NR28、Q86VZ4、P35247、O95544、Q14956、P18827、P10145、Q53H82、Q9BUD6、Q16820、Q9Y5K6、P41236、Q13275、Q96LA6、P19022、P00797、Q8N1Q1、Q07108、Q9UK05、O95841、Q9UEW3、P02462、P07204、P01241、A6NI73、Q01973、Q16773、P09467、P42830、Q9BQB4、Q76M96、P19971、Q92520、P07711、P04792、Q99523、P20711、O60496、P07911、Q13361、P00750、O75326、P23141、P22748、P55058、P01130、P13598、Q8NBP7、P15090、Q76LX8、P08833、P33151、Q16270、P54760、Q96AP7、P32942、P08118、Q06141、P01589、P07858、Q9UBP4、Q86U17、P04066、Q14767、Q9NQ79、O14798、Q5VY43、P48304、P15085、Q07507、P17931、P04275、P55808、Q03167、P14555、Q9Y275、P08581、Q9H2A7、Q9UM47、P07451、P09619、P80370、Q14162、Q99988、P04080、P02144、Q13822、P08236、Q01638、Q13740、P48960、P17813、P31146、P12111、P16581、P15086、Q15828、Q9NNX6、P04054、Q9H1U4、P19021、P48745、P20062、O75023、P18065、O00584、P19961、Q12860、Q13231、P39060、P25445、P23284、O15467、Q13867、Q13332、P19957、P35590、P09093、P46531、P04746、P78324、P04085、Q9HD89、O15031、P24158、P05107、P13987、O75594、Q12884、P05121、P00533、P13686、P02452、P20160、P42574、P10451、Q16769、Q14393、P42785、Q8TDL5、Q16663、Q8N423、P10586、Q9Y4L1、P15907、Q8NHL6、P43121、P00740、P12830、P15529、P13591、P12318、Q9UBR2、P18428、Q12794、P07478、P07359、P98160、P08887、P59665、P24821、P16109、Q14515、Q86VB7、O95998、P20023、Q9NZK5、Q13508、Q15485、P80188、P30530、Q99650、Q15113、O14786、Q96KN2、Q6EMK4、P19320、P00441、O75015、P07339、Q16853、P15144、P08174、O00533、P02786、P05556、P10646、Q9BXJ1、Q9NPY3、P10721、P14543、O95445、Q96H15、P08571、Q99969、A1L4H1、Q07654、P35443、P55774、Q9Y5C1、Q16627、P08709、P41222、P06681、P24592、Q15582、P36222、Q06033、Q9UGM5、P49747、Q92820、P00915、P13501、P05451、Q12805、P03950、P27487、P04070、P05362、P01034、P17936、P01033、P14902、Q14160、P12829、Q9BY49、Q9NZN3、Q96C92、Q5SW79、O75506、Q15477、P04141、P21817、A6BM72、O00291、Q8IZC4、O60701、O14958、E2RYF7、Q9NVZ3、P23634、Q9Y4C8、Q9Y623、P54709、Q07973、P48507、P06753、Q04695、P25391、Q15059、O00567、Q9NZJ5、P35228、Q13503、P08913、P33121、Q9BY32、P30049、P10109、P55011、Q01780、Q6UWF7、Q9Y3B8、A6NCE7、Q08499、P46783、Q96DA2、P49755、Q96HD9、B6SEH8、O43734、O95180、Q9H2M3、P06729、Q96IW2, P55769, Q9Y2W1, O95858, Q9H347, P78524, Q14353, Q15370, P20929, Q9BW61, Q5TA50, O15305, P05026, Q86UW2, P38935, Q14088, Q9Y2Y0, Q8WZ42 P12270, O75521, P05976, P14415, Q9UFP1, Q9BZC7, Q6NZY4, Q9NYX4, P16066, Q99707, Q8N8E3, P37058, Q92935, P21673, O43290, Q96K76, Q13296, Q6P4F 2、P05000、P57078、Q9UKX7、Q02127、Q6ZN66、Q9BV94、Q07075、P23511、Q96L B8、Q8NET8、Q9NV35、Q16774、Q16836、P54296、Q9BZL6、Q10587、A6NHS7、Q150 18、O00425、Q9UNN8、Q14807、P35606、P20382、Q96PU4、P00966、P48668、O00 327、O95670、P50461、Q3SXY8、Q03013、O43896、P59901、Q01484、P19838、P22 033, Q12986, Q01581, O94766, Q14781, Q96A35, Q58F21, Q8NFP7, P46926, Q9UBV2, Q5JTV8, Q8ND90, P32241, P35609, O75427, Q93052, Q86VR7, P41227, Q5 W0V3、Q86VP3、Q99598、Q13563、O75534、A6NDB9、Q5VVQ6、Q96EU7、P55010、Q 9Y2L6、P13224、P0C7L1、O15018、P10082、Q7Z7H5、Q16206、P29536、Q14324、Q 96ID5, P13929, P20645, P23327, Q9H173, Q9BTK6, P01225, Q8TER0, Q0VD83, O95980, Q13316, P50053, Q14457, Q99942, I3L3R5, Q99807, Q53T59, Q8N668 P55809, O75348, P11532, Q9Y5X3, P05305, Q8WZ75, Q8IVF2, P35914, Q14643, Q9BQI0, P36776, Q9H7C9, O14841, Q8WXC3, O75061, Q8NC42, Q8TAE8, Q5GAN6P35520、P30084、Q8WUF8、O43423、Q13137、O94979、Q16621、Q9H3K6、P07098、P21754、P07492、P20042、O60476、Q9NYZ4、Q09666、Q92835、P43487、Q6ZRY4、P07355、Q6YN16、Q9UJ70、O95825、Q24JP5、P02458、P09543、P50914、Q7L266、P01189、Q8NFL0、Q96DR5、Q9HB40、Q8IWT1、Q5FWE3、Q6UY14、Q9BV79、Q6UWR7、P07942、Q9NR61、P09681、P58107、Q12841、P0DPI2、P08590、Q86X76、Q96DC8、Q8TCD5、Q7Z7M9、Q00872、Q14914、Q9UBQ7、Q9BVM4、Q12982、P33681、Q6UW49、P51511、Q9BW04、O14933、Q8WWV6、P23919、O75711、Q6UXI7、P29692、P02008、Q9NQR4、Q9BQS7、Q9Y2E5、Q9H3S4、O43405、Q96C24、O60234、Q7Z304、P78539、Q9P2J2、Q8N4F0、P53674、P16035、Q8N436、Q13442、P14854、P23467、Q13428、O75223、O75154、Q6NUS6、Q96EM0、Q96FZ7、Q969H8、P98161、Q9BXD5、P54687、Q9BXN1、P51688、O14960、P23471、P32320、P08138、Q6PI73、Q8NDI1、P08582、P52209、O43681、P15502、Q969X0、Q96MK3、Q8IZF2、Q96AG4、Q7Z7K0、P07093、P62072、P61026、P45954、Q6ZMM2、P05413、Q15388、Q9UBR1、P49593、O00194、P13667、P23560、P30046、Q86TH1、P02730、P13796、Q9Y303、Q6H9L7、P07288、P16410、P40199、Q8N6C8、Q02817、P98095、P02461、Q6UWP8、Q6UVK1、P39059、Q9BYJ0、Q9HCU0、Q96CG8、Q96NZ9、P47972、P02818、Q8N114、Q6IBS0、P30405, P32971, Q9Y2Y8, P35579, P13727, P08575, O43280, Q9NRR1, O75339, Q9H2X3, Q9Y646, P10645, Q04721, O95965, Q9Y251, Q8TDY8, Q15063, P0821 7、Q9UQP3、P17900、P37837、Q8WWQ8、P55000、P12277、Q13510、P11279、P076 02、P17174、P61916、P19878、P40933、P11274、P52564、Q9UN19、P24394、Q6ZU J8、P01730、Q13241、P35613、P50452、O43915、O00253、P10147、Q92609、Q9G ZT9、Q9Y266、Q14242、Q12918、Q3KPI0、Q9NRM6、Q01344、P02745、Q9HBG7、O9 4992、Q08174、O60449、O15455、P22304、P43234、P14210、Q12866、P51671、P 42701、P09874、Q5R372、Q13459、O95760、P14784、Q8NHJ6、P01584、P60568、O 76038、O95715、Q8N6P7、P22301、Q9UPV0、P28838、O60934、P57771、Q03426、 O14904、Q9Y478、P20809、P05412、O43707、Q96PD4、P05112、P35225、Q96AX2 Q9NYY1, Q96P31, Q9NP70, Q13007, Q9HCU5, Q8WV07, Q9Y2J8, Q9Y3P8, Q8IU57, P30838, O14867, P19801, Q16552, Q7Z739, O60575, P26951, Q8TAD2, Q9P0M 4, Q7Z6M3, Q8TCS8, Q5T4W7, Q99748, P48061, Q04759, Q12933, P42768, O95379, Q13219, Q13574, P63241, O43736, O60542, P13693, P09038, Q9Y5A7, Q6U XK5, P01375, Q13651, Q96RJ3, P27540, Q969V3, Q9UHF4, Q06520, Q6UB28, Q0Z7S8, O60880, Q12968, P78362, P01903, P78410, O43521-2, P01583, P01579Q05084, Q7L8A9, P05113, O43597, Q13261, P12034, Q92844, O95644, P09919, Q9BXJ7, Q13291, P51617, Q12778, Q14435, P30048, P32456, P01591, P55957 Q12765, Q6ZMH5, Q8N8S7, Q9Y6K9, P18564, P58294, Q9HB29, P05231, P12872, Q96DB9, Q96LC7, O75475, P19474, B1AKI9, P13232, P13747, Q9UNK0, P3324 1, Q8WTT0, P13725, Q8IVG5, Q8TD46, Q9UHC6, P50995, Q6DN72, P23582, Q8NDB2, Q01151, P45984, Q9NRJ3, Q9NZN5, Q9HD26, P28827, P29965, P16455, Q9BT 73、Q8N608、P28845、Q9UNE0、P20849、Q9HCM2、P01588、P23229、P80098、O76 036、P01374、P42575、P24071、Q9NWZ3、Q6UXB4、P37235、Q9Y258、Q9UKX5、Q9H 0P0, P08727, P20340, Q9UIB8, P78310, P32970, Q29983_Q29980, O14788, Q9UDT6, Q9C035, P26022, Q07065, P80162, P20783, Q14773, Q16698, P50591, Q8 WXI8、O94856、P49771、Q14005、Q15517、O15169、Q9NQ25、Q9UMR7、O43561、P 10145、Q96SB3、P41217、P14317、Q9BZW8、Q16719、O00273、Q13478、O75077、Q 9UQV4、P24001、P36959、P30203、P20273、Q6UXB2、P68106、P12544、O95971、 P43489、P01137、Q15661、Q04637、P48023、P40259、Q03431、Q9Y6Q6、Q96LA5、 Q9BXN2、Q9H4D0、P29460、P42702、Q99616、P00813、P30044、O60884、P15692、 O43508、O15444、P10144、P01135、P78556、P03956、P49763、Q9BY76、O95750、O14836, P46109, Q03405, O43598, Q9HCB6, Q9NQ30, Q16651, O00468, P29350, Q07325, Q9UII2, Q9H008, P19876, Q6UWV6, P09341, Q9H3U7, Q92583, Q99538 、O00182、Q03403、P53634、Q5ZPR3、P55773、P25116、Q9NZC2、P34896、Q1538 9、O00626、O75888、P47712、Q15166、Q14118、Q9BZZ2、Q8NFT8、Q99685、O0058 5, P19256, P09326, Q5KU26, Q6GTX8, P09603, P51888, P16422, P01133, P02778, Q92484, Q7KYR7, O43291, Q9Y6N7, Q8WU39, P35625, O43639, O76096, O003 39、O75462、Q9UJU6、Q15109、Q08334、Q9HC38、P21709、P27930、Q9UJA9、Q9N ZV1、P15260、P21860、P09238、Q9NR12、P11684、P18510、Q14210、Q14116、P13 236, Q99435, P36941, P30613, P55145, O00241, P29279, P0DMV8, O00300, Q8WXD2, O14773, Q96KG7, Q4KMG0, O95866, P56470, O75563, P01127, Q96PL1, O9 5633、Q9UKU9、Q9BYZ8、P24387、Q99983、Q13232、Q9Y3D6、P19883、P15291、P 12532、Q9UHX3、Q9NQ76、Q6UXH1、Q99895、P07148、Q16363、Q92956、P22466、Q 8TEU8、Q8IYS5、O00175、P25942、P54317、Q9BU40、Q8N907、Q04609、P62834、 P10070、P61328、O00206、P46013、P24864、Q9H832、P85299、P49715、Q9Y4C1、 Q9NP95、P06401、Q6UXL0、Q9NP85、P15531、Q6UXM1、Q05329、P09693、P24928、 Q09472、Q9HBE5、P43378、Q92185、Q9BY41、P10767、Q01201、P41273、P03372、Q9UPW0, P25490, Q6R327, Q13190, Q6UXZ4, Q9NQI0, Q8WX93, Q16665, P15927, Q99665, O75365, P24530, Q04837, O00401, Q9NZS2, Q9NWV8, Q7Z698, Q9Y2I 7, Q9NRR2, Q96EB6, Q12888, Q9NY59, Q8WVV4, Q13114-2, Q6EBC2, Q14511, Q8NHP1, O95429, P36897, Q8IX19, Q6B9Z1, Q9Y3D3, Q8IYW5, Q96MM7, Q9UHN6, Q5 TBC7, P09923, Q96D71, Q92574, P98170, P36551, O00148, O43583, P48730, Q9NS62, Q8NI17, Q93062, Q9UIK4, Q8NBK3, P35219, O60447, Q8IV38, P54274 Q9GZN4, O75173, P17643, Q15465, Q7Z5L3, Q96EP0, Q9NQ66, Q5JS54, O43184, Q9UHA7, Q15697, P17050, Q9H0U9, P56645, P16671, P26436, Q96PX8, P29536 、P14902、Q14160、Q16718、Q15223、P11487、P48546、Q8WV28、O60500、Q96PL 5、Q86UE4、P20701、P52630、O43320、P81534、Q9Y2X7、Q15399、Q9H3T2、P552 11, Q99584, Q9UHI8, P23743, Q99062, O75688, Q5T2W1, P49757, P11234, P06213, O95157, O60437, Q9H7Z7, O15400, Q02880, P06730, Q15762, Q9BV40, P49 765、P32927、Q5QGZ9、Q9NZH8、Q16643、Q9H171、P09564、O60238、P24666、Q9 6PU5、O94916、O95835、Q8N556、P17301、Q8NEU8、Q02223、Q5JS37、Q6QNK2、Q 8N8U9、P20908、Q9BX67、P22303、Q9BUH6、O76074、Q9BRK3、Q9H7Y0、P01037、 P04083、Q8TEA8、Q9ULI3、P30040、Q96QR1、O75347、Q9UGN4、Q8NDA2、P06280、Q6P5S2、Q16378、P61457、Q9UI42、P27348、Q99574、P06132、Q5TDH0、Q9HCU4、P40197、P04406、Q9NS98、P00325、P55083、Q8TDQ7、Q9H939、Q9BUN1、O75190、P04090、O95393、O43399、Q96HD1、Q02952、Q4VCS5、P10912、Q9BXJ0、P30101、P0C862、P30047、Q15276、O14745、Q6FHJ7、Q58EX2、Q9HB71、P14091、O60279、P20155、A6NC86、A2VDF0、P49862、P09529、P25686、Q86SX6、P22455、Q53FA7、O96007、Q8NHV1、Q9NQ48、Q96AJ9、Q6UX06、Q96A49、P01210、Q92619、P21854、P52758、Q9Y4D1、Q9UK23、Q8IUZ5、Q9NRS6、Q5VTT5、P37840、P07311、Q04323、Q6P589、P32321、P56192、O15335、Q9P2T1、P35611、P54764、Q13976、Q8N0X7、Q6GMV3、P50502、P32119、P0DML2、P40925、P10599、Q92686、Q08830、Q13790、P07333、P00390、P0DJD7、P02748、P17927、P12821、Q6UXB8、P22897、P16233、P50552、P04180、O43493、O00602、Q93091、P12955、Q0ZGT2、P20061、P02652、P04278、P00742、P07307、P08294、Q86YW5、P09172、P27169、P16442、O95497、Q04756、P06276、Q9BXR6、P54108、P06396、P0DN86、P26927、P04040、P34096、P07998、Q01459、P06744、P04114、P01019、P0DUB6_P0DTE7_P0DTE8、P02765、O14791、P07358、P36980、P08185、P05543、P36955、Q96PD5、Q08380、P10909、P61769、P43652、Q92496、P35542、P0DOY2、P00748、P11226、P00734、P02649、Q9Y5Y7、P00746、P07225、P03952、Q16610、O75882-2、P05154、O00391、P02775、Q96IY4、P20742、P06727、P09871、P02776、P05452、Q14624、P05546、P02647、Q15848、P10643、P08519、P00736、P05090、P14151、P43251、Q9NZP8、P02763、P19827、P02654、P00751、P01009、P49908、P01031、P02750、P04196、P00747、P27918、P02787、P02751、P02766、P05160、P02774、P08697、P02671、P05155、P01008、P01024、O43866、P29622、P02743、P01011、P04217、P08603、P03951、P05156、A1E959、P06748、Q9NRG1、P58417、Q9H3R2、Q6UX27、Q15043、Q9HA65、O60242、Q06323、Q92765、P16278、Q9H1C3、Q9NR71、O95994、O60609、Q9HD42、Q16762、P61244、P09211、Q6UXK2、P28325、P78423、P29466、P07196、Q9Y2W6、O94985、Q6UWL6、Q9UHV9、P50135、Q9H3S3、O14763、O15496、Q9Y6Y9、O00559、P11464、P15509、P42892、Q96NZ8、P78333、P78560、O14618、P07306、P13500、P02533、Q9NRW1、Q6ZMC9、P53582、Q6ZMJ4、Q8IU54、Q8N2G4、P20936、P48047、Q9BXS1、Q4LE39、P05455、Q6P1M0、Q9NXA8、Q6PIL6、P53985、Q86VW0、P49023、P51531、P08758、P15018、P13385、P15336、O96013、Q96JP9、P45452、P10636、O14579、P09769、P53539、Q9HAW4、O95466、Q96FQ6、O00399、Q8WTV0、P22676、Q9UBB4、Q9H2W6、Q07812、Q16864、Q9BW30、P23276、Q92597、P26440、Q9UM07、Q9UJY5、P14625、Q5JZY3、Q8WUW1、Q15126、P30519、Q86Z14、Q9H0C8、O14523、O94813、P13647、Q9NS71、Q13308、Q9UKV5、Q9NSK7、P29474、P60484、Q15633、Q9Y5P4、Q9UPY6、P01375、P29475、P51452、P41208、O60240、Q9Y6E0、P30039、Q03393、Q13426、P39905、Q9UK53、Q02643、Q9Y2V2、Q9UJ72、Q8NBI3、P50453、P52798、Q92932、Q02083、Q9H5V8、Q16740、P52943、P36269、Q9Y6D9、P63098、O95256、Q9HAN9、Q96IU4、O43155、P07741、O00220、Q96PP9、Q92917、Q8IUN9、Q96B36、Q92851、O95817、O95630、Q2TAL6、Q9UBT3、O00308、P09466、Q08629、Q6UX15、Q53H47、P49789、P05231、P18031、P29218、P14868、O95727、P01138、Q9BZM5、Q8TCT1、Q2MKA7、Q9NQ88、Q9Y680、Q86WV1、O14944、Q13451、Q9Y4K4、P16444、O76070、Q15814、P08962、O15232、P48740、P38484、Q6ISS4、Q6P1N0、Q07011、Q8WWN9、O15197、Q9NPH6、P10145、Q9P0K1、Q9H477、Q8N474、P40222、P02771、Q02790、Q96GW7、Q9NZQ7、P01258、Q14108、P56279、Q9UL46、Q9BS40、P23588、P08134、Q8NFP4、P22079、O43557、Q6XZF7、P37023、O95407、Q06830、Q8WV92、Q92752、Q99426、P12644、Q9BYC5、O43927、P15151、Q08345、P06733、Q96GP6、P14384、O14594、Q96B86、Q6P4E1、O14625、P35237、P52823、Q08708、P01178、P17405、P50225、P15311、P13611、Q8NBJ7、Q8NBS9、P55291、P22223、O14737、P57087、P21757、Q9NP79、P05060、P78325、O94907、P30533、Q9P126、Q10589、P01236、Q92854、P28906、Q14739、P01732、Q99731、Q9H446、Q9UNZ2、P21217_Q11128、Q9BRF8、P14207、P22894、Q10588、P09104、Q6ZMJ2、Q16288、Q6NW40、Q01469、O00214、Q9HCK4、P13473、P12429、Q96CD2、P28908、P48052、P25774、O94779、O75509、P04216、P04233、Q9NZD4、Q14112、P01215、Q96RD9、Q9NPH3、P16871、Q9P232、Q8N6Q3、P06734、P00749、Q9Y240、Q8TCZ2、P20774、Q96FE7、Q6UXH9、Q969Z4、Q9NQ38、P20333、P14174、Q9UBW5、P12081、P04118、P30740、P23381、Q86T13、Q99972、Q15155、Q15818、Q9UGT4、P05164、P55103、P22692、P11215、P02749、P50895、P19440、P31948、O76076、P30086、Q96F46、P08254、P14778、Q04900、Q02487、O60462、O15389、P05067、Q6YHK3、P04179、P00352、O43278、O76061、Q14126、Q9Y279、P11717、P15289、Q99727、Q02747、O15117、Q9HCN6、Q5T2D2、P22105、P16284、Q9Y624、P30043、P14209、P01833、Q8IWV2、P22749、P14780、P04155、P28799、P00918、Q99497、P17538、Q9UKK9、Q08431、Q9UKJ1、Q16881、Q14696、P35442、Q8TDQ0、Q13093、P08648、P63313、P20827、P00995、Q8N149、P07108、P19438、P23280、Q6UXG3、O43505、O15394、Q5SXM8、Q96Q89、Q9NQW8、P26441、Q96GG9、P48169、P49792、Q03721、Q96L15、Q8N111、Q96P66、P35221、Q13936 、O95372、P04234、P25440、O95049、P08908、Q30KQ4、Q13509、Q8IVU1、Q6P5Q 4, Q9NZA1, P0C7M8, P63172, Q02641, Q14894, O76039, Q7RTW8, Q5U5Z8, P07954, Q8N8R5, O15350, Q9ULW2, Q9Y3E2, Q9BZ29, P47874, Q8IWB1, P23435, Q997 00、Q8WTQ1、Q9P2M7、Q14129、Q13002、Q9H492、Q6QNY1、P26378、Q86T26、Q99 653、Q6QEF8、Q12774、Q8WVC0、Q9NXH3、Q13084、P30419、Q8N967、Q8IZJ0、Q1 5042、Q96PV0、P05549、Q0VDD7、P28222、O15382、Q687X5、Q5T871、Q9UL42、O 95196、Q9C0A0、O15013、P12319、P54284、Q96I25、Q01432、Q9Y6X8、Q8WW22、P 04439、Q9NWU2、Q5SW96、O15240、Q92859、Q14627、Q9Y3B9、Q96A32、Q9NPB3、 Q9ULU8、Q6GQQ9、Q96HC4、Q9BV20、Q14184、P43007、Q68J44、P14649、Q96G03 Q9NY46, O75718, B6A8C7, Q9Y4C0, Q7L311, Q15735, P48431, Q6ZVN8, O43665, Q6Y7W6, Q8WXW3-4, Q9H4X1, Q16625, P78318, Q6ZVM7, Q9H6S1, P52179, Q99 250、O75631、Q96J84、Q9UKV0、Q19T08、Q9ULH4、O95295、A8MVW0、Q9BZE9、P6 1764、P68400、Q13224、O75312、P60880、A6NGG8、P05787、Q99856、Q96PH6、P 52788, Q9Y6U3, Q9Y285, Q9HCM4, P26715, Q6P9F5, Q8WXG9, O15020, P55283, P10523, O14503, O43396, Q8IY33, Q2M3V2, P48775, P61366, Q10571, Q8N2Q7Q12809、O15212、P61266、Q86UP6、P0DKB5、P11229、Q5VT99、Q16650、Q6ZUT3 、Q6UWJ8、P63211、P20823、Q16401、O75121、O15234、Q86UW9、Q86WK6、O15083 、Q53GD3、Q8NFZ4、P54105、Q93015、Q9UHL0、Q8N6Q1、Q16520、Q96KJ4、Q9UGI 9、Q53EL9、P04062、Q7L0J3、P01270、Q8IY31、P10301、P98164、Q92834、O1496 7, Q9H461, O43474, Q9NY72, P43220, P41732, Q17R60, Q9UKW4, Q9UJC5, Q96BJ3, O60218, P52799, P43146, P49069, P09683, Q7Z3D4, P24386, Q96QH8, Q9UM 54, P08651, Q9HCY8, O60662, Q96JB5, A6NFN3, Q9P2M1, Q9HBL6, P10746, Q2UY09, Q9NPD7, Q9BY14, Q9BZJ3, P48436, Q9NUG6, P09455, Q0P6D2, Q8WWM7, Q07 617, Q9BX66, Q15398, P35556, Q8IWP9, Q9UNY4, Q9UH03, P53367, P60763, P63010, O14578, O60939, P20916, Q9BX10, Q86TM3, Q15102, P54277, P50897, Q8 TC05、O95202、Q9UQ16、Q9H6Q3、Q15714、Q8IVM0、Q86YD3、P29536、O95954、Q 5T848、B2RUY7、Q14151、Q92888、Q96A00、Q15366、O60869、O75167、O95467、P 22102, Q07021, Q08378, O75792, O43432, Q9BUJ2, Q96PE7, Q6UWW0, P51687, Q01826, P04808, Q8WWV3, Q9H6H4, O60469, O75592, Q9H777, Q14197, Q5JSP0 O43776, P31751, P29377, P14902, Q99460, Q9BXI9, Q5T5Y3, Q6PKH6, Q13277, P21579, Q53GL0, O60890, O15269, Q9H1P3, Q96RU2, P01350, P78352, P26718O77932, Q92599, P13995, Q8N163, Q8IWY9, Q86VP1, Q9H251, Q14149, P43320, P51649, Q9NWM8, Q14160, Q9BUP0, O94830, Q8IWQ3, Q9NPG4, Q6UXV0, Q4ZHG4 、Q13938、P47813、Q9UKY0、P54315、Q15256、O14530、Q15025、P13861、P1009 2、Q9ULA0、P59780、Q8WYQ3、Q9NPE2、Q6UX71、Q6XQN6、P46937、Q9BXI3、Q9P2X 3、Q99704、Q6P1J6、O60235、Q03252、Q8TAT2、Q5F1R6、Q03014、Q6QNY0、Q9BP X1、P32455、Q9Y3C0、P62760、Q2L4Q9" 37, Q9BRJ6, Q7Z4V5, P08579, Q9BQT9, Q9UBC9, P0CG30, Q8IXS6, O75146, Q9NRY6, Q96CN9, Q9NXV2, P07320, P54252, Q8IXM2, A8MVW5, O60220, Q8TF65, P53 814、Q07817、Q00722、P40313、Q8N4C8、P54819、Q9NR46、P09110、O14713、Q1 3145、Q9NX58、O95498、Q07954、Q9BTE6、Q6UWW8、P01229、P05937、P08069、P0 0519、Q96I82、Q9BS26、Q14241、Q9HAV7、Q9BSL1、Q9C0C4、O00592、Q9UK85、Q 7Z5R6、P30041、P82980、P47992、Q9NSA1、Q9NTU7、Q14213_Q8NEV9、P13726、P 06756, P61218, Q96NA2, P50579, Q9UBG3, Q14790, P35637, Q13490, Q9UQB8, Q6EIG7, P80075, O00292, Q9BSG5, Q99075, Q9Y5W5, P42658, Q99717, O43699 Q86SJ6, P35318, P35813, Q7L5Y9, P01375, Q9Y265, P42331, P06850, Q8IUK5, Q9BSW2, O95388, Q2VWP7, Q00796, O95786, Q9UHF1, P14136, P31994, P55789P55273, Q9Y243, P22307, P43628, P31350, P39748, O14964, Q9NRA1, Q05516, P48643, P46060, O75569, Q6UX82, Q99683, P01242, Q08AG7, Q96DU3, P43629 、O43752、O60828、P35070、Q8IWL2、Q7Z7D3、P34130、Q9UKR0、Q6NXT1、P54727、Q6BAA4、Q92982、Q8NBZ7、P41586、O75787、Q15797、Q96NB1、Q07960、P5074 9、Q6PGN9、P06731、Q8IWL1、O14662、Q7Z6M1、Q9UQQ2、P25786、Q9H4P4、O754 93、Q9NS15、A4D1B5、P49788、P21810、Q7LG56、Q9P0J1、Q9Y5V3、Q8N5S9、Q7Z4 34, P07332, O15116, P43490, O75380, O60907, Q01543, Q9UKS7, Q06787, P04637, Q8WUX2, Q9Y6A5, P34949, Q8WYN0, Q96PQ0, P15121, P36888, Q9Y662, Q8T E58, Q9BYE9, P05231, Q7L5N7, P55008, P40198, Q9Y223, Q9Y5L3, P05783, Q8TD06, Q9Y2Z0, Q9P0V8, P51580, O43524, O75695, O00233, Q9GZY6, Q5VIR6, Q9 UJ71、Q86WD7、Q15427、P10606、P51692、P0CG37、Q9H4A9、P08473、Q9NUY8、P 17948、P10747、Q16772、Q9BUE0、O00186、Q3B7J2、Q6P2H3、O00221、Q9BQ51、O 94760、Q9UHD8、P30260、Q9Y639、O95831、Q6UXD5、O75054、Q9Y570、P07947、 P15848、Q11201、P55039、Q8IX05、Q12846、Q96RT1、O15357、P23515、P28907、 O60911, Q7Z5A7, P16870, O60760, Q96EK5, Q8N9I9, O60825, Q9UBM4, O60763, P07949, Q8N386, Q8NEZ2, P15514, P18627, Q86SF2, O00622, O75144, Q13576O00748, P58499, P26010, Q9UKR3, P49441, O43570, P37108, P38936, Q13561, O14828, P07948, Q9NZT2, P01275, P50583, Q9Y653, Q8N129, Q49AH0, P29317 、Q9Y5K8、O00451、P29459_P29460、Q99795、Q99536、Q9GZV9、Q9H156、P9807 3、Q9P0G3、O43715、Q9ULX7、Q86SR1、Q9C005、Q13421、Q15116、Q9UJM8、P0518 7, P25685, Q8WXI7, P10145, O43827, P39900, P09105, P13521, P50120, P09960, Q9HAV5, P05089, Q9H4F8, Q02742, O14558, Q14203, Q9Y336, P01303, Q9H6 B4、P47929、Q6UWN8、P40121、Q16595、Q8IXJ6、P80511、Q86SJ2、P98082、Q9B XY4" 397, Q7Z5L0, Q96JA1, Q16790, P09758, O60243, Q9NPH0, Q96I15, P16562, P27695, Q02246, Q9BZR6, P62166, Q10471, Q8WWY7, Q6PCB0, P51858, Q16775, Q8 TDQ1、P02760、Q9H3G5、Q496F6、P35052、P56159、P35475、P32926、Q96D42、P 20472、O15123、P29017、Q14508、O43895、Q9UBX1、P07237、Q6FI81、P41439、Q 5JTD0, O14974, P37173, Q15303, Q92832, Q96NY8, Q96J42, Q9H8J5, P21741, Q9BYH1, Q16543, Q9NZ53, P20851, P41271, P35916, P20138, Q96SM3, P35968 Q02763、P21589、O95721、P09486、Q9UP79、P32004、O43464、Q7Z4W1、Q9HAT2、 Q8NCC3、P21802、Q14512、Q9NP84、O00244、Q96PD2、P78552、P01298、P13688、P26447、O75629、P09958、P48307、P18084、P15328、O60259、Q9UJ68、P26842 、P06870、P12931、O43240、O95274、O00548、P49767、P04626、Q16674、Q9UBX7 、Q92876、Q6NT46、Q7Z460、Q7Z4W2、Q9UPY8、O00337、P23771、Q5VSG8、P4363 0、A7E2Y1、P30304、Q9P2D8、Q14123、P12004、O60941、O43504、P19075、O6050 2, Q13443, Q93033, Q8NHZ8, Q92499, O75781, Q6UWK7, Q9UMS0, Q96KB5, Q8ND71, O75330, Q14677, Q9GZT3, Q00994, Q86SQ0, Q9NPI5, Q9H9E1, O43889, P481 65, Q9UQE7, O15264, P06493, Q15652, P00167, Q8N130, Q6PUV4, Q9H741, P78395, Q3MIW9, Q9UJZ1, Q12836, A6NGN9, A8MVZ5, Q99259, Q8N4E4, Q8WZ55, P18 848, O00534, P31689, P31371, Q8IYV9, Q13017, Q12899, Q15014, Q9H867, Q14674, P17980, Q9NS37, O43739, Q9H293, Q587J8, Q4VC05, Q9NZQ9, O75794, Q8 NDC4、Q8TE77、Q13127、O94986、Q9GZP4、Q8TDX7、A8MTB9、Q9Y6I3、Q5VT06、P 01100、P21781、P29353、P08700、P34910、Q6P996、Q9UN42、Q6PH85、Q9UI15、O 75409, Q969P6, Q9Y3C4, Q12849, A8MYV0, P49756, Q7Z6A9, O43422, Q86TS9, P07992, Q96SD1, Q13087, Q9H0R8, Q9ULR5, Q13972, Q9NQP4, Q06609, Q9UBU8 Q14554、P27701、Q03111、O95793、Q14055、Q9UNP9、Q00653、Q9UHJ6、P42681、 Q68DV7、O60603、P01112、Q03518、O15078、Q99487、Q9UHL9、O43903、P47928、O14879、Q9H2G2、P49137、B0FP48、P43166、Q13445、P11310、P49662、Q6N021 、Q8TCU4、P59282、Q8WWU5、Q9UK41、A6NLU5、O75665、O15164、O95777、O43247 、P04183、Q16181、O15294、Q9HCM3、Q96F10、Q2M296、O60237、P51815、O9569 6、P07766、Q8TEW0、P19526、P10398、P78358、Q8WUY3、P22528、O15211、P1524 8, O75293, Q6UY09, O00422, Q5VUJ9, Q9Y4G8, Q14204, P25092, Q32MZ4, Q7Z6I6, P50454, Q96NB3, Q9UMX5, Q9ULD2, P43357, P78317, Q9Y2J4, Q495A1, O147 17, P30279, P26639, Q9UJ99, Q8IV48, P33764, Q9NYJ8, P53420, Q6UXC1, Q15796, Q92973, Q6ZVL6, Q9HC77, O43768, Q8WWF5, P51948, P49366, Q92817, P78 540, Q6P995, Q9NQ84, P09914, O95997, Q7Z6P3, Q9BW66, O43663, P18754, P20702, Q99963, Q96AT9, P49454, Q6P5Z2, O75843, Q9UHY7, P21695, O15213, Q9 P000、Q14641、O75460、Q8IYS2、P52848、Q96IQ7、Q6B8I1、Q9GZZ8、Q9HC57、Q 16891、P56851、P63146、Q5VX71、Q9BZW2、Q92485、P49354、P05091、Q9UKM9、O 95433, Q13287, Q06643, Q6PKG0, Q9Y5E8, P35249, Q8NEB7, Q9BU02, O60749, O75351, Q6UW88, P20807, Q8N5J2, Q96RE7, Q86TE4, Q8WUD1, Q9NUW8, A1KZ92 Q49A26, Q6UW15, P01266, Q9GZX6, Q8N0Z9, Q13410, P31785, Q8IY22, Q9BW85, P36873, Q14160, P29536, Q765P7, Q7Z692, O95166, P19525, Q14BN4, A6NM11Q9Y3L3, O75528, Q9H8Y8, Q96T91, P61812, P21912, Q96AQ6, Q8TBM8, Q9HD43, P84022, P43632, P43627, Q86U X2, Q8NG06, Q9UHP3, P13051, Q9H2R5, Q9BRQ6, Q92783, P54652, Q6WCQ1, Q86SQ7, Q9H6E4, Q9ULC4, P51808, Q 15599, Q05193, P33316, O60245, Q16549, P10415, O43805, Q96K21, Q6UW56, O60232, P11387, Q96A25, Q9BQE 9. Q9Y5Q6, Q03169, P53384, Q7Z4H3, P14902, Q16819, Q9P013, O00203, Q8IXQ3, O75940, Q68D85, Q9UH65, Q14 011, P17181, Q676U5, Q96RF0, Q15172, Q86UU1, O43312, P20700, Q02750, Q99733, O43653, Q6NUJ1, P54577, Q2WEN9, P42081, Q8WXX5, Q8IV16, Q63HQ2, Q9UIM3, P21128, O75830, Q14246, Q96DE0, Q9Y5S2, O75071, Q8TF6 4. Q15262, Q14258, P13284, Q674X7, Q92890, Q6PL24, Q7Z569, Q8N6M0, P53990, O94988, Q17RW2, P49223, Q9 9447, Q96BQ1, Q9H910, P17568, Q9H2K0, Q9HC56, Q9NPJ3, Q6BCY4, P62330, Q8IWZ8, P48060, Q96R05, O15182. ,
[0199] sample In various embodiments, this disclosure relates to obtaining test samples, such as biological test samples, like bodily fluid samples, from a subject for analysis of multiple target molecules therein (e.g., multiple peptides and / or cfDNA molecules). Samples according to embodiments of the invention can be collected in any clinically acceptable manner. Any sample suspected of containing multiple target molecules can be used in conjunction with the methods of the invention. In some embodiments, the sample may comprise bodily fluids. In some embodiments, biological samples are collected from healthy subjects. In some embodiments, biological samples are collected from subjects known to have a specific disease or disorder (e.g., a specific cancer or tumor). In some embodiments, biological samples are collected from subjects suspected of having a specific disease or disorder.
[0200] As used herein, the terms “body fluid” and “biofluid” refer to liquid materials derived from a subject (e.g., human or non-human mammal). Non-limiting examples of body fluids commonly used in conjunction with the methods of the present invention include mucus, blood, plasma, serum, serum derivatives, synovial fluid, lymph, bile, sputum, saliva, sweat, tears, phlegm, amniotic fluid, menstrual fluid, vaginal secretions, semen, urine, cerebrospinal fluid (CSF) (such as lumbar or ventricular CSF), gastric juice, liquid samples containing one or more substances derived from nasal, pharyngeal, or oral swabs, liquid samples containing one or more substances derived from irrigation procedures (such as peritoneal, gastric, thoracic, or ductal irrigation procedures), etc.
[0201] In some embodiments, the biofluid includes fluid blood, plasma, serum, urine, saliva, pleural fluid, pericardial fluid, cerebrospinal fluid (CSF), peritoneal fluid, or any combination thereof. In some embodiments, the biofluid includes blood, blood fractions, plasma, or serum. In some embodiments, the biofluid is plasma.
[0202] In some embodiments, the sample may contain fine-needle aspirate or biopsy tissue. In some embodiments, the sample may contain a culture medium containing cells or biological material. In some embodiments, the sample may contain a blood clot, for example, a blood clot obtained from whole blood after serum removal. In some embodiments, the sample may contain feces. In one embodiment, the sample is drawn whole blood. In one aspect, only a portion of the whole blood sample, such as plasma, red blood cells, white blood cells, and platelets, is used. In some embodiments, the sample is separated into two or more components in conjunction with the method of the invention. For example, in some embodiments, the whole blood sample is separated into plasma, red blood cell, white blood cell, and platelet components.
[0203] In some embodiments, the sample contains multiple peptides and / or nucleic acids (e.g., cfDNA) that are derived not only from the subject from whom the sample was obtained, but also from one or more other organisms, such as viral DNA / RNA present in the subject at the time of sampling.
[0204] Nucleic acids and / or peptides can be extracted from a sample using any suitable method known in the art, and the extracted nucleic acids can be used in conjunction with the methods described herein. In some embodiments, peptides are purified from the sample. In some embodiments, free nucleic acids (e.g., cfDNA) are extracted from the sample. In some embodiments, peptides are detected (and optionally quantified) without a protein extraction step. For example, peptides can be detected by directly contacting a detection reagent with a biological sample (e.g., a biofluid sample, such as serum or plasma).
[0205] In embodiments, samples are “matched” or “paired” samples. Generally, the terms “matched sample” and “paired sample” refer to a pair of samples collected from the same subject, preferably at approximately the same time (e.g., as part of a single procedure or outpatient visit, or on the same day). In embodiments, a pair of samples comprises two biological fluid samples, which may be identical or different. In embodiments, a pair of samples comprises two aliquots separated from a single original sample (e.g., two aliquots of plasma from a blood sample). These terms may also be used to refer to peptides and / or polynucleotides derived from the matched sample, or their sequencing reads. In embodiments, multiple paired samples are analyzed. Multiple paired samples may come from the same individual collected at different times (e.g., a paired sample from early-stage cancer and a paired sample from late-stage cancer), from different individuals at the same or different times, or a combination of these. In embodiments, matched samples come from different subjects. In embodiments, multiple matched samples come from subjects with the same cancer type and optionally the same cancer stage.
[0206] Cancer types The methods according to embodiments of this disclosure can be used to detect the presence or absence of cancer. In some embodiments, cancer staging is stage I, stage II, stage III, or stage IV cancer. In some embodiments, cancer staging is stage 0 cancer (e.g., carcinoma in situ).
[0207] In some embodiments, these methods involve detecting the presence or absence of cancers selected from the following, determining their stage, monitoring their progression, and / or classifying them: breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, urothelial carcinoma of the renal pelvis, non-urothelial renal cancer, prostate cancer, anorectal cancer, anal cancer, colorectal cancer, hepatobiliary cancer caused by hepatocellular carcinoma, hepatobiliary cancer caused by cells other than hepatocellular carcinoma, liver / cholecystitis, esophageal cancer, pancreatic cancer, gastric cancer, squamous cell carcinoma of the upper gastrointestinal tract, non-squamous upper gastrointestinal cancer, head and neck cancer, lung cancer, lung adenocarcinoma, small cell lung cancer, squamous cell lung cancer, and cancers other than adenocarcinoma or small cell lung cancer, neuroendocrine carcinoma, melanoma, thyroid cancer, sarcoma, plasma cell tumor, multiple myeloma, myeloid tumors, lymphoma, and leukemia. In some embodiments, the cancer is selected from one or more of the following: anorectal cancer, bladder cancer, colorectal cancer, esophageal cancer, head and neck cancer, liver cancer, bile duct cancer, lung cancer, ovarian cancer, pancreatic cancer, gastric cancer, breast cancer, prostate cancer, kidney cancer, cervical cancer, endometrial cancer, and blood cancer.
[0208] In some embodiments, the same assay is applied to detect any one of multiple cancer conditions (e.g., cancer types and / or cancer stages disclosed herein). For example, the assay according to embodiments can be used to detect the presence (and optionally stage) of breast cancer in a sample from a first subject, and repeated to detect the presence (and optionally stage) of lung cancer in a sample from a second subject, based on evaluating biomarkers for each condition in both samples. In embodiments, the same assay is repeated in multiple samples to identify the presence of at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 75, 100 or more cancer conditions. In embodiments, the same assay is repeated in multiple samples to identify the presence of at least 10 cancer conditions. In embodiments, the same assay is repeated in multiple samples to identify the presence of at least 20 cancer conditions. In embodiments, the same assay is repeated in multiple samples to identify the presence of at least 30 cancer conditions. In embodiments, the same assay is repeated in multiple samples to identify the presence of at least 50 cancer conditions.
[0209] Detection and treatment In yet another embodiment, information obtained from any of the methods described herein (e.g., an overall probability score) can be used to make or influence clinical decisions (e.g., cancer diagnosis, treatment selection, assessment of treatment effectiveness, etc.). For example, in one embodiment, if the overall probability score exceeds a threshold, a physician may prescribe appropriate treatment (e.g., surgical resection, radiation therapy, chemotherapy, and / or immunotherapy). In some embodiments, information such as likelihood or probability scores can be provided as readings to the physician or subject.
[0210] In one aspect, the method includes selecting a subject who has a cancer type or is at increased risk of developing that cancer type, and administering to the subject a treatment effective for that cancer type, wherein the selection includes (a) measuring the level of a first target molecule from a first sample from the subject, wherein these first target molecules comprise cell-free DNA (cfDNA) from multiple different target genomic regions that are differentially methylated in at least one of the multiple cancer types; (b) measuring the level of a second target molecule from a second sample from the subject, wherein these second target molecules comprise multiple different peptides differentially expressed in at least one of the multiple cancer types; (c) applying a trained classifier to the measured levels of these first and second target molecules to assign an overall probability score to the cancer; wherein applying the trained classifier includes: (i) applying a first trained model to the measured levels of these first target molecules to assign a first probability score to the cancer; (ii) applying a second trained model to the measured levels of these second target molecules to assign a second probability score to the cancer; and (iii) summing the first probability score and the second probability score; and (d) The cancer is detected by identifying an overall probability score that is above a threshold for the presence of the cancer; and the treatment includes surgical resection, radiation therapy, chemotherapy, and / or immunotherapy.
[0211] In some embodiments of any of the various methods described herein, the measurement and analysis of one of the first target molecule and the second target molecule are performed in parallel or sequentially in any order. In some embodiments, the analysis is performed sequentially. In some embodiments, the first target molecule is measured and a first trained model is applied to it before the level of the second target molecule is measured. In some embodiments, the second target molecule is measured and a second trained model is applied to it before the level of the first target molecule is measured. In some embodiments in which the measurement and analysis are performed sequentially, the second measurement and analysis step, and the step of summarizing the first probability score and the second probability score, are performed only if the confidence level of the result of the first measurement and analysis step is below a specified threshold. For example, the result of analyzing a first type of analyte (e.g., multiple different peptides) may produce a sufficiently low cancer presence probability score (to have high confidence in the absence of cancer) or a sufficiently high cancer presence probability score (to have high confidence in the presence of cancer) so that the presence or absence of cancer does not need to be detected by means of analysis of a second type of analyte (e.g., cfDNA). In this case, the sample with a probability score of a single type of analyte between the threshold is analyzed with the second type of analyte (e.g., cfDNA) to confidently determine the presence or absence of cancer. In some embodiments, both the first target molecule and the second target molecule are measured and a trained model is applied to them sequentially or in parallel, regardless of the outcome of either one.
[0212] In some embodiments, the analysis is performed sequentially, starting with the analysis of one or more (e.g., multiple) peptides, and cfDNA analysis is performed only if the probability score of cancer presence based on the peptide analysis is above a threshold. For example, when the probability score of cancer presence based on the peptide analysis is below a threshold, the sample can be scored as non-cancer; while if the probability score is equal to or above a threshold, cfDNA analysis is performed to determine whether the probability score based on cfDNA (and / or the overall score of both protein- and cfDNA-based analyses) is above another threshold. In this way, the analysis of one or more peptides can be used as a screening tool to reduce the number of samples undergoing cfDNA analysis. In some embodiments, a trained model as described herein is used to determine the probability scores based on both peptide and cfDNA analyses. In some embodiments, the peptide-based model has lower specificity for cancer detection than the cfDNA (or overall)-based model. Although lower specificity in peptide-only assays can lead to a higher number of false positives, pairing such a lower-specificity peptide assay with a higher-specificity cfDNA assay allows the second assay to reduce false positives while also reducing the number of samples undergoing potentially more expensive and resource-intensive sample processing and analysis steps. In some embodiments, peptide analysis uses a first trained model with a defined cancer detection specificity of 0.500 or higher (e.g., at least 0.600, 0.700, 0.800, 0.900, 0.950 or higher). In some embodiments, cfDNA analysis uses a second trained model with a defined cancer detection specificity of 0.900 or higher (e.g., at least 0.950, 0.975, 0.980, 0.985, 0.990, 0.995 or higher). In some embodiments, the second trained model aggregates the results of both protein and cfDNA analysis to assign a probability score for the presence of cancer, which is detected when the probability score is above a threshold.
[0213] Models and classifiers (as described herein) can be used to determine a probability score (e.g., an overall probability score or a second probability score) for a sample from a subject with cancer. In one embodiment, when the probability score exceeds a threshold, an appropriate treatment (e.g., surgical resection or therapeutic intervention) is prescribed. For example, in one embodiment, one or more appropriate treatments are prescribed if the probability score is greater than or equal to 60. In another embodiment, one or more appropriate treatments are prescribed if the overall probability score is greater than or equal to 65, greater than or equal to 70, greater than or equal to 75, greater than or equal to 80, greater than or equal to 85, greater than or equal to 90, or greater than or equal to 95. In other embodiments, the cancer log-odds ratio can indicate the effectiveness of cancer treatment. For example, an increase in the cancer log-odds ratio over time (e.g., at a second time post-treatment) can indicate treatment ineffectiveness. Similarly, a decrease in the cancer log-odds ratio over time (e.g., at a second time post-treatment) can indicate treatment success. In another embodiment, one or more appropriate treatments are prescribed if the cancer log-odds ratio is greater than 1, greater than 1.5, greater than 2, greater than 2.5, greater than 3, greater than 3.5, or greater than 4.
[0214] In some embodiments, the treatment is one or more cancer therapeutic agents selected from the group consisting of chemotherapeutic agents, targeted cancer therapeutic agents, differentiation therapeutic agents, hormone therapeutic agents, and immunotherapeutic agents. For example, the treatment may be one or more chemotherapeutic agents selected from the group consisting of alkylating agents, antimetabolites, anthracyclines, antitumor antibiotics, scaffold disruptors (taxanes), topoisomerase inhibitors, mitotic inhibitors, corticosteroids, kinase inhibitors, nucleotide analogs, platinum-based agents, and any combination thereof. In some embodiments, the treatment is one or more targeted cancer therapeutic agents selected from the group consisting of signal transduction inhibitors (e.g., tyrosine kinase and growth factor receptor inhibitors), histone deacetylase (HDAC) inhibitors, retinoic acid receptor agonists, proteasome inhibitors, angiogenesis inhibitors, and monoclonal antibody conjugates. In some embodiments, the treatment is one or more differentiation therapeutic agents, including retinoids such as retinoic acid, levamisole, and bexarotin. In some embodiments, the treatment is one or more hormonal therapeutic agents selected from the group consisting of: anti-estrogens, aromatase inhibitors, progestins, estrogens, anti-androgens, and GnRH agonists or analogues. In one embodiment, the treatment is one or more immunotherapeutic agents selected from the group consisting of: monoclonal antibody therapies (such as rituximab (Rituxan) and alemtuzumab (CAMPATH)), nonspecific immunotherapies and adjuvants (such as BCG, interleukin-2 (IL-2), and interferon-α), and immunomodulatory drugs (e.g., thalidomide and lenalidomide (Revlimid)). Experienced physicians or oncologists can select appropriate cancer therapeutic agents based on a variety of characteristics, such as tumor type, cancer stage, previous cancer treatments or agents, and other characteristics of the cancer.
[0215] Cancer and treatment monitoring In some embodiments, the first time point is before cancer treatment (e.g., before resection or treatment intervention), and the second time point is after cancer treatment (e.g., after resection or treatment intervention), and the method is used to monitor the effectiveness of the treatment. For example, if the overall probability score at the second time point decreases compared to the overall probability score at the first time point, the treatment can be considered effective. However, if the second overall probability score increases compared to the first overall probability score, the treatment can be considered ineffective. In other embodiments, both the first and second time points are before cancer treatment (e.g., before resection or treatment intervention). In still other embodiments, both the first and second time points are after cancer treatment (e.g., before resection or treatment intervention), and the method is used to monitor the effectiveness of the treatment or the loss of treatment effectiveness. In still other embodiments, cfDNA and peptide samples may be obtained from and analyzed from the cancer patient at the first and second time points, for example, to monitor cancer progression, determine whether the cancer is in remission (e.g., post-treatment), to monitor or detect residual disease or disease recurrence, or to monitor the effect of treatment (e.g., therapeutic).
[0216] Test samples can be obtained from cancer patients at any desired time point set and analyzed according to the method of the invention to monitor the patient's cancer status. In some embodiments, the duration of the interval between the first and second time points ranges from about 15 minutes to about 30 years, such as about 30 minutes, such as about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23 or about 24 hours, such as about 1, 2, 3, 4, 5, 10, 15, 20, 25 or about 30 days, or such as about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12 months, or such as about 1, 1.5, 2, 2.5, 3, 3. 5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, 10.5, 11, 11.5, 12, 12.5, 13, 13.5, 14, 14.5, 15, 15.5, 16, 16.5, 17, 17.5, 18, 18.5, 19, 19.5, 20, 20.5, 21, 21.5, 22, 22.5, 23, 23.5, 24, 24.5, 25, 25.5, 26, 26.5, 27, 27.5, 28, 28.5, 29, 29.5 or approximately 30 years. In other embodiments, the frequency of obtaining test samples from patients may be: at least once every 3 months, at least once every 6 months, at least once a year, at least once every 2 years, at least once every 3 years, at least once every 4 years, or at least once every 5 years.
[0217] Computer systems and devices In one aspect, this disclosure provides a computer system for implementing one or more steps of the methods disclosed herein. In another aspect, this disclosure provides a non-transitory computer-readable medium having computer-readable instructions stored thereon for implementing one or more steps of the methods disclosed herein.
[0218] The methods disclosed herein can be implemented using software, hardware, firmware, hardwiring, or any combination thereof. Features that implement the functionality can also be physically located in various locations, including distributed systems, such that some functions are implemented in different physical locations (e.g., the imaging device is in one room and the host workstation is in another room, or in different buildings, for example, via wireless or wired connections).
[0219] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, receiving or transferring data to, or both: one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks. Information carriers suitable for embodying computer program instructions and data include all forms of non-volatile memory, for example, semiconductor memory devices (e.g., EPROM, EEPROM, solid-state drives (SSDs), and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and optical disks (e.g., CDs and DVDs). Processors and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[0220] To provide interaction with a user, the subjects described herein can be implemented on a computer with I / O devices, such as CRT, LCD, LED, or projection devices for displaying information to the user, and input or output devices, such as keyboards and pointing devices (e.g., mice or trackballs), through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user. For example, feedback provided to the user can be any form of sensory feedback (e.g., visual, auditory, or tactile feedback), and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0221] The subject matter described herein can be implemented in a computing system that includes backend components (e.g., a data server), middleware components (e.g., an application server), or frontend components (e.g., a client computer with a graphical user interface or web browser through which a user can interact with an implementation of the subject matter described herein), or any combination of such backend, middleware, and frontend components. The components of the system can be interconnected via digital data communication networks (e.g., communication networks) of any form or medium. For example, a reference dataset can be stored at a remote location, and the computer can communicate across the network to access the reference dataset for comparison purposes. However, in other embodiments, the reference dataset can be stored locally on the computer, and the computer accesses the reference dataset within its CPU for comparison purposes. Examples of communication networks include, but are not limited to, cellular networks (e.g., 3G or 4G), local area networks (LANs), and wide area networks (WANs), such as the Internet.
[0222] The subject matter described herein can be implemented as one or more computer program products, such as one or more computer programs tangibly embodied in an information carrier (e.g., in a non-transitory computer-readable medium) for performing or controlling operations by a data processing device (e.g., a programmable processor, computer, or multiple computers). Computer programs (also referred to as programs, software, software applications, application programs, macros, or code) can be written in any form of programming language, including compiled or interpreted languages (e.g., C, C++, Perl), and can be deployed in any form, including as standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. The systems and methods disclosed herein may include instructions written in any suitable programming language known in the art, including but not limited to C, C++, Perl, Java, ActiveX, HTML5, Visual Basic, or JavaScript.
[0223] A computer program does not necessarily correspond to a file. A program can be stored in a file or a portion of a file that holds other programs or data, in a single file dedicated to that program, or in multiple coordinated files (e.g., a file that stores one or more modules, subroutines, or portions of code). A computer program can be deployed and executed on a single computer or multiple computers at a single site, or distributed across multiple sites and interconnected via a communications network.
[0224] Files can be digital files, for example, stored on hard drives, SSDs, CDs, or other tangible, non-transitory media. Files can be sent from one device to another over a network (e.g., as data packets from a server to a client, such as via a network interface card, modem, wireless network card, or similar).
[0225] Writing, as disclosed herein, involves transforming tangible, non-transitory computer-readable media, for example, by adding, removing, or rearranging particles (e.g., converting particles with net charge or dipole moment into magnetization modes via a read / write head), which subsequently represent a new combination of information about objective physical phenomena that is desired and useful to the user. In some embodiments, writing involves physically transforming the material in a tangible, non-transitory computer-readable medium (e.g., having certain optical properties so that an optical read / write device can read new and useful combinations of information, such as burning a CD-ROM). In some embodiments, writing includes transforming a physical flash memory device (such as a NAND flash memory device) and storing information by transforming the physical elements in a memory cell array made of floating-gate transistors. Methods of writing are well known in the art and can be invoked manually or automatically, for example, by a program or by a save command in software or a write command in a programming language.
[0226] Suitable computing devices typically include mass storage, at least one graphical user interface, at least one display device, and typically include communication between devices. Mass storage is a type of computer-readable medium, i.e., computer storage medium. Computer storage media can include volatile, non-volatile, removable, and non-removable media, implemented in any method or technology, for storing information such as computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage or other magnetic storage devices, radio frequency identification (RFID) tags or chips, or any other medium that can be used to store desired information and is accessible by a computing device.
[0227] The functionality described herein can be implemented using software, hardware, firmware, hardwiring, or any combination thereof. Any of the software components can be physically located in various places, including distributed systems, allowing different parts of the functionality to be implemented in different physical locations.
[0228] As will be recognized by those skilled in the art, the computer system used to implement some or all of the described inventive methods, which is necessary or most suitable for the execution of the methods disclosed herein, may include one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), or both), main memory, and static memory, which communicate with each other via a bus.
[0229] A processor will typically consist of chips, such as single-core or multi-core chips, to provide a central processing unit (CPU). Processors can be supplied by chips from Intel or AMD.
[0230] The memory may include one or more machine-readable means on which one or more instruction sets (e.g., software) are stored, which, when executed by one or more processors of any of the disclosed computers, can implement some or all of the methods or functions described herein. The software may also reside wholly or at least partially in main memory and / or processor during execution by the computer system. Preferably, each computer includes non-transitory memory, such as solid-state drives, flash drives, disk drives, hard disk drives, etc.
[0231] While in exemplary embodiments, a machine-readable device may be a single medium, the term "machine-readable device" should be understood to include a single or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more instructions and / or datasets. These terms should also be understood to include any and more media capable of storing, encoding, or preserving a set of instructions executable by a machine, and enabling the machine to perform any or more of the methods disclosed herein. Therefore, these terms should be understood to include, but are not limited to, one or more solid-state memories (e.g., a Subscriber Identity Module (SIM) card, a Secure Digital Card (SD card), a MicroSD card, or a Solid State Drive (SSD)), optical and magnetic media, and / or any other tangible storage media.
[0232] The computer disclosed herein will generally include one or more I / O devices, such as, for example, one or more video display units (e.g., liquid crystal display (LCD) or cathode ray tube (CRT)), alphanumeric input devices (e.g., keyboard), cursor control devices (e.g., mouse), disk drive units, signal generation devices (e.g., speakers), touch screens, accelerometers, microphones, cellular radio frequency antennas, and network interface devices (which may be, for example, network interface cards (NICs), Wi-Fi cards, or cellular modems).
[0233] Any part of the software can be physically located in various locations, including distributed systems, allowing different functionalities to be implemented in different physical locations.
[0234] Additionally, the system disclosed herein may be provided to include reference data. Any suitable genomic data may be stored and used in the system. Examples include, but are not limited to: comprehensive, multidimensional maps of key genomic changes in major types and subtypes of cancer from The Cancer Genome Atlas (TCGA); catalogs of genomic anomalies from the International Cancer Genome Consortium (ICGC); catalogs of somatic mutations in cancer from COSMIC; the latest versions of the human genome and other commonly used model organisms; the latest reference SNPs from dbSNP; gold standard insertions and deletions from the 1000 Genomes Project and the Broad Institute; exome capture kit annotations from Enomina, Agilent Technologies, Nimblegen, and Ion Torrent; transcriptome annotations; and small test data for pipeline experiments (e.g., for new users).
[0235] In some embodiments, data is available in the context of databases included in the system. Any suitable database structure can be used, including relational databases, object-oriented databases, and others. In some embodiments, reference data is stored in a relational database (such as a “non-SQL only” (NoSQL) database). In various embodiments, the system disclosed herein includes a graph database. It should also be understood that the term “database” as used herein is not limited to a single database; rather, the system may include multiple databases. For example, according to embodiments of this disclosure, databases may include two, three, four, five, six, seven, eight, nine, ten, fifteen, twenty, or more individual databases, including any integer number of such databases. For example, one database may contain public reference data, a second database may contain test data from patients, a third database may contain data from healthy subjects, and a fourth database may contain data from subjects with known conditions or disorders. It should be understood that any other database configurations with respect to the data contained therein are also considered through the methods described herein.
[0236] Illustrative Examples This disclosure provides the following illustrative examples.
[0237] Example 1: A method for detecting cancer in a subject, the method comprising: (a) Measure the level of a first target molecule from a first sample from the subject, wherein the first target molecule contains cell-free DNA (cfDNA) from multiple different target genomic regions that are differentially methylated in at least one of multiple cancer types; (b) Measure the level of a second target molecule from a second sample from the subject, wherein the second target molecule comprises multiple different peptides differentially expressed in at least one of the multiple cancer types; (c) Applying a trained classifier to the measurement levels of the first and second target molecules to assign an overall probability score to the cancer; wherein applying the trained classifier includes: (i) applying a first trained model to the measurement levels of the first target molecules to assign a first probability score to the cancer; (ii) applying a second trained model to the measurement levels of the second target molecules to assign a second probability score to the cancer; and (iii) summing the first and second probability scores; and (d) Detect cancer by identifying an overall probability score that is above the threshold for the presence of the cancer.
[0238] Example 2: The method as described in Example 1, wherein for reference samples from (1) reference subjects with known cancer and (2) reference subjects without cancer, the trained classifier is trained using a reference first probability score from the first trained model, a reference second probability score from the second trained model, and a reference overall probability score that summarizes the reference first probability scores and the reference second probability scores.
[0239] Example 3: The method as described in Example 1 or 2, wherein the trained classifier assigns an overall probability score to each of a plurality of different cancer types, and detecting the cancer includes identifying the cancer type as the cancer type with the highest overall probability score.
[0240] Example 4: The method as described in any one of Examples 1-3, wherein summarizing the first cancer probability score and the second cancer probability score includes combining the first probability score and the second probability score of the cancer in a linear model.
[0241] Example 5: The method as described in any one of Examples 1-4, wherein the first sample and the second sample are identical.
[0242] Example 6: The method as described in any one of Examples 1-5, wherein the plurality of different target genomic regions include at least 1,000, 5,000, 10,000, 20,000, or 30,000 target genomic regions.
[0243] Example 7: The method as described in any one of Examples 1-6, wherein the plurality of target genomic regions have a total collective length of at least 50 kb, 100 kb, 500 kb or 1000 kb.
[0244] Example 8: The method as described in any one of Examples 1-7, wherein each of the plurality of different target genomic regions contains at least five methylation sites.
[0245] Example 9: The method as described in any one of Examples 1-8, wherein measuring these first target molecules comprises sequencing transformed cfDNA or its amplified products from the plurality of different target genomic regions, wherein the transformed cfDNA comprises cfDNA treated with a deamination agent.
[0246] Example 10: The method as described in Example 9, further comprising treating the cfDNA with the deamination agent, optionally wherein the deamination agent is cytosine deaminase or bisulfite.
[0247] Example 11: The method as described in Example 9 or 10, wherein the sequencing produces at least 100,000 sequencing reads.
[0248] Example 12: The method as described in any one of Examples 8-11, wherein measuring these first target molecules includes enriching the transformed cfDNA or its amplification product to produce an enriched polynucleotide sample.
[0249] Example 13: The method as described in Example 12, wherein the enrichment includes capturing the transformed cfDNA or its amplification product with a plurality of corresponding bait oligonucleotides.
[0250] Example 14: The method as described in Example 13, wherein the plurality of different target genomic regions for enrichment by these bait oligonucleotides are genomic regions identified by the first trained model as differentially methylated in at least one of a plurality of cancer types relative to non-cancer tissue or relative to different types of cancer.
[0251] Example 15: The method as described in any one of Examples 1-14, wherein the plurality of different polypeptides comprises at least 5, 10, 25, 50, 100, 200, 500, 1000, 2000, 3000, 5000 or 7500 different polypeptides.
[0252] Example 16: The method as described in Example 15, wherein the plurality of different polypeptides includes: (a) a polypeptide for identifying proteins selected from List 1; (b) a polypeptide for identifying proteins selected from any of Lists 2-19; (c) a polypeptide for identifying proteins selected from List 20; or (d) a polypeptide for identifying one or more of CHAD, KRT19, MMP12, PTN, SERPINA3 and SPP1.
[0253] Example 17: The method as described in any one of Examples 1-16, wherein the trained classifier distinguishes subjects with cancer from subjects without cancer with specificity for each of the plurality of cancer types.
[0254] Example 18: The method as described in any one of Examples 1-17, wherein the trained classifier has a higher cancer detection sensitivity than each of the first trained model and the second trained model; optionally, wherein the trained classifier has a cancer detection specificity equal to or greater than each of the first trained model and the second trained model.
[0255] Example 19: The method as described in any one of Examples 1-18, wherein the trained classifier is a binary classifier, a hybrid model classifier, a multilayer perceptron model classifier, or a logistic regression classifier.
[0256] Example 20: The method as described in any one of Examples 1-19, wherein the first trained model and / or the second trained model is a binary classifier, a hybrid model classifier, a multilayer perceptron model classifier, or a logistic regression classifier.
[0257] Example 21: The method as described in any one of Examples 1-19, wherein the first trained model binarizes the measurement levels of the first target molecules by assigning a first value when the target genomic region is detected and a second value when the target genomic region is not detected.
[0258] Example 22: The method as described in any one of Examples 1-19, wherein the second trained model performs a logarithmic transformation on the measurement levels of these second target molecules normalized to control proteins present in known amounts.
[0259] Example 23: The method as described in any one of Examples 1-22, wherein (a) the first trained model is trained using measurement levels of the first target molecules for a first reference sample, (b) the second trained model is trained using measurement levels of the second target molecules for a second reference sample, and (c) the first and second reference samples include samples from reference subjects with known cancer and reference subjects without cancer.
[0260] Example 24: The method as described in any one of Examples 1-23, wherein the first sample and / or the second sample comprises a biological fluid; optionally, wherein the biological fluid comprises blood, plasma, serum, urine, saliva, pleural fluid, pericardial fluid, cerebrospinal fluid (CSF), peritoneal fluid, or any combination thereof.
[0261] Example 25: The method as described in Example 24, wherein the biofluid comprises blood, blood fractions, plasma, or serum.
[0262] Example 26: The method as described in Example 25, wherein the first sample and / or the second sample is a plasma sample.
[0263] Example 27: The method as described in any one of Examples 1-26, wherein the plurality of cancer types includes at least 10 cancer types.
[0264] Example 28: The method as described in any one of Examples 1-27, wherein the plurality of cancer types includes one or more of anorectal cancer, bladder cancer, colorectal cancer, esophageal cancer, head and neck cancer, liver cancer, bile duct cancer, lung cancer, ovarian cancer, pancreatic cancer, gastric cancer, breast cancer, prostate cancer, kidney cancer, cervical cancer, endometrial cancer, and hematologic malignancies.
[0265] Example 29: The method as described in any one of Examples 1-28, further comprising treating the subject with the type of cancer.
[0266] Example 30: The method as described in Example 29, wherein the treatment includes surgical resection, radiotherapy, chemotherapy, and / or immunotherapy.
[0267] Example 31: A method for treating a subject's cancer, the method comprising selecting a subject based on the results of a detection assay and treating the subject's cancer, wherein: (a) The detection method comprises the method described in any one of Examples 1-28; and (b) The treatment includes surgical resection, radiation therapy, chemotherapy and / or immunotherapy.
[0268] Example 32: A method for training a classifier to detect target molecules from cancer, the method comprising: (a) Receive a first measurement level of a first target molecule in a first sample for reference subjects, wherein (i) the first target molecules contain cell-free DNA (cfDNA) from multiple different target genomic regions that are differentially methylated in at least one of multiple cancer types, and (ii) the reference subjects include a first subject with a known cancer type and a second subject without cancer; (b) A first model is trained by applying a first machine learning algorithm to these first measurement levels to generate a first probability score of the presence of cancer in the subject; (c) Receive a second measurement level of a second target molecule in a second sample for these reference subjects, wherein the second target molecule comprises a plurality of different polypeptides differentially expressed in at least one of the plurality of cancer types; (d) A second model is trained to generate a second probability score of the presence of cancer in the subject by applying a second machine learning algorithm to these second measurement levels; (e) Use the trained first model to generate a reference first cancer probability score for these first samples; (f) Use the trained second model to generate reference second cancer probability scores for these second samples; (g) By aggregating the reference first cancer probability score and the reference second cancer probability score for each corresponding reference subject, a reference overall cancer probability score is generated for multiple reference subjects; and (h) A classifier is trained to generate the subject’s overall cancer probability score by applying a third machine learning algorithm to these reference first cancer probability scores, these reference second cancer probability scores and these reference overall cancer probability scores.
[0269] Example 33: The method as described in Example 32, wherein summarizing the first cancer probability score and the second cancer probability score includes combining the first probability score and the second probability score of the cancer in a linear model.
[0270] Example 34: The method as described in Example 32 or 33, wherein the first machine learning algorithm, the second machine learning algorithm, and / or the third machine learning algorithm is L1 regularized logistic regression, L2 regularized logistic regression, generalized linear model (GLM), random forest, multinomial logistic regression, multilayer perceptron, support vector machine, or neural network.
[0271] Example 35: The method as described in any one of Examples 32-34, wherein the first trained model binarizes the measurement levels of the first target molecules by assigning a first value when the target genomic region is detected and a second value when the target genomic region is not detected.
[0272] Example 36: The method as described in any one of Examples 32-35, wherein the second trained model performs a logarithmic transformation on the measurement levels of these second target molecules normalized to control proteins present in known amounts.
[0273] Example 37: The method as described in any one of Examples 32-36, wherein (a) the first trained model is trained using measurement levels of the first target molecules for a first reference sample, (b) the second trained model is trained using measurement levels of the second target molecules for a second reference sample, and (c) the first and second reference samples include samples from reference subjects with known cancer and reference subjects without cancer.
[0274] Example 38: The method as described in any one of Examples 32-37, wherein the third machine learning algorithm is logistic regression.
[0275] Example 39: The method as described in any one of Examples 32-38, wherein the plurality of different target genomic regions include at least 1,000, 5,000, 10,000, 20,000, or 30,000 target genomic regions.
[0276] Example 40: The method as described in any one of Examples 32-39, wherein the plurality of target genomic regions have a total collective length of at least 50 kb, 100 kb, 500 kb, or 1000 kb.
[0277] Example 41: The method as described in any one of Examples 32-40, wherein each of the plurality of different target genomic regions contains at least five methylation sites.
[0278] Example 42: The method as described in any one of Examples 32-41, wherein these first measurement levels include sequencing results of the cfDNA or its amplicon.
[0279] Example 43: The method as described in Example 42, wherein for each of these first samples, the sequencing results comprise at least 100,000 reads.
[0280] Example 44: The method of any one of Examples 32-43, wherein the plurality of different polypeptides comprises at least 5, 10, 25, 50, 100, 200, 500, 1000, 2000, 3000, 5000 or 7500 different polypeptides.
[0281] Example 45: The method as described in Example 44, wherein the plurality of different polypeptides includes: (a) a polypeptide for identifying proteins selected from List 1; (b) a polypeptide for identifying proteins selected from any of Lists 2-19; (c) a polypeptide for identifying proteins selected from List 20; or (d) a polypeptide for identifying one or more of CHAD, KRT19, MMP12, PTN, SERPINA3 and SPP1.
[0282] Example 46: The method as described in any one of Examples 32-45, wherein the plurality of cancer types includes at least 10 cancer types.
[0283] Example 47: The method as described in any one of Examples 32-46, wherein the plurality of cancer types includes one or more of anorectal cancer, bladder cancer, colorectal cancer, esophageal cancer, head and neck cancer, liver cancer, bile duct cancer, lung cancer, ovarian cancer, pancreatic cancer, gastric cancer, breast cancer, prostate cancer, kidney cancer, cervical cancer, endometrial cancer, and hematologic malignancies.
[0284] Example 48: A method for screening a subject for cancer, the method comprising: (a) Measure the level of a first target molecule from a first sample from the subject, wherein the first target molecule comprises multiple different polypeptides differentially expressed in at least one of multiple cancer types; (b) Applying a first trained model to the measurement level of these first target molecules to assign a first probability score to each of the plurality of cancer types; wherein (i) the first trained model has a first specificity for cancer detection, and (ii) the first probability score of at least one of these cancer types is above a first threshold for the presence of cancer. (c) Measure the level of a second target molecule from a second sample from the subject, wherein the second target molecule contains cell-free DNA (cfDNA) from multiple different target genomic regions that are differentially methylated in at least one of the multiple cancer types; (d) Applying a second trained model to the measurement levels of these second target molecules to assign a second probability score to the cancer; wherein (ii) the second trained model has a second specificity for cancer detection, and (ii) the second specificity is higher than the first specificity; and (e) Detect the cancer by identifying that the second probability score is above the threshold for the presence of the cancer.
[0285] Example 49: The method as described in Example 48, wherein (a) the first trained model is trained using measurement levels of the first target molecules for the first reference sample, (b) the second trained model is trained using measurement levels of the second target molecules for the second reference sample, and (c) the first and second reference samples include samples from reference subjects with known cancer and reference subjects without cancer.
[0286] Example 50: The method as described in Example 48 or 49, wherein the first sample and the second sample are identical.
[0287] Example 51: The method as described in any one of Examples 48-50, wherein the plurality of different target genomic regions include at least 1,000, 5,000, 10,000, 20,000, or 30,000 target genomic regions.
[0288] Example 52: The method as described in any one of Examples 48-51, wherein the plurality of target genomic regions have a total collective length of at least 50 kb, 100 kb, 500 kb or 1000 kb.
[0289] Example 53: The method as described in any one of Examples 48-52, wherein each of the plurality of different target genomic regions contains at least five methylation sites.
[0290] Example 54: The method as described in any one of Examples 48-53, wherein measuring these second target molecules comprises sequencing transformed cfDNA or its amplified products from the plurality of different target genomic regions, wherein the transformed cfDNA comprises cfDNA treated with a deamination agent.
[0291] Example 55: The method described in Example 54, wherein the sequencing produces at least 100,000 sequencing reads.
[0292] Example 56: The method as described in any one of Examples 54-55, wherein measuring these first target molecules includes enriching the transformed cfDNA or its amplification product to produce an enriched polynucleotide sample.
[0293] Example 57: The method as described in Example 56, wherein the enrichment includes capturing the transformed cfDNA or its amplification product with a plurality of corresponding bait oligonucleotides.
[0294] Example 58: The method as described in Example 57, wherein the plurality of different target genomic regions for enrichment by these bait oligonucleotides are identified by the second trained model as genomic regions that are differentially methylated in at least one of a plurality of cancer types relative to non-cancer tissue or relative to different types of cancer.
[0295] Example 59: The method as described in any one of Examples 48-58, wherein the plurality of different polypeptides comprises at least 5, 10, 25, 50, 100, 200, 500, 1000, 2000, 3000, 5000 or 7500 different polypeptides.
[0296] Example 60: The method as described in Example 59, wherein the plurality of different polypeptides includes: (a) a polypeptide for identifying proteins selected from List 1; (b) a polypeptide for identifying proteins selected from any of Lists 2-19; (c) a polypeptide for identifying proteins selected from List 20; or (d) a polypeptide for identifying one or more of CHAD, KRT19, MMP12, PTN, SERPINA3, and SPP1.
[0297] Example 61: The method as described in any one of Examples 48-60, wherein the first trained model and / or the second trained model is a neural network classifier, a binary classifier, a hybrid model classifier, a multilayer perceptron model classifier, or a logistic regression classifier.
[0298] Example 62: The method as described in any one of Examples 48-61, wherein the second trained model binarizes the measurement levels of the second target molecules by assigning a first value when the target genomic region is detected and a second value when the target genomic region is not detected.
[0299] Example 63: The method as described in any one of Examples 48-62, wherein the first trained model performs a logarithmic transformation on the measurement levels of these first target molecules normalized to a known amount of a control peptide.
[0300] Example 64: The method of any one of Examples 48-63, wherein (a) the first sample and / or the second sample comprises a biological fluid; optionally wherein the biological fluid comprises blood, plasma, serum, urine, saliva, pleural fluid, pericardial fluid, cerebrospinal fluid (CSF), peritoneal fluid or any combination thereof; and (b) the first sample and / or the second sample is a plasma sample.
[0301] Example 65: The method as described in any one of Examples 48-64, wherein (a) the plurality of cancer types includes at least 10 cancer types; and / or (b) the plurality of cancer types includes one or more of anorectal cancer, bladder cancer, colorectal cancer, esophageal cancer, head and neck cancer, liver cancer, bile duct cancer, lung cancer, ovarian cancer, pancreatic cancer, gastric cancer, breast cancer, prostate cancer, kidney cancer, cervical cancer, endometrial cancer, and hematologic malignancies.
[0302] Example 66: The method as described in any one of Examples 48-65, further comprising treating the subject with the type of cancer; optionally wherein the treatment comprises surgical resection, radiotherapy, chemotherapy and / or immunotherapy.
[0303] Example 67: A method for treating a subject's cancer, the method comprising selecting the subject based on the results of a screening assay and treating the subject's cancer, wherein: (a) The screening assay comprises the method as described in any one of Examples 48-65; and (b) The treatment includes surgical resection, radiation therapy, chemotherapy and / or immunotherapy.
[0304] Example 68: A non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, perform one or more steps of the method as described in any one of Examples 1-28 or 48-65.
[0305] Example 69: A non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, perform the method as described in any one of Examples 32-47. Example The following examples are provided to demonstrate to those skilled in the art how to make and use this specification in its entirety, and are not intended to limit the scope of the specification as the inventors see it, nor to represent that the following experiments are all or only those experiments performed. Efforts have been made to ensure the accuracy of the figures used (e.g., quantities, temperatures, etc.), but some experimental errors and biases should be taken into account. Example 1—Analysis of Combination Analytes for Cancer Detection To test the performance of combined cfDNA methylation and protein analysis in cancer detection, 1,538 cancer and 1,485 non-cancer circulating cell-free genome (CCGA) samples were analyzed. Cancer samples were selected based on their potential for improved mortality benefits, and these samples included pre-specified high-mortality cancers such as anal cancer, bladder cancer, colon cancer, rectal cancer, esophageal cancer, head and neck cancer, liver cancer, bile duct cancer, lung cancer, ovarian cancer, pancreatic cancer, and gastric cancer, as well as other solid tumor cancers with a large pool of available samples, such as breast cancer, prostate cancer, and kidney cancer. Non-cancer samples were age-matched to the cancer cohorts, and their selection criteria were designed to ensure that results with a specificity greater than 99% could be observed with 90% statistical power in the hold-out set. Samples were analyzed using an aggregate analysis of approximately 30,000 methylation biomarkers and approximately 3,000 protein biomarkers. The proteins were selected from the OLINK EXPLORE protein collection (see, for example, Eldjarn et al., Nature, October 2023;622(7982):348-358).
[0306] Then, a cancer detection model based on single analyte analysis was used to analyze data on cfDNA methylation and protein levels. Figure 2The study found that false positives in cancer detection based on cfDNA methylation or protein levels are orthogonal. Figure 3 Different false positives suggest the possibility of complementary analysis of the two analytes. Furthermore, considering both cfDNA methylation and protein levels, more samples may be detected as cancer samples without loss of specificity. Figure 3 ).
[0307] Based on these results, determining cfDNA methylation and protein levels can be combined for cancer detection with higher sensitivity without reducing specificity. An integrated multi-omics classifier model was developed. Figure 4 Two component models were trained: a cfDNA model and a protein model. For feature preprocessing, cfDNA methylation features were binarized, and protein features were log-transformed counts normalized relative to control counts. The model architectures of the two component models were identical. The processed features were used as input to an ensemble of eight logistic regressions, and the outputs of the ensemble members were averaged to produce a single cancer probability. The fusion model had three features: the cancer probability from the cfDNA model, the cancer probability from the protein model, and the product of these probabilities. This fusion model was a logistic regression and output a single cancer probability for binary cancer classification. The integrated multi-omics classifier model utilized cancer scores from each of the cfDNA methylation and protein levels from the single analyte models as input to the combined cancer scoring model. Analysis of CCGA samples with the integrated multi-omics classifier model increased multi-cancer detection, and higher cancer detection sensitivity and specificity were observed through the combined analysis of cfDNA methylation and protein levels compared to separate cfDNA methylation analysis. Figure 5A Compared to cfDNA methylation analysis alone, combined analysis of cfDNA methylation and protein levels also resulted in higher sensitivity with 99.4% specificity. Figure 5B ).
[0308] In summary, analysis of both methylated cfDNA and protein levels leads to higher sensitivity for detecting cancer in samples with high specificity compared to analyzing protein or methylated cfDNA alone. Combined analysis of methylated cfDNA and protein levels will allow for more accurate identification of true positive samples during cancer screening. Example 2—Biomarker Selection and Model Training The following describes an algorithm for generating biomarker selection and model training for cancer detection.
[0309] From an initial set of 2,907 biomarker proteins (the OLINK EXPLORE set), biomarkers with small variance are removed, and all biomarkers are screened using a fast univariate feature selection metric. An iterative greedy forward selection algorithm is used to select a specified number (K) of proteins. In short, the proteins that best predict cancer status in the model training data are identified, as measured by the area under the curve (AUC). In each subsequent iteration, the performance obtained by adding each candidate protein biomarker to the existing model (previous iteration) is evaluated, also measured by AUC. The protein biomarkers that achieve the best performance are added to the model for the next iteration. The iteration loop stops when the number of selected biomarkers reaches the pre-specified K. Next, a logistic regression is fitted based on the K selected biomarkers, and the model is subsequently fitted to predict cancer status.
[0310] The cross-validation performance of the models is then evaluated. The training dataset is divided into 6 distinct partitions. Biomarker selection is performed on 5 / 6 of the partitions using a biomarker selection and model training algorithm (described below), and the model generated from this process is used to predict the cancer status on the remaining 1 / 6 of the training data. This process is repeated 6 times, each time leaving a different 1 / 6. This produces 6 different models. Two additional cross-validation runs are performed, repeating step 1 twice more, each time randomly regrouping the training set into a different mixture of the 6 partitions.
[0311] For each set of proteins (K) with a specified number of proteins, the cross-validation performance of the models yielded performance estimates for 18 different models (6 partitions × 3 random shufflings). Table 1 includes an exemplary list of protein biomarkers for cancer detection determined based on biomarker selection and model training algorithms, including models for a set of 50 proteins (List 1) and models for a set of 20 proteins (Lists 2-20). The following six biomarkers were observed to be present in at least 75% of the 20 biomarker models generated from different iterations: CHAD, KRT19, MMP12, PTN, SERPINA3, and SPP1.
[0312] The overall probability of cancer is determined using a linear model that combines probabilities from both the cfDNA-based model and the protein-based model. The overall probability is calculated using the following formula: In the above formula: =The final probability of cancer output given DNA and protein p_DNA = Cancer prediction from DNA models p_prot = Cancer prediction from protein model β0 = Bias term, constant / intercept β1 = Learning coefficient of DNA probability β2 = Learning coefficient of protein probability β3 = Learning coefficient of the product of DNA probability and protein probability Figure 6 The results demonstrate that cancer detection performance, with an average sensitivity of 99.4% specificity, can be achieved by using a model based on cfDNA methylation biomarkers combined with analyses of varying numbers of protein biomarkers. Significant improvements were provided with as few as five protein biomarkers compared to cfDNA methylation biomarkers alone, further improved by using 10, and even further improved by using 20. The results also show that models using as few as 20 (or 30, 40, or 50) biomarkers performed approximately as well as models based on the full initial set of 2,907 proteins. Therefore, the results indicate that combined models can be flexible and significantly simplified while maintaining performance similar to that of cfDNA methylation analyses alone.
[0313] Table 1: Exemplary protein biomarkers.
[0314] While various embodiments of this disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many variations, modifications, and substitutions will now occur to those skilled in the art without departing from this disclosure. It should be understood that various alternatives to the embodiments of this disclosure described herein may be employed in the practice of this disclosure. It is intended that the following claims define the scope of this disclosure and thereby cover the methods and structures within the scope of these claims and their equivalents.
Claims
1. A method for detecting cancer in a subject, characterized in that: The method includes: (a) Measure the level of a first target molecule from a first sample from the subject, wherein the first target molecule contains free DNA from multiple different target genomic regions that are differentially methylated in at least one of multiple cancer types; (b) Measure the level of a second target molecule from a second sample from the subject, wherein the second target molecule comprises multiple different peptides differentially expressed in at least one of the multiple cancer types; (c) Applying a trained classifier to the measurement levels of the first and second target molecules to assign an overall probability score to the cancer; wherein applying the trained classifier includes: (i) applying a first trained model to the measurement levels of the first target molecules to assign a first probability score to the cancer; (ii) applying a second trained model to the measurement levels of the second target molecules to assign a second probability score to the cancer; and (iii) summing the first and second probability scores; and (d) Detect cancer by identifying an overall probability score that is above the threshold for the presence of the cancer.
2. The method as described in claim 1, characterized in that: For reference samples from (1) reference subjects with known cancer and (2) reference subjects without cancer, the trained classifier is trained using a reference first probability score from the first trained model, a reference second probability score from the second trained model, and a reference overall probability score that sums the reference first probability scores and the reference second probability scores.
3. The method as described in claim 1, characterized in that: The trained classifier assigns an overall probability score to each of several different cancer types, and detecting a cancer involves identifying that cancer type as the one with the highest overall probability score.
4. The method as described in claim 1, characterized in that: Summarizing the first probability score and the second probability score includes calculating the product of the first probability score and the second probability score for the cancer.
5. The method as described in claim 1, characterized in that: The first sample and the second sample are identical.
6. The method as described in claim 1, characterized in that: These multiple different target genome regions include at least 1,000, 5,000, 10,000, 20,000, or 30,000 target genome regions.
7. The method as described in claim 1, characterized in that: These multiple target genomic regions have a total collective length of at least 50 kb, 100 kb, 500 kb, or 1000 kb.
8. The method as described in claim 1, characterized in that: Each of these multiple different target genomic regions contains at least five methylation sites.
9. The method as described in claim 1, characterized in that: Measuring these first target molecules involves sequencing transformed cfDNA or its amplified products from the multiple different target genomic regions, wherein the transformed cfDNA includes cfDNA treated with a deamination agent.
10. The method as described in claim 9, characterized in that: Further, it includes treating the cfDNA with the deamination agent, optionally wherein the deamination agent is cytosine deaminase or bisulfite.
11. The method as described in claim 9, characterized in that: The sequencing yielded at least 100,000 sequencing reads.
12. The method as described in claim 8, characterized in that: Measuring these first target molecules involves enriching the transformed cfDNA or its amplification products to produce enriched polynucleotide samples.
13. The method as described in claim 12, characterized in that: The enrichment process involves capturing the transformed cfDNA or its amplified product using multiple corresponding bait oligonucleotides.
14. The method as described in claim 13, characterized in that: The multiple distinct target genomic regions used for enrichment by these bait oligonucleotides are genomic regions identified by the first trained model as differentially methylated in at least one of multiple cancer types relative to non-cancer tissue or relative to different types of cancer.
15. The method as described in claim 1, characterized in that: The multiple different polypeptides include at least 5, 10, 25, 50, 100, 200, 500, 1000, 2000, 3000, 5000, or 7500 different polypeptides.
16. The method as described in claim 15, characterized in that: The various polypeptides include: (a) polypeptides that identify proteins selected from List 1; (b) polypeptides that identify proteins selected from any of Lists 2-19; (c) polypeptides that identify proteins selected from List 20; or (d) polypeptides that identify one or more of CHAD, KRT19, MMP12, PTN, SERPINA3, and SPP1.
17. The method as described in claim 1, characterized in that: The trained classifier distinguishes subjects with cancer from those without cancer based on the specificity defined for each of the multiple cancer types.
18. The method as described in claim 1, characterized in that: The trained classifier has a higher cancer detection sensitivity than each of the first trained model and the second trained model; optionally, the trained classifier has a cancer detection specificity equal to or greater than that of each of the first trained model and the second trained model.
19. The method as described in claim 1, characterized in that: The trained classifier is a binary classifier, a hybrid model classifier, a multilayer perceptron model classifier, or a logistic regression classifier.
20. The method as described in claim 1, characterized in that: The first trained model and / or the second trained model is a binary classifier, a hybrid model classifier, a multilayer perceptron model classifier, or a logistic regression classifier.
21. The method as described in claim 1, characterized in that: The first trained model binarizes the measurement levels of these first target molecules by assigning a first value when the target genomic region is detected and a second value when the target genomic region is not detected.
22. The method as described in claim 1, characterized in that: The second trained model performs a logarithmic transformation on the measurement levels of these second target molecules, normalized to control proteins present in known amounts.
23. The method as described in claim 1, characterized in that: include: (a) The first trained model is trained using measurement levels of the first target molecules against the first reference sample, (b) The second trained model is trained using measurement levels of the second target molecules against the second reference sample, and (c) The first and second reference samples include samples from reference subjects with known cancer and reference subjects without cancer.
24. The method as described in claim 1, characterized in that: The first sample and / or the second sample includes a biological fluid; optionally, the biological fluid includes blood, plasma, serum, urine, saliva, pleural fluid, pericardial fluid, cerebrospinal fluid, peritoneal fluid, or any combination thereof.
25. The method as described in claim 24, characterized in that: This biofluid includes blood, blood fractions, plasma, or serum.
26. The method as described in claim 25, characterized in that: The first sample and / or the second sample is a plasma sample.
27. The method as described in claim 1, characterized in that: These multiple cancer types include at least 10 cancer types.
28. The method as described in claim 1, characterized in that: These multiple cancer types include one or more of the following: anorectal cancer, bladder cancer, colorectal cancer, esophageal cancer, head and neck cancer, liver cancer, bile duct cancer, lung cancer, ovarian cancer, pancreatic cancer, stomach cancer, breast cancer, prostate cancer, kidney cancer, cervical cancer, endometrial cancer, and blood cancer.
29. The method according to any one of claims 1-28, characterized in that: This further includes the type of cancer that was treated in the subject.
30. The method as described in claim 29, characterized in that: The treatment includes surgical resection, radiation therapy, chemotherapy, and / or immunotherapy.
31. A method for treating cancer in a subject, characterized in that: The method includes selecting subjects based on the results of detection measurements and treating the subject's cancer, wherein: (a) The detection method comprises the method as described in any one of claims 1-28; and (b) The treatment includes surgical resection, radiation therapy, chemotherapy and / or immunotherapy.
32. A method for training a classifier to detect target molecules from cancer, the method comprising: (a) Receiving a first measurement level of a first target molecule in a first sample for reference subjects, wherein (i) the first target molecules contain free DNA from multiple different target genomic regions that are differentially methylated in at least one of multiple cancer types, and (ii) the reference subjects include a first subject with a known cancer type and a second subject without cancer; (b) A first model is trained by applying a first machine learning algorithm to these first measurement levels to generate a first probability score of the presence of cancer in the subject; (c) Receive a second measurement level of a second target molecule in a second sample for these reference subjects, wherein the second target molecule comprises a plurality of different polypeptides differentially expressed in at least one of the plurality of cancer types; (d) A second model is trained to generate a second probability score of the presence of cancer in the subject by applying a second machine learning algorithm to these second measurement levels; (e) Use the trained first model to generate a reference first cancer probability score for these first samples; (f) Use the trained second model to generate reference second cancer probability scores for these second samples; (g) Generate a reference overall cancer probability score for multiple reference subjects by aggregating the reference first cancer probability score and the reference second cancer probability score for each corresponding reference subject; as well as (h) A classifier is trained to generate the subject’s overall cancer probability score by applying a third machine learning algorithm to these reference first cancer probability scores, these reference second cancer probability scores and these reference overall cancer probability scores.
33. The method as described in claim 32, characterized in that: Summarizing the first cancer probability score and the second cancer probability score includes calculating the product of the first probability score and the second probability score for the cancer.
34. The method as described in claim 32, characterized in that: The first machine learning algorithm, the second machine learning algorithm, and / or the third machine learning algorithm are L1 regularized logistic regression, L2 regularized logistic regression, generalized linear model, random forest, multinomial logistic regression, multilayer perceptron, support vector machine, or neural network.
35. The method as described in claim 32, characterized in that: The first trained model binarizes the measurement levels of these first target molecules by assigning a first value when the target genomic region is detected and a second value when the target genomic region is not detected.
36. The method as described in claim 32, characterized in that: The second trained model performs a logarithmic transformation on the measurement levels of these second target molecules, normalized to control proteins present in known amounts.
37. The method as described in claim 32, characterized in that: include: (a) The first trained model is trained using measurement levels of the first target molecules against the first reference sample, (b) The second trained model is trained using measurement levels of the second target molecules against the second reference sample, and (c) The first and second reference samples include samples from reference subjects with known cancer and reference subjects without cancer.
38. The method as described in claim 32, characterized in that: The third machine learning algorithm is logistic regression.
39. The method as described in claim 32, characterized in that: These multiple different target genome regions include at least 1,000, 5,000, 10,000, 20,000, or 30,000 target genome regions.
40. The method as described in claim 32, characterized in that: These multiple target genomic regions have a total collective length of at least 50 kb, 100 kb, 500 kb, or 1000 kb.
41. The method as described in claim 32, characterized in that: Each of these multiple different target genomic regions contains at least five methylation sites.
42. The method as described in claim 32, characterized in that: These first measurement levels include sequencing results of the cfDNA or its amplicon.
43. The method as described in claim 42, characterized in that: For each of these first samples, these sequencing results include at least 100,000 reads.
44. The method as described in claim 32, characterized in that: The multiple different polypeptides include at least 5, 10, 25, 50, 100, 200, 500, 1000, 2000, 3000, 5000, or 7500 different polypeptides.
45. The method as described in claim 44, characterized in that: The various polypeptides include: (a) polypeptides that identify proteins selected from List 1; (b) polypeptides that identify proteins selected from any of Lists 2-19; (c) polypeptides that identify proteins selected from List 20; or (d) polypeptides that identify one or more of CHAD, KRT19, MMP12, PTN, SERPINA3, and SPP1.
46. The method as described in claim 32, characterized in that: These multiple cancer types include at least 10 cancer types.
47. The method as described in claim 32, characterized in that: These multiple cancer types include one or more of the following: anorectal cancer, bladder cancer, colorectal cancer, esophageal cancer, head and neck cancer, liver cancer, bile duct cancer, lung cancer, ovarian cancer, pancreatic cancer, stomach cancer, breast cancer, prostate cancer, kidney cancer, cervical cancer, endometrial cancer, and blood cancer.
48. A method for screening a subject for cancer, characterized in that: The method includes: (a) Measure the level of a first target molecule from a first sample from the subject, wherein the first target molecule comprises multiple different polypeptides differentially expressed in at least one of multiple cancer types; (b) Applying a first trained model to the measurement level of these first target molecules to assign a first probability score to each of the plurality of cancer types; wherein (i) the first trained model has a first specificity for cancer detection, and (ii) the first probability score of at least one of these cancer types is above a first threshold for the presence of cancer. (c) Measure the level of a second target molecule from a second sample from the subject, wherein the second target molecule contains free DNA from multiple different target genomic regions that are differentially methylated in at least one of the multiple cancer types; (d) Applying a second trained model to the measurement levels of these second target molecules to assign a second probability score to the cancer; wherein (ii) the second trained model has a second specificity for cancer detection, and (ii) the second specificity is higher than the first specificity; and (e) Detect the cancer by identifying that the second probability score is above the threshold for the presence of the cancer.
49. The method as described in claim 48, characterized in that: include: (a) The first trained model is trained using measurement levels of the first target molecules against the first reference sample, (b) The second trained model is trained using measurement levels of the second target molecules against the second reference sample, and (c) The first and second reference samples include samples from reference subjects with known cancer and reference subjects without cancer.
50. The method as described in claim 48, characterized in that: The first sample and the second sample are identical.
51. The method as described in claim 48, characterized in that: These multiple different target genome regions include at least 1,000, 5,000, 10,000, 20,000, or 30,000 target genome regions.
52. The method as described in claim 48, characterized in that: These multiple target genomic regions have a total collective length of at least 50 kb, 100 kb, 500 kb, or 1000 kb.
53. The method as described in claim 48, characterized in that: Each of these multiple different target genomic regions contains at least five methylation sites.
54. The method as described in claim 48, characterized in that: Measuring these second target molecules involves sequencing transformed cfDNA or its amplified products from the multiple different target genomic regions, wherein the transformed cfDNA includes cfDNA treated with a deamination agent.
55. The method as described in claim 54, characterized in that: The sequencing yielded at least 100,000 sequencing reads.
56. The method as described in claim 54, characterized in that: Measuring these first target molecules involves enriching the transformed cfDNA or its amplification products to produce enriched polynucleotide samples.
57. The method as described in claim 56, characterized in that: The enrichment process involves capturing the transformed cfDNA or its amplified product using multiple corresponding bait oligonucleotides.
58. The method as described in claim 57, characterized in that: The multiple distinct target genomic regions used for enrichment by these bait oligonucleotides are identified by the second trained model as genomic regions that are differentially methylated in at least one of multiple cancer types relative to non-cancer tissue or relative to different types of cancer.
59. The method as described in claim 48, characterized in that: The multiple different polypeptides include at least 5, 10, 25, 50, 100, 200, 500, 1000, 2000, 3000, 5000, or 7500 different polypeptides.
60. The method as described in claim 59, characterized in that: The various polypeptides include: (a) polypeptides that identify proteins selected from List 1; (b) polypeptides that identify proteins selected from any of Lists 2-19; (c) polypeptides that identify proteins selected from List 20; or (d) polypeptides that identify one or more of CHAD, KRT19, MMP12, PTN, SERPINA3, and SPP1.
61. The method as described in claim 48, characterized in that: The first trained model and / or the second trained model is a neural network classifier, a binary classifier, a hybrid model classifier, a multilayer perceptron model classifier, or a logistic regression classifier.
62. The method as described in claim 48, characterized in that: The second trained model binarizes the measurement levels of these second target molecules by assigning a first value when the target genomic region is detected and a second value when the target genomic region is not detected.
63. The method as described in claim 48, characterized in that: The first trained model performs a logarithmic transformation on the measurement levels of these first target molecules, normalized to control peptides present in known amounts.
64. The method as described in claim 48, characterized in that: include: (a) The first sample and / or the second sample comprises a biological fluid; optionally, the biological fluid comprises blood, plasma, serum, urine, saliva, pleural fluid, pericardial fluid, cerebrospinal fluid, peritoneal fluid, or any combination thereof; (b) The first sample and / or the second sample is a plasma sample.
65. The method as described in claim 48, characterized in that: include: (a) The multiple cancer types include at least 10 cancer types; and / or (b) The multiple cancer types include one or more of the following: anorectal cancer, bladder cancer, colorectal cancer, esophageal cancer, head and neck cancer, liver cancer, bile duct cancer, lung cancer, ovarian cancer, pancreatic cancer, stomach cancer, breast cancer, prostate cancer, kidney cancer, cervical cancer, endometrial cancer, and blood cancer.
66. The method according to any one of claims 48-65, characterized in that: Further, it may include treating the subject for the type of cancer; optionally, the treatment may include surgical resection, radiation therapy, chemotherapy, and / or immunotherapy.
67. A method for treating cancer in a subject, characterized in that: The method includes selecting subjects based on the results of screening tests and treating the subject for the cancer, wherein: (a) The screening assay comprises the method as described in any one of claims 48-65; and (b) The treatment includes surgical resection, radiation therapy, chemotherapy and / or immunotherapy.
68. A non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, perform one or more steps of the method as claimed in any one of claims 1-28 or 48-65.
69. A non-transitory computer-readable medium having instructions stored thereon, characterized in that: These instructions, when executed by one or more processors, perform the method as described in any one of claims 32-47.
Citation Information
Patent Citations
Method for amplifying nucleic acid sequence and reagent kid therefor
JP1992262799A
Multiplexed Analyses of Test Samples
US20090042206A1
Selective oxidation of 5-methylcytosine by TET-family proteins
US20110236894A1
Composition and Methods Related to Modification of 5-Hydroxymethylcytosine (5-hmC)
US20110301045A1
Hyperthermophilic Polymerase Enabled Proximity Extension Assay
US20150044674A1