Determining methylation status of biological samples
By counting methylated CpGs and using machine learning, the method addresses sensitivity and invasiveness issues in BRCA1 promoter methylation detection, achieving accurate and non-invasive cancer diagnostics through cell-free DNA analysis.
Patent Information
- Application Number
- PCT/US2025/037016
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-04-22
- Filing Date
- 2025-07-09
- Publication Date
- 2026-01-15
AI Technical Summary
Current methods for detecting BRCA1 promoter methylation in cancer patients are limited by low sensitivity, require invasive tissue biopsies, and struggle with overestimation or underestimation due to sequencing noise and bisulfite conversion failures, making it difficult to accurately quantify methylation status from cell-free DNA.
A method that counts methylated CpGs per molecule and filters sequence reads to remove unconverted fragments, combined with machine learning models to estimate methylation status from cell-free DNA, enabling accurate detection of BRCA1 promoter hypermethylation without tissue biopsies.
The method provides sensitive and high-throughput quantification of BRCA1 promoter methylation, improving accuracy and sensitivity, allowing for non-invasive cancer diagnostics and treatment decisions based on cell-free DNA analysis.
Smart Images

Figure US2025037016_15012026_PF_FP_ABST
Abstract
Description
ATTORNEY DOCKET PATENT APPLICATION 087932.0118 1 of 119 Determining Methylation Status of Biological Samples TECHNICAL FIELD
[0001] This disclosure generally relates to DNA-based medical diagnostics, and in particular methods for calling BRCA1 promoter methylation from cfDNA and model-based estimations for determining the methylation status of cfDNA. BACKGROUND
[0002] Deoxyribonucleic acid (DNA) methylation plays an important role in regulating gene expression. Aberrant DNA methylation has been implicated in many disease processes, including cancer. DNA methylation profiling using methylation sequencing (e.g., whole genome bisulfite sequencing (WGBS) or targeted methylation sequencing) is increasingly recognized as a valuable diagnostic tool for detection, diagnosis, and / or monitoring of cancer. For example, specific patterns of differentially methylated regions and / or allele specific methylation patterns may be useful as molecular markers for non-invasive diagnostics using circulating cell-free (cf) DNA. As part of cancer classification, there remains a need to understand the effect covariate variables (or, more generally, variables that indicate cancer or non-cancer) may have on the human genome. Moreover, there remains a need to be able to distinguish variables that may indicate cancer and / or some other biological attribute such as age or sex.
[0003] BRCA1 promoter hypermethylation (BRCA1meth) is observed in up to 25% of triple negative breast cancers (TNBCs) and ovarian cancers (OvCas), is associated with worse outcomes compared to BRCA1 mutations (Menghi et al., PMID: 35857626), and, when found in individuals without cancer (i.e. constitutional BRCA1meth), is associated with increased risk of breast and ovarian cancer (Lonning, et al., PMID: 36074460). Therefore, identifying patients who have promoter methylation at BRCA1 can have therapeutic implications for both cancer patients, as well as risk implications in patients without cancer.
[0004] Several methods have been developed to assess locus specific methylation. Methylation specific PCR (MSP) uses methylation specific primers and provides a binary output (methylated vs. unmethylated) without any relative quantification (Ku et al., PMID: 21913069). The Methylation Sensitive- High Resolution Melting (MS-HRM) protocol exploits the difference in melting temperatures between GC-rich (i.e., methylated) vs. AT-rich (i.e., unmethylated) DNA molecules after methylation-nonspecific PCR amplification of bisulfiteACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 2 of 119 converted DNA (Javadmanesh et al., PMID: 35616708). Though both methods are easy to implement, they do not provide resolution at the level of individual CpG sites and also have limited sensitivity. In clonal bisulfite sequencing, Sanger sequencing of individual bisulfite converted DNA molecules cloned into recipient plasmids allows for the precise phasing of the methylation status of all CpGs within each fragment, and therefore precise determination of patterns of methylation. The requirement for cloning each fragment for Sanger sequencing significantly limits the throughput and cannot be easily used to quantify low methylation frequencies. Although phasing and representative quantification can be achieved through next generation sequencing (NGS) methodologies (Li et al., PMID: 21913068), current methodologies to determine the presence of BRCA1meth in a cancer sample require a tissue biopsy and have not been adopted for standard clinical assessment. On the other hand, a DNA (cfDNA)-based targeted methylation platform could provide a robust, biopsy-free, scalable assay that can distinguish cancer methylation patterns between different cancer signal origins. However, sensitive and specific methods are needed to identify patients with BRCA1 promoter methylation from a cfDNA test. SUMMARY OF PARTICULAR EMBODIMENTS
[0005] The presently disclosed subject matter provides methods for quantifying methylation at a single locus, that is sensitive (e.g., LoD95 of 0.0081), provides granular CpG- level information, and is high-throughput. Additionally, the method provides the advantage of using a cell-free DNA (cfDNA) approach to assess the summation of the BRCA1meth throughout the body rather than a single tissue. Furthermore, the presently disclosed subject matter in built on a clinically validated targeted-methylation platform, enabling interpretation alongside other metrics, such as tumor fraction (Melton et al., PMID: 38201510).
[0006] In particular embodiments, the method may measure a fraction of methylated molecules observed at the BRCA1 promoter region of a DNA sample (e.g., cfDNA). Although this disclosure describes counting the number methylated sequence reads with greater than a fixed number (e.g., 4 or 5) of methylated CpGs per molecule, this disclosure contemplates any suitable method for measuring the fraction of methylated molecules at the BRCA1 promoter region in any suitable manner.
[0007] In particular embodiments, the methylated sequence reads are obtained from a targeted methylation sequencing assay, or a whole genome bisulfite sequencing assay.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 3 of 119
[0008] In a certain embodiment, the method measures a fraction of methylated molecules at a BRCA1 promoter region by counting the number of methylated sequence reads with greater than 3, 4, or 5 methylated CpGs per molecule. In particular embodiments, the methylated sequence reads are filtered to remove the molecules when the CpGs have failed to undergo conversion during a bisulfite treatment and before the methylated sequence reads are counted.
[0009] Certain technical challenges exist for identifying patients who have BRCA1 promoter hypermethylation (BRCA1meth). One technical challenge may include overestimating or underestimating the BRCA1meth fraction (e.g., BRCA1meth fraction = methylated molecules / total molecules) due to noise (e.g., sequencing error, contamination). The solution presented by the embodiments disclosed herein to address this challenge may be counting the number of individual molecules with greater than a fixed number (e.g., 4 or 5) of methylated CpGs. Another technical challenge may include overestimating or underestimating the BRCA1meth fraction due to CpGs within the molecules failing to undergo bisulfite conversion. The solution presented by the embodiments disclosed herein to address this challenge may be filtering sequence reads to remove fragments with a high frequency of unconverted CHH / CHG per fragment. Still another technical challenge may include overestimating or underestimating rates of cfBRCA1 meth between patient samples due to variation in sequencing depth across samples. The solution presented by the embodiments disclosed herein to address this challenge may be restricting samples with a minimal coverage value in the cfBRCA1 meth region when comparing rates of cfBRCA1 meth across patient subsets. Another solution presented by the embodiments disclosed herein may address this challenge by downsampling the fragment files.
[0010] Certain embodiments disclosed herein may provide one or more technical advantages. A technical advantage of the embodiments may be the ability to detect BRCA1 promoter hypermethylation from a cfDNA sample. Another technical advantage of the embodiments may include increasing the accuracy and sensitivity of determining the BRCA1meth fraction in a patient sample.
[0011] In particular embodiments, an analytics system may estimate the methylation status of a single genomic site in circulating tumor DNA (ctDNA) using data modeling and machine learning, even when direct observation of the site is not possible due to low tumor DNA content in a liquid biopsy. By leveraging methylation patterns from multiple genomic regions observed in tissue or high-tumor-fraction plasma samples, the analytics system may train predictiveACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 4 of 119 models including classifiers and regression algorithms to infer the methylation state of a target site. This approach enables accurate epigenetic profiling from cell-free DNA (cfDNA), supporting cancer diagnostics and treatment decisions without requiring invasive tissue biopsies. The analytics system may also incorporate tumor methylation fraction estimation and model training using synthetic data, allowing for robust performance across diverse cancer types and sample conditions. Although this disclosure describes using particular systems to determine particular methylation status in particular manners, this disclosure contemplates using any suitable system for determining any suitable methylation status in any suitable manner.
[0012] In particular embodiments, the analytics system may access a cell-free DNA (cfDNA) extracted from a biological sample. The analytics system may then generate a plurality of sequence reads for the cfDNA. The analytics system may then identify, based on the sequence reads, methylation patterns at a plurality of genomic sites within the cfDNA. The analytics system may further execute one or more machine-learning models on the sequence reads and the identified methylation patterns to estimate a methylation status of a target genomic site. In particular embodiments, the one or more machine-learning models were trained using a set of reference data generated from tissue or plasma samples with identified methylation status at genomic sites.
[0013] Certain technical challenges exist for estimating single site methylation status. One technical challenge may include estimating the methylation status of a site when no ctDNA fragment covering that site is observed. In many clinical scenarios, especially early-stage cancer or during treatment, the fraction of ctDNA in a blood sample is extremely low, which makes it difficult or impossible to directly observe the methylation status of a specific genomic site using conventional methods. In addition, traditional methods like methylation-specific PCR or targeted sequencing require direct coverage of the site of interest, which may not be present in the sample due to stochastic sampling or low ctDNA abundance. The solution presented by the embodiments disclosed herein to address this challenge may include modeling methylation relationships and estimating single site methylation status without direct observation. The analytics system may use machine learning to learn the relationship between the methylation status of a target site and the methylation patterns of many other genomic sites and infer the methylation status of the target site based on the relationship between the methylation status of the target site and the methylation patterns of other genomic sites even when the target site is not directly observed in the cfDNA sample.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 5 of 119
[0014] Certain embodiments disclosed herein may provide one or more technical advantages. A technical advantage of the embodiments may be the ability to detect BRCA1 promoter hypermethylation from a cfDNA sample. Another technical advantage of the embodiments may include increasing the accuracy and sensitivity of determining the BRCA1 meth fraction in a patient sample. Another technical advantage of the embodiments may include addressing shortcomings of current attempts to assess tumor methylation status at a single genomic site in liquid biopsy. First, the embodiments disclosed herein may allow to assess the methylation site of multiple potential sites of interest with the same data, method, and assay. For many typical methods, individual site-specific reagents (for example, methylation-specific PCR primers or single-stranded pull-down probes) are needed to target a new site of interest. Second, the embodiments disclosed herein may enable estimates of the methylation status of a tumor even in the presence of few (down to no) ctDNA molecule in a sample that covers the site of interest, such that no direct read-out of the tumor methylation status is possible.
[0015] Certain embodiments disclosed herein may provide none, some, or all of the above technical advantages. One or more other technical advantages may be readily apparent to one skilled in the art in view of the figures, descriptions, and claims of the present disclosure.
[0016] In particular embodiments, the techniques described herein relate to a method for measuring a fraction of methylated molecules at a BRCA1 promoter region, including: (a) obtaining a plurality of nucleic acids from a biological sample; (b) generating a plurality of sequence reads from the nucleic acids; (c) determining a methylation status at a plurality of CpG sites within each sequence read of the plurality of sequence reads; (d) counting a number of methylated sequence reads and a number of unmethylated sequence reads, wherein the sequence read is counted as methylated if at least (n) number of methylated CpGs per molecule are present, and unmethylated if less than (n) number of methylated CpGs per molecule are present; and (e) calculating the fraction of methylated molecules by dividing the number of methylated sequence reads by a total number of methylated and unmethylated sequence reads.
[0017] In particular embodiments, the techniques described herein relate to a method for determining a methylation status of a BRCA1 promoter region in a subject, including: (a) obtaining a plurality of nucleic acids from a biological sample obtained from the subject; (b) generating a plurality of sequence reads from the nucleic acids; (c) determining a methylation status at a plurality of CpG sites within each sequence read of the plurality of sequence reads; (d) counting the sequence read as methylated, if the sequence read has at least (n) number ofACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 6 of 119 methylated CpGs per molecule, and counting the sequence read as unmethylated if the sequence read has less than (n) number of methylated CpGs per molecule; and (e) determining that the BRCA1 promoter region is methylated in the subject if there is at least one read counted as methylated according to step (d).
[0018] In particular embodiments, the techniques described herein relate to a method, wherein the (n) number of methylated CpGs per molecule is 3, 4, or 5.
[0019] In particular embodiments, the techniques described herein relate to a method, wherein the (n) number of methylated CpGs per molecule is 4.
[0020] In particular embodiments, the techniques described herein relate to a method, wherein the plurality of sequence reads are mapped to a reference genome.
[0021] In particular embodiments, the techniques described herein relate to a method, wherein the plurality of sequence reads of step (c) overlap with the BRCA1 promoter region.
[0022] In particular embodiments, the techniques described herein relate to a method, wherein the plurality of sequence reads of step (c) overlap with genomic coordinates, chr17: 41277134-41277486; GRCh37.
[0023] In particular embodiments, the techniques described herein relate to a method, wherein the plurality of sequence reads are filtered to remove molecules with fewer than 4 CpG sites prior to step (c).
[0024] In particular embodiments, the techniques described herein relate to a method, wherein the plurality of nucleic acids are subjected to a bisulfite conversion.
[0025] In particular embodiments, the techniques described herein relate to a method, wherein the plurality of sequence reads are generated by a targeted panel, whole-exome panel, or whole-genome panel.
[0026] In particular embodiments, the techniques described herein relate to a method, wherein the plurality of sequence reads are filtered to remove molecules that have failed to undergo conversion during bisulfite conversion prior to step (c).
[0027] In particular embodiments, the techniques described herein relate to a method, wherein the plurality of sequence reads are deduplicated prior to step (d).
[0028] In particular embodiments, the techniques described herein relate to a method, wherein the biological sample is tissue, blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal, saliva, sweat, tears, pleural fluid, pericardial fluid, peritoneal fluid, or a tumor biopsy.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 7 of 119
[0029] In particular embodiments, the techniques described herein relate to a method, wherein the biological sample is a liquid sample or a solid sample.
[0030] In particular embodiments, the techniques described herein relate to a method, wherein the plurality of nucleic acids are derived from cfDNA.
[0031] In particular embodiments, the techniques described herein relate to a method, further including determining a methylation status of a fragment, wherein the fragment is considered methylated if a corresponding sequencing read contains greater than 4 methylated CpG sites; and the sample is considered methylated if at least one fragment is considered methylated.
[0032] In particular embodiments, the techniques described herein relate to an assay for determining a subject's BRCA1 promoter methylation status including: (a) obtaining a plurality of nucleic acids from a biological sample obtained from the subject; (b) generating a plurality of sequence reads from said nucleic acids; (c) determining a methylation status at a plurality of CpG sites within each sequence read of the plurality of sequence reads; (d) counting the sequence read as methylated, if the sequence read has at least 4 methylated CpGs per molecule, and counting the sequence read as unmethylated if the sequence read has less than 4 methylated CpGs per molecule; and (e) determining that a BRCA1 promoter region is methylated in the subject if there is at least one read counted as methylated according to step (d).
[0033] In particular embodiments, the techniques described herein relate to an assay, further including determining a fraction of methylated fragments, wherein: a fragment is considered methylated if a corresponding sequence read contains greater than 4 methylated CpG sites; the fragment is considered unmethylated if the corresponding sequence read contains less than 4 methylated CpG sites; and the fraction of methylated fragments is calculated by dividing a number of methylated fragments by a total number of methylated and unmethylated fragments.
[0034] In particular embodiments, the techniques described herein relate to an assay, further including determining the fraction of methylated fragments, wherein the sample is considered methylated if at least one fragment is considered methylated.
[0035] In particular embodiments, the techniques described herein relate to an assay, wherein the plurality of sequence reads are generated by a targeted panel, whole-exome panel, or whole-genome panel.
[0036] In particular embodiments, the techniques described herein relate to an assay, wherein the biological sample is tissue, blood, whole blood, plasma, serum, urine,ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 8 of 119 cerebrospinal fluid, fecal, saliva, sweat, tears, pleural fluid, pericardial fluid, peritoneal fluid, or a tumor biopsy.
[0037] In particular embodiments, the techniques described herein relate to an assay, wherein the plurality of nucleic acids are derived from cfDNA.
[0038] In particular embodiments, the techniques described herein relate to an assay, wherein the plurality of sequence reads are generated as part of a multi-cancer early detection (MCED) test.
[0039] In particular embodiments, the techniques described herein relate to an assay, wherein the MCED test report includes a fraction of BRCA1 promoter methylation detected in the subject.
[0040] In particular embodiments, the techniques described herein relate to a method including, by one or more computing systems: accessing a cell-free DNA (cfDNA) extracted from a biological sample; generating a plurality of sequence reads for the cfDNA; identifying, based on the sequence reads, methylation patterns at a plurality of genomic sites within the cfDNA; and executing one or more machine-learning models on the sequence reads and the identified methylation patterns to estimate a methylation status of a target genomic site, wherein the one or more machine-learning models were trained using a set of reference data generated from tissue or plasma samples with identified methylation status at genomic sites.
[0041] In particular embodiments, the techniques described herein relate to a method, wherein at least one of the machine-learning models includes a supervised machine learning algorithm configured to associate methylation patterns at the plurality of genomic sites with the methylation status of the target genomic site.
[0042] In particular embodiments, the techniques described herein relate to a method, wherein the set of reference data is generated from tissue samples, and wherein the identified methylation status in the tissue samples are determined based on one or more of whole-genome bisulfite sequencing (WGBS), reduced representation bisulfite sequencing (RRBS), or a targeted methylation panel.
[0043] In particular embodiments, the techniques described herein relate to a method, wherein the set of reference data is generated from plasma samples, and wherein the identified methylation status in the plasma samples are determined based on tumor fractions.
[0044] In particular embodiments, the techniques described herein relate to a method, wherein the methylation status of the target genomic site is estimated in an absence of a direct observation of the target genomic site in the cfDNA.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 9 of 119
[0045] In particular embodiments, the techniques described herein relate to a method, further including: estimating tumor-derived fragments in the cfDNA based on applying an expectation maximization algorithm to the sequence reads associated with the cfDNA; and refining the estimation of the methylation status at the target genomic site based on the estimated tumor-derived fragments.
[0046] In particular embodiments, the techniques described herein relate to a method, wherein estimating the tumor-derived fragments includes: estimating the tumor-derived fragments separately for cancer types with methylated genomic sites and cancer types with unmethylated genomic sites.
[0047] In particular embodiments, the techniques described herein relate to a method, wherein refining the estimation of the methylation status at the target genomic site based on the estimated tumor-derived fragments includes estimating the methylation status at the target site as a cancer type with a highest amount of estimated tumor-derived fragments.
[0048] In particular embodiments, the techniques described herein relate to a method, wherein estimating the tumor-derived fragments includes: estimating a fraction of non-cancer cfDNA, a fraction of ctDNA for a cancer with methylated genomic sites, and a fraction of ctDNA with unmethylated genomic sites.
[0049] In particular embodiments, the techniques described herein relate to a method, wherein at least one of the machine-learning models include a regression model, the method further including: inputting the estimated tumor-derived fragments to the regression model.
[0050] In particular embodiments, the techniques described herein relate to a method, further including: detecting a presence of a circulating tumor DNA (ctDNA) in the cfDNA, wherein estimating the methylation status of the target genomic site is further based on the presence of the ctDNA.
[0051] In particular embodiments, the techniques described herein relate to a method, further including: predicting a presence of one or more cancer subtypes based on the sequence reads from the cfDNA, wherein estimating the methylation status of the target genomic site is further based on the presence of the one or more cancer subtypes.
[0052] In particular embodiments, the techniques described herein relate to a method, further including: determining, based on the one or more machine-learning models and the estimated methylation status of the target genomic site, a cancer prediction.
[0053] In particular embodiments, the techniques described herein relate to a method, wherein the cancer prediction includes one or more of a label indicating a particular cancerACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 10 of 119 state, a label indicating a particular cancer type, a label indicating a particular cancer stage, a likelihood of cancer, or a likelihood of a particular cancer type.
[0054] In particular embodiments, the techniques described herein relate to a method, wherein the target genomic site is associated with breast cancer gene 1 (BRCA1), and wherein the cancer prediction is associated with breast cancer.
[0055] In particular embodiments, the techniques described herein relate to a method, wherein the target genomic site is associated with tumor-associated calcium signal transducer 2 (Trop-2), and wherein the cancer prediction is associated with small cell lung cancer.
[0056] In particular embodiments, the techniques described herein relate to a method, wherein the target genomic site is a CpG site.
[0057] In particular embodiments, the techniques described herein relate to a method, wherein the biological sample is a blood draw or a liquid biopsy.
[0058] The embodiments disclosed herein are only examples, and the scope of this disclosure is not limited to them. Particular embodiments may include all, some, or none of the components, elements, features, functions, operations, or steps of the embodiments disclosed herein. Embodiments according to the invention are in particular disclosed in the attached claims directed to a method, a storage medium, a system and a computer program product, wherein any feature mentioned in one claim category, e.g. method, can be claimed in another claim category, e.g. system, as well. The dependencies or references back in the attached claims are chosen for formal reasons only. However, any subject matter resulting from a deliberate reference back to any previous claims (in particular multiple dependencies) can be claimed as well, so that any combination of claims and the features thereof are disclosed and can be claimed regardless of the dependencies chosen in the attached claims. The subject-matter which can be claimed comprises not only the combinations of features as set out in the attached claims but also any other combination of features in the claims, wherein each feature mentioned in the claims can be combined with any other feature or combination of other features in the claims. Furthermore, any of the embodiments and features described or depicted herein can be claimed in a separate claim and / or in any combination with any embodiment or feature described or depicted herein or with any of the features of the attached claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] FIG.1 illustrates an example overall workflow of cancer classification of a sample.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 11 of 119
[0060] FIG. 2A illustrates an example flowchart of devices for sequencing nucleic acid samples.
[0061] FIG. 2B illustrates an example block diagram of an analytics system 200 for processing DNA samples.
[0062] FIG.3 illustrates an example flowchart describing a process of sequencing nucleic acids.
[0063] FIG. 4A illustrates an example flowchart describing a process of sequencing a fragment of cfDNA to obtain a methylation state vector.
[0064] FIG.4B illustrates an example process of FIG.4A of sequencing a cfDNA molecule to obtain a methylation state vector.
[0065] FIG.5 illustrates example blocks of a reference genome.
[0066] FIG.6A illustrates example methylation features that can be derived from a single CpG site as a genomic region.
[0067] FIG.6B illustrates example methylation features that can be derived from multiple CpG sites as a genomic region.
[0068] FIG. 7A illustrates an example method of generating a data structure for a healthy control group with which the analytics system 200 may calculate p-value scores.
[0069] FIG. 7B illustrates an example method of calculating a p-value score with the generated data structure.
[0070] FIG. 8A illustrates an example flowchart describing a process of training a cancer classifier.
[0071] FIG. 8B illustrates an example generation of feature vectors used for training the cancer classifier.
[0072] FIG.9 illustrates various sources of tissues from which BRCA1 methylation can be detected.
[0073] FIG.10 illustrates methylation in various samples.
[0074] FIG.11 illustrates an example of methylation data at BRCA1.
[0075] FIG.12 illustrates a set of considerations for calling BRCA1 methylation.
[0076] FIG.13 illustrates a set of considerations for calling BRCA1 methylation.
[0077] FIG.14A illustrates the probability of incorrectly calling a molecule methylated.
[0078] FIG.14B illustrates the probability of incorrectly calling a sample methylated.
[0079] FIG.15 illustrates an example computer system.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 12 of 119
[0080] FIG. 16 illustrates a schematic for determining the level of BRCA1 promoter methylation.
[0081] FIG. 17A demonstrates the linearity of a BRCA1 promoter methylation assay (cfBRCA1meth).
[0082] FIG. 17B demonstrates the quantitative detection limit for a cell-free BRCA1 promoter methylation assay (cfBRCA1meth).
[0083] FIG.18 illustrates a schematic overview of the CCGA sample cohort.
[0084] FIGs. 19A-B illustrates the distributions of cfBRCA1meth fractions in cfBRCA1meth-positive individuals without cancer.
[0085] FIG. 20 shows the prevalence of cfBRCA1meth in the CCGA cohort. The prevalence of cfBRCA1meth was computed for females and males across the entire non-cancer and cancer cohorts and for each specific tumor types.
[0086] FIGs.21A-B show the prevalence of BRCA1meth across cancer types in the TCGA cohort. BRCA1meth prevalence in the TCGA cohort, with samples subgrouped based on tumor type and sex (A, females; B, males).
[0087] FIGs.22A-B show the distribution of cfBRCA1meth fraction and association with TMeF in female individuals with evidence of cfBRCA1meth. (A) Boxplots show the distribution of cfBRCA1meth fraction for OvCa and TNBC compared to other cancers and non-cancer. (B) Shows the correlation between cfBRCA1meth fraction and tumor methylated fraction (TMeF).
[0088] FIG. 23 shows the correlation between cfBRCA1meth fractions and tumor methylated fractions (TMeFs) in female individuals with cfBRCA1meth-positive cancers other than TNBC and OvCa (‘other cancers’).
[0089] FIG. 24 illustrates a graph depicting % translocation and other considerations regarding a cohort of samples used to evaluate feasibility of distinguishing t(11;14) status in multiple myeloma patients using a methylation based platform.
[0090] FIG. 25 illustrates a graph depicting tumor fraction and other considerations regarding a cohort of samples used to evaluate feasibility of distinguishing t(11;14) status in multiple myeloma patients using a methylation based platform.
[0091] FIG. 26 illustrates a graph showing an example segmentation of DMRs for distinguishing t(11;14) status in multiple myeloma patients using a methylation based platform.
[0092]
[0093] FIG.27 illustrates an example method for estimating single site methylation status.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 13 of 119 DESCRIPTION OF EXAMPLE EMBODIMENTS 1. Definitions
[0094] The term “cell free nucleic acid,” “cell free DNA,” or “cfNA” refers to nucleic acid fragments that circulate in bodily fluids in an individual’s body (e.g., blood, sweat, urine, or saliva) and originate from one or more healthy cells and / or from one or more unhealthy cells (e.g., cancer cells). The term “cell free DNA,” or “cfDNA” refers to deoxyribonucleic acid fragments that circulate in bodily fluids in an individual’s body (e.g., blood, sweat, urine, or saliva). Additionally, cfNAs or cfDNA in an individual’s body may come from other non- human sources.
[0095] The term “genomic nucleic acid,” “genomic DNA,” or “gDNA” refers to nucleic acid molecules or deoxyribonucleic acid molecules obtained from one or more cells. In particular embodiments, gDNA can be extracted from healthy cells (e.g., non-tumor cells) or from tumor cells (e.g., a biopsy sample). In particular embodiments, gDNA can be extracted from a cell derived from a blood cell lineage, such as a white blood cell.
[0096] The term “circulating tumor DNA” or “ctDNA” refers to nucleic acid fragments or deoxyribonucleic acid fragments that originate from tumor cells or other types of cancer cells, and which may be released into a bodily fluid of an individual (e.g., blood, sweat, urine, or saliva) as result of biological processes such as apoptosis or necrosis of dying cells or actively released by viable tumor cells.
[0097] The term “DNA fragment,” “fragment,” or “DNA molecule” may generally refer to any deoxyribonucleic acid fragments, i.e., cfDNA, gDNA, ctDNA, etc.
[0098] The term “anomalous fragment,” “anomalously methylated fragment,” or “fragment with an anomalous methylation pattern” refers to a fragment that has anomalous methylation of CpG sites. Anomalous methylation of a fragment may be determined using probabilistic models to identify unexpectedness of observing a fragment’s methylation pattern in a control group.
[0099] The term “unusual fragment with extreme methylation” or “UFXM” refers to a hypomethylated fragment or a hypermethylated fragment. A hypomethylated fragment and a hypermethylated fragment refers to a fragment with at least some number of CpG sites (e.g., 5) that have over some threshold percentage (e.g., 90%) of methylation or unmethylation, respectively.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 14 of 119
[0100] The term “anomaly score” refers to a score for a CpG site based on a number of anomalous fragments (or, in some embodiments, UFXMs) from a sample overlaps that CpG site. The anomaly score is used in context of featurization of a sample for classification.
[0101] The term “individual” can refer to any living organism, such as a human individual or an animal individual. The term “healthy individual” refers to an individual presumed to not have a cancer or disease.
[0102] The term “subject” refers to an individual whose DNA is being analyzed. A subject may be a test subject whose DNA is be evaluated using whole genome sequencing or a targeted panel as described herein to evaluate whether the person has a disease state (e.g., cancer, type of cancer, or cancer tissue of origin). A subject may also be part of a control group known not to have cancer or another disease. A subject may also be part of a cancer or other disease group known to have cancer or another disease. Control and cancer / disease groups may be used to assist in designing or validating the targeted panel.
[0103] The term “reference sample” refers to a sample obtained from a subject with a known disease state.
[0104] The term “training sample” refers to a sample obtained from a known disease state that can be used to generate sequence reads. Training samples may be applied to probability models to generate features that can be utilized for disease state classification.
[0105] The term “test sample” refers to a sample that may have an unknown disease state.
[0106] The term “sequence read” refers to a nucleotide sequence read from a sample obtained from an individual. Sequence reads may be generated from nucleic acid fragments in the sample. A sequence read can be a collapsed sequence read generated from a plurality of sequence reads derived from a plurality of amplicons from a single original nucleic acid molecule. In particular embodiments, the sequence read can be a deduplicated sequence read. Sequence reads can be obtained through various methods known in the art.
[0107] The term “disease state” refers to presence or non-presence of a disease, a type of disease, and / or a disease tissue of origin. In particular embodiments, the present disclosure provides methods, systems, and non-transitory computer readable medium for detecting cancer (i.e., presence or absence of cancer), a type of cancer, or a cancer tissue of origin.
[0108] The term “tissue of origin” or “TOO” or “Cancer Signal Origin” or “CSO” refers to the organ, organ group, body region or cell type from which a disease state may arise or originate. As an example and not by way of limitation, the identification of a tissue of originACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 15 of 119 or cancer cell type typically allows to identify appropriate next steps to further diagnose, stage, and decide on treatment.
[0109] The term “methylation” as used herein refers to a chemical process by which a methyl group is added to a DNA molecule. Two of DNA’s four bases, cytosine (“C”) and adenine (“A”) can be methylated. As an example and not by way of limitation, a hydrogen atom on the pyrimidine ring of a cytosine base can be converted to a methyl group, forming 5- methylcytosine. Methylation tends to occur at dinucleotides of cytosine and guanine referred to herein as “CpG sites.” In other instances, methylation may occur at a cytosine not part of a CpG site or at another nucleotide that is not cytosine; however, these are rarer occurrences. In this present disclosure, methylation is discussed in reference to CpG sites for the sake of clarity. However, the principles described herein are equally applicable for the detection of methylation in a non-CpG context, including non-cytosine methylation. As an example and not by way of limitation, Adenine methylation has been observed in bacteria, plant and mammalian DNA, although it has received considerably less attention.
[0110] In particular embodiments, the wet laboratory assay used to detect methylation may vary from those described herein as well known in the art. Further, the methylation state vectors may contain elements that are generally vectors of sites where methylation has or has not occurred (even if those sites are not CpG sites specifically). With that substitution, the remainder of the processes described herein are the same, and consequently the inventive concepts described herein are applicable to those other forms of methylation.
[0111] The term “CpG site” refers to a region of a DNA molecule where a cytosine nucleotide is followed by a guanine nucleotide in the linear sequence of bases along its 5′ to 3′ direction. “CpG” is a shorthand for 5′-C-phosphate-G-3′ that is cytosine and guanine separated by only one phosphate group; phosphate links any two nucleotides together in DNA. Cytosines in CpG dinucleotides can be methylated to form 5-methylcytosine.
[0112] The term “methylation site” refers to a single site of a DNA molecule where a methyl group can be added. “CpG” sites are the most common methylation site, but methylation sites are not limited to CpG sites. As an example and not by way of limitation, DNA methylation may occur in cytosines in CHG and CHH, where H is adenine, cytosine or thymine. Cytosine methylation in the form of 5-hydroxymethylcytosine may also assessed (see, e.g., WO 2010 / 037001 and WO 2011 / 127136, which are incorporated herein by reference), and features thereof, using the methods and procedures disclosed herein. The term “hypomethylated” or “hypermethylated” refers to a methylation status of a DNA moleculeACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 16 of 119 containing multiple CpG sites (e.g., more than 3, 4, 5, 6, 7, 8, 9, 10, etc.) where a high percentage of the CpG sites (e.g., more than 80%, 85%, 90%, or 95%, or any other percentage within the range of 50%-100%) are unmethylated or methylated, respectively.
[0113] As used herein, the term “about” or “approximately” can mean within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which can depend in part on how the value is measured or determined, e.g., the limitations of the measurement system. As an example and not by way of limitation, “about” can mean within 1 or more than 1 standard deviation, per the practice in the art. “About” can mean a range of ±20%, ±10%, ±5%, or ±1% of a given value. The term “about” or “approximately” can mean within an order of magnitude, within 5-fold, or within 2-fold, of a value. Where particular values are described in the application and claims, unless otherwise stated the term “about” meaning within an acceptable error range for the particular value should be assumed. The term “about” can have the meaning as commonly understood by one of ordinary skill in the art. The term “about” can refer to ±10%. The term “about” can refer to ±5%.
[0114] As used herein, the term “biological sample,” “patient sample,” or “sample” refers to any sample taken from a subject, which can reflect a biological state associated with the subject, and that includes cell-free DNA. Examples of biological samples include, but are not limited to, blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid of the subject. A biological sample can include any tissue or material derived from a living or dead subject. A biological sample can be a cell-free sample. A biological sample can comprise a nucleic acid (e.g., DNA or RNA) or a fragment thereof. The term “nucleic acid” can refer to deoxyribonucleic acid (DNA), ribonucleic acid (RNA) or any hybrid or fragment thereof. The nucleic acid in the sample can be a cell-free nucleic acid. A sample can be a liquid sample or a solid sample (e.g., a cell or tissue sample). A biological sample can be a bodily fluid, such as blood, plasma, serum, urine, vaginal fluid, fluid from a hydrocele (e.g., of the testis), vaginal flushing fluids, pleural fluid, ascitic fluid, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, discharge fluid from the nipple, aspiration fluid from different parts of the body (e.g., thyroid, breast), etc. A biological sample can be a stool sample. In particular embodiments, the majority of DNA in a biological sample that has been enriched for cell-free DNA (e.g., a plasma sample obtained via a centrifugation protocol) can be cell-free (e.g., greater than 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA can be cell-free). A biological sample can be treated to physically disrupt tissue or cell structure (e.g., centrifugation and / or cell lysis), thus releasingACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 17 of 119 intracellular components into a solution which can further contain enzymes, buffers, salts, detergents, and the like which can be used to prepare the sample for analysis.
[0115] As used herein, the terms “control,” “control sample,” “reference,” “reference sample,” “normal,” and “normal sample” describe a sample from a subject that does not have a particular condition, or is otherwise healthy. In an example, a method as disclosed herein can be performed on a subject having a tumor, where the reference sample is a sample taken from a healthy tissue of the subject. A reference sample can be obtained from the subject, or from a database. The reference can be, e.g., a reference genome that is used to map nucleic acid fragment sequences obtained from sequencing a sample from the subject. A reference genome can refer to a haploid or diploid genome to which nucleic acid fragment sequences from the biological sample and a constitutional sample can be aligned and compared. An example of a constitutional sample can be DNA of white blood cells obtained from the subject. For a haploid genome, there can be only one nucleotide at each locus. For a diploid genome, heterozygous loci can be identified; each heterozygous locus can have two alleles, where either allele can allow a match for alignment to the locus.
[0116] As used herein, the term “cancer” or “tumor” refers to an abnormal mass of tissue in which the growth of the mass surpasses and is not coordinated with the growth of normal tissue.
[0117] As used herein, the phrase “healthy,” refers to a subject possessing good health. A healthy subject can demonstrate an absence of any malignant or non-malignant disease. A “healthy individual” can have other diseases or conditions, unrelated to the condition being assayed, which can normally not be considered “healthy.”
[0118] As used herein, the term “methylation” refers to a modification of deoxyribonucleic acid (DNA) where a hydrogen atom on the pyrimidine ring of a cytosine base is converted to a methyl group, forming 5-methylcytosine. In particular, methylation tends to occur at dinucleotides of cytosine and guanine referred to herein as “CpG sites.” In other instances, methylation may occur at a cytosine not part of a CpG site or at another nucleotide that’s not cytosine; however, these are rarer occurrences. Anomalous cfDNA methylation can be identified as hypermethylation or hypomethylation, both of which may be indicative of cancer status. DNA methylation anomalies (compared to healthy controls) can cause different effects, which may contribute to cancer. The principles described herein are equally applicable for the detection of methylation in a CpG context and non-CpG context, including non-cytosine methylation. Further, the methylation state vectors may contain elements that are generallyACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 18 of 119 vectors of sites where methylation has or has not occurred (even if those sites are not CpG sites specifically).
[0119] As used interchangeably herein, the term “methylation fragment” or “nucleic acid methylation fragment” refers to a sequence of methylation states for each CpG site in a plurality of CpG sites, determined by a methylation sequencing of nucleic acids (e.g., a nucleic acid molecule and / or a nucleic acid fragment). In a methylation fragment, a location and methylation state for each CpG site in the nucleic acid fragment is determined based on the alignment of the sequence reads (e.g., obtained from sequencing of the nucleic acids) to a reference genome. A nucleic acid methylation fragment comprises a methylation state of each CpG site in a plurality of CpG sites (e.g., a methylation state vector), which specifies the location of the nucleic acid fragment in a reference genome (e.g., as specified by the position of the first CpG site in the nucleic acid fragment using a CpG index, or another similar metric) and the number of CpG sites in the nucleic acid fragment. Alignment of a sequence read to a reference genome, based on a methylation sequencing of a nucleic acid molecule, can be performed using a CpG index. As used herein, the term “CpG index” refers to a list of each CpG site in the plurality of CpG sites (e.g., CpG 1, CpG 2, CpG 3, etc.) in a reference genome, such as a human reference genome, which can be in electronic format. The CpG index further comprises a corresponding genomic location, in the corresponding reference genome, for each respective CpG site in the CpG index. Each CpG site in each respective nucleic acid methylation fragment is thus indexed to a specific location in the respective reference genome, which can be determined using the CpG index.
[0120] As used herein, the term “methylation variant” refers to a distinct methylation pattern in a set of k (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 25, etc.) contiguous CpG sites. The methylation variants may be used to distinguish DNA derived from a specific cancer type from non-cancer plasma. A given methylation variant has a single defined “variant state” which is a single specific and detectable pattern of methylation states. A “reference state” is any pattern of methylation states that is not a variant state. As one or more examples, a set of 5 contiguous CpGs can have a variant state of 5 methylated CpGs (e.g., MMMMM, where “M” denotes methylated), another variant state of 5 unmethylated CpGs (e.g., UUUUU, where “U” denotes unmethylated), or a third variant state including a mixture of methylation states (e.g., UMMMU). The reference states can include all other patterns not deemed to be one of the variant states.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 19 of 119
[0121] As used herein, the term “true positive” (TP) refers to a subject having a condition. “True positive” can refer to a subject that has a tumor, a cancer, a pre-cancerous condition (e.g., a pre-cancerous lesion), a localized or a metastasized cancer, or a non-malignant disease. “True positive” can refer to a subject having a condition and is identified as having the condition by an assay or method of the present disclosure. As used herein, the term “true negative” (TN) refers to a subject that does not have a condition or does not have a detectable condition. True negative can refer to a subject that does not have a disease or a detectable disease, such as a tumor, a cancer, a pre-cancerous condition (e.g., a pre-cancerous lesion), a localized or a metastasized cancer, a non-malignant disease, or a subject that is otherwise healthy. True negative can refer to a subject that does not have a condition or does not have a detectable condition, or is identified as not having the condition by an assay or method of the present disclosure.
[0122] As used herein, the term “reference genome” refers to any particular known, sequenced or characterized genome, whether partial or complete, of any organism or virus that may be used to reference identified sequences from a subject. Exemplary reference genomes used for human subjects as well as many other organisms are provided in the on-line genome browser hosted by the National Center for Biotechnology Information (“NCBI”) or the University of California, Santa Cruz (UCSC). A “genome” refers to the complete genetic information of an organism or virus, expressed in nucleic acid sequences. As used herein, a reference sequence or reference genome often is an assembled or partially assembled genomic sequence from an individual or multiple individuals. In particular embodiments, a reference genome is an assembled or partially assembled genomic sequence from one or more human individuals. The reference genome can be viewed as a representative example of a species’ set of genes. In particular embodiments, a reference genome comprises sequences assigned to chromosomes. Exemplary human reference genomes include but are not limited to NCBI build 34 (UCSC equivalent: hg16), NCBI build 35 (UCSC equivalent: hg17), NCBI build 36.1 (UCSC equivalent: hg18), GRCh37 (UCSC equivalent: hg19), and GRCh38 (UCSC equivalent: hg38).
[0123] As used herein, the term “sequence reads” or “reads” refers to nucleotide sequences produced by any sequencing process described herein or known in the art. Reads can be generated from one end of nucleic acid fragments (“single-end reads”), and sometimes are generated from both ends of nucleic acids (e.g., paired-end reads, double-end reads). In particular embodiments, sequence reads (e.g., single-end or paired-end reads) can be generatedACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 20 of 119 from one or both strands of a targeted nucleic acid fragment. The length of the sequence read is often associated with the particular sequencing technology. High-throughput methods, for example, provide sequence reads that can vary in size from tens to hundreds of base pairs (bp). In particular embodiments, the sequence reads are of a mean, median or average length of about 15 bp to 900 bp long (e.g., about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130, about 140 bp, about 150 bp, about 200 bp, about 450 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, or about 500 bp. In particular embodiments, the sequence reads are of a mean, median or average length of about 1000 bp, 2000 bp, 5000 bp, 10,000 bp, or 50,000 bp or more. Nanopore sequencing, for example, can provide sequence reads that can vary in size from tens to hundreds to thousands of base pairs. Illumina parallel sequencing can provide sequence reads that do not vary as much, for example, most of the sequence reads can be smaller than 200 bp. A sequence read (or sequencing read) can refer to sequence information corresponding to a nucleic acid molecule (e.g., a string of nucleotides). As an example and not by way of limitation, a sequence read can correspond to a string of nucleotides (e.g., about 20 to about 150) from part of a nucleic acid fragment, can correspond to a string of nucleotides at one or both ends of a nucleic acid fragment, or can correspond to nucleotides of the entire nucleic acid fragment. A sequence read can be obtained in a variety of ways, e.g., using sequencing techniques or using probes, e.g., in hybridization arrays or capture probes, or amplification techniques, such as the polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification.
[0124] As used herein, the terms “sequencing” and the like as used herein refers generally to any and all biochemical processes that may be used to determine the order of biological macromolecules such as nucleic acids or proteins. As an example and not by way of limitation, sequencing data can include all or a portion of the nucleotide bases in a nucleic acid molecule such as a DNA fragment.
[0125] As used herein, the term “sequencing depth,” is interchangeably used with the term “coverage” and refers to the number of times a locus is covered by a consensus sequence read corresponding to a unique nucleic acid target molecule aligned to the locus; e.g., the sequencing depth is equal to the number of unique nucleic acid target molecules covering the locus. The locus can be as small as a nucleotide, or as large as a chromosome arm, or as large as an entire genome. Sequencing depth can be expressed as “Yx”, e.g., 50x, 100x, etc., where “Y” refersACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 21 of 119 to the number of times a locus is covered with a sequence corresponding to a nucleic acid target; e.g., the number of times independent sequence information is obtained covering the particular locus. In particular embodiments, the sequencing depth corresponds to the number of genomes that have been sequenced. Sequencing depth can also be applied to multiple loci, or the whole genome, in which case Y can refer to the mean or average number of times a locus or a haploid genome, or a whole genome, respectively, is sequenced. When a mean depth is quoted, the actual depth for different loci included in the dataset can span over a range of values. Ultra-deep sequencing can refer to at least 100x in sequencing depth at a locus.
[0126] As used herein, the term “sensitivity” or “true positive rate” (TPR) refers to the number of true positives divided by the sum of the number of true positives and false negatives. Sensitivity can characterize the ability of an assay or method to correctly identify a proportion of the population that truly has a condition. As an example and not by way of limitation, sensitivity can characterize the ability of a method to correctly identify the number of subjects within a population having cancer. In another example, sensitivity can characterize the ability of a method to correctly identify the one or more markers indicative of cancer.
[0127] As used herein, the term “specificity” or “true negative rate” (TNR) refers to the number of true negatives divided by the sum of the number of true negatives and false positives. Specificity can characterize the ability of an assay or method to correctly identify a proportion of the population that truly does not have a condition. As an example and not by way of limitation, specificity can characterize the ability of a method to correctly identify the number of subjects within a population not having cancer. In another example, specificity characterizes the ability of a method to correctly identify one or more markers indicative of cancer.
[0128] As used herein, the term “subject” refers to any living or non-living organism, including but not limited to a human (e.g., a male human, female human, fetus, pregnant female, child, or the like), a non-human animal, a plant, a bacterium, a fungus, or a protist. Any human or non-human animal can serve as a subject, including but not limited to mammal, reptile, avian, amphibian, fish, ungulate, ruminant, bovine (e.g., cattle), equine (e.g., horse), caprine and ovine (e.g., sheep, goat), swine (e.g., pig), camelid (e.g., camel, llama, alpaca), monkey, ape (e.g., gorilla, chimpanzee), ursid (e.g., bear), poultry, dog, cat, mouse, rat, fish, dolphin, whale, and shark. In particular embodiments, a subject is a male or female of any stage (e.g., a man, a woman or a child). A subject from whom a sample is taken, or is treated by any of the methods or compositions described herein can be of any age and can be an adult, infant or child.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 22 of 119
[0129] As used herein, the term “tissue” can correspond to a group of cells that group together as a functional unit. More than one type of cell can be found in a single tissue. Different types of tissue may consist of different types of cells (e.g., hepatocytes, alveolar cells or blood cells), but also can correspond to tissue from different organisms (mother vs. fetus) or to healthy cells vs. tumor cells. The term “tissue” can generally refer to any group of cells found in the human body (e.g., heart tissue, lung tissue, kidney tissue, nasopharyngeal tissue, oropharyngeal tissue). In some aspects, the term “tissue” or “tissue type” or “originating tissue type” can be used to refer to a tissue from which a cell-free nucleic acid molecule or fragment originates. In one example, viral nucleic acid fragments can be derived from blood tissue. In another example, viral nucleic acid fragments can be derived from tumor tissue.
[0130] As used herein, the term “genomic” refers to a characteristic of the genome of an organism. Examples of genomic characteristics include, but are not limited to, those relating to the primary nucleic acid sequence of all or a portion of the genome (e.g., the presence or absence of a nucleotide polymorphism, indel, sequence rearrangement, mutational frequency, etc.), the copy number of one or more particular nucleotide sequences within the genome (e.g., copy number, allele frequency fractions, single chromosome or entire genome ploidy, etc.), the epigenetic status of all or a portion of the genome (e.g., covalent nucleic acid modifications such as methylation, histone modifications, nucleosome positioning, etc.), the expression profile of the organism’s genome (e.g., gene expression levels, isotype expression levels, gene expression ratios, etc.).
[0131] The terminology used herein is for the purpose of describing particular cases only and is not intended to be limiting. As used herein, the singular forms “a,” “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, to the extent that the terms “including,” “includes,” “having,” “has,” “with,” or variants thereof are used in either the detailed description and / or the claims, such terms are intended to be inclusive in a manner similar to the term “comprising”. 2. Overview
[0132] Early detection and classification of cancer is an important technology. Being able to detect cancer before it becomes symptomatic is beneficial to all parties involved, including patients, doctors, and loved ones. For patients, early cancer detection allows them a greater chance of a beneficial outcome; for doctors, early cancer detection allows more pathways ofACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 23 of 119 treatment that may lead to a beneficial outcome; for loved ones, early cancer detection increases the likelihood of not losing friends and family to the disease.
[0133] Recently, early cancer detection technology has progressed towards analyzing genetic fragments (e.g., DNA), for example, in a person’s blood to determine if any of those genetic fragments originate from cancer cells. These new techniques allow doctors to identify a cancer presence in a patient that may not be detectable otherwise, e.g., in conventional screening processes. For instance, consider the example of a person at high risk for breast cancer. Traditionally, this person will regularly visit their doctor for a mammogram, which creates an image of their breast tissue (e.g., taking x-ray images) that a doctor uses to identify cancerous tissue. Unfortunately, with even the highest resolution mammograms, doctors are only able to identify tumors once they are approximately a millimeter in size. This means that the cancer has been present for some time in the person and has gone undiagnosed and untreated. Visual determinations like this are typical for most cancers—that is, only identifiable once it has grown to a sufficient size to be detected by some sort of imaging technology.
[0134] Cancer detection using analysis of genetic fragments in a patient’s, e.g., blood alleviates this issue. To illustrate, cancer cells will start sloughing DNA fragments into a person’s bloodstream as soon as they form. This occurs when there are very few of the cancer cells, and before they would be visible with imaging techniques. With the appropriate methods, therefore, a system that analyzes DNA fragments in the bloodstream could identify cancer presence in a person based on sloughed cancer DNA fragments, and, more importantly, the system could do so before the cancer is identifiable using more traditional cancer detection techniques.
[0135] Cancer detection based on the analysis of DNA fragments is enabled by next- generation sequencing (“NGS”) techniques. NGS, broadly, is a group of technologies that allows for high throughput sequencing of genetic material. As discussed in greater detail herein, NGS largely consists of (1) sample preparation, (2) DNA sequencing, and (3) data analysis. Sample preparation is the laboratory methods necessary to prepare DNA fragments for sequencing, sequencing is the process of reading the ordered nucleotides in the samples, and data analysis is processing and analyzing the genetic information in the sequencing data to identify cancer presence.
[0136] While these steps of NGS may help enable early cancer detection, they also introduce their own complex, detrimental problems to cancer detection and, therefore, any improvements to sample preparation, DNA sequencing, and / or data analysis, including the pre-ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 24 of 119 processing, algorithmic processing, and summary or presentation of predications or conclusions, results in an improvement to cancer detection technologies and early cancer detection more generally.
[0137] To illustrate, as an example, problems introduced in (1) sample preparation include DNA sample quality, sample contamination, fragmentation bias, and accurate indexing. Remedying these problems would yield better genetic data for cancer detection.
[0138] Similarly, problems introduced in (2) sequencing include, for example, errors in accurate transcribing of fragments (e.g., reading an “A” instead of a “C”, etc.), incorrect or difficult fragment assembly and overlap, disparate coverage uniformity, sequencing depth vs. cost vs. specificity, and insufficient sequencing length. Again, remedying any of these problems would yield improved genetic data for cancer detection.
[0139] The problems in (3) data analysis are the most daunting and complex. The introduced challenges stem from the vast amounts of data created by NGS sequencing techniques. Sequencing data for a single sample can be on the order of hundreds of thousands (up to millions) of sequence reads, amounting to terabytes of data. Multiply that by the thousands (up to tens of thousands) of samples which are collected for use in training of the analytical models. Effectively and efficiently analyzing that amount of data is both procedurally and computationally demanding. For instance, analyzing NGS sequencing involves several baseline processing steps such as, e.g., aligning reads to one another, aligning and mapping reads to a reference genome, de-duping duplicative reads, detecting contamination of a sample, identifying and calling variant genes, identifying and calling abnormally methylated genes, generating functional annotations, etc. Performing any of these processes on terabytes of genetic data is computationally expensive for even the most powerful of computer architectures, and completely impossible for a normal human mind. Additionally, with the genetic sequencing data derived from the error-prone processes of sample preparation and sequence reading, large portions of the resulting genetic data may be low-quality or unusable for cancer identification. As an example and not by way of limitation, large amounts of the genetic data may include contaminated samples, transcription errors, mismatched regions, overrepresented regions, etc. and may be unsuitable for high accuracy cancer detection. Identifying and accounting for low quality genetic data across the vast amount of genetic data obtained from NGS sequencing is also procedurally and computationally rigorous to accomplish and is also not practically performable by a human mind. Overall, any process created that leads to more efficient processing of large array sequencing data would be anACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 25 of 119 improvement to cancer detection using NGS sequencing. Moreover, such processes were crafted as a solution to the various hurdles created in NGS sequences, and as such are non- routine and unconventional activity in the technical field of endeavor.
[0140] Particularly, under (3) data analysis, accurate identification of anomalous DNA from NGS data to identify a cancer presence is also a difficult task-at-hand. To be effective, algorithms are sought to compensate for, e.g., errors generated by sample preparation and sequencing, and to overcome the large-scale data analysis problems accompanying NGS techniques. That is, designing a machine learning model or models, or other computational processing algorithms, that enable early cancer detection based on next generation sequencing techniques must be configured to account for the problems that those techniques create. Some of those techniques and models are discussed hereinbelow and particular improvements to state-of-the-art techniques and models are further discussed. Furthermore, such techniques are non-routine and unconventional activity in the technical field of endeavor.
[0141] One of the problems in creating and appropriately applying a cancer detection model is, as described above, the vast amount of sequencing data to which the model may be applied. Within that sequencing data is a vast array of “cancer signal” and “separate signal,” where cancer signal indicates sequencing data that is indicative of cancer presence or cancer non-presence in a test subject and separate signal indicates sequencing data that is indicative of, e.g., an additional biological or clinical characteristic of the test subject (e.g., age, sex, smoker, etc.). To compound this issue, some of the separate signal may appear to share similar characteristics to the cancer signal, or some of the cancer signal may appear to share similar characteristics to the separate signal. Still further, some of the sequencing data may be both cancer signal and separate signal. As an example and not by way of limitation, the methylation patterns associated with increased age may appear similar to the methylation patterns for increased age, or, in a similar view, a specific methylation pattern may be indicative of both cancer and increased age.
[0142] This interplay of signal(s) in the sequencing data is troublesome because it may lead to error in cancer determinations. For instance, a cancer classifier that is inappropriately trained may determine a person has subject signals based on analysis of their cfDNA because methylation of some of the cfDNA appears “chronologically old and cancerous,” when, in fact, the subject is merely chronologically old. As such, great care should be taken in selecting sequencing data for training a cancer classifier when the sequencing data can indicate several closely related biological processes or clinical characteristics. Any methods that improveACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 26 of 119 cancer determination in these situations represent an improvement to the field of early cancer detection.
[0143] Additionally, carefully curating a collection of data which accurately indicates cancer and / or separate biological and / or clinical characteristics alleviates the complexity and corresponding computational expense of executing a machine-learned model to indicate or predicted a presence of call cancer (i.e., “call” cancer) in a sample. In effect, by carefully selecting genomic sites for processing by the model that are “strongly” associated with cancer and / or a biological or clinical characteristic, the computational load on a model will be reduced yielding higher speed and efficiency, which represents an improvement to the field of cancer determination. As an example and not by way of limitation, consider an example machine- learned model trained to identify cancer based on methylation described above. In this instance, however, each indicative feature is also associated with a “strength” of cancer indication for that site, or a “strength” of chronological age indication for that site. That is, abnormal methylation at the first site may strongly indicate cancer, while abnormal methylation at a second site may strongly, or weakly indicate chronological age, etc. In this case, having the machine-learned model process genomic sites that are merely “weakly” indicative of cancer introduces computational expense without providing a corresponding benefit to the accuracy or specificity in cancer determination.
[0144] To provide a quantitative example, applying the machine-learned model to 100 weakly indicative sites may not provide as much benefit as applying the machine-learned model to 1 strongly indicative site. As such, selecting the appropriate sites for feature sets for processing by a machine-learned model to identify cancer presence and / or biological or clinical characteristics greatly reduces processing cost without greatly reducing the accuracy or specificity of the model. In effect, it may lessen the analytical load from, e.g., millions of genomic sites to tens of thousands of genomic sites without sacrificing model accuracy. More succinctly, reducing the analytical load from millions of genomic sites to tens of thousands of sites, reduces the processing load and processing time by several orders of magnitude. This reduction enables faster cancer detections and more importantly, frees up computer resources for other models and classifications (e.g., processing on additional samples), improves the performance of the computer implementing the model, reduces the monetary cost of such systems, and improves the fields of public health, medicine, diagnostics, treatment, etc. by detecting cancer earlier than even possible by conventional methods.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 27 of 119
[0145] Another problem associated with the vast amounts of data generated by NGS sequencing is appropriately training a machine-learned model to identify cancer within the large amount of data. For instance, a machine-learned model may be trained to identify cancer by comparing a feature vector to genomic data. The “features” in the feature vector, as set forth below, may be any genomic site with a sufficient depth of abnormally methylated genomic locations that correspond to cancer presence. When building feature vectors across an entire genome, this can lead to, typically, tens of thousands of features, and, as laid out above, some of those features may be more indicative of cancer presence than others. With this context, selecting which features and corresponding genomic data to use in training a machine learned model is difficult. The machine-learned model should be trained and configured to accurately identify cancer presence, but the resulting model should not be overly expensive computationally. In other words, appropriately selecting data and features for training a machine-learned model improves early cancer-detection. 2.1. Cancer Classification Workflow
[0146] FIG. 1 illustrates an example overall workflow 100 of cancer classification of a sample. The workflow 100 is by one or more entities, e.g., including a healthcare provider, a sequencing device, an analytics system, etc. Objectives of the workflow include detecting and / or monitoring cancer in individuals. From a healthcare standpoint, the workflow 100 can serve to supplement other existing cancer diagnostic tools. The workflow 100 may serve to provide early cancer detection and / or routine cancer monitoring to better inform treatment plans for individuals diagnosed with cancer. The overall workflow 100 may include additional / fewer steps than those shown in FIG.1.
[0147] A healthcare provider performs sample collection 110. An individual to undergo cancer classification visits their healthcare provider. The healthcare provider collects the sample for performing cancer classification. Examples of biological samples include, but are not limited to, tissue biopsy, blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid of the subject. The sample includes genetic material belonging to the individual, which may be extracted and sequenced for cancer classification. Once the sample is collected, the sample is provided to a sequencing device. Along with the sample, the healthcare provider may collect other information relating to the individual, e.g., biological sex, age, ethnicity, smoking status, any prior diagnoses, etc.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 28 of 119
[0148] A sequencing device performs sample sequencing 120. A lab clinician may perform one or more processing steps to the sample in preparation of sequencing. Once prepared, the clinician loads the sample in the sequencing device. An example of devices utilized in sequencing is further described in conjunction with FIGS. 2A-2B. The sequencing device generally extracts and isolates fragments of nucleic acid that are sequenced to determine a sequence of nucleobases corresponding to the fragments. Sequencing may also include amplification of nucleic material. Different sequencing processes include Sanger sequencing, fragment analysis, and next-generation sequencing. Sequencing may be whole-genome sequencing or targeted sequencing with a target panel. In the context of DNA methylation, bisulfite sequencing (e.g., further described in FIGS.4A-4B) can determine methylation status through bisulfite conversion of unmethylated cytosines at CpG sites. Sample sequencing 120 yields sequences for a plurality of nucleic acid fragments in the sample. In one or more embodiments, the sequences may include methylation state vectors, wherein each methylation state vector describes the methylation statuses for CpG sites on a fragment.
[0149] An analytics system 200 performs pre-analysis processing 130. An example analytics system 200 is described in FIG.2B. Pre-analysis processing 130 may include, but not limited to, de-duplication of sequence reads, determining metrics relating to coverage, determining whether the sample is contaminated, removal of contaminated fragments, calling sequencing error, etc.
[0150] The analytics system 200 performs one or more analyses 140. The analyses are statistical analyses or application of one or more trained models to predict at least a cancer status of the individual from whom the sample is derived. Different genetic features may be evaluated and considered, such as methylation of CpG sites, single nucleotide polymorphisms (SNPs), insertions or deletions (indels), other types of genetic mutation, etc. In the context of methylation, analyses 140 may include tissue purity assessment 142 (e.g., further described in FIGS. 6A-6B), feature extraction 144, and applying a cancer classifier 146 to determine a cancer prediction (e.g., further described in FIGS. 8A-8B). Tissue purity assessment 142 involves applying a mixture model to deconvolute proportions of tissue components contributing DNA fragments to the sample. In general, tissue purity assessment 142 may be used to determine what proportion of methylation sequence reads for a sample were shed from cancerous tissue compared to a proportion of methylation sequence reads from non-cancer cells. Tissue purity assessment 142 may be particularly useful in deconvolving heterogeneous tumors that comprise multiple clonal populations with potentially distinct genetic signatures.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 29 of 119 The mixture model may also be applied to quantify cancer signal or grade cancer status based on the determined proportions. The cancer classifier 146 inputs the extracted features to determine a cancer prediction. The cancer prediction may be a label or a value. The label may indicate a particular cancer state, e.g., binary labels can indicate presence or absence of cancer, multiclass labels can indicate one or more cancer types from a plurality of cancer types that are screened for, cancer stage, etc. The value may indicate a likelihood of a particular cancer state, e.g., a likelihood of cancer, and / or a likelihood of a particular cancer type.
[0151] The analytics system 200 returns the prediction 150 to the healthcare provider. The prediction 150 may include binary prediction of presence or absence of cancer, a particular cancer type, cancer stage, tissue proportions, etc. The healthcare provider may establish or adjust a treatment plan based on the returned prediction 150. Optimization of treatment is further described herein. 2.2. Methylation Overview
[0152] According to the present description, cfDNA fragments from an individual are treated, for example by converting unmethylated cytosines to uracils, sequenced and the sequence reads compared to a reference genome to identify the methylation states at specific CpG sites within the DNA fragments. Each CpG site may be methylated or unmethylated. Identification of anomalously methylated fragments, in comparison to healthy individuals, may provide insight into a subject’s cancer status. As is well known in the art, DNA methylation anomalies (compared to healthy controls) can cause different effects, which may contribute to cancer. Various challenges arise in the identification of anomalously methylated cfDNA fragments. First off, determining a DNA fragment to be anomalously methylated can hold weight in comparison with a group of control individuals, such that if the control group is small in number, the determination loses confidence due to statistical variability within the smaller size of the control group. Additionally, among a group of control individuals, methylation status can vary which can be difficult to account for when determining a subject’s DNA fragments to be anomalously methylated. On another note, methylation of a cytosine at a CpG site can causally influence methylation at a subsequent CpG site. To encapsulate this dependency can be another challenge in itself.
[0153] Methylation can typically occur in deoxyribonucleic acid (DNA) when a hydrogen atom on the pyrimidine ring of a cytosine base is converted to a methyl group, forming 5- methylcytosine. In particular, methylation can occur at dinucleotides of cytosine and guanineACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 30 of 119 referred to herein as “CpG sites”. In other instances, methylation may occur at a cytosine not part of a CpG site or at another nucleotide that is not cytosine; however, these are rarer occurrences. In this present disclosure, methylation is discussed in reference to CpG sites for the sake of clarity. Anomalous DNA methylation can be identified as hypermethylation or hypomethylation, both of which may be indicative of cancer status. Throughout this disclosure, hypermethylation and hypomethylation can be characterized for a DNA fragment, if the DNA fragment comprises more than a threshold number of CpG sites with more than a threshold percentage of those CpG sites being methylated or unmethylated.
[0154] The principles described herein can be equally applicable for the detection of methylation in a non-CpG context, including non-cytosine methylation. In particular embodiments, the wet laboratory assay used to detect methylation may vary from those described herein. Further, the methylation state vectors discussed herein may contain elements that are generally sites where methylation has or has not occurred (even if those sites are not CpG sites specifically). With that substitution, the remainder of the processes described herein can be the same, and consequently the inventive concepts described herein can be applicable to those other forms of methylation. 2.3. Example Analytics System
[0155] FIG. 2A illustrates an example flowchart of devices for sequencing nucleic acid samples. This illustrative flowchart includes devices such as a sequencer 220 and an analytics system 200. The sequencer 220 and the analytics system 200 may work in tandem to perform one or more steps in the processes.
[0156] In particular embodiments, the sequencer 220 receives an enriched nucleic acid sample 210. As shown in FIG. 2A, the sequencer 220 can include a graphical user interface 225 that enables user interactions with particular tasks (e.g., initiate sequencing or terminate sequencing) as well as one more loading stations 230 for loading a sequencing cartridge including the enriched fragment samples and / or for loading necessary buffers for performing the sequencing assays. Therefore, once a user of the sequencer 220 has provided the necessary reagents and sequencing cartridge to the loading station 230 of the sequencer 220, the user can initiate sequencing by interacting with the graphical user interface 225 of the sequencer 220. Once initiated, the sequencer 220 performs the sequencing and outputs the sequence reads of the enriched fragments from the nucleic acid sample 210.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 31 of 119
[0157] In particular embodiments, the sequencer 220 is communicatively coupled with the analytics system 200. The analytics system 200 includes some number of computing devices used for processing the sequence reads for various applications such as assessing methylation status at one or more CpG sites, variant calling or quality control. The sequencer 220 may provide the sequence reads in a BAM file format to the analytics system 200. The analytics system 200 can be communicatively coupled to the sequencer 220 through a wireless, wired, or a combination of wireless and wired communication technologies. Generally, the analytics system 200 is configured with a processor and non-transitory computer-readable storage medium storing computer instructions that, when executed by the processor, cause the processor to process the sequence reads or to perform one or more steps of any of the methods or processes disclosed herein.
[0158] In particular embodiments, the sequence reads may be aligned to a reference genome using known methods in the art to determine alignment position information. Alignment position may generally describe a beginning position and an end position of a region in the reference genome that corresponds to a beginning nucleotide base and an end nucleotide base of a given sequence read. Corresponding to methylation sequencing, the alignment position information may be generalized to indicate a first CpG site and a last CpG site included in the sequence read according to the alignment to the reference genome. The alignment position information may further indicate methylation statuses and locations of all CpG sites in a given sequence read. A region in the reference genome may be associated with a gene or a segment of a gene; as such, the analytics system 200 may label a sequence read with one or more genes that align to the sequence read. In particular embodiments, fragment length (or size) can be determined from the beginning and end positions.
[0159] In particular embodiments, when a paired-end sequencing process is used, a sequence read is comprised of a read pair denoted as R_1 and R_2. As an example and not by way of limitation, the first read R_1 may be sequenced from a first end of a double-stranded DNA (dsDNA) molecule whereas the second read R_2 may be sequenced from the second end of the double-stranded DNA (dsDNA). Therefore, nucleotide base pairs of the first read R_1 and second read R_2 may be aligned consistently (e.g., in opposite orientations) with nucleotide bases of the reference genome. Alignment position information derived from the read pair R_1 and R_2 may include a beginning position in the reference genome that corresponds to an end of a first read (e.g., R_1) and an end position in the reference genome that corresponds to an end of a second read (e.g., R_2). In other words, the beginning positionACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 32 of 119 and end position in the reference genome can represent the likely location within the reference genome to which the nucleic acid fragment corresponds. An output file having SAM (sequence alignment map) format or BAM (binary) format may be generated and output for further analysis.
[0160] FIG. 2B illustrates an example block diagram of an analytics system 200 for processing DNA samples. The analytics system 200 implements one or more computing devices for use in analyzing DNA samples. The analytics system 200 includes a sequence processor 220, sequence database 245, model database 255, models 250, parameter database 265, and score engine 260. In particular embodiments, the analytics system 200 performs some or all of the processes described throughout this disclosure.
[0161] The sequence processor 220 generates methylation state vectors for fragments from a sample. At each CpG site on a fragment, the sequence processor 220 generates a methylation state vector for each fragment specifying a location of the fragment in the reference genome, a number of CpG sites in the fragment, and the methylation state of each CpG site in the fragment whether methylated, unmethylated, or indeterminate via the process of FIG.2A. The sequence processor 220 may store methylation state vectors for fragments in the sequence database 245. Data in the sequence database 245 may be organized such that the methylation state vectors from a sample are associated to one another.
[0162] Further, multiple different models 250 may be stored in the model database 255 or retrieved for use with test samples. Generally, a model receives an input and generates an output based on a function that operates on the input according to one or more parameters. In one example, a model is a trained mixture model for deconvolving component proportions based on a sample’s methylation signature. The mixture model may comprise various sub- models each with its own parameters and function(s). In another example, a model is a trained cancer classifier for determining a cancer prediction for a test sample using a feature vector derived from anomalous fragments. The training and use of the cancer classifier will be further discussed in conjunction with Section IV. Cancer Classifier for Determining Cancer. The analytics system 200 may train the one or more models 250 and store various trained parameters in the parameter database 265. The analytics system 200 stores the models 250 along with functions in the model database 255.
[0163] During inference, the score engine 260 uses the one or more models 250 to return outputs. The score engine 260 accesses the models 250 in the model database 255 along with trained parameters from the parameter database 265. According to each model, the score engineACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 33 of 119 receives an appropriate input for the model and calculates an output based on the received input, the parameters, and a function of each model relating the input and the output. In some use cases, the score engine 260 further calculates metrics correlating to a confidence in the calculated outputs from the model. In other use cases, the score engine 260 calculates other intermediary values for use in the model. 3. Identifying Informative Fragments
[0164] In particular embodiments, the analytics system 200 determines informative fragments for a sample using the sample’s methylation state vectors. As an example and not by way of limitation, for each nucleic acid molecule or fragment in a sample, the analytics system 200 determines whether the nucleic acid molecule or fragment is an informative molecule or fragment (via analysis of sequence reads derived therefrom), relative to an expected methylation state vector from a healthy sample using the methylation state vector corresponding to the nucleic acid molecule. In particular embodiments, the analytics system 200 calculates a p-value score for each methylation state vector describing a probability of observing that methylation state vector or other methylation state vectors even less probable in the healthy control group (as described, for example, in U.S. Pat. Appl. Pub. No. 2019 / 0287652, which is incorporated herein by reference). The process for calculating a p- value score will also be discussed below. The analytics system 200 may determine, and optionally filter out, sequence reads of nucleic acid molecules or fragments with a methylation state vector having below a threshold p-value score as informative fragments. In another embodiment, the analytics system 200 further labels fragments with at least some number of CpG sites that have over some threshold percentage of methylation or unmethylation as hypermethylated and hypomethylated fragments, respectively. A hypermethylated fragment or a hypomethylated fragment may also be referred to as an unusual fragment with extreme methylation (UFXM). In other embodiments, the analytics system 200 may implement various other probabilistic models for determining informative molecules or fragments. Examples of other probabilistic models include a mixture model, a deep probabilistic model, etc. In particular embodiments, the analytics system 200 may use any combination of the processes described below for identifying informative fragments. With the identified informative fragments, the analytics system 200 may filter the set of methylation state vectors for a sample for use in other processes, e.g., for use in training and deploying a cancer classifier.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 34 of 119 4. Assay Protocol
[0165] FIG. 3 illustrates an example flowchart describing a process 300 of sequencing nucleic acids. In particular embodiments, the process 300 is performed to generate the sequence reads as part of step 110 of the workflow 100 of FIG.1.
[0166] In step 310, a nucleic acid sample (e.g., DNA or RNA) is extracted from a subject. In the present disclosure, DNA and RNA can be used interchangeably unless otherwise indicated. That is, the embodiments described herein can be applicable to both DNA and RNA types of nucleic acid sequences. However, the examples described herein can focus on DNA for purposes of clarity and explanation. The sample can include nucleic acid molecules derived from any subset of the human genome, including the whole genome. The sample can include blood, plasma, serum, urine, fecal, saliva, other types of bodily fluids, or any combination thereof. In particular embodiments, methods for drawing a blood sample (e.g., syringe or finger prick) can be less invasive than procedures for obtaining a tissue biopsy, which can require surgery. The extracted sample can comprise cfDNA and / or ctDNA. If a subject has a disease state, such as cancer, cell free nucleic acids (e.g., cfDNA) in an extracted sample from the subject generally includes detectable level of the nucleic acids that can be used to assess a disease state.
[0167] In step 315, the extracted nucleic acids (e.g., including cfDNA fragments) are treated to convert unmethylated cytosines to uracils. In particular embodiments, the process 300 uses a bisulfite treatment of the samples which converts the unmethylated cytosines to uracils without converting the methylated cytosines. As an example and not by way of limitation, a commercial kit such as the EZ DNA MethylationTM– Gold, EZ DNA MethylationTM– Direct or an EZ DNA MethylationTM– Lightning kit (available from Zymo Research Corp (Irvine, CA)) is used for the bisulfite conversion. In another embodiment, the conversion of unmethylated cytosines to uracils is accomplished using an enzymatic reaction. As an example and not by way of limitation, the conversion can use a commercially available kit for conversion of unmethylated cytosines to uracils, e.g., APOBEC-Seq (NEBiolabs, Ipswich, MA).
[0168] In step 320, a sequencing library is prepared. In particular embodiments, the preparation includes at least two steps. In a first step, a ssDNA adapter is added to the 3ʹ-OH end of a bisulfite-converted ssDNA molecule using a ssDNA ligation reaction. In particular embodiments, the ssDNA ligation reaction uses CircLigase II (Epicentre) to ligate the ssDNAACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 35 of 119 adapter to the 3ʹ-OH end of a bisulfite-converted ssDNA molecule, wherein the 5ʹ-end of the adapter is phosphorylated and the bisulfite-converted ssDNA has been dephosphorylated (i.e., the 3ʹ end has a hydroxyl group). In another embodiment, the ssDNA ligation reaction uses Thermostable 5ʹ AppDNA / RNA ligase (available from New England BioLabs (Ipswich, MA)) to ligate the ssDNA adapter to the 3ʹ-OH end of a bisulfite-converted ssDNA molecule. In this example, the first UMI adapter is adenylated at the 5ʹ-end and blocked at the 3ʹ-end. In another embodiment, the ssDNA ligation reaction uses a T4 RNA ligase (available from New England BioLabs) to ligate the ssDNA adapter to the 3ʹ-OH end of a bisulfite-converted ssDNA molecule.
[0169] In a second step, a second strand DNA is synthesized in an extension reaction. As an example and not by way of limitation, an extension primer, that hybridizes to a primer sequence included in the ssDNA adapter, is used in a primer extension reaction to form a double-stranded bisulfite-converted DNA molecule. Optionally, in some embodiments, the extension reaction uses an enzyme that is able to read through uracil residues in the bisulfite- converted template strand.
[0170] Optionally, in a third step, a dsDNA adapter is added to the double-stranded bisulfite-converted DNA molecule. Then, the double-stranded bisulfite-converted DNA can be amplified to add sequencing adapters. As an example and not by way of limitation, PCR amplification using a forward primer that includes a P5 sequence and a reverse primer that includes a P7 sequence is used to add P5 and P7 sequences to the bisulfite-converted DNA. Optionally, during library preparation, unique molecular identifiers (UMI) can be added to the nucleic acid molecules (e.g., DNA molecules) through adapter ligation. The UMIs are short nucleic acid sequences (e.g., 4-10 base pairs) that are added to ends of DNA fragments during adapter ligation. In particular embodiments, UMIs are degenerate base pairs that serve as a unique tag that can be used to identify sequence reads originating from a specific DNA fragment. During PCR amplification following adapter ligation, the UMIs are replicated along with the attached DNA fragment, which provides a way to identify sequence reads that came from the same original fragment in downstream analysis.
[0171] In an optional step 325, the nucleic acids (e.g., fragments) can be hybridized. Hybridization probes (also referred to herein as “probes”) may be used to target, and pull down, nucleic acid fragments informative for disease states. For a given workflow, the probes can be designed to anneal (or hybridize) to a target (complementary) strand of DNA or RNA. The target strand can be the “positive” strand (e.g., the strand transcribed into mRNA, andACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 36 of 119 subsequently translated into a protein) or the complementary “negative” strand. The probes can range in length from 10s, 100s, or 1000s of base pairs. Moreover, the probes can cover overlapping portions of a target region.
[0172] In an optional step 330, the hybridized nucleic acid fragments are captured and can be enriched, e.g., amplified using PCR. In particular embodiments, targeted DNA sequences can be enriched from the library. This is used, for example, where a targeted panel assay is being performed on the samples. As an example and not by way of limitation, the target sequences can be enriched to obtain enriched sequences that can be subsequently sequenced. In general, any known method in the art can be used to isolate, and enrich for, probe-hybridized target nucleic acids. As an example and not by way of limitation, as is well known in the art, a biotin moiety can be added to the 5ʹ-end of the probes (i.e., biotinylated) to facilitate isolation of target nucleic acids hybridized to probes using a streptavidin-coated surface (e.g., streptavidin-coated beads).
[0173] In step 335, sequence reads are generated from the nucleic acid sample, e.g., enriched sequences. Sequencing data can be acquired from the enriched DNA sequences by known means in the art. As an example and not by way of limitation, the method can include next generation sequencing (NGS) techniques including synthesis technology (Illumina), pyrosequencing (454 Life Sciences), ion semiconductor technology (Ion Torrent sequencing), single-molecule real-time sequencing (Pacific Biosciences), sequencing by ligation (SOLiD sequencing), nanopore sequencing (Oxford Nanopore Technologies), or paired-end sequencing. In particular embodiments, massively parallel sequencing is performed using sequencing-by-synthesis with reversible dye terminators.
[0174] In step 340, the sequence processor 220 can generate methylation information using the sequence reads. A methylation state vector can then be generated using the methylation information determined from the sequence reads. FIG.7A is an illustration of the process 700, starting from process 300 of FIG.3 of sequencing a cfDNA molecule, to obtain a methylation state vector 452, according to an embodiment. As an example, the analytics system 200 receives a cfDNA molecule 412 that, in this example, contains three CpG sites. As shown, the first and third CpG sites of the cfDNA molecule 412 are methylated 414. During the treatment step 315, the cfDNA molecule 412 is converted to generate a converted cfDNA molecule 422. During the treatment 315, the second CpG site which was unmethylated has its cytosine converted to uracil. However, the first and third CpG sites were not converted.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 37 of 119
[0175] After conversion, a sequencing library 430 is prepared and sequenced generating a sequence read 442. The analytics system 200 aligns (not shown) the sequence read 442 to a reference genome 444. The reference genome 444 provides the context as to what position in a human genome the fragment cfDNA originates from. In this simplified example, the analytics system 200 aligns the sequence read 442 such that the three CpG sites correlate to CpG sites 23, 24, and 25 (arbitrary reference identifiers used for convenience of description). The analytics system 200 thus generates information both on methylation status of all CpG sites on the cfDNA molecule 412 and the position in the human genome that the CpG sites map to. As shown, the CpG sites on sequence read 442 which were methylated are read as cytosines. In this example, the cytosines appear in the sequence read 442 only in the first and third CpG site which allows one to infer that the first and third CpG sites in the original cfDNA molecule were methylated. Whereas, the second CpG site is read as a thymine (U is converted to T during the sequencing process), and thus, one can infer that the second CpG site was unmethylated in the original cfDNA molecule. With these two pieces of information, the methylation status and location, the analytics system 200 generates 200 a methylation state vector 452 for the fragment cfDNA 412. In this example, the resulting methylation state vector 452 is < M23, U24, M25 >, wherein M corresponds to a methylated CpG site, U corresponds to an unmethylated CpG site, and the subscript number corresponds to a position of each CpG site in the reference genome. 5. Methylation Sequencing of DNA Fragments
[0176] FIG.4A illustrates an example flowchart describing a process 400 of sequencing a fragment of cfDNA to obtain a methylation state vector. In order to analyze DNA methylation, an analytics system 200 first obtains 410 a sample from an individual comprising a plurality of cfDNA molecules. In additional embodiments, the process 400 may be applied to sequence other types of DNA molecules. The process 400 is an embodiment of sample sequencing 120 of FIG.1.
[0177] From the sample, the analytics system 200 can isolate 410 each cfDNA molecule. The cfDNA molecules can be treated 420 to convert unmethylated cytosines to uracils. In particular embodiments, the method uses a bisulfite treatment of the DNA which converts the unmethylated cytosines to uracils without converting the methylated cytosines. As an example and not by way of limitation, a commercial kit such as the EZ DNA MethylationTM– Gold, EZ DNA MethylationTM– Direct or an EZ DNA MethylationTM– Lightning kit (available from Zymo Research Corp (Irvine, CA)) is used for the bisulfite conversion. In another embodiment,ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 38 of 119 the conversion of unmethylated cytosines to uracils is accomplished using an enzymatic reaction. As an example and not by way of limitation, the conversion can use a commercially available kit for conversion of unmethylated cytosines to uracils, such as APOBEC-Seq (NEBiolabs, Ipswich, MA).
[0178] From the converted cfDNA molecules, a sequencing library can be prepared 430. During library preparation, unique molecular identifiers (UMI) can be added to the nucleic acid molecules (e.g., DNA molecules) through adapter ligation. The UMIs can be short nucleic acid sequences (e.g., 4-10 base pairs) that are added to ends of DNA fragments (e.g., DNA molecules fragmented by physical shearing, enzymatic digestion, and / or chemical fragmentation) during adapter ligation. UMIs can be degenerate base pairs that serve as a unique tag that can be used to identify sequence reads originating from a specific DNA fragment. During PCR amplification following adapter ligation, the UMIs can be replicated along with the attached DNA fragment. This can provide a way to identify sequence reads that came from the same original fragment in downstream analysis.
[0179] Optionally, the sequencing library may be enriched 435 for cfDNA molecules, or genomic regions, that are informative for cancer status using a plurality of hybridization probes. The hybridization probes are short oligonucleotides capable of hybridizing to particularly specified cfDNA molecules, or targeted regions, and enriching for those fragments or regions for subsequent sequencing and analysis. Hybridization probes may be used to perform a targeted, high-depth analysis of a set of specified CpG sites of interest to the researcher. Hybridization probes can be tiled across one or more target sequences at a coverage of 1X, 2X, 3X, 4X, 5X, 6X, 7X, 8X, 9X, 10X, or more than 10X. As an example and not by way of limitation, hybridization probes tiled at a coverage of 2X comprises overlapping probes such that each portion of the target sequence is hybridized to 2 independent probes. Hybridization probes can be tiled across one or more target sequences at a coverage of less than 1X.
[0180] In particular embodiments, the hybridization probes are designed to enrich for DNA molecules that have been treated (e.g., using bisulfite) for conversion of unmethylated cytosines to uracils. During enrichment, hybridization probes (also referred to herein as “probes”) can be used to target and pull down nucleic acid fragments informative for the presence or absence of cancer (or disease), cancer status, or a cancer classification (e.g., cancer class or tissue of origin). The probes may be designed to anneal (or hybridize) to a target (complementary) strand of DNA. The target strand may be the “positive” strand (e.g., the strand transcribed into mRNA, and subsequently translated into a protein) or the complementaryACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 39 of 119 “negative” strand. The probes may range in length from 10s, 100s, or 1000s of base pairs. The probes can be designed based on a methylation site panel. The probes can be designed based on a panel of targeted genes to analyze particular mutations or target regions of the genome (e.g., of the human or another organism) that are suspected to correspond to certain cancers or other types of diseases. Moreover, the probes may cover overlapping portions of a target region.
[0181] Once prepared, the sequencing library or a portion thereof can be sequenced 440 to obtain a plurality of sequence reads. The sequence reads may be in a computer-readable, digital format for processing and interpretation by computer software. The sequence reads may be aligned to a reference genome to determine alignment position information. The alignment position information may indicate a beginning position and an end position of a region in the reference genome that corresponds to a beginning nucleotide base and end nucleotide base of a given sequence read. Alignment position information may also include sequence read length, which can be determined from the beginning position and end position. A region in the reference genome may be associated with a gene or a segment of a gene. A sequence read can be comprised of a read pair denoted as ^^and ^^. As an example and not by way of limitation, the first read ^^may be sequenced from a first end of a nucleic acid fragment whereas the second readmay be sequenced from the second end of the nucleic acid fragment. Therefore, nucleotide base pairs of the first read ^^and second read ^^may be aligned consistently (e.g., in opposite orientations) with nucleotide bases of the reference genome. Alignment position information derived from the read pair ^^and ^^may include a beginning position in the reference genome that corresponds to an end of a first read (e.g., ^^) and an end position in the reference genome that corresponds to an end of a second read (e.g., ^^). In other words, the beginning position and end position in the reference genome can represent the likely location within the reference genome to which the nucleic acid fragment corresponds. An output file having SAM (sequence alignment map) format or BAM (binary) format may be generated and output for further analysis such as methylation state determination.
[0182] From the sequence reads, the analytics system 200 determines 450 a location and methylation state for each CpG site based on alignment to a reference genome. The analytics system 200 generates 460 a methylation state vector for each fragment specifying a location of the fragment in the reference genome (e.g., as specified by the position of the first CpG site in each fragment, or another similar metric), a number of CpG sites in the fragment, and the methylation state of each CpG site in the fragment whether methylated (e.g., denoted as M),ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 40 of 119 unmethylated (e.g., denoted as U), or indeterminate (e.g., denoted as I). Observed states can be states of methylated and unmethylated; whereas, an unobserved state is indeterminate. Indeterminate methylation states may originate from sequencing errors and / or disagreements between methylation states of a DNA fragment's complementary strands. The methylation state vectors may be stored in temporary or persistent computer memory for later use and processing. Further, the analytics system 200 may remove duplicate reads or duplicate methylation state vectors from a single sample. The analytics system 200 may determine that a certain fragment with one or more CpG sites has an indeterminate methylation status over a threshold number or percentage, and may exclude such fragments or selectively include such fragments but build a model accounting for such indeterminate methylation statuses.
[0183] FIG. 4B illustrates an example process 400 of FIG. 4A of sequencing a cfDNA molecule to obtain a methylation state vector. As an example, the analytics system 200 receives a cfDNA molecule 412 that, in this example, contains three CpG sites. As shown, the first and third CpG sites of the cfDNA molecule 412 are methylated 414. During the treatment step 420, the cfDNA molecule 412 is converted to generate a converted cfDNA molecule 422. During the treatment 420, the second CpG site which was unmethylated has its cytosine converted to uracil. However, the first and third CpG sites were not converted.
[0184] After conversion, a sequencing library 430 is prepared and sequenced 440 to generate a sequence read 442. The analytics system 200 aligns 450 the sequence read 442 to a reference genome 444. The reference genome 444 provides the context as to what position in a human genome the fragment cfDNA originates from. In this simplified example, the analytics system 200 aligns 450 the sequence read 442 such that the three CpG sites correlate to CpG sites 33, 24, and 25 (arbitrary reference identifiers used for convenience of description). The analytics system 200 can thus generate information both on methylation status of all CpG sites on the cfDNA molecule 412 and the position in the human genome that the CpG sites map to. As shown, the CpG sites on sequence read 442 which are methylated are read as cytosines. In this example, the cytosines appear in the sequence read 442 only in the first and third CpG site which allows one to infer that the first and third CpG sites in the original cfDNA molecule are methylated. Whereas, the second CpG site can be read as a thymine (U is converted to T during the sequencing process), and thus, one can infer that the second CpG site is unmethylated in the original cfDNA molecule. With these two pieces of information, the methylation status and location, the analytics system 200 generates 460 a methylation state vector 452 for the fragment cfDNA 412. In this example, the resulting methylation state vector 452 is < M23, U24, M25 >,ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 41 of 119 wherein M corresponds to a methylated CpG site, U corresponds to an unmethylated CpG site, and the subscript number corresponds to a position of each CpG site in the reference genome.
[0185] One or more alternative sequencing methods can be used for obtaining sequence reads from nucleic acids in a biological sample. The one or more sequencing methods can comprise any form of sequencing that can be used to obtain a number of sequence reads measured from nucleic acids (e.g., cell-free nucleic acids), including, but not limited to, high- throughput sequencing systems such as the Roche 454 platform, the Applied Biosystems SOLID platform, the Helicos True Single Molecule DNA sequencing technology, the sequencing-by-hybridization platform from Affymetrix Inc., the single-molecule, real-time (SMRT) technology of Pacific Biosciences, the sequencing-by-synthesis platforms from 454 Life Sciences, Illumina / Solexa and Helicos Biosciences, and the sequencing-by-ligation platform from Applied Biosystems. The ION TORRENT technology from Life technologies and Nanopore sequencing can also be used to obtain sequence reads from the nucleic acids (e.g., cell-free nucleic acids) in the biological sample. Sequencing-by-synthesis and reversible terminator-based sequencing (e.g., Illumina’s Genome Analyzer; Genome Analyzer II; HISEQ 3000; HISEQ 4500 (Illumina, San Diego Calif.)) can be used to obtain sequence reads from the cell-free nucleic acid obtained from a biological sample of a training subject in order to form the genotypic dataset. Millions of cell-free nucleic acid (e.g., DNA) fragments can be sequenced in parallel. In one example of this type of sequencing technology, a flow cell is used that contains an optically transparent slide with eight individual lanes on the surfaces of which are bound oligonucleotide anchors (e.g., adaptor primers). A cell-free nucleic acid sample can include a signal or tag that facilitates detection. The acquisition of sequence reads from the cell-free nucleic acid obtained from the biological sample can include obtaining quantification information of the signal or tag via a variety of techniques such as, for example, flow cytometry, quantitative polymerase chain reaction (qPCR), gel electrophoresis, gene-chip analysis, microarray, mass spectrometry, cytofluorimetric analysis, fluorescence microscopy, confocal laser scanning microscopy, laser scanning cytometry, affinity chromatography, manual batch mode separation, electric field suspension, sequencing, and combination thereof.
[0186] The one or more sequencing methods can comprise a whole-genome sequencing assay. A whole-genome sequencing assay can comprise a physical assay that generates sequence reads for a whole genome or a substantial portion of the whole genome which can be used to determine large variations such as copy number variations or copy number aberrations. Such a physical assay may employ whole-genome sequencing techniques or whole-exomeACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 42 of 119 sequencing techniques. A whole-genome sequencing assay can have an average sequencing depth of at least 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, at least 20x, at least 30x, or at least 40x across the genome of the test subject. In particular embodiments, the sequencing depth is about 30,000x. The one or more sequencing methods can comprise a targeted panel sequencing assay. A targeted panel sequencing assay can have an average sequencing depth of at least 50,000x, at least 55,000x, at least 60,000x, or at least 70,000x sequencing depth for the targeted panel of genes. The targeted panel of genes can comprise between 450 and 500 genes. The targeted panel of genes can comprise a range of 500±5 genes, a range of 500±10 genes, or a range of 500±25 genes.
[0187] The one or more sequencing methods can comprise paired-end sequencing. The one or more sequencing methods can generate a plurality of sequence reads. The plurality of sequence reads can have an average length ranging between 10 and 700, between 50 and 400, or between 100 and 300. The one or more sequencing methods can comprise a methylation sequencing assay. The methylation sequencing can be i) whole-genome methylation sequencing or ii) targeted DNA methylation sequencing using a plurality of nucleic acid probes. As an example and not by way of limitation, the methylation sequencing is whole- genome bisulfite sequencing (e.g., WGBS). The methylation sequencing can be a targeted DNA methylation sequencing using a plurality of nucleic acid probes targeting the most informative regions of the methylome, a unique methylation database and prior prototype whole-genome and targeted sequencing assays.
[0188] The methylation sequencing can detect one or more 5-methylcytosine (5mC) and / or 5-hydroxymethylcytosine (5hmC) in respective nucleic acid methylation fragments. The methylation sequencing can comprise conversion of one or more unmethylated cytosines or one or more methylated cytosines, in respective nucleic acid methylation fragments, to a corresponding one or more uracils. The one or more uracils can be detected during the methylation sequencing as one or more corresponding thymines. The conversion of one or more unmethylated cytosines or one or more methylated cytosines can comprise a chemical conversion, an enzymatic conversion, or combinations thereof.
[0189] As an example and not by way of limitation, bisulfite conversion involves converting cytosine to uracil while leaving methylated cytosines (e.g., 5-methylcytosine or 5- mC) intact. In some DNA, about 95% of cytosines may not be methylated in the DNA, and the resulting DNA fragments may include many uracils which are represented by thymines. Enzymatic conversion processes may be used to treat the nucleic acids prior to sequencing,ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 43 of 119 which can be performed in various ways. One example of a bisulfite-free conversion comprises a bisulfite-free and base-resolution sequencing method, TET-assisted pyridine borane sequencing (TAPS), for non-destructive and direct detection of 5-methylcytosine and 5- hydroxymethylcytosine without affecting unmodified cytosines. The methylation state of a CpG site in the corresponding plurality of CpG sites in the respective nucleic acid methylation fragment can be methylated when the CpG site is determined by the methylation sequencing to be methylated, and unmethylated when the CpG site is determined by the methylation sequencing to not be methylated.
[0190] A methylation sequencing assay (e.g., WGBS and / or targeted methylation sequencing) can have an average sequencing depth including but not limited to up to about 1,000x, 2,000x, 3,000x, 5,000x, 10,000x, 15,000x, 20,000x, or 30,000x. The methylation sequencing can have a sequencing depth that is greater than 30,000x, e.g., at least 40,000x or 50,000x. A whole-genome bisulfite sequencing method can have an average sequencing depth of between 20x and 50x, and a targeted methylation sequencing method has an average effective depth of between 100x and 1000x, where effective depth can be the equivalent whole- genome bisulfite sequencing coverage for obtaining the same number of sequence reads obtained by targeted methylation sequencing.
[0191] For further details regarding methylation sequencing (e.g., WGBS and / or targeted methylation sequencing), see, e.g., U.S. Patent Application No. 16 / 352,602, entitled “Methylation Fragment Anomaly Detection,” filed March 13, 2019, U.S. Patent Application No. 16 / 719,902, entitled “Systems and Methods for Estimating Cell Source Fractions Using Methylation Information,” filed December 18, 2019, and U.S. Patent Application No. 17 / 191,914, titled “Systems and Methods for Cancer Condition Determination Using Autoencoders,” filed March 4, 2021. Other methods for methylation sequencing, including those disclosed herein and / or any modifications, substitutions, or combinations thereof, can be used to obtain fragment methylation patterns. A methylation sequencing can be used to identify one or more methylation state vectors, as described, for example, in United States Patent Application No. 16 / 352,602, entitled “Anomalous Fragment Detection and Classification,” filed March 13, 2019, or according to any of the techniques disclosed in United States Patent Application No. 15 / 931,022, entitled “Model-Based Featurization and Classification,” filed May 13, 2020. Each reference is incorporated by reference in its entirety.
[0192] The methylation sequencing of nucleic acids and the resulting one or more methylation state vectors can be used to obtain a plurality of nucleic acid methylationACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 44 of 119 fragments. Each corresponding plurality of nucleic acid methylation fragments (e.g., for each respective genotypic dataset) can comprise more than 100 nucleic acid methylation fragments. An average number of nucleic acid methylation fragments across each corresponding plurality of nucleic acid methylation fragments can comprise 1000 or more nucleic acid methylation fragments, 5000 or more nucleic acid methylation fragments, 10,000 or more nucleic acid methylation fragments, 20,000 or more nucleic acid methylation fragments, or 30,000 or more nucleic acid methylation fragments. An average number of nucleic acid methylation fragments across each corresponding plurality of nucleic acid methylation fragments can be between 10,000 nucleic acid methylation fragments and 50,000 nucleic acid methylation fragments. The corresponding plurality of nucleic acid methylation fragments can comprise one thousand or more, ten thousand or more, 100 thousand or more, one million or more, ten million or more, 100 million or more, 500 million or more, one billion or more, two billion or more, three billion or more, four billion or more, five billion or more, six billion or more, seven billion or more, eight billion or more, nine billion or more, or 10 billion or more nucleic acid methylation fragments. An average length of a corresponding plurality of nucleic acid methylation fragments can be between 140 and 480 nucleotides. 6. Blocks of Reference Genome
[0193] FIG. 5 illustrates example blocks of a reference genome. The sequence processor 220 can partition a reference genome (or a subset of the reference genome) in one or more stages, e.g., for use cases involving a targeted methylation assay. For instance, the sequence processor 220 separates the reference genome into blocks of CpG sites. Each block is defined when there is a separation between two adjacent CpG sites that exceeds a threshold, e.g., greater than 200 base pairs (bp), 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, or 1,000 bp, among other values. Thus, blocks can vary in size of base pairs. For each block, the sequence processor 220 can subdivide the block into windows of a certain length, e.g., 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1,100 bp, 1,200 bp, 1,300 bp, 1,400 bp, or 1,500 bp, among other values. In other embodiments, the windows can be from 200 bp to 10 kilobase pairs (kbp), from 500 bp to 2 kbp, or about 1 kbp in length. Windows (e.g., that are adjacent) can overlap by a number of base pairs or a percentage of the length, e.g., 10%, 20%, 30%, 40%, 50%, or 60%, among other values. Windows can be separated between two adjacent CpG sites that exceeds a threshold, e.g., greater than 200 base pairs (bp), 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, or 1,000 bp, among other values.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 45 of 119
[0194] The sequence processor 220 can analyze sequence reads derived from DNA fragments using a windowing process. In particular, the sequence processor 220 scans through the blocks window-by-window and reads fragments within each window. The fragments can originate from tissue and / or high-signal cfDNA. High-signal cfDNA samples can be determined by a binary classification model, by cancer stage, or by another metric. By partitioning the reference genome (e.g., using blocks and windows), the sequence processor 220 can facilitate computational parallelization. Moreover, the sequence processor 220 can reduce computational resources to process a reference genome by targeting the sections of base pairs that include CpG sites, while skipping other sections that do not include CpG sites. 7. Deconvolution of Component Proportions
[0195] Deconvolution of component proportions involves determining a breakdown of component proportions where nucleic acid fragments in a sample originate from. The component proportions represent a breakdown of components contributing nucleic acid fragments to the sample. In one or more embodiments, the component proportions specify a percentage of the sample attributed to each component. As an example and not by way of limitation, the deconvolution process determines a first percentage of nucleic acid fragments in a sample originating from a first component, a second percentage of nucleic acid fragments in the sample originating from a second component, so on and so forth with remaining components. In other embodiments, the component proportions may specify a rank of the components according to the predicted proportions. As an example and not by way of limitation, component 3 is the largest contributor of nucleic acid fragments to the sample, followed by component 1, then component 2 (in an example with the mixture model predicting proportions for three components).
[0196] The deconvolution process models methylation signatures of various components. A methylation signature represents the methylation sequence reads of the component. In particular embodiments, the methylation signature comprises counts of methylation variants over a plurality of genomic regions. In other embodiments, the methylation signature may further comprise counts of reference states over the plurality of genomic regions. The reference state at a genomic region may include remaining methylation patterns not assigned to variant states. As for the components, each component may be a tissue type, a cell type, or further subtypes thereof.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 46 of 119
[0197] The deconvolution process may be applied at various stages of the overall cancer classification workflow 100 described in FIG. 1. In one example implementation, the deconvolution process may be utilized to assess sample purity, e.g., at step 142 of the workflow 100. In another example implementation, the deconvolution process may be utilized in the cancer classification 146 to determine a particular cancer type. As yet another example implementation, the deconvolution process may be utilized to quantify cancer signal, e.g., for use in monitoring cancer status and / or progression. Other example implementations may utilize the component proportions in other manners.
[0198] FIG.6A illustrates example methylation features that can be derived from a single CpG site as a genomic region. FIG.6A illustrates counting methylation variants from a single CpG site 605 as a genomic region, according to one or more embodiments. In a single CpG site genomic region, there generally are two methylation states, a reference state and a variant state (pertaining to the methylation variant). The analytics system 200 may count a first count of methylation sequence reads having the reference state and a second count of methylation sequence reads having the variant state.
[0199] As an example and not by way of limitation, in FIG. 6A, there are six fragments 610 that overlap the single CpG site 605. At each CpG site on a fragment, denoted as a diamond, the fragment has a methylation state. Methylation states may include methylated, shown as filled in, unmethylated shown as unfilled, and unknown shown with a diagonal hatch. Unknown methylation may include indetermined states caused from mutations or sequencing errors. As a single site, there are two possible methylation states at this genomic region (M or U). One methylation state is deemed the reference state while the other is deemed the variant state. As an example and not by way of limitation, at the CpG site 605 as the genomic region, the reference genome provides that this CpG site predominantly is methylated, thus the reference state is methylated. As for counts, the analytics system 200 may count four fragments (or methylation sequence reads) as having the reference state of methylated, and two fragments (or methylation sequence reads) as having the variant state of unmethylated.
[0200] FIG.6B illustrates example methylation features that can be derived from multiple CpG sites as a genomic region. FIG. 6B illustrates counting methylation variants from a genomic region 615 spanning multiple CpG sites 617, according to one or more embodiments. The CpG sites 617 include CpG sites 1, 2, and 3. As in FIG. 6A, a filled diamond indicates methylation, an unfilled diamond indicates unmethylation, and a diagonal hatch diamond indicates unknown. In the embodiment shown, fragment 1 620A may have a methylationACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 47 of 119 pattern comprising methylated at CpG site 1, unmethylated at CpG site 2, and methylated at CpG site 3. As an example, this may be a reference state. The analytics system 200 may count the number of fragments having the reference state of methylated, unmethylated, and methylated at the genomic region 615. In this example, four fragments have the reference state. Fragment 2620B has a methylation variant, that is different than the reference state. Fragment 2620B has a methylation pattern with all three CpG site 617 methylated. The analytics system 200 may count one total fragment (or methylation sequence read) as having this first methylation variant. Fragment 4620D has another methylation variant, that is different than the reference state and the first methylation variant. Fragment 4620D has a methylation pattern with all three CpG sites 617 unmethylated. The analytics system 200 may count one total fragment (or methylation sequence read) as having this second methylation variant. The analytics system 200 may also sum counts of all methylation variants. In such example, the count for all methylation variants is two, including fragment 2 620A having the first methylation variant and fragment 4620D having the second methylation variant. 8. Cancer Classifier for Determining Cancer
[0201] Cancer classification involves extraction genetic features and applying one or more models to the extracted features to determine a cancer prediction. The analytics system 200 aggregates extracted features into a feature vector which can then be input into a trained cancer classifier to determine a cancer prediction based on the input feature vector. The cancer prediction may comprise a label and / or a value. The label may be binary, indicating a presence or absence of cancer in the test subject, and / or multiclass, indicating one or more particular cancer types from a plurality of screened cancer types. The value may be a gradation of cancer status, e.g., a stage of cancer, or quantification of cancer signal. In particular, a cancer classifier may be a machine-learned model comprising a plurality of classification parameters and a function representing a relation between the feature vector as input and the cancer prediction as output. Inputting the feature vector into the function with the classification parameters yields the cancer prediction. In one or more embodiments, cancer classification utilizes a plurality of models. Some models may be used to determine one or more features from the methylation sequence reads for use in cancer classification. Prior to deployment of the cancer classifier, the analytics system 200 trains the cancer classifier.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 48 of 119 8.1. Identifying Anomalous Fragments
[0202] The analytics system 200 can determine anomalous fragments for a sample using the sample’s methylation state vectors. For each fragment in a sample, the analytics system 200 can determine whether the fragment is an anomalous fragment using the methylation state vector corresponding to the fragment. In particular embodiments, the analytics system 200 calculates a p-value score for each methylation state vector describing a probability of observing that methylation state vector or other methylation state vectors even less probable in the healthy control group. The analytics system 200 may determine fragments with a methylation state vector having below a threshold p-value score as anomalous fragments. In particular embodiments, the analytics system 200 further labels fragments with at least some number of CpG sites that have over some threshold percentage of methylation or unmethylation as hypermethylated and hypomethylated fragments, respectively. A hypermethylated fragment or a hypomethylated fragment may also be referred to as an unusual fragment with extreme methylation (UFXM). In other embodiments, the analytics system 200 may implement various other probabilistic models for determining anomalous fragments. Examples of other probabilistic models include a mixture model, a deep probabilistic model, etc. In particular embodiments, the analytics system 200 may use any combination of the processes described below for identifying anomalous fragments. With the identified anomalous fragments, the analytics system 200 may filter the set of methylation state vectors for a sample for use in other processes, e.g., for use in training and deploying a cancer classifier. P-Value Filtering
[0203] In particular embodiments, the analytics system 200 calculates a p-value score for each methylation state vector compared to methylation state vectors from fragments in a healthy control group. The p-value score can describe a probability of observing the methylation status matching that methylation state vector or other methylation state vectors even less probable in the healthy control group. In order to determine a DNA fragment to be anomalously methylated, the analytics system 200 can use a healthy control group with a majority of fragments that are normally methylated. When conducting this probabilistic analysis for determining anomalous fragments, the determination can hold weight in comparison with the group of control subjects that make up the healthy control group. To ensure robustness in the healthy control group, the analytics system 200 may select some threshold number of healthy individuals to source samples including DNA fragments. FIG.7AACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 49 of 119 illustrates an example method of generating a data structure for a healthy control group with which the analytics system 200 may calculate p-value scores. FIG. 7B illustrates an example method of calculating a p-value score with the generated data structure.
[0204] FIG.7A is an exemplary flowchart describing a process 700 of generating a control group data structure for determining anomalously methylated fragments, according to one or more embodiments. To create a healthy control group data structure, the analytics system 200 can receive a plurality of DNA fragments (e.g., cfDNA) from a plurality of healthy individuals. The analytics system 200 can generate 705 a methylation state vector for each fragment, for example via the process 400.
[0205] With each fragment’s methylation state vector, the analytics system 200 can subdivide 710 the methylation state vector into strings of CpG sites. In particular embodiments, the analytics system 200 subdivides 710 the methylation state vector such that the resulting strings are all less than a given length. As an example and not by way of limitation, a methylation state vector of length 11 may be subdivided into strings of length less than or equal to 3 would result in 9 strings of length 3, 10 strings of length 2, and 11 strings of length 1. In another example, a methylation state vector of length 7 being subdivided into strings of length less than or equal to 4 can result in 4 strings of length 4, 5 strings of length 3, 6 strings of length 2, and 7 strings of length 1. If a methylation state vector is shorter than or the same length as the specified string length, then the methylation state vector may be converted into a single string containing all of the CpG sites of the vector.
[0206] The analytics system 200 tallies 715 the strings by counting, for each possible CpG site and possibility of methylation states in the vector, the number of strings present in the control group having the specified CpG site as the first CpG site in the string and having that possibility of methylation states. As an example and not by way of limitation, at a given CpG site and considering string lengths of 3, there are 2^3 or 8 possible string configurations. At that given CpG site, for each of the 8 possible string configurations, the analytics system 200 tallies 810 how many occurrences of each methylation state vector possibility come up in the control group. Continuing this example, this may involve tallying the following quantities: < Mx, Mx+1, Mx+2 >, < Mx, Mx+1, Ux+2 >, . . ., < Ux, Ux+1, Ux+2 > for each starting CpG site x in the reference genome. The analytics system 200 creates 715 the data structure storing the tallied counts for each starting CpG site and string possibility.
[0207] There are several benefits to setting an upper limit on string length. First, depending on the maximum length for a string, the size of the data structure created by the analytics systemACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 50 of 119 200 can dramatically increase in size. For instance, maximum string length of 4 means that every CpG site has at the very least 2^4 numbers to tally for strings of length 4. Increasing the maximum string length to 5 means that every CpG site has an additional 2^4 or 16 numbers to tally, doubling the numbers to tally (and computer memory required) compared to the prior string length. Reducing string size can help keep the data structure creation and performance (e.g., use for later accessing as described below), in terms of computational and storage, reasonable. Second, a statistical consideration to limiting the maximum string length can be to avoid overfitting downstream models that use the string counts. If long strings of CpG sites do not, biologically, have a strong effect on the outcome (e.g., predictions of anomalousness that predictive of the presence of cancer), calculating probabilities based on large strings of CpG sites can be problematic as it uses a significant amount of data that may not be available, and thus can be too sparse for a model to perform appropriately. As an example and not by way of limitation, calculating a probability of anomalousness / cancer conditioned on the prior 100 CpG sites can use counts of strings in the data structure of length 100, ideally some matching exactly the prior 100 methylation states. If only sparse counts of strings of length 100 are available, there can be insufficient data to determine whether a given string of length of 100 in a test sample is anomalous or not.
[0208] FIG. 7B is an exemplary flowchart describing a process 730 of determining a fragment to be anomalously methylated based on the control group data structure, according to one or more embodiments. In process 730, the analytics system 200 generates 740 methylation state vectors from cfDNA fragments of the subject, e.g., via the process 400. The analytics system 200 can handle each methylation state vector as follows.
[0209] For a given methylation state vector, the analytics system 200 enumerates 745 all possibilities of methylation state vectors having the same starting CpG site and same length (i.e., set of CpG sites) in the methylation state vector. As each methylation state is generally either methylated or unmethylated there can be effectively two possible states at each CpG site, and thus the count of distinct possibilities of methylation state vectors can depend on a power of 2, such that a methylation state vector of length n would be associated with 2npossibilities of methylation state vectors. With methylation state vectors inclusive of indeterminate states for one or more CpG sites, the analytics system 200 may enumerate 730 possibilities of methylation state vectors considering only CpG sites that have observed states.
[0210] The analytics system 200 calculates 750 the probability of observing each possibility of methylation state vector for the identified starting CpG site and methylation stateACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 51 of 119 vector length by accessing the healthy control group data structure. In particular embodiments, calculating the probability of observing a given possibility uses a Markov chain probability to model the joint probability calculation. The Markov model can be trained, at least in part, based upon evaluation of a methylation state of each CpG site in the corresponding plurality of CpG sites of the respective fragment (e.g., nucleic acid methylation fragment) across those nucleic acid methylation fragments in a healthy noncancer cohort dataset that have the corresponding plurality of CpG sites. As an example and not by way of limitation, a Markov model (e.g., a Hidden Markov Model or HMM) is used to determine the probability that a sequence of methylation states (comprising, e.g., “M” or “U”) can be observed for a nucleic acid methylation fragment in a plurality of nucleic acid methylation fragments, given a set of probabilities that determine, for each state in the sequence, the likelihood of observing the next state in the sequence. The set of probabilities can be obtained by training the HMM. Such training can involve computing statistical parameters (e.g., the probability that a first state can transition to a second state (the transition probability) and / or the probability that a given methylation state can be observed for a respective CpG site (the emission probability)), given an initial training dataset of observed methylation state sequences (e.g., methylation patterns). HMMs can be trained using supervised training (e.g., using samples where the underlying sequence as well as the observed states are known) and / or unsupervised training (e.g., Viterbi learning, maximum likelihood estimation, expectation-maximization training, and / or Baum- Welch training). In other embodiments, calculation methods other than Markov chain probabilities are used to determine the probability of observing each possibility of methylation state vector. As an example and not by way of limitation, such calculation method can include a learned representation. The p-value threshold can be between 0.01 and 0.10, or between 0.03 and 0.06. The p-value threshold can be 0.05. The p-value threshold can be less than 0.01, less than 0.001, or less than 0.0001.
[0211] The analytics system 200 calculates 755 a p-value score for the methylation state vector using the calculated probabilities for each possibility. In particular embodiments, this includes identifying the calculated probability corresponding to the possibility that matches the methylation state vector in question. Specifically, this can be the possibility having the same set of CpG sites, or similarly the same starting CpG site and length as the methylation state vector. The analytics system 200 can sum the calculated probabilities of any possibilities having probabilities less than or equal to the identified probability to generate the p-value score.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 52 of 119
[0212] This p-value can represent the probability of observing the methylation state vector of the fragment or other methylation state vectors even less probable in the healthy control group. A low p-value score can, thereby, generally correspond to a methylation state vector which is rare in a healthy individual, and which causes the fragment to be labeled anomalously methylated, relative to the healthy control group. A high p-value score can generally relate to a methylation state vector is expected to be present, in a relative sense, in a healthy individual. If the healthy control group is a non-cancerous group, for example, a low p-value can indicate that the fragment is anomalous methylated relative to the non-cancer group, and therefore possibly indicative of the presence of cancer in the test subject.
[0213] As above, the analytics system 200 can calculate p-value scores for each of a plurality of methylation state vectors, each representing a cfDNA fragment in the test sample. To identify which of the fragments are anomalously methylated, the analytics system 200 may filter 765 the set of methylation state vectors based on their p-value scores. In particular embodiments, filtering is performed by comparing the p-values scores against a threshold and keeping only those fragments below the threshold. This threshold p-value score can be on the order of 0.1, 0.01, 0.001, 0.0001, or similar.
[0214] According to example results from the process 700, the analytics system 200 can yield a median (range) of 2,800 (1,500-12,000) fragments with anomalous methylation patterns for participants without cancer in training, and a median (range) of 3,000 (1,200-420,000) fragments with anomalous methylation patterns for participants with cancer in training. These filtered sets of fragments with anomalous methylation patterns may be used for the downstream analyses.
[0215] In particular embodiments, the analytics system 200 uses 760 a sliding window to determine possibilities of methylation state vectors and calculate p-values. Rather than enumerating possibilities and calculating p-values for entire methylation state vectors, the analytics system 200 can enumerate possibilities and calculates p-values for only a window of sequential CpG sites, where the window is shorter in length (of CpG sites) than at least some fragments (otherwise, the window would serve no purpose). The window length may be static, user determined, dynamic, or otherwise selected.
[0216] In calculating p-values for a methylation state vector larger than the window, the window can identify the sequential set of CpG sites from the vector within the window starting from the first CpG site in the vector. The analytics system 200 can calculate a p-value score for the window including the first CpG site. The analytics system 200 can then “slide” the windowACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 53 of 119 to the second CpG site in the vector, and calculates another p-value score for the second window. Thus, for a window size l and methylation vector length m, each methylation state vector can generate m–l+1 p-value scores. After completing the p-value calculations for each portion of the vector, the lowest p-value score from all sliding windows can be taken as the overall p-value score for the methylation state vector. In other embodiments, the analytics system 200 aggregates the p-value scores for the methylation state vectors to generate an overall p-value score.
[0217] Using the sliding window can help to reduce the number of enumerated possibilities of methylation state vectors and their corresponding probability calculations that would otherwise need to be performed. To give a realistic example, it can be for fragments to have upwards of 54 CpG sites. Instead of computing probabilities for 2^54 (~1.8×10^16) possibilities to generate a single p-score, the analytics system 200 can instead use a window of size 5 (for example) which results in 50 p-value calculations for each of the 50 windows of the methylation state vector for that fragment. Each of the 50 calculations can enumerate 2^5 (32) possibilities of methylation state vectors, which total results in 50×2^5 (1.6×10^3) probability calculations. This can result in a vast reduction of calculations to be performed, with no meaningful hit to the accurate identification of anomalous fragments.
[0218] In embodiments with indeterminate states, the analytics system 200 may calculate a p-value score summing out CpG sites with indeterminates states in a fragment’s methylation state vector. The analytics system 200 can identify all possibilities that have consensus with the all methylation states of the methylation state vector excluding the indeterminate states. The analytics system 200 may assign the probability to the methylation state vector as a sum of the probabilities of the identified possibilities. As an example, the analytics system 200 can calculate a probability of a methylation state vector of < M1, I2, U3 > as a sum of the probabilities for the possibilities of methylation state vectors of < M1, M2, U3> and < M1, U2, U3 > since methylation states for CpG sites 1 and 3 are observed and in consensus with the fragment’s methylation states at CpG sites 1 and 3. This method of summing out CpG sites with indeterminate states can use calculations of probabilities of possibilities up to 2^i, wherein i denotes the number of indeterminate states in the methylation state vector. In additional embodiments, a dynamic programming algorithm may be implemented to calculate the probability of a methylation state vector with one or more indeterminate states. Advantageously, the dynamic programming algorithm operates in linear computational time.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 54 of 119
[0219] In particular embodiments, the computational burden of calculating probabilities and / or p-value scores may be further reduced by caching at least some calculations. As an example and not by way of limitation, the analytic system 200 may cache in transitory or persistent memory calculations of probabilities for possibilities of methylation state vectors (or windows thereof). If other fragments have the same CpG sites, caching the possibility probabilities can allow for efficient calculation of p-score values without needing to re- calculate the underlying possibility probabilities. Equivalently, the analytics system 200 may calculate p-value scores for each of the possibilities of methylation state vectors associated with a set of CpG sites from vector (or window thereof). The analytics system 200 may cache the p-value scores for use in determining the p-value scores of other fragments including the same CpG sites. Generally, the p-value scores of possibilities of methylation state vectors having the same CpG sites may be used to determine the p-value score of a different one of the possibilities from the same set of CpG sites.
[0220] In particular embodiments, the p-value scores may be adjusted for multiple hypothesis testing according to a variety of suitable techniques. As an example, and without limitation, said techniques can include controlling for false positive rate, family-wise error rate, experiment-wise error rate, false discovery rate, etc. Known and applicable techniques for adjusting p-values include, without limitation, a Bonferroni procedure, a Holm procedure, a Hochberg procedure, a harmonic mean p-value procedure, a Benjamini-Hochberg procedure, a Benjamini-Yekutieli procedure, a Storey-Tibshirani procedure, etc. Adjusting p-value scores for multiple hypothesis testing can be used to improve the accuracy associated with basing positive detection calls based on the p-value and to reduce the incidence of false positives.
[0221] One or more nucleic acid methylation fragments can be filtered prior to training region models or cancer classifier. Filtering nucleic acid methylation fragments can comprise removing, from the corresponding plurality of nucleic acid methylation fragments, each respective nucleic acid methylation fragment that fails to satisfy one or more selection criteria (e.g., below or above one selection criteria). The one or more selection criteria can comprise a p-value threshold. The output p-value of the respective nucleic acid methylation fragment can be determined, at least in part, based upon a comparison of the corresponding methylation pattern of the respective nucleic acid methylation fragment to a corresponding distribution of methylation patterns of those nucleic acid methylation fragments in a healthy noncancer cohort dataset that have the corresponding plurality of CpG sites of the respective nucleic acid methylation fragment.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 55 of 119
[0222] Filtering a plurality of nucleic acid methylation fragments can comprise removing each respective nucleic acid methylation fragment that fails to satisfy a p-value threshold. The filter can be applied to the methylation pattern of each respective nucleic acid methylation fragment using the methylation patterns observed across the first plurality of nucleic acid methylation fragments. Each respective methylation pattern of each respective nucleic acid methylation fragment (e.g., Fragment One, …, Fragment N) can comprise a corresponding one or more methylation sites (e.g., CpG sites) identified with a methylation site identifier and a corresponding methylation pattern, represented as a sequence of 1’s and 0’s, where each “1” represents a methylated CpG site in the one or more CpG sites and each “0” represents an unmethylated CpG site in the one or more CpG sites. The methylation patterns observed across the first plurality of nucleic acid methylation fragments can be used to build a methylation state distribution for the CpG site states collectively represented by the first plurality of nucleic acid methylation fragments (e.g., CpG site A, CpG site B, …, CpG site ZZZ). Further details regarding processing of nucleic acid methylation fragments are disclosed in U.S. Provisional Patent Application No. 17 / 191,914, titled “Systems and Methods for Cancer Condition Determination Using Autoencoders,” filed March 4, 2021, which is hereby incorporated herein by reference in its entirety.
[0223] The respective nucleic acid methylation fragment may fail to satisfy a selection criterion in the one or more selection criteria when the respective nucleic acid methylation fragment has an anomalous methylation score that is less than an anomalous methylation score threshold. In this situation, the anomalous methylation score can be determined by a mixture model. As an example and not by way of limitation, a mixture model can detect an anomalous methylation pattern in a nucleic acid methylation fragment by determining the likelihood of a methylation state vector (e.g., a methylation pattern) for the respective nucleic acid methylation fragment based on the number of possible methylation state vectors of the same length and at the same corresponding genomic location. This can be executed by generating a plurality of possible methylation states for vectors of a specified length at each genomic location in a reference genome. Using the plurality of possible methylation states, the number of total possible methylation states and subsequently the probability of each predicted methylation state at the genomic location can be determined. The likelihood of a sample nucleic acid methylation fragment corresponding to a genomic location within the reference genome can then be determined by matching the sample nucleic acid methylation fragment to a predicted (e.g., possible) methylation state and retrieving the calculated probability of the predictedACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 56 of 119 methylation state. An anomalous methylation score can then be calculated based on the probability of the sample nucleic acid methylation fragment.
[0224] The respective nucleic acid methylation fragment can fail to satisfy a selection criterion in the one or more selection criteria when the respective nucleic acid methylation fragment has less than a threshold number of residues. The threshold number of residues can be between 10 and 50, between 50 and 100, between 100 and 150, or more than 150. The threshold number of residues can be a fixed value between 20 and 90. The respective nucleic acid methylation fragment may fail to satisfy a selection criterion in the one or more selection criteria when the respective nucleic acid methylation fragment has less than a threshold number of CpG sites. The threshold number of CpG sites can be 4, 5, 6, 7, 8, 9, or 10. The respective nucleic acid methylation fragment can fail to satisfy a selection criterion in the one or more selection criteria when a genomic start position and a genomic end position of the respective nucleic acid methylation fragment indicates that the respective nucleic acid methylation fragment represents less than a threshold number of nucleotides in a human genome reference sequence.
[0225] The filtering can remove a nucleic acid methylation fragment in the corresponding plurality of nucleic acid methylation fragments that has the same corresponding methylation pattern and the same corresponding genomic start position and genomic end position as another nucleic acid methylation fragment in the corresponding plurality of nucleic acid methylation fragments. This filtering step can remove redundant fragments that are exact duplicates, including, in some instances, PCR duplicates. The filtering can remove a nucleic acid methylation fragment that has the same corresponding genomic start position and genomic end position and less than a threshold number of different methylation states as another nucleic acid methylation fragment in the corresponding plurality of nucleic acid methylation fragments. The threshold number of different methylation states used for retention of a nucleic acid methylation fragment can be 1, 2, 3, 4, 5, or more than 5. As an example and not by way of limitation, a first nucleic acid methylation fragment having the same corresponding genomic start and end position as a second nucleic acid methylation fragment but having at least 1, at least 2, at least 3, at least 4, or at least 5 different methylation states at a respective CpG site (e.g., aligned to a reference genome) is retained. As another example, a first nucleic acid methylation fragment having the same methylation state vector (e.g., methylation pattern) but different corresponding genomic start and end positions as a second nucleic acid methylation fragment is also retained.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 57 of 119
[0226] The filtering can remove assay artifacts in the plurality of nucleic acid methylation fragments. The removal of assay artifacts can comprise removing sequence reads obtained from sequenced hybridization probes and / or sequence reads obtained from sequences that failed to undergo conversion during bisulfite conversion. The filtering can remove contaminants (e.g., due to sequencing, nucleic acid isolation, and / or sample preparation).
[0227] The filtering can remove a subset of methylation fragments from the plurality of methylation fragments based on mutual information filtering of the respective methylation fragments against the cancer state across the plurality of training subjects. As an example and not by way of limitation, mutual information can provide a measure of the mutual dependence between two conditions of interest sampled simultaneously. Mutual information can be determined by selecting an independent set of CpG sites (e.g., within all or a portion of a nucleic acid methylation fragment) from one or more datasets and comparing the probability of the methylation states for the set of CpG sites between two sample groups (e.g., subsets and / or groups of genotypic datasets, biological samples, and / or subjects). A mutual information score can denote the probability of the methylation pattern for a first condition versus a second condition at the respective region in the respective frame of the sliding window, thus indicating the discriminative power of the respective region. A mutual information score can be similarly calculated for each region in each frame of the sliding window as it progresses across the selected sets of CpG sites and / or the selected genomic regions. Further details regarding mutual information filtering are disclosed in U.S. Patent Application 17 / 119,606, titled “Cancer Classification using Patch Convolutional Neural Networks,” filed December 11, 2020, which is hereby incorporated herein by reference in its entirety. Hypermethylated Fragments and Hypomethylated Fragments
[0228] In particular embodiments, the analytics system 200 identifies 770 determines hypomethylated fragments or hypermethylated fragments from the filtered set as anomalous fragments. The analytics system 200 identifies hypermethylated fragments having over a threshold number of CpG sites and over a threshold percentage of the CpG sites methylated. The analytics system 200 identifies hypomethylated fragments having over the threshold number of CpG sites and over a threshold percentage of CpG sites unmethylated. Example thresholds for length of fragments (or CpG sites) include more than 3, 4, 5, 6, 7, 8, 9, 10, etc. Example percentage thresholds of methylation or unmethylation include more than 80%, 85%, 90%, or 95%, or any other percentage within the range of 50%-100%.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 58 of 119 8.2. Training of Cancer Classifier
[0229] FIG. 8A illustrates an example flowchart describing a process 800 of training a cancer classifier, according to an embodiment. The analytics system 200 obtains 810 a plurality of training samples each having a set of anomalous fragments and a label of a cancer type. The plurality of training samples can include any combination of samples from healthy individuals with a general label of “non-cancer,” samples from subjects with a general label of “cancer” or a specific label (e.g., “breast cancer,” “lung cancer,” etc.). The training samples from subjects for one cancer type may be termed a cohort for that cancer type or a cancer type cohort.
[0230] The analytics system 200 determines 820, for each training sample, a feature vector based on the set of anomalous fragments of the training sample. The analytics system 200 can calculate an anomaly score for each CpG site in an initial set of CpG sites. The initial set of CpG sites may be all CpG sites in the human genome or some portion thereof – which may be on the order of 104, 105, 106, 107, 108, etc. In particular embodiments, the analytics system 200 defines the anomaly score for the feature vector with a binary scoring based on whether there is an anomalous fragment in the set of anomalous fragments that encompasses the CpG site. In another embodiment, the analytics system 200 defines the anomaly score based on a count of anomalous fragments overlapping the CpG site. In one example, the analytics system 200 may use a trinary scoring assigning a first score for lack of presence of anomalous fragments, a second score for presence of a few anomalous fragments, and a third score for presence of more than a few anomalous fragments. As an example and not by way of limitation, the analytics system 200 counts 5 anomalous fragment in a sample that overlap the CpG site and calculates an anomaly score based on the count of 5. In one or more embodiments, the feature vector further includes one or more features based on the methylation variants.
[0231] Once all anomaly scores are determined for a training sample, the analytics system 200 can determine the feature vector as a vector of elements including, for each element, one of the anomaly scores associated with one of the CpG sites in an initial set. The analytics system 200 can normalize the anomaly scores of the feature vector based on a coverage of the sample. Here, coverage can refer to a median or average sequencing depth over all CpG sites covered by the initial set of CpG sites used in the classifier, or based on the set of anomalous fragments for a given training sample.
[0232] FIG. 8B illustrates an example generation of feature vectors used for training the cancer classifier. As an example, reference is now made to FIG. 8B illustrating a matrix ofACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 59 of 119 training feature vectors 822. In this example, the analytics system 200 has identified CpG sites [S] 826 for consideration in generating feature vectors for the cancer classifier. The analytics system 200 selects training samples [N] 824. The analytics system 200 determines a first anomaly score 828 for a first arbitrary CpG site [s1] to be used in the feature vector for a training sample [n1]. The analytics system 200 checks each anomalous fragment in the set of anomalous fragments. If the analytics system 200 identifies at least one anomalous fragment that includes the first CpG site, then the analytics system 200 determines the first anomaly score 828 for the first CpG site as 1, as illustrated in FIG. 8B. Considering a second arbitrary CpG site [s2], the analytics system 200 similarly checks the set of anomalous fragments for at least one that includes the second CpG site [s2]. If the analytics system 200 does not find any such anomalous fragment that includes the second CpG site, the analytics system 200 determines a second anomaly score 829 for the second CpG site [s2] to be 0, as illustrated in FIG.8B. Once the analytics system 200 determines all the anomaly scores for the initial set of CpG sites, the analytics system 200 determines the feature vector for the first training sample [n1] including the anomaly scores with the feature vector including the first anomaly score 828 of 1 for the first CpG site [s1] and the second anomaly score 829 of 0 for the second CpG site [s2] and subsequent anomaly scores, thus forming a feature vector [1, 0, …].
[0233] Additional approaches to featurization of a sample can be found in: U.S. Application No. 15 / 931,022 entitled “Model-Based Featurization and Classification;” U.S. Application No. 16 / 579,805 entitled “Mixture Model for Targeted Sequencing;” U.S. Application No.16 / 352,602 entitled “Anomalous Fragment Detection and Classification;” and U.S. Application No. 16 / 723,716 entitled “Source of Origin Deconvolution Based on Methylation Fragments in Cell-Free DNA Samples;” all of which are incorporated by reference in their entirety.
[0234] The analytics system 200 may further limit the CpG sites considered for use in the cancer classifier. The analytics system 200 computes 830, for each CpG site in the initial set of CpG sites, an information gain based on the feature vectors of the training samples. From step 820, each training sample has a feature vector that may contain an anomaly score all CpG sites in the initial set of CpG sites which could include up to all CpG sites in the human genome. However, some CpG sites in the initial set of CpG sites may not be as informative as others in distinguishing between cancer types, or may be duplicative with other CpG sites.
[0235] In particular embodiments, the analytics system 200 computes 830 an information gain for each cancer type and for each CpG site in the initial set to determine whether to includeACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 60 of 119 that CpG site in the classifier. The information gain is computed for training samples with a given cancer type compared to all other samples. As an example and not by way of limitation, two random variables ‘anomalous fragment’ (‘AF’) and ‘cancer type’ (‘CT’) are used. In particular embodiments, AF is a binary variable indicating whether there is an anomalous fragment overlapping a given CpG site in a given samples as determined for the anomaly score / feature vector above. CT is a random variable indicating whether the cancer is of a particular type. The analytics system 200 computes the mutual information with respect to CT given AF. That is, how many bits of information about the cancer type are gained if it is known whether there is an anomalous fragment overlapping a particular CpG site. In practice, for a first cancer type, the analytics system 200 computes pairwise mutual information gain against each other cancer type and sums the mutual information gain across all the other cancer types.
[0236] For a given cancer type, the analytics system 200 can use this information to rank CpG sites based on how cancer specific they are. This procedure can be repeated for all cancer types under consideration. If a particular region is commonly anomalously methylated in training samples of a given cancer but not in training samples of other cancer types or in healthy training samples, then CpG sites overlapped by those anomalous fragments can have high information gains for the given cancer type. The ranked CpG sites for each cancer type can be greedily added (selected) 840 to a selected set of CpG sites based on their rank for use in the cancer classifier.
[0237] In additional embodiments, the analytics system 200 may consider other selection criteria for selecting informative CpG sites to be used in the cancer classifier. One selection criterion may be that the selected CpG sites are above a threshold separation from other selected CpG sites. As an example and not by way of limitation, the selected CpG sites are to be over a threshold number of base pairs away from any other selected CpG site (e.g., 100 base pairs), such that CpG sites that are within the threshold separation are not both selected for consideration in the cancer classifier.
[0238] In particular embodiments, according to the selected set of CpG sites from the initial set, the analytics system 200 may modify 850 the feature vectors of the training samples as needed. As an example and not by way of limitation, the analytics system 200 may truncate feature vectors to remove anomaly scores corresponding to CpG sites not in the selected set of CpG sites.
[0239] With the feature vectors of the training samples, the analytics system 200 may train the cancer classifier in any of a number of ways. The feature vectors may correspond to theACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 61 of 119 initial set of CpG sites from step 820 or to the selected set of CpG sites from step 850. In particular embodiments, the analytics system 200 trains 860 a binary cancer classifier to distinguish between cancer and non-cancer based on the feature vectors of the training samples. In this manner, the analytics system 200 uses training samples that include both non-cancer samples from healthy individuals and cancer samples from subjects. Each training sample can have one of the two labels “cancer” or “non-cancer.” In this embodiment, the classifier outputs a cancer prediction indicating the likelihood of the presence or absence of cancer.
[0240] In another embodiment, the analytics system 200 trains 870 a multiclass cancer classifier to distinguish between many cancer types (also referred to as tissue of origin (TOO) or cancer signal origin CSO labels). Cancer types can include one or more cancers and may include a non-cancer type (may also include any additional other diseases or genetic disorders, etc.). To do so, the analytics system 200 can use the cancer type cohorts and may also include or not include a non-cancer type cohort. In this multi-cancer embodiment, the cancer classifier is trained to determine a cancer prediction (or, more specifically, a TOO or CSO prediction) that comprises a prediction value for each of the cancer types being classified for. The prediction values may correspond to a likelihood that a given training sample (and during inference, a test sample) has each of the cancer types. In one implementation, the prediction values are scored between 0 and 100, wherein the cumulation of the prediction values equals 100. As an example and not by way of limitation, the cancer classifier returns a cancer prediction including a prediction value for breast cancer, lung cancer, and non-cancer. As an example and not by way of limitation, the classifier can return a cancer prediction that a test sample is 65% likelihood of breast cancer, 25% likelihood of lung cancer, and 10% likelihood of non-cancer. The analytics system 200 may further evaluate the prediction values to generate a prediction of a presence of one or more cancers in the sample, also may be referred to as a TOO or CSO prediction indicating one or more TOO labels, e.g., a first TOO label with the highest prediction value, a second TOO label with the second highest prediction value, etc. Continuing with the example above and given the percentages, in this example the system may determine that the sample has breast cancer given that breast cancer has the highest likelihood.
[0241] In both embodiments, the analytics system 200 trains the cancer classifier by inputting sets of training samples with their feature vectors into the cancer classifier and adjusting classification parameters so that a function of the classifier accurately relates the training feature vectors to their corresponding label. The analytics system 200 may group the training samples into sets of one or more training samples for iterative batch training of theACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 62 of 119 cancer classifier. After inputting all sets of training samples including their training feature vectors and adjusting the classification parameters, the cancer classifier can be sufficiently trained to label test samples according to their feature vector within some margin of error. The analytics system 200 may train the cancer classifier according to any one of a number of methods. As an example, the binary cancer classifier may be a L2-regularized logistic regression classifier that is trained using a log-loss function. As another example, the multi- cancer classifier may be a multinomial logistic regression. In practice either type of cancer classifier may be trained using other techniques. These techniques are numerous including potential use of kernel methods, random forest classifier, a mixture model, an autoencoder model, machine learning algorithms such as multilayer neural networks, etc.
[0242] The classifier can include a logistic regression algorithm, a neural network algorithm, a support vector machine algorithm, a Naive Bayes algorithm, a nearest neighbor algorithm, a boosted trees algorithm, a random forest algorithm, a decision tree algorithm, a multinomial logistic regression algorithm, a linear model, or a linear regression algorithm. 8.3. Deployment of Cancer Classifier
[0243] During use of the cancer classifier, the analytics system 200 can obtain a test sample from a subject of unknown cancer type. The analytics system 200 may process the test sample comprised of DNA molecules with any combination of the processes 400 and 730 to achieve a set of anomalous fragments. The analytics system 200 can determine a test feature vector for use by the cancer classifier according to similar principles discussed in the process 800. The analytics system 200 can calculate an anomaly score for each CpG site in a plurality of CpG sites in use by the cancer classifier. As an example and not by way of limitation, the cancer classifier receives as input feature vectors inclusive of anomaly scores for 1,000 selected CpG sites. The analytics system 200 can thus determine a test feature vector inclusive of anomaly scores for the 1,000 selected CpG sites based on the set of anomalous fragments. The analytics system 200 can calculate the anomaly scores in a same manner as the training samples. In particular embodiments, the analytics system 200 defines the anomaly score as a binary score based on whether there is a hypermethylated or hypomethylated fragment in the set of anomalous fragments that encompasses the CpG site. The analytics system 200 may generate the test feature vector including one or more features based on the methylation variants.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 63 of 119
[0244] The analytics system 200 can then input the test feature vector into the cancer classifier. The function of the cancer classifier can then generate a cancer prediction based on the classification parameters trained in the process 900 and the test feature vector. In the first manner, the cancer prediction can be binary and selected from a group consisting of “cancer” or non-cancer;” in the second manner, the cancer prediction is selected from a group of many cancer types and “non-cancer.” In additional embodiments, the cancer prediction has predictions values for each of the many cancer types. Moreover, the analytics system 200 may determine that the test sample is most likely to be of one of the cancer types. Following the example above with the cancer prediction for a test sample as 65% likelihood of breast cancer, 25% likelihood of lung cancer, and 10% likelihood of non-cancer, the analytics system 200 may determine that the test sample is most likely to have breast cancer. In another example, where the cancer prediction is binary as 60% likelihood of non-cancer and 40% likelihood of cancer, the analytics system 200 determines that the test sample is most likely not to have cancer. In additional embodiments, the cancer prediction with the highest likelihood may still be compared against a threshold (e.g., 40%, 50%, 60%, 70%) in order to call the test subject as having that cancer type. If the cancer prediction with the highest likelihood does not surpass that threshold, the analytics system 200 may return an inconclusive result.
[0245] In additional embodiments, the analytics system 200 chains a cancer classifier trained in step 860 of the process 800 with another cancer classifier trained in step 870 or the process 800. The analytics system 200 can input the test feature vector into the cancer classifier trained as a binary classifier in step 860 of the process 800. The analytics system 200 can receive an output of a cancer prediction. The cancer prediction may be binary as to whether the test subject likely has or likely does not have cancer. In other implementations, the cancer prediction includes prediction values that describe likelihood of cancer and likelihood of non- cancer. As an example and not by way of limitation, the cancer prediction has a cancer prediction value of 85% and the non-cancer prediction value of 15%. The analytics system 200 may determine the test subject to likely have cancer. Once the analytics system 200 determines a test subject is likely to have cancer, the analytics system 200 may input the test feature vector into a multiclass cancer classifier trained to distinguish between different cancer types. The multiclass cancer classifier can receive the test feature vector and returns a cancer prediction of a cancer type of the plurality of cancer types. As an example and not by way of limitation, the multiclass cancer classifier provides a cancer prediction specifying that the test subject is most likely to have ovarian cancer. In another implementation, the multiclass cancer classifierACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 64 of 119 provides a prediction value for each cancer type of the plurality of cancer types. As an example and not by way of limitation, a cancer prediction may include a breast cancer type prediction value of 40%, a colorectal cancer type prediction value of 15%, and a liver cancer prediction value of 45%.
[0246] According to generalized embodiment of binary cancer classification, the analytics system 200 can determine a cancer score for a test sample based on the test sample’s sequencing data (e.g., methylation sequencing data, SNP sequencing data, other DNA sequencing data, RNA sequencing data, etc.). The analytics system 200 can compare the cancer score for the test sample against a binary threshold cutoff for predicting whether the test sample likely has cancer. The binary threshold cutoff can be tuned using TOO thresholding based on one or more TOO subtype classes. The analytics system 200 may further generate a feature vector for the test sample for use in the multiclass cancer classifier to determine a cancer prediction indicating one or more likely cancer types.
[0247] The classifier may be used to determine the disease state of a test subject, e.g., a subject whose disease status is unknown. The method can include obtaining a test genomic data construct (e.g., single time point test data), in electronic form, that includes a value for each genomic characteristic in the plurality of genomic characteristics of a corresponding plurality of nucleic acid fragments in a biological sample obtained from a test subject. The method can then include applying the test genomic data construct to the test classifier to thereby determine the state of the disease condition in the test subject. The test subject may not be previously diagnosed with the disease condition.
[0248] The classifier can be a temporal classifier that uses at least (i) a first test genomic data construct generated from a first biological sample acquired from a test subject at a first point in time, and (ii) a second test genomic data construct generated from a second biological sample acquired from a test subject at a second point in time.
[0249] The trained classifier can be used to determine the disease state of a test subject, e.g., a subject whose disease status is unknown. In this case, the method can include obtaining a test time-series data set, in electronic form, for a test subject, where the test time-series data set includes, for each respective time point in a plurality of time points, a corresponding test genotypic data construct including values for the plurality of genotypic characteristics of a corresponding plurality of nucleic acid fragments in a corresponding biological sample obtained from the test subject at the respective time point, and for each respective pair of consecutive time points in the plurality of time points, an indication of the length of timeACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 65 of 119 between the respective pair of consecutive time points. The method can then include applying the test genotypic data construct to the test classifier to thereby determine the state of the disease condition in the test subject. The test subject may not be previously diagnosed with the disease condition. 9. Cancer Assay Panel
[0250] In particular embodiments, the predictive cancer models described herein use samples enriched using a cancer assay panel comprising a plurality of probes or a plurality of probe pairs. A number of targeted cancer assay panels are known in the art, for example, as describe in WO 2019 / 195268 filed April 2, 2019, PCT / US2019 / 053509 filed September 27, 2019 and PCT / US2020 / 015082 filed January 24, 2020 (which are incorporated herein by reference). As an example and not by way of limitation, the cancer assay panel can be designed to include a plurality of probes (or probe pairs) that can capture fragments that can together provide information relevant to diagnosis of cancer. In particular embodiments, a panel includes at least 50, 100, 500, 1,000, 2,000, 2,500, 5,000, 6,000, 7,500, 10,000, 15,000, 20,000, 25,000, or 50,000 pairs of probes. In other embodiments, a panel includes at least 500, 1,000, 2,000, 5,000, 10,000, 12,000, 15,000, 20,000, 30,000, 40,000, 50,000, or 100,000 probes. The plurality of probes together can comprise at least 0.1 million, 0.2 million, 0.4 million, 0.6 million, 0.8 million, 1 million, 2 million, 3 million, 4 million, 5 million, 6 million, 7 million, 8 million, 9 million, or 10 million nucleotides. The probes (or probe pairs) are specifically designed to target one or more genomic regions differentially methylated in cancer and non- cancer samples. The target genomic regions can be selected to maximize classification accuracy, subject to a size budget (which is determined by sequencing budget and desired depth of sequencing).
[0251] Samples enriched using a cancer assay panel can be subject to targeted sequencing. Samples enriched using the cancer assay panel can be used to detect the presence or absence of cancer generally and / or provide a cancer classification such as cancer type, stage of cancer such as I, II, III, or IV, or provide the tissue of origin where the cancer is believed to originate. Depending on the purpose, a panel can include probes (or probe pairs) targeting genomic regions differentially methylated between general cancerous (pan-cancer) samples and non- cancerous samples, or only in cancerous samples with a specific cancer type (e.g., lung cancer- specific targets). Specifically, a cancer assay panel is designed based on bisulfite sequencingACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 66 of 119 data generated from the cell-free DNA (cfDNA) or genomic DNA (gDNA) from cancer and / or non-cancer individuals.
[0252] In particular embodiments, the cancer assay panel designed by methods provided herein comprises at least 1,000 pairs of probes, each pair of which comprises two probes configured to overlap each other by an overlapping sequence comprising a 30-nucleotide fragment. The 30-nucleotide fragment comprises at least five CpG sites, wherein at least 80% of the at least five CpG sites are either CpG or UpG. The 30-nucleotide fragment is configured to bind to one or more genomic regions in cancerous samples, wherein the one or more genomic regions have at least five methylation sites with an abnormal methylation pattern. Another cancer assay panel comprises at least 2,000 probes, each of which is designed as a hybridization probe complimentary to one or more genomic regions. Each of the genomic regions is selected based on the criteria that it comprises (i) at least 30 nucleotides, and (ii) at least five methylation sites, wherein the at least five methylation sites have an abnormal methylation pattern and are either hypomethylated or hypermethylated.
[0253] Each of the probes (or probe pairs) is designed to target one or more target genomic regions. The target genomic regions are selected based on several criteria designed to increase selective enriching of relevant cfDNA fragments while decreasing noise and non-specific bindings. As an example and not by way of limitation, a panel can include probes that can selectively bind and enrich cfDNA fragments that are differentially methylated in cancerous samples. In this case, sequencing of the enriched fragments can provide information relevant to diagnosis of cancer. Furthermore, the probes can be designed to target genomic regions that are determined to have an abnormal methylation pattern and / or hypermethylation or hypomethylation patterns to provide additional selectivity and specificity of the detection. As an example and not by way of limitation, genomic regions can be selected when the genomic regions have a methylation pattern with a low p-value according to a Markov model trained on a set of non-cancerous samples, that additionally cover at least 5 CpG’s, 90% of which are either methylated or unmethylated. In other embodiments, genomic regions can be selected utilizing mixture models, as described herein.
[0254] Each of the probes (or probe pairs) can target genomic regions comprising at least 25bp, 30bp, 35bp, 40bp, 45bp, 50bp, 60bp, 70bp, 80bp, or 90bp. The genomic regions can be selected by containing less than 20, 15, 10, 8, or 6 methylation sites. The genomic regions can be selected when at least 80, 85, 90, 92, 95, or 98% of the at least five methylation (e.g., CpG) sites are either methylated or unmethylated in non-cancerous or cancerous samples.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 67 of 119
[0255] Genomic regions may be further filtered to select only those that are likely to be informative based on their methylation patterns, for example, CpG sites that are differentially methylated between cancerous and non-cancerous samples (e.g., abnormally methylated or unmethylated in cancer versus non-cancer). For the selection, calculation can be performed with respect to each CpG site. In particular embodiments, a first count is determined that is the number of cancer-containing samples (cancer_count) that include a fragment overlapping that CpG, and a second count is determined that is the number of total samples containing fragments overlapping that CpG (total). Genomic regions can be selected based on criteria positively correlated to the number of cancer-containing samples (cancer_count) that include a fragment overlapping that CpG, and inversely correlated with the number of total samples containing fragments overlapping that CpG (total).
[0256] In particular embodiments, the number of non-cancerous samples (nnon-cancer) and the number of cancerous samples (ncancer) having a fragment overlapping a site arecounted. Then the probability that a sample is cancer is estimated, for example as (ncancer + 1) / (ncancer + nnon-cancer + 2). CpG sites by this metric are ranked and greedily added to a panel until is exhausted.
[0257] Depending on whether the assay is intended to be a pan-cancer assay or a single- cancer assay, or depending on what kind of flexibility is desired when picking which CpG sites are contributing to the panel, which samples are used for cancer-count can vary. A panel for diagnosing a specific cancer type (e.g., TOO or CSO) can be designed using a similar process. In this embodiment, for each cancer type, and for each CpG site, the information gain is computed to determine whether to include a probe targeting that CpG site. The information gain is computed for samples with a given cancer type compared to all other samples. As an example and not by way of limitation, two random variables, “AF” and “CT”. “AF” is a binary variable that indicates whether there is an abnormal fragment overlapping a particular CpG site in a particular sample (yes or no). “CT” is a binary random variable indicating whether the cancer is of a particular type (e.g., lung cancer or cancer other than lung). One can compute the mutual information with respect to “CT” given “AF.” That is, how many bits of information about the cancer type (lung vs. non-lung in the example) are gained if one knows whether there is an informative fragment overlapping a particular CpG site. This can be used to rank CpG's based on how specific they are for a particular cancer type (e.g., TOO or CSO). This procedure is repeated for a plurality of cancer types. As an example and not by way of limitation, if a particular region is commonly differentially methylated only in lung cancer (and not otherACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 68 of 119 cancer types or non-cancer), CpG’s in that region would tend to have high information gains for lung cancer. For each cancer type, CpG sites ranked by this information gain metric, and then greedily added to a panel until the size budget for that cancer type was exhausted.
[0258] Further filtration can be performed to select target genomic regions that have off- target genomic regions less than a threshold value. As an example and not by way of limitation, a genomic region is selected only when there are less than 15, 10 or 8 off-target genomic regions. In other cases, filtration is performed to remove genomic regions when the sequence of the target genomic regions appears more than 5, 10, 15, 20, 25, or 30 times in a genome. Further filtration can be performed to select target genomic regions when a sequence, 90%, 95%, 98% or 99% homologous to the target genomic regions, appear less than 15, 10 or 8 times in a genome, or to remove target genomic regions when the sequence, 90%, 95%, 98% or 99% homologous to the target genomic regions, appear more than 5, 10, 15, 20, 25, or 30 times in a genome. This is for excluding repetitive probes that can pull down off-target fragments, which are not desired and can impact assay efficiency.
[0259] In particular embodiments, fragment-probe overlap of at least 45bp was demonstrated to be required to achieve a non-negligible amount of pulldown (though this number can be different depending on assay details). Furthermore, it has been suggested that more than a 10% mismatch rate between the probe and fragment sequences in the region of overlap is sufficient to greatly disrupt binding, and thus pulldown efficiency. Therefore, sequences that can align to the probe along at least 45bp with at least a 90% match rate are candidates for off-target pulldown. Thus, the number of such regions may be scored. The best probes have a score of 1, meaning they match in only one place (the intended target region). Probes with a low score (say, less than 5 or 10) are accepted, but any probes above the score are discarded. Other cutoff values can be used for specific samples.
[0260] In particular embodiments, the selected target genomic regions can be located in various positions in a genome, including but not limited to exons, introns, intergenic regions, and other parts. In particular embodiments, probes targeting non-human genomic regions, such as those targeting viral genomic regions, can be added. 10. Applications
[0261] BRCA1 promoter methylation (BRCA1meth) is involved in the genesis of a significant fraction of triple negative breast (TNBC) and ovarian cancers (OvCa) and has therapeutic predictive value in these cancer types. BRCA1meth is established during earlyACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 69 of 119 embryonic development as a constitutional event and detection of BRCA1meth in the peripheral blood cells of unaffected women is associated with increased risk of developing breast and / or ovarian cancers. Thus, facile and sensitive assessment of BRCA1meth is an important component for both risk assessment and, once cancers are established, the therapeutic management of these cancer types. The present disclosure describes a targeted based methylation platform as a quantitative method to measure the level of BRCA1meth in circulating cell-free DNA (cfBRCA1meth) from plasma samples of individuals with and without cancer, allowing the detection of BRCA1 methylated alleles down to a frequency of ~ 0.001. Using this method, constitutional cfBRCA1meth was found in 4 % (113 / 2790) of individuals without cancer. Interestingly, the prevalence of cfBRCA1meth was significantly higher in females compared to males (4.8% vs.2.9%, p = 0.016), suggesting that the dynamics of BRCA1 promoter methylation and its consequences for the development of cancer may differ by sex. =. In a pan-cancer cohort comprising 2,958 patients and representing 15 different tumor types, only TNBC and OvCa had significantly higher BRCA1meth prevalence compared to non-cancer, at 15.2% (p = 0.001) and 13.9% (p = 0.014). Furthermore, the fraction of cfBRCA1meth in these cases correlated with tumor fraction as measured by targeted methylation platform. Thus, cfBRCA1meth assessment using the presently disclosed methods provides a single quantitative platform that can provide prognostic insight into cancer risk and therapeutic response.
[0262] In particular embodiments, the methods, analytics systems 200 and / or classifier of the present disclosure can be used to detect the presence of cancer, monitor cancer progression or recurrence, monitor therapeutic response or effectiveness, determine a presence or monitor minimum residual disease (MRD), or any combination thereof. As an example and not by way of limitation, as described herein, a classifier can be used to generate a probability score (e.g., from 0 to 100) describing a likelihood that a test feature vector is from a subject with cancer. In particular embodiments, the probability score is compared to a threshold probability to determine whether or not the subject has cancer. In other embodiments, the likelihood or probability score can be assessed at multiple different time points (e.g., before or after treatment) to monitor disease progression or to monitor treatment effectiveness (e.g., therapeutic efficacy). In still other embodiments, the likelihood or probability score can be used to make or influence a clinical decision (e.g., diagnosis of cancer, treatment selection, assessment of treatment effectiveness, etc.). In particular embodiments, if the probability score exceeds a threshold, a physician can prescribe an appropriate treatment.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 70 of 119 10.1. Early Detection of Cancer
[0263] In particular embodiments, the methods and / or classifier of the present disclosure are used to detect the presence or absence of cancer in a subject suspected of having cancer. As an example and not by way of limitation, a classifier can be used to determine a cancer prediction describing a likelihood that a test feature vector is from a subject that has cancer.
[0264] In particular embodiments, the methods and / or classifier of the present disclosure are used to predict the origin of the cancer. In particular embodiments, the cancer signal of origin (CSO) prediction is generated by the one or more classifiers associated with the analytics system 200 to predict cancer subtype from cfDNA.
[0265] In particular embodiments, a cancer prediction is a likelihood (e.g., scored between 0 and 100) for whether the test sample has cancer (i.e., binary classification). Thus, the analytics system 200 may determine a threshold for determining whether a test subject has cancer. As an example and not by way of limitation, a cancer prediction of greater than or equal to 60 can indicate that the subject has cancer. In still other embodiments, a cancer prediction greater than or equal to 65, greater than or equal to 70, greater than or equal to 75, greater than or equal to 80, greater than or equal to 85, greater than or equal to 90, or greater than or equal to 95 indicates that the subject has cancer. In other embodiments, the cancer prediction can indicate the severity of disease. As an example and not by way of limitation, a cancer prediction of 80 may indicate a more severe form, or later stage, of cancer compared to a cancer prediction below 80 (e.g., a probability score of 70). Similarly, an increase in the cancer prediction over time (e.g., determined by classifying test feature vectors from multiple samples from the same subject taken at two or more time points) can indicate disease progression or a decrease in the cancer prediction over time can indicate successful treatment.
[0266] In another embodiment, a cancer prediction comprises many prediction values, wherein each of a plurality of cancer types being classified (i.e., multiclass classification) for has a prediction value (e.g., scored between 0 and 100). The prediction values may correspond to a likelihood that a given training sample (and during inference, training sample) has each of the cancer types. The analytics system 200 may identify the cancer type that has the highest prediction value and indicate that the test subject likely has that cancer type. In other embodiments, the analytics system 200 further compares the highest prediction value to a threshold value (e.g., 50, 55, 60, 65, 70, 75, 80, 85, etc.) to determine that the test subject likely has that cancer type. In other embodiments, a prediction value can also indicate the severity ofACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 71 of 119 disease. As an example and not by way of limitation, a prediction value greater than 80 may indicate a more severe form, or later stage, of cancer compared to a prediction value of 60. Similarly, an increase in the prediction value over time (e.g., determined by classifying test feature vectors from multiple samples from the same subject taken at two or more time points) can indicate disease progression or a decrease in the prediction value over time can indicate successful treatment.
[0267] In particular embodiments, the methods and systems of the present disclosure can be trained to detect or classify multiple cancer indications. As an example and not by way of limitation, the methods, systems and classifiers of the present disclosure can be used to detect the presence of one or more, two or more, three or more, five or more, ten or more, fifteen or more, or twenty or more different types of cancer.
[0268] Examples of cancers that can be detected using the methods, systems and classifiers of the present disclosure include carcinoma, lymphoma, blastoma, sarcoma, and leukemia or lymphoid malignancies. More particular examples of such cancers include, but are not limited to, squamous cell cancer (e.g., epithelial squamous cell cancer), skin carcinoma, melanoma, lung cancer, including small-cell lung cancer, non-small cell lung cancer (“NSCLC”), adenocarcinoma of the lung and squamous carcinoma of the lung, cancer of the peritoneum, gastric or stomach cancer including gastrointestinal cancer, pancreatic cancer (e.g., pancreatic ductal adenocarcinoma), cervical cancer, ovarian cancer (e.g., high grade serous ovarian carcinoma), liver cancer (e.g., hepatocellular carcinoma (HCC)), hepatoma, hepatic carcinoma, bladder cancer (e.g., urothelial bladder cancer), testicular (germ cell tumor) cancer, breast cancer (e.g., HER2 positive, HER2 negative, and triple negative breast cancer), brain cancer (e.g., astrocytoma, glioma (e.g., glioblastoma)), colon cancer, rectal cancer, colorectal cancer, endometrial or uterine carcinoma, salivary gland carcinoma, kidney or renal cancer (e.g., renal cell carcinoma, nephroblastoma or Wilms’ tumor), prostate cancer, vulval cancer, thyroid cancer, anal carcinoma, penile carcinoma, head and neck cancer, esophageal carcinoma, and nasopharyngeal carcinoma (NPC). Additional examples of cancers include, without limitation, retinoblastoma, thecoma, arrhenoblastoma, hematological malignancies, including but not limited to non-Hodgkin's lymphoma (NHL), multiple myeloma and acute hematological malignancies, endometriosis, fibrosarcoma, choriocarcinoma, laryngeal carcinomas, Kaposi's sarcoma, Schwannoma, oligodendroglioma, neuroblastomas, rhabdomyosarcoma, osteogenic sarcoma, leiomyosarcoma, and urinary tract carcinomas.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 72 of 119
[0269] In particular embodiments, the cancer is one or more of anorectal cancer, bladder cancer, breast cancer, cervical cancer, colorectal cancer, esophageal cancer, gastric cancer, head & neck cancer, hepatobiliary cancer, leukemia, lung cancer, lymphoma, melanoma, multiple myeloma, ovarian cancer, pancreatic cancer, prostate cancer, renal cancer, thyroid cancer, uterine cancer, or any combination thereof.
[0270] In particular embodiments, the one or more cancer can be a “high-signal” cancer (defined as cancers with greater than 50% 5-year cancer-specific mortality), such as anorectal, colorectal, esophageal, head & neck, hepatobiliary, lung, ovarian, and pancreatic cancers, as well as lymphoma and multiple myeloma. High-signal cancers tend to be more aggressive and typically have an above-average cell-free nucleic acid concentration in test samples obtained from a patient. 10.2. Cancer and Treatment Monitoring
[0271] In particular embodiments, the cancer prediction can be assessed at multiple different time points (e.g., or before or after treatment) to monitor disease progression or to monitor treatment effectiveness (e.g., therapeutic efficacy). As an example and not by way of limitation, the present disclosure includes methods that involve obtaining a first sample (e.g., a first plasma cfDNA sample) from a cancer patient at a first time point, determining a first cancer prediction therefrom (as described herein), obtaining a second test sample (e.g., a second plasma cfDNA sample) from the cancer patient at a second time point, and determining a second cancer prediction therefrom (as described herein).
[0272] In particular embodiments, the first time point is before a cancer treatment (e.g., before a resection surgery or a therapeutic intervention), and the second time point is after a cancer treatment (e.g., after a resection surgery or therapeutic intervention), and the classifier is utilized to monitor the effectiveness of the treatment. As an example and not by way of limitation, if the second cancer prediction decreases compared to the first cancer prediction, then the treatment is considered to have been successful. However, if the second cancer prediction increases compared to the first cancer prediction, then the treatment is considered to have not been successful. In other embodiments, both the first and second time points are before a cancer treatment (e.g., before a resection surgery or a therapeutic intervention). In still other embodiments, both the first and the second time points are after a cancer treatment (e.g., after a resection surgery or a therapeutic intervention). In still other embodiments, cfDNA samples may be obtained from a cancer patient at a first and second time point and analyzed. e.g., toACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 73 of 119 monitor cancer progression, to determine if a cancer is in remission (e.g., after treatment), to monitor or detect residual disease or recurrence of disease, or to monitor treatment (e.g., therapeutic) efficacy.
[0273] Those of skill in the art will readily appreciate that test samples can be obtained from a cancer patient over any desired set of time points and analyzed according to the methods of the present disclosure to monitor a cancer state in the patient. In particular embodiments, the first and second time points are separated by an amount of time that ranges from about 15 minutes up to about 30 years, such as about 30 minutes, such as about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or about 24 hours, such as about 1, 2, 3, 4, 5, 10, 15, 20, 25 or about 50 days, or such as about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 months, or such as about 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, 10.5, 11, 11.5, 12, 12.5, 13, 13.5, 14, 14.5, 15, 15.5, 16, 16.5, 17, 17.5, 18, 18.5, 19, 19.5, 20, 20.5, 21, 21.5, 22, 22.5, 23, 23.5, 24, 24.5, 25, 25.5, 26, 26.5, 27, 27.5, 28, 28.5, 29, 29.5 or about 30 years. In other embodiments, test samples can be obtained from the patient at least once every 5 months, at least once every 6 months, at least once a year, at least once every 2 years, at least once every 3 years, at least once every 4 years, or at least once every 5 years. 10.3. Treatment
[0274] In particular embodiments, the cancer prediction can be used to make or influence a clinical decision (e.g., diagnosis of cancer, treatment selection, assessment of treatment effectiveness, etc.). As an example and not by way of limitation, if the cancer prediction (e.g., for cancer or for a particular cancer type) exceeds a threshold, a physician can prescribe an appropriate treatment (e.g., a resection surgery, radiation therapy, chemotherapy, and / or immunotherapy). The physician can prescribe an appropriate treatment based on analyses performed by the analytics system 200200, e.g., the analyses 140 of FIG.1.
[0275] BRCA1 deficiency induced by BRCA1 promoter methylation (BRCA1meth) is found in 20-27% of patients across ovarian cancer (OvCa) and triple negative breast cancer (TNBC) cohorts (Koboldt et al., PMID: 23000897) (Menghi et al., PMID: 35857626) (patch et al., PMID: 26017449) (Staaf et al., PMID: 31570822) and is considered causative in the development of these cancers (Lonning et al., PMID: 36074460). Furthermore, tumors with BRCA1meth and pathogenic mutations in BRCA1 (BRCA1mut) have identical genomic scars indicative of homologous recombination (HR) deficiency (Menghi et al., PMID: 35857626) (Glodzik et al., PMID: 32719340). However, whereas BRCA1mut tumors are highly sensitiveACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 74 of 119 to alkylating agents (such as platinum-based chemotherapy) and PARP inhibitors, BRCA1meth tumors appear to be associated with poorer outcomes in the clinic despite initial sensitivity (Menghi et al., PMID: 35857626) (Sun et al., PMID: 24788697) (Tutt et al., PMID: 29713086). This poor response to alkylating agents is caused by the allelic conversion from a homozygously methylated, silenced BRCA1 promoter to a heterozygously methylated promoter that restores BRCA1 expression and HR competency (Menghi et al., PMID: 35857626). This single locus switch is responsible for platinum resistance even after one treatment cycle. Thus, the presence of BRCA1 promoter methylation and its dynamic changes throughout treatment is a prognostic and predictive biomarker of therapeutic response for patients with TNBC and / or OvCa.
[0276] BRCA1 promoter has been found to be methylated in normal somatic cells from the peripheral blood in between 5-9% of women without cancer, a phenomenon known as “primary constitutional epimutation” (Lonning et al., PMID: 40179326). Constitutional BRCA1meth typically affects one allele (i.e. it is heterozygous) and it presents in a relatively small fraction of somatic cells: the median fraction of BRCA1 methylated allele ranges between 0.01-0.05 (on a scale between 0.0 to 1.0), with rare cases of high values of up to 0.3 (Lonning et al., PMID: 36074460) (Azzollini et al., PMID: 30634417) (Gupta et al., PMID: 25376744) (Lonning et al., PMID: 29335712) (Snell et al., PMID: 18269736) (Wong et al., PMID: 20978122). Importantly, it has been found that women with constitutional BRCA1meth have at least a 2.5- fold increased risk of developing TNBC and at least a 1.8-fold increased risk for developing OvCa, with both considered underestimates (Lonning et al., PMID: 36074460) (Lonning et al., PMID: 29335712) (Wong et al., PMID: 20978122). Since BRCA1 promoter methylation is detected in cord blood samples and in adult tissues from different germinal layers, and because of the intra-individual consistency of the affected allele, it is proposed that constitutional BRCA1meth occurs as a single event during early embryonic development and that it is maintained as mosaicism into adulthood, with individuals carrying cells both with and without BRCA1 promoter methylation. The direct role of constitutional BRCA1meth in the genesis of TNBC and OvCa was observed in two cases of familial breast and OvCa, where an inherited sequence mutation in the 5’ untranslated region of the BRCA1 gene leads to BRCA1 promoter methylation in cis, a phenomenon known as a secondary constitutional epimutation; i.e. an epigenetic change associated with cis-acting genetic mutation (Lonning et al., PMID: 40179326). This results in the silencing of the mutated BRCA1 allele and consequently haploinsufficiency at the BRCA1 locus. The pedigrees of these two families show a clearACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 75 of 119 inheritance pattern, whereby individuals affected by breast or ovarian cancers carried the 5’ UTR mutation and its associated BRCA1meth status. Allelic concordance of methylation between DNA isolated from blood, buccal mucosa and hair follicles, and from tumor tissue from the same individuals as well as loss of the non-methylated allele in the cancer samples examined points to BRCA1meth as the oncogenic driver in these families (Evans et al., PMID: 30075112). Given the impact of BRCA1 promoter methylation (BRCA1meth) on cancer risk assessment and prediction of therapeutic outcomes for patients with breast and ovarian cancers, a facile and quantitative detection of BRCA1meth that does not require tissue biopsy is an important biomarker in clinical cancer medicine.
[0277] Promoter methylation has been shown to silence the expression of genes highly relevant in oncology including MLH1, MLH2, CDKN2A, CDKN2B, RB1, E2F, RAD51C, TERT, MGMT, and BRCA1 (Geissler et al., PMID: 3829277). Several of these genes have definitive roles in defining optimal therapy: MLH1 and MLH2 silencing leads to increased tumor mutational burden and predicts response to immune checkpoint inhibitors, MGMT gene methylation is associated with improved response to alkylator therapy in glioblastomas; silencing of RAD51C and BRCA1 through promoter methylation is associated with homologous recombination deficiency (Glodzik et al., PMID: 32719340) and a likely increased response to certain alkylator chemotherapies and PARP inhibitors. The unique aspect of BRCA1 promoter methylation is that it accounts for a significant proportion of TNBC (~20%) and OvCa (~15%) cases. Despite the initial sensitivity to these agents, BRCA1meth has been associated with worse survival in both disorders likely due to the reversal of BRCA1 promoter methylation and restoration of BRCA1 expression after treatment (Menghi et al., PMID: 35857626). Given that constitutional BRCA1 methylation, even when it manifests at a low degree of mosaicism, is associated with an increased risk of incident breast and ovarian cancers (Lonning et al., PMID: 36074460), accurate, reproducible and quantitative tests for BRCA1meth will have clinical utility in risk assessment and therapeutic guidance.
[0278] When assessing cfBRCA1meth in cancer patients, a similar frequency of BRCA1meth was observed in CCGA samples compared to TCGA samples from both patients with TNBC (CCGA: 15.2%, TCGA 14%) and OvCa (CCGA: 13.9%, TCGA: 16%).The advantage of using cfDNA is that the assay does not require an invasive tissue biopsy. In agreement with the results from a broad analysis of the TCGA pancancer dataset, no cancer type other than TNBC and OvCa in the CCGA cancer cohort exhibited significantly higher BRCA1meth prevalence when compared to healthy individuals. The presently disclosedACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 76 of 119 cfBRCA1meth approach can be run in conjunction with a multi-cancer early detection (MCED) assay, allowing the analysis to be done in conjunction with TMeF, a metric of tumor fraction (Melton et al., PMID: 38201510). Although the constitutional BRCA1meth is mosaic and seen in allelic fractions usually lower than 0.05, cancer cells driven by BRCA1meth are typically homozygous or hemizygous for the silenced allele (Lonning et al., PMID: 36074460). Moreover, cancers are known to extrude DNA into the plasma at higher rates than normal cells commensurate with tumor mass (Stejskal et al., PMID: 36681803). These results indicate that circulating cell free DNA can be used to accurately and quantitatively detect clinically meaningful methylation at the BRCA1 promoter in TNBC and ovarian cancers. As BRCA1 disruptions appear to drive a small fraction of cases of other malignancies such as biliary tract, gastric, lung and pancreatic cancers (Momozawa et al., PMID: 35420638), assessing the frequency of cfBRCA1meth may identify a subset of these refractory cancers who have BRCA1 promoter methylation as the mechanism for BRCA1 silencing in their tumors.
[0279] Without being bound by theory, individuals with cfBRCA1meth may benefit from PARP inhibitors and some alkylating agents which have been effective in patients with BRCA1 deficient tumors (Brown et al., PMID: 34904809). Conversely, in TNBC and in OvCa, patients with BRCA1meth in tumors responds more poorly to platinum-based chemotherapy than those with disruptive mutations in the body of the BRCA1 gene (Menghi et al., PMID: 35857626), (Tutt et al., PMID: 29713086), (Bernards et al., PMID: 29233532). Therefore, the presence of cfBRCA1meth in cancer patients is an important prognostic marker for guiding treatment planning.
[0280] A classifier (as described herein) can be used to determine a cancer prediction that a sample feature vector is from a subject that has cancer. In particular embodiments, an appropriate treatment (e.g., resection surgery or therapeutic) is prescribed when the cancer prediction exceeds a threshold. As an example and not by way of limitation, if the cancer prediction is greater than or equal to 60 one or more appropriate treatments are prescribed. In another embodiment, if the cancer prediction is greater than or equal to 65, greater than or equal to 70, greater than or equal to 75, greater than or equal to 80, greater than or equal to 85, greater than or equal to 90, or greater than or equal to 95, one or more appropriate treatments are prescribed. In particular embodiments, the cancer prediction can indicate the severity of disease. An appropriate treatment matching the severity of the disease may then be prescribed.
[0281] In particular embodiments, the treatment is one or more cancer therapeutic agents selected from the group consisting of a chemotherapy agent, a targeted cancer therapy agent, aACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 77 of 119 differentiating therapy agent, a hormone therapy agent, and an immunotherapy agent. As an example and not by way of limitation, the treatment can be one or more chemotherapy agents selected from the group consisting of alkylating agents, antimetabolites, anthracyclines, anti- tumor antibiotics, cytoskeletal disruptors (taxans), topoisomerase inhibitors, mitotic inhibitors, corticosteroids, kinase inhibitors, nucleotide analogs, platinum-based agents and any combination thereof. In particular embodiments, the treatment is one or more targeted cancer therapy agents selected from the group consisting of signal transduction inhibitors (e.g. tyrosine kinase and growth factor receptor inhibitors), histone deacetylase (HDAC) inhibitors, retinoic receptor agonists, proteosome inhibitors, angiogenesis inhibitors, and monoclonal antibody conjugates. In particular embodiments, the treatment is one or more differentiating therapy agents including retinoids, such as tretinoin, alitretinoin and bexarotene. In particular embodiments, the treatment is one or more hormone therapy agents selected from the group consisting of anti-estrogens, aromatase inhibitors, progestins, estrogens, anti-androgens, and GnRH agonists or analogs. In particular embodiments, the treatment is one or more immunotherapy agents selected from the group comprising monoclonal antibody therapies such as rituximab (RITUXAN) and alemtuzumab (CAMPATH), non-specific immunotherapies and adjuvants, such as BCG, interleukin-2 (IL-2), and interferon-alfa, immunomodulating drugs, for instance, thalidomide and lenalidomide (REVLIMID). It is within the capabilities of a skilled physician or oncologist to select an appropriate cancer therapeutic agent based on characteristics such as the type of tumor, cancer stage, previous exposure to cancer treatment or therapeutic agent, and other characteristics of the cancer. 10.4. Multi-Cancer early detection (MCED) test report
[0282] In particular embodiments, a multi-cancer early detection (MCED) test report is generated from a patient sample. In particular embodiments, the patient sample can include blood, plasma, serum, urine, fecal, saliva, other types of bodily fluids, or any combination thereof. In particular embodiments, the extracted sample can comprise cfDNA and / or ctDNA.
[0283] In particular embodiments, a MCED test report includes the fraction of BRCA1 promoter methylation detected in a patient sample using the methods disclosed herein. In particular embodiments, a MCED test report can further include whether a cancer signal has been detected in a patient sample using the methods disclosed herein. In particular embodiments, a MCED test report can further include a cancer signal of origin (CSO) prediction in a patient sample using the methods disclosed herein. In particular embodiments,ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 78 of 119 the CSO prediction can be any cancer selected from the group consisting of anus, bladder, urothelial tract, bone and soft tissue, breast, cervix, colon, rectum, head and neck, hematopoietic and lymphoid organs, kidney, liver, bile duct, lung, melanocyte-containing tissues, skin, ovary, pancreas, gallbladder, prostate, stomach, esophagus, thyroid, and uterus.
[0284] In particular embodiments, a MCED test report includes the fraction of BRCA1 promoter methylation detected in a patient sample regardless of whether a cancer signal or cancer signal of origin (CSO) has been detected in the patient sample.
[0285] The presently disclosed method accurately detects cfBRCA1meth in 95% of plasma samples with a cfBRCA1meth fraction as low as 0.0085. In the analysis of 2,790 non-cancer individuals, cfBRCA1meth was detected in 4.8% and 2.9% of females and males respectively, with a per-sample cfBRCA1meth methylation allelic fraction ranging between 0.04 to 0.37. The prevalence of cfBRCA1meth across non-cancer individuals, the within-sample methylation frequency, and the statistical difference in cfBRCA1meth rates between males and females are highly similar to what has been described in recent studies applying targeted bisulfite next generation sequencing to DNA isolated from peripheral blood cells in healthy individuals (Lonning et al., PMID: 36074460) (Nikolaienko, et al., PMID: 38053165). This validates that cfBRCA1meth is a reasonable alternative to tissue-based bisulfite NGS for the assessment of constitutional BRCA1meth status in populations. This observation of an almost two-fold increased incidence of constitutional BRCA1meth in females over males is consistent with a study of newborns and their parents. Nikolaienko, et al. found the constitutive BRCA1meth rate to be 9% in newborn girls but 4.5% in newborn boys. When the parents of the girls were examined, no correlation was observed between the constitutional BRCA1meth status of the daughters with that of either their mothers or fathers, however, constitutional BRCA1meth was significantly more common in the mothers (8%) than in the fathers (3%). These results confirm sex differential but with larger numbers. While it has been speculated that constitutional BRCA1meth is higher in normal young individuals, the preset disclosure showed no association between age and frequency of constitutional cfBRCA1meth or age with the cfBRCAmeth fraction as a measure of mosaicism across a wide age span (19-86 years)
[0286] Although the majority of non-cancer cases that are positive for cfBRCA1meth (>80%) have frequencies lower than 0.05,a few individuals with higher frequencies, in the range of 0.10-0.37, which are comparable to the cfBRCA1meth frequencies observed in patients with OvCa and TNBC. Thus, a high fraction of cfBRCA1meth alone is not indicative of the presence of cancer. Nevertheless, it has value as a prognostic metric, especially when coupledACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 79 of 119 with cancer detection tools that can quantify tumor fraction. This is useful in the development of new therapeutics targeting homologous recombination deficient tumors, as current screening systems are focused on DNA sequencing and typically overlook BRCA1meth epigenetic silencing via promoter methylation.
[0287] Monitoring non-cancer individuals with high constitutive cfBRCA1meth frequencies will determine whether these individuals are at a greater risk of cancer due to a greater mosaic fraction of somatic cells with BRCA1 promoter methylation. Secondly, cfBRCA1meth data from large cohorts, including population scale studies that utilize the presently disclosed cfBRCA1meth approach were aggregated and will help identify at-risk subpopulations. Lastly, assessment of the presently disclosed cell-free approach alongside sample-level data including HRD status and treatment outcomes will provide prognostic insight into cfBRCA1meth in tumorigenesis and treatment response, respectively. 11. Circulating Cell-Free Genome Atlas Study
[0288] In particular embodiments, each predictive cancer model is trained using a set of training data derived from a training subset of patients of a circulating cell-free genome atlas (CCGA) study (see Clinical Trial.gov Identifier: NCT02889978 (https: / / www.clinicaltrials.gov / ct2 / show / NCT02889978)) and then subsequently tested using a set of testing or validation data derived from a testing or validation subset of patients from the CCGA study.
[0289] The predictive cancer models described herein were trained using a plurality of known cancer types from the circulating cell-free genome atlas (CCGA) study. The CCGA sample set included the following cancer types: breast, lung, prostate, colorectal, renal, uterine, pancreas, esophageal, lymphoma, head and neck, ovarian, hepatobiliary, melanoma, cervical, multiple myeloma, leukemia, thyroid, bladder, gastric, and anorectal. As such, a model can be a multi-cancer model (or a multi-cancer classifier) for detecting of one or more, two or more, three or more, four or more, five or more, ten or more, or 20 or more different types of cancer.
[0290] Predictive cancer models can be trained using a refined set of training data derived from a first subset of patients of the CCGA study and then subsequently tested using a refined set of testing data derived from a second subset of patients from the CCGA study.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 80 of 119 12. Working Example #1: Calling cfBRCA1 Promoter Methylation 12.1. Thresholding to determine BRCA1 promoter methylation
[0291] Appropriate thresholding is a key technical consideration when determining whether promoter methylation is present in a sample. A liberal threshold can lead to too many false positives, while a conservative threshold can lead to false negatives. about a primary concern is having robust thresholds that protect against technical errors that can occur in sample processing; these include sequencing errors due to incomplete bisulfite conversion, PCR errors, etc. A conservative upper bound is estimated for this error by calculating the total number of methylated CpGs observed in samples from a population of individuals without cancer, divided by the total number of CpGs observed in those samples, and denote this as e. This is an overestimate of the actual error rate, since a small proportion of individuals (~5%) will have true, biological methylation at these CpGs that is not a result of technical error. The probability of this type of error, e, is estimated in this way to be 0.002 per CpG.
[0292] Next, the probability of a single molecule being “methylated” due to technical error alone is estimated. A molecule will be considered methylated if at least k CpGs are methylated. The probability that at least k CpGs in a molecule with N CpGs are methylated due to error alone can be expressed using a binomial probability distribution. In other words, each CpG is analogous to a coin toss with probability equal to the error rate. We’re interested in the probability of getting at least k methylated CpGs with error out of N total CpGs in the molecule.
[0293] ^ ^^^_^^^^^ = ^ ∑^ ^^ ^^^^^^^^ (^ ≥ ^ ) =^ ^ ^^^^ (1 − ^)minimum methylated CpG thresholds and varying total number of CpGs on the fragment. Generally, the probability of observing CpG methylation due to technical error increases with the number of CpGs observed on a molecule and decreases as the minimum number of methylated CpGs required (k) increases.
[0295] From the per-molecule error analysis, we chose 25 total CpGs as an upper threshold of the number of CpGs we would observe in a fragment (based on maximum values observed in representative data). Using the per-molecule error rates estimated at this level, we explored the effect of these thresholds on a per-sample level.
[0296] A sample is considered to have methylation if it has at least T molecules methylated amongst M total molecules observed in the sample. We can express the probability of calling a sample methylated due to error using a binomial distribution, with probability equal toACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 81 of 119
[0297] ^^^ ^!" ^_^##^#.
[0298] ^ = ^(^ ≥ ' ) =(^()(1 −(−) ' ^^^ ^!" ^^^ ^!"$%^& ^_^##^# − − ^^ ^!" ^_^##^#
[0301] Generally, as the number of molecules increases in a sample, the probability of observing any molecules with error increases (FIG. 14B). When the threshold is set at >= 4 methylated CpGs per molecule, the probability of calling a sample methylated due to error remains below 10-4(dashed line). This corresponds to no more than 1 out of 10,000 samples being called false positives. Thus, we require at least 4 CpGs to be methylated in a molecule to consider a sample methylated. We expect this parameter to be robust across fragments of varying numbers of CpGs and samples with varying numbers of molecules, within a large cohort of samples. 12.2. Approach to calling cfBRCA1meth
[0302] FIG. 9 illustrates various sources of tissues from which BRCA1 methylation can be detected. There are no true labels in the Circulating Cell-free Genome Atlas (CCGA) data set (e.g., it is not known which patients have evidence of BRCA1 methylation). It is assumed that any detected molecules with methylated CpGs is evidence of methylation, as long as it is unlikely to be the result of assay error. Therefore, cfBRCA1meth will be easier to detect in the samples with deeper sequencing methods. This is analogous to small variant detection in cfDNA, where deeper sequencing provides more opportunities to observe the variant, or in the present case, the methylated molecules.
[0303] FIG. 10 illustrates methylation beta values of various samples across BRCA1 promoter CpG sites. An initial survey indicates some sample types are highly methylated.
[0304] FIG. 11 illustrates exemplary data of methylation across fragments. “Real” methylation signal appears to occur across the majority of CpGs in a given fragment, while noise appears to occur uniformly across fragments without clustering in particular fragments.
[0305] FIG.12 illustrates an exemplary approach for calling cfBRCA1 methylation based on median error modeling.
[0306] FIG.13 illustrates an exemplary approach for calling cfBRCA1 methylation based on max error modeling.
[0307] FIG.14A illustrates exemplary data for incorrectly calling a molecule methylated.
[0308] FIG.14B illustrates exemplary data for incorrectly calling a sample methylated.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 82 of 119
[0309] A set of considerations for determining cfBRCA1 methylation within a sample and for comparing the rates of cfBRCA1 methylation across a subset of samples may include the following:
[0310] 1. Assumption: If you have one error in a molecule, you are not more likely to have another error in that molecule (i.e. error rate is independent across CpGs in a given molecule)
[0311] 1.a. This is not true: molecules can sometimes “skip” bisulfite conversion due to post-bisulfite contamination or issues in the bisulfite reaction. However, these can be identified by looking at the frequency of unconverted CHH / CHG per fragment and can be excluded.
[0312] 2. It will be easier to detect cfBRCA1meth in samples with higher coverage at the BRCA1 promoter locus
[0313] 2.a. Sequencing depth varies across some samples
[0314] 2.b. Coverage is more complicated for a multi-probe panel: it may vary by cancer type / tumor fraction if probes from other parts of the panel are frequently bound
[0315] 2.c. We can restrict to samples with a minimum coverage value in the cfBRCA1meth region when comparing rates of cfBRCA1meth across patient subsets, or we can downsample fragment files (but this will underestimate the true rate of cfBRCA1meth). 12.3. cfBRAC1meth can be detected in a sensitive and quantitative manner
[0316] The presently disclosed cfBRCA1meth assay is a blood-based cfDNA next generation sequencing (NGS) assay that is based on a targeted bisulfite sequencing approach to detect differences in methylation patterns between cancer and non-cancer individuals. While the NGS assay queries over a million CpGs, the cfBRCA1meth assay is focused to one locus, the BRCA1 promoter as described in Methods and (FIG.16). In order to assess the sensitivity and quantitative limits of the cfBRCA1meth assay, a dilution series of commercially acquired contrived samples was analyzed. The linearity and detection limit of the cfBRCA1meth assay was assessed in this dilution series of biochemically methylated contrived samples. These contrived samples have an expected range of cfBRCA1meth fractions from 0.0003 to 1 (where 1 = 100% methylated). The assay shows a strong linear correlation between observed and expected values across the full range of cfBRCA1meth fractions (R = 0.96, p <2.2E-16, FIG. 17A).
[0317] The empirical limit of detection (LoD) for cfBRCA1meth from this dilution series was determined. Statistically, LoD (N%) represents the cfBRCA1meth fraction (as determined by the dilution level of the contrived sample) at which the assay detects N% of samples withACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 83 of 119 known BRCA1meth. The results show that the cfBRCA1meth assay has an LoD95 of 0.0081 cfBRCA1meth fraction (FIG. 17B) (cfBRCA1meth fraction = methylated molecules / total molecules). This states that samples with a cfBRCA1meth allelic fraction as low as 0.0081 would be determined as positive in 95% of the cases. Thus, the presently disclosed subject matter is both sensitive and quantitative in detecting BRCA1 promoter methylation in circulating cfDNA. 12.4. Constitutional cfBRCA1meth is detected in individuals without cancer
[0318] The cfBRCA1meth status of 2,790 individuals without cancer was analyzed using the presently disclosed methods. As part of the Circulating Cell-free Genome Atlas (CCGA) (NCT02889978) study, clinical data were recorded, and plasma samples were collected and processed on a targeted methylation platform from 2,790 individuals without cancer and from 2,849 patients with cancer prior to any treatment. Samples from patients with a variety of solid tumors and heme cancers were considered, including breast, ovarian, cervical, uterine, lung, prostate, and colorectal cancers, among others (FIG.18).
[0319] In this non-caner cohort, 4.1% (113 / 2,790) of all individuals harbored evidence of cfBRCA1meth (Table 1), a percentage that aligns with previous assessments of constitutional BRCA1meth using ultra-deep targeted sequencing of bisulfite-converted DNA from peripheral blood cells (PBCs) of adult individuals (Lonning et al., PMID: 36074460). The cfBRCA1meth fractions in cfBRCA1meth-positive samples ranged from 0.004 to 0.32 (median = 0.015, FIG. 19A), with the majority of cfBRCA1meth-positive individuals (92 / 113, 81%) having a cfBRCA1meth fraction lower than 0.05 (FIG. 19A). The fractional range of cfBRCA1meth is comparable to what was found in previous surveys using peripheral blood cell (PCB) DNA, confirming that constitutional BRCA1meth can be sensitively and accurately detected using cfDNA. A statistically lower incidence of meth positivity was observed in males compared to females, with (31 / 1080 (2.9%) cfBRCA1meth-positive cases in males vs 82 / 1710 (4.8%) cfBRCA1meth-positive cases in females, p=0.016, Table 1. Despite the difference in cfBRCA1meth prevalence, male and female individuals share similar distributions of cfBRCA1meth fractions (FIG. 19A). In particular, FIG. 19A shows the distribution of cfBRCA1meth fractions across individuals without cancer (left). Individual data points are coded to highlight frequencies lower and equal or higher than 0.05. Female and male cfBRCA1meth-positive samples share a similar distribution of cfBRCA1meth fractions (non significant p-value calculated by Wilcoxon signed-rank test). The percentages ofACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 84 of 119 cfBRCA1meth samples with a cfBRCA1meth fraction lower than 0.05 corresponding to the sample sets depicted on the left, are reported in the barplots on the right. FIG. 19B depicts a scatter plot of age (in years) and cfBRCA1meth fractions for all cfBRCA1meth-positive individuals without cancer. A linear regression line is shown and the R value and associated p- value from a Pearson’s correlation test is indicated. The significantly greater BRCA1meth rate in females as compared with males is consistent with data of constitutional BRCA1meth prevalence in umbilical cord blood of newborns which showed a higher constitutive BRCA1meth rate of 9% in females vs 4.5% of males (Nikolaienko, et al., PMID: 38053165). The possibility that cfBRCA1meth may be more prevalent in younger adults was further explored. However, there was no significant age-related change in the rate of cfBRCA1meth within the non-cancer cohort which includes individuals ranging from 20 to 85 years of age (median = 57 years), Table 1. This was true even when the analysis was restricted to each sex, Table 2. Moreover, there was no correlation between age and the fraction of cfBRCA1 meth in this non-cancer cohort, Table 3. Finally, the association between the prevalence of cfBRCA1meth and additional demographic features was explored, however, there was no significant association between cfBRCA1meth prevalence and self-reported race / ethnicity, nor smoking history, Table 1. Table 1: Association of cfBRCA1meth status across demographic characteristics in non- cancer cohort CCGA cohort: non-cancer ( n = 2,790) Characteristic cfBRCA1meth % p-value (No )ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 85 of 119 Black, non-Hispanic 3.9 (9 / 229) Pacific islander 8.1 (5 / 62)of individuals in parentheses. For race / ethnicity and smoking status, groups with fewer than 50 samples were excluded. A Pearson’s Chi-squared test for association was performed and the corresponding p-value is shown, not significant (NS). Table 2: Percent Methylated by Age Bin and Group Age Bin Female Female Male Male N N
[0321] For each group, the % of individuals with cfBRCA1meth is shown, with the number of individuals in parentheses. When prevalences are compared across groups, the Fisher’s Exact test for count data was applied if at least one of the cell counts was less than or equal to 5 or if the total count was less than or equal to 1,000. Otherwise, the Chi-squared test was used. Table 3: Percent Methylated by Age Bin (over / under 40) and Group Age Bin Female Female Male MaleACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 86 of 119 Fisher P- p = 0.01 p = 0.36 p = 1.00 p = 0.21 valueof individuals in parentheses. When prevalences are compared across groups, the Fisher’s Exact test for count data was applied if at least one of the cell counts was less than or equal to 5 or if the total count was less than or equal to 1,000. Otherwise, the Chi-squared test was used. 12.5. cfBRCA1meth is enriched in triple negative breast and ovarian cancers
[0323] The cfBRCA1meth assay was then applied to the subset of cancer patients from the same CCGA study cohort (N = 2,849). Because the rates of BRCA1meth differ between females and males without cancer, the prevalence of BRCA1meth positivity between controls and cancer patients were segregated by sex. It was found that the prevalence of cfBRCA1meth was significantly higher in women with cancer compared to their non-cancer counterpart (6.4% vs. 4.8%, p = 0.05, FIG. 20). On the other hand, the percentage of cfBRCA1meth in men with cancer was lower than what we observed in men without cancer (1.4% vs.2.9%, p = 0.01, FIG. 20). Further analysis found that cfBRCA1meth was more common in younger females with cancer (40 years old or younger, prevalence of cfBRCA1meth = 13%) compared to their older counterparts ( > 40 years old, 6%, Table 3), suggesting that cfBRCA1meth associates with early cancer onset. On the other hand, no association between the rate of cfBRCA1meth and age was found in the male cancer cohort. In addition, when each tumor type was individually assessed, the prevalence of cfBRCA1meth was found to be significantly higher only in female patients with triple negative breast cancers (TNBC) and ovarian cancers (OvCa) (TNBC: 17 / 112 (15.2%), p < 0.001; OvCa: 14 / 101 (13.9%), p < 0.001, FIG. 20). No other cancer had a significant enrichment of cfBRCA1meth patients over non-cancer controls (FIG. 20). Conversely, there was lower cfBRCA1meth rates in prostate cancer 4 / 484 or 0.8%, FIG. 20), which contributes to the reduced cfBRCA1meth rates in the male all-cancer cohort. In FIG.20, the % of individuals with cfBRCA1meth (x-axis) is shown, with error bars denoting the 95% binomial confidence interval. For visualization, error bars extending beyond 25% were truncated, and marked with an arrowhead. Prevalences for individual cancer groups were compared against the sex-matched reference group of non-cancer individuals. These differential rates of cfBRCA1meth across different tumor types are comparable to the intratumoral BRCA1 methylation frequencies found across the same tumor types in the CancerACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 87 of 119 Genome Atlas (TCGA) cohort where BRCA1meth is assessed directly in the tumor DNA using a methylation array technology (Menghi et al., PMID: 35857626) (FIG.21A-B). In FIGs.21A- B, only tumor types also represented in the CCGA dataset were included in this analysis for females (FIG. 21A) and males (FIG. 21B). Error bars represent the 95% binomial confidence intervals for the prevalences. Tumor types showing significant enrichment of BRCA1 methylation compared to the cohort-wide average (i.e., assuming equal prevalence across tumor types); p-value < 0.001, one-sided binomial test statistics.
[0324] It was further observed that cfBRCA1meth fractions in patients with TNBC and OvCa were ~3-fold higher than those observed in individuals without cancer (p < 0.001, FIG. 22A). Unlike the more traditional analysis of PBC DNA, the cfBRCA1meth signal detected in patients with cancer captures both the presence of constitutional BRCA1meth and the BRCA1meth signal originating from BRCA1meth-driven cancer cells. Therefore, it was assessed whether the cfBRCA1meth levels in TNBC and OvCa correlated with the tumor methylation fraction (TMeF), a previously described metric that aggregates the genome-wide signal to provide an estimate of tumor fraction in cfDNA, and which correlates with tumor mass (Melton et al., PMID: 38201510). It was discovered that the cfBRCA1meth fraction was indeed positively correlated with TMeF in both TNBC (R = 0.68, p = 0.005) and OvCa (R = 0.65, p = 0.03). , for TNBC and OvCa, respectively, FIG.22B). Only cfBRCA1meth-positive samples from female donors are included in the analysis shown in FIGs.22A-22B. FIG.22B shows the correlation between TMeF (x-axis, log10-transformed) and cfBRCA1meth fraction (y-axis, log10-transformed) is shown for cfBRCA1meth-positive OvCa and TNBC samples. A linear regression line is shown for each cancer type and the R value and associated p-value from a Pearson’s correlation test is also indicated. Only samples with TMeF above 100 ppm were included in the analysis. These results suggest that the larger fraction of cfBRCA1meth observed in patients with TNBC and OvCa originates from BRCA1meth cancer cells that release BRCA1meth DNA into the plasma commensurate with tumor mass. On the other hand, when all female cancers other than TNBC and OvCa were analyzed separately (i.e., ‘other cancers’), there was no significant association between cfBRCA1meth fraction and TMeF (FIG.23).
[0325] However, when comparing the distribution of cfBRCA1meth fractions between non-cancer and other-cancer females positive for cfBRCA1meth, the median cfBRCA1meth fraction was 1.5-fold higher in other cancer vs. non-cancer (p = 0.04, FIG.22A). This suggests that some women with cancers other than TNBC and OvCa either have higher constitutionalACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 88 of 119 cfBRCA1meth levels than non-cancer individuals or have occasional tumors driven by BRCA1 promoter methylation.
[0326] When the TNBC and ovarian cancer patients are excluded from the cfBRCA1meth analysis in the CCGA cancer cohort, several interesting observations arise. First, that women with cancers other than TNBC and ovarian cancers had a statistically higher cfBRCA1meth fraction than women without cancer. Nikolaienko, et al, found concordant tumor and blood BRCA1meth in 4 / 6 breast cancers with low tumor ER levels and in 3 / 220 breast cancers with high tumor ER. Thus, prevalent female cancers that are not TNBC and OvCa can also show tumor BRCA1meth and contribute to increasing the average cfBRCA1meth fraction in this subgroup. The more surprising observation, however, is that the cfBRCA1meth prevalence in the CCGA cohort was 1.4% (19 / 1,380) in males with cancer vs. 2.9% (31 / 1080) in males without cancer, which is a statistically significant difference. Notably, this differential is predominantly driven by the low rate of cfBRCA1meth observed in the prostate cancer subgroup, which is the most represented cancer type in the male CCGA cohort, accounting for 35% of all male cancer instances (484 / 1,380 cancer cases in the male cohort, of which only 4 (0.8%) are positive for cfBRCA1meth). The prevalence of BRCA1meth in prostate cancers from the TCGA cohort was also particularly low, with no BRCA1meth cases detected across the 502 cases examined (0%). Taken together, this suggests that while disruptive BRCA2 gene mutations are commonly found in prostate cancer, not only BRCA1meth is exceedingly rare in this tumor type (as are driver mutations in the BRCA1 gene), but that even constitutional BRCA1meth is underrepresented in prostate cancer patients.
[0327] In summary, the present disclosure provides a circulating cell-free DNA (cfDNA)- based targeted methylation platform that is a robust, biopsy-free, scalable assay that is derived from the previously validated next generation sequencing platform (Jamshidi et al., PMID: 3640018) (Klein et al., PMID: 34176681) (Liu et al., PMID: 33506766). Herein, the present disclosure demonstrates that the presently disclosed subject matter can be utilized for sensitive and quantitative cell-free BRCA1 promoter methylation (cfBRCA1meth) detection in large cohorts of both cancer patients and non-cancer individuals, with utility both as a predictive biomarker and a cancer risk assessment tool in precision oncology. 12.6. Methods CCGA study cohortACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 89 of 119
[0328] The Circulating Cell-Free Genome Atlas Study (CCGA, NCT02889978) is a previously reported multi-center, observational, longitudinal, case-control study. Cancer cases were required to have a confirmed diagnosis of invasive cancer. Non-cancer cases were required to not have current or prior cancer diagnosis nor any acute illness at the time of enrollment. Of the 4,077 participants in CCGA Substudy 3 and 6,689 participants in Substudy 2, cell-free DNA samples using version 2 of the assay were selected in order to maintain technical and analytical consistency in the BRCA1 promoter methylation data arising from these samples. Furthermore, cancer types with at least 30 samples and with matching cancer types in The Cancer Genome Atlas TCGA were selected. In total, 5,748 samples were selected for BRCA1 promoter methylation analysis, representing 2,958 patients with a range of cancers, and 2,790 individuals without cancer (FIG.18). TCGA pancancer cohort
[0329] Methylation beta-values from the Illumina Infinium HumanMethylation27 (27k) and HumanMethylation450 (450K) arrays were downloaded from the UCSC Xena Browser in January 2024. To avoid duplicates and potential biases, only primary tumors from unique patient donors were included in the analysis. In instances where both 27K and 450K array datasets were available for the same tumor sample, the 450K array dataset was prioritized. Only tumor types shared between the TCGA and the CCGA cohorts were included in the analysis, resulting in a total of 7,787 samples across 14 tumor types. For each tumor sample, the average beta-value of the probes mapping to the BRCA1 promoter region (chr17: 41277134-41277486; GRCh37) was computed and the threshold of 0.2 was selected to identify methylated cancers. Bisulfite sequencing
[0330] Briefly, whole-genome bisulfite sequencing: cfDNA was isolated from plasma, and whole-genome bisulfite sequencing (WGBS; 30x depth) was employed for analysis of cfDNA. cfDNA was extracted from two tubes of plasma (up to a combined volume of 10 ml) per patient using a modified QIAamp Circulating Nucleic Acid kit (Qiagen; Germantown, MD). Up to 75 ng of plasma cfDNA was subjected to bisulfite conversion using the EZ-96 DNA Methylation Kit (Zymo Research, D5003). Converted cfDNA was used to prepare dual indexed sequencing libraries using Accel-NGS Methyl-Seq DNA library preparation kits (Swift BioSciences; Ann Arbor, MI) and constructed libraries were quantified using KAPA Library Quantification Kit for Illumina Platforms (Kapa Biosystems; Wilmington, MA). Four libraries along with 10%ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 90 of 119 PhiX v3 library (Illumina, FC-110-3001) were pooled and clustered on an Illumina NovaSeq 7000 S2 flow cell followed by 150-bp paired-end sequencing (30x). Determining BRCA1 promoter methylation in circulating cell-free DNA (cfBRCAmeth)
[0331] A targeted methylation assay that detects cancer methylation signatures for multi- cancer early detection (MCED) was utilized (Liu et al., PMID: 33506766). While the MCED methylation assay captures over a million CpG sites genome-wide, the cfBRCA1meth assay focuses only on sequencing reads overlapping the BRCA1 promoter region ( chr17: 41277134- 41277486; GRCh37) and spanning at least 4 CpG sites (i.e., informative reads). After the targeted methylation panel was run and sequenced reads were generated, fragments were aligned to the BRCA1 promoter region and the alignment files were extracted. Fragments mapping to the BRCA1 promoter region and spanning at least 4 CpG sites (i.e., informative reads) were required to have at least two PCR duplicates to ensure robustness, however, fragments were then deduplicated so that each read represents a unique molecule of cfDNA in the sample. In both the cancer and non-cancer cohorts, the number of informative fragments per sample ranged from 21 to 237 (median = 92). A filtering step was applied to remove sequence reads obtained from molecules that failed to undergo conversion during bisulfite conversion based on the rate of CHH or CHG conversion. A fragment is considered informative if it maps and / or overlaps with the BRCA1 promoter region, spans at least 4 CpG sites, is a unique fragment (i.e., does not have a PCR duplicate), and has not failed to undergo bisulfite conversion. A fragment is considered methylated if at least 4 of its CpGs were methylated. Fragments contained a median of 11 CpGs (range 4-20) per sample. The distribution of CpGs per fragment were consistent across samples, independent of methylation status. Fragments tended to be either predominantly methylated (>75% CpGs methylated) or predominantly unmethylated (< 25% CpGs methylated). The threshold of 4 methylated CpG sites was chosen such that the theoretical probability of erroneously calling a sample methylated (due to sequencing or other technical error) is less than 0.0001. The methylation fraction for each sample is measured as the number of methylated fragments divided by the total number of methylated and un-methylated fragments at the BRCA1 promoter and is therefore a fraction between 0 and 1. Prevalence, reported as a percentage, refers to the proportion of samples in a group that exhibit evidence of cfBRCA1meth. At the individual sample level, a sample is considered methylated (i.e., cfBRCA1meth) if at least one informative fragment wasACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 91 of 119 methylated (FIG. 16). For each sample, the cfBRCA1meth frequency was computed as the fraction of methylated fragments over the total number of informative fragments, and it therefore ranges between 0 and 1. Percentages (range 0%-100%) can also be used when discussing the prevalence of cfBRCA1meth within groups of samples (FIG.16). Empirical limit of detection
[0332] The limit of cfBRCA1meth detection, was analyzed with seven serial dilutions of commercially available material with known levels of global contrived methylation, which is used as a proxy for the expected fraction of methylation at the BRCA1 promoter region. At contrived methylation levels of 0, 0.1, 0.9, and 1.0, samples were prepared by mixing biochemically methylated or unmethylated, sheared, and size-selected genomic DNA from DKO HCT116 cell lines, with 7–8 replicates per level. For lower contrived methylation levels (0.003–0.0243), Seraseq ctDNA reference material (SeraCare) was spiked into a background non-cancer sample, with 19–26 replicates per level. In all mixtures, the reported contrived methylation reflects only the spiked-in material and excludes background methylation. Linearity was assessed for samples with contrived methylation levels >0.001. Sensitivity was evaluated using the empirical LoD95, defined as the minimum cfBRCA1meth fraction (estimated from the contrived methylation level) at which 95% of replicates are classified as positive. Statistical analyses
[0333] All plotting and statistical analyses were performed in R version 4.3.2 (2023-10- 31). Where prevalence was measured within a group (e.g., based on cancer type, sex, race / ethnicity, age), the prevalence and a 95% exact binomial confidence interval was reported (computed with the Hmisc function of the binconf R package, using method = “exact”). When prevalences are compared across groups, the Fisher’s Exact test for count data was applied if at least one of the cell counts was less than or equal to 5 or if the total count was less than or equal to 1,000. Otherwise, the Chi-squared test was used.
[0334] As described, the presently disclosed subject matter demonstrates that the method, as a single blood-based assay, can sensitively detect and quantify BRCA1meth from cfDNA, which is informative of cancer susceptibility in non-cancer bearing individuals and has treatment implications for cancer patients.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 92 of 119 13. Working Example #2: Estimating Single Site Methylation Status in ctDNA 13.1. Evaluation of feasibility of distinguishing t(11;14) status in multiple myeloma patients using a methylation-based platform
[0335] A set of considerations for distinguishing t(11;14) status in multiple myeloma patients using a methylation based platform may include indirectly investigating the aggregating methylation signals associated with the translocation status (i.e., not directly assessing the translocation site). The advantages of using a methylation-based platform for distinguishing t(11;14) status includes the following:
[0336] 1. It is non-invasive (plasma from a blood draw may spare bone marrow biopsy).
[0337] 2. It provides an opportunity to learn more t(11;14) biology from panel-wide methylation signals.
[0338] 2.a. Translocation events are associated with changes in methylation.
[0339] 2.b. Cells that are positive for t(11;14) have a more B-cell-like phenotype, making it more likely that there are genome-wide methylation differentiation (Bal et al., 2022).
[0340] 2.c. Various transcriptional signatures that distinguish t(11;14) positive samples have been reported and differential methylation alongside transcription changes is to be expected (Hollein et al., 2022)
[0341] 3. cfDNA tumor fraction estimates and cancer detection status can be simultaneously reported.
[0342] Thus, the present disclosure provides a method to analyze nucleic acids from plasma extracted from a blood sample which is non-invasive, will lead to additional opportunities to learn about t(11;14) biology, and can complement the tumor fraction estimations and the cancer detection status already produced by the platform.
[0343] FIG. 24 illustrates a graph and other considerations regarding a cohort of samples used to evaluate feasibility of distinguishing t(11;14) status in multiple myeloma patients using a methylation based platform.
[0344] FIG. 25 illustrates another graph and other considerations regarding a cohort of samples used to evaluate feasibility of distinguishing t(11;14) status in multiple myeloma patients using a methylation based platform.
[0345] Differently methylated regions (DMRs) are identified between the positive and negative groups.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 93 of 119
[0346] 1. DMR analysis looks for significant differentially methylated regions (CpGs) between two group, here defined as positive and negative t(11;14).
[0347] 2. DMRs with corrected p-value <0.05 are considered significant.
[0348] 3. Principal Component Analysis (PCA)* on the significant DMRs allows for visualization of sample clustering. *PCA is a dimensional reduction method that finds the largest axes of variation in high-dimensional data. This allows for the capture of the majority of variation across many significant DMRs on just tow axes.
[0349] FIG. 26 illustrates a PCA analysis graph showing an example segmentation of DMRs for distinguishing t(11;14) status in multiple myeloma patients using a methylation based platform. Although one segmentation has been illustrated a variety of additional segmentations are expected to be discovered through further experimentation. FIG. 26 particularly provides an overview of DMR analysis using principal component analysis (PCA).
[0350] In summary, the present disclosure utilized a cohort of samples (CCGA) and identified a signal for distinguishing t(11;14) associated methylation changes. More samples will increase the power to detect methylation signals and validate findings. Analysis will be powered with a central test for t(11;14) status (CCGA results are a mix of local tests for validating t(11;14) status). Later stage samples, which have more signal will be included in future analysis (Most of the CCGA samples are early stage). With more high signal samples, other analysis will include exploring of t(11;14) methylation related biology and development of classifier. Therefore, taken together, the graphs and their associated analyses demonstrate that it is feasible to distinguish t(11;14) status in multiple myeloma patients using a targeted methylation-based analysis of nucleic acid fragments extracted from a sample comprising plasma extracted from a whole blood sample. 13.2. Estimating Single Site Methylation Status in ctDNA
[0351] In particular embodiments, the analytics system 200 may estimate the methylation status of a single genomic site in circulating tumor DNA (ctDNA) using data modeling and machine learning, even when direct observation of the site is not possible due to low tumor DNA content in a liquid biopsy. By leveraging methylation patterns from multiple genomic regions observed in tissue or high-tumor-fraction plasma samples, the analytics system 200 may train predictive models including classifiers and regression algorithms to infer the methylation state of a target site. This approach enables accurate epigenetic profiling from cell- free DNA (cfDNA), supporting cancer diagnostics and treatment decisions without requiringACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 94 of 119 invasive tissue biopsies. The analytics system 200 may also incorporate tumor methylation fraction estimation and model training using synthetic data, allowing for robust performance across diverse cancer types and sample conditions. Although this disclosure describes using particular systems to determine particular methylation status in particular manners, this disclosure contemplates using any suitable system for determining any suitable methylation status in any suitable manner.
[0352] In particular embodiments, the analytics system 200 may access a cell-free DNA (cfDNA) extracted from a biological sample. The analytics system 200 may then generate a plurality of sequence reads for the cfDNA. The analytics system 200 may then identify, based on the sequence reads, methylation patterns at a plurality of genomic sites within the cfDNA. The analytics system 200 may further execute one or more machine-learning models on the sequence reads and the identified methylation patterns to estimate a methylation status of a target genomic site. In particular embodiments, the one or more machine-learning models were trained using a set of reference data generated from tissue or plasma samples with identified methylation status at genomic sites.
[0353] Research, clinical studies, and companion diagnostics for targeted therapies can require the assessment of the methylation status of individual sites (for example, the CpG island at a transcription start site of a gene of interest). In particular embodiments, the biological sample may be a blood draw or a liquid biopsy. While the assessment of the methylation status of individual sites is typically available from tumor cells collected in a tissue biopsy or surgical procedure, in many clinical settings no tissue is available and it is desired to make this assessment from just a blood draw or liquid biopsy by assaying circulating tumor DNA (ctDNA), which is a part of the cell-free DNA (cfDNA). Examples for such a question include the determination of the methylation status of BRCA1 in breast cancer, or the methylation status of Trop-2 in SCLC.
[0354] DNA methylation status at a single site of interest in cfDNA can for example be obtained from methylation-specific PCR, WBGS, RRBS, or a targeted methylation assay. All of these methods are limited by the tumor fraction - the fraction of cfDNA that is actually ctDNA. In many clinical situations, especially early-stage cancer or on treatment, few or not a single tumor-derived cfDNA fragment at the site of interest might be contained or observable in a blood draw.
[0355] DNA methylation both causes and reflects cellular identity, regulation, and its dysregulation in cancer. This is represented in a wide number of methylation sites all acrossACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 95 of 119 the genome. Internal feedback loops, even in dysregulated cells, interconnect methylation status of different regions to establish a cell state.
[0356] In particular embodiments, at least one of the machine-learning models may include a supervised machine learning algorithm configured to associate methylation patterns at the plurality of genomic sites with the methylation status of the target genomic site. The analytics system 200 may use data modeling and / or supervised machine learning to identify the relationship between the methylation status in a target region of interest and many other observable methylation regions either from tissue samples of a cancer type of interest, or from plasma samples with a high tumor fraction that allow to observe the target site.
[0357] In particular embodiments, the methylation status of the target genomic site (e.g., a CpG site) may be estimated in an absence of a direct observation of the target genomic site in the cfDNA. Observations of many methylation sites in cfDNA can then be used to estimate the methylation state of tumor cells even if the site of interest is not observed, or not observed in a sufficient number of molecules in cfDNA.
[0358] As a result, the embodiments disclosed herein may have a technical advantage of addressing shortcomings of current attempts to assess tumor methylation status at a single genomic site in liquid biopsy. First, the embodiments disclosed herein may allow to assess the methylation site of multiple potential sites of interest with the same data, method, and assay. For many typical methods, individual site-specific reagents (for example, methylation-specific PCR primers or single-stranded pull-down probes) are needed to target a new site of interest. Second, the embodiments disclosed herein may enable estimates of the methylation status of a tumor even in the presence of few (down to no) ctDNA molecule in a sample that covers the site of interest, such that no direct read-out of the tumor methylation status is possible. 13.3. Supervised Machine Learning of Methylation Patterns
[0359] In particular embodiments, the analytics system 200 has the capability to detect the presence of ctDNA in cfDNA, and to identify a cancer signal origin, i.e., identify one of multiple learned, predefined methylation patterns indicative of the presence of different classes or types of cancer. Estimating the methylation status of the target genomic site may be further based on the presence of the ctDNA. The analytics system 200 can predict the organ or organ group afflicted with cancer, and predict cancer biology information underlying the cancer signal.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 96 of 119
[0360] The implementation of the embodiments disclosed herein may require either a selection of tissue samples of a selected target cancer type or plasma samples from patients with this cancer type where the target methylation status has been identified, independently, for example by a biopsy that is available for training and development of a method, but not for application.
[0361] In particular embodiments, the set of reference data may be generated from tissue samples. The identified methylation status in the tissue samples may be determined based on one or more of whole-genome bisulfite sequencing (WGBS), reduced representation bisulfite sequencing (RRBS), or a targeted methylation panel.
[0362] For tissue samples, the methylation status of tumor cells can be identified by whole- genome bisulfite sequencing (WGBS), reduced representation bisulfite sequencing (RRBS), or a targeted methylation panel. This can provide both the ground truth methylation status at a target region of interest and also the training data for the methylation pattern associated with this single site methylation status.
[0363] In particular embodiments, the set of reference data may be generated from plasma samples. The identified methylation status in the plasma samples may be determined based on tumor fractions.
[0364] For plasma samples, the methylation status of a target site can be identified in training samples with selected high tumor fraction, or get collected independently, for example from a biopsy that is not also sequenced with the target assay.
[0365] In particular embodiments, the analytics system 200 may predict a presence of one or more cancer subtypes based on the sequence reads from the cfDNA. Estimating the methylation status of the target genomic site may be further based on the presence of the one or more cancer subtypes. Using one or more classifiers associated with the analytics system 200 to predict cancer subtype from cfDNA, one CSO for the target tumor type with the target site methylated, and one CSO for the target tumor type with target site unmethylated can be created, and a respective classifier can be trained. For a given test sample, the detection of ctDNA presence and the prediction of presence of one of these subtypes may then provide the information on the methylation status of the target site in the cancer cells.
[0366] Modeling methylation relationships and estimating single site methylation status without direct observation may be an effective solution for addressing the technical challenge of estimating the methylation status of a site when no ctDNA fragment covering that site is observed. The analytics system 200 may use machine learning to learn the relationship betweenACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 97 of 119 the methylation status of a target site and the methylation patterns of many other genomic sites and infer the methylation status of the target site based on the relationship between the methylation status of the target site and the methylation patterns of other genomic sites even when the target site is not directly observed in the cfDNA sample. 13.4. Regression to Estimate Beta Values
[0367] In particular embodiments, the analytics system 200 may estimate tumor-derived fragments in the cfDNA based on applying an expectation maximization algorithm to the sequence reads associated with the cfDNA. The analytics system 200 may further refine the estimation of the methylation status at the target genomic site based on the estimated tumor- derived fragments.
[0368] A second implementation of the embodiments disclosed herein may combine the technology of identifying likely tumor-derived cfDNA molecules with a regression learned on tissue samples or high-signal cancer samples to predict the methylation status at the target site from the methylation site of other sites observable in cfDNA (the target methylation site may be one of the potential sites that contribute to this regression).
[0369] For the second implementation, two separate analyses may be combined. In one analysis, using existing tumor methylated fraction (TMeF) methods, informatively methylated regions may be identified in tissue samples of the target cancer type (separate for being mostly methylated or unmethylated at the target site) that can allow to estimate the presence of tumor- derived fragments in cfDNA. Using existing TMeF expectation maximization methods, this may allow to identify the presence and rate of tumor-derived fragments in cfDNA. Both abnormally methylated and abnormally unmethylated DNA fragments at the same site may contribute to this estimate. In another analysis, using the observed methylation state in tissue or high-signal plasma samples for both methylation states of the target site, a regression model may be learned that predicts methylation status or beta value at the target site from the methylation status or many other observed sites.
[0370] When analyzing a cfDNA sample, two different ways can use the information that is created during training and development above. In particular embodiments, as used for multiple cancer types, an expectation maximization method can estimate the TMeF separately for the cancer type with methylated target site, and once for the cancer type with unmethylated cancer site, and then pick the status as the cancer type with the highest estimated tumor methylation fraction. In other words, estimating the tumor-derived fragments may includeACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 98 of 119 estimating the tumor-derived fragments separately for cancer types with methylated genomic sites and cancer types with unmethylated genomic sites. Refining the estimation of the methylation status at the target genomic site based on the estimated tumor-derived fragments may include estimating the methylation status at the target site as a cancer type with a highest amount of estimated tumor-derived fragments.
[0371] In particular embodiments, estimating the tumor-derived fragments may include estimating a fraction of non-cancer cfDNA, a fraction of ctDNA for a cancer with methylated genomic sites, and a fraction of ctDNA with unmethylated genomic sites. The expectation maximization problem may be set up such that it estimates the fraction of normal non-cancer cfDNA, the fraction of ctDNA for the cancer with methylated target site, and the fraction of ctDNA with unmethylated target site in one expectation maximization problem. The estimated fraction of methylation at the target site may be then the ratio of the estimated tumor methylation fractions for the two types.
[0372] In particular embodiments, at least one of the machine-learning models may include a regression model. The analytics system 200 may input the estimated tumor-derived fragments to the regression model. All fragments identified as likely being tumor-derived based on the TMeF estimation algorithm may be input to the regression model that can predict methylation status at the target site from the observed methylation states of likely tumor-derived fragments.
[0373] Any of these or a combination of these methods may be used to identify the tumor DNA methylation status of a single target site when cfDNA with some fraction of ctDNA is available for observation. The embodiments disclosed herein may make use of the fact that it is much more likely to observe the methylation status at least some of multiple different sites in the tumor genome with an opportunity to then regress or otherwise estimate tumor methylation status than it would be to observe just the target methylation site itself, especially at low tumor fraction and with an uncertain background of methylation at this state in non- cancer cfDNA.
[0374] In particular embodiments, the analytics system 200 may determine, based on the one or more machine-learning models and the estimated methylation status of the target genomic site, a cancer prediction. The cancer prediction may include one or more of a label indicating a particular cancer state, a label indicating a particular cancer type, a label indicating a particular cancer stage, a likelihood of cancer, or a likelihood of a particular cancer type. As an example and not by way of limitation, the target genomic site is associated with breast cancer gene 1 (BRCA1). Accordingly, the cancer prediction is associated with breast cancer. AsACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 99 of 119 another example and not by way of limitation, the target genomic site is associated with tumor- associated calcium signal transducer 2 (Trop-2). Accordingly, the cancer prediction is associated with small cell lung cancer.
[0375] As a result, the embodiments disclosed herein may have technical advantage of the ability to detect BRCA1 promoter hypermethylation from a cfDNA sample. Another technical advantage of the embodiments may include increasing the accuracy and sensitivity of determining the BRCA1meth fraction in a patient sample.
[0376] FIG. 27 illustrates an example method 2700 for estimating single site methylation status. The method may begin at step 2710, where the analytics system 200 may access a cell- free DNA (cfDNA) extracted from a biological sample. The biological sample may be a blood draw or a liquid biopsy. At step 2720, the analytics system 200 may generate a plurality of sequence reads for the cfDNA. At step 2730, the analytics system 200 may detect a presence of a circulating tumor DNA (ctDNA) in the cfDNA. At step 2740, the analytics system 200 may predict a presence of one or more cancer subtypes based on the sequence reads from the cfDNA. At step 2750, the analytics system 200 may execute one or more machine-learning models on the sequence reads and the identified methylation patterns to estimate a methylation status of a target genomic site. The target genomic site may be a CpG site. The methylation status of the target genomic site may be estimated in an absence of a direct observation of the target genomic site in the cfDNA. In particular embodiments, estimating the methylation status of the target genomic site may be further based on the presence of the ctDNA and / or the presence of the one or more cancer subtypes. In particular embodiments, the one or more machine-learning models were trained using a set of reference data generated from tissue or plasma samples with identified methylation status at genomic sites. At least one of the machine- learning models may include a supervised machine learning algorithm configured to associate methylation patterns at the plurality of genomic sites with the methylation status of the target genomic site. At step 2760, the analytics system 200 may estimate tumor-derived fragments in the cfDNA based on applying an expectation maximization algorithm to the sequence reads associated with the cfDNA, comprising estimating the tumor-derived fragments separately for cancer types with methylated genomic sites and cancer types with unmethylated genomic sites or estimating a fraction of non-cancer cfDNA, a fraction of ctDNA for a cancer with methylated genomic sites, and a fraction of ctDNA with unmethylated genomic sites. At step 2770, the analytics system 200 may refine the estimation of the methylation status at the target genomic site based on the estimated tumor-derived fragments. At step 2780, the analytics system 200ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 100 of 119 may determine, based on the one or more machine-learning models and the estimated methylation status of the target genomic site, a cancer prediction. In particular embodiments, the cancer prediction may include one or more of a label indicating a particular cancer state, a label indicating a particular cancer type, a label indicating a particular cancer stage, a likelihood of cancer, or a likelihood of a particular cancer type. Particular embodiments may repeat one or more steps of the method of FIG.27, where appropriate. Although this disclosure describes and illustrates particular steps of the method of FIG.27 as occurring in a particular order, this disclosure contemplates any suitable steps of the method of FIG.27 occurring in any suitable order. Moreover, although this disclosure describes and illustrates an example method for estimating single site methylation status including the particular steps of the method of FIG. 27, this disclosure contemplates any suitable method for estimating single site methylation status including any suitable steps, which may include all, some, or none of the steps of the method of FIG. 27, where appropriate. Furthermore, although this disclosure describes and illustrates particular components, devices, or systems carrying out particular steps of the method of FIG. 27, this disclosure contemplates any suitable combination of any suitable components, devices, or systems carrying out any suitable steps of the method of FIG.27. 14. Machine Learning
[0377] In particular embodiments, the analytics system 200 may train one or more machine-learning models for different analytic tasks. As an example and not by way of limitation, the machine-learning models may comprise one or more of a support vector machine (SVM), random forest (RF), linear discriminant analysis (LDA), logistic regression, linear regression, a deep-learning model, a neural network, a kernel-based regression, an adaptive basis regression or classification, a Bayesian method, an ensemble method, a Gaussian process, a probabilistic model, or a probabilistic graphical model.
[0378] In particular embodiments, the analytics system 200 may train the machine-learning models based on data samples. The data samples may be divided into multiple datasets. In particular embodiments, the datasets may include training, validation, and test sets. A machine- learning model may be initially fit on a training dataset, which is a set of examples used to fit the parameters (e.g., weights of connections between neurons in artificial neural networks) of the model. The machine-learning model may be trained on the training dataset using a supervised learning method, for example using optimization methods such as gradient descent or stochastic gradient descent. In particular embodiments, the training dataset may include pairsACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 101 of 119 of an input vector (or scalar) and the corresponding output vector (or scalar), where the answer key is commonly denoted as the target (or label). The machine-learning model may be run with the training dataset and produce a result, which is then compared with the target, for each input vector in the training dataset. Based on the result of the comparison and the specific learning algorithm being used, the parameters of the machine-learning model may be adjusted. The model fitting can include both variable selection and parameter estimation.
[0379] In particular embodiments, the fitted model may be used to predict the responses for the observations in a second dataset called the validation data set. The validation dataset may provide an unbiased evaluation of a model fit on the training dataset while tuning the model’s hyperparameters (e.g., the number of hidden units—layers and layer widths—in a neural network). Validation datasets can be used for regularization by early stopping (stopping training when the error on the validation dataset increases, as this is a sign of over-fitting to the training dataset).
[0380] In particular embodiments, the test dataset may be used to provide an unbiased evaluation of a final model fit on the training dataset. If the data in the test dataset has never been used in training (for example in cross-validation), the test dataset may be also called a holdout dataset.
[0381] In particular embodiments, a dataset can be repeatedly split into several training and validation datasets, which is known as cross-validation. To confirm the model’s performance, an additional test dataset held out from cross-validation may be used.
[0382] In particular embodiments, the analytics system 200 may train the machine-learning models based on different approaches, e.g., supervised learning, unsupervised learning, and reinforcement learning.
[0383] In particular embodiments, the analytics system 200 may train various types of models and select the best model for a task. As an example and not by way of limitation, one type of model may include neural networks. The analytics system 200 may train the neural networks through empirical risk minimization. The analytics system 200 may optimize the network’s parameters to minimize the difference, or empirical risk, between the predicted output and the actual target values in a given dataset. Gradient based methods such as backpropagation may be also used to estimate the parameters of the network. During the training phase, the neural networks may learn from labeled training data by iteratively updating their parameters to minimize a defined loss function, which allows the network to generalize to unseen data.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 102 of 119 15. Systems and Methods
[0384] FIG. 15 illustrates an example computer system 1500. In particular embodiments, one or more computer systems 1500 perform one or more steps of one or more methods described or illustrated herein. In particular embodiments, one or more computer systems 1500 provide functionality described or illustrated herein. In particular embodiments, software running on one or more computer systems 1500 performs one or more steps of one or more methods described or illustrated herein or provides functionality described or illustrated herein. Particular embodiments include one or more portions of one or more computer systems 1500. Herein, reference to a computer system may encompass a computing device, and vice versa, where appropriate. Moreover, reference to a computer system may encompass one or more computer systems, where appropriate.
[0385] This disclosure contemplates any suitable number of computer systems 1500. This disclosure contemplates computer system 1500 taking any suitable physical form. As example and not by way of limitation, computer system 1500 may be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) (such as, for example, a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a mesh of computer systems, a mobile telephone, a personal digital assistant (PDA), a server, a tablet computer system, or a combination of two or more of these. Where appropriate, computer system 1500 may include one or more computer systems 1500; be unitary or distributed; span multiple locations; span multiple machines; span multiple data centers; or reside in a cloud, which may include one or more cloud components in one or more networks. Where appropriate, one or more computer systems 1500 may perform without substantial spatial or temporal limitation one or more steps of one or more methods described or illustrated herein. As an example and not by way of limitation, one or more computer systems 1500 may perform in real time or in batch mode one or more steps of one or more methods described or illustrated herein. One or more computer systems 1500 may perform at different times or at different locations one or more steps of one or more methods described or illustrated herein, where appropriate.
[0386] In particular embodiments, computer system 1500 includes a processor 1502, memory 1504, storage 1506, an input / output (I / O) interface 1508, a communication interface 1510, and a bus 1512. Although this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, thisACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 103 of 119 disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement.
[0387] In particular embodiments, processor 1502 includes hardware for executing instructions, such as those making up a computer program. As an example and not by way of limitation, to execute instructions, processor 1502 may retrieve (or fetch) the instructions from an internal register, an internal cache, memory 1504, or storage 1506; decode and execute them; and then write one or more results to an internal register, an internal cache, memory 1504, or storage 1506. In particular embodiments, processor 1502 may include one or more internal caches for data, instructions, or addresses. This disclosure contemplates processor 1502 including any suitable number of any suitable internal caches, where appropriate. As an example and not by way of limitation, processor 1502 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). Instructions in the instruction caches may be copies of instructions in memory 1504 or storage 1506, and the instruction caches may speed up retrieval of those instructions by processor 1502. Data in the data caches may be copies of data in memory 1504 or storage 1506 for instructions executing at processor 1502 to operate on; the results of previous instructions executed at processor 1502 for access by subsequent instructions executing at processor 1502 or for writing to memory 1504 or storage 1506; or other suitable data. The data caches may speed up read or write operations by processor 1502. The TLBs may speed up virtual-address translation for processor 1502. In particular embodiments, processor 1502 may include one or more internal registers for data, instructions, or addresses. This disclosure contemplates processor 1502 including any suitable number of any suitable internal registers, where appropriate. Where appropriate, processor 1502 may include one or more arithmetic logic units (ALUs); be a multi- core processor; or include one or more processors 1502. Although this disclosure describes and illustrates a particular processor, this disclosure contemplates any suitable processor.
[0388] In particular embodiments, memory 1504 includes main memory for storing instructions for processor 1502 to execute or data for processor 1502 to operate on. As an example and not by way of limitation, computer system 1500 may load instructions from storage 1506 or another source (such as, for example, another computer system 1500) to memory 1504. Processor 1502 may then load the instructions from memory 1504 to an internal register or internal cache. To execute the instructions, processor 1502 may retrieve the instructions from the internal register or internal cache and decode them. During or after execution of the instructions, processor 1502 may write one or more results (which may beACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 104 of 119 intermediate or final results) to the internal register or internal cache. Processor 1502 may then write one or more of those results to memory 1504. In particular embodiments, processor 1502 executes only instructions in one or more internal registers or internal caches or in memory 1504 (as opposed to storage 1506 or elsewhere) and operates only on data in one or more internal registers or internal caches or in memory 1504 (as opposed to storage 1506 or elsewhere). One or more memory buses (which may each include an address bus and a data bus) may couple processor 1502 to memory 1504. Bus 1512 may include one or more memory buses, as described below. In particular embodiments, one or more memory management units (MMUs) reside between processor 1502 and memory 1504 and facilitate accesses to memory 1504 requested by processor 1502. In particular embodiments, memory 1504 includes random access memory (RAM). This RAM may be volatile memory, where appropriate. Where appropriate, this RAM may be dynamic RAM (DRAM) or static RAM (SRAM). Moreover, where appropriate, this RAM may be single-ported or multi-ported RAM. This disclosure contemplates any suitable RAM. Memory 1504 may include one or more memories 1504, where appropriate. Although this disclosure describes and illustrates particular memory, this disclosure contemplates any suitable memory.
[0389] In particular embodiments, storage 1506 includes mass storage for data or instructions. As an example and not by way of limitation, storage 1506 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disc, a magneto-optical disc, magnetic tape, or a Universal Serial Bus (USB) drive or a combination of two or more of these. Storage 1506 may include removable or non-removable (or fixed) media, where appropriate. Storage 1506 may be internal or external to computer system 1500, where appropriate. In particular embodiments, storage 1506 is non-volatile, solid-state memory. In particular embodiments, storage 1506 includes read-only memory (ROM). Where appropriate, this ROM may be mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), or flash memory or a combination of two or more of these. This disclosure contemplates mass storage 1506 taking any suitable physical form. Storage 1506 may include one or more storage control units facilitating communication between processor 1502 and storage 1506, where appropriate. Where appropriate, storage 1506 may include one or more storages 1506. Although this disclosure describes and illustrates particular storage, this disclosure contemplates any suitable storage.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 105 of 119
[0390] In particular embodiments, I / O interface 1508 includes hardware, software, or both, providing one or more interfaces for communication between computer system 1500 and one or more I / O devices. Computer system 1500 may include one or more of these I / O devices, where appropriate. One or more of these I / O devices may enable communication between a person and computer system 1500. As an example and not by way of limitation, an I / O device may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet, touch screen, trackball, video camera, another suitable I / O device or a combination of two or more of these. An I / O device may include one or more sensors. This disclosure contemplates any suitable I / O devices and any suitable I / O interfaces 1508 for them. Where appropriate, I / O interface 1508 may include one or more device or software drivers enabling processor 1502 to drive one or more of these I / O devices. I / O interface 1508 may include one or more I / O interfaces 1508, where appropriate. Although this disclosure describes and illustrates a particular I / O interface, this disclosure contemplates any suitable I / O interface.
[0391] In particular embodiments, communication interface 1510 includes hardware, software, or both providing one or more interfaces for communication (such as, for example, packet-based communication) between computer system 1500 and one or more other computer systems 1500 or one or more networks. As an example and not by way of limitation, communication interface 1510 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI network. This disclosure contemplates any suitable network and any suitable communication interface 1510 for it. As an example and not by way of limitation, computer system 1500 may communicate with an ad hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or one or more portions of the Internet or a combination of two or more of these. One or more portions of one or more of these networks may be wired or wireless. As an example, computer system 1500 may communicate with a wireless PAN (WPAN) (such as, for example, a BLUETOOTH WPAN), a WI-FI network, a WI-MAX network, a cellular telephone network (such as, for example, a Global System for Mobile Communications (GSM) network), or other suitable wireless network or a combination of two or more of these. Computer system 1500 may include any suitable communication interface 1510 for any of these networks, where appropriate. Communication interface 1510 may include one or more communication interfaces 1510,ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 106 of 119 where appropriate. Although this disclosure describes and illustrates a particular communication interface, this disclosure contemplates any suitable communication interface.
[0392] In particular embodiments, bus 1512 includes hardware, software, or both coupling components of computer system 1500 to each other. As an example and not by way of limitation, bus 1512 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a front-side bus (FSB), a HYPERTRANSPORT (HT) interconnect, an Industry Standard Architecture (ISA) bus, an INFINIBAND interconnect, a low-pin-count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a serial advanced technology attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or another suitable bus or a combination of two or more of these. Bus 1512 may include one or more buses 1512, where appropriate. Although this disclosure describes and illustrates a particular bus, this disclosure contemplates any suitable bus or interconnect.
[0393] Herein, a computer-readable non-transitory storage medium or media may include one or more semiconductor-based or other integrated circuits (ICs) (such, as for example, field- programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical discs, optical disc drives (ODDs), magneto-optical discs, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM-drives, SECURE DIGITAL cards or drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these, where appropriate. A computer-readable non-transitory storage medium may be volatile, non-volatile, or a combination of volatile and non-volatile, where appropriate. 16. Recitation of Embodiments
[0394] Embodiment 1. A method for measuring a fraction of methylated molecules at a BRCA1 promoter region, comprising: (a) obtaining a plurality of nucleic acids from a biological sample; (b) generating a plurality of sequence reads from the nucleic acids; (c) determining a methylation status at a plurality of CpG sites within each sequence read of the plurality of sequence reads; (d) counting a number of methylated sequence reads and a number of unmethylated sequence reads, wherein the sequence read is counted as methylated if at least (n) number of methylated CpGs per molecule are present, and unmethylated if less than (n) number of methylated CpGs per molecule are present; and (e) calculating the fraction ofACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 107 of 119 methylated molecules by dividing the number of methylated sequence reads by a total number of methylated and unmethylated sequence reads.
[0395] Embodiment 2. A method for determining a methylation status of a BRCA1 promoter region in a subject, comprising: (a) obtaining a plurality of nucleic acids from a biological sample obtained from the subject; (b) generating a plurality of sequence reads from the nucleic acids; (c) determining a methylation status at a plurality of CpG sites within each sequence read of the plurality of sequence reads; (d) counting the sequence read as methylated, if the sequence read has at least (n) number of methylated CpGs per molecule, and counting the sequence read as unmethylated if the sequence read has less than (n) number of methylated CpGs per molecule; and (e) determining that the BRCA1 promoter region is methylated in the subject if there is at least one read counted as methylated according to step (d).
[0396] Embodiment 3. The method of Embodiment 1 or 2, wherein the (n) number of methylated CpGs per molecule is 3, 4, or 5.
[0397] Embodiment 4. The method of any one of Embodiments 1-3, wherein the (n) number of methylated CpGs per molecule is 4.
[0398] Embodiment 5. The method of any one of Embodiments 1-4, wherein the plurality of sequence reads are mapped to a reference genome.
[0399] Embodiment 6. The method of any one of Embodiments 1-5, wherein the plurality of sequence reads of step (c) overlap with the BRCA1 promoter region.
[0400] Embodiment 7. The method of any one of Embodiments 1-6, wherein the plurality of sequence reads of step (c) overlap with genomic coordinates, chr17: 41277134-41277486; GRCh37.
[0401] Embodiment 8. The method of any one of Embodiments 1-7, wherein the plurality of sequence reads are filtered to remove molecules with fewer than 4 CpG sites prior to step (c).
[0402] Embodiment 9. The method of any one of Embodiments 1-8, wherein the plurality of nucleic acids are subjected to a bisulfite conversion.
[0403] Embodiment 10. The method of any one of Embodiments 1-9, wherein the plurality of sequence reads are generated by a targeted panel, whole-exome panel, or whole-genome panel.
[0404] Embodiment 11. The method of Embodiment 9 or 10, wherein the plurality of sequence reads are filtered to remove molecules that have failed to undergo convers...
Claims
ATTORNEY DOCKET PATENT APPLICATION 087932.0118 113 of 119 CLAIMS What is claimed is:
1. A method for measuring a fraction of methylated molecules at a BRCA1 promoter region, comprising: (a) obtaining a plurality of nucleic acids from a biological sample; (b) generating a plurality of sequence reads from the nucleic acids; (c) determining a methylation status at a plurality of CpG sites within each sequence read of the plurality of sequence reads; (d) counting a number of methylated sequence reads and a number of unmethylated sequence reads, wherein the sequence read is counted as methylated if at least (n) number of methylated CpGs per molecule are present, and unmethylated if less than (n) number of methylated CpGs per molecule are present; and (e) calculating the fraction of methylated molecules by dividing the number of methylated sequence reads by a total number of methylated and unmethylated sequence reads.
2. A method for determining a methylation status of a BRCA1 promoter region in a subject, comprising: (a) obtaining a plurality of nucleic acids from a biological sample obtained from the subject; (b) generating a plurality of sequence reads from the nucleic acids; (c) determining a methylation status at a plurality of CpG sites within each sequence read of the plurality of sequence reads; (d) counting the sequence read as methylated, if the sequence read has at least (n) number of methylated CpGs per molecule, and counting the sequence read as unmethylated if the sequence read has less than (n) number of methylated CpGs per molecule; and (e) determining that the BRCA1 promoter region is methylated in the subject if there is at least one read counted as methylated according to step (d).
3. The method of Claim 1 or 2, wherein the (n) number of methylated CpGs per molecule is 3, 4, or 5.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 114 of 119 4. The method of any one of Claims 1-3, wherein the (n) number of methylated CpGs per molecule is 4.
5. The method of any one of Claims 1-4, wherein the plurality of sequence reads are mapped to a reference genome.
6. The method of any one of Claims 1-5, wherein the plurality of sequence reads of step (c) overlap with the BRCA1 promoter region.
7. The method of any one of Claims 1-6, wherein the plurality of sequence reads of step (c) overlap with genomic coordinates, chr17: 41277134-41277486; GRCh37.
8. The method of any one of Claims 1-7, wherein the plurality of sequence reads are filtered to remove molecules with fewer than 4 CpG sites prior to step (c).
9. The method of any one of Claims 1-8, wherein the plurality of nucleic acids are subjected to a bisulfite conversion.
10. The method of any one of Claims 1-9, wherein the plurality of sequence reads are generated by a targeted panel, whole-exome panel, or whole-genome panel.
11. The method of Claim 9 or 10, wherein the plurality of sequence reads are filtered to remove molecules that have failed to undergo conversion during bisulfite conversion prior to step (c).
12. The method of any one of Claims 1-11, wherein the plurality of sequence reads are deduplicated prior to step (d).
13. The method of any one of Claims 1-12, wherein the biological sample is tissue, blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal, saliva, sweat, tears, pleural fluid, pericardial fluid, peritoneal fluid, or a tumor biopsy.
14. The method of any one of Claims 1-13, wherein the biological sample is a liquid sample or a solid sample.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 115 of 119 15. The method of any one of Claims 1-14, wherein the plurality of nucleic acids are derived from cfDNA.
16. The method of any one of Claims 1-15, further comprising determining a methylation status of a fragment, wherein the fragment is considered methylated if a corresponding sequencing read contains greater than 4 methylated CpG sites; and the sample is considered methylated if at least one fragment is considered methylated.
17. An assay for determining a subject’s BRCA1 promoter methylation status comprising: (a) obtaining a plurality of nucleic acids from a biological sample obtained from the subject; (b) generating a plurality of sequence reads from said nucleic acids; (c) determining a methylation status at a plurality of CpG sites within each sequence read of the plurality of sequence reads; (d) counting the sequence read as methylated, if the sequence read has at least 4 methylated CpGs per molecule, and counting the sequence read as unmethylated if the sequence read has less than 4 methylated CpGs per molecule; and (e) determining that a BRCA1 promoter region is methylated in the subject if there is at least one read counted as methylated according to step (d).
18. The assay of Claim 17, further comprising determining a fraction of methylated fragments, wherein: a fragment is considered methylated if a corresponding sequence read contains greater than 4 methylated CpG sites; the fragment is considered unmethylated if the corresponding sequence read contains less than 4 methylated CpG sites; and the fraction of methylated fragments is calculated by dividing a number of methylated fragments by a total number of methylated and unmethylated fragments.
19. The assay of Claim 18, further comprising determining the fraction of methylated fragments, wherein the sample is considered methylated if at least one fragment is considered methylated.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 116 of 119 20. The assay of any one of Claims 17-19, wherein the plurality of sequence reads are generated by a targeted panel, whole-exome panel, or whole-genome panel.
21. The assay of any one of Claims 17-20, wherein the biological sample is tissue, blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal, saliva, sweat, tears, pleural fluid, pericardial fluid, peritoneal fluid, or a tumor biopsy.
22. The assay of any one of Claims 17-21, wherein the plurality of nucleic acids are derived from cfDNA.
23. The assay of any one of Claims 17-22, wherein the plurality of sequence reads are generated as part of a multi-cancer early detection (MCED) test.
24. The assay of Claim 23, wherein the MCED test report includes a fraction of BRCA1 promoter methylation detected in the subject.
25. A method comprising, by one or more computing systems: accessing a cell-free DNA (cfDNA) extracted from a biological sample; generating a plurality of sequence reads for the cfDNA; identifying, based on the sequence reads, methylation patterns at a plurality of genomic sites within the cfDNA; and executing one or more machine-learning models on the sequence reads and the identified methylation patterns to estimate a methylation status of a target genomic site, wherein the one or more machine-learning models were trained using a set of reference data generated from tissue or plasma samples with identified methylation status at genomic sites.
26. The method of Claim 25, wherein at least one of the machine-learning models comprises a supervised machine learning algorithm configured to associate methylation patterns at the plurality of genomic sites with the methylation status of the target genomic site.
27. The method of Claim 25 or 26, wherein the set of reference data is generated from tissue samples, and wherein the identified methylation status in the tissue samples are determined based on one or more of whole-genome bisulfite sequencing (WGBS), reduced representation bisulfite sequencing (RRBS), or a targeted methylation panel.ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 117 of 119 28. The method of any one of Claims 25-27, wherein the set of reference data is generated from plasma samples, and wherein the identified methylation status in the plasma samples are determined based on tumor fractions.
29. The method of any one of Claims 25-28, wherein the methylation status of the target genomic site is estimated in an absence of a direct observation of the target genomic site in the cfDNA.
30. The method of any one of Claims 25-29, further comprising: estimating tumor-derived fragments in the cfDNA based on applying an expectation maximization algorithm to the sequence reads associated with the cfDNA; and refining the estimation of the methylation status at the target genomic site based on the estimated tumor-derived fragments.
31. The method of any one of Claims 25-30, wherein estimating the tumor-derived fragments comprises: estimating the tumor-derived fragments separately for cancer types with methylated genomic sites and cancer types with unmethylated genomic sites.
32. The method of any one of Claims 25-31, wherein refining the estimation of the methylation status at the target genomic site based on the estimated tumor-derived fragments comprises estimating the methylation status at the target site as a cancer type with a highest amount of estimated tumor-derived fragments.
33. The method of any one of Claims 25-32, wherein estimating the tumor-derived fragments comprises: estimating a fraction of non-cancer cfDNA, a fraction of ctDNA for a cancer with methylated genomic sites, and a fraction of ctDNA with unmethylated genomic sites.
34. The method of any one of Claims 25-33, wherein at least one of the machine-learning models comprise a regression model, the method further comprising: inputting the estimated tumor-derived fragments to the regression model.
35. The method of any one of Claims 25-34, further comprising:ACTIVE 510017597.1ATTORNEY DOCKET PATENT APPLICATION 087932.0118 118 of 119 detecting a presence of a circulating tumor DNA (ctDNA) in the cfDNA, wherein estimating the methylation status of the target genomic site is further based on the presence of the ctDNA.
36. The method of any one of Claims 25-35, further comprising: predicting a presence of one or more cancer subtypes based on the sequence reads from the cfDNA, wherein estimating the methylation status of the target genomic site is further based on the presence of the one or more cancer subtypes.
37. The method of any one of Claims 25-36, further comprising: determining, based on the one or more machine-learning models and the estimated methylation status of the target genomic site, a cancer prediction.
38. The method of any one of Claims 25-37, wherein the cancer prediction comprises one or more of a label indicating a particular cancer state, a label indicating a particular cancer type, a label indicating a particular cancer stage, a likelihood of cancer, or a likelihood of a particular cancer type.
39. The method of any one of Claims 25-38, wherein the target genomic site is associated with breast cancer gene 1 (BRCA1), and wherein the cancer prediction is associated with breast cancer.
40. The method of any one of Claims 25-39, wherein the target genomic site is associated with tumor-associated calcium signal transducer 2 (Trop-2), and wherein the cancer prediction is associated with small cell lung cancer.
41. The method of any one of Claims 25-40, wherein the target genomic site is a CpG site.
42. The method of any one of Claims 25-41, wherein the biological sample is a blood draw or a liquid biopsy.ACTIVE 510017597.1
Citation Information
Patent Citations
Methylation Profile of Cancer
US20110190161A1
Methylation markers and targeted methylation probe panel
US20210025011A1
Determining tumor fraction for a sample based on methyl binding domain calibration data
US20210407623A1