Systems and methods for cell-free nucleic acid methylation assessment

By sequencing and analyzing differentially methylated regions in cell-free nucleic acids, the method improves lung cancer detection beyond genetic markers, utilizing epigenetic information for enhanced screening accuracy.

JP2026500198APending Publication Date: 2026-01-06THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
JP2025533204
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-08
Filing Date
2023-12-08
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

Current lung cancer screening methods, such as image-based screening, are inadequate, and while analysis of circulating tumor DNA (ctDNA) shows promise, they primarily focus on genetic alterations, neglecting the potential of epigenetic markers like DNA methylation for more accurate detection.

Method used

A method involving sequencing cell-free nucleic acids to identify differentially methylated regions, using nucleic acid probes to target specific methylation patterns, and converting nucleobases to determine methylation states, followed by high-throughput sequencing and computational analysis to assess the presence of cancer.

Benefits of technology

Enhances the detection of lung cancer by leveraging tissue-specific methylation signatures in cell-free DNA, providing a more accurate and reliable screening method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026500198000001_ABST
    Figure 2026500198000001_ABST
Patent Text Reader

Abstract

Provided are systems and methods for cell-free nucleic acid sequencing to assess a condition.Generally, cell-free nucleic acid samples are used to perform methyl sequencing, targeting specific regions associated with abnormal methylation.The methylation of cell-free nucleic acid molecules can be evaluated based on sequencing results.Various features can be derived from methylation evaluation and used in computational models to assess cell-free nucleic acid samples for a condition.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 386,557, filed December 8, 2022, entitled "Systems and Methods for Cell-Free DNA Methylation Assessment," the entire disclosure of which is incorporated herein by reference.

[0002] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT This invention was made with government support under contract NSF-1656518 awarded by the National Science Foundation. The government has certain rights in this invention.

[0003] Technical Field The present disclosure provides instructions for assessing methylated cell-free nucleic acids for the purpose of detecting a condition. [Background technology]

[0004] background Lung cancer screening remains an unmet clinical need. While image-based screening is the most common current screening method, analysis of circulating tumor DNA (ctDNA) is a promising alternative and complementary method. Previous studies have leveraged signatures of genetic alterations (e.g., single nucleotide variants and somatic copy number variants) found in cell-free DNA (cfDNA) to predict the lung cancer probability in plasma (Lung-CLiP) of a given sample (JJ Chabon et al., Nature. 2020 Apr;580(7802):245-251; the disclosure of which is incorporated herein by reference). However, in addition to this genetic information, cfDNA also reflects the epigenome of the cell from which it was derived. This means that tumor-derived cfDNA molecules (ctDNA) contain cancer-associated epigenetic signals that can be further exploited for the detection of malignancies. DNA methylation is a promising tumor biomarker. DNA methylation, a stable, heritable covalent modification of cytosines at CpG dinucleotides (CpG), is found at millions of loci across the genome and is known to contribute to chromatin conformation and the regulation of gene expression. Importantly, this means that methylation patterns vary significantly between cell types, meaning that the methylome is cell-type specific. Indeed, these tissue-specific methylation signatures have been used to deconvolute bulk DNA methylation data, work particularly relevant to cell-free DNA, which has been shown to include DNA from blood cells, liver, colon, and, to a lesser extent, other tissues. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] JJChabon et al., Nature.2020 Apr;580(7802):245-251 Summary of the Invention [Means for solving the problem]

[0006] overview In some embodiments, the method is for sequencing to identify differentially methylated regions associated with a state in cell-free nucleic acid.

[0007] In some embodiments, the methods include obtaining a cell-free nucleic acid sample comprising cell-free nucleic acid molecules.

[0008] In some embodiments, the methods involve extracting a subset of cell-free nucleic acid molecules from the cell-free nucleic acid sample using a panel of nucleic acid probes designed to hybridize to regions known to be differentially methylated in a condition.

[0009] In some embodiments, the method comprises converting the nucleobases of a subset of the cell-free nucleic acid molecules.

[0010] In some embodiments, the conversion of a nucleobase indicates the methylation state of that nucleobase.

[0011] In some embodiments, the method comprises sequencing a subset of the cell-free nucleic acid molecules via high-throughput sequencing.

[0012] In some embodiments, the condition is cancer.

[0013] In some embodiments, the cancer is non-small cell lung cancer.

[0014] In some embodiments, the regions known to be differentially methylated in a condition include at least 5% of the regions in Table 2.

[0015] In some embodiments, the regions known to be differentially methylated in a condition include at least 50% of the regions in Table 2.

[0016] In some embodiments, the nucleic acid probe panel excludes regions known to be associated with false discoveries.

[0017] In some embodiments, the nucleic acid probe panel excludes regions known to be differentially methylated in blood cells.

[0018] In some embodiments, the method comprises extracting a subset of cell-free nucleic acid molecules from the cell-free nucleic acid sample using a panel of nucleic acid probes designed to hybridize to regions known to correlate with factors associated with the condition.

[0019] In some embodiments, the condition is cancer and the regions known to correlate with factors associated with the condition include those in Table 1.

[0020] In some embodiments, the methods involve extracting a subset of cell-free nucleic acid molecules from the cell-free nucleic acid sample using a panel of nucleic acid probes designed to hybridize to regions known to be invariably hypermethylated or invariably hypomethylated.

[0021] In some embodiments, regions known to be invariably hypermethylated or invariably hypomethylated include regions in Table 1.

[0022] In some embodiments, converting the nucleobases of the subset of cell-free nucleic acid molecules comprises at least one of bisulfite treatment, TET2 oxidation and APOBEC3A conversion, or TET2 oxidation and pyridine borane treatment.

[0023] In some embodiments, converting the nucleobases of the subset of cell-free nucleic acid molecules comprises TET2 oxidation and APOBEC3A conversion.

[0024] In some embodiments, the cell-free nucleic acid sample is derived from a collection of blood, plasma, saliva, urine, stool, mucus, lymph, or another bodily fluid.

[0025] In some embodiments, the cell-free nucleic acid sample comprises at least 100,000 nucleic acid molecules.

[0026] In some embodiments, the cell-free nucleic acid of the cell-free nucleic acid sample is cell-free DNA.

[0027] In some embodiments, the cell-free nucleic acid of the cell-free nucleic acid sample comprises at least 1 ng of cell-free DNA.

[0028] In some embodiments, the cell-free nucleic acid of the cell-free nucleic acid sample comprises at least 15 ng of cell-free DNA.

[0029] In some embodiments, the method comprises attaching an adaptor to the cell-free nucleic acid molecule comprising the adaptor.

[0030] In some embodiments, the adaptor is resistant to the nucleobase conversion performed in the step of converting the nucleobases of a subset of the cell-free nucleic acid molecules.

[0031] In some embodiments, the nucleic acid probe panel comprises at least 50 unique probes.

[0032] In some embodiments, the method is for sequencing to enhance detection of differentially methylated regions for assessing the status of an individual.

[0033] In some embodiments, the method includes preparing a cell-free nucleic acid sample for targeted methyl-sequencing.

[0034] In some embodiments, the prepared cell-free nucleic acid sample has been collected from an individual and comprises at least 100,000 cell-free nucleic acid molecules from multiple regions known to be differentially methylated in a condition.

[0035] In some embodiments, the method includes sequencing the cell-free nucleic acid sample via a high-throughput sequencer to obtain sequencing results of cell-free nucleic acid molecules from multiple regions known to be differentially methylated in a condition.

[0036] In some embodiments, the method includes calculating a methylation metric using a computing device and the sequencing results, wherein the methylation metric is calculated for one region of a plurality of regions known to be differentially methylated in a condition.

[0037] In some embodiments, the methylation metric indicates the amount of methylation of cell-free nucleic acid molecules that align to the region.

[0038] In some embodiments, the method includes using a computing device to input the calculated methylation metrics as features into a computational model to obtain an assessment of the cell-free nucleic acid sample.

[0039] In some embodiments, the assessment indicates that the individual has the condition.

[0040] In some embodiments, the method includes aligning each cell-free nucleus molecular sequencing result across a region.

[0041] In some embodiments, the region is one of multiple regions that are differentially methylated in a state.

[0042] In some embodiments, the method comprises, for a set of cell-free nucleic acid molecules that are aligned across a region, determining the amount of methylation of each cell-free nucleic acid molecule of the set.

[0043] In some embodiments, the methylation metric is based on at least one cell-free molecule of the set.

[0044] In some embodiments, the method includes determining the number of cell-free nucleic acid molecules in the set that are methylated above a threshold value.

[0045] In some embodiments, the method comprises calculating the molecular methylation fraction (MMF) of the region.

[0046] In some embodiments,

number

[0047] In some embodiments,

number

[0048] In some embodiments,

number

[0049] In some embodiments, the threshold is 60% of the CpGs are methylated.

[0050] In some embodiments, calculating a methylation metric for the region further comprises identifying the most methylated cell-free nucleic acid molecule within the set of cell-free nucleic acid molecules that align across the region.

[0051] In some embodiments, the methylation metric is calculated as the amount of methylation of the most methylated cell-free nucleic acid molecule.

[0052] In some embodiments, each cell-free nucleic acid molecule of the set of cell-free nucleic acid molecules that align across a region has a number of CpGs that is greater than a threshold value.

[0053] In some embodiments, the method includes using a computing device and sequencing results to calculate a methylation metric for each region of at least 50 percent of a plurality of regions known to be differentially methylated in a condition.

[0054] In some embodiments, the method includes using a computing device to input each calculated methylation metric as a feature into a computational model to obtain an assessment of the cell-free nucleic acid sample.

[0055] In some embodiments, the method includes using a computing device and the sequencing results to calculate a methylation metric for each region of a plurality of regions known to be differentially methylated in a condition.

[0056] In some embodiments, the method includes using a computing device to input each calculated methylation metric as a feature into a computational model to obtain an assessment of the cell-free nucleic acid sample.

[0057] In some embodiments, the method includes using a computing device and the sequencing results to calculate a plurality of methylation metrics, each methylation metric being calculated for a region of a plurality of regions known to be differentially methylated in a condition.

[0058] In some embodiments, the method includes using a computing device and the plurality of methylation metrics to calculate a sample summary statistic combining the plurality of methylation metrics.

[0059] In some embodiments, the method includes using a computing device to input the sample summary statistics as features into a computational model to obtain an assessment of the cell-free nucleic acid sample.

[0060] In some embodiments, the sample summary statistic is a percentile of some region with a non-zero MMF.

[0061] In some embodiments, the sample summary statistic is a percentile of some range where the MMF is 0.

[0062] In some embodiments, the sample summary statistic is a percentile of some region where the MMF is greater than a threshold.

[0063] In some embodiments, the sample summary statistic is a percentile of the amount of methylation of the most methylated cell-free nucleic acid molecules.

[0064] In some embodiments, the sample summary statistic is the median amount of methylation of the most methylated cell-free nucleic acid molecules.

[0065] In some embodiments, the sample summary statistic is the skewness of the amount of methylation of the most methylated cell-free nucleic acid molecules.

[0066] In some embodiments, the plurality of regions known to be differentially methylated in a condition comprises at least 10 genomic regions associated with a condition.

[0067] In some embodiments, the plurality of regions known to be differentially methylated in a condition comprises at least 50 genomic regions associated with a condition.

[0068] In some embodiments, the condition is cancer. [Brief explanation of the drawings]

[0069] The description and claims will be more fully understood with reference to the following figures and data graphs, which are presented as exemplary embodiments and should not be construed as a complete recitation of the scope of the disclosure.

[0070]

Figure 1

[0071]

Figure 2

[0072]

Figure 3

[0073]

Figures 4A - B

Figures 4C - D

Figures 4E - G

[0074]

Figures 5A - B

Figures 5C - D

[0075]

Figures 6A - B

[0076]

Figures 7A - B

Figures 7C - E

[0077]

Figures 8A - C

[0078]

Figures 9A - B

Figure 9C

Figure 9D

[0079]

Figures 10A - B

[0080]

Figures 11A - B

Figures 11C - D

Figures 11E - F

[0081]

Figures 12A - C

Figures 12D - E

[0082]

Figures 13A - B

Figure 13C

Figure 13D

[0083]

Figures 14A - C

Figures 14D - G

[0084]

Figures 15A - B

Figures 15C - D

Figures 15E - F

[0085]

Figures 16A - C

Figures 16D - E

[0086] Detailed Description Referring now to the figures and data, in accordance with various embodiments herein, systems and methods for sequencing and analyzing methylated nucleic acids are described. The systems and methods provide a means for detecting cancer-derived cell-free nucleic acid (cfNA) in an individual. The systems and methods improve upon previous cfNA assessment methods and enhance the detection of early-stage cancers.

[0087] Cancer-derived methylation signals have been shown to be detectable in cell-free DNA (cfDNA). For example, aberrant promoter hypermethylation of a few tumor suppressor genes was first detected in the serum of non-small cell lung cancer (NSCLC) patients in 1999 by methylation-specific PCR (M. Esteller et al., Cancer Res. 1999 Jan. 1;59(1):67-70, the disclosure of which is incorporated by reference). More recently, broader high-throughput sequencing of the cell-free methylome has demonstrated success for cancer detection and localization (M. C. Liu et al., DNA. Ann. Oncol. 2020 Jun.;31(6):745-759; and S. Si Shen et al., Nature. 2018 Nov.;563(7732):579-583, the disclosure of which is incorporated by reference). While evidence suggests that some performance results for early lung cancer detection from published studies may be exaggerated, some of the largest studies to date suggest that cfDNA methylation profiling has powerful screening potential. However, despite promising existing methods, there is room for improving sensitivity, particularly in disease-focused approaches. Looking specifically at Grail results for lung cancer, stage I detection remains at approximately 20%, and when broken down by histology, the detection rate for stage I adenocarcinoma is 0–10% (X. Chen et al., Clin Cancer Res. 2021 Aug 1;27(15):4221–4229, the disclosure of which is incorporated herein by reference). This disclosure describes a lung cancer-focused cfDNA methylation assay that integrates both epigenetic and genetic signals to improve noninvasive detection of small tumors.

[0088] The disclosed system and method provides a means for evaluating the methylation of cfNA samples. Generally, the system and method can generate and sequence libraries using materials from cfNA samples. The system and method can be used in any sequencing technology to identify methylation patterns indicative of a condition. Abnormal methylation patterns are now recognized as biomarkers for various conditions, particularly cancer. Various embodiments of the present disclosure provide a means for detecting such conditions in cfNA samples using methylation patterns of various conditions.

[0089] Figure 1 provides an example of a method for sequencing a cell-free nucleic acid sample to detect differentially methylated regions. In this example, a cfNA sample is obtained and processed for targeted methyl-sequencing. This method can be useful for sequencing cfNA samples collected from individuals for the purpose of detecting conditions marked by abnormal methylation patterns. For example, this method can be utilized to sequence cfNA samples to detect abnormal methylation patterns associated with cancer. This method can be used for early detection screening (e.g., prior to any cancer diagnosis), treatment monitoring (e.g., progress and success of anti-cancer therapy), and / or minimal residual disease detection (e.g., detecting cancer recurrence after completion of treatment). Sequencing results can be utilized in downstream applications, such as performing computational analysis on the sequencing results to classify samples to detect conditions.

[0090] Method 100 can begin by obtaining a cfNA sample (101). The cfNA sample can include cell-free RNA (cfRNA) and / or cell-free DNA (cfDNA). In general, abnormal methylation patterns in cfDNA can be useful biomarkers for conditions such as cancer. cfNA can be collected from any extracellular source, such as (for example) blood, plasma, saliva, urine, stool, mucus, lymph, and / or other bodily fluids. Sometimes, cfNA samples are described as liquid biopsies and biological samples, although any description of a sample containing cfNA molecules is applicable. cfNA can be isolated and purified by any suitable means.

[0091] In general, human plasma samples typically contain 0.5–10 ng of cfDNA per mL, corresponding to 150–3,000 copies of the haploid human genome. Some conditions, such as cancer and donor transplant rejection, result in higher levels of circulating cfDNA, with levels exceeding 1,000 ng per mL having been detected. cfDNA typically circulates in fragments ranging from 120–220 bp, with a maximum peak of approximately 167 bp. Thus, plasma typically contains approximately 15–400 million cfDNA copies / mL, with some conditions exceeding 40 billion cfDNA copies / mL. Other biological samples vary in the amount of cfDNA, but generally have concentrations of several million copies per mL of sample, with concentrations higher when affected by conditions characterized by high necrosis and / or DNA leakage, such as cancer.

[0092] In the case of cfDNA, extracted and isolated cfDNA fragments can be used as source nucleic acid molecules.The collection of source nucleic acid molecules for sequencing reaction can have more than 10,000 nucleic acid molecules, more than 100,000 nucleic acid molecules, more than 1,000,000 nucleic acid molecules, more than 10,000,000 nucleic acid molecules, more than 100,000,000 nucleic acid molecules, or more than 1,000,000,000 nucleic acid molecules.

[0093] In some embodiments, a cfNA sample is obtained before any signs of cancer. In some embodiments, a cfNA sample is obtained to provide an early screen for detecting cancer before diagnosis. In some embodiments, a cfNA sample is obtained to detect whether residual cancer is present after treatment. In some embodiments, a cfNA sample is obtained during treatment to determine whether the treatment is producing the desired response. Screening for any particular cancer can be performed. In some embodiments, screening is performed to detect cancers that develop aberrant methylation patterns in stereotyped regions of the genome, such as lung cancer (for example). In some embodiments, screening is performed to detect cancers in which regions of aberrant methylation have been found using previously extracted cancer biopsies, which may be useful for monitoring treatment or detecting minimal residual disease.

[0094] In some embodiments, the cfNA sample is obtained from an individual with a determined risk of developing cancer, such as an individual with a family history of the disorder or a determined risk factor (e.g., exposure to a carcinogen). In many embodiments, the cfNA sample is obtained from any individual within the general population. In some embodiments, the cfNA sample is obtained from an individual within a specific age group at higher risk of cancer, such as an elderly individual over the age of 50. In some embodiments, the cfNA sample is obtained from an individual who has been diagnosed with and treated for cancer.

[0095] The method 100 can further generate a sequencing library that targets differentially methylated regions (103). Generally, targeted sequencing can be performed by capturing and / or specifically amplifying specific regions of the genome. In some embodiments, adapters and / or primers are attached to the cell-free nucleic acids to facilitate sequencing.

[0096] Any suitable amount of input cfDNA can be used in library preparation. The limit of detection (LOD) can be affected by the amount of input cfDNA. When evaluating deduplication depth, it has been found that only 1 ng of cfDNA can be used. However, in order to improve the detection sensitivity of differences in methylation patterns, more cfDNA is useful. In various embodiments, the amount of input cfDNA for library preparation is at least 1 ng, at least 2.5 ng, at least 5 ng, at least 10 ng, at least 15 ng, at least 20 ng, at least 25 ng, or at least 30 ng.

[0097] In some embodiments, targeted sequencing of specific genomic loci is to be performed, and thus specific sequences corresponding to specific loci are captured via hybridization prior to sequencing (e.g., capture sequencing). In some embodiments, capture sequencing is performed using a probe panel that pulls down (or captures) regions found to be differentially methylated for a particular cancer (e.g., lung cancer). In some embodiments, capture sequencing is performed using a probe panel that pulls down (or captures) regions found to be differentially methylated, as determined by methyl sequencing of cancer biopsies.

[0098] In various embodiments, the probe panel comprises at least 10 unique probes, at least 20 unique probes, at least 50 unique probes, at least 100 unique probes, at least 150 unique probes, at least 200 unique probes, at least 250 unique probes, at least 500 unique probes, or at least 1000 unique probes.

[0099] Table 2 provides a set of genomic loci for detecting aberrantly methylated regions in non-small cell lung cancer (NSCLC), particularly lung adenocarcinoma (LUAD) and LUSC. All or part of these regions can be used to evaluate NSCLC. Furthermore, these regions can be used to distinguish between LUAD and LUSC, which can be useful for determining treatment options.

[0100] In various embodiments, a panel of capture nucleic acid probes can be designed to hybridize to at least 5%, at least about 10%, at least 20%, at least 30%, at least about 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% of the genomic regions listed in Table 2. In various embodiments, the sequences of the probes are at least 50% complementary, at least 60% complementary, at least 70% complementary, at least 80% complementary, at least 90% complementary, at least 95% complementary, or at least 99% complementary to sequences within the genomic regions listed in Table 2. A standard genome reference, such as hg19, can be used to search for sequences in the genomic regions listed in Table 2.

[0101] If an individual is known to have cancer, methyl-sequencing of the cancer can be performed to identify regions that are differentially methylated in relation to cancer tissue. Nucleic acid probes can be designed to hybridize to these identified regions, so that methyl-sequencing cfNA can better detect the presence of cancer-derived cfNA in the individual's biological sample. This personalized method using probes designed to hybridize to identified differentially methylated regions can improve the ability to detect the presence of cancer when assessing therapeutic progress and / or detecting minimal residual disease. In various embodiments, the panel of capture nucleic acid probes can be designed to hybridize to at least 10 genomic regions, at least 20 genomic regions, at least 50 genomic regions, at least 100 genomic regions, at least 1500 genomic regions, at least 200 genomic regions, at least 250 genomic regions, at least 500 genomic regions, at least 750 genomic regions, or at least 1000 genomic regions. In various embodiments, the sequence of the probe is at least 50% complementary, at least 60% complementary, at least 70% complementary, at least 80% complementary, at least 90% complementary, at least 95% complementary, or at least 99% complementary to a sequence identified as differentially expressed in the individual's cancer.

[0102] In some embodiments, capture sequencing is performed using a panel of probes that pull down (or capture) regions associated with other useful information that may be useful in performing a diagnosis. For example, specific methylated regions may be useful in providing an indication of factors associated with cancer, such as regions where methylation patterns correlate with age, smoking history, and body mass index (BMI). Table 1 provides a number of regions associated with various factors, including age, BMI, cell type (cibersortX site), tissue origin (miniselector), multiple cancers, pan-cancer, smoking history, and BMI.

[0103] In various embodiments, a panel of capture nucleic acid probes can be designed to hybridize to at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, or about 100% of the genomic regions identified in Table 1. In various embodiments, the sequences of the probes are at least 50% complementary, at least 60% complementary, at least 70% complementary, at least 80% complementary, at least 90% complementary, at least 95% complementary, or at least 99% complementary to sequences within the genomic regions identified in Table 1. A standard genome reference, such as hg19, can be used to search for sequences of the genomic regions listed in Table 1.

[0104] In some embodiments, capture sequencing is performed using probe panels that pull down (or capture) regions that serve as controls, such as invariably hypermethylated and / or invariably hypomethylated regions. In some embodiments, capture sequencing is performed using probe panels that pull down (or capture) regions that are associated with other forms of molecular information useful for diagnosis, such as single nucleotide variants, insertions, deletions, and copy number variants.

[0105] A particular region that is differentially methylated in a particular condition may also be differentially methylated for other reasons unrelated to that condition. For example, a region may be differentially methylated between healthy individuals, upon environmental stimuli, in association with different cell types, etc. Differential methylation in blood cells is of particular concern because these cells are a high source of cfNAs. Detection of differential methylation in these regions may result in false positives. Therefore, in some embodiments, the probe panel excludes regions known to be associated with false negatives. Also, in some embodiments, the probe panel excludes regions known to be differentially methylated in blood cells.

[0106] In some embodiments, the adapters utilized for sequencing are resistant to conversion of nucleobases, for example, some adapters contain methylated cytosines that resist conversion via bisulfite treatment.

[0107] Method 100 can further convert nucleobases (105), which can distinguish methylated nucleobases from unmethylated nucleobases. Any method for converting nucleobases can be used. In some embodiments, chemical conversion is performed. In some embodiments, enzymatic conversion is performed. Various methodologies can be used to convert methylated nucleobases, including, but not limited to, bisulfite treatment, TET2 oxidation and APOBEC3A conversion, and TET2 oxidation and pyrazine borane treatment. In experiments performed and described in the Examples section, it was found that the TET2 oxidation and pyrazine borane treatment and TET2 oxidation and APOBEC3A conversion methods provided higher read mappability than bisulfite treatment. Furthermore, it was found that TET2 oxidation and APOBEC3A conversion provided better unique molecule recovery. Therefore, in some preferred embodiments, TET2 oxidation and APOBEC3A are used to convert methylated nucleobases. For further details regarding nucleobase conversion, including methods, data, and results, see the Examples section herein.

[0108] Although methylation is primarily described throughout, the systems and methods can be used with a variety of modified nucleobases, including, but not limited to, 5-methylcytosine (5-mC), 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5-carboxylcytosine (5-caC), depending on the biomarker for that condition. Various modification sequencing and detection methods are available, as reported, for example, in R.P. Darst, Curr Protoc Mol Biol. 2010 Jul; Chapter 7: Unit 7.9.1-17; C.X.Song et al., Nat Biotechnol. 2011 Jan; 29(1):68-72; J. Morrison et al., Epigenetics Chromatin. 2021 Jun 19; 14(1):28; F. Erger et al., Genome Med. 2020 Jun 24; 12(1):54; M.J. Booth et al., Nat Chem. 2014 May; 6(5):435-40; X. Lu et al., J Am Chem Soc. 2013 Jun 26; 135(25):9315-7; J. Xiong et al., Chem Sci. 2022 Aug. 11;13(34):9960-9972; and ABR McIntyre et al., Nat Commun. 2019 Feb 4;10(1):579 (the disclosures of each of which are incorporated by reference).

[0109] It should be noted that some sequencing platforms, such as nanopore sequencing technology, can detect methylation without conversion.If sequencing technology can directly detect methylation, the nucleic acid base conversion step can be omitted.Examples of sequencing platforms that can detect methylation include (but are not limited to) Oxford Nanopore Technologies PromethION, MinION, and GridION sequencing platforms (Oxford, UK) and Pacific Bioscience's single molecule real-time (SMRT) sequencing platform (Menlo Park, California).

[0110] Method 100: The generated and converted library is further sequenced (107) to detect the methylation status of the differentially methylated regions. Any suitable high-throughput sequencing technology capable of detecting converted nucleic acid bases can be utilized. High-throughput sequencing technologies include (but are not limited to) 454 sequencing, Illumina sequencing, SOLiD sequencing, Ion Torrent sequencing, single-read sequencing, paired-end sequencing, etc. High-throughput sequencing methods can simultaneously sequence at least about 10,000, at least about 100,000, at least about 1 million, at least about 10 million, at least about 100 million, or at least about 1 billion cfNA molecules.

[0111] Some embodiments are directed to using computational models to detect the presence of a condition using methyl-sequencing data from cfNA samples. Interpreting methyl-sequencing data results can be difficult when methylation distinctions in regions are not easily recognized. For example, detecting stage I lung cancer via cfDNA is challenging due to the low amount of cfDNA markers present in liquid biopsies. As described in the Examples section herein, characterizing methyl-sequencing results and utilizing these features in computational classifiers has been shown to improve the ability to detect stage I cancer in plasma samples. Classifying cfNA samples derived from cancer can allow clinical intervention for the individual.

[0112] Figure 2 provides an example of a computational method for classifying cfNAs based on methyl-sequencing results. Method 200 can begin by obtaining targeted methyl-sequencing results for a cfNA sample (201). Typically, the sequencing results are obtained via high throughput so that the methylation of a large number of cfNA molecules is determined. Sequencing can target specific regions of the genome, particularly regions known to be differentially methylated in certain regions. In some embodiments, lung cancer is evaluated, and at least some of the genomic regions identified in Table 2 are targeted for methyl-sequencing. In some embodiments, the cancer is sequenced to identify differentially methylated genomic regions so that subsequent cfNA sequencing targets those regions. In some embodiments, sequencing is performed as described with reference to Figure 1.

[0113] In various embodiments, the targeting sequence results cover at least 10 genomic regions, at least 20 genomic regions, at least 50 genomic regions, at least 100 genomic regions, at least 1500 genomic regions, at least 200 genomic regions, at least 250 genomic regions, at least 500 genomic regions, at least 750 genomic regions, or at least 1000 genomic regions. The genomic regions may include regions associated with a condition (e.g., cancer), regions correlated with factors associated with a condition, and / or control regions.

[0114] Using the sequencing results, method 200 can further assess the methylation of cfNA molecules (203). Methylation assessment can be performed in a variety of ways, but generally, the assessment can include the amount of cfNA molecules having at least one methylated nucleobase and / or the amount of methylated nucleobases per cfNA molecule. When assessing methylation, in various embodiments, at least 100 cfNA molecules of the sample are assessed, at least 1000 cfNA molecules of the sample are assessed, at least 10,000 cfNA molecules of the sample are assessed, at least 100,000 cfNA molecules of the sample are assessed, at least 1,000,000 molecules of the sample are assessed, or at least 10,000,000 molecules of the sample are assessed.

[0115] In some embodiments, methylation is evaluated by calculating a methylation metric that indicates the amount of methylation of cfNA molecules at a specific locus, taking into account the cfNA molecules that align to that locus.To perform this evaluation, a set of cfNA molecules that align to a region is used.In some embodiments, the region includes a CpG island.In various embodiments, each cfNA molecule used for evaluation contains at least 2 CpGs, at least 4 CpGs, at least 6 CpGs, at least 8 CpGs, at least 10 CpGs, at least 15 CpGs, or at least 20 CpGs in the region.In some embodiments, the methylation molecular fraction (MMF) is the number of cfNA molecules that have the amount of methylation in that region that exceeds a threshold value per total cfNA molecules evaluated for that region.

number

[0116] The MMF can be calculated for any or all regions sequenced in association with a particular trait (e.g., any or all regions differentially methylated in cancer). In various embodiments, the MMF is calculated for at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or all regions sequenced in association with a particular trait.

[0117] Method 200 optionally calculates sample summary statistics (205). Once methylation metrics for several regions have been determined, an overall sample summary statistic can be generated by combining the methylation metrics for multiple regions. In some embodiments, the sample summary statistic combines the methylation metrics for all regions associated with the sequenced trait. In some embodiments, the sample summary statistic combines the methylation metrics for a subset of regions associated with the sequenced trait. For example, several top informative regions can be combined to obtain a sample summary statistic. The most informative regions are those determined to have a greater association with or more predictive power for a condition (e.g., cancer) when compared to other regions evaluated. In one hypothetical example, the sample summary statistic can be determined by combining the methylation metrics for several regions with the greatest association with or the greatest predictive power for a condition. In various embodiments, summary sample statistics are obtained by combining at least 10% of the areas assessed, by combining at least 20% of the areas assessed, by combining at least 30% of the areas assessed, by combining at least 40% of the areas assessed, by combining at least 50% of the areas assessed, by combining at least 60% of the areas assessed, by combining at least 70% of the areas assessed, by combining at least 80% of the areas assessed, by combining at least 90% of the areas assessed, or by combining all areas assessed.

[0118] Sample summary statistics can be calculated via several different methods. In some embodiments, any statistic that can combine methylation metrics for multiple regions can be utilized. In some embodiments, percentiles of the distribution of MMF regions for a sample are determined, and the percentiles are determined by comparing to a cohort of samples. For example, the percentile of the number of regions with a non-zero MMF compared to the cohort is determined. In another example, the percentile of the number of regions with an MMF of 0 compared to the cohort. In another example, for the number of regions with an MMF > X% compared to the cohort, in some embodiments, X is any percentage between 0.01% and 1%, and in some embodiments, X is 0.1%. In another example, for each region, the cfNA molecules with the most methylated CpGs are further considered to determine the median methylated CpG content, and the X percentile methylated CpG content, where X is any percentage between 1% and 100%, is determined, and the skewness of the methylated CpG content is determined. In some embodiments, the sample summary statistics are normalized based on the length of the cfNA molecule. As will be appreciated from these examples, many other sample statistics can be determined.

[0119] The method 200 can further include inputting one or more methylation features (207) into the trained computational model to evaluate the sample, with the results indicating whether the sample is associated with a condition (e.g., cancer). The methylation feature includes a computational assessment of cfNA methylation. The feature can be based on methylation of a specific region associated with the condition being evaluated. The feature can be based on summary sample statistics of multiple regions associated with the condition being evaluated.

[0120] In some embodiments, the methylation assessment of a particular region is utilized as a feature. For example, a methylation metric calculated for a particular region can be utilized as a feature. In some embodiments, each feature of the plurality of features is based on a particular region, and the feature utilized in the model is based on the predictive ability of the feature or its association with the methylation of the region and the condition.

[0121] In some embodiments, multiple methylation assessments, each a sample summary statistic that combines the methylation assessment of a particular region, e.g., any sample statistic calculated in step 205, can be used as a feature.

[0122] Any suitable machine learning model and architecture can be utilized as the model. In some implementations, multiple trained machine models are utilized and / or combined (e.g., ensemble models). Implementable ML models include, but are not limited to, regression-based and / or classification-based models. Generally, regression-based models provide a score indicative of the likelihood of cancer, while classification-based models classify samples as likely to contain or not contain cancer. Regression-based models include, but are not limited to, lasso regression, ridge regression, k-nearest neighbors, elastic nets, least angle regression (LAR), and random forest regression. Classification-based models include, but are not limited to, support vector machines (SVMs), decision trees, random forests, and naive Bayes. In some embodiments, the regression-based model or classification-based model is regularized, while in various embodiments, the regression-based model or classification-based model is gradient boosted.

[0123] A computational model can be trained using a cohort of cfNA samples. For example, cfNA samples from a cohort can be used to train a lung cancer classifier. In some embodiments, a leave-one-out cross-validation (LOOCV) machine learning model is used to build and train the model. In each LOOCV round, the model is repeatedly trained on all samples except for one sample that is left out. Model performance can be evaluated on the left-out sample. LOOCV training is attractive because it reduces overfitting and provides a more accurate assessment of overall stability.

[0124] Method 200 may perform clinical intervention as needed (209) if the ML model indicates that the cfDNA sample contains cfDNA molecules derived from cancer. Clinical intervention may include further clinical evaluation or administration of treatment to the individual. In some embodiments, clinical procedures such as (e.g.,) blood tests, genetic testing, medical imaging, physical examination, tumor biopsy, or any combination thereof are performed. In some embodiments, a diagnosis is performed to determine the specific stage of the cancer. In some embodiments, treatment is performed, such as (e.g.,) chemotherapy, radiation therapy, chemoradiotherapy, immunotherapy, hormone therapy, targeted drug therapy, surgery, transplant, blood transfusion, medical monitoring, or any combination thereof. In some embodiments, the individual is evaluated and / or treated by a medical professional, such as a physician, internist, physician's assistant, nurse practitioner, nurse, caregiver, nutritionist, etc.

[0125] In some embodiments, non-limiting examples of treatments may include chemotherapy, radiation therapy, chemoradiotherapy, immunotherapy, adoptive cell therapy (e.g., chimeric antigen receptor (CAR) T-cell therapy, CAR NK cell therapy, modified T-cell receptor (TCR) T-cell therapy, etc.), hormone therapy, targeted drug therapy, surgery, transplantation, blood transfusion, or medical surveillance. Treating a subject's condition may include administering one or more therapeutic agents to the subject. The one or more therapeutic agents may be administered to the subject by one or more of the following: orally, intraperitoneally, intravenously, intraarterially, transdermally, intramuscularly, via liposomes, via local delivery by catheter or stent, subcutaneously, intraadipose, and intrathecally.

[0126] computer processing system The computing system for evaluating differentially methylated regions in cfNAs to detect conditions according to various methods of the present disclosure typically utilizes a processing system including one or more CPUs, GPUs, and / or neural processing engines. In some embodiments, the methyl sequencing results of cfNAs are processed and evaluated to detect conditions using the computing system. In some embodiments, the computing system is housed within a computing device associated with the sequencer. In some embodiments, the computing system is housed separately from the sequencer and receives the sequencing results. In certain embodiments, the computing system is implemented using a software application on a computing device such as, but not limited to, a mobile phone, a tablet computer, a wearable device (e.g., a watch), and / or a portable computer.

[0127] A computing system according to various embodiments of the present disclosure is illustrated in FIG. 3. The computing system 300 includes a processor system 302, an I / O interface 304, and a memory system 306. As can be readily appreciated, the processor system 302, the I / O interface 304, and the memory system 306 may be implemented using any of a variety of components suitable for the requirements of a particular application, including, but not limited to, a CPU, a GPU, an ISP, a DSP, a wireless modem (e.g., WiFi, Bluetooth modem), a serial interface, a depth sensor, an IMU, a pressure sensor, an ultrasonic sensor, volatile memory (e.g., DRAM), and / or non-volatile memory (e.g., SRAM and / or NAND flash). The memory system may store sequencing data 308, an application 310 for feature generation, and a computational model 312 for detecting conditions. The application may be downloaded and / or stored in non-volatile memory. When executed, the application for feature generation and / or the computational model for detecting a condition can configure the processing system to implement computational processes including, but not limited to, the computational processes described above and / or combinations and / or modified versions of the computational processes described above. In some embodiments, the application for feature generation 310 utilizes the sequence data 308 to generate features based on differentially methylated regions. The computational model for detecting a condition utilizes the generated features to determine whether a cfNA sample is from an individual with a condition, such as cancer. Intermediate data and / or final results can be temporarily stored in a memory system during processing and / or saved for use in downstream applications.

[0128] While a particular computing system is described above with reference to Figure 3, it should be readily understood that the computational and / or other processes utilized in providing for assessing differentially methylated regions of cfNAs according to various embodiments of the present disclosure can be implemented in any of a variety of processing devices, including combinations of processing devices. Accordingly, it should be understood that a computing apparatus according to the present disclosure is not limited to a particular computing system, but can be implemented using any combination of the systems described herein and / or modified versions of the systems described herein to perform the processes, combinations of processes, and / or modified versions of the processes described herein. [Example]

[0129] The systems and methods of the present disclosure will be better understood through some examples provided. Verification results are also provided.

[0130] Establishment of a cell-free DNA methylation detection assay Cancer personalized profiling by deep sequencing (CAPP-Seq), a method developed to detect tumor variants in cfDNA in a disease-specific manner, served as a template for designing a novel cfDNA methylation detection method. While most existing cfDNA methylation methods focus on broad genome coverage, current methodologies aim to target relatively small portions of the genome. This allows for high-depth coverage at relatively low sequencing costs and allows for the incorporation of barcoding and error suppression, as with CAPP-Seq. Furthermore, with the understanding that the initial blood-based detection test would best serve high-risk populations, we maintained the assay disease-specific. To achieve these goals, we designed a targeted sequencing panel covering a range of informative methylation states for lung cancer detection.

[0131] Design of a targeted methylation sequencing panel for NSCLC detection In SNV-based CAPP-Seq, the sequencing space is focused on regions containing recurrent mutations across patients to maximize the number of variants observed per patient while keeping sequencing costs (total bases covered) relatively low. Similarly, when designing a targeted methylation panel, the goal was to enrich for regions that distinguish lung tumor-derived ctDNA molecules from healthy background, primarily blood-derived DNA that constitutes the remainder of the cfDNA pool, and also distinguish lung cancer from normal lung tissue. To do this, we used publicly available Infinium HumanMethylation450 (450k) array data from lung tumors from The Cancer Genome Atlas (TCGA). 15、16 from, as well as blood from public datasets. 17~23 and normal lung 24、25The samples were downloaded and differentially methylated regions were identified. Comparison of lung adenocarcinoma (LUAD) with blood and normal lung identified numerous differentially methylated CpGs (Figures 4A and 4B) (EA Collisson et al., Nature. 2014 Jul 31;511(7511):543-50; PS Hammerman et al., Nature. 2012 Sep 27;489(7417):519-25; A. Arpon et al., Nutrients. 2017 Dec 23;10(1):15; M. V. Dogan et al., BMC Genomics. 2014 Feb 22;15:151; L. E. Reinius et al., PLoS One. 2012;7(7):e41361; S. A. Langie et al., PLoS One. 2016 Mar 21;11(3):e0151109; S. Horvath et al., Genome Biol. 2012 Oct 3;13(10):R97; W.P. Accomando et al., Enome Biol. 2014 Mar 5;15(3):R50; R.A. Harris et al., Inflamm Bowel Dis. 2012 Dec;18(12):2334-41; M.M. Bjaanaes et al., Mol Oncol. 2016 Feb;10(2):330-43; and J. Shi et al., Nat Commun. 2014 Feb 27;5:3365 (the disclosures of which are incorporated herein by reference). The top 500 most differentially methylated CpGs (DMCs) ranked by the absolute difference in mean beta values ​​(Δβ) between LUAD and normal tissues were selected, with a maximum Benjamini-Hochberg adjusted P value of 0.0001. By filtering this list to probes with low background in other normal tissues, including brain, liver, and skin, we obtained 412 DMCs for LUAD for inclusion in panel design (Figure 4B). Of these, 379 were hypermethylated in LUAD compared to blood and normal lung, while 33 were hypomethylated. A similar analysis of lung squamous cell carcinoma (LUSC) selected 370 differentially methylated CpGs (Figures 4C and 4D). Of these, 223 were hypermethylated in LUSC compared to blood and normal lung, while 147 were hypomethylated.131 differentially methylated CpGs were shared between LUAD and LUSC, and additional CpGs selected only for one histology showed methylation levels trending in the same direction in the other histology, suggesting that these sites may aid in detection for both histologies (Figure 4E). A total of 651 CpGs were selected as differentially methylated sites (NSCLC DMRs) for the sequencing panel. The majority of the 651 selected sites were located in CpG islands, as expected based on the composition of the 450K array (Figure 4F). Most NSCLC DMRs were associated with genes (481 / 651), many of which were located upstream or in promoters of genes (Figure 4G).

[0132] The results of this differential methylation analysis confirmed previous reports. For example, differentially methylated CpGs were identified in several homeobox and homeobox-associated genes, including HOXA10, HOXA11, HOXD12, OTX1, OTX2, PTX1, PAX6, and PAX9, consistent with previous NSCLC studies (M. Shiraishi et al., Oncogene. 2002 May 16;21(22):3659-62; JA Hwang et al., Oncotarget. 2013 Dec;4(12):2317-25; and T. Rauch et al., Proc Natl Acad Sci US A. 2007 Mar 27;104(13):5527-32, the disclosures of which are incorporated herein by reference). TP73 was identified as hypomethylated in LUSC, as previously observed (A. Daskalos et al., Cancer Lett. 2011 Jan 1;300(1):79-86, the disclosure of which is incorporated herein by reference). To further validate our selected markers, we downloaded 450k array methylation data from NSCLC cell lines and found that the cell lines displayed similar methylation status to TCGA primary tumors at our selected sites (Figure 5A) (K. Walter et al., Clin Cancer Res. 2012 Apr 15;18(8):2360-73, the disclosure of which is incorporated herein by reference).

[0133] Interestingly, when we examined the methylation status of 651 NSCLC DMRs in other cancers derived from TCGA, we found that many other cancer types shared similar methylation status in these regions with NSCLC, suggesting that these marks are not lung cancer-specific (Figure 5B) and may reflect a "cancer" signature. To maintain the ability to specifically distinguish lung cancer from other cancer types when evaluating patient cfDNA, we selected a subset of major cancer types from TCGA-LUAD, LUSC, bladder (BLCA), breast (BRCA), colorectal (COADREAD), B-cell lymphoma (DLBCL), hepatocellular carcinoma (LIHC), pancreatic (PAAD), and prostate (PRAD) and performed differential methylation analysis to identify cancer-type-specific methylation signals. Setting a minimum mean methylation difference (Δβ) of 0.3 between each cancer type and all other cancers and healthy blood (1 vs. others), we obtained a variable number of differentially methylated CpGs for each cancer type (range 1-33, median 15.5) (Figure 5C). Interestingly, comparisons between NSCLC and other cancer types failed to identify hypermethylated sites specific to lung cancer. However, we reasoned that assessing the methylation status of regions specific to other cancers would allow us to confidently assert lung cancer detection via the lack of signal in those regions.

[0134] In addition to regions that differed between lung tumors and the background of healthy blood and normal lung tissue, we wanted to include as controls regions whose methylation status was consistent across these different tissue types. To do this, we selected CpG sites from the 450k array data whose methylation levels were consistently either hypermethylated or hypomethylated across all sample types, with low variability. This analysis generated a list of 683 control CpGs, including 449 hypomethylated and 234 hypermethylated controls (Figure 5D, Table 1). We also included CpGs whose methylation status in blood has been shown to correlate with age, smoking history, and BMI (Table 2) (A.T. Lu et al., Nat Aging. 2023 Sep;3(9):1144-1166; S. Bocklandt et al., PLoS One. 2011;6(6):e14821; R. Joehanes et al., Circ Cardiovasc Genet. 2016 Oct;9(5):436-447; X. Gao et al., Clin Epigenetics. 2015 Oct 16;7:113; and S. Wahl et al., Nature. 2017 Jan 5;541(7635):81-86, the disclosures of which are incorporated herein by reference). Finally, we added regions containing the most commonly mutated genes in NSCLC, including TP53, KRAS, EFGR, and NFE2L2, among others, to maintain the possibility of performing SNV genotyping.

[0135] Validation of panel design in cell-free DNA A targeted sequencing panel was designed to detect the presence of lung tumor-derived DNA in a background of healthy cfDNA. Therefore, to mimic this scenario in a controlled manner, a mixture of sheared NCI-H441 lung cancer cell line DNA was generated from cfDNA from healthy donors at successively decreasing tumor fractions: 100% tumor (pure cell line), 5% tumor, 0.5% tumor, 0.05% tumor, 0.005% tumor, and 0% tumor (pure cfDNA). Sequencing libraries from each tumor fraction were generated in triplicate using a premethylation-CAPP-Seq (mCAPP-Seq) protocol. In this protocol, conversion-resistant methylated Y adapters were ligated to the DNA before bisulfite conversion and final PCR. Libraries were captured with the targeted sequencing panel and sequenced on a HiSeq4000. Alignment to the human genome and methylation calling was performed using Bismark (F. Krueger and SR Andrews; Bioinformatics. 2011 Jun 1;27(11):1571-2, the disclosure of which is incorporated herein by reference).

[0136] Regions hypermethylated in TCGA NSCLC (high DMR) were also found to be hypermethylated in NCI-H441 cell line DNA compared to healthy cfDNA (Figure 6A). Furthermore, methylation beta values ​​in high DMR tumors were substantially different from healthy cfDNA at a 5% tumor rate (Figure 6B).

[0137] The first bioinformatic approach for the detection of tumor-derived methylation signals in cfDNA DNA methylation states at adjacent CpGs are nonrandom and highly correlated. This observation has been exploited for the deconvolution of plasma cell-free DNA, with the hypothesis that fragment-level CpG patterns exhibit greater specificity than standard single-CpG beta values. Therefore, we sought to develop a method that similarly utilizes fragment-level methylation states. To do this, we first set a threshold that required molecules to cover at least 10 CpGs. Focusing on high DMRs, we then required that at least 80% of the CpGs in a fragment be methylated for the fragment to be considered highly methylated (and therefore likely tumor-derived). For each DMR region, we calculated the methylation "allele fraction" (AF) as the proportion of fragments in the region that had a highly methylated state.

[0138] Looking at the distribution of these DMR AFs in the cell line spiked experimental data, we observed that methylation AFs correlated with the intended spiked tumor fraction, confirming their biological validity (Figures 7A and 7B). Next, we sought to develop a methodology to determine whether a given DMR AF was elevated compared to that expected by chance. For a given DMR, we randomly sampled fragments from the selector's hypomethylated control region equal to the number of fragments in the DMR of interest. The AFs of highly methylated fragments in these control region fragments were calculated using the same threshold as above (a minimum of 10 CpGs with a minimum of 80% methylation), and this process was repeated 1,000 times to generate a distribution of background AFs. Comparing the AF of the DMR of interest to its background generated an empirical P value and the percentile at which the DMR was located relative to the background. This process was repeated for each high DMR, and DMRs falling in the 100th percentile were considered "above background." The total DMRs above background were then summed for each sample. Applying this methodology to cell line spiking experiments revealed a dose-response relationship between the number of DMRs above background and tumor fraction, suggesting that this metric is a good measure of plasma tumor burden (Figure 7C). All spike levels except the 0.005% condition had significantly higher DMRs above background counts than unspiked healthy cfDNA (0%) (Student's T-test, Figure 7C). As further validation of this method, we sequenced 12 additional healthy control samples and 12 advanced-stage (Stage IV) patient samples using the preliminary mCAPP-Seq protocol and calculated their DMRs above background. Confirming the results of the spiking experiments, patient cfDNA samples showed elevated DMRs above background compared to control cfDNA (Figure 7D).

[0139] Establishing the detection limit of the assay The controlled nature of the tumor fraction in the admixture experiment is useful for estimating the true limit of detection (LOD) of the assay, providing an early sense of the assay's performance in clinical samples. The total DMR above background was found to be significantly elevated at tumor fractions as low as 0.5% compared to 0% tumor (pure cfDNA) (Figure 7C, P<0.01, Student's T-test). A log-linear relationship was observed between spike levels of AF between 0.005% and 0.5% and DMR above background. Above that level, DMR above background saturated; below that, 0% AF was equivalent to 0.005% AF. The LOD was estimated by focusing on the linear portion of the curve. Because 0% and 0.005% were equivalent, those six samples (labeled 0.005% for the purposes of the LOD calculation) were combined and the mean + 3 standard deviations was calculated as a conservative estimate of the upper limit of the DMR metric above background. We then estimated the AF achieved so that the DMR value would be approximately 0.013% (Figure 7E). This analysis estimated that as little as 0.02% of tumors would be significantly detected above background. Previous analysis based on tumor information from the Lung-CLiP cohort showed that at least half of stage I NSCLCs had plasma ctDNA levels below 0.01%. This suggests that a significant portion of early-stage tumors may remain undetectable by the mCAPP-Seq assay. However, the orthogonality of methylation signals (compared to SNV signals) may still allow for the detection of cases missed by CLiP due to low mutation counts or low genome-wide copy number changes.

[0140] Optimizing molecular biology for targeted cfDNA methylation sequencing Given the limited amount of plasma that can be obtained from a given patient and the high sequencing depth required to resolve low tumor fraction events, it was important to optimize our molecular biology protocol for this particular application.

[0141] Comparison of molecular transformation methods for detecting methylated cytosines Sodium bisulfite conversion has long been the standard for detecting CpG methylation at base resolution. Bisulfite conversion works by deaminating unmethylated cytosines and converting them to uracil, while leaving 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC) intact. Thus, after sequencing, when comparing a given read to a reference genome, sites with a C in the reference but read as a T can be inferred to be unmethylated, while sites with a reference C read as a C can be inferred to be methylated.

[0142] However, bisulfite treatment is known to result in harsh temperature and pH conditions that can damage DNA and lead to material loss. Recently, enzymatic conversion approaches have emerged as alternatives to bisulfite, improving conversion and recovery. Therefore, we designed experiments to determine whether these alternatives outperform bisulfite for cfDNA applications and to compare bisulfite with two enzymatic approaches.

[0143] Two enzymatic alternatives are enzymatic methyl-seq (EM-Seq) and TET-assisted pyrazineborane sequencing (TAPS) (R. Vaisvila et al., Genome Res. 2021 Jul;31(7):1280-1289; and Y. Liu et al., Nat Biotechnol. 2019 Apr;37(4):424-429, the disclosures of which are incorporated herein by reference). EM-Seq uses the TET2 enzyme to first oxidize 5mC to 5caC, protecting it from further conversion by APOBEC3A. APOBEC3A then deaminates C and 5mC (but not 5caC), converting them to Ts. Thus, unmethylated cytosine is converted to thymine, and methylated cytosine is protected from conversion, resulting in the same sequence as bisulfite would produce. TAPS also begins with the TET-mediated oxidation of 5mC to 5caC. However, TAPS then proceeds with a chemical treatment with pyridine borane to reduce those 5caCs to dihydroxyuracil (DHU), which is then converted to Ts by PCR. Thus, unlike bisulfite or EM-Seq, TAPS converts methyl-Cs to Ts, resulting in a different final sequence. By converting only methylated Cs, the final TAPS sequence will have higher complexity. However, chemical treatments can also prove harsh on DNA. All three conversion methods were compared and evaluated for their performance using cfDNA.

[0144] The primary readouts for this experiment are molecular recovery (measured as specific sequencing depth), conversion efficiency, and mapping rate. Conversion efficiency can be measured using non-human or synthetic spike-in control DNA that is either fully unmethylated (for bisulfite and EM-Seq) or fully methylated (for TAPS). Mapping rate can be measured using any reads distributed across the human genome. However, specific molecular recovery will be best measured with high-depth sequencing, which requires a targeted sequencing approach. We designed a 44 kb "mini-selector" from the larger sequencing panel that covers the DMR region, control regions, and several tissue-specific regions. Because TAPS generates different final sequences than bisulfite or EM-Seq, separate probe pools for bisulfite / EM-Seq and TAPS were required. We also wanted to compare the three conversion methods with unconverted DNA as a reference. Therefore, we designed a third bait set for unconverted DNA that covers the same region.

[0145] We prepared methylation sequencing libraries from three healthy donor cfDNA samples using either bisulfite conversion, EM-Seq, TAPS, or no conversion with a fixed DNA input. We captured the libraries with the appropriate bait set and subjected the samples to sequencing. TAPS was found to have a higher mapping rate than either EM-Seq or bisulfite (Figure 8A). However, after deduplication, EM-Seq demonstrated higher unique molecule recovery than either of the other two conversion methods, despite equal DNA input (Figure 8B). Furthermore, EM-Seq demonstrated significantly higher conversion efficiency than bisulfite, as measured by unmethylated lambda control DNA (Figure 8C). EM-Seq was selected as the preferred conversion method for the mCAPP-Seq protocol moving forward.

[0146] Minimum input requirements for EM-Seq Next, given the precious and limited nature of cfDNA material as suggested above, it is desirable to optimize the amount of DNA that should be input into library preparation. As a first step, we desired to evaluate the lower limit of the input required to generate a successful sequencing library. To do this, we prepared libraries in duplicate using five different cfDNA inputs ranging from 30 ng to 1 ng. After sequencing, we found that deduplication depth correlated with DNA input and that libraries could be prepared with as little as 1 ng of cfDNA (Figure 9A). However, the loss of depth associated with this low input may not be suitable for low-AF detection. Therefore, we desired to find a sweet spot where sensitivity for detection would be maximized while preserving the remaining cfDNA for use in other applications.

[0147] Relationship between DNA input and LOD Another cell line spiking experiment was designed to examine the relationship between DNA input and LOD. Using NCI-H441 cell line DNA, cell line-healthy cfDNA mixtures were generated with four different tumor fractions (0.5%, 0.1%, 0.05%, and 0%), and three different DNA inputs (5 ng, 10 ng, and 30 ng) of each AF were tested in triplicate (36 libraries total). After sequencing, unique molecule recovery was found to correlate with the DNA input into the library preparation (Figure 9B). Furthermore, as seen in previous spiking experiments, a relationship was observed between tumor fraction and DMR above background (Figure 9C). Importantly, the DMR signal in the 0.05% AF sample was found to be clearly elevated above the 0% condition with a 30 ng input, but not with the 5 ng or 10 ng input conditions, demonstrating how LOD depends on DNA input.

[0148] However, an important caveat of this experiment was that our DNA input reflected the total DNA put into the library preparation, including any genomic DNA that may have contaminated the cfDNA. As previously mentioned, cfDNA consists of short, nucleosome-bound DNA molecules that exhibit a polymodal size distribution. If intact blood cells remaining in plasma after blood centrifugation remain, their genomic DNA can become the "cfDNA" eluate after DNA isolation. Although these large gDNA molecules count toward the total DNA quantification when measured by fluorescence, their size makes them unlikely to successfully complete library preparation. To adjust for this, we typically perform fragment analysis on all cfDNA samples to confirm their fragment size distribution and check for gDNA contamination. Using the fragment analyzer results, we can scale the DNA input by the percentage of fragments in the typical cfDNA size range (50-450 bp) to better control the input of molecules that will result in a sequenceable library. In the spike experiments described above, the input was not scaled in this manner, so the 5, 10, and 30 ng inputs included any gDNA present. Fragment analysis of the healthy cfDNA sample used as the denominator for the spike indicated that the sample indeed contained only approximately 50% of the molecules in the 50-450 bp size range. Therefore, the adjusted inputs for the spike experiments were closer to 2.5 ng, 5 ng, and 15 ng of DNA.

[0149] To further expand the LOD and test whether higher adjusted DNA inputs would improve sensitivity, we repeated the spike-in experiment with different inputs and AFs, again with each combination performed in triplicate: 10 ng, 15 ng, and 30 ng, and at 0.1%, 0.05%, 0.025%, and 0% (Figure 9D). We found that a 0.05% increase in DMR was detected compared to 0% for all three DNA inputs. While 0.025% did not reach statistical significance, two of the three replicates at 0.025% increased above 0% for both 15 ng and 30 ng inputs. We selected 15 ng (50–450 bp size range) as the appropriate input for mCAPP-Seq to maximize sensitivity while also preserving DNA for further use, including the Lung-CLiP assay.

[0150] Applying mCAPP-Seq to NSCLC patients and risk-matched controls Ultimately, the goal of mCAPP-Seq was to test its ability to detect the presence of lung cancer DNA in actual patient plasma samples. We planned to test its sensitivity and specificity for detection in early-stage (stages I-III) lung cancer cohorts and risk-matched controls. However, we first wanted to reconfirm with an optimized method that the expected signal was observed in high tumor burden settings, where significant ctDNA content can be expected. We applied the method to a pilot cohort of healthy control and stage IV NSCLC patient samples.

[0151] Targeted methylation signals in advanced disease We generated mCAPP-Seq libraries using the optimized EM-seq protocol from cfDNA from 12 healthy control and 12 stage IV NSCLC patient samples. These libraries were then captured with a targeted panel and sequenced at 100 million read pairs per sample. After mapping and deduplication, methylation calls were extracted for all CpGs, and DMRs above background were calculated as described above. Confirming previous observations using bisulfite sequencing (Figure 7D), the AF of DMR regions was higher in stage IV patients compared to controls (Figure 10A). Consequently, DMRs above background were also significantly higher in patients (Figure 10B). This test case again confirmed that this method works beyond the intended setting of adulterated samples in patient plasma. However, the detection of early-stage, low-ctDNA cases needs to be evaluated.

[0152] Application of mCAPP-Seq to early stage disease Construction of training and validation cohorts of early-stage NSCLC patients and high-risk controls Patient cohorts were constructed from NSCLC patients at Stanford, MGH, MSK, and Mayo Clinic who were treated with intent-to-cure and had plasma collected before treatment. Stanford and Mayo patients (n = 117) were assigned to the training cohort for the methylation assay. Control cohorts were collected from Stanford and MGH patients who underwent preventive low-dose CT screening for lung cancer based on a significant smoking history and were risk-matched to the patient cohort. Risk-matched healthy controls (n = 97) from both Stanford and MGH, as well as patients with benign granulomas (n = 10) collected at Mayo Clinic, were assigned to the training cohort. Having an independently collected validation cohort was considered important to test the performance of the assay in a robust and reliable manner. NSCLC patient samples collected at MGH and MSK (n = 84) were assigned to the validation cohort. Control samples (n=78) collected at MGH after 2018 were assigned to the validation cohort.

[0153] Availability of cfDNA in patient samples cfDNA yield was primarily determined by the amount of plasma available, which varied by institution and collection protocol. Across all extracted control and patient samples, a median of 6.7 ml of plasma was available for extraction (range 1.7–14.0 ml, Figure 11A). cfDNA isolation from all samples yielded a median of 59.9 ng for controls and 87.5 ng for NSCLC patients (Figure 11B). All samples were analyzed for genomic DNA contamination with an Agilent fragment analyzer, demonstrating a high cfDNA fraction (50–450 bp) for most samples (Figure 11C). After adjusting for gDNA contamination in plasma by considering only cfDNA in the 50–450 bp size range, total adjusted cfDNA yield remained, on average, higher in patients than in controls (Figure 11D). Even after normalizing for plasma volume, plasma cfDNA concentrations in ng / ml, considering only cfDNA between 50 and 450 bp, also remained higher in patients than in controls (Figure 11E).

[0154] Experimental design In addition to the goal of evaluating the performance of mCAPP-Seq for noninvasively detecting early-stage NSCLC, we also desired to compare mCAPP-Seq to a previously published method, Lung-CLiP (J.J. Chabon, Nature. 2020 Apr;580(7802):245-251, the disclosure of which is incorporated herein by reference). It was reasoned that while detection between the two methods may be correlated, there may be instances where detection by one method is not detected by the other, and that an integrated approach may be superior to either method alone. To accomplish this, both assays would need to be performed on the same sample, and we would require the 15 ng of adjusted cfDNA required for mCAPP-Seq and an additional minimum of 20 ng of cfDNA for CAPP-Seq input. Based on the available plasma volume and resulting cfDNA yield, depending on the cohort, only 55–69% of samples would have sufficient cfDNA for both assays (Figure 11F). Nevertheless, this will provide a pilot cohort for testing the integration of the two methods. The experimental design will be to test the methylation-only assay on the entire cohort as outlined above. Subset analysis will then be performed to test integration in patients for whom sufficient cfDNA was available to further perform CAPP-Seq. Ultimately, 73 patients and 56 controls will be run with both methods for the training cohort, and 58 patients and 45 controls will be run with both methods for the validation cohort.

[0155] initial results Samples from the training cohort were first analyzed. Libraries were prepared from 97 risk-matched controls, 10 granuloma controls, and 117 early-stage NSCLC cfDNA samples using a fixed input of 15 ng cfDNA ranging from 50 to 450 bp. Samples were ligated to UMI-containing methylated double-stranded adapters, and unmethylated cytosines were converted to uracils via EM-Seq, followed by 11 cycles of PCR amplification. Libraries were captured with a targeted sequencing panel specifically designed for NSCLC detection. After capture, libraries were sequenced with 80 million paired-end reads (i.e., 40 million read pairs).

[0156] The quality of the sequencing data was assessed. Despite aiming for equal library representation in each lane, libraries received a wide range of total read counts (median 37.5M reads vs. 14.8M-63M), but read counts did not differ between patients and controls (Figure 12A). After removing PCR duplicates using unique molecular identifiers (UMIs), median selector-wide depth was also similar between patients and controls (median 1256x across all training samples, range 263x-2467x, Figure 12B). As expected, there was a significant effect of unduplicated read count on deduplicated depth, suggesting that some low-depth samples could be rescued by adding additional sequencing reads if necessary (Figure 12C). Conversion efficiency, as measured by unmethylated lambda control DNA, was high across all samples.

[0157] DMR above background analysis was applied to this cohort of patient and control samples. The methylated allele fraction in high DMR regions was higher in patients than in controls (Figure 12C). DMR above background also tended to be higher in patients than in controls (Figure 12D), but the sensitivity of this metric at 95% specificity was found to be insufficient for early detection (Figure 12E). However, it was reasoned that DMR above background is a simple measure of signal and that it may be possible to improve detection performance through statistical learning.

[0158] Development of methylation-based classifiers for cancer detection As described above, the methylation "allele fraction"—or the proportion of fragments with a highly methylated state within a region—was calculated for each DMR of interest. Instead of summing the region above background as in previous analyses, these AFs were used as features in a machine learning model. A leave-one-out (LOO) framework was developed in which a LASSO logistic regression classifier was trained on all samples in the training cohort except one that had the DMR AF as a feature. The trained model was used to score holdout samples, and the process was repeated for each patient sample.

[0159] Because the initial definition of "highly methylated molecules" was based on heuristics, it was optimized to maximize LOO performance. A range of minimum CpGs per fragment was tested, and performance was measured using the minimum percent methylation across all pairwise combinations and LOO sensitivity at 95% specificity. For each iteration of LOO, a specificity of 95% was set for the training data, and holdout samples were called detected if their scores exceeded that threshold. True specificity was calculated as the percentage of controls correctly classified as non-cancer when held out. From this analysis, we found that a minimum methylation of 60%, in contrast to the previously used minimum methylation threshold of 80%, exhibited increased detection sensitivity across all minimum CpGs per fragment thresholds (Figure 13A). For holdout samples, true specificity was well calibrated at approximately 96% (Figure 13B).

[0160] Next, we explored how features beyond the DMR AF itself could be extracted for inclusion in the classifier. First, we explored ways to summarize the DMR AF distribution into higher-order summary statistics. It was hypothesized that these descriptive summary statistics could create a more robust and generalizable model. To do this, we visualized the distribution of DMR AF for each sample in the early-stage cohort as well as in the sequenced pre-stage IV patients (Figure 13C). Looking at the distribution of AF, we observed a highly visible upward shift in the distribution in stage IV patients compared to controls. In the early-stage cohort, the distribution was mixed, with some clearly elevated and others resembling control samples. We sought to summarize the distribution of AF and use these features as inputs to a machine learning model. To do this, we calculated percentiles of the AF distribution for each patient, ranging from the 99th percentile to the 50th percentile (median), and used them as features in a logistic regression model.

[0161] To explore additional signal sources that could further enhance detection sensitivity, we investigated fragment-level features of highly methylated molecules in DMR regions. We found that patients tended to have fragments with more methylated CpGs in high DMRs compared with controls (Figure 13D). To characterize this observation, for each sample, the maximum number of methylated CpGs in any fragment was identified per region. The distribution of these maximum values ​​across all DMR regions was summarized by taking the median, maximum, and 90th percentile of the distribution. The 50th percentile fragment methylation index (FMI) was used. 50 The median number of methylated CpGs, referred to as the median length (FMI), was found to be associated with stage, suggesting a tumor-associated signal consistent with ctDNA biology (Figure 14A). To further validate this metric and ensure that the increased FMI in patients was not due to differences in overall fragment size between patients and controls, the median fragment was calculated across all DMR fragments, and no differences were found between patients and controls (Figure 14B). There was also no difference between the proportion of all DMR fragments longer than 300 bp (Figure 14C). When considering only fragments with enough CpGs to meet the DMR AF filtering criteria (minimum 12 CpGs), no differences in median fragment length or proportion of fragments >300 bp were observed between patients and controls (Figures 14D and 14E). Finally, we considered fragments in DMR regions with the most CpGs, regardless of methylation status. The median length of these fragments did not differ between patients and controls (Figure 14F), and the median number of CpGs in these fragments was slightly higher in patients compared to controls, but was highly overlapping (Figure 14G). Taken together, these results suggest that there are no major differences in fragment size distribution between patients and controls who were driving FMI signals.

[0162] Using the different classes of identified features (AF features, AF distribution percentile features, and fragment methylation index features), we trained a series of statistical models in a leave-one-out (LOO) framework to classify samples as cancer versus control. Because this cohort was highly enriched for small stage I tumors, making their detection particularly challenging, it is important to consider detection sensitivity by stage. A correlation between stage and sensitivity is expected to be an important confirmation of the model's biological validity. A model using only DMR AFs detected stage IA tumors at 28%, with sensitivity increasing with increasing stage (Figure 15A). As further validation of the model, we trained a single model using all stage I-III samples and controls and used it to score holdout stage IV samples. Stage IV showed the highest sensitivity, again confirming biological validity (Figure 15A). Closer examination of the features selected by each LOO iteration of LASSO revealed that most models had four nonzero coefficients (Figure 15B), and a small number of regions were selected by most models, suggesting the importance of those features (Figure 15C). AF in these regions also correlated with patient stage (data not shown).

[0163] Next, we trained a model using the AF percentile features described above (Figure 15D). The stage-dependent sensitivity of that model also increased with stage, but the sensitivity was slightly lower than that of the AF-based model (Figure 15D). However, it remained to be determined whether the AF-based model was prone to overfitting compared to the summary statistics model. We then combined the DMR AF features with the fragment methylation index features described above, again achieving stage-related sensitivity (Figure 15E). In this model, both AF features and fragment-based features were selected across leave-one-out iterations (Figure 15F).

[0164] An additional model was trained using AF percentile and fragment methylation index features. This model also showed good stage-based sensitivity and specificity, calibrated to a target of 95% (Figure 16A), with a positive predictive value (PPV) of 89%. Interestingly, only fragment-level features were selected for these models (Figure 16B). This model showed a higher detection rate for lung squamous cell carcinoma than for adenocarcinoma (Figure 16C). To compare the results with the leading commercially available method, we calculated performance by stage for adenocarcinoma only and found 22% detection for stage I tumors, a three-fold improvement over the leading method (Figure 16D). (For details about the leading model, see X. Chen et al., Clin Cancer Res. 2021 Aug 1;27(15):4221-4229, the disclosure of which is incorporated herein by reference.) Finally, when considering how this method could be applied clinically, a lower specificity threshold was considered acceptable while still maintaining a high PPV when screening high-risk populations where the disease prevalence is higher than in the general population. Sensitivity was reported at 80% specificity, with increased detection of stage I and stage II tumors (Figure 16E). No increase in stage III detection was observed, which may be due to the low ctDNA levels in these patients.

[0165] conclusion Lung cancer screening has the potential to significantly improve patient outcomes, and blood-based assays are an attractive complement to imaging. Building on previous work in which a classifier was developed to determine lung cancer likelihood (Lung-CLiP) in plasma from genetic signatures of cell-free DNA, this study developed a method to test whether utilizing epigenetic cell-free DNA signatures could improve early lung cancer detection. To do this, we developed methyl-CAPP-Seq (mCAPP-Seq), a variant of the CAPP-Seq method that incorporates targeted deep methylation sequencing and bioinformatic tools to extract tumor-associated methylation status from cfDNA. NEB enzyme methyl-seq (EM-seq) was found to outperform the gold standard bisulfite conversion in both maintaining DNA integrity and properly converting unmethylated cytosines to thymines. We developed a framework to identify highly methylated reads and found that the signal correlated with the proportion of tumor DNA in plasma in both cell line admixture experiments and primary NSCLC patient samples. Finally, we developed a logistic regression classifier to distinguish healthy plasma samples from NSCLC samples via methylation signals and found that the model performance was better than conventional methods and was biologically plausible.

[0166] While early detection of lung cancer from blood remains a technical challenge, this study represents a technological advance in the field. First, this study is unique in its targeted sequencing approach, where high specific coverage across a small genomic space allows for efficient utilization of sequencing reads. From this data, previously undescribed cell-free DNA methylation signatures were identified that correlate with lung cancer tumor burden. Most notably, the model developed in this study demonstrated significantly higher sensitivity than the leading published method, which had a sensitivity of approximately 7% for stage I adenocarcinoma. Even a 15% increase in sensitivity for these small tumors would mean significantly improved outcomes for these patients.

[0167] method Differential methylation with 450K array data Illumina Infinium HumanMethylation450 (450k array) data were downloaded in processed form (beta values) from TCGA via UCSC Xena Browser or from publicly available datasets via GEO (accession numbers GSE32148, GSE41169, GSE54670, GSE73745, GSE53045, GSE107205, GSE35069 for blood samples; GSE52401 and GSE66836 for normal lung). After converting the beta values ​​to M values ​​(P. Du et al., BMC Bioinformatics. 2010 Nov 30;11:587, the disclosure of which is incorporated herein by reference), limma (MERitchie et al., Nucleic Acids Res. 2015 Apr 20;43(7):e47, the disclosure of which is incorporated herein by reference) was used to identify differentially methylated CpG sites between LUAD or LUSC samples and blood and normal lung samples (as a group). Limma P values ​​were adjusted for multiple hypothesis testing using the Benjamini-Hochberg method. CpG island and gene context annotation of selected CpGs was performed using the R package "annotatr" (R.G. Cavalcante and M.A. Sartor, Bioinformatics. 2017 Aug 1;33(15):2381-2383, the disclosure of which is incorporated herein by reference).

[0168] mCAPP-Seq library preparation Libraries were prepared from cfDNA or sheared genomic DNA using the KAPA HyperPrep kit. Unmethylated, sheared lambda phage DNA (approximately 1 pg) and methylated pUC19 DNA (approximately 0.3 pg) were added to each sample, and the sample volume was then brought to 50 μl in nuclease-free water. End repair and A-tailing (ER / AT) were performed according to the KAPA HyperPrep protocol. After ER / AT, a 100-fold molar excess of methylated partial Y adapters was added to each sample. The adapters contained the insert UMI but were methylated at all cytosine positions and resistant to conversion. Samples were ligated overnight at 4°C. After ligation, 1x SPRI bead cleanup was performed, and samples were eluted with the volume of water required for either the bisulfite or EM-Seq protocol.

[0169] Bisulfite conversion After ligation with methylated Y adapters, bisulfite conversion was performed using the Qiagen Epitect kit using the “Sodium bisulfite conversion of unmethylated cytosines in small amounts of fragmented DNA” protocol according to the manufacturer's instructions.

[0170] EM-Seq conversion After ligation to methylated adapters and bead cleanup, EM-Seq conversion was performed using the NEB EM-seq Conversion Module (NEB #E7125) according to the manufacturer's instructions, using formamide as the denaturant. After conversion, PCR was performed for 7 cycles as described in "Post-Conversion Grafting PCR." The library was further amplified by universal PCR for 4–7 cycles (4 cycles in most cases, depending on DNA input).

[0171] TAPS Conversion The library was ligated overnight at 4°C to a standard (non-methylated) partial Y adapter using KAPA HyperPrep and purified by 1× SPRI bead cleanup. The library was converted with the following modifications using the TAPS protocol as previously described (Y. Liu et al., Nat Biotechnol. 2019 Apr;37(4):424-429, the disclosure of which is incorporated herein by reference). After conversion, 7-cycle PCR was performed using a custom dual-index primer with KAPA HiFi Uracil+ master mix as described in "Converted graft PCR". After graft PCR, the library was further amplified by 6-cycle universal PCR.

[0172] Converted graft PCR After conversion by bisulfite, EM-Seq or TAPS, the library was amplified and indexed by graft PCR. Briefly, 25 ul of KAPA HiFi Uracil+ enzyme master mix and 2 ul of 12 uM forward + reverse index primer were added to each sample and amplified in a thermal cycler using the following protocol: 95°C for 2 minutes Repeat 7 cycles: 98°C for 30 seconds 60°C for 30 seconds 72°C for 4 minutes 72°C for 10 minutes Hold at 4°C

[0173] The primers contained a sample barcode with dual index. The PCR products were purified with 1× SPRI beads and eluted in 24 ul of nuclease-free water.

[0174] Universal PCR After index graft PCR, universal PCR was performed to further amplify the library. The number of cycles varied depending on the DNA input and conversion method. 25 μl of KAPA HiFi HotStart ReadyMix and 1 μl of 100 μM forward + reverse universal primers were added to 24 μl of the library. Samples were amplified in a thermal cycler using the following protocol: 98 °C for 45 seconds Repeat 4 - 7 cycles: 98 °C for 15 seconds 60 °C for 30 seconds 72 °C for 30 seconds 72 °C for 1 minute Hold at 4 °C

[0175] Sequencing Samples were sequenced on an Illumina HiSeq or NovaSeq 6000. All samples in the initial detection cohort were sequenced on a NovaSeq 6000 targeting 40 M read pairs per sample. The actual read counts ranged from 14.8 - 63 M read pairs, with a median of 37.5 M read pairs. The read counts did not differ significantly between patients and controls.

[0176] Data processing, alignment, and duplicate removal Sequencing data was demultiplexed using in-house scripts, and adapter read-through was trimmed with fastp (S. Chen et al., Bioinformatics. 2018 Sep 1;34(17):i884 - i890, the disclosure of which is incorporated herein by reference). Samples were then mapped to the human genome using Bismark (F. Krueger and S. R. Andrews, bioinformatics. 2011 Jun 1;27(11):1571 - 2, the disclosure of which is incorporated herein by reference). PCR duplicates were removed using in-house scripts. The methylation status of all CpGs was extracted with Bismark. Custom Python scripts were used to summarize the CpG status at the fragment level and region level.

[0177] Statistical Modeling and Analysis A normalized logistic regression model for predicting cancer status from cfDNA methylation was fitted in R using the cv.glmnet function with family = "logistic", alpha = 0 for ridge regression, alpha = 1 for LASSO regression, and other default parameters, using glmnet (J. Friedman, T. Hastie, and R. Tibshirani J Stat Softw. 2010;33(1):1 - 22, the disclosure of which is incorporated herein by reference). Model performance was summarized using pROC (X. Robin et al., Bioinformatics. 2011 Mar 17;12:77, the disclosure of which is incorporated herein by reference). All statistical analyses were performed in R 4.0.1 or 3.6.1.

Table 1 - 1

Table 1 - 2

Table 1 - 3

Table 1 - 4

Table 1 - 5

Table 1 - 6

Table 1 - 7

Table 1 - 8

Table 1 - 9

Table 1 - 10

Table 1-67

Table 1-90

Table 1-100

Claims

1. 1. A sequencing method for identifying differentially methylated regions associated with a state in cell-free nucleic acid, comprising: obtaining a cell-free nucleic acid sample comprising cell-free nucleic acid molecules; extracting a subset of the cell-free nucleic acid molecules from the cell-free nucleic acid sample using a panel of nucleic acid probes designed to hybridize to regions known to be differentially methylated in a certain condition; converting nucleobases of said subset of said cell-free nucleic acid molecules, said conversion of a nucleobase indicating the methylation status of that nucleobase; and sequencing said subset of said cell-free nucleic acid molecules via high-throughput sequencing.

2. The method of claim 1 , wherein the condition is cancer.

3. The method of claim 2, wherein the cancer is non-small cell lung cancer.

4. 4. The method of claim 3, wherein the regions known to be differentially methylated in a condition comprise at least 5% of the regions in Table 2.

5. 5. The method of claim 4, wherein the regions known to be differentially methylated in a condition comprise at least 50% of the regions in Table 2.

6. The method of any one of claims 1 to 5, wherein the nucleic acid probe panel excludes regions known to be associated with false discoveries.

7. The method of any one of claims 1 to 6, wherein the nucleic acid probe panel excludes regions known to be differentially methylated in blood cells.

8. 8. The method of any one of claims 1 to 7, further comprising extracting a subset of the cell-free nucleic acid molecules from the cell-free nucleic acid sample using a panel of nucleic acid probes designed to hybridize to regions known to correlate with factors associated with the condition.

9. 9. The method of claim 8, wherein the condition is cancer and the region known to correlate with a factor associated with the condition comprises a region in Table 1.

10. 10. The method of any one of claims 1 to 9, further comprising extracting a subset of the cell-free nucleic acid molecules from the cell-free nucleic acid sample using a panel of nucleic acid probes designed to hybridize to regions known to be invariably hypermethylated or invariably hypomethylated.

11. 11. The method of claim 10, wherein the regions known to be invariably hypermethylated or invariably hypomethylated comprise the regions of Table 1.

12. 12. The method of any one of claims 1-11, wherein converting the nucleobases of the subset of the cell-free nucleic acid molecules comprises at least one of bisulfite treatment, TET2 oxidation and APOBEC3A conversion, or TET2 oxidation and pyridine borane treatment.

13. 13. The method of claim 12, wherein converting the nucleobases of the subset of the cell-free nucleic acid molecules comprises TET2 oxidation and APOBEC3A conversion.

14. The method of any one of claims 1 to 13, wherein the cell-free nucleic acid sample is derived from a collection of blood, plasma, saliva, urine, stool, mucus, lymph or another bodily fluid.

15. The method of any one of claims 1 to 14, wherein the cell-free nucleic acid sample comprises at least 100,000 nucleic acid molecules.

16. The method of any one of claims 1 to 15, wherein the cell-free nucleic acid of the cell-free nucleic acid sample is cell-free DNA.

17. 17. The method of claim 16, wherein the cell-free nucleic acid of the cell-free nucleic acid sample comprises at least 1 ng of cell-free DNA.

18. 17. The method of claim 16, wherein the cell-free nucleic acid of the cell-free nucleic acid sample comprises at least 15 ng of cell-free DNA.

19. 19. The method of any one of claims 1 to 18, further comprising attaching an adaptor to the comprising cell-free nucleic acid molecules, wherein the adaptor is resistant to the nucleobase conversion performed in the step of converting the nucleobases of the subset of the cell-free nucleic acid molecules.

20. The method of any one of claims 1 to 19, wherein the nucleic acid probe panel comprises at least 50 unique probes.

21. 1. A sequencing method for enhancing detection of differentially methylated regions for assessing the status of an individual, comprising: preparing a cell-free nucleic acid sample for targeted methyl-sequencing, the prepared cell-free nucleic acid sample being collected from an individual and comprising at least 100,000 cell-free nucleic acid molecules derived from a plurality of regions known to be differentially methylated in a condition; sequencing the cell-free nucleic acid sample via a high-throughput sequencer to obtain sequencing results of the cell-free nucleic acid molecules from a plurality of regions known to be differentially methylated in a certain state; calculating a methylation metric using a computing device and the sequencing results, wherein the methylation metric is calculated for a region of the plurality of regions known to be differentially methylated in a state, the methylation metric indicating the amount of methylation of cell-free nucleic acid molecules that align to the region; and using the computing device to input the calculated methylation metrics as features into a computational model to obtain an assessment of the cell-free nucleic acid sample, wherein the assessment indicates that the individual has the condition.

22. Calculating a methylation metric for the region aligning each cell-free nuclear molecular sequencing result across a region, the region being one of the plurality of regions that are differentially methylated in a state; and for a set of cell-free nucleic acid molecules that align across the region, determining the amount of methylation of each cell-free nucleic acid molecule of the set; 22. The method of claim 21, wherein the methylation metric is based on at least one cell-free molecule of the set.

23. Calculating a methylation metric for the region determining the number of cell-free nucleic acid molecules in the set that are methylated above a threshold; calculating the molecular methylation fraction (MMF) of said region, wherein: [Equation 5] (In the formula, [Equation 6] is the number of cell-free nucleic acid molecules in the set that are determined to be methylated above a threshold, [Equation 7] is the total number of cell-free nucleic acid molecules in the set) Koto and 23. The method of claim 22, further comprising:

24. 24. The method of claim 23, wherein the threshold is 60% of CpGs methylated.

25. Calculating a methylation metric for the region 23. The method of claim 22, further comprising identifying a cell-free nucleic acid molecule that is most methylated within the set of cell-free nucleic acid molecules that align across the region, wherein the methylation metric is calculated as the amount of methylation of the cell-free nucleic acid molecule that is most methylated.

26. 26. The method of any one of claims 22 to 25, wherein each cell-free nucleic acid molecule of the set of cell-free nucleic acid molecules that align across the region has a number of CpGs that is greater than a threshold value.

27. using the computing device and the sequencing results to calculate a methylation metric for each region of at least 50 percent of the plurality of regions known to be differentially methylated in a condition; using the computing device to input each calculated methylation metric as a feature into the computational model to obtain an assessment of the cell-free nucleic acid sample, the assessment indicating that the individual has the condition; and The method of any one of claims 22 to 26, further comprising:

28. using the computing device and the sequencing results to calculate a methylation metric for each region of the plurality of regions known to be differentially methylated in a condition; using the computing device to input each calculated methylation metric as a feature into the computational model to obtain an assessment of the cell-free nucleic acid sample, the assessment indicating that the individual has the condition; and The method of any one of claims 22 to 26, further comprising:

29. calculating a plurality of methylation metrics using the computing device and the sequencing results, each methylation metric being calculated for a region of the plurality of regions known to be differentially methylated in a condition; calculating a sample summary statistic combining the plurality of methylation metrics using the computing device and the plurality of methylation metrics; using the computing device to input the sample summary statistics as features into a computational model to obtain the assessment of the cell-free nucleic acid sample, wherein the assessment indicates that the individual has the condition. The method according to any one of claims 21 to 28.

30. 30. The method of claim 29, wherein each methylation metric of the plurality of methylation metrics for calculating a sample summary statistic is calculated as follows: determining the number of cell-free nucleic acid molecules in the set that are methylated above a threshold; and calculating a methylated molecular fraction (MMF) for said region, wherein: [Equation 8] During the ceremony, [Equation 9] is the number of cell-free nucleic acid molecules in the set that are determined to be methylated above a threshold, [Equation 10] is the number of total cell-free nucleic acid molecules in the set.

31. 31. The method of claim 30, wherein the sample summary statistics are percentiles of several regions with non-zero MMF.

32. 31. The method of claim 30, wherein the sample summary statistic is a percentile of some region where the MMF is 0.

33. 31. The method of claim 30, wherein the sample summary statistic is a percentile of several regions where the MMF is greater than a threshold.

34. 31. The method of claim 30, wherein each methylation metric of the plurality of methylation metrics for calculating a sample summary statistic is calculated as follows: identifying a most methylated cell-free nucleic acid molecule within the set of cell-free nucleic acid molecules aligned across the region, wherein the methylation metric is calculated as the amount of methylation of the most methylated cell-free nucleic acid molecule.

35. 35. The method of claim 34, wherein the sample summary statistic is a percentile of the amount of methylation of the cell-free nucleic acid molecules that are most methylated.

36. 35. The method of Claim 34, wherein the sample summary statistic is the median amount of methylation of the cell-free nucleic acid molecules that are most methylated.

37. 35. The method of claim 34, wherein the sample summary statistic is the skewness of the amount of methylation of the cell-free nucleic acid molecule that is most methylated.

38. 38. The method of any one of claims 21 to 37, wherein the plurality of regions known to be differentially methylated in a condition comprises at least 10 genomic regions associated with a condition.

39. 39. The method of any one of claims 21 to 38, wherein the plurality of regions known to be differentially methylated in a condition comprises at least 50 genomic regions associated with a condition.

40. 40. The method of any one of claims 21 to 39, wherein the condition is cancer.

Citation Information

Cited By

  • Diester Compounds

    JPWO2022202703A1