Cell-free DNA biomarker for diagnosis and prognosis of diseases with degenerative processes

EP4724607A2Pending Publication Date: 2026-04-15RGT UNIV OF CALIFORNIA +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
RGT UNIV OF CALIFORNIA
Filing Date
2024-06-07
Publication Date
2026-04-15

AI Technical Summary

Technical Problem

Current methods for diagnosing and monitoring neurodegenerative diseases like ALS lack effective biomarkers, and existing technologies such as whole-genome bisulfite sequencing are costly and inefficient for clinical applications, limiting their utility in characterizing tissue-specific cell death and disease progression.

Method used

A method using cell-free DNA (cfDNA) as a non-invasive biomarker, involving the use of capture probes to target specific methylation sites, followed by methylation sequencing and machine learning algorithms to detect cell-type specific degeneration, which can be extracted from a standard blood draw and analyzed for ALS and other neurodegenerative diseases.

Benefits of technology

This approach provides a cost-effective and clinically relevant method for diagnosing and monitoring ALS and other neurodegenerative diseases, enabling the identification of tissue-specific cell death and disease progression, with a projected cost 10-fold lower than whole-genome bisulfite sequencing and high accuracy in predicting disease status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024033056_12122024_PF_FP_ABST
    Figure US2024033056_12122024_PF_FP_ABST
Patent Text Reader

Abstract

The development of a methylation capture panel that targets methylation sites that vary between tissues at high sequencing depth means that the projected cost for a sample run on the proposed capture technology is 10-fold smaller than with whole genome bisulfite sequencing (WGBS). This unexpected combination of methylation sequencing, cell type identification, machine learning, and capture technology enables a previously unattainable ability to predict disease status. These methods can be used to diagnose, prognose, and stage disease conditions including amyotrophic lateral sclerosis (ALS) and other degenerative conditions. In addition, it offers great advantages for other cfDNA-based detection applications, such as cancer and pregnancy-related conditions.
Need to check novelty before this filing date? Find Prior Art

Description

CELL-FREE DNA BIOMARKER FOR DIAGNOSIS AND PROGNOSIS OF DISEASES WITH DEGENERATIVE PROCESSES

[0001] This application claims benefit of United States provisional patent application number 63 / 506,899, filed June 8, 2023, the entire contents of which are incorporated by reference into this application. STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH

[0002] This invention was made with government support under NS122538 awarded by the National Institutes of Health. The government has certain rights in the invention. REFERENCE TO A SEQUENCE LISTING

[0003] The content of the XML file of the sequence listing named “UCLA293WOU1_seq”, which is 3 kb in size, created on June 6, 2024, and electronically submitted herewith the application, is incorporated herein by reference in its entirety. BACKGROUND

[0004] Cell-free DNA (cfDNA) is an emerging biomarker or biomarker candidate for multiple diseases, as it originates from dying tissues and can be non-invasively measured through a blood draw. CfDNA has been used in the detection of cancer (1–3), to identify fetal genetic abnormalities (4,5), to screen for infectious diseases (6,7), and to predict pregnancy complications (8). One underexplored domain for cfDNA, however, is neurodegenerative disease. Biomarkers for neurodegenerative diseases are critically needed for improving patient care and evaluating the efficiency of clinical trials (9). While the application of cfDNA to neurodegeneration is nascent, previous work (10–13) has shown alterations in the cell- free DNA and RNA of patients with neurodegeneration relative to healthy controls.

[0005] Amyotrophic Lateral Sclerosis (ALS) is a neurodegenerative disease with complex pathogenesis. There is no clinical biomarker for ALS. Our previous work focused on analyzing the cell-free DNA of ALS patients using whole-genome bisulfite sequencing (WGBS). However, about 80% of CpG sites are not variable between tissues, limiting their utility in characterizing tissue-specific cell death and its relevance to disease. Furthermore, at the required depths for accurate methylation estimation, WGBS can be prohibitively costly, especially for clinical applications.

[0006] There remains a need for a biomarker, as well as practical methods, to expedite diagnosis, monitor disease progression, or prioritize patients for clinical trials for ALS and other degenerative conditions.SUMMARY

[0007] The methods described herein provide cell-free DNA (cfDNA) as a non-invasive biomarker for ALS, ALS progression, detection and monitoring of other neurodegenerative diseases, and other conditions with degenerative processes. CfDNA is an ideal candidate because it is enriched in ALS patients, is informative about tissue-specific cell death, and can be extracted from a standard blood draw.

[0008] Described herein is a method of detecting cell type specific degeneration in a blood or other biological sample obtained from a subject. In some embodiments, the method comprises: (a) contacting circulating extracellular DNA extracted from the sample with a set of capture probes, wherein the capture probes are less than 50,000 in number, and wherein the capture probes specifically bind within 100 base pairs (bp) of methylated and unmethylated CpG sites on cell-free DNA (cfDNA), and wherein the methylation status of CpG sites is specific to a cell type of interest. The method further comprises: (b) performing methylation sequencing of cfDNA captured in step (a); (c) creating a methylation status dataset from the sequencing of step (b), wherein the methylation status dataset contains the methylation status of the cell type specific CpG sites (CTSCS); and (d) detecting cell type specific degeneration when the methylation status of the CTSCS is distinct from a reference methylation status dataset, wherein the reference methylation status dataset contains the methylation status for the CTSCS of all cell types. In some embodiments, the method further comprises inputting the methylation status dataset into a statistical model. Alternatively, the methylation status dataset has previously been inputted into the statistical model. The model maps the methylation status dataset to a reference methylation status dataset, wherein the reference methylation status dataset contains the methylation status for the CTSCS of all cell types.

[0009] In some embodiments, the CTSCS are identified using a supervised machine learning model and expectation maximization (EM) algorithm that estimates the proportion of the reference cell types for the subject and that estimates a contribution of unknown cell types not included in the reference methylation status dataset. In some embodiments, the set of capture probes is between 20 and 20,000 in number. In some embodiments, the set of capture probes is between 500 and 15,000 in number. In some embodiments, the set of capture probes is between 1,000 and 10,000 in number. In some embodiments, about 5,000 capture probes are included in the set. In some embodiments, the statistical model comprises a penalized regression model.

[0010] In some embodiments, the biological sample is a blood sample. In some embodiments, the biological sample is extracellular DNA that has been extracted from ablood sample, or other biological specimen. In some embodiments, the cfDNA is extracted from the biological sample as a separate step prior to contact with the capture probes. In some embodiments, the extracellular DNA extracted from the sample is treated with bisulfite, which introduces methylation-dependent sequence changes through selective chemical conversion of non-methylated cytosine to uracil. After treatment, all non-methylated cytosine bases are converted to uracil but all methylated cytosine bases remain cytosine. These methylation dependent C-to-T changes can subsequently be studied using conventional DNA analysis technologies.

[0011] In some embodiments, the cell type specific degeneration is amyotrophic lateral sclerosis (ALS). In some embodiments, the cell type specific degeneration is traumatic brain injury (TBI), cancer, Alzheimer’s disease, Parkinson’s disease, or a pregnancy-related phenotype. Examples of a pregnancy-related phenotype include, but are not limited to, gestational diabetes, pre-term birth, and preeclampsia.

[0012] In some embodiments, the methylation sequencing comprises high throughput sequencing.

[0013] In some embodiments, the CTSCS are specific to skeletal muscle, fibroblasts, neurons, and / or hematopoietic cells.

[0014] Also described is a computer implemented method of training a machine learning system to generate an estimator for identifying a cell type specific degeneration in a blood sample obtained from a subject. In some embodiments, the method comprises: (a) storing a set of data comprising a plurality of prospective patient records from more than 100 patients, each prospective patient record including a plurality of parameters and corresponding values for each patient included in the patient records, and a diagnostic indicator indicating whether or not the patient included in the patient records has been diagnosed with a specific condition associated with a cell type specific degeneration; (b) selecting a subset of the plurality of parameters for inputs into the machine learning system, wherein the subset comprises CTSCS, the subset further including at least one clinical parameter selected from age, gender, and smoking status; (c) randomly partitioning the set of data into training data and validation data; and (d) generating the estimator wherein the machine learning system is trained based on the training data and the subset of inputs.

[0015] In some embodiments, the machine learning system trained on the subset of inputs comprises a supervised machine learning model. In some embodiments, the estimator is trained with an expectation maximization (EM) algorithm that estimates the proportion of the reference cell types for the subject and that estimates a contribution of unknown cell types not included in the reference methylation status dataset, whereby the machine learningsystem is trained to generate the estimator. The estimator, when used with individual patient data, generates a composite algorithm value that is converted to a predictive score relative to a reference population. In some embodiments, the set of data comprising a plurality of prospective patient records includes records from more than 1,000 patients.

[0016] In some embodiments, the computer implemented method further comprises iteratively regenerating the estimator when the estimator does not meet a predetermined receiver operating curve (ROC) statistic, by using a different subset of inputs and / or by adjusting the associated weights of the inputs until the regenerated estimator meets the predetermined ROC statistic. In some embodiments, the computer implemented method further comprises generating a static configuration of the estimator when the machine learning system meets a predetermined ROC statistic. In some embodiments, the computer implemented method further comprises configuring a computing device accessible by a user with the static estimator; entering values for the subset of the plurality of parameters corresponding to the patient into the computing device; and estimating, using the static estimator, the patient into a category indicative of a likelihood of having the cell-type specific degeneration or into another category indicative of a likelihood of not having the cell-type specific degeneration.

[0017] In some embodiments, the computer implemented method further comprises obtaining test results from a diagnostic test that confirms or denies the presence of the cell- type specific degeneration; incorporating the test results into the training data for further training of the machine learning system; and generating an improved estimator by the machine learning system.

[0018] In some embodiments, the estimator functions as a classifier (e.g., when predicting ALS vs healthy). In some embodiments, it functions as a regression (e.g., when estimating a score that measures disease progress). BRIEF DESCRIPTION OF THE DRAWINGS

[0019] FIGS.1A-1D provide an overview of epigenetic cfDNA biomarker development approach.1A: Firstly, tissue informative markers (TIMs) were selected using WGBS data to capture CpG sites that were hypermethylated or hypomethylated in a tissue of interest (SEQ ID NO: 1).1B: Next, cfDNA was extracted from the blood plasma of ALS cases and controls. 1C: The cfDNA was bisulfite-treated, hybridized to capture probes, designed as complementary to TIMs, and then sequenced. Some off-target reads were also captured. 1D: Using computational approaches, we analyzed the tissue of origin of the cfDNA samples and performed machine learning to identify features of ALS.

[0020] FIGS.2A-2D show cohort demographic and clinical characteristics. For the UQ (n=43) and UCSF (n=42) ALS patients, (2A) the distribution of the age of onset of ALS disease symptoms, where the dashed line indicates the median age of onset, (2B) patient ALSFRS-R scores (2C) FVC, and (2D) the number of days between cfDNA collection and date ALS symptoms were observed. In the box plots, the center line of the box indicates the mean, the outer edges of the box indicate the upper and lower quartiles, and the whiskers indicate the maxima and minima of the distribution. Each dot indicates an individual.

[0021] FIGS.3A-3E illustrate the capture panel design. (3A) The panel was designed to capture both hypomethylated TIMs, which were CpG sites that were less methylated in a tissue of interest relative to other tissues, and hypermethylated TIMs, which were designed to capture sites more methylated in a tissue of interest than other tissues (SEQ ID NO: 1). (3B) The methylation proportion of reference tissues at either the site the TIM was selected for, or all other tissues. (3C) The distance hyper- or hypo-methylated TIMs are from the transcription start site of a gene. (3D) The number of hyper- and hypo-methylated TIMs in different genomic regions. (3E) For samples where the true genome-wide methylation proportion was between 0.0 and 1.0 (larger dots near center of boxes), the observed methylation proportion after capture and sequencing. For all box plots, the center line of the box indicates the mean, the outer edges of the box indicate the upper and lower quartiles, and the whiskers indicate the maxima and minima of the distribution. Each (smaller) dot indicates an individual.

[0022] FIGS.4A-4D demonstrate the capture panel performance on cfDNA data. (4A) The starting cfDNA concentration of ALS patients and controls for each cohort, where each point represents one individual. (4B) Coverage of the on-target and off-target CpG sites of each cohort, where each dot represents one sample. (4C) Correlation between the UQ and UCSF methylation proportions at on-target sites. A single point represents a TIM. (4D) The proportion of cfDNA from the controls and cases in each cohort that was estimated to originate from skeletal muscle. The oval indicates outlier control individuals discussed in Example 2. For all box plots, the center line of the box indicates the mean, the outer edges of the box indicate the upper and lower quartiles, and the whiskers indicate the maxima and minima of the distribution. Each circle indicates an individual.

[0023] FIGS.5A-5D show ALS disease classification with cfDNA epigenetic features. The false positive rate versus true positive rate for models trained and tested using CpG coverage, CpG methylation, and covariates as input features for (5A) ten-fold cross- validation within UQ samples (5B) ten-fold cross-validation within UCSF samples (5C) trained on UCSF data and tested on UQ data, and (5D) trained on UQ data and tested on UCSF data.

[0024] FIGS.6A-6G show the features selected by the elastic net algorithm. (6A) For each tissue the TIMs were selected for, and for the type of TIM, the total absolute β value. A larger absolute β sum indicated that the feature type contributed more to model predictions. The β values for the (6B) methylation proportion and (6C) the read coverage of individual TIMs selected to be hypermethylated, and the β values for the (6D) methylation proportion and (6E) read coverage of individual TIMs selected to be hypomethylated. (6F) The methylation proportion of cases and controls for each cohort for a hypermethylated TIM in the SHISA5 gene. (6G) The read coverage of cases and controls for each cohort for a hypermethylated TIM located in the XRCC6 gene. For all box plots, the center line of the box indicates the mean, the outer edges of the box indicate the upper and lower quartiles, and the whiskers indicate the maxima and minima of the distribution. Each individual dot indicates a cfDNA sample.

[0025] FIGS.7A-7C demonstrate predictive performance of cfDNA epigenetic features for ALS phenotypes. For a tenfold cross-validated model trained using cfDNA methylation proportion and coverage features the predicted versus true (7A) ALSFRS-R, (7B) FVC, and (7C) ALSFRS-R slope. Each point represents one ALS case.

[0026] FIGS.8A-8D illustrate cohort demographic characteristics. For the UQ and UCSF cohorts, (8A) the distribution of the age of the cases and controls, (8B) the percentage of the cohorts that are female, and the percentage of the (8C) ALS cases and (8D) controls that identify as five different racial / ethnic categories.

[0027] FIGS.9A-9B show properties of captured TIMs. (9A) The number of TIMs selected per chromosome and (9B) for the two types of TIMs, the distribution of distances between a TIM and a CpG island.

[0028] FIGS.10A-10C show deconvolution of validation data. The CelFiE estimates (10A) for sheared genomic DNA (n=2) samples taken from blood and (10B) healthy cfDNA (n=3). (10C) For cfDNA taken from one individual before and after exercise, the proportion of cfDNA estimated to be originating from neutrophils.

[0029] FIGS.11A-11E illustrate the on target percentage. The percentage of reads that were on-target (11A) before deduplication and (11B) after deduplication. For each cohort, (11C) the percentage of the total mapped starting reads before deduplication that remained after deduplication. The on-target saturation, defined as 1-(median depth on target after deduplication / median depth on target before deduplication) for (11D) the UCSF cohort and (11E) the UQ cohort.

[0030] FIGS.12A-12D show the cell-type decomposition estimates. The proportion of cfDNA estimated by CelFiE to originate from each tissue for each sample type in the (12A) UCSF cohort and (12B) UQ cohort.

[0031] FIGS.13A-13D show ALS classification using CpG coverage. The false positive rate versus true positive rate for models trained and tested using only CpG coverage as input features for (13A) ten-fold cross validation within UQ samples (13B) ten-fold cross validation within UCSF samples (13C) trained on UCSF data and tested on UQ data, and (13D) trained on UQ data and tested on UCSF data.

[0032] FIGS.14A-14D show ALS disease classification using CpG methylation. The false positive rate versus true positive rate for models trained and tested using only CpG methylation proportion as input features for (14A) ten-fold cross validation within UQ samples (14B) ten-fold cross validation within UCSF samples (14C) trained on UCSF data and tested on UQ data, and (14D) trained on UQ data and tested on UCSF data

[0033] FIGS.15A-15D show ALS disease classification using only covariate information. The false positive rate versus true positive rate for models trained and tested using only covariate information (age, sex, SIRE, starting cfDNA concentration, and total cfDNA input) as input features for (15A) ten-fold cross validation within UQ samples (15B) ten-fold cross validation within UCSF samples (15C) trained on UCSF data and tested on UQ data, and (15D) trained on UQ data and tested on UCSF data

[0034] FIG.16 demonstrates the relationship between read coverage and predictive performance. For UCSF cfDNA samples, the total number of reads was randomly downsampled to reduce overall on-target CpG coverage relative to the actual UCSF read coverage. The downsampled samples were then used as input for elastic net models trained using 10 fold cross validation to predict case-control status in the UCSF cohort and the AUC was recorded. The within-cohort UQ AUC is indicated by an X.

[0035] FIGS.17A-17D show ALS disease classification without skeletal muscle TIMS. The false positive rate versus true positive rate for models trained and tested using cfDNA CpG methylation, CpG coverage, and covariate information (age, sex, SIRE, starting cfDNA concentration, and total cfDNA input) for all TIMs besides those chosen for skeletal muscle as input features for (17A) ten-fold cross validation within UQ samples (17B) ten-fold cross validation within UCSF samples (17C) trained on UCSF data and tested on UQ data, and (17D) trained on UQ data and tested on UCSF data.

[0036] FIGS.18A-18B show ALS disease classification with off-target CpGs. The false positive rate versus true positive rate for models trained and tested using off target cfDNACpG methylation trained and tested used (18A) ten-fold cross validation within UCSF samples (18B) ten-fold cross validation within UQ samples. DETAILED DESCRIPTION

[0037] The methods described herein are based on the development of a methylation capture panel that targets methylation sites that vary between tissues at high sequencing depth. The projected cost for a sample run on the proposed capture technology is 10-fold smaller than with whole genome bisulfite sequencing (WGBS). This unexpected combination of methylation sequencing, cell type identification, machine learning, and capture technology enables a previously unattainable ability to predict disease status. These methods can be used to diagnose, prognose, and stage disease conditions including amyotrophic lateral sclerosis (ALS) and other degenerative conditions. In addition, it offers great advantages for other cfDNA-based detection applications, such as cancer and pregnancy-related conditions.

[0038] To study cfDNA in ALS with the potential to be clinically relevant, we developed a capture panel technology that has a projected cost of under $300 per sample. The panel was designed to capture methylation sites that are informative for tissue status. Using this technology, we captured over four thousand regions and performed next-generation sequencing of cases and controls. Sequencing data was combined with supervised machine learning to predict ALS disease status using the methylation status of the reads and clinical covariates. With this approach, we identified higher skeletal muscle cfDNA in ALS patients. Furthermore, our model could accurately predict ALS disease status at the level of clinical utility. These results demonstrate that cfDNA provides a valuable predictive biomarker in ALS. Definitions

[0039] All scientific and technical terms used in this application have meanings commonly used in the art unless otherwise specified. As used in this application, the following words or phrases have the meanings specified.

[0040] As used herein, a “control” or “reference” sample means a sample that is representative of normal measures of the respective marker, such as would be obtained from normal, healthy control subjects, or a baseline marker to be used for comparison. Typically, a baseline will be a measurement taken from the same subject or patient. The sample can be an actual sample used for testing, or a reference level or range, based on known normal measurements of the corresponding marker or based on sampling a broad group of subjects and / or cell types.

[0041] As used herein, an “increase”, “decrease”, or “distinct from” are terms used to indicate an observable difference relative to a comparison or reference value, as would be understood by a person skilled in the art. In some embodiments, the observable difference is a statistically significant difference.

[0042] As used herein, the term “probe” refers to an oligonucleotide, including DNA and RNA, naturally or synthetically produced, via recombinant methods or by PCR amplification, that hybridizes to at least part of another oligonucleotide of interest. A probe can be single- stranded or double-stranded. Optionally, the probe is labeled with a detectable marker, such as biotin.

[0043] As used herein, the term “active fragment” refers to a substantial portion of an oligonucleotide that is capable of performing the same function of specifically hybridizing to a target polynucleotide.

[0044] As used herein, "hybridizes," "hybridizing," and "hybridization" means that the oligonucleotide forms a noncovalent interaction with the target DNA molecule under standard conditions. Standard hybridizing conditions are those conditions that allow an oligonucleotide probe or primer to hybridize to a target DNA molecule. Such conditions are readily determined for an oligonucleotide probe or primer and the target DNA molecule using techniques well known to those skilled in the art. The nucleotide sequence of a target polynucleotide is generally a sequence complementary to the oligonucleotide primer or probe. The hybridizing oligonucleotide may contain nonhybridizing nucleotides that do not interfere with forming the noncovalent interaction. The nonhybridizing nucleotides of an oligonucleotide primer or probe may be located at an end of the hybridizing oligonucleotide or within the hybridizing oligonucleotide. Thus, an oligonucleotide probe or primer does not have to be complementary to all the nucleotides of the target sequence as long as there is hybridization under standard hybridization conditions.

[0045] As used herein, the term "subject" includes any human or non-human animal. The term "non-human animal" includes all vertebrates, e.g., mammals and non-mammals, such as non-human primates, horses, sheep, dogs, cows, pigs, chickens, and other veterinary subjects. In a typical embodiment, the subject is a human.

[0046] As used herein, “a” or “an” means at least one, unless clearly indicated otherwise.

[0047] As used herein, “Cell Free DNA Estimation via expectation-maximization” or “CelFiE” estimates the contribution of various cell types to the cfDNA of an individual via an EM optimization algorithm. The input to CelFiE is WGBS reference data consisting of T total cell types and WGBS cfDNA samples for N total individuals. Its output is the proportion of the reference cell types that make up each individual’s cfDNA, such that the proportion of all Tcell types sums to one for each individual. Notably, an arbitrary number of cell types can be missing, which addresses potential biases arising from estimating the proportions of cell types from a restricted reference panel. CelFiE also estimates the methylation values for each of the cell types included in the reference, which accommodates the currently noisy and low-coverage reference data sets. These developments are facilitated by CelFiE’s EM algorithm, which is a flexible framework for parameter estimation, even when there is missing data. CelFiE is detailed in Caggiano, C., et al. Comprehensive cell type decomposition of circulating cell-free DNA with CelFiE. Nat Commun 12, 2717 (2021).

[0048] Methods

[0049] The methods described herein provide a means to use cell-free DNA (cfDNA) as a non-invasive biomarker for ALS, ALS progression, detection and monitoring of other neurodegenerative diseases, and other conditions with degenerative processes. CfDNA is an ideal candidate because it is enriched in ALS patients, is informative about tissue-specific cell death, and can be extracted from a standard blood draw. The methods described herein can be used to distinguish between ALS patients and healthy controls or other neurodegenerative diseases, such as, for example, Alzheimer’s disease. The method can also be used to distinguish different subtypes of ALS, such as, for example, primary lateral sclerosis (PLS), which has a much different time-course and disease progression relative to other forms of ALS.

[0050] Described herein is a method of detecting cell type specific degeneration in a blood or other biological sample obtained from a subject. In some embodiments, the method comprises: (a) contacting circulating extracellular DNA extracted from a biological sample with a set of capture probes, wherein the capture probes are less than 50,000 in number, and wherein the capture probes specifically bind within 100 base pairs (bp) of methylated and unmethylated CpG sites on cell-free DNA (cfDNA), and wherein the methylation status of CpG sites is specific to a cell type of interest. The method further comprises: (b) performing methylation sequencing of cfDNA captured in step (a); and (c) creating a methylation status dataset from the sequencing of step (b). The methylation status dataset contains the methylation status of the cell type specific CpG sites (CTSCS). The method further comprises (d) detecting cell type specific degeneration when the methylation status of the CTSCS is distinct from a reference methylation status dataset, wherein the reference methylation status dataset contains the methylation status for the CTSCS of all cell types. In some embodiments, the method further comprises inputting the methylation status dataset into a statistical model. Alternatively, in some embodiments, the methylation status dataset has previously been inputted into the statistical model.

[0051] The model maps the methylation status dataset to a reference methylation status dataset. The reference methylation status dataset contains the methylation status for the CTSCS of all cell types. In some embodiments, the CTSCS are identified using a supervised machine learning model and expectation maximization (EM) algorithm that estimates the proportion of the reference cell types for the subject and that estimates a contribution of unknown cell types not included in the reference methylation status dataset.

[0052] The sites selected to capture and employ as input to the machine learning model are assigned by the model a weight for each of the sites. Some sites are assigned 0 weight because the model determined the site was not important. The data generated in this was informs which sites the model is picking up as important, whereby a higher weight is indicative of a site that is more important for discriminating between ALS and controls. The sites found to be informative of regions of the genome can be used to identify the tissues that are important in the cfDNA of ALS patients. For example, Fig.6A shows the average weight per tissue category for various tissue types. This can be used as a guide for identifying informative tissues, e.g., for discriminating between ALS and healthy patients.

[0053] In some embodiments, the set of capture probes is between 20 and 20,000 in number. In some embodiments, the set of capture probes is between 500 and 15,000 in number. In some embodiments, the set of capture probes is between 1,000 and 10,000 in number. In some embodiments, about 5,000 capture probes are included in the set. In some embodiments, the statistical model comprises a penalized regression model.

[0054] Also provided is a method of monitoring progression of a degenerative disease in a subject. In some embodiments, the method comprises performing the method of described above at a first time point on a biological sample obtained from the subject, and repeating the method at a subsequent time point, wherein an increase or decrease in the methylation status of the CTSCS is indicative of disease progression. One can look for changes in the sum of the weights across time points from one individual to see how their disease changes. Such analysis can be used to generate a “ALS disease risk score”. A change in slope of the risk score can inform regarding the speed of disease progression in that subject. In some embodiments, a separate algorithm is generated using training data in the form of progression as the input. The lasso approach described herein is employed to learn features of progression. One can assess changes in predicted progression by running the algorithm on subsequent samples. Changes in the amount of cfDNA from specific cell types and states are employed in the same manner, but tailored to disease progression. For ALS / PLS or ALS / other neurological disease, one can use the same machine learning approach. Because the outcome is different, the informative CpG sites selected by the model are different. Sincethe CpGs were included to represent different cell types, the model is learning about the different contributions of cell types to ALS and PLS cfDNA in this example.

[0055] In some embodiments, the biological sample is a blood sample. In some embodiments, the biological sample is extracellular DNA that has been extracted from a blood sample, or other biological specimen. In some embodiments, the cfDNA is extracted from the biological sample as a separate step prior to contact with the capture probes. In some embodiments, the extracellular DNA extracted from the sample is treated with bisulfite, which introduces methylation-dependent sequence changes through selective chemical conversion of non-methylated cytosine to uracil. After treatment, all non-methylated cytosine bases are converted to uracil, but all methylated cytosine bases remain cytosine. These methylation dependent C-to-T changes can subsequently be studied using conventional DNA analysis technologies.

[0056] In some embodiments, the cell type specific degeneration is amyotrophic lateral sclerosis (ALS). In some embodiments, the cell type specific degeneration is traumatic brain injury (TBI), cancer, Alzheimer’s disease, Parkinson’s disease, or a pregnancy-related phenotype. Examples of a pregnancy-related phenotype include, but are not limited to, gestational diabetes, pre-term birth, and preeclampsia.

[0057] In some embodiments, the methylation sequencing comprises high throughput sequencing.

[0058] In some embodiments, the CTSCS are specific to skeletal muscle, fibroblasts, neurons, and / or hematopoietic cells.

[0059] Computer Implementation & System

[0060] Also described is a computer implemented method of training a machine learning system to generate an estimator for identifying a cell type specific degeneration in a blood sample obtained from a subject. In some embodiments, the method comprises: (a) storing a set of data comprising a plurality of prospective patient records from more than 100 patients, each prospective patient record including a plurality of parameters and corresponding values for each patient included in the patient records, and a diagnostic indicator indicating whether or not the patient included in the patient records has been diagnosed with a specific condition associated with a cell type specific degeneration; (b) selecting a subset of the plurality of parameters for inputs into the machine learning system. The subset comprises CTSCS, and the subset further includes at least one clinical parameter selected from age, gender, and smoking status. The method further comprises (c) randomly partitioning the set of data into training data and validation data; and (d) generating the estimator. The machine learning system is trained based on the training data and the subset of inputs.

[0061] In some embodiments, the estimator is trained with a supervised machine learning model and expectation maximization (EM) algorithm that estimates the proportion of the reference cell types for the subject and that estimates a contribution of unknown cell types not included in the reference methylation status dataset, whereby the machine learning system is trained to generate the estimator. The estimator, when used with individual patient data, generates a composite algorithm value that is converted to a predictive score relative to a reference population. In some embodiments, the set of data comprising a plurality of prospective patient records includes records from more than 1,000 patients.

[0062] In some embodiments, the computer implemented method further comprises iteratively regenerating the estimator when the estimator does not meet a predetermined receiver operating curve (ROC) statistic, by using a different subset of inputs and / or by adjusting the associated weights of the inputs until the regenerated estimator meets the predetermined ROC statistic. In some embodiments, the computer implemented method further comprises generating a static configuration of the estimator when the machine learning system meets a predetermined ROC statistic.

[0063] In some embodiments, the computer implemented method further comprises configuring a computing device accessible by a user with the static estimator; entering values for the subset of the plurality of parameters corresponding to the patient into the computing device; and estimating, using the static estimator, the patient into a category indicative of a likelihood of having the cell-type specific degeneration or into another category indicative of a likelihood of not having the cell-type specific degeneration.

[0064] Also provided is a system comprising a computing device programmed to perform the methods described herein. In some embodiments, the system interfaces with tubes for receiving drawn blood, machinery that extracts cell free DNA, kits for processing DNA, sequencers, and computers for prediction. Such interface can include, for example, communicating instructions to initiate these tasks, and / or receiving input from other elements of the system.

[0065] In some embodiments, the computer implemented method further comprises obtaining test results from a diagnostic test that confirms or denies the presence of the cell- type specific degeneration; incorporating the test results into the training data for further training of the machine learning system; and generating an improved estimator by the machine learning system.

[0066] In some embodiments, the estimator functions as a classifier (e.g., when predicting ALS vs healthy). In some embodiments, it functions as a regression (e.g., when estimating a score that measures disease progress).EXAMPLES

[0067] The following examples are presented to illustrate the present invention and to assist one of ordinary skill in making and using the same. The examples are not intended in any way to otherwise limit the scope of the invention.

[0068] Example 1: Development of a non-invasive biomarker discovery in amyotrophic lateral sclerosis

[0069] This Example demonstrates the development of cfDNA as a biomarker for ALS, as well as for its use in identifying the tissue of origin, as well as providing a biomarker for other degenerative conditions. Circulating cell-free DNA (cfDNA) in the bloodstream comes from dying cells and can be used to learn about tissue health. However, cfDNA sequences themselves provide very little information about the tissue of origin.

[0070] Our preliminary study showed that cfDNA concentrations in plasma were significantly increased in ALS patients (n=28) relative to age-matched controls (n=25). We developed CelFiE to decompose cfDNA origins allowing noisy data and unknown cell types. Using CelFiE coupled with WGBS, we performed a preliminary study to characterize the tissue of origin of 16 ALS patients and 16 age matched controls. CelFiE estimated a higher skeletal muscle component in ALS patients than in controls, suggesting cfDNA could act as a quantitative biomarker.

[0071] WGBS data is expensive and most DNA methylation sites are not informative for tissue or disease state. Our solution is a capture technology to be used in parallel with statistical learning algorithms to learn about disease state. Probe capture can pull down a select number of chosen fragments in both their methylated and unmethylated states. We designed a probe set to only pull down cfDNA fragments containing highly informative CpGs In total, we designed ~5000 probes to target CpG sites that are informative for disease and tissue state.

[0072] We validated our capture panel and then ran it on an expanded cohort: 48 ALS patients and 48 controls. We confirmed our previous result of increased cfDNA originating from skeletal muscle in ALS patients in this larger cohort. We also observed overall methylation proportion differences in the probe sites between ALS patients and controls, a trend which can be seen using principle component analysis (PCA).

[0073] We wanted to know if we could accurately predict whether someone was an ALS patient or a control using cfDNA methylation alone. Our lasso (least absolute shrinkage and selection operator) model to predict binary ALS status performed well, and provided a far more effective tool than PCA. The lasso penalized logistic regression included 10 fold cross validationusing 43 case and 46 controls. The covariates were: cfDNA concentration, binary sex, binary ethnicity (white / non-white), and age.

[0074] Example 2: Tissue informative cell-free DNA methylation sites in amyotrophic lateral sclerosis

[0075] This Example demonstrates the new approach combining advances in molecular and computational technologies. First, statistical tools are developed to select tissue-informative DNA methylation sites relevant to a disease process of interest. A capture protocol is then employed to select these sites and perform targeted methylation sequencing. Multi-modal information about the DNA methylation patterns are then utilized in machine learning algorithms trained to predict disease status and disease progression. We applied our method to two independent cohorts of ALS patients and controls (n=192). Overall, we found that the targeted sites accurately predicted ALS status and replicated between cohorts. Additionally, we identified epigenetic features associated with ALS phenotypes, including disease severity. These findings highlight the potential of cfDNA as a non-invasive biomarker for ALS.

[0076] A limitation of whole genome epigenetic approaches is that the cost to achieve high sequencing coverage(14) is currently too expensive to be routinely applicable in clinical settings.(2,15) High sequencing coverage, however, is needed since certain cfDNA fragments may only be present in low quantities, which could be missed by shallow sequencing.(16) Furthermore, many methylation sites are not variable,(17) limiting their value in biomarker development.

[0077] To address these limitations, previous work has successfully used DNA methylation capture(18) to enrich for only relevant genomic regions, which can reduce sequencing costs while maintaining high coverage. Examples of DNA methylation capture in cfDNA applications include an approach to classify cancer types and to predict whether a patient develops preeclampsia. (3,8,19) A limitation of existing approaches, however, is that they are often optimized for a specific disease context. To the best of our knowledge, methylation capture of cfDNA has not yet been adapted for use in neurodegenerative disease.

[0078] In this Example, we developed an algorithm to identify regions of the epigenome that are informative for the presence of a tissue in the cfDNA. These regions can be used to learn about tissue death in a range of diseases, including neurodegeneration. We then developed algorithms that leveraged differences in the methylation of these tissue informative sites to classify patients by disease status based on their epigenetic cfDNAprofile. This methodology can be used to characterize the contribution of diverse tissues to ALS cfDNA, leading to a multidimensional picture of disease.

[0079] We applied this technology to two independent cohorts from The University of Queensland in Brisbane, Australia (UQ) and the University of California at San Francisco, United States (UCSF) comprising a total of 192 cfDNA samples from ALS patients, healthy controls, and patients with other neurological diseases. Together, these cohorts represent the largest application of cfDNA in the study of ALS to date. Consistent with our previous research, we found significantly elevated cfDNA concentrations in ALS patients in both cohorts. (10) Our machine learning model significantly predicted ALS disease status in both cohorts with high accuracy (UQ AUC=0.82, UCSF AUC=0.99). The model discriminated ALS cases from patients with other neurological diseases, such as frontotemporal degeneration. It also identified a previously unknown asymptomatic carrier of a pathogenic variant in C9orf72, which is the main genetic cause of ALS. Finally, we identified methylation sites associated with other ALS phenotypes, including disease severity. Together, these results suggest that epigenetic alterations in cfDNA are promising quantitative biomarker candidates for ALS, which can be used to non-invasively study the impacts of neurodegeneration.

[0080] Methods

[0081] Patient Recruitment and Clinical Data

[0082] A total of 192 participants were enrolled in a prospective manner at the UCSF ALS Clinic in San Francisco, California, USA, the Royal Brisbane and Women’s Hospital and Mater Hospital in Brisbane, Australia under neurologist supervision from 2018-2021. All participants provided written informed consent and the study received approval from the Human Research Ethics Committee at the Royal Brisbane and Women’s Hospital (HREC / 17 / QRBW / 299) and by the UCSF Committee on Human Research (IRB 10-05027).

[0083] Patients (with ALS / being assessed for ALS) and when possible, control (non-related, closely age-matched family members, caregivers or volunteers) were recruited. A second set of other neurological controls were recruited from a non-ALS outpatient clinic under neurologist supervision. Allocation to diagnostic groups was performed according to the latest available clinical information (clinical censor date October 2024).

[0084] For cases and controls, age, sex, and self-identified race / ethnicity (SIRE) were recorded. For ALS cases at the time of visit, FVC and ALSFRS-R were taken, and ALSFRS- R slope and FVC slope relative to the previous visit were calculated. The symptom onset site and date of first symptoms were also recorded.

[0085] To stabilize the cell-free DNA, all blood samples were collected in the PAXgene Blood ccfDNA Tubes following a clinic appointment. To ensure enough cfDNA was available for downstream applications 20 mL of whole blood from controls / OND and 10 mL of whole blood from cases were collected. Following laboratory receipt (typically within 24-48hrs of collection) blood was spun with the brake off (10mins, 1900g) before plasma was aliquoted and spun twice (10mins, 16000g) to remove any further debris. Plasma was then stored at - 80 before further processing.

[0086] Library Preparation and Sequencing

[0087] Using a harmonized protocol across two sites (UCSF and UQ) cfDNA was extracted and prepared for sequencing. Briefly, plasma was thawed at room temperature and cfDNA was extracted from all available plasma (range 2-8 ml) using the QIAGEN Circulating Nucleic Acid kit (Cat No: 55114) according to the manufacturer’s recommendations. Extracted cfDNA was quantified using Qubit dsDNA HS Assay and visualized using the cfDNA assay (Agilent - TapeStation 4200 (UCSF) and Agilent Bioanalyzer 2100 (HS kit) (UQ)). cfDNA was bisulfite converted using the Zymo Lightning kit (Zymo Research) and underwent library preparation using the Accel-NGS Methyl-Seq (Swift Biosciences) according to the manufacturer’s instructions with a major modification. Briefly, the denatured BS-converted cfDNA was subject to the adaptase, extension, and ligation reaction. Following the ligation purification, the DNA underwent primer extension (98C for 1 minute; 70C for 2 minutes; 65C for 5 minutes; 4C hold) using oligos containing random UMI and i5 barcodes. #The extension using a UMI-containing primer allows the tagging of each individual molecule in order to be able to remove PCR duplicates and correctly estimate DNA methylation levels.

[0088] Following exonuclease I treatment and subsequent purification, the libraries were then amplified using a universal custom P5 primer and custom i7-barcoded P7 primers (initial denaturation: 98C for 30 seconds; 15 cycles of: 98C for 10 seconds, 60C for 30 seconds, 68C for 60 seconds; final extension: 68C for 5 minutes; 4C hold). The resulting unique-dual indexed libraries were then purified, quantified using the Qubit HS-dsDNA assay, the quality checked using the D1000-HS assay (Agilent - TapeStation 4200), and grouped as 12-plex pools. Each pool was then subject to hybridization capture using the xGen Hybridization Capture Kit (IDT) using custom probes designed on approximately 5000 pre-selected regions.

[0089] For each top and bottom strands of the regions of interest, two probes were designed: one “unmethyl” probe with all G bases converted to A, and one “methyl” probe with all non-CpG G bases converted to A.

[0090] Following the hybridization capture, a final amplification PCR (initial denaturation: 98C for 30 seconds; 10 cycles of: 98C for 10 seconds, 60C for 30 seconds, 68C for 60 seconds; final extension: 68C for 5 minutes; 4C hold) has been performed, followed by SPRI beads purification and quantification as QC as previously described. To maximize consistency across sites, the same probes were used (shipped to Australia following UCSF library preparation).

[0091] The final pool of libraries was submitted for sequencing on an Illumina NovaSeq6000 (USA; UCSF Sequencing facility, Australia;UNSW Ramaciotti Sequencing facility) using identical run conditions (S4 lane - 150 PE, 8bases for i7, 17 bases for i5).

[0092] Tissue informative marker selection

[0093] TIMs were selected for 19 tissues and cell types: dendritic cells, endothelial cells, eosinophils, erythroblasts, macrophages, monocytes, neutrophils, T-cells, adipose, brain, fibroblast, heart, hepatocytes, lung, megakaryocytes, skeletal muscle, small intestine, placenta, and mammary epithelial cells. These tissues were determined based on our previous work to be relevant to ALS, or selected based on previous publications to be the primary contributors to cfDNA. At least two WGBS samples per reference dataset were obtained. The average methylation per CpG for the reference tissue replicates was calculated.

[0094] Per CpG, for one tissue at a time, the distance between the methylation proportion at that tissue and the mean methylation of all other tissues was calculated. The N sites per tissue with the greatest difference were kept as TIMs. If two tissues had the same CpG classified as a TIM, it was removed from both lists.

[0095] To begin, we selected 300 potential TIM sites and then performed quality control checks. To ensure that TIMs were sites that would be covered in cfDNA data, we used two WGBS cfDNA datasets and removed any CpG site that had less than an average of 10X coverage in both datasets. We also removed TIMs that overlapped a common SNP (minor allele frequency > 5%). Since we wanted to have the greatest diversity of regions targeted in the capture, if there were multiple TIM sites within 500bp of each other, we kept only the first site. Additional quality control was performed to remove TIMs overlapping repetitive regions and with low predicted target efficiency. After quality control, 4,994 TIMs remained.

[0096] Probe design

[0097] For each of the 4,994 TIMs, both a methylated and unmethylated probe were designed to bind to and capture both possible states of the targeted CpG. To increase the efficiency of the capture, 120 base pair probes were designed to target a window around theTIM. During bisulfite conversion, any cytosine base not protected by a methyl group in position 5 is converted into thymine.(60) Since methylation in humans primarily occurs at CpG sites, this means that all cytosines on the forward strand would be converted to thymine. Thus, to capture the unmethylated CpG state, the unmethylated probe was designed with all guanine bases converted to adenine. For the methylated state, where only cytosines in a CpG dinucleotide would be protected from the bisulfite treatment, only non- CpG guanine bases were converted to an adenine.

[0098] Bioinformatic processing

[0099] For data generated at UCSF, UMIs were first extracted from the index read and added to the header of the corresponding R1 and R2 fastq file using umi_tools.(61) This step was skipped for UQ samples since UMIs were not sequenced. For samples from both institutions, adapters were trimmed using trim_galore. Read alignment, processing, and methylation calling were performed using BsBolt v 1.6.1(62) in an adapted pipeline published in Morselli et at.(18) Reads were aligned to an hg38 bisulfite converted genome, which was generated using the BsBolt Index over an hg38 fasta file obtained from the UCSC genome browser. Reads were aligned using BsBolt Align in paired end mode with default parameters. To prepare for duplicate removal, aligned reads were subject to samtools fixmate and sorted.(63) Umi_tools(61) dedup in paired end mode was used to remove duplicate reads.

[0100] For both cohorts, CpG methylation was called using the command BsBolt CallMethylation-BG-CG-remove-ccgg. The CG parameter restricted to only CpG sites (ignoring non-CpG methylation), the the BG parameter sent the output to a bedgraph file and the -remove-ccgg parameter removed methylation calls in ccgg regions.

[0101] Genetic sex

[0102] As a quality control metric, we estimated the genetic sex of the samples and assessed how they corresponded to self-reported sex. We did this using scripts from Phung et al,(64) which calculates the number of reads mapped to chromosome 19 and compares them to the number of reads mapped to the X chromosome. In individuals assigned female at birth, the ratio of chromosome 19 reads should be approximately 1 since they have two X chromosomes and two chromosome 19. We removed one individual whose genetic sex did not match their reported clinical data.

[0103] Deconvolution

[0104] cfDNA deconvolution was performed using CelFiE, which is a supervised deconvolution algorithm that is designed for noisy read count data and missing referencetissues. Input sites for CelFiE were the on-target TIMs selected for capture, As demonstrated in the CelFiE publication, summing reads from adjacent CpGs can improve deconvolution performance by decreasing sampling noise. As such, reads were summed + / -250bp around the target CpG. Sites with no reads covering the CpG were set to have a read depth of zero.

[0105] Deconvolution was performed using tissues representing organs and hematopoietic cell types, selected for their relevance in cfDNA.(10,28,29) CelFiE can estimate an arbitrary number of unknown tissues. Since CelFiE learns from both the input and reference data, the number of samples influences the accuracy of unknown estimation. Based on simulation experiments published in the original CelFiE paper, 2 unknowns were chosen for the sample size of 96 total cfDNA input samples.

[0106] The reference panel for CelFiE consisted of 19 tissues over the same on-target TIMs as the input matrix. Reference samples were WGBS samples obtained from ENCODE(25) and Blueprint.(26,65) Reference samples were also summed in 500bp regions around the target CpG.

[0107] CelFiE was run over the UMI-deduplicated UQ samples, the UMI-deduplicated samples, and both cohorts combined. The CelFIE default of 10 random restarts was used.

[0108] After running deconvolution, differences in cell-type proportion between cases and controls were tested for one tissue at a time using the Python StatsModels package. A logistic regression model was run where the outcome was the binary case / control status and the input variable was the estimated tissue of origin proportion for a given tissue. Age, sex, and genetic ancestry were used as covariates.

[0109] Machine learning preprocessing

[0110] Samples with more than 10% of targeted CpGs missing, meaning that no reads were covering a CpG, were removed. Any site that had a median read coverage of 1 read or less was also removed. For the remaining sites and samples, the input matrix was made by dividing the number of methylated reads by the total number of reads. Imputation was performed per cohort over the methylation proportion matrix using SoftImpute, implemented in the Python package fancyImpute. For methylation coverage features, the coverage was normalized per sample by dividing the number of reads at a CpG by the total number of sequencing reads per individual.

[0111] Sex and SIRE were one-hot encoded and added as columns in the input matrix. Age, cfDNA starting concentration, and total cfDNA input were included as continuous covariates. Two separate matrices were kept, one for the ALS case / control status, and one for the methylation proportion and covariates.

[0112] Disease classification

[0113] Elastic net regression was performed in R using the BigStatsR package and big_spLogReg command.(42) ALS disease status served as the binary outcome variable, while the DNA methylation proportion at targeted CpGs and clinical variables served as predictors. We incorporated age, genetic sex, SIRE, cfDNA concentration (nanograms / microliter), and total input cfDNA quantity (nanograms) into the regression models as non-penalized variables.

[0114] Models were first trained on each cohort separately and then applied to the second cohort. The alpha parameter which controls model sparsity, was selected by performing ten- fold cross-validation on the training cohort and picking the optimal value. The BigStatsR package removes the manual selection of an optimal lambda value by introducing the Cross- Model Selection and Averaging (CMSA) procedure.(42,66) In brief, CMSA separates the training set into K folds and then performs cross-validation within the training set to obtain a set of vectors of predictions. This set of coefficients is averaged to produce the final coefficient value. For our model, we used the BigStatR default K value of 10. To standardize the weights produced per CpG site in each model, we scaled input value parameters to have mean zero and variance one. We scaled the test and training data separately.

[0115] Cohort-only models were trained only within a single cohort using ten-fold cross- validation. To evaluate the overall performance of the two cohorts, we trained a single model combining both sets of data and adding the cohort site as a non-penalized covariate. We used generalized linear models with a logit link function and additionally report area under the receiver operator curve (AUC).

[0116] Analysis of important features

[0117] To examine the importance of important DNA methylation and methylation coverage features in making model predictions, we obtained the weights, or β-values, at each feature from the combined cohort model. We merged the feature β-values with information on what tissue a TIM was selected for and whether it was hyper- or hypo-methylated. We used HOMER(67) to intersect a TIM with the closest gene to the TIM site.

[0118] To assess the relationship between the methylation or coverage at a specific TIM site, we performed a logistic regression, with the methylation value of the samples as the predictor and case-control status as the outcome. We used SIRE, age, sex, cfDNA concentration, and total cfDNA input as covariates.

[0119] ALS disease phenotype prediction

[0120] ALS disease prediction models were trained for ALSFRS, ALSFRS Slope, and FVC. The top 1000 methylation features and top 1000 coverage features from the combined case- control prediction model were used as input to the model along with age, sex, SIRE, input cfDNA concentration, and total cfDNA input as non-penalized covariates. Due to low sample sizes for the case-only analysis, we meta-analyzed the two cohorts and additionally added cohort as a non-penalized covariate. We trained the elastic net model using the BigStatsR package with the big_spLinReg command. Each of the three models were evaluated against an elastic net model trained on only the covariates.

[0121] Off target prediction models

[0122] Off target prediction models incorporated information for all CpGs obtained from high throughput sequencing. To do this, we found the union of all sites across all samples in a cohort. To maximize the number of off-target sites considered, we then removed sites with more than 5% missingness. Due to the lower coverage and increased number of sites, we did not impute missing sites. Since cohorts had differences in sequencing depth and on- target coverage, sites were analyzed separately. Case control status was then predicted using ten-foldcross validation in an elastic net model using the BigStatsR package in the same manner as the on-target models.

[0123] Downsampling simulations

[0124] To simulate samples with lower read depth, we used picard DownsampleSam68 to randomly remove reads at specified proportions of the total starting amount of reads to produce a bam file. We did this for each UCSF cfDNA sample. Then, methylation was re- called on the downsampled bam file using BsBolt to produce a new estimate of the methylation proportion and coverage of a CpG. We then subset to the on-target CpGs and individuals used in Fig.5B, imputing any missing values with SoftImpute. Then, an elastic net model was trained as described above. The ten-fold cross validated AUC was recorded for each set of downsampled samples.

[0125] Results

[0126] The approach was composed of four steps. First, we analyzed published whole- genome bisulfite sequencing (WGBS) tissue data to identify methylation sites with distinct patterns in a tissue of interest. We call these sites tissue-informative markers (TIMs). Previously published cfDNA WGBS data from diverse disease contexts was used to screen candidate TIMs for those actually observed in cfDNA. We then designed methylated and unmethylated probes complementary to these regions (Fig.1A). Next, cfDNA from our two cohorts was extracted (Fig.1B) and underwent methylation profiling on the TIM-enriched cfDNA (Fig.1C). Lastly, we analyzed the methylation status of the targeted regions anddeveloped statistical and machine learning approaches to learn about the disease status of the ALS patients and controls (Fig.1D).

[0127] Cohort characteristics

[0128] Our approach was applied to participants (n=192) who were recruited between 2018 and 2021 from two independent university-affiliated neurology clinics at UCSF and UQ (Table 1). The Revised El Escorial diagnostic criteria20 were used to classify cases (See Methods). Cases were composed of two groups of patients, those who had likely or probable ALS according to the criteria (referred to here as “ALS”), and those classified as possible ALS or primary lateral sclerosis (PLS) (referred to here as “PLS”), which is a related motor neuron disease.21,22

[0129] Table 1: Clinical characteristics.

[0130] The clinical and demographic characteristics per cohort. The number of total patients is shown, and the number of female patients is shown in parentheses.

[0131] The UCSF cohort comprised 42 ALS cases, 9 PLS cases, and 45 healthy age- matched controls consisting of unrelated partners or carers. At UQ, a total of 48 cases were enrolled (N=43 ALS and N=5 PLS). Forty-eight UQ controls were enrolled, consisting of both unrelated partners / carers (N=32) and patients with other neurological diseases (OND) (N=15). The UQ OND samples included a cross-section of neurological conditions, including diseases that share pathophysiology with ALS, like frontotemporal degeneration,(23) and other neurodegenerative diseases like Alzheimer’s disease (Table 2). Therefore, the UQ cohort represented a challenging real-world scenario for ALS biomarker development.

[0132] Table 2: Other neurological disease patients.

[0133] For each of the controls with other neurological diseases in the UQ cohort, the type of neurological disease (if known) and the number of patients with that disease.

[0134] There was heterogeneity of disease characteristics within and between cohorts. Both the UCSF and UQ cases had overlapping distributions in terms of age of onset, defined as the date the first ALS symptom was observed (Fig.2A). For each cohort, ALS severity was measured using the ALS Functional Rating Scale-Revised (ALSFRS-R)24 at the time of cfDNA collection, which is a qualitative measure of physical functioning on a scale from 0 (not functioning) to 48 (high functioning). The change in ALSFRS-R between visits, referred to as ALSFRS-R slope, was also calculated as a metric of disease progression. We found that cohorts were similar in the distribution of ALSFRS-R and ALSFRS-R slope, although the UCSF had slightly more progressed cases (Fig 2B). UQ samples had higher forced vital capacity (FVC) (t-test p-value=4.5×10-5), which is a measure of lung function, where a higher value indicates better function (Fig.2C). The two cohorts were also similar in the distribution of days between cfDNA collection and symptom onset (Fig.2D). We noted that patients in the UCSF cohort were slightly older (UQ mean age: 61.45 ± 8.17, UCSF mean age: 66.33 ± 9.96)) and that the UCSF cohort also contained patients from a larger variety of self-reported racial and ethnic (SIRE) backgrounds (Fig.8A-8D).

[0135] Selecting tissue informative markers

[0136] After collecting cfDNA, we turned to selecting methylation sites that are variable between tissues. In our previous work,(10) we introduced the concept of tissue informative markers (TIMs) as a method to identify methylation sites that vary between tissues and cell types. Briefly, a TIM is a site that is either hyper- or hypo-methylated relative to the average methylation proportion of all other tissues at that site (Fig.3A).

[0137] To find TIMs, we used reference WGBS methylomes that were obtained from two public reference consortiums, ENCODE(25) and Blueprint.(26) For this work, we focused on CpG sites as candidate TIMs, as most non-CpG sites are not methylated in adult tissues.27 We selected approximately 300 TIMs for 19 tissues (Table 3), which were prioritized based on deconvolution results from our previous work(10) and other recent works.(28,29) These tissues included several hematopoietic cell types, organs, epithelium, and brain (Table 3). We applied several filtering criteria to enrich for CpG sites that appear in previously published WGBS cfDNA data.

[0138] Table 3: TIM selection design.

[0139] Per tissue selected for capture, the number of hypermethylated TIMs selected, the number of hypomethylated TIMs selected, and the total number of final TIMs selected for capture.

[0140] An important property of cfDNA is that their fragmentation patterns are non- random.(30–32) cfDNA observed in blood generally are fragments approximately 160 basepairs long,(33) suggesting that cfDNA fragments are protected from degradation in the blood by the presence of tightly associated histone proteins. Since DNA from compacted chromatin is more likely to be protected and methylated, we chose to select a greater number of TIMs per tissue that were hypermethylated (Table 3) (Fig.3B).

[0141] After quality control, the final number of TIMS was 4,994. TIM sites were distributed throughout the genome (Fig.9A). Hypermethylated TIMs were closer, on average, to transcription start sites and CpG Islands than hypomethylated TIMs (Fig.3C; Fig.9B). Since at a hypermethylated TIM, all other tissues are predominantly unmethylated, this observation is consistent with the role of unmethylated CpGs in facilitating transcription.(34) Likewise, hypomethylated TIMs were more likely to be in intergenic and intronic regions (Fig.3D), suggesting that in most tissues, these sites did not have a strong regulatory function. Together, this suggests that hypermethylated and hypomethylated TIMs offer complementary types of genomic information.

[0142] Capture panel sequencing and validation

[0143] After designing the probes, we performed several validation experiments to ensure that probes could accurately profile the methylation state of the chosen TIMs. First, we used universal methylated DNA standards to create mixtures where the CpG sites were methylated 0, 25, 50, and 100% of the time. We captured the synthetic DNA mixtures with the probes and performed high-throughput sequencing. For each DNA mixture, we estimated the proportion of the time the captured CpG was methylated. We found that the observed methylation was highly concordant with the true methylation proportion (Fig.3E), suggesting that the probes were quantifying the methylation accurately.

[0144] Next, to examine how the capture panel might perform in real-world cfDNA scenarios, we validated the capture panel using sheared genomic DNA from blood (n=2), along with healthy cfDNA samples (n=3). After performing cell-type deconvolution with CelFiE,(10) we found that the sheared blood samples were estimated to be primarily composed of white blood cells (Fig.10A). The majority of cfDNA from healthy controls was also estimated to be originating from neutrophils and lymphocytes, consistent with published research (Fig. 10B).(35)

[0145] Lastly, we extracted cfDNA from a healthy control before and after vigorous exercise to examine the ability of the panel to measure tissue-specific changes in biological state. After capture and sequencing, we performed deconvolution of these two cfDNA samples. We found that cfDNA originating from neutrophils increased in the sample taken after exercise (Fig.10C), consistent with a recent report(36) studying the effect of exercise on cfDNA composition. Together, these experiments demonstrate that our approach fortargeting TIMs can correctly capture the methylation state of cfDNA and measure relevant tissue of origin effects.

[0146] cfDNA capture from ALS cases and controls

[0147] We next turned to examining the cfDNA epigenome of our disease cohorts. cfDNA was extracted from the blood plasma of cases and controls from both UQ and UCSF patients. We first confirmed our previous finding(10) of an increased concentration of cfDNA in the plasma of ALS patients relative to controls after correcting for age, sex, and SIRE (Fig. 4A) (logistic regression UQ: log odds ratio=7.5×10-3, p-value=1.8×10-2, UCSF: log odds ratio=2.4×10-2, p-value=6.0×10-3). Interestingly, cfDNA was also elevated in ALS patients relative to the OND controls (logistic regression log odds ratio=1.6×10-2, p-value=3.6×10-2), which had overall low levels of cfDNA. This suggests that the cfDNA generative processes of apoptosis and necrosis might differ between ALS and other types of neurological diseases.

[0148] After quantifying the amount of cfDNA, we performed high-throughput methylation sequencing on the captured regions. Since bisulfite treatment can degrade the already low quantity of input DNA, cfDNA sequencing experiments are prone to high duplication.(16) To address this, we used unique molecular identifiers (UMIs) to deduplicate reads. In total, after sequencing and deduplication, the average on-target coverage of UQ samples was 134 ± 166 reads per CpG and the average on-target coverage of UCSF samples was 195 ± 229 reads per CpG. The average methylation proportion at TIM sites was highly correlated between the two cohorts (Pearson’s R=0.98, p<1.0×10-16) (Fig.4C).

[0149] We noted that UCSF samples had a higher percentage of on-target reads (Fig.11), which likely contributed to differences in overall CpG read coverage. We also found that cfDNA starting concentration was a significant predictor of on target saturation after adjusting for total on target coverage (linear regression effect size=-1.7 × 10-3, p-value=9.0 × 10-3) (Fig.11D-11E).

[0150] Cell-type decomposition

[0151] Since TIMs were designed to be specific to a given tissue type, they can be used to estimate what tissues are contributing to the cfDNA in the context of neurodegeneration. To do this, we performed cfDNA cell-type decomposition with CelFiE.(10) CelFiE is a supervised decomposition algorithm that is designed to work with methylation read count data and missing or noisy reference data. As input, CelFiE takes the TIM read count data for each cfDNA sample and estimates the proportion of the cfDNA mixture originating from the tissues in the reference dataset, along with a specified number of unknown tissues.

[0152] We ran CelFiE with two unknown components using the methylation proportion of the captured sites as input (Fig.12A-12B). As with our prior ALS study, we observed elevated skeletal muscle in ALS patients in both cohorts relative to the healthy control samples (t-test p-value UCSF: 1.1×10-3, UQ: 4.7×10-2) (Fig.5D). This is consistent with muscle atrophy that occurs as part of their disease. We then tested the remaining tissue estimates for association with ALS disease status. At nominal significance (p<0.05) CelFiE estimated a depletion of cfDNA originating from eosinophils in ALS cases (t-test p-value UCSF: 1.0×10-2, UQ: 9.3×10-3)(Fig.12C-12D). While a preliminary result, previous studies have observed changes in granulocyte counts in the whole blood of ALS patients(37,38) and overall immune dysregulation is thought to be an important contributor to ALS etiology.(39)

[0153] Interestingly, we observed two UQ control samples with unusually high skeletal muscle components (an estimated 5.4% and 3.9% of their total cfDNA sample) (Fig.5D). One sample was an OND control with frontotemporal dementia, a disease that has substantial genetic and clinical overlap with ALS.(40) The other sample was originally classified as a healthy control. However, after further investigation into their clinical records, this individual had both a parent and sibling with ALS. Genetic testing revealed that this individual also tested positive for a C9orf72 repeat expansion, which is the most common genetic cause of ALS,(41) suggesting that the individual may be presymptomatic. Since the disease status of this patient was ambiguous, we reclassified them as OND.

[0154] Classification of ALS disease status

[0155] While muscle degeneration is a hallmark of ALS, it is not specific enough to serve as a diagnostic tool. Therefore, to further characterize the relationship between alterations in the cfDNA epigenome and disease, we developed a tissue-agnostic algorithm that utilized information from all TIM epigenetic profiles to predict whether a cfDNA sample was from an ALS patient or control. For these models, we did not consider PLS samples, but return to these samples below. Further models integrated all CpG sites, both on and off target.

[0156] We trained an elastic net prediction model(42) in four contexts to explore the generalizability of the results across the independent cohorts. First, a model was trained using ten-fold cross-validation within each cohort. Then, the transferability of the models was assessed by training a model on one cohort and applying it to the other. Since only the UQ cohort had OND and healthy controls, we combined the controls for this analysis, although we later examined the ability of the model to discriminate between the different sample types. Model parameters, including the elastic net mixing parameter, were selected by using a cross-model selection and averaging procedure within the training set.(42) Non-penalized covariates included age at the time of cfDNA sampling, sex, SIRE, cfDNA concentration, andtotal cfDNA input. We evaluated model performance with area under the receiver operating characteristic curve (AUC) and by testing whether the predictions could significantly predict true case-control status using a logistic regression model that included covariates.

[0157] To best characterize the different types of information that TIMs can provide we explored two classes of features for the prediction model, the methylation proportion and the coverage of the TIMs. Coverage was included because cfDNA fragmentation is non-random; we therefore reasoned that CpG coverage may also be informative of disease status. In total, we trained models using CpG coverage only, CpG methylation proportion only, and a combination of both as input features.

[0158] Overall, we found that tissue informative epigenetic features could significantly predict ALS case-control status in both cohorts (Fig.5, Fig.13-14, Table 4). The best- performing model incorporated both TIM coverage and methylation features (Fig.5). Within cohorts, the ten-fold cross-validated AUC was 0.82 within the UQ cohort (logistic regression odds ratio=2.34, p=2.32×10-7) and the UCSF AUC was 0.99 (logistic regression odds ratio=2.51, p-value<2.0×10-16). The models were more predictive than models trained using only covariate information (Fig.15). Importantly, even though the model was not trained to distinguish between ALS cases and OND, the AUC was high for both UQ models (within UQ: AUC=0.91, UCSF-UQ: AUC=0.76).

[0159] Table 4: Binary prediction model performance.

[0160] The AUC of predicting ALS vs all control samples for four models trained either within a cohort or trained in one cohort and tested on the remaining cohort. Models were trained with either only CpG coverage as input features, only CpG methylation, or both.

[0161] Models trained within one cohort replicated between cohorts. We noted that the prediction performance was higher for the UQ-trained and UCSF-tested model (AUC=0.91, logistic regression odds ratio=1.92, p=9.48 × 10-5) than the UCSF-trained model applied to the UQ samples (AUC=0.81, logistic regression odds ratio=2.46, p=4.24×10-4). Differences in model performance between cohorts were likely driven by a combination of factors, including cohort heterogeneity and technical variation. One likely contributing factor was the lower on-target coverage in the UQ cohort (Fig.4D, Fig.12). To test this, we randomly downsampled the number of reads in each UCSF cfDNA sample, which reduced effective on-target CpG coverage. Then, we re-ran the elastic net model within the UCSF cohort. We found that lower read coverage led to worse classification performance (Fig.16), suggesting that on-target CpG coverage is an important factor in prediction accuracy.

[0162] Importantly, the predictive performance of the elastic net models was stronger than using the CelFiE skeletal muscle estimate alone (Fig.4D). Indeed, models trained without any skeletal muscle TIMs, did not have reduced performance relative to the full model (Fig. 17), emphasizing the importance of combining information across tissue contributors.

[0163] We also noted that, despite methylation proportion being the more common feature considered in epigenetic cfDNA studies, the models trained only using CpG coverage also significantly predicted case-control status (Table 4, Fig.13). In fact, there was very similar performance within the UCSF cross-validated model (AUC=0.97) and the UQ cross- validated model (AUC=0.84). This suggests that disease-relevant information is contained in simply the observation of a given CpG in cfDNA sequencing data, providing an additional layer of information over the CpG methylation state alone. This information may be lost in other low-cost epigenetic assays, like methylation arrays, that only return methylation proportion values.

[0164] Lastly, we considered the elastic net models that incorporated off-target CpGs (UCSF total number of CpGs=32314, UQ total number of CpGs=49238). We found that the off- target models performed well (UCSF AUC=0.86, UQ AUC=0.76), even though this was a challenging setting as there were many more features than samples (Fig.18). Sites selected in these models could be chosen to refine TIM selection for capture panel development.

[0165] Biological significance of prediction features

[0166] Next, we sought to understand how different tissue informative sites contributed to predicting disease. An advantage of using a regularized regression model like an elastic net is that the model performs feature selection and assigns a higher weight, or absolute βvalue, to features that contribute more to accurately predicting the outcome. Features that do not contribute to the prediction will have an absolute β value near zero. Thus, to examine the overall contribution of different types of TIMs in making model predictions, we obtained the absolute β value for each TIM from an elastic net model trained on the entire UQ and UCSF cohorts (Fig.5C and Fig.5D). Then we examined how these values related to different characteristics of the TIMs.

[0167] We first analyzed whether TIMs selected for a given tissue type were more important in making predictions. As expected, skeletal muscle TIMs were highly important in making model predictions, especially for TIMs that were hypermethylated in skeletal muscle (Fig. 6A). Despite the importance of skeletal muscle TIMs, we noted that TIMs for every tissue type contributed to the model predictions (Fig.6). This again highlights the contribution of multiple tissues in neurodegeneration and the possibility of designing disease-specific biomarkers. For example, T-cell TIMs were highly important (Fig.6A), indicating that cfDNA originating from immune cell types may be relevant in ALS disease.

[0168] Overall, there were differences in the importance of each class of TIM. Hypermethylated TIMs generally had higher absolute β values than hypomethylated TIMs (Fig.6B-6E), which could be related to our previous observation that hypermethylated TIMs were more likely to be in promoter or genic regions (Fig.3). We also observed that there were differences in the distribution of absolute β values of methylation proportion and coverage features. For example, while methylation proportion features for fibroblast and epithelial cells had high absolute β values, coverage features for these tissues had relatively low absolute β values (Fig.6b-c). Instead, the coverage of TIMs for small intestine and T- cells were high, but close to zero as methylation proportion features. Together, this could mean that including both methylation proportion and coverage of tissue informative sites is useful for learning about disease in the context of cfDNA.

[0169] We next examined individual TIMs as an avenue for examining and generating hypotheses about individual epigenetic biomarker candidates. TIMs with a non-zero absolute β value were chosen for association with ALS case-control status, along with covariates and correcting for cohort. Multiple test correction was employed using false discovery rate at 10%. One of the most important methylation proportion features was a hypermethylated TIM selected for epithelium. We observed significantly increased methylation in ALS cases for this TIM (logistic regression odds ratio=14.09, q-value=8.06×10-2) suggesting that there was increased contribution from this gene in the cfDNA of ALS patients (Fig.6E). This TIM was located in the promoter region of the SHISA5 gene, which, along with p53, is involved in apoptosis.(43) Additionally, SHISA5 was found to be over-expressed in the spinal cord of ALS patients.(44)

[0170] We identified a similarly interesting hypermethylated TIM in the coverage features. While the TIM was selected for hepatocytes, it is in the XRCC6 gene, which was highly expressed in many tissues in bulk RNA-seq from the Genotype-Tissue Expression (GTEx) Project.(45) The TIM had significantly reduced coverage in ALS patients relative to controls (logistic regression odds ratio=-42.25, q-value=1.35 × 10-2) (Fig.6F), and while it is difficult to infer directly from cfDNA alone, this result could suggest potential dysregulation of this gene in cases. XRCC6 is involved in non-homologous end joining and DNA repair.(46) Disruption of non-homologous end joining has been previously linked to aging and ALS.(47,48)

[0171] ALS disease phenotypes

[0172] To further explore the value of tissue-specific methylation sites as a potential biomarker, we developed models to predict ALS disease phenotypes. To do this, we trained three linear elastic net models to predict ALSFRS-R (n=78), ALSFRS-R slope (n=60), and FVC (n=57) with ten-fold cross-validation. We hypothesized that high-weight features from the case-control analysis would also be associated with ALS phenotypes, and so, we chose the top 1000 coverage and top 1000 methylation features with the highest absolute β as input for the models. Since case-only numbers were relatively low in each cohort, we meta- analyzed the two cohorts, adding a non-penalized covariate for each cohort in the analysis, along with age, sex, cfDNA concentration, total cfDNA input quantity, and SIRE. To specifically evaluate the performance of cfDNA features over covariates, we separately trained an additional three models using only covariates.

[0173] We found the models based on cfDNA epigenetic features significantly predicted ALSFRS-R (Fig.7A) (Pearson’s R=0.66, p-value=3.71×10-9). This was statistically significantly higher (p-value=1.85×10-5) than the predictions from the model trained only on covariates (Pearson’s R=0.49, p-value=5.81×10-5). We found that the high predictive performance of the covariate-only model was largely attributed to cohort differences; within cohorts, the covariate model was not predictive (UQ: Pearson’s R=8.48×10-3p-value=0.95, UCSF: Pearson’s R=0.15, p-value=0.38) but epigenetics remained predictive (UQ: Pearson’s R=0.49, p-value=8.52×10-4UCSF: Pearson’s R=0.54 p-value=7.24×10-4).

[0174] We also found that the epigenetic models predicting FVC and ALSFRS-R slope were also significantly better than covariate-only models (FVC p-value=2.67×10-2, ALSFRS-R slope p-value=4.10×10-2), but more mild than the ALSFRS-R models (FVC Pearson’s R=0.50, p-value=3.78×10-3, ALSFRS-R slope Pearson’s R=0.28, p-value=2.81×10-2) (Fig. 7B-7C). Together, these results suggest that cfDNA epigenetic features are related to clinical traits used to measure ALS disease progression.

[0175] Lastly, we studied whether the same cfDNA epigenetic features that were associated with ALS disease phenotypes could differentiate between ALS and PLS cases. Due to the small sample number of PLS cases (n=15), we again combined the two cohorts and fit using 5-fold cross-validation with a non-penalized parameter for cohort. Although the analysis was underpowered, we observed a statistically significant difference between model predictions for ALS and PLS cases (AUC=0.74, linear regression effect size=36.61, p-value=1.9×10-2).

[0176] Discussion

[0177] Here, we presented a scalable cfDNA capture protocol that measures the methylation status of disease and tissue relevant CpG sites. We applied this capture technology to two independent cohorts of ALS patients and age-matched controls and examined the correlation with ALS disease status and progression. We then integrated both the read coverage and methylation proportion of the targeted sites in a machine-learning model. This model significantly discriminated between ALS patients and controls in two independent cohorts, including those with a variety of other neurological diseases. Together, our results suggest that a capture approach targeting tissue informative DNA methylation markers has value in quantitative biomarker development and that cfDNA has the potential to be a clinically relevant biomarker for ALS.

[0178] A key strength of using methylation markers informative of a broad variety of tissues is that it facilitates a comprehensive picture of a patient’s biological state and is not limited to a specific tissue or context. For example, neurofilament light chain is an exciting biomarker candidate for ALS.(49–51) However, neurofilament light chain also is elevated in other neurodegenerative diseases, which might limit its specificity for some applications.(52) By capturing and quantifying methylation levels at multiple tissue-informative CpG sites simultaneously, the panel has the potential to also learn about biological processes occurring in ALS outside of neurodegeneration. In particular, cfDNA is well-suited to measuring inflammation,(6,35) which has been of recent interest in ALS pathophysiology.(37,39) Future work could provide additional insight into how cfDNA relates to markers in ALS, providing a complementary avenue for investigation into disease mechanisms.

[0179] We observed differential performance between the UQ and UCSF cohorts. Specifically, the UCSF model outperformed the UQ model. Furthermore, the transferability was better when the UQ model was applied to the UCSF cohort. While it is likely a combination of factors, one explanation may be attributed to differences in sequencing depth. The UCSF cohort had higher on-target CpG coverage. Additional coverage may reduce noise, especially in analyses utilizing methylation proportion. In some cases, theoverall coverage is limited by the total amount of cfDNA available as input to the sequencing assay. This could be improved by recent high-throughput extraction technologies with the ability to increase cfDNA yield from a plasma sample.(53,54)

[0180] Model performance also may be affected by the slight differences in ALS patient characteristics between the cohorts. For example, the UCSF cohort had patients with lower ALSFRS-R scores and whose advanced condition may be easier to detect in cfDNA. ALS is also an extremely heterogeneous disease,(55) which can make designing biomarkers that generalize across patient populations difficult. It is also important to note that both cohorts were of majority European ancestry. Further exploration of how epigenetic cfDNA profiles differ between diverse subtypes of patients or change longitudinally as patients progress is now needed.

[0181] This Example only examined the performance of tissue informative markers in characterizing ALS. Since initiating these studies, other proposed blood-based biomarkers for ALS, like neurofilaments,(52,56,57) proteomics,(58) or miRNA(13), have demonstrated promise, and future studies can benchmark with at least one of these. Previous studies have also illustrated the benefit of combining different types of biomarkers to enhance predictive performance. Future work on cfDNA biomarker development in ALS could assay multiple biofluids simultaneously and include a range of cohorts (i.e. asymptomatic gene-positive carriers for diagnosis, multi-ancestry, neurological conditions presenting with weakness). Integration of these multiple measurements, along with information about existing patient genetic liability, would robustly test its potential context of use and may improve disease prediction models.

[0182] Lastly, there are numerous avenues for improving algorithms associated with the approach outlined here. While methylation capture arrays allow for a more cost-effective and focused analysis over relevant CpG sites, targeted capture also limits the coverage of the genome. This has the potential to miss important methylation changes occurring outside the targeted regions. Additionally, since we relied on published tissue methylation data sets that are low coverage and inherently noisy, TIM selection might be affected. Marker selection and overall algorithm performance might be improved by better, high-coverage reference data. Reference panel design for cfDNA applications is a robust area of current research, and incorporating new samples or biobanks into ALS disease prediction could be an area for future research. Finally, single-molecule(59) and nonlinear models(31) have shown recent promise in the analysis of cfDNA profiles.

[0183] Overall, the design of the cell-free DNA methylation capture panel and related prediction algorithms presented in this Example represents a significant advancement in thefield of ALS research. They demonstrate promising potential as a non-invasive and diagnostic tool for ALS, which could facilitate timely intervention and personalized treatment strategies. Further research and validation are necessary to refine the panel’s performance, assess its generalizability, and address practical considerations. Nonetheless, this study paves the way for the integration of DNA methylation biomarkers into the clinical management of ALS, bringing us closer to improved patient outcomes.

[0184] References

[0185] 1.^Baca, S.C., et al. (2023). Nat Med, 1–5.

[0186] 2.^Stackpole, M.L., et al. (2022). Nat Commun 13, 5566.

[0187] 3.^Liu, M.C., et al. (2020). Annals of Oncology 31, 745–759.

[0188] 4.^Zhang, J., et al. (2024). Nat Med, 1–10.

[0189] 5.^Lo, Y.M.D., et al. (2010). Science Translational Medicine 2, 61ra91–61ra91.

[0190] 6.^Cheng, A.P., et al. (2019). PNAS 116, 18738–18744.

[0191] 7.^Blauwkamp, T.A., et al. (2019). Nat Microbiol 4, 663–674.

[0192] 8.^De Borre, M., et al. (2023). at Med, 1–10.

[0193] 9.^Hansson, O. (2021). Nat Med 27, 954–963.

[0194] 10.^Caggiano, C., et al. (2021). Nat Commun 12, 2717.

[0195] 11.^Toden, S., et al. (2020). Science Advances 6, eabb1654.

[0196] 12. Lehmann-Werman, R., et al. (2016). PNAS 113, E1826–E1834.

[0197] 13.^Magen, I., et al. (2021). Nat Neurosci 24, 1534–1541.

[0198] 14.^Ziller, M.J., et al. (2015). Nat Methods 12, 230–232.

[0199] 15.^Hasegawa, K., et al. (2023). BMC Research Notes 16, 141.

[0200] 16.^Song, P., et al. (2022). Nat Biomed Eng 6, 232–245.

[0201] 17.^Ziller, M.J., et al. (2013). Nature 500, 477–481.

[0202] 18.^Morselli, M., et al. (2021). Methods 187, 13–27.

[0203] 19.^Fang, Q., et al. (2023). Clin Epigenet 15, 119.

[0204] 20.^Brooks, B.R., et al. (2000). Amyotrophic Lateral Sclerosis and Other Motor Neuron Disorders 1, 293–299.

[0205] 21.^Tartaglia, M.C., et al. (2007). Archives of Neurology 64, 232–236.

[0206] 22.^Gordon, P.H., et al. (2006). Neurology 66, 647–653.

[0207] 23.^Lomen-Hoerth, C., et al. (2002). Neurology 59, 1077–1079.

[0208] 24.^Cedarbaum, J.M., et al. (1999). J Neurol Sci 169, 13–21.

[0209] 25.^Consortium, T.E.P. (2012). Nature 489, 57–74.

[0210] 26.^Fernández, J.M., et al. (2016). Cell Syst 3, 491–495.e5.

[0211] 27.^Ziller, M.J., et al. (2011). PLOS Genetics 7, e1002389.

[0212] 28.^Moss, J., et al. (2018). Nat Commun 9, 1–12.

[0213] 29.^Sadeh, R., et al. (2021). Nature Biotechnology, 1–13.

[0214] 30.^Esfahani, M.S., et al. (2022). Nat Biotechnol, 1–13.

[0215] 31.^Zhou, Z., et al. (2023). PNAS 120, e2220982120.

[0216] 32.^Snyder, M.W., et al. (2016). Cell 164, 57–68.

[0217] 33.^Cristiano, S., et al. (2019). Nature 570, 385–389.

[0218] 34.Deaton, A.M., and Bird, A. (2011). Genes Dev 25, 1010–1022.

[0219] 35. Fox-Fisher, I., et al. (2021). eLife 10, e70520.

[0220] 36. Fridlich, O., et al. (2023). Cell Rep Med 4, 101074.

[0221] 37.^McCombe, P.A., and Henderson, R.D. (2011). Curr Mol Med 11, 246–254.

[0222] 38.^McCombe, P.A., et al. (2020). Frontiers in Neurology 11.

[0223] 39.^Hop, P.J., et al. (2022). Science Translational Medicine 14, eabj0264.

[0224] 40.^Ferrari, R., et al. (2011). Curr Alzheimer Res 8, 273–294.

[0225] 41.^Gijselinck, I., et al. (2018). The Genetics of C9orf72 Expansions. Cold Spring Harb Perspect Med 8, a026757.

[0226] 42.^Privé, F., et al. (2018). Bioinformatics 34, 2781–2787.

[0227] 43.^Bourdon, J.-C., et al. (2002). J Cell Biol 158, 235–246.

[0228] 44.^Andrés-Benito, P., et al. (2017). Aging (Albany NY) 9, 823–851.

[0229] 45.^Lonsdale, J., et al. (2013). Nat Genet 45, 580–585.

[0230] 46.^Wang, Z., et al. (2013). Autophagy 9, 925–927.

[0231] 47.^Cooper-Knock, J., et al. (2017). Front Mol Neurosci 10, 370.

[0232] 48.^Sama, R.R.K., et al. (2014). ASN Neuro 6, 1759091414544472.

[0233] 49.^Lu, C.-H., et al. (2015). Neurology 84, 2247–2257.

[0234] 50.^Zhou, Y., et al. (2021). Front Neurol 12, 712245.

[0235] 51.^Gaiani, A., et al. (2017). JAMA Neurology 74, 525–532.

[0236] 52.^Ashton, N.J., et al. (2021). Nat Commun 12, 3400.

[0237] 53.^Bettegowda, C., et al. (2014). Sci Transl Med 6, 224ra24.

[0238] 54.^Volik, S., et al. (2016). Cell-free DNA (cfDNA): Mol Cancer Res 14, 898–908.

[0239] 55.^Bendotti, C., et al. (2020). Amyotrophic Lateral Sclerosis and Frontotemporal Degeneration 21, 485–495.

[0240] 56.^Verde, F., et al. (2019). J Neurol Neurosurg Psychiatry 90, 157–164.

[0241] 57.^Bjornevik, K., et al. (2021). Neurology 97, e1466–e1474.

[0242] 58.^Hedl, T.J., et al. (2019). Front Neurosci 13, 548.

[0243] 59.^Unterman, I., et al. (2023). Multi-cell type deconvolution using a probabilistic model for single-molecule DNA methylation haplotypes (Bioinformatics)

[0244] 60.^Li, Y., and Tollefsbol, T.O. (2011). Methods Mol Biol 791, 11–21.

[0245] 61.^Smith, T.S., et al. (2017). Genome Res., gr.209601.116.

[0246] 62.^Farrell, C., et al. (2021). Gigascience 10, giab033.

[0247] 63.^Li, H., Handsaker, et al. (2009). Bioinformatics 25, 2078–2079.

[0248] 64.^Phung, T. (2024). SexChrLab / SexInference. (Sex Chromosome Lab).

[0249] 65.^Martens, J.H.A., and Stunnenberg, H.G. (2013). Haematologica 98, 1487–1489.

[0250] 66.^Privé, F., et al. (2019). Genetics 212, 65–74.

[0251] 67.^Heinz, S., et al. (2010). Mol Cell 38, 576–589. doi:10.1016 / j.molcel.2010.05.004.

[0252] 68.^Broad Institute Picard Tools.

[0253] Throughout this application various publications are referenced. The disclosures of these publications in their entireties are hereby incorporated by reference into this application in order to describe more fully the state of the art to which this invention pertains.

[0254] Those skilled in the art will appreciate that the conceptions and specific embodiments disclosed in the foregoing description may be readily utilized as a basis for modifying or designing other embodiments for carrying out the same purposes of the present invention. Those skilled in the art will also appreciate that such equivalent embodiments do not depart from the spirit and scope of the invention as set forth in the appended claims.

Claims

What is claimed is:

1. A method of detecting cell type specific degeneration in a biological sample obtained from a subject, the method comprising: (a) contacting extracellular DNA extracted from the biological sample with a set of capture probes, wherein the capture probes are less than 50,000 in number, and wherein the capture probes specifically bind within 100 base pairs (bp) of methylated and unmethylated CpG sites on cell-free DNA (cfDNA), and wherein the methylation status of CpG sites is specific to a cell type of interest; (b) performing methylation sequencing of cfDNA captured in step (a); (c) creating a methylation status dataset from the sequencing of step (b), wherein the methylation status dataset contains the methylation status of the cell type specific CpG sites (CTSCS); and (d) detecting cell type specific degeneration when the methylation status of the CTSCS is distinct from a reference methylation status dataset, wherein the reference methylation status dataset contains the methylation status for the CTSCS of all cell types.

2. The method of claim 1, wherein the methylation status dataset has been inputted into a statistical model, wherein the statistical model maps the methylation status dataset to the reference methylation status dataset.

3. The method of claim 2, wherein the CTSCS are identified using a supervised machine learning model and expectation maximization (EM) that estimates the proportion of the reference cell types for the subject and that estimates a contribution of unknown cell types not included in the reference methylation status dataset.

4. The method of claim 1, wherein the set of capture probes is between 20 and 20,000 in number.

5. The method of claim 1, wherein the cell type specific degeneration is amyotrophic lateral sclerosis (ALS).

6. The method of claim 2, wherein the statistical model comprises a penalized regression model.

7. The method of claim 1, wherein the cell type specific degeneration is traumatic brain injury (TBI), cancer, Alzheimer’s disease, Parkinson’s disease, or a pregnancy-related phenotype.

8. The method of claim 1, wherein the methylation sequencing comprises high throughput sequencing.

9. The method of claim 1, wherein the CTSCS are specific to skeletal muscle, fibroblasts, neurons, and / or hematopoietic cells.

10. A method of monitoring progression of a degenerative disease in a subject comprising performing the method of any one of claims 1-9 at a first time point on a biological sample obtained from the subject, and repeating the method at a subsequent time point, wherein an increase or decrease in the methylation status of the CTSCS is indicative of disease progression.

11. A computer implemented method of training a machine learning system to generate an estimator for identifying a cell type specific degeneration in a biological sample obtained from a subject, the method comprising: (a) storing a set of data comprising a plurality of prospective patient records from more than 100 patients, each prospective patient record including a plurality of parameters and corresponding values for each patient included in the patient records, and a diagnostic indicator indicating whether or not the patient included in the patient records has been diagnosed with a specific condition associated with a cell type specific degeneration; (b) selecting a subset of the plurality of parameters for inputs into the machine learning system, wherein the subset comprises CTSCS, the subset further including at least one clinical parameter selected from age, gender, and smoking status; (c) randomly partitioning the set of data into training data and validation data; (d) generating the estimator wherein the machine learning system is trained based on the training data and the subset of inputs; and wherein the estimator is trained with a supervised machine learning model and expectation maximization (EM) that estimates the proportion of the reference cell types for the subject and that estimates a contribution of unknown cell types not included in the reference methylation status dataset, whereby the machine learning system is trained to generate the estimator;wherein the estimator, when used with individual patient data, generates a composite algorithm value that is converted to a predictive score relative to a reference population.

12. The computer implemented method of claim 11, further comprising: (e) iteratively regenerating the estimator when the estimator does not meet a predetermined receiver operating curve (ROC) statistic, wherein the regenerating comprises selecting a different subset of inputs and / or adjusting associated weights of the inputs until the regenerated estimator meets a predetermined ROC statistic.

13. The computer implemented method of claim 11 or 12, further comprising: (f) generating a static configuration of the estimator when the machine learning system meets a predetermined ROC statistic.

14. The computer implemented method of claim 13, further comprising: (g) configuring a computing device accessible by a user with the static configuration of the estimator; (h) entering values for a subset of the plurality of parameters corresponding to the patient into the computing device; and (i) estimating, using the static estimator, the patient into a category indicative of a likelihood of having the cell-type specific degeneration or into another category indicative of a likelihood of not having the cell-type specific degeneration.

15. The computer implemented method of claim 11, further comprising incorporating test results from a diagnostic test that confirms or denies the presence of the cell-type specific degeneration into the training data.