Non-invasive detection of human diseases using cell-free DNA fragmentomes
By analyzing cfDNA fragment characteristics through low coverage whole genome sequencing, the method enables early and accurate detection of diseases like neurodegenerative diseases and chronic inflammation, facilitating timely treatment.
Patent Information
- Application Number
- PCT/US2025/018016
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-28
- Filing Date
- 2025-02-28
- Publication Date
- 2025-09-04
AI Technical Summary
Current methods for non-invasive detection of diseases such as neurodegenerative diseases, chronic inflammation, and non-neoplastic dysregulation of cellular growth are not sensitive enough to detect early changes in the fragmentome, epigenome, and transcriptome across the genome, limiting early diagnosis and treatment.
The method involves extracting cell-free DNA (cfDNA) from a subject's sample, conducting low coverage whole genome sequencing to generate genomic libraries, determining cfDNA fragmentome profiles by evaluating characteristics like fragment size, methylation frequencies, and comparing these profiles to reference profiles from healthy subjects to diagnose diseases, and treating the subject accordingly.
This approach allows for early and accurate detection of various diseases by analyzing cfDNA fragment characteristics, providing a basis for targeted treatment and improving patient outcomes.
Smart Images

Figure US2025018016_04092025_PF_FP_ABST
Abstract
Description
NON-INVASIVE DETECTION OF HUMAN DISEASES USING CELL-FREE DNA FRAGMENTOMES The present application claims the benefit of U.S. provisional application no.63 / 558,893 filed February 28, 2024, which is incorporated by reference herein in its entirety. STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH
[0001] This invention was made with government support under grants CA006973, CA121113, CA233259 and GM136577 awarded by the National Institutes of Health. The government has certain rights in the invention. FIELD
[0002] Non-invasive and ultrasensitive detection methods for analysis of single cell-free DNA (cfDNA) molecules to detect changes in the fragmentome, repeat landscape, epigenome, and transcriptome across the genome. In particular, these methods provide early detection of a disease condition allowing for treatment and / or improved patient-physician decision making early in the course of the disease. BACKGROUND
[0003] Cell-free DNA is found in the circulation after release from dying cells, and the size, distribution, and genomic content of cfDNA fragments reflect the underlying genomic, chromatin and epigenomic states of the cells from which they originate1,2,6,7. Assays that comprehensively analyze cfDNA fragments, or the cfDNA fragmentome, can capture the genome-wide structural, mutational, repeat-element, expression, and epigenetic changes that occur in cells across the body during the course of disease1–3,7,8. Through these approaches, such hallmarks of disease can be reliably detected with high performance in the plasma. SUMMARY
[0004] Plasma biomarkers based on cell-free DNA (cfDNA) molecules to detect changes in the fragmentation profiles across the genome are shown herein to specifically detect a range of clinical conditions including neurodegenerative disease, chronic inflammation, and non- neoplastic dysregulation of cellular growth. 1 168948691.1
[0005] Accordingly, embodiments are directed to detecting changes to the fragmentome across broad categories of disease phenotypes including neurodegeneration, chronic inflammation, and non-neoplastic dysregulation of cellular growth. Methods by which these changes can be measured and used in the early diagnosis of these illnesses are also provided.
[0006] In certain aspects, a method of early detection and treatment of diseases, comprises extracting cfDNA from a subject’s sample; isolating cell free DNA (cfDNA) from the subject’s biological samples; conducting low coverage whole genome sequencing of the subject’s isolated cfDNA to generate genomic libraries of sequenced cfDNA fragments; determining a cfDNA fragmentome profile of the subject’s sequenced DNA fragments comprising evaluating one or more cfDNA fragment characteristics; comparing the cfDNA fragmentome profiles of the subject to reference cfDNA fragmentome profiles from healthy subjects to detect and diagnose the disease; and, treating the subject with a disease specific therapy.
[0007] In certain embodiments, the one or more cfDNA fragment characteristics comprises calculating ratios of small to large cfDNA fragments, sequences of one or more cfDNAs, median cfDNA fragment sizes, fragment size distribution, mutant allele frequencies, fragment length, fragment size distribution, fragment end motifs, preferred end coordinates, breakpoint motifs, methylation frequencies, concentrations of cfDNA or combinations thereof. In certain embodiments, a small cfDNA fragment comprises about 80 base pairs (bp) to about 150 bp. In certain embodiments, a large cfDNA fragment comprises about 151 bp to about 300 bp.
[0008] In certain embodiments, the small to large cfDNA ratios are GC corrected. In certain embodiments, the cfDNA fragmentome profile comprises the sequence coverage of small cfDNA fragments in windows across the genome. In certain embodiments, the cfDNA fragmentome profile comprises the sequence coverage of cfDNA fragments in windows across the genome. In certain embodiments, the cfDNA fragmentome profile comprises the sequence coverage of small and large cfDNA fragments in windows across the genome.
[0009] In certain embodiments, the cfDNA fragmentome profiles in subjects identified and diagnosed early with a disease, are altered across the genome. In certain embodiments, the 2 168948691.1cfDNA fragmentome profiles in subjects identified and diagnosed early with a disease, have greater heterogeneity across the genome as compared to healthy subjects. In certain embodiments, the method further comprises identifying short nucleic acid sequences (kmers) in genomic or cell free DNA; selecting kmers occurring in a single repeat type and identifying unique kmers of repeat element types; wherein the unique kmers identify one or a plurality of repeat element types. In certain embodiments, genome wide differences in subjects identified and diagnosed early with a disease, have greater heterogeneity across the genome in repeat elements, long interspersed nuclear elements (LINEs), short interspersed nuclear elements (SINEs), long terminal repeats (LTRs), satellites, transposable elements, RNA elements or combinations thereof, as compared to healthy subjects.
[0010] In certain embodiments, the diseases comprise: neurodegenerative diseases, inflammatory diseases, cancer and non-neoplastic dysregulation of cellular growth. In certain embodiments, the cfDNA fragmentation profile for a specific disease is a biomarker profile diagnostic of a specific disease. In certain embodiments, the disease is Alzheimer’s disease.
[0011] In another aspect, a method of early detection and treatment of neurodegenerative diseases comprises extracting cfDNA from a subject’s sample; isolating cell free DNA (cfDNA) from the subject’s biological samples; conducting low coverage whole genome sequencing of the subject’s isolated cfDNA to generate genomic libraries of sequenced cfDNA fragments; determining a cfDNA fragmentome profile of the subject’s sequenced DNA fragments comprising evaluating one or more cfDNA fragment characteristics; comparing the cfDNA fragmentome profiles of the subject to reference cfDNA fragmentome profiles from healthy subjects to detect and diagnose the neurodegenerative disease; and, treating the subject.
[0012] In certain embodiments, the one or more cfDNA fragment characteristics comprises calculating ratios of small to large cfDNA fragments, sequences of one or more cfDNAs, median cfDNA fragment sizes, fragment size distribution, mutant allele frequencies, fragment length, fragment size distribution, fragment end motifs, preferred end coordinates, breakpoint motifs, methylation frequencies, concentrations of cfDNA or combinations thereof. In 3 168948691.1certain embodiments, concentrations of cfDNA are higher in subjects with a neurodegenerative disease as compared to as compared to healthy subjects.
[0013] In certain embodiments, the cfDNA fragmentome profiles in subjects identified and diagnosed early with a neurodegenerative disease, are altered across the genome. In certain embodiments, the cfDNA fragmentome profiles in subjects identified and diagnosed early with a disease, have greater heterogeneity across the genome as compared to healthy subjects.
[0014] In certain embodiments, the method further comprises identifying short nucleic acid sequences (kmers) in genomic or cell free DNA; selecting kmers occurring in a single repeat type and identifying unique kmers of repeat element types; wherein the unique kmers identify one or a plurality of repeat element types. In certain embodiments, genome wide differences in subjects identified and diagnosed early with a disease, have greater heterogeneity across the genome in repeat elements, long interspersed nuclear elements (LINEs), short interspersed nuclear elements (SINEs), long terminal repeats (LTRs), satellites, transposable elements, RNA elements or combinations thereof, as compared to healthy subjects.
[0015] In one embodiment, the neurodegenerative disease is selected from the group consisting of Alzheimer's disease, amyotrophic lateral sclerosis, Huntington's disease, and Parkinson's disease. In one embodiment, the neurodegenerative disease is Alzheimer's disease.
[0016] In another aspect, a method of early detection and treatment of inflammatory diseases, cancer or non-neoplastic dysregulation of cellular growth diseases comprises extracting cfDNA from a subject’s sample; isolating cell free DNA (cfDNA) from the subject’s biological samples; conducting low coverage whole genome sequencing of the subject’s isolated cfDNA to generate genomic libraries of sequenced cfDNA fragments; determining a cfDNA fragmentome profile of the subject’s sequenced DNA fragments comprising evaluating one or more cfDNA fragment characteristics; comparing the cfDNA fragmentome profiles of the subject to reference cfDNA fragmentome profiles from healthy subjects to detect and diagnose inflammatory diseases, cancer or non-neoplastic dysregulation of cellular growth diseases; and, treating the subject. 4 168948691.1
[0017] In additional aspects, methods are provided for detection and / or treatment of one or more liver diseases in a subject. In an aspect, methods are provided for detection and / or treatment of hepatitis in a subject. In an aspect, methods are provided for detection and / or treatment of non-alcoholic fatty liver disease in a subject. In an aspect, methods are provided for detection and / or treatment of metabolic associated steatotic liver disease in a subject. In an aspect, methods are provided for detection and / or treatment of steatosis in a subject. In a particular aspect, methods are provided for detection and / or treatment of fibrosis in a subject. In a particular aspect, methods are provided for detection and / or treatment of cirrhosis in a subject. In a particular aspect, methods are provided for detection and / or treatment of viral hepatitis in a subject.
[0018] Such methods suitably may comprise extracting cfDNA from a subject’s sample; isolating cell free DNA (cfDNA) from the subject’s biological samples; conducting low coverage whole genome sequencing of the subject’s isolated cfDNA to generate genomic libraries of sequenced cfDNA fragments; determining a cfDNA fragmentome profile of the subject’s sequenced DNA fragments comprising evaluating one or more cfDNA fragment characteristics; comparing the cfDNA fragmentome profiles of the subject to reference cfDNA fragmentome profiles from healthy subjects to detect and diagnose the disease or disorder such as a liver disease, hepatitis, non-alcoholic fatty liver disease, metabolic associated steatotic liver disease s or inflammatory diseases, cancer or non-neoplastic dysregulation of cellular growth diseases; and, treating the subject. A subject suitably may be identified as susceptible to susceptible to or known or suspected of suffering from one or more liver diseases, hepatitis including vial hepatitis, non-alcoholic fatty liver disease, metabolic associated steatotic liver disease, steatosis , fibrosis and / or cirrhosis.
[0019] In certain embodiments, the one or more cfDNA fragment characteristics comprises calculating ratios of small to large cfDNA fragments, sequences of one or more cfDNAs, median cfDNA fragment sizes, fragment size distribution, mutant allele frequencies, fragment length, fragment size distribution, fragment end motifs, preferred end coordinates, breakpoint motifs, methylation frequencies, concentrations of cfDNA, coverage of cfDNA fragments or combinations thereof. In certain embodiments, the cfDNA fragmentome profiles in 5 168948691.1subjects identified and diagnosed early with inflammatory diseases, cancer or non-neoplastic dysregulation of cellular growth diseases, are altered across the genome. In certain embodiments, the cfDNA fragmentome profiles in subjects identified and diagnosed early with inflammatory diseases, cancer or non-neoplastic dysregulation of cellular growth diseases, have greater heterogeneity across the genome as compared to healthy subjects.
[0020] In certain embodiments, the method further comprises identifying short nucleic acid sequences (kmers) in genomic or cell free DNA; selecting kmers occurring in a single repeat type and identifying unique kmers of repeat element types; wherein the unique kmers identify one or a plurality of repeat element types. In certain embodiments, genome wide differences in subjects identified and diagnosed early with inflammatory diseases, cancer or non-neoplastic dysregulation of cellular growth diseases, have greater heterogeneity across the genome in repeat elements, long interspersed nuclear elements (LINEs), short interspersed nuclear elements (SINEs), long terminal repeats (LTRs), satellites, transposable elements, RNA elements or combinations thereof, as compared to healthy subjects. In certain embodiments, the cfDNA fragmentation profile for a specific disease is a biomarker profile diagnostic of a specific disease, e.g., inflammatory diseases, cancer or non-neoplastic dysregulation of cellular growth diseases.
[0021] In certain embodiments, the method further comprises analysis of cfDNA fragment lengths and coverage at transcription factor binding sites or known genes, relative to genome-wide coverage. In certain embodiments, the method further comprises analysis of the amount and size of cfDNA fragments derived from mitochondrial DNA.
[0022] In additional aspects a method of determining morbidity in a subject suffering from a disease is provided and comprises extracting cfDNA from a biological sample obtained from the subject, conducting whole-genome sequencing of the subject’s isolated cfDNA, and generating fragmentation profiles comprising the size and coverage distribution of fragments in nonoverlapping regions, assessing the fragmentation profiles with respect to a morbidity index, and determining morbidity in the subject. 6 168948691.1
[0023] In certain preferred aspects, the whole-genome sequencing of the subject’s isolated cfDNA is conducted at an average of greater than 1.0 coverage such as at about 1.5x coverage.
[0024] In certain preferred aspects, fragmentation profiles comprising the size and coverage distribution of fragments are generated in nonoverlapping 5Mb regions.
[0025] In certain preferred aspects, the method may suitably further comprise, after extracting cfDNA from a biological sample obtained from the subject, constructing DNA libraries with size selection for cfDNA of nucleosomal origin, and thereafter conducting whole- genome sequencing of the subject’s isolated cfDNA.
[0026] In certain preferred aspects, the method may suitably further comprise, after conducting whole-genome sequencing of the subject’s isolated cfDNA, measuring cfDNA concentration and cfDNA lengths, and thereafter generating fragmentation profiles comprising the size and coverage distribution of fragments in nonoverlapping regions.
[0027] In certain preferred aspects, the morbidity index is a Charlson Comorbidity Index (CCI).
[0028] In certain embodiments, the cfDNA concentrations are measured by amount of extracted cfDNA in a nucleosomal length range excluding larger genomic DNA, wherein cfDNA concentrations are high as compared to healthy subjects.
[0029] In certain embodiments, high cfDNA concentrations correlated with a CCI 3+ or higher. In certain embodiments, shorter cfDNA fragments correlate with a high CCI (e.g.3+) as compared to those with low CCI (e.g.0-2).
[0030] In certain embodiments, a subject is identified as having a high concentration of cfDNA as compared to a healthy subject and / or has shorter cfDNA fragments as compared to a healthy individual is identified as having a higher morbidity.
[0031] In certain embodiments, the method further comprises generating a fragmentation morbidity index by utilizing a machine learning model using penalized logistic regression with 10-fold cross-validation to distinguish between subjects with high and low CCI from each other. In certain embodiments, the features utilized in the machine learning model include a fraction of 7 168948691.1short (from about 100-to about 150 bp) to long (from about 151- to about 220 bp) cfDNA fragments in each of the bins.
[0032] In certain embodiments, the fragmentation morbidity index is generated by z- scores representing chromosomal arm-level changes. In certain embodiments, the fragmentation morbidity index predicts survival. In certain embodiments, the fragmentation morbidity index predicts survival suitably after adjusting for one or more of for example one or more of clinical characteristics, such as one or more of age, inflammatory markers and CCI.
[0033] Definitions
[0034] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0035] As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, to the extent that the terms “including”, “includes”, “having”, “has”, “with”, or variants thereof are used in either the detailed description and / or the claims, such terms are intended to be inclusive in a manner similar to the term “comprising.” The term “about” or “approximately” means within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, “about” can mean within 1 or more than 1 standard deviation, per the practice in the art. Alternatively, “about” can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value or range. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude within 5-fold, and also within 2-fold, of a value. Where particular values are described in the application and claims, unless otherwise stated the term “about” meaning within an acceptable error range for the particular value should be assumed. 8 168948691.1
[0036] The term “agent” is used to describe a compound that has or may have a therapeutic or pharmacological activity. Agents include compounds that are known drugs, compounds for which therapeutic activity has been identified but which are undergoing further therapeutic evaluation, and compounds that are members of collections and libraries that are to be screened for a pharmacological activity.The terms “aligned”, “alignment”, “mapped” or “aligning”, “mapping” refer to one or more sequences that are identified as a match in terms of the order of their nucleic acid molecules to a known sequence from a reference genome. Such alignment can be done manually or by a computer algorithm, examples including the Efficient Local Alignment of Nucleotide Data (ELAND) computer program distributed as part of the Illumina Genomics Analysts pipeline. The matching of a sequence read in aligning can be a 100% sequence match or less than 100% (non-perfect match).
[0037] The term “cancer” as used herein is meant, a disease, condition, trait, genotype or phenotype characterized by unregulated cell growth or replication as is known in the art; including liver cancer (including hepatocellular carcinoma (HCC)), lung cancer (including non-small cell lung carcinoma), gastric cancer, colorectal cancer, as well as, for example, leukemias, e.g., acute myelogenous leukemia (AML), chronic myelogenous leukemia (CML), acute lymphocytic leukemia (ALL), and chronic lymphocytic leukemia, AIDS related cancers such as Kaposi's sarcoma; breast cancers; bone cancers such as Osteosarcoma, Chondrosarcomas, Ewing's sarcoma, Fibrosarcomas, Giant cell tumors, Adamantinomas, and Chordomas; Brain cancers such as Meningiomas, Glioblastomas, Lower- Grade Astrocytomas, Oligodendrocytomas, Pituitary Tumors, Schwannomas, and Metastatic brain cancers; cancers of the head and neck including various lymphomas such as mantle cell lymphoma, non-Hodgkins lymphoma, adenoma, squamous cell carcinoma, laryngeal carcinoma, gallbladder and bile duct cancers, cancers of the retina such as retinoblastoma, cancers of the esophagus, gastric cancers, multiple myeloma, ovarian cancer, uterine cancer, thyroid cancer, testicular cancer, endometrial cancer, melanoma, bladder cancer, prostate cancer, pancreatic cancer, sarcomas, Wilms' tumor, cervical cancer, head and neck cancer, skin cancers, nasopharyngeal carcinoma, liposarcoma, epithelial carcinoma, renal cell carcinoma, gallbladder adeno carcinoma, parotid adenocarcinoma, endometrial sarcoma, multidrug resistant 9 168948691.1cancers; and proliferative diseases and conditions, such as neovascularization associated with tumor angiogenesis.
[0038] The term “cell free nucleic acid,” “cell free DNA,” or “cfDNA” refers to nucleic acid fragments that circulate in an individual's body (e.g., bloodstream) and originate from one or more healthy cells and / or from one or more cancer cells. Additionally, cfDNA may come from other sources such as viruses, fetuses, etc.
[0039] The term “cfDNA sequence coverage” refers to the average number of cfDNA molecules overlapping a specific position.
[0040] The term “circulating tumor DNA” or “ctDNA” refers to nucleic acid fragments that originate from tumor cells or other types of cancer cells, which may be released into an individual's bloodstream as result of biological processes such as apoptosis or necrosis of dying cells or actively released by viable tumor cells.
[0041] The term “combination therapy”, as used herein, refers to those situations in which two or more different agents are administered in overlapping regimens so that the subject is simultaneously exposed to both agents. When used in combination therapy, two or more different agents may be administered simultaneously or separately. This administration in combination can include simultaneous administration of the two or more agents in the same dosage form, simultaneous administration in separate dosage forms, and separate administration. That is, two or more agents can be formulated together in the same dosage form and administered simultaneously. Alternatively, two or more agents can be simultaneously administered, wherein the agents are present in separate formulations. In another alternative, a first agent can be administered just followed by one or more additional agents. In the separate administration protocol, two or more agents may be administered a few minutes apart, or a few hours apart, or a few days apart.
[0042] The terms “determining”, “measuring”, “evaluating”, “detecting”, “assessing” and “assaying” are used interchangeably herein to refer to any form of measurement and include determining if an element is present or not. These terms include both quantitative and / or qualitative determinations. Assessing may be relative or absolute. “Assessing the presence of” 10 168948691.1includes determining the amount of something present, as well as determining whether it is present or absent. An “effective amount” as used herein, means an amount which provides a therapeutic or prophylactic benefit.
[0043] As used herein, the terms “fragmentation profile,” “fragmentome profile”, “position dependent differences in fragmentation patterns,” and “differences in fragment size and coverage in a position dependent manner across the genome” are equivalent and can be used interchangeably. In some embodiments, determining a cfDNA fragmentation profile in a mammal can be used for identifying a mammal as having cancer. For example, cfDNA fragments obtained from a mammal (e.g., from a sample obtained from a mammal) can be subjected to low coverage whole- genome sequencing, and the sequenced fragments can be mapped to the genome (e.g., in non- overlapping windows) and assessed to determine a cfDNA fragmentation profile. As described herein, a cfDNA fragmentation profile of a mammal having cancer is more heterogeneous (e.g., in fragment lengths) than a cfDNA fragmentation profile of a healthy mammal (e.g., a mammal not having cancer). As such, this disclosure also provides methods and materials for assessing, monitoring, and / or treating mammals (e.g., humans) having, or suspected of having, cancer. In some embodiments, this document provides methods and materials for identifying a mammal as having cancer. For example, a sample (e.g., a blood sample) obtained from a mammal can be assessed to determine the presence and, optionally, the tissue of origin of the cancer in the mammal based, at least in part, on the cfDNA fragmentation profile of the mammal. In some embodiments, methods and materials for monitoring a mammal as having cancer are provided. For example, a sample (e.g., a blood sample) obtained from a mammal can be assessed to determine the presence of the cancer in the mammal based, at least in part, on the cfDNA fragmentation profile of the mammal. In some embodiments, methods and materials for identifying a mammal as having cancer and administering one or more cancer treatments to the mammal to treat the mammal are provided. For example, a sample (e.g., a blood sample) obtained from a mammal can be assessed to determine if the mammal has cancer based, at least in part, on the cfDNA fragmentation profile of the mammal, and one or more cancer treatments can be administered to the mammal.
[0044] The term “genomic nucleic acid,” or “genomic DNA,” refers to nucleic acid including chromosomal DNA that originates from one or more healthy (e.g., non-tumor) cells or 11 168948691.1tumor cells. In various embodiments, genomic DNA can be extracted from a cell derived from a blood cell lineage, such as a white blood cell (WBC).
[0045] As used herein, “neurodegenerative diseases” comprise AIDS dementia complex, Alzheimer's disease, amyotrophic lateral sclerosis, adrenoleukodystrophy, Alexander disease, Alper's disease, ataxia telangiectasia, Batten disease, bovine spongiform encephalopathy (BSE), Canavan disease, corticobasal degeneration, Creutzfeldt-Jakob disease, dementia with Lewy bodies, fatal familial insomnia, frontotemporal lobar degeneration, Huntington's disease, Kennedy's disease, Krabbe disease, Lyme disease, Machado-Joseph disease, multiple sclerosis, multiple system atrophy, neuroacanthocytosis, Niemann-Pick disease, Parkinson's disease, Pick's disease, primary lateral sclerosis, progressive supranuclear palsy, Refsum disease, Sandhoff disease, diffuse myelinoclastic sclerosis, spinocerebellar ataxia, subacute combined degeneration of spinal cord, tabes dorsalis, Tay-Sachs disease, toxic encephalopathy, transmissible spongiform encephalopathy, and wobbly hedgehog syndrome.
[0046] A “neurological disorder” as used herein refers to a disease or disorder which affects the CNS and / or which has an etiology in the CNS. Exemplary CNS diseases or disorders include, but are not limited to, neuropathy, amyloidosis, cancer, an ocular disease or disorder, viral or microbial infection, inflammation, ischemia, neurodegenerative disease, seizure, behavioral disorders, and a lysosomal storage disease. For the purposes of this application, the CNS will be understood to include the eye, which is normally sequestered from the rest of the body by the blood-retina barrier. Specific examples of neurological disorders include, but are not limited to, neurodegenerative diseases (including, but not limited to, Lewy body disease, postpoliomyelitis syndrome, Shy-Draeger syndrome, olivopontocerebellar atrophy, Parkinson's disease, multiple system atrophy, striatonigral degeneration, tauopathies (including, but not limited to, Alzheimer disease and supranuclear palsy), prion diseases (including, but not limited to, bovine spongiform encephalopathy, scrapie, Creutzfeldt-Jakob syndrome, kuru, Gerstmann- Straussler-Scheinker disease, chronic wasting disease, and fatal familial insomnia), bulbar palsy, motor neuron disease, and nervous system heterodegenerative disorders (including, but not limited to, Canavan disease, Huntington's disease, neuronal ceroid-lipofuscinosis, Alexander's disease, Tourette's syndrome, Menkes kinky hair syndrome, Cockayne syndrome, Halervorden- 12 168948691.1Spatz syndrome, lafora disease, Rett syndrome, hepatolenticular degeneration, Lesch-Nyhan syndrome, and Unverricht-Lundborg syndrome), dementia (including, but not limited to, Pick's disease, and spinocerebellar ataxia), cancer (e.g. of the CNS and / or brain, including brain metastases resulting from cancer elsewhere in the body).
[0047] As used herein, “inflammatory diseases” and “dysregulation of cellular growth” refers to disorders that cause or are caused by inflammation and lead to changes in clinical outcomes. Exemplary inflammatory diseases or disorders include, but are not limited to include disorders of the liver, ovary, intestines, endometrium or nearby tissue. Specific examples of dysregulation of cellular growth include but are not limited to Liver Cirrhosis, Liver Fibrosis, NAFLD, NASH, Benign Adnexal Masses, Functional Ovarian Cysts, Endometriosis and other conditions related to proliferation of abnormal cells, formation of scar tissue, or otherwise aberrant growth.
[0048] “Optional” or “optionally” means that the subsequently described event or circumstance can or cannot occur, and that the description includes instances where the event or circumstance occurs and instances where it does not.
[0049] As used in this specification and the appended claims, the term “or” is generally employed in its sense including “and / or” unless the content clearly dictates otherwise.
[0050] The phrase “pharmaceutically acceptable carrier” refers to a carrier for the administration of a therapeutic agent. Exemplary carriers include saline, buffered saline, dextrose, water, glycerol, ethanol, and combinations thereof. For drugs administered orally, pharmaceutically acceptable carriers include, but are not limited to pharmaceutically acceptable excipients such as inert diluents, disintegrating agents, binding agents, lubricating agents, sweetening agents, flavoring agents, coloring agents and preservatives. Suitable inert diluents include sodium and calcium carbonate, sodium and calcium phosphate, and lactose, while corn starch and alginic acid are suitable disintegrating agents. Binding agents may include starch and gelatin, while the lubricating agent, if present, will generally be magnesium stearate, stearic acid or talc. If desired, the tablets may be coated with a material such as glyceryl monostearate or glyceryl distearate, to delay absorption in the gastrointestinal tract. 13 168948691.1
[0051] As used in this specification and the appended claims, the term “or” is generally employed in its sense including “and / or” unless the content clearly dictates otherwise.
[0052] “Parenteral” administration of an immunogenic composition includes, e.g., subcutaneous (s.c.), intravenous (i.v.), intramuscular (i.m.), or intrasternal injection, or infusion techniques.
[0053] The terms “patient” or “individual” or “subject” are used interchangeably herein, and refers to a mammalian subject to be treated, with human patients being preferred. In some embodiments, the methods of the invention find use in experimental animals, in veterinary application, and in the development of animal models for disease, including, but not limited to, rodents including mice, rats, and hamsters, and primates.
[0054] The term “reference genome” as used herein may refer to a digital or previously identified nucleic acid sequence database, assembled as a representative example of a species or subject. Reference genomes may be assembled from the nucleic acid sequences from multiple subjects, sample or organisms and does not necessarily represent the nucleic acid makeup of a single person. Reference genomes may be used to for mapping of sequencing reads from a sample to chromosomal positions. For example, a reference genome used for human subjects as well as many other organisms is found at the National Center for Biotechnology Information at ncbi.nlm.nih.gov.
[0055] The term “read segment” or “read” refers to any nucleotide sequences including sequence reads obtained from an individual and / or nucleotide sequences derived from the initial sequence read from a sample obtained from an individual.
[0056] The terms “sample,” “patient sample,” “biological sample,” and the like, encompass a variety of sample types obtained from a patient, individual, or subject and can be used in a diagnostic, prognostic and / or monitoring assay. The patient sample may be obtained from a healthy subject, a diseased patient, or a patient with lung cancer. In certain embodiments, a sample that is “provided” can be obtained by the person (or machine) conducting the assay, or it can have been obtained by another, and transferred to the person (or machine) carrying out the assay. Moreover, a sample obtained from a patient can be divided and only a portion may be used 14 168948691.1for diagnosis. Further, the sample, or a portion thereof, can be stored under conditions to maintain sample for later analysis. The definition specifically encompasses blood and other liquid samples of biological origin (including, but not limited to, peripheral blood, serum, plasma, cord blood, amniotic fluid, cerebrospinal fluid, urine, saliva, stool and synovial fluid), solid tissue samples such as a biopsy specimen or tissue cultures or cells derived therefrom and the progeny thereof. In certain embodiment, a sample comprises cerebrospinal fluid. In a specific embodiment, a sample comprises a blood sample. In another embodiment, a sample comprises a plasma sample. In yet another embodiment, a serum sample is used. The definition of “sample” also includes samples that have been manipulated in any way after their procurement, such as by centrifugation, filtration, precipitation, dialysis, chromatography, treatment with reagents, washed, or enriched for certain cell populations. The terms further encompass a clinical sample, and also include cells in culture, cell supernatants, tissue samples, organs, and the like. Samples may also comprise fresh-frozen and / or formalin-fixed, paraffin-embedded tissue blocks, such as blocks prepared from clinical or pathological biopsies, prepared for pathological analysis or study by immunohistochemistry.
[0057] The term “sequence reads” refers to nucleotide sequences read from a sample obtained from an individual. Sequence reads can be obtained through various methods known in the art.
[0058] As defined herein, a “therapeutically effective” amount of a compound or agent (i.e., an effective dosage) means an amount sufficient to produce a therapeutically (e.g., clinically) desirable result. The compositions can be administered from one or more times per day to one or more times per week, including once every other day. The skilled artisan will appreciate that certain factors can influence the dosage and timing required to effectively treat a subject, including but not limited to the severity of the disease or disorder, previous treatments, the general health and / or age of the subject, and other diseases present. Moreover, treatment of a subject with a therapeutically effective amount of the compounds of the invention can include a single treatment or a series of treatments.
[0059] As used herein, the terms “treat,” treating,” “treatment,” and the like refer to reducing or ameliorating a disorder and / or symptoms associated therewith. It will be appreciated 15 168948691.1that, although not precluded, treating a disorder or condition does not require that the disorder, condition or symptoms associated therewith be completely eliminated.
[0060] Genes: All genes, gene names, and gene products disclosed herein are intended to correspond to homologs from any species for which the compositions and methods disclosed herein are applicable. It is understood that when a gene or gene product from a particular species is disclosed, this disclosure is intended to be exemplary only, and is not to be interpreted as a limitation unless the context in which it appears clearly indicates. Thus, for example, for the genes or gene products disclosed herein, are intended to encompass homologous and / or orthologous genes and gene products from other species.
[0061] Ranges: throughout this disclosure, various aspects of the invention can be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 2.7, 3, 4, 5, 5.3, and 6. This applies regardless of the breadth of the range.
[0062] Any compositions or methods provided herein can be combined with one or more of any of the other compositions and methods provided herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawings will be provided by the Office upon request and payment of the necessary fee.
[0064] FIG.1 shows a schematic of study design and cohorts analyzed in the example.
[0065] FIGS.2A-2B demonstrate that cfDNA fragmentomes are altered in disease states. cfDNA fragmentation profiles (FIG. 2A) showed heterogeneity among patients with disease as 16 168948691.1compared to healthy individuals. The 5Mb coverage bins with the highest variance between samples differed between healthy and disease states (FIG. 2B) providing evidence that these fragmentomic alterations reflected underlying pathology that differed by disease state.
[0066] FIGS. 3A-3F are a series of graphs demonstrating the high performance for noninvasive detection of human diseases using cfDNA fragmentomes. The performance for detecting Alzheimer’s Disease (FIG. 3A), liver cirrhosis (FIG. 3B), and benign adnexal masses (FIG. 3C) are indicated in receiver operator curves (ROCs) using cross-validated ARTEMIS, DELFI or ARTEMIS-DELFI fragmentomic approaches for disease detection. Locked models had high performance for detecting only the disease type for which they were trained (FIGS.3D-3F), suggesting that cfDNA fragmentation reflected disease-specific changes in the circulation.
[0067] FIGS.4A-4D demonstrate that fragmentome changes in plasma from patients with Alzheimer’s Disease reflect disease-specific mechanisms. (FIG.4A) The increased percentage of mitochondrial-derived cfDNA fragments in patients with Alzheimer’s Disease may reflect stress- response, dysfunction and damage that occurs in the mitochondria resulting in release of DNA that can be assessed in the plasma. (FIG. 4B) Gene Set Enrichment Analysis of cfDNA fragment coverage across all known genes in patients with Alzheimer’s disease compared to in healthy patients revealed significant enrichment of two pathways related to Blood Brain Barrier integrity. (FIG.4C) The genes enriched in these pathways have numerous known protein interactions with genes in the canonical KEGG Alzheimer’s gene set, suggesting possible mechanisms by which these genes are implicated in Alzheimer’s disease. (FIG.4D) Top ten repeat elements with largest alterations to plasma representation between patients with and without Alzheimer’s disease suggested that increases in transposable element representation noted in Alzheimer’s brain tissue may be detectable in the plasma.
[0068] FIGS.5A-5C demonstrate that fragmentome changes in plasma from patients with and without liver cirrhosis suggest both tissue-specific and peripheral immune changes. (FIG.5A) Genes for which plasma cfDNA fragment coverage separated patients with and without cirrhosis with AUC > 0.75 included five key genes implicated in liver cirrhosis, and revealed protein interactions that may suggest mechanisms by which these alterations observed in the plasma are 17 168948691.1related to cirrhotic changes in the liver. (FIG. 5B) Alterations to plasma representation between patients with and without liver cirrhosis suggested that genomic and epigenomic changes to repeat elements may be present in liver cirrhosis and assessable in the plasma. (FIG.5C) Regions of the genome with activating histone marks in T-lymphocytes and repressive histone marks in B- lymphocytes, monocytes or neutrophils, had systematically higher coverage in the plasma of patients with cirrhosis than without, suggesting that cirrhosis related inflammatory changes that increase the representation of white blood cells other than T-lymphocytes can be assessed in the plasma (top row). Bottom row, no significant coverage differences were observed in regions of the genome repressed in T-lymphocytes as well as B-lymphocytes, monocytes or neutrophils suggesting that observed coverage changes were linked to cell count changes.
[0069] FIG. 6 is a series of plots showing the plasma cfDNA concentrations (ng / mL) in healthy individuals and disease patients.
[0070] FIG. 7 demonstrates that cfDNA kmer repeat landscapes showed heterogeneity among patients with disease as compared to healthy individuals.
[0071] FIG.8 demonstrates that the cfDNA fragmentome machine learning models were different among disease states. Relative contributions of cfDNA features to machine learning models differed by disease providing evidence that cfDNA fragmentation reflected disease- specific changes in the circulation.
[0072] FIG. 9 demonstrates differences in average cfDNA coverage in 100 Kb bins overlapping four canonical Alzheimer’s Disease related genes (APP, APOE, PSEN1 and MAPT) as compared to 50 random 100kb bins. Analyses of cfDNA in patients with Alzheimer’s Disease identified coverage differences at three of these loci, consistent with epigenetic and chromatin changes surrounding these genes.
[0073] FIGS.10A-10F. cfDNA fragmentomes as molecular biomarkers of disease. (FIG. 10A) Schematic of potential uses for fragmentomic liquid biopsies as disease-specific biomarkers. (FIG. 10B) Elevated cfDNA concentrations were observed in individuals with high Charlson Comorbidity Index (CCI) and various diseases. p-values indicate the Benjamini-Hochberg corrected p-value from Wilcoxon’s rank-sum test as compared to the low CCI group. (FIG.10C) 18 168948691.1Individuals with high CCI had an increase in short fragments less than ~170bp reflecting an increase in sub-mononucleosomal fragments. The line shows the mean difference in fragment length distribution, and shading indicates + / - 1SD. (FIG.10D) The cross-validated Fragmentation Morbidity Index (FMI) stratified patients by CCI. (FIG.10E) Individuals with FMI scores in the top decile had significantly shorter overall survival than individuals in the bottom decile. (FIG. 10F) In a Cox Proportional Hazards (CPH) multivariate analysis, FMI was an independent predictor of survival after adjustment for age, CCI and inflammatory markers. All inputs to the CPH model were centered and scaled to a mean of zero and a standard deviation of one prior to estimating hazard ratios.
[0074] FIGS. 11A-11F. |Fragmentomic liquid biopsy detects liver cirrhosis in the Discovery Cohort. (FIG.11A) Elevated cfDNA concentrations were observed in individuals with liver cirrhosis (n=66). p-value indicates the Benjamini-Hochberg corrected p-value from Wilcoxon’s rank-sum test as compared to the screening population (n=348). (FIG. 11B) Individuals with liver cirrhosis had an increase in short fragments less than ~170bp reflecting an increase in sub-mononucleosomal fragments. Each line shows the mean difference in fragment length distribution, and shading indicates + / - 1SD. (FIG.11C) Individuals in the presumed healthy population (n=348) had highly homogenous fragmentomes, while individuals with liver cirrhosis or advanced fibrosis (FIG.11D) showed patterns of aberration. Feature-level metrics are plotted in gray, with black and blue coloring used to indicate high variability in screening population or liver cirrhosis individuals respectively. In FIGS.11C and 11D the outer ring represents the ratio of short (<=150 bp) to long (>150 bp) fragments in 5Mb bins genome wide for each individual; lines are colored if Pearson’s correlation between the individual’s fragmentation profile and the median general screening population individual in the Discovery Cohort is less than 0.9. The second ring from the outside represents the percentage of individuals with a repeat element count outside the 95% confidence interval of all Discovery Cohort Individuals; elements are colored if this percentage is greater than 10%. The third ring from the outside represents the F-statistic comparing variation in 1Mb bins of high density of epigenetic marks for each group of individuals as compared to the individuals in the general screening population subset of the Discovery Cohort; Significant F-statistics (p<0.05) are colored. In FIG. 11D the center most ring represents total 19 168948691.1weighted feature importance for each class of features within the penalized logistic regression model for liver cirrhosis and advanced fibrosis detection. (FIG. 11E) The cross-validated ARTEMIS-DELFI LCr score generated from the classifier was predicted to be high in individuals with liver cirrhosis or Advanced Fibrosis as compared to those in the screening population and could be separated with high area under the receiver operating curve (AUC-ROC) (FIG.11F).
[0075] FIGS. 12A-12E. Locked fragmentome classifier demonstrates liver cirrhosis detection with limited cross-reactivity for other diseases in the Validation Cohort. The locked ARTEMIS-DELFI LCr detection classifier predicted high scores in the Validation Cohort (FIG. 12A) that separated individuals with liver cirrhosis and advanced fibrosis from those in screening populations with high performance. (FIG. 12B) At a locked ARTEMIS-DELFI LCr score threshold corresponding to 90% specificity in the Discovery Cohort, sensitivity and specificity remained high in the Validation Cohort demonstrating the generalizability of the classifier. (FIG. 12C) Comparison of performance of ARTEMIS-DELFI LCr score to Fibrosis-4 (FIB-4) index in Validation Cohort for a subset of individuals with liver cirrhosis where the FIB-4 metric is available (n=42). Gray lines show the clinically validated FIB-4 cut-off and the ARTEMIS-DELFI LCr score threshold from (B). (FIG.12D) Among individuals with indeterminate FIB-4 levels, the ARTEMIS-DELFI LCr score improves classification. The dotted gray line shows the 50% moderate specificity threshold used in FIG.22, which may be appropriate for use in an elevated- risk population. (FIG. 12E) The locked ARTEMIS-DELFI LCr classifier generated low scores when applied to individuals with other non-liver diseases, demonstrating the limited cross- reactivity of the approach. For comparison, the dotted gray line shows the 90% specificity threshold used in (FIG.12B).
[0076] FIGS.13A-13C. Fragmentomic changes in plasma reflect liver cirrhosis specific disease mechanisms. (FIG.13A) Coverage differences in the cfDNA across the gene body of all known genes in individuals with liver cirrhosis or Advanced Fibrosis as compared to screening population individuals reveal several genes with known roles in liver cirrhosis that are altered in both the liver cirrhosis Discovery and Validation Cohorts. When the same analysis is performed comparing the screening population individuals from the discovery cohort split into two random groups, there are no genes with significant coverage changes. Genes with significant differences 20 168948691.1(p<.05 after Bonferroni correction, threshold indicated by black horizontal line) are shown with darker shading, and blue indicates genes in the known liver cirrhosis gene set from the Monarch database. (FIG.13B) Regions of the genome with activating histone marks in T-lymphocytes and repressive histone marks in monocytes or neutrophils had systematically higher coverage in the plasma of patients with cirrhosis or advanced fibrosis than without, suggesting that cirrhosis related inflammatory changes that increase the representation of white blood cells other than T- lymphocytes can be assessed in the plasma. (FIG.13C) Deconvolution analysis of plasma cfDNA fragmentomes: for all tissues in the GTEx database, the Pearson R correlation was calculated between the effect size of relative coverage in transcription factor binding sites (TFBS) of patients with liver cirrhosis or advanced fibrosis vs the screening population and the effect size of transcription factor (TF) RNA expression of a given tissue vs whole blood. Due to the inverse relationship between cfDNA coverage in the circulation and gene expression in tissue, a negative correlation suggests a strong cfDNA contribution from that tissue. The median correlations for the top decile of number of TFBS were ranked with the most negative (strongest correlation) occurring first.
[0077] FIGS.14A-14B. Overview of MICA Morbidity Cohort. (FIG.14A) Individuals in the MICA Morbidity cohort were split into high and low Charlson Comorbidity Index groups and used for cfDNA fragmentation analyses and cross-validation of the Fragmentomic Morbidity Index. (FIG.14B) The samples from these individuals were prospectively collected as part of the Prospective Biomarker study MICA at Herlev and Gentofte Hospital, Cophenhagen University Hospital, Denmark between July 2016 and July 2019.
[0078] FIGS. 15A-15D. cfDNA fragmentomes can serve as molecular biomarkers of morbidity. (FIG.15A) Individuals in the MICA Morbidity Cohort with high Charlson Comorbidity Index (CCI) (CCI 3+) had higher levels of protein biomarkers associated with inflammation as compared to individuals with low CCI (CCI 0,1,2). (FIG.15B) Protein biomarkers of inflammation correlated with increased cfDNA concentrations. (FIG.15C) Individuals with high CCI had shorter average fragment lengths than individuals with low CCI, and the trend was maintained for subgroups of high CCI individuals with specific morbidities (Benjamini-Hochberg adjusted p- 21 168948691.1values are shown). (FIG. 15D) Individuals with FMI values above the mean had shorter overall survival than individuals with FMI below the mean.
[0079] FIG.16. Overview of liver cirrhosis Discovery and Validation Cohorts. Schematic of study design and cohorts analyzed.
[0080] FIG.17. Schematic of model architecture used in ARTEMIS-DELFI LCr detection model. The ARTEMIS-DELFI LCr score was generated as an ensemble of two components – a fragmentation score and a repeat element score. The repeat element score was itself an ensemble of scores from individual models corresponding to 6 feature families.
[0081] FIG. 18 ARTEMIS-DELFI LCr scores were not influenced by WGS Library Batch. For each type of individual and library batch, the median ARTEMIS-DELFI LCr score is plotted, demonstrating that disease scores were consistent across batches. Library batches were composed of individuals with and without disease, and Discovery and Validation cohorts were processed separately.
[0082] FIGS.19A-19B. Relationship between ARTEMIS-DELFI LCr scores and clinical covariates in the Discovery Cohort. ARTEMIS-DELFI LCr scores for (FIG. 19A) screening population individuals and (FIG.19B) individuals with liver cirrhosis in the Discovery Cohort as compared to various clinically relevant covariates.
[0083] FIGS.20A-20D. Disease specific changes in cfDNA fragmentomes of individuals with liver cirrhosis in the Validation Cohorts. (FIG. 20A) Elevated cfDNA concentrations were observed in individuals with liver cirrhosis (n=83). p-value indicates the Benjamini-Hochberg corrected p-value from Wilcoxon’s rank-sum test as compared to the screening population (n=121). (FIG.20B) Individuals with liver cirrhosis had an increase in short fragments less than ~170bp reflecting an increase in sub-mononucleosomal fragments. Each line shows the mean difference in fragment length distribution between these patient groups and the general screening population individuals in the Discovery Cohort, and shading indicates + / - 1SD. (FIG. 20C) Individuals in the screening population (n=121) had highly homogenous fragmentomes, while individuals with liver cirrhosis or advanced fibrosis (FIG. 20D) showed patterns of aberration. Feature-level metrics are plotted in gray, with black and blue coloring used to indicate high 22 168948691.1variability in screening population or liver cirrhosis individuals respectively. In FIGS. 20C and 20D the outer ring represents the ratio of short (<=150 bp) to long (>150 bp) fragments in 5Mb bins genome wide for each individual; lines are colored if Pearson’s correlation between the individual’s fragmentation profile and the median general screening population individual in the Discovery Cohort is less than 0.9. The second ring from the outside represents the percentage of individuals with a repeat element count outside the 95% confidence interval of all Discovery Cohort Individuals; elements are colored if this percentage is greater than 10%. The third ring from the outside represents the F-statistic comparing variation in 1Mb bins of high density of epigenetic marks for each group of individuals as compared to the individuals in the general screening population subset of the Discovery Cohort; Significant F-statistics (p<0.05) are colored. In FIG. 20D the center most ring represents total weighted feature importance for each class of features within the penalized logistic regression model for liver cirrhosis and advanced fibrosis detection.
[0084] FIGS.21A-21B. Relationship between ARTEMIS-DELFI LCr scores and clinical covariates in the Validation Cohort. ARTEMIS-DELFI LCr scores for (FIG. 21A) screening population individuals and (FIG.21B) individuals with liver cirrhosis in the validation cohort as compared to various clinically relevant covariates.
[0085] FIG. 22. Locked fragmentome classifier evaluated at moderate specificity threshold. At a locked ARTEMIS-DELFI LCr score threshold corresponding to 50% specificity in the Discovery Cohort, sensitivity remained high in the Validation Cohort demonstrating the generalizability of the classifier and its potential utility in a high-risk population.
[0086] FIGS. 23A-23B. Clinical utility of ARTEMIS-DELFI LCr Approach for liver cirrhosis Detection. Standard of Care (FIG.23A) and proposed (FIG.23B) clinical workflows for the detection of liver cirrhosis.
[0087] FIGS. 24A024B. Application of locked cfDNA lung or liver cancer detection classifiers to liver cirrhosis cohorts showed limited cross-reactivity. The locked cfDNA lung (FIG. 24A) or liver (FIG.24B) cancer models predicted low scores for individuals in the liver cirrhosis cohorts. 23 168948691.1
[0088] FIG. 25. Plasma coverage in patients with and without liver cirrhosis remains unchanged in regions of the genome without epigenetic difference between white blood cell types. Regions of the genome with repressive histone marks in both T-lymphocytes and monocytes or neutrophils showed no coverage differences in the plasma of patients with cirrhosis than without, highlighting that changes shown in FIG.13B may be due to immune-mediated changes to WBC populations in cirrhosis.
[0089] FIGS.26A-26E. Disease-specific changes in cfDNA fragmentomes of individuals with Alzheimer’s Disease (AD) and Liver Cirrhosis (LCr) in Discovery Cohorts. (FIG. 26A) Elevated cfDNA concentrations were observed in individuals with AD (n=71) or LCr (n=49). p- values indicate the Benjamini-Hochberg corrected p-value from Wilcoxon’s rank-sum test as compared to the healthy population (n=390). (FIG. 26B) Individuals with AD and LCr had an increase in short fragments less than ~170bp reflecting an increase in sub-mononucleosomal fragments. Each line shows the mean difference in fragment length distribution, and shading indicates + / - 1SD. (FIG.26C) Individuals in the presumed healthy population (n=390) had highly homogenous fragmentomes, while individuals with AD (FIG.26D) and LCr (FIG.26E) showed patterns of aberration. In general, feature-level metrics are plotted in gray, with black, blue and red coloring used to indicate high variability in healthy, LCr or AD individuals respectively. In panels c,d, and e the outer ring represents the ratio of short (<=150 bp) to long (>150 bp) fragments in 5Mb bins genome wide for each individual; lines are colored if Pearson’s correlation between the individual’s fragmentation profile and the median general screening population individual in the Discovery Cohort is less than 0.9. The second ring from the outside represents the percentage of individuals with a repeat element count outside the 95% confidence interval of all Discovery Cohort Individuals; elements are colored if this percentage is greater than 10%. The third ring from the outside represents the F-statistic comparing variation in 1Mb bins of high density of epigenetic marks for each group of individuals as compared to the individuals in the general screening population subset of the Discovery Cohort (n=293); Significant F-statistics (p<0.05) are colored. In FIGS.26D and 26E, the center most ring represents total weighted feature importance for each class of features within the penalized logistic regression models for LCr and AD detection, respectively. 24 168948691.1
[0090] FIGS.27A-27D. Cross-validated performance of fragmentome liquid biopsies for Alzheimer’s Disease (AD) and Liver Cirrhosis (LCr) detection in the Discovery Cohorts. The cross-validated cfDNA AD score generated from the classifier for AD detection was predicted to be high in individuals with AD as compared to those in healthy populations (FIG.27A) and could be separated with high area under the receiver operating curve (AUC-ROC) (FIG.27B). Similarly, the cross-validated cfDNA LCr score, generated from a separate classifier for LCr detection was high in individuals with LCr as compared to healthy populations (FIG. 27C) and separated individuals with and without LCr with high AUC-ROC (FIG.27D).
[0091] FIGS. 28A-28H. Locked fragmentome classifiers demonstrated Alzheimer’s Disease (AD) and Liver Cirrhosis (LCr) detection with limited cross-reactivity for other diseases in the Validation Cohorts. The locked AD detection classifier predicted high cfDNA AD scores in the Validation Cohort (FIG. 28A) that separated individuals with AD from those in healthy populations with high performance (FIG. 28B). At locked cfDNA AD score thresholds corresponding to fixed specificities in the Discovery Cohort, sensitivity and specificity remained high in the Validation Cohort demonstrating the generalizability of the classifier (FIG. 28C). Similarly, the locked LCr detection classifier predicted high cfDNA LCr scores in the Validation Cohort (FIG.28D) that separated individuals with LCr from those in healthy populations with and without metabolic risk factors (FIG. 28E), and locked score thresholds corresponding to fixed specificities in the Discovery Cohort maintained consistent, high sensitivity and specificity in the Validation Cohort (FIG.28F). The locked cfDNA AD (FIG.28G) and LCr (FIG.28H) classifiers generated low scores when applied to individuals with other diseases, demonstrating the limited cross-reactivity of the approach. For comparison, the dotted gray line shows the median Validation Cohort cfDNA AD Score for individuals with AD (FIG.28G), and the median Validation Cohort cfDNA LCr Score for individuals with LCr (FIG.28H).
[0092] FIGS. 29A-29C. Deconvolution analyses of plasma cfDNA fragmentomes revealed tissue-specific mechanisms of fragmentomic changes in Alzheimer’s Disease (AD) and Liver Cirrhosis (LCr). (FIG. 29A) The horizontal axis shows the quantile of number of transcription factor binding sites (TFBS) and the vertical axis shows the level of Pearson R correlation between the effect size of relative coverage in TFBS of patients with AD vs without 25 168948691.1and the effect size of transcription factor (TF) RNA expression of brain cortex tissue vs whole blood. Due to the inverse relationship between cfDNA coverage in the circulation and gene expression in tissue, the negative correlation suggested a strong brain-derived cfDNA contribution. This correlation was calculated and depicted across the full array of number of TFBS quantiles from 0 to 1. The gray shaded region identified the interquartile range (IQR) of correlations observed from random comparisons of individuals without AD. (FIG.29B) Similar graph to (FIG. 29A), individuals with and without LCr were compared to expression profiles of liver tissue vs whole blood. The negative correlation suggested a strong liver-derived cfDNA contribution. (FIG. 29C) These correlation comparisons were performed for individuals with AD (left) and LCr (right) as compared to all tissues in the GTEx database, and the correlations at the top quantile of TFBS were ranked with the most negative (strongest correlation) occurring first. This demonstrated an enrichment for brain related tissues in the cfDNA of patients with AD, which was not present for patients with LCr where the strongest contribution appears to be liver-derived.
[0093] FIGS.30A-30B. Overview of MICA Morbidity Cohort. (FIG. 30A) Individuals in the MICA Morbidity cohort were split into high and low Charlson Comorbidity Index groups and used for cfDNA fragmentation analyses and cross-validation of the Fragmentomic Morbidity Index. (FIG.30B) The samples from these individuals were prospectively collected as part of the Prospective Biomarker study MICA at Herlev and Gentofte Hospital, Copenhagen University Hospital, Denmark between July 2016 and July 2019.
[0094] FIGS. 31A-31D. cfDNA fragmentomes can serve as molecular biomarkers of morbidity. (FIG.31A) Individuals in the MICA Morbidity Cohort with high Charlson Comorbidity Index (CCI) (CCI 3+) had higher levels of protein biomarkers associated with inflammation as compared to individuals with low CCI (CCI 0,1,2). (FIG.31B) Protein biomarkers of inflammation correlated with increased cfDNA concentrations. (FIG.31C) Individuals with high CCI had shorter average fragment lengths than individuals with low CCI, and the trend was maintained for subgroups of high CCI individuals with specific morbidities (Benjamini-Hochberg adjusted p- values are shown). (FIG. 31D) Individuals with FMI values above the mean had shorter overall survival than individuals with FMI below the mean. 26 168948691.1
[0095] FIGS. 32A-32C. Overview of AD and LCr Discovery and Validation Cohorts. (FIG. 32A) Schematic of study design and cohorts analyzed. (FIGs. 32B and 32C) All plasma samples obtained are shown, with exclusions for clinical or quality reasons. This workflow led to the final Discovery and Validation Cohorts used for cfDNA liquid biopsy development.
[0096] FIGS. 33A-33B. Representative Tapestation traces of cfDNA extracted from plasma and DNA libraries constructed for whole genome sequencing (WGS). Tapestations of extracted cfDNA (top) and WGS libraries (bottom) are shown for a representative individual from the healthy screening population (FIG. 33A) and a representative individual with AD from the same source (FIG.33B). Tapestation results from extracted cfDNA show the assay’s lower marker in blue, and cfDNA peaks corresponding to nucleosomal cfDNA and a smaller peak of longer fragments potentially corresponding to genomic DNA in orange. The gray region shows the length range of nucleosomal cfDNA used to estimate concentration, which does not include genomic DNA. Tapestation results for WGS libraries show the assay’s lower and upper markers in blue, and nucleosomal cfDNA peaks in orange. This shows that size selection during library construction successfully removed the small amounts of genomic DNA that may have been present in the extracted DNA. The fragment length in WGS libraries includes annealed sequencing adaptors.
[0097] FIG.34. Concentration of cfDNA extracted from plasma by source, in the AD and LCr Discovery and Validation Cohorts. cfDNA concentrations were high in individuals with AD or LCr as compared to individuals without these diseases regardless of sample source. Concentrations above 200ng / mL are thresholded to 200ng / mL in this visualization.
[0098] FIG. 35. Application of Locked Fragmentation Morbidity Index (FMI) model in AD and LCr Discovery Cohorts. The locked FMI model predicted higher scores for individuals in the AD and LCr discovery cohorts who were estimated to have high CCI as compared to individuals with an estimated low CCI.
[0099] FIG.36. cfDNA scores were not influenced by WGS Library Batch. For each type of individual and library batch, the median cfDNA AD (top) or LCr (bottom) score is plotted, demonstrating that disease scores were consistent across batches. Library batches were composed 27 168948691.1of individuals with and without disease, and Discovery and Validation cohorts were processed separately. [000100] FIGS. 37A-37C. Relationship between cfDNA AD and LCr scores and clinical covariates in the Discovery Cohort. cfDNA scores for (FIG.37A) healthy individuals, (FIG.37B) individuals with AD, and (c) individuals with Liver Cirrhosis in the Discovery Cohort as compared to various clinically relevant covariates. [000101] FIGS. 38A-38C. cfDNA AD and cfDNA LCr scores stratified by disease status and source in Discovery and Validation Cohorts. Regardless of sample source, cfDNA AD (FIG. 38A) and LCr (FIG. 38B) scores were high for individuals with disease and low for healthy individuals in the Discovery and Validation cohorts. (FIG. 38C) Even when considering only samples from the same source, the cfDNA AD and LCr scores effectively separate individuals with and without disease. [000102] FIGS.39A-39E. Disease specific changes in cfDNA fragmentomes of individuals with AD and LCr in the Validation Cohorts. (FIG. 39A) Elevated cfDNA concentrations were observed in individuals with AD (n=42) and LCr (n=83). p-values indicate the Benjamini- Hochberg corrected p-value from Wilcoxon’s rank-sum test as compared to the healthy population (n=142). (FIG. 39B) Individuals with AD and LCr had an increase in short fragments less than ~170bp reflecting an increase in sub-mononucleosomal fragments. Each line shows the mean difference in fragment length distribution between these patient groups and a subset of healthy individuals in the Discovery Cohort, and shading indicates + / - 1SD. (FIG.39C) Individuals in the presumed healthy population (n=142) had highly homogenous fragmentomes, while individuals with AD (FIG.39D) and LCr (FIG.39E) showed patterns of aberration. In general, feature-level metrics are plotted in gray, with black, blue and red coloring used to indicate high variability in healthy, LCr or AD individuals respectively. In this figure all Validation Cohort patient groups, including healthy individuals, are compared to healthy individuals from the Discovery Cohort. In panels c,d, and e the outer ring represents the ratio of short (<=150 bp) to long (>150 bp) fragments in 5Mb bins genome wide for each individual; lines are colored if Pearson’s correlation between the individual’s fragmentation profile and the median general screening population individual from the Discovery Cohort is less than 0.9. The second ring from the outside represents the percentage 28 168948691.1of individuals with a repeat element count outside the 95% confidence interval of all Discovery Cohort individuals; elements are colored if this percentage is greater than 10%. The third ring from the outside represents the F-statistic comparing variation in 1Mb bins of high density of epigenetic marks for each group of individuals as compared to the individuals in the general screening population subset of the Discovery Cohort (n=293); Significant F-statistics (p<0.05) are colored. In panels d and e, the center most ring represents total weighted feature importance for each class of features within the penalized logistic regression models for LCr and AD detection, respectively. [000103] FIGS. 40A-40D. Relationship between cfDNA AD and LCr scores and clinical covariates. cfDNA scores for (FIG.40A) healthy individuals and (FIG.40B) individuals with AD in the validation cohort. (FIG. 40C) cfDNA AD and LCr scores demonstrate high performance for disease detection in Discovery and Validation Cohorts even when individuals in the healthy population are restricted to individuals with age 70 or above. (FIG.40D) Individuals with Liver Cirrhosis in the Validation Cohort as compared to various clinically relevant covariates. [000104] FIGS.41A-1B. Clinical utility of cfDNA AD Approach for Alzheimer’s Disease Detection. Standard of Care (FIG.41A) and proposed cfDNA AD (FIG.40B) clinical workflows for the detection of Alzheimer’s Disease. [000105] FIG. 42. cfDNA LCr Scores for high-risk, pre-cirrhotic liver diseases. In the Discovery and Validation Cohorts, individuals at high-risk for liver cirrhosis including those with Viral Hepatitis, Aflatoxin exposure, MASLD and Liver Fibrosis had predicted cfDNA LCr scores higher than those in the general healthy population but lower than those for patients with LCr, suggesting this approach may have utility for detecting fragmentomic changes in pre-cirrhotic liver conditions. [000106] FIG.43. Comparison of performance of cfDNA LCr score to Fibrosis-4 (FIB-4) index in Validation Cohort. For a subset of individuals with LCr in the validation cohort, the FIB- 4 metric is available (n=42). In this group, sensitivity for identifying individuals with LCr is 69% with the cfDNA LCr score locked at a threshold achieving 90% specificity in the Discovery Cohort. In contrast, at the FIB-4 threshold of 3.25, sensitivity is 55%. 29 168948691.1[000107] FIG. 44A-44B. Clinical utility of cfDNA LCr Approach for Liver Cirrhosis Detection. Standard of Care (FIG.44A) and proposed cfDNA LCr (FIG.44B) clinical workflows for the detection of Liver Cirrhosis. [000108] FIG. 45A-45E. Relationship between cfDNA AD and LCr scores and cfDNA concentration. cfDNA AD (FIG. 45A) and LCr (FIG. 45B) scores demonstrate weak to no correlations to cfDNA concentration, and cfDNA scores for individuals with AD and LCr were higher than for healthy indviduals regardless of cfDNA concentrations. The cfDNA AD (FIG. 45C) and LCr (FIG. 45D) scores demonstrate higher performance for detection than cfDNA concentration levels alone. Panels c and d show combined performance for the discovery and validation cohorts, with scores for discovery cohort individuals generated by cross-validation and scores for validation cohort individuals generated using locked models. (FIG. 45E) The locked cfDNA LCr model when applied to individuals with Alzheimer’s disease yields low scores with limited cross-reactivity across a range of cfDNA concentration levels. [000109] FIG. 46. Application of locked cfDNA lung cancer detection model to AD and LCr cohorts showed limited cross-reactivity. The locked cfDNA lung cancer model predicted low scores for individuals in the AD and LCr cohorts. [000110] FIGS. 47A-47B. Relationship between cfDNA coverage at transcription factor binding sites and differences in transcription factor expression between organ tissues and blood. (FIG.47A) Negative correlation between relative coverage at transcription factor binding sites in plasma from individuals with AD as compared to healthy individuals and transcription factor expression in brain cortex tissue as compared to whole blood suggests an increased proportion of brain-derived cfDNA in the blood of individuals with AD. (FIG. 47B) Negative correlation between relative coverage at transcription factor binding sites in plasma from individuals with LCr as compared to healthy individuals and transcription factor expression in liver tissue as compared to whole blood suggests an increased proportion of liver-derived cfDNA in the blood of individuals with LCr. From top to bottom, rows show this relationship for all transcription factors studied, all transcription factors above the 90thpercentile for number of binding sites, and for all transcription factors above the 99thpercentile for number of binding sites. 30 168948691.1[000111] FIGS. 48A-48B. Deconvolution analyses of plasma cfDNA fragmentomes revealed tissue-specific and immune mediated mechanisms of fragmentomic changes in Alzheimer’s Disease (AD) and Liver Cirrhosis (LCr). (FIG.48A) The horizontal axis shows the quantile of number of transcription factor binding sites (TFBS) and the vertical axis shows the level of Pearson R correlation between the effect size of relative coverage in TFBS of patients with AD vs without and the effect size of transcription factor (TF) RNA expression of several brain cell subpopulations vs whole blood. Due to the inverse relationship between cfDNA coverage in the circulation and gene expression in tissue, the negative correlation suggested a strong brain-derived cfDNA contribution. This correlation was calculated and depicted across the full array of number of TFBS quantiles from 0 to 1. The gray shaded region demonstrates the interquartile range (IQR) of correlations observed from random comparisons of individuals without AD. (FIG. 48B) This panel is a similar graph to (a); here individuals with and without AD or LCr were compared to expression profiles of immune cell subpopulations vs whole blood. Correlations outside the IQR observed in null permutations suggested systemic immune perturbations were evident in the circulation. [000112] FIG.49. Schematic of cohorts used in this study, annotated by sample source. All samples analyzed in the MICA Morbidity Cohort were obtained from a single source, distinct from those in the Disease Detection Cohorts. In the cohorts for AD and LCr detection, each source provided samples only for use in the Discovery or Validation Cohort, ensuring that the Validation Cohort contained only samples from sources external to those used in the Discovery Cohort to cross-validate the models. [000113] FIG.50. Principal Component Analysis (PCA) for fragmentomic features by site of collection. Each panel represents a separate PCA. Rows show PCAs performed on a subset of individuals, from top to bottom: healthy individuals, individuals at high-risk for LCr, individuals with LCr, and individuals with AD. Columns show PCAs based on different feature classes, from left to right: short to long fragment ratio features, repeat element features, and epigenetic features. Each point represents one individual and points are colored by sample source, demonstrating that within each set of individuals and features, source does not appear to drive significant variation. 31 168948691.1[000114] FIG.51. Schematic of model architecture used in AD and LCr detection models. The cfDNA AD or LCr score was generated as an ensemble of two components – a fragmentation score and a repeat element score. The repeat element score was itself an ensemble of scores from individual models corresponding to 6 feature families. DETAILED DESCRIPTION [000115] Plasma biomarkers based on cell-free DNA (cfDNA) have shown great promise for cancer detection, identification of response to therapy and tumor recurrence, and detection of exogenous DNA in prenatal testing and monitoring of organ transplantion1–8. Here their applicability in disease detection is demonstrated across a range of clinical conditions including neurodegenerative disease, chronic inflammation, and non-neoplastic dysregulation of cellular growth. [000116] In this study, evidence is presented for changes to the fragmentome across broad categories of disease phenotypes including neurodegeneration, chronic inflammation, and non- neoplastic dysregulation of cellular growth, using three representative examples for which there currently are no effective blood-based screening methods: Alzheimer’s disease (AD), liver cirrhosis, and benign adnexal masses. It is further demonstrated that the methodologies by which these changes can be measured and used in the diagnosis of these illnesses. [000117] Alzheimer’s disease (AD) is a devastating disease marked by progressive neurodegeneration and dementia. It is estimated that by 2050 approximately 152 million people worldwide will have AD and other dementias9. Less than 1% of AD cases have an autosomal dominant inheritance pattern and the vast majority have unclear etiologies10. AD is marked by beta amyloid plaques and neurofibrillary tau tangle accumulation in the brain, which drive neuronal death and cognitive impairment11. AD is a progressive disease with no cure, though recent drug approvals have raised hopes for a disease modifying treatment. In theory, preventing plaque deposits should slow or even prevent neurodegeneration, though this hinges on successful identification of patients and rapid treatment initiation, ideally in individuals who are presymptomatic or have early-stage disease, as these protein deposits are known to originate years before the onset of symptoms. Therefore, early detection of AD is a critical clinical unmet need. 32 168948691.1Today, the only confirmatory diagnosis of AD occurs through post-mortem histology, though presumed AD can be diagnosed, and treatment initiated based on cognitive impairment when other causes of dementia (i.e. vascular, frontotemporal, or Lewy body) are ruled out11. While PET CT imaging can identify amyloid and tau deposits, many cases of presumed AD occur in the absence of suggestive imaging findings, and this imaging modality is expensive and inaccessible for large sections of the global population. Given this clinical context, there is widespread interest in other biomarkers for AD which may be more informative and accessible. Many plasma and cerebrospinal fluid based protein biomarkers have been investigated, though their effectiveness has been limited and they remain far from widespread clinical adoption12. [000118] It stands to reason that changes to the cfDNA fragmentome would be evident in AD. First, AD patients have overall higher levels of cfDNA. The tissue specific origins of this increase are yet to be elucidated, but without wishing to be bound by theory, it was hypothesized that the neurotoxicity of beta-amyloid and tau deposits leading to neuronal cell-death and blood brain barrier degradation may increase the levels of brain-derived cell-free DNA present in the circulation of AD patients as compared to in healthy individuals. Second, AD is thought to involve a pro-inflammatory phenotype14,15. The majority of cfDNA typically originates from white blood cells, suggesting that fragmentomic signals could capture immune related changes. In the examples section which follows, the cfDNA fragmentome in AD was characterized and demonstrated its utility in a proof-of-concept liquid biopsy assay for disease detection. [000119] In certain embodiments, a method of early detection and treatment of neurodegenerative diseases comprises extracting cfDNA from a subject’s sample; isolating cell free DNA (cfDNA) from the subject’s biological samples; conducting low coverage whole genome sequencing of the subject’s isolated cfDNA to generate genomic libraries of sequenced cfDNA fragments; determining a cfDNA fragmentome profile of the subject’s sequenced DNA fragments comprising evaluating one or more cfDNA fragment characteristics; comparing the cfDNA fragmentome profiles of the subject to reference cfDNA fragmentome profiles from healthy subjects to detect and diagnose the neurodegenerative disease; and, treating the subject. 33 168948691.1[000120] In certain embodiments, the one or more cfDNA fragment characteristics comprises calculating ratios of small to large cfDNA fragments, sequences of one or more cfDNAs, median cfDNA fragment sizes, fragment size distribution, mutant allele frequencies, fragment length, fragment size distribution, fragment end motifs, preferred end coordinates, breakpoint motifs, methylation frequencies, concentrations of cfDNA or combinations thereof. In certain embodiments, concentrations of cfDNA are higher in subjects with a neurodegenerative disease as compared to as compared to healthy subjects. In certain embodiments, the cfDNA fragmentome profiles in subjects identified and diagnosed early with a neurodegenerative disease, are altered across the genome. In certain embodiments, the cfDNA fragmentome profiles in subjects identified and diagnosed early with a disease, have greater heterogeneity across the genome as compared to healthy subjects. [000121] In another example, liver cirrhosis, or scarring, occurs via a chronic pattern of cellular death and inflammation leading to irreversible liver damage. Liver disease, including cirrhosis, is the 11thleading cause of death globally. Among those with liver cirrhosis, there is a 4-12% annual risk of liver failure16, and a 2-6% annual risk of developing hepatocellular carcinoma (HCC)17. While liver cirrhosis is not reversible, its timely identification improves clinical management. Current clinical guidelines recommend that patients with cirrhotic livers should be regularly surveilled for development of HCC18. Moreover, behavioral and pharmaceutical interventions, such as treatment of alcohol use disorder and initiation of beta-blockers to treat portal hypertension, can preserve liver function and slow failure19. Despite the clear benefits to identification of cirrhosis, this remains a complicated clinical diagnosis. Liver stiffness measurements and fibrosis indices estimated from blood lab values can provide a diagnostic guide, though liver biopsy remains the gold-standard19. This creates challenges in early identification and longitudinal monitoring of cirrhosis. [000122] Widespread genomic changes have been identified in cirrhotic livers, including increased somatic mutations and structural variants as compared to normal liver20. Given the known detectability of such genomic changes in cfDNA and the high rates of cellular death present in cirrhotic livers21, in the examples section which follows, cirrhosis-related changes in the fragmentome were identified and incorporated into a liquid biopsy for detecting cirrhosis. 34 168948691.1[000123] In another example, benign adnexal masses, such as cysts and cystadenomas of the ovaries or fallopian tubes, are fairly common, and an estimated 10% of women will undergo surgery for removing such masses during their lifetime22. These masses, when large, can cause symptoms including ascites, abnormal bleeding, and abdominal pain23. Although these masses represent hyperproliferation of epithelial, germ, or stromal cells, only a minority of them (<20%) may be malignant22. For these reasons, identifying these lesions can facilitate reproductive health decision-making, surgical debulking if needed to alleviate symptoms, or initiation of cancer surveillance. [000124] Given the high incidence of these benign masses there is a clinical unmet need for a cost-efficient and accessible screening strategy. Imaging approaches such as ultrasound are typically only used in symptomatic women and their performance is operator dependent24. Protein-based blood biomarkers have been proposed as a screening strategy for ovarian cancer but there have been fewer efforts to identify benign lesions25. Non-invasive assays utilizing genome- wide cfDNA fragmentomes can detect subtle genomic changes from tissues not usually released in the circulation and may provide an accessible approach to identification of benign adnexal masses. [000125] Accordingly, in certain embodiments, a method of early detection and treatment of diseases, comprises extracting cfDNA from a subject’s sample; isolating cell free DNA (cfDNA) from the subject’s biological samples; conducting low coverage whole genome sequencing of the subject’s isolated cfDNA to generate genomic libraries of sequenced cfDNA fragments; determining a cfDNA fragmentome profile of the subject’s sequenced DNA fragments comprising evaluating one or more cfDNA fragment characteristics; comparing the cfDNA fragmentome profiles of the subject to reference cfDNA fragmentome profiles from healthy subjects to detect and diagnose the disease; and, treating the subject with a disease specific therapy. [000126] In certain embodiments, the one or more cfDNA fragment characteristics comprises calculating ratios of small to large cfDNA fragments, sequences of one or more cfDNAs, median cfDNA fragment sizes, fragment size distribution, mutant allele frequencies, 35 168948691.1fragment length, fragment size distribution, fragment end motifs, preferred end coordinates, breakpoint motifs, methylation frequencies, concentrations of cfDNA or combinations thereof. In certain embodiments, a small cfDNA fragment comprises about 80 base pairs (bp) to about 150 bp. In certain embodiments, a large cfDNA fragment comprises about 151 bp to about 300 bp. In certain embodiments, the small to large cfDNA ratios are GC corrected. [000127] In certain embodiments, the cfDNA fragmentome profile comprises the sequence coverage of small cfDNA fragments in windows across the genome. In certain embodiments, the cfDNA fragmentome profile comprises the sequence coverage of cfDNA fragments in windows across the genome. In certain embodiments, the cfDNA fragmentome profile comprises the sequence coverage of small and large cfDNA fragments in windows across the genome. [000128] In certain embodiments, the cfDNA fragmentome profiles in subjects identified and diagnosed early with a disease, are altered across the genome. In certain embodiments, the cfDNA fragmentome profiles in subjects identified and diagnosed early with a disease, have greater heterogeneity across the genome as compared to healthy subjects. In certain embodiments, the method further comprises identifying short nucleic acid sequences (kmers) in genomic or cell free DNA; selecting kmers occurring in a single repeat type and identifying unique kmers of repeat element types; wherein the unique kmers identify one or a plurality of repeat element types. [000129] In certain embodiments, genome wide differences in subjects identified and diagnosed early with a disease, have greater heterogeneity across the genome in repeat elements, long interspersed nuclear elements (LINEs), short interspersed nuclear elements (SINEs), long terminal repeats (LTRs), satellites, transposable elements, RNA elements or combinations thereof, as compared to healthy subjects. In certain embodiments, the diseases comprise: neurodegenerative diseases, inflammatory diseases, cancer and non-neoplastic dysregulation of cellular growth. In certain embodiments, the cfDNA fragmentation profile for a specific disease is a biomarker profile diagnostic of a specific disease. [000130] DNA Evaluation of Fragments for early Interception (DELFI) 36 168948691.1[000131] DNA Evaluation of Fragments for early Interception (DELFI) was previously developed, Cristiano S, Leal A, Phallen J, et al. Genome-wide cell-free DNA fragmentation in patients with cancer. Nature 2019;570:385-9 incorporated herein in its entirety, and used to evaluate genome-wide fragmentation patterns of cfDNA of 236 patients with breast, colorectal, lung, ovarian, pancreatic, gastric, or bile duct cancers as well as 245 healthy individuals. These analyses revealed that cfDNA profiles of healthy individuals reflected nucleosomal fragmentation patterns of white blood cells, while patients with cancer had altered fragmentation profiles. DELFI had sensitivities of detection ranging from 57% to >99% among the seven cancer types at 98% specificity and identified the tissue of origin of the cancers to a limited number of sites in 75% of embodiments. Assessing cfDNA (e.g., using DELFI) provide a screening approach for early detection of cancer, which can increase the chance for successful treatment of a patient having cancer. Assessing cfDNA (e.g., using DELFI) can also provide an approach for monitoring cancer, which can increase the chance for successful treatment and improved outcome of a patient having cancer. In addition, a cfDNA fragmentation profile can be obtained from limited amounts of cfDNA and using inexpensive reagents and / or instruments. [000132] Accordingly in certain embodiments, a method of early detection and treatment of diseases, comprises extracting cfDNA from a subject’s sample; isolating cell free DNA (cfDNA) from the subject’s biological samples; conducting low coverage whole genome sequencing of the subject’s isolated cfDNA to generate genomic libraries of sequenced cfDNA fragments; determining a cfDNA fragmentome profile of the subject’s sequenced DNA fragments comprising evaluating one or more cfDNA fragment characteristics; comparing the cfDNA fragmentome profiles of the subject to reference cfDNA fragmentome profiles from healthy subjects to detect and diagnose the disease; and, treating the subject with a disease specific therapy. [000133] In certain embodiments, the method further comprises subjecting the subject to a low dose helical computed tomography (LDCT). In certain embodiments, the method further comprises comparing clinical data between the subject diagnosed as having lung cancer and normal non-cancer subjects. In certain embodiments, the cfDNA fragment mean length and profiles are similar among non-cancer individuals. In certain embodiments, the cfDNA fragment 37 168948691.1profiles of cancer subjects vary. In certain embodiments, the serum levels of or one or more tumor antigens, cytokines or proteins are measured. [000134] In certain embodiments a DELFI score is generated, wherein the principal component analysis is incorporated into a machine learning predictive model to generate a score for each subject as an average over cross-validation repeats (DELFI score(s)). In certain embodiments, DELFI scores are generated for each disease, for example, Alzheimer’s disease, liver cirrhosis and the like. In certain embodiments, DELFI scores are generated for each Child- Pugh categorization of cirrhosis severity. In certain embodiments, DELFI scores are generated for stages of Alzheimer’s Disease cognitive stratification. [000135] In certain embodiments a DELFI score is generated, wherein the principal component analysis is incorporated into a machine learning predictive model to generate a score for each subject as an average over cross-validation repeats (DELFI score(s)). [000136] cfDNA Fragmentation Profiles: A cfDNA fragmentation profile can include one or more cfDNA fragmentation patterns. A cfDNA fragmentation pattern can include any appropriate cfDNA fragmentation pattern. Examples of cfDNA fragmentation patterns include, without limitation, median fragment size, fragment size distribution, ratio of small cfDNA fragments to large cfDNA fragments, and the coverage of cfDNA fragments. In some embodiments, a cfDNA fragmentation pattern includes two or more (e.g., two, three, or four) of median fragment size, fragment size distribution, ratio of small cfDNA fragments to large cfDNA fragments, and the coverage of cfDNA fragments. In some embodiments, cfDNA fragmentation profile can be a genome-wide cfDNA profile (e.g., a genome-wide cfDNA profile in windows across the genome). In some embodiments, cfDNA fragmentation profile can be a targeted region profile. A targeted region can be any appropriate portion of the genome (e.g., a chromosomal region). Examples of chromosomal regions for which a cfDNA fragmentation profile can be determined as described herein include, without limitation, a portion of a chromosome (e.g., a portion of 2q, 4p, 5p, 6q, 7p, 8q, 9q, 10q, 11q, 12q, and / or 14q) and a chromosomal arm (e.g., a chromosomal arm of 8q,13q, 11q, and / or 3p). In some embodiments, a cfDNA fragmentation profile can include two or more targeted region profiles. 38 168948691.1[000137] In some embodiments, a cfDNA fragmentation profile can be used to identify changes (e.g., alterations) in cfDNA fragment lengths. An alteration can be a genome-wide alteration or an alteration in one or more targeted regions / loci. A target region can be any region containing one or more cancer-specific alterations. In some embodiments, a cfDNA fragmentation profile can be used to identify (e.g., simultaneously identify) from about 10 alterations to about 500 alterations (e.g., from about 25 to about 500, from about 50 to about 500, from about 100 to about 500, from about 200 to about 500, from about 300 to about 500, from about 10 to about 400, from about 10 to about 300, from about 10 to about 200, from about 10 to about 100, from about 10 to about 50, from about 20 to about 400, from about 30 to about 300, from about 40 to about 200, from about 50 to about 100, from about 20 to about 100, from about 25 to about 75, from about 50 to about 250, or from about 100 to about 200, alterations). [000138] A cfDNA fragmentation profile can include a ratio of small cfDNA fragments to large cfDNA fragments and a correlation of fragment ratios to reference fragment ratios. As used herein, with respect to ratios of small cfDNA fragments to large cfDNA fragments, a small cfDNA fragment can be from about 100 bp in length to about 150 bp in length. As used herein, with respect to ratios of small cfDNA fragments to large cfDNA fragments, a large cfDNA fragment can be from about 151 bp in length to 220 bp in length. A cfDNA fragmentation profile can include coverage of all fragments. Coverage of all fragments can include windows (e.g., non- overlapping windows) of coverage. In some embodiments, coverage of all fragments can include windows of small fragments (e.g., fragments from about 100 bp to about 150 bp in length). In some embodiments, coverage of all fragments can include windows of large fragments (e.g., fragments from about 151 bp to about 220 bp in length). [000139] In some embodiments, methods and materials described herein also can include machine learning. For example, machine learning can be used for identifying an altered fragmentation profile (e.g., using coverage of cfDNA fragments, fragment size of cfDNA fragments, coverage of chromosomes, and mtDNA). [000140] ARTEMIS [000141] ARTEMIS (Analysis of RepeaT EleMents in dISease) was developed as an alignment-free, genome-wide approach for analyzing repeat landscapes in short read 39 168948691.1sequencing. ARTEMIS assesses over 1200 individual repeat types that occur genome-wide and span 57 subfamilies comprising 6 families (Satellites, RNA elements, transposable elements, LINEs, SINEs, LTRs).27ARTEMIS was used to show that repeat landscapes are enriched in genes commonly altered in human cancer and tumor-specific changes in repeats reflect a combination of structural and epigenetic changes in the cancer genome. In this study, ARTEMIS is used to evaluate neurodegenerative diseases (e.g. Alzheimer’s disease), inflammatory diseases, cancer, non-neoplastic dysregulation of cellular growth and the like. Non-neoplastic dysregulation of cellular growth diseases include Liver Cirrhosis and benign adnexal masses. [000142] Genome-wide repeat landscape analyses with ARTEMIS can be performed using low-coverage whole genome sequencing, permitting analysis of repeat landscapes in cfDNA for early diagnosis of neurodegenerative diseases (e.g. Alzheimer’s disease), inflammatory diseases, cancer, non-neoplastic dysregulation of cellular growth and the like. Non-neoplastic dysregulation of cellular growth diseases include Liver Cirrhosis and benign adnexal masses. [000143] Accordingly, in certain aspects, a method of identifying repeat element types, comprises extracting repeat sequences and coordinates from known repeat element types; identifying short nucleic acid sequences (kmers) in genomic or cell free DNA; selecting kmers occurring in a single repeat type and identifying unique kmers of repeat element types; wherein the unique kmers identify one or a plurality of repeat element types. In certain embodiments, repeat element types are excluded from families comprising low complexity, unknown, simple repeats or combinations thereof. In certain embodiments, elements from each family comprises tRNAs, srpRNAs, snRNAs, scRNAs, rRNAs, RNA elements, DNA, retroposons, or combinations thereof, are aggregated. [000144] In certain embodiments, one or a plurality of kmers are generated from one or more nucleic acid sequences. The nucleic acid sequence may be any sequence, such as a genomic sequence or a cfDNA. In one embodiment, kmer generation is performed by an automated algorithm that receives a nucleic acid sequence and generates multiple kmers from 40 168948691.1that sequence. [000145] The kmer may be produced from a single organism or species, or from multiple organisms or species. For example, kmer can be generated from the genomic sequences of common, possible or known contaminants to detect these contaminants during quality control analysis. In addition, k-mers can be generated from the genomic sequences of different organisms that may be in the composite sample to detect and analyze sequence data from multiple different organisms during quality control analysis. [000146] k-mers extracted from one or more nucleic acid sequences may have the same length or various different lengths. As just one example, the generated kmer may be about 14 bases, but longer and shorter k-mers are possible. Larger ks are less likely to find multiple hits in the reference genome, which can result in multiple types of k-mer annotations. For example, a small k maps to many parts of the genome and does not provide useful and unique annotation information. However, a large k also increases the amount of memory required to store the generated k-mer. A short k reduces the amount of memory required to store the generated kmer. Therefore, the determination of optimal k depends in particular on various factors such as computational power and memory size. [000147] In certain embodiments, the kmers occurring in single repeat type elements and not in non-repeat regions are selected. In certain embodiments, the kmer comprises 10 to 40 nucleotides. In certain embodiments, the kmer comprises 15 to 35 nucleotides. In certain embodiments, the kmer comprises 20 to 30 nucleotides. In certain embodiments, the kmer comprises 22 to 28 nucleotides. In certain embodiments, the kmer comprises 23 to 26 nucleotides. In certain embodiments, the kmer comprises about 24 nucleotides. [000148] Methods of Treatment [000149] In certain aspects, a method of early detection and treatment of diseases, comprises extracting cfDNA from a subject’s sample; isolating cell free DNA (cfDNA) from the subject’s biological samples; conducting low coverage whole genome sequencing of the subject’s isolated cfDNA to generate genomic libraries of sequenced cfDNA fragments; determining a cfDNA fragmentome profile of the subject’s sequenced DNA fragments 41 168948691.1comprising evaluating one or more cfDNA fragment characteristics; comparing the cfDNA fragmentome profiles of the subject to reference cfDNA fragmentome profiles from healthy subjects to detect and diagnose the disease; and, treating the subject with a disease specific therapy. [000150] In certain embodiments, the one or more cfDNA fragment characteristics comprises calculating ratios of small to large cfDNA fragments, sequences of one or more cfDNAs, median cfDNA fragment sizes, fragment size distribution, mutant allele frequencies, fragment length, fragment size distribution, fragment end motifs, preferred end coordinates, breakpoint motifs, methylation frequencies, concentrations of cfDNA or combinations thereof. [000151] In certain embodiments, a small cfDNA fragment comprises about 80 base pairs (bp) to about 150 bp. In certain embodiments, a large cfDNA fragment comprises about 151 bp to about 300 bp. In certain embodiments, the small to large cfDNA ratios are GC corrected. In certain embodiments, the cfDNA fragmentome profile comprises the sequence coverage of small cfDNA fragments in windows across the genome. [000152] In certain embodiments, the cfDNA fragmentome profile comprises the sequence coverage of cfDNA fragments in windows across the genome. In certain embodiments, the cfDNA fragmentome profile comprises the sequence coverage of small and large cfDNA fragments in windows across the genome. [000153] In certain embodiments, the cfDNA fragmentome profiles in subjects identified and diagnosed early with a disease, are altered across the genome. In certain embodiments, the cfDNA fragmentome profiles in subjects identified and diagnosed early with a disease, have greater heterogeneity across the genome as compared to healthy subjects. In certain embodiments, the method further comprises identifying short nucleic acid sequences (kmers) in genomic or cell free DNA; selecting kmers occurring in a single repeat type and identifying unique kmers of repeat element types; wherein the unique kmers identify one or a plurality of repeat element types. In certain embodiments, genome wide differences in subjects identified and diagnosed early with a disease, have greater heterogeneity across the genome in repeat elements, long interspersed nuclear elements (LINEs), short interspersed nuclear elements (SINEs), long 42 168948691.1terminal repeats (LTRs), satellites, transposable elements, RNA elements or combinations thereof, as compared to healthy subjects. [000154] In certain embodiments, the diseases comprise: neurodegenerative diseases (e.g. Alzheimer’s disease), inflammatory diseases, cancer, non-neoplastic dysregulation of cellular growth and the like. Non-neoplastic dysregulation of cellular growth diseases includes Liver Cirrhosis and benign adnexal masses. [000155] Examples of neurodegenerative diseases include AIDS dementia complex, Alzheimer's disease, amyotrophic lateral sclerosis, adrenoleukodystrophy, Alexander disease, Alper's disease, ataxia telangiectasia, Batten disease, bovine spongiform encephalopathy (BSE), Canavan disease, corticobasal degeneration, Creutzfeldt-Jakob disease, dementia with Lewy bodies, fatal familial insomnia, frontotemporal lobar degeneration, Huntington's disease, Kennedy's disease, Krabbe disease, Lyme disease, Machado-Joseph disease, multiple sclerosis, multiple system atrophy, neuroacanthocytosis, Niemann-Pick disease, Parkinson's disease, Pick's disease, primary lateral sclerosis, progressive supranuclear palsy, Refsum disease, Sandhoff disease, diffuse myelinoclastic sclerosis, spinocerebellar ataxia, subacute combined degeneration of spinal cord, tabes dorsalis, Tay-Sachs disease, toxic encephalopathy, transmissible spongiform encephalopathy, and wobbly hedgehog syndrome. [000156] As discussed, in certain embodiments, the diseases may comprise one or more liver diseases. In aspects, a disease may comprise hepatitis including viral hepatitis. In an aspect, a disease may comprise non-alcoholic fatty liver disease in a subject In an aspect, a disease may comprise metabolic associated steatotic liver disease. In an aspect, a disease may comprise steatosis. In a particular aspect, a disease may comprise liver fibrosis. In a particular aspect, a disease may comprise liver cirrhosis. [000157] [000158] Systems [000159] In some examples, the present disclosure provides systems, methods, or kits that can include data analysis realized in measurement devices (e.g., laboratory instruments, such as a sequencing machine), software code that executes on computing hardware. The software can be 43 168948691.1stored in memory and execute on one or more hardware processors. The software can be organized into routines or packages that can communicate with each other. A module can comprise one or more devices / computers, and potentially one or more software routines / packages that execute on the one or more devices / computers. For example, an analysis application or system can include at least a data receiving module, a data pre-processing module, a data analysis module (which can operate on one or more types of genomic data), a data interpretation module, or a data visualization module. [000160] The data receiving module can connect laboratory hardware or instrumentation with computer systems that process laboratory data. The data pre-processing module can perform operations on the data in preparation for analysis. Examples of operations that can be applied to the data in the pre-processing module include affine transformations, denoising operations, data cleaning, reformatting, or subsampling. The data analysis module, which can be specialized for analyzing genomic data from one or more genomic materials, can, for example, take assembled genomic sequences and perform probabilistic and statistical analysis to identify abnormal patterns related to a disease, pathology, state, risk, condition, or phenotype. The data interpretation module can use analysis methods, for example, drawn from statistics, mathematics, or biology, to support understanding of the relation between the identified abnormal patterns and health conditions, functional states, prognoses, or risks. The data analysis module and / or the data interpretation module can include one or more machine learning models, which can be implemented in hardware, e.g., which executes software that embodies a machine learning model. The data visualization module can use methods of mathematical modeling, computer graphics, or rendering to create visual representations of data that can facilitate the understanding or interpretation of results. The present disclosure provides computer systems that are programmed to implement methods of the disclosure. [000161] In some embodiments, the methods disclosed herein can include computational analysis on nucleic acid sequencing data of samples from an individual or from a plurality of individuals. An analysis can identify a variant inferred from sequence data to identify sequence variants based on probabilistic modeling, statistical modeling, mechanistic modeling, network modeling, or statistical inferences. Non-limiting examples of analysis methods include principal 44 168948691.1component analysis, autoencoders, singular value decomposition, Fourier bases, wavelets, discriminant analysis, regression, support vector machines, tree-based methods, networks, matrix factorization, and clustering. Non-limiting examples of variants include a germline variation or a somatic mutation. In some examples, a variant can refer to an already-known variant. The already- known variant can be scientifically confirmed or reported in literature. In some examples, a variant can refer to a putative variant associated with a biological change. A biological change can be known or unknown. In some examples, a putative variant can be reported in literature, but not yet biologically confirmed. Alternatively, a putative variant is never reported in literature, but can be inferred based on a computational analysis disclosed herein. In some examples, germline variants can refer to nucleic acids that induce natural or normal variations. [000162] In certain embodiments, the computer system includes a central processing unit (CPU, also “processor” and “computer processor” herein), which can be a single core or multi core processor, or a plurality of processors for parallel processing; memory (e.g., cache, random-access memory, read-only memory, flash memory, or other memory); electronic storage unit (e.g., hard disk), communication interface (e.g., network adapter) for communicating with one or more other systems; and peripheral devices, such as adapters for cache, other memory, data storage and / or electronic display. The memory, storage unit, interface and peripheral devices may be in communication with the CPU through a communication bus (solid lines), such as a motherboard. The storage unit can be a data storage unit (or data repository) for storing data. One or more analyte feature inputs can be entered from the one or more measurement devices. Example analytes and measurement devices are described herein. [000163] The computer system can be operatively coupled to a computer network (“network”) with the aid of the communication interface. The network can be the Internet, an internet and / or extranet, or an intranet and / or extranet that is in communication with the Internet. The network in some cases is a telecommunication and / or data network. The network can include one or more computer servers, which can enable distributed computing, such as cloud computing over the network (“the cloud”) to perform various aspects of analysis, calculation, and generation of the present disclosure, such as, for example, activation of a valve or pump to transfer a reagent or sample from one chamber to another or application of heat to a sample (e.g., during an 45 168948691.1amplification reaction), other aspects of processing and / or assaying a sample, performing sequencing analysis, measuring sets of values representative of classes of molecules, identifying sets of features and feature vectors from assay data, processing feature vectors using a machine learning model to obtain output classifications, and training a machine learning model (e.g., iteratively searching for optimal values of parameters of the machine learning model). Such cloud computing may be provided by cloud computing platforms such as, for example, Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform, and IBM cloud. The network, in some cases with the aid of the computer system, can implement a peer-to-peer network, which may enable devices coupled to the computer system to behave as a client or a server. [000164] The CPU can execute a sequence of machine-readable instructions, which can be embodied in a program or software. The instructions can be stored in a memory location, such as the memory. The instructions can be directed to the CPU, which can subsequently program or otherwise configure the CPU to implement methods of the present disclosure. The CPU an be part of a circuit, such as an integrated circuit. One or more other components of the system can be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC). [000165] The storage unit can store files, such as drivers, libraries and saved programs. The storage unit can store user data, e.g., user preferences and user programs. The computer system in some cases can include one or more additional data storage units that are external to the computer system, such as located on a remote server that is in communication with the computer system through an intranet or the Internet. [000166] The computer system can communicate with one or more remote computer systems through the network. For instance, the computer system can communicate with a remote computer system of a user. Examples of remote computer systems include personal computers (e.g., portable PC), slate or tablet PC's (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, Smart phones (e.g., Apple® iPhone, Android-enabled device, Blackberry®), or personal digital assistants. The user can access the computer system via the network. 46 168948691.1[000167] Methods as described herein can be implemented by way of machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system such as, for example, on the memory or electronic storage unit. The machine executable or machine-readable code can be provided in the form of software. During use, the code can be executed by the CPU. In some cases, the code can be retrieved from the storage unit and stored on the memory for ready access by the CPU. In some situations, the electronic storage unit can be precluded, and machine-executable instructions are stored on memory. [000168] The code can be pre-compiled and configured for use with a machine having a processer adapted to execute the code or can be compiled during runtime. The code can be supplied in a programming language that can be selected to enable the code to execute in a pre-compiled or as compiled fashion. [000169] Aspects of the systems and methods provided herein, such as the computer system, can be embodied in programming. Various aspects of the technology can be thought of as “products” or “articles of manufacture” typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine readable medium. Machine-executable code can be stored on an electronic storage unit, such as memory (e.g., read- only memory, random-access memory, flash memory) or a hard disk. “Storage” type media can include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer into the computer platform of an application server. Thus, another type of media that can bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links or the like, also can be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or 47 168948691.1machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution. [000170] Hence, a machine readable medium, such as computer-executable code, may take many forms, including but not limited to, a tangible storage medium, a carrier wave medium, or physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or the like, such as can be used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. [000171] Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include for example: a floppy disk, a flexible disk, hard disk, magnetic tape, any other magnetic medium, a CD-ROM, DVD or DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium with patterns of holes, a RAM, a ROM, a PROM and EPROM, a FLASH- EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media can be involved in carrying one or more sequences of one or more instructions to a processor for execution. [000172] The computer system can include or be in communication with an electronic display that comprises a user interface (UI) for providing, for example, a current stage of processing or assaying of a sample (e.g., a particular step, such as a lysis step, or sequencing step that is being performed). Inputs are received by the computer system from one or more measurement. Examples of UIs include, without limitation, a graphical user interface (GUI) and web-based user interface. The algorithm can, for example, process and / or assay a sample, perform sequencing analysis, measure sets of values representative of classes of molecules, identify sets of 48 168948691.1features and feature vectors from assay data, process feature vectors using a machine learning model to obtain output classifications, and train a machine learning model (e.g., iteratively search for optimal values of parameters of the machine learning model). [000173] In some embodiments, systems capable of executing one or more algorithms, e.g., laptops, desktops, iPads, mobile devices etc., for determining changes in cfDNA mutation profiles, frequency of mutations and / or fragmentation profiles classifies the subject as a cancer patient based on the cfDNA mutation profiles, frequency of mutations and / or fragmentation for the subject. These systems further execute machine learning algorithms that can be used to generate models such as, for example, high-risk populations and low-risk general populations (a penalized logistic regression with the Mathios et al. (Mathios D, Johansen JS, Cristiano S, Medina JE, Phallen J, Larsen KR, et al. Detection and characterization of lung cancer using cell-free DNA fragmentomes. Nat Commun 2021;12(1):5060) features as well as coverage from transcription factor binding sites. These models can be trained on the subject cohort with 5-fold cross validation with 10 repeats, and scores for each sample ae calculated by the mean across repeats and evaluated using AUC-ROC. For example, the first model used the high-risk non-cancer and HCC patients while the second used the non-cancer individuals without liver pathology. The locked high-risk model trained on the cohort was applied to a second and different cohort to generate cancer predictions on an external validation set. A “class label” can be applied to each sample indicating the classification of the sample for any number of input features. For example, the class labels for the set of cohorts could indicate the identity of cfDNA mutation profiles, frequency of mutations and / or fragmentation profiles based on genomic location etc. The resulting training sets are provided to machine learning unit, such as a neural network or a support vector machine. Using the training set, the machine learning unit may generate a model to classify the sample according to the cfDNA mutation profiles, frequency of mutations and / or fragmentation profile. [000174] In some embodiments, a method is provided for creating a trained classifier, comprising the steps of: (a) providing a plurality of different classes, wherein each class represents a set of subjects with a shared characteristic (e.g. from one or more cohorts); (b) providing a multi- parametric model representative of the cell-free DNA molecules from each of a plurality of samples belonging to each of the classes, thereby providing a training data set; and (c) training a 49 168948691.1learning algorithm on the training data set to create one or more trained classifiers, wherein each trained classifier classifies a test sample into one or more of the plurality of classes. [000175] As an example, a trained classifier may use a learning algorithm selected from the group consisting of a random forest, a neural network, a support vector machine, and a linear classifier. Each of the plurality of different classes may be selected from the group consisting of healthy, breast cancer, colon cancer, lung cancer, pancreatic cancer, prostate cancer, ovarian cancer, melanoma, and liver cancer. [000176] A trained classifier may be applied to a method of classifying a sample from a subject. This method of classifying may comprise: (a) providing a multi-parametric model representative of the cell-free DNA molecules from a test sample from the subject; and (b) classifying the test sample using a trained classifier. After the test sample is classified into one or more classes, a therapeutic intervention on the subject can be performed based on the classification of the sample. [000177] In some embodiments, training sets are provided to a machine learning unit, such as a neural network or a support vector machine. Using the training set, the machine learning unit may generate a model to classify the sample according to a treatment response to one or more therapeutic inventions. This is also referred to as “calling”. The model developed may employ information from any part of a test vector. [000178] In general, machine learning can be used to reduce a set of data generated from all (primary sample / analytes / test) combinations into an optimal predictive set of features, e.g., which satisfy specified criteria. In various examples statistical learning, and / or regression analysis can be applied. Simple to complex and small to large models making a variety of modeling assumptions can be applied to the data in a cross-validation paradigm. Simple to complex includes considerations of linearity to non-linearity and non-hierarchical to hierarchical representations of the features. Small to large models includes considerations of the size of basis vector space to project the data onto as well as the number of interactions between features that are included in the modelling process. 50 168948691.1[000179] Machine learning techniques can be used to assess the commercial testing modalities most optimal for cost / performance / commercial reach as defined in the initial question. A threshold check can be performed: If the method applied to a hold-out dataset that was not used in cross validation surpasses the initialized constraints, then the assay is locked, and production initiated. For example, a threshold for assay performance may include a desired minimum accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), or a combination thereof. For example, a desired minimum accuracy, PPV, NPV, clinical sensitivity, clinical specificity, or combination thereof may be at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. As another example, a desired minimum AUC may be at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.81, at least about 0.82, at least about 0.83, at least about 0.84, at least about 0.85, at least about 0.86, at least about 0.87, at least about 0.88, at least about 0.89, at least about 0.90, at least about 0.91, at least about 0.92, at least about 0.93, at least about 0.94, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99. A subset of assays may be selected from a set of assays to be performed on a given sample based on the total cost of performing the subset of assays, subject to the threshold for assay performance, such as desired minimum accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), and a combination thereof. If the thresholds are not met, then the assay engineering procedure can loop back to either the constraint setting for possible relaxation or to the wet lab to change the parameters in which data was acquired. Given the clinical question, biological constraints, budget, lab machines, etc., can constrain the problem. [000180] In certain embodiments, the computer processing of a machine learning technique can include method(s) of statistics, mathematics, biology, or any combination thereof. In various 51 168948691.1examples, any one of the computer processing methods can include a dimension reduction method, logistic regression, dimension reduction, principal component analysis, autoencoders, singular value decomposition, Fourier bases, singular value decomposition, wavelets, discriminant analysis, support vector machine, tree-based methods, random forest, gradient boost tree, logistic regression, matrix factorization, network clustering, statistical testing and neural network. [000181] In certain embodiments, the computer processing of a machine learning technique can include logistic regression, multiple linear regression (MLR), dimension reduction, partial least squares (PLS) regression, principal component regression, autoencoders, variational autoencoders, singular value decomposition, Fourier bases, wavelets, discriminant analysis, support vector machine, decision tree, classification and regression trees (CART), tree-based methods, random forest, gradient boost tree, logistic regression, matrix factorization, multidimensional scaling (MDS), dimensionality reduction methods, t-distributed stochastic neighbor embedding (t-SNE), multilayer perceptron (MLP), network clustering, neuro-fuzzy, neural networks (shallow and deep), artificial neural networks, Pearson product-moment correlation coefficient, Spearman's rank correlation coefficient, Kendall tau rank correlation coefficient, or any combination thereof. In some examples, the computer processing method is a supervised machine learning method including, for example, a regression, support vector machine, tree-based method, and neural network. In some examples, the computer processing method is an unsupervised machine learning method including, for example, clustering, network, principal component analysis, and matrix factorization. [000182] For supervised learning, training samples (e.g., in thousands) can include measured data (e.g., of various analytes) and known labels, which may be determined via other time- consuming processes, such as imaging of the subject and analysis by a trained practitioner. Example labels can include classification of a subject, e.g., discrete classification of whether a subject has cancer or not or continuous classifications providing a probability (e.g., a risk or a score) of a discrete value. A learning module can optimize parameters of a model such that a quality metric (e.g., accuracy of prediction to known label) is achieved with one or more specified criteria. Determining a quality metric can be implemented for any arbitrary function including the set of all risk, loss, utility, and decision functions. A gradient can be used in conjunction with a 52 168948691.1learning step (e.g., a measure of how much the parameters of the model should be updated for a given time step of the optimization process). [000183] As described above, examples can be used for a variety of purposes. For example, plasma (or other sample) can be collected from subjects symptomatic with a condition (e.g., known to have the condition) and healthy subjects. Genetic data (e.g., cfDNA) can be acquired analyzed to obtain a variety of different features, which can include features based on a genome wide analysis. These features can form a feature space that is searched, stretched, rotated, translated, and linearly or non-linearly transformed to generate an accurate machine learning model, which can differentiate between healthy subjects and subjects with the condition (e.g., identify a disease or non-disease status of a subject). Output derived from this data and model (which may include probabilities of the condition, stages (levels) of the condition, or other values), can be used to generate another model that can be used to recommend further procedures, e.g., recommend a biopsy or keep monitoring the subject condition. [000184] In some embodiments, DNA from a population of several individuals can be analyzed by a set of multiplexed arrays. The data for each multiplexed array may be self- normalized using the information contained in that specific array. This normalization algorithm may adjust for nominal intensity variations observed in the two-color channels, background differences between the channels, and possible crosstalk between the dyes. The behavior of each base position may then be modeled using a clustering algorithm that incorporates several biological heuristics on mutation profiles, frequency of mutations and / or fragmentation profiles. In cases where few cfDNA fragments are observed (e.g., due to low minor-allele frequency), locations and shapes of the missing sequences may be estimated using neural networks. Depending on the profiles and percent sequence identity, a statistical score may be devised (a Training score). A score such as GenCall Score is designed to mimic evaluations made by a human expert's visual and cognitive systems. In addition, it has been evolved using the genotyping data from top and bottom strands. This score may be combined with several penalty terms (e.g., low intensity, mismatch between existing and predicted cfDNA fragments) in order to make up the Training score. The Training score is saved for use by the calling algorithm. 53 168948691.1[000185] To call a therapeutic response, a calling algorithm may take the genetic information and treatment responses of a plurality of individuals having a disease or condition. The data may first be normalized (using the same procedure as for the clustering algorithm). The calling operation (classification) may be performed using, for example, a Bayesian model. The score for each call's Call Score can be the product of a Training Score and a data-to-model fit score. After scoring all the treatment responses, the application may compute a composite score. [000186] In some embodiments, a training dataset comprises clinical data selected from the group consisting of cancer stage, type of surgical procedure, age, tumor grading, depth of tumor infiltration, occurrence of post-operative complications, and the presence of venous invasion. In some embodiments, the training dataset is pre-processed, comprising transforming the provided data into class-conditional probabilities. [000187] Another embodiment uses machine learning techniques to train a statistical classifier, specifically a support vector machine, for each cancer stage category based on word occurrences in a corpus of histology reports for each patient. New reports can then be classified according to the most likely stage, facilitating the collection and analysis of population staging data. [000188] In some embodiments, a machine learning algorithm is selected from the group consisting of a supervised or unsupervised learning algorithm selected from support vector machine, random forest, nearest neighbor analysis, linear regression, binary decision tree, discriminant analyses, logistic classifier, and cluster analysis. [000189] In general, a system can comprise a report generator for reporting on cancer test results and treatment options. The report generator system can be a central data processing system configured to establish communications directly with a remote data site or laboratory, a medical practice / healthcare provider (treating professional) and / or a patient / subject through communication links. The laboratory can be medical laboratory, diagnostic laboratory, medical facility, medical practice, point-of-care testing device, or any other remote data site capable of generating subject clinical information. Subject clinical information includes but it is not limited to laboratory test data, X-ray data, examination and diagnosis. The healthcare provider or practice 54 168948691.126 includes medical services providers, such as doctors, nurses, home health aides, technicians and physician's assistants, and the practice is any medical care facility staffed with healthcare providers. In certain instances, the healthcare provider / practice is also a remote data site. In a cancer treatment embodiment, the subject may be afflicted with cancer, among others. [000190] Other clinical information for a cancer subject includes the results of laboratory tests, imaging or medical procedure directed towards the specific cancer that one of ordinary skill in the art can readily identify. The list of appropriate sources of clinical information for cancer includes but it is not limited to: CT scan, MRI scan, ultrasound scan, bone scan, PET Scan, bone marrow test, barium X-ray, endoscopy, lymphangiogram, IVU (Intravenous urogram) or IVP (IV pyelogram), lumbar puncture, cystoscopy, immunological tests (anti-malignin antibody screen), and cancer marker tests. [000191] The subject clinical information may be obtained from the laboratory manually or automatically. For simplicity of the system the information is obtained automatically at predetermined or regular time intervals. A regular time interval refers to a time interval at which the collection of the laboratory data is carried out automatically by the methods and systems described herein based on a measurement of time such as hours, days, weeks, months, years etc. In one embodiment of the invention, the collection of data and processing is carried out at least once a day. In one embodiment, the transfer and collection of data is carried out once every month, biweekly, or once a week, or once every couple of days. Alternatively, the retrieval of information may be carried out at predetermined but not regular time intervals. For instance, a first retrieval step may occur after one week and a second retrieval step may occur after one month. The transfer and collection of data can be customized according to the nature of the disorder that is being managed and the frequency of required testing and medical examinations of the subjects. [000192] In certain embodiments, a genetic report is generated from a subject’s sample, e.g. cfDNA. The polynucleotides in a sample can be sequenced, e.g., whole genome sequencing, NGS sequencing, producing a plurality of sequence reads. In some embodiments, genetic information comprises variables defining the genomic organization of cancer cells or the genomic organization 55 168948691.1of single disseminated cancer cells. In some embodiments, the genetic information comprises sequence or abundance data from one or more genetic loci in cell-free DNA from the individuals. [000193] cfDNA genetic information is processed. Genetic variants can also be identified. Genetic variants include sequence variants, copy number variants and nucleotide modification variants. A sequence variant is a variation in a genetic nucleotide sequence. A copy number variant is a deviation from wild type in the number of copies of a portion of a genome. Genetic variants include, for example, single nucleotide variations (SNPs), insertions, deletions, inversions, transversions, translocations, gene fusions, chromosome fusions, gene truncations, copy number variations (e.g., aneuploidy, partial aneuploidy, polyploidy, gene amplification), abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns and abnormal changes in nucleic acid methylation. The process then determines the frequency of genetic variants in the sample containing the genetic material. Since this process is noisy, the process separates information from noise. The sensitivity of detecting genetic variants can be increased by increasing read depth of polynucleotides (e.g., by sequencing to a greater read depth at in a sample from a subject at two or more time points). [000194] To increase the diagnosis confidence, a plurality of measurements can be taken. Or alternatively using measurements at a plurality of time points (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more time points) to determine whether cancer is advancing, in remission or stabilized. The diagnostic confidence can be used to identify disease states. For example, cell free polynucleotides taken from a subject can include polynucleotides derived from normal cells, as well as polynucleotides derived from diseased cells, such as cancer cells. Polynucleotides from cancer cells may bear genetic variants, such as somatic cell mutations and copy number variants. When cell free polynucleotides from a sample from a subject are sequenced, and cfDNA mutation profiles, frequency of mutations and / or fragmentation profiles can be produced as described in the examples section which follows. [000195] Numerous cancers may be detected using the methods and systems described herein. Cancers cells, as most cells, can be characterized by a rate of turnover, in which old cells die and replaced by newer cells. Generally dead cells, in contact with vasculature in a given subject, may release DNA or fragments of DNA into the blood stream. This is also true of cancer 56 168948691.1cells during various stages of the disease. Cancer cells may also be characterized, dependent on the stage of the disease, by various genetic aberrations such as copy number variation as well as mutations. This phenomenon may be used to detect the presence or absence of cancers individuals using the methods and systems described herein. [000196] Additionally, the systems and methods described herein may also be used to help characterize certain cancers. Genetic data produced from the system and methods of this disclosure may allow practitioners to help better characterize a specific form of cancer. Often times, cancers are heterogeneous in both composition and staging. Genetic profile data may allow characterization of specific sub-types of cancer that may be important in the diagnosis or treatment of that specific sub-type. This information may also provide a subject or practitioner clues regarding the prognosis of a specific type of cancer. [000197] The systems and methods provided herein may be used to monitor already known cancers, or other diseases in a particular subject. This may allow either a subject or practitioner to adapt treatment options in accord with the progress of the disease. In this example, the systems and methods described herein may be used to construct genetic cfDNA mutation profiles, frequency of mutations and / or fragmentation profiles of a particular subject of the course of the disease. In some instances, cancers can progress, becoming more aggressive and genetically unstable. In other examples, cancers may remain benign, inactive or dormant. The system and methods of this disclosure may be useful in determining disease progression. [000198] Further, the systems and methods described herein may be useful in determining the efficacy of a particular treatment option. In one example, certain treatment options may be correlated with genetic cfDNA mutation profiles, frequency of mutations and / or fragmentation profiles of cancers over time. This correlation may be useful in selecting a therapy. Additionally, if a cancer is observed to be in remission after treatment, the systems and methods described herein may be useful in monitoring residual disease or recurrence of disease. [000199] Further, the methods of the disclosure may be used to characterize the heterogeneity of an abnormal condition in a subject, the method comprising generating a cfDNA mutation profile, frequency of mutations and / or fragmentation profile of extracellular polynucleotides in the 57 168948691.1subject, wherein the cfDNA mutation profile comprises a plurality of data resulting from profile variation and mutation analyses. In some cases, including but not limited to cancer, a disease may be heterogeneous. Disease cells may not be identical. In the example of cancer, some tumors are known to comprise different types of tumor cells, some cells in different stages of the cancer. In other examples, heterogeneity may comprise multiple foci of disease. Again, in the example of cancer, there may be multiple tumor foci, perhaps where one or more foci are the result of metastases that have spread from a primary site (also known as distant metastases). [000200] The methods of this disclosure may be used to generate a profile, fingerprint, or set of data that is a summation of genetic information derived from different cells in a heterogeneous disease. This set of data may comprise copy number variation and mutation analyses alone or in combination. [000201] Further, these reports are submitted and accessed electronically via the internet. Analysis of data occurs at a site other than the location of the subject. The report is generated and transmitted to the subject's location. Via an internet enabled computer, the subject accesses the reports reflecting his tumor burden. [000202] The annotated information can be used by a health care provider to select other drug treatment options and / or provide information about drug treatment options to an insurance company. The method can include annotating the drug treatment options for a condition in, for example, the NCCN Clinical Practice Guidelines in Oncology™ or the American Society of Clinical Oncology (ASCO) clinical practice guidelines. [000203] Reports are generated, mapping genome positions and cfDNA mutation profile variation for the subject with cancer. These reports, in comparison to other profiles of subjects with known outcomes, can indicate that a particular cancer is aggressive and resistant to treatment. The subject is monitored for a period and retested. If at the end of the period, the cfDNA mutation profiles, frequency of mutations and / or fragmentation variation profile does not vary, this may indicate that the current treatment is not working. A comparison is done with cfDNA mutation profiles of other subjects. For example, if it is determined that a change in cfDNA mutation 58 168948691.1variation indicates that the cancer is advancing, then the original treatment regimen as prescribed is no longer treating the cancer and a new treatment is prescribed. [000204] In certain embodiments, the system receives genetic information from a DNA sequencer. The process then determines specific cfDNA alterations and frequencies thereof. These reports are submitted and accessed electronically via the internet. Analysis of data occurs at a site other than the location of the subject. The report is generated and transmitted to the subject's location. Via an internet enabled computer, the subject accesses the reports reflecting his tumor burden. [000205] While temporal information can be used to enhance the information for cfDNA mutation profiles and frequency of mutations, other consensus methods can be applied. In other embodiments, the historical comparison can be used in conjunction with other consensus cfDNA mutation profiles, frequency of mutations and / or fragmentation profiles. Consensus cfDNA mutation profiles and frequency of mutations can be normalized against control samples. Measures of molecules mapping to reference sequences can also be compared across a genome to identify areas in the genome in which cfDNA mutation profiles and frequency of mutations varies, or remains the same. Consensus methods include, for example, linear or non-linear methods of building consensus cfDNA mutation profiles and frequency of mutations (such as voting, averaging, statistical, maximum a posteriori or maximum likelihood detection, dynamic programming, Bayesian, hidden Markov or support vector machine methods, etc.) derived from digital communication theory, information theory, or bioinformatics. After the sequence read coverage has been determined, a stochastic modeling algorithm is applied to convert the normalized nucleic acid sequence read coverage for each window region to the discrete copy number states. In some cases, this algorithm may comprise one or more of the following: Hidden Markov Model, dynamic programming, support vector machine, Bayesian network, trellis decoding, Viterbi decoding, expectation maximization, Kalman filtering methodologies and neural networks. [000206] Artificial neural networks (NNets) mimic networks of “neurons” based on the neural structure of the brain. They process records one at a time, or in a batch mode, and “learn” 59 168948691.1by comparing their classification of the record (which, at the outset, is largely arbitrary) with the known actual classification of the record. In MLP-NNets, the errors from the initial classification of the first record is fed back into the network, and are used to modify the network's algorithm the second time around, and so on for many iterations. The neural networks use an iterative learning process in which data cases (rows) are presented to the network one at a time, and the weights associated with the input values are adjusted each time. [000207] After all cases are presented, the process often starts over again. During this learning phase, the network learns by adjusting the weights so as to be able to predict the correct class label of input samples. Neural network learning is also referred to as “connectionist learning,” due to connections between the units. Advantages of neural networks include their high tolerance to noisy data, as well as their ability to classify patterns on which they have not been trained. One neural network algorithm is back-propagation algorithm, such as Levenberg-Marquadt. Once a network has been structured for a particular application, that network is ready to be trained. To start this process, the initial weights are chosen randomly. Then the training, or learning, begins. [000208] The network processes the records in the training data one at a time, using the weights and functions in the hidden layers, then compares the resulting outputs against the desired outputs. Errors are then propagated back through the system, causing the system to adjust the weights for application to the next record to be processed. This process occurs over and over as the weights are continually tweaked. During the training of a network the same set of data is processed many times as the connection weights are continually refined. [000209] In an embodiment, the training step of the machine learning unit on the training data set may generate one or more classification models for applying to a test sample. These classification models may be applied to a test sample to predict the response of a subject to a therapeutic intervention. [000210] Comparison of sequence coverage to a control sample or reference sequence may aid in normalization across windows. In this embodiment, cell free DNAs are extracted and isolated from a readily accessible bodily fluid such as blood. For example, cell free DNAs can be extracted using a variety of methods known in the art, including but not limited to isopropanol 60 168948691.1precipitation and / or silica based purification. Cell free DNAs may be extracted from any number of subjects, such as subjects without cancer, subjects at risk for cancer, or subjects known to have cancer (e.g. through other means). [000211] Following the isolation / extraction step, any of a number of different sequencing operations may be performed on the cell free polynucleotide sample. Samples may be processed before sequencing with one or more reagents (e.g., enzymes, unique identifiers (e.g., barcodes), probes, etc.). In some cases, if the sample is processed with a unique identifier such as a barcode, the samples or fragments of samples may be tagged individually or in subgroups with the unique identifier. The tagged sample may then be used in a downstream application such as a sequencing reaction by which individual molecules may be tracked to parent molecules. [000212] The cell free polynucleotides can be tagged or tracked in order to permit subsequent identification and origin of the particular polynucleotide. The assignment of an identifier (e.g., a barcode) to individual or subgroups of polynucleotides may allow for a unique identity to be assigned to individual sequences or fragments of sequences. This may allow acquisition of data from individual samples and is not limited to averages of samples. In some examples, nucleic acids or other molecules derived from a single strand may share a common tag or identifier and therefore may be later identified as being derived from that strand. Similarly, all of the fragments from a single strand of nucleic acid may be tagged with the same identifier or tag, thereby permitting subsequent identification of fragments from the parent strand. In other cases, gene expression products (e.g., mRNA) may be tagged in order to quantify expression, by which the barcode, or the barcode in combination with sequence to which it is attached can be counted. In still other cases, the systems and methods can be used as a PCR amplification control. In such cases, multiple amplification products from a PCR reaction can be tagged with the same tag or identifier. If the products are later sequenced and demonstrate sequence differences, differences among products with the same identifier can then be attributed to PCR error. Additionally, individual sequences may be identified based upon characteristics of sequence data for the read themselves. For example, the detection of unique sequence data at the beginning (start) and end (stop) portions of individual sequencing reads may be used, alone or in combination, with the length, or number of base pairs of each sequence read unique sequence to assign unique identities to individual 61 168948691.1molecules. Fragments from a single strand of nucleic acid, having been assigned a unique identity, may thereby permit subsequent identification of fragments from the parent strand. This can be used in conjunction with bottlenecking the initial starting genetic material to limit diversity. [000213] Generally, the methods and systems provided herein are useful for preparation of cell free polynucleotide sequences to a down-stream application sequencing reaction. Often, a sequencing method is next generation sequencing (NGS), classic Sanger sequencing, whole- genome bisulfite sequencing (WGSB), small-RNA sequencing, low-coverage Whole-Genome Sequencing (lcWGS), etc. [000214] As used herein, the term “sequencing” refers to any of a number of technologies used to determine the sequence of a biomolecule, e.g., a nucleic acid such as DNA or RNA. Exemplary sequencing methods include, but are not limited to, targeted sequencing, single molecule real-time sequencing, exon sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy termination sequencing, whole-genome sequencing, sequencing by hybridization, pyrosequencing, capillary electrophoresis, gel electrophoresis, duplex sequencing, cycle sequencing, single-base extension sequencing, solid-phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification at lower denaturation temperature-PCR (COLD-PCR), multiplex PCR, sequencing by reversible dye terminator, paired-end sequencing, near-term sequencing, exonuclease sequencing, sequencing by ligation, short-read sequencing, single-molecule sequencing, sequencing-by-synthesis, real-time sequencing, reverse-terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, and a combination thereof. In some embodiments, sequencing can be performer by a gene analyzer such as, for example, gene analyzers commercially available from Illumina or Applied Biosystems. In some embodiments, the sequencing method can be massively parallel sequencing, that is, simultaneously (or in rapid succession) sequencing any of at least 100, 1000, 10,000, 100,000, 1 million, 10 million, 100 million, or 1 billion polynucleotide molecules. 62 168948691.1[000215] After sequencing, reads are assigned a quality score. A quality score may be a representation of reads that indicates whether those reads may be useful in subsequent analysis based on a threshold. In some cases, some reads are not of sufficient quality or length to perform the subsequent mapping step. Sequencing reads with a quality score at least 90%, 95%, 99%, 99.9%, 99.99% or 99.999% may be filtered out of the data set. In other cases, sequencing reads assigned a quality scored at least 90%, 95%, 99%, 99.9%, 99.99% or 99.999% may be filtered out of the data set. The genomic fragment reads that meet a specified quality score threshold are mapped to a reference genome, or a reference sequence that is known not to contain mutations. After mapping alignment, sequence reads are assigned a mapping score. A mapping score may be a representation or reads mapped back to the reference sequence indicating whether each position is or is not uniquely mappable. In instances, reads may be sequences unrelated to mutation analysis. For example, some sequence reads may originate from contaminant polynucleotides. Sequencing reads with a mapping score at least 90%, 95%, 99%, 99.9%, 99.99% or 99.999% may be filtered out of the data set. In other cases, sequencing reads assigned a mapping scored less than 90%, 95%, 99%, 99.9%, 99.99% or 99.999% may be filtered out of the data set. For each mappable base, bases that do not meet the minimum threshold for mappability, or low quality bases, may be replaced by the corresponding bases as found in the reference sequence. [000216] Numerous diseases may be detected. It was shown here that sensitive and specific fragmentomic assays, including those that measure the size, distribution of cfDNA fragments, and their repeat element content genome-wide, can be used to detect a range of disease states including those involving neurodegeneration, chronic inflammation, and dysregulation of cellular growth. It was further shown that these changes illuminate mechanisms of disease and reflect disease-specific genomic and epigenomic alterations. In certain embodiments, the diseases comprise: neurodegenerative diseases (e.g. Alzheimer’s disease), inflammatory diseases, cancer, non- neoplastic dysregulation of cellular growth and the like. Non-neoplastic dysregulation of cellular growth diseases includes Liver Cirrhosis and benign adnexal masses. [000217] Additionally, the systems and methods described herein may also be used to help characterize certain cancers. Genetic data produced from the system and methods of this disclosure may allow practitioners to help better characterize a specific form of cancer. Often times, cancers 63 168948691.1are heterogeneous in both composition and staging. Genetic profile data may allow characterization of specific sub-types of cancer that may be important in the diagnosis or treatment of that specific sub-type. This information may also provide a subject or practitioner clues regarding the prognosis of a specific type of cancer. [000218] The systems and methods provided herein may be used to monitor already known cancers, or other diseases in a particular subject. This may allow either a subject or practitioner to adapt treatment options in accord with the progress of the disease. In this example, the systems and methods described herein may be used to construct genetic profiles of a particular subject of the course of the disease. In some instances, cancers can progress, becoming more aggressive and genetically unstable. In other examples, cancers may remain benign, inactive or dormant. The system and methods of this disclosure may be useful in determining disease progression. [000219] Further, the systems and methods described herein may be useful in determining the efficacy of a particular treatment option. In one example, successful treatment options may actually increase the amount of copy number variation or mutations detected in subject's blood if the treatment is successful as more cancers may die and shed DNA. In other examples, this may not occur. In another example, perhaps certain treatment options may be correlated with genetic profiles of cancers over time. This correlation may be useful in selecting a therapy. Additionally, if a cancer is observed to be in remission after treatment, the systems and methods described herein may be useful in monitoring residual disease or recurrence of disease. [000220] The data is sent over a direct connection or over the internet to a computer for processing. The data processing aspects of the system can be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or in combinations of them. Data processing apparatus of the invention can be implemented in a computer program product tangibly embodied in a machine-readable storage device for execution by a programmable processor; and data processing method steps of the invention can be performed by a programmable processor executing a program of instructions to perform functions of the invention by operating on input data and generating output. The data processing aspects of the invention can be implemented advantageously in one or more computer programs that are executable on a programmable system 64 168948691.1including at least one programmable processor coupled to receive data and instructions from and to transmit data and instructions to a data storage system, at least one input device, and at least one output device. Each computer program can be implemented in a high-level procedural or object- oriented programming language, or in assembly or machine language, if desired; and, in any case, the language can be a compiled or interpreted language. Suitable processors include, by way of example, both general and special purpose microprocessors. Generally, a processor will receive instructions and data from a read-only memory and / or a random access memory. Storage devices suitable for tangibly embodying computer program instructions and data include all forms of nonvolatile memory, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM disks. Any of the foregoing can be supplemented by, or incorporated in, ASICs (application-specific integrated circuits). [000221] To provide for interaction with a user, the methods can be implemented using a computer system having a display device such as a monitor or LCD (liquid crystal display) screen for displaying information to the user and input devices by which the user can provide input to the computer system such as a keyboard, a two-dimensional pointing device such as a mouse or a trackball, or a three-dimensional pointing device such as a data glove or a gyroscopic mouse. The computer system can be programmed to provide a graphical user interface through which computer programs interact with users. The computer system can be programmed to provide a virtual reality, three-dimensional display interface. [000222] While various embodiments of the present disclosure have been described above, it should be understood that they have been presented by way of example only, and not limitation. Numerous changes to the disclosed embodiments can be made in accordance with the disclosure herein without departing from the spirit or scope of the invention. Thus, the breadth and scope of the present disclosure should not be limited by any of the above-described embodiments. [000223] EXAMPLES 65 168948691.1[000224] In the following Examples 1-3, certain of the protocols, materials (including plasma samples) are overlapping as indicated, but vary as specified including in collected and analyzed data. [000225] The following examples have been included to provide guidance to one of ordinary skill in the art for practicing representative embodiments of the presently disclosed subject matter. In light of the present disclosure and the general level of skill in the art, those of skill can appreciate that the following examples are intended to be exemplary only and that numerous changes, modifications, and alterations can be employed without departing from the scope of the presently disclosed subject matter. The synthetic descriptions and specific examples that follow are only intended for the purposes of illustration, and are not to be construed as limiting in any manner to make compounds of the disclosure by other methods. [000226] EXAMPLE 1: NON-INVASIVE DETECTION OF HUMAN DISEASES USING CELL- FREE DNA FRAGMENTOMES [000227] The fragmentome is evaluable using low-coverage whole-genome sequencing and machine learning and does not rely on prior knowledge of specific disease-associated mutations or epigenetic changes. Since the fragmentome reflects an individual’s physiological state – comprising cfDNA from both healthy cells and cells with disease-specific changes – it was reasoned herein that it could be useful for detecting illnesses beyond cancer. Ultimately, this signal could be used for non-invasive earlier disease diagnoses and may be well-suited for the study of diseases with complex etiologies and which lack effective biomarkers. [000228] In this study, evidence is presented for changes to the fragmentome across broad categories of disease phenotypes including neurodegeneration, chronic inflammation, and non- neoplastic dysregulation of cellular growth, using three representative examples for which there currently are no effective blood-based screening methods: Alzheimer’s disease (AD), liver cirrhosis, and benign adnexal masses. Further demonstrated are methodologies by which these changes can be measured and used in the diagnosis of these illnesses, and to elucidate relationships to both tissue-specific and immune-mediated mechanisms of disease. 66 168948691.1[000229] METHODS [000230] Study Populations [000231] Plasma samples for patients with Alzheimer’s Disease (n=70) and healthy controls (n=97) were obtained from ProteoGenex, Inc. Plasma samples from 78 patients with liver cirrhosis were described in our previous study2and included samples collected prospectively as part of HCC biomarker registry and the AIDS Linked to the IntraVenous Experience (ALIVE) study at the Johns Hopkins University School of Medicine, under protocols approved by the Johns Hopkins Institutional Review Board or retrospectively collected by BioIVT (Westbury, NY). Plasma samples from patients with benign adnexal masses (n=195) were obtained from the Netherlands Cancer Institute (NKI, originally collected for trial NL58253.031.16). Plasma samples from 293 healthy were described in our previous studies2,7. These samples were originally obtained from two screening clinical trial cohorts for colorectal cancer in Denmark (Endoscopy III) and the Netherlands (COCOS, Netherlands Trial Register ID NTR182946). The protocol for the Endoscopy III Project was approved by the Regional Ethics Committee and the Danish Data Protection Agency, and for the COCOS trial, ethical approval was obtained from the Dutch Health Council. The inclusion criteria for both the Dutch and the Danish cohorts were any individuals of age 50–75 eligible for colorectal cancer screening. All patients used had either a FIT negative test or a negative colonoscopy result. All patients provided written informed consent and the studies were performed according to the Declaration of Helsinki. [000232] cfDNA extraction and sequencing [000233] The protocols for blood collection, cfDNA extraction and sequencing have been previously described1,2. In brief, whole blood was collected in EDTA or Streck tubes, and plasma was separated by centrifugation and aliquoted into EDTA tubes. cfDNA was isolated from 2-4mL of plasma, and next-generation sequencing libraries were prepared using 15ng of cfDNA when available, or the entire purified amount if less than 15ng were available. All libraries underwent four cycles of PCR amplification. All cfDNA data used for modeling was sequenced at 1-10x coverage on the Illumina NovaSeq6000 platform using 100bp paired-end runs. 67 168948691.1[000234] Prior to alignment, adapter sequences were filtered from reads using the fastp software. Sequence reads were aligned against the hg19 human reference genome using Bowtie2 and duplicate reads were removed using Sambamba. Post-alignment, each aligned pair was converted to a genomic interval representing the sequenced DNA fragment using bedtools. Only reads with a MAPQ score of at least 30 or greater were retained. Read pairs were further filtered if overlapping the Duke Excluded Regions blacklist (genome.ucsc.edu / cgi- bin / hgTrackUi?db=hg19&g=wgEncodeMapability). To capture large-scale epigenetic differences in fragmentation across the genome estimable from low-coverage WGS, the hg19 reference genome was tiled into non-overlapping 5 Mb bins. Bins with an average GC content < 0.3 and an average mappability < 0.9 were excluded, leaving 473 bins spanning approximately 2.4 GB of the genome. Following Mathios et al.1, GC correction was performed independently for short (< 150 bp) and long (>= 150 bp) cfDNA fragments using an external panel of 20 individuals without cancer sequenced on a NovaSeq to generate a target distribution. [000235] Characterization of the cfDNA fragmentome [000236] DELFI was employed to compute genome-wide fragmentation features, as described in the Mathios et al. study1Briefly, after alignment to the hg19 reference genome and performance of a fragment level GC correction using an external panel of 20 individuals without cancer, the total coverage and the ratio of short to long fragments were calculated for 473 non- overlapping 5 MB bins across the genome. Z-scores representing arm gains / losses were calculated for the 39 autosomal chromosome arms. ARTEMIS, an alignment-free approach to localizing repeat elements27was also employed. Briefly, 1.2 billion kmers that map to 1280 unique repeat element types were counted in all sequencing reads for each sample. These kmer counts were aggregated for each repeat type and normalized to aligned coverage. This generated a kmer repeat landscape for each sample. Aligned coverage in 5611Mb bins genome-wide corresponding to regions with a high density of epigenetic marks was further calculated. [000237] Disease detection models [000238] Modeling approaches1,2,7,27to train models for detection of Alzheimer’s disease, benign adnexal masses and liver cirrhosis were used. For the DELFI model, the principal 68 168948691.1components of the fragmentation ratios representing greater than 90% of variance and the z-scores were used to train Lasso logistic regression models. For ARTEMIS, repeat element features were partitioned -into 6 families (LINEs, SINEs, Satellites, LTRs, RNA / DNA Elements, and high epigenetic density bins) and centered and scaled the features within each family in each sample. A penalized logistic regression (PLR) model was then trained for each family, and ensembled them together with another PLR. For the joint ARTEMIS-DELFI model, 3 models were further trained to represent the three families of fragmentation features (ratios, z-scores and coverage) and ensembled these with the ARTEMIS models. The models were ensembled with leave-one-out cross validation with nested 5-fold cross validation to train each learner. The scores obtained from the DELFI, ARTEMIS and ARTEMIS-DELFI models were evaluated for disease detection in the three example diseases studied (Alzheimer’s disease, benign adnexal masses and liver cirrhosis). Each locked model was applied to patients with the other diseases to demonstrate the disease- specific nature of each model. [000239] Gene Coverage Analyses [000240] We computed GC-corrected cfDNA coverage in 100kb bins overlapping all known genes and used these coverage metrics to generate AUCs for each gene for the classification of AD patients from healthy patients, and Liver Cirrhosis patients from healthy patients. We used these AUCs to rank the genes and performed GSEA analyses using KEGG gene sets41. We further analyzed protein interactions between genes using Cytoscape and STRING. [000241] Epigenetic Analyses [000242] We downloaded Histone CHIP-Seq data for white blood cells from ENCODE and identified 100kb regions with activating (H3K27ac) and repressive (H3K27me3) states in B lymphocytes, T lymphocytes, Monocytes and Neutrophils. We then computed GC corrected coverage in these bins from the plasma of each cfDNA sample. [000243] RESULTS [000244] Plasma was obtained from patients with Alzheimer’s disease (n=70), liver cirrhosis (n=78), or benign adnexal masses (n=195) as well as from healthy individuals without these conditions (n=390) (FIG 1). cfDNA fragments were extracted and whole-genome 69 168948691.1sequencing performed at 1X-10X coverage depth. The concentration of cfDNA found in the plasma was markedly higher for patients with AD as compared to patients without disease (p < 2.2e16, Wilcoxon’s test, FIG.6), providing evidence that AD pathology influences the composition of the fragmentome. [000245] Though cfDNA levels were not appreciably higher for patients with liver cirrhosis or benign adnexal masses, it was observed that genome-wide cfDNA fragmentation profiles were altered in all three disease states. Analysis of the fraction of short (100-150 bp) to long (151-220 bp) cfDNA fragments in each of 4735Mb bins genome-wide revealed greater heterogeneity among patients with liver cirrhosis and Alzheimer’s disease, as compared to healthy patients (24.4% of liver cirrhosis patients, 21.4% of Alzheimer’s disease patients, and 5.13% of patients with benign adnexal mass had a Pearson’s correlation coefficient to the median healthy of less than 0.9 as compared to 6.92% of individuals without these conditions; FIG.2A). Differences in the genome-wide representation of repeat elements were observed, including those from the SINE, LINE, LTR, Satellite, and RNA / DNA Element families, providing evidence that cfDNA changes related to disease may occur in diverse genomic elements including in repetitive regions outside of known genes (FIG.7). [000246] Notably, these cfDNA changes appeared to represent disease-specific signals, perhaps reflecting both target organ contributions to the cfDNA as well as immune mediated changes to white blood cell (WBC) content in cfDNA. Patterns of variance between individual’s cfDNA fragment coverage in 5 Mb bins genome wide varied by disease state and were overall higher in individuals with liver cirrhosis, benign adnexal masses and Alzheimer’s disease than in healthy individuals (mean bin-level variance 0.0717, 0.0326 and 0.0692 for patients with liver cirrhosis, benign adnexal mass or Alzheimer’s disease, as compared to 0.0298 healthy patients; FIG.2B). [000247] As a proof-of-concept for the use of cfDNA fragmentation as a biomarker for neurodegeneration, chronic inflammation, and dysregulation of cellular growth, the fragmentation profile (DELFI) and kmer repeat landscape (ARTEMIS) were used to train and cross-validate machine learning models for detection of each disease (FIGS.3A, 3B, 3C). ARTEMIS, DELFI and a joint ARTEMIS-DELFI model identified individuals with or without 70 168948691.1disease with high performance (area under the curve, AUCs >0.85) for all pathologies and models studied. Different families of features drove classification performance for each disease (FIG.8), and the locked disease-specific models did not identify patients with other diseases (FIGS.3D, 3E, 3F) suggesting that cfDNA fragmentation reflects disease-specific changes in the circulation. [000248] To understand the biological mechanisms by which fragmentomic changes occur, we performed several disease-specific analyses. Increased extracellular release of mtDNA (mitochondrial DNA) has been observed in AD, likely as a stress signal that initiates other immune-mediated sequelae26. Accordingly, we observed an increased fraction of mitochondrial DNA-derived cfDNA fragments in patients with AD as compared to healthy patients (p=1.41e11, FIG.4A). We next analyzed cfDNA coverage across canonical AD genes and found differences between plasma from patients with and without Alzheimer’s Disease at the APP, APOE and PSEN1 loci (p<1.0e-5, FIG.9). We and others2,3,27,28have previously shown that these coverage differences reflect genome-wide epigenetic and chromatin changes. In a gene-set enrichment analysis of coverage differences at all known genes, we observed an enrichment in genes part of the ether lipid and alpha linolenic acid metabolism pathways (FIG.4B). These pathways are highly relevant to Alzheimer’s Disease and have been hypothesized to be involved in maintenance of blood-brain-barrier integrity, amyloid protein processing, and other related functions29,30,31,32,33,34,35,36. Genes in these pathways that we inferred to be implicated in Alzheimer’s Disease based on cfDNA changes displayed numerous protein interactions with genes previously linked to Alzheimer’s Disease (FIG.4C). Finally, we observed changes to the genome-wide representation of repeat elements from diverse classes including LINEs, SINEs and Transposable Elements, which may be consistent with the general observation that the state of repeat elements may be altered in AD via Tau mediated pathways37. [000249] In patients with and without liver cirrhosis, we observed changes in cfDNA coverage at genes highly relevant to cirrhosis (FIG.5A) such as NOTCH1 and SLC11A1, and in many other genes with protein interactions with these known cirrhosis genes. We also observed repeat element changes in the plasma of patients with and without liver cirrhosis that demonstrated progressive changes between healthy patients and patients with viral hepatitis but 71 168948691.1no known cirrhosis, and then further changes between these high-risk patients and those with cirrhosis (FIG.5B). [000250] In addition to these liver disease specific changes, we also observed changes to cfDNA fragmentation reflecting immune-mediated disease mechanisms. cfDNA is known to be predominantly derived from cells of white blood cell (WBC) lineages, and pro-inflammatory phenotypes, such as cirrhosis, alter the proportions of various WBC types found in the blood. We measured cfDNA coverage at regions of activating and repressive histone marks in B and T lymphocytes, monocytes and neutrophils and observed changes suggesting proportional increases in neutrophils, monocytes and B lymphocytes as compared to T lymphocytes (FIG. 5C). This is concordant with observed T lymphopenia in Cirrhosis38, and elevations to inflammatory markers such as the neutrophil to lymphocyte ratio (NLR) and monocyte to lymphocyte ratio (MLR) that characterize inflammatory states like cirrhosis39,40. Taken together, these results suggest that changes in the fragmentomes of patients with diseases characterized by neurodegeneration, chronic inflammation, and dysregulation of cellular growth are a result of both tissue-specific and immune-mediated mechanisms of disease. [000251] DISCUSSION [000252] This study provides an initial demonstration of the ways in which disease pathology alters the cfDNA fragmentome. It was shown here that sensitive and specific fragmentomic assays, including those that measure the size, distribution of cfDNA fragments, and their repeat element content genome-wide, can be used to detect a range of disease states including those involving neurodegeneration, chronic inflammation, and dysregulation of cellular growth. It was further shown that these changes illuminate mechanisms of disease and reflect disease-specific genomic and epigenomic alterations. It is likely that these changes reflect a combination of disease-specific genomic alterations in target tissues and systemic changes to white blood cells. Before clinical use, these detection models will need to be externally validated in larger cohorts as well as in additional disease pathologies. Nevertheless, this study has shown that changes in cfDNA fragmentomes characterize diverse disease phenotypes including in non- neoplastic conditions. The genomic location, extent, and type of these alterations varies by disease state, and may be used for non-invasive, early detection of multiple pathologies. 72 168948691.1[000253] REFERENCES 1. D. Mathios, J. S. Johansen, S. Cristiano, J. E. Medina, J. Phallen, K. R. Larsen, D. C. Bruhm, N. Niknafs, L. Ferreira, V. Adleff, J. Y. Chiao, A. Leal, M. Noe, J. R. White, A. S. Arun, C. Hruban, A. V. Annapragada, S. Jensen, M. W. Ørntoft, A. H. Madsen, B. Carvalho, M. de Wit, J. Carey, N. C. Dracopoli, T. Maddala, K. C. Fang, A. R. Hartman, P. M. Forde, V. Anagnostou, J. R. Brahmer, R. J. A. Fijneman, H. J. Nielsen, G. A. Meijer, C. L. Andersen, A. Mellemgaard, S. E. Bojesen, R. B. Scharpf, V. E. Velculescu, Detection and characterization of lung cancer using cell-free DNA fragmentomes. Nat Commun 12, 5060 (2021). 2. Z. H. Foda, A. V. Annapragada, K. Boyapati, D. C. Bruhm, N. A. Vulpescu, J. E. Medina, D. Mathios, S. Cristiano, N. Niknafs, H. T. Luu, M. G. Goggins, R. A. Anders, J. Sun, S. H. Meta, D. L. Thomas, G. D. Kirk, V. Adleff, J. Phallen, R. B. Scharpf, A. K. Kim, V. E. Velculescu, Detecting Liver Cancer Using Cell-Free DNA Fragmentomes. Cancer Discov 13, 616–631 (2023). 3. D. C. Bruhm, D. Mathios, Z. H. Foda, A. V. Annapragada, J. E. Medina, V. Adleff, E. J. Chiao, L. Ferreira, S. Cristiano, J. R. White, S. A. Mazzilli, E. Billatos, A. Spira, A. H. Zaidi, J. Mueller, A. K. Kim, V. Anagnostou, J. Phallen, R. B. Scharpf, V. E. Velculescu, Single- molecule genome-wide mutation profiles of cell-free DNA for non-invasive detection of cancer. Nat. Genet.55, 1301–1310 (2023). 4. K. C. A. Chan, P. Jiang, K. Sun, Y. K. Y. Cheng, Y. K. Tong, S. H. Cheng, A. I. C. Wong, I. Hudecova, T. Y. Leung, R. W. K. Chiu, Y. M. D. Lo, Second generation noninvasive fetal genome analysis reveals de novo mutations, single-base parental inheritance, and preferred DNA ends. Proc National Acad Sci 113, E8159–E8168 (2016). 5. Q. Zhou, G. Kang, P. Jiang, R. Qiao, W. K. J. Lam, S. C. Y. Yu, M.-J. L. Ma, L. Ji, S. H. Cheng, W. Gai, W. Peng, H. Shang, R. W. Y. Chan, S. L. Chan, G. L. H. Wong, L. T. Hiraki, S. Volpi, V. W. S. Wong, J. Wong, R. W. K. Chiu, K. C. A. Chan, Y. M. D. Lo, Epigenetic analysis of cell-free DNA by fragmentomic profiling. Proc National Acad Sci 119, e2209852119 (2022). 6. M. W. Snyder, M. Kircher, A. J. Hill, R. M. Daza, J. Shendure, Cell-free DNA Comprises an In Vivo Nucleosome Footprint that Informs Its Tissues-Of-Origin. Cell 164, 57–68 (2016). 7. S. Cristiano, A. Leal, J. Phallen, J. Fiksel, V. Adleff, D. C. Bruhm, S. Ø. Jensen, J. E. Medina, C. Hruban, J. R. White, D. N. Palsgrove, N. Niknafs, V. Anagnostou, P. Forde, J. Naidoo, K. Marrone, J. Brahmer, B. D. Woodward, H. Husain, K. L. van Rooijen, M.-B. W. Ørntoft, A. H. Madsen, C. J. H. van de Velde, M. Verheij, A. Cats, C. J. A. Punt, G. R. Vink, N. C. T. van Grieken, M. Koopman, R. J. A. Fijneman, J. S. Johansen, H. J. Nielsen, G. A. Meijer, C. L. Andersen, R. B. Scharpf, V. E. Velculescu, Genome-wide cell-free DNA fragmentation in patients with cancer. Nature 570, 385–389 (2019). 73 168948691.18. J. Phallen, M. Sausen, V. Adleff, A. Leal, C. Hruban, J. White, V. Anagnostou, J. Fiksel, S. Cristiano, E. Papp, S. Speir, T. Reinert, M.-B. W. Orntoft, B. D. Woodward, D. Murphy, S. Parpart-Li, D. Riley, M. Nesselbush, N. Sengamalay, A. Georgiadis, Q. K. Li, M. R. Madsen, F. V. Mortensen, J. Huiskens, C. Punt, N. van Grieken, R. Fijneman, G. Meijer, H. Husain, R. B. Scharpf, L. A. DiazJr, S. Jones, S. Angiuoli, T. Ørntoft, H. J. Nielsen, C. L. Andersen, V. E. Velculescu, Direct detection of early-stage cancers using circulating tumor DNA. Sci Transl Med 9, eaan2415 (2017). 9. X. Li, X. Feng, X. Sun, N. Hou, F. Han, Y. Liu, Global, regional, and national burden of Alzheimer’s disease and other dementias, 1990–2019. Front. Aging Neurosci.14, 937486 (2022). 10. R. J. Bateman, P. S. Aisen, B. D. Strooper, N. C. Fox, C. A. Lemere, J. M. Ringman, S. Salloway, R. A. Sperling, M. Windisch, C. Xiong, Autosomal-dominant Alzheimer’s disease: a review and proposal for the prevention of Alzheimer’s disease. Alzheimer’s Res. Ther.3, 1 (2011). 11. D. S. Knopman, H. Amieva, R. C. Petersen, G. Chételat, D. M. Holtzman, B. T. Hyman, R. A. Nixon, D. T. Jones, Alzheimer disease. Nat. Rev. Dis. Prim.7, 33 (2021). 12. S. Gunes, Y. Aizawa, T. Sugashi, M. Sugimoto, P. P. Rodrigues, Biomarkers for Alzheimer’s Disease in the Current State: A Narrative Review. Int. J. Mol. Sci.23, 4962 (2022). 13. B. A. Santamaria, I. B. Luquin, J. Corroza, M. S. Miguel, M. Roldan, S. Zueco, I. E. Beiras, C. Cabello, M. Á. G. Agudo, I. Jericó, E. Erro, M. Mendioroz, Liquid biopsy shows differences in cfDNA fragmentation pattern between AD patients and controls. Alzheimer’s Dement.16 (2020), doi:10.1002 / alz.039748. 14. V. Haage, P. L. D. Jager, Neuroimmune contributions to Alzheimer’s disease: a focus on human data. Mol. Psychiatry 27, 3164–3181 (2022). 15. J. W. Kinney, S. M. Bemiller, A. S. Murtishaw, A. M. Leisgang, A. M. Salazar, B. T. Lamb, Inflammation as a central mechanism in Alzheimer’s disease. Alzheimer’s Dement.: Transl. Res. Clin. Interv.4, 575–590 (2018). 16. H. Devarbhavi, S. K. Asrani, J. P. Arab, Y. A. Nartey, E. Pose, P. S. Kamath, Global burden of liver disease: 2023 update. J. Hepatol.79, 516–537 (2023). 17. K. Onyirioha, S. Mittal, A. G. Singal, Is hepatocellular carcinoma surveillance in high- risk populations effective? Hepatic Oncol.7, HEP25 (2020). 18. A. G. Singal, E. Zhang, M. Narasimman, N. E. Rich, A. K. Waljee, Y. Hoshida, J. D. Yang, M. Reig, G. Cabibbo, P. Nahon, N. D. Parikh, J. A. Marrero, HCC surveillance 74 168948691.1improves early detection, curative treatment receipt, and survival in patients with cirrhosis: A meta-analysis. J. Hepatol.77, 128–139 (2022). 19. E. B. Tapper, N. D. Parikh, Diagnosis and Management of Cirrhosis and Its Complications. JAMA 329, 1589–1602 (2023). 20. S. F. Brunner, N. D. Roberts, L. A. Wylie, L. Moore, S. J. Aitken, S. E. Davies, M. A. Sanders, P. Ellis, C. Alder, Y. Hooks, F. Abascal, M. R. Stratton, I. Martincorena, M. Hoare, P. J. Campbell, Somatic mutations and clonal dynamics in healthy and cirrhotic human liver. Nature 574, 538–542 (2019). 21. S. Aizawa, G. Brar, H. Tsukamoto, Cell Death and Liver Disease. Gut Liver 14, 20–29 (2020). 22. B. Bullock, L. Larkin, L. Turker, K. Stampler, Management of the Adnexal Mass: Considerations for the Family Medicine Physician. Front. Med.9, 913549 (2022). 23. J. Carvalho, R. Moretti-Marques, A. Filho, Adnexal mass: diagnosis and management. Rev. Bras. Ginecol. e Obstet. RBGO - Gynecol. Obstet.42, 438–443 (2020). 24. D.-A. Clevert, C. Nyhsen, P. Ricci, P. S. Sidhu, C. Tziakouri, M. Radziņa, V. Cantisani, A. P. Brady, Position statement and best practice recommendations on the imaging use of ultrasound from the European Society of Radiology ultrasound subcommittee. Insights Imaging 11, 115 (2020). 25. P. Charkhchi, C. Cybulski, J. Gronwald, F. O. Wong, S. A. Narod, M. R. Akbari, CA125 and Ovarian Cancer: A Comprehensive Review. Cancers 12, 3730 (2020). 26. I. K. Gorham, R. C. Barber, H. P. Jones, N. R. Phillips, Mitochondrial SOS: how mtDNA may act as a stress signal in Alzheimer’s disease. Alzheimer’s Res. Ther.15, 171 (2023). 27. A. Annapragada, N. Niknafs, J. R. White, D. C. Bruhm, C. Cherry, J. E. Medina, V. Adleff, C. Hruban, D. Mathios, Z. H. Foda, J. Phallen, R. B. Scharpf, V. E. Velculescu, Genome-wide repeat landscapes in cancer and cell-free DNA. Science Translational Medicine, In Press . 28. S. Cristiano, A. Leal, J. Phallen, J. Fiksel, V. Adleff, D. C. Bruhm, S. Jensen, J. E. Medina, C. Hruban, J. R. White, D. N. Palsgrove, N. Niknafs, V. Anagnostou, P. Forde, J. Naidoo, K. Marrone, J. Brahmer, B. D. Woodward, H. Husain, K. L. van Rooijen, M. W. Ørntoft, A. H. Madsen, C. J. H. van de Velde, M. Verheij, A. Cats, C. J. A. Punt, G. R. Vink, N. C. T. van Grieken, M. Koopman, R. J. A. Fijneman, J. S. Johansen, H. J. Nielsen, G. A. Meijer, C. L. Andersen, R. B. Scharpf, V. E. Velculescu, Genome-wide cell-free DNA fragmentation in patients with cancer. Nature 570, 385–389 (2019). 75 168948691.129. Oh. Y. Kim, J. Song, Important roles of linoleic acid and α-linolenic acid in regulating cognitive impairment and neuropsychiatric issues in metabolic-related dementia. Life Sci. 337, 122356 (2024). 30. O. Binyamin, K. Nitzan, K. Frid, Y. Ungar, H. Rosenmann, R. Gabizon, Brain targeting of 9c,11t-Conjugated Linoleic Acid, a natural calpain inhibitor, preserves memory and reduces Aβ and P25 accumulation in 5XFAD mice. Sci. Rep.9, 18437 (2019). 31. F. Dakterzada, M. Jové, R. Huerto, A. Carnes, J. Sol, R. Pamplona, G. Piñol-Ripoll, Changes in Plasma Neutral and Ether-Linked Lipids Are Associated with The Pathology and Progression of Alzheimer’s Disease. Aging Dis.14, 1728–1738 (2023). 32. K. Huynh, W. L. F. Lim, C. Giles, K. S. Jayawardana, A. Salim, N. A. Mellett, A. A. T. Smith, G. Olshansky, B. G. Drew, P. Chatterjee, I. Martins, S. M. Laws, A. I. Bush, C. C. Rowe, V. L. Villemagne, D. Ames, C. L. Masters, M. Arnold, K. Nho, A. J. Saykin, R. Baillie, X. Han, R. Kaddurah-Daouk, R. N. Martins, P. J. Meikle, Concordant peripheral lipidome signatures in two large clinical studies of Alzheimer’s disease. Nat. Commun.11, 5698 (2020). 33. L. D. Spears, S. Adak, G. Dong, X. Wei, G. Spyropoulos, Q. Zhang, L. Yin, C. Feng, D. Hu, I. J. Lodhi, F.-F. Hsu, R. Rajagopal, K. K. Noguchi, C. M. Halabi, L. Brier, A. R. Bice, B. V. Lananna, E. S. Musiek, O. Avraham, V. Cavalli, J. K. Holth, D. M. Holtzman, D. F. Wozniak, J. P. Culver, C. F. Semenkovich, Endothelial ether lipids link the vasculature to blood pressure, behavior, and neurodegeneration. J. Lipid Res.62, 100079 (2021). 34. M. Jové, N. Mota-Martorell, È. Obis, J. Sol, M. Martín-Garí, I. Ferrer, M. Portero-Otin, R. Pamplona, Ether Lipid-Mediated Antioxidant Defense in Alzheimer’s Disease. Antioxidants 12, 293 (2023). 35. T. Wang, K. Huynh, C. Giles, N. A. Mellett, T. Duong, A. Nguyen, W. L. F. Lim, A. A. Smith, G. Olshansky, G. Cadby, J. Hung, J. Hui, J. Beilby, G. F. Watts, P. Chatterjee, I. Martins, S. M. Laws, A. I. Bush, C. C. Rowe, V. L. Villemagne, D. Ames, C. L. Masters, K. Taddei, V. Doré, J. Fripp, M. Arnold, G. Kastenmüller, K. Nho, A. J. Saykin, R. Baillie, X. Han, R. N. Martins, E. K. Moses, R. Kaddurah‐Daouk, P. J. Meikle, APOE ε2 resilience for Alzheimer’s disease is mediated by plasma lipid species: Analysis of three independent cohort studies. Alzheimer’s Dement.18, 2151–2166 (2022). 36. A. Leikin-Frenkel, M. S. Beeri, I. Cooper, How Alpha Linolenic Acid May Sustain Blood–Brain Barrier Integrity and Boost Brain Resilience against Alzheimer’s Disease. Nutrients 14, 5091 (2022). 37. C. Guo, H.-H. Jeong, Y.-C. Hsieh, H.-U. Klein, D. A. Bennett, P. L. D. Jager, Z. Liu, J. M. Shulman, Tau Activates Transposable Elements in Alzheimer’s Disease. Cell Rep.23, 2874–2880 (2018). 76 168948691.138. L. Lombardo, A. Capaldi, G. Poccardi, P. Vineis, Peripheral blood CD3 and CD4 T- lymphocyte reduction correlates with severity of liver cirrhosis. Int. J. Clin. Lab. Res.25, 153–156 (1995). 39. X. Li, J. Wu, W. Mao, Evaluation of the neutrophil‐to‐lymphocyte ratio, monocyte‐to‐ lymphocyte ratio, and red cell distribution width for the prediction of prognosis of patients with hepatitis B virus‐related decompensated cirrhosis. J. Clin. Lab. Anal.34, e23478 (2020). 40. D. Li, W. Sun, L. Chen, J. Gu, H. Wu, H. Xu, J. Gan, Utility of neutrophil–lymphocyte ratio and platelet–lymphocyte ratio in predicting acute-on-chronic liver failure survival. Open Life Sci.18, 20220644 (2023). 41. M. Kanehisa, S. Goto, KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic Acids Res 28, 27–30 (2000). [000254] EXAMPLE 2 [000255] To distinguish between individuals with high and low CCI, we trained a machine learning model on features from the DELFI approach, including the ratio of short (100-150 bp) to long (151-220 bp) cfDNA fragments in each of 473 genome-wide 5 Mb bins as well as chromosomal arm-level changes. Using 10-fold cross-validation, the resulting scores, referred to as the “fragmentation morbidity index”, provide a molecular measure of broad disease burden for each individual. The fragmentation morbidity index separated individuals with high and low CCI (median 0.64 for high CCI vs.0.5 for low CCI, p=9.2e-13, Wilcoxon rank-sum test) (FIG.10D) and correlated with overall survival (median overall survival 2.8 years vs. not reached for individuals in the top vs. bottom decile of scores, p<0.001; p=0.036 for individuals with scores above vs. below the mean) (FIG.10, FIG.15D). In a multivariate cox proportional hazards model, the fragmentation morbidity index was an independent predictor of survival even after adjusting for clinical characteristics, including age, inflammatory markers and CCI (FIG. 10F). The interaction between CCI and fragmentation morbidity index explained variation in survival as compared to CCI alone, regardless of adjustment for age (Analysis of deviance, X2> 2.0, degrees of freedom=0, p<2.2e-16 for tests with and without age adjustment). These analyses suggest that genome-wide cfDNA fragmentomes may serve as molecular biomarkers for a wide range of 77 168948691.1diseases not previously evaluated, and have the potential to provide information about disease burden that is complementary to clinical covariates. [000256] Development of disease-specific fragmentome biomarkers for detection of liver cirrhosis [000257] Given the pan-disease observations of genomic changes in cfDNA fragmentomes in the MICA cohort, we evaluated the potential of this approach for detection of liver cirrhosis and advanced fibrosis as these represent illnesses with diagnostic challenges and an urgent need for improved non-invasive biomarkers. We evaluated blood samples in a Discovery Cohort (n=426) of individuals who were healthy or who had liver disease that we previously analyzed as non- cancer controls for detection of hepatocellular carcinoma (3). These included 348 individuals in a screening population (F0 / F1 fibrosis staging), 12 individuals with advanced fibrosis (F3 staging), and 66 individuals with liver cirrhosis (F4 staging) (FIG.16, Table 2). A subset of individuals had documented metabolic risk factors (n=55, arteriosclerosis, hyperlipidemia, diabetes, hypertension, or hypercholesterolemia) or had high-risk conditions for liver cirrhosis (33) (n=55, viral hepatitis B or C infection). All patients with liver cirrhosis or fibrosis were diagnosed by expert clinical evaluation as well as imaging including liver elastography and ultrasound. For all participants, blood was collected, cfDNA was extracted, and genomic libraries were prepared and sequenced to an average coverage of 6.6x. The concentration of cfDNA and overall fraction of short cfDNA fragments was elevated in patients with cirrhosis as compared to patients in the screening population (median concentration 7.78 ng / ml vs. 5.17 ng / ml, p=0.02, Wilcoxon rank-sum test; peak increase in fragments at 140bp with a concomitant peak decrease at 177bp, p<2.2x10-16for both lengths, Wilcoxon rank sum test) (FIGS.11A and 11B). [000258] Given the expected changes in genomic and epigenomic changes in this disease, we utilized both the DELFI fragmentation (1) and ARTEMIS analyses of repeat and epigenetic landscapes (2) in genome-wide cfDNA sequencing data to train, cross-validate, and externally validate a noninvasive liquid biopsy classifier for liver cirrhosis and advanced fibrosis. We found that analyses of fragmentation profiles across the genome revealed greater heterogeneity among patients with liver cirrhosis, as compared to healthy populations (FIGS. 11C and 11D). (2, 34). Similarly, we observed widespread differences in the genome-wide representation of repeat 78 168948691.1elements (FIG.11C and 11D) and in cfDNA coverage at epigenetic hotspots in individuals with liver cirrhosis and advanced fibrosis compared to healthy populations (FIG 11, C and D). [000259] We designed an ARTEMIS-DELFI approach for detection of liver cirrhosis (LCr) and advanced fibrosis in the Discovery Cohort that ensembled machine learners for fragmentation profiles, as well as epigenetic and repeat element features (FIG. 11D, FIG. 17) using a nested cross-validation procedure with penalized logistic regression to favor models with lower feature dimensionality (FIG.18). We evaluated whether the resulting ARTEMIS-DELFI LCr scores were affected by clinical or demographic characteristics in individuals without disease. We found that scores did not differ by age, sex, or the presence of other comorbidities like hypertension and diabetes (FIG. 19A), suggesting that a disease-specific screening model would be robust to variation among healthy individuals. [000260] We next examined the relationship between ARTEMIS-DELFI LCr scores and disease status. Scores for the 348 individuals in the screening population without fibrosis or cirrhosis were low, regardless of the presence of metabolic risk factors or viral hepatitis (median scores of 0.06, 0.06, and 0.2 for those in the general, metabolic risk, and viral hepatitis subsets, respectively) (FIG.11E). In contrast, patients with liver cirrhosis had significantly higher scores (median scores of 0.54, p<2.2x10-16, Wilcoxon rank-sum test) and could be distinguished from individuals in the screening population (AUC=0.91, 95% CI 0.87-0.95) (FIG.11F). The cirrhosis signal was similar in both men and women and was not related to the potential presence of chromosomal changes linked to undiagnosed hepatic malignancy (FIG.19B). ARTEMIS-DELFI LCr scores were increased in individuals with higher Child-Pugh cirrhosis scores, consistent with representation of increased disease severity, although even liver cirrhosis patients with lower Child-Pugh scores had ARTEMIS-DELFI LCr scores that were significantly higher than in the screening population (p<6.0x10-7for each Child-Pugh stage, Wilcoxon rank-sum test) (FIG.11E, FIG. 19B). Encouragingly, advanced fibrosis (F3) could also be detected by the ARTEMIS- DELFI LCr score (median score 0.68, AUC=0.96, 95% CI 0.94-0.99), suggesting the approach may be useful for detection of pre-cirrhotic liver disease (26). [000261] External validation of disease-specific liquid biopsies for liver cirrhosis 79 168948691.1[000262] Following generation of a locked classifier using the entire Discovery Cohort, we selected score cutpoints corresponding to high (90%) and moderate (50%) specificity in the Discovery Cohort and then evaluated the classifier in a Validation Cohort. The Validation Cohort included patients with liver cirrhosis (n=83), advanced fibrosis (n=6), metabolic dysfunction– associated steatotic liver disease (MASLD) (n=24, S2 or S3 steatosis) (20, 35), or a screening population (n=121, including n=31 with metabolic risk factors). Using the locked ARTEMIS- DELFI LCr classifier trained in the Discovery Cohort, we found that patients in the Validation Cohort with liver cirrhosis exhibited the same patterns of altered cfDNA fragmentation (FIG.20). ARTEMIS-DELFI LCr scores were higher for patients with cirrhosis (median score 0.58) compared to individuals from the population in the general and metabolic risk group (LCr median scores of 0.07 and 0.09, respectively) (FIG.3A, FIG.21). Scores were also higher in individuals with advanced fibrosis (LCr median score=0.3, p=0.19, Wilcoxon signed rank-sum test) and MASLD (LCr median score=0.14, p=0.027, Wilcoxon signed rank-sum test) (FIG.12A). Overall, we observed a sensitivity of 71% in the Validation Cohort and a specificity of 81% at the high specificity cutpoint, corresponding to a modest attenuation from the 73% sensitivity in the Discovery cohort (FIG.12B). At the lower specificity cutpoint potentially more suitable for high- risk populations, we observed 93% sensitivity for cirrhosis detection and 43% specificity, consistent with the 97% sensitivity in the Discovery Cohort (FIG.22). [000263] For a subset of individuals in the Validation Cohort where blood lab values for the Fibrosis-4 (FIB-4) index (27) biomarker were available (n=71), we compared our ARTEMIS- DELFI LCr score to this clinically used test. Our ARTEMIS-DELFI LCr score identified 57.1% and 60% of individuals with liver cirrhosis and advanced fibrosis, respectively, while FIB-4 identified 54.8% and 20% at the clinically validated rule-in score cutpoint of 3.25 (27) (FIG.12C. We next examined the potential for the ARTEMIS-DELFI LCr score to improve non-invasive early detection of cirrhosis, among individuals with an indeterminate FIB-4 score (below 3.25, but above the rule-out cut-off of 1.45). Here we employed the moderate specificity ARTEMIS-DELFI LCr score cutpoint locked at a 50% specificity threshold in the discovery cohort, reasoning that this lower specificity would be appropriate for this group already deemed to be at elevated risk by FIB-4. Among this subset of FIB-4 indeterminate individuals, the ARTEMIS-DELFI LCr score 80 168948691.1correctly classified 62.5% and 50% of individuals in the general screening and metabolic factors groups as not having liver disease, while 100% and 93.3% of individuals were accurately identified as having advanced fibrosis or liver cirrhosis, respectively. A combined test improved the overall sensitivity for cirrhosis detection to 88.1%. These observations suggest that the ARTEMIS-DELFI LCr model has the potential to be used for screening of liver cirrhosis or the sensitive triage of at- risk individuals in populations with metabolic risk factors or pre-existing liver disease (FIG.23). [000264] To evaluate the cross-reactivity of our screening approaches in other conditions that may also alter cfDNA fragmentomes, we analyzed blood samples from patients with liver cancer (n=78), lung cancer (n=126), benign lung nodules (n=67), or chronic pancreatitis (n=18) (1, 3). Analysis of these individuals using the locked ARTEMIS-DELFI LCr classifier revealed scores that were significantly lower than those for individuals with liver cirrhosis in the Validation Cohort (median scores of 0.06, 0.05, and 0.06, for chronic pancreatitis, benign lung nodules, and lung cancer, respectively, compared to median score of 0.58 for liver cirrhosis; p<3x10-7for all diseases, Wilcoxon rank-sum test) (FIG. 12E). The ARTEMIS-DELFI LCr score was elevated in individuals with liver cancer (median score 0.37 for liver cancer vs.0.56 for liver cirrhosis, p=0.21, Wilcoxon rank-sum test), likely reflecting the cirrhosis that precedes many hepatocellular carcinomas. Conversely, use of our previous cfDNA fragmentome prototype classifiers to detect liver or lung cancers (2, 3) resulted in significantly lower scores for liver cirrhosis patients (median score 0.13 for liver cirrhosis vs. 0.76 for liver cancer, p=1.19x10-14; median score 0.05 for liver cirrhosis vs.0.89 for lung cancer, p<2.2x10-16; Wilcoxon rank-sum test) (FIG.24) that were similar to those of individuals without cancer. These observations highlight that the ARTEMIS-DELFI classifier incorporated features of cfDNA fragmentomes that reflected liver disease-specific changes in the circulation that may be used for liver cirrhosis detection. [000265] Disease-specific mechanisms of changes in cfDNA fragmentomes [000266] To understand the biological mechanisms by which fragmentomic changes occur, we performed disease-specific analyses of cfDNA for liver cirrhosis patients. We used three orthogonal approaches to analyze genome-wide cfDNA coverage differences which we and others (3, 6, 8, 9, 36–39) have previously shown to reflect epigenetic and chromatin changes. cfDNA coverage changes at individual loci have been shown to correlate with gene expression, as actively 81 168948691.1transcribing regions of the genome typically have more open or active chromatin and are consequently more susceptible to degradation in the circulation (3, 13, 38). First, in patients with liver cirrhosis in the Discovery Cohort, we observed an enrichment of significant differences in cfDNA coverage in genes related to liver cirrhosis (40–44) (FIG. 13A, p<0.05 after Bonferroni correction). A subset of these genes, including notch receptor receptor 1 (NOTCH1), oxidoreductase HSD3B7, acyltransferase AGPAT2, glucosylceramidase beta 1 (GBA), and TERF1-interacting nuclear factor 2 (TINF2), were observed to remain enriched in the plasma of patients with liver cirrhosis in the Validation Cohort, while no genes were enriched when comparing groups of randomly chosen healthy individuals (FIG.13A). The altered regulation of these genes as revealed through cfDNA analyses is consistent with mechanisms of cirrhosis development. As an example, NOTCH1 normally plays a key role in cellular reprogramming and differentiation, and its dysregulation may promote hepatic stellate cell activation, excessive extracellular matrix production, and fibrosis (45, 46), while AGPAT2 has been implicated in lipodystrophy syndromes and development of steatosis(47, 48) . [000267] Second, we observed changes in cfDNA fragmentation reflecting immune- mediated disease mechanisms. cfDNA is known to be predominantly derived from cells of white blood cell (WBC) lineages, and pro-inflammatory phenotypes, such as cirrhosis, may alter the proportions of various WBC types found in the blood. We measured cfDNA coverage at regions of activating and repressive histone marks in T lymphocytes, monocytes and neutrophils and observed increased coverage in regions of genome-wide chromatin that are active in T lymphocytes but repressed in monocytes or neutrophils. This change suggests a proportional increase in monocytes and neutrophils as compared to T lymphocytes, leading to increased representation of DNA originating from these cell lineages in the circulation of individuals with liver cirrhosis (FIG.13B, FIG. 25, p=5.3x10-7and p=2.2x10-5, respectively, Wilcoxon rank-sum test). This signal was maintained in the Validation Cohort (p=8.6x10-13and p=6.9x10-10for monocytes and neutrophils, Wilcoxon rank-sum test). This is concordant with observed T lymphopenia (49), and elevations to inflammatory markers such as the neutrophil to lymphocyte ratio (NLR) and monocyte to lymphocyte ratio (MLR) in patients with liver cirrhosis (50, 51). 82 168948691.1[000268] Finally, we used our Deconvolution of CfDNA sIgnals via transcription Factor- informed fragment Representation (DECIFER) method (52) to independently confirm the observations above of disease-mediated changes evident in the plasma. We used DECIFER to compare changes in expression profiles of transcription factors in various organ tissues (53) (n=53) as compared to whole blood, to cfDNA coverage at transcription factor binding sites (TFBS) in our cohorts. Given the inverse relationship expected between expression and plasma representation (3, 13, 38), we would expect a negative correlation (with lower coverage due to higher expression) between plasma and tissue profiles, if a given tissue is contributing fragments to the cfDNA. Among all 53 normal tissues in the GTEx database (53), when comparing individuals with liver cirrhosis or advanced fibrosis to screening population individuals, the strongest anticorrelation observed was to liver tissue (FIG. 13C). Taken together, our results suggest that changes in the fragmentomes of patients with liver cirrhosis are a result of disease-specific epigenetic, transcriptomic, and immune-mediated mechanisms of disease. [000269] DISCUSSION. This study provides an initial demonstration of the ways in which non-cancer disease states alter cfDNA fragmentomes. We have previously developed methods for characterizing an individual’s cfDNA fragmentomes (1–3, 6, 9, 13, 14), including through genome-wide analyses of size, coverage distribution, and repeat element content of cfDNA fragments, providing a summary of the chromatin structure of the cell from which they originate. Using these methods as well as novel disease-specific analyses of cfDNA fragmentation, we show that changes to the cfDNA fragmentome are present across different diseases, can serve as a measure of clinical morbidity, and may be used to noninvasively detect liver cirrhosis. We have further shown that these changes reflect both target organ contributions to the cfDNA as well as immune mediated changes to white blood cell (WBC) content in cfDNA. Our results suggest that the number and types of alterations in cfDNA fragmentation may be broader than previously considered and likely include a range of disease processes beyond those in which the fragmentome has historically been studied (4, 5, 9). The impending demographic shift towards aging populations in the United States and other countries, and the corresponding increase in age-related morbidities (54), underscores the clinical unmet need for accessible biomarkers to enhance early detection across a wide spectrum of human diseases. 83 168948691.1[000270] The liquid biopsy approach for detection of liver cirrhosis is presented here as an initial prototype, but its performance provides a promising avenue alongside imaging and other biomarkers. Estimates suggest that 1 in 4 adults globally may have MASLD (19), and 4 in 5 adults in the US have at least one metabolic risk factor for liver cirrhosis (21). This high prevalence of pre-cirrhotic liver disease and associated risk factors illustrates the need for developing accessible biomarkers for screening and surveillance (26) . The only currently available screens for liver fibrosis and cirrhosis are blood-based Fibrosis-4 index (FIB-4), AST to Platelet Ratio Index (APRI), and the NAFLD Fibrosis Score (NFS), all of which have limited sensitivities, or image- based liver elastography, either by ultrasound or MRI, which is not broadly accessible (27). Blood- based cfDNA fragmentomic biomarkers could enhance and expand current screening paradigms for liver cirrhosis. [000271] Despite the encouraging results for fragmentomic detection of disease, our study has several limitations. While the samples in our cohorts were largely obtained via prospective collections, a subset were obtained from case-control collections that may not accurately reflect the characteristics of these diseases in the general population. A potential concern could also be that these approaches detect diseases other than liver cirrhosis. Our cross-reactivity and tissue of origin analyses suggest that the model and underlying features are specific to liver disease, but further work is needed to validate the classifier in additional diseases and larger patient groups, and to further characterize the organ-specific origins of altered cfDNA fragmentomes. An important next step will be large prospective clinical trials of the intended use populations to validate these observations. [000272] Nevertheless, this study has shown that changes in cfDNA fragmentomes are evident in diverse disease phenotypes including liver cirrhosis and advanced fibrosis. The genomic location, extent, and type of these alterations varies by disease state, and may be used for non-invasive screening of multiple pathologies with high performance and limited cross-reactivity. The unique specificity of the approach and the ability to extract genome-wide fragmentome information related to multiple diseases from a single blood sample could enable a facile and accessible liquid biopsy test for non-invasive, multi-disease screening. [000273] METHODS 84 168948691.1[000274] Study population for MICA Morbidity Cohort [000275] Plasma samples from 570 individuals suspected of having serious illness were included in the MICA Morbidity cohort (FIG.14). These samples were originally obtained from individuals in the prospective MICA biomarker study performed at the Diagnostic Outpatient Clinic at Herlev and Gentofte Hospital, Copenhagen University Hospital, Denmark (29). The MICA study enrolled all individuals over the age of 18 presenting to the Diagnostic Outpatient Clinic via the Cancer Patient Pathway (n=757) from July 2016 to July 2019, though in this study we analyzed samples only from individuals without cancer (n=570). This pathway is for patients presenting with non-specific signs and symptoms of cancer and suspected of having cancer or other serious illness. The most common symptoms at time of referral were weight loss, fatigue, pain or nausea. All patients were examined by medical doctors from the Diagnostic Outpatient Clinic at Herlev and Gentofte Hospital, and cancer diagnosis was based on radiological findings and tissue biopsies. Comorbidities were assessed by Charlson comorbidity index (CCI). The MICA study was approved by the Regional Ethics Committee (H-7-2014-011) and the Danish Data Protection Agency (HEH-2014-105). All participants were informed and signed an informed consent declaration before participation. [000276] Study populations for liver cirrhosis cohorts [000277] We analyzed plasma samples from 660 individuals in separate Discovery and Validation Cohorts for liver cirrhosis detection (FIG. 16). Plasma samples from patients in the Discovery Cohort with liver cirrhosis (n=66), advanced fibrosis (n=12), viral hepatitis (n=55), metabolic risk factors (n=55) or in a general screening population (n=238) were described in our previous study (3). We utilized sequencing data from all individuals without cancer in the training set of that previous study, including samples collected prospectively as part of the HCC biomarker registry or AIDS Linked to the IntraVenous Experience (ALIVE) study at the Johns Hopkins University Hospital (Baltimore, MD, USA), retrospectively collected by BioIVT (Westbury, NY, USA), or obtained from two screening clinical trial cohorts for colorectal cancer in Denmark (Endoscopy III) and the Netherlands (COCOS, Netherlands Trial Register ID NTR182946). Samples collected at the Johns Hopkins Hospital were collected under protocols approved by the Johns Hopkins Institutional Review Board. The protocol for the Endoscopy III Project was 85 168948691.1approved by the Regional Ethics Committee and the Danish Data Protection Agency. For the COCOS trial, ethical approval was obtained from the Dutch Health Council. The inclusion criteria for both the Dutch and the Danish cohorts were any individuals of age 50–75 eligible for colorectal cancer screening. All patients used had either a FIT negative test or a negative colonoscopy result. Where available, clinical metadata was used to identify individuals in the presumed healthy populations with confirmed metabolic risk factors for liver cirrhosis (arteriosclerosis, diabetes, hypertension). [000278] For the Validation Cohort, samples were obtained from individuals without cancer enrolled in a case-control study with longitudinal follow-up in Guatemala (n=42 patients with liver cirrhosis, n=6 patients with advanced fibrosis, n=24 patients with MASLD, n=15 patients with metabolic risk factors, and n=9 patients in the general screening population), locally led by the Institute of Nutrition of Central America and Panama (INCAP) in Guatemala City, a WHO- research center. INCAP collaborates extensively with physicians at Hospital San Juan de Dios and Hospital Roosevelt, the two large public health referral hospitals in Guatemala City. The Validation Cohort also included samples that were prospectively obtained at the Fundeni Clinical Institute or associated hospitals in Bucharest, Romania under protocols approved by local review board (n=41 patients with liver cirrhosis) or from ProteoGenex (Inglewood, CA, USA) (n=81 patients in a general screening population, n=16 patients with metabolic risk factors). All patients provided written informed consent and the studies were performed according to the Declaration of Helsinki. [000279] Study populations for cross-reactivity analysis [000280] For analyses of cross-reactivity using the trained ARTEMIS-DELFI LCr detection model, we analyzed previously published samples from individuals with hepatocellular carcinoma (HCC, n= 78) prospectively collected through the HCC biomarker registry at the Johns Hopkins University School of Medicine, lung cancer (n=126) or benign lung nodules (n=67) from the prospectively collected LUCAS Cohort(1) at Bispebjerg Hospital in Copenhagen, Denmark, or chronic pancreatitis (n=18) in the BIOPAC cohort (Clinical trials.gov ID: NCT03311776). The BIOPAC study has been approved by the Danish Regional Ethics Committee (BIOPAC: KA- 20060113) and the Danish Data Protection Agency (BIOPAC: P-2020-834). 86 168948691.1[000281] cfDNA extraction and sequencing [000282] The protocols for blood collection, cfDNA extraction and sequencing have been previously described (1, 3). In brief, whole blood was collected during outpatient venous blood draws in EDTA tubes. All collection sites utilized a consistent protocol: within a maximum of two to four hours from collection, tubes were centrifuged at low speed (1000-3000g) for atleast 10 minutes. The plasma portion was transferred to a new tube and spun at 18,000g for 10 minutes. EDTA plasma was aliquoted and stored between -70 and -80 C for cfDNA analysis. All samples were transported on dry ice with included temperature monitor to ensure that temperature was maintained prior to transfer to -80 C freezers in our laboratory. For all samples, cfDNA was isolated from 1-4mL of plasma using the Qiagen QIAamp Circulating Nucleic Acids Kit. Extracted cfDNA was quantified on the Agilent 4150 TapeStation system using the Agilent Cell-free DNA ScreenTape assay. We estimated concentration of extracted cfDNA fragments between 50 and 800bp to quantify only DNA in a nucleosomal length range chosen to exclude potential genomic DNA contamination. [000283] Then, next-generation sequencing libraries were prepared with the NEBNext DNA Library Prep Kit for Illumina using a target of 15ng of cfDNA when available. For samples in the Endoscopy III and COCOS trials, all available cfDNA was used to prepare the library (median library input 23.13 ng). We modified manufacturer’s guidelines to utilize an on-bead AMPure XP approach to minimize sample loss during elution and tube transfer steps, and to size select for cfDNA within a nucleosomal length range (1). Sequencing libraries were quantified on the Agilent 4150 TapeStation system using the D1000 DNA ScreenTape assay. [000284] All libraries underwent four cycles of PCR amplification and were sequenced with estimated 1-10x coverage on the Illumina NovaSeq6000 platform using 100bp paired-end runs. After sequencing, we excluded samples with Q20<80%, Q30<80% or <3e9 bases sequenced. [000285] The MICA Morbidity Cohort was prospectively collected at a single site and processed and sequenced in library batches distinct from the liver cirrhosis cohort. To reduce preanalytical and technical variation in the liver cirrhosis cohorts, samples from patients with and without disease were processed together in next-generation sequencing (NGS) library batches and the processing and sequencing of the Discovery and Validation cohorts was temporally separated 87 168948691.1(FIG.18). Additionally, no samples collected from the same site were used in both the Discovery and Validation Cohorts, to allow for an external validation of the locked model. [000286] Characterization of the cfDNA fragmentomes [000287] We computed genome-wide fragmentation features as previously described (1). Briefly, we filtered adapter sequences from 100 bp paired-end reads using fastp, aligned to the hg19 reference genome using Bowtie2(55), and removed duplicate reads with Sambamba. After alignment, we converted each read pair to a genomic interval using bedtools, and excluded reads with low alignment quality (MAPQ < 30) or that overlapped the Duke Excluded Regions blacklist (https: / / genome.ucsc.edu / cgi-bin / hgTrackUi?db=hg19&g=wgEncodeMapability). We performed a fragment level GC correction against a reference panel of 20 healthy individuals. We then analyzed fragment length distributions in all samples, and summarized fragmentomic features by calculating the ratio of short (100-150 bp) to long (151-220 bp) fragments in 473 non-overlapping 5 MB bins across the genome. We also computed z-scores representing arm gains and losses for the 39 autosomal chromosome arms. [000288] We also computed genome-wide repeat landscapes using an alignment-free approach to localizing repeat elements we previously described (2). Briefly, we counted 1.2 billion kmers that map to 1280 unique repeat element types from the LINE (long interspersed nuclear element), SINE (short interspersed nuclear element), DNA / RNA TE (Transposable Element), LTR (Long terminal repeat) and Satellite families in all sequencing reads for each sample. We then aggregated these kmer counts for each repeat type and normalized to aligned coverage. This generated a kmer repeat landscape for each sample. We utilized a subset of the 1280 elements (n=599) that were observed to be highly reproducible in low coverage sequencing, as estimated by within sample coefficient of variation of the feature among sequencing replicates. [000289] We further calculated aligned coverage in 551 1Mb bins (2) genome-wide corresponding to regions with a high density of epigenetic marks that influence repeat element representation. These epigenetic hotspot bins were described previously; briefly, we analyzed ENCODE Chip-Seq data from lymphoblastoid cell line GM12878(56) (ENCFF001SUG, ENCFF001SUI, ENCFF001SUJ, ENCFF001SUE, ENCFF001SUL, ENCFF001SUF, ENCFF001SUN, ENCFF001SUO, ENCFF001SUP, ENCFF001SUQ). We further used ENCODE 88 168948691.1chromatin state definitions, grouping states 1-5 (Promoter, Enhancer 1, Enhancer 2, Transcription 5` 1, Transcription 5` 2) as Activating, States 7-9 (Transcription 3` 1, Transcription 3` 2, Transcription 3` 3) as 3` Transcription, and States 10-13 (PC Repressed 1, PC Repressed 2, Heterochromatin 1, Heterochromatin 2) as Repressed(57). The 1 Mb bins selected had either >90% of their bases covered by peaks of H3K27me3, H3K36me3, H3K9me3 or H4K20me1 from one of the Histone CHIP-Seq experiments above, or >30% of their bases covered by one of the three groups of chromatin states defined above. [000290] Fragmentation Morbidity Index model [000291] We cross-validated a penalized logistic regression (PLR) model to distinguish individuals with high CCI (3+) from individuals with low CCI (0,1,2). The model used fragmentomic features comprising the principal components capturing 90% of variation among the ratios of short to long fragments in 4735Mb bins genome wide and 39 chromosomal arm-level aneuploidy z-scores (1). We used 10-fold cross-validation to generate a score from the PLR for each sample in the MICA Morbidity Cohort, which we called the Fragmentation Morbidity Index (FMI). [000292] We used the cross-validated FMI in three analyses of overall survival. First, we compared survival curves for individuals with FMI in the top and bottom deciles, or above and below the mean using a log rank test. Second, we performed a cox proportional hazards multivariate analysis using FMI, Charlson Comorbidity Index (CCI), Age, and protein biomarker levels of IL6, CRP and YKL40. Third, to assess interaction between CCI and FMI with respect to survival, we performed an Analysis of deviance between models of survival incorporating CCI as compared to those incorporating an interaction between CCI and FMI, with (survival ~ Age + YKL40 + IL6 + CRP + CCI vs. survival ~ Age + YKL40 + IL6 + CRP + CCI:FMI) and without (survival ~ YKL40 + IL6 + CRP + CCI vs. survival ~ YKL40 + IL6 + CRP + CCI:FMI) age adjustment. [000293] ARTEMIS-DELFI LCr detection model [000294] We employed our previously published modeling approaches(1–3) to train a model for detection of liver cirrhosis and advanced fibrosis (FIG.17). For the fragmentation model, we trained a penalized logistic regression (PLR) model on the principal components capturing 95% 89 168948691.1of variation among the ratios of short to long fragments in 4735Mb bins genome wide(1). We also included 39 arm-level chromosomal aneuploidy zscores, motivated by observations of early structural genomic changes in cirrhosis prior to hepatocellular carcinoma (28). This model generated a fragmentation score. The PLR architecture employs lasso regression to prioritize the selection of parsimonious feature sets. [000295] For the repeat landscape analyses, we partitioned repeat element features into 5 families (LINEs, SINEs, Satellites, LTRs, TEs) and centered and scaled the features within each family in each sample (2). We also included centered and scaled coverage in high density epigenetic hotspot bins as a 6thfeature family. We then trained a PLR model to generate a score for each family. We then ensembled scores from these 6 feature family models using another PLR to generate a repeat element score. We then ensembled this repeat element score with the fragmentation score described above using another PLR to generate an ARTEMIS-DELFI score. We cross-validated the model using ten-fold cross validation, with nested 5-fold cross validation to train each ensemble component. [000296] To evaluate these models in the Validation Cohorts, we use the cross-validated scores to select two specificity cutpoints corresponding to the 50thand 90thpercentiles of scores for screening population individuals in the Discovery Cohort. We then re-trained the model using the full Discovery Cohort and locked the model. For this re-training we first used 5-fold cross- validation (without nesting) within the full Discovery Cohort to obtain scores for each feature family. These scores were used as inputs to train and lock the LCr ensemble model on the full Discovery Set. Together the locked ensemble model, and its locked component models were used to generate scores for each sample in the Validation Cohorts. [000297] The scores obtained from the cfDNA model were evaluated for disease detection in liver cirrhosis and advanced fibrosis at the fixed cutpoint determined in the Discovery Cohort. We further applied the locked model to patients with other diseases to demonstrate the limited cross-reactivity of the model. For cross-reactivity analyses we also utilized the same fragmentomic model architecture trained to detect lung cancer or liver cancer in our previous work (1, 2) to generate scores for patients in the liver cirrhosis cohorts. 90 168948691.1[000298] To compute feature importance for the model, we selected all ensemble components with non-zero coefficients in each locked PLR ensemble and multiplied each coefficient by the mean of the ensemble component score it corresponded to. This scaling normalized for differences in magnitude between components. We then used the scaled coefficients from the top-level ensemble model to determine the relative importance of the fragmentation vs. repeat element ensemble components. We then subdivided the portion of feature importance comprising the repeat element ensemble based on the coefficients corresponding to scores for the epigenetic features and each repeat element family. [000299] Gene coverage analyses [000300] We computed GC-corrected cfDNA coverage in 100kb bins overlapping all known genes and used these coverage metrics to compute p-values (Wilcoxon’s rank-sum test) and effect sizes (Cohen’s D) for coverage differences at each gene in patients with liver cirrhosis and advanced fibrosis vs. screening population individuals. We identified genes with coverage differences significant after multiple test correction and compared these enriched genes to genes with a known connection to cirrhosis (Monarch gene set HP:0001394) in the Monarch database (44). [000301] White blood cell epigenetic analyses [000302] We downloaded Histone CHIP-Seq data for white blood cells from ENCODE (56) and identified 100kb regions with activating (H3K27ac) and repressive (H3K27me3) states in T lymphocytes, Monocytes and Neutrophils. We then computed GC corrected coverage in these bins from the plasma of each cfDNA sample. [000303] Genome-wide transcription factor binding site analyses [000304] We utilized the DECIFER methodology (52) to compare observed plasma cfDNA coverage differences at transcription factor binding sites to known transcription factor expression profiles in various tissues. Briefly, Chromatin immunoprecipitation followed by sequencing (ChIP-Seq) peaks were downloaded from the ReMap 2020 database (58), and relative coverage in each plasma sample at each peak (+ / - 100bp) vs. downstream (+ / - 2500-3000bp from the peak) was computed. RNA expression values (gene-level transcripts per million (TPM) units) for the same transcription factors were downloaded from the UCSC Toil RNAseq Recompute 91 168948691.1Compendium(59) for all normal tissue types (n=53) and whole blood in the GTEx dataset (53). TF expression matrices in TPM units were aggregated across datasets, and log-transformed by calculating log2(TPM + 0.001). [000305] We defined the effect size as the standardized difference between the average relative coverage at TFBS among individuals with liver cirrhosis or Advanced Fibrosis and the average relative coverage at TFBS among individuals in the screening population (Cohen's d). Similarly, we compared the expression of each TF in each normal organ tissue to its expression in whole blood samples from healthy individuals in the GTEx dataset and calculated the associated effect size (Cohen's d). The association between tissue expression differences and plasma cfDNA TFBS relative coverage differences was quantified by calculating Pearson's r correlation coefficient between the two effect size sets. Following the rationale that coverage metrics are more robust for TFs with large numbers of TFBS (due to overall increased coverage available for analysis) we iteratively applied increasingly stringent filters for the minimum number of TFBS (using quantiles) and recalculated this correlation coefficient at each quantile threshold. We computed the association between plasma coverage differences and expression differences for each cohort (liver cirrhosis Discovery and Validation) and each tissue studied. 92 168948691.1168948691.1168948691.1DOCKET NO.: 348358.17810 [000306] References for Example 2: 1. D. Mathios, et al. Detection and characterization of lung cancer using cell-free DNA fragmentomes. Nat. Commun.12, 5060 (2021). 2. A. V. Annapragada, et al. Genome-wide repeat landscapes in cancer and cell-free DNA. Sci. Transl. Med.16, eadj9283 (2024). 3. Z. H. Foda, et al. Detecting liver cancer using cell-free DNA fragmentomes. Cancer Discov.13, 616–631 (2022). 4. C. Martin-Alonso, et al. Priming agents transiently reduce the clearance of cell-free DNA to improve liquid biopsies. Science 383, eadf2341 (2024). 5. K. C. A. Chan, et al. Second generation noninvasive fetal genome analysis reveals de novo mutations, single-base parental inheritance, and preferred DNA ends. Proc. Natl. Acad. Sci.113, E8159–E8168 (2016). 6. D. C. Bruhm, et al. Single-molecule genome-wide mutation profiles of cell-free DNA for non- invasive detection of cancer. Nat. Genet.55, 1301–1310 (2023). 7. Q. Zhou, et al. Epigenetic analysis of cell-free DNA by fragmentomic profiling. Proc. Natl. Acad. Sci.119, e2209852119 (2022). 8. M. W. Snyder, et al. Cell-free DNA Comprises an In Vivo Nucleosome Footprint that Informs Its Tissues-Of-Origin. Cell 164, 57–68 (2016). 9. S. Cristiano, et al. Genome-wide cell-free DNA fragmentation in patients with cancer. Nature 570, 385–389 (2019). 10. S. Y. Shen, et al. Sensitive tumour detection and classification using plasma cell-free DNA methylomes. Nature 563, 579–583 (2018). 11. A. Zviran, et al. Genome-wide cell-free DNA mutational integration enables ultra-sensitive cancer monitoring. Nat. Med.26, 1114–1124 (2020). 12. J. D. Cohen, et al. Detection and localization of surgically resectable cancers with a multi- analyte blood test. Science 359, 926–930 (2018). 13. M. Noë, et al. DNA methylation and gene expression as determinants of genome-wide cell- free DNA fragmentation. Nat. Commun.15, 6690 (2024). 168834212.1 168948691.114. J. Phallen, et al. Direct detection of early-stage cancers using circulating tumor DNA. Sci. Transl. Med.9, eaan2415 (2017). 15. J. E. Medina, et al. Cell-free DNA approaches for cancer early detection and interception. J. Immunother. Cancer 11, e006013 (2023). 16. J. Moss, et al. Comprehensive human cell-type methylation atlas reveals origins of circulating cell-free DNA in health and disease. Nat. Commun.9, 5068 (2018). 17. N. Loyfer, et al. A DNA methylation atlas of normal human cell types. Nature 613, 355–364 (2023). 18. Z. M. Younossi, et al. The global epidemiology of nonalcoholic fatty liver disease (NAFLD) and nonalcoholic steatohepatitis (NASH): a systematic review. Hepatology 77, 1335–1347 (2023). 19. W. Eskridge, et al. Metabolic Dysfunction-Associated Steatotic Liver Disease and Metabolic Dysfunction-Associated Steatohepatitis: The Patient and Physician Perspective. J. Clin. Med.12, 6216 (2023). 20. M. E. Rinella, et al. A multisociety Delphi consensus statement on new fatty liver disease nomenclature. J. Hepatol.79, 1542–1556 (2023). 21. J. Araújo, et al. Prevalence of Optimal Metabolic Health in American Adults: National Health and Nutrition Examination Survey 2009–2016. Metab. Syndr. Relat. Disord.17, 46–52 (2019). 22. H. Devarbhavi, et al. Global burden of liver disease: 2023 update. J. Hepatol.79, 516–537 (2023). 23. K. Onyirioha, et al. Is hepatocellular carcinoma surveillance in high-risk populations effective? Hepatic Oncol.7, HEP25 (2020). 24. E. B. Tapper, et al. Diagnosis and Management of Cirrhosis and Its Complications. JAMA 329, 1589–1602 (2023). 25. A. G. Singal, et al. HCC surveillance improves early detection, curative treatment receipt, and survival in patients with cirrhosis: A meta-analysis. J. Hepatol.77, 128–139 (2022). 26. R. K. Sterling, et al. AASLD Practice Guideline on blood-based noninvasive liver disease assessment of hepatic fibrosis and steatosis. Hepatology 81, 321–357 (2025). 96 168948691.127. I. Graupera, et al. Low Accuracy of FIB-4 and NAFLD Fibrosis Scores for Screening for Liver Fibrosis in the Population. Clin. Gastroenterol. Hepatol.20, 2567-2576.e6 (2022). 28. S. F. Brunner, et al. Somatic mutations and clonal dynamics in healthy and cirrhotic human liver. Nature 574, 538–542 (2019). 29. A. N. Videmark, et al. Combined plasma C-reactive protein, interleukin 6 and YKL-40 for detection of cancer and prognosis in patients with serious nonspecific symptoms and signs of cancer. Cancer Med.12, 6675–6688 (2023). 30. M. E. Charlson, et al. A new method of classifying prognostic comorbidity in longitudinal studies: Development and validation. J. Chronic Dis.40, 373–383 (1987). 31. A. R. Thierry, Circulating DNA fragmentomics and cancer screening. Cell Genom.3, 100242 (2023). 32. J. E. Medina, et al. , Early detection of ovarian cancer using cell-free DNA fragmentomes and protein biomarkers. In Press. Cancer Discovery. 33. U. H. Iloeje, et al. (the R.-H. S. Group, Predicting Cirrhosis Risk Based on the Level of Circulating Hepatitis B Viral Load. Gastroenterology 130, 678–686 (2006). 34. M. Salvucci, et al. Frequent microsatellite instability in post hepatitis B viral cirrhosis. Oncogene 13, 2681–5 (1996). 35. M.-H. Lee, et al. Chronic Viral Hepatitis B and C Outweigh MASLD in the Associated Risk of Cirrhosis and HCC. Clin. Gastroenterol. Hepatol.22, 1275-1285.e2 (2024). 36. A. Annapragada, et al. Genome-wide repeat landscapes in cancer and cell-free DNA. Science Translational Medicine (2024). 37. P. Ulz, et al. Inferring expressed genes by whole-genome sequencing of plasma DNA. Nat. Genet.48, 1273–1278 (2016). 38. M. S. Esfahani, et al. Inferring gene expression from cell-free DNA fragmentation profiles. Nat. Biotechnol.40, 585–597 (2022). 39. Y. M. D. Lo, et al. Epigenetics, fragmentomics, and topology of cell-free DNA in liquid biopsies. Science 372 (2021), doi:10.1126 / science.aaw3616. 40. R. T. Starosta, et al. Liver involvement in patients with Gaucher disease types I and III. Mol. Genet. Metab. Rep.22, 100564 (2020). 97 168948691.141. M. Romero-Gómez, et al. M. A. Montes-Cano, M. A. Otero-Fernández, B. Torres, D. Sánchez-Muñoz, F. Aguilar, N. Barroso, L. Gómez-Izquierdo, V. M. Castellano-Megias, A. Núñez- Roldán, J. Aguilar-Reina, M. F. González-Escribano, SLC11A1 promoter gene polymorphisms and fibrosis progression in chronic hepatitis C. Gut 53, 446 (2004). 42. F. Geisler, et al. Emerging roles of Notch signaling in liver disease. Hepatology 61, 382–392 (2015). 43. H. Xu, et al. The Role of Notch Signaling Pathway in Non-Alcoholic Fatty Liver Disease. Front. Mol. Biosci.8, 792667 (2021). 44. T. E. Putman, et al. The Monarch Initiative in 2024: an analytic platform integrating phenotypes, genes and diseases across species. Nucleic Acids Res.52, D938–D949 (2023). 45. J. Shi, et al. Matrine promotes hepatic oval cells differentiation into hepatocytes and alleviates liver injury by suppression of Notch signalling pathway. Life Sci.261, 118354 (2020). 46. Y. M. Yang, et al. Hyaluronan synthase 2–mediated hyaluronan production mediates Notch1 activation and liver fibrosis. Sci. Transl. Med.11 (2019), doi:10.1126 / scitranslmed.aat9284. 47. V. A. Cortés, et al. Molecular Mechanisms of Hepatic Steatosis and Insulin Resistance in the AGPAT2-Deficient Mouse Model of Congenital Generalized Lipodystrophy. Cell Metab.9, 165– 176 (2009). 48. H. Y. Mak, et al. AGPAT2 interaction with CDP-diacylglycerol synthases promotes the flux of fatty acids through the CDP-diacylglycerol pathway. Nat. Commun.12, 6877 (2021). 49. L. Lombardo, et al. Peripheral blood CD3 and CD4 T-lymphocyte reduction correlates with severity of liver cirrhosis. Int. J. Clin. Lab. Res.25, 153–156 (1995). 50. X. Li, et al. Evaluation of the neutrophil-to-lymphocyte ratio, monocyte-to-lymphocyte ratio, and red cell distribution width for the prediction of prognosis of patients with hepatitis B virus-related decompensated cirrhosis. J. Clin. Lab. Anal.34, e23478 (2020). 51. D. Li, et al. Utility of neutrophil–lymphocyte ratio and platelet–lymphocyte ratio in predicting acute-on-chronic liver failure survival. Open Life Sci.18, 20220644 (2023). 52. D. Mathios, et al. Detection of brain cancer using genome-wide cell-free DNA fragmentation profiles and repeat landscapes. Submitted . 98 168948691.153. J. Lonsdale, et al. The Genotype-Tissue Expression (GTEx) project. Nat. Genet.45, 580–585 (2013). 54. C. H. Jones, et al. Healthcare on the brink: navigating the challenges of an aging society in the United States. npj Aging 10, 22 (2024). 55. B. Langmead, et al. Fast gapped-read alignment with Bowtie 2. Nat. Methods 9, 357–359 (2012). 56. I. Dunham et al. An integrated encyclopedia of DNA elements in the human genome. Nature 489, 57–74 (2012). 57. J. W. K. Ho, et al. , Comparative analysis of metazoan chromatin organization. Nature 512, 449–452 (2014). 58. J. Chèneby, et al. ReMap 2020: a database of regulatory regions from an integrative analysis of Human and Arabidopsis DNA-binding sequencing experiments. Nucleic Acids Res.48, D180–D188 (2020). 59. J. Vivian, A., et al. Toil enables reproducible, open source, big biomedical data analyses. Nat. Biotechnol.35, 314–316 (2017). [000307] Example 3: cfDNA fragmentomes as a molecular biomarker of disease morbidity [000308] We first collected blood from 570 individuals enrolled in the prospective MICA study suspected of having serious illness presenting to the Diagnostic Outpatient Clinic at Herlev and Gentofte Hospital, Copenhagen University Hospital, Denmark between 2016 and 2019 (FIG. 30)31. Individuals in the MICA Cohort had a range of diseases and non-neoplastic morbidities including cardiovascular disease (n=139), prior history of cancer (n=90), diabetes (n=74), cerebral vascular disease (n=25), connective tissue disease (n=46), heart insufficiency (n=38), dementia (n=7) and mild chronic liver disease (n=9). The combined burden of these and other morbidities can be measured using the Charlson Comorbidity Index (CCI) which provides an age-dependent measure of comorbidity burden that reflects probability of 10-year survival for an individual32. Individuals in the MICA cohort had a median CCI of 3 (range 0-10), and included 244 individuals with low CCI (<=2) and 326 with high CCI (3+). For each individual, we isolated plasma from whole blood, extracted cfDNA, constructed DNA libraries with size selection for cfDNA of nucleosomal origin, and performed whole-genome sequencing at an average 1.5x coverage. 99 168948691.1[000309] cfDNA concentrations, as measured by the amount of extracted cfDNA in a nucleosomal length range that excludes larger genomic DNA, are typically low in healthy individuals, with a mean of 7 ng / mL of plasma14. The concentration of cfDNA in the plasma of individuals with low CCI was similar to this level but was elevated in individuals with high CCI (3+) (mean low CCI = 5.63 ng / mL compared to mean high CCI=10.5 ng / mL, p=3.6e-13, Wilcoxon rank-sum test), with increased levels observed for most individuals for the diseases studied (p<0.01 in all conditions, Wilcoxon rank sum test) (FIG.10B). Consistent with previous observations of higher levels of cfDNA in individuals with increased inflammation, patients with high CCI had elevated levels of well-known inflammatory markers (CRP, IL-6, and YKL40), suggesting that altered cfDNA concentrations in the plasma may reflect systemic immune processes in addition to tissue-specific cell death33–37(FIGS.32A and 32B). [000310] The length of cfDNA fragments largely reflects their nucleosomal origins, with the length of a fragment wrapped around a nucleosome of histones and a linker being approximately 167 bp8,9,38. Underlying genomic, chromatin and epigenomic changes can alter this arrangement, influencing DNA degradation during cell death and altering cfDNA fragment length overall as well as at disease-specific loci7,9,13,39. We found that the overall mean length of cell-free DNA fragments was shorter in individuals with a high CCI (3+) as compared to those with low CCI (0- 2) and reflected an increase in fragments of sub mononucleosomal length (peak increase in fragments at 165bp with a concomitant peak decrease at 180bp, p<7.0x10-6for both lengths, Wilcoxon rank sum test) (FIG.10C, FIG.31C). These observations suggested that a nonspecific increase in cfDNA concentration evident in disease may reflect an increase in cfDNA from cells with aberrant genomic changes. [000311] Given the observed changes in cfDNA fragment lengths and concentrations in these conditions, we evaluated whether a more comprehensive examination of the cfDNA fragmentome could be used to identify individuals with high burden of comorbidities. We have previously shown that changes in cfDNA fragmentomes underlie cancer and pre-malignant conditions1,3,9,40and that features including fragment length, coverage, and other changes across the genome capture disease-specific genomic fingerprints. Here, in the context of diverse non-neoplastic morbidities, we constructed fragmentation profiles comprising the size and coverage distribution 100 168948691.1of ~45 million fragments in 473 nonoverlapping 5Mb regions spanning approximately 2.4Gb of the genome. [000312] We trained a machine learning model using penalized logistic regression with 10- fold cross-validation to distinguish between individuals with high and low CCI from each other, based on the hypothesis that cfDNA fragmentomic changes would be evident in a wide range of diseases not previously evaluated and present genome-wide information beyond aggregate measures of concentration and overall fragment length. The features used in this model included the fraction of short (100-150 bp) to long (151-220 bp) cfDNA fragments in each of 4735Mb bins and z-scores representing chromosomal arm-level changes and were used to generate a “fragmentation morbidity index”, representing a molecular measure of broad disease burden, for each individual. The fragmentation morbidity index separated individuals with high and low CCI with high-performance (median 0.64 for high CCI vs.0.5 for low CCI, p=9.2e-13, Wilcoxon rank- sum test) (FIG. 10D) and correlated with overall survival (median overall survival 2.8 years vs. not reached for individuals in the top vs. bottom decile of scores, p<0.001; p=0.036 for individuals with scores above vs. below the mean) (FIG.10E, FIG.31D). In a multivariate cox proportional hazards model, the fragmentation morbidity index was an independent predictor of survival even after adjusting for clinical characteristics, including age, inflammatory markers and CCI (FIG. 10F). The interaction between CCI and fragmentation morbidity index explained variation in survival as compared to CCI alone, regardless of adjustment for age (Analysis of deviance, X2> 2.0, degrees of freedom=0, p<2.2e-16 for tests with and without age adjustment). This suggested that genome-wide cfDNA fragmentomes may serve as a molecular marker of disease, with the potential to provide information that is complementary to clinical covariates. [000313] Development of disease-specific fragmentome biomarkers for detection of Alzheimer’s Disease and liver cirrhosis [000314] Given the pan-disease prevalence of genomic changes in cfDNA fragmentomes, we sought to evaluate their potential as noninvasive disease-specific biomarkers. We chose two diseases with known diagnostic challenges and an unmet need for non-invasive biomarkers, Alzheimer’s Disease (AD) and Liver Cirrhosis (LCr). Blood samples were evaluated in Discovery Cohorts from patients with AD (n=71), with LCr (n=49), at high-risk for LCr with viral hepatitis41101 168948691.1(n=26), or from a screening population presumed to be generally healthy (n=390) ( FIG. 32). All AD patients were diagnosed with AD using clinical criteria, which require a cognitive deficit and the exclusion of other causes of dementia using imaging42,43. All patients with LCr were diagnosed by expert clinical evaluation, and as needed, imaging including liver elastography and ultrasound. For subsets of the healthy population, neurological deficit was specifically ruled out (n=97) or metabolic risk factors for LCr (arteriosclerosis, hyperlipidemia, diabetes, hypertension, or hypercholesterolemia) were confirmed (n=71). For all participants, blood collection, cfDNA extraction and genomic library preparation were conducted under uniform protocols similar to those used in the MICA cohort, and samples were sequenced at an average coverage of 6.6x (FIG. 33). We used the genome-wide cfDNA sequencing data to separately train, cross-validate, and externally validate a noninvasive liquid biopsy classifier for each disease. [000315] We hypothesized that changes to cfDNA fragmentation characteristics would be evident in AD and LCr. There have been previous observations that patients with dementia may have overall higher levels of cfDNA44,45, and changes to the epigenetic state of specific genes have been observed in the cfDNA of individuals with AD46–48. The tissue specific origins of this increase are yet to be elucidated, but we postulated that the neurotoxicity of protein deposits leading to neuronal cell-death and blood brain barrier degradation may increase the levels of brain- derived cell-free DNA present in the circulation of AD patients, in addition to other systemic changes46,49–51. Similarly, we hypothesized that fragmentomic changes would be evident in LCr, which is characterized by systemic inflammation and high rates of cellular death52. Widespread genomic changes have also been identified in cirrhotic livers, including increased somatic mutations and structural variants as compared to normal liver53. [000316] In keeping with this hypothesis, the concentration of cfDNA in a nucleosomal length range in patients with either AD and LCr in the Discovery Cohort was elevated compared to patients without disease (median 28.2 ng / ml and 10.4 ng / mL cfDNA for AD and LCr patients, respectively, compared to 5.2 ng / ml in individuals without disease, p<0.001 for both diseases, Wilcoxon rank-sum test) (FIG. 26A, FIG. 34). Similarly, the overall fraction of short cfDNA fragments was increased in individuals with AD or LCr compared to healthy individuals (median 102 168948691.1fragment lengths 164.8 and 164.1 bp, respectively, compared to 167.6, p<9.5x10-14for both diseases, Wilcoxon rank-sum test) (FIG. 26B). Consistent with these observations, the fragmentation morbidity index was higher in individuals in these cohorts who had an elevated estimated CCI (FIG.35). [000317] We examined whether these and other underlying changes in the cfDNA fragmentome could be used to develop disease specific classifiers for noninvasive detection of these illnesses. We found that analyses of the ratio of short (100-150 bp) to long (151-220 bp) cfDNA fragments in each of 473 5Mb bins across the genome revealed greater heterogeneity among patients with LCr or AD, as compared to healthy populations with different genomic regions characterizing each disease (mean Pearson’s correlation to median healthy individual 0.93 and 0.91 respectively, as compared to 0.95) (FIGS. 26A, 26D and 26E). A sizeable fraction of cfDNA fragments originate from genomic repeat elements, and structural and epigenetic changes of these elements and other regions of the genome have been implicated in human diseases, including AD and LCr2,51,54. We evaluated repeat and epigenetic landscapes in individuals in the AD or LCr Discovery cohorts across hundreds of repeat elements as well as using coverage in 1Mb bins containing a high density of epigenetic marks throughout the genome2. For both AD and LCr, we observed widespread differences as compared to healthy controls in the genome-wide representation of repeat elements (FIGS. 26A, 26D and 26E) and increased heterogeneity in cfDNA coverage at epigenetic hotspots (FIGS.26A, 26D and 26E). These observations suggested that disease-specific changes to cfDNA fragmentomes occurred in diverse genomic elements including in repetitive regions outside of known genes for both AD and LCr. [000318] As a proof-of-concept for the use of genome-wide cfDNA fragmentomes as biomarkers for AD or LCr, we designed two classifiers in the Discovery Cohort, one for AD and one for LCr, each ensembling machine learners for fragmentation profiles, and epigenetic and repeat element features. To characterize performance and mitigate the risk of over-fitting, the cfDNA AD and cfDNA LCr scores were obtained from a nested cross-validation procedure (FIG. 36) using penalized logistic regression to favor models with lower feature dimensionality. The types of features that drove classification performance were different in number, composition and 103 168948691.1importance for each disease (FIGS.26A, 26D and 26E). Epigenetic and repeat features were more influential for AD, while fragmentation changes were enriched in LCr (FIGS.26A, 26D and 26E). [000319] As clinical characteristics may affect disease biomarkers, we investigated whether clinical or demographic parameters such as age, sex or common illnesses were associated with cfDNA scores in individuals without either AD or LCr. For both models, scores among patients from the healthy population without AD and LCr did not differ by age, sex, or the presence of other comorbidities like hypertension and diabetes (FIG. 37A), suggesting that disease-specific screening models would be robust to variation among healthy individuals. [000320] We next examined the relationship between cfDNA AD scores and AD disease status. In the Discovery Cohort, scores for the 390 individuals in the healthy population were low, with median scores of 0.05 for those in both the general and no neurological deficit groups (FIG. 27A). In contrast, the patients with AD had significantly higher cfDNA scores, with a median value of 0.77 (p<2.2x10-16, Wilcoxon rank-sum test). A receiver operator characteristic curve (ROC) of the cfDNA approach to identify patients with AD revealed an area under the curve (AUC) of 0.98 (95% CI=0.97-1.0) in the general screening population and AUC of 0.99 (95% CI=0.97-1) in the healthy population without neurological deficit (FIG.27B). Importantly, scores were similar in both men and women with AD, and were high across a range of cognitive deficits and symptoms and regardless of sample source (FIGS.37A and 38). [000321] Evaluation of the cfDNA LCr model in the Discovery Cohort similarly predicted low scores for healthy individuals, regardless of the presence of metabolic risk factors (median scores of 0.04 and 0.05 for those with and without metabolic risk factors) (FIG.27C). In contrast, patients with LCr had significantly higher scores (median scores of 0.55, p<2.2x10-16, Wilcoxon rank-sum test) and could be distinguished with high performance from individuals in the healthy population with and without metabolic risk factors (AUC=0.97, 95% CI 0.94-1.0 and AUC=0.95, 95% CI=0.92-0.98, respectively) (FIG.3D). The cirrhosis signal was significantly higher in both men and women, across sample sources, and not related to the potential presence of chromosomal changes linked to undiagnosed hepatic malignancy (FIG.37C). cfDNA LCr scores were increased in individuals with higher Child-Pugh cirrhosis severity, though scores were still significantly 104 168948691.1higher than in healthy populations in all cases (p<3.2x10-6for Child-Pugh stages A, B and C, Wilcoxon rank-sum test) (FIG.37C). [000322] External validation of disease-specific liquid biopsies for Alzheimer’s Disease and Liver Cirrhosis [000323] To evaluate whether locked AD and LCr classifiers trained on the full Discovery Cohort would generalize to external Validation Sets, we first selected high specificity cutpoints corresponding to the 90th, 95th and 99th percentiles of cross-validated scores for healthy individuals from the ensemble models previously described. We then re-trained both the AD and LCr models on their full Discovery Cohorts, used the locked models to generate scores on the External Validation Cohorts, and evaluated these scores using the fixed cutpoints determined from the Discovery Cohorts. [000324] The Validation Cohorts comprised patients with AD (n=42) or LCr (n=83), as well as with known high-risk for LCr conditions including metabolic dysfunction–associated steatotic liver disease (MASLD)25,55, liver fibrosis56, or aflatoxin exposure57(n=54), or a healthy population (n=142) (FIG. 31). All healthy individuals were without neurological deficit, and subsets had a family history of AD (n=93) or metabolic risk factors for LCr (n=55). [000325] In the external AD Validation Cohort, patients with AD exhibited the same patterns of increased cfDNA fragmentation aberration observed in the Discovery Cohort (FIG.40). When the locked cfDNA AD model was applied to this external validation cohort, cfDNA scores remained high for patients with AD (median scores of 0.83 vs.0.04 for both healthy populations without neurological deficit or with a family history of AD; FIG.28A, FIGS.41A and 41B), and detected AD patients with high performance compared to healthy populations without neurological deficit (AUC=0.93, 95% CI=0.9-1.0) or with a family history of AD but without signs of disease (AUC=0.95, 95% CI=0.9-1.0) (FIG. 28B, FIG. 41C). Encouragingly, at a threshold yielding specificity of 90% and sensitivity of 96% in the Discovery Cohort, the sensitivity of our locked approach for detecting AD patients in the Validation Cohort was 88% with 90% specificity, confirming the generalizability of this classifier for AD and suggesting that a cfDNA fragmentomic approach could aid as a sensitive rule-in for Positron Emission Tomography (PET) 105 168948691.1imaging without triggering unnecessary diagnostic odysseys (FIG. 28C, FIG. 42), alone or in combination with blood and CSF protein biomarkers58,59. [000326] Similarly, in the LCr Validation Cohort, the locked cfDNA LCr model detected individuals with LCr with high performance compared to healthy populations with and without metabolic risk factors (AUC=0.95, 95% CI=0.91-0.99, and AUC=0.93, 95% CI=0.89-0.97, respectively) (FIGS.28A, 28D and 28E, FIGS.41A, 41C and 41D). Among individuals at high- risk for LCr with aflatoxin exposure, MASLD, or already present liver fibrosis, cfDNA LCr scores were elevated as compared to healthy individuals (median scores of 0.08 compared to 0.02, p<6.0x10-16, Wilcoxon rank-sum test), but remained significantly lower than for individuals with LCr (median scores 0.08 vs.0.6, p=8.0x10-5, Wilcoxon rank-sum test) (FIG.43). [000327] The locked cfDNA LCr model achieved a sensitivity of 78% and a specificity of 92% for individuals with LCr as compared to healthy individuals in the Validation Cohort locked at a threshold of 90% specificity and 90% sensitivity in the Discovery Cohort (FIG. 28F). This model outperforms two commonly used fibrosis indices estimated from blood lab values, APRI and FIB-460(FIG.44). These observations suggest that the cfDNA LCr model has the potential to be used for screening of LCr or the sensitive triage of at-risk individuals in populations with metabolic risk factors or pre-existing liver disease (FIG.45). [000328] To evaluate the cross-reactivity of our screening approaches in other conditions that may also alter cfDNA fragmentomes, we further analyzed blood samples from patients in the Validation Cohorts with AD or LCr as well as with patients with lung cancer (n=126), liver cancer (n=78), benign lung nodules (n=67), or chronic pancreatitis (n=18)1,3. Analysis of individuals with LCr using the locked AD model revealed scores that were significantly lower than those for individuals with AD (median 0.05 vs. 0.83, p<8x10-16, Wilcoxon rank-sum test) (FIG. 4G). Similarly, individuals with AD had significantly lower scores than individuals with LCr (median 0.04 vs. 0.6, p<2.0x10-12, Wilcoxon rank-sum test) when analyzed using the locked LCr model (FIG. 28H). Moreover, these models were not confounded by the presence of other diseases including other conditions involving tissue damage, inflammation, and cellular growth that may lead to non-specific systemic changes, such as chronic pancreatitis, benign lung nodules, or malignancies of the lung, or liver (p<4.0x10-4for all conditions, Wilcoxon rank-sum test) (FIG. 106 168948691.128G and 28H). cfDNA scores demonstrated greater disease-specificity than aggregate metrics like total cfDNA concentration (FIG.46). Conversely, use of our previous cfDNA prototype classifier to detect lung cancer1,2resulted in significantly lower scores for AD or LCr patients (p<2.2x10-16for both diseases, Wilcoxon rank-sum test) (FIG. 47) that were similar to those of individuals without cancer. These observations highlight that cfDNA fragmentomes reflect disease-specific changes in the circulation and the machine learning models developed for these may be used for sensitive and specific screening approaches for detection of AD and LCr. [000329] Disease-specific mechanisms of changes in cfDNA fragmentomes [000330] To understand the biological mechanisms by which fragmentomic changes occur, we performed disease-specific analyses of cfDNA for AD and LCr patients. We analyzed genome- wide cfDNA coverage differences at transcription factor binding sites (TFBS) which we and others have previously shown to reflect epigenetic and chromatin changes that influence gene expression and genome structure1,3,8,13,39,61–64. We used our Deconvolution of CfDNA sIgnals via transcription Factor-informed fragment Representation (DECIFER) method61to compare changes in expression profiles of transcription factors in various organ tissues65(n=53) as well as additional immune66(n=4) and brain67(n=6) cell types as compared to whole blood, to cfDNA coverage at TFBS in our cohorts. cfDNA coverage changes at individual loci have been shown to correlate with gene expression, as actively transcribing regions of the genome typically have more open or active chromatin and are consequently more susceptible to degradation in the circulation3,13,39. Consequently, we would expect a negative correlation (with lower coverage due to higher expression) between plasma and tissue profiles, if a given tissue is contributing fragments to the cfDNA. [000331] When comparing plasma samples from individuals with AD to healthy individuals, we observed a significant negative correlation of plasma TFBS coverage with the relative expression changes of brain tissue as compared to whole blood, indicating an increased contribution of brain-derived cfDNA (FIG. 29A, FIG. 48A). Given the genomic distinctions known to exist between different brain regions, we verified that for individuals with AD, the trend remained consistent across specific brain cell types (FIG. 49A). In contrast, when comparing individuals with LCr to healthy individuals, we observed a negative correlation between plasma 107 168948691.1TFBS coverage changes and the relative expression changes of liver tissue as compared to whole blood, with these changes being highest for transcription factors with the largest numbers of binding sites (FIG.29B, FIG.48B). When analyzing the correlations for all 53 normal tissues in the GTEx database65), we observed an enrichment for brain-related tissues in plasma from individuals with AD (FIG.29C, p=0.032, null permutation test). In contrast, this enrichment was not seen in individuals with LCr, where the strongest anticorrelation was to liver tissue (FIG.29). DECIFER analyses of white blood cell subsets revealed changes in cfDNA composition of immune cell signatures in individuals with these conditions (FIG. 49B). Taken together, these results suggest that fragmentomes of patients with these diseases are a result of epigenetic, transcriptomic and genomic as well as immune-cell mediated changes in cfDNA. [000332] DISCUSSION [000333] This Example provides an initial demonstration of the ways in which non-cancer disease states such as those involving neurodegeneration, chronic inflammation, and dysregulation of cellular growth alter cfDNA fragmentomes. We have previously developed methods for characterizing an individual’s cfDNA fragmentome1–3,6,9,13,14, including through genome-wide analyses of size, coverage distribution, and repeat element content of cfDNA fragments, providing a summary of the chromatin structure of the cell from which they originate. Using these methods as well as novel disease-specific analyses of cfDNA fragmentation, we show that changes to the cfDNA fragmentome are present across a wide spectrum of disease and may serve as a measure of clinical morbidity, and demonstrate that these may be used to detect Alzheimer’s Disease and liver cirrhosis. We have further shown that these changes reflect both target organ contributions to the cfDNA as well as immune mediated changes to white blood cell (WBC) content in cfDNA. Our results suggest that the number and types of alterations in cfDNA fragmentation may be broader than previously considered and likely include a range of disease processes beyond those in which the fragmentome has historically been studied4,5,9. The impending demographic shift towards aging populations in the United States and other countries, and the corresponding increase in age- related morbiditiess68, underscores the clinical unmet need for accessible biomarkers to enhance early detection of a wide spectrum of human disease. 108 168948691.1[000334] The examples of liquid biopsies in AD and LCr presented here are initial prototype approaches, but their high performance provides a promising avenue alongside imaging and other biomarkers. It is estimated that by 2050 approximately 152 million people worldwide will have AD and other dementias19. However, no broadly available screening method is currently in use, and the capacity of available PET scanners for definitive detection of amyloid and tau pathology is insufficient to meet the demand for screening21,22,69,70. Blood-based amyloid beta and tau measures have been proposed as screening approaches that could be used to direct patients to PET imaging, but their use remains experimental and adoption is not yet widespread, in part due to performance limitations58,59,71,72. cfDNA fragmentation-based biomarkers, like those described in our study, have the potential to be complementary to protein biomarkers including phosphorylated tau-217 and the beta-amyloid 42 / 40 ratio and may be useful alone or in combination to identify patients in need of further imaging. [000335] Liver disease, including LCr, is the 11thleading cause of death globally, and as many as 25-30% of the global population may have metabolic dysfunction–associated steatotic liver disease (MASLD)24, previously known as nonalcoholic fatty liver disease (NAFLD), which can lead to cirrhosis if untreated25. In the United States, upwards of 80% of adults have at least one metabolic risk factor for liver cirrhosis26. While liver cirrhosis is not reversible, its timely identification improves clinical management including initiation of Hepatocellular Carcinoma surveillance, and behavioral and pharmaceutical intervention29,30. For LCr, the only available screens for liver fibrosis and cirrhosis are blood-based Fibrosis-4 index (FIB-4), AST to Platelet Ratio Index (APRI), and the NAFLD Fibrosis Score (NFS), which have limited sensitivity, or image-based liver elastography, either by ultrasound or MRI, which is not broadly accessible60. Blood-based cfDNA fragmentomic biomarkers could enhance and expand current screening paradigms for liver cirrhosis. [000336] Despite the encouraging results for noninvasive detection of these conditions, our study has several limitations. While the described classifiers provide a proof-of-concept for the ability to obtain disease-specific signals from fragmentomic data, the disease diagnosis for the individuals in the AD cohorts was based on cognitive decline and imaging rule-out of other causes of dementia. A definitive diagnosis for AD requires the presence of amyloid or tau pathology, 109 168948691.1confirmed via PET imaging or postmortem histology, although in practice this is not widely available. Additionally, while the samples in our cohorts were largely obtained via prospective collections, a subset were obtained from case-control collections that may not accurately reflect the characteristics of these diseases in the general population. A potential concern could be that these approaches detect diseases other than LCr or AD. Our cross-reactivity and tissue of origin analyses suggest that the model and underlying features are specific to each disease. Further work is needed to validate these classifiers in additional diseases and larger patient groups, and to further characterize the organ-specific origins of altered cfDNA fragmentomes. An important next step will be separate large prospective clinical trials of the respective intended use populations to validate these observations. [000337] Nevertheless, this study has shown that changes in cfDNA fragmentomes are evident in diverse disease phenotypes including AD and LCr. The genomic location, extent, and type of these alterations varies by disease state, and may be used for non-invasive screening of multiple pathologies with high performance and limited cross-reactivity. The unique specificity of the approach and the ability to extract genome-wide fragmentome information related to multiple diseases from a single blood sample could enable a facile and accessible liquid biopsy test for multi-disease screening. [000338] METHODS [000339] Study Population for MICA Morbidity Cohort [000340] Plasma samples from 570 individuals suspected of having serious illness were included in the MICA Morbidity cohort (FIGS.30 and 49). These samples were originally obtained from individuals in the prospective biomarker study MICA “New biomarkers in patients referred because of suspected serious illness - are they giving new diagnostic information?” performed at the Diagnostic Outpatient Clinic at Herlev and Gentofte Hospital, Copenhagen University Hospital, Denmark31. The MICA study enrolled all individuals over the age of 18 presenting to the Diagnostic Outpatient Clinic via the Cancer Patient Pathway (n=757) from July 2016 to July 2019, though in this study we analyzed samples only from individuals without cancer (n=570). This pathway is for patients presenting with non-specific signs and symptoms of cancer and suspected 110 168948691.1of having cancer or other serious illness. The most common symptoms at time of referral were weight loss, fatigue, pain or nausea. All patients were examined by medical doctors from the Diagnostic Outpatient Clinic at Herlev and Gentofte Hospital, and cancer diagnosis was based on radiological findings and tissue biopsies. Comorbidities were assessed by Charlson comorbidity index (CCI). The MICA study is approved by the Regional Ethics Committee (H-7-2014-011) and the Danish Data Protection Agency (HEH-2014-105). All participants were informed and signed an informed consent declaration before participation. [000341] Study Populations for AD and LCr Cohorts [000342] We analyzed plasma samples from 857 individuals in separate Discovery and Validation Cohorts for AD and LCr detection (FIG.27 and FIG.50). [000343] Plasma samples from 293 healthy individuals used as a general, presumed healthy population in both the AD and LCr Discovery Cohorts were described in our previous studies3,9. These samples were originally obtained from two screening clinical trial cohorts for colorectal cancer in Denmark (Endoscopy III) and the Netherlands (COCOS, Netherlands Trial Register ID NTR182946). The protocol for the Endoscopy III Project was approved by the Regional Ethics Committee and the Danish Data Protection Agency, and for the COCOS trial, ethical approval was obtained from the Dutch Health Council. The inclusion criteria for both the Dutch and the Danish cohorts were any individuals of age 50–75 eligible for colorectal cancer screening. All patients used had either a FIT negative test or a negative colonoscopy result. All patients provided written informed consent and the studies were performed according to the Declaration of Helsinki. [000344] Plasma samples for patients with Alzheimer’s Disease (n=71) and healthy controls without neurological deficit or malignancy (n=97) were obtained from ProteoGenex (Inglewood, CA, USA) for use in the Discovery Cohort. We further obtained plasma from 42 patients with AD and 42 patients without neurological deficit from iSpecimen (Woburn, MA, USA) for use in the Validation Cohort. For the Validation Cohort, we obtained 100 patient samples without neurological deficit, including 93 with a family history of AD from Precision For Medicine (Frederick, MD, USA). For the diagnosis of Alzheimer’s Disease, we required that patients meet clinical criteria for a probable Alzheimer’s Diagnosis, with cognitive deficit and brain imaging 111 168948691.1(ideally MRI or CT) to rule out other causes of dementia. We included only individuals with Moderate or Severe dementia, as assessed by Mini Mental State Examination (scores of 0-10 for Severe dementia, scores of 10-20 for Moderate Dementia). To avoid confounding factors in these commercially obtained cohorts, we specifically selected samples without known cancers, and without obvious inflammatory conditions that carry potential pre-cancer risks (ie, pancreatitis, cholecystitis, gastritis). [000345] Plasma samples from patients in the Discovery Cohort with liver cirrhosis (n=49) or high risk for LCr with viral hepatitis (n=26) were described in our previous study3. For the Discovery Cohort, we utilized sequencing data from all individuals without cancer in the training set of that previous study, including samples collected prospectively as part of the HCC biomarker registry at the Johns Hopkins University School of Medicine (Baltimore, MD, USA), under protocols approved by the Johns Hopkins Institutional Review Board or retrospectively collected by BioIVT (Westbury, NY, USA). We excluded the subset of patients with AIDS or using injection drugs (n=58) from the AIDS Linked to the IntraVenous Experience (ALIVE) study as individuals with AIDS or who use injections drugs are known to have altered immune profiles which impact white blood cell lineages, a large contributor to cfDNA fragmentomes73–75. For the Validation Cohort, samples were obtained from individuals without cancer enrolled in a case-control study with longitudinal follow-up in Guatemala (n=42 patients with LCr, n=54 patients with high-risk for LCr conditions including MASLD, Fibrosis or Aflatoxin Exposure), locally led by the Institute of Nutrition of Central America and Panama (INCAP) in Guatemala City, a WHO-research center. INCAP collaborates extensively with physicians at Hospital San Juan de Dios and Hospital Roosevelt, the two large public health referral hospitals in Guatemala City. Also in the Validation Cohort, samples were prospectively obtained at the Fundeni Clinical Institute or associated hospitals in Bucharest, Romania under protocols approved by local review board (n=41 patients with LCr). [000346] Patients with AD were excluded from the LCr Discovery Cohort and patients with LCr or high-risk conditions were excluded from the AD Discovery Cohort for model training and cross-validation. Where available, clinical metadata was used to identify individuals in the 112 168948691.1presumed healthy populations with confirmed metabolic risk factors for LCr (arteriosclerosis, hyperlipidemia, diabetes, hypertension, or hypercholesterolemia). [000347] Study Populations for Cross-Reactivity Analyses [000348] For analyses of cross-reactivity using the trained AD and LCr detection models, we analyzed previously published samples from individuals with hepatocellular carcinoma (HCC, n= 78) prospectively collected through the HCC biomarker registry at the Johns Hopkins University School of Medicine, lung cancer (n=126) or benign lung nodules (n=67) from the prospectively collected LUCAS Cohort1at Bispebjerg Hospital in Copenhagen, Denmark, or chronic pancreatitis (n=18) in the BIOPAC cohort (Clinical trials.gov ID: NCT03311776). The BIOPAC study has been approved by the Danish Regional Ethics Committee (BIOPAC: KA-20060113) and the Danish Data Protection Agency (BIOPAC: P-2020-834). [000349] cfDNA extraction and sequencing [000350] The protocols for blood collection, cfDNA extraction and sequencing have been previously described1,3. In brief, whole blood was collected during outpatient venous blood draws in EDTA tubes. All collection sites utilized a consistent protocol: within a maximum of two to four hours from collection, tubes were centrifuged at low speed (1000-3000g) for atleast 10 minutes. The plasma portion was transferred to a new tube and spun at 18,000g for 10 minutes. EDTA plasma was aliquoted and stored between -70 and -80 C for cfDNA analysis. All samples were transported on dry ice with included temperature monitor to ensure that temperature was maintained prior to transfer to -80 C freezers in our laboratory. For all samples, cfDNA was isolated from 1-4mL of plasma using the Qiagen QIAamp Circulating Nucleic Acids Kit. Extracted cfDNA was quantified on the Agilent 4150 TapeStation system using the Agilent Cell-free DNA ScreenTape assay. We estimated concentration of extracted cfDNA fragments between 50 and 800bp to quantify only DNA in a nucleosomal length range chosen to exclude potential genomic DNA contamination. [000351] Then, next-generation sequencing libraries were prepared with the NEBNext DNA Library Prep Kit for Illumina using a target of 15ng of cfDNA when available. For samples in the Endoscopy III and COCOS trials, all available cfDNA was used to prepare the library (median 113 168948691.1library input 23.13 ng). We modified manufacturer’s guidelines to utilize an on-bead AMPure XP approach to minimize sample loss during elution and tube transfer steps, and to size select for cfDNA within a nucleosomal length range1. This ensures that long fragments of genomic DNA are excluded from the library preparation. Sequencing libraries were quantified on the Agilent 4150 TapeStation system using the D1000 DNA ScreenTape assay. [000352] All libraries underwent four cycles of PCR amplification and were sequenced with estimated 1-10x coverage on the Illumina NovaSeq6000 platform using 100bp paired-end runs. After sequencing, we excluded samples with Q20<80%, Q30<80% or <3e9 bases sequenced. [000353] The MICA Morbidity Cohort was prospectively collected at a single site and processed and sequenced in library batches entirely distinct from the AD and LCr cohorts (FIG. 49). To reduce preanalytical and technical variation in the AD and LCr cohorts, samples from patients with and without disease were processed together in next-generation sequencing (NGS) library batches, the processing and sequencing of the Discovery and Validation cohorts was temporally separated, and all sources of AD cases also provided samples from individuals without neurological deficit (FIGS. 50 and 51). Additionally, no samples collected from the same site were used in both the Discovery and Validation Cohorts, to allow for an external validation of the locked models. [000354] Characterization of the cfDNA fragmentome [000355] We computed genome-wide fragmentation features as previously described1. Briefly, we filtered adapter sequences from 100 bp paired-end reads using fastp, aligned to the hg19 reference genome using Bowtie276, and removed duplicate reads with Sambamba. After alignment, we converted each read pair to a genomic interval using bedtools, and excluded reads with low alignment quality (MAPQ < 30) or that overlapped the Duke Excluded Regions blacklist (https: / / genome.ucsc.edu / cgi-bin / hgTrackUi?db=hg19&g=wgEncodeMapability). We performed a fragment level GC correction against a reference panel of 20 healthy individuals. We then analyzed fragment length distributions in all samples, and summarized fragmentomic features by calculating the ratio of short (100-150 bp) to long (151-220 bp) fragments in 473 non-overlapping 114 168948691.15 MB bins across the genome. We also computed z-scores representing arm gains and losses for the 39 autosomal chromosome arms. [000356] We also computed genome-wide repeat landscapes using an alignment-free approach to localizing repeat elements we previously described2. Briefly, we counted 1.2 billion kmers that map to 1280 unique repeat element types from the LINE (long interspersed nuclear element), SINE (short interspersed nuclear element), DNA / RNA TE (Transposable Element), LTR (Long terminal repeat) and Satellite families in all sequencing reads for each sample. We then aggregated these kmer counts for each repeat type and normalized to aligned coverage. This generated a kmer repeat landscape for each sample. We utilized a subset of the 1280 elements (n=599) that were observed to be highly reproducible in low coverage sequencing, as estimated by within sample coefficient of variation of the feature among sequencing replicates. [000357] We further calculated aligned coverage in 551 1Mb bins2genome-wide corresponding to regions with a high density of epigenetic marks that influence repeat element representation. These epigenetic hotspot bins were described previously; briefly, we analyzed ENCODE Chip-Seq data from lymphoblastoid cell line GM1287877(ENCFF001SUG, ENCFF001SUI, ENCFF001SUJ, ENCFF001SUE, ENCFF001SUL, ENCFF001SUF, ENCFF001SUN, ENCFF001SUO, ENCFF001SUP, ENCFF001SUQ). We further used ENCODE chromatin state definitions, grouping states 1-5 (Promoter, Enhancer 1, Enhancer 2, Transcription 5` 1, Transcription 5` 2) as Activating, States 7-9 (Transcription 3` 1, Transcription 3` 2, Transcription 3` 3) as 3` Transcription, and States 10-13 (PC Repressed 1, PC Repressed 2, Heterochromatin 1, Heterochromatin 2) as Repressed78. The 1 Mb bins selected had either >90% of their bases covered by peaks of H3K27me3, H3K36me3, H3K9me3 or H4K20me1 from one of the Histone CHIP-Seq experiments above, or >30% of their bases covered by one of the three groups of chromatin states defined above. [000358] Fragmentation Morbidity Index Model [000359] We trained and cross-validated a penalized logistic regression (PLR) model to distinguish individuals with high CCI (3+) from individuals with low CCI (0,1,2). The model used fragmentomic features comprising the principal components capturing 90% of variation among 115 168948691.1the ratios of short to long fragments in 4735Mb bins genome wide and 39 chromosomal arm-level aneuploidy z-scores1. We used 10-fold cross-validation to generate a score from the PLR for each sample in the MICA Morbidity Cohort, which we called the Fragmentation Morbidity Index (FMI). [000360] We used the cross-validated FMI in three analyses of overall survival. First, we compared survival curves for individuals with FMI in the top and bottom deciles, or above and below the mean using a log rank test. Second, we performed a cox proportional hazards multivariate analysis using FMI, Charlson Comorbidity Index (CCI), Age, and protein biomarker levels of IL6, CRP and YKL40. Third, to assess interaction between CCI and FMI with respect to survival, we performed an Analysis of deviance between models of survival incorporating CCI as compared to those incorporating an interaction between CCI and FMI, with (survival ~ Age + YKL40 + IL6 + CRP + CCI vs. survival ~ Age + YKL40 + IL6 + CRP + CCI:FMI) and without (survival ~ YKL40 + IL6 + CRP + CCI vs. survival ~ YKL40 + IL6 + CRP + CCI:FMI) age adjustment. [000361] We retrained the FMI model using the full MICA Morbidity Cohort and validated the locked model by using it to generate scores for the AD and LCr Discovery Cohorts. We compared FMI scores between individuals estimated to have high CCI (3+) or low CCI (0,1,2). In these cohorts, we estimated a minimum CCI for each individual based on the available incomplete clinical metadata32. In these cohorts, the CCI estimate was made based on Age (+1 for each decade of age over 50), and the presence of Dementia (+1 for individuals with Alzheimer’s Disease) or Liver Disease (+1 for Hepatitis, +2 for Liver Cirrhosis). Of note, we assigned +2 for all individuals with Liver Cirrhosis, since a score of +1 corresponds to cirrhosis without portal hypertension and a score of +3 corresponds to cirrhosis with portal hypertension and we did not have information on portal hypertension. We knew all individuals in our cohorts did not have cancer (+0 for leukemia, lymphoma and solid tumors) or AIDS (+0). We incorporated knowledge of metabolic comorbidities where available (+1 for Diabetes Mellitus). We were unable to provide scores for the other components of CCI, however, based on our efforts to include only healthy individuals or individuals with AD or LCr who did not have cancer or other pre-cancer inflammatory conditions, 116 168948691.1we believe that the available metadata provides an approximate estimate of CCI, for the purpose of estimating the assignment of individuals to the low (0,1,2) or high (3+) CCI groups. [000362] Disease detection models [000363] We employed our previously published modeling approaches1–3to train models for detection of Alzheimer’s disease and liver cirrhosis (FIG. 51). For the fragmentation model, we trained a penalized logistic regression (PLR) model on the principal components capturing 95% of variation among the ratios of short to long fragments in 4735Mb bins genome wide1. In the case of the LCr model, we also included 39 arm-level chromosomal aneuploidy zscores, motivated by observations of early structural genomic changes in cirrhosis prior to hepatocellular carcinoma53. This model generated a fragmentation score. The PLR architecture employs lasso regression to prioritize the selection of parsimonious feature sets. [000364] For the repeat landscape analyses, we partitioned repeat element features into 5 families (LINEs, SINEs, Satellites, LTRs, TEs) and centered and scaled the features within each family in each sample2. We also included centered and scaled coverage in high density epigenetic hotspot bins as a 6thfeature family. We then trained a PLR model to generate a score for each family. We then ensembled scores from these 6 feature family models using another PLR to generate a repeat element score. We then ensembled this repeat element score with the fragmentation score described above using another PLR to generate a cfDNA score. We cross- validated two models (one for AD detection and one for LCr detection) using ten-fold cross validation, with nested 5-fold cross validation to train each ensemble component. [000365] To evaluate these models in the Validation Cohorts, we use the cross-validated scores to select high specificity cutpoints corresponding to the 90th, 95thand 99thpercentiles of scores for healthy individuals in the Discovery Cohorts for each model. We then re-trained each model using the full Discovery Cohort for each disease and locked the models. For this re-training we first used 5-fold cross-validation (without nesting) within the full Discovery Cohort to obtain scores for each feature family. These scores were used as inputs to train and lock the AD and LCr ensemble models on the full Discovery Set. Additionally we trained and locked models for each feature family on the full Discovery Set. Together the two locked ensemble models, and their 117 168948691.1locked feature family component models were used to generate scores for each sample in the Validation Cohorts. [000366] The scores obtained from the cfDNA models were evaluated for disease detection in the example diseases studied (Alzheimer’s disease and liver cirrhosis) at the fixed cutpoints determined in the Discovery Cohorts. We further applied each locked model to patients with other diseases to demonstrate the limited cross-reactivity of each model. For cross-reactivity analyses we also utilized the same fragmentomic model architecture trained to detect lung cancer in the LUCAS cohort1to generate scores for AD and LCr patients. [000367] To compute feature importance for each model, we selected all ensemble components with non-zero coefficients in each locked PLR ensemble and multiplied each coefficient by the mean of the ensemble component score it corresponded to. This scaling normalized for differences in magnitude between components. We then used the scaled coefficients from the top-level ensemble model to determine the relative importance of the fragmentation vs. repeat element ensemble components. We then subdivided the portion of feature importance comprising the repeat element ensemble based on the coefficients corresponding to scores for the epigenetic features and each repeat element family. [000368] Genome-wide transcription factor binding site analyses [000369] We utilized the DECIFER methodology61to compare observed plasma cfDNA coverage differences at transcription factor binding sites to known transcription factor expression profiles in various tissues. Briefly, Chromatin immunoprecipitation followed by sequencing (ChIP-Seq) peaks were downloaded from the ReMap 2020 database79, and relative coverage in each plasma sample at each peak (+ / - 100bp) vs. downstream (+ / - 2500-3000bp from the peak) was computed. RNA expression values (gene-level transcripts per million (TPM) units) for the same transcription factors were downloaded from the UCSC Toil RNAseq Recompute Compendium80for all normal tissue types (n=53) and whole blood in the GTEx dataset65. We also obtained cell-sorted datasets of brain67(SRP064454) and immune66(SRP045500) cell transcriptomes and processed them using the same pipeline as used in the GTEx toil dataset outlined at as outlined in https: / / toil.xenahubs.net:443. Samples from the immune cell 118 168948691.1subpopulation dataset were narrowed down to the subset from healthy controls, resulting in n=4 samples from each of CD4+ T-cell, CD8+ T-cell, monocytes, and neutrophil subpopulations. In the brain subpopulation dataset, we analyzed samples from adult brain tissue (astrocytes n=11, endothelial n=2, myeloid n=3, oligodendrocytes n=5, and whole cortex n=3). TF expression matrices in TPM units were aggregated across datasets, and log-transformed by calculating log2(TPM + 0.001). [000370] We defined the effect size as the standardized difference between the average relative coverage at TFBS among individuals with AD (or LCr) and the average relative coverage at TFBS among individuals in the general screening population (Cohen's d). Similarly, we compared the expression of each TF in each normal organ tissue, immune or brain cell subpopulation to its expression in whole blood samples from healthy individuals in the GTEx dataset and calculated the associated effect size (Cohen's d). The association between tissue expression differences and plasma cfDNA TFBS relative coverage differences was quantified by calculating Pearson's r correlation coefficient between the two effect size sets. Following the rationale that coverage metrics are more robust for TFs with large numbers of TFBS (due to overall increased coverage available for analysis) we iteratively applied increasingly stringent filters for the minimum number of TFBS (using quantiles) and recalculated this correlation coefficient at each quantile threshold. We computed the association between plasma coverage differences and expression differences for each cohort (AD and LCr, Discovery and Validation) and each tissue or cell subpopulation studied. As a negative control, we performed the same procedure for 1000 random divisions of the Discovery Cohort General Screening Population, and computed the IQR for correlations at each quantile. [000371] References for Example 3: 1. Mathios, D. et al. Detection and characterization of lung cancer using cell-free DNA fragmentomes. Nat. Commun.12, 5060 (2021). 2. Annapragada, A. V. et al. Genome-wide repeat landscapes in cancer and cell-free DNA. Sci. Transl. Med.16, eadj9283 (2024). 119 168948691.13. Foda, Z. H. et al. Detecting liver cancer using cell-free DNA fragmentomes. Cancer Discov.13, 616–631 (2022). 4. Martin-Alonso, C. et al. Priming agents transiently reduce the clearance of cell-free DNA to improve liquid biopsies. Science 383, eadf2341 (2024). 5. Chan, K. C. A. et al. Second generation noninvasive fetal genome analysis reveals de novo mutations, single-base parental inheritance, and preferred DNA ends. Proc. Natl. Acad. Sci.113, E8159–E8168 (2016). 6. Bruhm, D. C. et al. Single-molecule genome-wide mutation profiles of cell-free DNA for non- invasive detection of cancer. Nat. Genet.55, 1301–1310 (2023). 7. Zhou, Q. et al. Epigenetic analysis of cell-free DNA by fragmentomic profiling. Proc. Natl. Acad. Sci.119, e2209852119 (2022). 8. Snyder, M. W., Kircher, M., Hill, A. J., Daza, R. M. & Shendure, J. Cell-free DNA Comprises an In Vivo Nucleosome Footprint that Informs Its Tissues-Of-Origin. Cell 164, 57–68 (2016). 9. Cristiano, S. et al. Genome-wide cell-free DNA fragmentation in patients with cancer. Nature 570, 385–389 (2019). 10. Shen, S. Y. et al. Sensitive tumour detection and classification using plasma cell-free DNA methylomes. Nature 563, 579–583 (2018). 11. Zviran, A. et al. Genome-wide cell-free DNA mutational integration enables ultra-sensitive cancer monitoring. Nat. Med.26, 1114–1124 (2020). 12. Cohen, J. D. et al. Detection and localization of surgically resectable cancers with a multi- analyte blood test. Science 359, 926–930 (2018). 13. Noë, M. et al. DNA methylation and gene expression as determinants of genome-wide cell- free DNA fragmentation. Nat. Commun.15, 6690 (2024). 14. Phallen, J. et al. Direct detection of early-stage cancers using circulating tumor DNA. Sci. Transl. Med.9, eaan2415 (2017). 120 168948691.115. Medina, J. E. et al. Cell-free DNA approaches for cancer early detection and interception. J. Immunother. Cancer 11, e006013 (2023). 16. Moss, J. et al. Comprehensive human cell-type methylation atlas reveals origins of circulating cell-free DNA in health and disease. Nat. Commun.9, 5068 (2018). 17. Loyfer, N. et al. A DNA methylation atlas of normal human cell types. Nature 613, 355–364 (2023). 18. Leary, R. J. et al. Detection of Chromosomal Alterations in the Circulation of Cancer Patients with Whole-Genome Sequencing. Sci. Transl. Med.4, 162ra154 (2012). 19. Li, X. et al. Global, regional, and national burden of Alzheimer’s disease and other dementias, 1990–2019. Front. Aging Neurosci.14, 937486 (2022). 20. Knopman, D. S. et al. Alzheimer disease. Nat. Rev. Dis. Prim.7, 33 (2021). 21. Leuzy, A. et al. Blood‐based biomarkers for Alzheimer’s disease. EMBO Mol. Med. 14, e14408 (2022). 22. Henriksen, K. et al. The future of blood‐based biomarkers for Alzheimer’s disease. Alzheimer’s Dement.10, 115–131 (2014). 23. Gunes, S., Aizawa, Y., Sugashi, T., Sugimoto, M. & Rodrigues, P. P. Biomarkers for Alzheimer’s Disease in the Current State: A Narrative Review. Int. J. Mol. Sci.23, 4962 (2022). 24. Eskridge, W. et al. Metabolic Dysfunction-Associated Steatotic Liver Disease and Metabolic Dysfunction-Associated Steatohepatitis: The Patient and Physician Perspective. J. Clin. Med.12, 6216 (2023). 25. Rinella, M. E. et al. A multisociety Delphi consensus statement on new fatty liver disease nomenclature. J. Hepatol.79, 1542–1556 (2023). 26. Araújo, J., Cai, J. & Stevens, J. Prevalence of Optimal Metabolic Health in American Adults: National Health and Nutrition Examination Survey 2009–2016. Metab. Syndr. Relat. Disord.17, 46–52 (2019). 121 168948691.127. Devarbhavi, H. et al. Global burden of liver disease: 2023 update. J. Hepatol. 79, 516–537 (2023). 28. Onyirioha, K., Mittal, S. & Singal, A. G. Is hepatocellular carcinoma surveillance in high-risk populations effective? Hepatic Oncol.7, HEP25 (2020). 29. Singal, A. G. et al. HCC surveillance improves early detection, curative treatment receipt, and survival in patients with cirrhosis: A meta-analysis. J. Hepatol.77, 128–139 (2022). 30. Tapper, E. B. & Parikh, N. D. Diagnosis and Management of Cirrhosis and Its Complications. JAMA 329, 1589–1602 (2023). 31. Videmark, A. N. et al. Combined plasma C‐reactive protein, interleukin 6 and YKL‐40 for detection of cancer and prognosis in patients with serious nonspecific symptoms and signs of cancer. Cancer Med.12, 6675–6688 (2023). 32. Charlson, M. E., Pompei, P., Ales, K. L. & MacKenzie, C. R. A new method of classifying prognostic comorbidity in longitudinal studies: Development and validation. J. Chronic Dis. 40, 373–383 (1987). 33. Andargie, T. E. et al. Cell-free DNA maps COVID-19 tissue injury and risk of death, and can cause tissue injury. JCI Insight 6, e147610 (2021). 34. Mattox, A. K. et al. The Origin of Highly Elevated Cell-Free DNA in Healthy Individuals and Patients with Pancreatic, Colorectal, Lung, or Ovarian Cancer. Cancer Discov. 13, 2166–2179 (2023). 35. Humińska-Lisowska, K. et al. cfDNA Changes in Maximal Exercises as a Sport Adaptation Predictor. Genes 12, 1238 (2021). 36. Tanaka, A. et al. Increased levels of circulating cell-free DNA in COVID-19 patients with respiratory failure. Sci. Rep.14, 17399 (2024). 37. Jylhävä, J. et al. Aging is associated with quantitative and qualitative changes in circulating cell-free DNA: The Vitality 90+ study. Mech. Ageing Dev.132, 20–26 (2011). 122 168948691.138. Thierry, A. R. Circulating DNA fragmentomics and cancer screening. Cell Genom.3, 100242 (2023). 39. Esfahani, M. S. et al. Inferring gene expression from cell-free DNA fragmentation profiles. Nat. Biotechnol.40, 585–597 (2022). 40. Medina, J. E. et al. Early detection of ovarian cancer using cell-free DNA fragmentomes and protein biomarkers. In Press. Cancer Discovery. 41. Iloeje, U. H. et al. Predicting Cirrhosis Risk Based on the Level of Circulating Hepatitis B Viral Load. Gastroenterology 130, 678–686 (2006). 42. Jack, C. R. et al. Introduction to the recommendations from the National Institute on Aging- Alzheimer’s Association workgroups on diagnostic guidelines for Alzheimer’s disease. Alzheimer’s Dement.7, 257–262 (2011). 43. McKhann, G. M. et al. The diagnosis of dementia due to Alzheimer’s disease: Recommendations from the National Institute on Aging-Alzheimer’s Association workgroups on diagnostic guidelines for Alzheimer’s disease. Alzheimer’s Dement.7, 263–269 (2011). 44. Feger, D., Nidadavolu, L., Oh, E., Abadir, P. & Gross, A. Circulating Cell-Free DNA Is Associated With Cognitive Outcomes. Innov. Aging 4, 518–518 (2020). 45. Nidadavolu, L. S. et al. Circulating Cell-Free Genomic DNA Is Associated with an Increased Risk of Dementia and with Change in Cognitive and Physical Function. J. Alzheimer’s Dis. 89, 1233–1240 (2022). 46. Pollard, C., et al. Detection of neuron-derived cfDNA in blood plasma: a new diagnostic approach for neurodegenerative conditions. Front. Neurol.14, 1272960 (2023). 47. Macías, M. et al. Liquid Biopsy in Alzheimer’s Disease Patients Reveals Epigenetic Changes in the PRLHR Gene. Cells 12, 2679 (2023). 48. Giannini, L. A. A. et al. Distinctive cell‐free DNA methylation characterizes presymptomatic genetic frontotemporal dementia. Ann. Clin. Transl. Neurol.11, 744–756 (2024). 123 168948691.149. Kinney, J. W. et al. Inflammation as a central mechanism in Alzheimer’s disease. Alzheimer’s Dement.: Transl. Res. Clin. Interv.4, 575–590 (2018). 50. Gorham, I. K., et al. Mitochondrial SOS: how mtDNA may act as a stress signal in Alzheimer’s disease. Alzheimer’s Res. Ther.15, 171 (2023). 51. Guo, C. et al. Tau Activates Transposable Elements in Alzheimer’s Disease. Cell Rep. 23, 2874–2880 (2018). 52. Aizawa, S., Brar, G. & Tsukamoto, H. Cell Death and Liver Disease. Gut Liver 14, 20–29 (2020). 53. Brunner, S. F. et al. Somatic mutations and clonal dynamics in healthy and cirrhotic human liver. Nature 574, 538–542 (2019). 54. Salvucci, M. et al. Frequent microsatellite instability in post hepatitis B viral cirrhosis. Oncogene 13, 2681–5 (1996). 55. Lee, M.-H. et al. Chronic Viral Hepatitis B and C Outweigh MASLD in the Associated Risk of Cirrhosis and HCC. Clin. Gastroenterol. Hepatol.22, 1275-1285.e2 (2024). 56. Bataller, R. & Brenner, D. A. Liver fibrosis. J. Clin. Investig.115, 209–218 (2005). 57. Alvarez, C. S. et al. Aflatoxin B1 exposure and liver cirrhosis in Guatemala: a case–control study. BMJ Open Gastroenterol.7, e000380 (2020). 58. Janelidze, S. et al. Head-to-Head Comparison of 8 Plasma Amyloid-β 42 / 40 Assays in Alzheimer Disease. JAMA Neurol.78, 1375–1382 (2021). 59. Ashton, N. J. et al. Diagnostic Accuracy of a Plasma Phosphorylated Tau 217 Immunoassay for Alzheimer Disease Pathology. JAMA Neurol.81, 255–263 (2024). 60. Graupera, I. et al. Low Accuracy of FIB-4 and NAFLD Fibrosis Scores for Screening for Liver Fibrosis in the Population. Clin. Gastroenterol. Hepatol.20, 2567-2576.e6 (2022). 61. Mathios, D. et al. Detection of brain cancer using genome-wide cell-free DNA fragmentation profiles and repeat landscapes. Submitted. 124 168948691.162. Ulz, P. et al. Inferring expressed genes by whole-genome sequencing of plasma DNA. Nat. Genet.48, 1273–1278 (2016). 63. Ulz, P. et al. Inference of transcription factor binding from cell-free DNA enables tumor subtype prediction and early detection. Nat. Commun.10, 4666 (2019). 64. Lo, Y. M. D., et al. Epigenetics, fragmentomics, and topology of cell-free DNA in liquid biopsies. Science 372, (2021). 65. Lonsdale, J. et al. The Genotype-Tissue Expression (GTEx) project. Nat. Genet.45, 580–585 (2013). 66. Linsley, P. S., et al. Copy Number Loss of the Interferon Gene Cluster in Melanomas Is Linked to Reduced T Cell Infiltrate and Poor Patient Prognosis. PLoS ONE 9, e109760 (2014). 67. Zhang, Y. et al. Purification and Characterization of Progenitor and Mature Human Astrocytes Reveals Transcriptional and Functional Differences with Mouse. Neuron 89, 37–53 (2016). 68. Jones, C. H. & Dolsten, M. Healthcare on the brink: navigating the challenges of an aging society in the United States. npj Aging 10, 22 (2024). 69. Prato, F. S., et al. Screening for Dementia Caused by Modifiable Lifestyle Choices Using Hybrid PET / MRI. J. Alzheimer’s Dis. Rep.3, 31–45 (2019). 70. Gallach, M. et al. Addressing Global Inequities in Positron Emission Tomography-Computed Tomography (PET-CT) for Cancer Management: A Statistical Model to Guide Strategic Planning. Méd. Sci. Monit. : Int. Méd. J. Exp. Clin. Res.26, e926544-1-e926544-8 (2020). 71. Giudici, K. V. et al. Assessment of Plasma Amyloid-β42 / 40 and Cognitive Decline Among Community-Dwelling Older Adults. JAMA Netw. Open 3, e2028634 (2020). 72. Sterling, R. K. et al. Development of a simple noninvasive index to predict significant fibrosis in patients with HIV / HCV coinfection. Hepatology 43, 1317–1325 (2006). 73. Soder, H. E. et al. Elevated Neutrophil to Lymphocyte Ratio in Older Adults with Cocaine Use Disorder as a Marker of Chronic Inflammation. Clin. Psychopharmacol. Neurosci. 18, 32–40 (2020). 125 168948691.174. Quraishi, R., et al. R. Effect of chronic opioid use on the hematological and inflammatory markers: A retrospective study from North India. Indian J. Psychiatry 64, 252–256 (2022). 75. Tomescu, C. et al. Persons who inject drugs (PWID) retain functional NK cells, dendritic cell stimulation, and adaptive immune recall responses despite prolonged opioid use. J. Leukoc. Biol. 110, 385–396 (2021). 76. Langmead, B. & Salzberg, S. L. Fast gapped-read alignment with Bowtie 2. Nat. Methods 9, 357–359 (2012). 77. Dunham, I. et al. An integrated encyclopedia of DNA elements in the human genome. Nature 489, 57–74 (2012). 78. Ho, J. W. K. et al. Comparative analysis of metazoan chromatin organization. Nature 512, 449–452 (2014). 79. Chèneby, J. et al. ReMap 2020: a database of regulatory regions from an integrative analysis of Human and Arabidopsis DNA-binding sequencing experiments. Nucleic Acids Res.48, D180– D188 (2020). 80. Vivian, J. et al. Toil enables reproducible, open source, big biomedical data analyses. Nat. Biotechnol.35, 314–316 (2017). [000372] From the foregoing description, it will be apparent that variations and modifications may be made to the disclosure described herein to adopt it to various...
Claims
What is claimed:
1. A method of early detection and treatment of diseases, comprising: extracting cfDNA from a subject’s sample; isolating cell free DNA (cfDNA) from the subject’s biological samples; conducting low coverage whole genome sequencing of the subject’s isolated cfDNA to generate genomic libraries of sequenced cfDNA fragments; determining a cfDNA fragmentome and repeat landscape profiles of the subject’s sequenced DNA fragments comprising evaluating one or more cfDNA fragment characteristics; comparing the cfDNA fragmentome profiles of the subject to reference cfDNA fragmentome profiles from healthy subjects to detect and diagnose the disease; and, treating the subject with a disease specific therapy.
2. The method of claim 1, wherein the one or more cfDNA fragment characteristics comprises calculating ratios of small to large cfDNA fragments, sequences of one or more cfDNAs, median cfDNA fragment sizes, fragment size distribution, mutant allele frequencies, fragment length, fragment size distribution, fragment end motifs, preferred end coordinates, breakpoint motifs, methylation frequencies, repeat element representation in cfDNA, fragment coverage, amount of mitochondrial derived fragments, concentrations of cfDNA or combinations thereof.
3. The method of claim 2, wherein a small cfDNA fragment comprises about 80 base pairs (bp) to about 150 bp.
4. The method of claim 2 or 3, wherein a large cfDNA fragment comprises about 151 bp to about 300 bp.
5. The method of any of claims 2 to 4, wherein the small to large cfDNA ratios are GC corrected.
6. The method of any of claims 2 to 5, wherein the cfDNA fragmentome profile comprises the sequence coverage of small cfDNA fragments in windows across the genome. 127 168948691.
17. The method of any of claims 2 to 6, wherein the cfDNA fragmentome profile comprises the sequence coverage of cfDNA fragments in windows across the genome.
8. The method of any of claims 2 to 7, wherein the cfDNA fragmentome profile comprises the sequence coverage of small and large cfDNA fragments in windows across the genome.
9. The method of any of claims 1 to 8, wherein the cfDNA fragmentome profiles in subjects identified and diagnosed early with a disease, are altered across the genome.
10. The method of any of claims 1 to 9, wherein the cfDNA fragmentome profiles in subjects identified and diagnosed early with a disease, have greater heterogeneity across the genome as compared to healthy subjects.
11. The method of any of claims 1 to 10, further comprising identifying short nucleic acid sequences (kmers) in genomic or cell free DNA; selecting kmers occurring in a single repeat type and identifying unique kmers of repeat element types; wherein the unique kmers identify one or a plurality of repeat element types.
12. The method of claim 11, wherein genome wide differences in subjects identified and diagnosed early with a disease, have greater heterogeneity across the genome in repeat elements, long interspersed nuclear elements (LINEs), short interspersed nuclear elements (SINEs), long terminal repeats (LTRs), satellites, transposable elements, RNA elements or combinations thereof, as compared to healthy subjects.
13. The method of any one of claims 1-12, wherein the diseases comprise: neurodegenerative diseases, inflammatory diseases, cancer and non-neoplastic dysregulation of cellular growth.
14. The method of any one of claims 1-13, wherein the cfDNA fragmentation profile for a specific disease is a biomarker profile diagnostic of a specific disease.
15. A method of early detection and treatment of a neurodegenerative disease comprising: extracting cfDNA from a subject’s sample; isolating cell free DNA (cfDNA) from the subject’s biological samples; conducting low coverage whole genome sequencing of the subject’s isolated cfDNA to generate genomic libraries of sequenced cfDNA fragments; 128 168948691.1determining a cfDNA fragmentome profile of the subject’s sequenced DNA fragments comprising evaluating one or more cfDNA fragment characteristics; comparing the cfDNA fragmentome profiles of the subject to reference cfDNA fragmentome profiles from healthy subjects to detect and diagnose the neurodegenerative disease; and, treating the subject.
16. The method of claim 15, wherein the one or more cfDNA fragment characteristics comprises calculating ratios of small to large cfDNA fragments, sequences of one or more cfDNAs, median cfDNA fragment sizes, fragment size distribution, mutant allele frequencies, fragment length, fragment size distribution, fragment end motifs, preferred end coordinates, breakpoint motifs, methylation frequencies, fragment coverage, amount of mitochondrial derived fragments, concentrations of cfDNA or combinations thereof.
17. The method of claims 15 or 16, wherein concentrations of cfDNA are higher in subjects with a neurodegenerative disease as compared to as compared to healthy subjects.
18. The method of any of claims 15 to 17, wherein the cfDNA fragmentome profiles in subjects identified and diagnosed early with a neurodegenerative disease, are altered across the genome.
19. The method of any of claims 15 to 18, wherein the cfDNA fragmentome profiles in subjects identified and diagnosed early with a disease, have greater heterogeneity across the genome as compared to healthy subjects.
20. The method of any of claims 15 to 19, further comprising identifying short nucleic acid sequences (kmers) in genomic or cell free DNA; selecting kmers occurring in a single repeat type and identifying unique kmers of repeat element types; wherein the unique kmers identify one or a plurality of repeat element types.
21. The method of claim 20, wherein genome wide differences in subjects identified and diagnosed early with a disease, have greater heterogeneity across the genome in repeat elements, long interspersed nuclear elements (LINEs), short interspersed nuclear elements (SINEs), long 129 168948691.1terminal repeats (LTRs), satellites, transposable elements, RNA elements or combinations thereof, as compared to healthy subjects.
22. The method of any one of claims 15-21 wherein the neurodegenerative disease is Alzheimer’s disease.
23. A method of early detection and treatment of inflammatory diseases, cancer or non- neoplastic dysregulation of cellular growth diseases comprising: extracting cfDNA from a subject’s sample; isolating cell free DNA (cfDNA) from the subject’s biological samples; conducting low coverage whole genome sequencing of the subject’s isolated cfDNA to generate genomic libraries of sequenced cfDNA fragments; determining a cfDNA fragmentome profile of the subject’s sequenced DNA fragments comprising evaluating one or more cfDNA fragment characteristics; comparing the cfDNA fragmentome profiles of the subject to reference cfDNA fragmentome profiles from healthy subjects to detect and diagnose the neurodegenerative disease; and, treating the subject.
24. The method of claim 23, wherein the one or more cfDNA fragment characteristics comprises calculating ratios of small to large cfDNA fragments, sequences of one or more cfDNAs, median cfDNA fragment sizes, fragment size distribution, mutant allele frequencies, fragment length, fragment size distribution, fragment end motifs, preferred end coordinates, breakpoint motifs, methylation frequencies, concentrations of cfDNA or combinations thereof.
25. The method of claims 23 or 24, wherein the cfDNA fragmentome profiles in subjects identified and diagnosed early with inflammatory diseases, cancer or non-neoplastic dysregulation of cellular growth diseases, are altered across the genome.
26. The method of any of claims 23 to 25, wherein the cfDNA fragmentome profiles in subjects identified and diagnosed early with inflammatory diseases, cancer or non-neoplastic 130 168948691.1dysregulation of cellular growth diseases, have greater heterogeneity across the genome as compared to healthy subjects.
27. The method of any of claims 23 to 26, further comprising identifying short nucleic acid sequences (kmers) in genomic or cell free DNA; selecting kmers occurring in a single repeat type and identifying unique kmers of repeat element types; wherein the unique kmers identify one or a plurality of repeat element types.
28. The method of claim 27, wherein genome wide differences in subjects identified and diagnosed early with a disease, have greater heterogeneity across the genome in repeat elements, long interspersed nuclear elements (LINEs), short interspersed nuclear elements (SINEs), long terminal repeats (LTRs), satellites, transposable elements, RNA elements or combinations thereof, as compared to healthy subjects.
29. The method of any one of claims 23-28, wherein the cfDNA fragmentation profile for a specific disease is a biomarker profile diagnostic of a specific disease.
30. The method of any one of claims 23-29, wherein the inflammatory diseases, cancer or non-neoplastic dysregulation of cellular growth diseases include Liver Cirrhosis and benign adnexal masses.
31. A method of early detection and treatment of a liver disease, comprising: extracting cfDNA from a subject’s sample; isolating cell free DNA (cfDNA) from the subject’s biological samples; conducting low coverage whole genome sequencing of the subject’s isolated cfDNA to generate genomic libraries of sequenced cfDNA fragments; determining a cfDNA fragmentome profile of the subject’s sequenced DNA fragments comprising evaluating one or more cfDNA fragment characteristics; comparing the cfDNA fragmentome profiles of the subject to reference cfDNA fragmentome profiles from healthy subjects to detect and diagnose the Liver Cirrhosis; and, treating the subject. 131 168948691.
132. The method of claim 31, wherein the one or more cfDNA fragment characteristics comprises calculating ratios of small to large cfDNA fragments, sequences of one or more cfDNAs, median cfDNA fragment sizes, fragment size distribution, mutant allele frequencies, fragment length, fragment size distribution, fragment end motifs, preferred end coordinates, breakpoint motifs, methylation frequencies, concentrations of cfDNA or combinations thereof.
33. The method of claims 31 or 32, wherein the cfDNA fragmentome profiles in subjects identified and diagnosed early with inflammatory diseases, cancer or non-neoplastic dysregulation of cellular growth diseases, are altered across the genome.
34. The method of any of claims 31 to 33, wherein the cfDNA fragmentome profiles in subjects identified and diagnosed early with inflammatory diseases, cancer or non-neoplastic dysregulation of cellular growth diseases, have greater heterogeneity across the genome as compared to healthy subjects.
35. The method of any of claims 31 to 34 further comprising identifying short nucleic acid sequences (kmers) in genomic or cell free DNA; selecting kmers occurring in a single repeat type and identifying unique kmers of repeat element types; wherein the unique kmers identify one or a plurality of repeat element types.
36. The method of claim 35, wherein genome wide differences in subjects identified and diagnosed early with a disease, have greater heterogeneity across the genome in repeat elements, long interspersed nuclear elements (LINEs), short interspersed nuclear elements (SINEs), long terminal repeats (LTRs), satellites, transposable elements, RNA elements or combinations thereof, as compared to healthy subjects.
37. The method of any one of claims 31-36 wherein the cfDNA fragmentation profile for a specific disease is a biomarker profile diagnostic of a specific liver disease.
38. A method of early detection and treatment of hepatitis, non alcoholic fatty liver disease, metabolic associated steatotic liver disease, steatosis, fibrosis and / or cirrhosis, comprising: extracting cfDNA from a subject’s sample; isolating cell free DNA (cfDNA) from the subject’s biological samples; 132 168948691.1conducting low coverage whole genome sequencing of the subject’s isolated cfDNA to generate genomic libraries of sequenced cfDNA fragments; determining a cfDNA fragmentome profile of the subject’s sequenced DNA fragments comprising evaluating one or more cfDNA fragment characteristics; comparing the cfDNA fragmentome profiles of the subject to reference cfDNA fragmentome profiles from healthy subjects to detect and diagnose the hepatitis, non alcoholic fatty liver disease, metabolic associated steatotic liver disease, steatosis, fibrosis and / or cirrhosis; and, treating the subject.
39. The method of claim 38, wherein the one or more cfDNA fragment characteristics comprises calculating ratios of small to large cfDNA fragments, sequences of one or more cfDNAs, median cfDNA fragment sizes, fragment size distribution, mutant allele frequencies, fragment length, fragment size distribution, fragment end motifs, preferred end coordinates, breakpoint motifs, methylation frequencies, concentrations of cfDNA or combinations thereof.
40. The method of claims 38 or 39 wherein the cfDNA fragmentome profiles in subjects identified and diagnosed early with inflammatory diseases, cancer or non-neoplastic dysregulation of cellular growth diseases, are altered across the genome.
41. The method of any of claims 38 to 40 wherein the cfDNA fragmentome profiles in subjects identified and diagnosed early with inflammatory diseases, cancer or non-neoplastic dysregulation of cellular growth diseases, have greater heterogeneity across the genome as compared to healthy subjects.
42. The method of any of claims 38 to 41 further comprising identifying short nucleic acid sequences (kmers) in genomic or cell free DNA; selecting kmers occurring in a single repeat type and identifying unique kmers of repeat element types; wherein the unique kmers identify one or a plurality of repeat element types.
43. The method of claim 42 wherein genome wide differences in subjects identified and diagnosed early with a disease, have greater heterogeneity across the genome in repeat elements, 133 168948691.1long interspersed nuclear elements (LINEs), short interspersed nuclear elements (SINEs), long terminal repeats (LTRs), satellites, transposable elements, RNA elements or combinations thereof, as compared to healthy subjects.
44. The method of any one of claims 38-43 wherein the cfDNA fragmentation profile for a specific disease is a biomarker profile diagnostic of one or more of hepatitis, non alcoholic fatty liver disease, metabolic associated steatotic liver disease, steatosis, fibrosis and cirrhosis.
45. A method of determining morbidity in a subject suffering from a disease comprising: extracting cfDNA from a biological sample obtained from the subject, conducting whole-genome sequencing of the subject’s isolated cfDNA, generating fragmentation profiles comprising the size and coverage distribution of fragments in nonoverlapping regions, and assessing the fragmentation profiles with respect to a morbidity index, and determining morbidity in the subject.
46. The method of claim 45 wherein the whole-genome sequencing of the subject’s isolated cfDNA is conducted at an average of about 1.5x coverage, 47. The method of claim 45 or 46 wherein fragmentation profiles comprising the size and coverage distribution of fragments are generated in nonoverlapping 5Mb regions.
48. The method of any one of claims 45 to 47 further comprising, after extracting cfDNA from a biological sample obtained from the subject, constructing DNA libraries with size selection for cfDNA of nucleosomal origin, and thereafter conducting whole-genome sequencing of the subject’s isolated cfDNA.
49. The method of any one of claims 45 to 48 further comprising, after conducting whole- genome sequencing of the subject’s isolated cfDNA, measuring cfDNA concentration and cfDNA lengths, and thereafter generating fragmentation profiles comprising the size and coverage distribution of fragments in nonoverlapping regions. 134 168948691.1measuring cfDNA concentrations and cfDNA lengths, 50. The method of any one of claims 45 to 49 wherein the morbidity index is a Charlson Comorbidity Index (CCI).
51. The method of any one of claims 45 to 50 wherein cfDNA concentrations are measured by amount of extracted cfDNA in a nucleosomal length range excluding larger genomic DNA.
52. The method of any one of claims 45 to 51 wherein cfDNA concentrations are high as compared to healthy subjects.
53. The method of any one of claims 45 to 52 wherein high cfDNA concentrations correlated with a CCI 3+ or higher.
54. The method of any one of claims 45 to 53 wherein shorter cfDNA fragments correlate with a high CCI (3+) as compared to those with low CCI (0-2).
55. The method of any one of claims 45 to 54 wherein a subject is identified as having a high concentration of cfDNA as compared to a healthy subject and / or has shorter cfDNA fragments as compared to a healthy individual is identified as having a higher morbidity.
56. The method of any one of claims 45 to 55 further comprising generating a fragmentation morbidity index by utilizing a machine learning model using penalized logistic regression with 10-fold cross-validation to distinguish between subjects with high and low CCI from each other.
57. The method of claim 56 wherein features utilized in the machine learning model include a fraction of short (from about 100-to about 150 bp) to long (from about 151- to about 220 bp) cfDNA fragments in each of the bins 58. The method of claim 56 or 57 wherein the fragmentation morbidity index is generated by z-scores representing chromosomal arm-level changes. 135 168948691.
159. The method of any one of claims 56 to 58 wherein the fragmentation morbidity index predicts survival. 136 168948691.1
Citation Information
Patent Citations
Cell-free DNA for assessing and / or treating cancer
US20200131571A1