Detection of liver cancer using cell-free DNA fragmentation

The analysis of cfDNA fragmentation profiles using whole-genome sequencing and machine learning addresses the limitations of current liver cancer screening by offering a sensitive and cost-effective diagnostic tool for early detection.

JP2025538137APending Publication Date: 2025-11-26JOHNS HOPKINS UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025525596
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-06
Filing Date
2023-11-06
Publication Date
2025-11-26

AI Technical Summary

Technical Problem

Current screening methods for liver cancer, such as ultrasound imaging and blood-based tests, have limited sensitivity and specificity, especially in high-risk populations, and there is a need for non-invasive and cost-effective approaches to improve early detection.

Method used

A method involving the analysis of single cell-free DNA (cfDNA) fragmentation profiles using whole-genome sequencing and machine learning to identify changes in genome-wide fragmentation patterns, correlating with transcription factor binding and chromatin structure, to diagnose liver cancer.

Benefits of technology

This approach provides a highly sensitive and specific non-invasive method for early detection of liver cancer, improving diagnostic accuracy and reducing costs compared to existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025538137000001_ABST
    Figure 2025538137000001_ABST
Patent Text Reader

Abstract

The method for liver cancer detection uses a combination of genome-wide mutations and fragmentation characteristics of cfDNA to facilitate cancer screening. TIFF2025538137000008.tif97157
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 423,003, filed November 6, 2022, the entire contents of which are incorporated herein by reference.

[0002] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH This invention was made with United States government support under grants GM136577, CA121113, CA006973, and CA233259 awarded by the National Institutes of Health. The United States government reserves certain rights in this invention. [Background technology]

[0003] background Liver cancer contributes to staggering morbidity and mortality worldwide, with over 900,000 new cases diagnosed and over 800,000 deaths annually (1). In the United States, liver cancer is one of the few cancers showing an increase in incidence and mortality over the past 20 years. Ninety percent of liver cancer cases are hepatocellular carcinoma (HCC), and survival rates depend heavily on the stage of the disease at diagnosis. Five-year survival rates are 34% for localized disease (44% of patients), 12% for regional disease (27% of patients), and 3% for distant metastases (18% of patients) (2). There are large, well-defined populations at significantly increased risk for HCC, including individuals with chronic hepatitis B (HBV) infection or cirrhosis from various causes, including hepatitis C (HCV) (3), nonalcoholic fatty liver disease (NAFLD) (4), heavy alcohol use (5), aflatoxin, and other conditions (6). Worldwide, 350 million people have chronic viral hepatitis infection, and 50 million have cirrhosis (7). In the United States, 4.5 million people have chronic HCV, and 29 million have been diagnosed with NAFLD. One-third of people with cirrhosis and 25–40% of people with HBV will develop HCC during their lifetime, representing an annual risk of up to 8% for cirrhotic patients (8). NAFLD is a growing subset of people at risk for liver cancer, accounting for 29 million people in the United States, and 20% of HCC cases in this population occur without cirrhosis (9). Medical associations worldwide recommend screening for the highest-risk populations, currently using abdominal ultrasound imaging, regardless of the presence or absence of alpha-fetoprotein (AFP). However, overall compliance with international guidelines remains low, with fewer than one in five eligible individuals worldwide receiving any level of surveillance, and fewer than 2% following recommended screening (10–12). Many factors contribute to low adherence to screening guidelines, including the identification of high-risk individuals and the infrastructure and personnel requirements of imaging-based screening methods ( 11 ).Current screening tests, including ultrasound imaging with or without AFP, have demonstrated limited sensitivity, ranging from 47% to 84%, and specificity, ranging from 67% to over 90%. (13) Additionally, the lack of noninvasive diagnostic approaches for NAFLD suggests that there is an increasing population currently not included in HCC screening recommendations. Therefore, there is a strong need worldwide for the development of accessible and sensitive screening approaches for HCC.

[0004] One recent approach to overcoming these challenges is the development of novel blood-based cell-free DNA (cfDNA) biomarkers for cancer detection. Somatic mutation-based approaches have been used as biomarkers for liver cancer but are limited by the need for tissue-based mutation identification and the limited number of detectable changes in plasma (14). Site-specific and genome-wide methylation profiling, as well as copy number alterations, also offer feasible approaches for liver cancer detection, but their sensitivity in very early disease remains suboptimal (15-20). Recently developed multicancer early detection tests appear to be useful for detecting many cancers (including liver cancer) in average-risk cohorts (21), but there are no published reports of the use of these approaches in high-risk populations for HCC. Additionally, the cost of most cfDNA-based tests is far higher than estimated affordable costs for screening tests in the United States and worldwide (22). Combining these approaches with AFP has improved performance, but requires two separate tests and remains limited in early-stage disease (23). Summary of the Invention

[0005] overview Provided herein is the non-invasive and ultra-sensitive analysis of single cell-free DNA (cfDNA) molecule to detect the change in the fragmentation profile of genome-wide.It is found that liver cancer patients have the fragmentation profile that is changed compared with healthy people, because of the change in genome and chromatin in liver cancer, for example, from the region related to liver-specific transcription factor.

[0006] Therefore, in one aspect, the method for diagnosing liver disease or disorder in subject comprises: separating circulating cell-free DNA (cfDNA) from subject; carrying out whole genome sequencing of cfDNA molecules to create genome library and fragmentation profile; comparing this fragmentation profile with healthy and / or reference genome; and diagnosing whether this subject has liver disease or disorder.In certain embodiments, the fragmentation profile of the subject without cancer is consistent, but the fragmentation profile of the subject with liver cancer varies greatly.In certain embodiments, this method further comprises identifying the cellular origin of cfDNA fragmentation profile, and identifying the cellular origin of cfDNA fragmentation profile comprises comparing genome-wide fragmentome profile with high-throughput sequencing chromosome conformation capture (Hi-C).In certain embodiments, the cellular origin of the cfDNA fragmentation profile of healthy subject corresponds to lymphoblastoid cell as cellular origin. In certain embodiments, the cellular origin of the cfDNA fragmentation profile of the subject with liver cancer corresponds to the cfDNA fragmentation profile of the chromatin compartment of peripheral blood cells.In certain embodiments, this method further comprises determining whether the cfDNA fragmentation profile is correlated with the DNA binding change of transcription factor.In certain embodiments, the transcription factor DNA binding site is determined by calculating the total cfDNA coverage of all identified transcription factor DNA binding sites compared with the entire adjacent genome coverage, to create a single metric for each transcription factor per sample.In certain embodiments, the transcription factor DNA binding site identified in the subject with liver cancer is compared with the transcription factor DNA binding site in healthy individuals.In certain embodiments, the cfDNA fragmentation profile of the subject with liver cancer is correlated with the DNA binding change of transcription factor.In certain embodiments, this method further comprises determining the gain or loss of chromosome in the subject with liver cancer compared with healthy individuals.In certain embodiments, this method further comprises a machine learning model for determining the change in cfDNA fragmentation profile, and this machine learning model classifies this object as cancer patient according to the cfDNA fragmentation profile of object.In certain embodiments, this machine learning model generates the score for each object based on the combination of regional and large-scale fragmentation profile.In certain embodiments, the generated score is for diagnosing liver cancer stage.

[0007] In another aspect, systems and methods are disclosed for detecting cancer by performing whole genome sequencing of cfDNA molecules to generate a genomic library and a fragmentation profile and inputting the data into a computer memory; and running a machine learning model to determine changes in the cfDNA fragmentation profile, which classifies a subject as a cancer patient based on the subject's cfDNA fragmentation profile.

[0008] In another aspect, a method for diagnosing liver cancer comprises the following steps: isolate circulating cell-free DNA (cfDNA) from biological sample, and perform genome sequencing of cfDNA molecules to create genome library and fragmentation profile;Compare genome-wide fragmentome profile with high-throughput sequencing chromosome structure capture (Hi-C) to identify the cellular origin of the cfDNA fragmentation profile;Correlate the cfDNA fragmentation profile with the change of DNA binding of transcription factor;Compare with healthy individuals, determine the increase or decrease of chromosomes in liver cancer subjects;Compare with healthy individuals, classify the subject as cancer patients according to the subject's cfDNA fragmentation profile, and execute a machine learning model to determine the change in the cfDNA fragmentation profile;Thereby, diagnose liver cancer and administer cancer treatment to the subject.In certain embodiments, the cellular origin of the cfDNA fragmentation profile of healthy individuals corresponds to lymphoblastoid cells as the cellular origin.In certain embodiments, the cellular origin of the cfDNA fragmentation profile of the subject with liver cancer corresponds to the cfDNA fragmentation profile of the chromatin compartment of peripheral blood cells. In certain embodiments, transcription factor DNA binding sites are determined by calculating the total cfDNA coverage of all identified transcription factor DNA binding sites compared with the entire adjacent genome coverage, to generate a single metric for each transcription factor per sample.In certain embodiments, the identified transcription factor DNA binding sites in subjects with liver cancer are compared with the transcription factor DNA binding sites in healthy individuals.In certain embodiments, the cfDNA fragmentation profile of subjects with liver cancer correlates with the altered DNA binding of transcription factors.In certain embodiments, machine learning model generates a score for each subject based on the combination of regional and large-scale fragmentation profiles.In certain embodiments, the generated score is used to diagnose liver cancer stage.In certain embodiments, the generated score is used to diagnose hepatocellular carcinoma (HCC).In certain embodiments, the generated score is used for the differential diagnosis of liver disease, including cirrhosis.In certain embodiments, treatments include surgery, adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, hormone therapy, cytotoxic therapy, immunotherapy, adoptive T cell therapy, targeted therapy, and combinations thereof.

[0009] definition Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those skilled in the art to which this invention belongs. Furthermore, terms, for example, terms defined in commonly used dictionaries, should be interpreted to have a meaning consistent with the meaning in the context of the relevant art, and unless expressly defined herein, they should not be interpreted in an idealized or overly formal sense.

[0010] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly dictates otherwise. Furthermore, to the extent the terms "including," "includes," "having," "has," "with," or variations thereof, are used in the detailed description and / or claims, such terms are intended to be inclusive, similar to the term "comprising."

[0011] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one of ordinary skill in the art. This depends in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within one or more standard deviations, as practiced in the art. Alternatively, "about" can mean within 20%, 10%, 5%, or 1% of a given value or range. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude of within 5-fold of a value, and can also mean within 2-fold of a value. When specific values ​​are described in this application and claims, unless otherwise specified, the term "about" should be assumed to mean within an acceptable error range for the particular value.

[0012] The terms "aligned," "alignment," "mapped," or "aligning" or "mapping" refer to one or more sequences that are identified as matching, in terms of nucleic acid molecule order, with a known sequence from a reference genome. Such alignments can be performed manually or by computer algorithms. Examples include the Efficient Local Alignment of Nucleotide Data (ELAND) computer program distributed as part of the Illumina Genomics Analysts pipeline. The matching of sequence reads in an alignment can be a 100% sequence match or less than 100% (non-perfect match).

[0013] The term "cancer" as used herein refers to a disease, condition, trait, genotype, or phenotype characterized by unregulated cell proliferation or replication, as known in the art. The terms "neoplasm" and "tumor" refer to abnormal tissue that grows faster than normal due to cell proliferation and continues to grow even after the stimulus that caused the proliferation is removed. Such abnormal tissue shows partial or complete loss of structural organization and functional coordination with normal tissue, and can be benign (such as benign tumor) or malignant (such as malignant tumor). Examples of cancer include liver cancer (including hepatocellular carcinoma (HCC)), lung cancer (including non-small cell lung cancer), gastric cancer, colorectal cancer, and, for example, leukemias, such as acute myeloid leukemia (AML), chronic myeloid leukemia (CML), acute lymphocytic leukemia (ALL), and chronic lymphocytic leukemia, AIDS-related cancers, such as Kaposi's sarcoma; breast cancer; bone cancers, such as osteosarcoma, chondrosarcoma, Ewing's sarcoma, fibrosarcoma, giant cell tumor, adamantinoma, and chordoma; brain cancers, such as meningioma, glioblastoma, low-grade astrocytoma, oligodendrocytoma, pituitary tumor, cytomegalovirus (CYP2), thyroid cancer, and thyroid cancer. Cancers of the head and neck, including various lymphomas, e.g., mantle cell lymphoma, non-Hodgkin's lymphoma, adenoma, squamous cell carcinoma, laryngeal cancer, gallbladder and bile duct cancer, retinal cancer, e.g., retinoblastoma, esophageal cancer, gastric cancer, multiple myeloma, ovarian cancer, uterine cancer, thyroid cancer, testicular cancer, endometrial cancer, melanoma, bladder cancer, prostate cancer, pancreatic cancer, sarcoma, Wilms' tumor, cervical cancer, head and neck cancer, skin cancer, nasopharyngeal carcinoma, liposarcoma, epithelial carcinoma, renal cell carcinoma, gallbladder adenocarcinoma, parotid adenocarcinoma, endometrial sarcoma, multidrug resistant cancer; and proliferative diseases and conditions, e.g., angiogenesis associated with tumor angiogenesis.

[0014] The terms "cell-free nucleic acid," "cell-free polynucleotide," "cell-free DNA," or "cfDNA" refer to nucleic acid fragments circulating in an individual's body (e.g., in the bloodstream) and derived from one or more healthy cells and / or one or more cancer cells. Additionally, cfDNA may originate from other sources, such as viruses, fetuses, etc.

[0015] The term "circulating tumor DNA" or "ctDNA" refers to nucleic acid fragments derived from tumor cells or other types of cancer cells that may be released into an individual's bloodstream as a result of biological processes such as apoptosis or necrosis of dying cells, or that may be actively released by viable tumor cells.

[0016] As used herein, the terms "comprising," "comprise," or "comprised," and their conjugations, are intended to be inclusive or open-ended, allowing for additional elements with respect to a defined or described element of an item, composition, apparatus, method, process, system, etc., thereby indicating that the defined or described item, composition, apparatus, method, process, system, etc. includes the specified element, or equivalents thereof, as appropriate, and that other elements are included and still fall within the scope / definition of the defined item, composition, apparatus, method, process, system, etc.

[0017] "Diagnostic" or "diagnosed" means determining the presence or severity of a pathological condition. Diagnostic methods vary in their sensitivity and specificity. The "sensitivity" of a diagnostic assay is the percentage of diseased individuals who test positive (percent "true positives"). Diseased individuals not detected by the assay are "false negatives." Subjects who do not have the disease and test negative in the assay are called "true negatives." The "specificity" of a diagnostic assay is 1 minus the false positive rate, where the "false positive" rate is defined as the proportion of individuals without the disease who test positive. While a particular diagnostic method may not definitively diagnose a condition, it suffices if the method provides a positive indication that aids in diagnosis.

[0018] As used herein, "effective amount" means an amount that provides a therapeutic or prophylactic benefit.

[0019] As used herein, the terms "fragmentation profile," "position-dependent differences in fragmentation patterns," and "position-dependent differences in fragment size and coverage across the genome" are synonymous and can be used interchangeably. In some embodiments, determining a cfDNA fragmentation profile in a mammal can be used to identify a mammal as having cancer. For example, cfDNA fragments obtained from a mammal (e.g., derived from a sample obtained from the mammal) can be subjected to low-coverage whole genome sequencing, and the sequenced fragments can be mapped to the genome (e.g., in non-overlapping windows) and evaluated to determine a cfDNA fragmentation profile. As described herein, the cfDNA fragmentation profile of a mammal with cancer is more heterogeneous (e.g., in terms of fragment length) than the cfDNA fragmentation profile of a healthy mammal (e.g., a mammal without cancer). Thus, the present disclosure also provides methods and materials for evaluating, monitoring, and / or treating a mammal (e.g., a human) with or suspected of having cancer. In some embodiments, this document provides methods and materials for identifying a mammal as having cancer. For example, a sample (e.g., blood sample) obtained from a mammal can be evaluated to determine the presence of cancer in the mammal, and optionally the primary tissue of cancer, based at least in part on the cfDNA fragmentation profile of the mammal. In some embodiments, methods and materials are provided for monitoring whether a mammal has cancer. For example, a sample (e.g., blood sample) obtained from a mammal can be evaluated to determine the presence of cancer in the mammal, based at least in part on the cfDNA fragmentation profile of the mammal. In some embodiments, methods and materials are provided for determining whether a mammal has cancer and for treating the mammal by administering one or more cancer treatments to the mammal. For example, a sample (e.g., blood sample) obtained from a mammal can be evaluated to determine whether the mammal has cancer, based at least in part on the cfDNA fragmentation profile of the mammal, and one or more cancer treatments can be administered to the mammal.

[0020] As used herein, the "frequency" of a mutation is defined as the number of variants per million positions evaluated across all sequenced DNA molecules.

[0021] The term "genomic nucleic acid" or "genomic DNA" refers to nucleic acid, including chromosomal DNA, derived from one or more healthy (e.g., non-tumor) cells. In various embodiments, genomic DNA can be extracted from cells derived from blood cell lineages, such as white blood cells (WBCs).

[0022] The term " mutation profile " used herein refers to the type and frequency of mutation observed in the bottle across the genome.The mutation profile comparison between the genomic region that is more frequently changed in cancer and the mutation profile that is derived from the region that is more frequently mutated in normal cfDNA can be used to determine multi-region difference.

[0023] "Optional" or "optionally" means that the subsequently described event or circumstance may or may not occur, and that the description includes cases where the event or circumstance occurs and cases where it does not occur.

[0024] As used in this specification and the appended claims, the term "or" is generally used in its sense inclusive of "and / or" unless the content clearly dictates otherwise.

[0025] "Parenteral" administration of the immunogenic compositions includes, for example, subcutaneous (sc), intravenous (iv), intramuscular (im), or intrasternal injection or infusion techniques.

[0026] The terms "patient" or "individual" or "subject" are used interchangeably herein and refer to a mammalian subject to be treated, with human patients being preferred. In some embodiments, the methods of the present invention are used in the development of animal models of disease in laboratory animals, including, but not limited to, rodents, including mice, rats, and hamsters, and primates, for veterinary use.

[0027] As used herein, the term "reference genome" may refer to a digital or previously identified nucleic acid sequence database assembled as a representative of a species or subject. A reference genome may be assembled from nucleic acid sequences from multiple subjects, samples, or organisms and does not necessarily represent the nucleic acid makeup of a single individual. A reference genome may be used to map sequencing reads from a sample to chromosomal locations. For example, reference genomes for human subjects and many other organisms can be found at the National Center for Biotechnology Information's ncbi.nlm.nih.gov.

[0028] The term "read segment" or "read" refers to any nucleotide sequence comprising a sequence read obtained from an individual and / or a nucleotide sequence derived from an initial sequence read from a sample obtained from an individual.

[0029] The terms "sample," "patient sample," "biological sample," and the like encompass a variety of sample types obtained from a patient, individual, or subject and can be used in diagnostic, prognostic, and / or monitoring assays. Patient samples may be obtained from healthy individuals, diseased patients, or lung cancer patients. In certain embodiments, a "provided" sample may be obtained by the person (or machine) performing the assay, or may be obtained by another person (or machine) and transferred to the person (or machine) performing the assay. Furthermore, a sample obtained from a patient can be divided, and only a portion may be used for diagnosis. Furthermore, the sample, or a portion thereof, can be stored under conditions that preserve the sample for later analysis. This definition specifically encompasses blood and other liquid samples of biological origin (including, but not limited to, peripheral blood, serum, plasma, umbilical cord blood, amniotic fluid, cerebrospinal fluid, urine, saliva, feces, and synovial fluid), solid tissue samples, such as biopsy specimens or tissue cultures or cells derived therefrom and their progeny. In certain embodiments, the sample comprises cerebrospinal fluid. In certain embodiments, the sample comprises a blood sample. In other embodiments, the sample comprises a plasma sample. In yet another embodiment, a serum sample is used. The definition of "sample" also includes samples that have been manipulated in any way after procurement, for example, by centrifugation, filtration, precipitation, dialysis, chromatography, treatment with reagents, washed, or enriched for a particular cell population. These terms further encompass clinical samples, including cells in culture, cell supernatants, tissue samples, organs, and the like. Samples can also include fresh-frozen and / or formalin-fixed, paraffin-embedded tissue blocks, such as blocks prepared from clinical or pathological biopsies, or blocks prepared for pathological analysis or immunohistochemical studies.

[0030] The term "sequence read" refers to a nucleotide sequence read from a sample obtained from an individual. Sequence reads can be obtained by various methods known in the art.

[0031] As defined herein, a "therapeutically effective" amount (i.e., an effective dosage) of a compound or agent means an amount sufficient to produce a desired therapeutic (e.g., clinical) result. The composition can be administered from one or more times daily, including once every other day, to one or more times weekly. One of ordinary skill in the art will appreciate that certain factors, including, but not limited to, the severity of the disease or disorder, previous treatments, the overall health and / or age of the subject, and other diseases present, may influence the dosage and timing required to effectively treat a subject. Furthermore, treatment of a subject with a therapeutically effective amount of a compound of the invention can include a single treatment or a series of treatments.

[0032] As used herein, the terms "treat," "treating," "treatment," and the like refer to reducing or ameliorating a disorder and / or symptoms associated with a disorder. Although not excluded, it is understood that treating a disorder or condition does not require that the disorder, condition, or symptoms associated therewith be completely eliminated.

[0033] Gene: All genes, gene names, and gene products disclosed herein are intended to correspond to homologs from any species to which the compositions and methods disclosed herein are applicable. When a gene or gene product from a particular species is disclosed, it is understood that this disclosure is intended to be illustrative only and should not be construed as limiting unless clearly indicated by the relevant context. Thus, for example, a gene or gene product disclosed herein is intended to encompass homologous and / or orthologous genes and gene products from other species.

[0034] Ranges: Throughout this application, various aspects of the invention may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, the description of a range should be considered to have specifically disclosed all possible subranges as well as individual numerical values ​​within that range. For example, the description of a range such as 1 to 6 should be considered to have specifically disclosed subranges such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., as well as individual numerical values ​​within that range, e.g., 1, 2, 2.7, 3, 4, 5, 5.3, and 6. This applies regardless of the breadth of the range.

[0035] Any composition or method provided herein can be combined with one or more of any of the other compositions and methods provided herein. [Brief explanation of the drawings]

[0036] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0037] [Figure 1]Figure 1 (including Figures 1A–1C) is a plot demonstrating that genome-wide fragmentation profiles reflect the underlying chromatin structure. Figure 1A: Fragmentation profiles of 501 individuals across 473 non-overlapping 5-mb genomic regions. The fragmentation profiles of cancer individuals show significant heterogeneity compared with non-cancer individuals with and without liver disease. Figure 1B: Comparison of plasma fragmentation features with reference A / B compartments. Track 1 shows the A / B compartments extracted from liver cancer tissue (29). Track 2 shows the median liver cancer component extracted from HCC plasma samples of 10 liver patients with a high tumor fraction due to purulent CAN (57). Track 3 shows the median fragmentation profile in the plasma of these 10 HCC samples, and Track 4 shows the median fragmentation profile in 10 healthy plasma samples. Track 5 shows the A / B compartments of lymphoblastoid cells (29). These five tracks show an example of chromosome 22, with darker shading indicating informative regions of the genome where the two reference tracks differ in terms of domain (open / closed) or size. Figure 1C: Further results comparing plasma fragmentation features with the reference A / B compartments. [Figure 2]Figure 2 (including Figures 2A-2E) is a series of plots showing fragmentation profiles in HCC patients, highlighting liver-specific transcription factors. Figure 2A: Coverage of transcription factor binding sites (TFBSs) and their surroundings for nine transcription factors showing relative coverage at binding sites with the highest degree of HCC segregation from non-cancer samples. Mean values ​​are plotted for each group, with + / - 1 standard deviation (SD) shaded. These confidence intervals (CIs) indicate segregation, highlighting the fact that differences in coverage at TFBSs can provide information about cancer status. Figure 2B: Coverage of TFBSs and their surroundings for nine TFs with the lowest degree of HCC segregation from non-cancer samples in the US / EU cohort. These CIs largely overlap, reflecting their status as poorly discriminatory TFBSs. Gene set enrichment analysis of the TFs analyzed in both hepatocellular carcinoma and lung adenocarcinoma showed that the TFs were selectively enriched in numerous pathways associated with liver cancer and lung cancer, respectively ( Figure 2C ), such as adult liver cancer and lung adenocarcinoma ( Figure 2D , Figure 2E ). [Figure 3] Figure 3 (including Figures 3A-3C) demonstrates that high-dimensional fragmentation features reflect the biology of liver cancer and are incorporated into the DELFI machine learning approach. Figure 3A: Heatmap reflecting the complexity of genome-wide fragmentation and transcription factor binding site features utilized in the DELFI machine learning approach. Each row represents a sample, and each column represents an individual genomic feature. Figure 3B: Analysis of copy number alterations in tissue from 372 TCGA liver cancers and plasma from 501 individuals reflects biological consistency. TCGA-generated copy number changes (red = gain, blue = loss) were also observed at the chromosome arm level in HCC plasma but not in cancer-free individuals. Figure 3C: Heatmap showing the contribution of individual genomic regions to the final trained DELFI model. Fragmentation features were summarized as three principal components in the model, and aneuploidy was summarized as an arm-level z-score. The top, middle, and right panels show the fragmentation component, arm-level z-score, and transcription factor binding site coefficients in the model, respectively. CNV, copy number variation. [Figure 4] Figure 4 (including Figures 4A–4D) is a series of plots demonstrating the high sensitivity and specificity of the DELFI machine learning model for detecting liver cancer. Figure 4A: DELFI scores for the US / EU cohort across liver disease and cancer stages for the screening and surveillance model. Patients with cirrhosis have DELFI scores that are, on average, higher than those without cancer or those with viral hepatitis, but lower than those across all stages of liver cancer. Liver cancer patients at all stages have relatively high DELFI scores, with stage C individuals uniformly having the highest DELFI scores. Figure 4B: ROC analysis of the US / EU general population cohort and high-risk surveillance cohort. Figure 4C: ROC analysis of the US / EU general population cohort and surveillance cohort stratified by BCLC stage, demonstrating high sensitivity and specificity across disease stages. Figure 4D: ROC analysis for the fixed surveillance model applied to the Hong Kong cohort, which includes 90 individuals with HCC (85 with BCLC stage A cancer and 5 with BCLC stage B cancer), 101 individuals with cirrhosis and viral hepatitis, and 32 individuals without cancer or liver disease. [Figure 5] Figure 5 shows a heatmap showing the contribution of individual genomic regions to the final trained DELFI surveillance model. Fragmentation features were summarized in the model as three principal components, and aneuploidy was summarized as arm-level z-scores. The top and right panels show the variable importance of the fragmentation components and arm-level z-scores in the model, respectively. [Figure 6] Figures 6A and 6B are plots showing that DELFI scores in individuals without cancer are not affected by age or gender. DELFI scores by age (Figure 6A) or gender (Figure 6B) in individuals without cancer or liver disease (n=293, 45 women, 88 men). [Figure 7] FIG. 7 is a series of plots demonstrating that DELFI scores in cancer-free individuals do not differ between racial and ethnic groups. [Figure 8]Figure 8 is a plot demonstrating that the DELFI score in cirrhotic patients (n=40) with available information was not associated with the Child-Pugh score of cirrhosis severity. [Figure 9] FIG. 9 is a series of plots demonstrating that DELFI scores vary with BMI in individuals without cancer (n=53 patients with viral hepatitis and n=78 patients with cirrhosis). [Figure 10] Figures 10A and 10B show the performance of alternative DELFI models with and without the addition of TF binding site relative coverage features. ROC analysis of the surveillance model incorporating transcription factor binding site features from liver-derived CHIP-Seq analysis in the high-risk US / EU cohort, and ROC analysis of the screening model without transcription factor binding sites for the general population US / EU cohort (Figure 10A) and by BCLC stage (Figure 10B). [Figure 11] Figure 11 shows a plot demonstrating the correlation between rank-ordered DELFI scores derived from the surveillance and screening models. Among all HCC patients (n = 75), there is a high correlation between rank-ordered scores using our high-risk DELFI model and the screening DELFI model. [Figure 12] Figures 12A and 12B are a series of plots demonstrating that DELFI scores in HCC patients were related to the size and number of tumor lesions. DELFI scores in HCC patients (n=75) were positively correlated with increasing lesion diameter (Figure 12A) and showed a positive trend with increasing lesion number (Figure 12B). [Figure 13] Figure 13 is a graph demonstrating that DELFI scores for HCC patients (n=54) with resectable disease stages (BCLC 0, A, and B) did not differ by underlying liver disease etiology. Stage C patients (n=21) had an etiology significantly associated with viral hepatitis (n=14), and therefore were not included in this analysis. [Figure 14]Figure 14 is a plot demonstrating that DELFI scores correlate with AFP levels in HCC patients. Overall, 39 of 75 individuals with HCC had AFP levels above the recommended screening threshold of 20 ng / ml, indicated by the vertical dotted line. All but three of these individuals would have been detected at the DELFI threshold (0.26), indicated by the horizontal dotted line, corresponding to 80% specificity in the high-risk cohort. Among individuals not detected by AFP, 30 of 36 were detected by DELFI. [Figure 15] FIG. 15 is a plot demonstrating that the correlation of fragmentation profiles to median non-cancer profiles did not differ among non-cancer subgroups. [Figure 16] Figure 16 is a plot demonstrating that the genome-wide fragmentation profile in the Hong Kong validation cohort shows similarity to that of the US / EU cohort. Fragmentation profiles of 223 individuals in 473 non-overlapping 5 mb genomic regions. The fragmentation profiles of cancer individuals show significant heterogeneity compared to non-cancer individuals with or without liver disease, which is similar to the findings in the US / EU cohort. [Figure 17] Figure 17 shows that chromosomal copy number changes detected in plasma are biologically consistent between the US / EU and Hong Kong cohorts, as well as in TCGA tissue samples. Red indicates copy number gain, and blue indicates copy number loss. These changes are readily observed in both tissue and HCC plasma, but are not apparent in cancer-free individuals. CNV, copy number variation. [Figure 18]Figure 18 (including Figures 18A-18D) shows modeling of the implementation of DELFI in liver cancer screening. Figure 18A: The uncertainty in the sensitivity and specificity of DELFI screening, as well as ultrasound and AFP, was modeled for a theoretical population of 100,000 high-risk individuals. The predicted distributions of the number of liver cancers detected (Figure 18B), false-negative rate (Figure 18C), and negative predictive value (Figure 18D) among these individuals incorporated variation in both liver cancer prevalence and adherence to imaging and blood-based screening. The center line of the box plot indicates the median, the upper limit of the box plot indicates the third quartile (75th percentile), the lower limit of the box plot indicates the first quartile (25th percentile), the upper whisker indicates the maximum value of the data within 1.5 times the interquartile range above the 75th percentile, and the lower whisker indicates the minimum value of the data within 1.5 times the interquartile range below the 25th percentile. Figure 1 shows that chromosomal copy number changes detected in plasma are biologically consistent between the US / EU and Hong Kong cohorts, as well as in TCGA tissue samples. Red indicates copy number gain, and blue indicates copy number loss. These changes are readily observed in both tissue and HCC plasma, but are not apparent in cancer-free individuals. CNV, copy number variation. DETAILED DESCRIPTION OF THE INVENTION

[0038] Detailed Description DELFI (DNA Evaluation of Fragments for Early Intervention) utilizes genome-wide fragmentation profiles to provide a high-performance, low-cost approach for cancer detection (21, 22). Fragmentation and methylation information has also demonstrated the ability to distinguish liver cancer patients from cancer-free individuals (24), but such an approach requires two separate methods: cfDNA library preparation and analysis. To date, no studies have validated a genome-wide approach for detecting HCC in independent groups or across different high-risk populations.

[0039] Thus, in certain embodiments, a method for diagnosing a liver disease or disorder in a subject includes isolating circulating cell-free DNA (cfDNA) from the subject; performing whole-genome sequencing of the cfDNA molecules to generate a genomic library and a fragmentation profile; comparing the fragmentation profile to a healthy individual and / or a reference genome; and diagnosing whether the subject has a liver disease or disorder.

[0040] In another embodiment, a method for diagnosing liver cancer includes the following steps: isolating circulating cell-free DNA (cfDNA) from a biological sample and performing whole-genome sequencing of the cfDNA molecules to create a genomic library and a fragmentation profile; identifying the cellular origin of the cfDNA fragmentation profile, including comparing the genome-wide fragmentome profile with high-throughput sequencing chromosome structure capture (Hi-C); correlating the cfDNA fragmentation profile with changes in DNA binding of transcription factors; determining chromosomal gains or losses in liver cancer subjects compared to healthy individuals; implementing a machine learning model to determine changes in the cfDNA fragmentation profile, classifying the subject as a cancer patient based on the subject's cfDNA fragmentation profile; thereby diagnosing liver cancer and administering cancer treatment to the subject.

[0041] cfDNA fragmentation profile: The cfDNA fragmentation profile may include one or more cfDNA fragmentation patterns. The cfDNA fragmentation pattern may include any suitable cfDNA fragmentation pattern. Examples of cfDNA fragmentation patterns include, but are not limited to, median fragment size, fragment size distribution, ratio of small cfDNA fragments to large cfDNA fragments, and cfDNA fragment coverage. In some embodiments, the cfDNA fragmentation pattern includes two or more (e.g., two, three, or four) of median fragment size, fragment size distribution, ratio of small cfDNA fragments to large cfDNA fragments, and cfDNA fragment coverage. In some embodiments, the cfDNA fragmentation profile may be a genome-wide cfDNA profile (e.g., a genome-wide cfDNA profile in a genome-wide window). In some embodiments, the cfDNA fragmentation profile may be a target region profile. The target region may be any suitable portion of the genome (e.g., a chromosomal region). Examples of chromosomal regions that can determine cfDNA fragmentation profiles as described herein include, but are not limited to, chromosome portions (e.g., 2q, 4p, 5p, 6q, 7p, 8q, 9q, 10q, 11q, 12q, and / or portions of 14q) and chromosome arms (e.g., 8q, 13q, 11q, and / or 3p chromosome arms).In some embodiments, the cfDNA fragmentation profile can include two or more target region profiles.

[0042] In some embodiments, cfDNA fragmentation profile can be used to identify the change (for example, alteration) of cfDNA fragment length.The alteration can be genome-wide alteration, or the alteration of one or more target regions / locuses.The target region can be any region that contains one or more cancer-specific alterations. In some embodiments, the cfDNA fragmentation profile can be used to identify (e.g., simultaneously identify) between about 10 alterations and about 500 alterations (e.g., between about 25 and about 500, between about 50 and about 500, between about 100 and about 500, between about 200 and about 500, between about 300 and about 500, between about 10 and about 400, between about 10 and about 300, between about 10 and about 200, between about 10 and about 100, between about 10 and about 50, between about 20 and about 400, between about 30 and about 300, between about 40 and about 200, between about 50 and about 100, between about 20 and about 100, between about 25 and about 75, between about 50 and about 250, or between about 100 and about 200 alterations).

[0043] The cfDNA fragmentation profile can be obtained using any suitable method. In some embodiments, cfDNA from a mammal (e.g., a mammal with or suspected of having cancer) can be processed into a sequencing library, which can be subjected to whole genome sequencing (e.g., low-coverage whole genome sequencing), mapped to the genome, and analyzed to determine cfDNA fragment lengths. The mapped sequences can be analyzed in non-overlapping windows covering the genome. The windows can be of any suitable size. For example, the windows can be thousands to millions of bases in length. As a non-limiting example, the windows can be about 5 megabases (Mb) in length. Any suitable number of windows can be mapped. For example, tens to thousands of windows can be mapped in the genome. For example, hundreds to thousands of windows can be mapped in the genome. The cfDNA fragmentation profile can be determined within each window.

[0044] In some embodiments, the methods and materials described herein can also include machine learning.For example, machine learning can be used to identify mutation frequency, fragmentation profile change (for example, by using cfDNA fragment coverage, cfDNA fragment size, chromosome coverage and mtDNA).

[0045] In some embodiments, determining a cfDNA fragmentation profile in a mammal can be used to identify a mammal as having cancer. For example, cfDNA fragments obtained from a mammal (e.g., derived from a sample obtained from the mammal) can be subjected to low-coverage whole genome sequencing, and the sequenced fragments can be mapped to the genome and evaluated to determine a cfDNA fragmentation profile. As described herein, the cfDNA fragmentation profile of a mammal with cancer is more heterogeneous (e.g., in terms of fragment length) than the cfDNA fragmentation profile of a healthy mammal (e.g., a mammal without cancer). Therefore, methods and materials are also provided for evaluating, monitoring, and / or treating a mammal (e.g., a human) that has or is suspected of having cancer. In some embodiments, methods and materials are provided for identifying a mammal as having cancer. For example, a sample (e.g., a blood sample) obtained from a mammal can be evaluated to determine the presence of cancer in the mammal, and optionally the tissue of origin of the cancer, based at least in part on the cfDNA fragmentation profile of the mammal. In some embodiments, methods and materials are provided for monitoring a mammal as having cancer. For example, sample (for example, blood sample) obtained from mammal can be evaluated, and based at least in part on the cfDNA fragmentation profile of this mammal, determine the existence of cancer in this mammal.In some embodiments, provide methods and materials for determining that mammal has cancer, and administer one or more cancer treatments to this mammal to treat this mammal.For example, sample (for example, blood sample) obtained from mammal can be evaluated, and based at least in part on the cfDNA fragmentation profile of this mammal, determine whether this mammal has cancer, and then administer one or more cancer treatments to this mammal.

[0046] In some embodiments, cfDNA fragmentation profile can be used to detect tumor-derived DNA.For example, by comparing the cfDNA fragmentation profile of a mammal that has or is suspected of having cancer with a reference cfDNA fragmentation profile (e.g., the cfDNA fragmentation profile of a healthy mammal and / or the nucleosomal DNA fragmentation profile of a healthy cell from a mammal that has or is suspected of having cancer), cfDNA fragmentation profile can be used to detect tumor-derived DNA.In some embodiments, the reference cfDNA fragmentation profile is a pre-created profile from a healthy mammal.For example, the method provided herein can be used to determine a reference cfDNA fragmentation profile in a healthy mammal, and the reference cfDNA fragmentation profile can be stored (e.g., in a computer or other electronic storage medium) for future comparison with a test cfDNA fragmentation profile in a mammal that has or is suspected of having cancer.In some embodiments, the reference cfDNA fragmentation profile (e.g., the stored cfDNA fragmentation profile) of a healthy mammal is determined over the entire genome.In some embodiments, the reference cfDNA fragmentation profile (e.g., the stored cfDNA fragmentation profile) of a healthy mammal is determined over a subgenomic interval.

[0047] In some embodiments, a cfDNA fragmentation profile can be used to identify a mammal (e.g., a human) as having cancer (e.g., liver cancer, colorectal cancer, lung cancer, breast cancer, gastric cancer, pancreatic cancer, bile duct cancer, and / or ovarian cancer).

[0048] The cfDNA fragmentation profile can include a cfDNA fragment size pattern. The cfDNA fragments can be of any suitable size. For example, the cfDNA fragments can be about 50 base pairs (bp) to about 400 bp in length.

[0049] The cfDNA fragmentation profile can include a cfDNA fragment size distribution. As described herein, a mammal with cancer may have a cfDNA size distribution that is more variable than that of a healthy mammal. In some embodiments, the size distribution can be within the target region. A healthy mammal (e.g., a mammal without cancer) can have a target region cfDNA fragment size distribution of about 1 or less than about 1. In some embodiments, a mammal with cancer can have a target region cfDNA fragment size distribution that is longer (e.g., 10, 15, 20, 25, 30, 35, 40, 45, 50 bp, or more, or any number of base pairs between these values) than that of a healthy mammal. In some embodiments, a mammal with cancer can have a target region cfDNA fragment size distribution that is shorter (e.g., 10, 15, 20, 25, 30, 35, 40, 45, 50 bp, or more, or any number of base pairs between these values) than that of a healthy mammal. In some embodiments, the size distribution can be a genome-wide size distribution. Healthy mammals (e.g., mammals without cancer) may have very similar genome-wide distributions of short and long cfDNA fragments. In some embodiments, mammals with cancer may have genome-wide one or more changes (e.g., increases and decreases) in cfDNA fragment size. One or more changes can be in any suitable chromosomal region of the genome. For example, one change can be in one part of one chromosome. Examples of chromosomal parts that can contain one or more changes in cfDNA fragment size include, but are not limited to, parts 2q, 4p, 5p, 6q, 7p, 8q, 9q, 10q, 11q, 12q, and 14q. For example, the change can span one chromosomal arm (e.g., the entire chromosomal arm).

[0050] The cfDNA fragmentation profile can include the ratio of small cfDNA fragments to large cfDNA fragments, and the correlation of the fragment ratio to a reference fragment ratio. As used herein, with respect to the ratio of small cfDNA fragments to large cfDNA fragments, the small cfDNA fragments can be about 100 bp to about 150 bp in length. As used herein, with respect to the ratio of small cfDNA fragments to large cfDNA fragments, the large cfDNA fragments can be about 151 bp to about 220 bp in length. A mammal with cancer may have a lower fragment ratio correlation (e.g., the correlation of the cfDNA fragment ratio to a reference DNA fragment ratio, such as a DNA fragment ratio from one or more healthy mammals) than that in a healthy mammal (e.g., 1 / 2, 1 / 3, 1 / 4, 1 / 5, 1 / 6, 1 / 7, 1 / 8, 1 / 9, 1 / 10, or less). A healthy mammal (e.g., a mammal without cancer) may have a fragment ratio correlation (e.g., the correlation of cfDNA fragment ratios to a reference DNA fragment ratio, such as DNA fragment ratios from one or more healthy mammals) of about 1 (e.g., about 0.96). In some embodiments, a mammal with cancer may have a fragment ratio correlation (e.g., the correlation of cfDNA fragment ratios to a reference DNA fragment ratio, such as DNA fragment ratios from one or more healthy mammals) that is lower on average than the fragment ratio correlation in a healthy mammal (e.g., the correlation of cfDNA fragment ratios to a reference DNA fragment ratio, such as DNA fragment ratios from one or more healthy mammals).

[0051] The cfDNA fragmentation profile can include coverage of all fragments. Coverage of all fragments can include coverage windows (e.g., non-overlapping windows). In some embodiments, coverage of all fragments can include windows of small fragments (e.g., fragments with a length of about 100 bp to about 150 bp). In some embodiments, coverage of all fragments can include windows of large fragments (e.g., fragments with a length of about 151 bp to about 220 bp).

[0052] In certain embodiments, the cfDNA fragmentation profile can be used to identify the molecular origin of cfDNA in a patient and to identify genomic and chromatin features associated with altered fragmentation.

[0053] In some embodiments, cfDNA fragmentation profile can be used to identify the primary tissue of cancer (e.g., liver cancer, colorectal cancer, lung cancer, breast cancer, gastric cancer, pancreatic cancer, bile duct cancer, or ovarian cancer).For example, cfDNA fragmentation profile can be used to identify localized cancer.When cfDNA fragmentation profile includes target region profile, one or more changes described herein can be used to identify the primary tissue of cancer.In some embodiments, one or more changes in chromosomal region can be used to identify the primary tissue of cancer.

[0054] The cfDNA fragmentation profile can be obtained using any suitable method. In some embodiments, cfDNA from a mammal (e.g., a mammal having or suspected of having cancer) can be processed into a sequencing library, which can be subjected to whole genome sequencing (e.g., low-coverage whole genome sequencing), mapped to the genome, and analyzed to determine the cfDNA fragment length. The mapped sequences can be analyzed in non-overlapping windows covering the genome. The windows can be of any suitable size. For example, the windows can be thousands to millions of bases long. As a non-limiting example, the windows can be about 5 megabases (Mb) long. Any suitable number of windows can be mapped. For example, tens to thousands of windows can be mapped in the genome. For example, hundreds to thousands of windows can be mapped in the genome. The cfDNA fragmentation profile can be determined within each window.

[0055] In some embodiments, the methods and materials described herein can also include machine learning.For example, machine learning can be used to identify altered fragmentation profile (for example, by using cfDNA fragment coverage, cfDNA fragment size, chromosome coverage and mtDNA).

[0056] In some embodiments, the methods and materials described herein can be the only method used to identify mammals (for example, humans) have liver cancer.For example, determining cfDNA fragmentation profile can be the only method used to identify mammals have liver cancer.

[0057] In some embodiments, the methods and materials described herein can be used together with one or more additional methods for determining that mammals (for example, humans) have liver cancer.The example of the method used for determining that mammals have cancer includes, but is not limited to, determining one or more cancer-specific sequence alterations, determining one or more chromosomal alterations (for example, aneuploidy and rearrangement), and determining other cfDNA alterations.For example, determining cfDNA fragmentation profile can be used together with determining one or more cancer-specific mutations in mammalian genome to determine that mammals have liver cancer.For example, determining cfDNA fragmentation profile can be used together with determining one or more aneuploidy in mammalian genome to determine that mammals have liver cancer.

[0058] Treatment method The method included in this paper comprises the steps of determining that mammal has cancer.This method comprises the steps of: extracting cell-free DNA (cfDNA) from the biological sample of object; making a genome library from the extracted cfDNA; sequencing each cfDNA molecule to obtain fragmentation profile; comparing this fragmentation profile with healthy person and / or reference genome; and diagnosing whether this object has liver disease or disorder.

[0059] In certain embodiments, a method for diagnosing liver cancer includes the following steps: isolating circulating cell-free DNA (cfDNA) from a biological sample and performing whole-genome sequencing of the cfDNA molecules to create a genomic library and a fragmentation profile; identifying the cellular origin of the cfDNA fragmentation profile, including comparing the genome-wide fragmentome profile with high-throughput sequencing chromosome structure capture (Hi-C); correlating the cfDNA fragmentation profile with changes in DNA binding of transcription factors; determining chromosomal gains or losses in liver cancer subjects compared to healthy individuals; implementing a machine learning model to determine changes in the cfDNA fragmentation profile, classifying the subject as a cancer patient based on the subject's cfDNA fragmentation profile; thereby diagnosing liver cancer and administering cancer treatment to the subject.

[0060] In certain embodiments, the subject is diagnosed with cancer, e.g., early stage cancer. In certain embodiments, the type and stage of liver cancer is identified and the subject is treated with one or more cancer therapies.

[0061] In some embodiments, methods and materials are provided herein for evaluating, monitoring, and / or treating a mammal (e.g., a human) that has or is suspected of having cancer. In some embodiments, methods and materials are provided for identifying a mammal that has liver cancer. For example, a sample (e.g., a blood sample) obtained from a mammal can be evaluated to determine whether the mammal has cancer based at least in part on the cfDNA fragmentation profile of the mammal. In some embodiments, a sample (e.g., a blood sample) obtained from a mammal can be evaluated to determine the tissue of origin of cancer in the mammal based at least in part on the cfDNA fragmentation profile of the mammal. In some embodiments, methods and materials are provided for identifying a mammal that has liver cancer, administering one or more treatments to the mammal, and treating the mammal. For example, a sample (e.g., a blood sample) obtained from a mammal can be evaluated to determine whether the mammal has liver cancer based at least in part on the cfDNA fragmentation profile of the mammal, and then administering one or more cancer treatments to the mammal. In some embodiments, methods and materials are provided for treating a mammal that has cancer. For example, a mammal that is identified as having cancer (e.g., based at least in part on the cfDNA fragmentation profile of the mammal) can be administered one or more cancer treatments to treat the mammal.In some embodiments, during or after a course of cancer treatment (e.g., any of the cancer treatments described herein), the mammal can be monitored (or selected for intensive monitoring) and / or undergo further diagnostic testing.In some embodiments, monitoring can include, as described herein, evaluating a mammal that has or is suspected of having liver cancer, for example, by evaluating a sample (e.g., a blood sample) obtained from the mammal to determine the cfDNA fragmentation profile of the mammal, and the change in the cfDNA fragmentation profile over time can be used to identify the response to treatment and / or identify the mammal as having cancer (e.g., residual cancer).

[0062] Any suitable mammal can be evaluated, monitored and / or treated as described herein.Mammal can be the mammal with liver cancer.Mammal can be the mammal suspected of having liver cancer.The examples of mammal that can be evaluated, monitored and / or treated as described herein include but are not limited to humans, primates such as monkeys, dogs, cats, horses, cows, pigs, sheep, mice and rats.For example, the human that has or is suspected of having liver cancer can be evaluated as described herein to determine cfDNA fragmentation profile, and can optionally be treated with one or more cancer treatments as described herein.

[0063] Any suitable sample from a mammal can be evaluated (e.g., evaluated for DNA fragmentation patterns) as described herein. In some embodiments, the sample can include DNA (e.g., genomic DNA). In some embodiments, the sample can include cfDNA (e.g., circulating tumor DNA (ctDNA)). In some embodiments, the sample can be a bodily fluid sample (e.g., liquid biopsy). Examples of samples that can include DNA and / or polypeptides include, but are not limited to, blood (e.g., whole blood, serum, or plasma), amniotic membrane, tissue, urine, cerebrospinal fluid, saliva, sputum, bronchoalveolar lavage fluid, bile, lymph, cyst fluid, stool, ascites, Papanicolaou smear, breast milk, and exhaled breath condensate. For example, a plasma sample can be evaluated to determine a cfDNA fragmentation profile as described herein.

[0064] The sample from mammal that is evaluated as described herein (for example, evaluated for DNA fragmentation pattern) can contain any suitable amount of cfDNA.In some embodiments, sample can contain limited amount of DNA.For example, cfDNA fragmentation profile can be obtained from the sample that contains less DNA than is usually required for other cfDNA analysis methods, such as those described in, for example, Phallen et al., 2017 Sci Transl Med 9; Cohen et al., 2018 Science 359:926; Newman et al., 2014 Nat Med 20:548; and Newman et al., 2016 Nat Biotechnol 34:547.

[0065] In some embodiments, a sample can be processed (e.g., to isolate and / or purify DNA and / or polypeptides from the sample). For example, DNA isolation and / or purification can include cell lysis (e.g., using detergents and / or surfactants), protein removal (e.g., using proteases), and / or RNA removal (e.g., using RNases). As another example, polypeptide isolation and / or purification can include cell lysis (e.g., using detergents and / or surfactants), DNA removal (e.g., using DNases), and / or RNA removal (e.g., using RNases).

[0066] The cancer may be any stage of cancer. In some embodiments, the cancer may be an early stage cancer. In some embodiments, the cancer may be an asymptomatic cancer. In some embodiments, the cancer may be residual disease and / or recurrence (e.g., after surgical resection and / or cancer therapy). The cancer may be any type of cancer. Examples of types of cancer that may be assessed, monitored, and / or treated as described herein include, but are not limited to, colorectal cancer, lung cancer, breast cancer, gastric cancer, pancreatic cancer, bile duct cancer, and ovarian cancer.

[0067] When treating a mammal having or suspected of having liver cancer as described herein, the mammal may be administered one or more cancer treatments. The cancer treatment may be any suitable cancer treatment. One or more cancer treatments described herein may be administered to the mammal at any appropriate frequency (e.g., once or multiple times over a period ranging from several days to several weeks). Examples of cancer treatments include, but are not limited to, adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, hormone therapy, cytotoxic therapy, immunotherapy, adoptive T cell therapy (e.g., T cells having chimeric antigen receptors and / or wild-type or modified T cell receptors), targeted therapy such as administration of kinase inhibitors (e.g., kinase inhibitors that target specific genetic lesions such as translocations or mutations), (e.g., kinase inhibitors, antibodies, bispecific antibodies), signal transduction inhibitors, bispecific antibodies or antibody fragments (e.g., BiTEs), monoclonal antibodies, immune checkpoint inhibitors, surgery (e.g., surgical resection), or a combination of the above. In some aspects, a cancer treatment can reduce the severity of the cancer, decrease the symptoms of the cancer, and / or decrease the number of cancer cells present in the mammal.

[0068] In some embodiments, the cancer treatment may include an immune checkpoint inhibitor, non-limiting examples of which include nivolumab (Opdivo), pembrolizumab (Keytruda), atezolizumab (Tecentriq), avelumab (Bavencio), durvalumab (Imfinzi), and ipilimumab (Yervoy).

[0069] Cancer therapy also commonly involves a variety of combination therapies with chemical- and radiation-based treatments. Combination chemotherapy includes, for example, cisplatin (CDDP), carboplatin, procarbazine, mechlorethamine, cyclophosphamide, camptothecin, ifosfamide, melphalan, chlorambucil, busulfan, nitrosoureas, dactinomycin, daunorubicin, doxorubicin, bleomycin, plicomycin, mitomycin, etoposide (VP16), tamoxifen, raloxifene, estrogen receptor binding agents, taxol, gemcitabien, navelbine, famesyl-protein transferase inhibitor, transplatinum, 5-fluorouracil, vincristine, vinblastine, and methotrexate, temazolomide (aqueous form of DTIC), or any analog or derivative variant of the foregoing. The combination of chemotherapy and biological therapy is known as biochemotherapy. Chemotherapy may also be given in successive low doses, known as metronomic chemotherapy.

[0070] Still further combination chemotherapy includes, for example, alkylating agents, such as thiotepa and cyclosphosphamide; alkylsulfonates, such as busulfan, improsulfan, and piposulfan; aziridines, such as benzodopa, carboquone, meturedopa, and uredopa; altretamine, triethylenemelamine, triethylenephosphoramide, triethylenethiophosphoramide, and trimethylolmethamine. ethylenimines and methylamelamines, including trimethylololamimes; acetogenins (especially bullatacin and bullatacinone); camptothecins (including the synthetic analog topotecan); bryostatin; callystatin; CC-1065 (including its adozelesin, carzelesin, and biceresin synthetic analogs); cryptophycins (especially cryptophycin 1 and cryptophycin 2); dolastatins; duocarmycins (including synthetic analogs KW-2189 and CB1-TM1); eluterobin; pancratistatin; sarcodictyin; spongistatins; nitrogen mustards, e.g., chlorambucil, chlornaphazine, chlorophosphamide, estramustine, ifosfamide, mechlorethamine, mechlorethamine oxide hydrochloride, melphalan, novem novembichin, phenesterine, prednimustine, trofosfamide, uracil mustard; nitrosoureas, such as carmustine, chlorozotocin, fotemustine, lomustine, nimustine, and ranimustine; antibiotics, such as enediyne antibiotics (e.g., calicheamicin, especially calicheamicin gamma II and calicheamicin omega II; dynemicins, including dynemicin A); bisphosphonates, such as clodronate; esperamicin;and neocarzinostatin chromophore and related chromoprotein enediyne antiobiotic chromophores, aclacinomycin, actinomycin, authrarnycin, azaserine, bleomycin, cactinomycin, carabicin, carminomycin, carzinophilin, chromomycinis, dactinomycin, daunorubicin, detorubicin, 6-diazo-5-oxo-L-norleucine, doxorubicin (morpholino-doxol), doxorubicin, cyanomorpholino-doxorubicin, 2-pyrrolino-doxorubicin, and deoxydoxorubicin), epirubicin, esorubicin, idarubicin, marcelomycin, mitomycins, such as mitomycin C, mycophenolic acid, nogalamycin, olivomycin, peplomycin, potfiromycin, puromycin, quelamycin, rodorubicin, streptonigrin, streptozocin, tubercidin, ubenimex, zinostaphylococcus aureus, cyclosporin, cyclosporin, cyclosporin, cyclosporin-3, cyclosporin-4, cyclosporin-5, cyclosporin-6, cyclosporin-7, cyclosporin-8, cyclosporin-9, cyclosporin-10, cyclosporin-11, cyclosporin-12, cyclosporin-13, cyclosporin-14, cyclosporin-15, cyclosporin-16, cyclosporin-17, cyclosporin-18, cyclosporin-19, cyclosporin-20, cyclosporin-21, cyclosporin-22, cyclosporin-23, cyclosporin-24, cyclosporin-25, cyclosporin-26, cyclosporin-27, cyclosporin-28, cyclosporin-29, cyclosporin-30, cyclosporin-31, cyclosporin-32, cyclosporin-33, cyclosporin-34, cyclosporin-35, cyclosporin-36, cyclosporin-37, cyclosporin-38, cyclosporin-39, cyclosporin-40 antimetabolites such as methotrexate and 5-fluorouracil (5-FU); folic acid analogs such as denopterin, pteropterin, trimetrexate; purine analogs such as fludarabine, 6-mercaptopurine, thiamiprine, thioguanine; pyrimidine analogs such as ancitabine, azacitidine, 6-azauridine, carmofur, cytarabine, dideoxyuridine, doxifluridine, enocitabine, floxuridine; androgens such as calsterone, propionic acid Dromostanolone, epitiostanol, mepitiostane, testolactone; antiadrenal drugs, e.g., mitotane, trilostane; folic acid supplements, e.g., folinic acid; aceglatone; aldophosphamide glycoside; aminolevulinic acid; eniluracil; amsacrine; bestravcil; bisantrene; edatraxate; defofamine; demecolcine; diaziquone; elformithine; elliptinium acetate; epothilone; etoglucide; gallium nitrate; hydroxyurea;Lentinan; lonidynin; maytansinoids, such as maytansine and ansamitocin; mitoguazone; mitoxantrone; mopidanmol; nitraerine; pentostatin; phenamet; pirarubicin; losoxantrone; podophyllinic acid; 2-ethylhydrazide; procarbazine; PSK polysaccharide complex; razoxane; rhizoxin; schizophyllum. Sizofiran; spirogermanium; tenuazonic acid; triaziquone; 2,2',2''-trichlorotriethylamine; trichothecenes (especially T-2 toxin, verracurin A, roridin A, and anguidine); urethane; vindesine; dacarbazine; mannomustine; mitobronitol; mitolactol; pipobroman; gacytosine; arabinoside ("Ara-C") ;cyclophosphamide;taxoids, such as paclitaxel and docetaxel gemcitabine;6-thioguanine;mercaptopurine;platinum coordination complexes, such as cisplatin, oxaliplatin, and carboplatin;vinblastine;platinum;etoposide (VP-16);ifosfamide;mitoxantrone;vincristine;vinorelbine;novantrone;teniposide;edatrexate;daunomycin;aminopterin;xeloda;ibandronate;irinotecan (e.g. , CPT-11); the topoisomerase inhibitor RFS2000; difluoromethylornithine (DMFO); retinoids, such as retinoic acid; capecitabine; carboplatin, procarbazine, plicomycin, gemcitabien, navelbine, farnesyl-protein transferase inhibitors, transplatinum, and pharmaceutically acceptable salts, acids, or derivatives of any of the foregoing.

[0071] Immunotherapeutics generally rely on the use of immune effector cells and molecules to target and destroy cancer cells. The immune effector may be, for example, an antibody specific for some marker on the surface of tumor cells. The antibody may serve alone as the therapeutic effector or may recruit other cells that actually bring about cell death. Antibodies may also be conjugated to drugs or toxins (e.g., chemotherapeutic agents, radionuclides, ricin A chain, cholera toxin, pertussis toxin, etc.) and simply serve as targeting agents. Alternatively, the effector may be a lymphocyte carrying a surface molecule that interacts directly or indirectly with the tumor cell target. Various effector cells include cytotoxic T cells and NK cells, as well as genetically engineered variants of these cell types engineered to express chimeric antigen receptors.

[0072] Immunotherapy may involve the suppression of T regulatory cells (Tregs), myeloid-derived suppressor cells (MDSCs), and cancer-associated fibroblasts (CAFs). In some embodiments, immunotherapy is tumor vaccines (e.g., whole tumor cell vaccines, peptides, and recombinant tumor-associated antigen vaccines) or adoptive cell therapy (ACT) (e.g., T cells, natural killer cells, TILs, and LAK cells). T cells may be engineered with chimeric antigen receptors (CARs) or T cell receptors (TCRs) directed against specific tumor antigens. As used herein, chimeric antigen receptors (or CARs) can refer to any engineered receptor specific to an antigen of interest that, when expressed in T cells, confers the specificity of the CAR to the T cells. Once created using standard molecular techniques, T cells expressing the chimeric antigen receptor can be introduced into patients using techniques such as adoptive cell transfer. In some aspects, the T cells are CD4 and / or CD8 T cells that produce γ-IFN and / or activated CD4 and / or CD8 T cells in an individual characterized by enhanced cytolytic activity compared to before administration of the combination. The CD4 and / or CD8 T cells may exhibit increased release of cytokines selected from the group consisting of IFN-γ, TNF-α, and interleukins. The CD4 and / or CD8 T cells may be effector memory T cells. In certain embodiments, the CD4 and / or CD8 effector memory T cells are CD44 high CD62L low The present invention is characterized by the expression of

[0073] Immunotherapy may also be a cancer vaccine containing one or more cancer antigens, particularly proteins or immunogenic fragments thereof, DNA or RNA encoding the cancer antigens, particularly proteins or immunogenic fragments thereof, cancer cell lysates, and / or protein preparations from tumor cells. As used herein, a cancer antigen is an antigenic substance present in cancer cells. In principle, any protein produced in cancer cells with an abnormal structure due to mutation can serve as a cancer antigen. In principle, cancer antigens can be the products of mutated oncogenes and tumor suppressor genes, the products of other mutated genes, overexpressed or aberrantly expressed cellular proteins, cancer antigens produced by tumor viruses, carcinoembryonic antigens, altered cell surface glycolipids and glycoproteins, or cell type-specific differentiation antigens. Examples of cancer antigens include abnormal products of the ras and p53 genes. Other examples include tissue differentiation antigens, mutated protein antigens, tumor virus antigens, cancer-testis antigens, and vascular- or stroma-specific antigens. Tissue differentiation antigens are tissue differentiation antigens specific to a particular type of tissue.

[0074] The immunotherapy may be an antibody, for example, part of a polyclonal antibody preparation, or a monoclonal antibody. The antibody may be a humanized antibody, a chimeric antibody, an antibody fragment, a bispecific antibody, or a single-chain antibody. The antibodies disclosed herein include antibody fragments, including, but not limited to, Fab, Fab' and F(ab')2, Fd, single-chain Fvs (scFv), single-chain antibodies, disulfide-linked Fv (sdfv), and fragments containing either the VL or VH domain. In some aspects, the antibody or fragment thereof specifically binds to epidermal growth factor receptor (EGFR1, Erb-B1), HER2 / neu (Erb-B2), CD20, vascular endothelial growth factor (VEGF), insulin-like growth factor receptor (IGF-1R), TRAIL-receptor, epithelial cell adhesion molecule, carcinoembryonic antigen, prostate-specific membrane antigen, mucin-1, CD30, CD33, or CD40.

[0075] Examples of monoclonal antibodies include trastuzumab (anti-HER2 / neu antibody); pertuzumab (anti-HER2 mAb); cetuximab (chimeric monoclonal antibody against epidermal growth factor receptor-EGFR); panitumumab (anti-EGFR antibody); nimotuzumab (anti-EGFR antibody); zalutumumab (anti-EGFR mAb); necitumumab (anti-EGFR mAb); MDX-210 (humanized anti-HER-2 bispecific antibody); MDX-210 (humanized anti-HER-2 bispecific antibody); MDX-447 (humanized anti-EGF receptor bispecific antibody); rituximab (chimeric murine / human anti-CD20 mAb); obinutuzumab (anti-CD20 mAb); ofatumumab (anti-CD20 mAb); tositumomab-I131 (anti-CD20 mAb); ibritumomab tiuxetan (anti-CD20 mAb); bevacizumab (anti-VEGF mAb); ramucirumab (anti-VEGFR2 mAb); ranibizumab (anti-VEGF mAb); aflibercept (extracellular domains of VEGFR1 and VEGFR2 fused to IgG1 Fc); AMG386 (angiopoietin-1 and -2 binding peptide fused to IgG1 Fc); dalotuzumab (anti-IGF-1R mAb); gemtuzumab ozogamicin (anti-CD33 mAb); alemtuzumab (anti-Campath-1 / CD52 mAb); brentuximab vedotin (anti-CD30 mAb); catumaxomab (bispecific mAb targeting epithelial cell adhesion molecule and CD3); naptumomab (anti-5T4 mAb); Girentuximab (anticarbonic anhydrase IX); or Farletuzumab (antifolate receptor agonist).Other examples include Panorex™ (17-1A) (mouse monoclonal antibody); Panorex (MAb17-1A) (chimeric mouse monoclonal antibody); BEC2 (combined with an anti-idiotypic mAb mimicking the GD epitope) (with BCG); Oncolym (Lym-1 monoclonal antibody); SMART M195Ab, humanized 13'1 LYM-1 (Oncolym), Ovarex (B43.13, an anti-idiotypic mouse mAb); 3622W94 mAb that binds to the EGP40 (17-1A) pancarcinoma antigen present on adenocarcinoma; Zenapax (SMART anti-Tac (IL-2 receptor); SMART M195 Ab, humanized Ab, humanized); NovoMAb-G2 (pancarcinoma-specific Ab); TNT (chimeric mAb against histone antigen); TNT (chimeric mAb against histone antigen); Gliomab-H (monoclonal humanized Ab); GNI-250 Mab; EMD-72000 (chimeric EGF antagonist); LymphoCide (humanized IL.L.2 antibody); and antibodies such as MDX-260 bispecific targeting GD-2, ANA Ab, SMART IDIO Ab, SMART ABL 364 Ab, or ImmuRAIT-CEA.Further examples of antibodies include zanulimumab (anti-CD4 mAb), keliximab (anti-CD4 mAb); ipilimumab (MDX-101; anti-CTLA-4 mAb); tremilimumab (anti-CTLA-4 mAb); (daclizumab (anti-CD25 / IL-2R mAb); basiliximab (anti-CD25 / IL-2R mAb); MDX-1106 (anti-PD1 mAb); antibodies against GITR; GC1008 (anti-TGF-β antibody); metelimumab / CAT-192 (anti-TGF-β antibody); lerdelimumab / CAT-152 (anti-TGF-β antibody); ID11 (anti-TGF-β antibody); denosumab (anti-RANKL mAb); BMS-663513 (humanized anti-4-lBB mAb); SGN-40 (humanized anti-CD40 mAb); CP870,893 (human anti-CD40 mAb); infliximab (chimeric anti-TNF mAb); adalimumab (human anti-TNF mAb); certolizumab (humanized Fab anti-TNF); golimumab (anti-TNF); etanercept (extracellular domain of TNFR fused to IgG1 Fc); belatacept (extracellular domain of CTLA-4 fused to Fc); abatacept (extracellular domain of CTLA-4 fused to Fc); belimumab (anti-B lymphocyte stimulator); muromonab-CD3 (anti-CD3 mAb); otelixizumab (anti-CD3 mAb); teplizumab (anti-CD3 mAb); tocilizumab (anti-IL6R mAb); REGN88 (anti-IL6R mAb); ustekinumab (anti-IL-12 / 23 mAb); briakinumab (anti-IL-12 / 23 mAb); natalizumab (anti-α4 integrin); vedolizumab (anti-α4β7 integrin mAb); T1h (anti-CD6 mAb); epratuzumab (anti-CD22 mAb); efalizumab (anti-CD11a mAb); and atacicept (extracellular domain of transmembrane activator-calcium-regulating ligand interactor fused to Fc).

[0076] When monitoring a mammal having or suspected of having cancer (e.g., based at least in part on the mammal's cfDNA fragmentation profile) as described herein, the monitoring can be before, during, and / or after a course of cancer treatment. The monitoring methods provided herein can be used to determine the efficacy of one or more cancer treatments and / or select a mammal for intensive monitoring. In some embodiments, the monitoring can include determining a cfDNA fragmentation profile as described herein. For example, a cfDNA fragmentation profile can be obtained before administering one or more cancer treatments to a mammal having or suspected of having cancer, the mammal can be administered one or more cancer treatments, and one or more cfDNA fragmentation profiles can be obtained during the course of the cancer treatment. In some embodiments, the cfDNA fragmentation profile can change during the course of a cancer treatment (e.g., any of the cancer treatments described herein). For example, a cfDNA fragmentation profile indicating that a mammal has cancer can change to a cfDNA fragmentation profile indicating that the mammal does not have cancer. Such a change in the cfDNA fragmentation profile indicates that the cancer treatment is effective. Conversely, the cfDNA fragmentation profile may remain unchanged (e.g., the same or nearly the same) during the course of a cancer treatment (e.g., any of the cancer treatments described herein). Such an unchanged cfDNA fragmentation profile indicates that the cancer treatment is not working.

[0077] In some embodiments, monitoring may involve conventional techniques capable of monitoring one or more cancer treatments (e.g., the efficacy of one or more cancer treatments). In some embodiments, mammals selected for enhanced monitoring may be subjected to diagnostic testing (e.g., any of the diagnostic tests disclosed herein) at an increased frequency compared to mammals not selected for enhanced monitoring. For example, mammals selected for enhanced monitoring may be subjected to diagnostic testing twice daily, once daily, twice weekly, once weekly, twice monthly, once monthly, quarterly, twice yearly, once yearly, or any frequency therein. In some embodiments, mammals selected for enhanced monitoring may be subjected to one or more additional diagnostic tests compared to mammals not selected for enhanced monitoring. For example, mammals selected for enhanced monitoring may be subjected to two diagnostic tests, while mammals not selected for enhanced monitoring are subjected to only one diagnostic test (or no diagnostic test). In some embodiments, mammals selected for enhanced monitoring may also be selected for additional diagnostic tests. Once the presence of a tumor or cancer (e.g., cancer cells) has been identified (e.g., by any of the various methods disclosed herein), it may be beneficial for the mammal to undergo both intensive monitoring (e.g., to assess the progression of the tumor or cancer in the mammal and / or to assess the development of one or more cancer biomarkers, such as mutations) and further diagnostic testing (e.g., to determine the size and / or precise location (e.g., tissue of origin) of the tumor or cancer). In some embodiments, the mammal selected for intensive monitoring after cancer biomarkers are detected and / or after the mammal's cfDNA fragmentation profile has not improved or worsened can be administered one or more cancer treatments. Any of the cancer treatments disclosed herein or known in the art can be administered. For example, the mammal selected for intensive monitoring can be further monitored, and if the presence of cancer cells persists over the intensive monitoring period, a cancer treatment can be administered.Additionally or alternatively, the mammal that has been selected for enhanced monitoring can be subjected to cancer treatment, and further monitored as the cancer treatment progresses.In some embodiments, after the mammal that has been selected for enhanced monitoring is subjected to cancer treatment, the enhanced monitoring can reveal one or more cancer biomarkers (e.g., mutations).In some embodiments, such one or more cancer biomarkers can provide the reason for administering another cancer treatment (e.g., resistance mutations may occur in cancer cells during cancer treatment, and the cancer cells that contain this resistance mutation are resistant to the initial cancer treatment).

[0078] As described herein, when a mammal is identified as having cancer (for example, based at least in part on the cfDNA fragmentation profile of the mammal), this identification can be before and / or during a course of cancer treatment.The method provided herein for identifying a mammal as having cancer can be used as an initial diagnosis to identify the mammal (for example, as having cancer before any course of treatment) and / or select the mammal for further diagnostic testing.In some embodiments, once a mammal is determined to have cancer, the mammal can be subjected to further testing and / or selected for further diagnostic testing.In some embodiments, the method provided herein can be used to select the mammal for further diagnostic testing at a time before conventional techniques can diagnose the mammal as having early-stage cancer.For example, the method provided herein for selecting a mammal for further diagnostic testing can be used when the mammal has not been diagnosed as having cancer by conventional methods and / or when the mammal is not known to have cancer. In some embodiments, mammals selected for further diagnostic testing may be subjected to the diagnostic test (e.g., any of the diagnostic tests disclosed herein) at an increased frequency compared to mammals not selected for further diagnostic testing. For example, mammals selected for further diagnostic testing may be subjected to the diagnostic test twice daily, once daily, twice weekly, once weekly, twice monthly, once monthly, quarterly, twice yearly, once yearly, or any frequency therein. In some embodiments, mammals selected for further diagnostic testing may be subjected to one or more further diagnostic tests compared to mammals not selected for further diagnostic testing. For example, mammals selected for further diagnostic testing may be subjected to two diagnostic tests, while mammals not selected for further diagnostic testing are subjected to only one diagnostic test (or no diagnostic test).In some embodiments, the diagnostic test can determine the presence of a cancer of the same type (e.g., having the same tissue of origin) as the initially detected cancer (e.g., based at least in part on the mammal's cfDNA fragmentation profile). Additionally or alternatively, the diagnostic test can determine the presence of a cancer of a different type than the initially detected cancer. In some embodiments, the diagnostic test is a scan. In some embodiments, the scan is a computed tomography (CT), CT angiography (CTA), esophagography (barium swallow), barium enema, magnetic resonance imaging (MRI), PET scan, ultrasound (e.g., endobronchial ultrasound, endoscopic ultrasound), radiography, or DEXA scan.

[0079] In some embodiments, the diagnostic test is a physical examination, such as anoscopy, bronchoscopy (e.g., autofluorescence bronchoscopy, white light bronchoscopy, guided bronchoscopy), colonoscopy, breast digital tomosynthesis, endoscopic retrograde cholangiopancreatography (ERCP), esophagogastroduodenoscopy, mammography, Papanicolaou smear test, pelvic examination, or positron emission tomography-computed tomography (PET-CT) scan. In some embodiments, a mammal that has been selected for further diagnostic testing can also be selected for enhanced monitoring. Once the presence of a tumor or cancer (e.g., cancer cells) has been identified (e.g., by any of the various methods disclosed herein), it may be beneficial for the mammal to undergo both enhanced monitoring (e.g., to assess the progression of the tumor or cancer in the mammal and / or to assess the development of one or more cancer biomarkers, e.g., mutations) and further diagnostic testing (e.g., to determine the size and / or precise location of the tumor or cancer). In some embodiments, a mammal selected for further diagnostic testing after a cancer biomarker is detected and / or after the mammal's cfDNA fragmentation profile has not improved or worsened is administered a cancer treatment. Any cancer treatment disclosed herein or known in the art can be administered. For example, a mammal selected for further diagnostic testing can be subjected to further diagnostic testing, and if the presence of a tumor or cancer is confirmed, a cancer treatment can be administered. Additionally or alternatively, a mammal selected for further diagnostic testing can be administered a cancer treatment and further monitored as the cancer treatment progresses. In some embodiments, after a mammal selected for further diagnostic testing is administered a cancer treatment, additional testing can reveal one or more cancer biomarkers. In some embodiments, such one or more cancer biomarkers (e.g., mutations) can provide a reason for administering another cancer treatment (e.g., resistance mutations can occur in cancer cells during cancer treatment, and cancer cells containing the resistance mutations are resistant to the initial cancer treatment).

[0080] system In some examples, the present disclosure provides systems, methods, or kits that may include software code executing on data analysis and computational hardware implemented in a measurement device (e.g., laboratory equipment, e.g., a sequencing device). The software can be stored in memory and executed on one or more hardware processors. The software can be organized into routines or packages that can communicate with each other. Modules may comprise one or more software routines / packages executing on one or more devices / computers, and potentially on one or more devices / computers. For example, an analysis application or system may comprise at least a data reception module, a data preprocessing module, a data analysis module (which can operate on one or more types of genomic data), a data interpretation module, or a data visualization module.

[0081] The data receiving module can connect laboratory hardware or instrumentation with a computer system that processes laboratory data. The data preprocessing module can perform operations on data ready for analysis. Examples of operations that can be applied to data in the preprocessing module include affine transformations, denoising operations, data cleaning, reformatting, or subsampling. The data analysis module can specialize in analyzing genomic data from one or more genomic materials, for example, by taking assembled genomic sequences and performing probabilistic and statistical analyses to identify abnormal patterns associated with a disease, pathology, state, risk, condition, or phenotype. The data interpretation module can use analytical methods, e.g., derived from statistics, mathematics, or biology, to support an understanding of the relationship between identified abnormal patterns and health status, functional status, prognosis, or risk. The data analysis module and / or data interpretation module can comprise one or more machine learning models that can be executed in hardware, e.g., hardware executing software embodying the machine learning models. The data visualization module may use methods of mathematical modeling, computer graphics, or rendering to create a visual display of the data that can facilitate understanding or interpretation of the results. The present disclosure provides a computer system programmed to perform the methods of the present disclosure.

[0082] In some embodiments, the methods disclosed herein may include computational analysis of nucleic acid sequencing data from samples from one or more individuals. The analysis may identify variants inferred from sequence data, such as identifying sequence variants based on probability modeling, statistical modeling, mechanistic modeling, network modeling, or statistical inference. Non-limiting examples of analytical methods include principal component analysis, autoencoders, singular value decomposition, Fourier bases, wavelets, discriminant analysis, regression, support vector machines, tree-based methods, networks, matrix factorization, and clustering. Non-limiting examples of variants include germline variations or somatic mutations. In some cases, a variant may refer to a known variant. A known variant may be scientifically confirmed or reported in the literature. In some cases, a variant may refer to a putative variant associated with a biological change. The biological change may be known or unknown. In some cases, a putative variant may be reported in the literature but has not yet been biologically confirmed. Alternatively, a putative variant may never have been reported in the literature but may be inferred based on the computational analysis disclosed herein. In some instances, a germline variant may refer to a nucleic acid that induces natural or normal variation.

[0083] In certain embodiments, the computer system comprises a central processing unit (CPU, also referred to herein as "processor" and "computer processor"), which may be a single-core processor or a multi-core processor, or multiple processors for parallel processing; memory (e.g., cache, random access memory, read-only memory, flash memory, or other memory); electronic storage (e.g., hard disk), a communication interface (e.g., network adapter) for communicating with one or more other systems; and peripheral devices, e.g., adapters for cache, other memory, data storage, and / or electronic displays. The memory, storage, interfaces, and peripheral devices may be in communication with the CPU through a communication bus (solid line), e.g., a motherboard. The storage device may be a data storage device (or data repository) for storing data. One or more analytical feature inputs may be input from one or more measurement devices. Examples of analytes and measurement devices are described herein.

[0084] The computer system may be operably connected to a computer network ("network") via a communications interface. The network may be the Internet, an Internet and / or extranet, or an intranet and / or extranet in communication with the Internet. In some cases, the network is a telecommunications and / or data network. The network may include one or more computer servers that can enable distributed computing, e.g., cloud computing over a network ("cloud"), to perform various aspects of the analysis, calculation, and production of the present disclosure, such as activating valves or pumps to transfer reagents or samples from one chamber to another, or heating samples (e.g., during amplification reactions), processing samples and / or other aspects of assays, performing sequencing analysis, measuring sets of values ​​representative of molecular classes, identifying sets of features and feature vectors from assay data, processing the feature vectors using machine learning models to obtain output classifications, and training machine learning models (e.g., iteratively searching for optimal values ​​for machine learning model parameters). Such cloud computing may be provided by cloud computing platforms such as Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform, and IBM Cloud. The network may, in some cases with the help of a computer system, implement a peer-to-peer network, which may allow devices connected to the computer system to act as clients or servers.

[0085] The CPU may execute a sequence of machine-readable instructions, which may be embodied in the form of a program or software. The instructions may be stored in a memory location, such as a memory. The instructions may be directed to the CPU, which may then be programmed or otherwise configured to perform the methods of the present disclosure. The CPU may be part of a circuit, such as an integrated circuit. One or more other components of the system may be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).

[0086] The storage device can store files, such as drivers, libraries, and saved programs. The storage device can store user data, such as user preferences and user programs. The computer system may optionally include one or more additional data storage devices external to the computer system, for example, one or more additional data storage devices located on remote servers in communication with the computer system over an intranet or the Internet.

[0087] The computer system can communicate with one or more remote computer systems over a network. For example, the computer system can communicate with a user's remote computer system. Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad, a Samsung® Galaxy Tab), a telephone, a smartphone (e.g., an Apple® iPhone, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user can access the computer system over a network.

[0088] The methods described herein can be implemented as machine-executable code (e.g., a computer processor) stored in an electronic storage location, e.g., memory or electronic storage, of a computer system. The machine-executable or machine-readable code can be provided in the form of software. During use, the code can be executed by a CPU. In some cases, the code can be retrieved from storage and stored in memory for rapid access by the CPU. In some situations, electronic storage can be omitted, and machine-executable instructions are stored in memory.

[0089] The code may be pre-compiled and configured for use with a machine having a processor adapted to execute the code, or may be compiled at run-time. The code may be supplied in a programming language that can be selected to allow the code to be executed in a pre-compiled or as-compiled manner.

[0090] Aspects of the systems and methods provided herein, e.g., computer systems, can be embodied in programming. Various aspects of this technology can be considered "products" or "articles of manufacture," typically in the form of machine- (or processor-) executable code and / or associated data carried on or embodied in some type of machine-readable medium. The machine-executable code can be stored in electronic storage, e.g., memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. "Storage" type media can include any or all of a computer's tangible memory, processor, etc., or its associated modules, e.g., various semiconductor memories, tape drives, disk drives, etc., that can provide non-transitory storage at any time for software programming. All or portions of the software may sometimes be communicated over the Internet or various other telecommunications networks. Such communication may, for example, enable loading of software from one computer or processor to another, e.g., from a management server or host computer to an application server computer platform. Thus, other types of media that may have software elements include light waves, radio waves, and electromagnetic waves, e.g., those used across physical interfaces between local devices, through wired and optical landline networks, and across various air-links. The physical elements that carry such waves, e.g., wired or wireless links, optical links, etc., may also be considered media that have software. Unless limited to non-transitory tangible "storage" media, the term computer or machine "readable medium," as used herein, refers to any medium that participates in providing instructions to a processor for execution.

[0091] Thus, machine-readable media, such as computer-executable code, may take many forms, including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media include, for example, optical or magnetic disks, any storage device present in any computer, and can be used to execute databases, etc., as shown in the drawings. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system.

[0092] Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media thus include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punched cards, paper tape, any other physical storage media with patterns of holes, RAMs, ROMs, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carrier waves carrying data or instructions, cables or links carrying such carrier waves, or any other medium from which a computer can read programming code and / or data. Many of these types of computer-readable media can be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0093] The computer system may include, or be in communication with, an electronic display that includes a user interface (UI), for example, to indicate the current stage of sample processing or assay (e.g., a particular step, e.g., lysis step, or sequencing step being performed). Input is received by the computer system from one or more measurements. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces. The algorithm may, for example, process and / or assay the sample, perform a sequencing analysis, measure a set of values ​​representative of molecular classes, identify a set of features and feature vectors from the assay data, process the feature vectors using a machine learning model to obtain an output classification, and train the machine learning model (e.g., iteratively search for optimal values ​​for machine learning model parameters).

[0094] In some embodiments, a system (e.g., laptop, desktop, iPad, mobile device, etc.) capable of executing one or more algorithms to determine changes in cfDNA fragmentation profiles classifies a subject as a cancer patient based on the subject's cfDNA fragmentation profile. Additionally, these systems utilize machine learning algorithms (e.g., penalized logistic regression) that can be used to create models of high-risk and low-risk general populations, as described in Mathios et al. (Mathios D, Johansen JS, Cristiano S, Medina JE, Phallen J, Larsen KR, et al. Detection and characterization of lung cancer using cell-free DNA fragmentomes. Nat Commun. The models are run with features and coverage from transcription factor binding sites in the NIH (2021;12(1):5060). These models can be trained on the target cohort using 10 iterations of 5-fold cross-validation. Scores for each sample are calculated by averaging the multiple iterations and evaluated using AUC-ROC. For example, the first model used high-risk non-cancer and HCC patients, while the second model used non-cancer individuals without liver pathology. The locked high-risk model trained on the cohort was applied to a second, different cohort to perform cancer predictions on an external validation set. A "class label" indicating sample classification across any number of input features can be applied to each sample. For example, the class label for a cohort set may indicate the identity of the cfDNA fragmentation profile based on genomic location, etc. The resulting training set is then provided to a machine learning unit, such as a neural network or support vector machine. Using the training set, the machine learning unit can create a model to classify samples according to their cfDNA fragmentation profile.

[0095] In some embodiments, methods are provided for creating a trained classifier, the methods including: (a) preparing a plurality of distinct classes, each class representing a set of subjects (e.g., from one or more cohorts) that share a common characteristic; (b) preparing a multi-metric model that represents cell-free DNA molecules from each of a plurality of samples belonging to each class, thereby providing a training dataset; and (c) training a learning algorithm on the training dataset to create one or more trained classifiers, wherein each trained classifier classifies test samples into one or more of the plurality of classes.

[0096] As an example, the trained classifier may use a learning algorithm selected from the group consisting of random forests, neural networks, support vector machines, and linear classifiers. Each of the multiple different classes may be selected from the group consisting of healthy, breast cancer, colon cancer, lung cancer, pancreatic cancer, prostate cancer, ovarian cancer, melanoma, and liver cancer.

[0097] The trained classifier can be applied to a method for classifying a sample from a subject. The classification method can include: (a) preparing a multimetric model representing cell-free DNA molecules derived from a test sample from the subject; and (b) classifying the test sample using the trained classifier. After the test sample is classified into one or more classes, a therapeutic intervention can be administered to the subject based on the classification of the sample.

[0098] In some embodiments, the training set is provided to a machine learning unit, such as a neural network or support vector machine. Using the training set, the machine learning unit can create a model to classify samples according to treatment response to one or more therapeutic inventions. This is also called "calling." The developed model can use information from any portion of the test vector.

[0099] Generally, machine learning can be used to reduce a dataset created from all (primary sample / analyte / test) combinations to an optimal predictive set of features, for example, that meets specified criteria. In various examples, statistical learning and / or regression analysis can be applied. Simple to complex and small to large models making various modeling assumptions can be applied to the data in a cross-validation paradigm. Simple to complex includes considerations of linear to nonlinear and non-hierarchical to hierarchical representations of features. Small to large models include considerations of the size of the basis vector space for projecting the data onto the number of interactions between features included in the modeling process.

[0100] Machine learning methods can be used to evaluate commercial testing modalities that are optimal for cost / performance / commercial reach as defined in the initial questions. Threshold checks can be performed; that is, if the method applied to a holdout dataset not used in cross-validation is superior to the initialized constraints, the assay is locked and production begins. For example, assay performance thresholds may include a desired minimum accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), or a combination thereof. For example, the desired minimum accuracy, PPV, NPV, clinical sensitivity, clinical specificity, or a combination thereof may be at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. As another example, a desirable minimum AUC may be at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.81, at least about 0.82, at least about 0.83, at least about 0.84, at least about 0.85, at least about 0.86, at least about 0.87, at least about 0.88, at least about 0.89, at least about 0.90, at least about 0.91, at least about 0.92, at least about 0.93, at least about 0.94, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.A subset of assays may be selected from the set of assays to be performed on a particular sample based on the total cost of performing the subset of assays, subject to assay performance thresholds, such as desired minimum accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), and combinations thereof. If the thresholds are not met, the assay engineering procedure can return to the constraint settings for possible relaxation or to the wet lab to change the parameters under which the data were obtained. Given a clinical question, biological constraints, budget, lab machinery, etc. may constrain this problem.

[0101] In certain embodiments, the computational processing of the machine learning method may include methods of statistics, mathematics, biology, or any combination thereof. In various examples, any one of the computational methods may include dimensionality reduction methods, logistic regression, dimensionality reduction, principal component analysis, autoencoders, singular value decomposition, Fourier basis, singular value decomposition, wavelets, discriminant analysis, support vector machines, tree-based methods, random forests, gradient boost trees, logistic regression, matrix factorization, network clustering, statistical testing, and neural networks.

[0102] In certain embodiments, the computerized machine learning methods may include logistic regression, multiple linear regression (MLR), dimensionality reduction, partial least squares (PLS) regression, principal component regression, autoencoder, variational autoencoder, singular value decomposition, Fourier basis, wavelets, discriminant analysis, support vector machines, decision trees, classification and regression trees (CART), tree-based methods, random forests, gradient boosted trees, logistic regression, matrix factorization, multidimensional scaling (MDS), dimensionality reduction methods, t-distributed stochastic neighbor embedding (t-SNE), multilayer perceptron (MLP), network clustering, neuro-fuzzy, neural networks (shallow and deep), artificial neural networks, Pearson's product-moment correlation coefficient, Spearman's rank correlation coefficient, Kendall tau rank correlation coefficient, or any combination thereof. In some examples, the computational methods are supervised machine learning methods, including, for example, regression, support vector machines, tree-based methods, and neural networks. In some examples, the computational methods are unsupervised machine learning methods, including, for example, clustering, networks, principal component analysis, and matrix factorization.

[0103] In supervised learning, training samples (e.g., in the thousands) may include measurement data (e.g., of various analytes) and known labels that can be determined by other time-consuming processes, such as imaging the subject and analysis by a trained practitioner. Examples of labels may include subject classifications, such as a discrete classification of whether the subject has or does not have cancer, or a continuous classification indicating a discrete value of probability (e.g., risk or score). The learning module can optimize the model parameters so that a quality metric (e.g., the accuracy of prediction for a known label) is achieved by one or more specified criteria. Quality metric determinations can be performed for any arbitrary function, including the full set of risk, loss, utility, and decision functions. Gradients can be used in conjunction with the learning step (e.g., a measure of how much the model parameters should be updated for a given time step in the optimization process).

[0104] As noted above, examples can be used for a variety of purposes. For example, plasma (or other samples) can be collected from subjects symptomatic of a condition (e.g., known to have the condition) and healthy individuals. Genetic data (e.g., cfDNA) can be obtained and analyzed to derive a variety of different features. The features may include features based on genome-wide analysis. These features can form a feature space that is searched, stretched, rotated, translated, and linearly or non-linearly transformed to create an accurate machine learning model that can distinguish healthy individuals from subjects with the condition (e.g., identify a diseased or non-disease state of the subject). This data and output derived from the model (which may include a probability of the condition, stage (level) of the condition, or other value) can be used to create another model that can be used to recommend further procedures, such as recommending a biopsy, or to continue monitoring the subject's condition.

[0105] In some embodiments, DNA from a population of several individuals can be analyzed using a set of multiplexed arrays. The data from each multiplexed array can be self-normalized using the information contained in that particular array. This normalization algorithm can be adjusted for the nominal intensity variations observed in the two color channels, background differences between channels, and possible crosstalk between dyes. The behavior of each base position can then be modeled using a clustering algorithm that incorporates several biological heuristics for the fragmentation profile. When few cfDNA fragments are observed (e.g., due to low minor allele frequency), the location and shape of missing sequences can be estimated using a neural network. Depending on the profile and percent sequence identity information, a statistical score can be devised (training score). Scores such as the GenCall score are designed to mimic the evaluation made by the visual and cognitive systems of human experts. Furthermore, it has been evolved using genotyping data from the top and bottom strands. This score can be combined with several penalty terms (e.g., low intensity, mismatch between existing and predicted cfDNA fragments) to form a training score. The training scores are saved for use by the calling algorithm.

[0106] To call a treatment response, a calling algorithm can obtain genetic information and treatment responses for multiple individuals with a disease or condition. First, the data may be normalized (using the same procedure as a clustering algorithm). The calling operation (classification) may be performed, for example, using a Bayesian model. The score for each call, the Call Score, may be the product of the training score and the data-to-model fit score. After scoring all treatment responses, the application can calculate a composite score.

[0107] In some embodiments, the training dataset comprises clinical data selected from the group consisting of cancer stage, type of surgical procedure, age, tumor grading, depth of tumor invasion, occurrence of postoperative complications, and presence of venous invasion. In some embodiments, the training dataset is preprocessed, including converting the provided data into class-conditional probabilities.

[0108] In another embodiment, machine learning methods are used to train statistical classifiers, specifically support vector machines, for each cancer stage category based on word occurrence in the corpus of each patient's histology reports. New reports can then be classified according to their most likely stage, facilitating the collection and analysis of population staging data.

[0109] In some embodiments, the machine learning algorithm is selected from the group consisting of supervised or unsupervised learning algorithms selected from support vector machines, random forests, nearest neighbor analysis, linear regression, binary decision trees, discriminant analysis, logistic classifiers, and cluster analysis.

[0110] Generally, the system may include a report generator for reporting cancer test results and treatment options. The report generator system may be a central data processing system configured to establish communication with a remote data site or laboratory, a medical clinic / healthcare provider (treatment specialist), and / or a patient / subject directly via a communication link. The laboratory may be a medical laboratory, a diagnostic laboratory, a medical facility, a medical clinic, a point-of-care testing device, or any other remote data site capable of generating subject clinical information. The subject clinical information includes, but is not limited to, laboratory test data, X-ray data, examinations, and diagnoses. The medical provider or clinic 26 includes health care service providers, such as doctors, nurses, home health aides, technicians, and physician assistants, and the clinic is any medical facility staffed by a health care provider. In certain cases, the medical provider / clinic is also a remote data site. In cancer treatment embodiments, the subject may specifically be suffering from cancer.

[0111] Other clinical information for cancer subjects includes the results of laboratory tests, imaging, or medical treatments directed at the specific cancer, which can be easily identified by those skilled in the art. A list of appropriate sources of clinical information for cancer includes, but is not limited to, CT scan, MRI scan, ultrasound scan, bone scan, PET scan, bone marrow examination, barium X-ray, endoscopy, lymphangiogram, IVU (intravenous urogram) or IVP (IV pyelogram), lumbar puncture, cystoscopy, immunological tests (antimalignin antibody screen), and cancer marker tests.

[0112] The subject's clinical information may be obtained manually or automatically from the laboratory. To simplify the system, the information may be obtained automatically at predetermined or fixed time intervals. A fixed time interval refers to a time interval in which the method and system described herein automatically collects laboratory data based on a time measurement such as hours, days, weeks, months, or years. In one embodiment of the present invention, data collection and processing are performed at least once a day. In one embodiment, data transfer and collection are performed monthly, every two weeks, or once a week, or once every few days. Alternatively, information retrieval may be performed at predetermined but non-fixed time intervals. For example, the first retrieval step may be performed one week later, and the second retrieval step may be performed one month later. Data transfer and collection can be customized depending on the nature of the disorder being managed and the frequency of the subject's required tests and medical examinations.

[0113] In certain embodiments, genetic report is generated from target sample (for example, cfDNA).The polynucleotide in sample can be sequenced, for example, whole genome sequencing, NGS sequencing, to generate multiple sequence readings.In some embodiments, genetic information comprises the variable that defines the genome composition of multiple cancer cells or the genome composition of one scattered cancer cell.In some embodiments, genetic information comprises the sequence or abundance data from one or more loci in cell-free DNA from individual.

[0114] The cfDNA genetic information is processed (72). Genetic variants can also be identified. Genetic variants include sequence variants, copy number variants, and nucleotide modification variants. Sequence variants are changes in the nucleotide sequence of a gene. Copy number variants are variations in the copy number of a portion of the genome that deviate from the wild type. Genetic variants include, for example, single nucleotide variations (SNPs), insertions, deletions, inversions, transversions, translocations, gene fusions, chromosome fusions, gene truncations, copy number variations (e.g., aneuploidy, partial aneuploidy, polyploidy, gene amplification), abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid methylation. This process then determines the frequency of genetic variants in samples containing genetic material. Because this process is noisy, it separates information from noise (73). The sensitivity of detecting genetic variants can be increased by increasing the polynucleotide read depth (e.g., by sequencing samples from a subject at two or more time points to a greater read depth).

[0115] Multiple measurements can be performed to increase diagnostic accuracy. Alternatively, measurements at multiple time points (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more) can be used to determine whether the cancer is progressing, in remission, or stable. Diagnostic accuracy can be used to identify disease states. For example, cell-free polynucleotides collected from a subject can include polynucleotides derived from normal cells and polynucleotides derived from diseased cells, such as cancer cells. Polynucleotides derived from cancer cells can have genetic variants, such as somatic mutations and copy number variants. When sequencing cell-free polynucleotides derived from a sample from a subject, a cfDNA fragmentation profile can be generated as described in the Examples section below.

[0116] The methods and systems described herein can be used to detect a large number of cancers.Cancer cells, like most cells, can be characterized by a turnover rate, where old cells die and are replaced by new cells.In general, dead cells are in contact with the vasculature in a particular subject and may release DNA or DNA fragments into the bloodstream.This also applies to cancer cells in various stages of disease.Cancer cells can also be characterized by various genetic abnormalities, such as copy number variations and mutations, depending on the stage of disease.This phenomenon can be used to detect the presence or absence of cancer individuals using the methods and systems described herein.

[0117] In the early detection of cancer, any system or method described herein, including mutation detection or copy number variation detection, can be used to detect cancer.These systems and methods can be used to detect any number of genetic abnormalities that can cause cancer or can be attributed to cancer.These can include, but are not limited to, cfDNA fragmentation profile, mutation, mutation, indel, copy number variation, transversion, translocation, inversion, deletion, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structure change, gene fusion, chromosomal fusion, gene truncation, gene amplification, gene duplication, chromosomal damage, DNA damage, abnormal changes in nucleic acid chemical modification, abnormal changes in epigenetic pattern, abnormal changes in nucleic acid methylation, infectious disease and cancer.

[0118] Furthermore, the systems and methods described herein may also be used to help characterize a particular cancer. Genetic data generated from the systems and methods of the present disclosure may enable practitioners to better characterize a particular type of cancer. Cancers are often heterogeneous in both composition and staging. Genetic profile data may be used to characterize specific subtypes of cancer, which may be important in the diagnosis or treatment of certain subtypes. This information may also provide clues to a subject or practitioner regarding the prognosis of a particular type of cancer.

[0119] The systems and methods provided herein may also be used to monitor known cancer or other diseases in a particular subject. This may allow the subject or practitioner to adapt treatment options according to the progression of the disease. In this example, the systems and methods described herein may be used to build a genetic cfDNA fragmentation profile of a particular subject during the course of the disease. In some cases, cancer may progress, become aggressive, and become genetically unstable. In other cases, cancer may remain benign, inactive, or dormant. The systems and methods of the present disclosure may be useful in determining disease progression.

[0120] Furthermore, the system and method described herein can be useful in determining the efficacy of certain treatment options.In one example, certain treatment options can be correlated with cancer gene cfDNA fragmentation profile over time.This correlation can be useful in selecting therapy.In addition, if cancer is observed to be in remission after treatment, the system and method described herein can be useful in monitoring residual disease or disease recurrence.

[0121] Furthermore, the disclosed method may be used to characterize the heterogeneity of an abnormal condition in a subject, the method comprising generating a cfDNA fragmentation profile of extracellular polynucleotides in the subject, the cfDNA fragmentation profile comprising multiple data resulting from profile variation and mutation analysis. In some cases, including but not limited to cancer, diseases may be heterogeneous. Disease cells may not be identical. In the example of cancer, it is known that some tumors contain different types of tumor cells, and some cells are at different stages of cancer. In other cases, heterogeneity may comprise multiple disease foci. Again, in the example of cancer, there may be multiple tumor foci, and perhaps one or more foci are the result of metastasis that has spread from the primary site (also known as distant metastasis).

[0122] The methods of the present disclosure may be used to generate a profile, fingerprint, or set of data that is the sum of genetic information from different cells within a heterogeneous disease. This data set may include copy number variation and mutation analysis, alone or in combination.

[0123] Additionally, these reports are submitted and accessed electronically via the internet. Analysis of the data is performed at a location other than the subject's location. Reports are generated and transmitted to the subject's location. Subjects access reports reflecting their tumor burden via an internet-enabled computer.

[0124] The annotated information can be used by healthcare providers to select other drug treatment options and / or to provide information about drug treatment options to insurance companies. The method includes annotating drug treatment options for the condition, for example, in the NCCN Clinical Practice Guidelines in Oncology™ or the American Society of Clinical Oncology (ASCO) clinical practice guidelines.

[0125] Reports are generated that map the genomic location and cfDNA fragmentation profile variations of subjects with cancer. These reports, when compared to other profiles of subjects with known outcomes, may indicate that a particular cancer is aggressive and resistant to treatment. The subject is monitored and retested over a period of time. If the cfDNA fragmentation variation profile remains unchanged at the end of the period, this may indicate that the current treatment is not working. Comparisons are made with the cfDNA fragmentation profiles of other subjects. For example, if changes in cfDNA fragmentation variation are determined to indicate that the cancer is progressing, the initially prescribed treatment regimen is no longer treating the cancer, and a new treatment is prescribed.

[0126] In certain embodiments, the system receives genetic information from a DNA sequencer.This process then determines specific cfDNA fragmentation changes and their amounts.These reports are submitted and accessed electronically via the Internet.Data analysis is performed at a location other than the subject's location.The report is generated and transmitted to the subject's location.The subject accesses the report that reflects the subject's tumor burden via an Internet-enabled computer.

[0127] Although temporal information can be used to enhance the information of the cfDNA fragmentation profile, other consensus methods can be applied. In other embodiments, historical comparison can be used together with other consensus cfDNA fragmentation profiles. The consensus cfDNA fragmentation profile can be normalized against control samples. Molecular mapping measurements against reference sequences can also be compared across the genome to identify regions in the genome where the cfDNA fragmentation profile has changed or remains the same. Consensus methods include, for example, linear or nonlinear methods for generating consensus cfDNA fragmentation profiles derived from digital communication theory, information theory, or bioinformatics (e.g., voting, averaging, statistical, maximum a posteriori, or maximum likelihood detection, dynamic programming, Bayesian, hidden Markov, or support vector machine methods, etc.). After the sequence read coverage is determined, a probabilistic modeling algorithm is applied to convert the normalized nucleic acid sequence read coverage of each window region into discrete copy number states. In some cases, the algorithm may include one or more of the following: hidden Markov models, dynamic programming, support vector machines, Bayesian networks, trellis decoding, Viterbi decoding, expectation maximization, Kalman filtering, and neural networks.

[0128] Artificial neural networks (NNets) mimic networks of "neurons" based on the neural structure of the brain. They process records one at a time, or in batches, and "learn" the classification of the (initially almost arbitrary) record by comparing it with the known actual classification of the record. In MLP-NNets, the error from the initial classification of the first record is fed back into the network and used to improve the network's algorithm over multiple iterations, such as the second time. Neural networks use an iterative learning process in which data examples (rows) are presented to the network one at a time, and the weights associated with the input values ​​are adjusted each time.

[0129] After all examples have been presented, the process often begins again. During this training phase, the network learns by adjusting its weights so that it can predict the correct class label of input samples. Neural network learning is also called "connectionist learning" due to the connections between units. Advantages of neural networks include their high tolerance for noisy data and their ability to classify patterns for which they have never been trained. One neural network algorithm is the back-propagation algorithm, e.g., Levenberg-Marquadt. Once a network has been constructed for a particular application, it is ready to be trained. To begin this process, initial weights are randomly selected. Training, or learning, then begins.

[0130] The network uses the weights and functions in the hidden layer to process records in the training data one at a time, then compares the resulting output against the desired output. Errors are then backpropagated through the system, which causes the system to adjust the weights to apply to the next record to be processed. This process occurs many times, with the weights being successively fine-tuned. During network training, the same data set is processed many times, with the connection weights being continually refined.

[0131] In one embodiment, training the machine learning unit on the training dataset can generate one or more classification models for application to test samples, which can be applied to the test samples to predict a subject's response to a therapeutic intervention.

[0132] Comparison of sequence coverage with control samples or reference sequences can aid in window-wide normalization.In this embodiment, cell-free DNA is extracted and isolated from easily accessible body fluids such as blood.For example, cell-free DNA can be extracted using various methods known in the art, including but not limited to isopropanol precipitation and / or silica-based purification.Cell-free DNA can be extracted from any number of subjects, for example, cancer-free subjects, cancer risk subjects, or subjects known to have cancer (e.g., by other means).

[0133] After the isolation / extraction step, any of a number of different sequencing operations can be performed on the cell-free polynucleotide sample. The sample may be treated with one or more reagents (e.g., enzymes, unique identifiers (e.g., barcodes), probes, etc.) before sequencing. In some cases, once the sample has been treated with a unique identifier such as a barcode, the sample or fragments of the sample may be tagged individually or in subgroups with a unique identifier. The tagged sample may then be used in downstream applications, such as sequencing reactions, in which individual molecules can be traced back to their parent molecules.

[0134] Cell-free polynucleotides can be tagged or tracked to enable subsequent identification and origin of specific polynucleotides. Assigning identifiers (e.g., barcodes) to individual polynucleotides or subgroups of polynucleotides may allow for the assignment of unique identifying information to individual sequences or fragments of sequences. This may allow for data acquisition from individual samples, not limited to sample averages. In some cases, nucleic acids or other molecules derived from a single strand may share a common tag or identifier and thus may later be traced back to that strand. Similarly, all fragments derived from a single strand of nucleic acid may be tagged with the same identifier or tag, thereby later identifying the fragments derived from the parent strand. In other cases, gene expression products (e.g., mRNA) may be tagged to quantify expression, allowing for counting of barcodes or barcodes combined with sequences attached to the barcodes. In still other cases, the systems and methods can be used as PCR amplification controls. In such cases, multiple amplification products from a PCR reaction can be tagged with the same tag or identifier. If the products are later sequenced and show sequence differences, differences between products with the same identifier can be attributed to PCR errors. Additionally, individual sequences may be identified based on characteristics of the sequence data of the read itself. For example, the detection of unique sequence data at the beginning (start) and end (stop) of each sequencing read may be used alone or in combination with the length, i.e., number of base pairs, of each sequence read's unique sequence to assign unique identities to individual molecules. Fragments derived from a single strand of nucleic acid that have been assigned unique identities may then be used to identify fragments from the parent strand. This can be used in conjunction with bottlenecking the initial starting genetic material to limit diversity.

[0135] Generally, the methods and systems provided herein are useful for preparing cell-free polynucleotide sequences for downstream sequencing reactions.In most cases, sequencing methods include next-generation sequencing (NGS), classical Sanger sequencing, whole genome bisulfite sequencing (WGSB), small RNA sequencing, low-coverage whole genome sequencing (IcWGS), etc.

[0136] As used herein, the term "sequencing" refers to any of a number of techniques used to determine the sequence of a biomolecule, for example, a nucleic acid such as DNA or RNA. Exemplary sequencing methods include targeted sequencing, single-molecule real-time sequencing, exon sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxytermination sequencing, whole genome sequencing, sequencing by hybridization, pyrosequencing, capillary electrophoresis, gel electrophoresis, duplex sequencing, cycle sequencing, single-base extension sequencing, solid-phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification at lower denaturation temperature-PCR (COLD-PCR), multiplex PCR, sequencing by reversible dye terminators, and the like. terminator, paired-end sequencing, near-term sequencing, exonuclease sequencing, sequencing by ligation, short-read sequencing, single-molecule sequencing, sequencing-by-synthesis, real-time sequencing, reverse-terminator sequencingExamples of sequencing methods include, but are not limited to, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, and combinations thereof. In some embodiments, sequencing can be performed by a genetic analyzer, such as a commercially available genetic analyzer from Illumina or Applied Biosystems. In some embodiments, the sequencing method can be massively parallel sequencing, which simultaneously (or in rapid succession) sequences at least 100, 1000, 10,000, 100,000, 1 million, 10 million, 100 million, or 1 billion polynucleotide molecules.

[0137] After sequencing, a quality score is assigned to the read. The quality score can be a representation of the read that indicates whether the read can be useful in subsequent analysis based on a threshold. In some cases, some reads are not of sufficient quality or length to perform subsequent mapping steps. Sequencing reads with a quality score of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered from the dataset. In other cases, sequencing reads assigned a quality score of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered from the dataset. In step 306, genome fragment reads that meet a specified quality score threshold are mapped to a reference genome or a reference sequence that is known not to contain mutations. After mapping alignment, a mapping score is assigned to the sequence read. The mapping score can be a representation or read that is mapped back to a reference sequence, indicating whether each position is uniquely mappable. In some cases, the read may be a sequence that is not relevant to mutation analysis.For example, some sequence reads may be generated from contaminating polynucleotides.The sequencing reads that have a mapping score of at least 90%, 95%, 99%, 99.9%, 99.99% or 99.999% can be filtered from the data set.In other cases, the sequencing reads that are assigned a mapping score of less than 90%, 95%, 99%, 99.9%, 99.99% or 99.999% can be filtered from the data set.For each mappable base, the base that does not meet the minimum threshold of mappability, i.e., low-quality base, can be replaced by the corresponding base found in the reference sequence.

[0138] The method and system described herein can be used to detect a large number of cancers.Cancer cells, like most cells, can be characterized by a turnover rate, where old cells die and are replaced by new cells.In general, dead cells are in contact with the vasculature in a particular subject and may release DNA or DNA fragments into the bloodstream.This also applies to cancer cells in various stages of disease.Cancer cells may also be characterized by various genetic abnormalities, such as copy number variations and mutations, depending on the stage of disease.This phenomenon can be used to detect the presence or absence of cancer individuals using the method and system described herein.

[0139] The types and number of cancers that may be detected may include, but are not limited to, blood cancer, brain cancer, lung cancer, skin cancer, nasal cancer, pharyngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, solid tumors, heterogeneous tumors, homogeneous tumors, etc.

[0140] Furthermore, the systems and methods described herein may also be used to help characterize a particular cancer. Genetic data generated from the systems and methods of the present disclosure may enable practitioners to better characterize a particular type of cancer. Cancers are often heterogeneous in both composition and staging. Genetic profile data may be used to characterize specific subtypes of cancer, which may be important in the diagnosis or treatment of certain subtypes. This information may also provide clues to a subject or practitioner regarding the prognosis of a particular type of cancer.

[0141] The systems and methods provided herein may be used to monitor a known cancer or other disease in a particular subject. This may allow the subject or practitioner to adapt treatment options as the disease progresses. In this example, the systems and methods described herein may be used to build a genetic profile of a particular subject's disease progression. In some cases, the cancer may progress, become aggressive, or become genetically unstable. In other examples, the cancer may remain benign, inactive, or dormant. The systems and methods of the present disclosure may be useful in determining disease progression.

[0142] Furthermore, the systems and methods described herein may be useful in determining the efficacy of a particular treatment option. In one example, a successful treatment may kill more cancer cells and shed DNA, so that the success of the treatment option may actually increase the amount of copy number variations or mutations detected in the subject's blood. In other examples, this may not occur. In another example, perhaps a particular treatment option may correlate with the cancer's genetic profile over time. This correlation may be useful in selecting a therapy. Furthermore, if a cancer is observed to be in remission after treatment, the systems and methods described herein may be useful in monitoring residual disease or disease recurrence.

[0143] Data is transmitted to a computer for processing by a direct connection or over the Internet. The data processing aspects of the system may be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or combinations thereof. The data processing apparatus of the present invention may be implemented in a computer program product tangibly embodied in a machine-readable storage device for execution by a programmable processor, and the data processing method steps of the present invention may be performed by the programmable processor executing a program of instructions to perform the functions of the present invention by operating on input data and producing output. The data processing aspects of the present invention may conveniently be implemented in one or more computer programs executable on a programmable system comprising at least one programmable processor coupled to receive data and instructions from a data storage system, at least one input device, and at least one output device, and to transmit data and instructions to the data storage system, at least one input device, and at least one output device. Each computer program may be implemented in a high-level procedural or object-oriented programming language, and may be implemented in assembly or machine language if desired. In any case, the language may be a compiled or interpreted language. Suitable processors include, by way of example, general-purpose and special-purpose microprocessors. Generally, a processor receives instructions and data from a read-only memory and / or a random-access memory. Suitable storage devices for tangibly embodying computer program instructions and data include, by way of example, semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks and removable disks; magneto-optical disks; and non-volatile memory of all kinds, including CD-ROM disks. Any of the foregoing may be supplemented by, or incorporated in, ASICs (application-specific integrated circuits).

[0144] To interact with a user, the method can be implemented using a computer system having a display device, such as a monitor or LCD (liquid crystal display) screen, for displaying information to the user, and an input device, such as a keyboard, a two-dimensional pointing device, such as a mouse or trackball, or a three-dimensional pointing device, such as a data glove or gyroscopic mouse, that allows the user to input information into the computer system.The computer system can be programmed to provide a graphical user interface through which a computer program interacts with the user.The computer system can be programmed to provide a virtual reality three-dimensional display interface. [Example]

[0145] Example 1: Genomic analysis of clinical cohorts and cfDNA We tested plasma samples from 501 individuals, including 75 with HCC and 426 without cancer. Among the cancer-free individuals, 133 had conditions that increased HCC risk, such as cirrhosis of any cause or viral hepatitis without cirrhosis. Blood samples were prospectively collected from HCC patients at various cancer stages and from high-risk individuals at Johns Hopkins University, while the remaining samples were identified through screening activities at other US or EU hospitals (US / EU cohort) (Table 1). We isolated 0.5–5 ml of plasma from each of these individuals, created genomic libraries, and sequenced cfDNA fragments using low-coverage whole-genome sequencing (approximately 2.6x coverage) with an average of 49 million high-quality paired reads per sample, containing 9 Gb of sequence data (24, 25). In addition to the US / EU cohort, we examined whole-genome sequencing data of 223 patients from Hong Kong as a validation cohort, including patients with resectable early-stage HCC (n = 90, stage A = 85, B = 5), HBV (n = 66), and HBV-associated cirrhosis (n = 35), as well as healthy controls without liver disease (n = 32) (Hong Kong cohort) (Table 1) (15, 28).

[0146] Genome-wide cfDNA fragmentation profiles characterized by underlying chromatin structure We used the DELFI approach to evaluate the fragmentome and generate fragmentation profiles across the genome in 473 non-overlapping 5-Mb regions, each containing approximately 80,000 fragments and spanning approximately 2.4 GB of the genome (24). Fragmentation profiles were consistent among individuals without cancer but varied widely among HCC patients (Figure 1A). The profiles of cirrhosis patients were closer to those of non-cancer individuals without cirrhosis than to those of HCC patients without cirrhosis (Figure 1A). Similarly, patients with viral hepatitis had fragmentation profiles nearly identical to those of non-cancer individuals without liver disease (Figure 1A).

[0147] To investigate the origin of cfDNA fragmentation patterns, we compared genome-wide fragmentome profiles with the open (A) and closed (B) compartments of high-throughput sequencing chromosome structure capture (Hi-C). We found that the cfDNA patterns of healthy individuals were highly correlated with those of lymphoblastoid cells (Figure 1B). Analysis of cfDNA profiles from 10 HCC patients with high ctDNA levels revealed that their fragmentomes reflected two components: one resembling the profiles of non-cancer individuals and a distinct cfDNA component with high similarity to the A / B compartments previously estimated from liver cancer (Figure 1B) (29). In addition, when these two components were estimated, the predicted liver component cfDNA profile shared high similarity with the genome-wide A / B compartments of liver cancer, while the profiles of HCC patients shared only moderate similarity with liver cancer (Figure 1B, C). In contrast, the profiles of cancer-free individuals were closer to the A / B compartment of lymphoblastoid cells (Figures 1B and 1C). These analyses suggest that the cfDNA fragmentomes from individuals with HCC represent a mixture of cfDNA profiles from the chromatin compartments of peripheral blood- and liver cancer-derived cells.

[0148] Disease-specific transcription factors inferred from genome-wide cfDNA fragmentation Because chromatin organization reflects the underlying cellular transcriptional program (30-32), we investigated whether cfDNA fragmentation characteristics could reflect changes resulting from altered DNA binding of transcription factors (TFs) in liver cancer. To identify DNA binding sites for all known TFs, we analyzed 5,620 CHIP-seq experiments from the ReMap2020 database (33). For each TF, we calculated the total cfDNA coverage across all identified binding sites (4K-490K per sample) compared to the entire contiguous genome coverage, thereby creating a single metric for each TF in each sample. We compared these TFs in patients with and without HCC to identify TFs with the greatest and least differences in genome-wide binding site coverage in cfDNA (Figure 2A, B). Gene set enrichment analysis using the DisGeNET database of gene-disease associations revealed that differences in cfDNA TF binding coverage between HCC and cancer-free individuals were predicted to be associated with liver and other cancers (Figure 2C, D). In addition, the top-scoring individual TFs corresponded to those with known biological relevance to chromatin organization and liver cancer (Table 2). These included members of the activator protein 1 (AP1) complex, including the JUN, JUND, ATF2, and ATF7 genes, which integrate extracellular signals (34) and are associated with liver tumorigenesis (35, 36); transcriptional enhancer factor domain family member 4 (TEAD4), which has been shown to have a role in oncogenesis in HCC (37, 38); poly(C)-binding protein 2 (PCBP2) transcriptional coregulator, which, when overexpressed, is associated with a worse prognosis for HCC patients (39); prohibitin 2 (PHB), which promotes HCC progression (40); and AT-rich interacting domain 3A (ARID3A), an oncogenic transcription factor that, when upregulated, enhances liver cancer malignancy (41).Similar analysis of cfDNA fragmentation data from our recent study of patients in the LUCAS lung cancer diagnostic trial (24) showed differential enrichment of coverage in binding sites for lung cancer-associated TFs (Figure 2C, E). Overall, these findings suggest that changes in cfDNA fragmentation in patients with liver cancer and other cancers result from alterations in numerous transcriptional profiles present within cancer cells.

[0149] Genomic alterations in HCC revealed by cfDNA fragmentome Because the cfDNA fragmentome may contain alterations associated with large-scale genomic changes released from cancer cells (24, 25), we also examined circulating chromosome gains and losses in these patients. In addition to the genome-wide fragmentation profile attributed to chromatin and transcription factor alterations observed in liver cancer patients (Figure 3A), our analysis revealed altered representation of chromosome arms consistent with commonly gained or lost chromosomes in liver cancer reported in a previous TCGA large-scale genomic study of HCC (n=372) (Figure 3B). This included increased cfDNA representation of 1q, 7p, 7q, and 8q, and decreased levels of 4q, 8p, 9p, 13q, and 21q, all of which are known to be increased or decreased, respectively, in HCC (42, 43). Importantly, these alterations were observed in HCC patients but not in cancer-free individuals, even those with cirrhosis or chronic liver disease (Figure 3B).

[0150] DELFI model for HCC detection Given the direct relationship between genomic and chromatin alterations and cfDNA fragmentation in liver cancer, we used a machine learning approach to determine whether alterations in the cfDNA fragmentome could distinguish HCC patients from those without cancer. We previously used this approach to develop a robust classifier for lung cancer detection, which was externally validated in an independent population (24). We determined the performance of this classifier in the US / EU cohort by repeated 5-fold cross-validation and generated a score for each individual (DELFI score) averaged over 10 cross-validation iterations. The resulting model contained a combination of regional and large-scale fragmentation features that were optimal for identifying individuals with liver cancer (Figure 5). These features encompassed the majority of the informative chromosomal, chromatin, and regional alterations identified above, accounting for >90% of the variance in fragmentation profiles between samples.

[0151] Because clinical characteristics can influence tumor biomarkers, we investigated whether measurements of liver dysfunction or demographic parameters such as age, sex, race, or weight were associated with DELFI scores in cancer-free individuals for whom this information was available.

[0152] We observed no association between DELFI score and age (R = 0.18, p = 0.08, Spearman correlation) (Figure 6A) or differences in DELFI scores between men and women (p = 0.58, Wilcoxon test) (Figure 6B). Asians and African Americans have been shown to have a higher incidence of liver cancer diagnosed at more advanced stages (44), and we observed small differences in fragmentation scores among cancer-free high-risk individuals across these and other racial or ethnic groups. However, these analyses are limited by the lack of information on clinical covariates in some of these cases (p = 0.037 for patients with viral hepatitis and p = 0.026 for patients with cirrhosis, Kruskal-Wallis test) (Figure 7). Among individuals with cirrhosis, we found a correlation between the extent of liver disease, as measured by the Child-Pugh score, and the DELFI score (R = 0.58, p = 8.6e-5, Spearman correlation) (Figure 8). Increased body mass index (BMI), a risk factor for NAFLD and liver cancer, was not associated with changes in DELFI score in patients with viral hepatitis (R = 0.027, p = 0.85, Spearman correlation). However, lower BMI in patients with cirrhosis was associated with higher DELFI scores (R = -0.23, p = 0.043, Spearman correlation), likely due to cachexia in patients with severe cirrhosis (Figure 9).

[0153] We next examined the relationship between DELFI scores and the presence and stage of liver cancer in high-risk groups. The 133 cancer-free individuals had low DELFI scores, with median DELFI scores of 0.078 and 0.080 for individuals with viral hepatitis or cirrhosis, respectively. In contrast, the 75 HCC patients had significantly higher median DELFI scores across all BCLC stages, including stage 0 = 0.46, stage A = 0.61, stage B = 0.83, and stage C = 0.92 (p < 0.01 for stages 0, A, B, or C, Wilcoxon rank-sum test, Figure 4A). The receiver operating characteristic (ROC) curve for the DELFI approach to identify HCC patients showed an area under the curve (AUC) of 0.90 (95% CI = 0.86-0.94) among high-risk individuals. Performance for early-stage HCC remained robust, with AUCs of 0.9 and 0.81 for BCLC stages 0 and A. Individuals with advanced-stage HCC (BCLC C) were almost perfectly detected (AUC > 0.97) among those analyzed (Figure 4C).

[0154] To extend these analyses to individuals at low risk of developing liver cancer, we investigated the ability of the DELFI model to distinguish between individuals with cancer and those from the general population (n = 293) without viral hepatitis or cirrhosis. In this larger cohort, where additional features could be included in cross-validation training, we created a DELFI model for the general population using the model features described above and also including cfDNA coverage of CHIP-seq-derived transcription factor binding sites from liver cell lines available in the ReMAP database. This approach performed well for cancer detection among these individuals (AUC = 0.98). We evaluated the model's performance at a specificity of 99%, a threshold appropriate for average-risk populations (25), and observed an overall sensitivity of 80% in this setting (Figure 4B), with sensitivity exceeding 65% across all disease stages. Use of a model that did not incorporate transcription factor binding sites resulted in slightly reduced performance, and there was a high correlation between rank-ordered scores using our DELFI model for the high-risk and screening populations (R = 0.64, p < 1e-15) (Figures 10 and 11).

[0155] To investigate the relationship between fragmentation profile and liver cancer progression, we evaluated the size, number and characteristics of liver cancer lesions, and whether the etiology of neoplasia is associated with abnormal fragmentation profile, if this information is available.We found that tumor size and lesion number are positively correlated with DELFI score (R=0.42 and 0.31, respectively, p=0.00026 and p=0.0064, Spearman correlation) (Figure 12), which is consistent with the idea that fragmentation profile is associated with overall tumor burden.

[0156] Among patients with resectable stage (0, A, and B) liver cancer, cancer etiology, e.g., viral hepatitis or cirrhosis due to alcohol, NAFLD, or idiopathic causes, resulted in similar DELFI scores (p=0.43, Kruskal-Wallis test) (Figure 13). These findings suggest that the fragmentation profile was the result of ongoing tumor-associated cfDNA processes and was not influenced by early events in tumor development.

[0157] To examine the actual impact of this method in the context of HCC detection, we compared the performance of DELFI fragmentome with the current screening index, alpha-fetoprotein (AFP) levels. AFP levels were elevated above the recommended screening threshold of 20 ng / ml in 39 of 75 individuals (52%) with cancer, consistent with previous reports (45). Among individuals with AFP levels below 20 ng / ml who were not detected by this approach, DELFI detected 30 of 36 (83%). The use of AFP measurements was considered to detect 8 of 24 patients (33%) with stage 0 / A disease, 17 of 30 patients (57%) with stage B disease, and 14 of 21 patients (66%) with stage C disease (Figure 14). In contrast, the DELFI approach detected 19 of 24 patients (79%) with stage 0 / A disease, 25 of 30 patients (83%) with stage B disease, and 20 of 21 patients (95%) with stage C disease. Overall, genome-wide cfDNA fragmentation analysis has improved performance compared to AFP detection of HCC, and the combination of DELFI and AFP may provide improved detection over the DELFI approach alone, as we observed a combined sensitivity of 92% with a combined specificity of 80%.

[0158] External validation of the DELFI model in East Asian populations with HCC In addition to our own cross-validation analysis of the US / EU cohort, we tested the fixed DELFI model in 223 patients from the Hong Kong cohort. These included patients with mostly resectable early-stage HCC (n = 90, stage A = 85, B = 5) and 101 patients with cirrhosis or HBV infection. Although these samples were previously sequenced using a different sequencer (HiSeq 2000 vs. Novaseq; read length 76 bp vs. 100 bp), a different library preparation, and a higher number of PCR cycles (14 vs. 4 cycles), we observed genome-wide patterns similar to our previous analysis (Figure 15). The fragmentation profiles of patients with viral hepatitis and cirrhosis and healthy individuals were highly consistent across the genome, whereas those of HCC patients were variable and disordered (Figure 16). Furthermore, the chromosomal alterations observed in plasma from the Hong Kong cohort were similar to those in the earlier US / EU cohort, as were those from TCGA-derived cancers (Figure 17). Overall, in this validation cohort, the DELFI model distinguished HCC patients from high-risk disease patients with an AUC of 0.97 (Figure 4D). These findings suggest that the underlying characteristics of cfDNA fragmentation were similar in this cohort and that DELFI is a robust method for detecting HCC and is generalizable across various high-risk populations.

[0159] Simulation of DELFI performance on a population scale To evaluate how well our approach works for surveillance and detection in high-risk patients for liver cancer, we used Monte Carlo simulations to evaluate the DELFI model in a theoretical population of 100,000 high-risk individuals. Given the importance of early cancer detection, our modeling focused on detecting stage 0 / A disease. We compared the DELFI approach with the current standard of care, simultaneous ultrasonography and AFP, and modeled the uncertainty in the sensitivity and specificity of these surveillance modalities in this theoretical population with probability distributions centered on empirical estimates from our cohort or from a previous report (11, see Methods). Despite surveillance recommendations, compliance with HCC surveillance in the United States is poor, with even the most generous estimate of 39% compliance (46), resulting in an average of 40,042 individuals being tested in this theoretical population (95% CI, 21,320–61,890). Given the high availability and compliance of blood tests, with reported adherence rates of 80–90% for blood-based biomarkers (47, 48), we conservatively assumed that an average of 75% (95% CI, 60–90%) of this population could be tested using the DELFI approach. Because the prevalence of cirrhosis, viral hepatitis, and the co-occurrence of these comorbidities with HCC may vary by region, we used prior probability distributions to reflect our own uncertainty regarding the composition of these diseases and expected regional differences. Monte Carlo simulations using these probability distributions (Methods) revealed that ultrasound with AFP detected an average of 2,233 (95% CI, 1,088-3,699) individuals with liver cancer (Figure 18). Using DELFI, we detected an average of 2,794 additional liver cancer cases, a 2.46-fold increase (95% CI, 1.25-4.57-fold increase) compared with ultrasound with AFP alone (Figure 18).The DELFI approach is expected to not only substantially improve liver cancer detection, but also reduce the false-negative rate (FNR), i.e., the proportion of cancers missed at the time of testing, from 38% (95% CI, 25%-51.5%) for ultrasound with AFP to 24% (95% CI, 9%-42.6%) for DELFI. Furthermore, the negative predictive value (NPV) of the test is expected to increase from 95.7% (95% CI, 93.8%-97.3%) for ultrasound with AFP to 97.1% (95% CI, 94.8%-99.0%) for DELFI. These analyses suggest significant benefits for the population as a whole from using a highly specific blood-based early detection test as a tool for liver cancer detection.

[0160] Consideration Overall, in this study, we demonstrate the use of genome-wide cfDNA fragmentome features to detect HCC with high sensitivity and specificity. Furthermore, we show that fragmentation profiles capture genomic and chromatin features, including alterations known to be important in HCC. Our cfDNA fragmentome approach has robust performance in detecting HCC, including very early-stage disease, regardless of disease etiology. To our knowledge, this is the first genome-wide fragmentation analysis to be independently validated in distinct high-risk populations, with stable and robust performance across different racial and ethnic groups in the United States and Hong Kong.

[0161] Our results also demonstrated that disease-specific transcription factor signatures can be obtained through analysis of genome-wide cfDNA fragmentation profiles. While such analyses have been performed using specific transcription factors to distinguish small cell lung cancer from non-small cell lung cancer (24), this study suggests that analysis of disease-specific transcriptional regulation using genome-wide cfDNA fragmentation may improve the detection and identification of tissue of origin in cancer patients. With sufficient patient numbers, cfDNA transcriptional profiles may further improve machine learning algorithms for detecting HCC and other cancers.

[0162] Compared with other solid cancers, HCC is unique in that there exists a well-defined, large high-risk population with an average annual risk of developing HCC of 3–4%, for whom regular cancer screening every 6 months is recommended (49). Unfortunately, currently available tests have limited diagnostic utility, especially for early-stage disease (13). In our study, the sensitivity of AFP for HCC detection was 52%, consistent with the known performance of this biomarker (13). Ultrasound-based surveillance also has technical limitations, including operator dependence and low sensitivity in obese patients and those with cirrhosis (50). Most importantly, ultrasound has low compliance with established guidelines, less than 20% worldwide (10, 11), compared with the much higher adherence rates for blood tests for other conditions (47). Despite these challenges, HCC screening provides an overall survival benefit in patients with HBV (51) and cirrhosis (10), highlighting the significant need to improve current screening tests. The high performance of cfDNA fragmentome analysis in HCC detection, together with its cost-effectiveness, will enable DELFI to become an accessible screening test for HCC and will enable DELFI to increase screening rates beyond the current dismal level. An interesting aspect of HCC-specific cfDNA analysis is that transplantation is the most curative treatment for early- to intermediate-stage HCC, and studies of cfDNA in posttransplant patients have shown promise (52). HCC surveillance in posttransplant patients using a liquid biopsy approach may have a dual role in tracking recurrence and rejection.

[0163] Although this study represents a promising improvement over current screening approaches, it does have some limitations. For example, the sample size of individuals with HCC included in this study was relatively small. Although an independent validation cohort was conducted with preanalytical variations in the laboratory and sequencing method, the fact that the DELFI approach performed well in this population suggests that this method may eventually be available across a range of diagnostic laboratories. Larger validation studies will be necessary before this approach can become clinically useful. Nevertheless, the finding that scalable, cost-effective, noninvasive cfDNA fragmentome analysis can detect patients with liver cancer may provide an opportunity to screen high-risk and general populations worldwide.

[0164] material and method Test group For the US / EU cohort, samples from 208 patients, including 75 with HCC and 133 high-risk patients without HCC, were prospectively collected as part of the HCC Biomarker Registry and AIDS Linked to the IntraVenous Experience (ALIVE) study at the Johns Hopkins University School of Medicine under a protocol approved by the Johns Hopkins University Institutional Review Board. HCC was defined by histology or appropriate imaging features as defined by approved guidelines. Tumor staging was determined by the Barcelona Clinic Liver Cancer Staging System (BCLC). Detailed clinical data were extracted from electronic medical records.

[0165] High-risk patients were defined as individuals with cirrhosis of any etiology and / or chronic hepatitis B or C for whom regular HCC screening is recommended by professional society guidelines (50). We also included 38 patients with hepatitis B or cirrhosis retrospectively collected by BioIVT (Westbury, NY). AFP levels were quantified in a clinical laboratory by an affiliated institution using a US Food and Drug Administration-approved AFP test.

[0166] The US / EU cohort also included previously analyzed samples from 293 cancer-free individuals originally obtained from two colorectal cancer screening clinical trial cohorts in Denmark (Endoscopy III) and the Netherlands (COCOS, Dutch Clinical Trial Registration ID NTR182946) (24). The Endoscopy III project protocol was approved by the Regional Ethics Committee and the Danish Data Protection Agency, and ethical approval was obtained from the Dutch Health Council for the COCOS study protocol. The inclusion criterion for both the Dutch and Danish cohorts was any individual aged 50–75 years eligible for colorectal cancer screening. All patients included had either a negative FIT test or a negative colonoscopy result.

[0167] For the Hong Kong cohort, all enrolled subjects provided written informed consent, and the study was approved by the Joint Chinese University of Hong Kong and New Territories East Cluster Clinical Research Ethics Committee ( 15 , 28 ).

[0168] Specimen collection and storage Sample collection was performed as follows: venous peripheral blood was collected into one K2-EDTA tube and two serum gel tubes. Within 2 hours of collection, the tubes were centrifuged at 2330 g for 10 minutes at 4°C. The plasma was transferred to a new tube, and the sample was spun at 14,000 rpm (18,000 rcf) for 10 minutes at room temperature to pellet any remaining cellular debris. After centrifugation, the EDTA plasma was aliquoted and stored at -80°C for cfDNA analysis.

[0169] Sequencing library preparation For all plasma samples, circulating cell-free DNA was isolated from 2–4 ml of plasma using the Qiagen QIAamp Circulating Nucleic Acid Kit (Qiagen GmbH), eluted in 52 μl of RNase-free water containing 0.04% sodium azide (Qiagen GmbH), and stored in LoBind tubes (Eppendorf AG) at −20°C. cfDNA concentration and quality were assessed using a Bioanalyzer 2100 (Agilent Technologies).

[0170] Next-generation sequencing (NGS) cfDNA libraries for WGS were prepared using 15 ng of cfDNA when available, or the entire purified volume if less than 15 ng. Briefly, genomic libraries were prepared using the Illumina (New England Biolabs (NEB)) NEBNext DNA Library Preparation Kit, following the manufacturer's guidelines with four major modifications: (i) to minimize sample loss during the elution and tube transfer steps, the library purification step was performed after the on-bead AMPure XP (Beckman Coulter) approach; (ii) the volumes of NEBNext End Repair, A-tailing, and adapter ligation enzymes and buffers were adjusted appropriately to accommodate the on-bead AMPure XP purification; (iii) Illumina dual-index adapters were used in the ligation reaction; and (iv) the cfDNA libraries were amplified using Phusion hot-start polymerase. All samples underwent a four-cycle PCR amplification after the DNA ligation step.

[0171] Low-coverage whole genome sequencing and alignment Whole genome libraries from cancer patients and cancer-free individuals were prepared as described (24), with the modification that they were sequenced using 100-bp paired-end runs (200 cycles) on the Illumina NovaSeq platform at 1–2× coverage per genome. Before alignment, adaptor sequences were filtered from reads using fastp software (53). Sequence reads were aligned to the hg19 human reference genome using Bowtie2 (54), and duplicate reads were removed using Sambamba (55). After alignment, each aligned pair was converted to a genomic interval representing the sequenced DNA fragment using bedtools (56). Only reads with a MAPQ score of at least 30 were retained. Read pairs were further filtered if they overlapped with the Duke Excluded Regions blacklist (genome.ucsc.edu / cgi-bin / hgTrackUi?db=hg19&g=wgEncodeMapability). To capture large-scale epigenetic differences in fragmentation across the genome that can be estimated from low-coverage WGS, we tiled the hg19 reference genome into non-overlapping 5 Mb bins. Bins with an average GC content of less than 0.3 and an average mappability of less than 0.9 were excluded, leaving 473 bins spanning approximately 2.4 GB of the genome. Following Mathios et al. (24), GC correction was performed independently on short (<150 bp) and long (≥150 bp) cfDNA fragments using an external panel of 20 cancer-free individuals sequenced on NovaSeq to generate target distributions.

[0172] As reported (15, Jiang, 2018 #1645), Fastq files for patients in the Hong Kong cohort were obtained from the Chinese University of Hong Kong (CUHK) Circulating Nucleic Acids Research Group and processed to generate DELFI features as described above and in Mathios et al. GC correction was performed by normalizing to the target distribution provided at github.com / cancer-genomics / PlasmaToolsNovaseq.hg19; the same target distribution was also used for GC correction in the US / EU cohort. The validation set consisted of libraries constructed by 14-cycle PCR and sequenced on a HiSeq 2000. These libraries were normalized to the 4-cycle NovaSeq target distribution to facilitate comparisons between studies. One sample each from the cirrhosis and HBV groups was excluded because they were identified as having an HCC diagnosis.

[0173] Chromatin structure analysis The A / B compartments for liver cancer tissues and lymphoblastoid cells were obtained from github.com / Jfortin1 / TCGA_AB_Compartments and github.com / Jfortin1 / HiC_AB_Compartments as described in (29). Two reference tracks were compared to identify informative 100 kb bins, defined as bins in which chromatin domains differed between the two reference tracks, or significant differences in eigenvalues ​​corresponding to a z-score of >1.96 or <−1.96 (p=.05) across all eigenvalue differences.

[0174] We calculated the median fragmentation profiles of the 10 liver samples with the highest tumor fraction estimates by ichorCNA(57) and 10 randomly selected individuals without cancer. This information was used to extract the median estimated liver content in plasma, weighted by the ichor score of each individual plasma sample.

[0175] Genome-wide transcription factor analysis Chromatin immunoprecipitation sequencing (ChIP-Seq) peaks from 5620 experiments were downloaded from the ReMap 2020 database (33). This set was filtered for experiments with more than 4000 peaks, resulting in 4293 experiments. For each peak on an autosome, we defined the center of the peak as position 0.

[0176] The average coverage at each position (-3,000 to +3,000 relative to the center of each peak) was calculated across all peaks for each sample. For ROC curves, relative coverage for each sample was calculated as the average coverage within a + / - 100 bp window surrounding the center of the binding site divided by the average coverage within a + / - 250 bp window surrounding 2,750 bp upstream and downstream of the binding site. ROC curves were generated using pROC 1.16.2 (58). The AUC for each peak set was ranked. Each transcription factor was matched to its NCBI ID, leaving 797 unique transcription factors ranked by AUC. This ranked list was the input for the gseDGN function from the DOSE package in R. The output from this was ranked by normalized enrichment score (NSE).

[0177] Whole genome fragment features Fragmentation features were calculated as described by Mathios et al. (24). Briefly, the ratio of short to long fragments was calculated for 473 non-overlapping 5 MB bins across the genome, and z-scores representing arm gain / loss were calculated for autosomal arms. Principal components of the ratios and z-scores representing >90% of the variance were used to train a machine learning model.

[0178] Machine Learning and Cross-Validation Analysis Two machine learning models were developed: one for the high-risk population (GBM using the features of Mathios et al.) and one for the low-risk general population (penalized logistic regression using the features of Mathios et al. and coverage from transcription factor binding sites). Similar to Mathios et al. (24), these models were trained on the US / EU cohort using Caret with 10 replicates of 5-fold cross-validation. Each sample's score was calculated as the average across repeats and evaluated using AUC-ROC. The first model was used for high-risk non-cancer and HCC patients, and the second was used for non-cancer individuals without liver pathology. The locked high-risk model trained on the US / EU cohort was then applied to the Hong Kong cohort to perform cancer predictions on the external validation set.

[0179] TCGA analysis Copy number data from the TCGA HCC cancer cohort (LIHC n=372) were searched and analyzed using the RTCGA v1.16.0 package to determine the frequency of copy number gains and losses in 473 5-mb bins for this cohort. (24) Somatic copy number alteration (SCNA) thresholds used by Mathios et al. (24, 59) were used to determine gains and losses in the HCC cohort.

[0180] Association of clinical covariates with DELFI scores Potential associations between clinical covariates (for patients for whom this information was available) and DELFI scores were assessed using Spearman's rank correlation coefficient (continuous variables) and Kruskal-Wallis one-way analysis of variance (categorical variables).

[0181] simulation We compared the DELFI approach with ultrasound and AFP in a theoretical surveillance population using Monte Carlo simulations. We used 95% confidence interval estimates of sensitivity and specificity for DELFI and published 95% confidence intervals for ultrasound and AFP (13). The R package epiR was used to derive a priori predictive probability distributions (beta distributions) from these confidence intervals ((60) R package version 2.47, CRAN.R-project.org / package=Epi). Zhao et al. (2017) reported an adherence rate of 39% (95% CI: 21%-65%) for US and AFP surveillance. Because reported adherence rates for other noninvasive blood-based tests are greater than 75% (47, 48), we assumed adherence to DELFI to be 60% or greater with a probability of 0.975 or greater. Using these confidence interval estimates, epiR was used to derive a priori predictive beta distributions for adherence rates. We simulated multinomial probabilities from a Dirichlet distribution with parameters 230, 680, 60, 23, and 7 for the prevalence of hepatitis B, cirrhosis, hepatitis B + HCC, cirrhosis + HCC, and hepatitis B + cirrhosis + HCC, respectively. For a single Monte Carlo simulation for ultrasound and AFP examinations, we used (i) sampling the probability of compliance (η) from the ex ante predictive distribution; (ii) simulate the number of 100,000 individuals (S) participating in surveillance (S~Binomial(η, 100,000)); (iii) sample the probability of comorbidity (Dirichlet(230, 680, 60, 23, 7)); (iv) calculate the prevalence of HCC (θ); (v) Simulate HCC cases (P~Binomial(θ,S)) and calculate the number of cancer-free individuals (N=SP); (vi) sampling the sensitivity (se) and specificity (sp) from the corresponding prior distributions, and (vii) True positives (TP~Binomial(P, se)) and false positives (FP~Binomial(N, 1-sp)) were sampled. With TP and FP, we calculated NPV as (true negatives) / (true negatives + false negatives), where true negatives = N-FP and false negatives = P-TP. We repeated the above simulation 1000 times to obtain distributions of TP, FP, and NPV. Using the sensitivity, specificity, and adherence parameters for the DELFI approach, we repeated the same Monte Carlo analysis to allow comparison between these two surveillance methodologies.

[0182] Table 1. Patient demographics and clinical information TIFF2025538137000002.tif210124*To compare data from individuals with and without liver cancer, P values ​​were calculated for the following variables: mean age using Student's unpaired two-tailed t-test, gender distribution, cirrhosis etiology, and Child-Pugh stage using the X2 test. # Validation cohort data were obtained from Jiang et al., PNAS, 2015.

[0183] Table 2. Top-scoring transcription factors in US / EU cohort samples TIFF2025538137000003.tif45166*TF = transcription factor, HAT = histone acetyltransferase

[0184] References TIFF2025538137000004.tif142161TIFF2025538137000005.tif216160TIFF2025538137000006.tif216161TIFF2025538137000007.tif193160

[0185] Other Aspects While the present disclosure has been described in conjunction with its detailed description, the above description is intended to illustrate, but not limit, the scope of the disclosure, which is defined by the appended claims. Other aspects, advantages, and modifications are within the scope of the appended claims.

[0186] The patents and scientific literature referred to herein demonstrate knowledge available to those skilled in the art. All U.S. patents and published or unpublished U.S. patent applications cited herein are incorporated herein by reference. All published foreign patents and patent applications cited herein are incorporated herein by reference. All other published literature, documents, manuscripts, and scientific literature cited herein are incorporated herein by reference.

Claims

1. 1. A method of diagnosing a liver disease or disorder in a subject, comprising: isolating circulating cell-free DNA (cfDNA) from the subject; performing whole genome sequencing of the cfDNA molecules to generate a genomic library and a fragmentation profile; Comparing the fragmentation profile with that of a healthy individual; and Diagnosing whether the subject has a liver disease or disorder.

2. 10. The method of claim 1, wherein the fragmentation profile is consistent for subjects without cancer.

3. The method of claim 1, wherein the fragmentation profile is highly variable for subjects with liver cancer.

4. The method of any one of claims 1 to 3, further comprising identifying the cellular origin of the cfDNA fragmentation profile.

5. 5. The method of claim 4, wherein identifying the cellular origin of the cfDNA fragmentation profile comprises comparing a genome-wide fragmentome profile with high-throughput sequencing chromosome conformation capture (Hi-C).

6. 6. The method of claim 5, wherein the cellular origin of the cfDNA fragmentation profile of the healthy individual corresponds to lymphoblastoid cells as the cellular origin.

7. The method of claim 5, wherein the cellular origin of the cfDNA fragmentation profile of the subject with liver cancer corresponds to the cfDNA fragmentation profile of the chromatin compartment of peripheral blood cells.

8. 8. The method of any one of claims 1 to 7, further comprising determining whether the cfDNA fragmentation profile correlates with alterations in DNA binding of transcription factors.

9. 9. The method of claim 8, wherein transcription factor DNA binding sites are determined by calculating total cfDNA coverage across all identified transcription factor DNA binding sites compared to the entire contiguous genome coverage, creating a single metric for each transcription factor per sample.

10. The method of claim 9, wherein the transcription factor DNA binding sites identified in subjects with liver cancer are compared with transcription factor DNA binding sites in healthy individuals.

11. The method of claim 10, wherein the cfDNA fragmentation profile of the subject with liver cancer correlates with alterations in DNA binding of transcription factors.

12. The method of any one of claims 1 to 11, further comprising determining chromosomal gains or losses in liver cancer subjects compared to healthy individuals.

13. 13. The method of any one of claims 1 to 12, further comprising a machine learning model for determining alterations in the cfDNA fragmentation profile, wherein the method classifies the subject as a cancer patient based on the subject's cfDNA fragmentation profile.

14. 14. The method of claim 13, wherein the machine learning model generates a score for each subject based on a combination of regional and large-scale fragmentation profiles.

15. The method of claim 14, wherein the generated score is for diagnosing a stage of liver cancer.

16. 16. The method of any one of claims 1 to 15, further comprising administering a cancer treatment to the subject diagnosed with a liver disease or disorder.

17. 17. The method of claim 16, wherein the cancer treatment comprises surgery, adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, hormonal therapy, cytotoxic therapy, immunotherapy, adoptive T cell therapy, targeted therapy, and combinations thereof.

18. 1. A method for diagnosing liver cancer, comprising the steps of: isolating circulating cell-free DNA (cfDNA) from the biological sample and performing whole genome sequencing of the cfDNA molecules to generate a genomic library and fragmentation profile; identifying the cellular origin of the cfDNA fragmentation profile, including comparing the genome-wide fragmentome profile with high-throughput sequencing chromosome structure capture (Hi-C); correlating the cfDNA fragmentation profile with alterations in DNA binding of transcription factors; determining chromosomal gains or losses in liver cancer subjects compared to healthy individuals; running a machine learning model to determine an alteration in the cfDNA fragmentation profile that classifies the subject as a cancer patient based on the subject's cfDNA fragmentation profile; thereby diagnosing liver cancer in said subject and administering cancer treatment to said diagnosed subject.

19. 19. The method of claim 18, wherein the cellular origin of the cfDNA fragmentation profile of the healthy individual corresponds to lymphoblastoid cells as the cellular origin.

20. The method of claim 18, wherein the cellular origin of the cfDNA fragmentation profile of the subject with liver cancer corresponds to the cfDNA fragmentation profile of the chromatin compartment of peripheral blood cells.

21. 21. The method of any one of claims 18-20, wherein the transcription factor DNA binding sites are determined by calculating total cfDNA coverage across all identified transcription factor DNA binding sites compared to the entire contiguous genome coverage, creating a single metric for each transcription factor per sample.

22. 22. The method of claim 21, wherein the transcription factor DNA binding sites identified in subjects with liver cancer are compared with transcription factor DNA binding sites in healthy individuals.

23. 23. The method of claim 22, wherein the cfDNA fragmentation profile of the subject with liver cancer correlates with alterations in DNA binding of transcription factors.

24. 24. The method of any one of claims 18 to 23, wherein the machine learning model generates a score for each subject based on a combination of regional and large-scale fragmentation profiles.

25. 25. The method of claim 24, wherein the generated score is for diagnosing a stage of liver cancer.

26. 25. The method of claim 24, wherein the generated score is for diagnosing hepatocellular carcinoma (HCC).

27. 25. The method of claim 24, wherein the generated score is for the differential diagnosis of liver diseases, including cirrhosis.

28. 28. The method of any one of claims 18 to 27, wherein the treatment comprises surgery, adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, hormonal therapy, cytotoxic therapy, immunotherapy, adoptive T cell therapy, targeted therapy, and combinations thereof.