Detection of liver cancer using cell free DNA fragmentation

Through whole-genome sequencing and machine learning model analysis of free DNA of circulating cells, the DNA binding sites of specific transcription factor of liver cancer was identified, which solved the problem of insufficient sensitivity and specificity of existing liver cancer screening methods, and achieved early and low-cost liver cancer diagnosis.

CN120380165APending Publication Date: 2025-07-25JOHNS HOPKINS UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380083968.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-06
Filing Date
2023-11-06
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing liver cancer screening methods are insufficient in sensitivity and specificity, especially in early disease detection, and are costly, making it difficult to meet the needs of high-risk people around the world.

Method used

By whole-genome sequencing of subjects' circulating cell free DNA (cfDNA), fragmented profiles were analyzed and combined with machine learning models, the changes in the DNA binding sites of liver cancer-specific transcription factor were identified, and early diagnosis of liver cancer was achieved.

Benefits of technology

It provides a non-invasive and ultra-sensitive liver cancer screening method, which can detect liver cancer in the early stage, improves the sensitivity and specificity of the detection, reduces costs, and is suitable for a wide range of applications in high-risk groups.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120380165A_ABST
    Figure CN120380165A_ABST
Patent Text Reader

Abstract

A method for liver cancer detection uses a combination of whole genome mutations and fragmentation characteristics of cfDNA to facilitate cancer screening.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority benefit of U.S. Provisional Application No. 63 / 423,003, filed on November 6, 2022, which is hereby incorporated by reference in its entirety.

[0002] Statement Regarding Federally Sponsored Research

[0003] This invention was made with government support under grants GM136577, CA121113, CA006973, and CA233259 awarded by the National Institutes of Health. The government has certain rights in this invention.

[0004] Background

[0005] Liver cancer is the cause of an alarming amount of morbidity and mortality worldwide, with over 900,000 new diagnoses and over 800,000 deaths annually (1). In the United States, liver cancer is one of the few cancers with increasing incidence and mortality trends over the past 20 years. Ninety percent of liver cancer cases are hepatocellular carcinoma (HCC), and survival rates depend largely on the stage of the disease at diagnosis. When the cancer is localized (44% of patients), the five-year survival rate is 34%, when the cancer is regional (27% of patients), the five-year survival rate is 12%, and when distant disease is detected (18% of patients), the five-year survival rate is 3% (2). A large, well-defined population is at significantly increased risk of developing HCC, including individuals with chronic hepatitis B (HBV) infection or cirrhosis due to various causes, including hepatitis C (HCV) (3), non-alcoholic fatty liver disease (NAFLD) (4), heavy alcohol consumption (5), aflatoxin, and other conditions (6). Worldwide, 350 million individuals have chronic viral hepatitis infection, and 50 million individuals have cirrhosis (7). In the United States, 4.5 million individuals have chronic HCV, and 29 million individuals are diagnosed with NAFLD. Up to one-third of cirrhotic patients and 25 - 40% of HBV patients will develop HCC during their lifetimes, with an annual risk of up to 8% in cirrhotic patients (8). An increasing number of individuals at risk of liver cancer - including 29 million individuals in the United States - have NAFLD, and 20% of HCC cases in this group occur without cirrhosis (9). Medical societies around the world recommend screening high-risk groups, currently using abdominal ultrasound imaging (with or without alpha-fetoprotein (AFP)). However, overall compliance with international guidelines remains low, with less than one-fifth of eligible individuals globally receiving some level of surveillance, and less than 2% following recommended screening (10 - 12). Many factors contribute to low compliance with screening guidelines, including the identification of high-risk individuals, the infrastructure and personnel requirements of imaging-based screening methods (11). Current screening tests (including ultrasound imaging, with or without AFP) show limited sensitivity of 47 - 84% and specificities of 67% to over 90% (13). In addition, the lack of non-invasive diagnostic protocols for NAFLD indicates that the population currently not covered by HCC screening recommendations is increasing. Therefore, there is an urgent need globally to develop accessible and sensitive HCC screening programs.

[0006] One recent approach to overcoming these challenges has been to develop novel cell-free DNA (cfDNA)-based biomarkers for cancer detection. Somatic mutation-based assays have been used as biomarkers for liver cancer but are limited by the need for tissue-based mutation identification and the low levels of changes detectable in plasma (14). Methylation profiling and copy number alterations at specific loci and across the genome have also provided viable approaches for detecting liver cancer, but their detection sensitivity in very early disease is still not optimal (15-20). Multiple recently developed early cancer detection assays appear to be useful for detecting multiple cancers (including liver cancer) in average-risk populations (21), but there are no published reports on the use of these assays in high-risk HCC populations. In addition, the cost of most cfDNA-based assays is far higher than the estimated affordable cost of screening tests in the United States and globally (22). Combining these assays with AFP improves performance but requires two separate tests and still has limitations in early disease (23). SUMMARY OF THE INVENTION

[0007] Provided herein is a non-invasive and ultrasensitive analysis of single cell-free DNA (cfDNA) molecules to detect changes in fragmentation profiles across the genome. Compared to healthy subjects, liver cancer patients were found to have altered fragmentation profiles due to genomic and chromatin changes in liver cancer, including those from regions associated with liver-specific transcription factors.

[0008] Accordingly, in one aspect, a method for diagnosing a liver disease or disorder in a subject includes: isolating circulating cell-free DNA (cfDNA) from the subject; performing whole genome sequencing of the cfDNA molecules to generate a genomic library and a fragmentation profile; comparing the fragmentation profile with that of a healthy subject and / or a reference genome, and diagnosing whether the subject has a liver disease or disorder. In certain embodiments, the fragmentation profile is consistent for subjects without cancer, while the fragmentation profile varies greatly for subjects with liver cancer. In certain embodiments, the method further includes identifying the cellular origin of the cfDNA fragmentation profile, wherein the identification of the cellular origin of the cfDNA fragmentation profile includes comparing the whole genome fragmentome profile with high-throughput sequencing chromosome conformation capture (Hi-C). In certain embodiments, the cellular origin of the cfDNA fragmentation profile of a healthy subject is related to lymphoblastoid-like cells as the cellular origin. In certain embodiments, the cellular origin of the cfDNA fragmentation profile of a subject with liver cancer is related to the cfDNA fragmentation profile of the chromatin compartments of peripheral blood cells. In certain embodiments, the method further includes determining whether the cfDNA fragmentation profile is associated with altered binding of transcription factors to DNA. In certain embodiments, the transcription factor DNA binding sites are determined by calculating the aggregate cfDNA coverage across all identified transcription factor DNA binding sites relative to the overall adjacent genomic coverage to generate a single metric for each transcription factor in each sample. In certain embodiments, the transcription factor DNA binding sites identified in subjects with liver cancer are compared with those in healthy subjects. In certain embodiments, the cfDNA fragmentation profile of a subject with liver cancer is associated with altered binding of transcription factors to DNA. In certain embodiments, the method further includes determining chromosomal gains or losses in subjects with liver cancer compared to healthy subjects. In certain embodiments, the method further includes a machine learning model for determining changes in the cfDNA fragmentation profile, the model classifying the subject as a cancer patient based on the cfDNA fragmentation profile of the subject. In certain embodiments, the machine learning model generates a score for each subject based on a combination of regional and large-scale fragmentation profiles. In certain embodiments, the generated score can diagnose the stage of liver cancer.

[0009] In another aspect, systems and methods for detecting cancer are disclosed: performing whole genome sequencing of cfDNA molecules to generate a genomic library and a fragmentation profile and data input into a computer memory; executing a machine learning model for determining changes in the cfDNA fragmentation profile, the model classifying the subject as a cancer patient based on the cfDNA fragmentation profile of the subject.

[0010] In another aspect, a method for diagnosing liver cancer includes: isolating circulating cell-free DNA (cfDNA) from a biological sample and performing whole-genome sequencing of cfDNA molecules to generate a genomic library and a fragmentation profile; identifying the cellular origin of the cfDNA fragmentation profile including comparing the whole-genome fragmentome profile with high-throughput sequencing chromosome conformation capture (Hi-C); correlating the cfDNA fragmentation profile with changes in transcription factor binding to DNA; determining chromosomal gains or losses in liver cancer subjects compared to healthy subjects; performing a machine learning model for determining changes in the cfDNA fragmentation profile, the model classifying the subject as a cancer patient based on the subject's cfDNA fragmentation profile; thereby diagnosing liver cancer and administering cancer treatment to the subject. In certain embodiments, the cellular origin of the cfDNA fragmentation profile of a healthy subject is related to lymphoblastoid-like cells as the cellular origin. In certain embodiments, the cellular origin of the cfDNA fragmentation profile of a subject with liver cancer is related to the cfDNA fragmentation profile of the chromatin compartment of peripheral blood cells. In certain embodiments, the transcription factor DNA binding sites are determined by calculating the aggregate cfDNA coverage across all identified transcription factor DNA binding sites relative to the overall adjacent genomic coverage to generate a single metric for each transcription factor in each sample. In certain embodiments, the transcription factor DNA binding sites identified in subjects with liver cancer are compared with the transcription factor DNA binding sites in healthy subjects. In certain embodiments, the cfDNA fragmentation profile of a subject with liver cancer is related to changes in transcription factor binding to DNA. In certain embodiments, the machine learning model generates a score for each subject based on a combination of regional and large-scale fragmentation profiles. In certain embodiments, the generated score can diagnose the stage of liver cancer. In certain embodiments, the generated score can diagnose hepatocellular carcinoma (HCC). In certain embodiments, the generated score can perform differential diagnosis of liver diseases including cirrhosis. In certain embodiments, the treatment comprises: surgery, adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, hormone therapy, cytotoxic therapy, immunotherapy, adoptive T cell therapy, targeted therapy, and combinations thereof.

[0011] Definitions

[0012] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Further understanding, terms (such as those defined in common dictionaries) should be interpreted as having a meaning consistent with their meaning in the context of the relevant field and should not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0013] As used herein, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. Further, to the extent that the terms "comprises," "comprising," "has," "having," "includes," "including," "with," or variants thereof are used in the detailed description and / or the claims, such terms are intended to be inclusive in a manner similar to the term "comprising."

[0014] The terms "about" or "approximately" mean, as determined by one of ordinary skill in the art, within an acceptable error range of a particular value, which depends in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, in accordance with practice in the art, "about" can mean within one standard deviation or more than one standard deviation. Alternatively, "about" can mean within a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value or range. Alternatively, especially for biological systems or processes, the term can mean within an order of magnitude of within five-fold, and within two-fold, of a particular value. Where the application and claims describe a particular value, unless otherwise stated, the term "about" should be assumed to mean within an acceptable error range of the particular value.

[0015] The terms "align," "alignment," "map," or "mapping" mean that one or more sequences are identified as matching a known sequence from a reference genome in the order of their nucleic acid molecules. Such alignments can be done manually or by computer algorithms, examples of which include the Efficient Local Alignment of Nucleotide Data (ELAND) computer program distributed as part of the Illumina Genomics Analysts pipeline. The match of sequence reads in an alignment can be 100% sequence match or less than 100% (imperfect match).

[0016] As used herein, the term "cancer" refers to a disease, disorder, trait, genotype or phenotype known in the art that is characterized by dysregulated cell growth or replication. The terms "neoplasm" and "tumor" denote such abnormal tissues that grow by more rapid cell proliferation than normal and continue to grow after the stimulus that initiated the proliferation has been removed. Such abnormal tissues exhibit partial or complete lack of the structural organization and functional coordination of normal tissues and may be benign (such as a benign tumor) or malignant (such as a malignant tumor). Examples of cancers include liver cancer (including hepatocellular carcinoma (HCC)), lung cancer (including non-small cell lung cancer), gastric cancer, colorectal cancer, and, for example, leukemia, such as acute myeloid leukemia (AML), chronic myeloid leukemia (CML), acute lymphoblastic leukemia (ALL) and chronic lymphocytic leukemia, AIDS-related cancers such as Kaposi's sarcoma; breast cancer; bone cancer such as osteosarcoma, chondrosarcoma, Ewing's sarcoma, fibrosarcoma, giant cell tumor, ameloblastoma and chordoma; brain cancer such as meningioma, glioblastoma, low-grade astrocytoma, oligodendroglioma, pituitary tumor, schwannoma and metastatic brain cancer; cancers of the head and neck, including various lymphomas such as mantle cell lymphoma, non-Hodgkin lymphoma, adenoma, squamous cell carcinoma, laryngeal cancer, gallbladder and bile duct cancer, retinoblastoma such as retinoblastoma, esophageal cancer, gastric cancer, multiple myeloma, ovarian cancer, uterine cancer, thyroid cancer, testicular cancer, endometrial cancer, melanoma, bladder cancer, prostate cancer, pancreatic cancer, sarcoma, Wilms' tumor, cervical cancer, head and neck cancer, skin cancer, nasopharyngeal cancer, liposarcoma, epithelial cancer, renal cell cancer, gallbladder adenocarcinoma, parotid adenocarcinoma, endometrial sarcoma, multi-drug resistant cancer; and proliferative diseases and disorders such as neovascularization associated with tumor angiogenesis.

[0017] The term "cell-free nucleic acid", "cell-free polynucleotide", "cell-free DNA" or "cfDNA" denotes nucleic acid fragments that circulate in an individual's body (e.g., bloodstream) and are derived from one or more healthy cells and / or one or more cancer cells. In addition, cfDNA may be from other sources such as viruses, fetuses, etc.

[0018] The term "circulating tumor DNA" or "ctDNA" denotes nucleic acid fragments derived from tumor cells or other types of cancer cells that may be released into an individual's blood due to biological processes such as apoptosis or necrosis of dying cells, or actively released by living tumor cells.

[0019] As used herein, the terms "comprising", "including", or "contained" and variations thereof, when referring to the elements defined or described in an article, composition, device, method, process, system, etc., are meant to be inclusive or open-ended to allow for additional elements, such that the article, composition, device, method, process, system, etc. defined or described includes those specified elements (or, where appropriate, their equivalents), and further indicates that other elements may be included and still fall within the scope / definition of the article, composition, device, method, process, system, etc. defined or described.

[0020] "Diagnosing" or "being diagnosed" means identifying the presence or nature of a pathological condition. Diagnostic methods vary in terms of sensitivity and specificity. The "sensitivity" of a diagnostic assay is the percentage of diseased individuals with a positive test result (the percentage of "true positives"). Diseased individuals not detected by the assay are "false negatives". Subjects who are not diseased and have a negative test result in the assay are called "true negatives". The "specificity" of a diagnostic assay is equal to 1 minus the false positive rate, where the "false positive" rate is defined as the proportion of subjects who are not diseased but have a positive test result. Although a particular diagnostic method may not provide a definitive diagnosis of a condition, it meets the requirements if it provides positive indications that assist in the diagnosis.

[0021] As used herein, an "effective amount" means an amount that provides a therapeutic or prophylactic benefit.

[0022] The terms "fragmentation profile", "position-dependent differences in fragmentation mode", and "differences in fragment size and coverage across the genome in a position-dependent manner" as used herein are equivalent and may be used interchangeably. In certain embodiments, determining the cfDNA fragmentation profile in a mammal can be used to identify that the mammal has cancer. For example, cfDNA fragments obtained from a mammal (e.g., from a sample obtained from a mammal) can be subjected to low-coverage whole-genome sequencing, and the sequenced fragments can be mapped to the genome (e.g., in non-overlapping windows) and evaluated to determine the cfDNA fragmentation profile. As described herein, the cfDNA fragmentation profile of a mammal having cancer is more heterogeneous (e.g., in fragment length) than that of a healthy mammal (e.g., a mammal not having cancer). Accordingly, the present disclosure also provides methods and materials for evaluating, monitoring, and / or treating a mammal (e.g., a human) having or suspected of having cancer. In certain embodiments, the document provides methods and materials for identifying that a mammal has cancer. For example, a sample obtained from a mammal (e.g., a blood sample) can be evaluated to determine the presence of cancer and optionally the tissue of origin of the cancer in the mammal, at least in part based on the cfDNA fragmentation profile of the mammal. In certain embodiments, methods and materials for monitoring a mammal having cancer are provided. For example, a sample obtained from a mammal (e.g., a blood sample) can be evaluated to determine the presence of cancer in the mammal, at least in part based on the cfDNA fragmentation profile of the mammal. In certain embodiments, methods and materials for identifying that a mammal has cancer and administering one or more cancer treatments to the mammal to treat the mammal are provided. For example, a sample obtained from a mammal (e.g., a blood sample) can be evaluated to determine whether the mammal has cancer, at least in part based on the cfDNA fragmentation profile of the mammal, and one or more cancer treatments can be administered to the mammal.

[0023] The "frequency" of a mutation as used herein is defined as the number of variants per million evaluated positions in all sequenced DNA molecules.

[0024] The term "genomic nucleic acid" or "genomic DNA" refers to nucleic acids including chromosomal DNA derived from one or more healthy (e.g., non-tumor) cells. In various embodiments, genomic DNA can be extracted from cells of the blood cell lineage, such as white blood cells (WBCs).

[0025] As used herein, the term "mutation spectrum" refers to the types and frequencies of mutations observed in bins across the genome. By comparing the mutation spectra between genomic regions that are more frequently altered in cancer and those that are more frequently mutated in normal cfDNA, multi-regional differences can be determined.

[0026] "Optional" or "optionally" means that the subsequent described event or circumstance may or may not occur, and the description includes the case where the event or circumstance occurs and the case where the event or circumstance does not occur.

[0027] The term "or" as used in this specification and the appended claims is generally used in its inclusive sense of "and / or" unless the context clearly indicates otherwise.

[0028] "Parenteral" administration of an immunogenic composition includes, for example, subcutaneous (s.c.), intravenous (i.v.), intramuscular (i.m.) or intrasternal injection or infusion techniques.

[0029] The terms "patient" or "individual" or "subject" are used interchangeably herein and refer to a mammalian subject to be treated, with human patients being preferred. In certain embodiments, the methods of the present invention can be used in experimental animals, veterinary applications, and the development of animal disease models, including, but not limited to, rodents (including mice, rats, and hamsters) and primates.

[0030] The term "reference genome" as used herein can refer to a digital or previously identified database of nucleic acid sequences that is assembled as a representative example of a species or subject. A reference genome may be assembled from nucleic acid sequences from multiple subjects, samples, or organisms and does not necessarily represent the nucleic acid composition of a single individual. A reference genome can be used to map sequencing reads from a sample to chromosomal locations. For example, reference genomes for human subjects and many other organisms can be found at the National Center for Biotechnology Information (at ncbi.nlm.nih.gov).

[0031] The term "read" or "read length" refers to any nucleotide sequence, including sequence reads obtained from an individual and / or nucleotide sequences derived from the initial sequence reads of a sample obtained from an individual.

[0032] The terms "sample", "patient sample", "biological sample", etc. encompass various sample types obtained from a patient, individual, or subject, and can be used for diagnostic, prognostic, and / or monitoring assays. A patient sample can be obtained from a healthy subject, an affected patient, or a patient with lung cancer. In certain embodiments, the "provided" sample can be obtained by the person (or machine) performing the assay, or it can be obtained by another person and transferred to the person (or machine) performing the assay. Additionally, a sample obtained from a patient can be separated, and only a portion can be used for diagnosis. Further, a sample or a portion thereof can be stored under conditions that preserve the sample for subsequent analysis. This definition specifically encompasses blood and other liquid samples of biological origin (including but not limited to peripheral blood, serum, plasma, cord blood, amniotic fluid, cerebrospinal fluid, urine, saliva, feces, and synovial fluid), solid tissue samples (such as biopsy samples or tissue cultures or cells and their progeny derived therefrom). In certain embodiments, the sample comprises cerebrospinal fluid. In one specific embodiment, the sample comprises a blood sample. In another embodiment, the sample comprises a plasma sample. In yet another embodiment, a serum sample is used. The definition of "sample" also includes a sample that has been manipulated in any way after its acquisition (such as by centrifugation, filtration, precipitation, dialysis, chromatography, reagent treatment, washing, or enrichment for a specific cell population). The term further encompasses clinical samples, and also includes cultured cells, cell supernatants, tissue samples, organs, etc. The sample can also comprise fresh frozen and / or formalin-fixed, paraffin-embedded tissue blocks, such as those prepared from a clinical or pathological biopsy, prepared for pathological analysis or immunohistochemical studies.

[0033] The term "sequence read" refers to a nucleotide sequence read from a sample obtained from an individual. Sequence reads can be obtained by various methods known in the art.

[0034] As defined herein, a "therapeutically effective" amount (i.e., an effective dose) of a compound or reagent is an amount sufficient to produce a therapeutically (e.g., clinically) desired result. The composition can be administered once or more per day to once or more per week, including once every other day. Skilled artisans will appreciate that certain factors can affect the dosage and timing required to effectively treat a subject, including but not limited to the severity of the disease or disorder, previous treatment, the overall health and / or age of the subject, and the presence of other diseases. Additionally, treatment of a subject with a therapeutically effective amount of a compound of the invention can include a single treatment or a series of treatments.

[0035] As used herein, the term "treat" ("treat", "treating", "treatment", etc.) refers to alleviating or ameliorating a disorder and / or symptoms associated therewith. It should be understood that treating a disorder or condition does not necessarily require complete elimination of the disorder, condition or symptoms associated therewith, although such elimination is not excluded.

[0036] Gene: All genes, gene names and gene products disclosed herein are intended to correspond to homologs of any species to which the compositions and methods disclosed herein are applicable. It should be understood that when a gene or gene product from a particular species is disclosed, such disclosure is intended to be exemplary only and should not be construed as limiting, unless the context in which it appears clearly indicates otherwise. Thus, for example, with respect to the genes or gene products disclosed herein, homologs and / or orthologs from other species are intended to be encompassed.

[0037] Range: Throughout this disclosure, various aspects of the invention may be presented in range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as a rigid limitation on the scope of the invention. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as the individual numerical values within that range. For example, a description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6, etc., as well as the individual numbers within that range, for example, 1, 2, 2.7, 3, 4, 5, 5.3, and 6. This applies regardless of the breadth of the range.

[0038] Any composition or method provided herein can be combined with one or more of any other compositions and methods provided herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawings will be provided by the Patent Office upon request and payment of the necessary fee.

[0040] Figure 1 (including Figure 1A-1C) The figures show that the genome-wide fragmentation profiles reflect the underlying chromatin structure. Figure 1A: Fragmentation profiles in 473 non-overlapping 5 Mb genomic regions of 501 individuals. The fragmentation profiles of cancer individuals showed significant heterogeneity compared to non-cancer individuals with and without liver disease. Figure 1B: Comparison of plasma fragmentation features with reference A / B compartments. Track 1 shows the A / B compartments (29) extracted from liver cancer tissues. Track 2 shows the median liver cancer component (57) extracted from HCC plasma samples of 10 liver disease patients with high tumor fractions by ichor CAN. Track 3 shows the median fragmentation profile in the plasma of these 10 HCC samples, and Track 4 shows the median profile of 10 healthy plasma samples. Track 5 shows the A / B compartments of lymphoblasts (29). These five tracks show chromosome 22 as an example, where darker shading indicates informative regions of the genome, and the two reference tracks differ in domain (open / closed) or magnitude. Figure 1C: Shows further comparison results of plasma fragmentation features with reference A / B compartments.

[0041] Figure 2. A series of figures (including Figure 2A-2E ) show the fragmentation profiles in HCC patients, highlighting liver-specific transcription factors. Figure 2A: Coverage at and around transcription factor binding sites (TFBS) of 9 TFs, for which the relative coverage at the binding sites has the highest separation between HCC samples and non-cancer samples. The mean of each group is plotted, and ±1 standard deviation (SD) is shaded. These confidence intervals (CI) show separation, highlighting that the difference in coverage at TFBS can provide information about the cancer status. Figure 2B: Coverage at and around TFBS of 9 TFs, which has the lowest separation between HCC and non-cancer samples in the US / EU cohort. These CIs largely overlap, reflecting their poor discriminatory status as TFBS. Gene set enrichment analysis of TFs analyzed in both hepatocellular carcinoma and lung adenocarcinoma showed that TFs were selectively enriched in many pathways related to liver cancer and lung cancer respectively (Figure 2C), including adult liver cancer and lung adenocarcinoma (Figure 2D, Figure 2E).

[0042] Figure 3 (including Figure 3A-3C) shows that high-dimensional fragmentation features reflect liver cancer biology and are incorporated into the DELFI machine learning scheme. Figure 3A: Heatmap reflecting the complexity of whole-genome fragmentation and transcription factor binding site features used in the DELFI machine learning scheme. Each row represents a sample, while the columns show individual genomic features. Figure 3B: Analysis of copy number changes in 372 TCGA liver cancer tissues and 501 individual plasma samples reflects biological consistency. Copy number changes that occur in TCGA (red = increase, blue = decrease) are also present in HCC plasma at the chromosomal arm level, but not in individuals without cancer. Figure 3C: Heatmap depicting the contribution of individual genomic regions to the final trained DELFI model. Fragmentation features are summarized as three main components in the model, while aneuploidy is summarized as arm-level z-scores. The top, middle, and right panels depict the coefficients of the fragmentation component, arm-level z-score, and transcription factor binding site in the model, respectively. CNV, copy number variation.

[0043] Figure 4 (including Figure 4A-4D ) shows a series of figures indicating that the DELFI machine learning model detects liver cancer with high sensitivity and specificity. Figure 4A: DELFI scores for liver disease and cancer stages in the US / EU cohort in the screening and surveillance models. The average DELFI score for patients with cirrhosis is higher than that for individuals without cancer or with viral hepatitis, but lower than that for liver cancer at all stages. Patients with liver cancer at all stages have relatively high DELFI scores, with individuals in stage C consistently having the highest DELFI scores. Figure 4B: ROC analysis for the US / EU general population cohort and the high-risk surveillance cohort. Figure 4C: ROC analysis for the US / EU general population and surveillance cohorts by BCLC stage, showing high sensitivity and specificity at each stage. Figure 4D: ROC analysis for the fixed surveillance model applied to the Hong Kong, China cohort, which includes 90 HCC individuals with HCC (85 with BCLC stage A cancer and 5 with BCLC stage B cancer), 101 individuals with cirrhosis and viral hepatitis, and 32 individuals without cancer or liver disease.

[0044] Figure 5 The heatmap depicts the contribution of individual genomic regions to the final trained DELFI surveillance model. Fragmentation features are summarized as three main components in the model, while aneuploidy is summarized as arm-level z-scores. The top panel and the right panel depict the different importance of the fragmentation component and the arm-level z-score in the model, respectively.

[0045] Figure 6A and 6B show that the DELFI scores of individuals without cancer are not affected by age and gender. By age ( Figure 6A ) or gender ( Figure 6B) DELFI scores of non-cancer or liver disease individuals (n = 293, 45 females, 88 males) divided.

[0046] Figure 7 A series of graphs show that the DELFI scores of individuals without cancer do not differ among different racial or ethnic groups.

[0047] Figure 8 The graph shows that the DELFI scores of cirrhotic patients (n = 40) with available information are not related to the Child-Pugh score of cirrhosis severity.

[0048] Figure 9 A series of graphs show that the DELFI scores of individuals without cancer (n = 53 patients with viral hepatitis, n = 78 cirrhotic patients) vary with BMI.

[0049] Figure 10A and 10B show the performance of alternative DELFI models with and without TF binding site relative coverage features. ROC analysis of the monitoring model, combined with the transcription factor binding site features of liver-derived CHIP-Seq analysis in high-risk US / EU cohorts, and used for screening models without transcription factor binding sites in the general population US / EU cohorts ( Figure 10A ), and by BCLC stage ( Figure 10B ).

[0050] Figure 11 The graph shown shows the correlation between the ranked DELFI scores in the monitoring and screening models. Using our high-risk and screening DELFI models in all HCC patients (n = 75), there is a high correlation between the ranked scores.

[0051] Figure 12A and 12B A series of graphs show that the DELFI scores of HCC patients are related to the size and number of tumor lesions. The DELFI scores of HCC patients (n = 75) are positively correlated with the increasing lesion diameter ( Figure 12A ), and show a positive trend with the increasing number of lesions ( Figure 12B ).

[0052] Figure 13 The graph shows that the DELFI scores of HCC patients with resectable stage disease (BCLC 0, A, and B) (n = 54) do not show differences due to underlying liver disease etiologies. C-stage patients (n = 21) were not included in this analysis because these patients had etiologies mainly related to viral hepatitis (n = 14).

[0053] Figure 14The figure shows that the DELFI score in HCC patients is correlated with the AFP level. Overall, 39 out of 75 HCC individuals had an AFP level above the recommended screening threshold of 20 ng / ml, as indicated by the vertical dashed line. All but three of these individuals could be detected at the DELFI threshold, which corresponds to 80% specificity (0.26) in the high-risk cohort, as indicated by the horizontal dashed line. Among the individuals not detectable by AFP, DELFI detected 30 out of 36.

[0054] Figure 15 The figure shows that there was no difference in the correlation between the fragmentation profile and the median non-cancer profile in the non-cancer subgroup.

[0055] Figure 16 The figure shows that the genome-wide fragmentation profiles in the Hong Kong, China validation cohort showed similarity to the genome-wide fragmentation profiles in the US / EU cohort. Fragmentation profiles in 473 non-overlapping 5-Mb genomic regions in 223 individuals. The fragmentation profiles of cancer individuals showed significant heterogeneity compared to non-cancer individuals with and without liver disease, similar to the observations in the US / EU cohort.

[0056] Figure 17 Showed that chromosomal copy number changes detected in plasma were biologically consistent in the US / EU and Hong Kong, China cohorts and in TCGA tissue samples. Red indicates an increase in copy number, while blue indicates a decrease in copy number. These changes were readily observable in both tissue and HCC plasma but not in individuals without cancer. CNV, copy number variation.

[0057] Figure 18 (including Figure 18 A-18D) shows the modeling of DELFI implementation in HCC screening. Figure 18 A: In a theoretical population of 100,000 high-risk individuals, the uncertainty in the sensitivity and specificity of ultrasound, AFP, and DELFI screening was modeled. The number of HCCs detected in these individuals ( Figure 18 B), false negative rate ( Figure 18 C), and negative predictive value ( Figure 18The predicted distribution in panel (D) incorporates the variation in the prevalence of liver cancer and compliance with image- and blood-based screening. The center line in the box plot represents the median, the upper bound of the box represents the third quartile (75th percentile), and the lower bound of the box represents the first quartile (25th percentile). The upper whisker is the maximum value within 1.5 times the interquartile range above the 75th percentile, and the lower whisker is the minimum value within 1.5 times the interquartile range below the 25th percentile, indicating that the chromosomal copy number changes detected in plasma are biologically consistent in the US / EU and Hong Kong, China cohorts and in the TCGA tissue samples. Red indicates an increase in copy number, while blue indicates a decrease in copy number. These changes are readily observable in both tissue and HCC plasma but are not evident in individuals without cancer. CNV, copy number variation. Detailed implementation

[0058] DELFI (DNA fragment evaluation for early interception) uses the whole-genome fragmentation profile to provide a high-performance and cost-effective solution for cancer detection (21, 22). Fragmentation and methylation information has also been shown to distinguish patients with liver cancer from those without cancer (24), although this approach requires two different cfDNA library preparation and analysis methods. To date, no study has validated a whole-genome approach for detecting HCC in independent cohorts or different high-risk groups.

[0059] Thus, in certain embodiments, a method of diagnosing a liver disease or disorder in a subject includes: isolating cell-free circulating DNA (cfDNA) from the subject; performing whole-genome sequencing of the cfDNA molecules to generate a genomic library and a fragmentation profile; comparing the fragmentation profile to that of a healthy subject and / or a reference genome, and diagnosing whether the subject has a liver disease or disorder.

[0060] In another embodiment, a method of diagnosing liver cancer includes: isolating cell-free circulating DNA (cfDNA) from a biological sample and performing whole-genome sequencing of the cfDNA molecules to generate a genomic library and a fragmentation profile; identifying the cellular origin of the cfDNA fragmentation profile, which includes comparing the whole-genome fragmentome profile to high-throughput chromosome conformation capture (Hi-C); correlating the cfDNA fragmentation profile with changes in transcription factor binding to DNA; determining chromosomal gains or losses in liver cancer subjects compared to healthy subjects; performing a machine learning model for determining changes in the cfDNA fragmentation profile, the model classifying the subject as a cancer patient based on the subject's cfDNA fragmentation profile; thereby diagnosing liver cancer and administering cancer treatment to the subject.

[0061] cfDNA fragmentation profile

[0062] The cfDNA fragmentation profile can include one or more cfDNA fragmentation patterns. The cfDNA fragmentation pattern can include any suitable cfDNA fragmentation pattern. Examples of cfDNA fragmentation patterns include, but are not limited to, median fragment size, fragment size distribution, ratio of small cfDNA fragments to large cfDNA fragments, and coverage of cfDNA fragments. In certain embodiments, the cfDNA fragmentation pattern includes two or more (e.g., 2, 3, or 4) of median fragment size, fragment size distribution, ratio of small cfDNA fragments to large cfDNA fragments, and coverage of cfDNA fragments. In certain embodiments, the cfDNA fragmentation profile can be a genome-wide cfDNA profile (e.g., a genome-wide cfDNA profile in a window across the entire genome). In certain embodiments, the cfDNA fragmentation profile can be a targeted region profile. The targeted region can be any suitable part of the genome (e.g., a chromosomal region). Examples of chromosomal regions for which the cfDNA fragmentation profile can be determined as described herein include, but are not limited to, a part of a chromosome (e.g., a part of 2q, 4p, 5p, 6q, 7p, 8q, 9q, 10q, 11q, 12q, and / or 14q) and a chromosomal arm (e.g., the chromosomal arms of 8q, 13q, 11q, and / or 3p). In certain embodiments, the cfDNA fragmentation profile can include two or more targeted region profiles.

[0063] In certain embodiments, the cfDNA fragmentation profile can be used to identify changes (e.g., alterations) in the length of cfDNA fragments. The alteration can be a genome-wide alteration or an alteration in one or more targeted regions / loci. The target region can be any region containing one or more cancer-specific alterations. In certain embodiments, the cfDNA fragmentation profile can be used to identify (e.g., simultaneously identify) from about 10 alterations to about 500 alterations (e.g., from about 25 to about 500, from about 50 to about 500, from about 100 to about 500, from about 200 to about 500, from about 300 to about 500, from about 10 to about 400, from about 10 to about 300, from about 10 to about 200, from about 10 to about 100, from about 10 to about 50, from about 20 to about 400, from about 30 to about 300, from about 40 to about 200, from about 50 to about 100, from about 20 to about 100, from about 25 to about 75, from about 50 to about 250, or from about 100 to about 200 alterations).

[0064] Any suitable method can be used to obtain the cfDNA fragmentation profile. In certain embodiments, cfDNA from a mammal (e.g., a mammal having or suspected of having cancer) can be processed into a sequencing library, which can be subjected to whole genome sequencing (e.g., low coverage whole genome sequencing), mapped to the genome, and analyzed to determine the cfDNA fragment lengths. The mapped sequences can be analyzed in non-overlapping windows that cover the genome. The windows can be any suitable size. For example, the length of the window can be thousands to millions of bases. As a non-limiting example, the window can be approximately 5 megabases (Mb) in length. Any suitable number of windows can be mapped. For example, dozens to thousands of windows can be mapped in the genome. For example, hundreds to thousands of windows can be mapped in the genome. The cfDNA fragmentation profile can be determined within each window.

[0065] In certain embodiments, the methods and materials described herein can also include machine learning. For example, machine learning can be used to identify mutation frequencies, altered fragmentation profiles (e.g., using the coverage of cfDNA fragments, the fragment sizes of cfDNA fragments, the coverage of chromosomes, and mtDNA).

[0066] In certain embodiments, determining the cfDNA fragmentation profile in a mammal can be used to identify a mammal having cancer. For example, cfDNA fragments obtained from a mammal (e.g., from a sample obtained from a mammal) can be subjected to low-coverage whole-genome sequencing, and the sequenced fragments can be mapped to the genome and evaluated to determine the cfDNA fragmentation profile. As described herein, the cfDNA fragmentation profile of a mammal having cancer is more heterogeneous (e.g., in fragment length) than that of a healthy mammal (e.g., a mammal not having cancer). Accordingly, methods and materials are also provided for evaluating, monitoring, and / or treating a mammal (e.g., a human) having or suspected of having cancer. In certain embodiments, methods and materials are provided for identifying a mammal having cancer. For example, a sample obtained from a mammal (e.g., a blood sample) can be evaluated to determine the presence of cancer and optionally the tissue of origin of the cancer in the mammal, at least in part based on the cfDNA fragmentation profile of the mammal. In certain embodiments, methods and materials are provided for monitoring a mammal having cancer. For example, a sample obtained from a mammal (e.g., a blood sample) can be evaluated to determine the presence of cancer in the mammal, at least in part based on the cfDNA fragmentation profile of the mammal. In certain embodiments, methods and materials are provided for identifying a mammal having cancer and administering one or more cancer treatments to the mammal to treat the mammal. For example, a sample obtained from a mammal (e.g., a blood sample) can be evaluated to determine whether the mammal has cancer, at least in part based on the cfDNA fragmentation profile of the mammal, and one or more cancer treatments can be administered to the mammal.

[0067] In certain embodiments, cfDNA fragmentation profiles can be used to detect tumor-derived DNA. For example, cfDNA fragmentation profiles can be used to detect tumor-derived DNA by comparing the cfDNA fragmentation profile of a mammal with cancer or suspected of having cancer to a reference cfDNA fragmentation profile (e.g., the cfDNA fragmentation profile of a healthy mammal and / or the nucleosomal DNA fragmentation profile of healthy cells of a mammal with cancer or suspected of having cancer). In certain embodiments, the reference cfDNA fragmentation profile is a profile previously generated from a healthy mammal. For example, the methods provided herein can be used to determine the reference cfDNA fragmentation profile in a healthy mammal, and the reference cfDNA fragmentation profile can be stored (e.g., in a computer or other electronic storage medium) for future comparison with a test cfDNA fragmentation profile in a mammal with cancer or suspected of having cancer. In certain embodiments, the reference cfDNA fragmentation profile of a healthy mammal is determined across the entire genome (e.g., the stored cfDNA fragmentation profile). In certain embodiments, the reference cfDNA fragmentation profile of a healthy mammal is determined within sub-genomic intervals (e.g., the stored cfDNA fragmentation profile)

[0068] In certain embodiments, cfDNA fragmentation profiles can be used to identify a mammal (e.g., a human) as having cancer (e.g., liver cancer, colorectal cancer, lung cancer, breast cancer, gastric cancer, pancreatic cancer, cholangiocarcinoma, and / or ovarian cancer).

[0069] The cfDNA fragmentation profile can include a cfDNA fragment size pattern. The cfDNA fragments can be of any suitable size. For example, the cfDNA fragments can be from about 50 base pairs (bp) to about 400 bp in length.

[0070] The cfDNA fragmentation profile can include the cfDNA fragment size distribution. As described herein, the cfDNA size distribution of a mammal with cancer can be more variable than the cfDNA fragment size distribution of a healthy mammal. In certain embodiments, the size distribution can be within the targeted region. A healthy mammal (e.g., a mammal not having cancer) can have a targeted region cfDNA fragment size distribution of about 1 or less than about 1. In certain embodiments, the targeted region cfDNA fragment size distribution of a mammal with cancer can be longer than the targeted region cfDNA fragment size distribution in a healthy mammal (e.g., 10, 15, 20, 25, 30, 35, 40, 45, 50 or more bp longer, or any number of base pairs between these numbers). In certain embodiments, the targeted region cfDNA fragment size distribution of a mammal with cancer can be shorter than the targeted region cfDNA fragment size distribution in a healthy mammal (e.g., 10, 15, 20, 25, 30, 35, 40, 45, 50 or more bp shorter, or any number of base pairs between these numbers). In certain embodiments, the size distribution can be a genome-wide size distribution. A healthy mammal (e.g., a mammal not having cancer) can have a very similar distribution of short cfDNA fragments and long cfDNA fragments across the genome. In certain embodiments, a mammal with cancer can have one or more alterations (e.g., increases and decreases) in the cfDNA fragment sizes across the genome. The one or more alterations can be any suitable chromosomal region of the genome. For example, the alteration can be in a portion of a chromosome. Examples of chromosomal portions that can contain one or more alterations in cfDNA fragment size include, but are not limited to, portions of 2q, 4p, 5p, 6q, 7p, 8q, 9q, 10q, 11q, 12q, and 14q. For example, the alteration can span a chromosomal arm (e.g., an entire chromosomal arm).

[0071] The cfDNA fragmentation profile can include the ratio of small cfDNA fragments to large cfDNA fragments and the correlation of the fragment ratio with a reference fragment ratio. As used herein, with respect to the ratio of small cfDNA fragments to large cfDNA fragments, the small cfDNA fragments can be from about 100 bp in length to about 150 bp in length. As used herein, with respect to the ratio of small cfDNA fragments to large cfDNA fragments, the large cfDNA fragments can be from about 151 bp in length to 220 bp in length. A mammal having cancer can have a lower (e.g., 2-fold lower, 3-fold lower, 4-fold lower, 5-fold lower, 6-fold lower, 7-fold lower, 8-fold lower, 9-fold lower, 10-fold lower or more) fragment ratio correlation (e.g., the correlation of the cfDNA fragment ratio with a reference DNA fragment ratio such as the DNA fragment ratio from one or more healthy mammals) than a healthy mammal. A healthy mammal (e.g., a mammal not having cancer) can have a fragment ratio correlation (e.g., the correlation of the cfDNA fragment ratio with a reference DNA fragment ratio such as the DNA fragment ratio from one or more healthy mammals) of about 1 (e.g., about 0.96). In certain embodiments, a mammal having cancer can have, on average, a fragment ratio correlation (e.g., the correlation of the cfDNA fragment ratio with a reference DNA fragment ratio such as the DNA fragment ratio from one or more healthy mammals) that is lower than that of a healthy mammal.

[0072] The cfDNA fragmentation profile can include the coverage of all fragments. The coverage of all fragments can include windows of coverage (e.g., non-overlapping windows). In certain embodiments, the coverage of all fragments can include windows of small fragments (e.g., fragments from about 100 bp to about 150 bp in length). In certain embodiments, the coverage of all fragments can include windows of large fragments (e.g., fragments from about 151 bp to about 220 bp in length).

[0073] In certain embodiments, the cfDNA fragmentation profile can be used to identify the molecular origin of cfDNA in a patient and to identify genomic and chromatin features associated with fragmentation changes.

[0074] In certain embodiments, cfDNA fragmentation profiles can be used to identify the tissue of origin of cancer (e.g., liver cancer, colorectal cancer, lung cancer, breast cancer, gastric cancer, pancreatic cancer, cholangiocarcinoma, or ovarian cancer). For example, cfDNA fragmentation profiles can be used to identify localized cancer. When the cfDNA fragmentation profile includes a targeted region profile, one or more of the alterations described herein can be used to identify the tissue of origin of cancer. In certain embodiments, one or more alterations in chromosomal regions can be used to identify the tissue of origin of cancer.

[0075] Any suitable method can be used to obtain the cfDNA fragmentation profile. In certain embodiments, cfDNA from a mammal (e.g., a mammal having or suspected of having cancer) can be processed into a sequencing library, which can be subjected to whole-genome sequencing (e.g., low-coverage whole-genome sequencing), mapped to the genome, and analyzed to determine cfDNA fragment lengths. The mapped sequences can be analyzed in non-overlapping windows that cover the genome. The windows can be any suitable size. For example, the length of the window can be from thousands to millions of bases. As a non-limiting example, the window can be about 5 megabases (Mb) long. Any suitable number of windows can be mapped. For example, dozens to thousands of windows can be mapped in the genome. For example, hundreds to thousands of windows can be mapped in the genome. The cfDNA fragmentation profile can be determined within each window.

[0076] In certain embodiments, the methods and materials described herein can also include machine learning. For example, machine learning can be used to identify altered fragmentation profiles (e.g., using coverage of cfDNA fragments, fragment size of cfDNA fragments, coverage of chromosomes, and mtDNA).

[0077] In certain embodiments, the methods and materials described herein can be the only method for identifying a mammal (e.g., a human) having liver cancer. For example, determining the cfDNA fragmentation profile can be the only method for identifying a mammal having liver cancer.

[0078] In certain embodiments, the methods and materials described herein can be used in conjunction with one or more additional methods for identifying a mammal (e.g., a human) having liver cancer. Examples of methods for identifying a mammal having cancer include, but are not limited to, identifying one or more cancer-specific sequence alterations, identifying one or more chromosomal alterations (e.g., aneuploidy and rearrangements), and identifying other cfDNA alterations. For example, determining the cfDNA fragmentation profile can be used in conjunction with identifying one or more cancer-specific mutations in the genome of a mammal to identify the mammal as having liver cancer. For example, determining the cfDNA fragmentation profile can be used in conjunction with identifying one or more aneuploidies in the genome of a mammal to identify the mammal as having liver cancer.

[0079] Therapeutic method

[0080] The methods embodied herein include identifying a mammalian subject having cancer. The methods include: extracting cell-free DNA (cfDNA) from a biological sample of a subject; generating a genomic library from the extracted cfDNA; sequencing individual cfDNA molecules to obtain a fragmentation profile; comparing the fragmentation profile to that of a healthy subject and / or a reference genome, and diagnosing whether the subject has a liver disease or disorder.

[0081] In certain embodiments, a method of diagnosing liver cancer includes: isolating circulating cell-free DNA (cfDNA) from a biological sample and performing whole-genome sequencing of cfDNA molecules to generate a genomic library and a fragmentation profile; identifying the cellular origin of the cfDNA fragmentation profile, which includes comparing the whole-genome fragmentome profile to high-throughput chromosome conformation capture (Hi-C); correlating the cfDNA fragmentation profile with altered binding of transcription factors to DNA; determining chromosomal gains or losses in liver cancer subjects compared to healthy subjects; performing a machine learning model for determining changes in the cfDNA fragmentation profile, the model classifying the subject as a cancer patient based on the subject's cfDNA fragmentation profile; thereby diagnosing liver cancer and administering cancer treatment to the subject.

[0082] In certain embodiments, the subject is diagnosed as having cancer, such as early-stage cancer. In certain embodiments, the type and stage of the liver cancer are identified, and the subject is treated with one or more cancer therapies.

[0083] In certain embodiments, the methods and materials described herein are used to evaluate, monitor, and / or treat a mammal (e.g., a human) having or suspected of having cancer. In certain embodiments, methods and materials are provided for identifying that a mammal has liver cancer. For example, a sample obtained from a mammal (e.g., a blood sample) can be evaluated to determine whether the mammal has cancer, at least in part based on the cfDNA fragmentation profile of the mammal. In certain embodiments, a sample obtained from a mammal (e.g., a blood sample) can be evaluated to determine the tissue of origin of cancer in the mammal, at least in part based on the cfDNA fragmentation profile of the mammal. In certain embodiments, methods and materials are provided for identifying that a mammal has liver cancer and administering one or more treatments to the mammal to treat the mammal. For example, a sample obtained from a mammal (e.g., a blood sample) can be evaluated to determine whether the mammal has liver cancer, at least in part based on the cfDNA fragmentation profile of the mammal, and one or more cancer treatments can be administered to the mammal. In certain embodiments, methods and materials are provided for treating a mammal having cancer. For example, one or more cancer treatments can be administered to a mammal identified as having cancer (e.g., at least in part based on the cfDNA fragmentation profile of the mammal) to treat the mammal. In certain embodiments, during or after a cancer treatment (e.g., any cancer treatment described herein), the mammal can be monitored (or selected for increased monitoring) and / or further diagnostic tests. In certain embodiments, the monitoring can include evaluating a mammal having or suspected of having liver cancer, e.g., by evaluating a sample obtained from the mammal (e.g., a blood sample) to determine the cfDNA fragmentation profile of the mammal described herein, and changes in the cfDNA fragmentation profile over time can be used to identify response to treatment and / or identify that the mammal has cancer (e.g., residual cancer).

[0084] Any suitable mammal can be evaluated, monitored, and / or treated as described herein. The mammal can be a mammal having liver cancer. The mammal can be a mammal suspected of having liver cancer. Examples of mammals that can be evaluated, monitored, and / or treated as described herein include, but are not limited to, humans, primates such as monkeys, dogs, cats, horses, cows, pigs, sheep, mice, and rats. For example, a human having or suspected of having liver cancer can be evaluated to determine cfDNA fragmentation as analyzed herein, and alternatively, can be treated with one or more cancer treatments as described herein.

[0085] Any suitable sample from a mammal can be evaluated as described herein (e.g., to evaluate DNA fragmentation patterns). In certain embodiments, the sample can include DNA (e.g., genomic DNA). In certain embodiments, the sample can include cfDNA (e.g., circulating tumor DNA (ctDNA)). In certain embodiments, the sample can be a fluid sample (e.g., liquid biopsy). Examples of samples that can contain DNA and / or polypeptides include, but are not limited to, blood (e.g., whole blood, serum, or plasma), amniotic fluid, tissue, urine, cerebrospinal fluid, saliva, sputum, bronchoalveolar lavage fluid, bile, lymph fluid, cyst fluid, feces, ascites, Pap smear, breast milk, and exhaled breath condensate. For example, a plasma sample can be evaluated to determine cfDNA fragmentation as analyzed herein.

[0086] A sample from a mammal to be evaluated as described herein (e.g., to evaluate DNA fragmentation patterns) can include any suitable amount of cfDNA. In certain embodiments, the sample can contain a limited amount of DNA. For example, a cfDNA fragmentation profile can be obtained from a sample containing less DNA than typically required by other cfDNA analysis methods, such as those described in, for example, Phallen et al., 2017 Sci Transl Med 9; Cohen et al., 2018 Science 359:926; Newman et al., 2014 Nat Med 20:548; and Newman et al., 2016 Nat Biotechnol 34:547).

[0087] In certain embodiments, the sample can be processed (e.g., to isolate and / or purify DNA and / or polypeptides from the sample). For example, DNA isolation and / or purification can include cell lysis (e.g., using a detergent and / or surfactant), protein removal (e.g., using a protease), and / or RNA removal (e.g., using an RNase). As another example, polypeptide isolation and / or purification can include cell lysis (e.g., using a detergent and / or surfactant), DNA removal (e.g., using a DNase), and / or RNA removal (e.g., using an RNase).

[0088] The cancer can be cancer at any stage. In certain embodiments, the cancer can be early-stage cancer. In certain embodiments, the cancer can be asymptomatic cancer. In certain embodiments, the cancer can be residual disease and / or recurrence (e.g., after surgical resection and / or cancer treatment). The cancer can be any type of cancer. Examples of cancer types that can be evaluated, monitored, and / or treated as described herein include, but are not limited to, colorectal cancer, lung cancer, breast cancer, gastric cancer, pancreatic cancer, cholangiocarcinoma, and ovarian cancer.

[0089] When treating a mammal having or suspected of having liver cancer as described herein, one or more cancer treatments can be administered to the mammal. The cancer treatment can be any suitable cancer treatment. The one or more cancer treatments described herein can be administered to the mammal at any suitable frequency (e.g., once or multiple times over a period of days to weeks). Examples of cancer treatments include, but are not limited to, adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, hormone therapy, cytotoxic therapy, immunotherapy, adoptive T cell therapy (e.g., chimeric antigen receptor and / or T cells having wild-type or modified T cell receptors), targeted therapies such as administration of kinase inhibitors (e.g., kinase inhibitors targeting specific genetic lesions such as translocations or mutations) (e.g., kinase inhibitors, antibodies, bispecific antibodies), signal transduction inhibitors, bispecific antibodies or antibody fragments (e.g., BiTE), monoclonal antibodies, immune checkpoint inhibitors, surgery (e.g., surgical resection), or any combination of the foregoing. In certain embodiments, the cancer treatment can reduce the severity of the cancer, alleviate the symptoms of the cancer, and / or reduce the number of cancer cells present in the mammal.

[0090] In certain embodiments, the cancer treatment can include immune checkpoint inhibitors. Non-limiting examples of immune checkpoint inhibitors include nivolumab (Opdivo), pembrolizumab (Keytruda), atezolizumab (tecentriq), avelumab (bavencio), durvalumab (imfinzi), ipilimumab (yervoy).

[0091] Cancer treatment typically also includes a combination therapy of both multiple chemistries and radiation. Combination chemotherapy includes, for example, cisplatin (CDDP), carboplatin, procarbazine, mechlorethamine, cyclophosphamide, camptothecin, ifosfamide, melphalan, chlorambucil, busulfan, nitrosurea, dactinomycin, daunorubicin, doxorubicin, bleomycin, plicamycin, mitomycin, etoposide (VP16), tamoxifen, raloxifene, estrogen receptor binders, taxol, gemcitabine, navelbine, farnesyl-protein transferase inhibitors, transplatin, 5-fluorouracil, vincristine, vinblastine, and methotrexate, temozolomide (aqueous form of DTIC), or any analogs or derivative variants of the foregoing. A combination of chemotherapy and biotherapy is called chemobiotherapy. Chemotherapy can also be administered at low continuous doses, which is called metronomic chemotherapy.

[0092] Other combination chemotherapy regimens include, for example, alkylating agents such as thiotepa and cyclophosphamide; alkyl sulfonates such as busulfan, improsulfan and piposulfan; aziridines such as benzodepa, carboquone, meturedepa and uredepa; ethyleneimines and methylamelamines including altretamine, triethylenemelamine, triethylenephosphoramide, triethylenethiophosphoramide and trimethylolomelamine; acetogenins (especially bullatacin and bullatacinone); camptothecin (including the synthetic analogue topotecan); bryostatin; callystatin; CC-1065 (including its synthetic analogues adozelesin, carzelesin and bizelesin); cryptophycins (especially cryptophycin 1 and cryptophycin 8); dolastatin; duocarmycin (including the synthetic analogues KW-2189 and CB1-TM1); eleutherobin; pancratistatin; sarcodictyin; spongistatin; nitrogen mustards such as chlorambucil, chlornaphazine, cholophosphamide, estramustine, ifosfamide, mechlorethamine, mechlorethamine oxide hydrochloride, melphalan, novembichin, phenesterine, prednimustine, trofosfamide, uracil mustard; nitrosureas such as carmustine, chlorozotocin, fotemustine, lomustine, nimustine and ranimustine; antibiotics such as the enediyne antibiotics (e.g., calicheamicin, especially calicheamicin γ1I and calicheamicin ω1I; dynemicin, including dynemicin A; bisphosphonates, such as clodronate; esperamicin; as well as neocarzinostatin chromophore and related chromoproteins enediyne antibiotic chromophores, aclacinomysins, actinomycin, authramycin, azaserine, bleomycin, cactinomycin, carabicin, carminomycin, carzinophilin, chromomycin, dactinomycin, daunorubicin, detorubicin, 6-diazo-5-oxo-L-norleucine, doxorubicin (including morpholino-doxorubicin, cyanomorpholino-doxorubicin, 2-pyrrolino-doxorubicin and deoxydoxorubicin), epirubicin, esorubicin, idarubicin, marcellomycin, mitomycins such as mitomycin C, mycophenolic acid, nogalamycin, olivomycin, peplomycin, potfiromycin, puromycin, quelamycin, rodorubicin, streptozocin, tubercidin, ubenimex, zinostatin, zorubicin; antimetabolites such as methotrexate and 5-fluorouracil (5-FU); folic acid analogues such as decihydrofolic acid, pteropterin, trimetrexate; purine analogues such as fludarabine, 6-mercaptopurine, thiamiprine, thioguanine; pyrimidine analogues such as ancitabine, azacitidine, 6-azauridine, carmofur, cytarabine, dideoxyuridine, doxifluridine, enocitabine, floxuridine; androgens such as calusterone, dromostanolone propionate, epitiostanol, mepitiostane, testolactone;Anti-adrenal drugs such as mitotane, trilostane; folic acid supplements such as folinic acid; aceglucuronosone; aldophosphamide glycoside; aminolevulinic acid; eniluracil; amsacrine; bestrabucil; bisantrene; edatrexate; defosfamide; colcemid; diazocone; elformithine; elliptonium acetate; epothilone; etogluconate; gallium nitrate; hydroxyurea; lentinan; lonidamine; maytansine alkaloids such as maytansine and ansamitocin; mitotane Guanidinium; mitoxantrone; mopidanmol; nitraerine; pentostatin; methambucil; pirarubicin; losoxantrone; podophyllic acid; 2-ethylhydrazide; procarbazine; PSK polysaccharide complex; razoxane; rhizoxin; siroxan; spirogermanium; tricholomaric acid; triazoline quinone; 2,2',2"-trichlorotriethylamine; trichothecenes (particularly T-2 toxin, verrucosporin A, baculosporin A, and serpentin) ); urethan; vindesine; dacarbazine; mannomustine; dibromomannitol; dibromodulanol; pipobroman; gacytosine; cytarabine ("Ara-C"); cyclophosphamide; taxanes, e.g., paclitaxel and docetaxel gemcitabine; 6-thioguanine; mercaptopurine; platinum coordination complexes such as cisplatin, oxaliplatin, and carboplatin; vinblastine; platinum; etoposide (VP-16); ifosfamide; mitoxantrone; vincristine; Vinorelbine; Noxolin; Teniposide; Edatrexate; Daunomycin; Aminopterin; Xeloda; Ibandronate; Irinotecan (e.g., CPT-11); Topoisomerase inhibitor RFS2000; Difluoromethylornithine (DMFO); Retinoids such as retinoic acid; Capecitabine; Carboplatin, procarbazine, plicamycin, gemcitabine, Navelbine, farnesyl-protein transferase inhibitors, transplatinum; and any pharmaceutically acceptable salts, acids or derivatives thereof. ;

[0093] Immunotherapeutic agents usually rely on the use of immune effector cells and molecules to target and destroy cancer cells. Immune effectors can be antibodies specific to some markers on the surface of tumor cells, for example. The antibody itself can serve as the effector of treatment, or it can recruit other cells to actually achieve cell killing. Antibodies can also be conjugated to drugs or toxins (chemotherapeutic agents, radionuclides, ricin A chains, cholera toxin, pertussis toxin, etc.) and are only used as targeting agents. Alternatively, the effector can be a lymphocyte carrying a surface molecule that interacts directly or indirectly with a tumor cell target. Various effector cells include cytotoxic T cells and NK cells, as well as variants of these cell types modified to express chimeric antigen receptors through genetic engineering.

[0094] Immunotherapy can include inhibiting regulatory T cells (Tregs), myeloid-derived suppressor cells (MDSCs), and cancer-associated fibroblasts (CAFs). In certain embodiments, the immunotherapy is a tumor vaccine (e.g., whole tumor cell vaccine, peptide, and recombinant tumor-associated antigen vaccine) or adoptive cell therapy (ACT) (e.g., T cells, natural killer cells, TIL, and LAK cells). T cells can be engineered to target specific tumor antigens via a chimeric antigen receptor (CAR) or a T cell receptor (TCR). The chimeric antigen receptor (or CAR) used herein can refer to any engineered receptor that is specific for a target antigen, which, when expressed in a T cell, confers the specificity of the CAR to the T cell. Once created using standard molecular techniques, T cells expressing the chimeric antigen receptor can be introduced into a patient, just as in techniques such as adoptive cell transfer. In certain aspects, the T cells are activated CD4 and / or CD8 T cells in an individual, characterized by CD4 and / or CD8 T cells that produce γ-IFN and / or enhanced cytolytic activity relative to before administration of the combination. The CD4 and / or CD8 T cells can exhibit increased release of cytokines selected from the group consisting of IFN-γ, TNF-α, and interleukins. The CD4 and / or CD8 T cells can be effector memory T cells. In certain embodiments, the CD4 and / or CD8 effector memory T cells are characterized by having CD44 高 CD62L 低 expression.

[0095] The immunotherapy can be a cancer vaccine that includes one or more cancer antigens, particularly a protein or an immunogenic fragment thereof, DNA or RNA encoding the cancer antigen, particularly a protein or an immunogenic fragment thereof, cancer cell lysates, and / or a protein preparation from tumor cells. The cancer antigens used herein are antigenic substances present in cancer cells. In principle, any protein with an abnormal structure due to mutation produced in cancer cells can serve as a cancer antigen. In principle, cancer antigens can be products of mutated oncogenes and tumor suppressor genes, products of other mutated genes, overexpressed or abnormally expressed cellular proteins, cancer antigens produced by oncogenic viruses, carcinoembryonic antigens, altered cell surface glycolipids and glycoproteins, or cell type-specific differentiation antigens. Examples of cancer antigens include abnormal products of the ras and p53 genes. Other examples include tissue differentiation antigens, mutant protein antigens, oncogenic virus antigens, cancer-testis antigens, and vascular or stromal-specific antigens. Tissue differentiation antigens are antigens specific for a certain type of tissue.

[0096] The immunotherapy can be an antibody, such as part of a polyclonal antibody preparation, or can be a monoclonal antibody. The antibody can be a humanized antibody, a chimeric antibody, an antibody fragment, a bispecific antibody or a single-chain antibody. The antibodies disclosed herein include antibody fragments, such as, but not limited to, Fab, Fab', and F(ab')2, Fd, single-chain Fv (scFv), single-chain antibodies, disulfide-linked Fvs (sdfv), and fragments comprising VL or VH domains. In certain aspects, the antibody or fragment thereof specifically binds to epidermal growth factor receptor (EGFR1, Erb-B1), HER2 / neu (Erb-B2), CD20, vascular endothelial growth factor (VEGF), insulin-like growth factor receptor (IGF-1R), TRAIL-receptor, epithelial cell adhesion molecule, carcinoembryonic antigen, prostate-specific membrane antigen, mucin-1, CD30, CD33, or CD40.

[0097] Examples of monoclonal antibodies include, but are not limited to, trastuzumab (anti-HER2 / neu antibody); pertuzumab (anti-HER2 mAb); cetuximab (chimeric monoclonal antibody against epidermal growth factor receptor EGFR); panitumumab (anti-EGFR antibody); nimotuzumab (anti-EGFR antibody); zalutumumab (anti-EGFR mAb); necitumumab (anti-EGFR mAb); MDX-210 (humanized anti-HER-2 bispecific antibody); MDX-210 (humanized anti-HER-2 bispecific antibody); MDX-447 (humanized anti-EGF receptor bispecific antibody); rituximab (chimeric mouse / human anti-CD20 mAb); ofatumumab (anti-CD20 mAb); obinutuzumab (anti-CD20 mAb); tositumomab-I131 (anti-CD20 mAb); ibritumomab (anti-CD20 mAb); bevacizumab (anti-VEGF mAb); ramucirumab (anti-VEGFR2 mAb); ranibizumab (anti-VEGF mAb); aflibercept (extracellular domains of VEGFR1 and VEGFR2 fused to IgG1 Fc); AMG 386 (angiopoietin-1 and -2 binding peptide fused to IgG1 Fc); daratumumab (anti-IGF-1R mAb); gemtuzumab ozogamicin (anti-CD33 mAb); alemtuzumab (anti-Campath-1 / CD52 mAb); brentuximab vedotin (anti-CD30 mAb); catumaxomab (bispecific mAb targeting epithelial cell adhesion molecule and CD3); naprotumab (anti-5T4 mAb); girentuximab (anti-carbonic anhydrase ix); or farletuzumab (anti-folate receptor).Other examples include antibodies such as Panorex.TM. (17-1A) (murine monoclonal antibody); Panorex (MAb17-1A) (chimeric murine monoclonal antibody); BEC2 (anti-idiotypic mAb, mimicking GD epitope) (with BCG); Oncolym (Lym-1 monoclonal antibody); SMART M195 Ab, humanized 13'1LYM-1 (Oncolym), Ovarex (B43.13, anti-idiotypic murine mAb); 3622W94 mAb that binds to the EGP40 (17-1A) pan-carcinoma antigen on adenocarcinoma; Zenapax (SMART anti-Tac (IL-2 receptor); SMART M195 Ab, humanized Ab, humanized); NovoMAb-G2 (pan-carcinoma specific Ab); TNT (chimeric mAb against histone antigen); TNT (chimeric mAb against histone antigen); Gliomab-H (monoclonal - humanized Ab); GNI-250 Mab; EMD-72000 (chimeric - EGF antagonist); LymphoCide (humanized IL.L.2 antibody); and MDX-260 bispecific, targeting GD-2, ANA Ab, SMART IDIO Ab, SMART ABL 364Ab or ImmuRAIT-CEA.Other examples of antibodies include Zanulimumab (anti-CD4 mAb), keliximab (anti-CD4 mAb); ipilimumab (MDX-101; anti-CTLA-4 mAb); tremelimumab (anti-CTLA-4 mAb); daclizumab (anti-CD25 / IL-2R mAb); basiliximab (anti-CD25 / IL-2R mAb); MDX-1106 (anti-PD1 mAb); antibodies against GITR; GC1008 (anti-TGF-β antibody); metelimumab / CAT-192 (anti-TGF-β antibody); lededumab / CAT-152 (anti-TGF-β antibody); ID11 (anti-TGF-β antibody); denosumab (anti-RANKL mAb); BMS-663513 (humanized anti-4-1BB mAb); SGN-40 (humanized anti-CD40 mAb); CP870,893 (human anti-CD40 mAb); infliximab (chimeric anti-TNF mAb); adalimumab (human anti-TNF mAb); certolizumab (humanized Fab anti-TNF); golimumab (anti-TNF); etanercept (extracellular domain of TNFR fused to IgG1 Fc); belatacept (extracellular domain of CTLA-4 fused to Fc); abatacept (extracellular domain of CTLA-4 fused to Fc); belimumab (anti-B lymphocyte stimulator); muromonab-CD3 (anti-CD3 mAb); ocrelizumab (anti-CD3 mAb); tilizumab (anti-CD3 mAb); tocilizumab (anti-IL6R mAb); REGN88 (anti-IL6R mAb); ustekinumab (anti-IL-12 / 23 mAb); briakinumab (anti-IL-12 / 23 mAb); natalizumab (anti-α4 integrin); vedolizumab (anti-α4β7 integrin mAb); T1 h (anti-CD6 mAb); epratuzumab (anti-CD22 mAb); efalizumab (anti-CD11a mAb); and atacicept (extracellular domain of transmembrane activator and calcium modulator ligand interactor fused to Fc).

[0098] When monitoring a mammal having or suspected of having a cancer described herein (e.g., at least in part based on the cfDNA fragmentation profile of the mammal), the monitoring can be performed before, during, and / or after a cancer treatment process. The monitoring methods provided herein can be used to determine the efficacy of one or more cancer treatments and / or to select a mammal for enhanced monitoring. In certain embodiments, the monitoring can include identifying a cfDNA fragmentation profile as described herein. For example, a cfDNA fragmentation profile can be obtained before administering one or more cancer treatments to a mammal having or suspected of having a cancer, one or more cancer treatments can be administered to the mammal, and one or more cfDNA fragmentation profiles can be obtained during the cancer treatment process. In certain embodiments, the cfDNA fragmentation profile can change during a cancer treatment (e.g., any cancer treatment described herein). For example, a cfDNA fragmentation profile indicative of a mammal having a cancer can change to a cfDNA fragmentation profile indicative of a mammal not having a cancer. Such a change in the cfDNA fragmentation profile can indicate that the cancer treatment is working. Conversely, the cfDNA fragmentation profile can remain static (e.g., the same or substantially the same) during a cancer treatment (e.g., any cancer treatment described herein). Such a static cfDNA fragmentation profile can indicate that the cancer treatment is not working.

[0099] In certain embodiments, the monitoring can include conventional techniques capable of monitoring one or more cancer treatments (e.g., the efficacy of one or more cancer treatments). In certain embodiments, a mammal selected for enhanced monitoring can be administered a diagnostic test (e.g., any diagnostic test disclosed herein) at an increased frequency compared to a mammal not selected for enhanced monitoring. For example, a diagnostic test can be administered to a mammal selected for enhanced monitoring at a frequency of twice daily, daily, every two weeks, weekly, every two months, monthly, quarterly, semi-annually, annually, or any frequency therebetween. In certain embodiments, a mammal selected for enhanced monitoring can be administered one or more additional diagnostic tests compared to a mammal not selected for enhanced monitoring. For example, two diagnostic tests can be performed on a mammal selected for enhanced monitoring, while only one diagnostic test (or no diagnostic test) is performed on a mammal not selected for enhanced monitoring. In certain embodiments, a mammal selected for enhanced monitoring can also be selected for further diagnostic testing. Once the presence of a tumor or cancer (e.g., cancer cells) has been identified (e.g., by any of the various methods disclosed herein), it can be beneficial for the mammal to undergo enhanced monitoring (e.g., to assess the progression of the tumor or cancer in the mammal and / or to evaluate the development of one or more cancer biomarkers such as mutations) and further diagnostic testing (e.g., to determine the size and / or exact location of the tumor or cancer (e.g., the tissue of origin)). In certain embodiments, one or more cancer treatments can be administered to a mammal selected for enhanced monitoring after detection of a cancer biomarker and / or after the cfDNA fragmentation profile of the mammal has not improved or has deteriorated. Any cancer treatment disclosed herein or known in the art can be administered. For example, a mammal selected for enhanced monitoring can be further monitored, and if cancer cells are present throughout the enhanced monitoring period, cancer treatment can be performed. Additionally, or alternatively, a mammal selected for enhanced monitoring can be treated for cancer and further monitored as the cancer treatment progresses. In certain embodiments, enhanced monitoring will reveal one or more cancer biomarkers (e.g., mutations) after cancer treatment of a mammal selected for enhanced monitoring. In certain embodiments, such one or more cancer biomarkers will provide a reason to administer a different cancer treatment (e.g., resistant mutations may arise in cancer cells during cancer treatment, and cancer cells carrying such resistant mutations are resistant to the initial cancer treatment).

[0100] When a mammal is identified as having cancer as described herein (e.g., at least in part based on the cfDNA fragmentation profile of the mammal), the identification can be performed before and / or during cancer treatment. The methods provided herein for identifying that a mammal has cancer can be used as a first diagnosis to identify a mammal (e.g., having cancer prior to any course of treatment) and / or to select a mammal for further diagnostic testing. In certain embodiments, once it is determined that a mammal has cancer, further testing can be performed on the mammal and / or the mammal can be selected for further diagnostic testing. In certain embodiments, the methods provided herein can be used to select a mammal for further diagnostic testing during a time period prior to the time period when conventional techniques are able to diagnose that the mammal has early-stage cancer. For example, the methods provided herein for selecting a mammal for further diagnostic testing can be used when a mammal has not been diagnosed with cancer by conventional methods and / or when it is not known that the mammal carries cancer. In certain embodiments, a mammal selected for further diagnostic testing can be administered a diagnostic test (e.g., any diagnostic test disclosed herein) at an increased frequency compared to a mammal not selected for further diagnostic testing. For example, a mammal selected for further diagnostic testing can be administered a diagnostic test twice daily, daily, bi-weekly, weekly, bi-monthly, monthly, quarterly, semi-annually, annually, or at any frequency therebetween. In certain embodiments, a mammal selected for further diagnostic testing can be administered one or more additional diagnostic tests compared to a mammal not selected for further diagnostic testing. For example, a mammal selected for further diagnostic testing can be administered two diagnostic tests, while a mammal not selected for further diagnostic testing is administered only one diagnostic test (or no diagnostic test). In certain embodiments, the diagnostic test method can determine the presence of cancer of the same type as the initially detected cancer (e.g., having the same tissue or origin) (e.g., at least in part based on the cfDNA fragmentation profile of the mammal). Additionally, or alternatively, the diagnostic test method can determine the presence of cancer of a different type than the initially detected cancer. In certain embodiments, the diagnostic test method is a scan. In certain embodiments, the scan is computed tomography (CT), CT angiography (CTA), esophagography (barium swallow), barium enema, magnetic resonance imaging (MRI), PET scan, ultrasound (e.g., endobronchial ultrasound, endoscopic ultrasound), X-ray, DEXA scan.

[0101] In certain embodiments, the diagnostic test method is a physical examination, such as anoscopy, bronchoscopy (e.g., autofluorescence bronchoscopy, white light bronchoscopy, navigational bronchoscopy), colonoscopy, digital breast tomosynthesis, endoscopic retrograde cholangiopancreatography (ERCP), esophagogastroduodenoscopy, mammography, Pap smear, pelvic examination, positron emission tomography and computed tomography (PET-CT) scan. In certain embodiments, a mammal that has been selected for further diagnostic testing may also be selected for enhanced monitoring. Once the presence of a tumor or cancer (e.g., cancer cells) has been identified (e.g., by any of the various methods disclosed herein), it may be beneficial for the mammal to undergo enhanced monitoring (e.g., to assess the progression of the tumor or cancer in the mammal and / or to assess the development of one or more cancer biomarkers such as mutations) and further diagnostic testing (e.g., to determine the size and / or exact location of the tumor or cancer). In certain embodiments, a cancer treatment is administered to a mammal selected for further diagnostic testing after a cancer biomarker has been detected and / or after the cfDNA fragmentation profile of the mammal has not improved or has deteriorated. Any cancer treatment disclosed herein or known in the art may be administered. For example, a further diagnostic test may be administered to a mammal selected for further diagnostic testing, and if the presence of a tumor or cancer is confirmed, a cancer treatment is administered. Additionally, or alternatively, a cancer treatment may be administered to a mammal selected for further diagnostic testing, and the mammal may be further monitored as the cancer treatment progresses. In certain embodiments, additional testing will reveal one or more cancer biomarkers after a cancer treatment has been administered to a mammal selected for further diagnostic testing. In certain embodiments, such one or more cancer biomarkers (e.g., mutations) will provide a reason to administer a different cancer treatment (e.g., resistant mutations may occur in cancer cells during cancer treatment, where cancer cells carrying the resistant mutation are resistant to the initial cancer treatment).

[0102] System

[0103] In some embodiments, the present disclosure provides systems, methods, or kits that may include data analysis implemented in a measurement device (e.g., a laboratory instrument such as a sequencer), and software code executed on computing hardware. The software may be stored in a memory and executed on one or more hardware processors. The software may be organized into routines or packages that can communicate with each other. A module may include one or more devices / computers, and potentially one or more software routines / packages executed on the one or more devices / computers. For example, an analysis application or system may at least include a data receiving module, a data preprocessing module, a data analysis module (which may operate on one or more types of genomic data), a data interpretation module, or a data visualization module.

[0104] The data receiving module may connect laboratory hardware or instruments to a computer system that processes laboratory data. The data preprocessing module may operate on data prepared for analysis. Examples of operations that may be applied to the data in the preprocessing module include affine transformation, denoising operations, data cleaning, reformatting, or subsampling. The data analysis module (which may be specialized for analyzing genomic data from one or more genomic materials) may, for example, obtain an assembled genomic sequence and perform probabilistic and statistical analyses to identify abnormal patterns associated with a disease, pathology, state, risk, condition, or phenotype. The data interpretation module may use analytical methods, for example, taken from statistics, mathematics, or biology, to support understanding the relationship between the identified abnormal patterns and a health condition, functional state, prognosis, or risk. The data analysis module and / or the data interpretation module may include one or more machine learning models, which may be implemented in hardware, for example, the hardware executing software embodying the machine learning model. The data visualization module may use methods of mathematical modeling, computer graphics, or rendering to create a visual representation of the data, which may facilitate understanding or interpretation of the results. The present disclosure provides computer systems programmed to implement the methods of the present disclosure.

[0105] In certain embodiments, the methods disclosed herein can include computational analysis of nucleic acid sequencing data from samples of one or more individuals. The analysis can identify variants inferred from the sequence data to identify sequence variants based on probabilistic modeling, statistical modeling, mechanistic modeling, network modeling, or statistical inference. Non-limiting examples of analysis methods include principal component analysis, autoencoders, singular value decomposition, Fourier bases, wavelets, discriminant analysis, regression, support vector machines, tree-based methods, networks, matrix factorization, and clustering. Non-limiting examples of variants include germline variants or somatic mutations. In certain embodiments, the variants can represent known variants. Known variants can be scientifically validated or reported in the literature. In certain embodiments, the variants can represent putative variants associated with biological changes. The biological changes can be known or unknown. In certain embodiments, putative variants can be reported in the literature but have not been biologically validated. Alternatively, putative variants have never been reported in the literature but can be inferred based on the computational analysis disclosed herein. In certain embodiments, germline variants can represent nucleic acids that induce natural or normal variation.

[0106] In certain embodiments, the computer system includes a central processing unit (CPU, also referred to herein as “processor” and “computer processor”), which can be a single-core or multi-core processor, or multiple processors for parallel processing; a memory (e.g., cache, random access memory, read-only memory, flash memory, or other memory); an electronic storage unit (e.g., hard disk), a communication interface for communicating with one or more other systems (e.g., network adapter); and peripheral devices, such as adapters for caching, other memories, data storage, and / or electronic displays. The memory, storage unit, interface, and peripheral devices can communicate with the CPU via a communication bus (solid lines) (such as a motherboard). The storage unit can be a data storage unit (or data repository) for storing data. One or more analyte feature inputs can be input from one or more measurement devices. Example analytes and measurement devices are described herein.

[0107] A computer system can be operatively coupled to a computer network ("network") via a communication interface. The network can be the Internet, an intranet and / or an extranet, or an intranet and / or an extranet that communicates with the Internet. In some cases, the network is a telecommunications and / or data network. The network can include one or more computer servers that can implement distributed computing, such as cloud computing via the network ("cloud"), to perform various aspects of the analysis, calculation, and generation of the present disclosure, such as, for example, activating a valve or pump to transfer a reagent or sample from one chamber to another or applying heat to a sample (e.g., during an amplification reaction), processing and / or assaying other aspects of the sample, performing sequencing analysis, measuring a set of values representative of molecular classes, identifying a set of features and feature vectors from the assay data, processing the feature vectors using a machine learning model to obtain an output classification, and training a machine learning model (e.g., iteratively searching for the optimal values of the parameters of the machine learning model). Such cloud computing can be provided by a cloud computing platform, such as, for example, Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform, and IBM Cloud. In some cases, via the computer system, the network can implement a peer-to-peer network, which can enable the devices coupled to the computer system to act as clients or servers.

[0108] The CPU can execute a series of machine-readable instructions, which can be embodied in a program or software. The instructions can be stored in a memory location, such as a memory. The instructions can be directed to the CPU, which can then be programmed or otherwise configured to implement the methods of the present disclosure. The CPU can be part of a circuit, such as an integrated circuit. One or more other components of the system can be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).

[0109] The storage unit can store files, such as drivers, libraries, and saved programs. The storage unit can store user data, such as user preferences and user programs. In some cases, the computer system can include one or more additional data storage units located outside the computer system, such as on a remote server that communicates with the computer system via an intranet or the Internet.

[0110] The computer system can communicate with one or more remote computer systems via the network. For example, the computer system can communicate with the user's remote computer system. Examples of remote computer systems include personal computers (e.g., portable PCs), tablets (e.g., iPad, Galaxy Tab), phones, smartphones (e.g., iPhone, Android-supported devices, ) or a personal digital assistant. The user can access the computer system through the network.

[0111] The methods described herein can be implemented by machine (e.g., computer processor) executable code stored on an electronic storage location of a computer system (such as, for example, a memory or an electronic storage unit). The machine executable or machine readable code can be provided in the form of software. During use, the code can be executed by a CPU. In some cases, the code can be retrieved from the storage unit and stored in the memory for convenient access by the CPU. In some cases, instead of using an electronic storage unit, the machine executable instructions can be stored in the memory.

[0112] The code can be pre-compiled and configured to be used with a machine having a processor suitable for executing the code, or can be compiled at runtime. The code can be provided in a programming language, and the programming language can be selected to enable the code to be executed in a pre-compiled or compiled manner.

[0113] Aspects of the systems and methods provided herein (such as computer systems) can be embodied in programming. Aspects of the technology can be considered a “product” or “manufactured article” typically in the form of machine (or processor) executable code and / or associated data carried or embodied on a type of machine readable medium. The machine executable code can be stored on an electronic storage unit such as a memory (e.g., read only memory, random access memory, flash memory) or a hard disk. “Storage” type media can include any or all tangible memories or their associated modules such as various semiconductor memories, tape drives, disk drives, and the like of a computer, a processor, or the like, which can provide non-transitory storage for software programming at any time. All or part of the software can sometimes be communicated via the Internet or various other telecommunications networks. For example, such communication can enable the software to be loaded from one computer or processor to another, e.g., from a management server or a host computer to a computer platform of an application server. Thus, another type of media that can carry software elements includes light waves, radio waves, and electromagnetic waves, such as used via wired and optical transmission line networks between local devices and across physical interfaces via various air links. Physical elements carrying such waves, such as wired or wireless links, fiber optic links, or the like, can also be considered media carrying software. As used herein, unless limited to non-transitory, tangible “storage” media, terms such as computer or machine “readable media” represent any media that participates in providing instructions to a processor for execution.

[0114] Thus, machine-readable media (such as computer-executable code) can take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media includes, for example, optical or magnetic disks, such as any storage device in any computer or the like, such as may be used to implement databases shown in the figures. Volatile storage media includes dynamic memory, such as the main memory of such a computer platform. Tangible transmission media includes coaxial cables; copper wire and fiber optics, including the wires that make up a bus within a computer system.

[0115] Carrier transmission media can take the form of electrical or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example: floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROM, DVD or DVD-ROM, any other optical media, punched cards, paper tape, any other physical storage media with a pattern of holes, RAM, ROM, PROM, and EPROM, FLASH-EPROM, any other storage chip or cartridge, a carrier wave that transports data or instructions, a cable or link that transports such a carrier wave, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer-readable media can participate in transporting one or more sequences of one or more instructions to a processor for execution.

[0116] The computer system can include or communicate with an electronic display that includes a user interface (UI) for providing, for example, the current stage of sample processing or assay (e.g., a particular step, such as a lysis step, or an ongoing sequencing step). The computer system receives input from one or more measurements. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces. For example, the algorithm can process and / or assay a sample, perform sequencing analysis, measure a set of values representative of molecular classes, identify a set of features and feature vectors from the assay data, process the feature vectors using a machine learning model to obtain an output classification, and train a machine learning model (e.g., iteratively search for the optimal values of the parameters of the machine learning model).

[0117] In certain embodiments, a system (e.g., a laptop, desktop, iPad, mobile device, etc.) capable of executing one or more algorithms (for determining cfDNA fragmentation profile variations) classifies the subject as a cancer patient based on the subject's cfDNA fragmentation profile. These systems further execute machine learning algorithms, which can be used to generate models, such as, for example, high-risk groups and low-risk general groups (penalized logistic regression, using the features of Mathios et al. (Mathios D, Johansen JS, Cristiano S, Medina JE, Phallen J, Larsen KR, et al. Detection and characterization of lung cancer using cell-free DNA fragmentomes. Nat Commun 2021;12(1):5060) and the coverage of transcription factor binding sites. These models can be trained using 5-fold cross-validation with 10 repeats on a cohort of subjects, and the score for each sample is calculated as the mean of the repeats and evaluated using AUC-ROC. For example, the first model uses high-risk non-cancer and HCC patients, while the second model uses non-cancer individuals without liver pathology. The locked high-risk model trained on this cohort is applied to a second and different cohort to generate cancer predictions on an external validation set. A "class label" can be applied to each sample, which indicates the classification of the sample for any number of input features. For example, the class labels for a set of cohorts can indicate the identity of the cfDNA fragmentation profile based on genomic location, etc. The resulting training set is provided to a machine learning unit, such as a neural network or a support vector machine. Using the training set, the machine learning unit can generate a model to classify samples based on the cfDNA fragmentation profile.

[0118] In certain embodiments, a method of creating a trained classifier is provided, comprising the steps of: (a) providing a plurality of different classes, where each class represents a group of subjects (e.g., from one or more cohorts) having common characteristics; (b) providing a multi-parameter model representing cell-free DNA molecules from each of a plurality of samples belonging to each class, thereby providing a training data set; and (c) training a learning algorithm on the training data set to create one or more trained classifiers, where each trained classifier classifies a test sample into one or more of the plurality of classes.

[0119] As an example, the trained classifier can use a learning algorithm selected from the group consisting of random forest, neural network, support vector machine, and linear classifier. Each of the plurality of different classes can be selected from the group consisting of: healthy, breast cancer, colon cancer, lung cancer, pancreatic cancer, prostate cancer, ovarian cancer, melanoma, and liver cancer.

[0120] The trained classifier can be applied to a method for classifying a sample from a subject. Such a classification method can include: (a) providing a multi-parameter model representative of cell-free DNA molecules in a test sample from a subject; and (b) using the trained classifier to classify the test sample. After classifying the test sample into one or more categories, a treatment intervention can be performed on the subject based on the classification of the sample.

[0121] In certain embodiments, a training set is provided to a machine learning unit, such as a neural network or a support vector machine. Using the training set, the machine learning unit can generate a model that classifies samples according to the treatment response to one or more treatment inventions. This is also referred to as "calling". The developed model can utilize information from any part of the test vector.

[0122] Generally, machine learning can be used to reduce a data set generated from all (primary sample / analyte / test) combinations into a set of optimal feature prediction sets, e.g., which meet specified criteria. In various embodiments, statistical learning and / or regression analysis can be applied. Models ranging from simple to complex and from small to large that make various modeling assumptions can be applied to data in a cross-validation paradigm. From simple to complex includes considerations of the features from linear to non-linear and from non-hierarchical to hierarchical representations. Small to large models include considerations of the size of the basis vector space to which the data is mapped and the number of interactions between the features included in the modeling process.

[0123] Machine learning techniques can be used to evaluate the best commercial assay patterns for a cost / performance / business scope defined in an initial problem. Threshold checks can be performed: if the method applied to a held-out dataset not used in cross-validation exceeds an initialized constraint, the assay is locked and production begins. For example, thresholds for assay performance can include a desired minimum accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), or a combination thereof. For example, the desired minimum accuracy, PPV, NPV, clinical sensitivity, clinical specificity, or a combination thereof can be at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. As another example, the desired minimum AUC can be at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.81, at least about 0.82, at least about 0.83, at least about 0.84, at least about 0.85, at least about 0.86, at least about 0.87, at least about 0.88, at least about 0.89, at least about 0.90, at least about 0.91, at least about 0.92, at least about 0.93, at least about 0.94, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99. Based on the total cost of performing a subset of assays, the subset of assays can be selected from a set of assays to be performed on a given sample, while adhering to thresholds for assay performance, such as the desired minimum accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), and combinations thereof. If the thresholds are not met, the assay engineering process can loop back to the constraint setting for possible relaxation, or loop back to the wet lab to change the parameters for obtaining data. Given the clinical problem, biological constraints, budget, laboratory machines, etc. can limit the problem.

[0124] In certain embodiments, computer processing of machine learning techniques can include statistical, mathematical, biological methods, or any combination thereof. In various embodiments, any of the computer processing methods can include dimensionality reduction methods, logistic regression, dimensionality reduction, principal component analysis, autoencoders, singular value decomposition, Fourier basis, singular value decomposition, wavelets, discriminant analysis, support vector machines, tree-based methods, random forests, gradient boosting trees, logistic regression, matrix factorization, network clustering, statistical tests, and neural networks.

[0125] In certain embodiments, computer processing of machine learning techniques can include logistic regression, multiple linear regression (MLR), dimensionality reduction, partial least squares (PLS) regression, principal component regression, autoencoders, variational autoencoders, singular value decomposition, Fourier basis, wavelets, discriminant analysis, support vector machines, decision trees, classification and regression trees (CART), tree-based methods, random forests, gradient boosting trees, logistic regression, matrix factorization, multidimensional scaling (MDS), dimensionality reduction methods, t-distributed stochastic neighbor embedding (t-SNE), multi-layer perceptrons (MLP), network clustering, neuro-fuzzy, neural networks (shallow and deep), artificial neural networks, Pearson product-moment correlation coefficient, Spearman's rank correlation coefficient, Kendall tau rank correlation coefficient, or any combination thereof. In certain embodiments, the computer processing method is a supervised machine learning method, including, for example, regression, support vector machines, tree-based methods, and neural networks. In certain embodiments, the computer processing method is an unsupervised machine learning method, including, for example, clustering, network, principal component analysis, and matrix factorization.

[0126] For supervised learning, training samples (e.g., thousands) can include measured data (e.g., measurement data of various analytes) and known labels, which can be determined by other time-consuming processes, such as imaging a subject and analyzing by a trained practitioner. Example labels can include classification of the subject, e.g., a discrete classification of whether the subject has cancer or a continuous classification providing a probability (e.g., risk or score) of a discrete value. The learning module can optimize the parameters of the model to achieve a quality metric (e.g., prediction accuracy for known labels) through one or more specified criteria. The quality metric can be implemented for any function (including the set of all risk, loss, utility, and decision functions). Gradients can be used in conjunction with the learning step (e.g., a measure of how much the model parameters should be updated for a given time step of the optimization process).

[0127] As described above, embodiments can be used for a variety of purposes. For example, plasma (or other samples) can be collected from subjects presenting symptoms of a disorder (e.g., known to have the disorder) and healthy subjects. Genetic data (e.g., cfDNA) can be obtained and analyzed to derive a variety of different features, which can include features based on whole genome analysis. These features can form a feature space that is searched, stretched, rotated, translated, and linearly or non-linearly transformed to generate an accurate machine learning model that can distinguish between healthy subjects and subjects having the disorder (e.g., identify the diseased or non-diseased state of a subject). Outputs derived from this data and model (which may include probabilities of a disorder, stages (levels) of a disorder, or other values) can be used to generate another model that can be used to recommend further procedures, e.g., recommend a biopsy or continued monitoring of the subject's condition.

[0128] In certain embodiments, DNA from a population of multiple individuals can be analyzed by a set of multiplexed arrays. Data for each multiplexed array can be self-normalized using information contained in that particular array. The normalization algorithm can adjust for nominal intensity variations observed in the two-color channels, background differences between channels, and possible crosstalk between dyes. A clustering algorithm that incorporates several bio-inspired features on the fragmentation profile can then be used to model the behavior at each base position. In cases where a small number of cfDNA fragments are observed (e.g., due to low minor allele frequency), a neural network can be used to estimate the location and shape of the missing sequences. Based on the profile and the percentage of sequence identity, a statistical score (training score) can be designed. Scores such as the GenCall score are involved in mimicking the evaluations made by the visual and cognitive systems of human experts. In addition, it has evolved using genotyping data from the top and bottom strands. This score can be combined with several penalty terms (e.g., low intensity, mismatches between existing and predicted cfDNA fragments) to form the training score. The training score is saved for use by the calling algorithm.

[0129] To call a treatment response, the calling algorithm can obtain the genetic information and treatment responses of multiple individuals having a disease or disorder. The data can first be normalized (using the same procedure as the clustering algorithm). A calling operation (classification) can be performed using, for example, a Bayesian model. The score for each call (Call score) can be the product of the training score and the data-to-model fit score. After scoring all treatment responses, the application can calculate a composite score.

[0130] In certain embodiments, the training data set comprises clinical data selected from the group consisting of cancer stage, type of surgery, age, tumor grade, depth of tumor invasion, occurrence of postoperative complications, and presence of venous invasion. In certain embodiments, the training data set is pre-processed, including converting the provided data into class conditional probabilities.

[0131] Another embodiment uses machine learning techniques to train a statistical classifier, particularly a support vector machine, for each cancer stage category based on the word occurrences in the histological report corpus of each patient. New reports can then be classified according to the most likely stage, thus facilitating the collection and analysis of population staging data.

[0132] In certain embodiments, the machine learning algorithm is selected from supervised or unsupervised learning algorithms, which are selected from support vector machines, random forests, nearest neighbor analysis, linear regression, binary decision trees, discriminant analysis, logistic classifiers, and clustering analysis.

[0133] Generally, the system may include a report generator for reporting cancer trial results and treatment options. The report generator system may be a central data processing system configured to establish direct communication via a communication link with: remote data sites or laboratories, medical institutions / healthcare providers (treatment professionals), and / or patients / subjects. The laboratory may be a medical laboratory, diagnostic laboratory, healthcare facility, medical institution, point-of-care testing device, or any other remote data site capable of generating clinical information of the subject. Subject clinical information includes, but is not limited to, laboratory test data, X-ray data, examinations, and diagnoses. Healthcare providers or institutions 26 include healthcare service providers such as physicians, nurses, home health aides, technicians, and physician assistants, and the institution is any healthcare facility equipped with healthcare providers. In some cases, the healthcare provider / institution is also a remote data site. In cancer treatment embodiments, the subject may be suffering from cancer, etc.

[0134] Other clinical information of cancer subjects includes the results of laboratory tests, imaging, or medical procedures for specific cancers that can be readily identified by a person of ordinary skill in the art. A suitable list of sources of cancer clinical information includes, but is not limited to: CT scans, MRI scans, ultrasound scans, bone scans, PET scans, bone marrow tests, barium X-rays, endoscopy, lymphangiography, IVU (intravenous urography) or IVP (IV pyelography), lumbar puncture, cystoscopy, immunological tests (anti-oncogenic antibody screening), and cancer marker tests.

[0135] Subject clinical information can be obtained from a laboratory manually or automatically. For the simplicity of the system, information is obtained automatically at a predetermined or fixed time interval. A fixed time interval refers to the time interval for automatically collecting laboratory data based on time measurements (such as hours, days, weeks, months, years, etc.) by the methods and systems described herein. In one embodiment of the present invention, data collection and processing are performed at least once a day. In one embodiment, data transmission and collection are performed monthly, bi-weekly, weekly, or every few days. Alternatively, information retrieval can be performed at a predetermined rather than regular time interval. For example, the first retrieval step can occur after one week, and the second retrieval step can occur after one month. Data transmission and collection can be customized according to the nature of the disorder being managed and the frequency of tests and medical examinations required by the subject.

[0136] In certain embodiments, a genetic report is generated from a subject sample (e.g., cfDNA). The polynucleotides in the sample can be sequenced, e.g., whole genome sequencing, NGS sequencing, to generate multiple sequence reads. In certain embodiments, the genetic information includes variables that define the genomic organization of cancer cells or the genomic organization of a single disseminated cancer cell. In certain embodiments, the genetic information includes sequence or abundance data of one or more genetic loci in cell-free DNA from an individual.

[0137] Process cfDNA genetic information (72). Genetic variations can also be identified. Genetic variations include sequence variations, copy number variations, and nucleotide modification variations. A sequence variation is a variation in the genetic nucleotide sequence. A copy number variation is a deviation in the copy number of a part of the genome from the wild type. Genetic variations include, for example, single nucleotide variations (SNPs), insertions, deletions, inversions, transversions, translocations, gene fusions, chromosome fusions, gene truncations, copy number variations (e.g., aneuploidy, segmental aneuploidy, polyploidy, gene amplification), abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid methylation. Then, the process determines the frequency of genetic variations in a sample containing genetic material. Since the process has noise, the process separates the information from the noise (73). The sensitivity of detecting genetic variations can be increased by increasing the read depth of the polynucleotides (e.g., by sequencing samples from a subject at a greater read depth at two or more time points).

[0138] Multiple measurements can be made to increase diagnostic confidence. Alternatively, measurements at multiple time points (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more time points) can be used to determine whether the cancer is progressing, in remission, or stable. Diagnostic confidence can be used to identify disease states. For example, cell-free polynucleotides obtained from a subject can include polynucleotides from normal cells, as well as polynucleotides from diseased cells (such as cancer cells). Polynucleotides from cancer cells may carry genetic variations, such as somatic mutations and copy number variations. When sequencing cell-free polynucleotides from a sample of a subject, a cfDNA fragmentation profile can be generated as described in the Examples section below.

[0139] Multiple cancers can be detected using the methods and systems described herein. Like most cells, cancer cells can be characterized by a turnover rate, where older cells die and are replaced by newer cells. Generally, dead cells release DNA or DNA fragments into the blood when in contact with the vasculature of a given subject. This is also true for cancer cells at all stages of the disease. Depending on the stage of the disease, cancer cells can also be characterized by various genetic aberrations, such as copy number variations and mutations. This phenomenon can be used to detect the presence or absence of an individual with cancer using the methods and systems described herein.

[0140] In the early detection of cancer, any of the systems or methods described herein (including mutation detection or copy number variation detection) can be used to detect cancer. These systems and methods can be used to detect any number of genetic aberrations that may cause or be caused by cancer. These can include, but are not limited to, cfDNA fragmentation profiles, mutations, insertions, deletions, copy number variations, transversions, translocations, inversions, deletions, aneuploidy, segmental aneuploidy, polyploidy, chromosomal instability, chromosomal structural alterations, gene fusions, chromosomal fusions, gene truncations, gene amplifications, gene duplications, chromosomal damage, DNA damage, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, nucleic acid methylation infections, and cancer abnormalities.

[0141] In addition, the systems and methods described herein can also be used to help characterize certain cancers. The genetic data generated by the systems and methods of the present disclosure can allow practitioners to better characterize specific forms of cancer. Often times, cancers are heterogeneous in composition and stage. Gene profiling data can allow for the characterization of specific subtypes of cancer, which may be important for the diagnosis or treatment of that specific subtype. This information can also provide clues to the subject or practitioner about the prognosis of a specific type of cancer.

[0142] The systems and methods provided herein can be used to monitor known cancers or other diseases in a particular subject. This can allow the subject or practitioner to adjust treatment options based on the progression of the disease. In this embodiment, the systems and methods described herein can be used to construct a genetic cfDNA fragmentation profile of a particular subject during the course of a disease. In some cases, cancer can progress and become more aggressive and genetically unstable. In other embodiments, the cancer may remain benign, inactive, or dormant. The systems and methods of the present disclosure are useful in determining disease progression.

[0143] In addition, the systems and methods described herein can be used to determine the efficacy of a particular treatment option. In one embodiment, certain treatment options can be correlated with the genetic cfDNA fragmentation profile of cancer over time. This correlation can be used to select a therapy. Further, if cancer is observed to be in remission after treatment, the systems and methods described herein can be used to monitor for residual disease or disease recurrence.

[0144] In addition, the methods of the present disclosure can be used to characterize the heterogeneity of an abnormal condition in a subject, the method comprising: generating a cfDNA fragmentation profile of extracellular polynucleotides of the subject, wherein the cfDNA fragmentation profile comprises a plurality of data obtained from profile variation and mutation analysis. In some cases, including but not limited to cancer, a disease can be heterogeneous. Diseased cells can be not identical. Taking cancer as an example, it is known that some tumors contain different types of tumor cells, and some cells are at different stages of cancer. In other embodiments, the heterogeneity can comprise multiple disease foci. Similarly, in the case of cancer, there can be multiple tumor foci, and one or more of the foci may be the result of metastases that have spread from the primary site (also referred to as distant metastases).

[0145] The methods of the present disclosure can be used to generate a profile, fingerprint, or dataset that is the sum of genetic information from different cells in a heterogeneous disease. The set of data can comprise copy number variation and mutation analysis, either alone or in combination.

[0146] In addition, these reports are submitted and accessed electronically via the Internet. The data analysis occurs at a location other than the subject's location. Reports are generated and transmitted to the subject's location. The subject accesses, via a networked computer, a report that reflects their tumor burden.

[0147] Healthcare providers can use the annotation information to select other drug treatment options and / or provide information about drug treatment options to insurance companies. The method can include annotating drug treatment options for a particular disease, such as in the NCCN Clinical Practice Guidelines in Oncology TMor in the American Society of Clinical Oncology (ASCO) Clinical Practice Guidelines.

[0148] Generate reports that map the genomic locations and cfDNA fragmentation profile variations of subjects with cancer. Compared to the other profiles of subjects with known outcomes, these reports can indicate that a particular cancer is aggressive and resistant to treatment. Monitor the subject over a period of time and retest. If at the end of that period, the cfDNA fragmentation variant profile has not changed, this can indicate that the current treatment is not working. Compare with the cfDNA fragmentation profiles of other subjects. For example, if it is determined that changes in the cfDNA fragmentation variants indicate that the cancer is progressing, the original treatment regimen employed is no longer treating the cancer, and a new treatment is adopted.

[0149] In certain embodiments, the system receives genetic information from a DNA sequencer. The process then determines specific cfDNA fragmentation alterations and their amounts. These reports are submitted and accessed electronically via the Internet. Data analysis occurs at a location outside the subject's location. Reports are generated and transmitted to the subject's location. The subject accesses the reports reflecting their tumor burden via a networked computer.

[0150] Although time information can be used to enhance the information of the cfDNA fragmentation profile, other consensus methods can be applied. In other embodiments, historical comparisons can be used in combination with other consensus cfDNA fragmentation profiles. The consensus cfDNA fragmentation profiles can be normalized against control samples. Measurements of molecules mapped to a reference sequence can also be compared across the genome to identify regions where the cfDNA fragmentation profile has changed or remained unchanged. Consensus methods include, for example, linear or non-linear methods (such as voting, averaging, statistical, maximum a posteriori or maximum likelihood detection, dynamic programming, Bayesian, hidden Markov or support vector machine methods, etc.) for constructing a consensus cfDNA fragmentation profile derived from digital communication theory, information theory or bioinformatics. After determining the sequence read coverage, a stochastic modeling algorithm is applied to convert the normalized nucleic acid sequence read coverage of each window region into a discrete copy number state. In some cases, the algorithm may include one or more of the following: hidden Markov models, dynamic programming, support vector machines, Bayesian networks, grid decoding, Viterbi decoding, expectation maximization, Kalman filtering methods, and neural networks.

[0151] Artificial neural networks (NNets) simulate a network of "neurons" based on the structure of the brain. They process one record at a time, or process records in batch mode, and "learn" by comparing their classification of the records (which is largely arbitrary at the beginning) with the actual known classification of the records. In MLP-NNets, the error from the initial classification of the first record is fed back into the network and used to modify the network's algorithm for the second time, and so on for multiple iterations. Neural networks use an iterative learning process, where data cases (rows) are presented to the network one at a time, and the weights associated with the input values are adjusted each time.

[0152] After all cases have been presented, the process usually starts again. During this learning phase, the network learns by adjusting the weights so that it can predict the correct class label for the input samples. Due to the connections between units, neural network learning is also known as "associative learning". The advantages of neural networks include their high tolerance for noisy data and their ability to classify untrained patterns. One neural network algorithm is the backpropagation algorithm, such as Levenberg-Marquadt. Once the network has been constructed for a specific application, the network can be trained. Initial weights are randomly selected to start this process. Then the training or learning begins.

[0153] The network processes the records in one training data one at a time using the weights and functions in the hidden layer, and then compares the resulting output with the desired output. Then, the error propagates back through the system, causing the system to adjust the weights applied to the next record to be processed. This process occurs continuously as the weights are fine-tuned. During network training, the same set of data is processed multiple times as the connection weights are continuously improved.

[0154] In one embodiment, the training step of the machine learning unit on the training dataset can generate one or more classification models to be applied to test samples. These classification models can be applied to test samples to predict the response of a subject to a treatment intervention.

[0155] Comparison of sequence coverage to a control sample or reference sequence may assist in normalization across windows. In this embodiment, cell-free DNA is extracted and isolated from an easily accessible body fluid (such as blood). For example, cell-free DNA can be extracted using a variety of methods known in the art, including but not limited to isopropanol precipitation and / or silica-based purification. Cell-free DNA can be extracted from any number of subjects, such as subjects without cancer, subjects at risk of cancer, or subjects known to have cancer (e.g., by other means).

[0156] After the isolation / extraction step, the cell-free polynucleotide sample can be subjected to any of a variety of different sequencing operations. Prior to sequencing, the sample can be treated with one or more reagents (e.g., enzymes, unique identifiers (e.g., barcodes), probes, etc.). In some cases, if the sample is treated with a unique identifier such as a barcode, the sample or sample fragments can be labeled individually or in subgroups with the unique identifier. The labeled sample can then be used in downstream applications such as a sequencing reaction through which individual molecules can be traced to the parental molecule.

[0157] Cell-free polynucleotides can be labeled or traced to allow subsequent identification of that particular polynucleotide and its source. Assigning an identifier (e.g., a barcode) to an individual polynucleotide or subgroup of polynucleotides can allow a unique identity to be assigned to an individual sequence or sequence fragment. This can allow data to be obtained from a single sample and is not limited to the average of the sample. In some embodiments, nucleic acids or other molecules from a single strand may share a common tag or identifier and can thus be later identified as originating from that strand. Similarly, all fragments from a single strand of nucleic acid can be labeled with the same identifier or tag, allowing subsequent identification of fragments from the parental strand. In other cases, gene expression products (e.g., mRNA) can be labeled to quantify expression, whereby the barcodes or combinations of barcodes and the sequences to which they are attached can be counted. In other cases, the systems and methods can be used as PCR amplification controls. In such cases, multiple amplification products in a PCR reaction can be labeled with the same tag or identifier. If the products are subsequently sequenced and sequence differences are confirmed, the differences between products with the same identifier can be attributed to PCR errors. In addition, individual sequences can be identified based on characteristics of the sequence data of the read length itself. For example, detecting unique sequence data at the start (initiation) and end (termination) portions of a single sequencing read can be used alone or in combination with the length or number of base pairs of a unique sequence for each sequence read to assign a unique identity to an individual molecule. Fragments from a single strand of nucleic acid (which have been assigned a unique identity) can thus allow subsequent identification of fragments from the parental strand. This can be used in combination with restricting the initial starting genetic material to limit diversity.

[0158] Generally, the methods and systems provided herein can be used to prepare cell-free polynucleotide sequences for downstream application sequencing reactions. Typically, the sequencing methods are next-generation sequencing (NGS), classical Sanger sequencing, whole-genome bisulfite sequencing (WGSB), small RNA sequencing, low-coverage whole-genome sequencing (lcWGS), etc.

[0159] As used herein, the term "sequencing" refers to any of a variety of techniques for determining the sequence of a biomolecule (e.g., a nucleic acid such as DNA or RNA). Exemplary sequencing methods include, but are not limited to, targeted sequencing, single molecule real-time sequencing, exon sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy termination sequencing, whole genome sequencing, hybridization sequencing, pyrosequencing, capillary electrophoresis, gel electrophoresis, duplex sequencing, cycle sequencing, single base extension sequencing, solid phase sequencing, high throughput sequencing, massively parallel signature sequencing, emulsion PCR, low denaturation temperature co-amplification-PCR (COLD-PCR), multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, near-term sequencing, exonuclease sequencing, ligation sequencing, short read sequencing, single molecule sequencing, synthetic sequencing, real-time sequencing, reverse terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD TM sequencing, MS-PET sequencing, and combinations thereof. In certain embodiments, sequencing can be performed by a genetic analyzer, such as, for example, a genetic analyzer commercially available from Illumina or Applied Biosystems. In certain embodiments, the sequencing method can be massively parallel sequencing, i.e., simultaneous (or rapid consecutive) sequencing of any one of at least 100, 1000, 10,000, 100,000, 1 million, 10 million, 100 million, or 1 billion polynucleotide molecules.

[0160] After sequencing, a quality score is assigned to a read. The quality score can be an indication of the read, which indicates whether those reads are useful in subsequent analysis based on a threshold. In some cases, the quality or length of some reads is insufficient to perform subsequent mapping steps. Sequencing reads with a quality score of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered out of the dataset. In other cases, sequencing reads assigned a quality score of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered out of the dataset. In step 306, genomic fragment reads that meet the specified quality score threshold are mapped to a reference genome, or a reference sequence known not to contain mutations. After the mapping alignment, a mapping score is assigned to the sequence reads. The mapping score can be an indication of the mapping back to the reference sequence or the read, which indicates whether each position is uniquely mappable. In some cases, the read can be a sequence that is irrelevant to mutation analysis. For example, some sequence reads can be derived from contaminating polynucleotides. Sequencing reads with a mapping score of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered out of the dataset. In other cases, sequencing reads assigned a mapping score less than 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered out of the dataset. For each mappable base, bases that do not meet the minimum threshold for mappability or low-quality bases can be replaced by the corresponding bases found in the reference sequence.

[0161] Multiple cancers can be detected using the methods and systems described herein. Like most cells, cancer cells can be characterized by a turnover rate, where old cells die and are replaced by newer cells. Generally, dead cells can release DNA or DNA fragments into the blood when in contact with the vasculature of a given subject. This is also true for cancer cells at various stages of the disease. Depending on the stage of the disease, cancer cells can also be characterized by various genetic aberrations such as copy number variations and mutations. This phenomenon can be used to detect the presence or absence of cancer individuals using the methods and systems described herein.

[0162] The types and numbers of cancers that can be detected can include but are not limited to blood cancers, brain cancers, lung cancers, skin cancers, nasal cancers, laryngeal cancers, liver cancers, bone cancers, lymphomas, pancreatic cancers, skin cancers, intestinal cancers, rectal cancers, thyroid cancers, bladder cancers, kidney cancers, oral cancers, stomach cancers, solid tumors, heterogeneous tumors, homogeneous tumors, etc.

[0163] In addition, the systems and methods described herein can also be used to help characterize certain cancers. The genetic data generated by the systems and methods of the present disclosure can enable practitioners to better characterize specific forms of cancer. Often times, cancers are heterogeneous in composition and stage. Gene profile data can characterize specific subtypes of cancer, which can be important for the diagnosis or treatment of that specific subtype. This information can also provide clues to the prognosis of a specific type of cancer for a subject or practitioner.

[0164] The systems and methods provided herein can be used to monitor known cancers or other diseases in a specific subject. This can allow a subject or practitioner to adjust treatment options based on the progression of the disease. In this embodiment, the systems and methods described herein can be used to construct a gene profile of a specific subject during the course of a disease. In some cases, cancer can progress and become more aggressive and genetically unstable. In other embodiments, the cancer may remain benign, inactive, or dormant. The systems and methods of the present disclosure can be used to determine disease progression.

[0165] In addition, the systems and methods described herein can be used to determine the efficacy of a specific treatment option. In one embodiment, if a treatment is successful, a successful treatment option can actually increase the amount of copy number variations or mutations detected in a subject's blood because more cancer cells can die and shed DNA. In other embodiments, this may not occur. In another embodiment, certain treatment options may be correlated with the gene profile of a cancer over time. This correlation can be used to select a therapy. Additionally, if a cancer is observed to be in remission after treatment, the systems and methods described herein can be used to monitor for residual disease or disease recurrence.

[0166] Data is sent to a computer for processing either through a direct connection or the Internet. The data processing aspect of the system can be implemented by digital electronic circuitry, or by computer hardware, firmware, software, or combinations thereof. The data processing apparatus of the present invention can be implemented in the form of a computer program product tangibly embodied in a machine-readable storage device for execution by a programmable processor; and the data processing method steps of the present invention can be executed by a programmable processor executing an instruction program, the programmable processor performing the functions of the present invention by operating on input data and generating output. The data processing aspect of the present invention can be advantageously implemented in one or more computer programs executable on a programmable system including at least one programmable processor coupled to receive data and instructions from, and to transmit data and instructions to, a data storage system, at least one input device, and at least one output device. Each computer program can be implemented in a high-level procedural or object-oriented programming language, or, if desired, in assembly or machine language; and, in any case, the language can be a compiled or interpreted language. Suitable processors include, by way of example, both general and special purpose microprocessors. Generally, a processor will receive instructions and data from a read-only memory and / or a random access memory. Storage devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including by way of example semiconductor memory devices (such as EPROM, EEPROM, and flash memory devices); magnetic disks (such as internal hard disks and removable disks); magneto-optical disks; and CD-ROM disks. Any of the foregoing may be supplemented by, or incorporated in, an ASIC (Application Specific Integrated Circuit).

[0167] To provide interaction with a user, the method can be implemented using a computer system having a display device (such as a monitor or LCD (Liquid Crystal Display) screen) for displaying information to the user and an input device by which the user can provide input to the computer system, the input device such as a keyboard, a two-dimensional pointing device (such as a mouse or trackball), or a three-dimensional pointing device (such as a data glove or gyroscopic mouse). The computer system can be programmed to provide a graphical user interface through which the computer program interacts with the user. The computer system can be programmed to provide a virtual reality, three-dimensional display interface.

[0168] Embodiments

[0169] Embodiment 1: Genomic Analysis of Clinical Cohorts and cfDNA

[0170] We examined plasma samples from 501 individuals, including 75 individuals with HCC and 426 individuals without cancer. Among the individuals without cancer, 133 had conditions associated with an increased risk of HCC, including cirrhosis of all causes or viral hepatitis without cirrhosis. Blood samples were prospectively collected from HCC patients at different cancer stages and from high-risk individuals at Johns Hopkins Hospital, while the remaining samples were identified through screening efforts at other US or EU hospitals (US / EU cohort) (Table 1). We isolated 0.5 to 5 ml of plasma from each of these individuals, generated genomic libraries, and sequenced cfDNA fragments using low-coverage whole-genome sequencing (approximately 2.6-fold coverage), with an average of 49 million high-quality paired reads per sample, containing 9 Gb of sequence data (24, 25). In addition to the US / EU cohort, we also examined whole-genome sequence data from 223 patients from Hong Kong, China, as a validation cohort, including patients with resectable early HCC (n = 90, stage A = 85, B = 5), HBV (n = 66), and HBV-related cirrhosis (n = 35), as well as healthy individuals without liver disease (n = 32) (Hong Kong, China cohort) (Table 1) (15, 28).

[0171] Understanding the whole-genome cfDNA fragmentation profile through underlying chromatin structure

[0172] We evaluated the fragmentome using the DELFI protocol (24) and generated a fragmentation profile of the entire genome in 473 non-overlapping 5-MB regions, each containing approximately 80,000 fragments and spanning approximately 2.4 GB of the genome. In individuals without cancer, the fragmentation profiles were consistent, but in patients with HCC, the fragmentation profiles varied widely (Figure 1A). The profiles of patients with cirrhosis were closer to those of non-cancer individuals without cirrhosis compared to the profiles of patients with HCC (Figure 1A). Similarly, the fragmentation profiles of patients with viral hepatitis were almost identical to those of non-cancer individuals without liver disease (Figure 1A).

[0173] To examine the origin of cfDNA fragmentation patterns, we compared whole-genome fragmentome profiles with high-throughput sequencing chromatin conformation capture (Hi-C) open (A) and closed (B) compartments. We found that the cfDNA patterns of healthy individuals were highly correlated with those of lymphoblastoid-like cells (Figure 1B). Analysis of the cfDNA profiles of 10 HCC patients with high ctDNA levels revealed that their fragmentomes reflected two components, one similar to the profiles of individuals without cancer, and a separate cfDNA component with high similarity to the A / B compartments previously evaluated from liver cancers (Figure 1B)(29). In addition, when evaluating these two components, the predicted cfDNA profiles of the liver component were highly similar to the whole-genome A / B compartments of liver cancers, while the profiles of HCC patients had moderate similarity to liver cancers (Figure 1B, C). In contrast, the profiles of individuals without cancer were closer to the A / B compartments of lymphoblastoid-like cells (Figure 1B, 1C). These analyses suggest that the cfDNA fragmentome of HCC individuals represents a mixture of the cfDNA profiles of chromatin compartments of peripheral blood cells and those from liver cancers.

[0174] Disease-specific transcription factors inferred from whole-genome cfDNA fragmentation

[0175] Since chromatin organization reflects underlying cellular transcriptional programs (30 - 32), we examined whether cfDNA fragmentation signatures might reflect changes resulting from altered DNA binding by transcription factors (TFs) in liver cancer. To identify DNA binding sites for all known TFs, we analyzed 5,620 CHIP-seq experiments from the ReMap 2020 database (33). For each TF, we calculated the aggregated cfDNA coverage across all identified binding sites (4K - 490K per sample), which was compared to the overall adjacent genomic coverage to generate a single metric for each TF in each sample. We compared these TFs in patients with and without HCC to identify those with the largest and smallest differences in genome-wide binding site coverage of cfDNA (Figures 2A, B). Gene set enrichment analysis using the DisGeNET database of gene-disease associations indicated that differences in cfDNA TF binding coverage between HCC and cancer-free individuals were expected to be related to liver cancer and other cancers (Figures 2C, D). Additionally, the top-scoring TFs represented those with known biological relevance to chromatin organization and liver cancer (Table 2). These included members of the activator protein 1 (AP1) complex, including the JUN, JUND, ATF2, and ATF7 genes (which integrate extracellular signals (34) and have been associated with liver tumorigenesis (35, 36)), the transcriptional enhancer domain family member 4 (TEAD4), which has been shown to have an oncogenic role in HCC (37, 38), the poly(C)-binding protein 2 (PCBP2) transcriptional co-regulator, which is associated with poor prognosis in HCC patients when overexpressed (39), prohibitin 2 (PHB), which promotes HCC progression (40), and the AT-rich interaction domain 3A (ARID3A), an oncogenic transcription factor that promotes liver cancer malignancy when upregulated (41). Similar analysis of cfDNA fragmentation data from our recent study of patients in the LUCAS lung cancer diagnostic trial (24) revealed enrichment of differences in binding site coverage of TFs associated with lung cancer (Figures 2C, E). Collectively, these observations suggest that cfDNA fragmentation changes in patients with liver cancer and other cancers originate from the large number of altered transcriptional profiles present in cancer cells.

[0176] Uncovering Genomic Alterations in HCC from the cfDNA Fragmentome

[0177] Since the cfDNA fragmentome may contain changes associated with large-scale genomic alterations released by cancer cells (24, 25), we also examined chromosomal gains and losses in the circulation of these patients. In addition to the genome-wide fragmentation profiles resulting from chromatin and transcription factor changes observed in patients with liver cancer (Figure 3A), our analysis revealed altered representations of chromosomal arms, which matched those of the common gains or losses reported in liver cancer in the previous TCGA large-scale genomic study of HCC (n = 372) (Figure 3B). These included increased cfDNA representations of 1q, 7p, 7q, 8q, and decreased levels of 4q, 8p, 9p, 13q, and 21q, all of which are known to be increased or lost in HCC, respectively (42, 43). Importantly, these alterations were observed in patients with HCC but not in individuals without cancer, even if they had cirrhosis or chronic liver disease (Figure 3B).

[0178] The DELFI model for HCC detection

[0179] Given the direct link between genomic and chromatin changes in liver cancer and cfDNA fragmentation, we used a machine learning protocol to determine whether changes in the cfDNA fragmentome could distinguish between patients with HCC and those without cancer. We previously developed a robust classifier for lung cancer detection using this protocol, which was externally validated in an independent cohort (24). We determined the performance of this classifier in the US / EU cohort by repeated five-fold cross-validation, generating a score for each individual that was the mean of ten cross-validation repeats (DELFI score). The resulting model included a combination of regional and large-scale fragmentation features that were optimal for identifying individuals with liver cancer ( Figure 5 ). These features incorporated most of the informative chromosomes, chromatin, and local changes identified above and accounted for more than 90% of the differences in fragmentation profiles between samples.

[0180] Since clinical characteristics may affect tumor biomarkers, we investigated whether indicators of liver dysfunction or demographic parameters such as age, gender, race, or weight were associated with the DELFI score in cancer-free individuals for whom this information was available.

[0181] We observed no association between the DELFI score and age (R = 0.18, p = 0.08, Spearman correlation) ( Figure 6A ) and no difference in DELFI scores between men and women (p = 0.58, Wilcoxon test) ( Figure 6B)。It has been confirmed that Asians and African Americans have a higher incidence of hepatocellular carcinoma diagnosed at a later stage (44), and we observed that there were subtle differences in fragmentation scores among high-risk individuals without cancer in these or other racial or ethnic groups, although these analyses were limited by the lack of information on clinical covariates in some of these cases (p = 0.037 in patients with viral hepatitis and p = 0.026 in patients with cirrhosis, Kruskal-Wallis test)( Figure 7 )。Among individuals with cirrhosis, we observed a correlation between the degree of liver disease measured by the Child-Pugh score and the DELFI score (R = 0.58, p = 8.6e-5, Spearman correlation)( Figure 8 )。Increased body mass index (BMI), a risk factor for NAFLD and hepatocellular carcinoma, was not associated with changes in the DELFI score in patients with viral hepatitis (R =.027, p =.85, Spearman correlation); however, lower BMI in patients with cirrhosis was associated with higher DELFI scores, possibly due to cachexia in patients with severe cirrhosis (R = -.23, p =.043, Spearman correlation)( Figure 9 )。

[0182] Next, we examined the association between the DELFI score and the presence and stage of hepatocellular carcinoma in high-risk groups for hepatocellular carcinoma. The DELFI scores of 133 cancer-free individuals were lower, with median DELFI scores of 0.078 or 0.080 for patients with viral hepatitis or cirrhosis, respectively. In contrast, 75 HCC patients had significantly higher median DELFI scores in all BCLC stages, including stage 0 = 0.46, stage A = 0.61, stage B = 0.83, and stage C = 0.92 (p < 0.01 for stage 0, A, B, or C, Wilcoxon rank-sum test, Figure 4A). The receiver operating characteristic (ROC) curve for the DELFI protocol used to identify HCC patients showed an area under the curve (AUC) of 0.90 (95% CI = 0.86 - 0.94) for high-risk individuals. For early HCC, the performance remained strong, with AUCs of 0.9 and.81 for BCLC stage 0 and stage A, respectively. Among the individuals analyzed, late HCC (BCLC C) individuals were detected almost perfectly (AUC > 0.97) (Figure 4C).

[0183] To extend these analyses to individuals at low risk of developing HCC, we examined the ability of the DELFI model to distinguish individuals with cancer from those in a general population without viral hepatitis or cirrhosis (n = 293). In this larger cohort where additional features could be incorporated into cross-validation training, we used the features of the model described above and also included cfDNA coverage of CHIP-seq-derived transcription factor binding sites in hepatocyte cell lines available in the ReMAP database to create a DELFI model applicable to the general population. This protocol had high-performance cancer detection (AUC = 0.98) in these individuals. We evaluated the performance of the model at 99% specificity, which is the threshold suitable for the average-risk population (25), and observed an overall sensitivity of 80% in this setting (Figure 4B), with sensitivities for all stages higher than 65%. Application of the model without incorporation of transcription factor binding sites resulted in slightly decreased performance, and there was a high correlation between the ranked scores using our DELFI model for high-risk and screening populations (R =.64, p < 1e-15) (Figures 10 and 11).

[0184] To examine the relationship between the fragmentation profile and HCC progression, we evaluated whether the size, number, and characteristics of HCC lesions and the tumor etiology were associated with an abnormal fragmentation profile (when this information was available). We found that tumor size and number of lesions were positively correlated with DELFI scores (R = 0.42 and.31, p = 0.00026 and p = 0.0064, Spearman correlation), respectively) (Figure 12), which is consistent with the view that the fragmentation profile is related to the overall tumor burden.

[0185] Among HCC patients in the resectable stage (0, A, and B), the cancer etiology (including cirrhosis due to viral hepatitis or alcohol, NAFLD, or idiopathic causes) produced similar DELFI scores (p = 0.43, Kruskal-Wallis test)( Figure 13 ). These observations suggest that the fragmentation profile is the result of ongoing tumor-related cfDNA processes and is not affected by early events in tumorigenesis.

[0186] To examine the practical impact of this method in the context of HCC detection, we compared the performance of the DELFI fragmentome with current screening measurements of alpha-fetoprotein (AFP) levels. In 39 of 75 cancer individuals (52%), the AFP level was above the recommended screening threshold of 20 ng / ml, which is consistent with previous reports (45). Among individuals with an AFP level below 20 ng / ml who were not detected by this protocol, DELFI detected 30 of 36 (83%). The application of AFP measurement detected 8 / 24 (33%) stage 0 / A patients, 17 / 30 (57%) stage B patients, and 14 / 21 (66%) stage C patients ( Figure 14 ). In contrast, the DELFI protocol detected 19 / 24 (79%) stage 0 / A patients, 25 / 30 (83%) stage B patients, and 20 / 21 (95%) stage C patients. Overall, whole-genome cfDNA fragmentation analysis has improved performance compared to AFP detection for HCC, and the combination of DELFI and AFP provides improved detection compared to the DELFI protocol alone, as we observed a combined sensitivity of 92% and a combined specificity of 80%.

[0187] External validation of the DELFI model in East Asian populations with HCC

[0188] In addition to our cross-validation analysis of the US / EU cohort, we also tested the fixed DELFI model in 223 patients from the Hong Kong, China cohort. These included patients with mostly resectable early HCC (n = 90, stage A = 85, B = 5) and 101 patients with cirrhosis or HBV infection. These samples had been previously sequenced using different sequencers (HiSeq 2000 vs. Novaseq; 76 bp vs. 100 bp read lengths), different library preparations, and a higher number of PCR cycles (14 vs. 4 cycles), but we observed a similar whole-genome pattern to our previous analysis ( Figure 15 ). The fragmentation profiles of patients with viral hepatitis and cirrhosis, as well as healthy individuals, had a highly consistent profile across the genome, while the fragmentation profiles of HCC patients were variable and jumbled ( Figure 16 ). In addition, chromosomal changes observed in plasma from the Hong Kong, China cohort were similar to those in the original US / EU cohort and in cancers from TCGA ( Figure 17)。Overall, in this validation cohort, the DELFI model distinguished HCC patients from patients with high-risk diseases with an AUC of 0.97 (Figure 4D). These observations suggest that the underlying features of cfDNA fragmentation are similar in this cohort and that DELFI is a powerful method for detecting HCC and can be generalized to different high-risk groups.

[0189] Simulation of DELFI performance at population scale

[0190] To evaluate the surveillance and detection performance of our assay in high-risk patients with liver cancer, we used Monte Carlo simulation to evaluate the DELFI model in a theoretical cohort of 100,000 high-risk individuals. Given the importance of early cancer detection, we focused our modeling on the detection of stage 0 / A disease. We compared the DELFI assay to the current standard of care, concurrent ultrasound and AFP, and modeled the uncertainty in the sensitivity and specificity of these surveillance modalities in this theoretical cohort using probability distributions centered on estimates from our cohort or previously reported experience ((11), see Methods). Despite surveillance recommendations, compliance with HCC surveillance in the United States is low, with the best estimate indicating 39% compliance (46), resulting in an average of 40,042 individuals being tested in this theoretical cohort (95% CI, 21,320 - 61,890). Since blood tests offer high accessibility and compliance, with reported adherence rates for blood-based biomarkers of 80 - 90% (47, 48), we conservatively estimated that an average of 75% (95% CI, 60 - 90%) of this cohort would be tested using the DELFI assay. Since the prevalence of cirrhosis, viral hepatitis, and the co-occurrence rates of these comorbidities with HCC may vary by region, we used prior probability distributions to reflect our uncertainty about the composition of these diseases and possible regional differences. Monte Carlo simulation from these probability distributions (Methods) showed that ultrasound with AFP detected an average of 2,233 (95% CI, 1,088 - 3,699) individuals with liver cancer ( Figure 18 ). Using DELFI, we would detect an average of 2,794 additional cases of liver cancer, or a 2.46-fold increase compared to ultrasound alone with AFP (95% CI, 1.25 - 4.57-fold increase) ( Figure 18)。The DELFI assay not only will significantly improve the detection of liver cancer, but also is expected to reduce the false negative rate (FNR), or the proportion of cancers missed in the test, from 38% (95% CI, 25% - 51.5%) for AFP ultrasound to 24% (95% CI, 9% - 42.6%) for DELFI. Additionally, the negative predictive value (NPV) of the test is expected to increase from 95.7% (95% CI, 93.8% - 97.3%) for AFP ultrasound to 97.1% (95% CI, 94.8% - 99.0%) for DELFI. These analyses suggest that using a highly specific blood-based early detection assay as a tool for detecting liver cancer has significant benefits for the overall population.

[0191] Discussion

[0192] Overall, in this study, we demonstrated the use of whole-genome cfDNA fragmentome signatures for the detection of HCC with high sensitivity and specificity. Additionally, we demonstrated that the fragmentation profile captures genomic and chromatin features, including alterations known to be important in HCC. Our cfDNA fragmentome assay performs well in detecting HCC, including very early-stage disease, and is not affected by etiology. To our knowledge, this is the first whole-genome fragmentation analysis independently validated in a separate high-risk population with stable and robust performance across diverse racial and ethnic groups in the United States and Hong Kong, China.

[0193] Our results also suggest that disease-specific transcription factor signatures can be obtained by analyzing the whole-genome cfDNA fragmentation profile. Although such an analysis has been performed using specific transcription factors to distinguish small cell lung cancer from non-small cell lung cancer (24), this study suggests that using whole-genome cfDNA fragmentation to analyze disease-specific transcriptional regulation can improve the detection and identification of primary tissues in cancer patients. With a sufficient number of patients, the cfDNA transcriptional profile can further improve machine learning algorithms for detecting HCC and other cancers.

[0194] Compared with other solid cancers, HCC is unique in that there are a large number of well-defined high-risk groups, with an average annual risk of developing HCC of 3-4% (49), and regular cancer screening every six months is recommended. Unfortunately, currently available tests have limited diagnostic utility, especially for early disease (13). In our study, AFP had a sensitivity of 52% in detecting HCC, which is consistent with the known performance of this biomarker (13). Ultrasound-based monitoring also has operator-dependent technical limitations and lower sensitivity in obese and cirrhotic patients (50). Most importantly, compared with much higher compliance of blood tests for other conditions (47), ultrasound has low compliance with established guidelines, less than 20% globally (10, 11). Despite these challenges, HCC screening still provides an overall survival benefit for patients with HBV (51) and cirrhosis (10), highlighting the urgent need to improve current screening tests. The high performance and cost-effective characteristics of cfDNA fragmentome analysis in HCC detection will make DELFI an accessible HCC screening test and raise the screening rate above the current abysmal levels. An interesting aspect of cfDNA analysis for HCC is that transplantation is the most effective treatment for early to mid-stage HCC, and using a liquid biopsy protocol for HCC monitoring in post-transplant patients can play a dual role in tracking recurrence and rejection, as studies of cfDNA in post-transplant patients have shown promise (52).

[0195] Although this study represents a potential improvement over current screening programs, there are still certain limitations. For example, this study included a relatively small sample of HCC individuals. Despite pre-analytical differences in laboratory and sequencing methods for the conduct of the independent validation cohort, the fact that the DELFI protocol performed well in this cohort suggests that the method will ultimately be able to be applied in a range of different diagnostic laboratories. Larger validation studies will be needed before the protocol can be clinically applied. Nevertheless, the observation that scalable and cost-effective non-invasive cfDNA fragmentome analysis can detect liver cancer patients can provide opportunities for screening high-risk and general populations worldwide.

[0196] Materials and Methods

[0197] Study Population

[0198] For the US / EU cohort, samples from 208 patients, including 75 HCC patients and 133 high-risk patients without HCC, were prospectively collected at Johns Hopkins University School of Medicine under a protocol approved by the Johns Hopkins Institutional Review Board as part of the HCC Biomarker Registry and AIDS Linked to the IntraVenous Experience (ALIVE) study. HCC was defined by histological examination or appropriate imaging features defined by accepted guidelines. Tumor stage was determined by the Barcelona Clinic Liver Cancer staging system (BCLC). Detailed clinical data were extracted from electronic medical records.

[0199] High-risk patients were defined as individuals with cirrhosis of any etiology and / or individuals with chronic hepatitis B or C recommended for routine HCC screening by expert society guidelines (50). In addition, we included 38 patients with hepatitis B or cirrhosis retrospectively collected by BioIVT (Westbury, NY). The collaborating centers quantified AFP levels using a Food and Drug Administration-approved AFP assay in their clinical laboratories.

[0200] The US / EU cohort also included samples from 293 cancer-free individuals previously analyzed (24), originally from two screening clinical trial cohorts for colorectal cancer in Denmark (Endoscopy III) and the Netherlands (COCOS, Dutch Trial Register ID NTR182946). The protocol for the Endoscopy III Project was approved by the Regional Ethics Committee and the Danish Data Protection Agency, and for the COCOS trial, ethical approval was obtained from the Dutch Health Council. The inclusion criteria for both the Dutch and Danish cohorts were any individual aged 50 - 75 years eligible for colorectal cancer screening. All patients used had a negative FIT test or negative colonoscopy results.

[0201] For the Hong Kong, China cohort, written informed consent was obtained from all recruited subjects, and the study was approved by the Joint Chinese University of Hong Kong and the New Territories East Cluster Clinical Research Ethics Committee (15, 28).

[0202] Sample collection and storage

[0203] Sample collection was performed as follows: Peripheral venous blood was collected into one K2-EDTA tube and two serum gel tubes. Within two hours of blood collection, the tubes were centrifuged at 2330 g for 10 min at 4 °C, the plasma was transferred to new tubes, and the samples were spun at 14,000 rpm (18,000 rcf) for 10 min at room temperature to pellet any remaining cell debris. After centrifugation, the EDTA plasma was aliquoted and stored at -80 °C for cfDNA analysis.

[0204] Sequencing library preparation

[0205] Circulating cell-free DNA was isolated from 2–4 ml plasma using the Qiagen QIAamp Circulating Nucleic Acids Kit (Qiagen GmbH), eluted in 52 μl RNase-free water containing 0.04% sodium azide (Qiagen GmbH), and stored at -20 °C in LoBind tubes (Eppendorf AG). The concentration and quality of cfDNA were evaluated using a Bioanalyzer 2100 (Agilent Technologies).

[0206] Next-generation sequencing (NGS) cfDNA libraries were prepared using 15 ng of cfDNA (when available) or the entire purified amount (when less than 15 ng) for whole-genome sequencing (WGS). Briefly, genomic libraries were prepared using Illumina's NEBNext DNA Library Prep Kit (New England Biolabs (NEB)), with four major modifications to the manufacturer's guidelines: (i) Library purification steps followed the AMPure XP on beads (Beckman Coulter) protocol to minimize sample loss during elution and tube transfer steps; (ii) NEBNext end repair, A-tailing, and adapter ligation enzyme and buffer volumes were appropriately adjusted to accommodate AMPure XP purification on beads; (iii) Illumina dual-index adapters were used in the ligation reaction; and (iv) the cfDNA library was amplified with Phusion Hot Start polymerase. All samples underwent 4 cycles of PCR amplification after the DNA ligation step.

[0207] Low-coverage whole-genome sequencing and alignment

[0208] Whole-genome libraries from cancer patients and cancer-free individuals were prepared as in (24), with the following modifications: They were sequenced at 1–2× coverage per genome using a 100-bp paired-end run (200 cycles) on the Illumina Novaseq platform. Prior to alignment, adapter sequences were filtered from the reads using the fastp software (53). Sequences reads were aligned to the hg19 human reference genome using Bowtie2 (54), and duplicate reads were removed using Sambamba (55). After alignment, each aligned pair was converted to a genomic interval representing the sequenced DNA fragment using bedtools (56). Only reads with a MAPQ score of at least 30 or greater were retained. Reads pairs were further filtered if they overlapped with the Duke Excluded Regions blacklist (genome.ucsc.edu / cgi-bin / hgTrackUi?db=hg19&g=wgEncodeMapability). To capture large-scale epigenetic differences in fragmentation across the genome that are estimable from low-coverage WGS, we tiled the hg19 reference genome into non-overlapping 5-Mb bins. Bins with an average GC content of <0.3 and an average mappability of <0.9 were excluded, leaving 473 bins spanning approximately 2.4 GB of the genome. Following Mathios et al. (24), an external panel of 20 cancer-free individuals sequenced on the NovaSeq was used to perform GC correction independently on short (<150 bp) and long (>=150 bp) cfDNA fragments to generate the target distribution.

[0209] The Fastq files of patients in the Hong Kong cohort were obtained from the Circulating Nucleic Acids Research Group at The Chinese University of Hong Kong (CUHK) as reported in (15, Jiang, 2018 #1645) and processed as described above and by Mathios et al. to generate DELFI features. GC correction was performed by normalizing to the target distribution provided at github.com / cancer-genomics / PlasmaToolsNovaseq.hg19, which is the same as the target distribution used for GC correction in the US / EU cohort. The validation set consisted of libraries constructed with 14 PCR cycles and sequenced on the HiSeq 2000. These libraries were normalized to the 4-cycle NovaSeq target distribution for ease of comparison between studies. One sample each from the cirrhosis group and the HBV group was excluded because they were identified as having an HCC diagnosis.

[0210] Chromatin structure analysis

[0211] The A / B compartments of liver cancer tissues and lymphoblastoid-like cells were obtained from github.com / Jfortin1 / TCGA_AB_Compartments and github.com / Jfortin1 / HiC_AB_Compartments as described in (29). Two reference tracks were compared to identify informative 100-kb bins, which were defined as bins in which the chromatin domains differed between the two reference tracks or the magnitude difference of the eigenvalue corresponded to a z-score greater than 1.96 or less than -1.96 (p =.05) (for all eigenvalue differences).

[0212] The median fragmentation profiles of 10 liver samples with the highest estimated tumor scores according to ichorCNA (57) and 10 randomly selected cancer-free individuals were calculated. This information was used to extract the estimated median liver component in plasma, which was weighted by the ichor score of each plasma sample.

[0213] Genome-wide transcription factor analysis

[0214] Chromatin immunoprecipitation and subsequent sequencing (ChIP-Seq) peaks from 5620 experiments were downloaded from the ReMap 2020 database (33). This set of experiments with more than 4000 peaks was filtered, resulting in 4293 experiments. For each peak in the autosomes, we defined the center of the peak as position 0.

[0215] For each sample, the average coverage at each position (-3,000 to +3,000 relative to the center of each peak) was calculated across all peaks. For the ROC curve, the relative coverage for each sample was calculated as the average coverage within a ±100 bp window around the center of the binding site divided by the average coverage within a ±250 bp window around 2,750 bp upstream and downstream of the binding site. The ROC curve was generated using pROC 1.16.2 (58). The AUCs for each peak set were ranked. Each transcription factor was matched to its NCBI ID, and the remaining 797 unique transcription factors were ranked by AUC. This ranked list was used as input to the gseDGN function from the DOSE package in R. Its output was ranked by the normalized enrichment score (NSE).

[0216] Genome-wide fragment features

[0217] Fragmentation features were calculated as described by Mathios et al. (24). Briefly, the ratio of short to long fragments was calculated for 473 non-overlapping 5 MB bins in the genome, and z-scores representing arm gains / losses representing autosomal chromosomal arms were calculated. The principal components and z-scores for ratios representing >90% differences were used to train machine learning models.

[0218] Machine learning and cross-validation analysis

[0219] Two machine learning models were developed, one for the high-risk group (GBM using Mathios et al. features) and one for the low-risk general group (penalized logistic regression using Mathios et al. features and transcription factor binding site coverage). These models were trained on the US / EU cohorts in Caret, with 5-fold cross-validation repeated 10 times, and scores for each sample were calculated as the mean of the repeats and evaluated using AUC-ROC, as described by Mathios et al. (24). The first model used high-risk non-cancer and HCC patients, while the second model used non-cancer individuals without a history of liver disease. The locked high-risk model trained on the US / EU cohort was applied to the Hong Kong, China cohort to generate cancer predictions on an external validation set.

[0220] TCGA analysis

[0221] Copy number data for the HCC cancer cohort (LIHC n = 372) in TCGA were retrieved using the RTCGA v1.16.0 package. Analyses were performed to determine the frequency of copy number gains and losses in 473 5 mb bins for this cohort (24). The somatic copy number alteration (SCNA) thresholds used by Mathios et al. were used to determine gains and losses in the HCC cohort (24, 59).

[0222] Association between clinical covariates and DELFI scores

[0223] The potential associations between clinical covariates (for those patients for whom this information was available) and DELFI scores were evaluated using Spearman's rank correlation coefficient (for continuous variables) and Kruskal-Wallis one-way analysis of variance (for categorical variables).

[0224] Simulation

[0225] Monte Carlo simulation was used to compare the DELFI protocol with ultrasound and AFP in a theoretical surveillance population. We used the 95% confidence intervals for the estimated sensitivity and specificity of DELFI and the publicly available 95% confidence intervals for ultrasound and AFP (13). The prior predictive probability distributions (beta distributions) were derived from these confidence intervals using the R package epiR ((60) R package version 2.47, CRAN.R-project.org / package=Epi). Zhao et al. (2017) reported that compliance with US and AFP surveillance in the United States was 39% (95% CI: 21%-65%). Since other non-invasive blood-based tests had reported compliance of over 75% (47, 48), we assumed that compliance with DELFI would reach 60% or higher with a probability of 0.975 or higher. Using these confidence estimates, epiR was used to derive the beta prior predictive distribution of compliance. We simulated the multinomial probabilities of the prevalence of hepatitis B, cirrhosis, hepatitis B + HCC, cirrhosis + HCC, and hepatitis B + cirrhosis + HCC using Dirichlet with parameters 230, 680, 60, 23, 7, respectively. For a single Monte Carlo simulation of the ultrasound and AFP tests, we

[0226] (i) Drew the compliance probability (η) from the prior predictive distribution,

[0227] (ii) Simulated the number of individuals (S) participating in surveillance (S ~ binomial(η, 100,000)),

[0228] (iii) Drew the probability of comorbidity (Dirichlet(230, 680, 60, 23, 7)),

[0229] (iv) Calculated the prevalence of HCC (θ),

[0230] (v) Simulated HCC cases (P ~ binomial(θ, S)) and calculated the number of individuals without cancer (N = S - P),

[0231] (vi) Drew the sensitivity (se) and specificity (sp) from the corresponding prior predictive distributions, and

[0232] (vii) Extract true positives (TP ~ binomial(P, se)) and false positives (FP ~ binomial(N, 1 - sp)).

[0233] Given TP and FP, we calculate the NPV as (true negatives) / (true negatives + false negatives), where true negatives = N - FP and false negatives = P - TP. We repeated the above simulation 1000 times to obtain the distributions of TP, FP, and NPV. Using the parameters of sensitivity, specificity, and compliance of the DELFI protocol, we repeated the same Monte Carlo analysis to allow comparison of the two monitoring methods.

[0234] Table 1. Patient demographics and clinical information

[0235]

[0236] * P-values were calculated by comparing data from individuals with and without liver cancer for the following variables: mean age using the Student's unpaired two-tailed t-test, sex distribution, etiology of cirrhosis, and Child-Pugh stage using the chi-square test (χ 2 test).

[0237] # Validation cohort data were obtained from Jiang et al., PNAs, 2015.

[0238] Table 2. Top-ranked transcription factors in US / EU cohort samples

[0239]

[0240] * TF = transcription factor, HAT = histone acetyltransferase

[0241] References

[0242] 1. Sung H, Ferlay J, Siegel RL, Laversanne M, Soerjomataram I, Jemal A, et al. Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA: a cancer journal for clinicians 2021;71(3):209 - 49.

[0243] 2. Siegel RL, Miller KD, Fuchs HE, Jemal A. Cancer statistics, 2021. CA: a cancer journal for clinicians 2021; 71(1): 7 - 33.

[0244] 3. Di Bisceglie AM. Hepatitis C and hepatocellular carcinoma. Hepatology 1997; 26(3 Suppl 1): 34S - 8S doi 10.1002 / hep.510260706.

[0245] 4. Pinyopompanish K, Khoudari G, Saleh MA, Angkurawaranon C, Pinyopornpanish K, Mansoor E, et al. Hepatocellular carcinoma in non - alcoholic fatty liver disease with or without cirrhosis: a population - based study. BMC gastroenterology 2021; 21(1): 1 - 7.

[0246] 5. Donato F, Tagger A, Gelatti U, Parrinello G, Boffetta P, Albertini A, et al. Alcohol and hepatocellular carcinoma: the effect of lifetime intake and hepatitis virus infections in men and women. Am J Epidemiol 2002; 155(4): 323 - 31 doi 10.1093 / aje / 155.4.323.

[0247] 6. Waly Raphael S, Yangde Z, Yuxiang C. Hepatocellular carcinoma: focus on different aspects of management. ISRN Oncol 2012; 2012: 421673 doi 10.5402 / 2012 / 421673.

[0248] 7. Asrani SK, Devarbhavi H, Eaton J, Kamath PS. Burden of liver diseases in the World. J Hepatol 2019;70(1):151-71 doi 10.1016 / j.jhep.2018.09.014.

[0249] 8. Frenette CT, Isaacson AJ, Bargellini I, Saab S, Singal AG. A Practical Guideline for Hepatocellular Carcinoma Screening in Patients at Risk. Mayo Clin Proc Innov Qual Outcomes 2019;3(3):302-10 doi 10.1016 / j.mayocpiqo.2019.04.005.

[0250] 9. Kanwal F, Singal AG. Surveillance for hepatocellular carcinoma: current best practice and future direction. Gastroenterology 2019;157(1):54-64.

[0251] 10. Singal AG, Pillai A, Tiro J. Early detection, curative treatment, and survival rates for hepatocellular carcinoma surveillance in patients with cirrhosis: a meta-analysis. PLoS medicine 2014;11(4):e1001624.

[0252] 11. Singal AG, Yopp A, Skinner CS, Packer M, Lee WM, Tiro JA. Utilization of hepatocellular carcinoma surveillance among American patients: a systematic review. Journal of general internal medicine 2012;27(7):861-7.

[0253] 12. Singal AG, Li X, Tiro J, Kandunoori P, Adams-Huet B, Nehra MS, et al. Racial, social, and clinical determinants of hepatocellular carcinoma surveillance. The American journal of medicine 2015;128(1):90.e1-.e7.

[0254] 13. Tzartzeva K, Obi J, Rich NE, Parikh ND, Marrero JA, Yopp A, et al. Surveillance imaging and alpha fetoprotein for early detection of hepatocellular carcinoma in patients with cirrhosis: a meta-analysis. Gastroenterology 2018;154(6):1706-18.e1.

[0255] 14. Benesova L, Belsanova B, Suchanek S, Kopeckova M, Minarikova P, Lipska L, et al. Mutation-based detection and monitoring of cell-free tumor DNA in peripheral blood of cancer patients. Analytical biochemistry 2013;433(2):227-34.

[0256] 15. Jiang P, Chan CW, Chan KC, Cheng SH, Wong J, Wong VW, et al. Lengthening and shortening of plasma DNA in hepatocellular carcinoma patients. Proc Natl Acad Sci U S A 2015;112(11):E1317-25 doi 10.1073 / pnas.1500076112.

[0257] 16. Cai J, Chen L, Zhang Z, Zhang X, Lu X, Liu W, et al. Genome-wide mapping of 5-hydroxymethylcytosines in circulating cell-free DNA as a non-invasive approach for early detection of hepatocellular carcinoma. Gut 2019;68(12):2195-205 doi 10.1136 / gutjnl-2019-318882.

[0258] 17. Xu RH, Wei W, Krawczyk M, Wang W, Luo H, Flagg K, et al. Circulating tumour DNA methylation markers for diagnosis and prognosis of hepatocellular carcinoma. Nat Mater 2017;16(11):1155-61 doi 10.1038 / nmat4997.

[0259] 18. Wang Y, Zhou K, Wang X, Liu Y, Guo D, Bian Z, et al. Multiple-level copy number variations in cell-free DNA for prognostic prediction of HCC with radical treatments. Cancer Sci 2021;112(11):4772-84 doi 10.1111 / cas.15128.

[0260] 19. Kisiel JB, Dukek BA, R VSRK, Ghoz HM, Yab TC, Berger CK, et al. Hepatocellular Carcinoma Detection by Plasma Methylated DNA: Discovery, Phase I Pilot, and Phase II Clinical Validation. Hepatology 2019;69(3):1180 - 92 doi 10.1002 / hep.30244.

[0261] 20. Chalasani NP, Ramasubramanian TS, Bhattacharya A, Olson MC, Edwards VD, Roberts LR, et al. A Novel Blood - Based Panel of Methylated DNA and Protein Markers for Detection of Early - Stage Hepatocellular Carcinoma. Clin Gastroenterol Hepatol 2021;19(12):2597 - 605e4 doi 10.1016 / j.cgh.2020.08.065.

[0262] 21. Klein E, Richards D, Cohn A, Tummala M, Lapham R, Cosgrove D, et al. Clinical validation of a targeted methylation - based multi - cancer early detection test using an independent validation set. Annals of Oncology 2021;32(9):1167 - 77.

[0263] 22. Cancer Screening Cost with Galleri. <https: / / www.galleri.com / the - galleri - test / cost>.

[0264] 23. Chalasani NP, Porter K, Bhattacharya A, Book AJ, Neis BM, Xiong KM, et al. Validation of a Novel Multitarget Blood Test Shows High Sensitivity to Detect Early Stage Hepatocellular Carcinoma. Clin Gastroenterol H 2022; 20(1): 173 - 82.e7.

[0265] 24. Mathios D, Johansen JS, Cristiano S, Medina JE, Phallen J, Larsen KR, et al. Detection and characterization of lung cancer using cell-free DNA fragmentomes. Nat Commun 2021; 12(1): 5060 doi 10.1038 / s41467 - 021 - 24994 - w.

[0266] 25. Cristiano S, Leal A, Phallen J, Fiksel J, Adleff V, Bruhm DC, et al. Genome-wide cell-free DNA fragmentation in patients with cancer. Nature 2019; 570(7761): 385 - 9 doi 10.1038 / s41586 - 019 - 1272 - 6.

[0267] 26. Zhang X, Wang Z, Tang W, Wang X, Liu R, Bao H, et al. Ultra-Sensitive and Affordable Assay for Early Detection of Primary Liver Cancer Using Plasma cfDNA Fragmentomics. Hepatology 2021.

[0268] 27. Chen L, Abou-Alfa GK, Zheng B, Liu JF, Bai J, Du LT, et al. Genome-scale profiling of circulating cell-free DNA signatures for early detection of hepatocellular carcinoma in cirrhotic patients. Cell Res 2021; 31(5): 589-92 doi 10.1038 / s41422-020-00457-7.

[0269] 28. Jiang P, Sun K, Tong YK, Cheng SH, Cheng THT, Heung MMS, et al. Preferred end coordinates and somatic variants as signatures of circulating tumor DNA associated with hepatocellular carcinoma. Proc Natl Acad Sci USA 2018; 115(46): E10925-E33 doi 10.1073 / pnas.1814616115.

[0270] 29. Fortin J-P, Hansen KD. Reconstructing A / B compartments as revealed by Hi-C using long-range correlations in epigenetic data. Genome biology 2015; 16(1): 1-23.

[0271] 30. Choi JK, Kim YJ. Intrinsic variability of gene expression encoded in nucleosome positioning sequences. Nat Genet 2009; 41(4): 498-503 doi 10.1038 / ng.319.

[0272] 31. Ulz P, Perakis S, Zhou Q, Moser T, Belic J, Lazzeri I, et al. Inference of transcription factor binding from cell-free DNA enables tumor subtype prediction and early detection. Nat Commun 2019;10(1):4666 doi 10.1038 / s41467-019-12714-4.

[0273] 32. Snyder MW, Kircher M, Hill AJ, Daza RM, Shendure J. Cell-free DNA Comprises an In Vivo Nucleosome Footprint that Informs Its Tissues-Of-Origin. Cell 2016;164(1-2):57-68 doi 10.1016 / j.cell.2015.11.050.

[0274] 33. Cheneby J, Menetrier Z, Mestdagh M, Rosnet T, Douida A, Rhalloussi W, et al. ReMap 2020: a database of regulatory regions from an integrative analysis of Human and Arabidopsis DNA-binding sequencing experiments. Nucleic Acids Res 2020;48(D1):D180-D8 doi 10.1093 / nar / gkz945.

[0275] 34. Bejjani F, Evanno E, Zibara K, Piechaczyk M, Jariel-Encontre I. The AP-1 transcriptional complex: Local switch or remote command? Biochim Biophys Acta Rev Cancer 2019;1872(1):11-23 doi 10.1016 / j.bbcan.2019.04.003.

[0276] 35. Gozdecka M, Lyons S, Kondo S, Taylor J, Li Y, Walczynski J, et al. JNK suppresses tumor formation via a gene-expression program mediated by ATF2. Cell Rep 2014;9(4):1361-74 doi 10.1016 / j.celrep.2014.10.043.

[0277] 36. Yan P, Zhou B, Ma Y, Wang A, Hu X, Luo Y, et al. Tracking the important role of JUNB in hepatocellular carcinoma by single-cell sequencing analysis. Oncol Lett 2020;19(2):1478-86 doi 10.3892 / ol.2019.11235.

[0278] 37. Coto-Llerena M, Tosti N, Taha-Mehlitz S, Kancherla V, Paradiso V, Gallon J, et al. Transcriptional Enhancer Factor Domain Family member 4 Exerts an Oncogenic Role in Hepatocellular Carcinoma by Hippo-Independent Regulation of Heat Shock Protein 70 Family Members. Hepatol Commun 2021;5(4):661-74 doi 10.1002 / hep4.1656.

[0279] 38. Zhang Z, Fang X, Xie G, Zhu J. GATA3 is downregulated in HCC and accelerates HCC aggressiveness by transcriptionally inhibiting slug expression Corrigendum in / 10.3892 / ol.2021.12836. Oncology letters 2021;21(3):1-.

[0280] 39. Zhang X, Hua L, Yan D, Zhao F, Liu J, Zhou H, et al. Overexpression of PCBP2 contributes to poor prognosis and enhanced cell growth in human hepatocellular carcinoma. Oncol Rep 2016; 36(6): 3456 - 64.

[0281] 40. Xiang X, Fu Y, Zhao K, Miao R, Zhang X, Ma X, et al. Cellular senescence in hepatocellular carcinoma induced by a long non - coding RNA - encoded peptide PINT87aa by blocking FOXM1 - mediated PHB2. Theranostics 2021; 11(10): 4929 - 44 doi 10.7150 / thno.55672.

[0282] 41. Shen M, Li S, Zhao Y, Liu Y, Liu Z, Huan L, et al. Hepatic ARID3A facilitates liver cancer malignancy by cooperating with CEP131 to regulate an embryonic stem cell - like gene signature. Cell Death & Disease 2022; 13(8): 1 - 13.

[0283] 42. Marchio A, Pineau P, Meddeb M, Terris B, Tiollais P, Bernheim A, et al. Distinct chromosomal abnormality pattern in primary liver cancer of non - B, non - C patients. Oncogene 2000; 19(33): 3733 - 8 doi 10.1038 / sj.onc.1203713.

[0284] 43. Longerich T, Mueller MM, Breuhahn K, Schirmacher P, Benner A, Heiss C. Oncogenetic tree modeling of human hepatocarcinogenesis. Int J Cancer 2012; 130(3): 575 - 83.

[0285] 44. Stewart SL, Kwong SL, Bowlus CL, Nguyen TT, Maxwell AE, Bastani R, et al. Racial / ethnic disparities in hepatocellular carcinoma treatment and survival in California, 1988 - 2012. World journal of gastroenterology 2016; 22(38): 8584 - 95 doi 10.3748 / wjg.v22.i38.8584.

[0286] 45. Gupta S, Bent S, Kohlwes J. Test characteristics of alpha - fetoprotein for detecting hepatocellular carcinoma in patients with hepatitis C. A systematic review and critical analysis. Ann Intern Med 2003; 139(1): 46 - 50 doi 10.7326 / 0003 - 4819 - 139 - 1 - 200307010 - 00012.

[0287] 46. Zhao C, Jin M, Le RH, Le MH, Chen VL, Jin M, et al. Poor adherence to hepatocellular carcinoma surveillance: a systematic review and meta - analysis of a complex issue. Liver International 2018; 38(3): 503 - 14.

[0288] 47. Bokhorst LP, Alberts AR, Rannikko A, Valdagni R, Pickles T, Kakehi Y, et al. Compliance Rates with the Prostate Cancer Research International Active Surveillance (PRIAS) Protocol and Disease Reclassification in Noncompliers. Eur Urol 2015;68(5):814 - 21 doi 10.1016 / j.eururo.2015.06.012.

[0289] 48. Duffy MJ, van Rossum LG, van Turenhout ST, Malminiemi O, Sturgeon C, Lamerz R, et al. Use of faecal markers in screening for colorectal neoplasia: a European group on tumor markers position paper. Int J Cancer 2011;128(1):3 - 11.

[0290] 49. Singal AG, Lampertico P, Nahon P. Epidemiology and surveillance for hepatocellular carcinoma: New trends. J Hepatol 2020;72(2):250 - 61.

[0291] 50. Heimbach JK, Kulik LM, Finn RS, Sirlin CB, Abecassis MM, Roberts LR, et al. AASLD guidelines for the treatment of hepatocellular carcinoma. Hepatology 2018;67(1):358 - 80 doi 10.1002 / hep.29086.

[0292] 51. Zhang B-H, Yang B-H, Tang Z-Y. Randomized controlled trial of screening for hepatocellular carcinoma. Journal of cancer research and clinical oncology 2004;130(7):417-22.

[0293] 52. Goh SK, Do H, Testro A, Pavlovic J, Vago A, Lokan J, et al. The Measurement of Donor-Specific Cell-Free DNA Identifies Recipients With Biopsy-Proven Acute Rejection Requiring Treatment After Liver Transplantation. Transplant Direct 2019;5(7):e462 doi 10.1097 / txd.0000000000000902.

[0294] 53. Chen S, Zhou Y, Chen Y, Gu J. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 2018;34(17):i884-i90 doi 10.1093 / bioinformatics / bty560.

[0295] 54. Langmead B, Salzberg SL. Fast gapped-read alignment with Bowtie2. Nat Methods 2012;9(4):357-9 doi 10.1038 / nmeth.1923.

[0296] 55. Tarasov A, Vilella AJ, Cuppen E, Nijman IJ, Prins P. Sambamba: fast processing of NGS alignment formats. Bioinformatics 2015;31(12):2032-4 doi 10.1093 / bioinformatics / btv098.

[0297] 56. Quinlan AR, Hall IM. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics 2010;26(6):841-2 doi 10.1093 / bioinformatics / btq033.

[0298] 57. Adalsteinsson VA, Ha G, Freeman SS, Choudhury AD, Stover DG, Parsons HA, et al. Scalable whole-exome sequencing of cell-free DNA reveals high concordance with metastatic tumors. Nat Commun 2017;8(1):1324 doi 10.1038 / s41467-017-00965-y.

[0299] 58. Robin X, Turck N, Hainard A, Tiberti N, Lisacek F, Sanchez JC, et al. pROC: an open-source package for R and S+ to analyze and compare ROC curves. Bmc Bioinformatics 2011;12:77 doi 10.1186 / 1471-2105-12-77.

[0300] 59. Davoli T, Uno H, Wooten EC, Elledge SJ. Tumor aneuploidy correlates with markers of immune evasion and with reduced response to immunotherapy. Science 2017;355(6322):eaaf8399 doi 10.1126 / science.aaf8399.

[0301] 60. Bendix Carstensen M, Esa L, Michael H. Epi: A package for statistical analysis in epidemiology. R package version 2016;20.

[0302] Other embodiments

[0303] Although the present disclosure has been described in connection with its detailed description, the foregoing description is intended to illustrate and not limit the scope of the present disclosure, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.

[0304] The patents and scientific literature referred to herein establish the knowledge available to those of ordinary skill in the art. All U.S. patents and published and unpublished U.S. patent applications cited herein are incorporated by reference. All published foreign patents and patent applications cited herein are hereby incorporated by reference. All other published references, documents, manuscripts, and scientific literature cited herein are hereby incorporated by reference.

Claims

1. A method for diagnosing liver disease or disorder in a subject, the method comprising: Isolating circulating cell-free DNA (cfDNA) from the subject; Performing whole-genome sequencing of cfDNA molecules to generate a genomic library and a fragmentation profile; Comparing the fragmentation profile with that of a healthy subject, and Diagnosing whether the subject has a liver disease or disorder.

2. The method of claim 1, wherein the fragmentation profile is consistent for subjects without cancer.

3. The method of claim 1, wherein the fragmentation profile is highly variable for subjects with liver cancer.

4. The method according to any one of claims 1-3, further comprising identifying the cellular origin of the cfDNA fragmentation profile.

5. The method of claim 4, wherein identifying the cellular origin of the cfDNA fragmentation profile comprises comparing the whole-genome fragmentome profile with high-throughput sequencing chromosome conformation capture (Hi-C).

6. The method of claim 5, wherein the cellular origin of the cfDNA fragmentation profile of a healthy subject is related to lymphoblastoid-like cells as the cellular origin.

7. The method of claim 5, wherein the cellular origin of the cfDNA fragmentation profile of a subject with liver cancer is related to the cfDNA fragmentation profile of the chromatin compartment of peripheral blood cells.

8. The method according to any one of claims 1-7, further comprising determining whether the cfDNA fragmentation profile is associated with altered binding of transcription factors to DNA.

9. The method of claim 8, wherein the transcription factor DNA binding sites are determined by calculating the aggregate cfDNA coverage across all identified transcription factor DNA binding sites relative to the overall adjacent genomic coverage to generate a single metric for each transcription factor in each sample.

10. The method of claim 9, wherein the transcription factor DNA binding sites identified in subjects with liver cancer are compared with those in healthy subjects.

11. The method of claim 10, wherein the cfDNA fragmentation profile of a subject with liver cancer is associated with altered binding of transcription factors to DNA.

12. The method according to any one of claims 1-11, further comprising determining chromosomal gains or losses in subjects with liver cancer compared to healthy subjects.

13. The method according to any one of claims 1-12, further comprising a machine learning model for determining changes in the cfDNA fragmentation profile, the model classifying the subject as a cancer patient based on the cfDNA fragmentation profile of the subject.

14. The method of claim 13, wherein the machine learning model generates a score for each subject based on a combination of regional and large-scale fragmentation profiles.

15. The method of claim 14, wherein the generated score diagnoses the stage of liver cancer.

16. The method according to any one of claims 1-15, further comprising administering cancer treatment to a subject diagnosed with a liver disease or disorder.

17. The method according to claim 16, wherein the cancer treatment comprises: Surgery, adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, hormone therapy, cytotoxic therapy, immunotherapy, adoptive T cell therapy, targeted therapy, and combinations thereof.

18. A method for diagnosing liver cancer, the method comprising: Isolating circulating cell-free DNA (cfDNA) from a biological sample and performing whole-genome sequencing of cfDNA molecules to generate a genomic library and a fragmentation profile; Identifying the cellular origin of the cfDNA fragmentation profile includes comparing the whole-genome fragmentome profile with high-throughput sequencing chromosome conformation capture (Hi-C); Associating the cfDNA fragmentation profile with changes in transcription factor binding to DNA; Determining chromosomal gains or losses in liver cancer subjects compared to healthy subjects; Executing a machine learning model for determining changes in the cfDNA fragmentation profile, the model classifying the subject as a cancer patient based on the cfDNA fragmentation profile of the subject; whereby, Diagnosing liver cancer in the subject and administering cancer treatment to the diagnosed subject.

19. The method of claim 18, wherein the cellular origin of the cfDNA fragmentation profile of a healthy subject is related to lymphoblastoid-like cells as the cellular origin.

20. The method of claim 18, wherein the cellular origin of the cfDNA fragmentation profile of a subject with liver cancer is related to the cfDNA fragmentation profile of the chromatin compartment of peripheral blood cells.

21. The method of any one of claims 18-20, wherein the transcription factor DNA binding sites are determined by calculating the aggregate cfDNA coverage across all identified transcription factor DNA binding sites relative to the overall adjacent genomic coverage to generate a single metric for each transcription factor in each sample.

22. The method of claim 21, wherein the transcription factor DNA binding sites identified in a subject with liver cancer are compared with the transcription factor DNA binding sites in a healthy subject.

23. The method of claim 22, wherein the cfDNA fragmentation profile of a subject with liver cancer is associated with changes in transcription factor binding to DNA.

24. The method of any one of claims 18-23, wherein the machine learning model generates a score for each subject based on a combination of regional and large-scale fragmentation profiles.

25. The method of claim 24, wherein the generated score diagnoses the stage of liver cancer.

26. The method of claim 24, wherein the generated score diagnoses hepatocellular carcinoma (HCC).

27. The method of claim 24, wherein the generated score differentiates liver diseases including cirrhosis.

28. The method according to any one of claims 18-27, wherein the treatment comprises: Surgery, adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, hormone therapy, cytotoxic therapy, immunotherapy, adoptive T cell therapy, targeted therapy, and combinations thereof.

Citation Information

Patent Citations

  • Spray nozzle

    CA121113A

  • Tip for billiard cues

    CA233259A