Cell-free DNA for assessing and / or treating cancer

By analyzing the cfDNA fragment profiles of mammals, the problem of insufficient early diagnosis of cancer in existing technologies has been solved, and early detection and personalized treatment with high sensitivity and high specificity have been achieved.

CN120608151APending Publication Date: 2025-09-09JOHNS HOPKINS UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510688706.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-01-23
Filing Date
2019-05-17
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing cancer diagnosis and treatment methods lack effective early detection methods, resulting in most cancers being discovered in the late stages and having poor treatment effects.

Method used

By performing low-coverage whole-genome sequencing and mapping analysis of mammalian cell-free DNA (cfDNA) fragmentation profiles, comparing cfDNA fragmentation profiles with reference profiles, cancers can be identified and personalized treatments can be performed.

Benefits of technology

It achieves early detection and localization of various cancers, improves detection sensitivity and specificity, increases the chances of successful treatment, and provides a non-invasive monitoring method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005421316090000351
    Figure BDA0005421316090000351
  • Figure BDA0005421316090000361
    Figure BDA0005421316090000361
  • Figure BDA0005421316090000381
    Figure BDA0005421316090000381
Patent Text Reader

Abstract

The present invention relates to cell-free DNA for assessing and / or treating cancer, in particular, methods and materials for assessing, monitoring and / or treating mammals (e.g., humans) having cancer. For example, methods and materials are provided for identifying that a mammal has cancer (e.g., topical cancer). For example, methods and materials are provided for assessing, monitoring and / or treating mammals with cancer.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This invention is a divisional application of the PCT patent application entered into China with Chinese patent application number 201980047828.3, invention name “Cell-free DNA for evaluating and / or treating cancer”, and international application date May 17, 2019.

[0002] Related applications

[0003] This application claims priority to U.S. patent application 62 / 673,516, filed May 18, 2018, and to U.S. patent application 62 / 795,900, filed January 23, 2019, the disclosures of which are considered part of the disclosure of this application (and are incorporated herein by reference).

[0004] Government authorization

[0005] This invention was made with government support from the National Institutes of Health under Grant No. CA 121113. The U.S. Government has certain rights in this invention. Technical Field

[0006] The present invention relates to methods and materials for evaluating and / or treating mammals (e.g., humans) suffering from cancer. For example, the present invention provides methods and materials for identifying mammals suffering from cancer (e.g., localized cancer). For example, the present invention provides methods and materials for monitoring and / or treating mammals suffering from cancer. Background Art

[0007] Most of the morbidity and mortality from human cancer worldwide is the result of late diagnosis of the disease, for which treatment is less effective (Torre et al., 2015 CA Cancer J Clin 65:87; and World Health Organization, 2017 Guide to Cancer Early Diagnosis). Unfortunately, clinically validated biomarkers that can be used to diagnose and treat a wide range of patients are not widely available (Mazzucchelli, 2000 Advances in clinical pathology 4:111; Ruibal Morell, 1992 The International journal of biological markers 7:160; Galli et al., 2013 Clinical chemistry and laboratory medicine 51:1369; Sikaris, 2011 Heart, lung & circulation 20:634; Lin et al., 2016 in Screening for Colorectal Cancer: A Systematic Review for the US Preventive Services Task Force. (Rockville, MD); Wanebo et al., 1978 N Engl J Med 299:448; and Zauber, 2015 Dig Dis Sci 60:681). Summary of the Invention

[0008] Recent analyses of cell-free DNA suggest that such approaches may provide new avenues for early diagnosis (Phallen et al., 2017 Sci Transl Med 9; Cohen et al., 2018 Science 359:926; Alix-Panabiere et al., 2016 Cancer discovery 6:479; Siravegna et al., 2017 Nature reviews. Clinical oncology 14:531; Haber et al., 2014 Cancer discovery 4:650; Husain et al., 2017 JAMA 318:1272; and Wan et al., 2017 Nat Rev Cancer 17:223).

[0009] The present invention provides methods and materials for determining a cell-free DNA (cfDNA) fragmentation profile in a mammal (e.g., in a sample obtained from a mammal). In some cases, determining the cfDNA fragmentation profile in a mammal can be used to identify whether the mammal has cancer. For example, cfDNA fragments obtained from a mammal (e.g., a sample obtained from a mammal) can be subjected to low-coverage whole-genome sequencing, and the sequenced fragments can be mapped to the genome (e.g., in non-overlapping windows) and evaluated to determine the cfDNA fragmentation profile. The present invention also provides methods and materials for assessing and / or treating a mammal (e.g., a human) suffering from or suspected of having cancer. In some cases, the present invention provides methods and materials for identifying a mammal suffering from cancer. For example, a sample (e.g., a blood sample) obtained from a mammal can be assessed based on (at least in part based on) a cfDNA fragmentation profile to determine whether the mammal has cancer. In some cases, the present invention provides methods and materials for monitoring and / or treating a mammal suffering from cancer. For example, one or more cancer treatments can be given to a mammal identified as having cancer (e.g., based on or at least in part based on a cfDNA fragmentation profile) to treat the mammal.

[0010] This specification describes a non-invasive method for the early detection and localization of cancer. cfDNA (cfDNA) in the blood can provide a non-invasive diagnostic approach for cancer patients. As described herein, the DNA Evaluation for Early Intercept Fragmentation (DELFI) assay was developed and used to assess genome-wide fragmentation patterns in 236 individuals with breast, colorectal, lung, ovarian, pancreatic, gastric, or bile duct cancer, as well as 245 healthy individuals. These analyses demonstrated that the cfDNA profiles of healthy individuals mirrored the nucleosomal fragmentation profiles of white blood cells, whereas the fragmentation profiles of cancer patients were altered. DELFI demonstrated a sensitivity of 57% to >99% across seven cancer types, a specificity of 98% across all seven cancers, and in 75% of cases, identified the tissue of origin of the cancer to a few defined sites. Assessing cfDNA (e.g., using DELFI) could provide a screening method for early detection of cancer, which could increase the chances of successful treatment for cancer patients. Assessing cfDNA (e.g., using DELFI) could also provide a method for monitoring cancer, which could increase the chances of successful treatment and improve outcomes for patients with cancer. Additionally, cfDNA fragmentation profiles can be obtained from limited amounts of cfDNA using inexpensive reagents and / or instrumentation.

[0011] In general, one aspect of this specification is characterized by determining the method for the cfDNA fragmentation map of a mammal. The method may include or be essentially composed of the following: the cfDNA fragments obtained from the sample obtained from the mammal are processed into a sequencing library, the sequencing library is subjected to whole genome sequencing (e.g., low coverage whole genome sequencing) to obtain sequencing fragments, the sequencing fragments are mapped to the genome to obtain a window of a mapping sequence, and the window of the mapping sequence is analyzed to determine the cfDNA fragment length. The mapping sequence may include tens to thousands of windows. The window of the mapping sequence may be a non-overlapping window. The window of the mapping sequence may each contain approximately 5 million base pairs. The cfDNA fragment map may be determined within each window. The cfDNA fragment map may include a median fragment size. The cfDNA fragment map may include a fragment size distribution. The cfDNA fragment map may include a ratio of small cfDNA fragments to large cfDNA fragments in the mapping sequence window. The cfDNA fragment map may cover the entire genome. The cfDNA fragment map may span a subgenomic interval (e.g., an interval in a portion of a chromosome).

[0012] On the other hand, the present specification is characterized in that, it is a method for identifying a mammal with cancer. The method may include or consist essentially of the following steps: determining a cell-free DNA (cfDNA) fragmentation profile in a sample obtained from a mammal, comparing the cfDNA fragmentation profile with a reference cfDNA fragmentation profile, and identifying the mammal as having cancer when the cfDNA fragmentation profile obtained from the mammal is different from the reference cfDNA fragmentation profile. The reference cfDNA fragmentation profile may be a cfDNA fragmentation profile of a healthy mammal. The reference cfDNA fragmentation profile may be generated by determining the cfDNA fragmentation profile in a sample obtained from a healthy mammal. The reference DNA fragmentation pattern may be a fragmentation profile of a reference nucleosomal cfDNA. The cfDNA fragmentation profile may include a median fragment size, and the median fragment size of the cfDNA fragmentation profile may be shorter than the median fragment size of the reference cfDNA fragmentation profile. The cfDNA fragmentation profile may include a fragment size distribution, and the fragment size distribution of the cfDNA fragmentation profile may differ by at least 10 nucleotides compared to the fragment size distribution of the reference cfDNA fragmentation profile. The cfDNA fragmentation profile may include position-dependent differences in the fragmentation pattern, including the ratio of small cfDNA fragments to large cfDNA fragments, wherein the length of the small cfDNA fragments may be 100 base pairs (bp) to 150bp, the length of the large cfDNA fragments may be 151bp to 220bp, and the correlation of the fragment ratios in the cfDNA fragment profile may be lower than the correlation of the fragment ratios in the reference cfDNA fragment profile. The cfDNA fragment profile may include coverage sequences of small cfDNA fragments, large cfDNA fragments, or cfDNA fragments of both sizes throughout the genome. Cancer may be colorectal cancer, lung cancer, breast cancer, bile duct cancer, pancreatic cancer, gastric cancer, and ovarian cancer. The step of comparing may include comparing the cfDNA fragment profile with the reference cfDNA fragment profile in a window across the genome. The step of comparing may include comparing the cfDNA fragment profile with the reference cfDNA fragment profile within a subgenomic interval (e.g., an interval in a portion of a chromosome). The mammal may have undergone cancer therapy in advance to treat cancer. Cancer treatment can be surgery, adjuvant chemotherapy, neoadjuvant chemotherapy, radiotherapy, hormone therapy, cytotoxic therapy, immunotherapy, adoptive T cell therapy, targeted therapy or its any combination. The method can also include giving cancer treatment (such as surgery, adjuvant chemotherapy, neoadjuvant chemotherapy, radiotherapy, hormone therapy, cytotoxic therapy, immunotherapy, adoptive T cell therapy, targeted therapy or its any combination) to mammal. After giving cancer treatment, it is possible to monitor whether mammal has cancer.

[0013] In another aspect, the present invention features a method for treating a mammal suffering from cancer. The method may include or consist essentially of the following steps: identifying a mammal suffering from cancer, wherein the identification includes determining a cfDNA fragmentation profile in a sample obtained from the mammal, comparing the cfDNA fragmentation profile to a reference cfDNA fragmentation profile, and identifying the mammal as suffering from cancer when the cfDNA fragmentation profile obtained from the mammal differs from the reference cfDNA fragmentation profile; and treating the mammal for cancer. The mammal may be human. The cancer may be colorectal cancer, lung cancer, breast cancer, gastric cancer, pancreatic cancer, bile duct cancer, or ovarian cancer. The cancer treatment may be surgery, adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, hormonal therapy, cytotoxic therapy, immunotherapy, adoptive T cell therapy, targeted therapy, or a combination thereof. The reference cfDNA fragmentation profile may be a cfDNA fragmentation profile of a healthy mammal. The reference cfDNA fragmentation profile may be generated by determining a cfDNA fragmentation profile in a sample obtained from a healthy mammal. The reference DNA fragmentation pattern may be a reference nucleosomal cfDNA fragmentation profile. The cfDNA fragmentation profile may include a median fragment size, wherein the median fragment size of the cfDNA fragmentation profile is shorter than the median fragment size of the reference cfDNA fragmentation profile. The cfDNA fragmentation profile may include a fragment size distribution, wherein the fragment size distribution of the cfDNA fragmentation profile differs from the fragment size distribution of the reference cfDNA fragmentation profile by at least 10 nucleotides. The cfDNA fragmentation profile may include a ratio of small cfDNA fragments to large cfDNA fragments in a mapped sequence window, wherein the length of the small cfDNA fragments is 100 bp to 150 bp, wherein the length of the large cfDNA fragments is 151 bp to 220 bp, and the correlation of the fragment ratios in the cfDNA fragmentation profile is lower than the correlation of the fragment ratios in the reference cfDNA fragmentation profile. The cfDNA fragmentation profile may include sequence coverage of small cfDNA fragments across the entire genomic window. The cfDNA fragmentation profile may include sequence coverage of large cfDNA fragments across the entire genomic window. The cfDNA fragmentation profile may include sequence coverage of small and large cfDNA fragments across the entire genomic window. The comparing step may include comparing the cfDNA fragmentation profile with the reference cfDNA fragmentation profile across the entire genome. The comparing step can include comparing the cfDNA fragmentation profile to a reference cfDNA fragmentation profile within a subgenomic interval. The mammal can have previously received cancer therapy to treat the cancer. The cancer therapy can be surgery, adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, hormone therapy, cytotoxic therapy, immunotherapy, adoptive T cell therapy, targeted therapy, or a combination thereof. The method can also include monitoring the mammal for the presence of cancer after administering the cancer therapy.

[0014] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which the invention pertains. Although methods and materials similar to or equivalent to those described herein can be used to implement the present invention, suitable methods and materials are described below. All publications, patent applications, patents, and other references mentioned herein are incorporated herein by reference in their entirety. In the event of a conflict, the present specification (including definitions) shall control. In addition, the materials, methods, and examples are illustrative only and are not intended to be limiting.

[0015] The details of one or more embodiments of the invention are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the invention will become apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 Figure 2 is a schematic diagram of an exemplary DELFI method. Blood was collected from a cohort of healthy individuals and cancer patients. Nucleosome-protected cfDNA was extracted from the plasma fraction, processed into sequencing libraries, detected by whole-genome sequencing, mapped to the genome, and analyzed to determine cfDNA fragmentation profiles in different windows across the genome. Machine learning methods were used to classify individuals as healthy or with cancer, and the genome-wide cfDNA fragmentation patterns were used to identify the tumor tissue of origin.

[0017] Figure 2 This study simulates noninvasive cancer detection based on analyzing the number of alterations and the distribution of tumor-derived cfDNA fragments. Monte Carlo simulations were performed using varying numbers of tumor-specific alterations to estimate the probability of detecting a cancer-occurring alteration in cfDNA in a given fraction of tumor-derived molecules. Simulations were performed under the assumption that cfDNA contains an average of 2,000 genome equivalents and that five or more alterations are required to observe any alteration. These analyses suggest that increasing the number of tumor-specific alterations improves the sensitivity of circulating tumor DNA detection.

[0018] Figure 3 Figure 1 is the distribution of tumor-derived cfDNA fragments. The cumulative density function of cfDNA fragment lengths at 95% confidence levels (blue) is shown for 42 loci harboring tumor-specific alterations from 30 patients with breast, colorectal, lung, or ovarian cancer. At these loci, the lengths of mutant cfDNA fragments differ significantly compared to wild-type cfDNA fragments (red).

[0019] Figure 4A and 4B Represents the GC content and fragment length of tumor-derived cfDNA. Figure 4A It was shown that the GC content of the mutant and non-mutant fragments was similar. Figure 4B This indicates that the GC content is independent of the fragment length.

[0020] Figure 5 is the fragment distribution of germline cfDNA. Shown is the cumulative density function of fragment lengths for 44 loci harboring germline alterations (non-tumor origin) from 38 patients with breast, colorectal, lung, or ovarian cancer, at a 95% confidence level. Fragments harboring germline mutations (blue) are comparable in length to wild-type cfDNA fragment lengths (red).

[0021] Figure 6 Figure 2 is the fragment distribution of hematopoietic cfDNA. The cumulative density function of fragment lengths for 41 loci containing hematopoietic alterations (non-tumor origin) from 28 patients with breast, colorectal, lung, or ovarian cancer is shown, at a 95% confidence level. After correction for multiple testing, the size distributions of mutant hematopoietic cfDNA fragments (blue) and wild-type cfDNA fragments (red) did not differ significantly (α = 0.05).

[0022] 7A to 7F is the cfDNA fragmentation profile in healthy individuals and cancer patients. Figure 7A The whole-genome cfDNA fragmentation profiles (defined as the ratio of short to long fragments) from approximately 9x whole-genome sequencing in 30 healthy individuals (top) and 8 lung cancer patients (bottom) are shown (bottom) in 5 Mb bins. Figure 7B The fragmentation and lymphocyte profiles of chromosome 1 were analyzed at 1 Mb resolution from cfDNA of a healthy individual (top), cfDNA of a lung cancer patient (center), and a healthy lymphocyte (bottom). The healthy lymphocyte profile was measured with a standard deviation equal to the median healthy individual cfDNA profile. While the cfDNA profile of a healthy individual closely resembles that of a healthy lymphocyte, the profile of cfDNA from a lung cancer patient is more diverse and distinct from both healthy individuals and healthy lymphocytes. Figure 7C The first eigenvector of the genomic contact matrix, obtained from a previously reported Hi-C analysis of lymphoblasts (bottom), is the smoothed median distance between adjacent nucleosomes, centered at zero, using 100 kb bins of cfDNA from healthy individuals (top) and nuclease-digested healthy lymphocytes (center). The nucleosome distances in healthy cfDNA are very similar to those in nuclease-digested lymphocytes and lymphoblasts in Hi-C analysis. The fragmentation profiles of cfDNA from healthy individuals (n = 30) show a high correlation with the fragmentation profiles of lymphocytes (D), cfDNA from healthy individuals (E), and the median nucleosome distances in lymphocytes (F), while the correlation is low in patients with lung cancer.

[0023] Figure 8 is the density of cfDNA fragment length in healthy individuals and lung cancer patients. cfDNA fragment length in healthy individuals (n=30, gray) and lung cancer patients (n=8, blue) is shown.

[0024] Figure 9A and 9B It is a subsampling of whole-genome sequencing data used to analyze cfDNA fragmentation profiles. Figure 9A The high-multiplicity (9x) whole-genome sequencing data were subsampled to 2x, 1x, 0.5x, 0.2x, and 0.1x coverage. For each subsample coverage, the average genome-wide fragment profiles of 30 healthy individuals and 8 lung cancer patients in 5Mb bins are depicted, with the median plot shown in blue. Figure 9B It is the Pearson correlation between the sub-sampling maps and the initial sampling maps of healthy individuals and lung cancer patients at 9 times coverage.

[0025] Figure 10 Figure 3 is the cfDNA fragmentation profile and sequence changes during treatment. Targeted sequencing (top) and whole-genome fragmentation profiles (bottom) were used to detect and monitor cancer in the blood of a series of patients with non-small cell lung cancer (NSCLC) who were treated with targeted tyrosine kinase inhibitors (black arrows). For each case, the vertical axis of the lower figure shows -1 times the correlation of each sample with the median of the cfDNA fragmentation profile of a healthy individual. The error bars depict the confidence intervals of the mutant allele fraction from the binomial test and the confidence intervals of the whole-genome fragmentation profile calculated using the Fisher transformation. Although these methods analyze different aspects of cfDNA (whole genome versus specific alterations), targeted sequencing and fragmentation profiles are similar in patients who respond to treatment and those with stable or progressive disease. Because fragmentation profiles reflect changes in the genome and epigenomics, while mutant allele fragments reflect only individual mutations, mutant allele fragments alone may not reflect the absolute level of correlation of the fragmentation profile of a healthy individual.

[0026] Figures 11A to 11C is the cfDNA fragmentation profile in healthy individuals and cancer patients. Figure 11A Figure 2. Segment profiles (bottom) in the context of tumor copy number variation (top) from parallel analysis of tumor tissue in colorectal cancer patients. The segment mean and the distribution of integer copy number are shown in the upper right corner in the indicated colors. Altered segment profiles are present in copy-neutral regions of the genome and are further affected in regions of copy number variation. Figure 11BGC-adjusted fragment profiles from 1-2x whole-genome sequencing for healthy individuals and patients with cancer are depicted for each cancer type using a 5-Mb window. The median profile for healthy individuals is shown in black, while the 98% confidence interval is shown in gray. For patients with cancer, individual profiles are colored according to their relative relative to the median for healthy individuals. Figure 11C , if more than 10% of cancer patient samples have fragment ratios that are more than three standard deviations away from the median healthy individual fragment ratio, the window is indicated in orange.These analyses highlight alterations associated with numerous locations throughout the cfDNA genome of cancer individuals.

[0027] Figure 12A and Figure 12B Figure 3. cfDNA fragment length profiles in copy-neutral regions in healthy individuals and a patient with colorectal cancer. Figure 12A Figure 2: Segment profiles of 25 randomly selected healthy individuals (grey) in 211 copy-neutral windows on chromosomes 1 to 6. For a patient with colorectal cancer (CGCRC291) with an estimated mutant allele prevalence of 20%, the cancer segment length profile was diluted to approximately 10% tumor contribution (blue). Figure 12A and Figure 12B , although the marginal densities of the fragment profiles of healthy samples and cancer patients showed substantial overlap ( Figure 12A , right), but as the fragment map visualized ( Figure 12A , left), the fragment profiles were different, and samples from colorectal cancer patients and healthy individuals were separated in principal component analysis ( Figure 12B ).

[0028] Figure 13A and 13B is the genome-wide GC correction of cfDNA fragments. To estimate and control for the effect of GC content on sequencing coverage, coverage was calculated in non-overlapping 100 kb genomic windows across autosomes. For each window, the average GC content of the aligned fragments was calculated. Figure 13A, Loess-smoothed raw coverage (upper row) of aneuploidy detection (PA score <2.35) in two randomly selected healthy subjects (CGPLH189 and CGPLH380) and two cancer patients (CGPLLU161 and CGPLBR24). After subtracting the mean coverage predicted by the Loess model, the residuals were rescaled to the median autosomal coverage (lower row). Because the length of the fragment can also lead to coverage bias, this GC correction procedure was performed separately for short fragments (≤150bp) and long fragments (≥151bp). Although the 100kb bin on chromosome 19 (blue dot) always had less coverage than predicted by the Loess model, we did not implement chromosome-specific correction because this approach would eliminate the effect of chromosome copy number on coverage. Figure 13B Overall, there was limited correlation between corrected short or long fragment coverage and GC content in healthy subjects and cancer patients with PA scores < 3.

[0029] Figure 14 Figure 1 is a schematic diagram of a machine learning model. Gradient tree-enhanced machine learning was used to examine whether cfDNA could be classified as having features of a cancer patient or a healthy individual. The machine learning model included features of fragment size and coverage across the entire genome window, as well as chromosome arm and mitochondrial DNA copy number. A 10-fold cross-validation approach was used to randomly assign each sample to a fold, with 9 folds (90% of the data) used for training and one fold (10% of the data) used for testing. The prediction accuracy of a single cross-validation was the average of the 10 possible combinations of the test and training sets. Because this prediction accuracy could reflect bias in the initial randomization of the patients, the entire process was repeated, including randomizing the patients 10 times. For all cases, feature selection and model estimation were performed on the training data and validated on the test data, and the test data was never used for feature selection. Ultimately, a DELFI score was obtained that can be used to classify individuals as likely healthy or having cancer.

[0030] Figure 15 is the AUC distribution in repeated 10-fold cross validation. The dashed lines indicate the 25th, 50th, and 75th percentiles of 100 AUCs in a cohort of 215 healthy individuals and 208 cancer patients.

[0031] Figure 16A and 16B is a genome-wide analysis of chromosome arm copy number changes and mitochondrial genome representation. Figure 16AThe Z-scores for each autosomal arm are depicted for healthy individuals (n=215) and cancer patients (n=208). The vertical axis represents normal copies at zero, with positive and negative values ​​representing increases and decreases in the arm, respectively. Z-scores greater than 50 or less than -50 are thresholded at that indicator value. Figure 16B Depicted are the fractions of sequenced reads that map to the mitochondrial genome for healthy individuals and cancer patients.

[0032] Figure 17A and 17B , using DELFI to detect cancer. Figure 17A , for a cohort of 215 healthy individuals and 208 cancer patients (DELFI, AUC = 0.94), cfDNA fragmentation profiles and other genome-wide features were used in a machine learning approach as receiver operator characteristics for detecting cancer, with a specificity of ≥ 95% indicated in blue shading. Machine learning analyses of chromosome arm copy number (Chr copy number (ML)) and mitochondrial genome copy number (mtDNA) are shown in the indicated colors. Figure 17B , the AUCs for analyzing individual cancer types using the combined DELFI approach ranged from 0.86 to >0.99.

[0033] Figure 18 DELFI detects cancer by stage. cfDNA fragmentation profiles and other genome-wide features were used as receiver operating characteristics for detecting cancer in a machine learning approach, depicting a cohort of 215 healthy individuals and 208 cancer patients at each stage, with ≥95% specificity indicated in blue shading.

[0034] Figure 19 , DELFI tissue-of-origin prediction. This study describes the receiver operating characteristics (ROI) of DELFI tissue prediction for cholangiocarcinoma, breast cancer, colorectal cancer, gastric cancer, lung cancer, ovarian cancer, and pancreatic cancer. To increase the sample size within each cancer type category, cases detected at 90% specificity were also included, and the lung cancer cohort was supplemented with baseline cfDNA data from 18 previously treated lung cancer patients (see, e.g., Shen et al., 2018 Nature, 563:579–583).

[0035] Figure 20Cancer detection using DELFI and mutation-based cfDNA methods. In a cohort of 126 patients with breast, bile duct, colorectal, gastric, lung, or ovarian cancer, DELFI (green) and targeted sequencing (blue) were performed for mutation identification. For the DELFI assay, the number of individuals detected by each method and by the combined methods was 98% specific, >99% specific for targeted sequencing, and 98% for the combined method. ND indicates not detected. DETAILED DESCRIPTION

[0036] The present invention provides methods and materials for determining the cfDNA fragmentation profile of a mammal (e.g., a sample obtained from a mammal). As used herein, the terms "fragmentation profile," "position-dependent differences in fragmentation patterns," and "differences in fragment size and coverage in a position-dependent manner across the genome" are equivalent and can be used interchangeably. In some cases, determining the cfDNA fragmentation profile in a mammal can be used to identify a mammal with cancer. For example, cfDNA fragments obtained from a mammal (e.g., a sample obtained from a mammal) can be subjected to low-coverage whole-genome sequencing, and the sequenced fragments can be mapped to the genome (e.g., in non-overlapping windows) and evaluated to determine the cfDNA fragmentation profile. As described in this specification, the cfDNA fragmentation profile of a mammal with cancer is more heterogeneous (e.g., fragment length) than the cfDNA fragmentation profile of a healthy mammal (e.g., a mammal without cancer). Therefore, this specification also provides methods and materials for assessing, monitoring, and / or treating a mammal (e.g., a human) suffering from or suspected of having cancer. In some cases, this specification provides methods and materials for identifying mammals with cancer. For example, a sample (e.g., a blood sample) taken from a mammal can be evaluated based at least in part on a mammal's cfDNA fragment analysis to determine the presence of cancer in the mammal and optionally determine the tissue of origin of the cancer. In some cases, this specification provides methods and materials for monitoring a mammal with cancer. For example, a sample (e.g., a blood sample) obtained from a mammal can be evaluated based at least in part on a mammal's cfDNA fragment profile to determine the presence of cancer in the mammal. In some cases, this specification provides methods and materials for identifying a mammal with cancer and performing one or more cancer treatments on the mammal to treat the mammal. For example, a sample (e.g., a blood sample) obtained from a mammal can be evaluated to determine whether the mammal has cancer based at least in part on a mammal's cfDNA fragment profile, and one or more cancer treatments can be given to the mammal.

[0037] The cfDNA fragmentation pattern may include one or more cfDNA fragmentation patterns. A cfDNA fragmentation pattern may include any appropriate cfDNA fragmentation pattern. The example of a cfDNA fragmentation pattern includes but is not limited to the coverage of median fragment size, fragment size distribution, the ratio of small cfDNA fragments to large cfDNA fragments and cfDNA fragments. In some cases, the cfDNA fragmentation pattern includes two or more (e.g., two, three or four) median fragment size, fragment size distribution, the ratio of small cfDNA fragments to large cfDNA fragments and cfDNA fragments. In some cases, the cfDNA fragmentation pattern may be a full genome cfDNA pattern (e.g., a full genome cfDNA pattern across a genome window). In some cases, the cfDNA fragmentation pattern may be a target region pattern. The targeted region may be any appropriate portion of a genome (e.g., a chromosome region). Examples of chromosome regions that can be determined by cfDNA fragmentation maps described herein may include, but are not limited to, a portion of a chromosome (e.g., 2q, 4p, 5p, 6q, 7p, 8q, 9q, 10q, 11q, a portion of 12q and / or 14q) and a chromosome arm (e.g., 8q, 13q, 11q, and / or a chromosome arm of 3p). In some cases, a cfDNA fragmentation map may include two or more target region maps.

[0038] In some cases, the cfDNA fragmentation profile can be used to identify changes (e.g., alterations) in the length of cfDNA fragments. The alterations can be whole genome alterations or alterations in one or more targeted regions / positions. The target region can be any region containing one or more cancer-specific alterations. Examples of cancer-specific alterations and their chromosomal locations include, but are not limited to, those shown in Table 3 (Appendix C) and Table 6 (Appendix F). In some cases, the cfDNA fragmentation profile can be used to identify (e.g., simultaneously identify) from about 10 changes to about 500 changes (e.g., from about 25 to about 500, from about 50 to about 500, from about 100 to about 500, about 200 to about 500, about 300 to about 500, about 10 to about 400, about 10 to about 300, about 10 to about 200, about 10 to about 100, about 10 to about 50, about 20 to about 400, about 30 to about 300, about 40 to about 200, about 50 to about 100, about 20 to about 100, about 25 to about 75, about 50 to about 250, or about 100 to 200, etc.).

[0039] In some cases, cfDNA fragmentation profiles can be used to detect tumor-derived DNA. For example, tumor-derived DNA is detected by comparing the cfDNA fragmentation profiles of a mammal suffering from or suspected of having cancer with a reference cfDNA fragmentation profile (e.g., cfDNA fragmentation profiles of healthy mammals and / or nucleosomal DNA fragmentation profiles of healthy cells from a mammal suffering from or suspected of having cancer). In some cases, the reference cfDNA fragmentation profile is a profile previously generated from a healthy mammal. For example, the method provided herein can be used to determine a reference cfDNA fragmentation profile in a healthy mammal, and the reference cfDNA fragmentation profile can be stored (e.g., in a computer or other electronic storage medium) for future comparison with a test cfDNA fragmentation profile of a mammal suffering from or suspected of having cancer. In some cases, a reference cfDNA fragmentation profile (e.g., a stored cfDNA fragmentation profile) of a healthy mammal is determined across the entire genome. In some cases, a reference cfDNA fragmentation profile (e.g., a stored cfDNA fragmentation profile) of a healthy mammal is determined within a subgenomic interval.

[0040] In certain instances, cfDNA fragmentation profiles can be used to identify mammals (e.g., humans) having cancer (e.g., colorectal cancer, lung cancer, breast cancer, gastric cancer, pancreatic cancer, bile duct cancer, and / or ovarian cancer).

[0041] The cfDNA fragment distribution graph can include a cfDNA fragment size pattern. The cfDNA fragment can be any suitable size. For example, the length of the cfDNA fragment can be about 50 base pairs (bp) to about 400bp. As described in this specification, a mammal suffering from cancer can have a cfDNA fragment size pattern comprising a median cfDNA fragment size shorter than the cfDNA fragment median in a healthy mammal. A healthy mammal (e.g., a mammal without cancer) can have a cfDNA fragment size, and its cfDNA fragment median size is about 166.6bp to about 167.2bp (e.g., about 166.9bp). In some cases, the cfDNA fragment size of a mammal suffering from cancer is on average about 1.28bp to about 2.49bp (e.g., about 1.88bp) shorter than the cfDNA fragment size in a healthy mammal. For example, a mammal suffering from cancer can have a cfDNA fragment size, and its cfDNA fragment median size is about 164.11bp to about 165.92bp (e.g., about 165.02bp).

[0042] The cfDNA fragmentation profile may include a cfDNA fragment size distribution. As described herein, a mammal with cancer may have a more variable cfDNA size distribution than the cfDNA fragment size distribution in a healthy mammal. In some cases, the size distribution may be located within a targeted region. The targeted region cfDNA fragment size distribution of a healthy mammal (e.g., a mammal without cancer) may be about 1 or less than about 1. In some cases, the target region cfDNA fragment size distribution of a mammal with cancer may be longer (e.g., 10, 15, 20, 25, 30, 35, 40, 45, 50bp or longer base pairs, or any base pairs between these numbers) than the target region cfDNA fragment size distribution in a healthy mammal. In some cases, the targeted region cfDNA fragment size distribution of a mammal with cancer may be shorter (e.g., 10, 15, 20, 25, 30, 35, 40, 45, 50 or shorter base pairs, or any base pairs between these numbers) than the target region cfDNA fragment size distribution in a healthy mammal. In some cases, the target region cfDNA fragment size distribution of a mammal with cancer is shorter than about 47bp and longer than about 30bp than the target region cfDNA fragment size distribution in a healthy mammal. In some cases, the target region cfDNA fragment size distribution average length difference of a mammal with cancer is 10, 11, 12, 13, 14, 15, 15, 17, 18, 19, 20bp or more. For example, the size distribution average length of the target region cfDNA fragments that a mammal with cancer may have differs by about 13bp. In some cases, the size distribution can be a genome-wide size distribution. Healthy mammals (e.g., mammals without cancer) are very similar in the distribution of long and short cfDNA fragments in the whole genome. In some cases, a mammal with cancer can have one or more changes (e.g., increases and decreases) in the size of cfDNA fragments in the whole genome. One or more changes can be any suitable chromosome region of the genome. For example, the change can be in a part of a chromosome. Examples of portions of chromosomes that may contain one or more changes in cfDNA fragment size include, but are not limited to, portions of 2q, 4p, 5p, 6q, 7p, 8q, 9q, 10q, 11q, 12q, and 14q. For example, the change may span a chromosome arm (e.g., an entire chromosome arm).

[0043] The cfDNA fragment distribution graph may include the ratio of small cfDNA fragments to large cfDNA fragments and the correlation of the fragment ratio to the reference fragment ratio. As used in the present invention, with respect to the ratio of small cfDNA fragments to large cfDNA fragments, the length of the small cfDNA fragments may be from about 100bp to about 150bp. As used in the present invention, with respect to the ratio of small cfDNA fragments to large cfDNA fragments, the length of the large cfDNA fragments may be from about 151bp to 220bp. As described herein, mammals with cancer may have a lower fragment ratio (e.g., 2 times lower, 3 times lower, 4 times lower, 5 times lower, 6 times lower, 7 times lower, 8 times lower, 9 times lower, 10 times lower or more) than healthy mammals (e.g., the correlation of the cfDNA fragment ratio with the reference DNA fragment ratio from one or more healthy mammals). Healthy mammals (e.g., mammals without cancer) may have a fragment ratio correlation of about 1 (e.g., about 0.96) (e.g., the correlation of the cfDNA fragment ratio with, for example, the reference DNA fragment ratio from one or more healthy mammals). In some cases, the fragmentation ratio correlation (e.g., correlation of cfDNA fragmentation ratios with, e.g., fragmentation ratios of a reference DNA from one or more healthy mammals) in mammals having cancer is, on average, about 0.19 to about 0.30 (e.g., about 0.25) lower than the correlation of fragmentation ratios in healthy mammals (e.g., correlation of cfDNA fragmentation ratios with, e.g., fragmentation ratios of a reference DNA from one or more healthy mammals).

[0044] The cfDNA fragmentation map may include coverage of all fragments. The coverage of all fragments may include a window of coverage (e.g., a non-overlapping window). In some cases, the coverage of all fragments may include a window of small fragments (e.g., a fragment with a length of about 100bp to about 150bp). In some cases, the coverage of all fragments may include a window of large fragments (e.g., a fragment with a length of from about 151bp to about 220bp).

[0045] In some cases, the cfDNA fragmentation profile can be used to identify the tissue of origin of a cancer (e.g., colorectal cancer, lung cancer, breast cancer, gastric cancer, pancreatic cancer, bile duct cancer, or ovarian cancer). For example, the cfDNA fragmentation profile can be used to identify localized cancers. When the cfDNA fragmentation profile includes a targeted region profile, one or more changes described in this specification (e.g., Table 3 (Appendix C) and / or Table 6 (Appendix F)) can be used to identify the tissue of origin of the cancer. In some cases, one or more changes in a chromosomal region can be used to identify the tissue of origin of the cancer.

[0046] Any appropriate method can be used to obtain a cfDNA fragmentation profile. In some cases, cfDNA from a mammal (e.g., a mammal having cancer or suspected of having cancer) can be processed into a sequencing library, which is then subjected to whole genome sequencing (e.g., low coverage whole genome sequencing) and mapped to the genome and analyzed to determine the cfDNA fragment lengths. The mapped sequences can be analyzed in non-overlapping windows covering the genome. The windows can be of any suitable size. For example, the length of the window can be from thousands to millions of bases. As a non-limiting example, a window can be about 5 megabases (Mb) long. Any number of windows can be mapped. For example, tens to thousands of windows can be mapped in the genome. For example, hundreds to thousands of windows can be mapped in the genome. The cfDNA fragment profile can be determined within each window. In some cases, a cfDNA fragment profile can be obtained as described in Example 1. In some cases, the cfDNA fragment profile can be obtained as described in Example 1. Figure 1 The cfDNA fragmentation profile was obtained as shown.

[0047] In some cases, the methods and materials described herein can also include machine learning. For example, machine learning can be used to identify altered fragment profiles (e.g., using coverage of cfDNA fragments, fragment size of cfDNA fragments, coverage of chromosomes, and mtDNA).

[0048] In some cases, the methods and materials described herein can be a single method for identifying a mammal (e.g., a human) having cancer (e.g., colorectal cancer, lung cancer, breast cancer, gastric cancer, pancreatic cancer, bile duct cancer, and / or ovarian cancer). For example, determining a cfDNA fragmentation profile can be a single method for identifying a mammal having cancer.

[0049] In some cases, the methods and materials described herein can be used together with one or more other methods for identifying mammals (e.g., humans) suffering from cancer (e.g., colorectal cancer, lung cancer, breast cancer, gastric cancer, pancreatic cancer, bile duct cancer, and / or ovarian cancer). Examples of methods for identifying mammals suffering from cancer include, but are not limited to, identifying one or more cancer-specific (cancer-specific) sequence changes, identifying one or more chromosome changes (e.g., aneuploidies and rearrangements), and identifying other cfDNA changes. For example, determining that a cfDNA fragmentation profile can be used together with identifying one or more cancer-specific mutations in a mammalian genome to identify a mammal suffering from cancer. For example, determining that a cfDNA fragmentation profile can be used together with identifying one or more aneuploidies in a mammalian genome to identify a mammal suffering from cancer.

[0050] In certain aspects, the present specification also provides methods and materials for assessing, monitoring and / or treating mammals (e.g., humans) that have or are suspected of having cancer. In some cases, the present specification provides methods and materials for identifying mammals that have cancer. For example, a sample (e.g., a blood sample) obtained from a mammal can be assessed to determine whether the mammal has cancer based at least in part on the cfDNA fragments of the mammal. In some cases, the present specification provides methods and materials for identifying the location (e.g., anatomical site or tissue of origin) of cancer in a mammal. For example, a sample (e.g., a blood sample) obtained from a mammal can be assessed based at least in part on the cfDNA fragment profile of the mammal to determine the tissue of origin of the cancer in the mammal. In some cases, the present invention provides methods and materials for identifying a mammal that has cancer and performing one or more cancer treatments on the mammal to treat the mammal. For example, a sample (e.g., a blood sample) obtained from a mammal can be assessed to determine whether the mammal has cancer based at least in part on the cfDNA fragment profile of the mammal, and one or more cancer treatments can be performed on it. In some cases, the present specification provides methods and materials for treating a mammal that has cancer. For example, a mammal identified as having cancer can be subjected to one or more cancer treatments (e.g., based at least in part on the cfDNA fragmentation profile of the mammal) to treat the mammal. In some cases, during or after a cancer treatment (e.g., any cancer treatment described herein), the mammal may be monitored (or selected for increased monitoring) and / or further diagnostic testing. In some cases, monitoring may include assessing a mammal having or suspected of having cancer by, for example, assessing a sample (e.g., a blood sample) obtained from the mammal to determine the cfDNA fragmentation of the mammal as described herein. Changes in the cfDNA fragmentation profile over time can be used to identify a response to treatment and / or identify a mammal having cancer (e.g., a residual cancer).

[0051] Any suitable mammal can be assessed, monitored, and / or treated as described herein. The mammal can be a mammal suffering from cancer. The mammal can be a mammal suspected of suffering from cancer. Examples of mammals that can be assessed, monitored, and / or treated as described herein include, but are not limited to, humans, primates such as monkeys, dogs, cats, horses, cattle, pigs, sheep, mice, and rats. For example, a human suffering from or suspected of suffering from cancer can be assessed to determine if they have a cfDNA fragmentation profile as described herein, and optionally, they can be treated with one or more cancer treatment methods described herein.

[0052] Any suitable sample from a mammal can be assessed as described herein (e.g., to assess DNA fragmentation patterns). In some cases, the sample can comprise DNA (genomic DNA). In some cases, the sample can comprise cfDNA (e.g., circulating tumor DNA (ctDNA)). In some cases, the sample can be a fluid sample (e.g., a liquid biopsy). Examples of samples that can contain DNA and / or polypeptides include, but are not limited to, blood (e.g., whole blood, serum, or plasma), amniotic membrane, tissue, urine, cerebrospinal fluid, saliva, sputum, bronchoalveolar lavage fluid, bile, lymphatic fluid, cyst fluid, feces, ascites, cervical smear, breast milk, and exhaled breath condensate. For example, a plasma sample can be assessed to determine a cfDNA fragmentation profile as described herein.

[0053] As described herein, a sample from a mammal to be evaluated (e.g., to evaluate DNA fragmentation patterns) may include any suitable amount of cfDNA. In some cases, the sample may contain a limited amount of DNA. For example, a cfDNA fragmentation profile may be obtained from a sample containing less DNA than is typically required for other cfDNA analysis methods. For example, as described in Phallen et al., 2017 Sci Transl Med 9; Cohen et al., 2018 Science 359:926; Newman et al., 2014 Nat Med 20:548; and Newman et al., 2016 Nat Biotechnol 34:547.

[0054] In some cases, the sample can be processed (e.g., to isolate and / or purify DNA and / or polypeptides from the sample). For example, DNA isolation and / or purification can include cell lysis (e.g., using detergents and / or surfactants), protein removal (e.g., using proteases), and / or RNA removal (e.g., using RNases). As another example, polypeptide isolation and / or purification can include cell lysis (e.g., using detergents and / or surfactants), DNA removal (e.g., using DNases), and / or RNA removal (e.g., using RNases).

[0055] The methods and materials described herein can be used to assess (e.g., determine cfDNA fragmentation profiles) mammals suffering from (or suspected of having) any appropriate type of cancer and / or treat (e.g., by subjecting the mammal to one or more cancer treatments). The cancer can be cancer at any stage. In some cases, the cancer can be early-stage cancer. In some cases, the cancer can be asymptomatic cancer. In some cases, the cancer can be residual disease and / or recurrence (e.g., after surgical resection and / or after cancer treatment). The cancer can be any type of cancer. Examples of cancer types that can be assessed, monitored, and / or treated as described herein include, but are not limited to, colorectal cancer, lung cancer, breast cancer, gastric cancer, pancreatic cancer, bile duct cancer, and ovarian cancer.

[0056] When treating a mammal suffering from or suspected of having cancer as described herein, one or more cancer treatments can be performed on the mammal.Cancer treatment can be any appropriate cancer treatment.One or more cancer treatments described in this specification can be administered to a mammal at any suitable frequency (e.g., once or multiple times over a period of several days to weeks).Examples of cancer treatments include, but are not limited to, adjuvant chemotherapy, neoadjuvant chemotherapy, radiotherapy, hormone therapy, cytotoxic therapy, immunotherapy, adoptive T cell therapy (e.g., chimeric antigen receptors and / or T cells with wild-type or modified T cell receptors), targeted therapy such as administering kinase inhibitors (e.g., kinase inhibitors targeting specific genetic lesions, such as translocations or mutations), (e.g., kinase inhibitors, antibodies, bispecific antibodies), signal transduction inhibitors, bispecific antibodies or antibody fragments (e.g., BiTEs), monoclonal antibodies, immune checkpoint inhibitors, surgery (e.g., surgical resection) or any combination thereof. In some cases, cancer treatment can reduce the severity of cancer, alleviate the symptoms of cancer and / or reduce the number of cancer cells present in a mammal.

[0057] In some cases, cancer treatment may include immune checkpoint inhibitors. Non-limiting examples of immune checkpoint inhibitors include nikolamonab (Opdivo), pembrolizumab (Keytruda), atuzumab (tecentriq), avaxinumab (bavencio), durvalumab (imfinzi), ipilimumab (yervoy). See, for example, Pardoll (2012) Nat. Rev Cancer 12:252-264; Sun et al. (2017) Eur Rev Med Pharmacol Sci 21(6):1198-1205; Hamanishi et al. (2015) J. Clin. Oncol. 33(34):4015-22; Brahmer et al. (2012) N Engl J Med 366(26):2455-65; Ricciuti et al. (2017) J. Thorac Oncol. 12 (5): e51-e55; Ellis et al. (2017) Clin Lung Cancer pii: S1525-7304 (17) 30043-8; Zou and Awad (2017) Ann Oncol 28(4):685-687;Sorscher(2017)N Engl J Med 376(10:996-7;Huiet al.(2017)Ann Oncol 28(4):874-881;Vansteenkiste et al.(2017)Expert OpinBiolTher 17(6):781-789;Hellmann et al. al.(2017)Lancet Oncol.18(1):31-41; Chen(2017)J.Chin Med Assoc 80(1):7-14.

[0058] In some cases, cancer treatment can be adoptive T cell therapy (e.g., chimeric antigen receptors and / or T cells with wild-type or modified T cell receptors). See, for example, Rosenberg and Restifo (2015) Science 348(6230):62-68; Chang and Chen (2017) Trends Mol Med 23(5):430-450; Yee and Lizee (2016) Cancer J. 23(2):144-148; Chen et al. (2016) Oncoimmunology 6(2):e1273302; US2016 / 0194404; US2014 / 0050788; US2014 / 0271635; US 9,233,125; herein incorporated by reference in their entirety.

[0059] In some cases, the cancer treatment may be a chemotherapeutic agent.Non-limiting examples of chemotherapeutic agents include amlodipine, azacitidine, axathioprine, bevacizumab (or an antigen-binding fragment thereof), bleomycin, busulfan, carboplatin, capecitabine, chlorambucil, cisplatin, cyclophosphamide, cytarabine, dacarbazine, daunorubicin, docetaxel, doxifluridine hydrochloride, doxorubicin, epirubicin, erlotinib hydrochloride, and dapoxetine. hydrochlorides, etoposide, fiudarabine, floxuridine, fludarabine, fluorouracil, gemcitabine, hydroxyurea, idarubicin, ifosfamide, irinotecan, lomustine, mechlorethamine, melphalan, mercaptopurine, methotrexate, mitomycin, mitoxantrone, oxaliplatin, paclitaxel, pemetrexed, procarbazine, all-trans retinoic acid acid, streptozocin, tafluposide, temozolomide, teniposide, tioguanine, topotecan, uramustine, valrubicin, vinblastine, vincristine, vindesine, vinorelbine, and combinations thereof.Other examples of anticancer therapies are known in the art.See, for example, the treatment guidelines of the American Society of Clinical Oncology (ASCO), the European Society of Medical Oncology (ESMO), or the National Comprehensive Cancer Network (NCCN).

[0060] When monitoring a mammal suffering from or suspected of having a cancer as described herein (e.g., based at least in part on the mammal's cfDNA fragmentation profile), monitoring can be performed before, during, and / or after the cancer treatment process. The monitoring methods provided herein can be used to determine the efficacy of one or more cancer treatments and / or select mammals to enhance monitoring. In some cases, monitoring may include identifying a cfDNA fragmentation profile as described herein. For example, a cfDNA fragmentation profile can be obtained before administering one or more cancer treatments to a mammal suffering from or suspected of having cancer, one or more cancer treatments can be administered to the mammal, and one or more cfDNA fragmentation profiles can be obtained during the cancer treatment process of the mammal. In some cases, the cfDNA fragmentation profile can change during the course of cancer treatment (e.g., any cancer treatment described herein). For example, a cfDNA fragmentation profile indicating that a mammal has cancer can be changed to a cfDNA fragmentation profile indicating that the mammal does not have cancer. Such changes in cfDNA fragmentation profiles may indicate that the cancer treatment is working. In contrast, during cancer treatment (e.g., any cancer treatment described herein), the cfDNA fragmentation profile may remain static (e.g., the same or approximately the same). Such a static cfDNA fragmentation profile may indicate that the cancer treatment is ineffective. In some cases, monitoring can include the conventional techniques that can monitor one or more cancer treatments (for example, the effect of one or more cancer treatments).In some cases, compared with the mammal not yet selected for increasing monitoring, the frequency that the mammal selected for increasing monitoring can increase is carried out diagnostic test (for example, any diagnostic test disclosed in this specification sheets).For example, the mammal selected for increasing monitoring can be carried out diagnostic test with twice a day, once a day, once every two weeks, once a week, once every two months, once a month, once a quarter, once every six months, once a year, or any frequency therein.In some cases, compared with the mammal not yet selected for enhancing monitoring, one or more other diagnostic tests can be carried out to the mammal selected for enhancing monitoring.For example, two diagnostic tests can be carried out to the mammal selected for enhancing monitoring, and the mammal not yet selected for enhancing monitoring only carries out single diagnostic test (or does not carry out diagnostic test).In some cases, the mammal selected for enhancing monitoring can be selected for further diagnostic test. Once the presence of a tumor or cancer (e.g., cancer cells) has been identified (e.g., by any of the various methods disclosed herein), it may be beneficial to perform enhanced monitoring of the mammal (e.g., assessing the progression of the tumor or cancer in the mammal and / or assessing the development of one or more cancer biomarkers (e.g., mutations)) and perform further diagnostic testing (e.g., determining the size and / or exact location (e.g., tissue of origin) of the tumor or cancer).In some cases, one or more cancer treatments may be performed on a mammal selected for enhanced monitoring after cancer biomarkers are detected and / or after the mammal's cfDNA fragmentation profile has not improved or worsened. Any cancer treatment disclosed in this specification or known in the art may be given. For example, a mammal selected for increased monitoring may be further monitored, and if the persistence of cancer cells persists throughout the enhanced monitoring period, cancer treatment may be performed. Additionally or alternatively, a mammal selected for increased monitoring may be treated for cancer and further monitored as the cancer treatment proceeds. In some cases, after cancer treatment is performed on a mammal selected for enhanced monitoring, enhanced monitoring will reveal one or more cancer biomarkers (e.g., mutations). In some cases, such one or more cancer biomarkers will provide a basis for performing different cancer treatments (e.g., drug-resistant mutations may be generated in cancer cells during cancer treatment, and such cancer cells with drug-resistant mutations are resistant to the original cancer treatment).

[0061] When a mammal is identified as having cancer as described herein (e.g., at least in part based on a mammal's cfDNA fragmentation profile), the identification can be performed before and / or during cancer treatment. The method for identifying a mammal having cancer provided by the present invention can be used as a first diagnosis for identifying the mammal (e.g., having cancer before any treatment process) and / or selecting the mammal for further diagnostic testing. In some cases, once it is determined that a mammal has cancer, the mammal can be further examined and / or selected for further diagnostic testing. In some cases, the method provided herein can be used to select a mammal so that a certain period before the period before conventional technology can diagnose a mammal with early-stage cancer is further diagnosed. For example, when a mammal is not diagnosed with cancer by conventional methods and / or when a mammal does not have cancer, the method for further diagnostic testing of a selected mammal provided by this specification can be used. In some cases, a mammal selected for further diagnostic testing can be subjected to diagnostic testing (e.g., any diagnostic test disclosed in this specification) at an increased frequency compared to a mammal that has not been selected for further diagnostic testing. For example, a mammal selected for further diagnostic testing may undergo diagnostic testing twice a day, daily, biweekly, weekly, bimonthly, monthly, quarterly, semi-annually, annually, yearly, or any frequency thereof. In some cases, a mammal selected for further diagnostic testing may be subjected to one or more additional diagnostic tests compared to a mammal that has not yet been selected for further diagnostic testing. For example, a mammal selected for further diagnostic testing may undergo two diagnostic tests, while a mammal that has not yet been selected for further diagnostic testing undergoes only a single diagnostic test (or no diagnostic test). In some cases, a diagnostic test method may determine the presence of the same type of cancer (e.g., of the same tissue or origin) as the cancer that was initially detected (e.g., based at least in part on the mammal's cfDNA fragmentation profile). Additionally or alternatively, a diagnostic test method may determine the presence of a different type of cancer than the cancer that was initially detected. In some cases, the diagnostic test method is a scan. In some cases, the scan is computed tomography (CT), CT angiography (CTA), esophagography (barium swallow), barium enema, magnetic resonance imaging (MRI), PET scan, ultrasound (e.g., endobronchial ultrasound, endoscopic ultrasound), X-ray, or DEXA scan.In some cases, the diagnostic test method is a physical examination, such as anoscopy, bronchoscopy (e.g., autofluorescence bronchoscopy, white light bronchoscopy, navigational bronchoscopy), colonoscopy, digital mammography, endoscopic retrograde cholangiopancreatography (ERCP), endoscopy, duodenoscopy, nipple smear, pelvic examination, positron emission tomography and computed tomography (PET-CT) scan. In some cases, the mammal that has been selected for further diagnostic testing can also be selected to enhance monitoring. Once the presence of a tumor or cancer (e.g., cancer cells) has been identified (e.g., by any of the various methods disclosed herein), it may be beneficial to perform enhanced monitoring of the mammal (e.g., to assess the progression of the tumor or cancer in the mammal and / or to assess the development of one or more cancer biomarkers (e.g., mutations)) and perform further diagnostic testing (e.g., to determine the size and / or exact location of the tumor or cancer). In some cases, after the detection of cancer biomarkers and / or after the cfDNA fragmentation profile of the mammal has not improved or worsened, cancer treatment is performed on the mammal selected for further diagnostic testing. Any cancer treatment disclosed in this specification or known in the art can be given. For example, further diagnostic tests can be performed on a mammal that has been selected for further diagnostic testing, and if the presence of a tumor or cancer is confirmed, cancer treatment can be performed. Additionally or alternatively, cancer treatment can be performed on a mammal that has been selected for further diagnostic testing, and it can be further monitored as the cancer treatment progresses. In some cases, after cancer treatment is performed on a mammal that has been selected for further diagnostic testing, other tests will reveal one or more cancer biomarkers (e.g., mutations). In some cases, one or more cancer biomarkers (e.g., mutations) will become the basis for implementing different cancer treatments (e.g., cancer cells may develop drug-resistant mutations during cancer treatment, and cancer cells with drug-resistant mutations are resistant to the initial cancer treatment).

[0062] The present invention will be further described in the following examples, which do not limit the scope of the invention described in the claims.

[0063] [Example]

[0064] Example 1: Fragmentation of Cell-Free DNA from Cancer Patients

[0065] Analysis of cell-free DNA has primarily focused on targeted sequencing of specific genes. While such studies can detect a small number of tumor-specific alterations in cancer patients, not all patients, particularly those with early-stage disease, have detectable alterations. Whole-genome sequencing of cell-free DNA can identify chromosomal abnormalities and rearrangements in cancer patients, but detecting such alterations has been challenging, in part due to the difficulty in distinguishing small abnormalities from changes in normal chromosomes (Leary et al., 2010 Sci Transl Med 2:20ra14; and Leary et al., 2012 Sci Transl Med 4:162ra154). Other studies have shown that nucleosome patterns and chromatin structure may differ between cancer and normal tissues, and that cfDNA from cancer patients may result in abnormalities in cfDNA fragment size and position (Snyder et al., 2016 Cell 164:57; Jahr et al., 2001 Cancer Res 61:1659; Ivanov et al., 2015 BMC Genomics 16(Suppl 13):S1). However, the amount of sequencing required for nucleosome footprinting analysis of cfDNA is impractical for routine analysis.

[0066] The sensitivity of any cell-free DNA method depends on the number of potential changes examined and the technical and biological limits of detecting such changes. Since a typical blood sample contains approximately 2000 genome equivalents of cfDNA per milliliter of plasma (Phallen et al., 2017 Sci Transl Med 9), the limit of detecting a single variant may theoretically be no more than one wild-type molecule among thousands of mutants. Methods that detect a large number of changes in the same number of genome equivalents will be more sensitive for detecting cancer in the circulation. Monte Carlo simulations have shown that increasing the number of potential abnormalities detected from just a few to tens or hundreds can potentially increase the detection limit by several orders of magnitude, similar to recent probabilistic analyses of multiple methylation changes in cfDNA ( Figure 2 ).

[0067] This study presents a new method, called DELFI, for detecting cancer and further identifying the tissue of origin using whole-genome sequencing ( Figure 1The method uses cfDNA fragmentation profiles and machine learning to distinguish patterns in DNA from healthy blood cells and tumor-derived DNA and identify primary tumor tissue. DELFI was used to retrospectively analyze cfDNA from 245 healthy individuals and 236 patients with breast, colorectal, lung, ovarian, pancreatic, gastric, or bile duct cancer, most of whom presented with localized disease. Assuming that this method has a sensitivity of ≥0.80 for distinguishing cancer patients from healthy individuals at a specificity of 0.95, a study of at least 200 cancer patients would be able to estimate the true sensitivity with a margin of error of 0.06 at an expected specificity of 0.95 or greater.

[0068] Materials and methods

[0069] Patient and sample characteristics

[0070] Plasma samples from healthy individuals and plasma and tissue samples from patients with breast, lung, ovarian, colorectal, bile duct, or gastric cancer were obtained from ILS Bio / Bioreclamation, Aarhus University, Herlev Hospital, University of Copenhagen, Hvidovre Hospital, University Medical Center Utrecht, Academic Medical Center, University of Amsterdam, Netherlands Cancer Institute, and University of California, San Diego. All samples were obtained according to protocols approved by the Institutional Review Board and informed consent was obtained for research at participating institutions. Plasma samples were obtained from healthy individuals during routine screening (including colonoscopy or Pap smear). They were considered healthy if they had no previous history of cancer and a negative screening result.

[0071] Plasma samples were obtained from individuals with breast, colorectal, gastric, lung, ovarian, pancreatic, and bile duct cancer at diagnosis, before tumor resection, or before treatment. Changes in cfDNA fragmentation profiles at multiple time points were analyzed in 19 lung cancer patients who were receiving anti-EGFR or anti-ERBB2 therapy (see, e.g., Phallen et al., 2019 Cancer Research 15, 1204-1213). Table 1 (Appendix A) lists the clinical data of all patients in the study. Gender was confirmed by the representation of X and Y chromosomes in genomic analysis. Pathological staging of gastric cancer patients was performed after neoadjuvant therapy. Samples with unknown tumor stage were indicated as stage X or unknown.

[0072] Nucleosomal DNA purification

[0073] Live frozen lymphocytes were elutriated from leukocytes of males (C0618) and females (D0808-L) obtained from healthy individuals (Advanced Biotechnologies Inc., Eldersburg, MD). 1 x 10 6 Aliquots of cells were used for nucleosomal DNA purification. Initially, cells were treated with 100 μl Nuclei Prep Buffer and incubated on ice for 5 minutes. After centrifugation at 200 g for 5 minutes, the supernatant was discarded and the precipitated nuclei were treated twice with 100 μl Atlantis digestion buffer or 100 μl micrococcal nuclease (MN) digestion buffer. Finally, the cellular nucleic acid DNA was split with 0.5 U Atlantis dsDNase at 42 degrees C for 20 minutes and with 1.5 U MNase at 37 degrees C for 20 minutes. The reaction was terminated using 5X MN termination buffer and cleaved using a Zymo-Spin TM DNA was purified using an IIC column, and the concentration and quality of the eluted cellular nucleic acid DNA were analyzed using a Bioanalyzer 2100 (Agilent Technologies, Santa Clara, CA).

[0074] Sample preparation and cfDNA sequencing

[0075] For the three cancer patients participating in the surveillance analysis, whole blood was collected in EDTA tubes and processed immediately or within one day after storage at 4°C, or collected in Streck tubes and processed within two days. Plasma and cellular components were separated by centrifugation at 800g for 10 minutes at 4°C. Plasma was centrifuged a second time at 18,000g at room temperature to remove any residual cellular debris and stored at -80°C until DNA extraction. DNA was isolated from plasma using the Qiagen circulating nucleic acid kit (Qiagen GmbH) and eluted in LoBind tubes (Eppendorf AG). The concentration and quality of cfDNA were assessed using a Bioanalyzer 2100 (Agilent Technologies).

[0076] NGS cfDNA libraries for whole-genome sequencing and targeted sequencing were prepared using 5 to 250 ng of cfDNA as described elsewhere (see, e.g., Phallen et al., 2017 Sci Transl Med 9:eaan2415). Briefly, genomic libraries were prepared using the NEBNext DNA Library Preparation Kit for Illumina [New England Biolabs (NEB)], with four major modifications to the manufacturer's guidelines: (i) the library purification step used the on-bead AMPure XP method to minimize sample loss during elution and tube transfer steps (see, e.g., Fisher et al., 2011 Genome Biol 12:R1); (ii) the volumes of NEBNext End Repair, A-tailing, and adapter ligases and buffers were adjusted appropriately to accommodate the on-bead AMPure XP purification strategy; (iii) eight unique Illumina dual-index adapters with 8-base-pair (bp) barcodes were used in the ligation reaction, replacing the standard Illumina single- or dual-index adapters with 6- or 8-bp barcodes, respectively; and (iv) the cfDNA libraries were amplified with Phusion hot-start polymerase.

[0077] The whole genome library was sequenced directly. For the targeted library, Agilent SureSelect reagent and a customized hybridization probe set for 58 genes (e.g., see Phallen et al., 2017 Sci Transl Med 9: eaan2415) were used for capture according to the manufacturer's guidelines. The captured library was amplified with Phusion hot start polymerase (NEB). The concentration and quality of the captured cfDNA library were assessed on a Bioanalyzer 2100 using a DNA1000 kit (Agilent Technologies). The target library was sequenced using 100bp paired end sequencing on an Illumina HiSeq2000 / 2500 (Illumina).

[0078] Targeted sequencing data analysis of cfDNA

[0079] The targeted NGS data of cfDNA samples were analyzed as described elsewhere (see, e.g., Phallen et al., 2017 Sci Transl Med 9: eaan2415). In brief, the main processing, including the shielding of multi-index and dual-index adapter sequences, was completed using Illumina CASAVA (Consistency Assessment of Sequences and Variations) software (version 1.8). Sequence fragments were aligned to the human reference genome (hg18 or hg19 version) using NovoAlign, and additional alignments were performed on selected regions using the Needleman-Wunsch method (see, e.g., Jones et al., 2015 Sci Transl Med 7: 283ra53). The position of the sequence changes was not affected by different genome builds. Candidate mutations consisting of point mutations, small insertions, and deletions were identified within the targeted region of interest using VariantDx (see, e.g., Jones et al., 2015 Sci Transl Med 7: 283ra53) (Personal Genome Diagnostics, Baltimore, MD).

[0080] To analyze the fragment length of cfDNA molecules, the Phred quality score of each sequencing fragment pair (read pair) of the cfDNA molecule sequencing fragment was required to be ≥30. All duplicate ctDNA fragments were deleted, which were defined as having the same start, end, and index barcodes. For each mutation, only fragments of one or two sequencing fragment pairs containing mutant (or wild-type) bases at a given position were included. This analysis was completed using the R software package Rsamtools and GenomicAlignments.

[0081] For each genomic locus identified somatic mutation, the length of the fragment comprising the mutant allele is compared with the length of the fragment of the wild-type allele. If more than 100 mutant fragments are identified, the average fragment length is compared using Welch's two sample t tests. For the locus less than 100 mutant fragments, a bootstrap procedure was implemented. Specifically, N fragments of the replacement comprising the wild-type allele were sampled, where N represents the number of fragments with mutation. For the self-service replication of each wild-type fragment, their median length was calculated. The p value is estimated as the ratio replicated by the self-service procedure, where the length of the median wild-type fragment is equal to or greater than the length of the observed median mutant fragment.

[0082] Whole-genome sequencing data analysis of cfDNA

[0083] Initial processing of whole-genome NGS data from cfDNA samples, including demultiplexing and masking of multi-index adapter sequences, was performed using Illumina CASAVA (Consensus Assessment of Sequences and Variants) software (version 1.8.2). Sequence reads were aligned to the human reference genome (hg19 version) using ELAND.

[0084] Sequencing fragments and PCR duplicates with MAPQ scores below 30 were deleted. The hg19 autosomes were tiled into 26,236 adjacent, non-overlapping 100 kb bins. Low mappable regions were removed according to the bin indication of the lowest 10% coverage (see, e.g., Fortin et al., 2015 Genome Biol 16:180), as well as sequencing fragments that fell into the Duke blacklist region (see, e.g., hgdownload.cse.ucsc.edu / goldenpath / hg19 / encodeDCC / wgEncodeMap pability / ). Using this approach, 361 Mb (13%) of the hg19 reference genome were excluded, including centromere and telomeric regions. Short fragments were defined as having a length between 100 and 150 bp, and long fragments were defined as having a length between 151 and 220 bp.

[0085] In order to illustrate the coverage deviation due to genomic GC content, a local weighted smoother loess with a span of 3 / 4 is applied to a scatter plot of average fragment GC, rather than the coverage calculated for each 100kb bin. For short and long fragments, Loess regression is performed respectively to solve the possible differences (see, e.g., Benjamini et al., 2012 Nucleic Acids Res 40:e72) in how fragment length affects plasma coverage. The predicted values ​​of the short-term and long-term coverage explained by GC in the Loess model are subtracted to obtain short-term and long-term residuals that are unrelated to GC. By adding the medium, short-term and long-term coverage estimates across the entire genome, residuals can be restored to their original levels. This process is repeated for each sample to illustrate that there may be differences in the impact of GC on coverage between samples. In order to further reduce feature space (featurespace) and noise, the total coverage adjusted by GC in the 5Mb bin is calculated.

[0086] To compare the fragment length variability of cancer patients with that of healthy subjects, the standard deviation of the long fragment distribution profile was calculated for each individual. The standard deviations of the two groups were compared using the Wilcoxon rank sum test.

[0087] Analysis of chromosome arm copy number changes

[0088] In order to develop arm-level statistics for copy number changes, a method for aneuploidy detection in plasma described in other literature was used (see, for example, Leary et al., 2012 Sci Transl Med 4:162ra154). This method divides the genome into non-overlapping 50KB bins and obtains a GC-corrected log2 sequencing fragment depth after correction with a loess span of 3 / 4. This loess-based correction method is comparable to the above method, but is evaluated on a log2 scale to improve the robustness of outliers in smaller bins and is not stratified by fragment length. In order to obtain an arm-specific Z value for copy number changes, the average GC-adjusted sequencing fragment depth of each group of arms (GR) is centered and the healthy samples are calibrated by the mean and standard deviation of the GR values ​​obtained from 50 independent groups.

[0089] Mitochondrial alignment analysis of cfDNA

[0090] The whole genome sequence fragments originally mapped to the mitochondrial genome were extracted from the bam file and aligned to the hg19 reference genome in an end-to-end mode using Bowtie2, as described elsewhere (see, for example, Langmead et al., 2012 Nat Methods 9:357-359). The resulting aligned sequence fragments were filtered so that both pairs were aligned to the mitochondrial genome with a MAPQ >= 30. The number of fragments mapped to the mitochondrial genome was counted and converted to a percentage of the total number of fragments in the original bam file.

[0091] Predictive models for cancer classification

[0092] In order to distinguish healthy patients from cancer patients using fragment maps, a stochastic gradient boosting model (gbm) was used; see, for example, Friedman et al., 2001 Ann Stat 29: 1189-1232; and Friedman et al., 2002 Comput Stat Data An 38: 367-378). The sum of the short fragment coverages corrected for GC for all 504 bins was centered and scaled so that the mean value for each sample was 0 and the unit standard deviation was zero. Other features included the Z value of each of the 39 autosomal arms and mitochondrial representation (log10 conversion ratio of sequencing fragments mapped to mitochondria). In order to estimate the prediction error of the method, 10-fold cross validation was used as described elsewhere (see, for example, Efron et al., 1997 J Am Stat Assoc 92, 548-560). Only the feature selection performed on the training data in each cross validation run removed bins with high correlation (correlation> 0.9) or variance close to zero. Machine learning with stochastic gradient boosting was performed using the R package gbm with parameters n.trees = 150, interaction.depth = 3, shrinkage = 0.1, and n.minobsinside = 10. A 10-fold cross-validation procedure was repeated 10 times to average the prediction errors across folds from randomization of patients. Confidence intervals for sensitivity, fixed at 98%, and specificity, at 95%, were obtained from bootstrap replications performed in 2000.

[0093] Predictive model for tumor tissue origin classification

[0094] For samples correctly classified as cancer patients with 90% specificity (n=174), a separate stochastic gradient boosting model was trained to classify the tissue of origin. To address the small number of lung cancer samples used for prediction, 18 cfDNA baseline samples from patients with advanced lung cancer were included from the surveillance analysis. The performance characteristics of the model were evaluated by 10-fold cross-validation repeated 10 times. The gbm model was trained using the same features as the cancer classification model. As mentioned above, during cross-validation, features that showed correlations higher than 0.9 with each other or had variances close to zero were removed from each training dataset. The tissue class probabilities were averaged across the 10 replicates for each patient, and the class with the highest probability was used as the predicted tissue.

[0095] Nucleosomal DNA analysis of human lymphocytes and cfDNA

[0096] As described for whole genome cfDNA analysis, from the lymphocytes treated with nuclease, fragment size is analyzed in 5Mb bin. A whole genome map of nucleosome position is constructed from the lymphocyte line treated with nuclease. The method determines the local deviation within the coverage of the cycle fragments, showing that the region can be avoided from degradation. "Window positioning score" (WPS) is used to score each base pair in the genome (see, for example, Snyder et al., 2016Cell 164:57). Using a 60bp sliding window centered on each base, WPS is calculated as the number of fragments that completely span the window minus the number of fragments with only one end in the window. Since the median length of the fragments produced by nucleosomes is 167bp, high WPS represents the possible nucleosome position. Using continuous median, WPS score is centered at zero and smoothed using Kolmogorov-Zurbenko filter (for example, see Zurbenko, The spectral analysis of time series.North-Hollandseries in statistics and probability; Elsevier, New York, NY, 1986). For positive WPS spans between 50 and 450 bp, the nucleosome peak was defined as the set of base pairs with a WPS above the median within that window. Nucleosome positions were calculated for cfDNA from 30 healthy individuals using the same method as for lymphocyte DNA, with 9x sequence coverage. To ensure representative nucleosomes in healthy cfDNA, a consensus nucleosome track was defined, consisting only of nucleosomes identified in two or more individuals. The median distance between adjacent nucleosomes was calculated from the consensus track.

[0097] Simulation of Monte Carlo detection sensitivity

[0098] Monte Carlo simulation is used to estimate the possibility of detecting molecules with tumor-derived changes. In short, 1 million molecules were generated from a multinomial distribution. For a simulation with m changes, wild-type molecules were simulated with probability p and m tumor changes were simulated with probability (1-p) / m. Next, g*m molecules were randomly sampled and replaced, where g represents the number of genome equivalents in 1ml of plasma. If the tumor change was sampled s times or more, the sample was classified as cancer-derived. The simulation was repeated 1000 times to estimate the possibility of the sample being correctly classified as cancer in silico by the average value of the cancer indicator. If g=2000 and s=5, the number of tumor changes changes from 1 to 256 in powers of 2, and the proportion of tumor-derived molecules changes from 0.0001% to 1%.

[0099] Statistical analysis

[0100] All statistical analyses were performed using R version 3.4.3. R packages caret (version 6.0-79) and gbm (version 2.1-4) were used to classify healthy individuals from cancer patients and tissues of origin. Confidence intervals for model output were obtained using the pROC (version 1.13) R package (e.g., see Robin et al., 2011 BMC bioinformatics 12:77). Assuming a high prevalence of undiagnosed cancer in this population (1 or 2 cases per 100 healthy individuals), a genomic assay with a specificity of 0.95 and a sensitivity of 0.8 would have a useful operating characteristic (positive predictive value of 0.25 and negative predictive value close to 1). Power calculations indicate that, by analyzing more than 200 cancer patients and an approximately equal number of healthy controls, sensitivity can be estimated with an error margin of 0.06 at an expected specificity of 0.95 or higher.

[0101] Data and code availability

[0102] The sequence data used in this study have been deposited in the European Genome-Phenome Archive under accession numbers EGAS00001003611 and EGAS00001002577. The analysis code is available at github.com / Cancer-Genomics / delfi_scripts.

[0103] result

[0104] DELFI allows for the simultaneous analysis of a large number of abnormalities in cfDNA through genome-wide fragmentation pattern analysis. This approach is based on low-coverage whole-genome sequencing and analysis of isolated cfDNA. Mapped sequences are analyzed in non-overlapping windows covering the genome. Conceptually, window sizes can range from thousands to millions of bases, resulting in hundreds to thousands of windows across the genome. 5-Mb windows are used to assess cfDNA fragmentation patterns, providing over 20,000 sequenced fragments per window, even at limited 1–2x genome coverage. Within each window, the coverage and size distribution of cfDNA fragments are examined. This approach was used to assess variations in genome-wide fragment distributions in healthy and cancer populations (Table 1; Appendix A). An individual's genome-wide pattern can be compared with a reference population to determine whether the pattern is likely to be derived from a healthy individual or from a cancerous individual. Because genome-wide mapping reveals locational differences associated with specific tissues that may be missed in the overall fragment size distribution, these patterns may also indicate the tissue of origin of the cfDNA.

[0105] The fragment size of cfDNA has been of interest because it has been found that cfDNA molecules derived from cancer may vary in size more than cfDNA derived from noncancerous cells. cfDNA fragments from targeted regions of patients with breast, colorectal, lung, or ovarian cancer (Table 1 (Appendix A), Table 2 (Appendix B), and Table 1) were captured and sequenced at high coverage (total coverage 43,706, differential coverage 8,044), and Table 3 (Appendix C) were initially examined. Analysis of 165 tumor-specific altered sites from 81 patients (range of 1-7 alterations per patient) revealed that mutant and wild-type cfDNA fragments ( Figure 3 , Table 3 (Appendix C)) was 6.5 bp (95% CI, 5.4-7.6 bp). Compared with the wild-type sequence in these regions, the median size of mutant cfDNA fragments ranged from 30 bases smaller at 41,266,124 on chromosome 3 to 47 bases larger at 108,117,753 on chromosome 11 (Table 3; Appendix C). The GC content of mutant and unmutated fragments was similar (Figure 4a), and there was no correlation between GC content and fragment length (Figure 4b). Similar analysis of 44 germline alterations from 38 patients identified median cfDNA size differences between allele fragment lengths of less than 1 bp ( Figure 5 , Table 3 (Appendix C). In addition, 41 alterations associated with clonal hematopoiesis were identified by previous DNA sequence comparison of plasma, buffy coat, and tumors from the same individuals. Unlike the fragments derived from tumors, no significant differences were found between the fragments associated with hematopoiesis alterations and the wild-type fragments ( Figure 6 , Table 3 (Appendix C). Overall, cancer-derived cfDNA fragment lengths were more variable compared with non-cancer cfDNA fragments in certain genomic regions (p < 0.001, variance ratio test). Hypothesizing that these differences may be due to changes in higher-order chromatin structure and other genomic and epigenomic abnormalities in cancer, cfDNA fragmentation in a location-specific manner could serve as a unique biomarker for cancer detection.

[0106] Because targeted sequencing can only analyze a limited number of loci, a large-scale, whole-genome analysis was performed to detect additional abnormalities in cfDNA fragments. cfDNA was isolated from approximately 4 ml of plasma from 8 patients with stage I to III lung cancer and 30 healthy individuals (Table 1 (Appendix A), Table 4 (Appendix D), and Table 5 (Appendix E)). The cfDNA was converted into next-generation sequencing libraries using a highly efficient method and whole-genome sequenced at approximately 9-fold coverage (Table 4; Appendix D). Total cfDNA fragment length was larger in healthy individuals, with an average fragment size of 167.3 bp, compared with an average fragment size of 163.8 in patients with cancer (p < 0.01, Welch's t-test) (Table 5; Appendix E). To examine differences in fragment size and coverage across the genome, sequenced fragments were mapped to their genomic origin, and fragment length was assessed in 504 5-Mb windows, covering approximately 2.6 Gb of the genome. For each window, the ratio of small cfDNA fragments (100 to 150 bp in length) to larger cfDNA fragments (151 to 220 bp) and the total coverage were determined and used to obtain a genome-wide fragmentation profile for each sample.

[0107] Healthy individuals have very similar fragment distribution patterns throughout the genome (Figures 7 and Figure 8 ). In order to examine the origin of the fragmentation pattern commonly observed in cfDNA, nuclei were isolated from elutriated lymphocytes of two healthy individuals and treated with DNA nuclease to obtain nucleosomal DNA fragments. Analysis of the observed cfDNA pattern of healthy individuals showed that it was highly correlated with the lymphocyte nucleosomal DNA fragmentation profile (Figures 7b and 7d) and nucleosomal distance (Figures 7c and 7f). As revealed using the Hi-C method, the median distance between nucleosomes in lymphocytes is associated with the open (A) and closed (B) compartments of lymphoblasts (see, for example, Lieberman-Aiden et al., 2009 Science 326: 289-293; and Fortin et al., 2015 Genome Biol 16: 180) for examining the three-dimensional structure of the genome (Figure 7c). These analyses show that the fragmentation pattern of normal cfDNA is the result of the nucleosomal DNA pattern, which largely reflects the chromatin structure of normal blood cells.

[0108] Compared with healthy individuals, cancer patients showed multiple genomic variations in fragment length across different regions (Figures 7a and 7b). Similar to our observations from targeted analysis, cancer patients also had greater genome-wide fragment length variation compared with healthy individuals.

[0109] To determine whether cfDNA fragment length patterns could be used to distinguish cancer patients from healthy individuals, genome-wide correlation analysis of the ratio of long to short cfDNA fragments was performed for each sample (Figures 7a, 7b, and 7e) and compared with the median fragment length profiles calculated from healthy individuals. Although cfDNA fragmentation profiles in healthy individuals were highly concordant (median correlation of 0.99), the median correlation of genome-wide fragment ratios in cancer patients was 0.84 (0.15 lower, 95% CI 0.07-0.50, p < 0.001, Wilcoxon rank sum test; Table 5 (Appendix E)). Similar differences were observed when the fragmentation profiles of cancer patients were compared with the fragmentation profiles or nucleosome distances in healthy lymphocytes (Figures 7c, 7d, and 7f). To account for potential bias attributable to GC content, a locally weighted smoother was applied to each sample separately, and it was found that after this adjustment, the differences in fragment profiles between healthy individuals and cancer patients remained (median correlation between cancer patients and healthy people = 0.83) (Table 5; Appendix E).

[0110] The whole genome sequence data were subsampled and analyzed at 9x coverage of cfDNA from cancer patients, and the genome coverage was and The fragment profiles that were determined to be susceptible to change were identified even at 0.5x genomic coverage (Figure 9). Based on these observations, whole-genome sequencing at 1-2x coverage was performed to assess whether the fragment profiles might change during targeted therapy in a manner similar to monitoring sequence changes. cfDNA was evaluated in 19 patients with NSCLC who were undergoing anti-EGFR or anti-ERBB2 therapy, including 5 with local radiographic responses, 8 with stable disease, 4 with progressive disease, and 2 with unmeasurable disease (Table 6; Appendix F). Figure 10 As shown, the degree of abnormality in the fragment profile during treatment closely matched the levels of EGFR or ERBB2 mutant allele fragments determined using targeted sequencing (Spearman correlation between mutant allele and fragment profile = 0.74). These correlations were significant because the genome-wide and mutation-based approaches were orthogonal and examined different cfDNA alterations that may have been suppressed in these patients due to previous treatment. Notably, all cases whose tumors did not progress and who survived for six months or longer showed decreased or very low levels of ctDNA after initial treatment as determined by the fragment profile, while cases with poor clinical outcomes had increased ctDNA. These results demonstrate the feasibility of using fragment analysis to detect tumor-derived cfDNA and suggest that such analyses may also be useful for quantitative monitoring of cancer patients during treatment.

[0111] In patients for whom tumor tissue was available for parallel analysis, fragmentation profiles were examined in the context of known copy number alterations. These analyses revealed altered fragmentation profiles in copy-neutral genomic regions, and that these fragments could be further affected in regions of copy number alterations (Figures 11a and 12a). Position-dependent differences in fragmentation patterns could be used to distinguish cancer-derived cfDNA from cfDNA from healthy individuals in these regions (Figures 12a, b), whereas overall cfDNA fragment size measurements would miss these differences (Figure 12a).

[0112] These analyses were extended to independent cohorts of cancer patients and healthy individuals. Whole-genome sequencing of cfDNA at 1–2× coverage was performed in a total of 208 cancer patients with breast (n = 54), colorectal (n = 27), lung (n = 12), ovarian (n = 28), pancreatic (n = 54), gastric (n = 27), or bile duct (n = 26) cancer, as well as 215 individuals without cancer (Table 1 (Appendix A) and Table 4 (Appendix D)). All cancer patients were treatment-naive, and the majority had resectable disease (n = 183). After GC adjustment for coverage of short and long cfDNA fragments (Figure 13a), the coverage and size characteristics of fragments were examined across the entire genomic window (Figure 11b, Table 4 (Appendix D), and Table 7 (Appendix G)). Genome-wide correlations of coverage with GC content were limited, and no differences in these correlations were observed between cancer patients and healthy individuals (Figure 13b). Healthy individuals have highly consistent fragmentation profiles, whereas cancer patients have higher variability and reduced correlation with the median healthy profile (Table 7; Appendix G). Analysis of the most frequently altered fragmentation windows in the genomes of cancer patients revealed a median of 60 affected windows across the analyzed cancer types, highlighting the numerous position-dependent changes in cfDNA fragmentation in cancer patients (Figure 11c).

[0113] To determine whether position-dependent segment changes can be used to detect individuals with cancer, a gradient tree boosting machine learning model was implemented to examine whether cfDNA could be classified as having characteristics of cancer patients or healthy individuals, and the performance characteristics of this method were evaluated by a 10-fold cross-validation method with ten replicates ( Figure 14 and 15). The machine learning model included features of GC-adjusted short and long fragment coverage across the entire genomic window. A machine learning classifier was also developed based on features related to copy number changes in chromosome arms rather than a single score value (Figure 16a and Table 8 (Appendix H)) and including mitochondrial copy number changes (Figure 16b) to help distinguish between cancer and healthy individuals. Using this implementation of DELFI, a score value was obtained that can be used to classify patients as healthy or with cancer. Of 208 patients with cancer, 152 were detected (sensitivity 73%, 95% CI 67%-79%), while 4 of 215 healthy individuals were misclassified (specificity 98%) (Table 9). At a specificity threshold of 95%, 80% of patients with cancer were detected (95% CI, 74%-85%), including 79% of patients with resectable (stage I–III) disease (145 of 183 patients) and 82% of patients with metastatic (stage IV) disease (18 of 22 patients) (Table 9). The receiver operating characteristic analysis for detecting cancer patients had an AUC of 0.94 (95% CI 0.92–0.96), ranging from 0.86 for pancreatic cancer to ≥0.99 for lung and ovarian cancer ( Figures 17a and 17b ), and an AUC of ≥0.92 for all stages ( Figure 18 There were no differences in DELFI classifier scores with age in either cancer patients or healthy individuals (Table 1; Appendix A).

[0114] Table 9. DELFI performance results for cancer detection

[0115]

[0116] To evaluate the contribution of fragment size and coverage, chromosome arm copy number, or mitochondrial mapping to the model's prediction accuracy, a repeated 10-fold cross-validation procedure was implemented to evaluate the performance characteristics of these features individually. It was observed that the fragment coverage feature alone (AUC = 0.94) was almost identical to the classifier combining all features (AUC = 0.94) (Figure 17a). In contrast, the analysis of chromosome copy number changes had lower performance (AUC = 0.88), but was still more predictive than copy number changes based on a single score value (AUC = 0.78) or mitochondrial mapping (AUC = 0.72) (Figure 17a). These results indicate that fragment coverage is the main contributor to our classifier. Since information about cancer patients can be obtained from the same genomic sequence data, including all features in the prediction model may have a complementary effect on the detection of cancer patients.

[0117] Because the fragmentation profiles revealed regional differences in fragmentation between tissues, a similar machine learning approach was used to examine whether cfDNA patterns could identify the tissue of origin of these tumors. The method was found to be 61% accurate (95% CI 53%-67%) for breast cancer, 76% for bile duct cancer, 44% for colorectal cancer, 71% for gastric cancer, 53% for lung cancer, 48% for ovarian cancer, and 50% for pancreatic cancer. Figure 19 , Table 10). When considering assigning patients with abnormal cfDNA to one of the two sites of origin, accuracy increased to 75% (95% CI 69%-81%) (Table 10). For all tumor types, classification of tissue of origin by DELFI was significantly superior to classification determined by random assignment (p < 0.01, binomial test, Table 10).

[0118] Table 10. DELFI tissue origin prediction

[0119]

[0120] *Detected patients based on 90% specificity of the DELFI assay. The lung cohort includes other patients with previously treated lung cancer.

[0121] Because cancer-specific sequence alterations can be used to identify patients with cancer, we evaluated whether combining DELFI with this approach could improve the sensitivity of cancer detection ( Figure 20 Analysis of cfDNA from a subset of untreated cancer patients using DELFI and targeted sequencing revealed altered fragment profiles in 82% (103 of 126) of patients, and sequence alterations in 66% (83 of 126) of patients. DELFI detected more than 89% of cases with a mutant allele fraction >1%, compared with 80% of cases with a mutant allele fraction <1%, including those that could not be detected using targeted sequencing (Table 7; Appendix G). When these methods were used together, the combined test increased its sensitivity to 91% (115 of 126 patients) and specificity to 98% ( Figure 20 ).

[0122] Overall, genome-wide cfDNA fragmentation profiles differ between cancer patients and healthy individuals. Fragment length and coverage vary across the genome in a position-dependent manner, potentially explaining previously contradictory observations from cfDNA analyses of specific loci or total fragment size. In patients with cancer, heterogeneous fragmentation patterns in cfDNA appear to result from a mixture of nucleosomal DNA from blood and neoplastic cells. This study provides a method for simultaneously analyzing tiny amounts of cfDNA for tens to hundreds of tumor-specific abnormalities, thereby overcoming the potential limitations of more sensitive cfDNA analyses. Compared to previous cfDNA analysis methods that focused on sequence or overall fragment size, DELFI analysis detected a higher proportion of cancer patients (see, for example, Phallen et al., 2017 Sci Transl Med 9:eaan2415; Cohen et al., 2018 Science 359:926; Newman et al., 2014 Nat Med 20:548; Bettegowda et al., 2014 Sci Transl Med 6:224ra24; Newman et al., 2016 Nat Biotechnol 34:547). As demonstrated in this example, combining DELFI with other analyses of cfDNA alterations can further improve detection sensitivity. Because fragmentation patterns appear to correlate with nucleosomal DNA patterns, DELFI can be used to determine the primary source of tumor-derived cfDNA. By including clinical features, other biomarkers (including methylation alterations), and other diagnostic methods, the identification of the origin of circulating tumor DNA in more than half of the patients analyzed could be further improved (Ruibal Morell, 1992 The International Journal of Biological Markers 7:160; Galli et al., 2013 Clinical Chemistry and Laboratory Medicine 51:1369; Sikaris, 2011 Heart, Lung & Circulation 20:634; Cohen et al., 2018 Science 359:926). Finally, this approach requires only a small amount of whole-genome sequencing, rather than the deep sequencing required by typical approaches that focus on specific alterations. The performance characteristics and limited sequencing required for DELFI suggest that our approach could be widely applicable to the screening and management of cancer patients.

[0123] Our results suggest that genome-wide cfDNA fragmentation profiles differ between cancer patients and healthy individuals. Therefore, cfDNA fragmentation profiles may have important implications for future research and noninvasive methods for detecting human cancer.

[0124] Other embodiments

[0125] It should be understood that although the invention has been described in conjunction with the detailed description of the invention, the foregoing description is intended to illustrate rather than limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the appended claims.

[0126]

[0127]

[0128]

[0129]

[0130]

[0131]

[0132]

[0133]

[0134]

[0135]

[0136]

[0137]

[0138]

[0139]

[0140]

[0141]

[0142]

[0143]

[0144]

[0145]

[0146] Appendix-D: Table 4. Summary of Whole-Genome cfDNA Analysis

[0147]

[0148]

[0149]

[0150]

[0151]

[0152]

[0153]

[0154]

[0155]

[0156]

[0157]

[0158]

[0159]

[0160]

[0161]

[0162]

[0163]

Claims

1. A non-diagnostic or non-therapeutic method for identifying a mammal as having cancer by determining the cell-free DNA (cfDNA) fragmentation profile of the mammal, the method comprising: processing cfDNA fragments obtained from a sample obtained from the mammal into a sequencing library; Performing low-coverage whole-genome sequencing on the sequencing library to obtain sequencing fragments; Mapping the sequenced fragments to the genome to obtain a mapping sequence window; analyzing the window of mapped sequences to determine a cfDNA fragmentation profile; and The cfDNA fragmentation profile is compared to a reference cfDNA fragmentation profile, wherein an increased variability in the cfDNA fragmentation profile obtained from the mammal compared to the reference cfDNA fragmentation profile indicates that the mammal has cancer.

2. The method according to claim 1, wherein The mapping sequence includes tens to thousands of windows.

3. The method according to claim 1, wherein The windows are non-overlapping windows.

4. The method according to any one of claims 1 to 3, wherein The windows each comprised approximately 5 million base pairs.

5. The method according to any one of claims 1 to 3, wherein the cfDNA fragmentation profile is determined within each window.

6. The method according to any one of claims 1 to 3, wherein The cfDNA fragmentation profile includes a median fragment size.

7. The method according to any one of claims 1 to 3, wherein The cfDNA fragmentation profile includes fragment size distribution.

8. The method of any one of claims 1 to 3, wherein the cfDNA fragmentation profile comprises the ratio of small cfDNA fragments to large cfDNA fragments in the mapping sequence window.

9. The method according to any one of claims 1 to 3, wherein The cfDNA fragmentation profile includes sequence coverage of small cfDNA fragments in windows across the genome.

10. The method according to any one of claims 1 to 3, wherein The cfDNA fragmentation profile includes sequence coverage of large cfDNA fragments in windows across the genome.

11. The method according to any one of claims 1 to 3, wherein The cfDNA fragmentation profile includes sequence coverage of small and large cfDNA fragments across the entire genomic window.

12. The method according to any one of claims 1 to 3, wherein The cfDNA fragmentation profile is across the genome.

13. The method according to any one of claims 1 to 3, wherein The cfDNA fragmentation profiles span subgenomic intervals.

Citation Information

Patent Citations

  • Improvement in sounding-plates for musical instruments

    US144148A

  • Micronized placental tissue compositions and methods of making and using the same

    US20140050788A1

  • Treatment of cancer using humanized Anti-CD19 chimeric antigen receptor

    US20140271635A1

  • Compositions and Methods for Treatment of Cancer

    US20160194404A1

  • Universal anti-tag chimeric antigen receptor-expressing T cells and methods of treating cancer

    US9233125B2