Incorporating clinical risk into biomarker-based assessments for cancer pre-screening
By integrating clinical risk factors with genomic signatures, the method enhances cancer screening efficiency and accuracy by reducing unnecessary screenings and improving risk prediction.
Patent Information
- Application Number
- JP2025519685
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-07
- Filing Date
- 2023-10-06
- Publication Date
- 2025-11-12
AI Technical Summary
Existing cancer screening methods relying solely on genomic signatures fail to account for individual clinical risk factors, leading to inefficiencies in identifying subjects at high risk of cancer.
Incorporating clinical risk factors, such as age, sex, and smoking history, with genomic risk scores derived from cell-free DNA fragment size density data to enhance cancer screening accuracy.
Improves the identification of subjects likely to have cancer by reducing the number of screenings required and increasing the specificity and sensitivity of cancer prediction.
Smart Images

Figure 2025536890000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority under 35 U.S.C. §119(e) to U.S. Provisional Patent Application No. 63 / 414,370, filed October 7, 2022. The disclosure of the prior application is considered part of the disclosure of this application and is incorporated herein by reference in its entirety.
[0002] FIELD OF THE INVENTION The present invention relates generally to cancer pre-screening, and more specifically to improving cancer pre-screening results by incorporating clinical risk factors into the analysis of cell-free DNA ("cfDNA"). [Background technology]
[0003] Background information Blood-based biomarker assessments that identify genomic signatures of cancer have the potential to improve early cancer detection. In particular, cancer pre-screening using blood samples in which cfDNA fragments are sequenced and aligned to the genome can provide information such as the composition of the cfDNA population, the genomic location of the cfDNA fragments, their physical characteristics such as fragment size and fragment ends, and the presence of alterations indicative of cancer, such as copy number alterations, microsatellite instability, or other known cancer-causing genetic mutations. Summary of the Invention
[0004] The present invention builds on the pioneering discovery that incorporating individual-level clinical risk with a genomic signature of cancer improves identification of subjects most likely to have cancer detected by screening. In particular, the present disclosure demonstrates that incorporating clinical risk factors for lung cancer into cfDNA test analysis improves identification of subjects most likely to have a positive lung cancer diagnosis following standard low-dose computed tomography ("LDCT") lung cancer screening.
[0005] In clinical practice, cancer genomic signatures are usually interpreted using a cutoff point, above which a result is considered positive and below which a result is considered negative. However, relying solely on the genomic signature ignores the underlying clinical risk factors associated with the subject. The present disclosure describes a method for pre-screening for cancer based on a blood sample. Individual-level clinical risk is matched with the cancer genomic signature, thereby improving the identification of subjects who are most likely to have cancer detected by standard cancer screening methods.
[0006] In one embodiment, the present invention provides a method for predicting a subject's cancer status, comprising calculating a clinical risk score for the subject, calculating a genomic risk score for the subject, and combining the clinical risk score with the genomic risk score, thereby predicting the subject's cancer status. In one aspect, the present invention provides a method in which the clinical score includes the subject's age, sex, and / or race. In a further aspect, the genomic risk score includes cell-free DNA (cfDNA) fragment size density data from the subject. In a particular aspect, the cfDNA is obtained from a blood sample from the subject.
[0007] In certain embodiments, calculating cfDNA fragment size density data for a subject includes processing a sample from the subject containing cfDNA fragments into a library, subjecting the library to low-coverage whole-genome sequencing to obtain sequenced fragments, mapping the sequenced fragments to the genome to obtain mapped sequence windows, analyzing the mapped sequence windows to calculate the lengths of the cfDNA fragments, and generating the cfDNA fragment size density data.
[0008] In certain embodiments, cfDNA fragment size density data is calculated for one or more subgenomic interval(s). In further embodiments, a cfDNA fragmentation profile is calculated for each subgenomic interval. In further embodiments, the cfDNA fragment size density data comprises a curve. In some such embodiments, the cfDNA fragment size density curve from the subject is compared to cfDNA fragment size density curves from known healthy subjects and / or known cancer patients. In more detailed embodiments, the cfDNA fragmentation profile includes the most frequent fragment sizes. In further embodiments, the cfDNA fragmentation profile comprises a fragment size distribution with various frequencies of fragment sizes. In some embodiments, the cfDNA fragmentation profile comprises sequence coverage of small cfDNA fragments within a genome-wide window. In further embodiments, the cfDNA fragmentation profile comprises sequence coverage of large cfDNA fragments within a genome-wide window. In other embodiments, the cfDNA fragmentation profile comprises sequence coverage of large and small cfDNA fragments within a genome-wide window. In certain embodiments, the mapped cfDNA fragment sequences comprise tens to thousands of genome windows. In some such embodiments, these windows are non-overlapping windows. In other embodiments, these windows each comprise about 5 million base pairs. In further embodiments, the cfDNA fragmentation profile covers the entire genome.
[0009] In another embodiment, combining a clinical risk score with a genomic risk score results in an increased number of positive cancer diagnoses per subject screen compared to using either a clinical risk score or a genomic risk score alone. In certain embodiments, the number of subject screenings required to achieve a single positive cancer diagnosis is reduced by at least about 5%, 15%, 25%, 35%, 45%, 55%, 65%, 75%, or more on average compared to using a clinical risk score alone. In additional embodiments, the number of subject screenings required to achieve a single positive cancer diagnosis is reduced by at least about 5%, 10%, 15%, 20%, 25%, or more on average compared to using a genetic risk score alone. In a further embodiment, combining a clinical risk score with a genomic risk score results in improved discrimination between subjects predicted to have a high risk of cancer and subjects predicted to have a low risk of cancer. In another embodiment, combining a clinical risk score with a genomic risk score results in higher specificity of cancer prediction compared to using a clinical risk score alone or a genomic risk score alone. In further embodiments, the sensitivity of cancer prediction is at least about 50%, 60%, 70%, 80%, 90% or more.
[0010] In a further embodiment, the cancer is lung cancer. In some such embodiments, the subject's clinical risk score is calculated from data including the subject's age, sex, race, smoking status, number of pack-years, and smoking duration. In certain embodiments, the subject's clinical risk score is calculated from data including the Bach lung cancer development model described in Bach, PB, et al. J NATL CANCER INST. 95(6):470-8 (2003). The description of the Bach lung cancer development model is incorporated herein. In a further embodiment, combining the clinical risk score and the genomic risk score results in a composite score that increases as the subject's cancer risk increases. In certain embodiments, a cancer treatment is administered to a subject at increased risk of cancer. [Brief explanation of the drawings]
[0011] [Figure 1] A flow diagram showing the distribution of participants is shown. [Figure 2] Demographic and clinical characteristics of participants. "IQR" stands for "interquartile range," and "n / a" stands for "not applicable." [Figure 3] Dichotomized clinical risk by cancer status is shown. [Figure 4] The distribution of clinical risk by cancer status is shown. The lines within the rectangular boxes represent the median (line) and IQR (box), respectively. [Figure 5] The distribution of simulated genomic risk by clinical risk status is shown. The lines within the rectangular boxes represent the median (line) and IQR (box), respectively. [Figure 6] The number of CT scans needed to detect one case of lung cancer is shown for each type of risk estimate. [Figure 7] The specificity (95% CI) of the model to detect lung cancer with a sensitivity of 80% is shown. [Figure 8] Predicted probability of lung cancer diagnosis using clinical risk. [Figure 9] This shows the predicted probability of being diagnosed with lung cancer using clinical risk and genomic risk. [Figure 10] 8 illustrates an example of a computer 800 that may be used in predicting a subject's cancer status. DETAILED DESCRIPTION OF THE INVENTION
[0012] Detailed Description of the Invention The present invention builds on the pioneering discovery that incorporating individual-level clinical risk with a genomic signature improves identification of subjects most likely to have cancer detected by screening. In particular, the present disclosure demonstrates that incorporating clinical risk factors for lung cancer into the analysis of cfDNA tests improves identification of subjects most likely to have a positive lung cancer diagnosis from low-dose computed tomography ("LDCT") lung cancer screening.
[0013] Before describing the compositions and methods of the present invention, it is to be understood that this invention is not limited to the particular compositions, methods, and experimental conditions described, as such compositions, methods, and conditions may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the invention will be limited only in the appended claims.
[0014] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Thus, for example, it will be apparent to one of skill in the art upon reading this disclosure and so forth that reference to "the method" includes one or more methods, and / or steps of the type described herein.
[0015] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.
[0016] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. In the practice or testing of the present invention, any methods and materials similar or equivalent to those described herein can be used, but it is understood that modifications and variations are within the spirit and scope of the present disclosure. Preferred methods and materials are now described.
[0017] In one embodiment, the present invention provides a method for predicting a cancer status of a subject, comprising calculating a clinical risk score for the subject; calculating a genomic risk score for the subject; and combining the clinical risk score with the genomic risk score, thereby predicting the cancer status of the subject.
[0018] In one embodiment, calculating a clinical risk score comprises estimating the subject's one-year lung cancer risk. In one embodiment, the subject's one-year lung cancer risk is calculated using the Bach lung cancer incidence model. This Bach lung cancer incidence model is described in Bach, PB, et al. J NATL CANCER INST. 95(6):470-8 (2003), which is incorporated herein for its description of the Bach lung cancer incidence model. In one embodiment, this clinical risk score is calculated based on the subject's age, sex, asbestos exposure history, and smoking history. In one embodiment, estimating the subject's one-year lung cancer risk comprises classifying the subject's cancer risk into low clinical risk or high clinical risk. In one embodiment, the 25th percentile of clinical risk can be used to distinguish between low clinical risk and high clinical risk. In one aspect, calculating the clinical risk score includes questioning the subject about the subject's age, sex, smoking history, asbestos exposure history, history of obstructive pulmonary disease, brands of cigarettes smoked, types of asbestos exposure, chest x-ray findings, and exposure to radon or secondhand smoke, or any combination thereof; calculating the subject's cancer risk based on the answers provided by the subject using the Bach lung cancer incidence model; and assigning the subject a clinical risk score.
[0019] In one aspect, the invention provides a method wherein the clinical score includes the age, sex, and / or race of the subject.
[0020] In a further embodiment, the genomic risk score comprises cell-free DNA (cfDNA) fragment size density data from the subject.In certain embodiments, calculating the cfDNA fragment size density data of the subject comprises: processing the sample from the subject that contains cfDNA fragments into a library; subjecting the library to low-coverage whole genome sequencing to obtain sequenced fragments; mapping the sequenced fragments to genome to obtain mapped sequence windows; analyzing the mapped sequence windows to calculate the length of cfDNA fragments; and generating cfDNA fragment size density data.
[0021] In a further embodiment, genomic risk score is calculated based on the cfDNA fragmentation profile of object.In one embodiment, the cfDNA fragmentation profile can be calculated by: obtaining and isolating the cfDNA fragments from object, sequencing the cfDNA fragments to obtain sequenced fragments, mapping the sequenced fragments to genome to obtain mapped sequence window, and analyzing the mapped sequence window to calculate the length of cfDNA fragments, and generate cfDNA fragmentation profile.
[0022] In certain embodiments, the cfDNA is obtained from a blood sample from the subject.
[0023] In some embodiments, calculating cfDNA fragment size density data for a subject includes processing a sample from the subject containing cfDNA fragments into a library, subjecting the library to low-coverage whole-genome sequencing to obtain sequenced fragments, mapping the sequenced fragments to the genome to obtain mapped sequence windows, analyzing the mapped sequence windows to calculate the lengths of the cfDNA fragments, and generating cfDNA fragment size density data.
[0024] In some embodiments, the cfDNA fragmentation profile can be calculated by obtaining and isolating cfDNA fragments from a subject, sequencing the cfDNA fragments to obtain sequenced fragments, mapping the sequenced fragments to the genome to obtain mapped sequence windows, and analyzing the mapped sequence windows to calculate the lengths of the cfDNA fragments and generate a cfDNA fragmentation profile.
[0025] The methodology of the present invention is based on low-coverage whole-genome sequencing and analysis of isolated cfDNA. In one embodiment, the data used to develop the methodology of the present invention is based on shallow whole-genome sequence data (1-2x coverage).
[0026] In some embodiments, the mapped sequences are analyzed within non-overlapping windows covering the genome. Conceptually, the window size ranges from thousands to millions of bases, resulting in hundreds to thousands of windows across the genome. Because even a limited amount of 1-2x genome coverage yields more than 20,000 reads per window, 5Mb windows were used to evaluate cfDNA fragmentation patterns. Within each window, the coverage and size distribution of cfDNA fragments were examined. In some embodiments, the genome-wide patterns obtained from an individual can be compared to a reference population to determine whether the pattern is likely to be healthy or cancer-related.
[0027] In certain embodiments, the mapped sequences comprise tens to thousands of genomic windows, such as 10, 50, 100 to 1,000, 5,000, 10,000 or more windows. Such windows may be non-overlapping or overlapping and may comprise about 1 million, 2 million, 3 million, 4 million, 5 million, 6 million, 7 million, 8 million, 9 million, or 10 million base pairs.
[0028] In various embodiments, cfDNA fragmentation profile is calculated in each window.Thus, the present invention provides a method for calculating the cfDNA fragmentation profile in a subject (for example, a sample collected from a subject).
[0029] In some embodiments, cfDNA fragmentation profile can be used to identify the change (for example, mutation) in cfDNA fragment length.Mutation can be the mutation of the whole genome, or the mutation in one or more target regions / locus.Target region can be any region that contains one or more cancer-specific mutations. In some embodiments, the fragmentation profile of cfDNA can be used to identify (e.g., simultaneously identify) about 10 mutations to about 500 mutations (e.g., about 25 to about 500, about 50 to about 500, about 100 to about 500, about 200 to about 500, about 300 to about 500, about 10 to about 400, about 10 to about 300, about 10 to about 200, about 10 to about 100, about 10 to about 50, about 20 to about 400, about 30 to about 300, about 40 to about 200, about 50 to about 100, about 20 to about 100, about 25 to about 75, about 50 to about 250, or about 100 to about 200).
[0030] In various embodiments, the cfDNA fragmentation profile can include a cfDNA fragment size pattern. The cfDNA fragments can be of any appropriate size. For example, in some embodiments, the length of the cfDNA fragments can be from about 50 base pairs (bp) to about 400 bp. As described herein, a subject with cancer can have a cfDNA fragment size pattern that includes a median cfDNA fragment size that is shorter than the median cfDNA fragment size in healthy subjects. A healthy subject (e.g., a subject without cancer) can have a cfDNA fragment size with a median cfDNA fragment size of about 166.6 bp to about 167.2 bp (e.g., about 166.9 bp). In some embodiments, a subject with cancer can have a cfDNA fragment size that is, on average, about 1.28 bp to about 2.49 bp (e.g., about 1.88 bp) shorter than the cfDNA fragment size in healthy subjects. For example, a subject with cancer can have a cfDNA fragment size with a median cfDNA fragment size of about 164.11 bp to about 165.92 bp (e.g., about 165.02 bp).
[0031] In some embodiments, dinucleosomal cfDNA fragments can be about 230 base pairs (bp) to about 450 bp in length. As described herein, subjects with cancer can have a dinucleosomal cfDNA fragment size pattern that includes a median dinucleosomal cfDNA fragment size that is shorter than the median dinucleosomal cfDNA fragment size in healthy subjects. In some embodiments, on average, subjects without cancer have longer cfDNA fragments in the dinucleosomal range (average size 334.75 bp), while subjects with cancer have shorter dinucleosomal cfDNA fragments (average size 329.6 bp). Thus, healthy subjects (e.g., subjects without cancer) can have a median dinucleosomal cfDNA fragment size of about 334.75 bp. In some embodiments, subjects with cancer can have a dinucleosomal cfDNA fragment size that is shorter than the dinucleosomal cfDNA fragment size in healthy subjects. For example, a subject with cancer may have a dinucleosomal cfDNA fragment size with a median cfDNA fragment size of about 329.6 bp.
[0032] The cfDNA fragmentation profile can include a cfDNA fragment size distribution. As described herein, a subject with cancer may have a cfDNA size distribution that is more variable than the cfDNA fragment size distribution in a healthy subject. In some embodiments, the size distribution can exist within a target region. A healthy subject (e.g., a subject without cancer) may have a cfDNA fragment size distribution in a target region that is about 1 or less than about 1. In some embodiments, a subject with cancer may have a cfDNA fragment size distribution in a target region that is longer (e.g., 10, 15, 20, 25, 30, 35, 40, 45, 50 bp or more longer, or any number of base pairs longer than the cfDNA fragment size distribution in a target region in a healthy subject. In some embodiments, a subject with cancer may have a cfDNA fragment size distribution in a target region that is shorter (e.g., 10, 15, 20, 25, 30, 35, 40, 45, 50 bp or more shorter, or any number of base pairs shorter than the cfDNA fragment size distribution in a target region in a healthy subject. In some embodiments, a subject with cancer may have a cfDNA fragment size distribution in the target region that is about 47 bp shorter to about 30 bp longer than the cfDNA fragment size distribution in the target region of a healthy subject.In some embodiments, a subject with cancer may have, on average, a difference of 10, 11, 12, 13, 14, 15, 15, 17, 18, 19, 20 bp or more in the cfDNA fragment length in the cfDNA fragment size distribution in the target region.For example, a subject with cancer may have, on average, a difference of about 13 bp in the cfDNA fragment length in the cfDNA fragment size distribution in the target region.In some embodiments, the size distribution may be a genome-wide size distribution.
[0033] The cfDNA fragmentation profile can include the ratio of small cfDNA fragments to large cfDNA fragments, and the correlation of the fragment ratio with a reference fragment ratio. As used herein, with respect to the ratio of small cfDNA fragments to large cfDNA fragments, the small cfDNA fragments can be about 100 bp to about 150 bp in length. As used herein, with respect to the ratio of small cfDNA fragments to large cfDNA fragments, the large cfDNA fragments can be about 151 bp to 220 bp in length. As described herein, a subject with cancer can have a fragment ratio correlation (e.g., a correlation of a cfDNA fragment ratio with a reference DNA fragment ratio, such as a DNA fragment ratio from one or more healthy subjects) that is lower (e.g., 2-fold lower, 3-fold lower, 4-fold lower, 5-fold lower, 6-fold lower, 7-fold lower, 8-fold lower, 9-fold lower, 10-fold lower, or more) than a healthy subject. Healthy subjects (e.g., subjects without cancer) can have a fragment ratio correlation (e.g., the correlation of cfDNA fragment ratios to a reference DNA fragment ratio, such as DNA fragment ratios from one or more healthy subjects) of about 1 (e.g., about 0.96). In some embodiments, subjects with cancer can have a fragment ratio correlation (e.g., the correlation of cfDNA fragment ratios to a reference DNA fragment ratio, such as DNA fragment ratios from one or more healthy subjects) that is, on average, about 0.19 to about 0.30 (e.g., about 0.25) lower than the fragment ratio correlation in healthy subjects (e.g., the correlation of cfDNA fragment ratios to a reference DNA fragment ratio, such as DNA fragment ratios from one or more healthy subjects).
[0034] In certain embodiments, cfDNA fragment size density data is calculated for one or more subgenomic intervals. In further embodiments, a cfDNA fragmentation profile is calculated for each subgenomic interval.
[0035] In further embodiments, the cfDNA fragment size density data comprises a curve. In some such embodiments, the cfDNA fragment size density curve from the subject is compared to cfDNA fragment size density curves from known healthy subjects and / or known cancer patients.
[0036] In more embodiments, the cfDNA fragmentation profile comprises the most frequent fragment sizes. In further embodiments, the cfDNA fragmentation profile comprises a fragment size distribution having various frequencies of fragment sizes.
[0037] In some embodiments, the cfDNA fragmentation profile comprises the sequence coverage of small cfDNA fragments in a genome-wide window.In further embodiments, the cfDNA fragmentation profile comprises the sequence coverage of large cfDNA fragments in a genome-wide window.In other embodiments, the cfDNA fragmentation profile comprises the sequence coverage of large and small cfDNA fragments in a genome-wide window.In certain embodiments, the sequence of the mapped cfDNA fragments comprises tens to thousands of genome windows.In some such embodiments, these windows are non-overlapping windows.In other embodiments, each of these windows comprises about 5 million base pairs.In further embodiments, the cfDNA fragmentation profile covers the entire genome.
[0038] In certain embodiments, the method further comprises generating a cell-free DNA (cfDNA) fragmentation profile for predicting the cancer state of the subject. In certain embodiments, generating a cell-free DNA (cfDNA) fragmentation profile for predicting the cancer state of the subject comprises obtaining a sample from the subject, processing the sample to obtain a plasma fraction, extracting and purifying nucleosome-protected cfDNA fragments from the plasma fraction, processing the cfDNA fragments obtained from the sample obtained from the subject into a sequencing library, and subjecting the sequencing library to whole genome sequencing to obtain sequencing fragments, wherein the genome coverage is about 9x to 0.1x.
[0039] In another aspect, combining a clinical risk score with a genomic risk score results in an increased number of positive cancer diagnoses per subject screened compared to using only a clinical or genomic risk score.
[0040] In certain embodiments, the number of screenings required in a subject to achieve a single positive cancer diagnosis is reduced by at least about an average of 5%, 15%, 25%, 35%, 45%, 55%, 65%, 75% or more compared to using a clinical risk score alone.
[0041] In additional embodiments, the number of subject screenings required to achieve a single positive cancer diagnosis is reduced by at least about an average of 5%, 10%, 15%, 20%, 25%, or more compared to using genetic risk scores alone.
[0042] In an additional embodiment, the incorporation of a clinical risk score and a genomic risk score results in improved discrimination between subjects predicted to have a high risk of cancer and subjects predicted to have a low risk of cancer.
[0043] In another embodiment, the combination of a clinical risk score and a genomic risk score results in a higher specificity of cancer prediction compared to using either the clinical risk score alone or the genomic risk score alone, hi a further embodiment, the sensitivity of cancer prediction is at least about 50%, 60%, 70%, 80%, 90% or more.
[0044] In further embodiments, the cancer is lung cancer. In some such embodiments, a subject's clinical risk score is calculated from data including the subject's age, sex, race, smoking status, number of pack-years, and smoking duration.
[0045] In certain embodiments, the clinical risk score for a subject is calculated from data that includes the Bach lung cancer model described in Bach, PB, et al. J NATL CANCER INST. 95(6):470-8 (2003), which is incorporated herein for its description of the Bach lung cancer model.
[0046] In further embodiments, the cancer may be any stage of cancer. In some embodiments, the cancer may be an early stage cancer. In some embodiments, the cancer may be an asymptomatic cancer. In some embodiments, the cancer may be residual disease and / or recurrence (e.g., after surgical resection and / or cancer treatment). The cancer may be any type of cancer. Examples of cancer types that can be assessed, monitored, and / or treated as described herein include, but are not limited to, lung cancer, colorectal cancer, prostate cancer, breast cancer, pancreatic cancer, bile duct cancer, liver cancer, CNS cancer, gastric cancer, esophageal cancer, gastrointestinal stromal tumor (GIST), uterine cancer, and ovarian cancer. Additional cancer types include, but are not limited to, myeloma, multiple myeloma, B-cell lymphoma, follicular lymphoma, lymphocytic leukemia, leukemia, myeloid leukemia, etc. In some embodiments, the cancer is a solid tumor. In some embodiments, the cancer is a sarcoma, carcinoma, or lymphoma. In some embodiments, the cancer is lung cancer, colorectal cancer, prostate cancer, breast cancer, pancreatic cancer, bile duct cancer, liver cancer, CNS cancer, gastric cancer, esophageal cancer, gastrointestinal stromal tumor (GIST), uterine cancer, or ovarian cancer. In some embodiments, the cancer is a hematological cancer. In some embodiments, the cancer is myeloma, multiple myeloma, B-cell lymphoma, follicular lymphoma, lymphocytic leukemia, leukemia, or myeloid leukemia.
[0047] In an additional embodiment, the clinical risk score and the genomic risk score are combined to result in a composite score that increases as the subject's risk of cancer increases.
[0048] In certain embodiments, a subject at increased risk of cancer is administered a cancer treatment.
[0049] The cancer treatment can be any suitable cancer treatment. One or more cancer treatments described herein can be administered to a subject at any appropriate frequency (e.g., one or more times over a period of several days to several weeks). Examples of cancer treatments include, but are not limited to, surgical intervention, adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, hormonal therapy, cytotoxic therapy, immunotherapy, adoptive T cell therapy (e.g., T cells having chimeric antigen receptors and / or wild-type or modified T cell receptors), targeted therapy (e.g., kinase inhibitors, antibodies, bispecific antibodies) such as administration of kinase inhibitors (e.g., kinase inhibitors that target specific genetic abnormalities such as translocations or mutations), signal transduction inhibitors, bispecific antibodies or antibody fragments (e.g., BiTEs), monoclonal antibodies, immune checkpoint inhibitors, surgery (e.g., surgical resection), or any combination of the above. In some embodiments, the cancer treatment can reduce the severity of the cancer, reduce the symptoms of the cancer, and / or reduce the number of cancer cells present in the subject.
[0050] In some embodiments, the cancer therapeutic agent can be a chemotherapeutic agent. Non-limiting examples of chemotherapeutic agents include amsacrine, azacitidine, azathioprine, bevacizumab (or an antigen-binding fragment thereof), bleomycin, busulfan, carboplatin, capecitabine, chlorambucil, cisplatin, cyclophosphamide, cytarabine, dacarbazine, daunorubicin, docetaxel, doxifluridine, doxorubicin, epirubicin, erlotinib hydrochloride, etoposide, fludarabine, floxuridine, fludarabine, fluorouracil, gemcitabine, hydroxyurea, and the like.
[0013] Additional examples of anti-cancer drug treatments include rhea, idarubicin, ifosfamide, irinotecan, lomustine, mechlorethamine, melphalan, mercaptopurine, methotrexate, mitomycin, mitoxantrone, oxaliplatin, paclitaxel, pemetrexed, procarbazine, all-trans retinoic acid, streptozocin, tafluposide, temozolomide, teniposide, thioguanine, topotecan, uramustine, valrubicin, vinblastine, vincristine, vindesine, vinorelbine, and combinations thereof. Additional examples of anti-cancer drug treatments are known in the art; see, for example, the American Society of Clinical Oncology (ASCO), the European Society for Medical Oncology (ESMO), or the National Comprehensive Cancer Network (NCCN) treatment guidelines.
[0051] In various embodiments, DNA is present in a biological sample collected from a subject and used in the methodology of the present invention. The biological sample can be virtually any type of biological sample containing DNA. The biological sample is typically a liquid, such as whole blood or a portion thereof containing circulating cfDNA. In embodiments, the sample contains DNA from a tumor or a liquid biopsy, including, but not limited to, amniotic fluid, aqueous humor, vitreous humor, blood, whole blood, fractionated blood, plasma, serum, breast milk, cerebrospinal fluid (CSF), earwax, chyle, oozing fluid, endolymph, perilymph, feces, exhaled breath, gastric acid, gastric juice, lymph, mucus (including nasal discharge and sputum), pericardial fluid, ascites, pleural fluid, pus, ocular discharge, saliva, exhaled breath condensate, sebum, semen, sputum, sweat, synovial fluid, tears, vomit, prostatic fluid, nipple aspirate, tears, sweat, buccal swab, cell lysate, gastrointestinal fluid, biopsy tissue, urine, or other biological fluids. In one embodiment, the sample comprises DNA from circulating tumor cells.
[0052] As disclosed above, the biological sample can be a blood sample. The blood sample can be collected using methods known in the art, such as a finger prick or venous blood draw. Preferably, the blood sample is about 0.1-20 ml, or about 1-15 ml, with a blood volume of about 10 ml. Smaller amounts of blood may also be used, as may free circulating DNA in the blood. Microsampling, needle biopsy, catheter sampling, stool samples, and the generation of bodily fluids containing DNA are also potential sources of biological samples.
[0053] The methods and systems of the present disclosure utilize nucleic acid sequence information and therefore may include any method or sequencing device for performing nucleic acid sequencing, including nucleic acid amplification, polymerase chain reaction (PCR), nanopore sequencing, 454 sequencing, and insertion tag sequencing. In some embodiments, the methodology or system of the present disclosure utilizes a system such as those provided by Illumina, Inc. (including but not limited to HiSeq™ X10, HiSeq™ 1000, HiSeq™ 2000, HiSeq™ 2500, Genome Analyzers™, MiSeq™, NextSeq, and NovaSeq 6000 systems), Applied Biosystems Life Technologies (SOLiD™ System, ion Proton™ Sequencer, ion Proton™ Sequencer), or Genapsys or BGI MGI, and other systems. Nucleic acid analysis can also be performed with systems provided by Oxford Nanopore Technologies (GridiON™, MinION™) or Pacific Biosciences (Pacbio™ RS II or Sequel I or II).
[0054] The present invention, including systems for performing the steps of the disclosed methods, is described in part in terms of functional components and various process steps. Such functional components and process steps can be realized by any number of components, operations, and techniques configured to perform the specified functions and achieve various results. For example, the present invention can employ various biological samples, biomarkers, elements, materials, computers, data sources, storage systems and media, information collection techniques and processes, data processing standards, statistical analyses, regression analyses, and the like, which can perform various functions.
[0055] Accordingly, the present invention further provides a system for predicting a cancer status of a subject, which in various embodiments includes (a) a sequencer configured to generate a low-coverage whole-genome sequence dataset for a sample, and (b) a computer system and / or processor having functionality for performing the methods of the present invention.
[0056] In some embodiments, the computer system further comprises one or more additional modules. For example, the system may comprise one or more extraction and / or isolation units operable to perform suitable genetic component analysis, e.g., to select cfDNA fragments of a particular size.
[0057] In some embodiments, the computer system further comprises a visual display device, which may be operable to display the curve fit line, the reference curve fit line, and / or a comparison of the two.
[0058] The method for predicting a subject's cancer status according to various aspects of the present invention may be implemented in any suitable manner, such as by using a computer program running on a computer system. As discussed herein, exemplary systems according to various aspects of the present invention may be implemented in combination with a computer system, e.g., a conventional computer system including a processor and random access memory, such as a remotely accessible application server, network server, personal computer, or workstation. The computer system also preferably includes additional storage or information storage systems, such as a mass storage system, and a user interface, such as a conventional monitor, keyboard, and tracking device. However, the computer system may include any suitable computer system and associated equipment and may be configured in any suitable manner. In one embodiment, the computer system is a stand-alone system. In another embodiment, the computer system is part of a network of computers, including a server and a database.
[0059] The software necessary to receive, process, and analyze information may be implemented in a single device or multiple devices. This software may be accessible over a network so that information storage and processing is remote to the user. Systems and their various components according to various aspects of the present invention provide functions and operations that facilitate detection and / or analysis, such as data collection, processing, analysis, reporting, and / or diagnosis. For example, in this aspect, a computer system executes a computer program capable of receiving, storing, retrieving, analyzing, and reporting information related to the human genome or regions thereof. This computer program may include multiple modules that perform various functions or operations, such as a processing module that processes raw data to generate supplemental data and an analysis module that analyzes the raw data and supplemental data to generate quantitative assessments of disease state models and / or diagnostic information.
[0060] The procedures performed by the system may include any suitable process for facilitating analysis and / or cancer diagnosis. In one embodiment, the system is configured to establish a disease state model and / or calculate a disease state in the patient. Calculating or identifying a disease state may include generating any useful information regarding the patient's condition related to disease, such as making a diagnosis, providing information useful in a diagnosis, assessing the stage or progression of a disease, identifying conditions that may indicate susceptibility to a disease, identifying whether further testing is recommended, predicting and / or evaluating the effectiveness of one or more treatment programs, or otherwise assessing the patient's disease state, likelihood of disease, or other health aspects.
[0061] 10 illustrates an example of a computer 800 that may be used in predicting a subject's cancer status. For example, in some embodiments, the computer 800 may include a machine learning system that trains a machine learning model for predicting a subject's cancer status, as described above, or a portion or combination thereof. The computer 800 may be any electronic device that executes software applications derived from compiled instructions, including, but not limited to, a personal computer, a server, a smartphone, a media player, an electronic tablet, a gaming console, an email device, etc. In some implementations, the computer 800 may include one or more processors 802, one or more input devices 804, one or more display devices 806, one or more network interfaces 808, and one or more computer-readable media 812. Each of these components may be coupled by a bus 810, or in some embodiments, these components may be distributed across multiple physical locations and coupled by a network.
[0062] The display device 806 may be of any known display technology, including, but not limited to, a display device using liquid crystal display (LCD) or light-emitting diode (LED) technology. The processor(s) 802 may use any known processor technology, including, but not limited to, graphics processors and multi-core processors. The input device(s) 804 may be of any known input device technology, including, but not limited to, a keyboard (including a virtual keyboard), a mouse, a trackball, a camera, a touch-sensitive pad, or a display. The bus 810 may be of any known internal or external bus technology, including, but not limited to, ISA, EISA, PCI, PCI Express, USB, Serial ATA, or FireWire. The computer-readable medium 812 may be any non-transitory medium involved in providing instructions to the processor(s) 804 for execution, including, but not limited to, a non-volatile storage medium (e.g., optical disks, magnetic disks, flash drives, etc.) or a volatile medium (e.g., SDRAM, ROM, etc.).
[0063] The computer-readable medium 812 may include various instructions 814 for implementing an operating system (e.g., Mac OS, Windows, Linux). The operating system may be multi-user, multi-processing, multi-tasking, multi-threading, real-time, etc. The operating system may perform basic tasks including, but not limited to, recognizing input from the input device(s) 804, sending output to the display device 806, keeping track of files and directories on the computer-readable medium 812, controlling peripheral devices (e.g., disk drives, printers, etc.) that are controllable directly or through I / O controllers, and managing traffic on the bus 810. The network communication instructions 816 may establish and maintain network connections (e.g., software for implementing communication protocols such as TCP / IP, HTTP, Ethernet, telephony, etc.).
[0064] Machine learning instructions 818 may include instructions that enable computer 800 to function as a machine learning system and / or train a machine learning model to generate DMS values as described herein. Application(s) 820 may be applications that use or implement the processes described herein and / or other processes. The processes may also be implemented in operating system 814. For example, application 820 and / or the operating system may create tasks in the application as described herein.
[0065] The features described herein may be implemented in one or more computer programs executable on a programmable system including at least one programmable processor connected to receive data and instructions from, and transmit data and instructions to, a data storage system, at least one input device, and at least one output device. A computer program is a set of instructions that can be used, directly or indirectly, in a computer to perform a particular activity or bring about a particular result. Computer programs may be written in any type of programming language (e.g., Objective-C, Java), including compiled or interpreted languages, and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0066] Processors suitable for executing a program of instructions may include, by way of example, general-purpose and special-purpose microprocessors, the sole processor, or one of multiple processors or cores of any type of computer. Generally, a processor can receive instructions and data from a read-only memory, a random-access memory, or both. The basic elements of a computer can include a processor for executing instructions and one or more memories for storing instructions and data. Generally, a computer can include, or be operatively coupled to communicate with, one or more mass storage devices for storing data files; such devices include magnetic, such as internal hard disks or removable disks, magneto-optical, and optical disks. Storage devices suitable for tangibly embodying computer program instructions and data include, by way of example, semiconductor memory devices such as EPROMs, EEPROMs, and flash memory devices; magnetic, such as internal hard disks or removable disks; and non-volatile memory in all forms, such as CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, application-specific integrated circuits (ASICs).
[0067] To provide for user interaction, these functions may be implemented on a computer that has a display device, such as an LED or LCD monitor, to display information to the user, and a keyboard and pointing device, such as a mouse or trackball, to allow the user to provide input to the computer.
[0068] This functionality may be implemented within computer systems including back-end components such as data servers, computer systems including middleware components such as application servers or Internet servers, computer systems including front-end components such as client computers having graphical user interfaces or Internet browsers, or any combination thereof. The components of the system may be connected by any form or medium of digital data communication, such as a communications network. Examples of communications networks include, for example, the telephone network, a LAN, a WAN, and the computers and networks forming the Internet.
[0069] A computer system may include clients and servers. Clients and servers may generally be remote from each other and typically interact through a network. The relationship of client and server may arise by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0070] One or more features or steps of the disclosed embodiments can be implemented using an API, which can define one or more parameters passed between a calling application and other software code (e.g., an operating system, library routine, function) that provides a service, provides data, or performs an operation or calculation.
[0071] An API may be implemented as one or more calls in program code that receive or send one or more parameters through a parameter list or other structure according to a calling convention defined in an API specification. A parameter may be a constant, a key, a data structure, an object, an object class, a variable, a data type, a pointer, an array, a list, or another call. API calls and parameters may be implemented in any programming language. This programming language may define the vocabulary and calling conventions that programmers employ to access functions that support the API.
[0072] In some implementations, calls to the API may report to the application the capabilities of the device on which the application is running, such as input capabilities, output capabilities, processing capabilities, power capabilities, communication capabilities, etc.
[0073] While various embodiments have been described above, it should be understood that they have been presented by way of example, and not limitation. It will be apparent to one skilled in the relevant art(s) that various changes in form and detail can be made without departing from the spirit and scope. Indeed, after reading the above description, it will be apparent to one skilled in the relevant art(s) how to implement alternative embodiments. For example, other steps may be added to or deleted from the described flows, and other components may be added to or deleted from the described systems. Accordingly, other implementations are within the scope of the following claims.
[0074] Furthermore, it should be understood that any diagrams highlighting features and advantages are presented for illustrative purposes only, and that the disclosed methodologies and systems are each sufficiently flexible and configurable that they may be utilized in ways other than those illustrated.
[0075] In this specification, claims, and drawings, the term "at least one" is often used, but terms such as "a," "an," "the," and "said" also mean "at least one" or "said at least one" in this specification, claims, and drawings.
[0076] Finally, it is Applicant's intention that only those claims which include the phrase "means for" or "step for" be construed under 35 U.S.C. 112(f). Any claim which does not expressly include the phrase "means for" or "step for" shall not be construed under 35 U.S.C. 112(f).
[0077] The presently described methods and systems are useful for detecting, predicting, treating, and / or monitoring the status of cancer in a subject. Any suitable subject, such as a mammal, can be evaluated, monitored, and / or treated as described herein. Some examples of mammals that can be evaluated, monitored, and / or treated as described herein include, but are not limited to, humans, primates such as monkeys, dogs, cats, horses, cows, pigs, sheep, mice, and rats. For example, a human with or suspected of having cancer can be evaluated using the methods described herein and, optionally, treated with one or more cancer treatments described herein.
[0078] The following examples are provided to further illustrate embodiments of the present invention and are not intended to limit the scope of the invention, which is typical of methods that may be used, although other procedures, methodologies, or techniques known to those skilled in the art may alternatively be used. [Example]
[0079] Example 1 - Incorporating clinical risk into a biomarker-based assessment used as a pre-screen for LDCT lung cancer screening Because populations with a smoking history have different probabilities of benefiting from screening, personalized risk assessment may improve the overall benefit of LDCT screening. Lung cancer risk can be estimated from clinical factors such as age and smoking history. However, blood-based biomarkers have shown the potential to significantly improve risk estimation beyond clinical risk. Blood-based biomarker assessments that identify genomic signatures of lung cancer may improve the efficiency of LDCT screening if used as prescreening.
[0080] Among these eligible individuals, such assessment can distinguish between those with a high and low likelihood of lung cancer detection by LDCT. In clinical use, genomic signatures are typically interpreted using cutpoints, above which results are considered positive and below which results are considered negative. However, relying solely on genomic signatures does not account for differences in underlying clinical risk factors such as age and smoking history. Here, we show that integrating individual-level clinical risk with a lung cancer genomic signature improves identification of those most likely to have lung cancer detected by screening.
[0081] research participants The data for this study include participants from the National Lung Screening Trial (NLST). A total of 53,452 participants were enrolled in the NLST (see Figure 1, top panel). 26,730 participants were randomly assigned to the radiographic examination treatment group and 26,722 to the spiral CT examination group. The former group (the radiographic examination treatment group) was excluded from this analysis. Additionally, 1,620 participants in the spiral CT examination treatment group were excluded from this analysis because they lacked necessary clinical data. In total, 25,102 participants were included in this analysis (see Figure 1, bottom panel).
[0082] Clinical risk calculation For eligible participants, the 1-year lung cancer risk for each participant was estimated using the Bach lung cancer incidence model described in Bach, PB, et al. J NATL CANCER INST. 95(6):470-8 (2003). The description of the Bach lung cancer incidence model is incorporated herein. The 25th percentile of clinical risk was selected as the cutpoint separating low and high clinical risk. Observed 1-year lung cancer diagnoses were predicted using logistic regression models based on the genomic signature score alone, clinical risk categories alone, and both genomic and clinical risk categories. These models were compared for specificity, assuming a sensitivity of 80%, and the number of CT scans needed to detect one lung cancer case at an overall prevalence of 1%. Wilson score confidence intervals were estimated.
[0083] Genomic risk calculation A genomic signature score was simulated for each participant. The score was derived from the distribution of the cohort, stratified by cancer status and stage, and evaluated using the DELFI technology, which evaluates the fragmentation profile of cell-free DNA present in the blood and uses supervised machine learning to detect cancer signals.
[0084] Statistical analysis Clinical risk was considered as a continuous predictor of 1-year observed lung cancer risk in two additional logistic regression models: one using clinical risk alone and one incorporating both genomic and clinical risk. Predicted lung cancer outcome probabilities were derived from the logistic regression models with continuous predictors. A threshold of 80% sensitivity was calculated to stratify the predicted probabilities from the multivariate models. These models were compared for specificity at 80% sensitivity and the number of CT scans needed to detect one lung cancer case at an overall prevalence of 1%. 95% confidence intervals ("CI") were calculated using bootstrap sampling.
[0085] result The analysis included 25,102 subjects, of whom 254 (1.0%) were diagnosed with lung cancer within 1 year (see Figure 2). The median clinical risk (interquartile range) was 0.39% (0.23-0.63). Therefore, the cutpoint for separating low and high clinical risk was 0.23%. The median lung cancer risk was 0.15% (low-risk group) and 0.49% (high-risk group).
[0086] Simulated genomic risk, clinical risk, and models incorporating both were all significantly associated with lung cancer diagnosis (p<0.001). For example, Figure 3 shows the binary clinical risk by cancer status, while Figure 4 shows the distribution of clinical risk by cancer status. The latter (Figure 4) shows that the median clinical risk was 0.60% (0.37-0.93) in the lung cancer group and 0.38% (0.23-0.63) in the non-cancer group. The lines within the rectangular boxes in Figure 4 represent the median (line) and IQR (rectangular box), respectively.
[0087] Figure 5 shows the distribution of simulated genomic risk by clinical risk status. For both the low and high clinical risk groups, the simulated genomic risk scores ranged from 0 to 1. Again, in Figure 5, the lines within the rectangular boxes represent the median (line) and IQR (rectangular box), respectively.
[0088] Figure 6 shows the number of CT scans needed to detect one case of lung cancer by type of risk estimate. The observation rate of CT scans needed to detect one case of lung cancer was calculated from the prevalence of lung cancer in the NLST CT-treated group. Compared with using categorical clinical risk alone, using genomic risk alone reduced the number of CT scans needed by 32%, from 95 to 65. Combining genomic risk and categorical clinical risk reduced the number by 37%, from 95 to 60.
[0089] Figure 7 shows that incorporating clinical risk into genomic risk improved specificity from 56% (95% CI 0.55-0.57) to 59% (95% CI 0.58-0.60) at 80% sensitivity, reduced the number of CT scans needed to detect one lung cancer from 65 (with genomic risk alone) to 60, and reduced the number of screenings needed with LDCT by 7%. For comparison, without pre-screening assessment, the number of screenings needed to detect one lung cancer was approximately 100.
[0090] Figure 8 shows the predicted probability of lung cancer diagnosis using clinical risk, and Figure 9 shows the predicted probability of lung cancer diagnosis using clinical and genomic risk. The 80% sensitivity threshold separating low and high predicted probabilities of lung cancer diagnosis using the combined clinical (continuous) and genomic risk was set at 0.005 (dotted line). Incorporating clinical risk into genomic risk allowed for further differentiation between patients below and above the threshold.
Claims
1. 1. A method for predicting cancer status in a subject, comprising: a) calculating a clinical risk score for said subject; b) calculating a genomic risk score for the subject; and c) combining the clinical risk score with the genomic risk score to thereby predict the cancer status of the subject; A method comprising:
2. The method of claim 1 , wherein the clinical score includes the age, sex, and / or race of the subject.
3. 2. The method of claim 1, wherein the genomic risk score comprises cell-free DNA (cfDNA) fragment size density data from the subject.
4. calculating the cfDNA fragment size density data for the subject, a) processing a sample from said subject containing cfDNA fragments into a library; b) subjecting said library to low-coverage whole genome sequencing to obtain sequenced fragments; c) mapping the sequenced fragments to the genome to obtain a window of mapped sequences; d) analyzing said windows of mapped sequences to calculate the lengths of cfDNA fragments; e) generating said cfDNA fragment size density data; The method of claim 3, comprising:
5. 5. The method of claim 4, wherein the cfDNA fragment size density data comprises a curve.
6. 6. The method of claim 5, wherein the cfDNA fragment size density curve from the subject is compared to cfDNA fragment size density curves from known healthy subjects and / or known cancer patients.
7. 2. The method of claim 1, wherein combining the clinical risk score and the genomic risk score results in an increased number of positive cancer diagnoses per subject screened compared to using either the clinical risk score or the genomic risk score alone.
8. 8. The method of claim 7, wherein the number of times the subject needs to be screened to achieve a positive cancer diagnosis is reduced by at least about 5%, 15%, 25%, 35%, 45%, 55%, 65%, 75% or more on average compared to using a clinical risk score alone.
9. 8. The method of claim 7, wherein the number of times the subject needs to be screened to achieve a positive cancer diagnosis is reduced by at least about 5%, 10%, 15%, 20%, 25% or more, on average, compared to using the genetic risk score alone.
10. 8. The method of claim 7, wherein the sensitivity of the cancer prediction is at least about 50%, 60%, 70%, 80%, 90% or more.
11. 11. The method of claim 10, wherein combining the clinical risk score and the genomic risk score results in a more specific cancer prediction compared to using only the clinical risk score or only the genomic risk score.
12. 10. The method of claim 1, wherein combining the clinical risk score and the genomic risk score results in improved discrimination between subjects predicted to be at high risk for cancer and subjects predicted to be at low risk for cancer.
13. The method according to any one of claims 1 to 12, wherein the cancer is lung cancer.
14. 14. The method of claim 13, wherein the subject's clinical risk score is calculated from data including the subject's age, sex, race, smoking status, number of pack-years, and smoking duration.
15. 15. The method of claim 14, wherein the subject's clinical risk score is calculated from data comprising the Bach lung cancer development model.
16. 10. The method of claim 1, wherein the clinical risk score and the genomic risk score result in a combined score that increases as the subject's risk for cancer increases.
17. 10. The method of claim 1, wherein the clinical risk score and the genomic risk score result in a combined score that decreases as the subject's risk for cancer decreases.
18. 5. The method of claim 4, wherein the cfDNA fragment size density data is calculated for a subgenomic interval.
19. The method of claim 1, wherein cancer treatment is administered to a subject predicted to have cancer.
20. 4. The method of claim 3, wherein the cfDNA is obtained from a blood sample from the subject.
21. The method of claim 4 , wherein the mapped sequence comprises tens to thousands of windows.
22. The method of claim 4 , wherein the windows are non-overlapping windows.
23. 5. The method of claim 4, wherein the windows each comprise about 5 million base pairs.
24. 5. The method of claim 4, wherein a cfDNA fragmentation profile is calculated within each window.
25. 25. The method of claim 24, wherein the cfDNA fragmentation profile comprises the most frequent fragment sizes.
26. 25. The method of Claim 24, wherein the cfDNA fragmentation profile comprises a fragment size distribution having different frequencies of fragment sizes.
27. 25. The method of claim 24, wherein the cfDNA fragmentation profile comprises a ratio of small to large cfDNA fragments within the window of mapped sequence.
28. 25. The method of claim 24, wherein the cfDNA fragmentation profile comprises the sequence coverage of small cfDNA fragments within the genome-wide window.
29. 25. The method of claim 24, wherein the cfDNA fragmentation profile comprises sequence coverage of large cfDNA fragments within the genome-wide window.
30. 25. The method of claim 24, wherein the cfDNA fragmentation profile comprises the sequence coverage of small and large cfDNA fragments within the genome-wide window.
31. 25. The method of claim 24, wherein the cfDNA fragmentation profile spans the entire genome.
32. 25. The method of claim 24, wherein the cfDNA fragmentation profile spans a subgenomic interval.