Including clinical risk into biomarker-based cancer pre-screening assessment

By combining individual clinical risk scores and genomic risk scores, the size and density data of cell free DNA fragments are analyzed, and the problem of ignoring clinical risk factors in existing cancer pre-screening is solved, which improves the sensitivity and specificity of cancer pre-screening and reduces the number of LDCT screening times.

CN120359307APending Publication Date: 2025-07-22DELFI DIAGNOSTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380070413.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-07
Filing Date
2023-10-06
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing cancer prescreening methods rely on genomic characteristics and ignore individual clinical risk factors, resulting in insufficient screening efficiency and accuracy.

Method used

Combining individual clinical risk scores and genomic risk scores, cfDNA fragment size and density data are generated by analyzing cell free DNA fragment size and density data to improve the accuracy and efficiency of cancer pre-screening.

Benefits of technology

It improves the sensitivity and specificity of cancer prescreening, reduces the number of standard low-dose computed tomography (LDCT) screenings, and improves the recognition rate of potential cancer individuals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120359307A_ABST
    Figure CN120359307A_ABST
Patent Text Reader

Abstract

The present invention discloses a method for cancer pre-screening of a subject. Genetic risks associated with genetic characteristics are determined by sequencing and analyzing cell free DNA ("cfDNA") fragments present in a blood sample of the subject. The clinical risk is determined based on factors such as age, gender, and race. In some cases, clinical factors specific to certain cancers, such as smoking conditions, are combined. By incorporating clinical risks into genomic risk analysis, improved lung cancer pre-screening results are provided, enabling the same number of positive cancer detections to be achieved using a small number of LDCT lung cancer screening times.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of priority of U.S. Provisional Patent Application Serial No. 63 / 414,370, filed on October 7, 2022, under 35 U.S.C. § 119(e). The disclosure of the prior application is considered part of the disclosure of this application and is incorporated herein by reference in its entirety. Background Art Technical Field

[0003] The present invention generally relates to cancer pre-screening, and more particularly to improving cancer pre-screening results by incorporating clinical risk factors into the analysis of cell-free DNA ("cfDNA"). Background Art

[0005] Blood-based biomarker assays identify genomic features of cancer and have the potential to improve early detection of cancer. In particular, cancer pre-screening using blood samples in which cfDNA fragments are sequenced and aligned to the genome can provide information such as the composition of the cfDNA population, the genomic location of cfDNA fragments, physical characteristics (such as fragment size and fragment ends), and the presence of changes indicative of cancer (such as copy number changes, microsatellite instability, or other known oncogenic genetic variations). Summary of the Invention

[0006] The present invention is based on the groundbreaking discovery that combining clinical risk at the individual level with genomic features of cancer improves the identification of subjects most likely to have cancer detected by screening. Among other things, the present disclosure demonstrates that incorporating clinical risk factors for lung cancer into the analysis of cfDNA tests improves the identification of subjects most likely to have a positive confirmation of lung cancer by standard low-dose computed tomography ("LDCT") lung cancer screening.

[0007] In clinical use, genomic features of cancer are typically interpreted using a cut-off point, above which the result is positive and below which the result is negative. However, relying solely on genomic features ignores potential clinical risk factors associated with the subject. The present disclosure describes methods for cancer pre-screening based on blood samples. Clinical risk at the individual level is matched with genomic features of cancer, thereby improving the identification of subjects most likely to have cancer detected by standard cancer screening methods.

[0008] In one embodiment, the present invention provides a method for predicting the cancer status of a subject, the method comprising determining a clinical risk score of the subject; determining a genomic risk score of the subject; and combining the clinical risk score with the genomic risk score to thereby predict the cancer status of the subject. In one aspect, the present invention provides a method, wherein the clinical score comprises the age, sex, and / or race of the subject. In a further aspect, the genomic risk score comprises cell-free DNA (cfDNA) fragment size density data from the subject. In certain aspects, the cfDNA is obtained from a blood sample from the subject.

[0009] In certain aspects, determining the cfDNA fragment size density data of the subject comprises: processing a sample from the subject comprising cfDNA fragments into a library; performing low-coverage whole-genome sequencing on the library to obtain sequenced fragments; mapping the sequenced fragments to a genome to obtain windows of mapped sequences; analyzing the windows of mapped sequences to determine cfDNA fragment lengths; and generating cfDNA fragment size density data.

[0010] In certain aspects, the cfDNA fragment size density data is calculated for one or more sub-genomic intervals. In additional embodiments, a cfDNA fragmentation profile is determined for each sub-genomic interval. In a further aspect, the cfDNA fragment size density data comprises a curve. In some such aspects, the cfDNA fragment size density curve from the subject is compared to the cfDNA fragment size density curves from known healthy subjects and / or known cancer patients. In more aspects, the cfDNA fragmentation profile comprises the fragment size of the maximum frequency. In a further aspect, the cfDNA fragmentation profile comprises a fragment size distribution of fragment sizes with varying frequencies. In some aspects, the cfDNA fragmentation profile comprises sequence coverage in windows across the genome for small cfDNA fragments. In a further aspect, the cfDNA fragmentation profile comprises sequence coverage in windows across the genome for large cfDNA fragments. In other aspects, the cfDNA fragmentation profile comprises sequence coverage in windows across the genome for both small and large cfDNA fragments. In certain aspects, the mapped cfDNA fragment sequences comprise dozens to thousands of genomic windows. In some such aspects, the windows are non-overlapping windows. In other aspects, each window comprises approximately five million base pairs. In a further aspect, the cfDNA fragmentation profile covers the entire genome.

[0011] On the other hand, compared to using either a clinical risk score or a genomic risk score alone, combining a clinical risk score and a genomic risk score results in a greater number of positive cancer diagnoses per subject screening. In some aspects, compared to using a clinical risk score alone, the number of subject screenings required to achieve one positive cancer diagnosis is reduced by at least about 5%, 15%, 25%, 35%, 45%, 55%, 65%, 75% or more, on average. In additional aspects, compared to using a genetic risk score alone, the number of subject screenings required to achieve one positive cancer diagnosis is reduced by at least about 5%, 10%, 15%, 20%, 25% or more, on average. In additional aspects, combining a clinical risk score with a genomic risk score results in an improved discrimination between subjects predicted to have a high cancer risk and those predicted to have a low cancer risk. On the other hand, compared to using a clinical risk score alone or a genomic risk score alone, combining a clinical risk score with a genomic risk score results in a higher specificity for cancer prediction. In further aspects, the sensitivity of cancer prediction is at least about 50%, 60%, 70%, 80%, 90% or higher.

[0012] In further aspects, the cancer is lung cancer. In some such aspects, the clinical risk score of a subject is determined from data including the subject's age, sex, race, smoking status, pack-years, and smoking duration. In certain aspects, the clinical risk score of a subject is determined from data including the Bach lung cancer incidence model, as described in Bach, P.B., et al. JNATL CANCER INST. 95(6):470-8 (2003), the description of which Bach lung cancer incidence model is incorporated herein by reference. In additional aspects, combining a clinical risk score with a genomic risk score results in a combined score that increases as the cancer risk of the subject increases. In certain aspects, a cancer treatment is administered to a subject having an increased cancer risk. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 A flow chart of participant allocation is shown.

[0014] Figure 2 Demographics and clinical characteristics of the participants are provided. "IQR" stands for "interquartile range" and "n / a" stands for "not applicable".

[0015] Figure 3 Binary clinical risk according to cancer status is shown.

[0016] Figure 4 The distribution of clinical risk according to cancer status is shown. The lines within the rectangular boxes represent the median (line) and IQR (rectangular box), respectively.

[0017] Figure 5Shows the distribution of simulated genomic risk according to clinical risk status. The lines within the rectangular boxes represent the median (line) and IQR (rectangular box), respectively.

[0018] Figure 6 Shows the number of CT scans required to detect one case of lung cancer according to the type of risk estimate.

[0019] Figure 7 Shows the model specificity (95% CI) for detecting lung cancer at 80% sensitivity.

[0020] Figure 8 Shows the predicted probability of lung cancer diagnosis using clinical risk.

[0021] Figure 9 Shows the predicted probability of lung cancer diagnosis using clinical risk and genomic risk.

[0022] Figure 10 Shows an example computer 800 that can be used to predict the cancer status of a subject. Detailed Description

[0023] The present invention is based on the following groundbreaking discovery, namely that combining individual-level clinical risk with genomic characteristics improves the identification of subjects most likely to have cancer detected by screening. Among other things, the present disclosure demonstrates that incorporating clinical risk factors for lung cancer into the analysis of cfDNA testing improves the identification of subjects most likely to have a positive confirmation of lung cancer by low-dose computed tomography (“LDCT”) lung cancer screening.

[0024] Before describing the compositions and methods of the present invention, it should be understood that the present invention is not limited to the specific compositions, methods, and experimental conditions described, as such compositions, methods, and conditions may vary. It should also be understood that the terms used herein are for the purpose of describing specific embodiments only and are not intended to be limiting, as the scope of the present invention is limited only in the appended claims.

[0025] As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to “the method” includes one or more methods and / or steps of the type described herein and / or steps that will become apparent to those skilled in the art after reading the present disclosure and the like.

[0026] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the same extent as if each individual publication, patent, or patent application were specifically and individually indicated to be incorporated by reference.

[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, it should be understood that modifications and variations are covered within the spirit and scope of this disclosure. Preferred methods and materials are now described.

[0028] In one embodiment, the present invention provides a method for predicting the cancer status of a subject, the method comprising determining a clinical risk score of the subject; determining a genomic risk score of the subject; and combining the clinical risk score with the genomic risk score to thereby predict the cancer status of the subject.

[0029] In one aspect, determining the clinical risk score includes estimating the 1-year lung cancer risk of the subject. In one aspect, the Bach lung cancer incidence model is used to determine the 1-year lung cancer risk of the subject. The Bach lung cancer incidence model is described in Bach, P.B., et al. J NATL C ANCER I NST .95(6):470 - 8(2003), the description of which regarding the Bach lung cancer incidence model is incorporated herein. In one aspect, the clinical risk score is determined based on the subject's age, gender, asbestos exposure history, and smoking history. In one aspect, estimating the 1-year lung cancer risk of the subject includes classifying the subject's cancer risk as low clinical risk or high clinical risk. In one aspect, the 25th percentile of the clinical risk can be used to distinguish low clinical risk from high clinical risk. In one aspect, determining the clinical risk score includes: asking the subject about their age, gender, smoking history, asbestos exposure, obstructive lung disease history, brand of cigarettes smoked, type of asbestos exposure, chest x-ray findings, and exposure to radon or secondhand smoke or any combination thereof; using the Bach lung cancer incidence model to determine the subject's cancer risk based on the responses provided by the subject; and assigning a clinical risk score to the subject.

[0030] In one aspect, the present invention provides a method wherein the clinical score includes the age, gender, and / or race of the subject.

[0031] In a further aspect, the genomic risk score includes cell-free DNA (cfDNA) fragment size density data from the subject. In certain aspects, determining the cfDNA fragment size density data of the subject includes: processing a sample from the subject comprising cfDNA fragments into a library; performing low-coverage whole-genome sequencing on the library to obtain sequenced fragments; mapping the sequenced fragments to the genome to obtain windows of mapped sequences; analyzing the windows of mapped sequences to determine the cfDNA fragment lengths; and generating cfDNA fragment size density data.

[0032] In a further aspect, a genomic risk score is determined based on the cfDNA fragmentation profile of a subject. In one aspect, the cfDNA fragmentation profile can be determined by: obtaining and isolating cfDNA fragments from the subject; sequencing the cfDNA fragments to obtain sequenced fragments; mapping the sequenced fragments to the genome to obtain windows of mapped sequences; and analyzing the windows of mapped sequences to determine cfDNA fragment lengths and generate a cfDNA fragmentation profile.

[0033] In certain aspects, the cfDNA is obtained from a blood sample from the subject.

[0034] In some aspects, determining the cfDNA fragment size density data of a subject includes: processing a sample from the subject comprising cfDNA fragments into a library; performing low-coverage whole-genome sequencing on the library to obtain sequenced fragments; mapping the sequenced fragments to the genome to obtain windows of mapped sequences; analyzing the windows of mapped sequences to determine cfDNA fragment lengths; and generating cfDNA fragment size density data.

[0035] In some aspects, the cfDNA fragmentation profile can be determined by: obtaining and isolating cfDNA fragments from the subject; sequencing the cfDNA fragments to obtain sequenced fragments; mapping the sequenced fragments to the genome to obtain windows of mapped sequences; and analyzing the windows of mapped sequences to determine cfDNA fragment lengths and generate a cfDNA fragmentation profile.

[0036] The methods of the invention are based on low-coverage whole-genome sequencing and analysis of isolated cfDNA. In one aspect, the data used to develop the methods of the invention are based on shallow whole-genome sequence data (1 to 2x coverage).

[0037] In some aspects, the mapped sequences are analyzed in non-overlapping windows that cover the genome. Conceptually, the window size can range from thousands to millions of bases, resulting in hundreds to thousands of windows across the genome. 5Mb windows are used to evaluate the cfDNA fragmentation pattern because these will provide over 20,000 reads per window even at limited 1 to 2x genome coverage. Within each window, the coverage and size distribution of cfDNA fragments are examined. In some aspects, the whole-genome pattern from an individual can be compared to a reference population to determine whether the pattern is likely to be of healthy or cancer origin.

[0038] In some aspects, the mapped sequences include dozens to thousands of genomic windows, such as 10, 50, 100 to 1,000, 5,000, 10,000 or more windows. Such windows can be non - overlapping or overlapping and include about 1 million, 2 million, 3 million, 4 million, 5 million, 6 million, 7 million, 8 million, 9 million or 10 million base pairs.

[0039] In various aspects, a cfDNA fragmentation profile is determined within each window. Accordingly, the present invention provides methods for determining a cfDNA fragmentation profile in a subject (e.g., in a sample obtained from a subject).

[0040] In some aspects, the cfDNA fragmentation profile can be used to identify changes (e.g., alterations) in cfDNA fragment lengths. The alterations can be genome - wide alterations or alterations in one or more targeted regions / loci. The target region can be any region containing one or more cancer - specific alterations. In some aspects, the cfDNA fragmentation profile can be used to identify (e.g., simultaneously identify) from about 10 alterations to about 500 alterations (e.g., from about 25 to about 500, from about 50 to about 500, from about 100 to about 500, from about 200 to about 500, from about 300 to about 500, from about 10 to about 400, from about 10 to about 300, from about 10 to about 200, from about 10 to about 100, from about 10 to about 50, from about 20 to about 400, from about 30 to about 300, from about 40 to about 200, from about 50 to about 100, from about 20 to about 100, from about 25 to about 75, from about 50 to 250 or from about 100 to about 200 alterations).

[0041] In various aspects, the cfDNA fragmentation profile can include a cfDNA fragment size pattern. The cfDNA fragments can be of any suitable size. For example, in some aspects, the length of the cfDNA fragments can be from about 50 base pairs (bp) to about 400 bp. As described herein, the cfDNA fragment size pattern of a subject with cancer can include a median cfDNA fragment size that is shorter than the median cfDNA fragment size of a healthy subject. A healthy subject (e.g., a subject without cancer) can have a cfDNA fragment size with a median of from about 166.6 bp to about 167.2 bp (e.g., about 166.9 bp). In some aspects, a subject with cancer can have a cfDNA fragment size that is on average about 1.28 bp to about 2.49 bp (e.g., about 1.88 bp) shorter than that of a healthy subject. For example, a subject with cancer can have a cfDNA fragment size with a median of from about 164.11 bp to about 165.92 bp (e.g., about 165.02 bp).

[0042] In some aspects, the length of the dinucleosome cfDNA fragments can be from about 230 base pairs (bp) to about 450 bp. As described herein, the dinucleosome cfDNA fragment size pattern of a subject with cancer can include a median dinucleosome cfDNA fragment size that is shorter than the median dinucleosome cfDNA fragment size of a healthy subject. In some aspects, on average, cancer-free subjects have longer cfDNA fragments within the dinucleosome range (average size of 334.75 bp), while subjects with cancer have shorter dinucleosome cfDNA fragments (average size of 329.6 bp). Accordingly, a healthy subject (e.g., a subject without cancer) can have a dinucleosome cfDNA fragment size with a median of about 334.75 bp. In some aspects, a subject with cancer can have a dinucleosome cfDNA fragment size that is shorter than that of a healthy subject. For example, a subject with cancer can have a dinucleosome cfDNA fragment size with a median of about 329.6 bp.

[0043] The cfDNA fragmentation profile can include the cfDNA fragment size distribution. As described herein, a subject with cancer can have a more variable cfDNA size distribution than that of a healthy subject. In some aspects, the size distribution can be within a targeted region. A healthy subject (e.g., a subject without cancer) can have a targeted region cfDNA fragment size distribution of about 1 or less than about 1. In some aspects, a subject with cancer can have a targeted region cfDNA fragment size distribution that is longer (e.g., 10, 15, 20, 25, 30, 35, 40, 45, 50 or more bp longer, or any number of base pairs between these numbers) than that of a healthy subject. In some aspects, a subject with cancer can have a targeted region cfDNA fragment size distribution that is shorter (e.g., 10, 15, 20, 25, 30, 35, 40, 45, 50 or more bp shorter, or any number of base pairs between these numbers) than that of a healthy subject. In some aspects, a subject with cancer can have a targeted region cfDNA fragment size distribution that is about 47 bp smaller to about 30 bp longer than that of a healthy subject. In some aspects, a subject with cancer can have a targeted region cfDNA fragment size distribution in which the lengths of the cfDNA fragments differ on average by 10, 11, 12, 13, 14, 15, 15, 17, 18, 19, 20 or more bp. For example, a subject with cancer can have a targeted region cfDNA fragment size distribution in which the lengths of the cfDNA fragments differ on average by about 13 bp. In some aspects, the size distribution can be a genome-wide size distribution.

[0044] The cfDNA fragmentation profile can include the ratio of small cfDNA fragments to large cfDNA fragments and the correlation of the fragment ratio to a reference fragment ratio. As used herein, with respect to the ratio of small cfDNA fragments to large cfDNA fragments, the length of the small cfDNA fragments can be from about 100 bp to about 150 bp. As used herein, with respect to the ratio of small cfDNA fragments to large cfDNA fragments, the length of the large cfDNA fragments can be from about 151 bp to 220 bp. As described herein, a subject with cancer can have a lower (e.g., 2-fold lower, 3-fold lower, 4-fold lower, 5-fold lower, 6-fold lower, 7-fold lower, 8-fold lower, 9-fold lower, 10-fold lower or more) correlation of the fragment ratio (e.g., the correlation of the cfDNA fragment ratio to a reference DNA fragment ratio such as the DNA fragment ratio from one or more healthy subjects) than a healthy subject. A healthy subject (e.g., a subject without cancer) can have a correlation of the fragment ratio (e.g., the correlation of the cfDNA fragment ratio to a reference DNA fragment ratio such as the DNA fragment ratio from one or more healthy subjects) of about 1 (e.g., about 0.96). In some aspects, a subject with cancer can have a correlation of the fragment ratio (e.g., the correlation of the cfDNA fragment ratio to a reference DNA fragment ratio such as the DNA fragment ratio from one or more healthy subjects) that is on average about 0.19 to about 0.30 (e.g., about 0.25) lower than that of a healthy subject.

[0045] In certain aspects, cfDNA fragment size density data is calculated for one or more sub-genomic intervals. In additional embodiments, a cfDNA fragmentation profile is determined for each sub-genomic interval.

[0046] In further aspects, the cfDNA fragment size density data includes a curve. In some such aspects, the cfDNA fragment size density curve from a subject is compared to the cfDNA fragment size density curves from known healthy subjects and / or known cancer patients.

[0047] In more aspects, the cfDNA fragmentation profile includes the fragment size of the maximum frequency. In further aspects, the cfDNA fragmentation profile includes a fragment size distribution of fragment sizes with varying frequencies.

[0048] In some aspects, the cfDNA fragmentation profile includes sequence coverage of small cfDNA fragments in windows across the genome. In further aspects, the cfDNA fragmentation profile includes sequence coverage of large cfDNA fragments in windows across the genome. In other aspects, the cfDNA fragmentation profile includes sequence coverage of small cfDNA fragments and large cfDNA fragments in windows across the genome. In certain aspects, the mapped cfDNA fragment sequences include dozens to thousands of genomic windows. In some such aspects, the windows are non-overlapping windows. In other aspects, each window includes approximately five million base pairs. In further aspects, the cfDNA fragmentation profile covers the entire genome.

[0049] Certain aspects further include preparing a cell-free DNA (cfDNA) fragmentation profile to predict the cancer status of a subject. In certain aspects, preparing a cell-free DNA (cfDNA) fragmentation profile to predict the cancer status of a subject can include: obtaining a sample from the subject; processing the sample to obtain a plasma fraction; extracting and purifying nucleosome-protected cfDNA fragments from the plasma fraction; processing the cfDNA fragments obtained from the sample obtained from the subject into a sequencing library; and performing whole-genome sequencing on the sequencing library to obtain sequenced fragments, wherein the genomic coverage is from about 9-fold to 0.1-fold.

[0050] In another aspect, combining a clinical risk score and a genomic risk score results in a greater number of positive cancer diagnoses per subject screening compared to using the clinical risk score or the genomic risk score alone.

[0051] In certain aspects, compared to using the clinical risk score alone, the number of subject screenings required to achieve one positive cancer diagnosis is reduced by at least about 5%, 15%, 25%, 35%, 45%, 55%, 65%, 75% or more, on average.

[0052] In additional aspects, compared to using the genetic risk score alone, the number of subject screenings required to achieve one positive cancer diagnosis is reduced by at least about 5%, 10%, 15%, 20%, 25% or more, on average.

[0053] In additional aspects, combining the clinical risk score and the genomic risk score results in improved discrimination between subjects predicted to have a high cancer risk and subjects predicted to have a low cancer risk.

[0054] In another aspect, combining the clinical risk score and the genomic risk score results in higher specificity of cancer prediction compared to using the clinical risk score alone or the genomic risk score alone. In further aspects, the sensitivity of cancer prediction is at least about 50%, 60%, 70%, 80%, 90% or higher.

[0055] In a further aspect, the cancer is lung cancer. In some such aspects, the clinical risk score of the subject is determined from data including the subject's age, gender, race, smoking status, pack-years, and duration of smoking.

[0056] In certain aspects, the clinical risk score of the subject is determined from data including the Bach lung cancer incidence model, as described in Bach, P.B., et al. J NATL CANCER INST. 95(6):470-8 (2003), the description of which regarding the Bach lung cancer incidence model is incorporated herein.

[0057] In a further aspect, the cancer can be cancer at any stage. In some aspects, the cancer can be early-stage cancer. In some aspects, the cancer can be asymptomatic cancer. In some aspects, the cancer can be residual disease and / or recurrence (e.g., after surgical resection and / or after cancer therapy). The cancer can be any type of cancer. Examples of cancer types that can be assessed, monitored, and / or treated as described herein include, but are not limited to, lung cancer, colorectal cancer, prostate cancer, breast cancer, pancreatic cancer, cholangiocarcinoma, liver cancer, CNS cancer, gastric cancer, esophageal cancer, gastrointestinal stromal tumor (GIST), uterine cancer, and ovarian cancer. Additional cancer types include, but are not limited to, myeloma, multiple myeloma, B-cell lymphoma, follicular lymphoma, lymphocytic leukemia, leukemia, and myeloid leukemia. In some aspects, the cancer is a solid tumor. In some aspects, the cancer is a sarcoma, carcinoma, or lymphoma. In some aspects, the cancer is lung cancer, colorectal cancer, prostate cancer, breast cancer, pancreatic cancer, cholangiocarcinoma, liver cancer, CNS cancer, gastric cancer, esophageal cancer, gastrointestinal stromal tumor (GIST), uterine cancer, or ovarian cancer. In some aspects, the cancer is a blood cancer. In some aspects, the cancer is myeloma, multiple myeloma, B-cell lymphoma, follicular lymphoma, lymphocytic leukemia, leukemia, or myeloid leukemia.

[0058] In another aspect, combining the clinical risk score with the genomic risk score results in a combined score that increases as the cancer risk of the subject increases.

[0059] In certain aspects, a cancer treatment is administered to a subject having an increased cancer risk.

[0060] Cancer treatment can be any suitable cancer treatment. One or more cancer treatments described herein can be administered to a subject at any suitable frequency (e.g., one or more times over a period of days to weeks). Examples of cancer treatments include, but are not limited to, surgical intervention, adjuvant chemotherapy, neoadjuvant chemotherapy, radiotherapy, hormone therapy, cytotoxic therapy, immunotherapy, adoptive T cell therapy (e.g., T cells with chimeric antigen receptors and / or wild-type or modified T cell receptors), targeted therapy, such as administration of kinase inhibitors (e.g., kinase inhibitors that target specific genetic lesions such as translocations or mutations) (e.g., kinase inhibitors, antibodies, bispecific antibodies), signal transduction inhibitors, bispecific antibodies or antibody fragments (e.g., BiTEs), monoclonal antibodies, immune checkpoint inhibitors, surgery (e.g., surgical resection), or any combination of the foregoing. In some aspects, cancer treatment can reduce the severity of cancer, alleviate the symptoms of cancer, and / or reduce the number of cancer cells present in a subject.

[0061] In some aspects, the cancer treatment can be a chemotherapeutic agent. Non-limiting examples of chemotherapeutic agents include: amsacrine, azacitidine, axathioprine, bevacizumab (or its antigen-binding fragment), bleomycin, busulfan, carboplatin, capecitabine, chlorambucil, cisplatin, cyclophosphamide, cytarabine, dacarbazine, daunorubicin, docetaxel, doxifluridine, doxorubicin, epirubicin, erlotinib hydrochloride, etoposide, fiudarabine, floxuridine, fludarabine, fluorouracil, gemcitabine, hydroxyurea, idarubicin, ifosfamide, irinotecan, lomustine, mechlorethamine, melphalan, mercaptopurine, methotrxate, mitomycin, mitoxantrone, oxaliplatin, paclitaxel, pemetrexed, procarbazine, all-trans retinoic acid, streptozocin, tafluposide, temozolomide, teniposide, tioguanine, topotecan, uramustine, valrubicin, vinblastine, vincristine, vindesine, vinorelbine, and combinations thereof.Additional examples of anti-cancer therapies are known in the art; see, for example, the therapy guidelines from the American Society of Clinical Oncology (ASCO), the European Society for Medical Oncology (ESMO), or the National Comprehensive Cancer Network (NCCN).

[0062] In various aspects, DNA is present in a biological sample taken from a subject and used in the methods of the present invention. The biological sample can actually be any type of biological sample that includes DNA. The biological sample is typically a fluid, such as whole blood or a portion thereof that has circulating cfDNA. In embodiments, the sample includes DNA from a tumor or a liquid biopsy, such as but not limited to amniotic fluid, aqueous humor, vitreous humor, blood, whole blood, fractionated blood, plasma, serum, breast milk, cerebrospinal fluid (CSF), cerumen (earwax), chyle, corn, endolymph, perilymph, feces, breath, gastric acid, gastric juice, lymph fluid, mucus (including nasal drainage and phlegm), pericardial fluid, peritoneal fluid, pleural fluid, pus, nasal mucus, saliva, exhaled breath condensate, sebum, semen, sputum, sweat, synovial fluid, tears, vomit, prostatic fluid, nipple aspirate fluid, tears, sweat, buccal swab, cell lysate, gastrointestinal fluid, biopsy tissue, and urine or other biological fluids. In one aspect, the sample includes DNA from circulating tumor cells.

[0063] As disclosed above, the biological sample can be a blood sample. The blood sample can be obtained using methods known in the art, such as finger prick or venipuncture. Suitably, the blood sample is about 0.1 to 20 ml, or alternatively about 1 to 15 ml, where the volume of blood is about 10 ml. Smaller amounts can also be used, as well as the circulating free DNA in the blood. Microsampling and sampling by needle biopsy, catheter, excretion, or production of a body fluid containing DNA are also potential sources of biological samples.

[0064] The methods and systems of the present disclosure utilize nucleic acid sequence information and can thus include any method or sequencing device for performing nucleic acid sequencing, including nucleic acid amplification, polymerase chain reaction (PCR), nanopore sequencing, 454 sequencing, tagmentation sequencing. In some aspects, the methods or systems of the present disclosure utilize systems such as those provided by Illumina, Inc. (including but not limited to HiSeq TM X10, HiSeq TM 1000, HiSeq TM 2000, HiSeq TM 2500, Genome Analyzers TM 、MiSeq TM、NextSeq, NovaSeq 6000 systems), systems provided by Applied Biosystems Life Technologies (SOLiD TM systems, Ion PGM TM sequencer, ion Proton TM sequencer) or systems provided by Genapsys or BGI MGI, and other systems. Nucleic acid analysis can also be performed by systems provided by the following companies: Oxford Nanopore Technologies (GridiON TM , MiniON TM ) or Pacific Biosciences (Pacbio TM RS II or Sequel I or II).

[0065] The present invention includes systems for performing the steps of the disclosed methods and is described in part in terms of functional components and various processing steps. Such functional components and processing steps can be implemented by any number of components, operations, and techniques configured to perform the specified functions and achieve various results. For example, the present invention can employ various biological samples, biomarkers, elements, materials, computers, data sources, storage systems and media, information collection techniques and processes, data processing standards, statistical analysis, regression analysis, etc. that can perform various functions.

[0066] Accordingly, the present invention further provides a system for predicting the cancer status of a subject. In various aspects, the system includes: (a) a sequencer configured to generate a low-coverage whole-genome sequencing data set of a sample; and (b) a computer system and / or a processor having the function of performing the method of the present invention.

[0067] In some aspects, the computer system further includes one or more additional modules. For example, the system can include one or more of an extraction unit and / or a separation unit operable to select a suitable genomic component for analysis (e.g., cfDNA fragments of a specific size).

[0068] In some aspects, the computer system further includes a visual display device. The visual display device can be operable to display a curve-fitting line, a reference curve-fitting line, and / or a comparison of both.

[0069] Methods for predicting a subject's cancer status according to various aspects of the present invention can be implemented in any suitable manner, such as using a computer program operating on a computer system. As discussed herein, exemplary systems according to various aspects of the present invention can be implemented in conjunction with a computer system, which is, for example, a conventional computer system including a processor and a random access memory, such as a remotely accessible application server, a web server, a personal computer, or a workstation. The computer system also suitably includes additional memory devices or information storage systems, such as a mass storage system and a user interface, such as a conventional monitor, keyboard, and tracking device. However, the computer system can include any suitable computer system and associated devices, and can be configured in any suitable manner. In one embodiment, the computer system comprises a stand-alone system. In another embodiment, the computer system is part of a computer network including a server and a database.

[0070] Software that is required for receiving, processing, and analyzing information can be implemented in a single device or in multiple devices. The software can be accessed via a network such that the storage and processing of information occur remotely relative to the user. The systems according to various aspects of the present invention and their various elements provide functions and operations that facilitate detection and / or analysis, such as data collection, processing, analysis, reporting, and / or diagnosis. For example, in this aspect, the computer system executes a computer program that can receive, store, search, analyze, and report information related to the human genome or a region thereof. The computer program can include multiple modules that perform various functions or operations, such as a processing module for processing raw data and generating supplementary data, and an analysis module for analyzing the raw data and the supplementary data to generate a quantitative assessment of a disease state model and / or diagnostic information.

[0071] The program executed by the system can include any suitable processes for facilitating analysis and / or cancer diagnosis. In one embodiment, the system is configured to establish a disease state model and / or determine the disease state of a patient. Determining or identifying the disease state can include generating any useful information regarding the patient's condition relative to the disease, such as performing a diagnosis, providing information that aids in the diagnosis, assessing the stage or progression of the disease, identifying conditions that may indicate susceptibility to the disease, identifying whether further tests will be recommended, predicting and / or assessing the efficacy of one or more treatment procedures, or otherwise assessing the patient's disease state, the likelihood of the disease, or other health aspects.

[0072] Figure 10An example computer 800 that can be used to predict the cancer status of a subject is shown. For example, computer 800 can include a machine learning system that trains a machine learning model to predict the cancer status of a subject or a part or combination thereof in some embodiments as described above. Computer 800 can be any electronic device that runs a software application derived from compiled instructions, including but not limited to personal computers, servers, smart phones, media players, electronic tablets, game consoles, email devices, etc. In some specific implementations, computer 800 can include one or more processors 802, one or more input devices 804, one or more display devices 806, one or more network interfaces 808, and one or more computer-readable media 812. Each of these components can be coupled via a bus 810, and in some embodiments, these components can be distributed among multiple physical locations and coupled via a network.

[0073] Display device 806 can be any known display technology, including but not limited to display devices using liquid crystal display (LCD) or light emitting diode (LED) technology. Processor 802 can use any known processor technology, including but not limited to graphics processors and multi-core processors. Input device 804 can be any known input device technology, including but not limited to keyboards (including virtual keyboards), mice, trackballs, cameras, and touch-sensitive pads or displays. Bus 810 can be any known internal or external bus technology, including but not limited to ISA, EISA, PCI, PCI Express, USB, Serial ATA, or FireWire. Computer-readable media 812 can be any non-transitory medium that participates in providing instructions to processor 804 for execution, including but not limited to non-volatile storage media (e.g., optical discs, magnetic disks, flash drives, etc.) or volatile media (e.g., SDRAM, ROM, etc.).

[0074] Computer-readable media 812 can include various instructions 814 for implementing an operating system (e.g., Mac Linux). The operating system can be a multi-user, multi-processing, multi-tasking, multi-threaded, real-time operating system, etc. The operating system can perform basic tasks, including but not limited to: identifying input from input device 804; sending output to display device 806; tracking files and directories on computer-readable media 812; controlling peripheral devices (e.g., disk drives, printers, etc.) that can be controlled directly or through an I / O controller; and managing traffic on bus 810. Network communication instructions 816 can establish and maintain a network connection (e.g., software for implementing communication protocols such as TCP / IP, HTTP, Ethernet, telephone, etc.).

[0075] Machine learning instruction 818 can include instructions that enable computer 800 to act as a machine learning system and / or train a machine learning model to generate the DMS values described herein. Application 820 can be an application that uses or implements the processes described herein and / or other processes. The process can also be implemented in operating system 814. For example, application 820 and / or the operating system can create tasks in the applications described herein.

[0076] The described features can be implemented in one or more computer programs that are executable on a programmable system including at least one programmable processor coupled to receive data and instructions from, and to send data and instructions to, a data storage system, at least one input device, and at least one output device. A computer program is a set of instructions that can be used, directly or indirectly, in a computer to perform certain activities or bring about a certain result. A computer program may be written in any form of programming language, including compiled or interpreted languages (e.g., Objective-C, Java), and it may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0077] By way of example, a processor suitable for executing a program of instructions may include a general-purpose microprocessor and a special-purpose microprocessor as well as the sole processor or one of multiple processors or cores of any kind of computer. Generally, a processor may receive instructions and data from a read-only memory or a random access memory or both. Essential elements of a computer may include a processor for executing instructions and one or more memories for storing instructions and data. Generally, a computer may also include one or more mass storage devices for storing data files, or be operatively coupled to communicate with them; such devices include magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and optical disks. Storage devices suitable for tangibly embodying computer program instructions and data may include all forms of non-volatile memory, by way of example, including semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, an ASIC (application specific integrated circuit).

[0078] To provide for interaction with a user, the features can be implemented on a computer having a display device (such as an LED or LCD monitor) for displaying information to the user and a keyboard and a pointing device (such as a mouse or trackball) by which the user can provide input to the computer.

[0079] The features can be implemented on a computer system, which includes backend components such as a data server, or middleware components such as an application server or an Internet server, or frontend components such as a client computer having a graphical user interface or an Internet browser, or any combination thereof. The components of the system can be connected by any form of digital data communication (such as a communication network) or a medium of digital data communication. Examples of communication networks include, for example, a telephone network, a LAN, a WAN, and the network of computers that form the Internet.

[0080] The computer system can include a client and a server. The client and the server can generally be far from each other and usually interact via a network. The relationship between the client and the server can be created by computer programs running on the respective computers and having a client-server relationship with each other.

[0081] One or more features or steps of the disclosed embodiments can be implemented using an API. The API can define one or more parameters that are passed between a calling application and other software code (such as an operating system, a library routine, a function) that provides a service, provides data, or performs an operation or calculation.

[0082] The API can be implemented as one or more calls in program code, and the one or more calls send or receive one or more parameters through a parameter list or other structures based on the calling convention defined in the API specification document. The parameters can be constants, keys, data structures, objects, object classes, variables, data types, pointers, arrays, lists, or another call. The API calls and parameters can be implemented in any programming language. The programming language can define the vocabulary and calling convention that a programmer will use to access the functions that support the API.

[0083] In some specific implementations, the API call can report the capabilities of the device running the application to the application, such as input capabilities, output capabilities, processing capabilities, power capabilities, communication capabilities, etc.

[0084] Although various embodiments have been described above, it should be understood that the embodiments are presented by way of example and not limitation. It will be apparent to those skilled in the relevant art that various changes in form and detail can be made without departing from the spirit and scope. In fact, after reading the above specification, it will be apparent to those skilled in the relevant art how to implement alternative embodiments. For example, other steps can be provided, or steps can be eliminated from the described process, and other components can be added to or removed from the described system. Therefore, other specific embodiments are within the scope of the following claims.

[0085] In addition, it should be understood that any figures highlighting features and advantages are presented for illustrative purposes only. The disclosed methods and systems are each sufficiently flexible and configurable such that the methods and systems can be utilized in ways different from those shown.

[0086] Although the term "at least one" may often be used in the specification, claims, and drawings, the terms "a", "an", "the", "said", etc. also denote "at least one" or "the at least one" in the specification, claims, and drawings.

[0087] Finally, the applicant's intention is that, under 35 U.S.C. 112(f), only claims that include the express language "means for" or "step for" can be construed. Under 35 U.S.C. 112(f), claims that do not expressly include the phrase "means for" or "step for" will not be construed.

[0088] The methods and systems described herein can be used to detect, predict, treat, and / or monitor the cancer status in a subject. Any suitable subject, such as a mammal, can be evaluated, monitored, and / or treated as described herein. Some examples of mammals that can be evaluated, monitored, and / or treated as described herein include, but are not limited to, humans, primates such as monkeys, dogs, cats, horses, cows, pigs, sheep, mice, and rats. For example, a human having or suspected of having cancer can be evaluated using the methods described herein and optionally treated with one or more cancer treatments as described herein.

[0089] The following examples are provided to further illustrate embodiments of the invention but are not intended to limit the scope of the invention. While these are typical methods that may be used, other procedures, methods, or techniques known to those skilled in the art may alternatively be used.

[0090] Examples

[0091] Example 1 - Incorporating Clinical Risk into Biomarker - Based Assessment for Prescreening Before LDCT Lung Cancer Screening

[0092] Personalized risk assessment can improve the net benefit of LDCT screening because the probability of screening benefit varies among groups with a history of smoking. The risk of lung cancer can be estimated from clinical factors, including age and smoking history. However, blood-based biomarkers have shown promise in greatly improving risk estimation beyond clinical risk. Blood-based biomarker assessment identifies genomic features of lung cancer, which, if used as a pre-screen, can improve the efficiency of LDCT screening.

[0093] Among those eligible assessments, such assessments can distinguish individuals more and less likely to have lung cancer detected by LDCT. In clinical use, genomic features are typically interpreted using a cutoff value, with results above the cutoff value being positive and results below the cutoff value being negative. However, relying solely on genomic features ignores potential differences in clinical risk factors, such as age and smoking history. It is shown here that the integration of individual-level clinical risk with genomic features of lung cancer improves the identification of individuals most likely to have lung cancer detected by screening.

[0094] Study Participants

[0095] Data for the current study included participants in the National Lung Screening Trial (NLST). A total of 53,452 NLST participants participated in the NLST. (See Figure 1 , the figure above). 26,730 participants were randomly assigned to the x-ray study group, while 26,722 participants were randomly assigned to the spiral CT study group. The former group (x-ray study group) was excluded from the current analysis. Additionally, 1,620 participants from the spiral CT study group were excluded from the current analysis because they lacked necessary clinical data. A total of 25,102 participants met the criteria for the current analysis. (See Figure 1 , the figure below).

[0096] Determining Clinical Risk

[0097] For eligible participants, the Bach lung cancer incidence model was used to estimate the 1-year lung cancer risk for each participant, as described in Bach, P.B., et al. J NATL CANCER INST. 95(6):470-8 (2003), the description of which regarding the Bach lung cancer incidence model is incorporated herein. The 25th percentile of clinical risk was selected as the cutoff value to separate low and high clinical risk. The 1-year observed lung cancer diagnoses were predicted in logistic regression models: one using only the genomic feature score, one using only the clinical risk category, and one combining both the genomic risk and clinical risk category. These models were compared in terms of specificity at 80% sensitivity and the number of CT scans required to detect one case of lung cancer at a total prevalence of 1%. Wilson score confidence intervals were estimated.

[0098] Determining Genomic Risk

[0099] Genomic feature scores were simulated for each participant. The scores were from the cohort distribution assessed using the DELFI technology, stratified by cancer status and disease stage. The DELFI technology evaluates the fragmentation profile of cell-free DNA present in the blood and uses supervised machine learning to detect signals of cancer.

[0100] Statistical Analysis

[0101] In two additional logistic regression models, clinical risk was considered as a continuous predictor of lung cancer observed over one year: one using only clinical risk and the other combining both genomic risk and clinical risk. The predicted probability of lung cancer prognosis was determined from the logistic regression models with continuous predictors. To stratify the predicted probabilities of the multivariable models, the threshold was calculated at 80% sensitivity. These models were compared in terms of the specificity at 80% sensitivity and the number of CT scans required to detect one case of lung cancer at a total prevalence of 1%. The 95% confidence interval ("CI") was calculated using bootstrapping.

[0102] Results

[0103] The analysis included 25,102 subjects, of whom 254 (1.0%) were diagnosed with lung cancer within one year. (See Figure 2 ). The median (interquartile range) clinical risk was 0.39% (0.23 to 0.63). Thus, the cut-off value separating low and high clinical risk was 0.23%. The median risk of lung cancer was 0.15% (low-risk group) and 0.49% (high-risk group).

[0104] The models combining simulated genomic risk, clinical risk, and both were all significantly associated with lung cancer diagnosis. (p <.0001). For example, Figure 3 shows the binary clinical risk according to cancer status, while Figure 4 shows the distribution of clinical risk according to cancer status. The latter ( Figure 4 ) indicates that the median clinical risk in the lung cancer group was 0.60% (0.37 to 0.93), and the median clinical risk in the non-cancer group was 0.38% (0.23 to 0.63). Figure 4 The lines within the rectangular boxes in

[0105] Figure 5 show the median (line) and IQR (rectangular box) respectively. Figure 5 shows the distribution of simulated genomic risk according to clinical risk status. In both the low clinical risk group and the high clinical risk group, the simulated genomic risk score ranged from 0 to 1. Here again, Figure 5 the lines within the rectangular boxes show the median (line) and IQR (rectangular box) respectively.

[0106] Figure 6Shows the number of CT scans required to detect one case of lung cancer according to the type of risk estimate. The observed rate of CT scans required to detect one case of lung cancer was calculated based on the prevalence of lung cancer in the NLST CT cohort. Using genomic risk alone reduced the number of CT scans required by 32%, from 95 to 65, compared to using classification clinical risk alone. Combining genomic risk and classification clinical risk reduced the number by 37%, from 95 to 60.

[0107] Figure 7 Shows that at 80% sensitivity, incorporating clinical risk into genomic risk increased specificity from 56% (95% CI 0.55 to 0.57) to 59% (95% CI 0.58 to 0.60), and reduced the number of CT scans required to detect one case of lung cancer from 65 (genomic risk alone) to 60, a 7% reduction in the number required for LDCT screening. For reference, the number of screens required to detect one case of lung cancer without any pre-screening assessment is approximately 100.

[0108] Figure 8 Shows the predicted probability of lung cancer diagnosis using clinical risk, while Figure 9 Shows the predicted probability of lung cancer diagnosis using clinical risk and genomic risk. At 80% sensitivity, the threshold separating the low and high predicted probabilities of lung cancer diagnosis when combining clinical risk (continuous) and genomic risk was set at 0.005 (dashed line). Incorporating clinical risk into genomic risk allowed further discrimination between those below and above the threshold.

Claims

1. A method for predicting the cancer status of a subject, the method comprising: a) determining a clinical risk score of the subject; b) determining a genomic risk score of the subject; c) combining the clinical risk score with the genomic risk score to thereby predict the cancer status of the subject.

2. The method according to claim 1, wherein the clinical score comprises the age, gender, and / or race of the subject.

3. The method according to claim 1, wherein the genomic risk score comprises cell-free DNA (cfDNA) fragment size density data from the subject.

4. The method according to claim 3, wherein determining the cfDNA fragment size density data of the subject comprises: a) processing a sample containing cfDNA fragments from the subject into a library; b) performing low-coverage whole-genome sequencing on the library to obtain sequenced fragments; c) mapping the sequenced fragments to a genome to obtain windows of mapped sequences; d) analyzing the windows of the mapped sequences to determine cfDNA fragment lengths; and e) generating the cfDNA fragment size density data.

5. The method according to claim 4, wherein the cfDNA fragment size density data comprises a curve.

6. The method according to claim 5, wherein the cfDNA fragment size density curve from the subject is compared with the cfDNA fragment size density curves from known healthy subjects and / or known cancer patients.

7. The method according to claim 1, wherein combining the clinical risk score and the genomic risk score results in a greater number of positive cancer diagnoses per subject screening compared to using the clinical risk score or the genomic risk score alone.

8. The method according to claim 7, wherein the number of subject screenings required to achieve one positive cancer diagnosis is reduced by at least about 5%, 15%, 25%, 35%, 45%, 55%, 65%, 75% or more, on average, compared to using the clinical risk score alone.

9. The method according to claim 7, wherein the number of subject screenings required to achieve one positive cancer diagnosis is reduced by at least about 5%, 10%, 15%, 20%, 25% or more, on average, compared to using the genetic risk score alone.

10. The method according to claim 7, wherein the sensitivity of cancer prediction is at least about 50%, 60%, 70%, 80%, 90% or higher.

11. The method according to claim 10, wherein combining the clinical risk score and the genomic risk score results in a higher specificity of cancer prediction compared to using the clinical risk score or the genomic risk score alone.

12. The method according to claim 1, wherein combining the clinical risk score and the genomic risk score results in an improved discrimination between subjects predicted to have a high cancer risk and subjects predicted to have a low cancer risk.

13. The method according to any one of claims 1 to 12, wherein the cancer is lung cancer.

14. The method according to claim 13, wherein the clinical risk score of the subject is determined from data including the age, gender, race, smoking status, pack-years, and smoking duration of the subject.

15. The method according to claim 14, wherein the clinical risk score of the subject is determined from data including the Bach lung cancer incidence model.

16. The method according to claim 1, wherein the clinical risk score and the genomic risk score result in a combined score that increases as the cancer risk of the subject increases.

17. The method according to claim 1, wherein the clinical risk score and the genomic risk score result in a combined score that decreases as the cancer risk of the subject decreases.

18. The method according to claim 4, wherein the cfDNA fragment size density data is calculated for sub-genomic intervals.

19. The method according to claim 1, wherein a cancer treatment is administered to a subject predicted to have cancer.

20. The method according to claim 3, wherein the cfDNA is obtained from a blood sample from the subject.

21. The method according to claim 4, wherein the mapped sequence includes dozens to thousands of windows.

22. The method according to claim 4, wherein the windows are non-overlapping windows.

23. The method according to claim 4, wherein each window includes approximately five million base pairs.

24. The method according to claim 4, wherein a cfDNA fragmentation profile is determined within each window.

25. The method according to claim 24, wherein the cfDNA fragmentation profile includes the fragment size of the maximum frequency.

26. The method according to claim 24, wherein the cfDNA fragmentation profile includes the fragment size distribution of fragment sizes with varying frequencies.

27. The method according to claim 24, wherein the cfDNA fragmentation profile includes the ratio of small cfDNA fragments to large cfDNA fragments in the windows of the mapped sequence.

28. The method according to claim 24, wherein the cfDNA fragmentation profile includes the sequence coverage of small cfDNA fragments in the windows across the genome.

29. The method according to claim 24, wherein the cfDNA fragmentation profile includes the sequence coverage of large cfDNA fragments in the windows across the genome.

30. The method according to claim 24, wherein the cfDNA fragmentation profile includes the sequence coverage of small cfDNA fragments and large cfDNA fragments in the windows across the genome.

31. The method according to claim 24, wherein the cfDNA fragmentation profile is across the entire genome.

32. The method according to claim 24, wherein the cfDNA fragmentation profile is on sub-genomic intervals.