Detection and treatment of ovarian cancer

WO2025213107A3PCT designated stage Publication Date: 2026-01-15JOHNS HOPKINS UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/023274
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-04
Filing Date
2025-04-04
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Current ovarian cancer screening methods are ineffective in detecting the disease in its early stages, leading to late-stage diagnoses and high mortality rates, and existing biomarkers lack the sensitivity and specificity needed for accurate differentiation between benign and malignant ovarian masses.

Method used

A method combining cell-free DNA fragmentome analysis and protein biomarker assessment, utilizing low-coverage whole-genome sequencing and machine learning models to evaluate cfDNA fragmentation profiles and protein concentrations, such as CA-125 and HE4, to detect ovarian cancer and differentiate between cancer subtypes and benign lesions.

Benefits of technology

The integrated approach achieves high specificity and sensitivity in detecting ovarian cancer across stages I-IV, with sensitivity ranging from 72% to 100% and specificity greater than 99%, and accurately distinguishes between benign and malignant ovarian masses, providing a non-invasive screening and diagnostic tool.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025023274_15012026_PF_FP_ABST
    Figure US2025023274_15012026_PF_FP_ABST
Patent Text Reader

Abstract

Methods for detecting cancer in the early stages incorporate we used whole-genome cell free DNA (cfDNA) fragmentome and protein biomarker, for example, CA-125 and HE4).
Need to check novelty before this filing date? Find Prior Art

Description

DOCKET NO.: 348358.18302 DETECTION AND TREATMENT OF OVARIAN CANCER CROSS REFERENCE TO RELATED APPLICATIONS

[0001] This Application claims the benefit of U.S. Provisional Application 63 / 574,641 filed on April 4, 2024. The entire contents of this application are incorporated herein by reference in its entirety. STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH

[0002] This invention was made with government support under grants CA121113, CA006973, CA233259, CA062924, CA271896 and CA228991, awarded by the National Institutes of Health, and under grant W81XWH-22-1-0852, awarded by the Army Medical Research and Development Command. The government has certain rights in the invention. FIELD

[0003] The present disclosure relates in general to methods of early detection of cancer. The disclosure relates in particular to methods that evaluate cell-free DNA fragmentomes and protein biomarkers to detect ovarian cancer allowing for treating subjects in the early stages of in the early stages of the disease. BACKGROUND

[0004] Ovarian cancer is a leading cause of death in women worldwide with more than 300,000 new cases and nearly 200,000 deaths globally each year (1). In the United States during 2024, approximately 19,600 new cases will be diagnosed and 12,700 women will succumb to ovarian cancer (2). The most common form of ovarian cancer is epithelial ovarian cancer with the four major subtypes including serous, clear cell, mucinous and endometrioid carcinomas. According to the Surveillance, Epidemiology, and End Results (SEER) database, for individuals with detected invasive epithelial ovarian cancer, the estimated five-year survival is 93% and 75% for localized (Stage I) or regional disease (Stage II or Stage IIIA1 with regional lymph node involvement), respectively, compared to 31% for distant disease (remaining Stage III or Stage IV) (3,4). Unfortunately, ovarian cancer is usually detected in advanced stages (Stage III and IV) due to nonspecific clinical symptoms of the disease at earlier stages and the lack of an effectiveDOCKET NO.: 348358.18302 screening approach (3). Consequently, there is a clear unmet clinical need for the development of highly specific and sensitive assays to detect ovarian cancer in its earliest stages.

[0005] Ovarian cancer screening trials such as the Prostate, Lung, Colorectal, and Ovarian (PLCO) trial (5), the UK Collaborative Trial of Ovarian Cancer Screening (UKCTOCS)(6), and the Normal Risk Ovarian Screening Study (NROSS) (7) have shown that existing biomarkers, including cancer antigen 125 (CA-125), may provide a shift toward detection of earlier stages of cancer but not a survival benefit, likely because of suboptimal detection of all ovarian cancers. These analyses open the door to new and more effective approaches aimed at identifying combinations of biomarkers with improved performance for early ovarian cancer detection. Such approaches would need to be affordable, accessible, and have high performance for high-grade serous ovarian carcinoma (HGSOC) which is more aggressive, typically detected in late stages, and responsible for the majority of ovarian cancer deaths (8).

[0006] A secondary clinical need also exists in determining whether women presenting with ovarian masses have benign or malignant lesions. In this setting, pre-operative malignancy classification is challenging and may lead to unnecessary procedures. A number of biomarkers have been proposed in this setting, including CA-125 and human epididymis protein 4 (HE4) (9– 11), prediction models using a combination of multiple protein biomarkers as well as age and menopausal status (12), the Risk of Malignancy Index (RMI) (13), and other ultrasound classifications (International Ovarian Tumor Analysis (IOTA)) (14) but these vary in accuracy, performance, and ease of use in a clinical setting. SUMMARY

[0007] There is an unmet need for effective ovarian cancer screening and diagnostic approaches that enable earlier stage cancer detection where therapy may have greater impact.

[0008] In one aspect, we now provide a high-performing cost-efficient approach that evaluates cell-free DNA fragmentomes and protein biomarkers to detect cancer early in subject, allowing for early medical intervention.

[0009] We have found the present methods and systems can detect ovarian cancer (including for any of stages I–IV of ovarian cancer), in human subjects with high specificity andDOCKET NO.: 348358.18302 sensitivity. Additionally, we have demonstrated in human subjects differentiating benign masses from ovarian cancers with high accuracy.

[0010] Such results show that the present integrated cfDNA fragmentome and protein analyses detect ovarian cancers with high performance, enabling inter alia a new accessible approach for noninvasive ovarian cancer screening, diagnostic evaluation and / or treatment.

[0011] In one aspect, a method of early detection of cancer and treatment of a subject comprises (i) determining a cell free DNA (cfDNA) fragmentation profile of the subject, the method comprising: extracting and enriching cell free DNA (cfDNA) from a subject’s biological sample; conducting low coverage whole genome sequencing of the subject’s isolated cfDNA to generate genomic libraries of sequenced cfDNA fragments; determining a cfDNA fragmentome profile of the subject’s sequenced DNA fragments comprising evaluating one or more cfDNA fragment characteristics; (ii) comparing the cfDNA fragmentome profiles of the subject to reference cfDNA fragmentome profiles from healthy subjects to detect and diagnose the cancer; (iii) assaying the subject’s biological sample to detect and quantify at least one biomarker; (iv) comparing the subject’s cfDNA fragmentation profile and protein biomarkers with normal reference non-cancer subjects; and, treating the subject diagnosed with cancer with a cancer specific therapy.

[0012] In certain embodiments, the one or more cfDNA fragment characteristics comprises calculating ratios of small to large cfDNA fragments, sequences of one or more cfDNAs, median cfDNA fragment sizes, fragment size distribution, mutant allele frequencies, fragment length, fragment size distribution, fragment end motifs, preferred end coordinates, breakpoint motifs, methylation frequencies, concentrations of cfDNA or combinations thereof.

[0013] In certain embodiments, a small cfDNA fragment comprises about 80 base pairs (bp) to about 150 bp. In certain embodiments, a large cfDNA fragment comprises about 151 bp to about 300 bp. In certain embodiments, the small to large cfDNA ratios are GC corrected.

[0014] In certain embodiments, the cfDNA fragmentome profile comprises the sequence coverage of small cfDNA fragments in windows across the genome. In certain embodiments, the cfDNA fragmentome profile comprises the sequence coverage of cfDNA fragments in windows across the genome. In certain embodiments, the windows are non-overlapping. In certain embodiments, the windows are overlapping. In certain embodiments, the cfDNA fragmentome profile comprises the sequence coverage of small and large cfDNA fragments in windows acrossDOCKET NO.: 348358.18302 the genome. In certain embodiments, the cfDNA fragmentome profiles in subjects identified and diagnosed early with cancer, are altered across the genome. In certain embodiments, the cfDNA fragmentome profiles in subjects identified and diagnosed early with cancer, have greater heterogeneity across the genome as compared to healthy subjects.

[0015] In certain embodiments, the cancer is ovarian cancer or an adnexal mass.

[0016] In certain embodiments, the windows each comprise about 5 million base pairs. In certain embodiments, a cfDNA fragmentation profile is determined within each window. In certain embodiments, the cfDNA fragmentation profile comprises the sequence coverage of small cfDNA fragments in windows across the genome. In certain embodiments, the cfDNA fragmentation profile comprises the sequence coverage of large cfDNA fragments in windows across the genome. In certain embodiments, the cfDNA fragmentation profile comprises the sequence coverage of small and large cfDNA fragments in windows across the genome. In certain embodiments, the windows are non-overlapping. In certain embodiments, the windows are overlapping. In certain embodiments, the cfDNA fragmentation profile is over the whole genome. In certain embodiments, the cfDNA fragmentation profile is over a subgenomic interval. In certain embodiments, the subgenomic interval comprises specific locations in the genome.

[0017] In certain embodiments, the method further comprises assaying for chromosomal gains and losses in the subject’s genome as compared to a normal reference genome.

[0018] In another aspect, the method further comprises a machine learning model wherein the model incorporates genome-wide fragmentation profiles, chromosomal arm-level changes, and the concentrations of ovarian cancer biomarkers. In certain embodiments, the ovarian cancer biomarkers comprise Carbohydrate Antigen 125 (CA-125), Osteopontin (OPN), Kallikreins (KLKs), Bikunin, Human Epididymis Protein 4 (HE4), Vascular Endothelial Growth Factor (VEGF), Prostasin (PSN), Creatine Kinase B (CKB), Mesothelin, Apolipoprotein A-I (apoA-I), Transthyretin (TTR), Transferrin or combinations thereof. In certain embodiments, the ovarian cancer biomarkers comprise Carbohydrate Antigen 125 (CA-125), Human Epididymis Protein 4 (HE4) or the combination thereof. In certain embodiments, the model generates a DELFI protein (DELFI-Pro) score. In certain embodiments, the DELFI-Pro score is diagnostic of cancer. In certain embodiments, the DELFI-Pro score is diagnostic of the stage of cancer. In certain embodiments, the DELFI-Pro score is diagnostic of the subtype of ovarian cancer. InDOCKET NO.: 348358.18302 certain embodiments, the DELFI-Pro score is diagnostic of an adnexal mass. In certain embodiments, subjects with low median DELFI-Pro scores are ovarian cancer free or do not have an adnexal mass. In certain embodiments, a low median DELFI-Pro score is in a range of about 0.0001 to about 0.1. In certain embodiments, subjects with high median DELFI-Pro scores are diagnosed as having ovarian cancer or an adnexal mass. In certain embodiments, ranges of high median DELFI-Pro scores are diagnostic of stages of ovarian cancer. In certain embodiments, the higher median DELFI-Pro scores comprise greater than about 0.7.

[0019] In another aspect, a method of distinguishing between ovarian cancer, ovarian cancer subtypes and benign cancers in subjects comprises (i) determining a cell free DNA (cfDNA) fragmentation profile of the subject, the method comprising: extracting and enriching cell free DNA (cfDNA) from a subject’s biological sample; conducting low coverage whole genome sequencing of the subject’s isolated cfDNA to generate genomic libraries of sequenced cfDNA fragments; determining a cfDNA fragmentome profile of the subject’s sequenced DNA fragments comprising evaluating one or more cfDNA fragment characteristics; (ii) comparing the cfDNA fragmentome profiles of the subject to reference cfDNA fragmentome profiles from healthy subjects to detect and differentially diagnose between ovarian cancer and ovarian cancer subtypes; (iii) assaying the subject’s biological sample to detect and quantify at least one biomarker; (iv) comparing the subject’s cfDNA fragmentation profile and protein biomarkers with normal reference non-cancer subjects; and, treating the subject diagnosed with ovarian cancer or ovarian cancer subtypes with a cancer specific therapy.

[0020] In certain embodiments, the one or more cfDNA fragment characteristics comprises calculating ratios of small to large cfDNA fragments, sequences of one or more cfDNAs, median cfDNA fragment sizes, fragment size distribution, mutant allele frequencies, fragment length, fragment size distribution, fragment end motifs, preferred end coordinates, breakpoint motifs, methylation frequencies, concentrations of cfDNA or combinations thereof. In certain embodiments, a small cfDNA fragment comprises about 80 base pairs (bp) to about 150 bp and a large cfDNA fragment comprises about 151 bp to about 300 bp. In certain embodiments, the cfDNA fragmentome profile comprises the sequence coverage of small cfDNA fragments in windows across the genome. In certain embodiments, the cfDNA fragmentome profile comprises the sequence coverage of cfDNA fragments in windows across the genome. In certain embodiments, the cfDNA fragmentome profile comprises the sequence coverage of smallDOCKET NO.: 348358.18302 and large cfDNA fragments in windows across the genome. In certain embodiments, the cfDNA fragmentome profiles in subjects identified and differentially diagnosed early between ovarian cancer and ovarian cancer subtypes, are altered across the genome. In certain embodiments, the cfDNA fragmentome profiles in subjects identified and differentially diagnosed early between ovarian cancer and ovarian cancer subtypes, have greater heterogeneity across the genome as compared to healthy subjects. In certain embodiments, the ovarian cancer subtypes comprise high-grade serous (HGSOC), low-grade serous (LGSOC), clear cell, mucinous, or endometroid ovarian cancers. In certain embodiments, the windows each comprise about 5 million base pairs. In certain embodiments, a cfDNA fragmentation profile is determined within each window. In certain embodiments, the cfDNA fragmentation profile comprises the sequence coverage of small cfDNA fragments in windows across the genome. In certain embodiments, the cfDNA fragmentation profile comprises the sequence coverage of large cfDNA fragments in windows across the genome. In certain embodiments, the cfDNA fragmentation profile comprises the sequence coverage of small and large cfDNA fragments in windows across the genome. In certain embodiments, the cfDNA fragmentation profile is over the whole genome. In certain embodiments, the cfDNA fragmentation profile is over a subgenomic intervals.

[0021] In certain embodiments, the method further comprises assaying for chromosomal gains and losses in the subject’s genome as compared to a normal reference genome.

[0022] In certain embodiments, the method further comprises a machine learning model wherein the model incorporates genome-wide fragmentation profiles, chromosomal arm-level changes, and the concentrations of ovarian cancer and ovarian cancer subtype biomarkers. In certain embodiments, the method further comprises the ovarian cancer and ovarian cancer subtype biomarkers comprise Carbohydrate Antigen 125 (CA-125), Osteopontin (OPN), Kallikreins (KLKs), Bikunin, Human Epididymis Protein 4 (HE4), Vascular Endothelial Growth Factor (VEGF), Prostasin (PSN), Creatine Kinase B (CKB), Mesothelin, Apolipoprotein A-I (apoA-I), Transthyretin (TTR), Transferrin or combinations thereof. In certain embodiments, the ovarian cancer biomarkers comprise Carbohydrate Antigen 125 (CA-125), Human Epididymis Protein 4 (HE4) or the combination thereof. In certain embodiments, the model generates a DELFI protein (DELFI-Pro) score. In certain embodiments, the DELFI-Pro score is differentially diagnostic of ovarian cancer and ovarian cancer subtypes. In certain embodiments, the DELFI-Pro score is diagnostic of the stage of ovarian cancer and ovarian cancer subtypes. InDOCKET NO.: 348358.18302 certain embodiments, the DELFI-Pro score is diagnostic of an adnexal mass. In certain embodiments, subjects with low median DELFI-Pro scores are ovarian cancer and ovarian cancer subtype free or do not have an adnexal mass. In certain embodiments, a low median DELFI-Pro score is in a range of about 0.0001 to about 0.1. In certain embodiments, subjects with high median DELFI-Pro scores are diagnosed as having ovarian cancer, or an ovarian cancer subtype, or an adnexal mass. In certain embodiments, ranges of high median DELFI-Pro scores are diagnostic of stages of ovarian cancer and ovarian cancer subtypes. In certain embodiments, the higher median DELFI-Pro scores comprise scores greater than about 0.7.

[0023] Definitions

[0024] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0025] As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, to the extent that the terms “including”, “includes”, “having”, “has”, “with”, or variants thereof are used in either the detailed description and / or the claims, such terms are intended to be inclusive in a manner similar to the term “comprising.”

[0026] The term “about” or “approximately” means within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, “about” can mean within 1 or more than 1 standard deviation, per the practice in the art. Alternatively, “about” can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value or range. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude within 5-fold, and also within 2-fold, of a value. Where particular values are described in the application and claims, unless otherwise stated the term “about” meaning within an acceptable error range for the particular value should be assumed.DOCKET NO.: 348358.18302

[0027] The terms “aligned”, “alignment”, “mapped” or “aligning”, “mapping” refer to one or more sequences that are identified as a match in terms of the order of their nucleic acid molecules to a known sequence from a reference genome. Such alignment can be done manually or by a computer algorithm, examples including the Efficient Local Alignment of Nucleotide Data (ELAND) computer program distributed as part of the Illumina Genomics Analysts pipeline. The matching of a sequence read in aligning can be a 100% sequence match or less than 100% (non- perfect match).

[0028] The term “cancer” as used herein is meant, a disease, condition, trait, genotype or phenotype characterized by unregulated cell growth or replication as is known in the art. In certain embodiments, the cancer comprises ovarian cancer and ovarian cancer subtypes. In certain embodiments, the ovarian cancer subtypes comprise high-grade serous (HGSOC), low-grade serous (LGSOC), clear cell, mucinous, or endometroid ovarian cancers. Other examples of cancer include liver cancer (including hepatocellular carcinoma (HCC)), lung cancer (including non-small cell lung carcinoma), gastric cancer, colorectal cancer, as well as, for example, leukemias, e.g., acute myelogenous leukemia (AML), chronic myelogenous leukemia (CML), acute lymphocytic leukemia (ALL), and chronic lymphocytic leukemia, AIDS related cancers such as Kaposi's sarcoma; breast cancers; bone cancers such as Osteosarcoma, Chondrosarcomas, Ewing's sarcoma, Fibrosarcomas, Giant cell tumors, Adamantinomas, and Chordomas; Brain cancers such as Meningiomas, Glioblastomas, Lower- Grade Astrocytomas, Oligodendrocytomas, Pituitary Tumors, Schwannomas, and Metastatic brain cancers; cancers of the head and neck including various lymphomas such as mantle cell lymphoma, non-Hodgkins lymphoma, adenoma, squamous cell carcinoma, laryngeal carcinoma, gallbladder and bile duct cancers, cancers of the retina such as retinoblastoma, cancers of the esophagus, gastric cancers, multiple myeloma, ovarian cancer, uterine cancer, thyroid cancer, testicular cancer, endometrial cancer, melanoma, bladder cancer, prostate cancer, pancreatic cancer, sarcomas, Wilms' tumor, cervical cancer, head and neck cancer, skin cancers, nasopharyngeal carcinoma, liposarcoma, epithelial carcinoma, renal cell carcinoma, gallbladder adeno carcinoma, parotid adenocarcinoma, endometrial sarcoma, multidrug resistant cancers; and proliferative diseases and conditions, such as neovascularization associated with tumor angiogenesis.DOCKET NO.: 348358.18302

[0029] The term “cell free nucleic acid,” “cell free DNA,” or “cfDNA” refers to nucleic acid fragments that circulate in an individual's body (e.g., bloodstream) and originate from one or more healthy cells and / or from one or more cancer cells. Additionally, cfDNA may come from other sources such as viruses, fetuses, etc.

[0030] The term “cfDNA sequence coverage” refers to the average number of cfDNA molecules overlapping a specific position.

[0031] The term “circulating tumor DNA” or “ctDNA” refers to nucleic acid fragments that originate from tumor cells or other types of cancer cells, which may be released into an individual's bloodstream as result of biological processes such as apoptosis or necrosis of dying cells or actively released by viable tumor cells.

[0032] As used herein, the terms “comprising,” “comprise” or “comprised,” and variations thereof, in reference to defined or described elements of an item, composition, apparatus, method, process, system, etc. are meant to be inclusive or open ended, permitting additional elements, thereby indicating that the defined or described item, composition, apparatus, method, process, system, etc. includes those specified elements--or, as appropriate, equivalents thereof--and that other elements can be included and still fall within the scope / definition of the defined item, composition, apparatus, method, process, system, etc.

[0033] “Diagnostic” or “diagnosed” means identifying the presence or nature of a pathologic condition. Diagnostic methods differ in their sensitivity and specificity. The “sensitivity” of a diagnostic assay is the percentage of diseased individuals who test positive (percent of “true positives”). Diseased individuals not detected by the assay are “false negatives.” Subjects who are not diseased and who test negative in the assay, are termed “true negatives.” The “specificity” of a diagnostic assay is 1 minus the false positive rate, where the “false positive” rate is defined as the proportion of those without the disease who test positive. While a particular diagnostic method may not provide a definitive diagnosis of a condition, it suffices if the method provides a positive indication that aids in diagnosis.

[0034] An “effective amount” as used herein, means an amount which provides a therapeutic or prophylactic benefit.DOCKET NO.: 348358.18302

[0035] As used herein, the terms “fragmentation profile,” “fragmentome profile”, “position dependent differences in fragmentation patterns,” and “differences in fragment size and coverage in a position dependent manner across the genome” are equivalent and can be used interchangeably. In some embodiments, determining a cfDNA fragmentation profile in a mammal can be used for identifying a mammal as having cancer. For example, cfDNA fragments obtained from a mammal (e.g., from a sample obtained from a mammal) can be subjected to low coverage whole- genome sequencing, and the sequenced fragments can be mapped to the genome (e.g., in non- overlapping windows) and assessed to determine a cfDNA fragmentation profile. As described herein, a cfDNA fragmentation profile of a mammal having cancer is more heterogeneous (e.g., in fragment lengths) than a cfDNA fragmentation profile of a healthy mammal (e.g., a mammal not having cancer). As such, this disclosure also provides methods and materials for assessing, monitoring, and / or treating mammals (e.g., humans) having, or suspected of having, cancer. In some embodiments, this document provides methods and materials for identifying a mammal as having cancer. For example, a sample (e.g., a blood sample) obtained from a mammal can be assessed to determine the presence and, optionally, the tissue of origin of the cancer in the mammal based, at least in part, on the cfDNA fragmentation profile of the mammal. In some embodiments, methods and materials for monitoring a mammal as having cancer are provided. For example, a sample (e.g., a blood sample) obtained from a mammal can be assessed to determine the presence of the cancer in the mammal based, at least in part, on the cfDNA fragmentation profile of the mammal. In some embodiments, methods and materials for identifying a mammal as having cancer and administering one or more cancer treatments to the mammal to treat the mammal are provided. For example, a sample (e.g., a blood sample) obtained from a mammal can be assessed to determine if the mammal has cancer based, at least in part, on the cfDNA fragmentation profile of the mammal, and one or more cancer treatments can be administered to the mammal.

[0036] The term “genomic nucleic acid,” or “genomic DNA,” refers to nucleic acid including chromosomal DNA that originates from one or more healthy (e.g., non-tumor) cells or tumor cells. In various embodiments, genomic DNA can be extracted from a cell derived from a blood cell lineage, such as a white blood cell (WBC).DOCKET NO.: 348358.18302

[0037] “Optional” or “optionally” means that the subsequently described event or circumstance can or cannot occur, and that the description includes instances where the event or circumstance occurs and instances where it does not.

[0038] As used in this specification and the appended claims, the term “or” is generally employed in its sense including “and / or” unless the content clearly dictates otherwise.

[0039] “Parenteral” administration of an immunogenic composition includes, e.g., subcutaneous (s.c.), intravenous (i.v.), intramuscular (i.m.), or intrasternal injection, or infusion techniques.

[0040] The terms “patient” or “individual” or “subject” are used interchangeably herein, and refers to a mammalian subject to be treated, with human patients being preferred. In some embodiments, the methods of the invention find use in experimental animals, in veterinary application, and in the development of animal models for disease, including, but not limited to, rodents including mice, rats, and hamsters, and primates.

[0041] The term “reference genome” as used herein may refer to a digital or previously identified nucleic acid sequence database, assembled as a representative example of a species or subject. Reference genomes may be assembled from the nucleic acid sequences from multiple subjects, sample or organisms and does not necessarily represent the nucleic acid makeup of a single person. Reference genomes may be used to for mapping of sequencing reads from a sample to chromosomal positions. For example, a reference genome used for human subjects as well as many other organisms is found at the National Center for Biotechnology Information at ncbi.nlm.nih.gov.

[0042] The term “read segment” or “read” refers to any nucleotide sequences including sequence reads obtained from an individual and / or nucleotide sequences derived from the initial sequence read from a sample obtained from an individual.

[0043] The terms “sample,” “patient sample,” “biological sample,” and the like, encompass a variety of sample types obtained from a patient, individual, or subject and can be used in a diagnostic, prognostic and / or monitoring assay. The patient sample may be obtained from a healthy subject, a diseased patient, or a patient with lung cancer. In certain embodiments, a sample that is “provided” can be obtained by the person (or machine) conducting the assay, or itDOCKET NO.: 348358.18302 can have been obtained by another, and transferred to the person (or machine) carrying out the assay. Moreover, a sample obtained from a patient can be divided and only a portion may be used for diagnosis. Further, the sample, or a portion thereof, can be stored under conditions to maintain sample for later analysis. The definition specifically encompasses blood and other liquid samples of biological origin (including, but not limited to, peripheral blood, serum, plasma, cord blood, amniotic fluid, cerebrospinal fluid, urine, saliva, stool and synovial fluid), solid tissue samples such as a biopsy specimen or tissue cultures or cells derived therefrom and the progeny thereof. In certain embodiment, a sample comprises cerebrospinal fluid. In a specific embodiment, a sample comprises a blood sample. In another embodiment, a sample comprises a plasma sample. In yet another embodiment, a serum sample is used. The definition of “sample” also includes samples that have been manipulated in any way after their procurement, such as by centrifugation, filtration, precipitation, dialysis, chromatography, treatment with reagents, washed, or enriched for certain cell populations. The terms further encompass a clinical sample, and also include cells in culture, cell supernatants, tissue samples, organs, and the like. Samples may also comprise fresh-frozen and / or formalin-fixed, paraffin-embedded tissue blocks, such as blocks prepared from clinical or pathological biopsies, prepared for pathological analysis or study by immunohistochemistry.

[0044] The term “sequence reads” refers to nucleotide sequences read from a sample obtained from an individual. Sequence reads can be obtained through various methods known in the art.

[0045] As defined herein, a “therapeutically effective” amount of a compound or agent (i.e., an effective dosage) means an amount sufficient to produce a therapeutically (e.g., clinically) desirable result. The compositions can be administered from one or more times per day to one or more times per week, including once every other day. The skilled artisan will appreciate that certain factors can influence the dosage and timing required to effectively treat a subject, including but not limited to the severity of the disease or disorder, previous treatments, the general health and / or age of the subject, and other diseases present. Moreover, treatment of a subject with a therapeutically effective amount of the compounds of the invention can include a single treatment or a series of treatments.

[0046] As used herein, the terms “treat,” treating,” “treatment,” and the like refer to reducing or ameliorating a disorder and / or symptoms associated therewith. It will be appreciatedDOCKET NO.: 348358.18302 that, although not precluded, treating a disorder or condition does not require that the disorder, condition or symptoms associated therewith be completely eliminated.

[0047] Genes: All genes, gene names, and gene products disclosed herein are intended to correspond to homologs from any species for which the compositions and methods disclosed herein are applicable. It is understood that when a gene or gene product from a particular species is disclosed, this disclosure is intended to be exemplary only, and is not to be interpreted as a limitation unless the context in which it appears clearly indicates. Thus, for example, for the genes or gene products disclosed herein, are intended to encompass homologous and / or orthologous genes and gene products from other species.

[0048] Ranges: throughout this disclosure, various aspects of the invention can be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 2.7, 3, 4, 5, 5.3, and 6. This applies regardless of the breadth of the range.

[0049] Any compositions or methods provided herein can be combined with one or more of any of the other compositions and methods provided herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawings will be provided by the Office upon request and payment of the necessary fee.

[0051] FIG. 1 is a schematic of ovarian cancer detection in screening and diagnostic models combining DELFI and protein biomarkers. Individuals undergo a blood collection, plasma is extracted and constructed genomic libraries undergo whole-genome sequencing at low coverage (~2x). Using blood samples from the same collection, proteins are quantified enabling the combined assessment of genome-wide fragmentation profiles and protein biomarkers. TheseDOCKET NO.: 348358.18302 features are evaluated in a machine learning model that classifies cancer and non-cancer individuals.

[0052] FIGS.2A-2D demonstrate the characteristics of cfDNA fragmentation for ovarian cancer detection. (FIG.2A) Fragmentation profiles where each line represents one participant and is colored according to that participant’ correlation to the median genome-wide profile for women without cancer. (FIG.2B) Heatmap of fragmentation and protein features show marked heterogeneity in the cfDNA fragmentome among individuals with ovarian cancer compared to those with benign lesions or without disease. In the heatmap, individuals are split into disease groups, and then successively ordered by DELFI-Pro, CA-125, and HE4, and for cancers, Stage and Subtype. Fragmentation features are clustered in columns. The top bar indicates the feature family containing the short to long ratio of fragment sizes (ratio) and chromosomal arm representation (z-scores). (FIG.2C) Chromosomal gains (red) and losses (blue) characteristic of ovarian cancer tumor tissue evaluated in TCGA were observed in cfDNA fragmentation data in patients with ovarian cancer and absent from those with benign lesions or without disease (red represents gains, while losses are blue, and purple are no changes in chromosomal representation, respectively). (FIG. 2D) Feature importance, as measured by scaled coefficients from the penalized logistic regression locked screening model for ovarian cancer, demonstrates contributions of cfDNA fragmentation (fragment length and aneuploidy) and proteins (CA-125 and HE4) to high performance.

[0053] FIGS. 3A-3D demonstrate that DELFI-Pro detects ovarian cancer with high sensitivity and specificity. (FIG.3A) In the Discovery Cohort, patients with ovarian cancer across all stages have elevated DELFI-Pro scores in HGSOC as well as other ovarian subtypes. (FIG. 3B) ROC analyses of the Discovery Cohort show high performance across stages and in HGSOC. (FIGS. 3C, 3D) The locked DELFI-Pro model at locked thresholds (for example, for 99% specificity, DELFI-Pro score >0.66) showed similar performance in the Validation Cohort.

[0054] FIGS.4A-4D show a modelling the implementation of DELFI for ovarian cancer screening. (FIG.4A) The proposed approach integrates the use of cell-free DNA fragmentation and protein analyses from a blood draw. Women with a positive result would undergo a transvaginal ultrasound and if positive would subsequently have a diagnostic cancer workup. A negative result at any step in this continuum would remove patients from subsequent steps andDOCKET NO.: 348358.18302 lead to annual screening. (FIG.4B) Modelling a theoretical population of 100,000 women based on existing performances for CA-125, HE4, as well those observed for DELFI-Pro. Predictive distributions for the (FIG.4C) positive predictive value and (FIG.4D) false positive rate highlight the potential benefit of implementing DELFI-Pro as compared to existing biomarkers.

[0001] FIGS.5A-5C are a series of plots demonstrating the correlation between screening DELFI-Pro score and comorbidities in individuals without cancer. Analyses were performed for individuals where clinical information was available (n=158). (FIG. 5A) DELFI-Pro scores for women with diabetes were not significantly different (p = 1, Wilcoxon, Bonferroni corrected). (FIG.5B) Women with arteriosclerosis, or (FIG.5C) hypertension did have significant differences in their DELFI-Pro scores (none vs arteriosclerosis, p=0.12; none vs hypertension, p=0.066; Wilcoxon, Bonferroni corrected).

[0002] FIGS. 6A-6C are a series of plots demonstrating DELFI-Pro score evaluation in available clinical characteristics of patients with cancer. (FIG.6A) Age at diagnosis and DELFI- Pro scores were not correlated. (FIG.6B) Comparing women with that presented clinically with asymptomatic or with symptoms as well as those (FIG. 6C) post-menopause to pre-menopausal status were not significant (asymptomatic vs symptomatic, p = 0.61; postmenopausal vs premenopausal, p=0.36; Wilcoxon).

[0003] FIG.7 is a series of plots demonstrating the detection of ovarian cancer subtypes using DELFI-Pro screening model. Robust performance across the main subtypes of ovarian cancer, including high-grade (also shown in FIG. 3B) and low-grade serous as well as endometrioid, mucinous, and clear cell were observed in the Discovery Cohort.

[0004] FIG. 8 is a plot demonstrating the detection of ovarian cancer subtypes using DELFI-Pro at high specificity. Partial AUCs across all cancers, high-grade serous ovarian cancers, and at different stages shows that DELFI-Pro performs better than either protein biomarker alone at high specificity.

[0005] FIG.9 is a plot demonstrating that genome-wide fragmentation profiles are altered in patients in the Validation Cohort with ovarian cancer. Fragmentation profiles for women with cancer (n=40) were disorganized genome-wide compared to those without cancer (n=22). The shade of blue indicates the level of correlation to patients without cancer in the Validation Cohort.DOCKET NO.: 348358.18302

[0006] FIGS. 10A-10C are a series of schematics showing Analyses of chromosomal changes in Discovery and Validation cohorts. Chromosomal gains (red) and losses (blue) characteristic of ovarian cancer tumor tissue evaluated in TCGA (FIG. 10A) were observed in cfDNA fragmentation data in patients with ovarian cancer and absent from those with benign lesions or without disease (red represents gains, while losses are blue, and purple are no changes in chromosomal representation, respectively) in Discovery (FIG.10B) or Validation (FIG.10C) cohorts.

[0007] FIGS.11A-11B are a series of plots demonstrating the assessment of DELFI-Pro in women with benign lesions. (FIG. 11A) Women with benign lesions that were asymptomatic compared to women that were symptomatic did not have a significant difference in DELFI-Pro scores (p = 0.97, Wilcoxon). (FIG. 11B) Evaluation of benign lesion sizes (cm) in women with masses less than 5 cm, 5-10 cm or greater than 10 cm did not identify differences in DELFI-Pro scores (< 5cm vs 5-10cm, p = 0.67; less than 5cm vs >10cm, p= 0.4; Wilcoxon).

[0008] FIGS.12A-12D are a series of lots demonstrating the performance of DELFI-Pro distinguishing ovarian cancer from benign masses. (FIG. 12A) Evaluation of DELFI-Pro diagnostic model score across women with benign adnexal masses (n=203) and different cancer stages as well as subtypes in Discovery Cohort. (FIG.12B) Evaluation of diagnostic model across all cancers in the Discovery Cohort (AUC 0.88), high-grade serous ovarian cancer (HGSOC) (AUC 0.96), and cancers by stages. (FIG.12C) The locked diagnostic model was evaluated in the Validation Cohort using different specificity thresholds (for example, for 80% specificity, DELFI- Pro score >0.25) for sensitivity and (FIG.12D) revealed high performance for all cancers as well as for HGSOC detection.

[0009] FIG. 13 is a series of plots demonstrating the performance of DELFI-Pro for distinguishing ovarian cancer subtypes from benign masses. Robust performance across the main subtypes of ovarian cancer, including high-grade (also shown in FIGS.12A-12D) and low-grade serous as well as endometrioid, mucinous, and clear cell were observed in the Discovery Cohor.

[0010] FIGS.14A-14B is a series of plots showing a stability analysis across fold repeats and collection source. (FIG. 14A) DELFI-Pro scores in the Discovery Screening model acrossDOCKET NO.: 348358.18302 different fold repeats for non-cancers and cancers. (FIG. 14B) DELFI-Pro scores separated by sample collection source.

[0011] FIGS.15A-15B show an ROC analyses of asymptomatic individuals in the using screening or diagnostic models in the Discovery Cohort. The high performance for classification of ovarian cancer is maintained in the Discovery (FIG.15A) screening or (FIG.15B) diagnostic model for individuals with an asymptomatic presentation.

[0012] FIGS.16A-16D are a series of plots demonstrating performance of ichorCNA and median cfDNA fragment lengths in the Discovery Cohorts. Evaluation of ichorCNA or median cfDNA fragment lengths in a (FIG.16A) screening or (FIG.16B) diagnostic applications reveal low performance. (FIGS. 16C, 16D) Comparison of DELFI-Pro scores with ichorCNA reveals that DELFI-Pro scores detected cancers that were below the ichorCNA limit of detection.

[0013] FIGS.17A-17D are a series of plots demonstrating performance of DELFI-Pro for detection of ovarian cancer in the Validation Cohort. (FIG. 17A) Boxplot demonstrating the distribution of DELFI-Pro screening model scores in the Validation Cohort as well as ROCs by (FIG.17B) stage and (FIG.17C) subtype.

[0014] FIG.18 is a plot showing the correlation of rank ordered DELFI-Pro scores for the Screening and Diagnostic models. DELFI-Pro models demonstrate a high correlation among rank ordered scores for all ovarian cancer patients in the Discovery Cohort (n = 94).

[0015] FIGS.19A-19C are a series of plots demonstrating performance of DELFI-Pro for distinguishing ovarian cancer subtypes from benign masses in Validation Cohort. (FIG.19A) The distribution of DELFI-Pro diagnostic model scores in the Validation Cohort as well as ROCs for (FIG.19B) stage and (FIG.19C) subtype.

[0016] FIGS. 20A-20B are a series of plots showing an evaluation of CA125 and HE4 blood concentrations measured at different centers. (FIG. 20A) Protein quantities of CA-125 measured for individuals with cancer (red) and benign masses (light blue) at the NKI in serum and at JHU in plasma demonstrated a high-correlation (R=0.85, Pearson correlation test), with one patient with cancer as an outlier (CA-125 value at JHU >4,000 U / ml and approximately 500 U / ml at NKI, not shown). (FIG.20B) Comparison of HE4 levels in plasma (pM) at NKI and JHU sites had very high correlation (R=1, Pearson correlation test).DOCKET NO.: 348358.18302 DETAILED DESCRIPTION

[0017] Ovarian cancer is a leading cause of death from gynecological malignancies worldwide. No effective screening methods exist resulting in tumor detection at advanced stages where treatment is less effective.

[0018] As disclosed herein, we demonstrate inter alia whole-genome cell-free DNA (cfDNA) fragmentome and protein biomarker analyses to evaluate 591 women from the European Union or the United States with ovarian cancer, benign adnexal masses, or without ovarian lesions. Using a machine learning model that incorporated multi-analyte fragmentome data and protein measurements, ovarian cancer was detected with high specificity >99% and sensitivity of 72%, 69%, 87%, and 100% for stages I–IV, respectively (AUC=0.96, 95% CI: 0.94-0.99), including 90% of high grade serous ovarian cancers. At the same specificity, CA-125 alone detected 34%, 62%, 63%, and 100% of ovarian cancers for stages I–IV (p=0.001, two- sided test of equal proportions). Additionally, this approach differentiated benign masses from ovarian cancers with high accuracy (AUC of 0.88, 95% CI=0.83-0.92). These results were validated in an independent population. These findings show that integrated cfDNA fragmentome and protein analyses detect ovarian cancers with high performance, enabling a new accessible approach for noninvasive ovarian cancer screening and / or diagnostic evaluation as well as treatment of the detected cancer.

[0019] In one aspect, a method of early detection of cancer and treatment of a subject comprises (i) determining a cell free DNA (cfDNA) fragmentation profile of the subject, the method comprising: extracting and enriching cell free DNA (cfDNA) from a subject’s biological sample; conducting low coverage whole genome sequencing of the subject’s isolated cfDNA to generate genomic libraries of sequenced cfDNA fragments; determining a cfDNA fragmentome profile of the subject’s sequenced DNA fragments comprising evaluating one or more cfDNA fragment characteristics; (ii) comparing the cfDNA fragmentome profiles of the subject to reference cfDNA fragmentome profiles from healthy subjects to detect and diagnose the cancer; (iii) assaying the subject’s biological sample to detect and quantify at least one biomarker; (iv) comparing the subject’s cfDNA fragmentation profile and protein biomarkers with normal reference non-cancer subjects; and, treating the subject diagnosed with cancer with a cancer specific therapy.DOCKET NO.: 348358.18302

[0020] In another aspect, a method of distinguishing between ovarian cancer, ovarian cancer subtypes and benign cancers in subjects comprises (i) determining a cell free DNA (cfDNA) fragmentation profile of the subject, the method comprising: extracting and enriching cell free DNA (cfDNA) from a subject’s biological sample; conducting low coverage whole genome sequencing of the subject’s isolated cfDNA to generate genomic libraries of sequenced cfDNA fragments; determining a cfDNA fragmentome profile of the subject’s sequenced DNA fragments comprising evaluating one or more cfDNA fragment characteristics; (ii) comparing the cfDNA fragmentome profiles of the subject to reference cfDNA fragmentome profiles from healthy subjects to detect and differentially diagnose between ovarian cancer and ovarian cancer subtypes; (iii) assaying the subject’s biological sample to detect and quantify at least one biomarker; (iv) comparing the subject’s cfDNA fragmentation profile and protein biomarkers with normal reference non-cancer subjects; and, treating the subject diagnosed with ovarian cancer or ovarian cancer subtypes with a cancer specific therapy.

[0021] DNA Evaluation of Fragments for early Interception (DELFI)

[0022] DNA Evaluation of Fragments for early Interception (DELFI) was previously developed, Cristiano S, Leal A, Phallen J, et al. Genome-wide cell-free DNA fragmentation inpatients with cancer. Nature 2019;570:385-9 incorporated herein in its entirety, and used toevaluate genome-wide fragmentation patterns of cfDNA of 236 patients with breast, colorectal, lung, ovarian, pancreatic, gastric, or bile duct cancers as well as 245 healthy individuals. These analyses revealed that cfDNA profiles of healthy individuals reflected nucleosomal fragmentation patterns of white blood cells, while patients with cancer had altered fragmentation profiles. DELFI had sensitivities of detection ranging from 57% to >99% among the seven cancer types at 98% specificity and identified the tissue of origin of the cancers to a limited number of sites in 75% of embodiments. Assessing cfDNA (e.g., using DELFI) provide a screening approach for early detection of cancer, which can increase the chance for successful treatment of a patient having cancer. Assessing cfDNA (e.g., using DELFI) can also provide an approach for monitoring cancer, which can increase the chance for successful treatment and improved outcome of a patient having cancer. In addition, a cfDNA fragmentation profile can be obtained from limited amounts of cfDNA and using inexpensive reagents and / or instruments.

[0023] A DELFI score can be generated, wherein the principle component analysis is incorporated into a machine learning predictive model to generate a score for each subject as anDOCKET NO.: 348358.18302 average over cross-validation repeats (DELFI score(s)). In certain embodiments, the DELFI scores for non-cancer individuals are less than about 0.3. In certain embodiments, the DELFI scores for stage I cancer are between about 0.3 to less than 0.5. In certain embodiments, the DELFI scores for stage II cancer are between about 0.5 to less than 0.8. In certain embodiments, the DELFI scores for stage III cancer are between about 0.8 to less than 0.99. In certain embodiments, the DELFI scores for stage IV cancer are about 0.99 or greater. In certain embodiments, the DELFI score for stage I cancer is about 0.35. In certain embodiments, the DELFI score for stage II cancer is about 0.75. In certain embodiments, the DELFI score for stage III cancer is about 0.9. In certain embodiments, the DELFI score for stage IV cancer is about 0.99.

[0024] In certain embodiments, a method of diagnosing cancer in a subject, comprises extracting cell free (cfDNA) from the subject’s biological sample; generating genomic libraries from the extracted cfDNA and whole genome sequencing of cfDNA fragments; mapping of the cfDNA fragments to a genomic origin and evaluating fragment length and obtaining genome- wide fragmentation profiles for each sample; identifying protein biomarkers of the subject; comparing the subject’s cfDNA fragmentation profile and protein biomarkers with normal reference non-cancer subjects.

[0025] DELFI-PRO: As disclosed in detail on the examples section, two machine learning models that incorporate multi-analyte fragmentome data and protein measurements, to predict the presence of ovarian cancer in (i) a screening setting, and (ii) a diagnostic setting. Both models used Penalized logistic regression and features included fragmentation profiles, chromosomal arm-level changes, as well as the protein biomarkers CA-125 and HE4. The models were trained and cross-validated using data from (i) individuals in the subset of the Discovery group with ovarian cancer or without any known ovarian lesions for the screening model, and (ii) Individuals in the subset of the Discovery group with ovarian cancer or benign adnexal masses for the diagnostic model. The principal components of the ratios representing greater than 90% of variance and the z-scores (21,22), along with levels of the protein biomarkers CA-125 and HE4, were used to train machine learning models. Training was performed with 10 repeats of 5-fold cross validation, generating a DELFI-Pro score for every individual in the Discovery Cohort, that was the average over 10 cross-validation repeats. For the Validation Cohorts, DELFI-Pro scores were generated using the locked models. Performance ofDOCKET NO.: 348358.18302 the models was assessed using receiver-operator curve analyses, and at fixed score thresholds for set specificities in the Discovery Cohort.

[0026] Other biomarkers include (note, the cancer indications indicated represent non- limiting examples): aminopeptidase N (CD13), annexin Al, B7-H3 (CD276, various cancers), CA125 (ovarian cancers), CA15-3 (carcinomas), CA19-9 (carcinomas), L6 (carcinomas), Lewis Y (carcinomas), Lewis X (carcinomas), alpha fetoprotein (carcinomas), CA242 (colorectal cancers), placental alkaline phosphatase (carcinomas), prostate specific antigen (prostate), prostatic acid phosphatase (prostate), epidermal growth factor (carcinomas), CD2 (Hodgkin's disease, NHL lymphoma, multiple myeloma), CD3 epsilon (T cell lymphoma, lung, breast, gastric, ovarian cancers, autoimmune diseases, malignant ascites), CD19 (B cell malignancies), CD20 (non-Hodgkin's lymphoma, B-cell neoplasmas, autoimmune diseases), CD21 (B-cell lymphoma), CD22 (leukemia, lymphoma, multiple myeloma, SLE), CD30 (Hodgkin's lymphoma), CD33 (leukemia, autoimmune diseases), CD38 (multiple myeloma), CD40 (lymphoma, multiple myeloma, leukemia (CLL)), CD51 (metastatic melanoma, sarcoma), CD52 (leukemia), CD56 (small cell lung cancers, ovarian cancer, Merkel cell carcinoma, and the liquid tumor, multiple myeloma), CD66e (carcinomas), CD70 (metastatic renal cell carcinoma and non- Hodgkin lymphoma), CD74 (multiple myeloma), CD80 (lymphoma), CD98 (carcinomas), CD123 (leukemia), mucin (carcinomas), CD221 (solid tumors), CD227 (breast, ovarian cancers), CD262 (NSCLC and other cancers), CD309 (ovarian cancers), CD326 (solid tumors), CEACAM3 (colorectal, gastric cancers), CEACAM5 (CEA, CD66e) (breast, colorectal and lung cancers), DLL4 (A-like-4), EGFR (various cancers), CTLA4 (melanoma), CXCR4 (CD 184, heme-oncology, solid tumors), Endoglin (CD 105, solid tumors), EPCAM (epithelial cell adhesion molecule, bladder, head, neck, colon, NHL prostate, and ovarian cancers), ERBB2 (lung, breast, prostate cancers), FCGR1 (autoimmune diseases), FOLR (folate receptor, ovarian cancers), FGFR (carcinomas), GD2 ganglioside (carcinomas), G-28 (a cell surface antigen glycolipid, melanoma), GD3 idiotype (carcinomas), heat shock proteins (carcinomas), HER1 (lung, stomach cancers), HER2 (breast, lung and ovarian cancers), HLA-DR10 (NHL), HLA- DRB (NHL, B cell leukemia), human chorionic gonadotropin (carcinomas), IGF1R (solid tumors, blood cancers), IL-2 receptor (T-cell leukemia and lymphomas), IL-6R (multiple myeloma, RA, Castleman's disease, IL6 dependent tumors), integrins (ανβ3, α5β1, α6β4, α11β3, α5β5, ανβ5, for various cancers), MAGE-1 (carcinomas), MAGE-2 (carcinomas), MAGE-3DOCKET NO.: 348358.18302 (carcinomas), MAGE 4 (carcinomas), anti-transferrin receptor (carcinomas), p97 (melanoma), MS4A1 (membrane-spanning 4-domains subfamily A member 1, Non-Hodgkin's B cell lymphoma, leukemia), MUC1 (breast, ovarian, cervix, bronchus and gastrointestinal cancer), MUC16 (CA125) (ovarian cancers), CEA (colorectal cancer), gp100 (melanoma), MARTI (melanoma), MPG (melanoma), MS4A1 (membrane-spanning 4-domains subfamily A, small cell lung cancers, NHL), nucleolin, Neu oncogene product (carcinomas), P21 (carcinomas), nectin-4 (carcinomas), paratope of anti-(N- glycolylneuraminic acid, breast, melanoma cancers), PLAP- like testicular alkaline phosphatase (ovarian, testicular cancers), PSMA (prostate tumors), PSA (prostate), ROB04, TAG 72 (tumor associated glycoprotein 72, AML, gastric, colorectal, ovarian cancers), T cell transmembrane protein (cancers), Tie (CD202b), tissue factor, TNFRSF10B (tumor necrosis factor receptor superfamily member 10B, carcinomas), TNFRSF13B (tumor necrosis factor receptor superfamily member 13B, multiple myeloma, NHL, other cancers, RA and SLE), TPBG (trophoblast glycoprotein, renal cell carcinoma), TRAIL-R1 (tumor necrosis apoptosis inducing ligand receptor 1, lymphoma, NHL, colorectal, lung cancers), VCAM-1 (CD106, Melanoma), VEGF, VEGF-A, VEGF-2 (CD309) (various cancers). Some other tumor associated antigens have been reviewed (Gerber, et al, mAbs 20091:247-253; Novellino et al, Cancer Immunol Immunother.200554:187-207, Franke, et al, Cancer Biother Radiopharm. 2000, 15:459-76, Guo, et al., Adv Cancer Res.2013; 119: 421–475, Parmiani et al. J Immunol. 2007178:1975-9). Examples of these antigens include Cluster of Differentiations (CD4, CD5, CD6, CD7, CD8, CD9, CD10, CDl la, CDl lb, CDl lc, CD12w, CD14, CD15, CD16, CDwl7, CD18, CD21, CD23, CD24, CD25, CD26, CD27, CD28, CD29, CD31, CD32, CD34, CD35, CD36, CD37, CD41, CD42, CD43, CD44, CD45, CD46, CD47, CD48, CD49b, CD49c, CD53, CD54, CD55, CD58, CD59, CD61, CD62E, CD62L, CD62P, CD63, CD68, CD69, CD71, CD72, CD79, CD81, CD82, CD83, CD86, CD87, CD88, CD89, CD90, CD91, CD95, CD96, CD100, CD103, CD105, CD106, CD109, CD117, CD120, CD127, CD133, CD134, CD135, CD138, CD141, CD142, CD143, CD144, CD147, CD151, CD152, CD154, CD156, CD158, CD163, CD166, .CD168, CD184, CDwl86, CD195, CD202 (a, b), CD209, CD235a, CD271, CD303, CD304), annexin Al, nucleolin, endoglin (CD105), ROB04, amino-peptidase N, -like-4 (DLL4), VEGFR-2 (CD309), CXCR4 (CD184), Tie2, B7-H3, WT1, MUC1, LMP2, HPV E6 E7, EGFRvIII, HER-2 / neu, idiotype, MAGE A3, p53 nonmutant, NY-ESO-1, GD2, CEA, MelanA / MARTl, Ras mutant, gp100, p53 mutant, proteinase3 (PR1), bcr-abl, tyrosinase,DOCKET NO.: 348358.18302 survivin, hTERT, sarcoma translocation breakpoints, EphA2, PAP, ML-IAP, AFP, EpCAM, ERG (TMPRSS2 ETS fusion gene), NA17, PAX3, ALK, androgen receptor, cyclin B l, polysialic acid, MYCN, RhoC, TRP-2, GD3, fucosyl GMl , mesothelin, PSCA, MAGE Al, sLe(a), CYPIB I, PLACl, GM3, BORIS, Tn, GloboH, ETV6-AML, NY-BR-1, RGS5, SART3, STn, carbonic anhydrase IX, PAX5, OY-TES1, sperm protein 17, LCK, HMWMAA, AKAP-4, SSX2, XAGE 1, B7H3, legumain, Tie 2, VEGFR2, MAD- CT-1, FAP, PDGFR-β, MAD-CT-2, Notch1, ICAM1 and Fos-related antigen 1.

[0027] In certain embodiments a DELFI-Pro score is generated for each subject wherein the DELFI-Pro score is diagnostic of cancer. In certain embodiments the DELFI-Pro score is diagnostic of the stage of cancer. In certain embodiments the DELFI-Pro score is diagnostic of the subtype of ovarian cancer. In certain embodiments the DELFI-Pro score is diagnostic of an adnexal mass. In certain embodiments subjects with low median DELFI-Pro scores are ovarian cancer free or do not have an adnexal mass. In certain embodiments a low median DELFI-Pro score is in a range of about 0.0001 to about 0.1. In certain embodiments subjects with high median DELFI-Pro scores are diagnosed as having ovarian cancer or an adnexal mass. In certain embodiments ranges of high median DELFI-Pro scores are diagnostic of stages of ovarian cancer. In certain embodiments the higher median DELFI-Pro scores comprise greater than about 0.7.

[0028] cfDNA Fragmentation Profiles: A cfDNA fragmentation profile can include one or more cfDNA fragmentation patterns. A cfDNA fragmentation pattern can include any appropriate cfDNA fragmentation pattern. Examples of cfDNA fragmentation patterns include, without limitation, median fragment size, fragment size distribution, ratio of small cfDNA fragments to large cfDNA fragments, and the coverage of cfDNA fragments. In some embodiments, a cfDNA fragmentation pattern includes two or more (e.g., two, three, or four) of median fragment size, fragment size distribution, ratio of small cfDNA fragments to large cfDNA fragments, and the coverage of cfDNA fragments. In some embodiments, cfDNA fragmentation profile can be a genome-wide cfDNA profile (e.g., a genome-wide cfDNA profile in windows across the genome). In some embodiments, cfDNA fragmentation profile can be a targeted region profile. A targeted region can be any appropriate portion of the genome (e.g., a chromosomal region). Examples of chromosomal regions for which a cfDNA fragmentation profile can be determined as described herein include, without limitation, a portion of a chromosome (e.g., a portion of 2q,DOCKET NO.: 348358.18302 4p, 5p, 6q, 7p, 8q, 9q, 10q, 11q, 12q, and / or 14q) and a chromosomal arm (e.g., a chromosomal arm of 8q,13q, 11q, and / or 3p). In some embodiments, a cfDNA fragmentation profile can include two or more targeted region profiles.

[0029] In some embodiments, a cfDNA fragmentation profile can be used to identify changes (e.g., alterations) in cfDNA fragment lengths. An alteration can be a genome-wide alteration or an alteration in one or more targeted regions / loci. A target region can be any region containing one or more cancer-specific alterations. In some embodiments, a cfDNA fragmentation profile can be used to identify (e.g., simultaneously identify) from about 10 alterations to about 500 alterations (e.g., from about 25 to about 500, from about 50 to about 500, from about 100 to about 500, from about 200 to about 500, from about 300 to about 500, from about 10 to about 400, from about 10 to about 300, from about 10 to about 200, from about 10 to about 100, from about 10 to about 50, from about 20 to about 400, from about 30 to about 300, from about 40 to about 200, from about 50 to about 100, from about 20 to about 100, from about 25 to about 75, from about 50 to about 250, or from about 100 to about 200, alterations).

[0030] In some embodiments, a cfDNA fragmentation profile can be used to detect tumor-derived DNA. For example, a cfDNA fragmentation profile can be used to detect tumor- derived DNA by comparing a cfDNA fragmentation profile of a mammal having, or suspected of having, cancer to a reference cfDNA fragmentation profile (e.g., a cfDNA fragmentation profile of a healthy mammal and / or a nucleosomal DNA fragmentation profile of healthy cells from the mammal having, or suspected of having, cancer). In some embodiments, a reference cfDNA fragmentation profile is a previously generated profile from a healthy mammal. For example, methods provided herein can be used to determine a reference cfDNA fragmentation profile in a healthy mammal, and that reference cfDNA fragmentation profile can be stored (e.g., in a computer or other electronic storage medium) for future comparison to a test cfDNA fragmentation profile in mammal having, or suspected of having, cancer. In some embodiments, a reference cfDNA fragmentation profile (e.g., a stored cfDNA fragmentation profile) of a healthy mammal is determined over the whole genome. In some embodiments, a reference cfDNA fragmentation profile (e.g., a stored cfDNA fragmentation profile) of a healthy mammal is determined over a subgenomic interval.DOCKET NO.: 348358.18302

[0031] In some embodiments, a cfDNA fragmentation profile can be used to identify a mammal (e.g., a human) as having cancer (e.g., a colorectal cancer, a lung cancer, a breast cancer, a gastric cancer, a pancreatic cancer, a bile duct cancer, and / or an ovarian cancer).

[0032] A cfDNA fragmentation profile can include a cfDNA fragment size pattern. cfDNA fragments can be any appropriate size. For example, cfDNA fragment can be from about 50 base pairs (bp) to about 400 bp in length. As described herein, a mammal having cancer can have a cfDNA fragment size pattern that contains a shorter median cfDNA fragment size than the median cfDNA fragment size in a healthy mammal. A healthy mammal (e.g., a mammal not having cancer) can have cfDNA fragment sizes having a median cfDNA fragment size from about 166.6 bp to about 167.2 bp (e.g., about 166.9 bp). In some embodiments, a mammal having cancer can have cfDNA fragment sizes that are, on average, about 1.28 bp to about 2.49 bp (e.g., about 1.88 bp) shorter than cfDNA fragment sizes in a healthy mammal. For example, a mammal having cancer can have cfDNA fragment sizes having a median cfDNA fragment size of about 164.11 bp to about 165.92 bp (e.g., about 165.02 bp).

[0033] A cfDNA fragmentation profile can include a cfDNA fragment size distribution. As described herein, a mammal having cancer can have a cfDNA size distribution that is more variable than a cfDNA fragment size distribution in a healthy mammal. In some embodiments, a size distribution can be within a targeted region. A healthy mammal (e.g., a mammal not having cancer) can have a targeted region cfDNA fragment size distribution of about 1 or less than about 1. In some embodiments, a mammal having cancer can have a targeted region cfDNA fragment size distribution that is longer (e.g., 10, 15, 20, 25, 30, 35, 40, 45, 50 or more bp longer, or any number of base pairs between these numbers) than a targeted region cfDNA fragment size distribution in a healthy mammal. In some embodiments, a mammal having cancer can have a targeted region cfDNA fragment size distribution that is shorter (e.g., 10, 15, 20, 25, 30, 35, 40, 45, 50 or more bp shorter, or any number of base pairs between these numbers) than a targeted region cfDNA fragment size distribution in a healthy mammal. In some embodiments, a mammal having cancer can have a targeted region cfDNA fragment size distribution that is about 47 bp smaller to about 30 bp longer than a targeted region cfDNA fragment size distribution in a healthy mammal. In some embodiments, a mammal having cancer can have a targeted region cfDNA fragment size distribution of, on average, a 10, 11, 12, 13, 14, 15, 15, 17, 18, 19, 20 or more bp difference in lengths of cfDNA fragments. For example, a mammal having cancer canDOCKET NO.: 348358.18302 have a targeted region cfDNA fragment size distribution of, on average, about a 13 bp difference in lengths of cfDNA fragments. In some embodiments, a size distribution can be a genome-wide size distribution. A healthy mammal (e.g., a mammal not having cancer) can have very similar distributions of short and long cfDNA fragments genome-wide. In some embodiments, a mammal having cancer can have, genome-wide, one or more alterations (e.g., increases and decreases) in cfDNA fragment sizes. The one or more alterations can be any appropriate chromosomal region of the genome. For example, an alteration can be in a portion of a chromosome. Examples of portions of chromosomes that can contain one or more alterations in cfDNA fragment sizes include, without limitation, portions of 2q, 4p, 5p, 6q, 7p, 8q, 9q, 10q, 11q, 12q, and 14q. For example, an alteration can be across a chromosome arm (e.g., an entire chromosome arm).

[0034] A cfDNA fragmentation profile can include a ratio of small cfDNA fragments to large cfDNA fragments and a correlation of fragment ratios to reference fragment ratios. As used herein, with respect to ratios of small cfDNA fragments to large cfDNA fragments, a small cfDNA fragment can be from about 100 bp in length to about 150 bp in length. As used herein, with respect to ratios of small cfDNA fragments to large cfDNA fragments, a large cfDNA fragment can be from about 151 bp in length to 220 bp in length. As described herein, a mammal having cancer can have a correlation of fragment ratios (e.g., a correlation of cfDNA fragment ratios to reference DNA fragment ratios such as DNA fragment ratios from one or more healthy mammals) that is lower (e.g., 2-fold lower, 3-fold lower, 4-fold lower, 5-fold lower, 6-fold lower, 7-fold lower, 8-fold lower, 9-fold lower, 10-fold lower, or more) than in a healthy mammal. A healthy mammal (e.g., a mammal not having cancer) can have a correlation of fragment ratios (e.g., a correlation of cfDNA fragment ratios to reference DNA fragment ratios such as DNA fragment ratios from one or more healthy mammals) of about 1 (e.g., about 0.96). In some embodiments, a mammal having cancer can have a correlation of fragment ratios (e.g., a correlation of cfDNA fragment ratios to reference DNA fragment ratios such as DNA fragment ratios from one or more healthy mammals) that is, on average, about 0.19 to about 0.30 (e.g., about 0.25) lower than a correlation of fragment ratios (e.g., a correlation of cfDNA fragment ratios to reference DNA fragment ratios such as DNA fragment ratios from one or more healthy mammals) in a healthy mammal.DOCKET NO.: 348358.18302

[0035] A cfDNA fragmentation profile can include coverage of all fragments. Coverage of all fragments can include windows (e.g., non-overlapping windows) of coverage. In some embodiments, coverage of all fragments can include windows of small fragments (e.g., fragments from about 100 bp to about 150 bp in length). In some embodiments, coverage of all fragments can include windows of large fragments (e.g., fragments from about 151 bp to about 220 bp in length).

[0036] A cfDNA fragmentation profile can be obtained using any appropriate method. In some embodiments, cfDNA from a mammal (e.g., a mammal having, or suspected of having, cancer) can be processed into sequencing libraries which can be subjected to whole genome sequencing (e.g., low-coverage whole genome sequencing), mapped to the genome, and analyzed to determine cfDNA fragment lengths. Mapped sequences can be analyzed in non-overlapping windows covering the genome. Windows can be any appropriate size. For example, windows can be from thousands to millions of bases in length. As one non-limiting example, a window can be about 5 megabases (Mb) long. Any appropriate number of windows can be mapped. For example, tens to thousands of windows can be mapped in the genome. For example, hundreds to thousands of windows can be mapped in the genome. A cfDNA fragmentation profile can be determined within each window.

[0037] In some embodiments, methods and materials described herein also can include machine learning. For example, machine learning can be used for identifying an altered fragmentation profile (e.g., using coverage of cfDNA fragments, fragment size of cfDNA fragments, coverage of chromosomes, etc).

[0038] Methods of Treatment

[0039] The methods embodied herein, include identifying a mammal as having cancer. The methods include, extracting cell-free DNA (cfDNA) from a subject’s biological sample; generating genomic libraries from the extracted cfDNA; sequencing individual cfDNA molecules ; and administering a cancer treatment to the subject.

[0040] In certain embodiments, a subject is diagnosed as having cancer, e.g. early stage cancer. In certain embodiments, the type of cancer is identified, and the cancer is treated by various therapeutics, including therapeutics specific for the type of cancer. In certain embodiments, the cancer comprises ovarian cancer, or subtypes thereof.DOCKET NO.: 348358.18302

[0041] In one aspect, a method of early detection of cancer and treatment of a subject comprises (i) determining a cell free DNA (cfDNA) fragmentation profile of the subject, the method comprising: extracting and enriching cell free DNA (cfDNA) from a subject’s biological sample; conducting low coverage whole genome sequencing of the subject’s isolated cfDNA to generate genomic libraries of sequenced cfDNA fragments; determining a cfDNA fragmentome profile of the subject’s sequenced DNA fragments comprising evaluating one or more cfDNA fragment characteristics; (ii) comparing the cfDNA fragmentome profiles of the subject to reference cfDNA fragmentome profiles from healthy subjects to detect and diagnose the cancer; (iii) assaying the subject’s biological sample to detect and quantify at least one biomarker; (iv) comparing the subject’s cfDNA fragmentation profile and protein biomarkers with normal reference non-cancer subjects; and, treating the subject diagnosed with cancer with a cancer specific therapy.

[0042] In another aspect, a method of distinguishing between ovarian cancer, and benign cancers in subjects comprises (i) determining a cell free DNA (cfDNA) fragmentation profile of the subject, the method comprising: extracting and enriching cell free DNA (cfDNA) from a subject’s biological sample; conducting low coverage whole genome sequencing of the subject’s isolated cfDNA to generate genomic libraries of sequenced cfDNA fragments; determining a cfDNA fragmentome profile of the subject’s sequenced DNA fragments comprising evaluating one or more cfDNA fragment characteristics; (ii) comparing the cfDNA fragmentome profiles of the subject to reference cfDNA fragmentome profiles from healthy subjects to detect and differentially diagnose between ovarian cancer and ovarian cancer subtypes; (iii) assaying the subject’s biological sample to detect and quantify at least one biomarker; (iv) comparing the subject’s cfDNA fragmentation profile and protein biomarkers with normal reference non-cancer subjects; and, treating the subject diagnosed with ovarian cancer or ovarian cancer subtypes with a cancer specific therapy.

[0043] The cancer treatment can be surgery, adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, hormone therapy, cytotoxic therapy, immunotherapy, adoptive T cell therapy, targeted therapy, or any combinations thereof. The method also can include administering to the mammal a cancer treatment (e.g., surgery, adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, hormone therapy, cytotoxic therapy,DOCKET NO.: 348358.18302 immunotherapy, adoptive T cell therapy, targeted therapy, or any combinations thereof). The mammal can be monitored for the presence of cancer after administration of the cancer treatment.

[0044] Cancer therapies in general also include a variety of combination therapies with both chemical and radiation-based treatments. Combination chemotherapies include, for example, cisplatin (CDDP), carboplatin, procarbazine, mechlorethamine, cyclophosphamide, camptothecin, ifosfamide, melphalan, chlorambucil, busulfan, nitrosurea, dactinomycin, daunorubicin, doxorubicin, bleomycin, plicomycin, mitomycin, etoposide (VP16), tamoxifen, raloxifene, estrogen receptor binding agents, taxol, gemcitabien, navelbine, famesyl-protein transferase inhibitors, transplatinum, 5-fluorouracil, vincristine, vinblastine and methotrexate, Temazolomide (an aqueous form of DTIC), or any analog or derivative variant of the foregoing. The combination of chemotherapy with biological therapy is known as biochemotherapy. The chemotherapy may also be administered at low, continuous doses which is known as metronomic chemotherapy.

[0045] Yet further combination chemotherapies include, for example, alkylating agents such as thiotepa and cyclosphosphamide; alkyl sulfonates such as busulfan, improsulfan and piposulfan; aziridines such as benzodopa, carboquone, meturedopa, and uredopa; ethylenimines and methylamelamines including altretamine, triethylenemelamine, trietylenephosphoramide, triethiylenethiophosphoramide and trimethylolomelamine; acetogenins (especially bullatacin and bullatacinone); a camptothecin (including the synthetic analogue topotecan); bryostatin; callystatin; CC-1065 (including its adozelesin, carzelesin and bizelesin synthetic analogues); cryptophycins (particularly cryptophycin 1 and cryptophycin 8); dolastatin; duocarmycin (including the synthetic analogues, KW-2189 and CB1-TM1); eleutherobin; pancratistatin; a sarcodictyin; spongistatin; nitrogen mustards such as chlorambucil, chlornaphazine, cholophosphamide, estramustine, ifosfamide, mechlorethamine, mechlorethamine oxide hydrochloride, melphalan, novembichin, phenesterine, prednimustine, trofosfamide, uracil mustard; nitrosureas such as carmustine, chlorozotocin, fotemustine, lomustine, nimustine, and ranimnustine; antibiotics such as the enediyne antibiotics (e.g., calicheamicin, especially calicheamicin gammall and calicheamicin omegall; dynemicin, including dynemicin A; bisphosphonates, such as clodronate; an esperamicin; as well as neocarzinostatin chromophoreDOCKET NO.: 348358.18302 and related chromoprotein enediyne antiobiotic chromophores, aclacinomysins, actinomycin, authrarnycin, azaserine, bleomycins, cactinomycin, carabicin, carminomycin, carzinophilin, chromomycinis, dactinomycin, daunorubicin, detorubicin, 6-diazo-5-oxo-L-norleucine, doxorubicin (including morpholino-doxorubicin, cyanomorpholino-doxorubicin, 2-pyrrolino- doxorubicin and deoxydoxorubicin), epirubicin, esorubicin, idarubicin, marcellomycin, mitomycins such as mitomycin C, mycophenolic acid, nogalarnycin, olivomycins, peplomycin, potfiromycin, puromycin, quelamycin, rodornbicin, streptonigrin, streptozocin, tubercidin, ubenimex, zinostatin, zombicin; anti-metabolites such as methotrexate and 5-fluorouracil (5- FU); folic acid analogues such as denopterin, pteropterin, trimetrexate; purine analogs such as fludarabine, 6-mercaptopurine, thiamiprine, thioguanine; pyrimidine analogs such as ancitabine, azacitidine, 6-azauridine, carmofur, cytarabine, dideoxyuridine, doxifluridine, enocitabine, floxuridine; androgens such as calusterone, dromostanolone propionate, epitiostanol, mepitiostane, testolactone; anti-adrenals such as rnitotane, trilostane; folic acid replenisher such as frolinic acid; aceglatone; aldophosphamide glycoside; arninolevulinic acid; eniluracil; arnsacrine; bestrabucil; bisantrene; edatraxate; defofarnine; demecolcine; diaziquone; elformithine; elliptinium acetate; an epothilone; etoglucid; gallium nitrate; hydroxyurea; lentinan; lonidainine; maytansinoids such as maytansine and ansamitocins; mitoguazone; mitoxantrone; mopidanmol; nitraerine; pentostatin; phenamet; pirarubicin; losoxantrone; podophyllinic acid; 2-ethylhydrazide; procarbazine; PSK polysaccharide complex; razoxane; rhizoxin; sizofiran; spirogermanium; tenuazonic acid; triaziquone; 2,2’,2”-trichlorotriethylamine; trichothecenes (especially T-2 toxin, verracurin A, roridin A and anguidine); urethan; vindesine; dacarbazine; mannomustine; mitobronitol; mitolactol; pipobroman; gacytosine; arabinoside (“Ara-C”); cyclophosphamide; taxoids, e.g., paclitaxel and docetaxel gemcitabine; 6- thioguanine; mercaptopurine; platinum coordination complexes such as cisplatin, oxaliplatin and carboplatin; vinblastine; platinum; etoposide (VP-16); ifosfamide; mitoxantrone; vincristine; vinorelbine; novantrone; teniposide; edatrexate; daunomycin; aminopterin; xeloda; ibandronate; irinotecan (e.g., CPT-11); topoisomerase inhibitor RPS 2000; difluorometlhylornithine (DMFO); retinoids such as retinoic acid; capecitabine; carboplatin, procarbazine, plicomycin, gemcitabien, navelbine, farnesyl-protein transferase inhibitors, transplatinum; and pharmaceutically acceptable salts, acids or derivatives of any of the above.DOCKET NO.: 348358.18302

[0046] Immunotherapeutics, generally, rely on the use of immune effector cells and molecules to target and destroy cancer cells. The immune effector may be, for example, an antibody specific for some marker on the surface of a tumor cell. The antibody alone may serve as an effector of therapy, or it may recruit other cells to actually effect cell killing. The antibody also may be conjugated to a drug or toxin (chemotherapeutic, radionuclide, ricin A chain, cholera toxin, pertussis toxin, etc.) and serve merely as a targeting agent. Alternatively, the effector may be a lymphocyte carrying a surface molecule that interacts, either directly or indirectly, with a tumor cell target. Various effector cells include cytotoxic T cells and NK cells as well as genetically engineered variants of these cell types modified to express chimeric antigen receptors.

[0047] The immunotherapy may comprise suppression of T regulatory cells (Tregs), myeloid derived suppressor cells (MDSCs) and cancer associated fibroblasts (CAFs). In some embodiments, the immunotherapy is a tumor vaccine (e.g., whole tumor cell vaccines, peptides, and recombinant tumor associated antigen vaccines), or adoptive cellular therapies (ACT) (e.g., T cells, natural killer cells, TILs, and LAK cells). The T cells may be engineered with chimeric antigen receptors (CARs) or T cell receptors (TCRs) to specific tumor antigens. As used herein, a chimeric antigen receptor (or CAR) may refer to any engineered receptor specific for an antigen of interest that, when expressed in a T cell, confers the specificity of the CAR onto the T cell. Once created using standard molecular techniques, a T cell expressing a chimeric antigen receptor may be introduced into a patient, as with a technique such as adoptive cell transfer. In some aspects, the T cells are activated CD4 and / or CD8 T cells in the individual which are characterized by γ-1FN- producing CD4 and / or CD8 T cells and / or enhanced cytolytic activity relative to prior to the administration of the combination. The CD4 and / or CD8 T cells may exhibit increased release of cytokines selected from the group consisting of IFN-γ, TNF-a and interleukins. The CD4 and / or CD8 T cells can be effector memory T cells. In certain embodiments, the CD4 and / or CDS effector memory T cells are characterized by having the expression of CD44highCD62Llow.

[0048] The immunotherapy may be a cancer vaccine comprising one or more cancer antigens, in particular a protein or an immunogenic fragment thereof, DNA or RNA encoding said cancer antigen, in particular a protein or an immunogenic fragment thereof, cancer cellDOCKET NO.: 348358.18302 lysates, and / or protein preparations from tumor cells. As used herein, a cancer antigen is an antigenic substance present in cancer cells. In principle, any protein produced in a cancer cell that has an abnormal structure due to mutation can act as a cancer antigen. In principle, cancer antigens can be products of mutated Oncogenes and tumor suppressor genes, products of other mutated genes, overexpressed or aberrantly expressed cellular proteins, cancer antigens produced by oncogenic viruses, oncofetal antigens, altered cell surface glycolipids and glycoproteins, or cell type-specific differentiation antigens. Examples of cancer antigens include the abnormal products of ras and p53 genes. Other examples include tissue differentiation antigens, mutant protein antigens, oncogenic viral antigens, cancer-testis antigens and vascular or stromal specific antigens. Tissue differentiation antigens are those that are specific to a certain type of tissue. Mutant protein antigens are likely to be much more specific to cancer cells because normal cells shouldn’t contain these proteins. Normal cells will display the normal protein antigen on their MHC molecules, whereas cancer cells will display the mutant version. Some viral proteins are implicated in forming cancer, and some viral antigens are also cancer antigens. Cancer-testis antigens are antigens expressed primarily in the germ cells of the testes, but also in fetal ovaries and the trophoblast. Some cancer cells aberrantly express these proteins and therefore present these antigens, allowing attack by T-cells specific to these antigens. Exemplary antigens of this type are CTAG1 B and MAGEA1 as well as Rindopepimut, a 14-mer intradermal injectable peptide vaccine targeted against epidermal growth factor receptor vlll (EGFRvlll; deletion of exons 2–7) variant. Rindopepimut is particularly suitable for treating glioblastoma when used in combination with an inhibitor of the CD95 / CD95L signaling system as described herein. Also, proteins that are normally produced in very low quantities, but whose production is dramatically increased in cancer cells, may trigger an immune response. An example of such a protein is the enzyme tyrosinase, which is required for melanin production. Normally tyrosinase is produced in minute quantities but its levels are very much elevated in melanoma cells. Oncofetal antigens are another important class of cancer antigens. Examples are alphafetoprotein (AFP) and carcinoembryonic antigen (CEA). These proteins are normally produced in the early stages of embryonic development and disappear by the time the immune system is fully developed. Thus, self-tolerance does not develop against these antigens. Abnormal proteins are also produced by cells infected with oncoviruses, e.g. EBV and HPV. Cells infected by these viruses contain latent viral DNA which is transcribed, and the resulting protein produces an immune response. ADOCKET NO.: 348358.18302 cancer vaccine may include a peptide cancer vaccine, which in some embodiments is a personalized peptide vaccine. In some embodiments. the peptide cancer vaccine is a multivalent long peptide vaccine, a multi-peptide vaccine, a peptide cocktail vaccine, a hybrid peptide vaccine, or a peptide-pulsed dendritic cell vaccines.

[0049] The immunotherapy may be an antibody, such as part of a polyclonal antibody preparation, or may be a monoclonal antibody. The antibody may be a humanized antibody, a chimeric antibody, an antibody fragment, a bispecific antibody or a single chain antibody. An antibody as disclosed herein includes an antibody fragment, such as, but not limited to, Fab, Fab’ and F(ab’)2, Fd, single-chain Fvs (scFv), single-chain antibodies, disulfide-linked Fvs (sdfv) and fragments including either a VL or VH domain. In some aspects, the antibody or fragment thereof specifically binds epidermal growth factor receptor (EGFR1, Erb-B1), HER2 / neu (Erb- B2), CD20, Vascular endothelial growth factor (VEGF), insulin-like growth factor receptor (IGF- 1R), TRAIL-receptor, epithelial cell adhesion molecule, carcinoembryonic antigen, Prostate- specific membrane antigen, Mucin-1, CD30, CD33, or CD40.

[0050] Examples of monoclonal antibodies include, without limitation, trastuzumab (anti-HER2 / neu antibody); Pertuzumab (anti-HER2 mAb); cetuximab (chimeric monoclonal antibody to epidermal growth factor receptor EGFR); panitumumab (anti-EGFR antibody); nimotuzumab (anti-EGFR antibody); Zalutumumab (anti-EGFR mAb); Necitumumab (anti- EGFR mAb); MDX-210 (humanized anti-HER-2 bispecific antibody); MDX-210 (humanized anti-HER-2 bispecific antibody); MDX-447 (humanized anti-EGF receptor bispecific antibody); Rituximab (chimeric murine / human anti-CD20 mAb); Obinutuzumab (anti-CD20 mAb); Ofatumumab (anti-CD20 mAb); Tositumumab-I131 (anti-CD20 mAb); lbritumomab tiuxetan (anti-CD20 mAb); Bevacizumab (anti-VEGF mAb); Ramucirumab (anti-VEGFR2 mAb); Ranibizumab (anti-VEGF mAb); Aflibercept (extracellular domains of VEGFR1 and VEGFR2 fused to IgG1 Fc); AMG386 (angiopoietin-1 and -2 binding peptide fused to IgG1 Fc); Dalotuzumab (anti-IGF-1R mAb); Gemtuzumab ozogamicin (anti-CD33 mAb); Alemtuzumab (anti-Campath-1 / CD52 mAb); Brentuximab vedotin (anti-CD30 mAb); Catumaxomab (bispecific mAb that targets epithelial cell adhesion molecule and CD3); Naptumomab (anti-5T4 mAb); Girentuximab (anti-Carbonic anhydrase ix); or Farletuzumab (anti-folate receptor). Other examples include antibodies such as Panorex™ (17-lA) (murine monoclonal antibody); PanorexDOCKET NO.: 348358.18302 (MAb17-lA) (chimeric murine monoclonal antibody); BEC2 (ami-idiotypic mAb, mimics the GD epitope) (with BCG); Oncolym (Lym-1 monoclonal antibody); SMART M195 Ab, humanized 13’ 1 LYM-1 (Oncolym), Ovarex (B43.13, anti-idiotypic mouse mAb); 3622W94 mAb that binds to EGP40 (17-1A) pancarcinoma antigen on adenocarcinomas; Zenapax (SMART Anti-Tac (IL-2 receptor); SMART M195 Ab, humanized Ab, humanized); NovoMAb- G2 (pancarcinoma specific Ab); TNT (chimeric mAb to histone antigens); TNT (chimeric mAb to histone antigens); Gliomab-H (Monoclonals-Humanized Abs); GNI-250 Mab; EMD-72000 (chimeric-EGF antagonist); LymphoCide (humanized IL.L.2 antibody); and MDX-260 bispecific, targets GD-2, ANA Ab, SMART IDIO Ab, SMART ABL 364 Ab or ImmuRAIT- CEA. Further examples of antibodies include Zanulimumab (anti-CD4 mAb), Keliximab (anti- CD4 mAb); Ipilimumab (MDX-101; anti-CTLA-4 mAb); Tremilimumab (anti-CTLA-4 mAb); (Daclizumab (anti-CD25 / IL-2R mAb); Basiliximab (anti-CD25 / IL-2R mAb); MDX-1106 (anti-PDl mAb); antibody to GITR; GC1008 (anti-TGF-β antibody); metelimumab / CAT-192 (anti- TGF-β antibody); lerdelimumab / CAT-152 (anti-TGF-β antibody); ID11 (anti-TGF-β antibody); Denosumab (anti-RANKL mAb); BMS-663513 (humanized anti-4-1BB mAb); SGN- 40 (humanized anti-CD40 mAb); CP870,893 (human anti-CD40 mAb); Infliximab (chimeric anti-TNF mAb; Adalimumab (human anti-TNF mAb); Certolizumab (humanized Fab anti-TNF); Golimumab (anti-TNF); Etanercept (Extracellular domain of TNFR fused to IgG1 Fc); Belatacept (Extracellular domain of CTLA-4 fused to Fe); Abatacept (Extracellular domain of CTLA-4 fused to Fe); Belimumab (anti-B Lymphocyte stimulator); Muromonab-CD3 (anti-CD3 mAb); Otelixizumab (anti-CD3 mAb); Teplizumab (anti-CD3 mAb); Tocilizumab (anti-IL6R mAb); REGN88 (anti-IL6R mAb); Ustekinumab (anti-IL-12 / 23 mAb); Briakinumab (anti-IL- 12 / 23 mAb); Natalizumab (anti-α4 integrin); Vedolizumab (anti-α4 β7 integrin mAb); T1 h (anti-CD6 mAb); Epratuzumab (anti-CD22 mAb); Efalizumab (anti-CD11a mAb); and Atacicept (extracellular domain of transmembrane activator and calcium-modulating ligand interactor fused with Fc).

[0051] Systems

[0052] In some examples, the present disclosure provides systems, methods, or kits that can include data analysis realized in measurement devices (e.g., laboratory instruments, such as a sequencing machine), software code that executes on computing hardware. The software can beDOCKET NO.: 348358.18302 stored in memory and execute on one or more hardware processors. The software can be organized into routines or packages that can communicate with each other. A module can comprise one or more devices / computers, and potentially one or more software routines / packages that execute on the one or more devices / computers. For example, an analysis application or system can include at least a data receiving module, a data pre-processing module, a data analysis module (which can operate on one or more types of genomic data), a data interpretation module, or a data visualization module.

[0053] The data receiving module can connect laboratory hardware or instrumentation with computer systems that process laboratory data. The data pre-processing module can perform operations on the data in preparation for analysis. Examples of operations that can be applied to the data in the pre-processing module include affine transformations, denoising operations, data cleaning, reformatting, or subsampling. The data analysis module, which can be specialized for analyzing genomic data from one or more genomic materials, can, for example, take assembled genomic sequences and perform probabilistic and statistical analysis to identify abnormal patterns related to a disease, pathology, state, risk, condition, or phenotype. The data interpretation module can use analysis methods, for example, drawn from statistics, mathematics, or biology, to support understanding of the relation between the identified abnormal patterns and health conditions, functional states, prognoses, or risks. The data analysis module and / or the data interpretation module can include one or more machine learning models, which can be implemented in hardware, e.g., which executes software that embodies a machine learning model. The data visualization module can use methods of mathematical modeling, computer graphics, or rendering to create visual representations of data that can facilitate the understanding or interpretation of results. The present disclosure provides computer systems that are programmed to implement methods of the disclosure.

[0054] In some embodiments, the methods disclosed herein can include computational analysis on nucleic acid sequencing data of samples from an individual or from a plurality of individuals. An analysis can identify a variant inferred from sequence data to identify sequence variants based on probabilistic modeling, statistical modeling, mechanistic modeling, network modeling, or statistical inferences. Non-limiting examples of analysis methods include principal component analysis, autoencoders, singular value decomposition, Fourier bases, wavelets,DOCKET NO.: 348358.18302 discriminant analysis, regression, support vector machines, tree-based methods, networks, matrix factorization, and clustering. Non-limiting examples of variants include a germline variation or a somatic mutation. In some examples, a variant can refer to an already-known variant. The already- known variant can be scientifically confirmed or reported in literature. In some examples, a variant can refer to a putative variant associated with a biological change. A biological change can be known or unknown. In some examples, a putative variant can be reported in literature, but not yet biologically confirmed. Alternatively, a putative variant is never reported in literature, but can be inferred based on a computational analysis disclosed herein. In some examples, germline variants can refer to nucleic acids that induce natural or normal variations.

[0055] In certain embodiments, the computer system includes a central processing unit (CPU, also “processor” and “computer processor” herein), which can be a single core or multi core processor, or a plurality of processors for parallel processing; memory (e.g., cache, random-access memory, read-only memory, flash memory, or other memory); electronic storage unit (e.g., hard disk), communication interface (e.g., network adapter) for communicating with one or more other systems; and peripheral devices, such as adapters for cache, other memory, data storage and / or electronic display. The memory, storage unit, interface and peripheral devices may be in communication with the CPU through a communication bus (solid lines), such as a motherboard. The storage unit can be a data storage unit (or data repository) for storing data. One or more analyte feature inputs can be entered from the one or more measurement devices. Example analytes and measurement devices are described herein.

[0056] The computer system can be operatively coupled to a computer network (“network”) with the aid of the communication interface. The network can be the Internet, an internet and / or extranet, or an intranet and / or extranet that is in communication with the Internet. The network in some cases is a telecommunication and / or data network. The network can include one or more computer servers, which can enable distributed computing, such as cloud computing over the network (“the cloud”) to perform various aspects of analysis, calculation, and generation of the present disclosure, such as, for example, activation of a valve or pump to transfer a reagent or sample from one chamber to another or application of heat to a sample (e.g., during an amplification reaction), other aspects of processing and / or assaying a sample, performing sequencing analysis, measuring sets of values representative of classes of molecules, identifyingDOCKET NO.: 348358.18302 sets of features and feature vectors from assay data, processing feature vectors using a machine learning model to obtain output classifications, and training a machine learning model (e.g., iteratively searching for optimal values of parameters of the machine learning model). Such cloud computing may be provided by cloud computing platforms such as, for example, Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform, and IBM cloud. The network, in some cases with the aid of the computer system, can implement a peer-to-peer network, which may enable devices coupled to the computer system to behave as a client or a server.

[0057] The CPU can execute a sequence of machine-readable instructions, which can be embodied in a program or software. The instructions can be stored in a memory location, such as the memory. The instructions can be directed to the CPU, which can subsequently program or otherwise configure the CPU to implement methods of the present disclosure. The CPU can be part of a circuit, such as an integrated circuit. One or more other components of the system can be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).

[0058] The storage unit can store files, such as drivers, libraries and saved programs. The storage unit can store user data, e.g., user preferences and user programs. The computer system in some cases can include one or more additional data storage units that are external to the computer system, such as located on a remote server that is in communication with the computer system through an intranet or the Internet.

[0059] The computer system can communicate with one or more remote computer systems through the network. For instance, the computer system can communicate with a remote computer system of a user. Examples of remote computer systems include personal computers (e.g., portable PC), slate or tablet PC's (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, Smart phones (e.g., Apple® iPhone, Android-enabled device, Blackberry®), or personal digital assistants. The user can access the computer system via the network.

[0060] Methods as described herein can be implemented by way of machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system such as, for example, on the memory or electronic storage unit. The machine executable or machine-readable code can be provided in the form of software. During use, the code can beDOCKET NO.: 348358.18302 executed by the CPU. In some cases, the code can be retrieved from the storage unit and stored on the memory for ready access by the CPU. In some situations, the electronic storage unit can be precluded, and machine-executable instructions are stored on memory.

[0061] The code can be pre-compiled and configured for use with a machine having a processer adapted to execute the code or can be compiled during runtime. The code can be supplied in a programming language that can be selected to enable the code to execute in a pre-compiled or as compiled fashion.

[0062] Aspects of the systems and methods provided herein, such as the computer system , can be embodied in programming. Various aspects of the technology can be thought of as “products” or “articles of manufacture” typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine readable medium. Machine-executable code can be stored on an electronic storage unit, such as memory (e.g., read- only memory, random-access memory, flash memory) or a hard disk. “Storage” type media can include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer into the computer platform of an application server. Thus, another type of media that can bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links or the like, also can be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.

[0063] Hence, a machine readable medium, such as computer-executable code, may take many forms, including but not limited to, a tangible storage medium, a carrier wave medium, or physical transmission medium. Non-volatile storage media include, for example, optical orDOCKET NO.: 348358.18302 magnetic disks, such as any of the storage devices in any computer(s) or the like, such as can be used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system.

[0064] Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include for example: a floppy disk, a flexible disk, hard disk, magnetic tape, any other magnetic medium, a CD-ROM, DVD or DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium with patterns of holes, a RAM, a ROM, a PROM and EPROM, a FLASH- EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media can be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0065] The computer system can include or be in communication with an electronic display that comprises a user interface (UI) for providing, for example, a current stage of processing or assaying of a sample (e.g., a particular step, such as a lysis step, or sequencing step that is being performed). Inputs are received by the computer system from one or more measurement. Examples of UIs include, without limitation, a graphical user interface (GUI) and web-based user interface. The algorithm can, for example, process and / or assay a sample, perform sequencing analysis, measure sets of values representative of classes of molecules, identify sets of features and feature vectors from assay data, process feature vectors using a machine learning model to obtain output classifications, and train a machine learning model (e.g., iteratively search for optimal values of parameters of the machine learning model).

[0066] In some embodiments, systems capable of executing one or more algorithms, e.g., laptops, desktops, iPads, mobile devices etc., for determining changes in cfDNA mutation profiles, frequency of mutations and / or fragmentation profiles classifies the subject as a cancer patient based on the cfDNA mutation profiles, frequency of mutations and / or fragmentation for theDOCKET NO.: 348358.18302 subject. These systems further execute machine learning algorithms that can be used to generate models such as, for example, high-risk populations and low-risk general populations (a penalized logistic regression with the Mathios et al. (Mathios D, Johansen JS, Cristiano S, Medina JE, Phallen J, Larsen KR, et al. Detection and characterization of lung cancer using cell-free DNA fragmentomes. Nat Commun 2021;12(1):5060) features as well as coverage from transcription factor binding sites. These models can be trained on the subject cohort with 5-fold cross validation with 10 repeats, and scores for each sample ae calculated by the mean across repeats and evaluated using AUC-ROC. For example, the first model used the high-risk non-cancer and HCC patients while the second used the non-cancer individuals without liver pathology. The locked high-risk model trained on the cohort was applied to a second and different cohort to generate cancer predictions on an external validation set. A “class label” can be applied to each sample indicating the classification of the sample for any number of input features. For example, the class labels for the set of cohorts could indicate the identity of cfDNA mutation profiles, frequency of mutations and / or fragmentation profiles based on genomic location etc. The resulting training sets are provided to machine learning unit, such as a neural network or a support vector machine. Using the training set, the machine learning unit may generate a model to classify the sample according to the cfDNA mutation profiles, frequency of mutations and / or fragmentation profile.

[0067] In some embodiments, a method is provided for creating a trained classifier, comprising the steps of: (a) providing a plurality of different classes, wherein each class represents a set of subjects with a shared characteristic (e.g. from one or more cohorts); (b) providing a multi- parametric model representative of the cell-free DNA molecules from each of a plurality of samples belonging to each of the classes, thereby providing a training data set; and (c) training a learning algorithm on the training data set to create one or more trained classifiers, wherein each trained classifier classifies a test sample into one or more of the plurality of classes.

[0068] As an example, a trained classifier may use a learning algorithm selected from the group consisting of: a random forest, a neural network, a support vector machine, and a linear classifier. Each of the plurality of different classes may be selected from the group consisting of healthy, breast cancer, colon cancer, lung cancer, pancreatic cancer, prostate cancer, ovarian cancer, melanoma, and liver cancer. In certain embodiments, the cancer comprises ovarian cancer, ovarianDOCKET NO.: 348358.18302 cancer subtypes, including high-grade serous (HGSOC), low-grade serous (LGSOC), clear cell, mucinous, and endometroid ovarian cancers.

[0069] A trained classifier may be applied to a method of classifying a sample from a subject. This method of classifying may comprise: (a) providing a multi-parametric model representative of the cell-free DNA molecules from a test sample from the subject; and (b) classifying the test sample using a trained classifier. After the test sample is classified into one or more classes, a therapeutic intervention on the subject can be performed based on the classification of the sample.

[0070] In some embodiments, training sets are provided to a machine learning unit, such as a neural network or a support vector machine. Using the training set, the machine learning unit may generate a model to classify the sample according to a treatment response to one or more therapeutic inventions. This is also referred to as “calling”. The model developed may employ information from any part of a test vector.

[0071] In general, machine learning can be used to reduce a set of data generated from all (primary sample / analytes / test) combinations into an optimal predictive set of features, e.g., which satisfy specified criteria. In various examples statistical learning, and / or regression analysis can be applied. Simple to complex and small to large models making a variety of modeling assumptions can be applied to the data in a cross-validation paradigm. Simple to complex includes considerations of linearity to non-linearity and non-hierarchical to hierarchical representations of the features. Small to large models includes considerations of the size of basis vector space to project the data onto as well as the number of interactions between features that are included in the modelling process.

[0072] Machine learning techniques can be used to assess the commercial testing modalities most optimal for cost / performance / commercial reach as defined in the initial question. A threshold check can be performed: If the method applied to a hold-out dataset that was not used in cross validation surpasses the initialized constraints, then the assay is locked, and production initiated. For example, a threshold for assay performance may include a desired minimum accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), or a combination thereof. For example, a desiredDOCKET NO.: 348358.18302 minimum accuracy, PPV, NPV, clinical sensitivity, clinical specificity, or combination thereof may be at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. As another example, a desired minimum AUC may be at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.81, at least about 0.82, at least about 0.83, at least about 0.84, at least about 0.85, at least about 0.86, at least about 0.87, at least about 0.88, at least about 0.89, at least about 0.90, at least about 0.91, at least about 0.92, at least about 0.93, at least about 0.94, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99. A subset of assays may be selected from a set of assays to be performed on a given sample based on the total cost of performing the subset of assays, subject to the threshold for assay performance, such as desired minimum accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), and a combination thereof. If the thresholds are not met, then the assay engineering procedure can loop back to either the constraint setting for possible relaxation or to the wet lab to change the parameters in which data was acquired. Given the clinical question, biological constraints, budget, lab machines, etc., can constrain the problem.

[0073] In certain embodiments, the computer processing of a machine learning technique can include method(s) of statistics, mathematics, biology, or any combination thereof. In various examples, any one of the computer processing methods can include a dimension reduction method, logistic regression, dimension reduction, principal component analysis, autoencoders, singular value decomposition, Fourier bases, singular value decomposition, wavelets, discriminant analysis, support vector machine, tree-based methods, random forest, gradient boost tree, logistic regression, matrix factorization, network clustering, statistical testing and neural network.

[0074] In certain embodiments, the computer processing of a machine learning technique can include logistic regression, multiple linear regression (MLR), dimension reduction, partial least squares (PLS) regression, principal component regression, autoencoders, variationalDOCKET NO.: 348358.18302 autoencoders, singular value decomposition, Fourier bases, wavelets, discriminant analysis, support vector machine, decision tree, classification and regression trees (CART), tree-based methods, random forest, gradient boost tree, logistic regression, matrix factorization, multidimensional scaling (MDS), dimensionality reduction methods, t-distributed stochastic neighbor embedding (t-SNE), multilayer perceptron (MLP), network clustering, neuro-fuzzy, neural networks (shallow and deep), artificial neural networks, Pearson product-moment correlation coefficient, Spearman's rank correlation coefficient, Kendall tau rank correlation coefficient, or any combination thereof. In some examples, the computer processing method is a supervised machine learning method including, for example, a regression, support vector machine, tree-based method, and neural network. In some examples, the computer processing method is an unsupervised machine learning method including, for example, clustering, network, principal component analysis, and matrix factorization.

[0075] For supervised learning, training samples (e.g., in thousands) can include measured data (e.g., of various analytes) and known labels, which may be determined via other time- consuming processes, such as imaging of the subject and analysis by a trained practitioner. Example labels can include classification of a subject, e.g., discrete classification of whether a subject has cancer or not or continuous classifications providing a probability (e.g., a risk or a score) of a discrete value. A learning module can optimize parameters of a model such that a quality metric (e.g., accuracy of prediction to known label) is achieved with one or more specified criteria. Determining a quality metric can be implemented for any arbitrary function including the set of all risk, loss, utility, and decision functions. A gradient can be used in conjunction with a learning step (e.g., a measure of how much the parameters of the model should be updated for a given time step of the optimization process).

[0076] As described above, examples can be used for a variety of purposes. For example, plasma (or other sample) can be collected from subjects symptomatic with a condition (e.g., known to have the condition) and healthy subjects. Genetic data (e.g., cfDNA) can be acquired analyzed to obtain a variety of different features, which can include features based on a genome wide analysis. These features can form a feature space that is searched, stretched, rotated, translated, and linearly or non-linearly transformed to generate an accurate machine learning model, which can differentiate between healthy subjects and subjects with the condition (e.g., identify a diseaseDOCKET NO.: 348358.18302 or non-disease status of a subject). Output derived from this data and model (which may include probabilities of the condition, stages (levels) of the condition, or other values), can be used to generate another model that can be used to recommend further procedures, e.g., recommend a biopsy or keep monitoring the subject condition.

[0077] In some embodiments, DNA from a population of several individuals can be analyzed by a set of multiplexed arrays. The data for each multiplexed array may be self- normalized using the information contained in that specific array. This normalization algorithm may adjust for nominal intensity variations observed in the two-color channels, background differences between the channels, and possible crosstalk between the dyes. The behavior of each base position may then be modeled using a clustering algorithm that incorporates several biological heuristics on mutation profiles, frequency of mutations and / or fragmentation profiles. In cases where few cfDNA fragments are observed (e.g., due to low minor-allele frequency), locations and shapes of the missing sequences may be estimated using neural networks. Depending on the profiles and percent sequence identity, a statistical score may be devised (a Training score). A score such as GenCall Score is designed to mimic evaluations made by a human expert's visual and cognitive systems. In addition, it has been evolved using the genotyping data from top and bottom strands. This score may be combined with several penalty terms (e.g., low intensity, mismatch between existing and predicted cfDNA fragments) in order to make up the Training score. The Training score is saved for use by the calling algorithm.

[0078] To call a therapeutic response, a calling algorithm may take the genetic information and treatment responses of a plurality of individuals having a disease or condition. The data may first be normalized (using the same procedure as for the clustering algorithm). The calling operation (classification) may be performed using, for example, a Bayesian model. The score for each call's Call Score can be the product of a Training Score and a data-to-model fit score. After scoring all the treatment responses, the application may compute a composite score.

[0079] In some embodiments, a training dataset comprises clinical data selected from the group consisting of cancer stage, type of surgical procedure, age, tumor grading, depth of tumor infiltration, occurrence of post-operative complications, and the presence of venous invasion. In some embodiments, the training dataset is pre-processed, comprising transforming the provided data into class-conditional probabilities.DOCKET NO.: 348358.18302

[0080] Another embodiment uses machine learning techniques to train a statistical classifier, specifically a support vector machine, for each cancer stage category based on word occurrences in a corpus of histology reports for each patient. New reports can then be classified according to the most likely stage, facilitating the collection and analysis of population staging data.

[0081] In some embodiments, a machine learning algorithm is selected from the group consisting of: a supervised or unsupervised learning algorithm selected from support vector machine, random forest, nearest neighbor analysis, linear regression, binary decision tree, discriminant analyses, logistic classifier, and cluster analysis.

[0082] In general, a system can comprise a report generator for reporting on cancer test results and treatment options. The report generator system can be a central data processing system configured to establish communications directly with a remote data site or laboratory, a medical practice / healthcare provider (treating professional) and / or a patient / subject through communication links. The laboratory can be medical laboratory, diagnostic laboratory, medical facility, medical practice, point-of-care testing device, or any other remote data site capable of generating subject clinical information. Subject clinical information includes but it is not limited to laboratory test data, X-ray data, examination and diagnosis. The healthcare provider or practice includes medical services providers, such as doctors, nurses, home health aides, technicians and physician's assistants, and the practice is any medical care facility staffed with healthcare providers. In certain instances, the healthcare provider / practice is also a remote data site. In a cancer treatment embodiment, the subject may be afflicted with cancer, among others.

[0083] Other clinical information for a cancer subject includes the results of laboratory tests, imaging or medical procedure directed towards the specific cancer that one of ordinary skill in the art can readily identify. The list of appropriate sources of clinical information for cancer includes but it is not limited to: CT scan, MRI scan, ultrasound scan, bone scan, PET Scan, bone marrow test, barium X-ray, endoscopy, lymphangiogram, IVU (Intravenous urogram) or IVP (IV pyelogram), lumbar puncture, cystoscopy, immunological tests (anti-malignin antibody screen), and cancer marker tests.DOCKET NO.: 348358.18302

[0084] The subject clinical information may be obtained from the laboratory manually or automatically. For simplicity of the system the information is obtained automatically at predetermined or regular time intervals. A regular time interval refers to a time interval at which the collection of the laboratory data is carried out automatically by the methods and systems described herein based on a measurement of time such as hours, days, weeks, months, years etc. In one embodiment of the invention, the collection of data and processing is carried out at least once a day. In one embodiment, the transfer and collection of data is carried out once every month, biweekly, or once a week, or once every couple of days. Alternatively, the retrieval of information may be carried out at predetermined but not regular time intervals. For instance, a first retrieval step may occur after one week and a second retrieval step may occur after one month. The transfer and collection of data can be customized according to the nature of the disorder that is being managed and the frequency of required testing and medical examinations of the subjects.

[0085] In certain embodiments, a genetic report is generated from a subject’s sample, e.g. cfDNA. The polynucleotides in a sample can be sequenced, e.g., whole genome sequencing, NGS sequencing, producing a plurality of sequence reads. In some embodiments, genetic information comprises variables defining the genomic organization of cancer cells or the genomic organization of single disseminated cancer cells. In some embodiments, the genetic information comprises sequence or abundance data from one or more genetic loci in cell-free DNA from the individuals.

[0086] Genetic variants can also be identified. Genetic variants include sequence variants, copy number variants and nucleotide modification variants. A sequence variant is a variation in a genetic nucleotide sequence. A copy number variant is a deviation from wild type in the number of copies of a portion of a genome. Genetic variants include, for example, single nucleotide variations (SNPs), insertions, deletions, inversions, transversions, translocations, gene fusions, chromosome fusions, gene truncations, copy number variations (e.g., aneuploidy, partial aneuploidy, polyploidy, gene amplification), abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns and abnormal changes in nucleic acid methylation. The process then determines the frequency of genetic variants in the sample containing the genetic material. Since this process is noisy, the process separates information from noise (73). The sensitivity of detecting genetic variants can be increased by increasing read depthDOCKET NO.: 348358.18302 of polynucleotides (e.g., by sequencing to a greater read depth at in a sample from a subject at two or more time points).

[0087] To increase the diagnosis confidence, a plurality of measurements can be taken. Or alternatively using measurements at a plurality of time points (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more time points) to determine whether cancer is advancing, in remission or stabilized. The diagnostic confidence can be used to identify disease states. For example, cell free polynucleotides taken from a subject can include polynucleotides derived from normal cells, as well as polynucleotides derived from diseased cells, such as cancer cells. Polynucleotides from cancer cells may bear genetic variants, such as somatic cell mutations and copy number variants. When cell free polynucleotides from a sample from a subject are sequenced, and cfDNA mutation profiles, frequency of mutations and / or fragmentation profiles can be produced as described in the examples section which follows.

[0088] Numerous cancers may be detected using the methods and systems described herein. Cancers cells, as most cells, can be characterized by a rate of turnover, in which old cells die and replaced by newer cells. Generally dead cells, in contact with vasculature in a given subject, may release DNA or fragments of DNA into the blood stream. This is also true of cancer cells during various stages of the disease. Cancer cells may also be characterized, dependent on the stage of the disease, by various genetic aberrations such as copy number variation as well as mutations. This phenomenon may be used to detect the presence or absence of cancers individuals using the methods and systems described herein.

[0089] In the early detection of cancers, any of the systems or methods herein described, including mutation detection or copy number variation detection may be utilized to detect cancers. These system and methods may be used to detect any number of genetic aberrations that may cause or result from cancers. These may include but are not limited to cfDNA mutation profiles, frequency of mutations, cfDNA fragmentation profiles, mutations, mutations, indels, copy number variations, transversions, translocations, inversion, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structure alterations, gene fusions, chromosome fusions, gene truncations, gene amplification, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, abnormal changes in nucleic acid methylation infection and cancer.DOCKET NO.: 348358.18302

[0090] Additionally, the systems and methods described herein may also be used to help characterize certain cancers. Genetic data produced from the system and methods of this disclosure may allow practitioners to help better characterize a specific form of cancer. Often times, cancers are heterogeneous in both composition and staging. Genetic profile data may allow characterization of specific sub-types of cancer that may be important in the diagnosis or treatment of that specific sub-type. This information may also provide a subject or practitioner clues regarding the prognosis of a specific type of cancer.

[0091] The systems and methods provided herein may be used to monitor already known cancers, or other diseases in a particular subject. This may allow either a subject or practitioner to adapt treatment options in accord with the progress of the disease. In this example, the systems and methods described herein may be used to construct genetic cfDNA mutation profiles, frequency of mutations and / or fragmentation profiles of a particular subject of the course of the disease. In some instances, cancers can progress, becoming more aggressive and genetically unstable. In other examples, cancers may remain benign, inactive or dormant. The system and methods of this disclosure may be useful in determining disease progression.

[0092] Further, the systems and methods described herein may be useful in determining the efficacy of a particular treatment option. In one example, certain treatment options may be correlated with genetic cfDNA mutation profiles, frequency of mutations and / or fragmentation profiles of cancers over time. This correlation may be useful in selecting a therapy. Additionally, if a cancer is observed to be in remission after treatment, the systems and methods described herein may be useful in monitoring residual disease or recurrence of disease.

[0093] Further, the methods of the disclosure may be used to characterize the heterogeneity of an abnormal condition in a subject, the method comprising generating a cfDNA mutation profile, frequency of mutations and / or fragmentation profile of extracellular polynucleotides in the subject, wherein the cfDNA mutation profile comprises a plurality of data resulting from profile variation and mutation analyses. In some cases, including but not limited to cancer, a disease may be heterogeneous. Disease cells may not be identical. In the example of cancer, some tumors are known to comprise different types of tumor cells, some cells in different stages of the cancer. In other examples, heterogeneity may comprise multiple foci of disease. Again, in the example ofDOCKET NO.: 348358.18302 cancer, there may be multiple tumor foci, perhaps where one or more foci are the result of metastases that have spread from a primary site (also known as distant metastases).

[0094] The methods of this disclosure may be used to generate a profile, fingerprint, or set of data that is a summation of genetic information derived from different cells in a heterogeneous disease. This set of data may comprise copy number variation and mutation analyses alone or in combination.

[0095] Further, these reports are submitted and accessed electronically via the internet. Analysis of data occurs at a site other than the location of the subject. The report is generated and transmitted to the subject's location. Via an internet enabled computer, the subject accesses the reports reflecting his tumor burden.

[0096] The annotated information can be used by a health care provider to select other drug treatment options and / or provide information about drug treatment options to an insurance company. The method can include annotating the drug treatment options for a condition in, for example, the NCCN Clinical Practice Guidelines in Oncology™ or the American Society of Clinical Oncology (ASCO) clinical practice guidelines.

[0097] Reports are generated, mapping genome positions and cfDNA mutation profile variation for the subject with cancer. These reports, in comparison to other profiles of subjects with known outcomes, can indicate that a particular cancer is aggressive and resistant to treatment. The subject is monitored for a period and retested. If at the end of the period, the cfDNA mutation profiles, frequency of mutations and / or fragmentation variation profile does not vary, this may indicate that the current treatment is not working. A comparison is done with cfDNA mutation profiles of other subjects. For example, if it is determined that a change in cfDNA mutation variation indicates that the cancer is advancing, then the original treatment regimen as prescribed is no longer treating the cancer and a new treatment is prescribed.

[0098] In certain embodiments, the system receives genetic information from a DNA sequencer. The process then determines specific cfDNA alterations and frequencies thereof. These reports are submitted and accessed electronically via the internet. Analysis of data occurs at a site other than the location of the subject. The report is generated and transmitted to the subject'sDOCKET NO.: 348358.18302 location. Via an internet enabled computer, the subject accesses the reports reflecting his tumor burden.

[0099] While temporal information can be used to enhance the information for cfDNA mutation profiles and frequency of mutations, other consensus methods can be applied. In other embodiments, the historical comparison can be used in conjunction with other consensus cfDNA mutation profiles, frequency of mutations and / or fragmentation profiles. Consensus cfDNA mutation profiles and frequency of mutations can be normalized against control samples. Measures of molecules mapping to reference sequences can also be compared across a genome to identify areas in the genome in which cfDNA mutation profiles and frequency of mutations varies, or remains the same. Consensus methods include, for example, linear or non-linear methods of building consensus cfDNA mutation profiles and frequency of mutations (such as voting, averaging, statistical, maximum a posteriori or maximum likelihood detection, dynamic programming, Bayesian, hidden Markov or support vector machine methods, etc.) derived from digital communication theory, information theory, or bioinformatics. After the sequence read coverage has been determined, a stochastic modeling algorithm is applied to convert the normalized nucleic acid sequence read coverage for each window region to the discrete copy number states. In some cases, this algorithm may comprise one or more of the following: Hidden Markov Model, dynamic programming, support vector machine, Bayesian network, trellis decoding, Viterbi decoding, expectation maximization, Kalman filtering methodologies and neural networks. [000100] Artificial neural networks (NNets) mimic networks of “neurons” based on the neural structure of the brain. They process records one at a time, or in a batch mode, and “learn” by comparing their classification of the record (which, at the outset, is largely arbitrary) with the known actual classification of the record. In MLP-NNets, the errors from the initial classification of the first record is fed back into the network, and are used to modify the network's algorithm the second time around, and so on for many iterations. The neural networks use an iterative learning process in which data cases (rows) are presented to the network one at a time, and the weights associated with the input values are adjusted each time. [000101] After all cases are presented, the process often starts over again. During this learning phase, the network learns by adjusting the weights so as to be able to predict the correctDOCKET NO.: 348358.18302 class label of input samples. Neural network learning is also referred to as “connectionist learning,” due to connections between the units. Advantages of neural networks include their high tolerance to noisy data, as well as their ability to classify patterns on which they have not been trained. One neural network algorithm is back-propagation algorithm, such as Levenberg-Marquadt. Once a network has been structured for a particular application, that network is ready to be trained. To start this process, the initial weights are chosen randomly. Then the training, or learning, begins. [000102] The network processes the records in the training data one at a time, using the weights and functions in the hidden layers, then compares the resulting outputs against the desired outputs. Errors are then propagated back through the system, causing the system to adjust the weights for application to the next record to be processed. This process occurs over and over as the weights are continually tweaked. During the training of a network the same set of data is processed many times as the connection weights are continually refined. [000103] In an embodiment, the training step of the machine learning unit on the training data set may generate one or more classification models for applying to a test sample. These classification models may be applied to a test sample to predict the response of a subject to a therapeutic intervention. [000104] Comparison of sequence coverage to a control sample or reference sequence may aid in normalization across windows. In this embodiment, cell free DNAs are extracted and isolated from a readily accessible bodily fluid such as blood. For example, cell free DNAs can be extracted using a variety of methods known in the art, including but not limited to isopropanol precipitation and / or silica based purification. Cell free DNAs may be extracted from any number of subjects, such as subjects without cancer, subjects at risk for cancer, or subjects known to have cancer (e.g. through other means). [000105] Following the isolation / extraction step, any of a number of different sequencing operations may be performed on the cell free polynucleotide sample. Samples may be processed before sequencing with one or more reagents (e.g., enzymes, unique identifiers (e.g., barcodes), probes, etc.). In some cases, if the sample is processed with a unique identifier such as a barcode, the samples or fragments of samples may be tagged individually or in subgroups with the uniqueDOCKET NO.: 348358.18302 identifier. The tagged sample may then be used in a downstream application such as a sequencing reaction by which individual molecules may be tracked to parent molecules. [000106] The cell free polynucleotides can be tagged or tracked in order to permit subsequent identification and origin of the particular polynucleotide. The assignment of an identifier (e.g., a barcode) to individual or subgroups of polynucleotides may allow for a unique identity to be assigned to individual sequences or fragments of sequences. This may allow acquisition of data from individual samples and is not limited to averages of samples. In some examples, nucleic acids or other molecules derived from a single strand may share a common tag or identifier and therefore may be later identified as being derived from that strand. Similarly, all of the fragments from a single strand of nucleic acid may be tagged with the same identifier or tag, thereby permitting subsequent identification of fragments from the parent strand. In other cases, gene expression products (e.g., mRNA) may be tagged in order to quantify expression, by which the barcode, or the barcode in combination with sequence to which it is attached can be counted. In still other cases, the systems and methods can be used as a PCR amplification control. In such cases, multiple amplification products from a PCR reaction can be tagged with the same tag or identifier. If the products are later sequenced and demonstrate sequence differences, differences among products with the same identifier can then be attributed to PCR error. Additionally, individual sequences may be identified based upon characteristics of sequence data for the read themselves. For example, the detection of unique sequence data at the beginning (start) and end (stop) portions of individual sequencing reads may be used, alone or in combination, with the length, or number of base pairs of each sequence read unique sequence to assign unique identities to individual molecules. Fragments from a single strand of nucleic acid, having been assigned a unique identity, may thereby permit subsequent identification of fragments from the parent strand. This can be used in conjunction with bottlenecking the initial starting genetic material to limit diversity. [000107] Generally, the methods and systems provided herein are useful for preparation of cell free polynucleotide sequences to a down-stream application sequencing reaction. Often, a sequencing method is next generation sequencing (NGS), classic Sanger sequencing, whole- genome bisulfite sequencing (WGSB), small-RNA sequencing, low-coverage Whole-Genome Sequencing (lcWGS), etc.DOCKET NO.: 348358.18302 [000108] As used herein, the term “sequencing” refers to any of a number of technologies used to determine the sequence of a biomolecule, e.g., a nucleic acid such as DNA or RNA. Exemplary sequencing methods include, but are not limited to, targeted sequencing, single molecule real-time sequencing, exon sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy termination sequencing, whole-genome sequencing, sequencing by hybridization, pyrosequencing, capillary electrophoresis, gel electrophoresis, duplex sequencing, cycle sequencing, single-base extension sequencing, solid-phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification at lower denaturation temperature-PCR (COLD-PCR), multiplex PCR, sequencing by reversible dye terminator, paired-end sequencing, near-term sequencing, exonuclease sequencing, sequencing by ligation, short-read sequencing, single-molecule sequencing, sequencing-by-synthesis, real-time sequencing, reverse-terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, and a combination thereof. In some embodiments, sequencing can be performer by a gene analyzer such as, for example, gene analyzers commercially available from Illumina or Applied Biosystems. In some embodiments, the sequencing method can be massively parallel sequencing, that is, simultaneously (or in rapid succession) sequencing any of at least 100, 1000, 10,000, 100,000, 1 million, 10 million, 100 million, or 1 billion polynucleotide molecules. [000109] After sequencing, reads are assigned a quality score. A quality score may be a representation of reads that indicates whether those reads may be useful in subsequent analysis based on a threshold. In some cases, some reads are not of sufficient quality or length to perform the subsequent mapping step. Sequencing reads with a quality score at least 90%, 95%, 99%, 99.9%, 99.99% or 99.999% may be filtered out of the data set. In other cases, sequencing reads assigned a quality scored at least 90%, 95%, 99%, 99.9%, 99.99% or 99.999% may be filtered out of the data set. The genomic fragment reads that meet a specified quality score threshold are mapped to a reference genome, or a reference sequence that is known not to contain mutations. After mapping alignment, sequence reads are assigned a mapping score. A mapping score may be a representation or reads mapped back to the reference sequence indicating whether each position is or is not uniquely mappable. In instances, reads may be sequences unrelated to mutation analysis.DOCKET NO.: 348358.18302 For example, some sequence reads may originate from contaminant polynucleotides. Sequencing reads with a mapping score at least 90%, 95%, 99%, 99.9%, 99.99% or 99.999% may be filtered out of the data set. In other cases, sequencing reads assigned a mapping scored less than 90%, 95%, 99%, 99.9%, 99.99% or 99.999% may be filtered out of the data set. For each mappable base, bases that do not meet the minimum threshold for mappability, or low quality bases, may be replaced by the corresponding bases as found in the reference sequence. [000110] Numerous cancers may be detected using the methods and systems described herein. Cancers cells, as most cells, can be characterized by a rate of turnover, in which old cells die and replaced by newer cells. Generally dead cells, in contact with vasculature in a given subject, may release DNA or fragments of DNA into the blood stream. This is also true of cancer cells during various stages of the disease. Cancer cells may also be characterized, dependent on the stage of the disease, by various genetic aberrations such as copy number variation as well as mutations. This phenomenon may be used to detect the presence or absence of cancers individuals using the methods and systems described herein. [000111] The types and number of cancers that may be detected may include but are not limited to blood cancers, brain cancers, lung cancers, skin cancers, nose cancers, throat cancers, liver cancers, bone cancers, lymphomas, pancreatic cancers, skin cancers, bowel cancers, rectal cancers, thyroid cancers, bladder cancers, kidney cancers, mouth cancers, stomach cancers, solid state tumors, heterogeneous tumors, homogenous tumors and the like. In certain embodiments, the cancer comprises ovarian cancer, ovarian cancer subtypes, including high-grade serous (HGSOC), low-grade serous (LGSOC), clear cell, mucinous, and endometroid ovarian cancers. [000112] Additionally, the systems and methods described herein may also be used to help characterize certain cancers. Genetic data produced from the system and methods of this disclosure may allow practitioners to help better characterize a specific form of cancer. Often times, cancers are heterogeneous in both composition and staging. Genetic profile data may allow characterization of specific sub-types of cancer that may be important in the diagnosis or treatment of that specific sub-type. This information may also provide a subject or practitioner clues regarding the prognosis of a specific type of cancer or subtype of cancer.DOCKET NO.: 348358.18302 [000113] The systems and methods provided herein may be used to monitor already known cancers, or other diseases in a particular subject. This may allow either a subject or practitioner to adapt treatment options in accord with the progress of the disease. In this example, the systems and methods described herein may be used to construct genetic profiles of a particular subject of the course of the disease. In some instances, cancers can progress, becoming more aggressive and genetically unstable. In other examples, cancers may remain benign, inactive or dormant. The system and methods of this disclosure may be useful in determining disease progression. [000114] Further, the systems and methods described herein may be useful in determining the efficacy of a particular treatment option. In one example, successful treatment options may actually increase the amount of copy number variation or mutations detected in subject's blood if the treatment is successful as more cancers may die and shed DNA. In other examples, this may not occur. In another example, perhaps certain treatment options may be correlated with genetic profiles of cancers over time. This correlation may be useful in selecting a therapy. Additionally, if a cancer is observed to be in remission after treatment, the systems and methods described herein may be useful in monitoring residual disease or recurrence of disease. [000115] The data is sent over a direct connection or over the internet to a computer for processing. The data processing aspects of the system can be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or in combinations of them. Data processing apparatus of the invention can be implemented in a computer program product tangibly embodied in a machine-readable storage device for execution by a programmable processor; and data processing method steps of the invention can be performed by a programmable processor executing a program of instructions to perform functions of the invention by operating on input data and generating output. The data processing aspects of the invention can be implemented advantageously in one or more computer programs that are executable on a programmable system including at least one programmable processor coupled to receive data and instructions from and to transmit data and instructions to a data storage system, at least one input device, and at least one output device. Each computer program can be implemented in a high-level procedural or object- oriented programming language, or in assembly or machine language, if desired; and, in any case, the language can be a compiled or interpreted language. Suitable processors include, by way of example, both general and special purpose microprocessors. Generally, a processor will receiveDOCKET NO.: 348358.18302 instructions and data from a read-only memory and / or a random access memory. Storage devices suitable for tangibly embodying computer program instructions and data include all forms of nonvolatile memory, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM disks. Any of the foregoing can be supplemented by, or incorporated in, ASICs (application-specific integrated circuits). [000116] To provide for interaction with a user, the methods can be implemented using a computer system having a display device such as a monitor or LCD (liquid crystal display) screen for displaying information to the user and input devices by which the user can provide input to the computer system such as a keyboard, a two-dimensional pointing device such as a mouse or a trackball, or a three-dimensional pointing device such as a data glove or a gyroscopic mouse. The computer system can be programmed to provide a graphical user interface through which computer programs interact with users. The computer system can be programmed to provide a virtual reality, three-dimensional display interface. [000117] While various embodiments of the present disclosure have been described above, it should be understood that they have been presented by way of example only, and not limitation. Numerous changes to the disclosed embodiments can be made in accordance with the disclosure herein without departing from the spirit or scope of the invention. Thus, the breadth and scope of the present disclosure should not be limited by any of the above-described embodiments. [000118] The following examples have been included to provide guidance to one of ordinary skill in the art for practicing representative embodiments of the presently disclosed subject matter. In light of the present disclosure and the general level of skill in the art, those of skill can appreciate that the following examples are intended to be exemplary only and that numerous changes, modifications, and alterations can be employed without departing from the scope of the presently disclosed subject matter. The descriptions and specific examples that follow are only intended for the purposes of illustration, and are not to be construed as limiting in any manner to make compounds of the disclosure by other methods.DOCKET NO.: 348358.18302 EXAMPLES [000119] EXAMPLE 1: EARLY DETECTION OF OVARIAN CANCER USING CELL-FREE DNA FRAGMENTOMES AND PROTEIN BIOMARKERS [000120] Analyses of circulating cell-free DNA (cfDNA) provide another approach for early cancer detection in the screening or diagnostic settings. Approaches for ovarian cancer have included identification of tumor-specific mutations (15,16), or alterations in DNA methylation (17) or specific repeat sequences (18,19), however these approaches have had limited sensitivities for early-stage disease, may be confounded by alterations in white blood cells (20), and have not been validated for clinical use. An emerging approach of cfDNA analyses have focused on the “cfDNA fragmentome”, defined as the genome-wide compendium of cfDNA fragments in the circulation, providing an integrated view of the genome, epigenome, chromatin, and transcriptome states of normal and cancer cells of an individual. Recent cfDNA fragmentome analyses using low-coverage whole genome sequencing combined with machine learning using DNA evaluation of fragments for early interception (DELFI) have demonstrated high sensitivity for early detection across lung (21), liver (22), and other cancer types (23–26) using an accessible, cost-efficient approach (27) that is not confounded by clonal hematopoiesis (20,28). [000121] In this study, a method is presented to detect ovarian cancer using cfDNA fragmentomes combined with protein biomarkers. This multi-analyte combination has the benefit of utilizing genome-wide multi-feature fragmentation analyses together with complementary protein biomarkers CA-125 and HE4 from the same blood draw that may have utility in both the screening and diagnostic settings. [000122] RESULTS [000123] Clinical cohorts [000124] Blood samples in the Discovery Cohort were collected from women with ovarian cancer (n=94), benign adnexal masses (n=203), or without any known ovarian lesions (n=182), who were part of previously reported prospective diagnostic or screening efforts at hospitals in the Netherlands and Denmark (9,21,23,29) (Table 1, Supplementary Table S1). For the Validation Cohort, we analyzed samples from patients prospectively collected at the University of Pennsylvania or through a commercial source in the US (n=40 patients with ovarian cancer,DOCKET NO.: 348358.18302 n=50 patients with benign ovarian masses, n=22 without known ovarian lesions) (Table 1, Supplementary Table S1). The patients analyzed were largely representative of ovarian cancer subtypes, including high-grade serous (HGSOC), low-grade serous (LGSOC), clear cell, mucinous, and endometroid ovarian cancers, across all FIGO (International Federation of Gynecology and Obstetrics) stages (Table 1). [000125] For all participants, we isolated plasma, extracted cfDNA, created genomic libraries, and performed next generation whole-genome sequencing of cfDNA fragments at ~2x coverage. An average of 3 ml of plasma per sample were used and all samples were successfully processed, without any sample or technical failures (Supplementary Table S2, S3). For all patients, we quantified levels of CA-125 and HE4 using clinical grade immunoassay measurements from the same blood samples that were used for genomic analyses or from serum samples of the same patients (Supplementary Table S4). [000126] Table 1. Patient characteristics for Discovery and Validation CohortsDOCKET NO.: 348358.18302 [000128] cfDNA fragmentation profiles that captured fragment size and coverage distributions in 473 nonoverlapping genome-wide 5-Mb regions, covering 2.4Gb of genome (21,22) (FIG.1) were evaluated. Fragmentation profiles were homogenous among individuals without cancer and showed limited changes in individuals with benign adnexal masses (FIG. 2A). In contrast, fragmentation profiles from patients with cancer showed marked heterogeneity both between patients and across different regions of the genome for the same individual, consistent with changes in chromatin landscapes that affect cfDNA fragmentation (FIGS.2A, 2D). [000129] Ovarian tumors are known for having marked large-scale genomic changes (29– 32). As cfDNA fragmentome may reflect large-scale genomic alterations contained in DNA fragments released from tumor cells, we also examined chromosomal copy number changes in the circulation of these individuals. In addition to changes in genome-wide cfDNA fragmentation (Fig.2A-B), we observed chromosomal gains and losses consistent with those expected from prior analyses of ovarian tumors in The Cancer Genome Atlas (TCGA) (n=597) (30) as well as from genomic analyses of early ovarian cancer precursors (32), including gains of 3q, 8q, 12p, 20p and 20q, and losses of 4q, 5q, 6q, 8p, 13q, 17p and 22q (Fig.2C). These gains and losses were not observed in individuals without cancer or with benign adnexal masses, consistent with the notion that while ovarian tumors and benign lesions may share similar anatomic locations, the observed changes in cfDNA were cancer-specific. [000130] Detection of ovarian cancer using cfDNA fragmentome and protein analyses [000131] Given the concordance between genomic changes and cfDNA fragmentation in ovarian cancer, we applied a machine learning approach to ascertain if alterations in cfDNA fragmentomes could distinguish individuals in the Discovery Cohort with ovarian cancer from those without ovarian lesions. The model incorporated genome-wide fragmentation profiles, chromosomal arm-level changes, and the concentrations of protein biomarkers CA-125 and HE4. We previously utilized similar approaches to construct high-performance classifiers for lung and liver cancer detection that were externally validated (21,22). These approaches utilized penalized logistic regression (PLR) due to its parsimonious model architecture, interpretability, and robustness to overfitting. Here, we determined the performance of this classifier using repeated 5-fold cross validation, producing a score for each patient as an average of 10 cross-validation repeats (DELFI Protein (DELFI-Pro) score) (Supplementary Table S5). The DELFI-Pro classifierDOCKET NO.: 348358.18302 utilized both fragmentomic features and proteomic measurements for detecting individuals with ovarian cancer (FIG.2D). The DELFI-Pro classifier utilized a PLR model to retain only the most informative features, including fragmentation characteristics reflecting chromatin and chromosomal changes alongside conventional protein biomarkers. [000132] Since clinical characteristics can influence biomarker profiles evident in the circulation, we examined the relationship between the DELFI-Pro score and demographic parameters such as age or common comorbidities such as diabetes, hypertension, or atherosclerosis in individuals without ovarian disease where this information was available. We observed either no or limited association between DELFI-Pro scores and these conditions, although this conclusion is limited by incomplete availability of clinical information (FIGS.5A- 5C). [000133] We then evaluated the relationship between DELFI-Pro scores and the presence and stage of ovarian cancer. The cross-validated DELFI-Pro scores, spanning a possible range from 0 to 1, for 182 women who were free of ovarian disease were low, with median scores of 0.07. In contrast, women with ovarian cancers had significantly higher median scores across all FIGO stages, including stage I = 0.93, stage II = 0.93, stage III = 1.00, and stage IV = 1.00 (p<0.0001 across all tumor stages, Wilcoxon rank sum test, Fig.3A). Scores did not differ by age (p=0.95, Pearson correlation test) and were not different among women with cancer who were symptomatic or asymptomatic (p=0.61, Wilcoxon signed-rank test) or who were pre- or post-menopausal (p=0.36, Wilcoxon signed-rank test) (FIGS.6A-6C). [000134] DELFI-Pro detected patients with ovarian cancer with an area under the receiver operator characteristic (AUC) of 0.96 (95% confidence interval (CI), 0.93-0.99) (Fig.3B). Among early-stage ovarian cancers, performance remained robust, with AUCs of 0.96 (95% CI=0.92-0.99) and 0.94 (95% CI=0.87-1.00) for stage I (n=32) and II (n=26), respectively (Fig. 3B). Individuals with advanced stage [stages III (n=30) and IV (n=2)] ovarian cancer were detected with high sensitivity among the individuals analyzed (AUCs 0.99 (95% CI=0.98-1.00) and 1.00 (95% CI=1.00-1.00), respectively). Stability analyses of the cross-validated model revealed highly consistent DELFI-Pro scores for non-cancers and cancers regardless of the held- out fold or source of sample collection (FIGS.14A-14B). Performance among patients with HGSOC (n=39) was high with an AUC=0.99 (95% CI=0.99-1.00), as well as in other ovarian cancers, including LGSOC (n=7), endometrioid (n=14), mucinous (n=12), clear cell (n=11) orDOCKET NO.: 348358.18302 other (n=11) subtypes [AUCs of 0.99 (95% CI=0.98-1.00), 0.97 (95% CI=0.94-1.00), 0.94 (95% CI=0.88-1.00), 0.84 (95% CI=0.65-1.00) and 0.96 (95% CI=0.87-1.00), respectively] (FIG.7). High performance was also observed when assessing only individuals who were asymptomatic (AUC 0.99 (95% CI= 0.97-1)) (FIGS.15A-15B). Other genome-wide analyses, such as ichorCNA, which only includes copy number changes, and analyses of overall median cfDNA fragment lengths provided substantially weaker performance with overall AUCs of 0.71 (95% CI = 0.64-0.78) and AUC 0.59 (95% CI= 0.52 – 0.66), respectively (FIGS.16A-16D). [000135] Given the low incidence of ovarian cancer (10.3 out of 100,000 age adjusted women in the US population) (33), any screening test would need to have high specificity in order to give a high positive predictive value and minimize the absolute number of false positive results leading to potentially unnecessary procedures or prolonged diagnostic odysseys. At a specificity >99%, the cross validated sensitivity in this setting was 72%, 69%, 87%, and 100% for stages I–IV, respectively (Table 2). High grade serous ovarian cancers typically had high DELFI-Pro scores, with 90% detected at this threshold (83%, 88%, 91%, and 100% for stages I- IV, respectively). Analysis of CA-125 alone in this population revealed a significantly lower fraction that was detected [34%, 62%, 63%, and 100% of ovarian cancers for stages I–IV (p=0.001, two-sided test of equal proportions] at the same specificity (FIG.8). [000136] In addition to the cross-validated analysis of the Discovery Cohort of EU patients, we evaluated the locked DELFI-Pro classifier in a Validation Cohort of 62 patients from the US. The Validation Cohort included patients across different ovarian cancer subtypes as well as individuals without ovarian cancer (Supplementary Table S1). Similar to the observations from the Discovery Cohort, the fragmentation profiles of women without ovarian cancer in the Validation Cohort were highly uniform across the genome, while patients with ovarian cancer were heterogenous (FIG.9). The chromosomal changes observed in the cfDNA of the US Validation Cohort patients resembled those observed in the EU Discovery group and in ovarian tumor tissue from TCGA (FIGS.10A-10C). The DELFI-Pro model detected patients with cancer in the Validation Cohort with high performance (AUC=0.93, 95% CI=0.87-1.00), including patients with HGSOC (AUC=1.00, 95% CI=1.00-1.00). At a fixed score threshold selected to achieve >99% specificity in the Discovery Cohort, we detected 73% ovarian cancers overall and 81% of HGSOC (FIGS.3C, 3D, FIGS.17A-17C, Supplementary Table S6). These results revealed the shared biological features of cfDNA fragmentation across cohorts and demonstratedDOCKET NO.: 348358.18302 the robustness and generalizability of DELFI-Pro in the detection of ovarian cancer in different populations. [000137] Distinguishing Ovarian Cancer from Benign Masses [000138] We examined whether our approach could be useful in a diagnostic setting for distinguishing between patients with ovarian cancer and patients with benign adnexal masses, which can be difficult to differentiate clinically using ultrasound -based prediction models. We observed that genome-wide fragmentation profiles were different between patients with cancer compared to those with benign lesions (Fig.2A). We trained and cross validated a DELFI-Pro machine learning model in the EU cohort to distinguish ovarian cancers from benign lesions. This machine learning model was similar to that developed for the Screening setting, with the rank-ordered DELFI-Pro scores being highly correlated across all Discovery Cohort ovarian cancer samples (n=94) (R= 0.78, P < 2.2e-16; FIG.18). Using this model, benign lesions had low median scores of 0.17, while patients with cancer had a stage-dependent increase in DELFI-Pro scores. Individuals with benign lesions had similar low scores regardless of lesion size or whether the patient was symptomatic or asymptomatic (FIGS.11A-11B). The model had strong performance in identifying patients with cancer as compared to those with benign lesions, with a ROC AUC of 0.88 (95% CI=0.83-0.92), ranging from 0.82 (95% CI=0.74-0.90) to 1.00 (95% CI=1.00-1.00) for stage I to IV. Patients with HGSOC, LGSOC or endometrioid cancers were more easily distinguished from benign lesions [AUCs of 0.96 (95% CI=0.93-1.00), 0.84 (95% CI=0.67-1.00), and 0.91 (95% CI=0.85-0.98), respectively] than those with mucinous or clear cell subtypes [AUC = 0.65 (95% CI=0.51-0.79) and 0.77 (95% CI=0.62-0.92)] (FIGS.12A-12D, 13). [000139] In the setting of a patient with a mass suspicious for ovarian cancer, the clinical pathway generally involves referral to a gynecological oncologist for surgical staging. A noninvasive test could help inform referral decisions, plan the extent of surgical resection, or even avoid resection in young or frail patients. In these scenarios, high sensitivity is critical to tailoring response, and a moderate specificity may be acceptable, because of the importance of not missing patients with ovarian cancer while at the same time avoiding anxiety and unnecessary surgeries for patients who would not need further follow-up (34). Consistent with this approach, at an 80% specificity in the Discovery Cohort, we distinguished 95% of patients with HGSOC from those individuals with benign masses. Evaluation of the locked model in theDOCKET NO.: 348358.18302 Validation Cohort resulted an AUC of 0.81 (95% CI=0.72-0.91) (FIGS.12C-12D, FIGS.19A- 19C), and at the score threshold achieving 80% specificity in the Discovery Cohort, we maintained a relatively high sensitivity, identifying 81% of patients with HGSOC at a specificity of 82% (Supplementary Table S6). In this setting, the DELFI-Pro scores appeared related to overall tumor burden, as we observed a positive correlation between the DELFI-Pro scores and the sum of reported lesion diameters where these data were available (R=0.65, p=0.03, FIGS. 20A-20B). [000140] Simulating the Performance of DELFI-Pro at Population Scale [000141] To examine how DELFI-Pro would perform for ovarian cancer screening, we used Monte Carlo simulations to evaluate a theoretical screening population of 100,000 women (Fig.4A). We compared DELFI-Pro to two other proposed clinical tests: CA-125 at a cutoff of 30 U / mL and HE4 at a cutoff of 70 pM. For both CA-125 and HE4, we evaluated the cut-point both using performance estimates reported in the literature (35,36) as well as those observed in our cohort, while for DELFI-Pro we evaluated performance at the cut-point achieving greater >99% specificity in our analyses. We blended sensitivity estimates according to the stage distribution in the UKCTOCS trial (36), and modeled the degree of uncertainty of sensitivities and specificities of these tests in our theoretical population, based on a 0.0037 prevalence of ovarian cancer (Fig.4B) (33,37). Monte Carlo simulations from these predicted probability distributions demonstrated that the positive predictive value (PPV) for DELFI-Pro was high (median 23.6%, 95% CI 8.73% - 68.5%), while all other modalities had a median PPV estimate of 9.17% or lower (Fig.4C). Given the low prevalence of ovarian cancer and the risks of exploratory surgery, a PPV greater than 10% (38,39) is needed to justify population-wide screening for ovarian cancer in a way that balances the benefits of early detection against potential harm from unnecessary surgical procedures in healthy women misdiagnosed with cancer. In addition, the cut-off chosen for DELFI-Pro at >99% specificity (no false positives in either the Discovery or Validation Cohorts), led to a predicted low False Positive Rate (median 0.95%, 95% CI 0.14% - 3.1%) as compared to the other four scenarios simulated (range of FPR medians, 3.12% to 20.60%) (Fig.4D). These analyses suggest that an accessible, high adherence, sensitive and specific assay like DELFI-Pro could enable population-wide ovarian cancer screening.DOCKET NO.: 348358.18302 [000142] DISCUSSION [000143] There is a clinical unmet need for an approach that improves detection of early stage ovarian cancer as well as provides guidance toward differentiating between benign or malignant ovarian masses. In this study, we demonstrate that cell-free DNA fragmentomes combined with existing protein biomarkers can noninvasively detect early stage ovarian cancer and distinguish these lesions from benign ovarian masses. [000144] The performance of our multi-analyte and multi-feature approach for detection of ovarian cancer was high and suggests that the combination of cfDNA and protein measurements is complementary, providing more information than either alone. This is particularly important for early stage disease, especially for high grade serous cancer, where intervention is thought to be most useful (32,40). The validation of this approach in a fully independent cohort suggests that the method is robust and likely generalizable across different populations. [000145] The survival benefit of current methods used in screening for ovarian cancer remains unclear (6,41). However, the development of a new classifier like DELFI-Pro that provides high performance at higher specificity than obtained in previous studies opens a new avenue for detection of individuals who may benefit most from subsequent diagnostic work-up or intervention. Our population-scale simulations suggest that the improved performance of DELFI-Pro in comparison to either CA-125 or HE4 alone would increase the positive predictive value (PPV) and decrease the predicted false positive rate (FPR), thereby improving the overall impact and benefit-to-risk ratio of this approach in a screening setting, especially when the disease prevalence is low. Recent literature suggests that early diagnoses of cancers reduce treatment costs (42), thereby potentially decreasing overall societal health care costs while improving outcomes. [000146] Although the study was performed in a sizeable European diagnostic cohort, it was subsequently validated in a modest sized but fully external US cohort. Additional and larger prospective studies, including one already underway (NCT04971421), will be needed to validate this approach for clinical use. In other cancer types we have previously associated DELFI fragmentation scores with survival outcome data (21), and future efforts are needed to evaluate the prognostic potential of the DELFI-Pro score in ovarian cancer. The performance for detecting some subtypes of ovarian cancer (i.e. clear cell or mucinous) was lower and inclusion of other protein biomarkers in our assay that are tailored to these cell types of origin mayDOCKET NO.: 348358.18302 increase performance in the future. Assays using a larger number of proteins have shown promising initial results (43) but are not yet broadly available for research or clinical use. The use of cfDNA fragmentation to distinguish between different cancer subtypes (21) may be feasible for differentiating among ovarian cancer subtypes and enabling personalized therapeutic approaches. [000147] Ultimately, evaluation of survival outcomes will be important to demonstrate the benefit of population scale screening with this approach, as stage shift alone may not result in an effective alternative measure of survival for ovarian cancer (6,44). The use of both cfDNA and protein measurements may initially appear to be complex, but both types of analytes can be assessed from the same sample of blood, and optimized methods suggest that this combined approach would be cost-efficient and accessible. Overall, this study provides a new accessible approach for early detection of ovarian cancer that may overcome current challenges for ovarian cancer screening and reduce the morbidity and mortality of this disease. [000148] EXAMPLE 2: METHODS [000149] Study population and design [000150] Liquid biopsies from 591 healthy individuals or individuals with ovarian cancer or benign adnexal masses were prospectively collected at University of Pennsylvania (Penn BioTrust Collection: RRID SCR_022387), from previously reported diagnostic or screening studies at the Netherlands Cancer Institute (trial NL58253.031.16) (9,29), the Danish Endoscopy III trial (21,23,45), or the Netherlands COCOS trial (Netherlands trial register ID NTR1829) (21,23), or through a commercial provider of biobanked research specimens (BioIVT). All samples were obtained under Institutional Review Board approved protocols with written informed consent from all participants for research use at participating institutions, and the studies were performed according to the Declaration of Helsinki. Liquid biopsies from healthy individuals were obtained at the time of routine clinical appointments. Individuals were considered healthy if they had no prior history of cancer. Individuals with symptoms indicating clinical follow-up or at high risk for development of ovarian cancer were assessed using imaging of the pelvic region to identify ovarian masses. Depending on size and estimated risk of malignancy, patients received either an exploratory laparotomy with frozen section (and staging when confirmed malignant or debulking if unexpected higher stage) or, if lesion was expected to be benign, laparoscopic cystectomy / adnectomy. Liquid biopsies from patients with ovarianDOCKET NO.: 348358.18302 cancer or benign adnexal masses were obtained at the time of diagnosis, prior to surgical resection or therapeutic intervention. Of the total 591 women included in the study, 204 were healthy, 253 had a benign adnexal mass, and 134 had ovarian cancer. All stages of ovarian cancer were represented in the study population including 46, 31, 41, and 8 individuals with stage I, II, III, and IV cancer, respectively (n=8, stage unknown). The cancer cohort was comprised largely of high grade serous ovarian cancer (n=55) with a subset of individuals having Low grade serous or Serous (n=12), clear cell (n=13), mucinous (n=17), endometrioid (n=19) or another (n=18) histopathological ovarian cancer diagnosis. Clinical data were completely de- identified for all individuals included in this study and are listed in Supplementary Table S1. [000151] This study was designed to provide proof-of-concept for noninvasive detection of ovarian cancer using a genome-wide fragmentomics-based approach. For cfDNA analyses, all liquid biopsies were processed to separate blood plasma from which cfDNA was extracted and processed to create genomic libraries for whole genome sequencing at ~2x coverage. The study population was subset to assess a Discovery Cohort to train and cross-validate a machine learning model for ovarian cancer detection, followed by application of the trained model to the subset of the population remaining as the Validation Cohort. Prediction of ovarian cancer was assessed in two clinical scenarios: 1. Screening model (ovarian cancer vs. no ovarian lesion), and 2. Diagnostic model (ovarian cancer vs. benign mass). The Discovery Cohorts were defined to include: (i) healthy individuals with no history of prior cancer and patients with ovarian cancer or (ii) individuals with a benign adnexal mass and patients with ovarian cancer for the screening and diagnostic models, respectively (n=479 Discovery, n=112 Validation). [000152] Liquid biopsy collection and extraction of cfDNA [000153] We collected venous peripheral blood in K2-EDTA or Streck tubes and, within two hours, centrifuged tubes at 800 × g at 4°C for 10 minutes. Then the plasma fraction was transferred to new tubes and spun at 18,000 x g for 10 minutes at room temperature to pellet remaining cellular debris. EDTA tubes from the Endoscopy III trial were centrifuged at low speed (3000 g) for 10 min within two hours from blood collection. The plasma portion from the first spin was spun a second time for 10 min. Plasma was subsequently aliquoted and stored at - 80°C. cfDNA was isolated from ~4-5ml of plasma using the Qiagen QIAamp Circulating Nucleic Acids Kit (Qiagen GmbH). Extracted cfDNA was eluted in 52ul into LoBind tubes (Eppendorf AG) and quantified using the Bioanalyzer 2100 (Agilent Technologies).DOCKET NO.: 348358.18302 [000154] Genomic library construction [000155] cfDNA libraries for next-generation whole-genome sequencing were prepared with 15 ng of cfDNA when available or the entire purified amount when less than 15 ng (Supplementary Table S2) (21–23). The genomic libraries were prepared using the NEBNext DNA Library Prep Kit for Illumina (New England Biolab) with four main modifications to the manufacturer’s guidelines: (i) the library purification steps followed the on-bead AMPure XP (Beckman Coulter) approach to minimize sample loss during elution and tube transfer steps; (ii) NEBNext End Repair, A-tailing, and adapter ligation enzyme and buffer volumes were adjusted as appropriate to accommodate on-bead AMPure XP purification; (iii) Illumina dual index adapters were used in the ligation reaction; and (iv) cfDNA libraries were amplified with Phusion Hot Start Polymerase. All samples underwent a 4-cycle PCR amplification after the DNA ligation step. [000156] Both genomic sequencing and protein measurements were performed in batches that included samples from individuals with or without cancer, including from other studies, to reduce the possibility that differences between patients with or without cancer were not due to batch variability (Supplementary Table S2). [000157] Whole-genome sequencing and alignment [000158] Whole-genome libraries were sequenced using 100-bp paired-end runs (200 cycles) on the Illumina HiSeq2500 platform at 1-2× coverage per genome (21–23). Before alignment, adapter sequences were filtered from reads using FASTP software (46). Sequence reads were then aligned to the hg19 human reference genome with Bowtie2 (47), duplicate reads were removed using Sambamba (48), and each aligned pair was converted to a genomic interval representing the sequenced DNA fragment using bedtools (49). Reads with a MAPQ score of less than 30 or that overlapped the Duke Excluded Regions blacklist (genome. ucsc.edu / cgi- bin / hgTrackUi?db=hg19&g=wgEncodeMapability) were excluded. To construct fragmentation profiles from low-coverage whole-genome sequencing that reflected large-scale epigenetic differences in fragmentation across the genome, we partitioned the hg19 reference genome into nonoverlapping 5-Mb bins. Bins with mean GC base content < 0.3 or mean mappability < 0.9 were excluded, leaving 473 bins spanning approximately 2.4 Gb of the genome. A fragment level GC correction was performed independently for short (<150 bp) and long (≥150 bp) cfDNADOCKET NO.: 348358.18302 fragments using an external reference panel of individuals without cancer to generate a target distribution, as previously described (21,22). [000159] Genome-wide fragmentome analyses [000160] Fragmentation features were calculated as the ratio of short to long fragments in 473 nonoverlapping 5-Mb bins across the genome, and as z-scores representing arm gains / losses for autosomal chromosome arms as described in our previous publications (21,22). [000161] Analyses of publicly available TCGA data [000162] Copy-number data from the OVCA cancer cohort in TCGA [ovarian cancer (OVCA) n = 597] were retrieved using the package RTCGA v1.16.0 and were analyzed to determine the frequency of copy-number gains and losses in the 4735-Mb bins for this cohort (21,22). A somatic copy-number alteration threshold was used to call gains and losses in the ovarian cohorts (21,50). [000163] Proteins [000164] Protein analyses were conducted on matched serum of the same patient or plasma from the same aliquot used for cell-free DNA isolation. The proteins CA-125 U / mL and HE4 pM (Roche Elecsys II) were measured using the Roche Cobas e 602 immunoassay analyzer in EDTA or Streck collected plasma (n = 435) and serum (n =70). Evaluation of protein measurements showed high correlation between current analyses and replicate measurements performed at other institutions (FIGS.20-20B). A subset of plasma from trial NL58253.031.16 was not available for protein analyses but had been previously evaluated for CA-125 (serum) and HE4 (plasma) (n=86). Assessment of proteins from multiple centers were batched by biospecimen type, source of collection, prior CA-125 data availability as well as cancer stage or non-cancer status and contained a set of technical replicates across batches. CA-125 and HE4 were measured at the Johns Hopkins Clinical Chemistry Research Laboratory, Department of Pathology, Division of Clinical Chemistry. [000165] Machine learning and cross-validation [000166] Two machine learning models were developed to predict the presence of ovarian cancer in (i) a screening setting, and (ii) a diagnostic setting. Both models used Penalized logistic regression and features included fragmentation profiles, chromosomal arm-level changes, as well as the protein biomarkers CA-125 and HE4. The models were trained and cross-validated using data from (i) individuals in the subset of the Discovery group with ovarianDOCKET NO.: 348358.18302 cancer or without any known ovarian lesions for the screening model, and (ii) Individuals in the subset of the Discovery group with ovarian cancer or benign adnexal masses for the diagnostic model. The principal components of the ratios representing greater than 90% of variance and the z-scores (21,22), along with levels of the protein biomarkers CA-125 and HE4, were used to train machine learning models. Training was performed with 10 repeats of 5-fold cross validation, generating a DELFI-Pro score for every individual in the Discovery Cohort, that was the average over 10 cross-validation repeats. For the Validation Cohorts, DELFI-Pro scores were generated using the locked models. Performance of the models was assessed using receiver- operator curve analyses, and at fixed score thresholds for set specificities in the Discovery Cohort. [000167] Association of clinical covariates and DELFI score [000168] Potential associations between clinical covariates and the DELFI-Pro score were assessed with Spearman rank correlation coefficient (continuous variables), Wilcoxon signed- rank test (two categorical variables) and Kruskal–Wallis one-way analysis of variance (>2 categorical variables). [000169] Modeling of DELFI performance in screening and diagnostic settings [000170] Monte Carlo simulations were used to compare the DELFI-Pro approach to other proposed biomarkers (CA-125 and HE4) in a theoretical surveillance population. For CA-125, we used published sensitivities and specificities for CA-125 at 30 U / mL from the UKCTOCs trial (36), and estimated sensitivity and specificity in our cohort using the same threshold. For HE4, we used published sensitivities and specificities for HE4 at 70 pM from (35), and estimated sensitivity and specificity in our cohort using the same threshold. For DELFI-Pro, we used sensitivity based on the score threshold yielding >99% specificity in the Discovery Cohort. We blended by-stage sensitivity estimates according to the stage distribution of cancers in UKCTOCS (36), and drew a 95% binomial confidence interval around each sensitivity and specificity estimate. As noninvasive blood-based tests have a reported adherence of more than 75% (51,52), we assumed a point estimate of 75% adherence to a blood based biomarker test, with a 95% confidence interval of 60% to 90%. [000171] The R package epiR was used to construct prior predictive probability distributions (beta distributions) from these CIs (R package version 2.47, epiR; RRID:SCR_021673) for sensitivity, specificity and adherence. We estimated prevalence ofDOCKET NO.: 348358.18302 ovarian cancer as 0.0037 using SEER (37) and US Census data (53) as follows: In 2020 there were 236,511 women with ovarian cancer in the United States, and 2022 census data indicated 63,757,324 women, age 50+. For a single Monte Carlo simulation for DELFI-Pro, we: [000172] Sampled the probability of adherence (η) from the prior predictive distribution, [000173] Simulated the number of 100,000 individuals (S) who participated in screening (S ~ Binomial(η,100,000)), [000174] Sampled prevalence of ovarian cancer [θ ~ Beta(236511, 63520813)] [000175] Simulated ovarian cancer cases (P ~ Binomial(θ, S)) and computed the number of individuals without cancer (N ¬= S P), [000176] Sampled the sensitivity (se) and specificity (sp) from the corresponding prior predictive distributions, and [000177] sampled the true positives (TP ~ Binomial(P, se)) and false positives (FP ~ Binomial(N,1− sp)). [000178] Given TP and FP, we calculated the Positive Predictive Value (PPV) as (TP) / (TP + FP) and the False Positive Rate (FPR) as FP / N. We repeated the above simulation 1,000 times, obtaining a distribution of PPV and FPR. Using parameters for sensitivity, specificity, and adherence for the CA-125 and HE4 scenarios, we repeated the same Monte Carlo analysis to allow comparisons between the different proposed screening methodologies. [000179] Bioinformatic and statistical software [000180] All statistical analyses were performed using R version 4.1.2. Trimming of adapter sequences was performed using fastp (0.20.0). We used Bowtie2 (2.3.0) to align paired- end reads to the hg19 reference genome. PCR duplicates were removed using Sambamba (0.6.8), and the remaining aligned read pairs were converted to a bed format using Bedtools (2.29.0). We used the R package data.table (1.12.8) for manipulation of tabular data and binning fragments in 5-Mb windows across the genome. The R package Caret (6.0.84) was used to implement the classification by penalized logistic regression and resampling. [000181] Statistics, reproducibility, and data availability [000182] The code and data needed for generating figures and results are available at https: / / github.com / cancer-genomics / delfipro2024. Code needed to run the DELFI pipeline and generate features used in modelling is available at github.com / cancer-genomics / delfi3. Sequence data and clinical variables generated in this study have been deposited at the database ofDOCKET NO.: 348358.18302 European Genome-Phenome Archive (EGA) under accession code EGAS00001005340 and EGAS50000000484.20381203220DOCKET NO.: 348358.18302 [000187] REFERENCES 1. Sung H, Ferlay J, Siegel RL, Laversanne M, Soerjomataram I, Jemal A, et al. Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. Wiley; 2021;71:209–49. 2. Siegel RL, Giaquinto AN, Jemal A. Cancer statistics, 2024. CA Cancer J Clin. 2024;74:12–49. 3. American Cancer Society. Ovarian Cancer Survival Rates [Internet]. Ovarian Cancer Early Detection, Diagnosis, and Staging. [cited 2024 Mar 12]. Available from: https: / / www.cancer.org / cancer / types / ovarian-cancer / detection-diagnosis-staging / survival- rates.html 4. Ruhl J, Callaghan C, Schussler N. Summary Stage 2018: Codes and Coding Instructions. Bethesda, MD: National Cancer Institute; 2023. 5. Prorok PC, Andriole GL, Bresalier RS, Buys SS, Chia D, Crawford ED, et al. Design of the Prostate, Lung, Colorectal and Ovarian (PLCO) Cancer Screening Trial. Control Clin Trials. 2000;21:273S-309S. 6. Menon U, Gentry-Maharaj A, Burnell M, Singh N, Ryan A, Karpinskyj C, et al. Ovarian cancer population screening and mortality after long-term follow-up in the UK Collaborative Trial of Ovarian Cancer Screening (UKCTOCS): a randomised controlled trial. Lancet. 2021;397:2182–93. 7. Han CY, Lu KH, Corrigan G, Perez A, Kohring SD, Celestino J, et al. Normal Risk Ovarian Screening Study: 21-Year Update. J Clin Oncol.2024;JCO2300141. 8. Kim J, Park EY, Kim O, Schilder JM, Coffey DM, Cho C-H, et al. Cell Origins of High- Grade Serous Ovarian Cancer. Cancers.2018;10:433. 9. Lof P, van de Vrie R, Korse CM, van Gent MDJM, Mom CH, Rosier-van Dunné FMF, et al. Can serum human epididymis protein 4 (HE4) support the decision to refer a patient with an ovarian mass to an oncology hospital? Gynecol Oncol.2022;166:284–91.10. Schummer M, Ng WV, Bumgarner RE, Nelson PS, Schummer B, Bednarski DW, et al. Comparative hybridization of an array of 21,500 ovarian cDNAs for the discovery of genes overexpressed in ovarian carcinomas. Gene.1999;238:375–85. 11. Chen F, Shen J, Wang J, Cai P, Huang Y. Clinical analysis of four serum tumor markers in 458 patients with ovarian tumors: diagnostic value of the combined use of HE4, CA125, CA19- 9, and CEA in ovarian tumors. Cancer Manag Res.2018;10:1313–8. 12. Reilly G, Bullock RG, Greenwood J, Ure DR, Stewart E, Davidoff P, et al. Analytical Validation of a Deep Neural Network Algorithm for the Detection of Ovarian Cancer. JCO Clin Cancer Inform.2022;6:e2100192. 13. Jacobs I, Oram D, Fairbanks J, Turner J, Frost C, Grudzinskas JG. A risk of malignancy index incorporating CA 125, ultrasound and menopausal status for the accurate preoperative diagnosis of ovarian cancer. Br J Obstet Gynaecol.1990;97:922–9. 14. Timmerman D, Valentin L, Bourne TH, Collins WP, Verrelst H, Vergote I, et al. Terms, definitions and measurements to describe the sonographic features of adnexal tumors: a consensus opinion from the International Ovarian Tumor Analysis (IOTA) Group. Ultrasound Obstet Gynecol.2000;16:500–5. 15. Phallen J, Sausen M, Adleff V, Leal A, Hruban C, White J, et al. Direct detection of early- stage cancers using circulating tumor DNA. Science Translational Medicine [Internet].2017;9. Available from: https: / / www.ncbi.nlm.nih.gov / pubmed / 28814544 16. Lennon AM, Buchanan AH, Kinde I, Warren A, Honushefsky A, Cohain AT, et al. Feasibility of blood testing combined with PET-CT to screen for cancer and guide intervention. Science [Internet].2020;369. Available from: http: / / dx.doi.org / 10.1126 / science.abb9601 17. Klein EA, Richards D, Cohn A, Tummala M, Lapham R, Cosgrove D, et al. Clinical validation of a targeted methylation-based multi-cancer early detection test using an independent validation set. Ann Oncol.2021;32:1167–77. 18. Taylor MS, Wu C, Fridy PC, Zhang SJ, Senussi Y, Wolters JC, et al. Ultrasensitive Detection of Circulating LINE-1 ORF1p as a Specific Multicancer Biomarker. Cancer Discov. 2023;13:2532–47.19. Sato S, Gillette M, de Santiago PR, Kuhn E, Burgess M, Doucette K, et al. LINE-1 ORF1p as a candidate biomarker in high grade serous ovarian carcinoma. Sci Rep. 2023;13:1537. 20. Leal A, van Grieken NCT, Palsgrove DN, Phallen J, Medina JE, Hruban C, et al. White blood cell and cell-free DNA analyses for detection of residual disease in gastric cancer. Nat Commun.2020;11:525. 21. Mathios D, Johansen JS, Cristiano S, Medina JE, Phallen J, Larsen KR, et al. Detection and characterization of lung cancer using cell-free DNA fragmentomes. Nat Commun. 2021;12:5060. 22. Foda ZH, Annapragada AV, Boyapati K, Bruhm DC, Vulpescu NA, Medina JE, et al. Detecting Liver Cancer Using Cell-Free DNA Fragmentomes. Cancer Discov.2023;13:616–31. 23. Cristiano S, Leal A, Phallen J, Fiksel J, Adleff V, Bruhm DC, et al. Genome-wide cell- free DNA fragmentation in patients with cancer. Nature.2019;570:385–9. 24. Bruhm DC, Mathios D, Foda ZH, Annapragada AV, Medina JE, Adleff V, et al. Single- molecule genome-wide mutation profiles of cell-free DNA for non-invasive detection of cancer. Nature Genetics.2023;55:1301–10. 25. Annapragada AV, Niknafs N, White JR, Bruhm DC, Cherry C, Medina JE, et al. Genome- wide repeat landscapes in cancer and cell-free DNA. Sci Transl Med.2024;16:eadj9283. 26. Noë M, Mathios D, Annapragada AV, Koul S, Foda ZH, Medina JE, et al. DNA methylation and gene expression as determinants of genome-wide cell-free DNA fragmentation. Nat Commun.2024;15:6690. 27. Medina JE, Dracopoli NC, Bach PB, Lau A, Scharpf RB, Meijer GA, et al. Cell-free DNA approaches for cancer early detection and interception. J Immunother Cancer [Internet]. 2023;11. Available from: http: / / dx.doi.org / 10.1136 / jitc-2022-006013 28. van ’t Erve I, Medina JE, Leal A, Papp E, Phallen J, Adleff V, et al. Metastatic Colorectal Cancer Treatment Response Evaluation by Ultra-Deep Sequencing of Cell-Free DNA and Matched White Blood Cells. Clin Cancer Res.2023;29:899–909.29. Gaillard DHK, Lof P, Sistermans EA, Mokveld T, Horlings HM, Mom CH, et al. Evaluating the effectiveness of pre-operative diagnosis of ovarian cancer using minimally invasive liquid biopsies by combining serum human epididymis protein 4 and cell-free DNA in patients with an ovarian mass. Int J Gynecol Cancer.2024;34:713–21. 30. Cancer Genome Atlas Research Network. Integrated genomic analyses of ovarian carcinoma. Nature.2011;474:609–15. 31. Macintyre G, Goranova TE, De Silva D, Ennis D, Piskorz AM, Eldridge M, et al. Copy number signatures and mutational processes in ovarian carcinoma. Nat Genet.2018;50:1262–70. 32. Labidi-Galy SI, Papp E, Hallberg D, Niknafs N, Adleff V, Noe M, et al. High grade serous ovarian carcinomas originate in the fallopian tube. Nat Commun.2017;8:1093. 33. SEER Cancer Stat Facts: Ovarian Cancer [Internet]. SEER Cancer Stat Facts: Ovarian Cancer. National Cancer Institute. Bethesda, MD. [cited 2024 Mar 16]. Available from: https: / / seer.cancer.gov / statfacts / html / ovary.html 34. Lof P, Engelhardt EG, van Gent MDJM, Mom CH, Rosier-van Dunné FMF, van Baal WM, et al. Psychological impact of referral to an oncology hospital on patients with an ovarian mass. Int J Gynecol Cancer.2022;33:74–82. 35. Jacob F, Meier M, Caduff R, Goldstein D, Pochechueva T, Hacker N, et al. No benefit from combining HE4 and CA125 as ovarian tumor markers in a clinical setting. Gynecol Oncol. 2011;121:487–91. 36. Menon U, Ryan A, Kalsi J, Gentry-Maharaj A, Dawnay A, Habib M, et al. Risk Algorithm Using Serial Biomarker Measurements Doubles the Number of Screen-Detected Cancers Compared With a Single-Threshold Rule in the United Kingdom Collaborative Trial of Ovarian Cancer Screening. J Clin Oncol.2015;33:2062–71. 37. SEER*Explorer: An interactive website for SEER cancer statistics [Internet] [Internet]. Surveillance Research Program, National Cancer Institute.2023 [cited 2024 Mar 17]. Available from: https: / / seer.cancer.gov / statistics-network / explorer / 38. OVARIAN CANCER SCREENING. American College of Medical Genetics; 1999 [cited 2024 Mar 17]; Available from: https: / / www.ncbi.nlm.nih.gov / books / NBK56952 / 39. Mathieu KB, Bedi DG, Thrower SL, Qayyum A, Bast RC Jr. Screening for ovarian cancer: imaging challenges and opportunities for improvement. Ultrasound Obstet Gynecol. 2018;51:293–303. 40. Menon U, Gentry-Maharaj A, Burnell M, Ryan A, Singh N, Manchanda R, et al. Tumour stage, treatment, and survival of women with high-grade serous tubo-ovarian cancer in UKCTOCS: an exploratory analysis of a randomised controlled trial. Lancet Oncol. 2023;24:1018–28. 41. Buys SS, Partridge E, Black A, Johnson CC, Lamerato L, Isaacs C, et al. Effect of screening on ovarian cancer mortality: the Prostate, Lung, Colorectal and Ovarian (PLCO) Cancer Screening Randomized Controlled Trial. JAMA.2011;305:2295–303. 42. Connal S, Cameron JM, Sala A, Brennan PM, Palmer DS, Palmer JD, et al. Liquid biopsies: the future of cancer early detection. J Transl Med.2023;21:118. 43. Cohen JD, Li L, Wang Y, Thoburn C, Afsari B, Danilova L, et al. Detection and localization of surgically resectable cancers with a multi-analyte blood test. Science. 2018;359:926–30. 44. Bach PB. Late-Stage Cancer End Points to Speed Cancer Screening Clinical Trials-Not So Fast. JAMA.2024. page 1894–5. 45. Rasmussen L, Wilhelmsen M, Christensen IJ, Andersen J, Jørgensen LN, Rasmussen M, et al. Protocol Outlines for Parts 1 and 2 of the Prospective Endoscopy III Study for the Early Detection of Colorectal Cancer: Validation of a Concept Based on Blood Biomarkers. JMIR Res Protoc.2016;5:e182. 46. Chen S, Zhou Y, Chen Y, Gu J. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics.2018;34:i884–90. 47. Langmead B, Salzberg SL. Fast gapped-read alignment with Bowtie 2. Nat Methods. Nature Publishing Group; 2012;9:357–9. 48. Tarasov A, Vilella AJ, Cuppen E, Nijman IJ, Prins P. Sambamba: fast processing of NGS alignment formats. Bioinformatics.2015;31:2032–4.49. Quinlan AR, Hall IM. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics.2010;26:841–2. 50. Davoli T, Uno H, Wooten EC, Elledge SJ. Tumor aneuploidy correlates with markers of immune evasion and with reduced response to immunotherapy. Science [Internet].2017;355. Available from: https: / / www.ncbi.nlm.nih.gov / pubmed / 28104840 51. Bokhorst LP, Alberts AR, Rannikko A, Valdagni R, Pickles T, Kakehi Y, et al. Compliance Rates with the Prostate Cancer Research International Active Surveillance (PRIAS) Protocol and Disease Reclassification in Noncompliers. Eur Urol.2015;68:814–21. 52. Duffy MJ, van Rossum LGM, van Turenhout ST, Malminiemi O, Sturgeon C, Lamerz R, et al. Use of faecal markers in screening for colorectal neoplasia: a European group on tumor markers position paper. Int J Cancer.2011;128:3–11. 53. Day JC. Population Projections of the United States by Age, Sex, Race, and Hispanic Origin: 1995 to 2050, U.S. Bureau of the Census, Current Population Reports. Washington, DC: U.S. Government Printing Office; page 25–1130. OTHER EMBODIMENTS [000188] From the foregoing description, it will be apparent that variations and modifications may be made to the disclosure described herein to adopt it to various usages and conditions. Such embodiments are also within the scope of the following claims. [000189] All citations to sequences, patents and publications in this specification are herein incorporated by reference to the same extent as if each independent patent and publication was specifically and individually indicated to be incorporated by reference. By their citation of various references in this document, Applicants do not admit any particular reference is “prior art” to their disclosure.

Claims

What is claimed:

1. A method of early detection of cancer and treatment of a subject comprising: (i) determining a cell free DNA (cfDNA) fragmentation profile of the subject, the method comprising: extracting and enriching cell free DNA (cfDNA) from a subject’s biological sample; conducting low coverage whole genome sequencing of the subject’s isolated cfDNA to generate genomic libraries of sequenced cfDNA fragments; determining a cfDNA fragmentome profile of the subject’s sequenced DNA fragments comprising evaluating one or more cfDNA fragment characteristics; (ii) comparing the cfDNA fragmentome profiles of the subject to reference cfDNA fragmentome profiles from healthy subjects to detect and diagnose the cancer; (iii) assaying the subject’s biological sample to detect and quantify at least one biomarker; (iv) comparing the subject’s cfDNA fragmentation profile and protein biomarkers with normal reference non-cancer subjects; and, treating the subject diagnosed with cancer with a cancer specific therapy.

2. The method of claim 1, wherein the one or more cfDNA the one or more cfDNA fragment characteristics comprises calculating ratios of small to large cfDNA fragments, sequences of one or more cfDNAs, median cfDNA fragment sizes, fragment size distribution, mutant allele frequencies, fragment length, fragment size distribution, fragment end motifs, preferred end coordinates, breakpoint motifs, methylation frequencies, concentrations of cfDNA or combinations thereof.

3. The method of claim 2, wherein a small cfDNA fragment comprises about 80 base pairs (bp) to about 150 bp.

4. The method of claim 2 or 3, wherein a large cfDNA fragment comprises about 151 bp to about 300 bp.

5. The method of any of claims 2 to 4, wherein the small to large cfDNA ratios are GC corrected.

6. The method of any of claims 2 to 5, wherein the cfDNA fragmentome profile comprises the sequence coverage of small cfDNA fragments in windows across the genome.

7. The method of any of claims 2 to 6, wherein the cfDNA fragmentome profile comprises the sequence coverage of cfDNA fragments in windows across the genome.

8. The method of any of claims 2 to 7, wherein the cfDNA fragmentome profile comprises the sequence coverage of small and large cfDNA fragments in windows across the genome.

9. The method of any of claims 1 to 8, wherein the cfDNA fragmentome profiles in subjects identified and diagnosed early with cancer, are altered across the genome.

10. The method of any of claims 1 to 9, wherein the cfDNA fragmentome profiles in subjects identified and diagnosed early with cancer, have greater heterogeneity across the genome as compared to healthy subjects.

11. The method of claims 9 or 10 wherein the cancer is ovarian cancer or an adnexal mass.

12. The method of any one of claims 7-11, wherein the windows each comprise about 5 million base pairs.

13. The method of any one of claims 7-12, wherein a cfDNA fragmentation profile is determined within each window.

14. The method of any one of claims 1-13, wherein the cfDNA fragmentation profile comprises the sequence coverage of small cfDNA fragments in windows across the genome.

15. The method of any one of claims 1-13, wherein the cfDNA fragmentation profile comprises the sequence coverage of large cfDNA fragments in windows across the genome.

16. The method of any one of claims 1-15, wherein the cfDNA fragmentation profile comprises the sequence coverage of small and large cfDNA fragments in windows across the genome.

17. The method of any one of claims 1-16, wherein the cfDNA fragmentation profile is over the whole genome.

18. The method of any one of claims 1-17, wherein the cfDNA fragmentation profile is over a subgenomic interval.

19. The method of any one of claims 1-18, further comprising assaying for chromosomal gains and losses in the subject’s genome as compared to a normal reference genome.

20. The method of any one of claims 1-19, further comprising a machine learning model wherein the model incorporates genome-wide fragmentation profiles, chromosomal arm-level changes, and the concentrations of ovarian cancer biomarkers.

21. The method of claim 20, wherein the ovarian cancer biomarkers comprise Carbohydrate Antigen 125 (CA-125), Osteopontin (OPN), Kallikreins (KLKs), Bikunin, Human Epididymis Protein 4 (HE4), Vascular Endothelial Growth Factor (VEGF), Prostasin (PSN), Creatine Kinase B (CKB), Mesothelin, Apolipoprotein A-I (apoA-I), Transthyretin (TTR), Transferrin or combinations thereof.

22. The method of claim 21, wherein the ovarian cancer biomarkers comprise Carbohydrate Antigen 125 (CA-125), Human Epididymis Protein 4 (HE4) or the combination thereof.

23. The method of claim 20, wherein the model generates a DELFI protein (DELFI-Pro) score.

24. The method of claim 23, wherein the DELFI-Pro score is diagnostic of cancer.

25. The method of claim 23, wherein the DELFI-Pro score is diagnostic of the stage of cancer.

26. The method of claim 23, wherein the DELFI-Pro score is diagnostic of the subtype of ovarian cancer.

27. The method of claim 23, wherein the DELFI-Pro score is diagnostic of an adnexal mass.

28. The method of any one of claims 23-27, wherein subjects with low median DELFI-Pro scores are ovarian cancer free or do not have an adnexal mass.

29. The method of claim 28, wherein a low median DELFI-Pro score is in a range of about 0.0001 to about 0.

1.

30. The method of any one of claims 23-27, wherein subjects with high median DELFI-Pro scores are diagnosed as having ovarian cancer or an adnexal mass.

31. The method of claim 30, wherein ranges of high median DELFI-Pro scores are diagnostic of stages of ovarian cancer.

32. The method of claims 34 or 35, wherein the higher median DELFI-Pro scores comprise greater than about 0.

7.

33. A method of distinguishing between ovarian cancer, ovarian cancer subtypes or benign masses in subjects, comprising: (i) determining a cell free DNA (cfDNA) fragmentation profile of the subject, the method comprising: extracting and enriching cell free DNA (cfDNA) from a subject’s biological sample; conducting low coverage whole genome sequencing of the subject’s isolated cfDNA to generate genomic libraries of sequenced cfDNA fragments; determining a cfDNA fragmentome profile of the subject’s sequenced DNA fragments comprising evaluating one or more cfDNA fragment characteristics; (ii) comparing the cfDNA fragmentome profiles of the subject to reference cfDNA fragmentome profiles from healthy subjects to detect and differentially diagnose between ovarian cancer and ovarian cancer subtypes; (iii) assaying the subject’s biological sample to detect and quantify at least one biomarker; (iv) comparing the subject’s cfDNA fragmentation profile and protein biomarkers with normal reference non-cancer subjects; and, treating the subject diagnosed with ovarian cancer, ovarian cancer subtypes or benign masses with a cancer specific therapy.

34. The method of claim 33, wherein the one or more cfDNA the one or more cfDNA fragment characteristics comprises calculating ratios of small to large cfDNA fragments, sequences of one or more cfDNAs, median cfDNA fragment sizes, fragment size distribution, mutant allele frequencies, fragment length, fragment size distribution, fragment end motifs, preferred end coordinates, breakpoint motifs, methylation frequencies, concentrations of cfDNA or combinations thereof.

35. The method of claim 35, wherein a small cfDNA fragment comprises about 80 base pairs (bp) to about 150 bp and a large cfDNA fragment comprises about 151 bp to about 300 bp.

36. The method of any of claims 34 to 35, wherein the cfDNA fragmentome profile comprises the sequence coverage of small cfDNA fragments in windows across the genome.

37. The method of any of claims 34 to 35, wherein the cfDNA fragmentome profile comprises the sequence coverage of cfDNA fragments in windows across the genome.

38. The method of any of claims 34 to 35, wherein the cfDNA fragmentome profile comprises the sequence coverage of small and large cfDNA fragments in windows across the genome.

39. The method of any of claims 33 to 38, wherein the cfDNA fragmentome profiles in subjects identified and differentially diagnosed early between ovarian cancer and ovarian cancer subtypes, are altered across the genome.

40. The method of any of claims 33 to 39, wherein the cfDNA fragmentome profiles in subjects identified and differentially diagnosed early between ovarian cancer and ovarian cancer subtypes, have greater heterogeneity across the genome as compared to healthy subjects.

41. The method of claim 33, wherein the ovarian cancer subtypes comprise high-grade serous (HGSOC), low-grade serous (LGSOC), clear cell, mucinous, or endometroid ovarian cancers.

42. The method of any one of claims 36-38, wherein the windows each comprise about 5 million base pairs.

43. The method of any one of claims 36-38, wherein a cfDNA fragmentation profile is determined within each window.

44. The method of any one of claims 33-43, wherein the cfDNA fragmentation profile comprises the sequence coverage of small cfDNA fragments in windows across the genome.

45. The method of any one of claims 33-43, wherein the cfDNA fragmentation profile comprises the sequence coverage of large cfDNA fragments in windows across the genome.

46. The method of any one of claims 33-43, wherein the cfDNA fragmentation profile comprises the sequence coverage of small and large cfDNA fragments in windows across the genome.

47. The method of any one of claims 33-43, wherein the cfDNA fragmentation profile is over the whole genome.

48. The method of any one of claims 33-43, wherein the cfDNA fragmentation profile is over a subgenomic interval.

49. The method of any one of claims 33-48, further comprising assaying for chromosomal gains and losses in the subject’s genome as compared to a normal reference genome.

50. The method of any one of claims 33-49, further comprising a machine learning model wherein the model incorporates genome-wide fragmentation profiles, chromosomal arm-level changes, and the concentrations of ovarian cancer biomarkers.

51. The method of claim 50, wherein the ovarian cancer biomarkers comprise Carbohydrate Antigen 125 (CA-125), Osteopontin (OPN), Kallikreins (KLKs), Bikunin, Human Epididymis Protein 4 (HE4), Vascular Endothelial Growth Factor (VEGF), Prostasin (PSN), Creatine Kinase B (CKB), Mesothelin, Apolipoprotein A-I (apoA-I), Transthyretin (TTR), Transferrin or combinations thereof.

52. The method of claim 51, wherein the ovarian cancer biomarkers comprise Carbohydrate Antigen 125 (CA-125), Human Epididymis Protein 4 (HE4) or the combination thereof.

53. The method of claim 50, wherein the model generates a DELFI protein (DELFI-Pro) score.

54. The method of claim 53, wherein the DELFI-Pro score is differentially diagnostic of ovarian cancer and ovarian cancer subtypes.

55. The method of claim 23, wherein the DELFI-Pro score is diagnostic of the stage of ovarian cancer and ovarian cancer subtypes.

56. The method of claim 53, wherein the DELFI-Pro score is diagnostic of an adnexal mass.

57. The method of any one of claims 53-56, wherein subjects with low median DELFI-Pro scores are ovarian cancer and ovarian cancer subtype free or do not have an adnexal mass.

58. The method of claim 57, wherein a low median DELFI-Pro score is in a range of about 0.0001 to about 0.

1.

59. The method of any one of claims 53-56, wherein subjects with high median DELFI-Pro scores are diagnosed as having ovarian cancer, or an ovarian cancer subtype, or an adnexal mass.

60. The method of claim 59, wherein ranges of high median DELFI-Pro scores are diagnostic of stages of ovarian cancer and ovarian cancer subtypes.

61. The method of claims 59-60, wherein the higher median DELFI-Pro scores comprise scores greater than about 0.7.

Citation Information

Patent Citations

  • Disease Detection in Liquid Biopsies

    US20230042332A1

  • Detection of lung cancer using cell-free DNA fragmentation

    WO2022140386A1