Systems and methods for multiple biomarker analysis in cancer
The integration of fragmentomics profiling with gene expression and transcription factor analysis in cell-free nucleic acids provides accurate cancer detection and prognosis, overcoming the limitations of invasive biopsies and improving treatment planning.
Patent Information
- Application Number
- PCT/US2025/023176
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-08
- Filing Date
- 2025-04-04
- Publication Date
- 2025-10-09
AI Technical Summary
Current cancer detection methods, particularly tissue biopsies, are invasive, costly, and may not be feasible for all patients, while non-invasive methods lack the accuracy needed for effective treatment planning and prognosis.
A system and method utilizing fragmentomics profiling of cell-free nucleic acids, combined with gene expression and transcription factor activation status analysis, to generate disease forecast characteristics, including cancer type, subtype, prognosis, and treatment response, through nucleosome profiling and multiple biomarker analysis.
Enhances cancer detection accuracy and prognosis by integrating fragmentomics, gene expression, and transcription factor data, allowing for precise treatment recommendations and improved patient outcomes.
Smart Images

Figure US2025023176_09102025_PF_FP_ABST
Abstract
Description
WSGR Docket No. 59987-717.601 SYSTEMS AND METHODS FOR MULTIPLE BIOMARKER ANALYSIS IN CANCER CROSS-REFERENCE
[0001] This application claims the benefit of U.S. Application No. 63 / 575,321, filed April 5,2024, and U.S. Application No. 63 / 631,131, filed April 8, 2024, each of which is incorporated by reference herein in its entirety. BACKGROUND
[0002] Cancer is a leading cause of deaths worldwide. Detection of cancer in individuals maybe critical for providing treatment and improving patient outcomes. Cancer may be caused by genetic aberration which may lead to unregulated growth of calls. Detection of the genetic aberrations may be important for the detection of cancer. Sequencing of nucleic acids in a sample from a patient may be used to detect genetic aberrations. SUMMARY
[0003] Provided herein are systems and methods for detection of the presence or absence ofcancer in a subject. The systems and methods provided herein comprises assaying polynucleotides to identify biomarkers of cancers in a subject. Detection of a type of cancer or the specific biomarkers for a given cancer may allow an effective treatment to be provided to an individual and may result in improved outcomes. For multiple types of cancer, the particular biomarkers that indicate a particular cancer type (or subtype) may be used to identify a prognosis for an individual suffering from the cancer. In order to provide accurate detection and prognosis for a cancer, multiple markers or analytes may be examined. By analyzing an increased number of analytes, and sets of biomarkers from the analytes), the detection of a cancer (or cancer parameter) may be improved and may allow for the recommendation of an effective treatment, and may also allow for the prognosis to be more accurate. For example, analysis of methylation, fragmentomics, transcriptomics, and other data on various biomarkers may allow for the determination of a type of cancer, a prognosis for a subject, or other various parameters of disease. The use of multiple analytes and biomarkers may allow for more accurate forecasting of disease progression or allow for improved (or more appropriate) treatment regimens and improved patient outcomes.
[0004] In an aspect, the present disclosure provides a method of determining one or moredisease forecast characteristics of a subject having or suspected of having cancer, the method comprising: (a) generating one or more fragmentomics profiles of cell free nucleic acids obtained or derived from the subject, wherein a fragmentomics profile is generated relative to a selectedWSGR Docket No. 59987-717.601 biomarker; (b) determining an expression of one or more genes of the subject, or determining an activation status of one or more transcription factors; (c) based at least in part on (i) the one or more fragmentomics profiles and (ii) the expression of one or more genes or the activation status of one or more transcription factors, determining one or more disease forecast characteristics in said subject.
[0005] In some embodiments, the one or more fragmentomics profile is a nucleosome profile.In some embodiments, the fragmentomics profile is generated based at least in part on distance of one or more selected subsets of the cell free nucleic acids to the selected biomarker In some embodiments, the distance comprises the number of base pairs between each of the one or more selected subsets of the cell free nucleic acids and the selected biomarker. In some embodiments, the fragmentomics profile is further generated based at least in part on the coverage of the one or more selected subsets of the cell free nucleic acids. In some embodiments, the nucleosome profile is further generated based at least in part on the coverage of the one or more selected subsets of cell free nucleic acids associated with the distance of the one or more selected subsets of the cell free nucleic acids from the selected biomarker. In some embodiments, the one or more selected subsets of the cell free nucleic acids comprise fragments of the cell free nucleic acids. In some embodiments, the method further comprises determining a fragmentomics profiling abnormality score. In some embodiments, the determining a fragmentomics profiling abnormality score comprises relating the one or more fragmentomics profiles to one or more reference fragmentomics profiles. In some embodiments, the determining the fragmentomics profiling abnormality score further comprises comparing the coverage of the one or more selected subsets of cell free nucleic acids to one or more coverage values of the one or more reference fragmentomics profiles. In some embodiments, the determining the fragmentomics profiling abnormality score further comprises determining a Z-score of the coverage of the one or more selected subsets of the cell free nucleic acids. In some embodiments, determining the fragmentomics profiling abnormality score further comprises mapping the coverage Z-score of the one or more selected subsets of the cell free nucleic acids to a coverage Z-score of each of the one or more reference fragmentomics profiles. In some embodiments, (c) comprises determining one or more disease forecast characteristics in said subject, in part by processing said fragmentomics profiling abnormality score. In some embodiments, generating the disease forecast further comprises mapping the nucleosome profiling abnormality score to the one or more disease forecast characteristics. In some embodiments, the one or more reference nucleosome profiles are associated with one or more disease forecast characteristics. In some embodiments, the one or more reference nucleosome profiles are associated with a likelihood of occurrence of one or more disease forecast characteristics. In some embodiments, one or more ofWSGR Docket No. 59987-717.601 the reference nucleosome profiles comprise data from a sample not having the disease (normal). In some embodiments, the selected biomarker comprises androgen receptor binding sites. In some embodiments, the one or more disease forecast characteristics comprise a cancer prognosis relating to the subject. In some embodiments, the one or more disease forecast characteristics comprise one or more of: an estimated survival time of the subject without a treatment intervention, an estimated survival time of the subject with a treatment intervention, determination of a type of the cancer, determination of a subtype of the cancer, determination of one or more clinical outcomes, or predicted treatment response of the subject to one or more treatments, or any combination thereof. In some embodiments, the determining an expression of one or more genes comprises determining one or more expression levels of the one or more genes.
[0006] In some embodiments, (b) comprises assaying nucleic acids derived from the subject.In some embodiments, the nucleic acids derived from the subject are RNA molecules. In some embodiments, the RNA molecules are cell-free RNA molecules. In some embodiments, assaying nucleic acids derived from subject comprises nucleic acid sequencing. In some embodiments, the nucleic acid sequencing is an RNA sequencing assay. In some embodiments, the nucleic acid sequencing is a whole transcriptome sequencing assay or a targeted sequencing assay. In some embodiments, the biological sample comprises cell-free deoxyribonucleic acid (cfDNA) molecules. In some embodiments, the biological sample comprises one or more of: a plasma sample, a serum sample, a red blood cell sample, a urine sample, urine cell pellet sample, a saliva sample, pleural fluid sample, peritoneal fluid sample, amniotic fluid sample, cerebrospinal fluid sample, lymphatic fluid sample, sweat sample, tear sample, semen sample, or any derivative thereof, and any combination thereof. In some embodiments, the biological sample comprises the plasma sample. In some embodiments, the biological sample comprises the urine sample or urine cell pellet sample. In some embodiments, the cfDNA molecules are obtained or derived from a single biological sample of the subject. In some embodiments, the cfDNA molecules are obtained or derived from different biological samples of the subject. In some embodiments, the biological sample is obtained or derived from the subject using an ethylenediaminetetraacetic acid (EDTA) collection tube, a cell-free RNA collection tube, or a cell-free deoxyribonucleic acid (DNA) collection tube, other blood collection tube, and CTC collection tubes. In some embodiments, the method further comprises assaying the biological sample to generate the one or more fragmentomics profile. In some embodiments, assaying the biological sample comprises subjecting said biological sample to conditions that are sufficient to isolate, enrich, or extract the cfDNA molecules. In some embodiments, the method further comprising fractionating a whole blood sample of the subject to obtain the cfDNA molecules.WSGR Docket No. 59987-717.601
[0007] In some embodiments, assaying the biological sample further comprises assaying thecfDNA molecules using nucleic acid sequencing to produce nucleic acid sequencing reads. In some embodiments, the nucleic acid sequencing further comprises DNA sequencing. In some embodiments, the DNA sequencing comprises one or more of: next-generation sequencing, whole genome sequencing, low-pass sequencing, targeted sequencing, whole exome sequencing, methylation-aware sequencing, or bisulfite sequencing, or a combination thereof. In some embodiments, DNA sequencing comprises low-pass whole genome sequencing. In some embodiments, the DNA sequencing comprises whole exome sequencing. In some embodiments, the DNA sequencing further comprises nucleic acid amplification. In some embodiments, the nucleic acid amplification comprises polymerase chain reaction (PCR) or isothermal amplification. In some embodiments, at least one of the cfDNA molecules are assayed using a polymerase chain reaction (PCR) assay, microarray, or a isothermal amplification. In some embodiments, the type of the cancer of the subject comprises one or more of: lung cancer, brain cancer, spinal cancer, colorectal cancer, melanoma, bladder cancer, non-Hodgkin lymphoma, kidney cancer, endometrial cancer, leukemia, pancreatic cancer, thyroid cancer, or liver cancer, or any combination thereof. In some embodiments, the type of the cancer of the subject comprises prostate cancer. In some embodiments, the subtype of the prostate cancer comprises one or more of: hormone sensitive prostate cancer (HSPC), castration-resistant prostate cancer (CRPC), androgen receptor-dependent prostate cancer (ARPC), metastatic prostate cancer, metastatic castration-resistant prostate cancer (mCRPC), neuroendocrine prostate cancer (NEPC), or any combination thereof. In some embodiments, the subject is asymptomatic for the cancer. In some embodiments, the one or more transcription factors is an androgen receptor. In some embodiments, the determining the activation status of the one or more transcription factors comprises processing the one or more fragmentomics profiles.
[0008] In some embodiments, the selected biomarker comprises a selected binding site. Insome embodiments, the selected biomarker comprises a transcription factor binding site. In some embodiments, the selected biomarker comprises one or more of: androgen receptor binding sites (ARBS). In some embodiments, the method further comprises mapping DNA methylation patterns of the biological sample. In some embodiments, the method further comprising associating the mapped DNA methylation patterns with the nucleosome profile data of the biological sample.
[0009] In some embodiments, the one or more selected subsets of the cell free nucleic acidsare selected based at least in part on DNA methylation data mapped to the cell free nucleic acids. In some embodiments, the method further comprises generating the disease forecast based at least in part on copy number variation data, or sequencing mutation data, or both, associated withWSGR Docket No. 59987-717.601 the sample. In some embodiments, the method further comprises determining a tumor fraction of the cell free nucleic acids. In some embodiments, the method further comprises determining a fragmentomics profiling abnormality score, comprising relating the fragmentomics profile of the subject sample to one or more reference fragmentomics profile and associating the tumor fraction with the fragmentomics profiling abnormality score.
[0010] In another aspect, the present disclosure provides a method for estimating a responseof a subject having cancer to one or more treatments, comprising: (a) generating a fragmentomics profile relating to cfDNA molecules derived from a biological sample obtained or derived from the subject; (b) determining expression of one or more genes, or determining an activation status of one or more transcription factors of the subject; (c) determining one or more characteristics of the cancer of the subject based at least in part on the fragmentomics profile of the subject; and (d) generating a treatment response determination for the subject based at least in part on the one or more characteristics of the cancer of the subject.
[0011] In some embodiments, the treatment response determination comprises one or moreof: a treatment plan, a value representing likelihood of one or more treatment responses, a binary treatment response indicator, a probability value for each treatment response, or any combination thereof. In some embodiments, the one or more characteristics of the cancer of the subject comprises a cancer type, a cancer subtype, an estimate of cancer progression, a prognosis of the subject without treatment intervention, a prognosis of the subject with treatment intervention, or any combination thereof.
[0012] In another aspect, the present disclosure provides A system for determining a diseaseforecast of a subject having cancer, the system comprising: (a) a memory; and (b) one or more processors configured to execute machine-readable instructions which, when executed, cause the one or more processors to perform a method comprising: (c) generating a fragmentomics profile of a biological sample obtained or derived from the subject, (d) mapping the fragmentomics profile to one or more reference fragmentomics profiles, (e) determining expression of one or more genes or determining an activation status of one or more transcription factor of the subject; (f) determining one or more disease forecast characteristics of the biological sample based at least in part on the mapping and the expression one or more genes or activation status of transcription factors, and (g) generating the disease forecast based at least in part on the one or more disease forecast characteristics.
[0013] In some embodiments, the one or more disease forecast characteristics comprise oneor more of: a type of the cancer, a subtype of the cancer, a prognosis of the subject, an estimated survival time of the subject without treatment intervention, an estimated survival time of theWSGR Docket No. 59987-717.601 subject with treatment intervention, an estimation of the subject’s response to one or more treatments, or any combination thereof.
[0014] In another aspect, the present disclosure provide a method for identifying presence oran absence of cancer in a subject, comprising: (a) assaying nucleic acid molecules from a firstbiological sample obtained or derived from said subject at a first time point; (b) detecting a set of biomarkers from said nucleic acid molecules based at least in part on said assaying of (a), wherein said set of biomarkers comprise differentially expressed markers or variants; (c) obtaining a plurality of probe nucleic acids that are customized for said subject, wherein said probe nucleic acids comprises sequences of at least a subset of said set of biomarkers; (d) using said plurality of probe nucleic acids, sequencing cell free nucleic acids (cfNA) from a second biological sample obtained or derived from said subject at a second time point to detect the presence or absence of said subset of said set of biomarkers, wherein said second biological sample comprises a cerebrospinal fluid sample; (e) computer processing said subset of said set of biomarkers to detect said cancer in said subject. In some embodiments, the first biological sample is selected from the group consisting of: a cell-free deoxyribonucleic acid (cfDNA) sample, a cell-free ribonucleic acid (cfRNA) sample, a plasma sample, a serum sample, a buffy coat sample, a peripheral blood mononuclear cell (PBMC) sample, a red blood cell sample, a urine sample, a urine cell pellet sample, a saliva sample, tissue biopsy, pleural fluid sample, peritoneal fluid sample, amniotic fluid sample, cerebrospinal fluid sample, lymphatic fluid sample, sweat sample, tear sample, semen sample, or any derivative thereof, and any combination thereof. In some embodiments, the first biological sample comprises said plasma sample. In some embodiments, the first biological sample comprises said urine sample. In some embodiments, the first biological sample comprises said tumor tissue sample. In some embodiments, the first or second biological sample is obtained or derived from said subject using an ethylenediaminetetraacetic acid (EDTA) collection tube, a cell-free RNA collection tube, or a cell-free deoxyribonucleic acid (DNA) collection tube, other blood collection tube, and CTC collection tubes. In some embodiments, the cfNA molecules comprise cell-free DNA (cfDNA) molecules. In some embodiments, (a) comprises subjecting said first or second biological sample to conditions that are sufficient to isolate, enrich, or extract said nucleic acid molecules or cfNA molecules. In some embodiments, the method further comprises fractionating said first biological sample of said subject to obtain said nucleic acid molecules, wherein said first biological sample is a whole blood sample. In some embodiments, at least one of said nucleic acid molecules are assayed using sequencing to produce nucleic acid sequencing reads. In some embodiments, the sequencing comprises whole exome sequencing. In some embodiments, the method further comprises filtering at least a subset of said nucleic acid sequencing reads based on a qualityWSGR Docket No. 59987-717.601 score. In some embodiments, the method further comprises performing error correction on said nucleic acid sequencing reads using sample barcodes or molecular barcodes attached to at least one of said DNA molecules. In some embodiments, the method further comprises performing at least one of single-stranded consensus calling and double-stranded consensus calling on said nucleic acid sequencing reads, thereby suppressing sequencing and PCR errors in said nucleic acid sequencing reads. In some embodiments, the sequencing of (d) is performed at a depth of at least 100x. In some embodiments, the sequencing of (d) is performed at a depth of at least 1,000x. In some embodiments, the sequencing of (d) is performed at a depth of at least 10,000x. In some embodiments, the sequencing of (d) is performed at a depth of at least 100,000x. In some embodiments, the assaying of (a) or sequencing of (d) comprises nucleic acid amplification. In some embodiments, the nucleic acid amplification comprises polymerase chain reaction (PCR) or isothermal amplification. In some embodiments, the cancer is a brain cancer or spine cancer. In some embodiments, the cancer comprises a glioma. In some embodiments, the subject is asymptomatic for said cancer. In some embodiments, the method comprises detecting said presence or absence of cancer in said subject at an accuracy of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some embodiments, the method comprises detecting said presence or absence of cancer in said subject at a sensitivity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some embodiments, the method comprises detecting said presence or absence of cancer in said subject at a specificity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some embodiments, the method comprises detecting said presence or absence of cancer in said subject at a positive predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some embodiments, the method comprises detecting said presence or absence of cancer in said subject in said subject at a negative predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some embodiments, the first biological sample is obtained or derived from said subject prior to said subject receiving a therapy for said cancer. In some embodiments, the biological sample is obtained or derived from said subject during a therapy for said cancer. In some embodiments, the biological sample is obtained or derived from said subject after receiving a therapy for said cancer. In some embodiments, the therapy is selected from the group consisting of: surgical resection, chemotherapy, radiotherapy,WSGR Docket No. 59987-717.601 immunotherapy, cell therapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and a combination thereof. In some embodiments, the method further comprises identifying a clinical intervention for said subject based at least in part on said detected presence or said absence of said cancer. In some embodiments, the clinical intervention is selected from a plurality of clinical interventions. In some embodiments, the clinical intervention is selected from the group consisting of: surgical resection, chemotherapy, radiotherapy, immunotherapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and a combination thereof. In some embodiments, the method further comprises administering said clinical intervention to said subject. In some embodiments, the plurality of probes comprise nucleic acid primers. In some embodiments, the plurality of probes comprise nucleic acid capture probes. In some embodiments, the plurality of probes have sequence complementarity with at least a portion of nucleic acid sequences of said set of biomarkers. In some embodiments, the plurality of probes comprise at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 different probes. In some embodiments, (d) further comprises sequencing using a fixed plurality of probes wherein the probes of the fixed plurality of probes comprises probes that do not comprise sequences of said subset of said set of biomarkers In some embodiments, the method further comprises determining a likelihood of said determination of said presence or said absence of said cancer in said subject. In some embodiments, the method further comprises monitoring said presence or said absence of said cancer in said subject, wherein said monitoring comprises assessing said presence or said absence of said cancer in said subject at each of a plurality of time points. In some embodiments, a difference in said assessment of said presence or said absence of said cancer in said subject among said plurality of time points is indicative of one or more clinical indications selected from the group consisting of: (i) a diagnosis of said cancer, (ii) a prognosis of said cancer, and (iii) an efficacy or non-efficacy of a course of treatment for treating said cancer of said subject. In some embodiments, the prognosis comprises an expected progression-free survival (PFS) or overall survival (OS). In some embodiments, the set of biomarkers from said cfNA molecules comprise tumor-associated alterations selected from the group consisting of: single nucleotide variants (SNVs), insertions or deletions (indels), and rearrangements. In some embodiments, the method further comprises determining, among said set of biomarkers, a mutant allele frequency of a set of somatic mutations. In some embodiments, the method further comprises determining a circulating tumor DNA (ctDNA) fraction of said cancer of said subject based at least in part on said set of mutant allele frequencies. In some embodiments, the further comprises determining a tumor mutational burden (TMB) of said cancer of said subject. In some embodiments, the method further comprises determining anWSGR Docket No. 59987-717.601 abnormality score of said cancer of said subject based at least in part on said set of mutant allele frequencies.
[0015] In another aspect the present disclosure provides a method for detecting a presence oran absence of a glioma in a subject, comprising: (a) assaying cell-free deoxyribonucleic acid (cfDNA) molecules from a biological sample obtained or derived from said subject, wherein the biological sample comprises cerebrospinal fluid sample; (b) detecting a set of biomarkers from said cfDNA molecules wherein said set of biomarkers comprise differentially expressed markers or variants; (c) computer processing said set of biomarkers to detect said presence or said absence of said cancer in said subject.
[0016] Another aspect of the present disclosure provides a non-transitory computer readablemedium comprising machine executable code that, upon execution by one or more computer processors, implements any of the methods above or elsewhere herein.
[0017] Another aspect of the present disclosure provides a system comprising one or morecomputer processors and computer memory coupled thereto. The computer memory comprises machine executable code that, upon execution by the one or more computer processors, implements any of the methods above or elsewhere herein.
[0018] Additional aspects and advantages of the present disclosure will become readily apparentto those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive. INCORPORATION BY REFERENCE
[0019] All publications, patents, and patent applications mentioned in this specification areherein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The novel features of the invention are set forth with particularity in the appended claims.A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, inWSGR Docket No. 59987-717.601 which the principles of the invention are utilized, and the accompanying drawings (also “figure” and “FIG.” herein), of which:
[0021] FIG. 1A-2E shows data relating to fragmentomics at AR binding sites.
[0022] FIG. 2A-2D shows data relating to cfDNA fragmentomics and chromatin level.
[0023] FIG. 3A-3D show data relating to gene expression and cfDNA fragmentomics
[0024] FIG 4 show these genome-wide CNV profiles of two mCRPC patients
[0025] FIG. 5A-D show data obtained from sequencing assays for a first individual.
[0026] FIG. 6A-D show data obtained from sequencing assays for a second individual.
[0027] FIG. 7A-C show data obtained from sequencing assays for a third individual.
[0028] FIG. 8A-D show data obtained from sequencing assays for a fourth individual.
[0029] FIG. 9A-D show data obtained from sequencing assays for a fifth individual.
[0030] FIG. 10 shows a computer system that is programmed or otherwise configured toimplement methods provided herein. DETAILED DESCRIPTION
[0031] While various embodiments of the invention have been shown and described herein, itwill be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.
[0032] Terms and Definitions
[0033] As used herein, the singular forms “a,” “an,” and “the” include plural references unlessthe context clearly dictates otherwise. Any reference to “or” herein is intended to encompass “and / or” unless otherwise stated.
[0034] As used herein, the phrases “at least one,” “one or more,” and “and / or” are open-endedexpressions that are both conjunctive and disjunctive in operation. For example, each of the expressions “at least one of A, B and C,” “at least one of A, B, or C,” “one or more of A, B, and C”, “one or more of A, B, or C” and “A, B, and / or C” means A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B and C together. As used herein, the phrase “at most three” can mean less than one, one, two, or three.
[0035] Reference throughout this specification to “some embodiments,” “further embodiments,”or “a particular embodiment,” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase “in some embodiments,” or “in further embodiments,” or “in a particular embodiment” in various places throughout this specification are not necessarily allWSGR Docket No. 59987-717.601 referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0036] The terms "subject," "individual," and "patient" may be used interchangeably and refer tohumans, as well as non-human mammals (e.g., non-human primates, canines, equines, felines, porcines, bovines, ungulates, lagomorphs, rodents, and the like). In various embodiments, the subject can be a human (e.g., adult male, adult female, adolescent male, adolescent female, male child, female child) under the care of a physician or other health worker in a hospital, as an outpatient, or other clinical context.
[0037] As used herein, “treatment” or “treating” refers to an approach for obtaining beneficial ordesired results with respect to a disease, disorder, or medical condition including, but not limited to, a therapeutic benefit and / or a prophylactic benefit. In certain embodiments, treatment or treating involves administering a therapeutic to a subject. A therapeutic benefit may include the eradication or amelioration of the underlying disorder being treated. Also, a therapeutic benefit may be achieved with the eradication or amelioration of one or more of the physiological symptoms associated with the underlying disorder, such as observing an improvement in the subject, notwithstanding that the subject may still be afflicted with the underlying disorder.
[0038] Next-generation sequencing has revolutionized cancer genomic research in the last 10+years. NGS technologies commercially available for guiding treatment plans for cancer patients that have been FDA cleared or approved for use in processing DNA from patient tissue or blood samples. A next-generation sequencing (NGS) assay on cfDNA can enable an accurate detection of genomic alterations, including single nucleotide variant (SNV), insertion and deletion (Indel), Copy number variation (CNV), and DNA re-arrangement.
[0039] In addition to staging and grading of a patient’s tumor, tissue biopsy often represent thegold standard in guiding treatment for cancer patients. However, depending on the location of the tumor or the condition of the patient, tumor biopsies can be painful, and the patient may incur risk of complication, whereby medical treatment can become costly. In some cases, tissue biopsy may not be feasible. Less invasive sampling methods such as sampling using plasma, urine, or cerebrospinal fluid, or other more accessible bodily fluid may be advantageous for a patient, and may result in accurate diagnoses without requiring invasive biopsies.
[0040] Analytes that can be used for tumor diagnosis from urine, plasma, or cerebrospinal fluid(CSF) can include cfDNA, cfRNA, non-coding-RNA, exfoliated tumor cells and proteins. During tumor destruction therapies or during apoptotic and necrotic processes, both healthy and diseased cells may release cfDNA fragments that are typically 100-200 base pairs in length. In patients without disease, the phagocytic cells englobe cellular debris and necrotic cells, and thus there are very low levels of cfDNA. In the case of patients with disease, phagocytosis is compromised,WSGR Docket No. 59987-717.601 DNA digestion is minimal, and the DNA fragments have a random dimension that could exceed 10,000 base pairs. Therefore, the cfDNA level in a patient with disease is often elevated.
[0041] For example, cfDNA (e.g., urine cfDNA, plasma cfDNA, or CSF cfDNA) can beextracted from a bodily fluid from a subject and cfDNA can be subjected to various reactions to allow for sequencing of the cfDNA. Library construction of cfDNA can comprise amplification, ligation of adapter or additional sequences and / or labeling with barcodes to generate a sequencing library. Additionally, a cfDNA or cfDNA library can be subjected to enrichment using capture probes or amplification primers to enrich for specific sequences of interest from the cfDNA. The library can then be subjected to sequencing reactions to generate sequencing data.
[0042] Similarly cfRNA or other RNA can also be extracted from a bodily fluid and assayedsimilarly. For example, libraries may be constructed from RNA. For example, cDNA derived from RNA may be generated and processed to form a sequencing library. RNA may be assayed by performing a whole transcriptome sequencing assay. The RNA may be assayed to determine the expression of one or more genes (e.g., an expression level of one or more genes). The expression level may be a relative level (e.g., as compared to a reference sample or a reference gene, such as a housekeeping gene) or an absolute level of expression.
[0043] Provided herein are systems and methods for determining one or more diseasecharacteristics or parameters of subject having or suspected of having cancer. For example, the systems and methods may detect the presence or absence of cancer in a subject. For example, the systems and methods may determine one or more disease forecast characteristics in a subject. The systems and methods provided herein comprises assaying polynucleotides to identify biomarkers of cancers in a subject. The biomarkers may be processed in order to identify the presence or absence of cancer. The methods described herein may process analytes to determine a presence or absence of cancer. The analytes may comprise cfDNA or other analytes that can be provided via non-invasive methods. By analyzing analytes obtained by non-invasive methods, the methods may allow for improved or similar detection or determination of a prognosis as compared to methods that use tumor or tissue biopsy.
[0044] The systems and methods of the present disclosure may use multiple different assays incombination to determine characteristics of a disease and may allow for determination of disease forecast characteristics. The use of multiple different assay may improve the accuracy as compared to the use of fewer assays. By combining the information gathered by the multiple assays, a more appropriate or effective treatment may be determined for a given individual. For example, a method may comprises generating a nucleosome profile and performing a whole transcriptome sequencing assay. The combined data on the nucleosome profile and the gene expression may allow for a more accurate (or sensitive, specific) result than the use of aWSGR Docket No. 59987-717.601 nucleosome profile without gene expression information. The combined data may synergistically improve the resulting forecast or determination.
[0045] The systems and methods of the present disclosure may use different sequencing assaysto observe different markers in the same individual. For example, the assays may detect epigenetic markers (e.g., methylation, nucleosome binding) or other changes to nucleic acids are unrelated to the primary sequence of a nucleic acid. In conjunction with assays that detect epigenetic markers, the assays mays detect changes to the copy number of a sequence (e.g. copy number gain or loss) or may detect changes to the sequence of genes (e.g., deletions, insertions, inversions, substitutions). The systems and methods can synthesize (e.g., via the use of a machine learning algorithm) the data generated from the multiple assay to make a determination regarding a disease or disease state (e.g., the presence of a cancer, the presence of a subtype of cancer, a prognosis of a disease).
[0046] In an aspect, disclosed herein is a method for determining a disease forecast of a subjecthaving cancer. In some embodiments, the method comprises generating one or more fragmentomics profiles. In some embodiments, the method can comprise generating a nucleosome profile relating to a biological sample obtained or derived from the subject. In some cases, the nucleosome profile can be generated relative to a selected biomarker. In some embodiments, the method can further comprise determining a nucleosome profiling abnormality score. In some cases, determining a nucleosome profiling abnormality score can comprise relating the nucleosome profile of the subject sample to one or more reference nucleosome profiles. In some embodiments, the method can further comprise generating the disease forecast based at least in part on the nucleosome profiling abnormality score. In some cases, the disease forecast can comprise one or more disease forecast characteristics.
[0047] In some embodiments, the one or more disease forecast characteristics can comprise acancer prognosis relating to the subject. The prognosis may comprise expected progression-free survival (PFS), overall survival (OS), or other metrics relating the severity or survivability of a cancer. In some embodiments, the disease forecast characteristics can comprise one or more of: an estimated survival time of the subject without a treatment intervention, an estimated survival time of the subject with a treatment intervention, determination of a type of the cancer, determination of a subtype of the cancer, determination of one or more clinical outcomes, or predicted treatment response of the subject to one or more treatments, or any combination thereof.
[0048] As describe in this disclosure, sequencing of nucleic acids may be performed andsequencing reads can be generated from the sequencing assays. Genetic alterations, such as single nucleotide variations (SNVs), indels, DNA rearrangements, and copy number variationsWSGR Docket No. 59987-717.601 (CNVs) can be identified by bioinformatic analysis of the sequencing data. Generally, a bioinformatic pipeline can utilize the raw sequencing data (e.g., BCL files) and output mutational calls. The pipeline can perform various tasks to analyze the sequencing data such as adapter trimming, barcode checking, or error correction. Cleaned paired files (e.g. FASTQ files) can be aligned to human reference genome using an alignment tool such as BWA alignment tool. Consensus sequences can then be derived by merging paired-end reads that originated from the same molecules as single strand fragments. Single strand fragments from the same double strand DNA molecules can be further merged as double stranded. These processes can allow for sequencing and PCR errors to be corrected.
[0049] The subject may be a suspected of a suffering from a cancer. The cancer may be specificor originating from an organ or other area of the subject. For example, the cancer may be breast cancer, lung cancer, prostate cancer, colorectal cancer, melanoma, bladder cancer, non-Hodgkin lymphoma, kidney cancer, endometrial cancer, leukemia, pancreatic cancer, thyroid cancer, and liver cancer, and any combination thereof. The cancer may hormone sensitive prostate cancer (HSPC), castration-resistant prostate cancer (CRPC), androgen receptor-dependent prostate cancer (ARPC), metastatic prostate cancer, metastatic castration-resistant prostate cancer (mCRPC), neuroendocrine prostate cancer (NEPC), or any combination thereof. The cancer may be a cancer of a tissue or cell of the genitourinary tract. For example, the cancer may be a bladder cancer, kidney cancer, or prostate cancer. The cancer may comprise biomarkers that are specific to a particular cancer. The specific biomarkers may indicate a presence of a particular cancer. For example, biomarker may indicate that a castrate-resistant prostate cancer is present. The identification of the presence of a type of cancer may allow the determination of a treatment option or recommendation.
[0050] In some cases, the subject may be asymptomatic for cancer. For example, the cancer maynot exhibit any symptoms and the subject may be unaware of the presence of cancer. The methods described herein may allow a cancer to be identified at an earlier stage than otherwise. The identification of the presence of the cancer at an earlier stage may allow a treatment option or recommendation to be determined at an earlier stage and may allow the subject to have an improved prognosis.
[0051] The biological sample may comprise nucleic acids. The biological sample be a cell-freedeoxyribonucleic acid (cfDNA) sample or a cell-free ribonucleic acid (cfRNA) sample. The biological sample may comprise genomic DNA or germline DNA(gDNA). The nucleic acid may be a DNA (e.g. double-stranded DNA, single-stranded DNA, single-stranded DNA hairpins, cDNA, genomic DNA, germline DNA, circulating tumor DNA (ctDNA), cell-free DNA (cfDNA)), an RNA (e.g. cfRNA, mRNA, cRNA, miRNA, siRNA, miRNA, snoRNA, piRNA,WSGR Docket No. 59987-717.601 tiRNA, snRNA), or a DNA / RNA hybrids. The biological sample may be a derived from or contain a biological fluid. For example, the biological sample may be a plasma sample, a serum sample, a buffy coat sample, a peripheral blood mononuclear cell (PBMC) sample, a red blood cell sample, a urine sample, a saliva sample, or other body fluid sample. The biological sample may comprise or be a pleural fluid sample, peritoneal fluid sample, amniotic fluid sample, cerebrospinal fluid sample, lymphatic fluid sample, sweat sample, tear sample, semen sample, or any combination of biological fluid. The biological sample may comprise a urine sample.
[0052] The biological sample may be collected, obtained, or derived from the subject using acollection tube. The collection tube may be an ethylenediaminetetraacetic acid (EDTA) collection tube, a cell-free RNA collection tube, or a cell-free deoxyribonucleic acid (DNA) collection tube and CTC collection tubes, or other blood collection tube. The collection tube may comprise additional reagents for stabilizing the nucleic acid molecules or blood cells. The collection tube may allow the nucleic acid or blood cells to be stable such to minimize degradation of the biological sample prior to assaying. The additional reagents may comprise buffer salts or chelators.
[0053] The biological sample may be obtained or derived from a subject at various times. Thebiological sample may be obtained or derived from a subject prior to the subject receiving a therapy for cancer. The biological sample may be obtained or derived from a subject during receiving a therapy for cancer. The biological sample may be obtained or derived from a subject after receiving a therapy for cancer. The biological sample may be obtained or derived from said subject via a transurethral resection of bladder tumor. The biological sample may be obtained or derived from said subject after performing a transurethral resection of bladder tumor.
[0054] The biological sample may be collected over 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 15, 20,30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 or time points. The time points may occur over a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 or more hour period. The time points may occur over a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 or more day period. The time points may occur over a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 or more week period. The time points may occur over a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 or more month period. The time points may occur over a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 or more year period.
[0055] In various aspects as described herein, a clinical intervention or a therapy may beidentified at least in part based on the identification of the presences of cancer, or the presence ofWSGR Docket No. 59987-717.601 a parameter of cancer. The clinical intervention may be a plurality of clinical interventions. The clinical intervention may be selected from a plurality of clinical interventions. The clinical intervention may be a surgical resection, chemotherapy, radiotherapy, immunotherapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, or a combination thereof. In some cases, the clinical interventions may be administered to the subject. After administration of the clinical intervention, a sample may be obtained or derived from the subject such to monitor the cancer or cancer parameters. As such, the methods and systems disclosed herein may be performed iteratively such that monitoring of a cancer can be performed. Additionally, by performing the methods or systems iteratively, therapies or clinical interventions may be updated based on the results of the methods. The monitoring of the cancer may include an assessment as well as a difference in assessment from a previously generated assessment. The difference in an assessment of cancer in the subject among a plurality of time points (or samples) may be indicative of one or more clinical indications such as a diagnosis of the cancer, a prognosis of the cancer, or an efficacy or non-efficacy of a course of treatment for treating the cancer of the subject. The prognosis may comprise expected progression-free survival (PFS), overall survival (OS), or other metrics relating the severity or survivability of a cancer.
[0056] The biological samples may be subjected to additional reactions or conditions prior toassaying. For example, the biological sample may be subjected to conditions that are sufficient to isolate, enrich, or extract nucleic acids, such cfDNA molecules.
[0057] The methods disclosed herein may comprise conducting one or more enrichmentreactions on one or more nucleic acid molecules in a sample. The enrichment reactions may comprise contacting a sample with one or more beads or bead sets. The enrichment reactions may comprise one or more hybridization reactions. For example, the enrichment reactions may comprise contacting a sample with one or more capture probes or bait molecules that hybridize to a nucleic acid molecule of the biological sample. The enrichment reaction may comprise differential amplification of a set of nucleic acid molecules. The enrichment reaction may enrich for a plurality of genetic loci or sequences corresponding to genetic loci. The enrichment reactions may comprise the use of primers or probes that may complementarity to sequences (or sequences upstream or downstream) of a sequence that is to be enriched. For example, a capture probe may comprise sequence complementarity to a set of genomic loci and allow the enrichment of the genomic loci. The enrichments reactions may comprise a plurality of probes or primers. For example, a capture probe may comprise sequence complementarity to one or more genes. A plurality of probes may comprise 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, 230, 235, 240, 245, 250, 255, 260,WSGR Docket No. 59987-717.601 265, 270, 275, 280, 285, 290, 295, 300, 305, 310, 315, 320, 325, 330, 335, 340, 345, 350, 355, 360, 365, 370, 375, 380, 385, 390, 395, 400, 405, 410, 415, 420, 425, 430, 435, 440, 445, 450, 455, 460, 465, 470, 475, 480, 485, 490, 495, 500, 505, 510, 515, 520, 525, 530, 535, 540, 545, 550, 555, 560, 565, 570, 575, 580, 585, 590, 595, or 600 different probes.
[0058] The methods disclosed herein may comprise conducting one or more isolation orpurification reactions on one or more nucleic acid molecules in a sample. The isolation or purification reactions may comprise contacting a sample with one or more beads or bead sets. The isolation or purification reaction may comprise one or more hybridization reactions, enrichment reactions, amplification reactions, sequencing reactions, or a combination thereof. The isolation or purification reaction may comprise the use of one or more separators. The one or more separators may comprise a magnetic separator. The isolation or purification reaction may comprise separating bead bound nucleic acid molecules from bead free nucleic acid molecules. The isolation or purification reaction may comprise separating capture probe hybridized nucleic acid molecules from capture probe free nucleic acid molecules. The isolation reactions may comprises removing or separating a group of nucleic acid molecules from another group of nucleic acids.
[0059] The methods disclosed herein may comprise conduction extraction reactions on one ormore nucleic acids in a biological sample. The extraction reactions may lyse cells or disrupt nucleic acid interactions with the cell such that the nucleic acids may be isolated, purified, enriched, or subjected to other reactions.
[0060] The methods disclosed herein may comprise amplification or extension reactions. Theamplification reactions may comprise polymerase chain reaction. The amplification reaction may comprise PCR-based amplifications, non-PCR based amplifications, or a combination thereof. The one or more PCR-based amplifications may comprise PCR, qPCR, nested PCR, linear amplification, or a combination thereof. The one or more non-PCR based amplifications may comprise multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, circle-to-circle amplification or a combination thereof. The amplification reactions may comprise an isothermal amplification.
[0061] The method disclosed herein may comprise a barcoding reaction. A barcoding reactionmay comprise the additional of a barcode or tag to the nucleic acid. The barcode may be a molecular barcode or a sample barcode. For example, a barcode nucleic acid may comprise a barcode sequence which may be a degenerate n-mer. The sequence may be randomly generated or generated such to synthesize a specific barcode sequence. The barcode nucleic acid may be added to a sample such to label the nucleic acid molecules in the sample. The barcodes may beWSGR Docket No. 59987-717.601 specific to a sample. For example, a plurality of barcode nucleic acids may be added to a sample in which the barcode sequence is the same. Upon barcoding of the nucleic acids, those originating from a same sample may have a same barcode sequence, and may allow a nucleic acid to be identified as belonging to a particular or given sample. A molecular barcode may also be used such that each molecule (or a plurality of molecules) in a same volume have a different molecular barcode. This barcode may be subjected to amplification such that all amplicons derived from a molecule have the same barcode. In this way, molecules originating from a same molecule may be identified. The sequences reads may be processed based on the barcode sequences. For example, the processing may reduce errors or allow a molecule to be tracked. Barcode sequences may be appended or otherwise added or incorporated into a sequence by various reactions, for example an amplification, extension, or ligation reaction, and may be performed enzymatically using a nucleic acid polymerase or ligase. The ligation may be an overhang or blunt end ligation and the barcodes may comprise complementarity to nucleic acids to be barcoded. This complementarity may be a sequence derived from the sample from the subject or may be constant sequence generated via a reaction performed on the nucleic acids in the sample.
[0062] In some cases, the biological sample may comprise multiple components. For example,the biological sample may be a whole blood sample. The biological sample may be subjected to reactions such to separate or fractionate a biological sample. For example, a whole blood sample may be a fractionated and cell free nucleic acids may be obtained. The whole blood sample may be fractionated using centrifugation such that blood cells may be separated from the plasma (which may contain cell free nucleic acid). A sample may be subjected to multiple rounds of separation or fractionation.
[0063] In various aspects described throughout the disclosure, the nucleic acids may be subjectedto sequencing reactions. The sequencing the reactions may be used on DNA, RNA or other nucleic acid molecules. Example of a sequencing reaction that may be used include capillary sequencing, next generation sequencing, Sanger sequencing, sequencing by synthesis, single molecule nanopore sequencing, sequencing by ligation, sequencing by hybridization, sequencing by nanopore current restriction, or a combination thereof. Sequencing by synthesis may comprise reversible terminator sequencing, processive single molecule sequencing, sequential nucleotide flow sequencing, or a combination thereof. Sequential nucleotide flow sequencing may comprise pyrosequencing, pH-mediated sequencing, semiconductor sequencing or a combination thereof. The sequencing reactions may comprise whole genome sequencing, whole exome sequencing, low-pass whole genome sequencing, targeted sequencing, methylation-aware sequencing, enzymatic methylation sequencing, bisulfite methylation sequencing. The sequencing reactionWSGR Docket No. 59987-717.601 may be a transcriptome sequencing, mRNA-seq, totalRNA-seq, smallRNA-seq, exosome sequencing, or combinations thereof. Combinations of sequencing reactions may be used in the methods described elsewhere herein. For example, a sample may be subjected to whole genome sequencing and whole transcriptome sequencing. For example, the methods may comprise using a targeted sequencing and a transcriptome sequencing (e.g. whole transcriptome sequencing). The methods may use a combination of methylation aware sequencing, targeted sequencing, whole genome sequencing, and whole transcriptome sequencing. As the samples may comprise multiple types of nucleic acids (e.g. RNA and DNA), sequencing reactions specific to DNA or RNA may be used such to obtain sequence reads relating to the nucleic acid type.
[0064] The sequencing reactions can be performed at various sequencing depths. The sequencingdepths of a sequencing reaction may be selected or modulated. The sequencing reactions may comprise sequencing at a region a depth of at least 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 11x ,12x, 13x, 14x, 15x, 16x, 17x, 18x, 19x, 20x, 25x, 30x, 35x, 40x, 45x, 50x, 60x, 70x, 80x, 90x, 100x, 200x, 300x, 400x, 500x, 600x, 700x, 800x, 900x,1000x, 2000x, 3000x, 4000x, 5000x, 6000x, 7000x, 8000x, 9000x, 10,000x, 20,000x, 30,000x, 40,000x, 50,000x, 60,000x, 70,000x, 80,000x, 90,000, 100,000x, or more. The sequencing reactions may comprise sequencing a region at a depth of no more than 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 11x ,12x, 13x, 14x, 15x, 16x, 17x, 18x, 19x, 20x, 25x, 30x, 35x, 40x, 45x, 50x, 60x, 70x, 80x, 90x, 100x, 200x, 300x, 400x, 500x, 600x, 700x, 800x, 900x,1000x, 2000x, 3000x, 4000x, 5000x, 6000x, 7000x, 8000x, 9000x, 10,000x, 20,000x, 30,000x, 40,000x, 50,000x, 60,000x, 70,000x, 80,000x, 90,000, 100,000x, or less.
[0065] In various embodiments, a whole genome sequencing is used to sequence nucleic acids.As described in this disclosure, the sequencing may be performed at various depths. For example, the whole genome sequencing may be a low pass whole genome or a high depth sequencing. The low pass whole genome sequence may be performed at an average sequencing depth of at least 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, or more. The low pass whole genome sequence may be performed at an average sequencing depth of no more than 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, or less. The low pass whole genome sequencing may be performed at an average depth of between 1x and 2x. The low pass whole genome sequencing may be used to determine copy number variation, while maintaining a smaller sequencing footprint.
[0066] In various embodiments, a sequencing reaction may be performed using a set ofpersonalized or customized probes. The sequencing reaction using a set of personalized or customized probes may be a deep sequencing reaction or ultra-deep sequencing reaction. For example, the sequencing reaction using a set of personalized or customized probes may be performed at an sequencing depth of 50x, 60x, 70x, 80x, 90x, 100x, 200x, 300x, 400x, 500x,WSGR Docket No. 59987-717.601 600x, 700x, 800x, 900x,1000x, 2000x, 3000x, 4000x, 5000x, 6000x, 7000x, 8000x, 9000x, 10,000x, 20,000x, 30,000x, 40,000x, 50,000x, 60,000x, 70,000x, 80,000x, 90,000, 100,000x, or more.
[0067] In various embodiments, a whole exome sequencing is used to sequence nucleic acids ofa subject. The whole exome sequencing may be performed at a non-uniform depth. For example, certain areas of the exome may be boosted or otherwise sequenced at a greater depth than other regions, or at a greater depth than the average depth of the whole exome sequencing. By sequencing certain regions at a higher depth, genes or regions that are of more interest may be analyzed with higher sensitivity, accuracy, and / or precision. Genes or regions associated with or related to cancer can be sequenced at a greater depth. For example, at least 100, 200, 300, 400, 500, 600, 700, 800, 900, or more genes can be sequenced at a higher depth than the rest of the exome (e.g. average depth of the whole exome sequencing).
[0068] In various embodiments, a targeted sequencing is used to sequence nucleic acids of asubject. For example the targeted sequencing may be a promoter targeted sequencing. For example, the sequencing may be targeted to sequence regions associated with a promoter. This promoter targeted sequence may allow detecting sequence that are expected to be actively transcribed. The targeted sequencing may be performed at a non-uniform depth. For example, certain areas of the target sequencing may be boosted or otherwise sequenced at a greater depth than other regions, or at a greater depth than the average depth of the targeted sequencing.
[0069] In various embodiments, a methylation aware sequencing is used to sequence nucleicacids of a subject. For example, a bisulfite sequencing may be used. In another example, sequences that are methylated may be preferentially converted or enriched to generate sequencing reads that correspond to methylated regions.
[0070] In various embodiments, sequencing may be used to determine the expression orexpression level of one or more genes. For example, the expression may be determined whole transcriptome sequencing, or targeted transcriptome sequencing. The sequencing may comprise sequencing RNA molecules (e.g., cfRNA, mRNA), or derivatives thereof. The sequencing may comprise converting one or more RNA molecules into one or more DNA molecules and sequencing the one or more DNA molecules, or derivatives thereof.
[0071] The sequencing may allow for determination of the activation of one or moretranscription factors. For example, the sequencing may allow for detection of nucleic acid fragments and may be used to generate fragmentomics profile. The one or more fragmentomics profiles may be generated at one or more transcription factor binding sites. Detection of activation of the transcription factor may then be determined by the fragmentomics at the transcription factor binding sites.WSGR Docket No. 59987-717.601
[0072] The sequencing of nucleic acids may generate sequencing read data. The sequencingreads may be processed such to generate data of improved quality. The sequencing reads may be generated with a quality score. The quality score may indicate an accuracy of a sequence read or a level or signal above a nose threshold for a given base call. The quality scores may be used for filtering sequencing reads. For example, sequencing reads may be removed that do not meet a particular quality score threshold. The sequencing reads may be processed such to generate a consensus sequence or consensus base call. A given nucleic acid (or nucleic acid fragment) may be sequenced and errors in the sequence may be generated due to reactions prior or during sequencing. For example, amplification or PCR may generate error in amplicons such that the sequences are not identical to a parent sequence. Using sample barcodes or molecular barcodes, error correction may be performed. Error correction may include identifying sequence reads that do not corroborate with other sequences from a same sample or same original parent molecules. The use of barcodes may allow the identification or a same parent or sample. Additionally, the sequence reads may be processed by performing single strand consensus calling or double stranded consensus call, thereby reducing or suppressing error.
[0073] In various aspects, the expression or expression levels of one or more genes may bedetermined. The expression or expression levels of one or more genes may be determined by assaying a sample from an individual. The assaying may include nucleic acid sequencing, microarrays, PCR based assays (e.g., qPCR, or TaqMan), other hybridization-based assay, or other techniques that can analyze the expression or expression levels of genes. The expression levels may be analyzed with other data (e.g., sequencing data or transcription activation data) to determine a disease forecast parameter.
[0074] The methods as disclosed herein may comprise determining allele frequency or othercancer related metric. The methods may comprise a mutant allele frequency of a set of somatic mutation among a set of biomarkers. The mutant allele frequency may be used to determine a circulating tumor DNA (ctDNA) fraction of a cancer of a subject. A plasma tumor mutational burden (pTMB) of a cancer of the subject may be determined based at least in part on the set of mutant allele frequencies. Detection of microsatellite instability may also be used to determine the presence or absence of a cancer or cancer metric. Methylation states may be determined using methods described herein and may be used to identify a presence of a cancer or cancer parameter.
[0075] In various aspects, sets of biomarkers are processed and data corresponding to thebiomarkers are generated. The sets of biomarkers may comprise quantitative measures from a set of cancer-associated genomic loci. The cancer-associated genomic loci may correspond to a set of genes.WSGR Docket No. 59987-717.601
[0076] The sets of biomarkers may correspond to genetic aberration of a genetic locus. Thegenetic aberration may a tumor associated alteration. The genetic aberration may comprise copy number alterations (CNAs), copy number losses (CNLs), single nucleotide variants (SNVs), insertions or deletions (indels), and / or rearrangements. The set of biomarkers may be identified in a variety of nucleic acid types. For example, the tumor associated alteration may be identified in cfDNA. The tumor associated alteration may comprise changes in allelic expression, or gene expression. Methods and systems disclosed herein may allow for gene expression profiling and identification of changes to the expression levels of gene.
[0077] In various aspects, the presence or absence of a biomarker may be identified. Forexample, an assay (e.g., a sequencing assay) may detect that a biomarker is present by detecting a sequence corresponding to the biomarker. Similarly, an assay (e.g., a sequencing assay) may be used to detect the absence of a biomarker by identifying that the sequences do not correspond to the biomarker. For example, a wild type sequence may be identified in a sequence read and a wild type biomarker may be identified as present. In another example, in a plurality of sequence reads, only a wild type sequence is observed. The absence of a mutation in the wild type may be determined based on the lack of sequences generated that contain the mutation. As described elsewhere in the disclosure, the presence or absence of the biomarker may then be used to determine the presence or absence of a disease, or other disease parameter (e.g., disease subtype), or disease forecast characteristic. The biomarker may be the activation status of a transcription factor. For example, the methods and systems may allow for the determination that a transcription factor is activated or has an activity that is greater than or less than a threshold or reference value. The activated status of one or more transcription factors may be indicative of a disease (e.g., cancer) or disease subtype (e.g., cancer subtype). For example, a cancer may be associated with a high (e.g., higher than a healthy individual or individual without cancer) activity of a transcription factor. The activation status of transcription factor may be determined via processing of sequencing data (e.g., whole genome sequencing data, transcriptome sequencing data, methylation-aware sequencing data) or may be used in conjunction with sequencing data and processed to determine a disease forecast parameter.
[0078] In various aspects, the methods may comprise identifying the presence of a cancer or acancer parameter. The methods may comprises determining a probability or a likelihood of the presence of cancer or a cancer parameter. For example, instead of a binary output indicating a presence or absence, an output may be generated that indicates a probability that subject has cancer. This probability may be determined based on algorithms as described elsewhere herein. Similarly, a probability or likely of response to a particular treatment or a probability of relapse may be outputted. In some embodiments, the selected biomarker can comprise a selected bindingWSGR Docket No. 59987-717.601 site. In some embodiments, the selected biomarker can comprise a transcription factor binding site. In some embodiments, the selected biomarker can comprise one or more of: androgen receptor binding sites (ARBS), Achaete-Scute Family BHLH Transcription Factor binding sites, estrogen receptor (ER) binding sites, or ErbB receptor binding sites, or any combination thereof.
[0079] In various embodiments, fragmentomics profiles can be generated. The fragment size andlocation of various cfDNA fragments may be analyzed to generate fragmentomics profiles. Fragmentomics profiles may can be used to determine disease parameters and disease forecast characteristics. In some embodiments, the fragmentomics profiles are nucleosome profiles. Nucleosome profiles may indicate the location and activity of nucleosomes, and can be used to determine the chromatin state or activity of a region of the genome. In some embodiments, the nucleosome profile can be generated based at least in part on distance of one or more selected subsets of the cfDNA molecules to the selected biomarker. In some embodiments, the distance can comprise the number of base pairs between each of the one or more selected subsets of the cfDNA molecules and the selected biomarker. In some embodiments, the nucleosome profile can be further generated based at least in part on the coverage of the one or more selected subsets of the cfDNA molecules. In some embodiments, the nucleosome profile can be further generated based at least in part on the coverage of the one or more selected subsets of the cfDNA molecules associated with the distance of the one or more selected subsets of the cfDNA molecules from the selected biomarker. In some embodiments, determining the nucleosome profiling abnormality score can further comprise comparing the coverage of the one or more selected subsets of the cfDNA molecules to one or more coverage values of the one or more reference nucleosome profiles. In some embodiments, determining the nucleosome profiling abnormality score can further comprise determining a Z-score of the coverage of the one or more selected subsets of the cfDNA molecules. In some embodiments, determining the nucleosome profiling abnormality score can further comprise mapping the coverage Z-score of the one or more selected subsets of the cfDNA molecules to a coverage Z-score of each of the one or more reference nucleosome profiles.
[0080] In various aspects, fragmentomics profiles (e.g., nucleosome profiles) may be generatedand analyzed. The fragmentomics profiles may be generated based at least in part on sequencing data, or the fragmentomics profiles may analyzed in conjunction to sequencing data, such as those sequencing data generated using sequencing as described throughout the disclosure. For example, the fragmentomics profiles may use or be analyzed in conjunction with methylation- aware sequencing data. The fragmentomics profiles may be analyzed to determine the activation of transcription factors. For example, the fragmentomics profiles may be analyzed or generated at transcription binding sites (e.g., using sequence reads that align to one or more transcriptionWSGR Docket No. 59987-717.601 binding sites). For example, the fragmentomics profiles may generated or analyzed in relation to Androgen receptor binding sites (ARBS). The fragmentomics profile, methylation data, expression data, or other data can be combined to determine an activation of a transcription factor. For example, methylation data at one or more transcription binding sites may be used to determine activation of a transcription factor.
[0081] In some embodiments, the one or more selected subsets of the cfDNA molecules cancomprise fragments of the cfDNA molecules.
[0082] In some embodiments, generating the disease forecast can further comprise mapping thenucleosome profiling abnormality score to the one or more disease forecast characteristics.
[0083] In some embodiments, the one or more reference nucleosome profiles can be associatedwith one or more disease forecast characteristics. In some embodiments, the one or more reference nucleosome profiles can be associated with a likelihood of occurrence of one or more disease forecast characteristics. In some embodiments, one or more of the reference nucleosome profiles can comprise data from a sample not having the disease (normal). In some embodiments, the method can further comprise mapping DNA methylation patterns of the biological sample.
[0084] In some embodiments, the method can further comprise associating the mapped DNAmethylation patterns with the nucleosome profile data of the biological sample. In some embodiments, the one or more selected subsets of the cfDNA molecules can be selected based at least in part on DNA methylation data mapped to the cfDNA molecules.
[0085] In some embodiments, the method can further comprise generating the disease forecastbased at least in part on copy number variation data, or sequencing mutation data, or both, associated with the sample.
[0086] In some cases, a clinical intervention or a therapy may be identified at least in part basedon the identification of the types or subtypes cancer, or one or more disease forecast characteristics of the subject. The clinical intervention may be a plurality of clinical interventions. The therapy may be a plurality of therapies. The therapy may be selected from a plurality of clinical interventions. The clinical intervention or therapy may be a surgical resection, chemotherapy, radiotherapy, immunotherapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, or any combination thereof. In some cases, the clinical intervention or therapy may be administered to the subject.
[0087] Methods of Determining Disease Subtype
[0088] In yet another aspect, disclosed herein is a method for determining a subtype of a cancerof a subject. In some embodiments, the method can comprise generating a fragmentomics profile (e.g., a nucleosome profile) relating to cfDNA molecules derived from a biological sample obtained or derived from the subject. In some cases, the fragmentomics profile (e.g., aWSGR Docket No. 59987-717.601 nucleosome profile) can be generated relative to one or more selected biomarkers. In some embodiments, the method can further comprise mapping one or more characteristics of the fragmentomics profile (e.g., a nucleosome profile) of the biological sample to one or more corresponding characteristics of one or more reference the fragmentomics profiles (e.g., nucleosome profiles) to generate a selected subset of reference nucleosome profiles.
[0089] In some embodiments, the method can further comprise obtaining data relating to thecancer subtypes associated with each reference fragmentomics profile (e.g., reference nucleosome profile) of the selected subset of reference fragmentomics profiles. In some embodiments, the method can further comprise determining the subtype of the cancer of the subject based at least in part on the cancer subtypes associated with each of the reference fragmentomics profiles. The method may further comprise using data related to expression or expression levels of one or more genes. For example, data on a subject’s gene expression may be generated. A determination of cancer subtype may be made based at least upon the fragmentomics profiles and data related to expression or expression levels of one or more genes.
[0090] In some embodiments, the method can further comprise determining a tumor fraction ofthe cfDNA molecules.
[0091] In some embodiments, the method can further comprise associating the tumor fractionwith a fragmentomics profiling abnormality score (e.g. a nucleosome profiling abnormality score). In some embodiments, the fragmentomics abnormality score can be associated with the one or more selected biomarkers. In some embodiments, the one or more selected biomarkers can comprise a binding site associated with a type of the cancer.
[0092] In some embodiments, the selected biomarker can comprise a transcription factor bindingsite associated with a type of the cancer. In some embodiments, the selected biomarker can comprise one or more of: androgen receptor binding sites (ARBS), Achaete-Scute Family BHLH Transcription Factor (ASCL- ) binding sites, estrogen receptor (ER) binding sites, or ErbB receptor binding sites, or any combination thereof.
[0093] In some embodiments, at least a portion of the reference fragmentomics profiles cancomprise non-disease profiles (normal). In some embodiments, at least a portion of the reference nucleosome profiles can comprise a cancer condition having a cancer type matching the type of the cancer of the subject.
[0094] Methods for Estimating Treatment Response
[0095] In yet another aspect, disclosed herein is a method for estimating a response of a subjecthaving cancer to one or more treatments. In some embodiments, the method can comprise generating a fragmentomics profile (e.g., a nucleosome profile) relating to cfDNA molecules derived from a biological sample obtained or derived from the subject. In some embodiments, theWSGR Docket No. 59987-717.601 method can further comprise determining one or more characteristics of the cancer of the subject based at least in part on the fragmentomics profile (e.g., a nucleosome profile) of the subject. The method may determine one or more characteristics of the cancer of the subject based at least in part on the fragmentomics profile (e.g., a nucleosome profile) of the subject and data relating to the expression of genes of the subject. In some embodiments, the method can further comprise generating a treatment response determination for the subject based at least in part on the one or more characteristics of the cancer of the subject.
[0096] In some embodiments, the treatment response determination can comprise one or moreof: a treatment plan, a value representing likelihood of one or more treatment responses, a binary treatment response indicator, a probability value for each treatment response, or any combination thereof.
[0097] In some embodiments, the one or more characteristics of the cancer of the subject cancomprise a cancer type, a cancer subtype, an estimate of cancer progression, a prognosis of the subject without treatment intervention, a prognosis of the subject with treatment intervention, or any combination thereof.
[0098] Systems for Determining Disease Forecast Characteristics
[0099] In yet another aspect, disclosed herein is a system for determining a disease forecast of asubject having cancer. In some embodiments, the system can comprise a memory and one or more processors. In some embodiments, the one or more processors can be configured to execute machine-readable instructions which, when executed, cause the one or more processors to perform a method. In some embodiments, the method performed by the one or more processors can comprise generating a fragmentomics profile (e.g., nucleosome profile) of a biological sample obtained or derived from the subject. In some embodiments, the method performed by the one or more processors can further comprise mapping the fragmentomics profile (e.g., nucleosome profile to one or more reference fragmentomics profiles. In some embodiments, the method performed by the one or more processors can further comprise determining one or more disease forecast characteristics of the biological sample. In some cases, the one or more disease forecast characteristics can be determined based at least in part on the mapping. In some embodiments, the method performed by the one or more processors can further comprise generating the disease forecast based at least in part on the one or more disease forecast characteristics.
[0100] In some embodiments, the one or more disease forecast characteristics cancomprise one or more of: a type of the cancer, a subtype of the cancer, a prognosis of the subject, an estimated survival time of the subject without treatment intervention, an estimated survivalWSGR Docket No. 59987-717.601 time of the subject with treatment intervention, an estimation of the subject’s response to one or more treatments, or any combination thereof.
[0101] In various aspects, the nucleosome profile of the subject, or one or more referencenucleosome profiles, or both, are processed using one or more algorithms. In some embodiments, the one or more processors can further comprise one or more software modules or models configured to operate utilizing the one or more algorithms. The one or more algorithms can comprise machine learning (ML) or artificial intelligence (AI) algorithms. The AI / ML algorithms may be trained algorithms. The trained algorithms may utilize the selected biomarkers, or the nucleosome profile of the sample, or both, as an input. The AI / ML algorithms can generate an output relating to the one or more disease forecast characteristics of a cancer. The output may be specific to a type of cancer or subtype of cancer. The output can comprise determining the type of cancer or the subtype of cancer of the sample. For example, the output may indicate the presence of a castrate-resistant prostate cancer. For example, the output may indicate the presence of ER- or ER+ breast cancer.
[0102] In various aspects, the sets of biomarkers are processed using an algorithm. Thealgorithm may be a trained algorithm. The trained algorithms may use the sets of biomarkers as an input and generate an output regarding the presence or absence of a cancer. The output may be specific to a type of cancer or subtype of cancer. For example, the output may indicate the presence of bladder cancer.
[0103] The trained algorithm may be trained on multiple samples. For example, the trainedalgorithm may be trained using at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 300, 400, 500 , 600 ,700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or more independent training samples. The trained algorithm may be trained using no more 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 300, 400, 500 , 600 ,700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or less, independent training samples. The training samples may be associated with a presence or an absence of the cancer. The training samples may be associated with a relapse of cancer. The training samples may be associated with cancer that is resistant to a particular drug or treatment. An individual training sample may be positive for a particular cancer. An individual training sample may be negative for a particular cancer. By using training samples, the trained algorithm may be able to detect a cancer, determine a probability of recurrence or relapse of a cancer, or determine if a cancerWSGR Docket No. 59987-717.601 comprises a set of biomarkers may be resistant to a treatment. The training sample may be associated with additional clinical health data of a subject. For example, additional clinical health data may comprise the gender, weight, height, or levels of metabolites or antibodies in a subjects. Additional clinical health data may comprise indication of other diseases, disorders, or diseases conditions.
[0104] The trained algorithms may be trained using multiple sets of training samples. The setsmay comprise training samples as described elsewhere herein. For example, the training may be performed using a first set of independent training samples associated with a presence of the cancer and a second set of independent training samples associated with an absence of the cancer. Similarly, a first set may be associated with relapse and a second sample may be associated with the absence of relapse.
[0105] The trained algorithm may also process additional clinical health data of the subject.For example, additional clinical health data may comprise the gender, weight, height, or levels of metabolites or antibodies in a subject. Additional clinical health data may comprise indication of other diseases, disorders, or diseases conditions that the subject may suffer from. By using the additional clinical health data, in conjunction with the biomarkers, the trained algorithm may output a presence or absences of cancer, probability of relapse, or resistance to drug treatment, which may be different from the output of an algorithm that does not process additional clinical health.
[0106] The trained algorithm may be an unsupervised machine learning algorithm. Forexample, the unsupervised machine learning algorithm may utilize cluster analysis to identify attributes of interest. The trained algorithm may be a supervised machine learning algorithm. For example, the trained algorithm may be trained with training data such to generate an expected or desired output. The supervised learning algorithm may comprise a deep learning algorithm, a support vector machine (SVM), a neural network, or a Random Forest. Via the machine learning algorithm, the trained algorithm may be able to identify relationships of biomarkers to a particular cancer prognosis or diagnosis. Without the trained algorithm, it may otherwise be difficult to identify relationships of the biomarkers to accurately identify the presence of a cancer or other parameters associated with the cancer.
[0107] Via the machine learning algorithm, the trained algorithm may be able to identifyrelationships of fragmentomics profiles (e.g., nucleosome profiles) to particular cancer prognoses or types or subtypes. Without the trained algorithm, it may otherwise be difficult to identify relationships of the fragmentomics profiles to cancer types or subtypes. Without the trained algorithm, it may otherwise be difficult to identify relationships of the fragmentomics profiles to prognosis or disease forecast characteristics.WSGR Docket No. 59987-717.601
[0108] In various aspects, the systems and methods may comprise a accuracy, sensitivity, orspecificity of determination of type or subtype of cancer or prognosis or disease forecast. For example, the methods or systems may comprise in the subject determination of type or subtype of cancer or prognosis or disease forecast at an accuracy of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. The methods or systems may comprise determination of type or subtype of cancer or prognosis or disease forecast in the subject at a sensitivity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. The methods or systems may comprise determination of type or subtype of cancer or prognosis or disease forecast in the subject at a specificity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.The methods or systems may comprise determination of type or subtype of cancer or prognosis or disease forecast in the subject at a positive predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. The methods or systems may comprise determination of type or subtype of cancer or prognosis or disease forecast in the subject at a negative predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
[0109] In some embodiments, performing nucleosome profiling at methylation-loss regions,can increase the sensitivity and specificity of the assays. In some embodiments, performing nucleosome profiling at methylation-loss regions, can increase the sensitivity and specificity of the assays by isolating and eliminating fragments from analysis that do not show methylation changes between cancer samples and normal samples.
[0110] In various aspects, the systems and methods may comprise an accuracy,sensitivity, or specificity of detection of the cancer or a parameter of the cancer. For example, the methods or systems may comprise detecting the presence or the absence of cancer (or the presence of a parameter of the cancer, such as recurrence, relapse, or drug resistance) in the subject at an accuracy of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. The methods or systems may comprise detecting the presence or the absence of cancer (or the presence of a parameter of the cancer, such as recurrence, relapse, or drug resistance) in the subject at a sensitivity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least aboutWSGR Docket No. 59987-717.601 98%, or at least about 99%. The methods or systems may comprise detecting the presence or the absence of cancer (or the presence of a parameter of the cancer, such as recurrence, relapse, or drug resistance) in the subject at a specificity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.The methods or systems may comprise detecting the presence or the absence of cancer (or the presence of a parameter of the cancer, such as recurrence, relapse, or drug resistance) in the subject at a positive predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. The methods or systems may comprise detecting the presence or the absence of cancer (or the presence of a parameter of the cancer, such as recurrence, relapse, or drug resistance) in the subject at a negative predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. Computer control systems
[0111] The present disclosure provides computer systems that are programmed toimplement methods of the disclosure. FIG. 10 shows a computer system 1001 that is programmed or otherwise configured to perform analysis or operations of the methods, for example determine a likelihood of the presence of a cancer based on a set of biomarkers of an individual or run an algorithm. The computer system 1001 can regulate various aspects of methods and systems of the present disclosure, such as, for example, perform an algorithm, input training data, analyze sets of biomarker, or output a result for the user as to the presence or absence of cancer. The computer system 1001 can be an electronic device of a user or a computer system that is remotely located with respect to the electronic device. The electronic device can be a mobile electronic device.
[0112] The computer system 1001 includes a central processing unit (CPU, also“processor” and “computer processor” herein) 1005, which can be a single core or multi core processor, or a plurality of processors for parallel processing. The computer system 1001 also includes memory or memory location 1010 (e.g., random-access memory, read-only memory, flash memory), electronic storage unit 1015 (e.g., hard disk), communication interface 1020 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 1025, such as cache, other memory, data storage and / or electronic display adapters. The memory 1010, storage unit 1015, interface 1020 and peripheral devices 1025 are in communication with the CPU 1005 through a communication bus (solid lines), such as a motherboard. The storage unit 1015 can be a data storage unit (or data repository) for storing data. The computer system 1001 can be operatively coupled to a computer network (“network”) 1030 with the aid of theWSGR Docket No. 59987-717.601 communication interface 1020. The network 1030 can be the Internet, an internet and / or extranet, or an intranet and / or extranet that is in communication with the Internet. The network 1030 in some cases is a telecommunication and / or data network. The network 1030 can include one or more computer servers, which can enable distributed computing, such as cloud computing. The network 1030, in some cases with the aid of the computer system 1001, can implement a peer-to- peer network, which may enable devices coupled to the computer system 1001 to behave as a client or a server.
[0113] The CPU 1005 can execute a sequence of machine-readable instructions, whichcan be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 1010. The instructions can be directed to the CPU 1005, which can subsequently program or otherwise configure the CPU 1005 to implement methods of the present disclosure. Examples of operations performed by the CPU 1005 can include fetch, decode, execute, and writeback.
[0114] The CPU 1005 can be part of a circuit, such as an integrated circuit. One or moreother components of the system 1001 can be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).
[0115] The storage unit 1015 can store files, such as drivers, libraries and savedprograms. The storage unit 1015 can store user data, e.g., user preferences and user programs. The computer system 1001 in some cases can include one or more additional data storage units that are external to the computer system 1001, such as located on a remote server that is in communication with the computer system 1001 through an intranet or the Internet.
[0116] The computer system 1001 can communicate with one or more remote computersystems through the network 1030. For instance, the computer system 1001 can communicate with a remote computer system of a user (e.g., a medical professional or patient). Examples of remote computer systems include personal computers (e.g., portable PC), slate or tablet PC’s (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, Smart phones (e.g., Apple® iPhone, Android-enabled device, Blackberry®), or personal digital assistants. The user can access the computer system 1001 via the network 1030.
[0117] Methods as described herein can be implemented by way of machine (e.g.,computer processor) executable code stored on an electronic storage location of the computer system 1001, such as, for example, on the memory 1010 or electronic storage unit 1015. The machine executable or machine readable code can be provided in the form of software. During use, the code can be executed by the processor 1005. In some cases, the code can be retrieved from the storage unit 1015 and stored on the memory 1010 for ready access by the processorWSGR Docket No. 59987-717.601 1005. In some situations, the electronic storage unit 1015 can be precluded, and machine- executable instructions are stored on memory 1010.
[0118] The code can be pre-compiled and configured for use with a machine having aprocesser adapted to execute the code, or can be compiled during runtime. The code can be supplied in a programming language that can be selected to enable the code to execute in a pre- compiled or as-compiled fashion.
[0119] Aspects of the systems and methods provided herein, such as the computer system401, can be embodied in programming. Various aspects of the technology may be thought of as “products” or “articles of manufacture” typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine readable medium. Machine-executable code can be stored on an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. “Storage” type media can include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer into the computer platform of an application server. Thus, another type of media that may bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.
[0120] Hence, a machine readable medium, such as computer-executable code, may takemany forms, including but not limited to, a tangible storage medium, a carrier wave medium or physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or the like, such as may be used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radioWSGR Docket No. 59987-717.601 frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include for example: a floppy disk, a flexible disk, hard disk, magnetic tape, any other magnetic medium, a CD-ROM, DVD or DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium with patterns of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0121] The computer system 1001 can include or be in communication with an electronicdisplay 435 that comprises a user interface (UI) 1040 for providing, for example, an input of biomarkers or sequencing data, or an visual output relating to a detection, diagnosis, or prognosis. Examples of UI’s include, without limitation, a graphical user interface (GUI) and web-based user interface.
[0122] Methods and systems of the present disclosure can be implemented by way of oneor more algorithms. An algorithm can be implemented by way of software upon execution by the central processing unit 1005. The algorithm can, for example, determine a presence or absence of a cancer or cancer parameter based on a set of input sequencing data from a sample derived from a subject.
[0123] EXAMPLES
[0124] Example 1: Use of fragmentomics profiles, DNA methylation, and wholetranscriptome sequencing for disease detection and forecasting
[0125] Liquid biopsies based on cell-free DNA (cfDNA) analysis provide non-invasiveclinical diagnostic insights. Recently, cfDNA fragmentomics has emerged as a tool for inferring epigenomic and transcriptional information from tumor-derived cfDNA. A study was conducted using comprehensive profiling of cfDNA fragmentomics from promoter-targeted panel, high depth whole genome sequencing, genome-wide DNA methylation, and whole transcriptome sequencing. In this example, two metastatic castration-resistant prostate cancer (mCRPC) samples were used to demonstrate the utility of multi-marker assays and systematically investigate molecular aberrations in mCRPC.
[0126] Using methods and systems of the present disclosure, nucleosome profiling wasutilized for disease forecasting by assaying samples. As shown in FIG. 1A, a sample such as blood-derived plasma was assayed using plasma isolation and DNA extraction methods. Library preparation was then performed. Sequencing was performed on the extracted DNA of the sample. As shown, a sequencing using a promoter targeted panel, high depth whole genome sequencingWSGR Docket No. 59987-717.601 (150x), whole genome DNA methylation sequencing (30x) and whole transcriptome sequencing were used.
[0127] These different assays were applied to two mCRPC patients and normal plasmabackground. FIG. 1B-E shows data for fragments at AR binding sites and show nucleosome depletion and hypo-methylation in both mCRPC patients. As illustrated in FIG. 1B, nucleosome profiling at androgen receptor (AR) binding sites was performed and show nucleosome depletion in two mCRPC patients, which are inferred by genome-wide DNA methylation profile., with some patients having metastatic castration-resistant prostate cancer (mCRPC), and nucleosome profiles at the AR binding sites of normal patients not having cancer. Distance from AR binding sites, measured as number of base pairs, was correlated with coverage of the cfDNA fragments of the sample.
[0128] FIG. 1C demonstrate that the AR binding sites have hypo-methylation in twomCRPC patients, compared with normal plasma background. Figure 1D demonstrate the nucleosome profiling in relation to AR binding sites and show nucleosome depletion in two mCRPC patients, inferred by high depth WGS. FIG. 1E provides a chart of DNA methylation and ARBS profiling. DNA methylation and ARBS nucleosome profiling abnormality score were negatively correlated in 54 mCRPC prostate cancer samples.
[0129] A sample-level ARBS score was quantified by comparing its centric fragmentcoverage (±60bp) with normal plasma background using standard Z-score. A higher score indicates higher nucleosome depletion and high AR signaling. A promoter-targeted score was also quantified by comparing its normalized fragment size entropy with normal plasma background using standard Z-score. A Higher score indicates higher transcriptional activity.
[0130] Using cfDNA fragmentomics, chromatin activity can be also ascertained. Figure2A-D demonstrate the use of cfDNA fragmentomics at AR enhancer region which indicates high chromatin activity in both mCRPC patients. FIG. 2A shows fragment size distribution at AR enhancer (20kb region, at 600kb upstream of AR promoter), inferred from promoter-targeted panel, which covers cancer specific genes and AR enhancer. FIG. 2B shows fragment size distribution at AR enhancer, inferred from high depth WGS. FIG. 2C shows Fragment size distribution were highly coordinate from promoter-targeted panel and high depth WGS. Figure 2D shows that AR enhancers have hypo-methylation in two mCRPC patients, compared with healthy male plasma background.
[0131] FIG. 3 demonstrates the correlation of gene expression and cfDNA fragmentomicsinferred by promoter-targeted panel. FIG. 3A provides a scatterplot to show the relationshipbetween normalized fragment size entropy score and RNA gene expression inferred by WTS for normal plasma cfDNA. FIG. 3B shows A boxplot to compare normalized fragment size entropyWSGR Docket No. 59987-717.601 score of 14 key highly expressed prostate cancer genes from TCGA and literature, between two mCRPC patients and normal plasma background, demonstrating a statistical difference in fragment size in the prostate cancer sample compared to the normal sample. Figure 3C shows another a boxplot to compare normalized fragment size entropy score of lowly expressed prostate cancer genes from TCGA, between two mCRPC patients and normal plasma background. Similarly, the normal and prostate cancer samples show a significant difference. FIG. 3D shows a bar plot to show normalized fragment size entropy score of the 14 highly expressed prostate cancer genes for two mCRPC patients.
[0132] Using the genome sequencing, CNV profiles were generated for each of the twopatients. FIG. 4 shows these genome-wide CNV profiles of two mCRPC patients.
[0133] Based on the differing epigenomic and transcriptomic outputs, detection of thecancer and the subtype of the cancer could be determined. Table 1 shows a summary of the two samples.
[0134] Table 1. Summary of epigenomic and transcriptional dysregulation of twomCRPC patients
[0135] Collectively, these findings indicated that Sample A may represent androgenreceptor-dependent prostate cancer (ARPC), while Sample B may have developed resistance to androgen therapy, possibly indicating neuroendocrine prostate cancer (NEPC).
[0136] ConclusionWSGR Docket No. 59987-717.601
[0137] This study, to our knowledge, is the first to combine promoter-targeted panel,high-depth WGS, genome-wide DNA methylation and whole transcriptome sequencing, to systematically study epigenomic and transcriptional dysregulation in mCRPC. The integration of mutation, copy number variation, fragmentomics, and DNA methylation profiling of tumor cfDNA can have clinical utility for inferring prostate cancer subtypes, facilitating patient stratification, and guiding treatment selection.
[0138] Example 2: Use of cfDNA from cerebrospinal fluid for cancer detection andmonitoring.
[0139] Assaying of cancer can benefit from the use of cfDNA. Specifically, some cancersmay be difficult to assess via biopsy. For example, accurate assessment of glioma treatment response remains a challenge. Cell-free DNA (cfDNA), a promising liquid biopsy biomarker, has shown utility in diverse types of cancers. While CSF-derived cfDNA has been explored in gliomas, the dynamics of cfDNA abundance and its genetic profile throughout treatment and recurrence are unclear. To address this gap, data is presented relating to an expanding biobank of longitudinal CSF samples collected from glioma patients using CSF access devices. This study demonstrates the use of CSF cfDNA as a monitoring tool to evaluate treatment response in gliomas.
[0140] Methods
[0141] Longitudinal CSF specimens were acquired via Ommaya reservoirs(NCT04692337) or ventriculoperitoneal shunts (NCT04692324) in patients with gliomas. cfDNA was extracted from 1-5 mL of CSF based on available sample volumes and quantified via Qubit. Next-Generation Sequencing or low-pass whole genome sequencing (LPWGS) was performed by Predicine, Inc., depending on the amount of extractable cfDNA.
[0142] Results
[0143] Data from 5 patients are presented here, including 29 samples tested with LPWGSand representative SNV results from samples with sufficient cfDNA for the PredicineCARE assay. Tumor resection increased the quantified cfDNA (2.97x, range: 1.58-5.26x), consistent with the impact of increased parenchymal disruption and closer contact with CSF. Thereafter, CSF cfDNA copy number burden (CNB) decreased in response to standard-of-care and experimental therapies, such as pembrolizumab (patients 78 and 79), consistent with the patients’ radiographic or clinical course during those treatments. An empirical CNB score cut-off of 7.5 was identified by clinical experts to represent measurable copy number alterations based on a review of LPWGS plots. In one patient with glioblastoma with a hypermutated phenotype (patient 79), CSF cfDNA revealed genomic alterations including MAP2K1, KIT, and PDGFRA, among other variants of unknown significance, most of which had been detected in sequencing ofWSGR Docket No. 59987-717.601 the tissue acquired at resection and subsequently decreased with chemoradiation. In another patient with known progressive GBM that was EGFR amplified, over 200 new variants were identified by the patient’s final sample, and EGFR copy number increased from 2 to over 30.
[0144] FIG. 5 show various data relating to a first patient. A female in her 50s underwentresection for a recurrent GBM with a known EGFR amplification, as well as Ommaya placement for research, including cfDNA testing (A). Further progression was noted by POD26, prompting initiation of lomustine. The patient then went on bevacizumab after progression on lomustine by POD75. CNB increased throughout progression (B), as did the EGFR and TERT variant allele frequencies (VAFs) (C). Whole genome CNV plot demonstrated +7 / -10 chromosomal alterations with progression (D).
[0145] FIG. 6 shows various data relating to a second patient. A female in her 30sunderwent resection for primary GBM, as well as Ommaya placement for research purposes. Tissue NGS revealed a highly mutated tumor (“hypermutant”). LPWGS analysis of all samples from this patient did not reveal any copy number change signals (data not shown). CSF cfDNA was highly abundant and increased following resection (A), as observed with patient 78. True to the hypermutated nature of her tumor, over 400 variants were detected in her cfDNA (B). Chemoradiation and pembrolizumab reduced the cfDNA tumor fraction. Radiographic progression was suspected around POD300, prompting the initiation of bevacizumab, to which there was a radiographic response. The evolution of established and new variant alleles from NGS is shown in Figure D.
[0146] FIG. 7 shows various data relating to a third patient. A male in his 60s underwentresection and Ommaya placement for a primary GBM. Resection increased the cfDNA yield and CNB (A-B). The patient then underwent chemoradiation, during which cfDNA yield and CNB decreased, as did many of the amplifications / deletions detected on the whole genome plot (C), in addition to VAFs (not shown). Suspected progression around POD260 led to initiation of regorafenib. The patient began receiving bevacizumab after further progression was noted 2-3 months after starting regorafenib.
[0147] FIG. 8 shows various data relating to a fourth patient. A male in his 30sunderwent resection and chemoradiation for an astrocytoma, IDH-mutant, grade 4. An EVD was present before resection and a VP shunt was placed after resection and sampled during treatment. CSF cfDNA yield increased with resection (A), while CNB and tumor fraction decreased (B-C). Radiographic change was noted at POD214. Consistent with this finding, IDH1 VAF and D-2- HG in CSF both increased (D).
[0148] FIG. 9 shows various data relating to a fifth patient. A male in his 40s underwentresection for an astrocytoma, IDH-mutant, grade 4 before undergoing chemoradiation and anWSGR Docket No. 59987-717.601 immunotherapy trial wherein he was randomized to either IL-7 or placebo. While CSF cfDNA yield was variable (A), the CNB trended downward in CSF obtained after chemoradiation during adjuvant TMZ throughout observation (POD146-413) (B). IDH1 VAF also decreased concurrently with decrease in D-2-hydroxyglutarate (D-2-HG), the oncometabolite of IDH- mutant tumors (C), as did alterations in the whole genome plot (D). This contrasted with the patient’s imaging that was equivocal for pseudoprogression versus tumor progression.
[0149] Conclusion
[0150] Analysis of CSF cfDNA over a patient’s disease course, including resection andtreatment with standard-of-care and experimental therapies, may be useful for disease monitoring and treatment response. Further work is needed to evaluate these findings in a larger cohort of patients and to determine the sensitivity of changes in mutations, CNB, and cfDNA quantity for glioma disease burden. From this study, multiple takeaway were found. Specifically, it was found that copy number burden can be used a more reliable indicator of changes in tumor burden than cfDNA yield. Additionally, in the study, a greater number of mutations can be detected in CSF cfDNA after resection. CSF D-2-hydroxyglutarate levels correlated well with changes in IDH1 cfDNA, which do not always correlate with radiographic findings.
[0151] As shown, it was possible to monitor over time the evolution of genetic variantsand their allelic frequencies within a patient’s CSF cfDNA. For example, detected amplifications and deletions in whole genome plots disappear throughout radiographically successful treatment. Based on the study, it was determined that longitudinal intracranial CSF can be acquired to determine the impact of treatment on CSF cfDNA.
[0152] While preferred embodiments of the present invention have been shown anddescribed herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it shall be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is therefore contemplated that the invention shall also cover any such alternatives, modifications, variations or equivalents. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
WSGR Docket No. 59987-717.601 CLAIMS WHAT IS CLAIMED IS:
1. A method of determining one or more disease forecast characteristics of a subject havingor suspected of having cancer, the method comprising: (a) generating one or more fragmentomics profiles of cell free nucleic acids obtainedor derived from the subject, wherein a fragmentomics profile is generated relative to a selected biomarker; (b) determining an expression of one or more genes of the subject, or determining anactivation status of one or more transcription factors; (c) based at least in part on (i) the one or more fragmentomics profiles and (ii) theexpression of one or more genes or the activation status of one or more transcription factors, determining one or more disease forecast characteristics in said subject.
2. The method of claim 1, wherein the one or more fragmentomics profile is a nucleosomeprofile.
3. The method of any one of claims 1 to 2, wherein the fragmentomics profile is generatedbased at least in part on distance of one or more selected subsets of the cell free nucleic acids to the selected biomarker.
4. The method of claim 3, wherein the distance comprises the number of base pairs betweeneach of the one or more selected subsets of the cell free nucleic acids and the selected biomarker.
5. The method of any of claims 1 to 4, wherein the fragmentomics profile is furthergenerated based at least in part on the coverage of the one or more selected subsets of the cell free nucleic acids.
6. The method of any of claims 1 to 5, wherein the nucleosome profile is further generatedbased at least in part on the coverage of the one or more selected subsets of cell free nucleic acids associated with the distance of the one or more selected subsets of the cell free nucleic acids from the selected biomarker.
7. The method of any one of claims 3 to 6, wherein the one or more selected subsets of thecell free nucleic acids comprise fragments of the cell free nucleic acids.
8. The method of any one of claims 1 or 2, further comprising determining a fragmentomicsprofiling abnormality score.
9. The method of any one of claims 3, wherein the determining a fragmentomics profilingabnormality score comprises relating the one or more fragmentomics profiles to one or more reference fragmentomics profiles.WSGR Docket No. 59987-717.60110. The method of claim 9, wherein determining the fragmentomics profiling abnormalityscore further comprises comparing the coverage of the one or more selected subsets of cell free nucleic acids to one or more coverage values of the one or more reference fragmentomics profiles.
11. The method of claim 9 or 10, wherein determining the fragmentomics profilingabnormality score further comprises determining a Z-score of the coverage of the one or more selected subsets of the cell free nucleic acids.
12. The method of claim 11, wherein determining the fragmentomics profiling abnormalityscore further comprises mapping the coverage Z-score of the one or more selected subsets of the cell free nucleic acids to a coverage Z-score of each of the one or more reference fragmentomics profiles.
13. The method of any one of claims 8 to 12, wherein (c) comprises determining one or moredisease forecast characteristics in said subject, in part by processing said fragmentomics profiling abnormality score.
14. The method of any one of claims 8 to 12, wherein generating the disease forecast furthercomprises mapping the nucleosome profiling abnormality score to the one or more disease forecast characteristics.
15. The method of any one of claims 1 to 14, wherein the one or more reference nucleosomeprofiles are associated with one or more disease forecast characteristics.
16. The method of claim 15, wherein the one or more reference nucleosome profiles areassociated with a likelihood of occurrence of one or more disease forecast characteristics.
17. The method of any one of claims 15 or 16, wherein one or more of the referencenucleosome profiles comprise data from a sample not having the disease (normal).
18. The method of any one of claims 1 to 13, wherein the selected biomarker comprisesandrogen receptor binding sites.
19. The method of any one of claims 1 to 18, wherein the one or more disease forecastcharacteristics comprise a cancer prognosis relating to the subject.
20. The method of any one of claims 1 to 19, wherein the one or more disease forecastcharacteristics comprise one or more of: an estimated survival time of the subject without a treatment intervention, an estimated survival time of the subject with a treatment intervention, determination of a type of the cancer, determination of a subtype of the cancer, determination of one or more clinical outcomes, or predicted treatment response of the subject to one or more treatments, or any combination thereof.
21. The method of any one of claims 1 to 20, wherein the determining an expression of one ormore genes comprises determining one or more expression levels of the one or more genes.WSGR Docket No. 59987-717.60122. The method of any one of claims 1 to 21,wherein (b) comprises assaying nucleic acidsderived from the subject.
23. The method of claim 22, wherein the nucleic acids derived from the subject are RNAmolecules.
24. The method of claim 23, wherein the RNA molecules are cell-free RNA molecules.
25. The method of any one of claim 22 to 24, wherein assaying nucleic acids derived fromsubject comprises nucleic acid sequencing.
26. The method of claim 25,wherein the nucleic acid sequencing is an RNA sequencingassay.
27. The method of claim 25, wherein the nucleic acid sequencing is a whole transcriptomesequencing assay or a targeted sequencing assay.
28. The method of any one of claims 1 to 27, wherein the biological sample comprises cell-free deoxyribonucleic acid (cfDNA) molecules.
29. The method of any one of claims 1 to 28, wherein the biological sample comprises one ormore of: a plasma sample, a serum sample, a red blood cell sample, a urine sample, a saliva sample, pleural fluid sample, peritoneal fluid sample, amniotic fluid sample, cerebrospinal fluid sample, lymphatic fluid sample, sweat sample, tear sample, semen sample, or any derivative thereof, and any combination thereof.
30. The method of any one of claims 1 to 29, wherein the biological sample comprises theplasma sample.
31. The method of any one of claims 1 to 30, wherein the biological sample comprises theurine sample.
32. The method of any one of claims 28 to 31, wherein the cfDNA molecules are obtained orderived from a single biological sample of the subject.
33. The method of any one of claims 28 to 31, wherein the cfDNA molecules are obtained orderived from different biological samples of the subject.
34. The method of any one of claims 1 to 33, wherein the biological sample is obtained orderived from the subject using an ethylenediaminetetraacetic acid (EDTA) collection tube, a cell- free RNA collection tube, or a cell-free deoxyribonucleic acid (DNA) collection tube, other blood collection tube, and CTC collection tubes.
35. The method of any one of claims 1 to 34, further comprising assaying the biologicalsample to generate the one or more fragmentomics profile.
36. The method of claim 35, wherein assaying the biological sample comprises subjectingsaid biological sample to conditions that are sufficient to isolate, enrich, or extract the cfDNA molecules.WSGR Docket No. 59987-717.60137. The method of claim 35 or 36, further comprising fractionating a whole blood sample ofthe subject to obtain the cfDNA molecules.
38. The method of any one of claims 35 to 37, wherein assaying the biological sample furthercomprises assaying the cfDNA molecules using nucleic acid sequencing to produce nucleic acid sequencing reads.
39. The method of claim 38, wherein the nucleic acid sequencing further comprises DNAsequencing.
40. The method of claim 39, wherein the DNA sequencing comprises one or more of: next-generation sequencing, whole genome sequencing, low-pass sequencing, targeted sequencing, whole exome sequencing, methylation-aware sequencing, or bisulfite sequencing, or a combination thereof.
41. The method of claim 39, wherein the DNA sequencing comprises low-pass wholegenome sequencing.
42. The method of claim 39, wherein the DNA sequencing comprises whole exomesequencing.
43. The method of any one of claims 39 to 42, wherein the DNA sequencing furthercomprises nucleic acid amplification.
44. The method of claim 43, wherein the nucleic acid amplification comprises polymerasechain reaction (PCR) or isothermal amplification.
45. The method of any one of claims 28 to 44 wherein at least one of the cfDNA moleculesare assayed using a polymerase chain reaction (PCR) assay, microarray, or a isothermal amplification.
46. The method of any one of claims 1 to 45, wherein the type of the cancer of the subjectcomprises one or more of: lung cancer, brain cancer, spinal cancer, colorectal cancer, melanoma, bladder cancer, non-Hodgkin lymphoma, kidney cancer, endometrial cancer, leukemia, pancreatic cancer, thyroid cancer, or liver cancer, or any combination thereof.
47. The method of any one of claims 1 to 46, wherein the type of the cancer of the subjectcomprises prostate cancer.
48. The method of claim 47, wherein the subtype of the prostate cancer comprises one ormore of: hormone sensitive prostate cancer (HSPC), castration-resistant prostate cancer (CRPC), androgen receptor-dependent prostate cancer (ARPC), metastatic prostate cancer, metastatic castration-resistant prostate cancer (mCRPC), neuroendocrine prostate cancer (NEPC), or any combination thereof.
49. The method of any one of claims 1 to 48, wherein the subject is asymptomatic for thecancer.WSGR Docket No. 59987-717.60150. The method of any one of claims 1 to 49, wherein the one or more transcription factors isan androgen receptor.
51. The method of any one of claim 1 to 50, wherein the determining the activation status ofthe one or more transcription factors comprises processing the one or more fragmentomics profiles.
52. The method of any one of claims 1 to 51, wherein the selected biomarker comprises aselected binding site.
53. The method of any one of claims 1 to 52, wherein the selected biomarker comprises atranscription factor binding site.
54. The method of claim 52 or 53, wherein the selected biomarker comprises one or more of:androgen receptor binding sites (ARBS).
55. The method of any one of claims 1 to 54, further comprising mapping DNA methylationpatterns of the biological sample.
56. The method of claim 55, further comprising associating the mapped DNA methylationpatterns with the nucleosome profile data of the biological sample.
57. The method of any one of claims 3 to 56, wherein the one or more selected subsets of thecell free nucleic acids are selected based at least in part on DNA methylation data mapped to the cell free nucleic acids.
58. The method of any one of claims 1 to 57, further comprising generating the diseaseforecast based at least in part on copy number variation data, or sequencing mutation data, or both, associated with the sample.
59. The method of any of claims 1 to 58, further comprising determining a tumor fraction ofthe cell free nucleic acids.
60. The method of claim 59, further comprising determining a fragmentomics profilingabnormality score, comprising relating the fragmentomics profile of the subject sample to one or more reference fragmentomics profile and associating the tumor fraction with the fragmentomics profiling abnormality score.
61. A method for estimating a response of a subject having cancer to one or more treatments,comprising: (a) generating a fragmentomics profile relating to cfDNA molecules derived from abiological sample obtained or derived from the subject; (b) determining expression of one or more genes, or determining an activation statusof one or more transcription factors of the subject; (c) determining one or more characteristics of the cancer of the subject based at leastin part on the fragmentomics profile of the subject; andWSGR Docket No. 59987-717.601 (d) generating a treatment response determination for the subject based at least in parton the one or more characteristics of the cancer of the subject.
62. The method of claim 61, wherein the treatment response determination comprises one ormore of: a treatment plan, a value representing likelihood of one or more treatment responses, a binary treatment response indicator, a probability value for each treatment response, or any combination thereof.
63. The method of claim 61 or 62, wherein the one or more characteristics of the cancer ofthe subject comprises a cancer type, a cancer subtype, an estimate of cancer progression, a prognosis of the subject without treatment intervention, a prognosis of the subject with treatment intervention, or any combination thereof.
64. A system for determining a disease forecast of a subject having cancer, the systemcomprising: (a) a memory; and(b) one or more processors configured to execute machine-readable instructionswhich, when executed, cause the one or more processors to perform a method comprising: (c) generating a fragmentomics profile of a biological sample obtained or derivedfrom the subject, (d) mapping the fragmentomics profile to one or more reference fragmentomicsprofiles, (e) determining expression of one or more genes or determining an activation statusof one or more transcription factor of the subject; (f) determining one or more disease forecast characteristics of the biological samplebased at least in part on the mapping and the expression one or more genes or activation status of transcription factors, and (g) generating the disease forecast based at least in part on the one or more diseaseforecast characteristics.
65. The system of claim 64, wherein the one or more disease forecast characteristics compriseone or more of: a type of the cancer, a subtype of the cancer, a prognosis of the subject, an estimated survival time of the subject without treatment intervention, an estimated survival time of the subject with treatment intervention, an estimation of the subject’s response to one or more treatments, or any combination thereof.
66. A method for identifying presence or an absence of cancer in a subject, comprising:(a) assaying nucleic acid molecules from a first biological sample obtained or derived from said subject at a first time point;WSGR Docket No. 59987-717.601 (b) detecting a set of biomarkers from said nucleic acid molecules based at least in part on said assaying of (a), wherein said set of biomarkers comprise differentially expressed markers or variants; (c) obtaining a plurality of probe nucleic acids that are customized for said subject, wherein said probe nucleic acids comprises sequences of at least a subset of said set of biomarkers; (d) using said plurality of probe nucleic acids, sequencing cell free nucleic acids (cfNA) from a second biological sample obtained or derived from said subject at a second time point to detect the presence or absence of said subset of said set of biomarkers, wherein said second biological sample comprises a cerebrospinal fluid sample; (e) computer processing said subset of said set of biomarkers to detect said cancer in said subject.
67. The method of claim 66, wherein said first biological sample is selected from the groupconsisting of: a cell-free deoxyribonucleic acid (cfDNA) sample, a cell-free ribonucleic acid (cfRNA) sample, a plasma sample, a serum sample, a buffy coat sample, a peripheral blood mononuclear cell (PBMC) sample, a red blood cell sample, a urine sample, a urine cell pellet sample, a saliva sample, tissue biopsy, pleural fluid sample, peritoneal fluid sample, amniotic fluid sample, cerebrospinal fluid sample, lymphatic fluid sample, sweat sample, tear sample, semen sample, or any derivative thereof, and any combination thereof.
68. The method of any of claims 66 or 67, wherein said first biological sample comprises saidplasma sample.
69. The method of any of claims 66 or 67, wherein said first biological sample comprises saidurine sample.
70. The method of any of claims 66 or 67, wherein said first biological sample comprises saidtumor tissue sample.
71. The method of any of claims 66 to 70, wherein said first or second biological sample isobtained or derived from said subject using an ethylenediaminetetraacetic acid (EDTA) collection tube, a cell-free RNA collection tube, or a cell-free deoxyribonucleic acid (DNA) collection tube, other blood collection tube, and CTC collection tubes.
72. The method of any of claims 66 to 71, wherein said cfNA molecules comprise cell-freeDNA (cfDNA) molecules.
73. The method of any of claims 66 to 72, wherein (a) comprises subjecting said first orsecond biological sample to conditions that are sufficient to isolate, enrich, or extract said nucleic acid molecules or cfNA molecules.WSGR Docket No. 59987-717.60174. The method of any of claims 66 to 73, further comprising fractionating said firstbiological sample of said subject to obtain said nucleic acid molecules, wherein said first biological sample is a whole blood sample.
75. The method of any of claims 66 to 74, wherein at least one of said nucleic acid moleculesare assayed using sequencing to produce nucleic acid sequencing reads.
76. The method of claim 75, wherein said sequencing comprises whole exome sequencing.
77. The method of any of claims 75 or 76, further comprising filtering at least a subset of saidnucleic acid sequencing reads based on a quality score.
78. The method of any of claims 75 to 77, further comprising performing error correction onsaid nucleic acid sequencing reads using sample barcodes or molecular barcodes attached to at least one of said DNA molecules.
79. The method of any of claims 75 to 78, further comprising performing at least one ofsingle-stranded consensus calling and double-stranded consensus calling on said nucleic acid sequencing reads, thereby suppressing sequencing and PCR errors in said nucleic acid sequencing reads.
80. The method of any of claims 66 to 79, wherein said sequencing of (d) is performed at adepth of at least 100x.
81. The method of any of claims 66 to 80, wherein said sequencing of (d) is performed at adepth of at least 1,000x.
82. The method of any of claims 66 to 81, wherein said sequencing of (d) is performed at adepth of at least 10,000x.
83. The method of any of claims 66 to 82, wherein said sequencing of (d) is performed at adepth of at least 100,000x.
84. The method of any of claims 66 to 83, wherein said assaying of (a) or sequencing of (d)comprises nucleic acid amplification.
85. The method of claim 84, wherein said nucleic acid amplification comprises polymerasechain reaction (PCR) or isothermal amplification.
86. The method of any of claims 66 to 85, wherein said cancer is a brain cancer or spinecancer.
87. The method of claim 86, wherein said cancer comprises a glioma.
88. The method of any of claims 66 to 87, wherein said subject is asymptomatic for saidcancer.
89. The method of any of claims 66 to 88, wherein the method comprises detecting saidpresence or absence of cancer in said subject at an accuracy of at least about 60%, at least aboutWSGR Docket No. 59987-717.601 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
90. The method any of claims 66 to 89, wherein the method comprises detecting saidpresence or absence of cancer in said subject at a sensitivity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
91. The method any of claims 66 to 90, wherein the method comprises detecting saidpresence or absence of cancer in said subject at a specificity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
92. The method of any of claims 66 to 91, wherein the method comprises detecting saidpresence or absence of cancer in said subject at a positive predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
93. The method of any of claims 66 to 92, wherein the method comprises detecting saidpresence or absence of cancer in said subject in said subject at a negative predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
94. The method of any of claims 66 to 93, wherein said first biological sample is obtained orderived from said subject prior to said subject receiving a therapy for said cancer.
95. The method of any of claims 66 to 94, wherein said biological sample is obtained orderived from said subject during a therapy for said cancer.
96. The method of any of claims 66 to 95, wherein said biological sample is obtained orderived from said subject after receiving a therapy for said cancer.
97. The method of any one of claims 94 to 96, wherein said therapy is selected from thegroup consisting of: surgical resection, chemotherapy, radiotherapy, immunotherapy, cell therapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and a combination thereof.
98. The method of any of claims 66 to 97, further comprising identifying a clinicalintervention for said subject based at least in part on said detected presence or said absence of said cancer.
99. The method of claim 98, wherein said clinical intervention is selected from a plurality ofclinical interventions.WSGR Docket No. 59987-717.601100. The method of any of claims 98 or 99, wherein said clinical intervention is selected fromthe group consisting of: surgical resection, chemotherapy, radiotherapy, immunotherapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and a combination thereof.
101. The method of any of claims 98 to 100, further comprising administering said clinicalintervention to said subject.
102. The method of any of claims 66 to 101, wherein said plurality of probes comprise nucleicacid primers.
103. The method of any of claims 66 to 102, the plurality of probes comprise nucleic acidcapture probes.
104. The method of any of claims 66 to 103, wherein said plurality of probes have sequencecomplementarity with at least a portion of nucleic acid sequences of said set of biomarkers.
105. The method of any of claims 66 to 104, wherein said plurality of probes comprise at least2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 different probes.
106. The method of any of claims 66 to 105, wherein (d) further comprises sequencing using afixed plurality of probes wherein the probes of the fixed plurality of probes comprises probes that do not comprise sequences of said subset of said set of biomarkers.
107. The method of any of claims 66 to 106, further comprising determining a likelihood ofsaid determination of said presence or said absence of said cancer in said subject.
108. The method of any of claims 66 to 107, further comprising monitoring said presence orsaid absence of said cancer in said subject, wherein said monitoring comprises assessing said presence or said absence of said cancer in said subject at each of a plurality of time points.
109. The method of claim 108, wherein a difference in said assessment of said presence or saidabsence of said cancer in said subject among said plurality of time points is indicative of one or more clinical indications selected from the group consisting of: (i) a diagnosis of said cancer, (ii) a prognosis of said cancer, and (iii) an efficacy or non-efficacy of a course of treatment for treating said cancer of said subject.
110. The method of claim 109, wherein said prognosis comprises an expected progression-freesurvival (PFS) or overall survival (OS).
111. The method of any of claims 66 to 110, wherein said set of biomarkers from said cfNAmolecules comprise tumor-associated alterations selected from the group consisting of: single nucleotide variants (SNVs), insertions or deletions (indels), and rearrangements.
112. The method of any of claims 66 to 111, further comprising determining, among said setof biomarkers, a mutant allele frequency of a set of somatic mutations.WSGR Docket No. 59987-717.601113. The method of claim 112, further comprising determining a circulating tumor DNA(ctDNA) fraction of said cancer of said subject based at least in part on said set of mutant allele frequencies.
114. The method of any of claims 112 or 113, further comprising determining a tumormutational burden (TMB) of said cancer of said subject.
115. The method of any of claims 112 to 114, further comprising determining an abnormalityscore of said cancer of said subject based at least in part on said set of mutant allele frequencies.
116. A method for detecting a presence or an absence of a glioma in a subject, comprising:(a) assaying cell-free deoxyribonucleic acid (cfDNA) molecules from a biological sample obtained or derived from said subject, wherein the biological sample comprises cerebrospinal fluid sample; (b) detecting a set of biomarkers from said cfDNA molecules wherein said set of biomarkers comprise differentially expressed markers or variants; (c) computer processing said set of biomarkers to detect said presence or said absence of said cancer in said subject.
Citation Information
Patent Citations
Systems and methods for multi-analyte detection of cancer
WO2022212590A1
Cell-free DNA sequence data analysis method to examine nucleosome protection and chromatin accessibility
WO2022217096A2