System and method for estimating cell origin fraction using methylation information

The method addresses the challenge of estimating cell source fractions in biological samples by combining methylation data with sequencing data to map and classify cell-free fragments, improving cancer detection and monitoring.

JP7775192B2Active Publication Date: 2025-11-25GRAIL INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022530797
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-12-18
Filing Date
2020-12-18
Publication Date
2025-11-25
Estimated Expiration
2040-12-18

AI Technical Summary

Technical Problem

Existing methods lack robust techniques for determining the cell source fraction, such as tumor fraction, in biological samples using cell-free DNA, which limits the diagnostic power of epigenetic patterns in cancer detection.

Method used

A method combining methylation data with whole genome or targeted genome sequencing data to estimate the cell source fraction by mapping cell-free fragments to sequence groups and using a classifier to assign cancer states, followed by calculating representative values to determine the proportion of interest.

Benefits of technology

Enhances diagnostic power by accurately estimating the cell source fraction, enabling more precise cancer detection and monitoring through methylation patterns in cell-free DNA.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007775192000027
    Figure 0007775192000027
  • Figure 0007775192000028
    Figure 0007775192000028
  • Figure 0007775192000029
    Figure 0007775192000029
Patent Text Reader

Abstract

A method for identifying a plurality of features for estimating a cell source fraction of a subject is provided. For each training subject in a plurality of training subjects, a corresponding methylation pattern for each cell-free fragment in a corresponding plurality of training cell-free fragments and a cancer indication for the corresponding subject are obtained. Each cell-free fragment is mapped to a bin within a plurality of bins, each bin representing a portion of a human reference genome. The corresponding methylation pattern for each cell-free fragment is input to a classifier, where a cancer state of the cell-free fragment is assigned to each cell-free fragment as a function of the classifier. A measure of association between the cancer state of the subject and the cancer state of the cell-free fragment is determined for each bin. A plurality of features for estimating a cell source fraction of the subject are identified as a subset of the plurality of bins.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent Application No. 62 / 950,071, filed December 18, 2019, entitled "Systems and Methods for Estimating Cell Source Fractions using Methylation Information," the contents of which are incorporated herein by reference in their entirety for all purposes. [Technical Field]

[0002] This specification describes the use of nucleic acids of a subject, particularly cell-free nucleic acid samples, to estimate the proportion of cell origin, eg, tumor proportion, in a biological sample obtained from a subject. [Background technology]

[0003] Increasing knowledge of the molecular basis of cancer and rapid development of next-generation sequencing technologies have advanced the study of early molecular changes involved in cancer development in body fluids. Large-scale sequencing technologies such as next-generation sequencing (NGS) offer the opportunity to achieve sequencing at a cost of less than $1 USD per million bases, with actual costs reaching less than 10 US cents. Specific genetic and epigenetic changes associated with the development of these cancers have been identified in cell-free DNA (cfDNA) from plasma, serum, and urine. These changes may serve as diagnostic biomarkers for several classes of cancer (see Salvi et al., 2016, Onco Targets Ther. 9:6549-6559).

[0004] Cell-free DNA (cfDNA) is found in serum, plasma, urine, and other body fluids (Chan et al., 2003, Ann Clin Biochem. 40(Pt 2):122-130), making it a "liquid biopsy" that represents the circulating profile of certain diseases (see De Mattos-Arruda and Caldas, 2016, Mol Oncol. 10(3):464-474). This represents a potential non-invasive method for screening for various cancers.

[0005] The existence of cfDNA was demonstrated several decades ago by Mandel and Metais (Mandel and Metais, 1948, CR Seances Soc Biol Fil. 142(3-4):241-243). cfDNA originates from necrotic or apoptotic cells and is generally released by all cell types. Stroun et al. further demonstrated that specific cancer-specific alterations can be found in patient cfDNA (see Stroun et al., 1989 Oncology 1989 46(5):318-322). Many subsequent publications confirmed that cfDNA contains specific tumor-associated alterations, such as mutations, methylation, and copy number variations (CNVs), confirming the existence of circulating tumor DNA (ctDNA) (Goessl et al., 2000 Cancer Res. 60(21):5941-5945 and Frenel et al., 2015, Clin Cancer Res. 21(20):4586-4596).

[0006] While cfDNA in plasma or serum has been well characterized, urinary cfDNA (ucfDNA) has traditionally been less well characterized, but recent studies have demonstrated that ucfDNA may also be a promising biomarker source (e.g., Casadio et al. 31(8):1744-1750).

[0007] In blood, apoptosis is a frequent event that determines the amount of cfDNA. However, in cancer patients, cfDNA quantity also appears to be influenced by necrosis (see Hao et al., 2014, Br J Cancer 111(8):1482-1489 and Zonta et al., 2015 Adv Clin Chem.70:197-246). Because apoptosis appears to be the primary release mechanism, circulating cfDNA has a size distribution that reveals an enrichment of short fragments of approximately 167 base pairs, corresponding to nucleosomes generated by apoptotic cells (see Heitzer et al., 2015, Clin Chem.61(1):112-123 and Lo et al., 2010, Sci Transl Med.2(61):61ra91).

[0008] The amount of circulating cfDNA in serum and plasma appears to be significantly higher in tumor patients than in healthy controls, especially in patients with advanced-stage tumors than in early-stage tumors (see Sozzi et al., 2003, J Clin Oncol. 21(21):3902-3908, Kim et al., 2014, Ann Surg Treat Res. 86(3):136-142; and Shao et al., 2015, Oncol Lett. 10(6):3478-3482). The amount of circulating cfDNA varies more in cancer patients than in healthy individuals (see Heitzer et al., 2013, Int J Cancer. 133(2):346-356), and the amount of circulating cfDNA is affected by several physiological and pathological conditions, including pro-inflammatory diseases (see Raptis and Menard, 1980, J Clin Invest. 66(6):1391-1399, and Shapiro et al., 1983, Cancer 51(11):2116-2120).

[0009] Methylation status and other epigenetic modifications are known to correlate with the presence of several disease states, including cancer (see Jones, 2002, Oncogene 21:5358-5360). Furthermore, specific patterns of methylation have been determined to be associated with specific cancer states (see Paska and Hudler, 2015, Biochemia Medica 25(2):161-176). Warton and Samimi demonstrated that methylation patterns can also be observed in cell-free DNA (Warton and Samimi, 2015, Front Mol Biosci, 2(13) doi: 10.3389 / fmolb.2015.00013).

[0010] Given the promise of circulating cfDNA, and other forms of genotype data, as diagnostic indicators, there is a need in the art for methods to evaluate such data to identify epigenetic patterns. Summary of the Invention

[0011] The present disclosure addresses the shortcomings identified in the background by providing a robust technique for determining the cell source fraction, such as tumor fraction, in a biological sample obtained from a subject using cfDNA. The combination of methylation data with whole genome or targeted genome sequencing data provides additional diagnostic power beyond conventional screening methods.

[0012] Technical solutions (eg, computer systems, methods, and non-transitory computer-readable storage media) for addressing the above-identified problems related to analyzing datasets are provided in the present disclosure.

[0013] The following presents a summary of the invention in order to provide a basic understanding of some aspects of the invention. This summary is not an extensive overview of the invention. It is not intended to identify key / critical elements of the invention or to delineate the scope of the invention. Its sole purpose is to present some concepts of the invention in a simplified form as a prelude to the more detailed description that is presented later.

[0014] A. Embodiments in which the cell source fraction is estimated based at least in part on a subset of sequence groups identified by the proportion of cancer-derived fragments in each sequence group.

[0015] One aspect of the present disclosure provides a method for identifying multiple features for estimating a subject's cell-origin fraction. The method includes acquiring a training dataset in electronic format in a computer system having one or more processors and a memory storing one or more programs for execution by the one or more processors. The training dataset includes, for each training subject of a plurality of training subjects, a corresponding methylation pattern for each cell-free fragment among a corresponding plurality of training cell-free fragments, and b) a cancer indication for each training subject. The corresponding methylation pattern for each cell-free fragment is determined by (i) methylation sequencing of one or more nucleic acid samples containing each fragment in a corresponding biological sample obtained from each training subject, and (ii) includes a methylation state for each CpG site among a corresponding plurality of CpG sites in each fragment. The cancer state of the subject is one of a first cancer state and a second cancer state. The method further includes mapping each cell-free fragment of each of the plurality of cell-free fragments to one of a plurality of sequence groups. wherein each sequence group of the plurality of sequence groups represents a corresponding portion of the human reference genome, thereby obtaining a plurality of training sets of cell-free fragments, each training set of cell-free fragments mapped to a different sequence group of the plurality of sequence groups. The method further includes assigning a cell-free fragment cancer state to each cell-free fragment in each training set of cell-free fragments in the plurality of training sets of cell-free fragments, the cell-free fragment cancer state being a function of the output of a classifier when the methylation pattern of each cell-free fragment is input to the classifier. The cell-free fragment cancer state is one of a first cancer state and a second cancer state. The method further includes, for each sequence group of the plurality of sequence groups, determining a corresponding measure of association between (a) the subject cancer state of each training subject of the plurality of training subjects and (b) the cell-free fragment cancer state of each cell-free fragment in the corresponding training set of cell-free fragments mapped to each sequence group. In some embodiments, the association method is a correlation calculation.In some embodiments, the association method is a mutual information calculation. In some embodiments, the association method is by calculating a distance metric (e.g., Manhattan distance, maximum, normalized Euclidean distance, normalized Manhattan distance, Dice coefficient, cosine distance, or Jacquard coefficient, etc.). The method proceeds by identifying a plurality of features for estimating the proportion of the cell source of interest as a subset of a plurality of sequences. Each sequence in the subset of a plurality of sequences satisfies a selection criterion based on the corresponding relatedness measure for each sequence. For example, in some embodiments, a sequence that ranks top in relatedness measure relative to all other sequence groups is considered to satisfy the selection criterion.

[0016] In some embodiments, the method further includes estimating the cell-origin fraction for the test subject by a procedure including obtaining, in electronic form, a corresponding methylation pattern for each cell-free fragment of the test plurality of cell-free fragments. The corresponding methylation pattern for each cell-free fragment is determined by (i) methylation sequencing of one or more nucleic acid samples containing each fragment in a biological sample obtained from the test subject, and (ii) includes a methylation state for each CpG site of the corresponding plurality of CpG sites in each fragment. Each cell-free fragment of the test plurality of cell-free fragments is mapped to one of the plurality of sequence sets, thereby obtaining a plurality of test sets of cell-free fragments, each test set of cell-free fragments being mapped to a different sequence set of the plurality of sequence sets. A cancer state for each cell-free fragment in each test set of cell-free fragments is assigned to the cell-free fragment in the plurality of test sets of cell-free fragments, the cancer state being a function of the output of the classifier when the methylation pattern of the cell-free fragment is input to the classifier. A first representative value for the number of cell-free fragments is calculated from the test subject assigned a first cancer status in each test set of cell-free fragments across the subset of the plurality of sequence groups. A second representative value for the number of cell-free fragments is calculated from the test subject in each test set of cell-free fragments across the subset of the plurality of sequence groups. The cell-source fraction of the test subject is then estimated using the first and second representative values.

[0017] In some embodiments, the second cancer condition is the absence of cancer and the test subject cell source proportion comprises the test subject cell source proportion.

[0018] In some embodiments, the classifier comprises a classifier having a formula:

number

[0019] In some embodiments, the measure of relevance I is:

number

[0020] In some embodiments, the measure of association is a correlation. In some embodiments, the correlation is a Pearson correlation coefficient. In some embodiments, the correlation is performed using an adjusted correlation coefficient, a weighted correlation coefficient, a reflected correlation coefficient, or a scaled correlation coefficient.

[0021] In some embodiments, the plurality of sequence groups consists of 1,000 to 100,000 sequence groups. In some embodiments, the plurality of sequence groups consists of 15,000 to 80,000 sequence groups. In some embodiments, each sequence group of the plurality of sequence groups has, on average, 10 to 1,200 residues. In some embodiments, each sequence group of the plurality of sequence groups has, on average, 10 to 10,000 residues.

[0022] In some embodiments, the first representative value is the arithmetic mean, weighted mean, mid-range, mid-hinge, trimean, winsorized mean, average, or mode of the numbers of cell-free fragments from multiple test subjects assigned the first cancer status in each test set of cell-free fragments across a subset of multiple sequence groups.

[0023] In some embodiments, the second representative value is the arithmetic mean, weighted mean, mid-range, mid-hinge, trinomial mean, winsorized mean, average, or mode of the numbers of cell-free fragments from multiple test subjects in each test set of cell-free fragments across a subset of multiple sequence groups.

[0024] In some embodiments, estimating the cell source fraction comprises dividing the first representative value by the second representative value.

[0025] In some embodiments, the plurality of training subjects consists of between 10 training subjects and 1000 training subjects.

[0026] In some embodiments, the selection criteria specify the selection of sequences having one of the top N relatedness measures, where N is a positive integer greater than or equal to 50. In some embodiments, N is between 500 and 5000. In some embodiments, N is between 800 and 1500.

[0027] In some embodiments, the methylation sequencing is paired-end sequencing. In some embodiments, the methylation sequencing is single-read sequencing. In some embodiments, the corresponding training cell-free fragments have an average length of less than 500 nucleotides.

[0028] In some embodiments, the first cancer condition is cancer and the second cancer condition is the absence of cancer.

[0029] In some embodiments, the first cancer condition is one of adrenal cancer, biliary tract cancer, bladder cancer, bone / bone marrow cancer, brain cancer, breast cancer, cervical cancer, colon cancer, esophageal cancer, gastric cancer, head / neck cancer, hepatobiliary cancer, renal cancer, liver cancer, lung cancer, ovarian cancer, pancreatic cancer, pelvic cancer, pleural cancer, prostate cancer, kidney cancer, skin cancer, stomach cancer, testicular cancer, thymic cancer, thyroid cancer, uterine cancer, lymphoma, melanoma, multiple myeloma, or leukemia, and the second cancer condition is the absence of cancer.

[0030] In some embodiments, the first cancer condition is one of a stage of adrenal cancer, a stage of biliary tract cancer, a stage of bladder cancer, a stage of bone / bone marrow cancer, a stage of brain cancer, a stage of breast cancer, a stage of cervical cancer, a stage of colon cancer, a stage of esophageal cancer, a stage of gastric cancer, a stage of head / neck cancer, a stage of hepatobiliary cancer, a stage of renal cancer, a stage of liver cancer, a stage of lung cancer, a stage of ovarian cancer, a stage of pancreatic cancer, a stage of pelvic cancer, a stage of pleural cancer, a stage of prostate cancer, a stage of renal cancer, a stage of skin cancer, a stage of stomach cancer, a stage of testicular cancer, a stage of thymic cancer, a stage of thyroid cancer, a stage of uterine cancer, a stage of lymphoma, a stage of melanoma, a stage of multiple myeloma, or a stage of leukemia, and the second cancer condition is the absence of cancer.

[0031] In some embodiments, the methylation sequencing is whole-genome methylation sequencing. In some embodiments, the methylation sequencing is targeted sequencing using a plurality of nucleic acid probes, and each sequence group of the plurality of sequence groups is associated with at least one nucleic acid probe of the plurality of nucleic acid probes.

[0032] In some embodiments, the plurality of nucleic acid probes comprises 1,000 or more nucleic acid probes, 2,000 or more nucleic acid probes, 3,000 or more nucleic acid probes, 5,000 or more nucleic acid probes, 10,000 or more nucleic acid probes, or between 1,000 and 30,000 nucleic acid probes.

[0033] In some embodiments, each sequence in the plurality of sequences comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more CpG sites. In some embodiments, each sequence in the plurality of sequences comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more contiguous CpG sites. In some embodiments, each sequence in the plurality of sequences consists of between 2 and 100 contiguous CpG sites in the human reference genome.

[0034] In some embodiments, the corresponding biological sample is a liquid biological sample. In some embodiments, the corresponding biological sample is a blood sample. In some embodiments, the corresponding biological sample comprises blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid of the training subject. In some embodiments, the corresponding biological sample consists of blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid of the training subject.

[0035] In some embodiments, the methylation state of each CpG site among the corresponding plurality of CpG sites in each fragment is a methylated state if methylation sequencing determines that each CpG site is methylated, an unmethylated state if methylation sequencing determines that each CpG site is unmethylated, or is flagged as "other" if the methylation state of each CpG site cannot be called as methylated or unmethylated.

[0036] In some embodiments, methylation sequencing detects one or more 5-methylcytosines (5mC) and / or 5-hydroxymethylcytosines (5hmC) in each fragment.

[0037] In some embodiments, methylation sequencing comprises converting one or more unmethylated cytosines or one or more methylated cytosines into one or more corresponding uracils in the sequence reads of each fragment.In some embodiments, one or more uracils are detected as one or more corresponding thymines during methylation sequencing.In some embodiments, the conversion of one or more unmethylated cytosines or one or more methylated cytosines comprises chemical conversion, enzymatic conversion, or a combination thereof.

[0038] In some embodiments, the first model is a first mixed model including a first plurality of sub-models, and the second model is a second mixed model including a second plurality of sub-models, each sub-model of the first and second plurality of sub-models representing an independent corresponding methylation model of a source of cell-free fragments in a corresponding biological sample.

[0039] In some embodiments, each of the independent corresponding methylation models is one of a binomial model, a beta-binomial model, an independent site model, or a Markov model.

[0040] In some embodiments, two or more sub-models of the first plurality of sub-models are independent region models, and two or more sub-models of the second plurality of sub-models are independent region models.

[0041] In some embodiments, the method further comprises applying one or more filter conditions to the plurality of cell-free fragments.

[0042] In some embodiments, one of the filter conditions of the one or more filter conditions is applying a p-value threshold to a corresponding methylation pattern of each cell-free fragment of the plurality of cell-free fragments, where the p-value threshold is representative of the frequency with which the methylation pattern is observed in a cohort of non-cancer subjects.

[0043] In some embodiments, the p-value threshold is between 0.001 and 0.20.

[0044] In some embodiments, the cohort comprises at least 20 subjects and the plurality of cell-free fragments comprises at least 10,000 different corresponding methylation patterns.

[0045] In some embodiments, the p-value threshold is met for a methylation pattern from a subject if the corresponding methylation pattern for each cell-free fragment among the plurality of cell-free fragments has a p-value of 0.10 or less, 0.05 or less, or 0.01 or less.

[0046] In some embodiments, one of the one or more filter conditions is to apply a requirement that each cell-free fragment of the plurality of cell-free fragments is represented by a threshold number of sequence reads in a corresponding plurality of sequence reads measured from one or more nucleic acid samples that included each fragment in a corresponding biological sample.

[0047] In some embodiments, the threshold number is 2, 3, 4, 5, 6, 7, 8, 9, 10, or an integer between 10 and 100.

[0048] In some embodiments, one of the one or more filter conditions is to apply the requirement that each cell-free fragment of the plurality of cell-free fragments is represented by a threshold number of cell-free nucleic acids in one or more nucleic acid samples that contain the respective fragment in the corresponding biological sample.

[0049] In some embodiments, the threshold number is 2, 3, 4, 5, 6, 7, 8, 9, 10, or an integer between 10 and 100.

[0050] In some embodiments, one of the filter conditions of the one or more filter conditions is to apply a requirement that each cell-free fragment of the plurality of cell-free fragments has a threshold number of CpG sites.

[0051] In some embodiments, the threshold number of CpG sites is at least 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 CpG sites.

[0052] In some embodiments, one of the filter conditions of the one or more filter conditions requires that each cell-free fragment of the plurality of cell-free fragments has a length that is less than a threshold number of base pairs.

[0053] In some embodiments, the threshold number of base pairs is 1,000, 2,000, 3,000, or 4,000 contiguous base pairs in length.

[0054] In some embodiments, the method includes repeating the obtaining, mapping, assigning, calculating first and second representative values, and estimating a cell source proportion for the test subject at each of a plurality of time points across an epoch, thereby obtaining a corresponding cell source proportion for the test subject at each time point across the plurality of cell source proportions, and using the plurality of cell source proportions to determine a state or progression of the disease state of the test subject during the epoch in terms of an increase or decrease in the first cell source proportion across the epoch.

[0055] In some embodiments, an epoch is a period of several months, and each time point in the plurality of time points is a different time point within the period of several months.

[0056] In some embodiments, the period of several months is less than four months.

[0057] In some embodiments, the epoch is a period of several years, and each of the plurality of time points is a different time point within the period of several years.

[0058] In some embodiments, the period of several years is between 2 and 10 years.

[0059] In some embodiments, an epoch is a period of several hours, and each time point in the plurality of time points is a different time point within the period of several hours.

[0060] In some embodiments, the period of several hours is between 1 hour and 6 hours.

[0061] In some embodiments, the method further comprises altering the diagnosis of the test subject if the subject's first cell source proportion is observed to change by a threshold amount over the epoch.

[0062] In some embodiments, the method further comprises altering the prognosis of the test subject if the subject's first cell source proportion is observed to change by a threshold amount over the epoch.

[0063] In some embodiments, the method further comprises altering the treatment of the test subject if the subject's first cell source proportion is observed to change by a threshold amount over the epoch.

[0064] In some embodiments, the threshold is greater than 10%, greater than 20%, greater than 30%, greater than 40%, greater than 50%, greater than 2-fold, greater than 3-fold, or greater than 5-fold.

[0065] In some embodiments, the tumor ratio tested is between 0.003 and 1.0.

[0066] In some embodiments, the method further comprises administering a treatment regimen to the test subject based at least in part on the value of the cell source fraction of the test subject.

[0067] In some embodiments, the treatment regimen comprises administering a cancer drug to the test subject.

[0068] In some embodiments, the cancer drug is a hormone, immunotherapy, x-ray, or anti-cancer drug.

[0069] In some embodiments, the cancer agent is lenalidomide, pembrolizumab, trastuzumab, bevacizumab, rituximab, ibrutinib, human papillomavirus tetravalent (types 6, 11, 16, and 18) vaccine, pertuzumab, pemetrexed, nilotinib, denosumab, abiraterone acetate, promacta, imatinib, everolimus, palbociclib, erlotinib, bortezomib, or generic equivalents thereof.

[0070] In some embodiments, the test subject is being treated with a cancer drug, and the method further comprises using the cell origin percentage of the test subject to assess the test subject's response to the cancer drug.

[0071] In some embodiments, the cancer drug is a hormone, immunotherapy, x-ray, or anti-cancer drug.

[0072] In some embodiments, the cancer agent is lenalidomide, pembrolizumab, trastuzumab, bevacizumab, rituximab, ibrutinib, human papillomavirus tetravalent (types 6, 11, 16, and 18) vaccine, pertuzumab, pemetrexed, nilotinib, denosumab, abiraterone acetate, promacta, imatinib, everolimus, palbociclib, erlotinib, bortezomib, or generic equivalents thereof.

[0073] In some embodiments, the test subject is being treated with a cancer drug, and the method further includes using the cell origin percentage of the test subject to determine whether to increase or discontinue the cancer drug in the test subject.

[0074] In some embodiments, the test subject has undergone surgical intervention to address cancer, and the method further comprises using the test subject's cell source fraction to assess the test subject's condition in response to the surgical intervention.

[0075] In some embodiments, the sequences in the plurality of sequences correspond to genomic regions listed in one or more of Tables 1-24 of International Patent Application No. PCT / US2019 / 025358 (published as WO 2019 / 195268), Lists 1-8 of International Patent Application No. PCT / US2019 / 053509 (published as WO 2020 / 069350), and / or Lists 1-16 of International Patent Application No. PCT / US2020 / 015082 (published as WO 2020 / 154682), each of which is incorporated by reference in its entirety.

[0076] In some embodiments, the sequences in the plurality of sequences map to at least 30% of the genomic regions listed in one or more of Tables 1-24 of International Patent Application No. PCT / US2019 / 025358 (published as WO 2019 / 195268), Lists 1-8 of International Patent Application No. PCT / US2019 / 053509 (published as WO 2020 / 069350), and / or Lists 1-16 of International Patent Application No. PCT / US2020 / 015082 (published as WO 2020 / 154682).

[0077] In some embodiments, the sequences in the plurality of sequences map to at least 50-95% of the genomic regions listed in one or more of Tables 1-24 of International Patent Application No. PCT / US2019 / 025358 (published as WO 2019 / 195268), Lists 1-8 of International Patent Application No. PCT / US2019 / 053509 (published as WO 2020 / 069350), and / or Lists 1-16 of International Patent Application No. PCT / US2020 / 015082 (published as WO 2020 / 154682).

[0078] In some embodiments, the sequences in the plurality of sequences are mapped to 1-10 unique corresponding genomic regions in one or more of Tables 1-24 of International Patent Application No. PCT / US2019 / 025358 (published as WO 2019 / 195268), Lists 1-8 of International Patent Application No. PCT / US2019 / 053509 (published as WO 2020 / 069350), and / or Lists 1-16 of International Patent Application No. PCT / US2020 / 015082 (published as WO 2020 / 154682).

[0079] In some embodiments, the sequences in the plurality of sequences map to a single, unique corresponding genomic region in one or more of Tables 1-24 of International Patent Application No. PCT / US2019 / 025358 (published as WO 2019 / 195268), Lists 1-8 of International Patent Application No. PCT / US2019 / 053509 (published as WO 2020 / 069350), and Lists 1-16 of International Patent Application No. PCT / US2020 / 015082 (published as WO 2020 / 154682).

[0080] In some embodiments, for each training subject of the plurality of training subjects, the training plurality of cell-free fragments comprises at least 100,000 cell-free fragments.

[0081] In some embodiments, for each training subject of the plurality of training subjects, the training plurality of cell-free fragments comprises at least 100,000 cell-free fragments.

[0082] In some embodiments, for each training subject of the plurality of training subjects, the training plurality of cell-free fragments comprises at least 1 million cell-free fragments.

[0083] In some embodiments, each sequence group of the plurality of sequence groups consists of fewer than 100 nucleic acid residues, fewer than 500 nucleic acid residues, fewer than 1000 nucleic acid residues, fewer than 2500 nucleic acid residues, fewer than 5000 nucleic acid residues, fewer than 10,000 nucleic acid residues, fewer than 25,000 nucleic acid residues, fewer than 50,000 nucleic acid residues, fewer than 100,000 nucleic acid residues, fewer than 250,000 nucleic acid residues, or fewer than 500,000 nucleic acid residues.

[0084] Another aspect of the present disclosure provides a computer system for estimating a subject's cell-origin fraction. The computer system includes one or more processors and a memory storing one or more programs executed by the one or more processors. The one or more programs include instructions in electronic form for obtaining a training dataset. The training dataset includes, for each training subject of a plurality of training subjects, a) a corresponding methylation pattern for each cell-free fragment among a corresponding plurality of training cell-free fragments, and b) a cancer indication for each training subject. The corresponding methylation pattern for each cell-free fragment is determined by (i) methylation sequencing of one or more nucleic acid samples containing each fragment in a corresponding biological sample obtained from each training subject, and (ii) includes a methylation state for each CpG site among a corresponding plurality of CpG sites in each fragment. The subject's cancer state is one of a first cancer state and a second cancer state. The one or more programs further include instructions for mapping each cell-free fragment of each of the plurality of cell-free fragments to one of a plurality of sequence groups. wherein each sequence group of the plurality of sequence groups represents a corresponding portion of the human reference genome, thereby resulting in a plurality of training sets of cell-free fragments, each training set of cell-free fragments mapped to a different sequence group of the plurality of sequence groups. The one or more programs further include instructions for assigning a cell-free fragment cancer state to each cell-free fragment in each training set of cell-free fragments of the plurality of training sets of cell-free fragments, the cell-free fragment cancer state being a function of the output of the classifier when the methylation pattern of each cell-free fragment is input to the classifier. The cell-free fragment cancer state is one of a first cancer state and a second cancer state. The one or more programs further include instructions for, for each sequence group of the plurality of sequence groups, determining a corresponding measure of association I between (a) the subject cancer state of each training subject of the plurality of training subjects and (b) the cancer state of each cell-free fragment in the corresponding training set of cell-free fragments mapped to each sequence group.The one or more programs further comprise instructions for identifying a plurality of features for estimating a cell source fraction of a subject as a subset of the plurality of sequences, each sequence in the subset of the plurality of sequences satisfying a selection criterion based on a corresponding measure of relatedness for each sequence.

[0085] Another aspect of the present disclosure provides a computer system as disclosed above, wherein the one or more programs further include instructions for performing any of the methods disclosed herein, either singly or in combination.

[0086] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium storing one or more programs for estimating a subject's cell-origin fraction. The one or more programs are configured to be executed by a computer. The one or more programs include instructions in electronic form for obtaining a training dataset. The training dataset includes, for each training subject of a plurality of training subjects, a) a corresponding methylation pattern for each cell-free fragment among a corresponding plurality of training cell-free fragments, and b) a cancer indication for each training subject. The corresponding methylation pattern for each cell-free fragment is (i) determined by methylation sequencing of one or more nucleic acid samples containing each fragment in a corresponding biological sample obtained from each training subject, and (ii) includes a methylation state for each CpG at a corresponding plurality of CpG sites in each fragment. The subject's cancer state is one of a first cancer state and a second cancer state. The one or more programs include instructions for mapping each cell-free fragment of each of the plurality of cell-free fragments to one of a plurality of sequence groups. wherein each sequence group of the plurality of sequence groups represents a corresponding portion of the human reference genome, thereby resulting in a plurality of training sets of cell-free fragments, each training set of cell-free fragments mapped to a different sequence group of the plurality of sequence groups. The one or more programs further include instructions for assigning a cell-free fragment cancer state to each cell-free fragment in each training set of cell-free fragments of the plurality of training sets of cell-free fragments, the cell-free fragment cancer state being a function of the output of the classifier when the methylation pattern of each cell-free fragment is input to the classifier. The cell-free fragment cancer state is one of a first cancer state and a second cancer state. The one or more programs further include instructions for, for each sequence group of the plurality of sequence groups, determining a corresponding measure of association I between (a) the subject cancer state of each training subject of the plurality of training subjects and (b) the cell-free fragment cancer state of each cell-free fragment of the corresponding training set of cell-free fragments mapped to each sequence group.The one or more programs include instructions for identifying a plurality of features for estimating a cell source fraction of interest as a subset of a plurality of sequences, each sequence in the subset of the plurality of sequences meeting a selection criterion based on a corresponding measure of relatedness for each sequence.

[0087] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium as disclosed above, wherein one or more programs further include instructions for performing any of the methods disclosed herein, singly or in combination.

[0088] B. An embodiment directed to determining the cell origin fraction of a test subject using methylation data obtained from cell-free DNA.

[0089] Another aspect of the present disclosure provides a method for estimating a subject's cell-origin fraction. The method includes, in a computer system having one or more processors and a memory storing one or more programs for execution by the one or more processors, obtaining, in an electronic format, a corresponding methylation pattern for each cell-free fragment among a plurality of cell-free fragments. The corresponding methylation pattern for each cell-free fragment (i) is determined by methylation sequencing of one or more nucleic acid samples containing each fragment in a biological sample obtained from the subject, and (ii) includes a methylation state for each CpG site among a corresponding plurality of CpG sites in each fragment. The method also includes mapping each cell-free fragment of the plurality of cell-free fragments to one of a plurality of sequence groups, thereby obtaining a plurality of cell-free fragment sets. Each cell-free fragment set maps to a different sequence group among the plurality of sequence groups. The method also includes assigning a cancer state to each cell-free fragment in each of the plurality of cell-free fragment sets, the cancer state being a function of an output of a classifier when the methylation pattern of each cell-free fragment is input to the classifier. The cancer state of the cell-free fragments is one of a first cancer state and a second cancer state. The method proceeds by calculating a first representative value of the number of cell-free fragments from subjects assigned the first cancer state for each cell-free fragment set across the plurality of sequence groups, and calculating a second representative value of the number of cell-free fragments from subjects for each cell-free fragment set across the plurality of sequence groups. The method further includes estimating the cell origin fraction of the subject using the first representative value and the second representative value.

[0090] In some embodiments, the plurality of sequence groups consists of between 1000 sequence groups, In some embodiments, the plurality of sequence groups consists of between 15,000 sequence groups and 80,000 sequence groups.

[0091] In some embodiments, each of the plurality of sequence groups has, on average, 10 to 1200 residues. In some embodiments, each of the plurality of sequence groups has, on average, 10 to 10000 residues.

[0092] In some embodiments, the first representative value is the arithmetic mean, weighted mean, mid-range, mid-hinge, trinomial mean, winsorized mean, average, or mode of the number of cell-free fragments from subjects assigned the first cancer status in each set of cell-free fragments across the plurality of sequence groups. In some embodiments, the second representative value is the arithmetic mean, weighted mean, mid-range, mid-hinge, trinomial mean, winsorized mean, average, or mode of the number of cell-free fragments from subjects in each set of cell-free fragments across the plurality of sequence groups.

[0093] In some embodiments, estimating the cell source fraction comprises dividing the first representative value by the second representative value.

[0094] In some embodiments, the methylation sequencing is paired-end sequencing. In some embodiments, the methylation sequencing is single-read sequencing.

[0095] In some embodiments, the plurality of cell-free fragments has an average length of less than 500 nucleotides.

[0096] In some embodiments, the first cancer condition is cancer and the second cancer condition is the absence of cancer.

[0097] In some embodiments, the first cancer condition is one of adrenal cancer, biliary tract cancer, bladder cancer, bone / bone marrow cancer, brain cancer, breast cancer, cervical cancer, colon cancer, esophageal cancer, gastric cancer, head / neck cancer, hepatobiliary cancer, renal cancer, liver cancer, lung cancer, ovarian cancer, pancreatic cancer, pelvic cancer, pleural cancer, prostate cancer, kidney cancer, skin cancer, stomach cancer, testicular cancer, thymic cancer, thyroid cancer, uterine cancer, lymphoma, melanoma, multiple myeloma, or leukemia, and the second cancer condition is the absence of cancer.

[0098] In some embodiments, the first cancer condition is one of a stage of adrenal cancer, a stage of biliary tract cancer, a stage of bladder cancer, a stage of bone / bone marrow cancer, a stage of brain cancer, a stage of breast cancer, a stage of cervical cancer, a stage of colon cancer, a stage of esophageal cancer, a stage of gastric cancer, a stage of head / neck cancer, a stage of hepatobiliary cancer, a stage of renal cancer, a stage of liver cancer, a stage of lung cancer, a stage of ovarian cancer, a stage of pancreatic cancer, a stage of pelvic cancer, a stage of pleural cancer, a stage of prostate cancer, a stage of renal cancer, a stage of skin cancer, a stage of stomach cancer, a stage of testicular cancer, a stage of thymic cancer, a stage of thyroid cancer, a stage of uterine cancer, a stage of lymphoma, a stage of melanoma, a stage of multiple myeloma, or a stage of leukemia, and the second cancer condition is the absence of cancer.

[0099] In some embodiments, the methylation sequencing is whole genome methylation sequencing. In some embodiments, the methylation sequencing is targeted sequencing using a plurality of nucleic acid probes, and each sequence group of the plurality of sequence groups is associated with at least one corresponding nucleic acid probe of the plurality of nucleic acid probes.

[0100] In some embodiments, the plurality of nucleic acid probes comprises 1,000 or more nucleic acid probes, 2,000 or more nucleic acid probes, 3,000 or more nucleic acid probes, 5,000 or more nucleic acid probes, 10,000 or more nucleic acid probes, or between 1,000 and 30,000 nucleic acid probes.

[0101] In some embodiments, each sequence in the plurality of sequences comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more CpG sites. In some embodiments, each sequence in the plurality of sequences comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more contiguous CpG sites. In some embodiments, each sequence in the plurality of sequences consists of between 2 and 100 contiguous CpG sites in the human reference genome.

[0102] In some embodiments, the biological sample is a liquid biological sample. In some embodiments, the biological sample is a blood sample. In some embodiments, the biological sample comprises blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid from a subject. In some embodiments, the biological sample consists of blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid from a subject.

[0103] In some embodiments, the methylation status of each CpG site among the corresponding plurality of CpG sites in each fragment is: methylated if methylation sequencing determines that each CpG site is methylated; unmethylated if methylation sequencing determines that each CpG site is unmethylated; or flagged as "other" if the methylation status of each CpG site cannot be called methylated or unmethylated.

[0104] In some embodiments, methylation sequencing detects one or more 5-methylcytosines (5mC) and / or 5-hydroxymethylcytosines (5hmC) in each fragment.

[0105] In some embodiments, methylation sequencing comprises converting one or more unmethylated cytosines or one or more methylated cytosines into one or more corresponding uracils in the sequence reads of each fragment.In some embodiments, one or more uracils are detected as one or more corresponding thymines during methylation sequencing.In some embodiments, the conversion of one or more unmethylated cytosines or one or more methylated cytosines comprises chemical conversion, enzymatic conversion, or a combination thereof.

[0106] In some embodiments, the first model is a first mixed model including a first plurality of sub-models, and the second model is a second mixed model including a second plurality of sub-models, each sub-model of the first and second plurality of sub-models representing an independent corresponding methylation model for a cell-free fragment source in a corresponding biological sample.

[0107] In some embodiments, each of the independent corresponding methylation models is one of a binomial model, a beta-binomial model, an independent site model, or a Markov model.

[0108] In some embodiments, two or more submodels in the first plurality of submodels are independent site models and two or more submodels in the second plurality of submodels are independent site models.

[0109] In some embodiments, the method further comprises applying one or more filter conditions to the plurality of cell-free fragments.

[0110] In some embodiments, one of the filter conditions of the one or more filter conditions is applying a p-value threshold to a corresponding methylation pattern of each cell-free fragment of the plurality of cell-free fragments, where the p-value threshold is representative of the frequency with which the methylation pattern is observed in a cohort of non-cancer subjects.

[0111] In some embodiments, the p-value threshold is between 0.001 and 0.20. In some embodiments, the p-value threshold is between 0.01 and 0.10. In some embodiments, the p-value threshold is greater than 0.001, greater than 0.005, greater than 0.010, greater than 0.020, greater than 0.030, greater than 0.040, greater than 0.050, greater than 0.060, greater than 0.070, greater than 0.080, greater than 0.090, or greater than 0.010.

[0112] In some embodiments, the cohort comprises at least 20, at least 30, at least 50, at least 100, at least 500, or at least 1000 subjects. In some embodiments, the plurality of cell-free fragments comprises at least 300, at least 500, at least 1000, at least 5000, at least 8000, or at least 10000 different corresponding methylation patterns.

[0113] In some embodiments, the p-value threshold is met for a methylation pattern from a subject if the corresponding methylation pattern for each cell-free fragment among the plurality of cell-free fragments has a p-value of 0.10 or less, 0.05 or less, or 0.01 or less.

[0114] In some embodiments, one of the one or more filter conditions is to require that each cell-free fragment of the plurality of cell-free fragments be represented by a threshold number of sequence reads in a corresponding plurality of sequence reads measured from one or more nucleic acid samples containing the fragment in a corresponding biological sample, in some embodiments, the threshold number is 2, 3, 4, 5, 6, 7, 8, 9, 10, or an integer between 10 and 100.

[0115] In some embodiments, one of the one or more filter conditions is to require that each cell-free fragment of the plurality of cell-free fragments be represented by a threshold number of cell-free nucleic acids in one or more nucleic acid samples that contain the respective fragment in a corresponding biological sample, in some embodiments, the threshold number is 2, 3, 4, 5, 6, 7, 8, 9, 10, or an integer between 10 and 100.

[0116] In some embodiments, the filter condition in the one or more filter conditions is to require that each cell-free fragment of the plurality of cell-free fragments have a threshold number of CpG sites. In some embodiments, the threshold number of CpG sites is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 CpG sites.

[0117] In some embodiments, the filter condition in the one or more filter conditions is a requirement that each cell-free fragment of the plurality of cell-free fragments have a length of less than a threshold number of base pairs, in some embodiments, the threshold number of base pairs is 1000, 2000, 3000, or 4000 contiguous base pairs in length.

[0118] In some embodiments, a single filter condition is applied. In some embodiments, two filter conditions are applied. In some embodiments, three filter conditions are applied. In some embodiments, four filter conditions are applied.

[0119] In some embodiments, the method further includes repeating the obtaining, mapping, assigning, calculating the first and second representative values, and estimating a cell source fraction for the test subject at each of the plurality of time points across the epoch to obtain a corresponding cell source fraction for the test subject at each time point across the plurality of cell source fractions, which in some embodiments are used to determine the status or progression of the disease state of the test subject during the epoch in terms of an increase or decrease in the first cell source fraction across the epoch.

[0120] In some embodiments, each epoch is several months in duration, and each of the plurality of time points is a different time point within the several month period. In some embodiments, the several month period is less than four months. In some embodiments, each epoch is one month in length. In some embodiments, each epoch is two months in length. In some embodiments, each epoch is three months in length. In some embodiments, each epoch is four months in length. In some embodiments, each epoch is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 months in length.

[0121] In some embodiments, the epoch is a period of several years, and each of the plurality of time points is a different time point within the period of several years. In some embodiments, the period of several years is 1 to 10 years. In some embodiments, the period of several years is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 years. In some embodiments, the epoch is 1 to 30 years.

[0122] In some embodiments, an epoch is a period of several hours, and each of the plurality of time points is a different time point within the period of several hours. In some embodiments, the period of several hours is 1 hour to 24 hours. In some embodiments, the period of several hours is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 hours.

[0123] In some embodiments, the method further includes changing the diagnosis of the subject if the test subject's first cell source proportion is observed to change by a threshold amount over epochs. For example, in some embodiments, the diagnosis is changed from having cancer to being in remission. As another example, in some embodiments, the diagnosis is changed from not having cancer to having cancer. As another example, in some embodiments, the diagnosis is changed from having first stage cancer to having second stage cancer. As another example, in some embodiments, the diagnosis is changed from having second stage cancer to having third stage cancer. As yet another example, in some embodiments, the diagnosis is changed from having third stage cancer to having fourth stage cancer. As yet another example, in some embodiments, the diagnosis is changed from having cancer that has not metastasized to having cancer that has metastasized.

[0124] In some embodiments, the method further includes altering the subject's prognosis when the test subject's first cell source proportion is observed to change by a threshold amount over epochs. For example, in some embodiments, the prognosis includes a life expectancy, and the prognosis is altered from a first life expectancy to a second life expectancy, the first and second life expectancies differing in duration. In some embodiments, the altered prognosis increases the subject's life expectancy. In some embodiments, the altered prognosis decreases the subject's life expectancy.

[0125] In some embodiments, the method further includes modifying the subject's treatment if the test subject's first cell source proportion is observed to change by a threshold amount over time. In some embodiments, the modification of treatment includes initiating a cancer drug therapy, increasing the dosage of the cancer drug therapy, stopping the cancer drug therapy, or decreasing the dosage of the cancer drug therapy. In some embodiments, the modification of treatment includes starting or terminating treatment of the subject with lenalidomide, pembrolizumab, trastuzumab, bevacizumab, rituximab, ibrutinib, human papillomavirus tetravalent (types 6, 11, 16, and 18) vaccine, pertuzumab, pemetrexed, nilotinib, denosumab, abiraterone acetate, promacta, imatinib, everolimus, palbociclib, erlotinib, bortezomib, or generic equivalents thereof. In some embodiments, the treatment change includes increasing or decreasing the dosage of lenalidomide, pembrolizumab, trastuzumab, bevacizumab, rituximab, ibrutinib, human papillomavirus tetravalent (types 6, 11, 16, and 18) vaccine, pertuzumab, pemetrexed, nilotinib, denosumab, abiraterone acetate, promacta, imatinib, everolimus, palbociclib, erlotinib, bortezomib, or generic equivalents thereof administered to the subject. In some embodiments, the threshold is greater than 10%, greater than 20%, greater than 30%, greater than 40%, greater than 50%, greater than 2-fold, greater than 3-fold, or greater than 5-fold.

[0126] In some embodiments, the tumor fraction of the test subject is between 0.003 and 1.0. In some embodiments, the tumor fraction of the test subject is between 0.005 and 0.80. In some embodiments, the tumor fraction of the test subject is between 0.01 and 0.70. In some embodiments, the tumor fraction of the test subject is between 0.05 and 0.60.

[0127] In some embodiments, the method further comprises administering a therapeutic regimen to the test subject based at least in part on the value of the cell origin fraction of the test subject. In some embodiments, the therapeutic regimen comprises administering a cancer drug to the test subject. In some embodiments, the cancer drug is a hormone, immunotherapy, x-ray, or anticancer drug. In some embodiments, the cancer drug is lenalidomide, pembrolizumab, trastuzumab, bevacizumab, rituximab, ibrutinib, human papillomavirus tetravalent (types 6, 11, 16, and 18) vaccine, pertuzumab, pemetrexed, nilotinib, denosumab, abiraterone acetate, promacta, imatinib, everolimus, palbociclib, erlotinib, bortezomib, or generic equivalents thereof.

[0128] In some embodiments, the test subject is being treated with a cancer drug, and the method further includes using the cell origin percentage of the test subject to evaluate the test subject's response to the cancer drug. In some embodiments, the cancer drug is a hormone, immunotherapy, x-ray, or anticancer drug. In some embodiments, the cancer drug is lenalidomide, pembrolizumab, trastuzumab, bevacizumab, rituximab, ibrutinib, human papillomavirus tetravalent (types 6, 11, 16, and 18) vaccine, pertuzumab, pemetrexed, nilotinib, denosumab, abiraterone acetate, promacta, imatinib, everolimus, palbociclib, erlotinib, bortezomib, or a generic equivalent thereof.

[0129] In some embodiments, the test subject is being treated with a cancer drug, and the method further includes using the cell source fraction of the test subject to determine whether to increase or discontinue the cancer drug in the test subject. For example, in some embodiments, observing a cell source fraction of at least a threshold value (e.g., greater than 0.05, greater than 0.10, greater than 0.15, greater than 0.20, greater than 0.25, or greater than 0.30, etc.) is used as a basis for increasing the cancer drug (e.g., increasing the dosage or increasing the radiation level of radiotherapy) in the test subject. In some embodiments, observing a cell source fraction below a threshold value (e.g., less than 0.05, 0.10, 0.15, 0.20, 0.25, or 0.30, etc.) is used as a basis for discontinuing the use of the cancer drug in the test subject.

[0130] In some embodiments, the test subject has undergone surgical intervention to address cancer, and the method further comprises using the cell source fraction of the test subject to assess the test subject's status in response to the surgical intervention, hi some embodiments, the status is a metric based on the cell source fraction calculated using the methods provided herein.

[0131] In some embodiments, a group of sequences in the plurality of groups of sequences corresponds to a single genomic region listed in one or more of Tables 1-24 of International Patent Application No. PCT / US2019 / 025358 (published as WO 2019 / 195268), Lists 1-8 of International Patent Application No. PCT / US2019 / 053509 (published as WO 2020 / 069350), and / or Lists 1-16 of International Patent Application No. PCT / US2020 / 015082 (published as WO 2020 / 154682), each of which is incorporated by reference in its entirety.

[0132] In some embodiments, the sequences in the plurality of sequences correspond to combinations of genomic regions listed in one or more of Tables 1-24 of International Patent Application No. PCT / US2019 / 025358 (published as WO 2019 / 195268), Lists 1-8 of International Patent Application No. PCT / US2019 / 053509 (published as WO 2020 / 069350), and / or Lists 1-16 of International Patent Application No. PCT / US2020 / 015082 (published as WO 2020 / 154682), each of which is incorporated by reference in its entirety. For example, in some embodiments, the groups of sequences in the plurality of groups of sequences include one, two, three, four, five, or more than five regions listed in Tables 1-24 of WO 2019 / 195268, Lists 1-8 of WO 2020 / 069350, and / or Lists 1-16 of WO 2020 / 154682.

[0133] In some embodiments, the sequences in the plurality of sequences map to at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99% or 100% of the genomic regions listed in one or more of Tables 1-24 of WO2019 / 195268, Lists 1-8 of WO2020 / 069350, and / or Lists 1-16 of WO2020 / 154682.

[0134] In some embodiments, the sequences in the plurality of sequences map to at least 50-95% of the genomic regions listed in one or more of Tables 1-24 of WO 2019 / 195268, Lists 1-8 of WO 2020 / 069350, and / or Lists 1-16 of WO 2020 / 154682.

[0135] In some embodiments, the sequences in the plurality of sequences map to 1 to 10 unique corresponding genomic regions in one or more of Tables 1-24 of WO 2019 / 195268, Lists 1-8 of WO 2020 / 069350, and / or Lists 1-16 of WO 2020 / 154682.

[0136] In some embodiments, sequences in the plurality of sequences map to a single, unique corresponding genomic region in one or more of Tables 1-24 of WO 2019 / 195268, Lists 1-8 of WO 2020 / 069350, and / or Lists 1-16 of WO 2020 / 154682.

[0137] In some embodiments, for each subject, the plurality of cell-free fragments comprises at least 10,000, 15,000, 20,000, 25,000, 50,000, 100,000, 200,000, 300,000, 500,000, or 1 million cell-free fragments. In some embodiments, for each subject, the plurality of cell-free fragments comprises at least 1 million cell-free fragments.

[0138] In some embodiments, each sequence group of the plurality of sequence groups contains fewer than 100 nucleic acid residues, fewer than 500 nucleic acid residues, fewer than 1000 nucleic acid residues, fewer than 2500 nucleic acid residues, fewer than 5000 nucleic acid residues, fewer than 10,000 nucleic acid residues, fewer than 25,000 nucleic acid residues, fewer than 50,000 nucleic acid residues, fewer than 100,000 nucleic acid residues, fewer than 250,000 nucleic acid residues, or fewer than 500,000 nucleic acid residues.

[0139] In some embodiments, each sequence group of the plurality of sequence groups comprises (i) 100 nucleic acid residues. (ii) contains 500, 1000, 2500, 5000, 10,000, 25,000, 50,000, 100,000, 250,000, or 500,000 nucleic acid residues.

[0140] Another aspect of the present disclosure provides a computer system for estimating a subject's cell-origin fraction. The computer system includes one or more processors and a memory storing one or more programs executed by the one or more processors. The one or more programs include instructions for obtaining, in an electronic format, a corresponding methylation pattern for each cell-free fragment among a plurality of cell-free fragments. The corresponding methylation pattern for each cell-free fragment (i) is determined by methylation sequencing of one or more nucleic acid samples containing each fragment in a biological sample obtained from the subject, and (ii) includes a methylation state for each CpG site among a corresponding plurality of CpG sites in each fragment. The one or more programs further include instructions for mapping each cell-free fragment of the plurality of cell-free fragments to one of a plurality of sequence groups, thereby obtaining a plurality of cell-free fragment sets. Each cell-free fragment set maps to a different sequence group among the plurality of sequence groups. The one or more programs further include instructions for assigning a cell-free fragment cancer status to each cell-free fragment in each of the plurality of cell-free fragment sets. The cancer state of the cell-free fragments is one of a first cancer state and a second cancer state that is a function of the output of the classifier when the methylation pattern of each cell-free fragment is input into the classifier. The one or more programs further include instructions for calculating a first representative value of the number of cell-free fragments from subjects assigned the first cancer state in each cell-free fragment set across the plurality of sequence groups, and instructions for calculating a second representative value of the number of cell-free fragments from subjects in each cell-free fragment set across the plurality of sequence groups. The one or more programs further include instructions for estimating the cell origin fraction of the subject using the first representative value and the second representative value.

[0141] Another aspect of the present disclosure provides a computer system as disclosed above, wherein the one or more programs further include instructions for performing any of the methods disclosed above, either singly or in combination.

[0142] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium storing one or more programs for estimating a subject's cell-origin fraction. The one or more programs are configured to be executed by a computer. The one or more programs include instructions for obtaining, in electronic form, a corresponding methylation pattern for each cell-free fragment among a plurality of cell-free fragments. The corresponding methylation pattern for each cell-free fragment is (i) determined by methylation sequencing of one or more nucleic acid samples containing each fragment in a biological sample obtained from the subject, and (ii) includes a methylation state of each CpG site among a corresponding plurality of CpG sites in each fragment. The one or more programs include instructions for mapping each cell-free fragment of the plurality of cell-free fragments to one of a plurality of sequence groups, thereby obtaining a plurality of cell-free fragment sets, each cell-free fragment set being mapped to a different sequence group among the plurality of sequence groups. The one or more programs further include instructions for assigning a cancer state to each cell-free fragment in each of the plurality of cell-free fragment sets, the cancer state being a function of the output of a classifier when the methylation pattern of each cell-free fragment is input to the classifier. The cancer state of the cell-free fragments is one of a first cancer state and a second cancer state. The one or more programs further include instructions for calculating a first representative value of the number of cell-free fragments from subjects assigned the first cancer state for each set of cell-free fragments across the plurality of sequence groups, and calculating a second representative value of the number of cell-free fragments from subjects for each set of cell-free fragments across the plurality of sequence groups. The one or more programs include instructions for estimating the cell origin fraction of the subject using the first representative value and the second representative value.

[0143] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium as disclosed above, wherein one or more programs further include instructions for performing any of the disclosed methods, either singly or in combination.

[0144] Various embodiments of the systems, methods, and apparatus within the scope of the appended claims each have several aspects, none of which is solely responsible for the desirable attributes described herein. Without limiting the scope of the appended claims, some prominent features are described herein. After considering this discussion, and particularly after reading the section entitled "Detailed Description," one will be able to understand how the features of the various embodiments may be used.

[0145] [Incorporation by Reference] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.

[0146] Embodiments disclosed herein are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings, in which like reference numerals refer to corresponding parts throughout the several views of the drawings. [Brief explanation of the drawings]

[0147] [Figure 1] FIG. 1 is an exemplary block diagram illustrating a computing device according to some embodiments of the present disclosure. [Figure 2] 2A and 2B are diagrams illustrating an exemplary flowchart of a method for identifying multiple features for estimating a percentage of a target cell origin according to some embodiments of the present disclosure, where dashed boxes represent optional steps. [Figure 3]3A and 3B are exemplary flowcharts illustrating methods for estimating the cell origin fraction of a subject according to some embodiments of the present disclosure, where dashed boxes represent optional steps. [Figure 4] FIG. 12 plots ctDNA rates as a function of cancer stage in subjects with any of the listed cancers, according to some embodiments of the present disclosure. [Figure 5] FIG. 1 shows a flowchart of a method for preparing a nucleic acid sample for sequencing, according to some embodiments of the present disclosure. [Figure 6] FIG. 1 is a graphical representation of a process for obtaining sequence reads according to some embodiments of the present disclosure. [Figure 7] 7 shows a comparison of tumor fraction estimates based on whole-genome bisulfite sequencing data with known tumor fractions obtained from tissue-based whole-genome sequencing data according to some embodiments of the present disclosure. Specifically, the WGBS-estimated tumor fraction includes the ratio of the average number of abnormal fragments to the average total number of fragments (e.g., each fragment mapped to a specific sequence group or region of the reference genome). FIG. 7 is based on sequencing information for 495 subjects. For known tissue tumor fractions > 0.01, the Spearman correlation for the WGBS tumor fraction estimates is 0.86. For known tissue tumor fractions > 0.005, the Spearman correlation for the WGBS tumor fraction estimates is 0.90. For known tissue tumor fractions > 0.001, the Spearman correlation for the WGBS tumor fraction estimates is 0.89. For known tissue tumor fractions > 0.0001, the Spearman correlation for the WGBS tumor fraction estimates is 0.74. This indicates that the WGBS-based tumor fraction estimates are correlated with the known tissue tumor fractions. [Figure 8] FIG. 1 illustrates a mutual information measure used in accordance with some embodiments of the present disclosure for feature identification. DETAILED DESCRIPTION OF THE INVENTION

[0148] Reference will now be made in detail to the embodiments illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0149] The embodiments described herein provide various technical solutions for determining an estimated cell origin fraction for a subject. In exemplary embodiments, nucleic acid fragments are obtained from a biological sample of the subject. The biological sample includes cell-free nucleic acids. Thus, the nucleic acid fragments are cell-free nucleic acids. The nucleic acid fragments are assessed for methylation status at a predefined set of methylation sites, and each is assigned a score based on the methylation status. The methylation status scores are converted into counts and compared with corresponding methylation scores for each methylation site of the predefined set of methylation sites. The corresponding methylation scores are obtained from an analysis of the methylation pattern in the cell source. This comparison determines a methylation frequency for the subject, which can be used to estimate the cell origin fraction for the cell source.

[0150] [Definition] As used herein, the term "about" or "approximately" means within an acceptable error range for a particular value as would be expected by one of ordinary skill in the art, which depends in part on how the value is measured or determined, e.g., the limitations of the measurement system. For example, in some embodiments, "about" means within one or more standard deviations, according to practice in the art. In some embodiments, "about" means within ±20%, ±10%, ±5%, or ±1% of a given value. In some embodiments, the term "about" or "approximately" means within an order of magnitude, within five-fold, or within two-fold of a value. When particular values ​​are described in the present application and claims, unless otherwise specified, the term "about" should be assumed to mean within an acceptable error range for the particular value. The term "about" may have the meaning commonly understood by one of ordinary skill in the art. In some embodiments, the term "about" refers to ±10%. In some embodiments, the term "about" refers to ±5%.

[0151] As used herein, the term "assay" refers to a technique for determining the characteristics of a substance, such as a nucleic acid, a protein, a cell, a tissue, or an organ. An assay (e.g., a first assay or a second assay) can include techniques for determining the copy number variation of a nucleic acid in a sample, the methylation status of a nucleic acid in a sample, the fragment size distribution of a nucleic acid in a sample, the mutation status of a nucleic acid in a sample, or the fragmentation pattern of a nucleic acid in a sample. Any assay known to those skilled in the art can be used to detect any of the nucleic acid characteristics mentioned herein. Nucleic acid characteristics can include sequence, genomic identity, copy number, methylation status at one or more nucleotide positions, nucleic acid size, the presence or absence of a nucleic acid mutation at one or more nucleotide positions, and the fragmentation pattern of a nucleic acid (e.g., the nucleotide positions at which the nucleic acid fragments). An assay or method can have a particular sensitivity and / or specificity, and its relative usefulness as a diagnostic tool can be measured using the ROC-AUC statistic.

[0152] As used herein, the terms "biological sample," "patient sample," and "sample" are used interchangeably and refer to any sample obtained from a subject and may reflect a biological state associated with the subject. In some embodiments, such samples contain cell-free nucleic acids, such as cell-free DNA. In some embodiments, such samples contain nucleic acids other than or in addition to cell-free nucleic acids. Examples of biological samples include, but are not limited to, a subject's blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, or peritoneal fluid. In some embodiments, a biological sample consists of a subject's blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, or peritoneal fluid. In such embodiments, the biological sample is limited to a subject's blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, or peritoneal fluid, and does not include other components of the subject (e.g., solid tissue, etc.). A biological sample can include any tissue or substance derived from a living or dead subject. A biological sample can be a cell-free sample. A biological sample can include nucleic acids (e.g., DNA or RNA) or fragments thereof. A sample can be a liquid sample or a solid sample (e.g., a cell or tissue sample). A biological sample can be a bodily fluid such as blood, plasma, serum, urine, vaginal fluid, edema (e.g., scrotal) fluid, vaginal washings, pleural fluid, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple discharge, aspirates from different parts of the body (e.g., thyroid, breast), etc. A biological sample can be a stool sample. In various embodiments, the majority of the DNA in a biological sample enriched for cell-free DNA (e.g., a plasma sample obtained via a centrifugation protocol) can be cell-free (e.g., more than 50%, more than 60%, more than 70%, more than 80%, more than 90%, more than 95%, or more than 99% of the DNA can be cell-free). Biological samples can be treated to physically disrupt tissue or cellular structures (e.g., by centrifugation and / or cell lysis) to release intracellular components into a solution that may further contain enzymes, buffers, salts, detergents, etc., which can be used to prepare the sample for analysis.A biological sample can be obtained from a subject invasively (eg, by surgical means) or non-invasively (eg, by drawing blood, swabbing, or collecting an excreted sample).

[0153] In some embodiments, the biological sample is derived from one tissue type (e.g., a single organ such as breast, lung, prostate, colon, kidney, uterus, pancreas, esophagus, lymph, ovary, cervix, epidermis, thyroid, bladder, or stomach). In some embodiments, the biological sample is derived from two or more tissue types (e.g., a combination of tissues from two or more organs). In some embodiments, the biological sample is derived from one or more cell types (e.g., cells from a single organ or a given set of organs).

[0154] As disclosed herein, the terms "nucleic acid" and "nucleic acid molecule" are used interchangeably. These terms refer to nucleic acids in any form, such as deoxyribonucleic acid (DNA, e.g., complementary DNA (cDNA), genomic DNA (gDNA), etc.), ribonucleic acid (RNA, e.g., message RNA (mRNA), small interfering RNA (siRNA), ribosomal RNA (rRNA), transfer RNA (tRNA), microRNA, RNA highly expressed in the fetus or placenta, etc.), and / or DNA or RNA analogs (e.g., containing base analogs, sugar analogs, and / or non-natural backbones, etc.), RNA / DNA hybrids, and polyamide nucleic acids (PNAs), all of which may be in single-stranded or double-stranded form. Unless specifically limited, nucleic acids may contain known analogs of natural nucleotides, some of which can function in a manner similar to naturally occurring nucleotides. Nucleic acids may be in any form (e.g., linear, circular, supercoiled, single-stranded, double-stranded, etc.) useful for carrying out the processes herein. In some embodiments, the nucleic acid may be derived from a single chromosome or a fragment thereof (e.g., a nucleic acid sample may be derived from one chromosome of a sample obtained from a diploid organism). In certain embodiments, the nucleic acid comprises a nucleosome, or a fragment or portion of a nucleosome or nucleosome-like structure. The nucleic acid may also comprise proteins (e.g., histones, DNA-binding proteins, etc.). Nucleic acids analyzed by the processes described herein may be substantially isolated and substantially free of association with proteins or other molecules. Nucleic acids also include derivatives, variants, and analogs of RNA or DNA synthesized, replicated, or amplified from single-stranded ("sense" or "antisense," "plus" or "minus" strand, "forward" or "reverse" reading frame) and double-stranded polynucleotides. Deoxyribonucleotides include deoxyadenosine, deoxycytidine, deoxyguanosine, and deoxythymidine. In the case of RNA, the base cytosine is replaced by uracil, and the 2' position of the sugar contains a hydroxyl moiety. The nucleic acid can be prepared using nucleic acid obtained from a subject as a template.

[0155] As used herein, the term "cell-free nucleic acid" refers to nucleic acid molecules that may exist outside cells in bodily fluids, such as a subject's blood, whole blood, plasma, serum, urine, cerebrospinal fluid, stool, saliva, sweat, tears, pleural fluid, pericardial fluid, or ascites. Cell-free nucleic acids are derived from one or more healthy cells and / or one or more cancer cells. Cell-free nucleic acids are used interchangeably as circulating nucleic acids. Examples of cell-free nucleic acids include, but are not limited to, RNA, mitochondrial DNA, or genomic DNA. As used herein, the terms "cell-free nucleic acid," "cell-free DNA," and "cfDNA" are used interchangeably. As used herein, the term "circulating tumor DNA" or "ctDNA" refers to nucleic acid fragments derived from tumor cells or other types of cancer cells, which may be released into fluids from an individual's body (e.g., the bloodstream) as a result of biological processes such as apoptosis or necrosis of dying cells, or may be actively released by viable tumor cells. Examples of cell-free nucleic acids include, but are not limited to, RNA, mitochondrial DNA, or genomic DNA.

[0156] As disclosed herein, the term "circulating tumor DNA" or "ctDNA" refers to nucleic acid fragments derived from abnormal tissue, such as cells of a tumor or other type of cancer, that may be released into a subject's bloodstream as a result of biological processes such as apoptosis or necrosis of dying cells, or may be actively released by viable tumor cells.

[0157] As disclosed herein, the term "reference genome" refers to any specific, known, sequenced, or characterized, partial or complete genome of any organism or virus that can be used to reference an identifying sequence from a subject. Exemplary reference genomes used for human test subjects and many other organisms are provided in online genome browsers hosted by the National Center for Biotechnology Information (NCBI) or the University of California, Santa Cruz (UCSC). "Genome" refers to the complete genetic information of an organism or virus, represented by nucleic acid sequences. As used herein, a reference sequence or reference genome is often a compiled or partially compiled genome sequence from an individual or multiple individuals. In some embodiments, a reference genome is a compiled or partially compiled genome sequence from one or more human individuals. A reference genome can be considered a representative set of genes for a species. In some embodiments, a reference genome includes sequences assigned to chromosomes. Exemplary human reference genomes include, but are not limited to, NCBI build 34 (UCSC equivalent: hg16), NCBI build 35 (UCSC equivalent: hg17), NCBI build 36.1 (UCSC equivalent: hg18), GRCh37 (UCSC equivalent: hg19), and GRCh38 (UCSC equivalent: hg38).

[0158] As disclosed herein, the terms "region of a reference genome," "genomic region," or "chromosomal region" refer to any portion, contiguous or non-contiguous, of a reference genome. They may also be referred to, for example, as a group of sequences, a partition, a genome portion, a portion of a reference genome, a portion of a chromosome, etc. In some embodiments, a genome section is based on a particular length of genome sequence. In some embodiments, a method can include analysis of multiple nucleic acid fragments mapped to multiple genome regions. Genomic regions can be approximately the same length, or genome sections can be different lengths. In some embodiments, genomic regions are approximately the same length. In some embodiments, genomic regions of different lengths are adjusted or weighted. In some embodiments, a genomic region is about 10 kilobases (kb) to about 500 kb, about 20 kb to about 400 kb, about 30 kb to about 300 kb, about 40 kb to about 200 kb, and in some cases, about 50 kb to about 100 kb. In some embodiments, a genomic region is about 100 kb to about 200 kb. A genomic region is not limited to a contiguous run of sequences. Thus, a genomic region may be composed of contiguous and / or non-contiguous sequences. A genomic region is not limited to a single chromosome. In some embodiments, a genomic region comprises all or part of one chromosome, or all or part of two or more chromosomes. In some embodiments, a genomic region may span one, two, or more entire chromosomes. Furthermore, a genomic region may span shared or non-shared portions of multiple chromosomes.

[0159] As used herein, the term "fragment" is used interchangeably with "nucleic acid fragment" (e.g., DNA fragment) and refers to a portion of a polynucleotide or polypeptide sequence comprising at least three consecutive nucleotides. In the context of sequencing cell-free nucleic acid molecules found in a biological sample, the terms "fragment" and "nucleic acid fragment" interchangeably refer to a cell-free nucleic acid molecule found in the biological sample or a representation thereof. In this context, sequencing data (e.g., sequence reads from whole genome sequencing, targeted sequencing, etc.) are used to derive one or more copies of all or part of such a nucleic acid fragment. Such sequence reads may actually be obtained from sequencing PCR replicates of the original nucleic acid fragment and thus "representative" or "supporting" the nucleic acid fragment. There may be multiple sequence reads, each representative of or supporting a particular nucleic acid fragment in a biological sample (e.g., PCR replicates). In some embodiments, a nucleic acid fragment can be considered a cell-free nucleic acid. In some embodiments, sequence reads from PCR replicates may be misleading; for example, when the abundance level of a particular cell-free nucleic acid molecule needs to be determined. In such embodiments, only one copy of a nucleic acid fragment is used to represent the original cell-free nucleic acid molecule (e.g., duplicates are removed by molecular identifiers attached to the cell-free nucleic acid molecules during the library preparation process). In some embodiments, methylation sequencing data can be used to further distinguish these nucleic acid fragments. For example, two nucleic acid fragments that share identical or nearly identical sequences may still correspond to different original cell-free nucleic acid molecules if each possesses a different methylation pattern.

[0160] In some embodiments, two fragments are considered to share a nearly identical nucleic acid sequence if the fragment sequences differ from each other by fewer than 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides. In some embodiments, two fragments are considered to share a nearly identical sequence if the fragment sequences differ from each other by fewer than 1% of the total nucleotides, fewer than 2% of the total nucleotides, fewer than 3% of the total nucleotides, fewer than 4% of the total nucleotides, or fewer than 5% of the total nucleotides.

[0161] In some embodiments, a first fragment from each (e.g., first or second) plurality of nucleic acid fragments is aligned to a first position in the reference genome, and a second fragment from each (e.g., first or second) plurality of nucleic acid fragments is aligned to a second position in the reference genome. In some embodiments, the first position and the second position correspond to different regions in the reference genome. In some embodiments, the first position and the second position are the same position (e.g., the first position and the second position correspond to the same region of the reference genome). In some embodiments, the first and second positions overlap in the reference genome by at least 1 residue, at least 2 residues, at least 3 residues, at least 4 residues, at least 5 residues, at least 6 residues, at least 7 residues, at least 8 residues, at least 9 residues, at least 10 residues, at least 11 residues, at least 12 residues, at least 13 residues, at least 14 residues, at least 15 residues, at least 16 residues, at least 17 residues, at least 18 residues, at least 19 residues, at least 20 residues, at least 30 residues, at least 40 residues, at least 50 residues, at least 60 residues, at least 70 residues, at least 80 residues, at least 90 residues, or at least 100 residues. In some embodiments, the first and second positions overlap in the reference genome by between 1 and 50 residues.

[0162] In some embodiments, each fragment is mapped to at least a first location and a second location in the reference genome (e.g., the nucleic acid sequence corresponding to each fragment is present in at least two different locations in the reference genome). In some embodiments, each fragment is mapped to at least three locations, at least four locations, at least five locations, at least six locations, at least seven locations, at least eight locations, at least nine locations, at least ten locations, at least 11 locations, at least 12 locations, at least 13 locations, at least 14 locations, at least 15 locations, at least 16 locations, at least 17 locations, at least 18 locations, at least 19 locations, or at least 20 locations in the reference genome. In some embodiments, the at least two mapped locations in the reference genome are separated from each other in the reference genome by at least 1 residue, at least 5 residues, at least 10 residues, at least 25 residues, at least 50 residues, at least 100 residues, at least 200 residues, at least 300 residues, at least 400 residues, at least 500 residues, at least 600 residues, at least 700 residues, at least 800 residues, at least 900 residues, or at least 1000 residues. In some embodiments, at least two mapped locations comprise different genes in the reference genome, hi some embodiments, at least two mapped locations are located on different chromosomes of the reference genome.

[0163] Nucleic acid fragments can retain the biological activity and / or some characteristics of the parent polynucleotide. As an example, nasopharyngeal carcinoma cells can accumulate fragments of Epstein-Barr virus (EBV) DNA in the bloodstream of a subject, e.g., a patient. These fragments can contain one or more BamHI-W sequence fragments, which can be used to detect the level of tumor-derived DNA in plasma. The BamHI-W sequence fragments correspond to sequences that can be recognized and / or digested using the BamHI restriction enzyme. The BamHI-W sequence can refer to the sequence 5'-GGATCC-3'.

[0164] Furthermore, polynucleotides can be divided or fragmented into multiple segments by natural processes, such as cfDNA fragments that may naturally exist in biological samples, or by in vitro manipulation.Various methods for fragmenting nucleic acids are well known in the art.These methods can be, for example, chemical, physical, or enzymatic in nature.Enzymatic fragmentation can include partial degradation by DNAse; partial depurination by acid; the use of restriction enzymes; intron-encoded endonucleases; DNA-based cleavage methods such as triplex and hybridization methods that rely on specific hybridization of nucleic acid segments to localize cleavage agents to specific locations on nucleic acid molecules; or other enzymes or compounds that cleave polynucleotides at known or unknown locations.Physical fragmentation methods can include subjecting polynucleotides to high shear rates.High shear rates can be achieved, for example, by moving DNA through a chamber or channel with pits or spikes, or by forcing a DNA sample through a channel with a limited size, for example, an aperture with cross-sectional dimensions in the micron or submicron range. Other physical methods include sonication and nebulization.A combination of physical and chemical fragmentation methods, such as thermal fragmentation and ion-mediated hydrolysis, can also be employed.See, for example, Sambrook et al., "Molecular Cloning: A Laboratory Manual", 3rd Ed.Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2001) ("Sambrook et al."), which is incorporated herein by reference for all purposes.These methods can be optimized to digest nucleic acid into fragments of selected size ranges.

[0165] As used herein, the term "sequence read" or "read" refers to a nucleotide sequence generated by any sequencing process described herein or known in the art. A read can be generated from one end of a nucleic acid fragment (a "single-end read"), or sometimes from both ends of a nucleic acid (e.g., a paired-end read, a double-end read). In some embodiments, a sequence read (e.g., a single-end read or a paired-end read) can be generated from one or both strands of a target nucleic acid fragment. The length of a sequence read is often related to a particular sequencing technology. For example, high-throughput methods provide sequence reads that can vary in size from tens to hundreds of base pairs (bp). In some embodiments, sequence reads are between about 15 bp and 900 bp in length (e.g., about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, or about 500 bp). In some embodiments, sequence reads are between about 1000 bp, 2000 bp, 5000 bp, 10,000 bp, or 50,000 bp, or longer. The length may be the average, median, or mean length. Nanopore sequencing can provide sequence reads that can vary in size, for example, from tens to hundreds to thousands of base pairs. Illumina parallel sequencing can provide sequence reads that are less variable, for example, the majority of sequence reads can be smaller than 200 bp. A sequence read (or sequencing read) can refer to sequence information corresponding to a nucleic acid molecule (e.g., a stretch of nucleotides). For example, a sequence read can correspond to a stretch of nucleotides (e.g., about 20 to about 150) from a portion of a nucleic acid fragment, a stretch of nucleotides at one or both ends of a nucleic acid fragment, or the nucleotides of the entire nucleic acid fragment.Sequence reads can be obtained in a variety of ways, for example, using sequencing techniques, or for example, using probes or capture probes in hybridization arrays, or using amplification techniques, for example, polymerase chain reaction (PCR) or linear or isothermal amplification using a single primer.

[0166] As disclosed herein, the terms "sequencing" and "sequencing" as used herein generally refer to any and all biochemical processes that can be used to determine the order of biological macromolecules, such as nucleic acids or proteins. For example, sequencing data can include all or a portion of the nucleotide bases in a nucleic acid molecule, such as a DNA fragment.

[0167] As disclosed herein, the term "single nucleotide variant" or "SNV" refers to a substitution of one nucleotide with a different nucleotide at a position (e.g., site) in a nucleotide sequence, e.g., a sequence read from an individual. A substitution of a first nucleobase X with a second nucleobase Y may be represented as "X>Y." For example, an SNV from cytosine to thymine may be represented as "C>T."

[0168] As used herein, the term "methylation profile" (also referred to as methylation status) can include information related to DNA methylation for a region. Information related to DNA methylation can include the methylation index of CpG sites, the methylation density of CpG sites within a region, the distribution of CpG sites across a contiguous region, the methylation pattern or level for individual CpG sites within a region containing multiple CpG sites, and non-CpG methylation. The methylation profile of a significant portion of a genome can be considered equivalent to a methylome. "DNA methylation" in mammalian genomes can refer to the addition of a methyl group to the 5-position of the cytosine heterocycle in CpG dinucleotides (e.g., to produce 5-methylcytosine). Cytosine methylation can occur at cytosine in other sequence contexts, such as 5'-CHG-3' and 5'-CHH-3' (where H is adenine, cytosine, or thymine). Cytosine methylation can also occur in the form of 5-hydroxymethylcytosine. DNA methylation can also include methylation of non-cytosine nucleotides, such as N6-methyladenine.

[0169] As used herein, "methylome" can be a measure of the amount of DNA methylation at multiple sites or loci in a genome. The methylome can correspond to the entire genome, a significant portion of the genome, or a relatively small portion of the genome. A "tumor methylome" can be the methylome of a tumor in a subject (e.g., a human). The tumor methylome can be determined using cell-free tumor DNA in tumor tissue or plasma. A tumor methylome can be an example of a methylome of interest. A methylome of interest can be the methylome of an organ (e.g., the methylome of brain cells, bone, lung, heart, muscle, kidney, etc.) that can provide nucleic acids, e.g., DNA, in bodily fluids. The organ can also be a transplanted organ.

[0170] As used herein, the term "methylation index" for each genomic site (e.g., a CpG site, a DNA region in which a cytosine nucleotide is followed by a guanine nucleotide in the linear sequence of bases along its 5' → 3' direction) can refer to the ratio of nucleic acid fragments that show methylation at the site to the total number of nucleic acid fragments that cover the site. The "methylation density" of a region can be the number of reads at sites within the region that show methylation divided by the total number of reads that cover sites within the region. A site can have specific characteristics (e.g., the site can be a CpG site). The "CpG methylation density" of a region can be the number of reads that show CpG methylation divided by the total number of reads that cover CpG sites within the region (e.g., a specific CpG site, CpG sites within a CpG island, or a larger region). For example, the methylation density of each 100 kb sequence group in the human genome can be determined from the total number of unconverted cytosines (which may correspond to methylated cytosines) at CpG sites as a percentage of all CpG sites covered by nucleic acid fragments mapped to the 100 kb region. In some embodiments, this analysis is performed on other sequence group sizes, such as 50 kb or 1 MB. In some embodiments, the region is the entire genome or a chromosome or portion of a chromosome (e.g., a chromosome arm). The methylation index of a CpG site can be the same as the methylation density of the region if the region contains only that CpG site. The "percentage of methylated cytosines" can refer to, for example, the number of cytosine sites "C" that are shown to be methylated (e.g., unconverted after bisulfite conversion) relative to the total number of analyzed cytosine residues in the region, including cytosines outside of CpG contexts. The methylation index, methylation density, and percentage of methylated cytosines are examples of "methylation levels."

[0171] As used herein, "plasma methylome" can be a methylome determined from the plasma or serum of an animal (e.g., a human). Because plasma and serum can contain cell-free DNA, the plasma methylome can be an example of a cell-free methylome. Because the plasma methylome can be a mixture of tumor / patient methylomes, it can be an example of a mixed methylome. A "cell methylome" can be a methylome determined from cells (e.g., blood cells or tumor cells) of a subject, e.g., a patient. The methylome of blood cells can be referred to as the blood cell methylome (or blood methylome).

[0172] As used herein, the term "aberrant methylation pattern" or "atypical methylation pattern" refers to a methylation state vector, methylation pattern, or methylation state of a DNA molecule having said methylation state vector that is expected to be found in a sample at a frequency lower than a threshold value. In certain embodiments provided herein, the expectation of finding a particular methylation state vector in a healthy control group including healthy individuals is represented by a p-value. In some embodiments, the p-value of a methylation state vector is determined as described in Example 5 of PCT / US2020 / 034317, entitled "Systems and Methods for Determining Whether a Subject Has a Cancer Condition Using Transfer Learning," filed May 22, 2020, and U.S. Patent Application No. 16 / 352,602, entitled "Anomalous fragment detection and classification," filed March 13, 2019, and now published as US2019 / 0287652, each of which is incorporated herein by reference in its entirety. A low p-value score generally corresponds to a relatively less expected methylation state vector compared to other methylation state vectors in samples from healthy individuals in the healthy control group. A high p-value score generally corresponds to a relatively more expected methylation state vector compared to other methylation state vectors found in samples from healthy individuals in the healthy control group. A methylation state vector with a p-value lower than a threshold value (e.g., 0.1, 0.01, 0.001, 0.0001, etc.) can be defined as an abnormal methylation pattern. Various methods known in the art can be used to calculate the p-value or expectation of a methylation pattern or methylation state vector. An exemplary method provided herein includes the use of Markov chain probability, which assumes that the methylation state of a CpG site depends on the methylation state of adjacent CpG sites.An alternative method provided herein calculates the expected probability of observing a particular methylation state vector in a healthy individual by utilizing a mixture model containing multiple mixture components (each of which is an independent site model in which the methylation at each CpG site is assumed to be independent of the methylation state at other CpG sites). The method provided herein uses genomic regions with atypical methylation patterns. If cfDNA fragments corresponding to or derived from a genomic region have a methylation state vector that appears less frequently than a threshold in a reference sample, the genomic region can be determined to have an atypical methylation pattern. The reference sample can be a sample from a control subject or a healthy subject. The frequency of a methylation state vector appearing in a reference sample can be expressed as a p-value score. If the cfDNA fragments corresponding to or derived from a genomic region do not have a single uniform methylation state vector, the genomic region can have multiple p-value scores for multiple methylation state vectors. In this case, the multiple p-value scores can be summed or averaged before being compared with a threshold. The p-value score corresponding to a genomic region can be compared to a threshold using various methods known in the art, including, but not limited to, the arithmetic mean, geometric mean, harmonic mean, median, and mode.

[0173] As used herein, the term "relative abundance" can refer to the ratio of a first amount of nucleic acid fragments having a particular characteristic (e.g., ending at one or more particular coordinate / end positions, aligning to a particular region of the genome, or having a particular methylation state, a particular length) to a second amount of nucleic acid fragments having a particular characteristic (e.g., ending at one or more particular coordinate / end positions or aligning to a particular region of the genome, a particular length). In one example, relative abundance can refer to the ratio of the number of DNA fragments ending at a first set of genomic locations to the number of DNA fragments ending at a second set of genomic locations. In some aspects, "relative abundance" can be a type of separation value that relates the amount of cell-free DNA molecules ending within a window of genomic locations (one value) to the amount of cell-free DNA molecules ending within another window of genomic locations (another value). The two windows can overlap but can be different sizes. In other embodiments, the two windows cannot overlap. Furthermore, in some embodiments, the window is one nucleotide wide and is therefore equivalent to one genomic location.

[0174] As used herein, the term "methylation" refers to a modification of deoxyribonucleic acid (DNA) in which a hydrogen atom on the pyrimidine ring of a cytosine base is converted to a methyl group to form 5-methylcytosine. In particular, methylation tends to occur at cytosine and guanine dinucleotides (referred to herein as "CpG sites"). In other instances, methylation can occur at cytosine that is not part of a CpG site or at other nucleotides other than cytosine; however, these occurrences are rare. In this disclosure, methylation is discussed with reference to CpG sites for clarity. Aberrant cfDNA methylation can be identified as hypermethylation or hypomethylation, both of which may indicate a cancerous state. As is well known in the art, aberrant DNA methylation (compared to healthy controls) can cause differential effects and contribute to cancer.

[0175] Identifying aberrantly methylated cfDNA fragments poses several challenges. First, determining whether a subject's cfDNA is aberrantly methylated is only meaningful when compared with a control group, and when the control group is small, the determination becomes unreliable. Furthermore, within a control group, the methylation status of subjects varies, which can be difficult to consider when determining whether a subject's cfDNA is aberrantly methylated. Furthermore, cytosine methylation at one CpG site is known to be causally related to the methylation of subsequent CpG sites.

[0176] Those skilled in the art will appreciate that the principles described herein are equally applicable to detecting methylation in non-CpG contexts, including non-cytosine methylation.

[0177] As disclosed herein, the term "subject" refers to any living or non-living organism, including, but not limited to, humans (e.g., male humans, female humans, fetuses, pregnant women, or children), non-human animals, plants, bacteria, fungi, or protists. Any human or non-human animal can serve as a subject, including mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, bovines (e.g., cows), equines (e.g., horses), caprines and ovines (e.g., sheep, goats), porcines (e.g., pigs), camelids (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), ursines (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, and sharks. The terms "subject" and "patient" are used interchangeably herein and refer to a human or non-human animal known to have or potentially having a medical condition or disorder, such as cancer. In some embodiments, the subject is male or female (eg, a man, woman, or child) at any stage.

[0178] The subject from whom a sample is taken or treated by any of the methods or compositions described herein can be of any age, and can be an adult, an infant, or a child. In some cases, the subject, e.g., patient, may be 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 years old, or within a range therein (e.g., from about 2 years old to about 20 years old, from about 20 years old to about 40 years old, or from about 40 years old to about 90 years old). A particular class of subjects, e.g., patients, that can benefit from the methods of the present disclosure are subjects, e.g., patients, over the age of 40.

[0179] Another particular class of subjects, e.g., patients, that can benefit from the methods of the present disclosure are pediatric patients, who may be at higher risk for chronic cardiac conditions. Furthermore, the subjects, e.g., patients, from whom samples are taken or who are treated by any of the methods or compositions described herein can be male or female.

[0180] As used herein, the term " normalization " refers to converting a value or a set of values ​​into a common reference frame for comparison purposes.For example, when diagnostic ctDNA level is " normalized " with baseline ctDNA level, diagnostic ctDNA level is compared with baseline ctDNA level, so that diagnostic ctDNA level can determine the amount that differs from baseline ctDNA level.

[0181] As used herein, the term "cancer" or "tumor" refers to an abnormal mass of tissue whose growth exceeds and is uncoordinated with that of normal tissue. Cancers or tumors can be defined as "benign" or "malignant" depending on the following characteristics: degree of cellular differentiation, including morphology and functionality, growth rate, local invasion, and metastasis. "Benign" tumors can be characterized by being well differentiated, growing slowly compared to malignant tumors, and being confined to the site of origin. Furthermore, benign tumors may lack the ability to invade, invade, or metastasize to distant sites. "Malignant" tumors can be characterized by being poorly differentiated (anaplastic) and growing rapidly while progressively infiltrating, invading, and destroying surrounding tissues. Furthermore, malignant tumors may have the ability to metastasize to distant sites.

[0182] As used herein, the term "level of cancer" refers to the presence or absence of cancer (e.g., presence or absence), the stage of cancer, tumor size, presence or absence of metastasis, total tumor burden in the body, and / or other indicators of the seriousness of cancer (e.g., cancer recurrence). Cancer levels can be numbers or other indicators, such as symbols, alphabetic letters, and colors. The level can be zero. Cancer levels can also include premalignant or precancerous conditions (statuses) associated with mutations or the number of mutations. Cancer levels can be used in a variety of ways. For example, screening can determine whether cancer is present in a person not previously known to have cancer. Evaluation can investigate a person diagnosed with cancer to monitor the progression of cancer over time, investigate the effectiveness of treatment, or determine a prognosis. In one embodiment, a prognosis can be expressed as the probability that a subject will die from cancer, or that the cancer will progress after a certain period or time, or that the cancer will metastasize. Detection can include "screening," or can include determining whether a person with suggestive features of cancer (e.g., symptoms or other positive tests) has cancer.

[0183] The terms "cancer burden," "tumor burden," "cancer burden," and "tumor burden" are used interchangeably herein to refer to the concentration or presence of tumor-derived nucleic acids in a test sample. Thus, the terms "cancer burden," "tumor burden," "cancer burden," and "tumor burden" are non-limiting examples of cell source fraction or tumor fraction in a biological sample. In some embodiments, tumor fraction is a specific version of cell source fraction.

[0184] As used herein, the term "tissue" refers to a group of cells organized as a functional unit. Multiple types of cells may be found in a single tissue. Different types of tissue may consist of different types of cells (e.g., liver cells, lung cells, or blood cells), but may also correspond to tissues from different organisms (maternal versus fetal) or healthy versus tumor cells. The term "tissue" can generally refer to any group of cells found in the human body (e.g., cardiac tissue, lung tissue, kidney tissue, nasopharyngeal tissue, oropharyngeal tissue). In some embodiments, the term "tissue" or "tissue type" can be used to refer to the tissue from which cell-free nucleic acid is derived. In one example, viral nucleic acid fragments may be derived from blood tissue. In another example, viral nucleic acid fragments may be derived from tumor tissue.

[0185] As used herein, the term "untrained classifier" refers to a classifier that has not been trained on a target dataset. However, an untrained classifier may be partially trained on a primary dataset (e.g., a small and / or reference dataset). It will be understood that the term "untrained classifier" does not exclude the possibility that transfer learning techniques may be used in such training of an untrained classifier. See, e.g., Fernandes et al., 2017, "Transfer Learning with Partial Observability Applied to Cervical Cancer Screening," Pattern Recognition and Image Analysis: 8 thIberian Conference Proceedings, pp. 243-250 (incorporated herein by reference) provides a non-limiting example of such transfer learning. When transfer learning is used, an untrained classifier is provided with additional data beyond that of the primary training dataset. Typically, this additional data takes the form of coefficients (e.g., regression coefficients) learned from another auxiliary training dataset. Furthermore, while a single auxiliary training dataset is disclosed, it will be understood that there is no limit to the number of auxiliary training datasets that may be used to supplement the primary training dataset when training an untrained classifier in this disclosure. For example, in some embodiments, two or more auxiliary training datasets, three or more auxiliary training datasets, four or more auxiliary training datasets, or five or more auxiliary training datasets are used to supplement the primary training dataset via transfer learning, each of which is different from the primary training dataset. In such embodiments, any method of transfer learning may be used. For example, consider a case where, in addition to the primary training dataset, there are a first auxiliary training dataset and a second auxiliary training dataset. The coefficients learned from the first auxiliary training data set (by applying a classifier such as regression to the first auxiliary training data set) are applied to the second auxiliary training data set using transfer learning techniques (e.g., the two-dimensional matrix multiplication described above), resulting in a trained intermediate classifier whose coefficients are then applied to the primary training data set, which is then applied to the untrained classifier along with the primary training data set itself.Alternatively, a first set of coefficients learned from the first auxiliary training dataset (by applying a classifier, such as regression on the first auxiliary training dataset) and a second set of coefficients learned from the second auxiliary training dataset (by applying a classifier, such as regression on the second auxiliary training dataset) may each be applied separately to separate instances of the primary training dataset (e.g., by separate, independent matrix multiplications), or both such applications of coefficients to separate instances of the primary training dataset, along with the primary training dataset itself (or some reduced form of the primary training dataset, such as principal components or regression coefficients learned from the primary training dataset), may then be applied to an untrained classifier to train the untrained classifier. In either example, knowledge of cell source (e.g., cancer type) from the first and second auxiliary training datasets is used, along with the cell source-labeled primary training dataset, to train the untrained classifier.

[0186] The term "classification" may refer to any number or other letter associated with a particular characteristic of a sample. For example, a "+" sign (or the word "positive") may indicate that the sample is classified as having a deletion or amplification. In another example, the term "classification" refers to the amount of tumor tissue in the subject and / or sample, the size of the tumor in the subject and / or sample, the stage of the tumor in the subject, the tumor burden in the subject and / or sample, and the presence of tumor metastasis in the subject. In some embodiments, classifications are binary (e.g., positive or negative) or have more classification levels (e.g., a scale of 1 to 10 or 0 to 1). In some embodiments, the terms "cutoff" and "threshold" refer to a predetermined numerical value used in an operation. In one example, cutoff size refers to the size at which fragments are excluded. In some embodiments, threshold refers to the value above or below which a particular classification applies. Any of these terms can be used in any of these contexts.

[0187] As used herein, the terms "cancer-associated alteration" or "cancer-specific alteration" can include cancer-derived mutations (including single nucleotide mutations, nucleotide deletions or insertions, gene or chromosomal segment deletions, translocations, and inversions), gene amplifications, virus-associated sequences (e.g., viral episomes, viral inserts, viral DNA that infects and then is released from cells, circulating or cell-free viral DNA), abnormal methylation profiles or tumor-specific methylation signatures, abnormal cell-free nucleic acid (e.g., DNA) size profiles, abnormal histone modification marks and other epigenetic modifications, and the location of the ends of cancer-associated or cancer-specific cell-free DNA fragments.

[0188] As used herein, the terms "control," "control sample," "reference," "reference sample," "normal," and "normal sample" refer to a sample from a subject who does not have a particular condition, i.e., is normally healthy. In one example, a method as disclosed herein can be performed on a subject with a tumor, where the reference sample is a sample taken from the subject's healthy tissue. The reference sample can be obtained from the subject or from a database. The reference can be, for example, a reference genome used to map nucleic acid fragments obtained from a sample from a subject. The reference genome can refer to a haploid or diploid genome to which nucleic acid fragments from a biological sample and a constitutional sample can be aligned and compared. An example of a constitutional sample can be DNA from white blood cells obtained from a subject. In the case of a haploid genome, only one nucleotide can be present at each locus. In the case of a diploid genome, heterozygous loci can be identified; each heterozygous locus can have two alleles, and either allele can be aligned to the locus.

[0189] Some aspects are described below with reference to the application of illustrative examples. It should be understood that numerous specific details, relationships, and methods are described to provide a thorough understanding of the features described herein. However, those skilled in the relevant art will readily recognize that the features described herein can be implemented without one or more of the specific details or using other methods. The features described herein are not limited by the depicted order of acts or events, as some acts may occur in different orders and / or simultaneously with other acts or events. Furthermore, not all depicted acts or events are required to implement a methodology in accordance with the features described herein.

[0190] Exemplary System Embodiments

[0191] Details of an exemplary system will now be described in conjunction with FIG. 1. FIG. 1 is a block diagram illustrating a system 100 according to some implementations. In some implementations, the device 100 includes one or more processing units CPUs 102 (also called processors or processing cores), one or more network interfaces 104, a user interface 106, non-persistent memory 111, persistent memory 112, and one or more communication buses 114 for interconnecting these components. The one or more communication buses 114 optionally include circuitry (sometimes called a chipset) that interconnects and controls communications between the system components. The non-persistent memory 111 typically includes high-speed random access memory such as DRAM, SRAM, DDR RAM, ROM, EEPROM, flash memory, etc., while the persistent memory 112 typically includes a CD-ROM, a digital versatile disk (DVD) or other optical storage, a magnetic cassette, a magnetic tape, a magnetic disk storage or other magnetic storage device, a magnetic disk storage device, an optical disk storage device, a flash memory device, or other non-volatile solid-state storage device. The persistent memory 112 optionally includes one or more storage devices located remotely from the CPU 102. The persistent memory 112 and the non-volatile memory devices in the non-persistent memory 112 constitute a non-transitory computer-readable storage medium. In some implementations, the non-persistent memory 111 or alternatively the non-transitory computer-readable storage medium stores the following programs, modules and data structures, or a subset thereof, sometimes in combination with the persistent memory 112: · an optional operating system 116 containing procedures for handling various basic system services and for performing hardware-dependent tasks; · an optional network communications module (or instructions) 118 for connecting the system 100 with other devices or communications networks; · a cell source fraction estimation module 120 for determining a cell source fraction 158 of the test subject in a biological sample of the test subject 140; a training dataset 122 including, for each training subject 124 (e.g., 124-1, . . . , 124-Z, where Z is a positive integer greater than 1), for each cell-free fragment 126 (e.g., 126-1-X, . . . , 126-1-Y, where X and any positive integer and Y is greater than X) of each training subject, at least (i) a corresponding methylation pattern 128 (e.g., 128-1-X) determined from each methylation state of each CpG site 130 (e.g., 130-1-XA, . . . , 130-1-XQ) of each cell-free fragment; and (ii) a corresponding subject cancer indication 136 for each training subject 136. For each of a plurality of cell-free fragments 142 (e.g., 142-G, . . . , 142-H (where G and H are positive integers and H is greater than G) derived from the biological sample to be tested), (i) at least the methylation state of each CpG site 148 (e.g., 146-GM, . . . , 146-GN, . . . , 146-HO, . . . 146H-P (where M, N, O, and P are positive integers)) of each of the cell-free fragments a test dataset 140 including (i) each methylation pattern 144 (e.g., 144-G, ..., 144-H) determined from the methylation pattern 144; (ii) each sequence group mapping 148 (e.g., 148-G, ..., 148H); and (iii) each predicted cell-free fragment cancer state 150 (e.g., 150-G, ..., 150-H), further including a first representative value 152, a second representative value 154, and an estimated cell-origin fraction 156.

[0192] In accordance with the present disclosure, a corresponding sequence group mapping 132 (e.g., 132-1-X) for each cell-free fragment and an assignment of a cell-free fragment cancer state 134 (e.g., 134-1-X) for each cell-free fragment are made. For convenience and ease of interpretation, these data constructs are shown as being in the training dataset. However, in typical embodiments, such data constructs are calculated from the methylation patterns of cell-free fragments in the training set and are not part of the original dataset. In other embodiments, the sequence group mapping 132 and cell-free fragment cancer state are part of the resulting training dataset 122.

[0193] According to some implementations, one or more of the above-identified elements are stored in one or more of the memory devices mentioned above and correspond to sets of instructions for performing the functions described above. The above-identified modules, data, or programs (e.g., sets of instructions) need not be implemented as separate software programs, procedures, data sets, or modules; thus, various subsets of these modules and data may be combined or otherwise rearranged in various implementations. In some implementations, non-persistent memory 111 optionally stores a subset of the above-identified modules and data structures. Additionally, in some embodiments, the memory stores additional modules and data structures not described above. In some embodiments, one or more of the above-identified elements are stored in a computer system separate from, and addressable by, visualization system 100, so that visualization system 100 can retrieve all or a portion of such data when needed.

[0194] While Figure 1 depicts "system 100," this diagram is intended as a functional illustration of various features that may be present in a computer system, rather than as a structural schematic of the implementations described herein. In practice, and as those skilled in the art will recognize, items shown separately may be combined and some items may be separated. Additionally, while Figure 1 depicts certain data and modules in non-persistent memory 111, some or all of these data sets and / or modules may reside in persistent memory 112.

[0195] Having disclosed a system according to the present disclosure with reference to Figure 1, methods according to the present disclosure will now be described in detail with reference to Figures 2A and 2B and Figures 3A and 3B. It will be understood that any of the disclosed methods may utilize or interface with any of the assays or algorithms disclosed in U.S. Patent Application No. 15 / 793,830, filed October 25, 2017, and / or International Patent Publication No. PCT / US17 / 58099, having an international filing date of October 24, 2017 (each of which is incorporated herein by reference) to determine a cancerous state in a test subject or the likelihood that the test subject has a cancerous state.

[0196] Identifying features to estimate cell source fraction

[0197] Block 202. One aspect of the present disclosure provides a method for identifying a plurality of features for estimating a cell source fraction for a subject, executed in a computer system having one or more processors and a memory storing one or more programs for execution by the one or more processors.

[0198] In some embodiments, the cell source fraction in block 202 of FIG. 2A corresponds to a first cancer state of a common primary site. In some embodiments, the cell source fraction corresponds to a tumor fraction or fraction of a cancer type. In some embodiments, the cell source fraction corresponds to a tumor fraction of a given stage of a first cancer state. In some embodiments, the cell source fraction is derived from one or more types of human cells.

[0199] Subject and cancer condition

[0200] Block 204. In block 204 of Figure 2A, the method proceeds by obtaining a training dataset in electronic form, the training dataset including, for each training subject of a plurality of training subjects, at least: a) a corresponding methylation pattern for each cell-free fragment of a corresponding plurality of training cell-free fragments; and b) a subject cancer indication for each training subject, where the subject's cancer state is one of a first cancer state and a second cancer state.

[0201] In accordance with block 206, in some embodiments, the plurality of training objects consists of between 10 and 1000 training objects. In some embodiments, the plurality of training objects consists of at least 10 training objects, at least 25 training objects, at least 50 training objects, at least 100 training objects, at least 250 training objects, at least 500 training objects, at least 750 training objects, at least 1000 training objects, or at least 1500 training objects. In some embodiments, the plurality of training objects consists of between 10 and 100,000 training objects, between 100 and 50,000 training objects, or between 100 and 10,000 training objects.

[0202] In some embodiments, the number of training subjects having a first cancer condition and a second cancer condition is balanced in the plurality of training subjects (e.g., the plurality of training subjects includes a substantially equal number of training subjects having each cancer condition). For example, if the plurality of training subjects includes at least 50 training subjects having a first cancer condition, the plurality of training subjects also includes at least 50 training subjects having a second cancer condition, or if the plurality of training subjects includes at least 500 training subjects having a first cancer condition, the plurality of training subjects also includes at least 500 training subjects having a second cancer condition. In some embodiments, between 5% and 95% of the training subjects have the first cancer condition, and the remainder have the second cancer condition. In some embodiments, between 20% and 80% of the training subjects have the first cancer condition, and the remainder have the second cancer condition. In some embodiments, between 30% and 70% of the training subjects have the first cancer condition, and the remainder have the second cancer condition. In some embodiments, between 40 percent and 60 percent of the training subjects have the first cancer condition and the remainder have the second cancer condition, hi some embodiments, between 45 percent and 55 percent of the training subjects have the first cancer condition and the remainder have the second cancer condition.

[0203] Referring to block 208, in some embodiments, the first cancer condition consists of cancer and the second cancer condition is the absence of cancer. In some embodiments, the first cancer condition is adrenal cancer, biliary tract cancer, bladder cancer, bone / bone marrow cancer, brain cancer, breast cancer, cervical cancer, colon cancer, esophageal cancer, gastric cancer, head / neck cancer, hepatobiliary cancer, renal cancer, liver cancer, lung cancer, ovarian cancer, pancreatic cancer, pelvic cancer, pleural cancer, prostate cancer, kidney cancer, skin cancer, stomach cancer, testicular cancer, thymic cancer, thyroid cancer, uterine cancer, lymphoma, melanoma, multiple myeloma, or leukemia, and the second cancer condition is the absence of cancer. In some embodiments, the first cancer condition is one of a stage of adrenal cancer, a stage of biliary tract cancer, a stage of bladder cancer, a stage of bone / bone marrow cancer, a stage of brain cancer, a stage of breast cancer, a stage of cervical cancer, a stage of colon cancer, a stage of esophageal cancer, a stage of gastric cancer, a stage of head / neck cancer, a stage of hepatobiliary cancer, a stage of renal cancer, a stage of liver cancer, a stage of lung cancer, a stage of ovarian cancer, a stage of pancreatic cancer, a stage of pelvic cancer, a stage of pleural cancer, a stage of prostate cancer, a stage of renal cancer, a stage of skin cancer, a stage of stomach cancer, a stage of testicular cancer, a stage of thymic cancer, a stage of thyroid cancer, a stage of uterine cancer, a stage of lymphoma, a stage of melanoma, a stage of multiple myeloma, a stage of leukemia, and the second cancer condition is the absence of cancer.

[0204] In some embodiments, the second cancer condition is one of adrenal gland cancer, biliary tract cancer, bladder cancer, bone / bone marrow cancer, brain cancer, breast cancer, cervical cancer, colon cancer, esophageal cancer, gastric cancer, head / neck cancer, hepatobiliary cancer, renal cancer, liver cancer, lung cancer, ovarian cancer, pancreatic cancer, pelvic cancer, pleural cancer, prostate cancer, kidney cancer, skin cancer, stomach cancer, testicular cancer, thymic cancer, thyroid cancer, uterine cancer, lymphoma, melanoma, multiple myeloma, or leukemia. In some embodiments, the second cancer condition is one of a stage of adrenal cancer, a stage of biliary tract cancer, a stage of bladder cancer, a stage of bone / bone marrow cancer, a stage of brain cancer, a stage of breast cancer, a stage of cervical cancer, a stage of colon cancer, a stage of esophageal cancer, a stage of gastric cancer, a stage of head / neck cancer, a stage of hepatobiliary cancer, a stage of renal cancer, a stage of liver cancer, a stage of lung cancer, a stage of ovarian cancer, a stage of pancreatic cancer, a stage of pelvic cancer, a stage of pleural cancer, a stage of prostate cancer, a stage of renal cancer, a stage of skin cancer, a stage of stomach cancer, a stage of testicular cancer, a stage of thymic cancer, a stage of thyroid cancer, a stage of uterine cancer, a stage of lymphoma, a stage of melanoma, a stage of multiple myeloma, or a stage of leukemia.

[0205] In some embodiments, the subject's cancer state is one of a first cancer state, a second cancer state, and a third cancer state. In some embodiments, the cancer state of each subject of each of the plurality of training subjects is individually selected from the plurality of cancer states. In some such embodiments, the plurality of training subjects includes at least a minimum number of training subjects having each cancer state of the plurality of cancer states. In some embodiments, the minimum number of training subjects having each cancer state is at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, or at least 500 training subjects.

[0206] In some embodiments, the plurality of cancer conditions comprises at least 5, at least 10, or at least 20 unique cancer conditions. In some embodiments, the plurality of cancer conditions comprises 22 unique cancer conditions.

[0207] In some embodiments, each cancer condition of the plurality of cancer conditions is one of adrenal cancer, biliary tract cancer, bladder cancer, bone / bone marrow cancer, brain cancer, breast cancer, cervical cancer, colorectal cancer, esophageal cancer, gastric cancer, head / neck cancer, hepatobiliary cancer, renal cancer, liver cancer, lung cancer, ovarian cancer, pancreatic cancer, pelvic cancer, pleural cancer, prostate cancer, kidney cancer, skin cancer, stomach cancer, testicular cancer, thymic cancer, thyroid cancer, uterine cancer, lymphoma, melanoma, multiple myeloma, or leukemia. In some embodiments, each cancer condition of the plurality of cancer conditions is one of a stage of adrenal cancer, a stage of biliary tract cancer, a stage of bladder cancer, a stage of bone / bone marrow cancer, a stage of brain cancer, a stage of breast cancer, a stage of cervical cancer, a stage of colorectal cancer, a stage of esophageal cancer, a stage of gastric cancer, a stage of head / neck cancer, a stage of hepatobiliary cancer, a stage of renal cancer, a stage of liver cancer, a stage of lung cancer, a stage of ovarian cancer, a stage of pancreatic cancer, a stage of pelvic cancer, a stage of pleural cancer, a stage of prostate cancer, a stage of renal cancer, a stage of skin cancer, a stage of stomach cancer, a stage of testicular cancer, a stage of thymic cancer, a stage of thyroid cancer, a stage of uterine cancer, a stage of lymphoma, a stage of melanoma, a stage of multiple myeloma, or a stage of leukemia.

[0208] Cell-free fragment acquisition and methylation sequencing

[0209] Referring again to block 204, for each training subject, the corresponding methylation pattern of each cell-free fragment of each corresponding training plurality of cell-free fragments is determined by (i) methylation sequencing of one or more nucleic acid samples comprising each fragment in a corresponding biological sample obtained from each training subject, and (ii) includes the methylation state of each CpG site of the corresponding plurality of CpG sites in each fragment.

[0210] In some embodiments, the corresponding biological sample is a liquid biological sample. In some embodiments, the corresponding biological sample is a blood sample. In some embodiments, the corresponding biological sample comprises blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid of the training subject. In some embodiments, the corresponding biological sample comprises blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid of the training subject.

[0211] In some embodiments, one or more nucleic acid samples in the corresponding biological sample from the training subject are cell-free nucleic acid samples (e.g., obtained from a liquid biological sample). In some embodiments, the cell-free nucleic acid obtained from the biological sample is any form of nucleic acid as defined in the present disclosure, or a combination thereof. For example, in some embodiments, the cell-free nucleic acid obtained from the biological sample is a mixture of RNA and DNA.

[0212] In some embodiments, where each training subject's corresponding plurality of cell-free training fragments are derived from cell-free nucleic acids from a biological sample (e.g., a liquid biological sample), it is advantageous for the cell-free nucleic acids to exhibit an evaluable fractional cell source. In some embodiments, the fractional cell source for the first or second cancer condition for the corresponding training subject is at least 2 percent, at least 5 percent, at least 10 percent, at least 15 percent, at least 20 percent, at least 25 percent, at least 50 percent, at least 75 percent, at least 90 percent, at least 95 percent, or at least 98 percent.

[0213] In some embodiments, biological samples are processed to extract cell-free nucleic acids in preparation for sequencing analysis. As a non-limiting example, in some embodiments, cell-free nucleic acid fragments are extracted from a biological sample (e.g., a blood sample) collected from a subject in a K2 EDTA tube. If the biological sample is blood, the sample is processed within two hours of collection by first double-spinning the biological sample at 1000 g for 10 minutes and then spinning the resulting plasma at 2000 g for 10 minutes. The plasma is then stored in 1 mL aliquots at -80°C. In this manner, an appropriate amount of plasma (e.g., 1-5 mL) is prepared from the biological sample for cell-free nucleic acid extraction. In some such embodiments, cell-free nucleic acids are extracted using a QIAamp Circulating Nucleic Acid Kit (Qiagen) and eluted in DNA suspension buffer (Sigma). In some embodiments, the purified cell-free nucleic acids are stored at -20°C until use. See, e.g., Swanton, et al., 2017, "Phylogenetic ctDNA analysis depicts early stage lung cancer evolution," Nature, 545(7655):446-451, which is incorporated herein by reference. Other equivalent methods can be used to prepare cell-free nucleic acids from biological methods for sequencing purposes, and all such methods are within the scope of the present disclosure.

[0214] In one embodiment, cell-free nucleic acid fragments are treated to convert unmethylated cytosine to uracil. In one embodiment, the method uses bisulfite treatment of DNA, which converts unmethylated cytosine to uracil without converting methylated cytosine. For example, a commercially available kit such as EZ DNA Methylation™-Gold, EZ DNA Methylation™-Direct, or EZ DNA Methylation™-Lightning Kit (available from Zymo Research Corp, Irvine, CA) is used for bisulfite conversion. In another embodiment, the conversion of unmethylated cytosine to uracil is achieved using an enzymatic reaction. For example, the conversion can be performed using a commercially available kit for converting unmethylated cytosine to uracil, such as APOBEC-Seq (NEBiolabs, Ipswich, MA).

[0215] A sequencing library is prepared from the converted cell-free nucleic acid fragments. Optionally, the sequencing library is enriched for cell-free nucleic acid fragments or genomic regions that are informative about cellular origin using multiple hybridization probes. Hybridization probes are short oligonucleotides that hybridize to specifically designated cell-free nucleic acid fragments or target regions, and enrich for these fragments or regions for subsequent sequencing and analysis. In some embodiments, hybridization probes are used to perform targeted deep analysis of a set of specific CpG sites that are informative about cellular origin. Once prepared, the sequencing library or a portion thereof is sequenced to obtain multiple sequence reads.

[0216] In some embodiments, sequence reads obtained from a subject's biological sample are normalized to a reference set (e.g., obtained from multiple reference subjects, such as a control cohort of healthy subjects). U.S. Patent Application No. 2019-0287649, published September 19, 2019, entitled "Method and system for selecting, managing, and analyzing data of high dimensionality," which is incorporated herein by reference in its entirety, discloses several methods of normalization.

[0217] In some embodiments, the plurality of sequence reads comprises at least 100, at least 500, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10,000, at least 20,000, at least 50,000, at least 100,000, or at least 1 million sequence reads. In some embodiments, the plurality of sequence reads comprises at least 5 million, at least 10 million, or at least 100 million sequence reads.

[0218] In some embodiments, the plurality of cell-free fragments for training for each training subject of the plurality of training subjects includes at least 100, at least 500, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10,000, at least 20,000, at least 50,000, at least 100,000, at least 1 million, at least 5 million, or at least 10 million cell-free fragments. In some embodiments, the plurality of cell-free fragments for training for each training subject of the plurality of training subjects includes at least 100, at least 500, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10,000, at least 20,000, at least 50,000, at least 100,000, at least 1 million, at least 5 million, or at least 10 million cell-free fragments.

[0219] In some embodiments, a first training object in the plurality of training objects has a first corresponding plurality of cell-free fragments that includes a first number of cell-free fragments, and a second training object in the plurality of training objects has a second corresponding plurality of cell-free fragments that includes a second number of cell-free fragments that is different from the first number (e.g., in some embodiments, each training object has a different training plurality of cell-free fragments).

[0220] In some embodiments, each of the corresponding plurality of training cell-free fragments has an average length of less than 500 nucleotides, hi some embodiments, each of the corresponding plurality of training cell-free fragments has an average length of less than 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides.

[0221] In some embodiments, the sequencing comprises methylation sequencing.

[0222] In some embodiments, methylation sequencing detects one or more 5-methylcytosines (5mC) and / or 5-hydroxymethylcytosines (5hmC) in each fragment. In some such embodiments, methylation sequencing further includes converting one or more unmethylated cytosines or one or more methylated cytosines to one or more corresponding uracils in the sequence reads of each fragment. In some embodiments, one or more uracils are detected as one or more corresponding thymines during methylation sequencing. In some embodiments, the conversion of one or more unmethylated cytosines or one or more methylated cytosines comprises chemical conversion, enzymatic conversion, or a combination thereof. In some embodiments, cytosine conversion is performed as described in U.S. Patent Application No. 62 / 877,755, entitled "Systems and Methods for Determining Tumor Fraction," filed July 23, 2019, which is incorporated herein by reference.

[0223] In some embodiments, the methylation state of each CpG site in each fragment's corresponding plurality of CpG sites is (i) a methylated state if methylation sequencing determines that each CpG site is methylated, (ii) an unmethylated state if methylation sequencing determines that each CpG site is unmethylated, and / or (iii) flagged as "other" if the methylation state of each CpG site cannot be called methylated or unmethylated.

[0224] In some embodiments, the methylation sequencing (e.g., used to determine methylation patterns) is paired-end sequencing. In some embodiments, the methylation sequencing is single-read sequencing. In some embodiments, the methylation sequencing is whole-genome methylation sequencing (e.g., whole-genome bisulfite sequencing).

[0225] A whole-genome sequencing assay is a physical assay that generates sequence reads for the entire genome or a significant portion of the genome and can be used to determine large variations, such as copy number variations or copy number abnormalities. Such physical assays can employ whole-genome sequencing or whole-exome sequencing technologies.

[0226] In some embodiments, whole genome methylation sequencing identifies one or more methylation status vectors, e.g., as described in U.S. Patent Application No. 16 / 352,602, entitled "Anomalous fragment detection and classification," filed March 13, 2019, and now published as US2019 / 0287652, which is incorporated herein by reference in its entirety.

[0227] In some embodiments, sequencing includes any form of sequencing that can be used to obtain several sequence reads measured from nucleic acids (e.g., cell-free nucleic acids), including, but not limited to, high-throughput sequencing systems such as the Roche 454 platform, the Applied Biosystems SOLiD platform, Helicos True Single Molecule DNA sequencing technology, Affymetrix Inc.'s sequencing-by-hybridization platform, Pacific Biosciences' single molecule real-time (SMRT) technology, 54 Life Sciences, Illumina / Solexa and Helicos Biosciences' sequencing-by-synthesis platforms, Applied Biosystems' sequencing-by-ligation platform, Life technologies' ION TORRENT technology, and / or nanopore sequencing. In some embodiments, sequencing includes sequencing-by-synthesis and reversible terminator-based sequencing (e.g., Illumina's Genome Analyzer; Genome Analyzer II; HISEQ 2000; HISEQ 2500 (Illumina, San Diego, CA)).

[0228] In some embodiments, whole genome methylation sequencing is used to sequence a portion of a genome. In some embodiments, the portion of the genome is at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or all of the genome (e.g., human reference genome). In some embodiments, whole genome methylation sequencing generates a plurality of sequence reads, and each sequence read in the plurality of sequence reads has a sequence length of 1000 base pairs or less. In some embodiments, whole genome methylation sequencing obtains a sequencing coverage of the portion of the genome that is at least 5 times, at least 10 times, at least 15 times, at least 20 times, at least 25 times, at least 30 times, at least 50 times, at least 100 times, or at least 200 times across the portion of the genome. In some embodiments, whole-genome methylation sequencing obtains at least 5x, at least 10x, at least 15x, at least 20x, at least 25x, at least 30x, at least 50x, at least 100x, or at least 200x sequencing coverage across the genome.

[0229] In some embodiments, the methylation sequencing is targeted sequencing using a plurality of nucleic acid probes, wherein each sequence group of the plurality of sequences (e.g., a genomic region of interest) is associated with at least one nucleic acid probe of the plurality of nucleic acid probes.

[0230] In some embodiments, targeted sequencing targets a portion of a genome (e.g., a human reference genome) using multiple nucleic acid probes, and the targeted sequencing obtains at least 5x, at least 10x, at least 15x, at least 20x, at least 25x, at least 30x, at least 50x, at least 100x, at least 250x, at least 500x, or at least 1000x sequencing coverage of the targeted portion of the genome (e.g., to which the probes map). In some embodiments, targeted sequencing obtains at least 100x, at least 200x, at least 500x, at least 1,000x, at least 2,000x, at least 3,000x, at least 4,000x, at least 5,000x, at least 10,000x, at least 15,000x, at least 20,000x, at least 25,000x, at least 30,000x, at least 40,000x, or at least 50,000x sequencing coverage across a selected region in the genome of a subject.

[0231] In some embodiments, targeted panel sequencing is beneficial because it is more efficient (e.g., in terms of material use for sequencing, length of time required for sequencing, etc.) than, for example, whole genome sequencing, while still obtaining meaningful information about regions of interest in a subject's reference genome. In other words, in some embodiments, targeted panel sequencing helps obtain as much information as possible from the underlying data (e.g., both at the cell-free nucleic acid level and across genomic regions) while making the task of determining tumor fraction (and / or tumor origin) for a subject computationally tractable. For example, a reference genome (e.g., the human reference genome) contains approximately 28 million CpG sites, while a targeted methylation panel directed at a reference genome contains fewer CpG sites (e.g., 10,000-5 million CpG sites, 100,000-3 million CpG sites).

[0232] In some embodiments, at least one probe of the plurality of probes is designed to bind to and enrich a nucleic acid in a biological sample that contains at least one predetermined CpG site. In some implementations, each probe of the plurality of probes is designed to bind to and enrich a nucleic acid in a biological sample that contains at least one predetermined CpG site.

[0233] In some embodiments, each probe of multiple probes is designed to target the nucleic acid that has a certain number of predetermined CpG sites.For example, in some embodiments, one or more probes in multiple probes are designed to bind and concentrate in the nucleic acid in biological sample that comprises 50 or less predetermined CpG sites, 40 or less predetermined CpG sites, 30 or less predetermined CpG sites, 25 or less predetermined CpG sites, 22 or less predetermined CpG sites, 20 or less predetermined CpG sites, 18 or less predetermined CpG sites, 15 or less predetermined CpG sites, 12 or less predetermined CpG sites, 10 or less predetermined CpG sites, 5 or less predetermined CpG sites, 3 or less predetermined CpG sites.

[0234] In some embodiments, for targeted methylation sequencing, the plurality of probes comprises 1,000 to 2,000,000 probes. In some embodiments, the plurality of probes comprises 1,000 or more probes, 2,000 or more probes, 3,000 or more probes, 4,000 or more probes, 5,000 or more probes, 10,000 or more probes, 20,000 or more probes, or 30,000 or more probes. In some embodiments, the plurality of probes is 1,000 to 30,000 probes. In some embodiments, the plurality of probes comprises at least 5,000, at least 10,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 100,000, at least 200,000, at least 300,000, at least 400,000, at least 500,000, at least 600,000, at least 700,000, at least 800,000, at least 900,000, or at least 1,000,000 probes.

[0235] It should be understood that the plurality of probes may include other numbers of probes, non-limiting examples of which include up to 1,500,000 probes, up to 1,400,000 probes, up to 1,300,000 probes, up to 1,200,000 probes, up to 1,100,000 probes, up to 1,000,000 probes, up to 900,000 probes, up to 800,000 probes, up to 700,000 probes, up to 600,000 probes, up to 500,000 probes, up to 400,000 probes, up to 300,000 probes, up to 200,000 probes, The number of probes may include up to 100,000 probes, up to 90,000 probes, up to 80,000 probes, up to 70,000 probes, up to 60,000 probes, up to 50,000 probes, up to 40,000 probes, up to 30,000 probes, up to 20,000 probes, up to 10,000 probes, up to 9,000 probes, up to 8,000 probes, up to 7,000 probes, up to 6,000 probes, up to 5,000 probes, up to 4,000 probes, up to 4,000 probes, up to 2,000 probes, or up to 1,000 probes.

[0236] In some embodiments, the plurality of probes targets a plurality of gene targets (e.g., a portion of a reference genome and / or a panel of gene targets) that collectively cover 0.5 to 50 megabases of the reference genome. In some embodiments, the plurality of gene targets of the plurality of probes collectively cover 5 to 40 megabases of the reference genome, 10 to 30 megabases of the reference genome, 15 to 35 megabases of the reference genome, 20 to 30 megabases of the reference genome, 25 to 35 megabases of the reference genome, or 30 to 40 megabases of the reference genome.

[0237] In some embodiments, the plurality of probes is a targeted cancer assay panel. Several targeted cancer assay panels are known in the art and are described, for example, in International Patent Application No. PCT / US2019 / 025358, filed April 2, 2019, published as WO 2019 / 195268, entitled "Methylated Markers and Targeted Methylation Probe Panels," International Patent Application No. PCT / US2019 / 053509, filed September 27, 2019, published as WO 2020 / 069350, entitled "Methylated Markers and Targeted Methylation Probe Panel," and International Patent Application No. PCT / US2020 / 015082, filed January 24, 2020, published as WO 2020 / 154682, entitled "Detecting Cancer, Cancer Tissue or Origin, or Cancer Type," each of which is incorporated herein by reference in its entirety. For example, in some embodiments, a targeted cancer assay panel comprises a plurality of probes (or probe pairs) capable of capturing fragments (cell-free nucleic acids) that together can provide information relevant to determining tumor incidence and / or diagnosing cancer. In some embodiments, the plurality of probes in the targeted cancer assay panel comprises at least 50, 100, 500, 1,000, 2,000, 2,500, 5,000, 6,000, 7,500, 10,000, 15,000, 20,000, 25,000, or 50,000 pairs of probes. In other embodiments, the plurality of probes in the targeted cancer assay panel comprises at least 500, 1,000, 2,000, 5,000, 10,000, 12,000, 15,000, 20,000, 30,000, 40,000, 50,000, or 100,000 probes. In some embodiments, the plurality of probes collectively comprises at least 100,000, 200,000, 400,000, 600,000, 800,000, 1,000,000, 2,000,000, 3,000,000, 4,000,000, 5,000,000, 6,000,000, 7,000,000, 8,000,000, 9,000,000, or 10,000,000 nucleotides.In some embodiments, the probes (or probe pairs) are specifically designed to target one or more genomic regions that are differentially methylated in cancer and non-cancer samples.

[0238] For example, the multiple probes in the targeted cancer assay panel can comprise probes that can selectively bind and enrich the cfDNA fragments that are differentially methylated in cancerous samples.In this case, sequencing of enriched fragments can provide information relevant to determining tumor proportion or diagnosing cancer.In addition, probes can be designed to target the genomic regions that are determined to have abnormal methylation patterns and / or hypermethylation or hypomethylation patterns, thereby providing further selectivity and specificity of detection.

[0239] In some embodiments, a probe (or probe pair) in the plurality of probes targets a genomic region comprising at least 25 bp, 30 bp, 35 bp, 40 bp, 45 bp, 50 bp, 60 bp, 70 bp, 80 bp, or 90 bp. In some embodiments, a probe in the plurality of probes targets a genomic region comprising at least five methylation sites. In some embodiments, a probe in the plurality of probes targets a genomic region comprising less than 20, 15, 10, 8, or 6 methylation sites. In some embodiments, a probe in the plurality of probes targets a genomic region having at least 80, 85, 90, 92, 95, or 98% methylated (e.g., CpG) sites, either methylated or unmethylated, in a non-cancerous or cancerous sample.

[0240] Filtering cell-free fragments

[0241] In some embodiments, the method further includes applying one or more filter conditions to the plurality of cell-free fragments. Thus, in some embodiments, not all cell-free fragments obtained from methylation sequencing of one or more nucleic acid samples are used to identify multiple features for estimating the subject's cell-origin fraction and / or to estimate the subject's cell-origin fraction. In some embodiments, this is due to the fact that nucleic acid fragments (e.g., cell-free nucleic acids) vary in information content, and in some embodiments, only nucleic acid fragments with the desired information content are retained for feature identification and / or cell-origin fraction estimation (e.g., fragments that do not provide relevant information are discarded). In some embodiments, features are determined from cell-free fragments that satisfy one or more filter conditions in the plurality of filter conditions (e.g., each filter condition evaluates the information content of the fragment). Several filtering methods are described in detail, for example, in International Patent Application No. PCT / US2020 / 034317, entitled "Systems and Methods for Determining Whether a Subject Has a Cancer Condition Using Transfer Learning," filed on May 22, 2020, and U.S. Patent Application No. 16 / 352,602, entitled "Anomalous fragment detection and classification," filed on March 13, 2019, and now published as US2019 / 0287652 (each of which is incorporated herein by reference). Non-limiting examples of filter conditions are provided below.

[0242] P-value filtering based on methylation vectors

[0243] In some embodiments, a filter condition in the plurality of filter conditions is a requirement that each cell-free fragment of the plurality of cell-free fragments have a corresponding p-value that is less than or equal to a threshold value, where the p-value is determined by p-value filtering as described in Example 5 of International Patent Application No. PCT / US2020 / 034317, entitled "Systems and Methods for Determining Whether a Subject Has a Cancer Condition Using Transfer Learning," filed May 22, 2020, and U.S. Patent Application No. 16 / 352,602, entitled "Anomalous Fragment Detection and Classification," filed March 13, 2019, now published as US2019 / 0287652, each of which is incorporated herein by reference in its entirety. The purpose of such a filter condition is to accept and use aberrantly methylated cell-free fragments based on their corresponding methylation state vectors. For example, for each cell-free fragment in a sample, a determination is made as to whether the fragment is aberrantly methylated relative to an expected methylation state vector (e.g., through analysis of sequence reads obtained therefrom) using the methylation state vector corresponding to the fragment (e.g., where the expected methylation state vector is determined from sequence analysis of a cohort of healthy subjects). Generation of such cell-free fragment methylation state vectors is disclosed, for example, in U.S. Patent Application Publication No. 2019 / 0287652, which is incorporated herein by reference in its entirety.

[0244] In some embodiments, the healthy cohort includes at least 20 subjects, and the plurality of cell-free fragments includes at least 10,000 different corresponding methylation patterns. In some embodiments, the healthy cohort includes at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100 subjects. In some embodiments, the healthy cohort includes 1-10, 10-50, 50-100, 100-500, 500-1000, or more than 1000 subjects. In some embodiments, the plurality of cell-free fragments comprises 1 to 1000, 1000 to 2000, 2000 to 4000, 4000 to 6000, 6000 to 8000, 8000 to 10,000, 10,000 to 20,000, 20,000 to 50,000, or more than 50,000 different corresponding methylation patterns.

[0245] In some embodiments, the p-value threshold is between 0.001 and 0.20. In some embodiments, the threshold is 0.01 (e.g., in such embodiments, p must be <0.01). In some embodiments, the threshold is 0.001, 0.005, 0.01, 0.015, 0.02, 0.05, or 0.10. In some embodiments, the threshold is between 0.0001 and 0.20. In some embodiments, the p-value threshold is met for a methylation pattern from a subject if the corresponding methylation pattern for each cell-free fragment among the plurality of cell-free fragments has a p-value of 0.10 or less, 0.05 or less, or 0.01 or less.

[0246] In such embodiments, only cell-free fragments having a p-value below a threshold contribute to feature identification and / or cell-source fraction estimation. For example, in some embodiments, the plurality of cell-free fragments is filtered by removing from the plurality of cell-free fragments each cell-free fragment whose corresponding methylation pattern (e.g., methylation state vector) across the corresponding plurality of CpG sites in each fragment has a p-value that does not satisfy a p-value threshold.

[0247] In some embodiments, aberrant fragments are identified as fragments that exceed a threshold number of CpG sites and a threshold percentage of methylated CpG sites (hypermethylated), or fragments that exceed a threshold percentage of unmethylated CpG sites (hypomethylated). See, for example, the filter criteria based on minimum CpG sites and / or fragment length, described below. In some embodiments, the threshold percentage of methylated and / or unmethylated CpG sites is at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, or at least 95%. In some embodiments, the threshold percentage of methylated and / or unmethylated CpG sites is between 50% and 100%.

[0248] In some embodiments, a Markov model (e.g., a hidden Markov model "HMM") is used to provide a set of probabilities for determining the likelihood of observing the next state in a sequence, and for each state of the methylation pattern of each cell-free fragment, to determine the probability that a sequence of methylation states (e.g., including "M" for methylation and "U" for unmethylated) will be observed for each cell-free fragment. In some embodiments, the set of probabilities is obtained by training the HMM. Such training involves providing an initial training dataset of observed methylation state sequences (e.g., methylation patterns) obtained from a cohort of non-cancer subjects and calculating statistical parameters (e.g., the probability that a first state transitions to a second state (transition probability) and / or the probability that a given methylation state will be observed for each CpG site (output probability)). In some embodiments, the HMM is trained using supervised training (e.g., using samples where the underlying sequences are known, as are the observed states). In some alternative embodiments, the HMM is trained using unsupervised training (e.g., Viterbi learning, maximum likelihood estimation, expectation-maximization training, and / or Baum-Welch training). For example, expectation-maximization algorithms, such as the Baum-Welch algorithm, estimate transition and output probabilities from observed sample sequences to generate a parameterized probabilistic model that best describes the observed sequences. Such algorithms iteratively calculate the likelihood function until the expected number of correctly predicted states is maximized. See, e.g., Yoon, 2009, "Hidden Markov Models and Their Applications in Biological Sequence Analysis," Curr. Genomics. Sep; 10(6):402-415, doi: 10.2174 / 138920209789177575

[0249] Minimum bag size

[0250] In some embodiments, the filter condition in the plurality of filter conditions is a requirement that each cell-free fragment has a bag size greater than a threshold integer. In other words, in some embodiments, the filter condition in one or more filter conditions is a requirement that each cell-free fragment of the plurality of cell-free fragments be represented by a threshold number of sequence reads in a corresponding plurality of sequence reads measured from one or more nucleic acid samples containing each fragment in a corresponding biological sample. For example, if the threshold integer is 1, the filter condition is a requirement that each cell-free fragment be represented by more than one sequence read in a corresponding plurality of sequence reads measured from the biological sample. In some embodiments, the threshold integer is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or an integer between 10 and 100. In some embodiments, the threshold integer is 1 to 10, 10 to 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 90, or 90 to 100. In some embodiments, the threshold integer is between 100 and 500, between 500 and 1000, or greater than 1000.

[0251] In some embodiments, a filter condition in the plurality of filter conditions is a requirement that each cell-free fragment have a bag size greater than a threshold integer, where the sequence reads in each bag (e.g., representing each cell-free fragment) are obtained from sequencing a plurality of cell-free nucleic acids. For example, in some embodiments, a filter condition in one or more filter conditions is a requirement that each cell-free fragment of the plurality of cell-free fragments is represented by a threshold number of cell-free nucleic acids in one or more nucleic acid samples containing each fragment in a corresponding biological sample. In some embodiments, the threshold integer is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or an integer between 10 and 100. In some embodiments, the threshold integer is 1 to 10, 10 to 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 90, or 90 to 100. In some embodiments, the threshold integer is between 100 and 500, between 500 and 1000, or greater than 1000.

[0252] Minimum number of CpG sites

[0253] In some embodiments, the filter condition in the one or more filter conditions is to require that each cell-free fragment of the plurality of cell-free fragments have a threshold number of CpG sites. In some embodiments, the threshold number of CpG sites is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 CpG sites. In some embodiments, the threshold number of CpG sites is 1 to 10, 10 to 20, 20 to 30, 30 to 40, 40 to 50, or more than 50 CpG sites.

[0254] In some embodiments, the filter condition in the one or more filter conditions requires that each cell-free fragment of the plurality of cell-free fragments be less than a threshold number of base pairs in length. In some embodiments, the threshold number of base pairs is 1000, 2000, 3000, or 4000 base pairs. In some embodiments, the threshold number of base pairs is 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 base pairs. In some embodiments, the threshold number of base pairs is 1000, 2000, 3000, or 4000 consecutive base pairs in length. In some embodiments, the threshold number of base pairs is 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 consecutive base pairs in length.

[0255] In some embodiments, a filter condition in the plurality of filter conditions requires that each cell-free fragment cover a first threshold number of CpG sites and be less than a second threshold in length in terms of base pairs. For example, if the first threshold is 1 CpG site and the second threshold is 1000 base pairs, each cell-free fragment must cover more than one CpG site and be less than 1000 base pairs in length. In some embodiments, each cell-free fragment must cover at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 CpG sites within a specific fragment length (e.g., the second threshold length). In some embodiments, each cell-free fragment must span a specific number of CpG sites (e.g., the first threshold) while being less than 500, 1000, 2000, 3000, or 4000 consecutive base pairs in length. In other words, for example, in some embodiments, the filter condition in the plurality of filter conditions requires that each cell-free fragment contains at least 1 CpG site, at least 2 CpG sites, at least 3 CpG sites, at least 4 CpG sites, at least 5 CpG sites, at least 6 CpG sites, at least 7 CpG sites, at least 8 CpG sites, at least 9 CpG sites, at least 10 CpG sites, at least 11 CpG sites, at least 12 CpG sites, at least 13 CpG sites, at least 14 CpG sites, or at least 15 CpG sites within less than 500 consecutive nucleotides of the reference genome.

[0256] Hypermethylation or hypomethylation

[0257] In some embodiments, the filter condition in the plurality of filter conditions requires that each cell-free fragment is hypermethylated. In some embodiments, the filter condition in the plurality of filter conditions requires that each cell-free fragment is hypomethylated. In some embodiments, the filter condition depends on the genomic region (e.g., a group of sequences). For example, some regions of the human genome having a hypermethylated state associated with one or more cancer conditions, and some regions of the human genome having a hypomethylated state associated with one or more cancer conditions, are described in International Patent Application No. PCT / US2019 / 025358, filed April 2, 2019, published as WO 2019 / 195268, entitled "Methylated Markers and Targeted Methylation Probe Panels," International Patent Application No. PCT / US2020 / 015082, filed January 24, 2020, published as WO 2020 / 154682, entitled "Detecting Cancer, Cancer Tissue or Origin, or Cancer Type," and International Patent Application No. PCT / US2020 / 015082, filed September 27, 2019 ... and International Patent Application No. PCT / US2019 / 053509, published as WO 2020 / 069350, entitled "Antibody-Based Immunoglobulin-Conjugated Immunoglobulin-Related Immunoglobulin-Specific ...Thus, in some embodiments of the present disclosure, each of the one or more groups of sequences in the plurality of genomic regions represents a corresponding genomic region in the regions disclosed in WO 2019 / 195268, WO 2020 / 154682, and / or WO 2020 / 069350, and the filter condition in the plurality of filter conditions is (a) a region associated with one or more cancer conditions as indicated by WO 2019 / 195268, WO 2020 / 154682, and / or WO 2020 / 069350; (b) selecting fragments that map to sequences representing human genomic regions with a hypomethylated state at CpG sites associated with one or more cancer conditions as set forth in WO 2019 / 195268, WO 2020 / 154682, and / or WO 2020 / 069350 requires selecting cell-free nucleic acids that are hypermethylated.

[0258] In some embodiments, the plurality of filter conditions require that a p-value threshold be met and that the cell-free fragments be hypermethylated. In some embodiments, the plurality of filter conditions require that a p-value threshold be met and that the cell-free fragments be hypomethylated. In some embodiments, the plurality of filter conditions are different for each sequence group. For example, for one sequence group of the plurality of sequence groups, the plurality of filter conditions require that a p-value threshold be met and that the cell-free fragments be hypomethylated, while for a second sequence group of the plurality of sequence groups, the plurality of filter conditions require that a p-value threshold be met and that the cell-free fragments be hypermethylated.

[0259] Cancer status

[0260] In some embodiments, a filter condition in the plurality of filter conditions requires that each cell-free fragment meets a threshold for a cancer state (e.g., the probability associated with each cancer state of each cell-free fragment exceeds a predetermined threshold). In some embodiments, each cancer state has a different predefined threshold. For example, as described in U.S. Patent Application No. 63 / 003,087, entitled "Systems and Methods for Using Neural Networks to Determine a Cancer State," filed March 31, 2020, which is incorporated by reference in its entirety, a trained neural network (e.g., trained on multiple reference subjects) is used to determine the probability of cancer for each genomic region (e.g., a group of sequences).

[0261] In some such embodiments, for each sequence group of the plurality of sequence groups, for each cell-free fragment of the plurality of cell-free fragments mapped to each sequence group, a corresponding trained neural network calculates a predicted value, which is the probability that the cell-free fragment is associated with a cancerous state (e.g., the presence of cancer), based on the methylation pattern of each cell-free fragment. Accordingly, in some such embodiments, the methylation pattern of each cell-free fragment is scored using a trained neural network, wherein the score output by the trained neural network includes a calculation based on the probability that the cell-free fragment has a cancerous state and / or the probability that the cell-free fragment is associated with a cancerous state (e.g., the presence of cancer). Each cell-free fragment passes the filter condition (e.g., is selected for use in identifying features for estimating cell-origin fraction and / or is selected for use in estimating cell-origin fraction) if the obtained score satisfies the condition defined above (e.g., a probability above a fixed threshold value). Each cell-free fragment does not pass the filter condition (e.g., is discarded) if the obtained score does not satisfy the condition defined above (e.g., a probability below a fixed threshold value).

[0262] In some such embodiments, the threshold is positive or negative. In some embodiments, the threshold is 0.1 to 1, 1 to 5, 5 to 10, 10 to 50, 50 to 100, or greater than 100. In some embodiments, the threshold is -0.1 to -1, -1 to -5, -5 to -10, -10 to -50, -50 to -100, or less than -100. In some embodiments, the threshold is zero. In some embodiments, each set of sequences has a respective threshold for each cancer state (e.g., each subset of sequences is associated with each cancer state).

[0263] In some embodiments, any combination of the disclosed filter conditions is imposed. In some embodiments, the plurality of cell-free fragments includes one or more cell-free fragments whose methylation patterns satisfy one or more filter conditions disclosed herein.

[0264] Mapping fragments and sequences

[0265] Block 210. In block 210, the method proceeds by mapping each cell-free fragment of each of the plurality of cell-free fragments to a set of sequences from a plurality of sets of sequences, thereby obtaining a plurality of training sets of cell-free fragments. Each set of sequences from the plurality of sets of sequences represents a corresponding portion of the human reference genome. Each training set of cell-free fragments is mapped to a different set of sequences from the plurality of sets of sequences.

[0266] In some embodiments, mapping is performed using a Smith-Waterman gapped alignment, e.g., as implemented in Arioc, or a Burrows-Wheeler transformation, e.g., as implemented in Bowtie. Other suitable alignment programs include, but are not limited to, BarraCUDA, BBMap, BFAST, BigBWA, BLASTN, BLAT, BWA, BWA-PSSM, and CASHX. See, for example, Langmead and Salzberg, 2012, Nat Methods 9, pp. 357-359; Li and Durbin, 2009, "Fast and accurate short read alignment with Burrows-Wheeler transform," Bioinformatics 25(14), 1754-1760; and Smith and Yun, 2017, "Evaluation alignment and variant-calling software for mutation identification in C. elegans by whole-genome sequencing," PLOS ONE, doi.org / 10.1371 / jouranal.pone.0174446 (each of which is incorporated herein by reference). In some embodiments, mismatches may occur by mapping each cell-free fragment to a sequence group within a plurality of sequence groups. In some embodiments, the mapping includes at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or more than 10 mismatches.

[0267] In some embodiments, referring to block 212, the plurality of sequences consists of or includes between 1,000 and 100,000 sequences. In some embodiments, the plurality of sequences consists of or includes between 15,000 and 80,000 sequences. In some embodiments, the plurality of sequences consists of or includes between 25,000 and 65,000 sequences. In some embodiments, the plurality of sequences consists of or includes between 45,000 and 65,000 sequences.

[0268] In some embodiments, the plurality of sequences comprises at least 1000 sequences, at least 2500 sequences, at least 5000 sequences, at least 10,000 sequences, at least 20,000 sequences, at least 30,000 sequences, at least 40,000 sequences, at least 50,000 sequences, at least 60,000 sequences, at least 70,000 sequences, at least 80,000 sequences, at least 90,000 sequences, at least 100,000 sequences, or at least 110,000 sequences.

[0269] Further, in some embodiments, according to block 214 of FIG. 2A , each sequence group of the plurality of sequence groups has, on average, 10 to 1200 residues (e.g., each sequence group corresponds to a portion of the human reference genome consisting of 10 to 1200 nucleotides). In some embodiments, each sequence group of the plurality of sequence groups has, on average, 10 to 10,000 residues. In some embodiments, each sequence group of the plurality of sequence groups has, on average, 10 to 500 residues. In some embodiments, each sequence group of the plurality of sequence groups has, on average, 10 to 100 residues. In some embodiments, each sequence group of the plurality of sequence groups has, on average, 25 to 100 residues. In some embodiments, each sequence group of the plurality of sequence groups has, on average, 5000 to 10000 residues.

[0270] In some embodiments, each sequence group of the plurality of sequence groups contains fewer than 10 residues, fewer than 20 residues, fewer than 30 residues, fewer than 40 residues, fewer than 50 residues, fewer than 60 residues, fewer than 70 residues, fewer than 80 residues, fewer than 90 residues, fewer than 100 residues, fewer than 200 residues, fewer than 300 residues, fewer than 400 residues, fewer than 500 residues, fewer than 600 residues, fewer than 700 residues, fewer than 800 residues, fewer than 900 residues, fewer than 1000 residues, fewer than 2000 residues, fewer than 3000 residues, fewer than 4000 residues, fewer than 5000 residues, fewer than 6000 residues, fewer than 7000 residues, fewer than 8000 residues, or fewer than 9000 residues.

[0271] Referring to block 216, in some embodiments, each sequence group of the plurality of sequences comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more CpG sites. In some embodiments, each sequence group of the plurality of sequences comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more contiguous CpG sites. In some embodiments, each sequence group of the plurality of sequences consists of between 2 and 100 contiguous CpG sites in the human reference genome. In some embodiments, each sequence group of the plurality of sequences consists of between 2 and 50 contiguous CpG sites. In some embodiments, each sequence group of the plurality of sequences consists of between 50 and 100 contiguous CpG sites. In some embodiments, each sequence group of the plurality of sequences consists of at least two consecutive CpG sites.

[0272] In some embodiments, the plurality of sequences is constructed by dividing all or a portion of a reference genome (e.g., mammalian, human, etc.) into equal-sized sequence groups, where each sequence group represents a unique, equal-sized portion of the reference genome. In some embodiments, the plurality of sequences is constructed by dividing all or a portion of a reference genome (e.g., mammalian, human, etc.) into equal or unequal-sized sequence groups, where each sequence group represents a unique portion of the reference genome.

[0273] In some embodiments, the plurality of sequence groups is constructed by dividing all or a portion of a reference genome (e.g., a mammal, a human, etc.) into equal or unequal sized sequence groups, where each sequence group represents a corresponding portion of the reference genome. In such embodiments, a corresponding portion of the reference genome represented by one sequence group in the plurality of sequence groups may overlap with a corresponding portion of the reference genome represented by another sequence group in the plurality of sequence groups. In some such embodiments, the plurality of sequence groups is constructed by dividing all of a reference genome (e.g., a mammal, a human, etc.) into equal or unequal sized sequence groups, where each sequence group represents a corresponding overlapping or non-overlapping portion of the reference genome. In some embodiments, the plurality of sequence groups is constructed by dividing a portion of a reference genome (e.g., a mammal, a human, etc.) into equal or unequal sized sequence groups, where each sequence group represents an overlapping or non-overlapping portion of the reference genome.

[0274] In some embodiments, the plurality of sequences is constructed such that at least a portion of a region of the human genome implicated in the absence or presence of cancer is represented by the plurality of sequences, while other regions of the reference genome are not represented by the plurality of sequences. Regardless of the approach, each sequence represents a unique portion of the reference genome. In some embodiments, the size of such sequence groups ranges from 30 bps to 5000 bps, 30 bps to 4000 bps, 30 bps to 3000 bps, 30 bps to 2000 bps, 30 bps to 1000 bps, or 40 bps to 800 bps of the reference genome. In alternative embodiments, the size of such sequences ranges from 10,000 bps to 100,000 bps, 20,000 bps to 300,000 bps, 30,000 bps to 500,000 bps, 40,000 bps to 1,000,000 bps, 50,000 bps to 5,000,000 bps, or 100,000 to 25,000,000 bps of the reference genome.

[0275] In some embodiments, the portion of the reference genome is chromosomes 1 to 22 of the reference genome, or at least 25 percent, at least 30 percent, at least 35 percent, at least 40 percent, at least 45 percent, at least 50 percent, at least 55 percent, at least 60 percent, at least 65 percent, at least 70 percent, at least 75 percent, at least 80 percent, at least 85 percent, at least 90 percent, at least 95 percent, or at least 99 percent of the reference genome. In some such embodiments, each group of sequences represents between 10,000 and 100,000 bases, between 20,000 and 300,000 bases, between 30,000 and 500,000 bases, between 40,000 and 1,000,000 bases, between 50,000 and 5,000,000 bases, or between 100,000 and 25,000,000 bases of the reference genome.

[0276] In some embodiments, each group of sequences represents a particular site in the reference genome that has been identified as being associated with cancer.

[0277] In some embodiments, each group of sequences represents a specific region of the reference genome identified as associated with cancer by cancer- and / or tissue-specific methylation patterns in cfDNA compared to non-cancer controls.

[0278] In some embodiments, each group of sequences represents all or part of an enhancer, promoter, 5'UTR, exon, exon / inhibitor boundary, intron, intron / exon boundary, 3'UTR region, CpG shelf, CpG shore, or CpG island in a reference genome. For suitable definitions of such regions, see, e.g., Cavalcante and Santor, 2017, "annotatr: genomic regions in context," Bioinformatics 33(15) 2381-2383, where such annotations are documented for several different species.

[0279] In some embodiments, genomic regions with high variability or low mappability are excluded from representation in the plurality of sequence groups, for example, using the methods disclosed in Jensen et al., 2013, PLoS One 8; e57381. Also see Li and Freudenberg, 2014, Front. Genet. 5, p. 318 for mappability analysis.

[0280] Selection of human genome regions to use in the sequence collection

[0281] In some embodiments of the present disclosure, each sequence group of the plurality of sequence groups is derived from a panel of genomic regions designed for target selection of cancer-specific methylation patterns. In some embodiments, each such genomic region is derived from Table 2 of International Patent Application No. PCT / US2020 / 015082, entitled "Detecting Cancer, Cancer Tissue or Origin, or Cancer Type," published as International Publication No. WO 2020 / 154682, filed January 24, 2020, which is incorporated herein by reference, including the sequence listings referenced therein. SEQ ID NOs 452,706-483,478 of PCT / US2020 / 015082 provide further information regarding specific hypermethylated or hypomethylated target genomic regions. Records of these SEQ ID NOs identify target genomic regions that may be differentially methylated in samples from a specific combination of cancer types. The target genomic regions of SEQ ID NOs. 452, 706-483, and 478 in PCT / US2020 / 015082 are derived from List 6 of PCT / US2020 / 015082. Many of the same target genomic regions are also found in Lists 1-5 and 7-16 of PCT / US2020 / 015082. Each SEQ ID entry indicates the chromosomal location of the target genomic region relative to hg19, whether the cfDNA fragments enriched from that region are hypermethylated or hypomethylated, the sequence of one DNA strand of the target genomic region, and the combination or combinations of cancer types that are differentially methylated at that genomic region. Because the methylation status of some target genomic regions distinguishes multiple combinations of cancer types, each entry identifies a first cancer type and one or more second cancer types as set forth in Table 3 of PCT / US2020 / 015082 (including the sequence listing referenced therein).

[0282] In some embodiments, the plurality of sequences of the present disclosure includes a separate sequence group for each of at least 200, 500, 1,000, 5,000, 10,000, 15,000, 20,000, 30,000, 40,000, or 50,000 target genomic regions in any one of Lists 1-16, Lists 1-3, Lists 13-16, List 12, List 4, or Lists 8-11 of PCT / US2020 / 015082. In some embodiments, the plurality of sequences of the present disclosure includes a separate sequence group for each of at least 200, 500, 1,000, 5,000, 10,000, 15,000, 20,000, 30,000, 40,000, or 50,000 target genomic regions in any combination of one or more Lists 1-16 of PCT / US2020 / 015082 (e.g., Lists 1-3, Lists 13-16, List 12, List 4, or Lists 8-11, etc.).

[0283] In some embodiments, the plurality of sequences of the present disclosure includes a distinct group of sequences for each of at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% of the target genomic regions in any one of Lists 1-16 of PCT / US2020 / 015082. In some embodiments, the plurality of sequences of the present disclosure includes a distinct group of sequences for each of at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% of the target genomic regions in any combination of one or more Lists 1-16 (e.g., Lists 1-3, Lists 13-16, List 12, List 4, or Lists 8-11, etc.) of PCT / US2020 / 015082.

[0284] Additional selection of human genomic regions to be used in the sequence collection

[0285] In some embodiments of the present disclosure, each sequence group of the plurality of sequence groups is derived from a panel of genomic regions designed for targeted selection of cancer-specific methylation patterns. In some embodiments, each such genomic region is derived from Table 2 of International Patent Application No. PCT / US2019 / 053509, published as WO 2020 / 069350, filed September 27, 2019, entitled "Methylated Markers and Targeted Methylation Probe Panel," which is incorporated herein by reference, including the sequence listings referenced therein.

[0286] The sequence listing in WO 2020 / 069350 includes the following information: (1) SEQ ID NO, (2) (a) the chromosome or contig in which the CpG site is located, (b) a sequence identifier identifying the start and stop positions of the region, (3) the sequence corresponding to (2), and (4) whether the region is included based on its hypermethylation or hypomethylation score. The chromosome number, start, and stop positions are provided relative to the known human reference genome, GRCh37 / hgl9. The sequence of GRCh37 / hgl9 is available from the National Center for Biotechnology Information (NCBI), the Genome Reference Consortium, Genome Browser, provided by Santa Cruz Genomics Institute.

[0287] Generally, the group of sequences may include any of the CpG sites contained within the start / stop range of any of the target regions included in lists 1-8 of WO2020 / 069350.

[0288] In some embodiments, the plurality of sequences of the present disclosure includes a distinct group of sequences for each of at least 200, 500, 1,000, 5,000, 10,000, 15,000, 20,000, 30,000, 40,000, or 50,000 target genomic regions in any of Lists 1-8 of WO 2020 / 069350. In some embodiments, the plurality of sequences of the present disclosure includes a distinct group of sequences for each of at least 200, 500, 1,000, 5,000, 10,000, 15,000, 20,000, 30,000, 40,000, or 50,000 target genomic regions in any combination of Lists 1-8 of WO 2020 / 069350.

[0289] In some embodiments, the plurality of sequences of the present disclosure includes a distinct group of sequences for each of at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% of the target genomic regions in any one of Lists 1-8 of WO 2020 / 069350. In some embodiments, the plurality of sequences of the present disclosure includes a distinct group of sequences for each of at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% of the target genomic regions of any combination of Lists 1-8 of WO 2020 / 069350.

[0290] In some embodiments of the present disclosure, each sequence group of the plurality of sequence groups is derived from a panel of genomic regions designed for targeted selection of cancer-specific methylation patterns. In some embodiments, each such sequence group corresponds to a genomic region in any of Tables 1-24 of International Patent Application No. PCT / US2019 / 025358, published as WO 2019 / 195268, entitled "Methylated Markers and Targeted Methylation Probe Panels," filed April 2, 2019, which is incorporated by reference in its entirety.

[0291] In some embodiments, each group of sequences of the present disclosure maps to a genomic region set forth in one or more of Tables 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23 and / or 24 of WO 2019 / 195268.

[0292] In some embodiments, the plurality of sequence groups of the present disclosure are configured in their entirety to map to at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% of the genomic regions in one or more of Tables 1-24 of WO 2019 / 195268. In some such embodiments, a group of sequences in the plurality of sequence groups maps to a single unique corresponding genomic region in any of Tables 1-24 of WO 2019 / 195268. In some such embodiments, a group of sequences in the plurality of sequence groups of the present disclosure maps to 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 unique corresponding genomic regions in any combination of Tables 1-24 of WO 2019 / 195268.

[0293] In some such embodiments, a group of sequences in the plurality of sequences of the present disclosure map to a single unique corresponding genomic region in either Tables 2-10 or 16-24 of WO 2019 / 195268. In some such embodiments, a group of sequences in the plurality of sequences map to 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 unique corresponding genomic regions in any combination of Tables 2-10 or 16-24 of WO 2019 / 195268.

[0294] In some embodiments, one or more groups of sequences in the plurality of groups of sequences of the present disclosure are configured together to map to at least 30%, 40%, 50%, 60%, 70%, 80%, 90% or 95% of the genomic regions in Tables 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23 and / or 24 of WO2019 / 195268.

[0295] Assignment of cell-free fragments to cancer status

[0296] Block 218. Referring to block 218 of Figure 2B, the method proceeds by assigning a cell-free fragment cancer state to each cell-free fragment in each training set of cell-free fragments among the plurality of training sets of cell-free fragments, where the cell-free fragment cancer state is one of a first cancer state and a second cancer state that is a function of the output of the classifier when the methylation pattern of each cell-free fragment is input to the classifier.

[0297] In some embodiments, the classifier has the following form:

number

[0298] In some such embodiments, JPEG0007775192000007.jpg1096 is the first model of the first cancer condition.

[0299] In some such embodiments JPEG0007775192000008.jpg1096 is a second model of a second cancer state. In some embodiments, with respect to the first and second models, "fragment" refers to the methylation pattern of each cell-free fragment. In some embodiments, the cancer state of each fragment cell-free fragment is assigned the first cancer state if R(fragment) meets a threshold. In some embodiments, the threshold is any value between 1 and 10. In some embodiments, the threshold is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10.

[0300] In some embodiments, the first model is a first mixed model including a first plurality of sub-models, and the second model is a second mixed model including a second plurality of sub-models, each sub-model of the first and second plurality of sub-models representing an independent corresponding methylation model for a source of cell-free fragments in a corresponding biological sample.

[0301] In some embodiments, the subject's cancer state is one of a plurality of cancer states (e.g., where the plurality of cancer states includes N cancer states). In some such embodiments, the classifier has the following format:

number

[0302] In some such embodiments JPEG0007775192000010.jpg1163 is a third model of a third cancer condition in the plurality of cancer conditions. JPEG0007775192000011.jpg13102 is the Nth model for the Nth cancer state among multiple cancer states.

[0303] Examples of mixture models for use in accordance with embodiments herein are described in U.S. Patent Application No. 62 / 847,223, entitled "Model-Based Featurization and Classification," filed May 13, 2019, which is incorporated herein by reference in its entirety.

[0304] In some embodiments, each independent corresponding methylation model is one of a binomial model, a beta-binomial model, an independent-site model, or a Markov model. In some embodiments, two or more submodels in the first plurality of submodels are independent-site models, and two or more submodels in the second plurality of submodels are independent-site models.

[0305] For example, U.S. Patent Application No. 62 / 983,443, filed February 28, 2020, entitled "Identifying Methylation Patterns that Discriminate or Indicate a Cancer Condition," which is incorporated herein by reference in its entirety, discloses methods for identifying methylation patterns that identify a particular cancer state of a subject. Specifically, in some embodiments, each cancer state (e.g., cancer of origin) of a group of cancer states corresponds to a respective pattern of aberrant methylation (e.g., modified methylation patterns) across a reference genome or a subset of a reference genome (e.g., assessed by targeted panel sequencing). To determine the cancer state of a particular subject, the method evaluates multiple genomic regions of interest and, for each genomic region of the multiple genomic regions, generates a corresponding count of fragments having a methylation pattern that maps to the respective genomic region (e.g., there is a respective count of fragments for each possible methylation pattern identified in the fragments that map to each genomic region). The method then compares the fragment counts across the plurality of genomic regions for the subject to a database (e.g., a library) of methylation patterns corresponding to different cancer states (e.g., each cancer state has a corresponding fragment count for each subset of genomic regions within the plurality of genomic regions) to determine a probable cancer state for the subject, where the cancer state corresponds to cancer versus non-cancer, cancer type, and / or tissue of origin. In some embodiments, the method is used to identify the cancer state of a subject for input into downstream applications (e.g., to estimate the subject's tumor fraction and / or determine minimal residual disease). In some embodiments, the plurality of sequences used in the present disclosure are selected to represent a portion of the genome identified in U.S. Patent Application No. 62 / 983,443, including methylation patterns associated with any single or any combination of cancers evaluated in U.S. Patent Application No. 62 / 983,443.

[0306] As another example, U.S. Patent Application No. 15 / 931,022, entitled "Model-Based Featurization and Classification," filed May 13, 2020, which is incorporated herein by reference in its entirety, discloses the development of probabilistic models using the methylation states of genomic regions (e.g., determined from fragments represented by sequence reads that map to the genomic regions) to identify methylation signatures corresponding to distinct cancer states. In some embodiments, the plurality of sequences used in this disclosure are selected to represent portions of the genome identified in U.S. Patent Application No. 15 / 931,022, including methylation patterns associated with any single or any combination of cancers evaluated in U.S. Patent Application No. 15 / 931,022.

[0307] Other methods for cancer classification in nucleic acid fragments include those disclosed in, for example, U.S. patent application Ser. No. 62 / 948,129, entitled "Cancer Classification using Patch Convolutional Neural Networks," filed December 13, 2019; U.S. patent application Ser. No. 16 / 352,739, entitled "Method and System for Selecting, Managing, and Analyzing Data of High Dimensionality," filed March 13, 2019; U.S. patent application Ser. No. 16 / 428,575, entitled "Convolutional Neural Network Systems and Methods for Data Classification," filed May 31, 2019; and U.S. patent application Ser. No. 62 / 985,258, entitled "Systems and Methods for Cancer Condition Determination using Autoencoders," filed March 4, 2020, each of which is incorporated herein by reference in its entirety.

[0308] In some embodiments, the classifier is a multivariate logistic regression, a neural network, a convolutional neural network, a support vector machine (SVM), a decision tree, a regression algorithm, or a supervised clustering model.

[0309] Logistic regression algorithms, including multivariate logistic regression, are disclosed in Agresti, An Introduction to Categorical Data Analysis, 1996, Chapter 5, pp. 103-144, John Wiley & Son, New York, which is incorporated herein by reference.

[0310] Neural network algorithms, including convolutional neural network algorithms, are disclosed in Vincent et al., 2010, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion,” J Mach Learn Res 11, pp. 3371-3408; Larochelle et al., 2009, “Exploring strategies for training deep neural networks,” J Mach Learn Res 10, pp. 1-40; and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of which is incorporated herein by reference.

[0311] The SVM algorithm is described in Cristianini and Shawe-Taylor, 2000, “An Introduction to Support Vector Machines,” Cambridge University Press, Cambridge; Boser et al., 1992, “A training algorithm for optimal margin classifiers,” in Proceedings of the 5 th Annual ACM Workshop on Computational Learning Theory, ACM Press, Pittsburgh, Pa., pp. 142-152; Vapnik, 1998, Statistical Learning Theory, Wiley, New York; Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY; Duda, Pattern Classification, Second Edition, 2001, John Wiley & Sons, Inc., pp. 259, 262-265; and Hastie, 2001, The Elements of Statistical Learning, Springer, New York; and Furey et al., 2000, Bioinformatics 16, 906-914, each of which is incorporated herein by reference in its entirety. When used for classification, SVM separates a given set of binary labeled data training set (for example, by tumor proportion value) using a hyperplane that is maximally far from the labeled data. When linear separation is not possible, SVMs can be combined with techniques called "kernels" that automatically achieve a nonlinear mapping to the feature space: the hyperplane that the SVM finds in the feature space corresponds to a nonlinear decision boundary in the input space.

[0312] Decision trees are generally described by Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 395-396, which is incorporated herein by reference. Tree-based methods divide the feature space into a set of rectangles and then fit a model (such as a constant) to each. In some embodiments, the decision tree is a random forest regression. One specific algorithm that can be used is a classification and regression tree (CART). Other specific decision tree algorithms include, but are not limited to, ID3, C4.5, MART, and random forest. CART, ID3, and C4.5 are described in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 396-408 and pp. 411-412, which are incorporated herein by reference. CART, MART, and C4.5 are described in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, Chapter 9, which is incorporated herein by reference in its entirety. Random forests are described in Breiman, 1999, "Random Forests--Random Features," Technical Report 567, Statistics Department, UC Berkeley, September 1999, which is incorporated herein by reference in its entirety.

[0313] Clustering is described in Duda and Hart, Pattern Classification and Scene Analysis, 1973, John Wiley & Sons, Inc., New York (hereinafter "Duda 1973"), pages 211-256, which is incorporated herein by reference in its entirety. As described in section 6.7 of Duda 1973, the clustering problem is described as finding natural groupings within a data set. To identify natural groupings, two problems are addressed. First, a method for measuring the similarity (or dissimilarity) between two samples is determined. This metric (similarity metric) is used to ensure that samples in one cluster are more similar to each other than to samples in other clusters. Second, a mechanism for dividing the data into clusters using the similarity metric is determined.

[0314] Similarity measures are discussed in Section 6.7 of Duda 1973, which states that one way to begin investigating clustering is to define a distance function and calculate a matrix of distances between all pairs of samples in the training set. If distance is a good measure of similarity, the distance between reference entities in the same cluster will be significantly smaller than the distance between reference entities in different clusters. However, as noted on page 215 of Duda 1973, clustering does not need to use a distance metric. For example, to compare two vectors x and x', it is possible to use a nonmetric similarity function s(x,x'). Conventionally, s(x,x') is a symmetric function whose value is large if x and x' are "similar" in some way. An example of a nonmetric similarity function s(x,x') is given on page 218 of Duda 1973.

[0315] After choosing a method for measuring "similarity" or "dissimilarity" between points in a data set, clustering requires a criterion function that measures the clustering quality of any partition of the data. The partition of the data set that maximizes the criterion function is used to cluster the data. See Duda 1973, page 217. Criterion functions are described in Duda 1973, section 6.8.

[0316] Recently, Duda et al., Pattern Classification, 2 ndPublished by John Wiley & Sons, Inc. New York. Pages 537-563 provide details on clustering. More detailed information on clustering techniques can be found in Kaufman and Rousseeuw, 1990, Finding Groups in Data: An Introduction to Cluster Analysis, Wiley, New York, NY; Everitt, 1993, Cluster analysis (3rd ed.), Wiley, New York, NY; and Backer, 1995, Computer-Assisted Reasoning in Cluster Analysis, Prentice Hall, Upper Saddle River, New Jersey, each of which is incorporated herein by reference. Specific exemplary clustering techniques that can be used in the present disclosure include, but are not limited to, hierarchical clustering (agglomerative clustering using nearest neighbor, farthest neighbor, average linkage, centroid, or sum-of-squares algorithms), k-means clustering, fuzzy k-means clustering, and Jarvis-Patrick clustering. Such clustering may be on a first set of features {p,...,pN-K} (or principal components derived from the first set of features). In some embodiments, the clustering includes unsupervised clustering, in which no predetermined opinion is imposed on what clusters should be formed when the training set is clustered.

[0317] Feature Identification

[0318] Block 220. Referring to block 220 of Figure 2B, the method proceeds by determining, for each sequence group of the plurality of sequence groups, a corresponding measure of association I between (a) the subject's cancer state for each training subject of the plurality of training subjects and (b) the cancer state of each cell-free fragment of the corresponding training set of cell-free fragments mapped to each sequence group.

[0319] In some embodiments, with respect to block 222, the measure of association is correlation. With reference to block 224, in some embodiments, the correlation is a Pearson correlation coefficient. With reference to block 226, in some embodiments, the correlation is performed using an adjusted correlation coefficient, a weighted correlation, a reflective correlation coefficient, or a scaled correlation coefficient.

[0320] In some embodiments, the measure of relatedness is a calculation of mutual information. See, e.g., Song et al., 2012, "Comparison of co-expression measures: mutual information, correlation, and model based indices," BMC Bioinformatics 13, 328. For example, in some embodiments, the mutual information is calculated according to Figure 8. As described in Figure 8, the mutual information between the training subject label Y (cancer type A or B in the case of two cancer types) and the sequence group feature X is calculated by mutual information. In fact, Figure 8 provides a method for calculating mutual information under the assumption that the subject has the same probability of having either cancer type A or B (P(Y=A)=P(Y=B)). In some specific embodiments, the measure of relatedness is mutual information calculated as follows:

number

[0321] In some such embodiments, i and j are independent indices for a set of cancer states (e.g., a first and a second cancer state).i is the number of training subjects having cancer status i among the plurality of training subjects (e.g., i is a first cancer status, or alternatively, i is a second cancer status, etc.). In some embodiments, y j is the number of training subjects assigned cancer state j among multiple training subjects having one or more cell-free fragments that map to each sequence group (e.g., j is the first cancer state, or alternatively, j is the second cancer state, etc.). For two cancer states, this association measure takes the form:

number

[0322] In some such embodiments, the measure of association is determined based on at least a) the number of training subjects who have a first cancer state and who have one or more cell-free fragments within each sequence group assigned to the first cancer state, b) the number of training subjects who have the first cancer state but who have one or more cell-free fragments within each sequence group assigned to the second cancer state, c) the number of training subjects who have the second cancer state and who have one or more cell-free fragments within each sequence group assigned to the second cancer state, and d) the number of training subjects who have the second cancer state but who have one or more cell-free fragments within each sequence group assigned to the first cancer state.

[0323] In one embodiment, the function

number

number

[0324] In some embodiments, when there are two possible cancer states, the measure of relatedness is a distance metric. Table 1 provides examples of such distance metrics.

[0325] [Table 1] JPEG0007775192000018.jpg124153

[0326] In some embodiments, the calculation of the measure of association determines a measure of association for each sequence group of the plurality of sequence groups, where each training subject in the plurality of training subjects has one of the plurality of cancer conditions. In some such embodiments, the measure of association is calculated as follows:

number

[0327] JPEG0007775192000020.jpg81153

[0328] Block 228. Referring to block 228 of FIG. 2B, the method continues by identifying a plurality of features for estimating the proportion of cell origin of interest as a subset of a plurality of sequences, each sequence in the subset of a plurality of sequences meeting a selection criterion based on a corresponding measure of relatedness for each sequence.

[0329] In some embodiments, the selection criteria specify the selection of sequences having one of the top N relatedness measures, where N is a positive integer greater than or equal to 50. In some embodiments, N is between 500 and 5000. In some embodiments, N is between 800 and 1500. In some embodiments, N is at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, at least 1200, at least 1300, at least 1400, or at least 1500.

[0330] In some embodiments, referring to block 230, the selection criteria specify the selection of sequences having one of the top N measures of relatedness, where N is a positive integer greater than or equal to 50 (e.g., at least 50 sequences having the highest measures of relatedness are selected as features).

[0331] In some embodiments, the plurality of features includes at least 10, at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, at least 1200, at least 1300, at least 1400, or at least 1500 features. In some embodiments, the plurality of features includes 500-5000, 800-1500, or more than 1500 features.

[0332] Estimation of cell source fraction

[0333] In some embodiments, after identifying a plurality of features (e.g., a subset of sequences) for estimating the cell origin fraction of a subject, the method further includes estimating the cell origin fraction of the test subject based on at least the plurality of features.

[0334] In some embodiments, the method performs estimation of cell origin or tumor fraction by a procedure that includes obtaining, in electronic form, a corresponding methylation pattern for each cell-free fragment of a plurality of test cell-free fragments (e.g., from a test subject for which cancer classification is desired), wherein the corresponding methylation pattern for each cell-free fragment (i) is determined by methylation sequencing of one or more nucleic acid samples containing each fragment in a biological sample obtained from the test subject, and (ii) includes a methylation state of each CpG site of a corresponding plurality of CpG sites in each fragment. The procedure further includes mapping each cell-free fragment of the plurality of test cell-free fragments to a sequence group in a plurality of sequence groups, thereby obtaining a plurality of test sets of cell-free fragments, each test set of cell-free fragments mapped to a different sequence group in the plurality of sequence groups. The procedure continues by assigning a cancer state to each cell-free fragment in each test set of cell-free fragments in the plurality of test sets of cell-free fragments, the cancer state being a function of the output of a classifier when the methylation pattern of each cell-free fragment is input to the classifier. The method includes calculating a first representative value of the number of cell-free fragments from test subjects assigned a first cancer status in each test set of cell-free fragments across the subset of the plurality of sequence groups, and calculating a second representative value of the number of cell-free fragments from test subjects in each test set of cell-free fragments across the subset of the plurality of sequence groups. The method uses the first and second representative values ​​to estimate a cell-source fraction for the test subject.

[0335] In some embodiments, the second cancer status comprises the absence of cancer, and the estimated cell source fraction for the test subject comprises a tumor fraction for the test subject.

[0336] For example, in some embodiments, an estimate of tumor fraction is calculated based on the assumption that one or more methylation status patterns in a test biological sample (e.g., cfDNA and / or plasma) are tumor-derived, and that the frequency of such tumor-derived methylation patterns is directly proportional to the proportion of cancer cells to normal cells (e.g., tumor fraction).

[0337] There are various methods for determining such fractions, some of which are described in U.S. patent application Ser. No. 16 / 719,902, filed December 18, 2019, entitled "Systems and Methods for Estimating Cell Source Fractions using Methylation Information," and U.S. patent application Ser. No. 16 / 850,634, filed April 16, 2020, entitled "Systems and Methods for Tumor Fraction Estimation from Small Variants," both of which are incorporated herein by reference in their entireties.

[0338] In some embodiments, the first representative value is the arithmetic mean, weighted mean, mid-range, mid-hinge, trinomial mean, winsorized mean, average, or mode of the number of cell-free fragments from the plurality of test subjects assigned the first cancer state in each test set of cell-free fragments across the subset of the plurality of sequence groups. In some embodiments, the second representative value is the arithmetic mean, weighted mean, mid-range, mid-hinge, trinomial mean, winsorized mean, average, or mode of the number of cell-free fragments from the plurality of test subjects in each test set of cell-free fragments across the subset of the plurality of sequence groups. In some embodiments, estimating the fraction of cell origin comprises dividing the first representative value by the second representative value. In some embodiments, the cancer state of each subject of each training subject among the plurality of training subjects is selected from a plurality of cancer states. In some embodiments, a corresponding measure of central tendency is determined for each cancer state of the plurality of cancer states. In some such embodiments, estimating the fraction of cell origin comprises dividing the first representative value by the sum of each of the other measures of central tendency.

[0339] In some embodiments, the tumor fraction of the test subject is between 0.003 and 1.0. In some embodiments, the tumor fraction of the test subject is within the range of 0.001 and 1.0. In some embodiments, the tumor fraction of the test subject is at least 0.001, at least 0.005, at least 0.01, at least 0.05, at least 0.1, at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, at least 0.8, at least 0.9, or at least 1.0.

[0340] In some embodiments, determining the subject's cell source (e.g., tumor) proportion further identifies the subject's cancer of origin. In some embodiments, the first and / or second cancer status includes the tissue of origin (e.g., where the cancer is believed to have arisen). In some embodiments, the first and / or second cancer status includes the stage of the cancer (e.g., stage I, II, III, or IV).

[0341] In some embodiments, the cancer of origin comprises a first cancerous condition selected from the group consisting of non-cancer, breast cancer, lung cancer, prostate cancer, colorectal cancer, renal cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head / neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, gastric cancer, nasopharyngeal cancer, liver cancer, or a combination thereof.

[0342] In some embodiments, the cancer of origin includes at least a first cancer condition and a second cancer condition, each selected from the group consisting of breast cancer, lung cancer, prostate cancer, colorectal cancer, renal cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head / neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, gastric cancer, nasopharyngeal cancer, liver cancer, or a combination thereof.

[0343] In some embodiments, the first and / or second cancer condition comprises a stage of breast cancer, a stage of lung cancer, a stage of prostate cancer, a stage of colorectal cancer, a stage of renal cancer, a stage of uterine cancer, a stage of pancreatic cancer, a stage of esophageal cancer, a stage of lymphoma, a stage of head / neck cancer, a stage of ovarian cancer, a stage of hepatobiliary cancer, a stage of melanoma, a stage of cervical cancer, a stage of multiple myeloma, a stage of leukemia, a stage of thyroid cancer, a stage of bladder cancer, a stage of gastric cancer, a stage of nasopharyngeal cancer, a stage of liver cancer, or a combination thereof.

[0344] In some embodiments, determining the cell source (e.g., tumor) fraction of the test subject further includes providing a treatment recommendation (e.g., cancer treatment) to the test subject, the treatment recommendation being based at least in part on the cell source fraction (e.g., how advanced the disease is) and the cancer of origin.

[0345] In some embodiments, the method further comprises determining the cell source (e.g., tumor) fraction of the test subject at one or more time points (e.g., before or after treatment) to monitor disease progression or to monitor treatment effectiveness (e.g., treatment efficacy). For example, in some embodiments, an increase in tumor fraction over time (e.g., at a second, later time point) indicates disease progression; conversely, in some embodiments, a decrease in tumor fraction over time (e.g., at a second, later time point) indicates successful treatment.

[0346] For example, in some embodiments, the method further includes administering a therapeutic regimen to the test subject based at least in part on the value of the cell origin fraction of the test subject. In some embodiments, the therapeutic regimen includes administering a cancer drug to the test subject. In some embodiments, the cancer drug is a hormone, immunotherapy, x-ray, or anticancer drug. In some embodiments, the cancer drug is lenalidomide, pembrolizumab, trastuzumab, bevacizumab, rituximab, ibrutinib, human papillomavirus tetravalent (types 6, 11, 16, and 18) vaccine, pertuzumab, pemetrexed, nilotinib, denosumab, abiraterone acetate, promacta, imatinib, everolimus, palbociclib, erlotinib, bortezomib, or generic equivalents thereof.

[0347] In some embodiments, the test subject has been treated with a cancer drug, and the method further includes assessing the test subject's response to the cancer drug using the cell origin percentage of the test subject. In some embodiments, the cancer drug is a hormone, immunotherapy, x-ray, or anticancer drug. In some embodiments, the cancer drug is lenalidomide, pembrolizumab, trastuzumab, bevacizumab, rituximab, ibrutinib, human papillomavirus tetravalent (types 6, 11, 16, and 18) vaccine, pertuzumab, pemetrexed, nilotinib, denosumab, abiraterone acetate, promacta, imatinib, everolimus, palbociclib, erlotinib, bortezomib, or generic equivalents thereof.

[0348] In some embodiments, the test subject has been treated with a cancer drug and the method further comprises using the cell source fraction of the test subject to determine whether to increase or discontinue the cancer drug in the test subject. In some embodiments, the test subject has undergone a surgical intervention to address cancer and the method further comprises using the cell source fraction of the test subject to assess the test subject's status in response to the surgical intervention.

[0349] In some embodiments, the method is repeated at each of a plurality of time points (e.g., two or more time points, three or more time points, four or more time points) across an epoch, thereby obtaining a corresponding cell source (e.g., tumor) proportion in the plurality of cell source (e.g., tumor) proportions for the test subject at each time point, and using the plurality of cell source (e.g., tumor) proportions to determine the state or progression of the subject's disease state during the epoch in terms of an increase or decrease in the first cell source (e.g., tumor) proportion across the epoch.

[0350] In some such embodiments, the epoch is a period of several months, and each of the plurality of time points is a different time point within the period of several months. In some embodiments, the period of several months is 1 to 4 months, 4 to 8 months, 8 to 12 months, 12 to 18 months, 18 to 24 months, or more than 24 months. In some embodiments, the period of several months is less than 4 months.

[0351] In some embodiments, the epoch is a period of several years, and each of the plurality of time points is a different time point within the period of several years. In some embodiments, the period of several years is between 2 and 10 years. In some embodiments, the period of years is between 1 and 5 years, between 5 and 10 years, between 10 and 15 years, between 15 and 20 years, or more than 20 years.

[0352] In some embodiments, an epoch is a period of several hours, and each of the plurality of time points is a different time point within the period of several hours. In some embodiments, the period of several hours is between 1 hour and 6 hours. In some embodiments, the period of several hours is between 1 hour and 3 hours, between 3 hours and 6 hours, between 6 hours and 9 hours, between 9 hours and 12 hours, between 12 hours and 18 hours, between 18 hours and 24 hours, or greater than 24 hours.

[0353] In some embodiments, the method further comprises altering the diagnosis of the test subject if the subject's first cell source (e.g., tumor) proportion is observed to change by a threshold amount over the epoch. In some embodiments, the method further comprises altering the subject's prognosis if the subject's first cell source (e.g., tumor) proportion is observed to change by a threshold amount over the epoch. In some embodiments, the method further comprises altering the subject's treatment if the subject's first cell source (e.g., tumor) proportion is observed to change by a threshold amount over the epoch. In some of the foregoing embodiments, the threshold is greater than 1 percent, greater than 5 percent, greater than 10 percent, greater than 20 percent, greater than 30 percent, greater than 40 percent, or greater than 50 percent. In some embodiments, the threshold is greater than 2-fold, greater than 3-fold, greater than 4-fold, or greater than 5-fold.

[0354] In certain embodiments, the method is performed at a first time point before cancer treatment (e.g., before resection surgery or therapeutic intervention) and at a second time point after cancer treatment (e.g., after resection surgery or therapeutic intervention), and the disclosed method is used to monitor the effectiveness of the treatment by comparing the cell source (e.g., tumor) proportion determined by the disclosed method at each time point. For example, if the tumor proportion at the second time point is reduced compared to the tumor proportion at the first time point, the treatment is considered successful. However, if the tumor proportion at the second time point is increased compared to the tumor proportion at the first time point, the treatment is considered unsuccessful. In other embodiments, both the first and second time points are before cancer treatment (e.g., before resection surgery or therapeutic intervention). In yet other embodiments, both the first and second time points are after cancer treatment (e.g., before resection surgery or therapeutic intervention), and the method is used to monitor the effectiveness or loss of effectiveness of the treatment. In yet other embodiments, biological samples (cfDNA samples) may be obtained from a test subject (e.g., a cancer patient) at first and second time points and analyzed, for example, to monitor cancer progression, determine whether the cancer is in remission (e.g., after treatment), monitor or detect residual disease or disease recurrence, or monitor treatment (e.g., therapeutic) effectiveness.

[0355] One skilled in the art will readily appreciate that biological samples can be obtained from a test subject (e.g., a cancer patient) over any number of time points and analyzed according to the methods of the present disclosure to monitor the patient's cancer status (e.g., by tumor prevalence). In some embodiments, the first and second time points are for an amount of time ranging from about 15 minutes up to about 30 years, e.g., about 30 minutes, e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or about 24 hours, e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, or about 30 days, e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 months, or e.g., about 1, 1.5, 2, 2.5, 3, 3 They are divided by the amount of time: 0.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, 10.5, 11, 11.5, 12, 12.5, 13, 13.5, 14, 14.5, 15, 15.5, 16, 16.5, 17, 17.5, 18, 18.5, 19, 19.5, 20, 20.5, 21, 21.5, 22, 22.5, 23, 23.5, 24, 24.5, 25, 25.5, 26, 26.5, 27, 27.5, 28, 28.5, 29, 29.5 or approximately 30 years. In other embodiments, biological samples may be obtained from a patient at least once every three months, at least once every six months, at least once a year, at least once every two years, at least once every three years, at least once every four years, or at least once every five years.

[0356] Determining the estimated cell source percentage of the test subject

[0357] Block 302. Referring to block 302 of FIG. 3A, a method of estimating a cell source fraction of a subject (e.g., a test subject) is provided. In some embodiments, the subject is a human. In some embodiments, the subject is a human being (e.g., a male, female, or child) at any stage of life. In some embodiments, the cell source fraction of the subject is derived from a single cell source. In some embodiments, the cell source fraction of the subject is derived from two or more cell sources. In some embodiments, the cell source fraction is as described with respect to block 202 above.

[0358] Block 304. Referring to block 304, the method continues by obtaining, in electronic form, a corresponding methylation pattern for each cell-free fragment of a plurality of cell-free fragments (e.g., the plurality of cell-free fragments is derived from a biological sample of the subject), where the corresponding methylation pattern for each cell-free fragment (i) is determined by methylation sequencing of one or more nucleic acid samples comprising each fragment in a biological sample obtained from the subject, and (ii) includes the methylation state of each CpG site among the corresponding plurality of CpG sites in each fragment. In some embodiments, referring to block 306, the plurality of cell-free fragments have an average length of less than 500 nucleotides. In some embodiments, the cell-free fragments are derived from a biological sample, as described above with respect to block 204.

[0359] In some embodiments, the biological sample comprises or consists of a subject's blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid. In such embodiments, the biological sample may include a subject's blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid, as well as other components of the subject (e.g., solid tissue, etc.).

[0360] Such biological samples contain cell-free nucleic acid fragments (e.g., cfDNA fragments). In some embodiments, the biological sample is processed to extract cell-free nucleic acids in preparation for sequencing analysis. As a non-limiting example, in some embodiments, cell-free nucleic acid fragments are extracted from a biological sample (e.g., a blood sample) collected from a subject in a K2 EDTA tube. If the biological sample is blood, the sample is processed by first double-spinning the biological sample at 1000 g for 10 minutes and then spinning the resulting plasma at 2000 g for 10 minutes within 2 hours of collection. The plasma is then stored at -80°C in 1 mL aliquots. In this manner, an appropriate amount of plasma (e.g., 1-5 mL) is prepared from the biological sample for cell-free nucleic acid extraction. In some such embodiments, cell-free nucleic acids are extracted using a QIAamp Circulating Nucleic Acid Kit (Qiagen) and eluted in DNA suspension buffer (Sigma). In some embodiments, the purified cell-free nucleic acids are stored at -20°C until use. See, e.g., Swanton, et al., 2017, "Phylogenetic ctDNA analysis depicts early stage lung cancer evolution," Nature, 545(7655): 446-451, which is incorporated herein by reference. Other equivalent methods can be used to prepare cell-free nucleic acids from biological methods for sequencing purposes, and all such methods are within the scope of the present disclosure.

[0361] In some embodiments, the cell-free nucleic acid fragments obtained from the biological sample are any form of nucleic acid as defined herein, or a combination thereof. For example, in some embodiments, the cell-free nucleic acid obtained from the biological sample is a mixture of RNA and DNA.

[0362] In one embodiment, cell-free nucleic acid fragments are treated to convert unmethylated cytosines to uracil. In one embodiment, the method uses bisulfite treatment of DNA, which converts unmethylated cytosines to uracil without converting methylated cytosines. For example, a commercially available kit such as EZ DNA Methylation™ - Gold, EZ DNA Methylation™ - Direct, or an EZ DNA Methylation™ - Lightning kit (available from Zymo Research Corp, Irvine, CA) is used for bisulfite conversion. In another embodiment, the conversion of unmethylated cytosines to uracil is achieved using an enzymatic reaction. For example, the conversion can be performed using a commercially available kit for converting unmethylated cytosines to uracil, such as APOBEC-Seq (NEBiolabs, Ipswich, MA).

[0363] A sequencing library is prepared from the converted cell-free nucleic acid fragments.Optionally, the sequencing library is enriched for cell-free nucleic acid fragments or genomic regions that are informative about cellular origin using multiple hybridization probes.Hybridization probes are short oligonucleotides that hybridize to specifically designated cell-free nucleic acid fragments or target regions, and enrich for these fragments or regions for subsequent sequencing and analysis.In some embodiments, hybridization probes are used to perform targeted deep analysis of a set of specific CpG sites that are informative about cellular origin.Once prepared, the sequencing library or a portion thereof is sequenced to obtain multiple sequence reads.

[0364] In some embodiments, the sequencing comprises methylation sequencing. In some embodiments, the methylation sequencing is paired-end sequencing. In some embodiments, the methylation sequencing is single-read sequencing. In some embodiments, the methylation sequencing is whole-genome methylation sequencing. In some embodiments, the methylation sequencing is targeted sequencing using a plurality of nucleic acid probes, wherein each sequence group of the plurality of sequences is associated with at least one corresponding nucleic acid probe in the plurality of nucleic acid probes. In some embodiments, each sequence group of the plurality of sequences is associated with at least two corresponding nucleic acid probes among the plurality of nucleic acid probes.

[0365] In some embodiments, the plurality of nucleic acid probes (e.g., probes used for target sequencing) includes 1,000 or more nucleic acid probes, 2,000 or more nucleic acid probes, 3,000 or more nucleic acid probes, 4,000 or more nucleic acid probes, 5,000 or more nucleic acid probes, 10,000 or more nucleic acid probes, 20,000 or more nucleic acid probes, or 30,000 or more nucleic acid probes. In some embodiments, the plurality of nucleic acid probes is between 1,000 nucleic acid probes and 30,000 nucleic acid probes.

[0366] In some embodiments, methylation sequencing (e.g., performed according to any methylation sequencing method described herein or known in the art) detects one or more 5-methylcytosines (5mC) and / or 5-hydroxymethylcytosines (5hmC) in each fragment.

[0367] In some embodiments, methylation sequencing comprises converting one or more unmethylated cytosines or one or more methylated cytosines into one or more corresponding uracils in the sequence reads of each fragment.In some embodiments, one or more uracils are detected as one or more corresponding thymines during methylation sequencing.In some embodiments, the conversion of one or more unmethylated cytosines or one or more methylated cytosines comprises chemical conversion, enzymatic conversion, or a combination thereof.

[0368] In some embodiments, the methylation status of each CpG site among the corresponding plurality of CpG sites in each fragment is: a) a methylated status if methylation sequencing determines that each CpG site is methylated, b) an unmethylated status if methylation sequencing determines that each CpG site is unmethylated, or c) flagged as "other" if the methylation status of each CpG site cannot be called methylated or unmethylated.

[0369] Block 308. Referring to block 308, the method maps each cell-free fragment of the plurality of cell-free fragments to a sequence group in the plurality of sequence groups, thereby obtaining a plurality of sets of cell-free fragments, each set of cell-free fragments mapping to a different sequence group in the plurality of sequence groups.

[0370] In some embodiments, referring to block 310, the plurality of sequence groups consists of 1,000 to 100,000 sequence groups. In some embodiments, the plurality of sequence groups consists of 15,000 to 80,000 sequence groups. In some embodiments, the plurality of sequence groups consists of any number of sequence groups, as described with respect to block 210 above.

[0371] Referring to block 312, in some embodiments, each sequence group in the plurality of sequence groups has, on average, 10 to 1200 residues. In some embodiments, each sequence group in the plurality of sequence groups has, on average, 10 to 10,000 residues. In some embodiments, each sequence group in the plurality of sequence groups has, on average, 10 to 500 residues. In some embodiments, each sequence group in the plurality of sequence groups has, on average, 10 to 100 residues. In some embodiments, each sequence group in the plurality of sequence groups has, on average, 25 to 100 residues. In some embodiments, each sequence group in the plurality of sequence groups has, on average, 5000 to 10,000 residues.

[0372] Further, with respect to block 314, in some embodiments, each sequence group of the plurality of sequence groups comprises or consists of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more CpG sites. In some embodiments, each sequence group of the plurality of sequence groups consists of 2 to 100 contiguous CpG sites in the human reference genome. In some embodiments, each sequence group of the plurality of sequence groups consists of 2 to 50 contiguous CpG sites. In some embodiments, each sequence group of the plurality of sequence groups consists of 50 to 100 contiguous CpG sites. In some embodiments, each sequence group of the plurality of sequence groups consists of at least two contiguous CpG sites.

[0373] Block 316. With reference to block 316, the method continues by assigning a cell-free fragment cancer state to each cell-free fragment in each training set of cell-free fragments among the plurality of training sets of cell-free fragments, where the cell-free fragment cancer state is one of a first cancer state and a second cancer state that is a function of the output of the classifier when the methylation pattern of each cell-free fragment is input to the classifier. With reference to block 318, in some embodiments, the first cancer state is cancer and the second cancer state is the absence of cancer. In some embodiments, the first cancer state is cancer and the second cancer state is the absence of cancer. In some embodiments, the cell-free fragment cancer state is one of a plurality of cancer states (e.g., as described above with reference to block 206).

[0374] In some embodiments, the classifier used to assign the cell-free fragment status includes a first model for a first cancer status and a second model for a second cancer status, where the first model is a first mixture model including a first plurality of submodels, and the second model is a second mixture model including a second plurality of submodels, each submodel of the first and second plurality of submodels representing an independent corresponding methylation model for a source of cell-free fragments in a corresponding biological sample. In some embodiments, the classifier has the form of Equation (1) or Equation (3):

[0375] Block 320. Referring to block 320 of FIG. 3B, the method further includes calculating a first representative value of the number of cell-free fragments from subjects assigned the first cancer state in each set of cell-free fragments across the plurality of sequence groups. In some embodiments, referring to block 322, the first representative value is the arithmetic mean, weighted mean, mid-range, mid-hinge, trinomial mean, winsorized mean, average, or mode of the number of cell-free fragments from subjects assigned the first cancer state in each set of cell-free fragments across the plurality of sequence groups.

[0376] Block 324. Referring to block 324, the method further includes calculating a second representative value of the number of cell-free fragments from subjects assigned the second cancer state in each set of cell-free fragments across the plurality of sequence groups. In some embodiments, referring to block 326, the second representative value is the arithmetic mean, weighted mean, mid-range, mid-hinge, trinomial mean, winsorized mean, average, or mode of the number of cell-free fragments from subjects assigned the first cancer state in each set of cell-free fragments across the plurality of sequence groups.

[0377] Block 328. Referring to block 328, the method proceeds by estimating a cell origin fraction of the subject using the first representative value and the second representative value. In some embodiments, the cell origin fraction includes a tumor fraction. Regarding block 330, in some embodiments, estimating the tumor fraction includes dividing the first representative value by the second representative value.

[0378] In some embodiments, the cell source fraction is used as a basis or partial basis for determining treatment options for treating a cell source-related disease (e.g., cancer) in a test subject. In some embodiments, the cell source fraction is used as a basis for monitoring treatment. In some embodiments, given a subject's estimated cell source fraction, it is possible to determine that a particular treatment option is ineffective or will not be effective for the subject. For example, checkpoint immunotherapy will be ineffective if cytotoxic T cells are dysfunctional and undergo apoptosis. Such a situation is indicated, for example, when multiple fragments from a subject's biological sample are determined to be derived from cytotoxic T cells in the blood. In some embodiments, the estimated cell source fraction aids in monitoring minimal residual disease burden.

[0379] Those skilled in the art will recognize that any of the embodiments disclosed in the previous section (see, e.g., "Identifying Features for Estimating Cellular Source Fraction") can be applied in any combination to the methods and embodiments described herein for determining an estimated cell source fraction of a test subject. [Example]

[0380] Example 1 - Median increase in ctDNA rate by cancer stage

[0381] Referring to Figure 4, subjects are grouped by cancer stage I, II, III, and IV, regardless of the type of cancer they have. In Figure 4, the x-axis indicates which cancer stage each subject has, and the y-axis indicates the observed ctDNA fraction for each subject. The method used to calculate each subject's cfDNA fraction includes obtaining a first plurality of nucleic acid fragment sequences in electronic format from a biological sample of each subject in the cohort, the biological sample including cell-free nucleic acid molecules.

[0382] Figure 4 shows how ctDNA fraction varies by cancer stage, regardless of cancer type, among subjects with cell-free sequence reads indicative of underlying cancer. Thus, Figure 4 demonstrates that as disease becomes more severe, as determined by clinical staging (Stages 1–4), more evidence of cellular origin fraction (greater ctDNA fraction) is found in cfDNA. Figure 4 demonstrates that while this is generally the case across the CCGA cohort (see Example 3 for details on the CCGA cohort), violations of this trend (outliers) do exist. Such outliers in Figure 4 are suggestive and are best explained by clinical misclassification. Thus, Figure 4 illustrates the typical expected cellular origin fraction fraction in cfDNA, a fundamental component of underlying disease. Figure 4 also demonstrates that some individuals in Stage 4 exhibit very low shedding rates, suggesting the existence of distinct substates within Stage 4.

[0383] Figure 4 shows that the shedding rate (ctDNA rate) can be used as a basis for setting meaningful informative thresholds.

[0384] Example 2 - Obtaining Multiple Sequence Reads

[0385] 5 is a flowchart of a method 500 for preparing a nucleic acid sample for sequencing, according to one embodiment. Method 500 includes, but is not limited to, the following steps: For example, any step of method 500 may include a quantification substep for quality control or other laboratory assay procedures known to those skilled in the art.

[0386] In block 502, a nucleic acid sample (DNA or RNA) is extracted from a subject. The sample may be any subset of the human genome, including the entire genome. The sample may be extracted from a subject known to have cancer or suspected of having cancer. The sample may include blood, plasma, serum, urine, feces, saliva, other types of bodily fluids, or any combination thereof. In some embodiments, methods for obtaining a blood sample (e.g., syringe or finger prick) may be less invasive than procedures for obtaining a tissue biopsy, which may require surgery. The extracted sample may contain cfDNA and / or ctDNA. In healthy individuals, the human body can naturally remove cfDNA and other cellular debris. If the subject has cancer or disease, ctDNA in the extracted sample may be present at detectable levels for diagnostic purposes.

[0387] In block 504, a sequencing library is prepared. During library preparation, unique molecular identifiers (UMIs) are added to nucleic acid molecules (e.g., DNA molecules) through adapter ligation. UMIs are short nucleic acid sequences (e.g., 4-10 base pairs) that are added to the ends of DNA fragments during adapter ligation. In some embodiments, UMIs are degenerate base pairs that function as unique tags that can be used to identify sequence reads derived from specific DNA fragments. During PCR amplification following adapter ligation, the UMIs are replicated along with the added DNA fragments. This allows sequence reads derived from the same original fragment to be identified in downstream analysis.

[0388] At block 506, target DNA sequences are enriched from the library. During enrichment, hybridization probes (also referred to herein as "probes") are used to target and pull down nucleic acid fragments that are informative about the presence or absence of cancer (or disease), the state of the cancer, or the classification of the cancer (e.g., cancer class or tissue of origin). In a given workflow, probes may be designed to anneal (or hybridize) to a target (complementary) strand of DNA. The target strand may be the "positive" strand (e.g., the strand that is transcribed into mRNA and then translated into protein) or the complementary "negative" strand. Probe lengths can range from 10, 100, or 1000 base pairs. In one embodiment, probes are designed based on a methylation site panel. In one embodiment, probes are designed based on a panel of target genes to analyze specific mutations or target regions of the genome (e.g., of a human or other organism) suspected to correspond to a particular cancer or other type of disease. Additionally, probes may cover overlapping portions of the target regions. In block 408, these probes are used to perform a general sequence read of the nucleic acid sample.

[0389] FIG. 6 is a graphical diagram illustrating a process for obtaining sequence reads according to one embodiment. FIG. 6 depicts an example of a nucleic acid segment 800 from a sample. Here, the nucleic acid segment 600 can be a single-stranded nucleic acid segment, such as a single strand. In some embodiments, the nucleic acid segment 600 is a double-stranded cfDNA segment. In the illustrated example, three regions 605A, 605B, and 605C of the nucleic acid segment are depicted that can be targeted by different probes. Specifically, each of the three regions 605A, 605B, and 605C includes overlapping positions on the nucleic acid segment 600. An example of an overlapping position is depicted in FIG. 5 as a cytosine (“C”) nucleotide base 602. The cytosine nucleotide base 602 is located near a first end of region 605A, a center of region 605B, and a second end of region 605C.

[0390] In some embodiments, one or more (or all) of the probes are designed based on a gene panel or methylation site panel to analyze specific mutations or target regions of a genome (e.g., of a human or another organism) suspected of corresponding to a particular cancer or other type of disease. By using a targeted gene panel or methylation site panel rather than sequencing all expressed genes in the genome, also known as "whole exome sequencing," method 600 can be used to increase the sequencing depth of the target region, where depth refers to the count of the number of times a given target sequence in a sample is sequenced. Increasing sequencing depth can reduce the required input amount of nucleic acid sample.

[0391] Hybridization of nucleic acid sample 600 with one or more probes yields target sequence 670. As shown in FIG. 6, target sequence 670 is the nucleotide base sequence of region 605 targeted by the hybridization probe. Target sequence 670 can also be referred to as a hybridized nucleic acid fragment. For example, target sequence 670A corresponds to region 605A targeted by a first hybridization probe, target sequence 670B corresponds to region 605B targeted by a second hybridization probe, and target sequence 670C corresponds to region 605C targeted by a third hybridization probe. Given that cytosine nucleotide base 602 is located at a different position within each region 605A-C targeted by a hybridization probe, each target sequence 670 contains a nucleotide base corresponding to cytosine nucleotide base 602 at a specific position on target sequence 670.

[0392] After the hybridization step, the hybridized nucleic acid fragments can be captured and amplified using PCR. For example, target sequence 670 can be enriched to obtain enriched sequence 680, which can then be sequenced. In some embodiments, each enriched sequence 680 is replicated from target sequence 670. Enriched sequences 680A and 680C, amplified from target sequences 670A and 670C, respectively, also contain a thymine nucleotide base located near the end of each sequence read 680A or 680C. As used below, a variant nucleotide base (e.g., a thymine nucleotide base) in enriched sequence 680 that is mutated relative to a reference allele (e.g., a cytosine nucleotide base 602) is considered an alternative allele. Furthermore, each enriched sequence 680B amplified from target sequence 670B contains a cytosine nucleotide base located near or in the center of each enriched sequence 680B.

[0393] In block 508 of Figure 5, sequence reads are generated from the enriched DNA sequences, e.g., enriched sequence 680 shown in Figure 6. Sequencing data may be obtained from the enriched DNA sequences by means known in the art. For example, method 600 may include next-generation sequencing (NGS) technologies, including synthesis technology (Illumina), pyrosequencing (454 Life Sciences), ion semiconductor technology (Ion Torrent sequencing), single-molecule real-time sequencing (Pacific Biosciences), sequencing by ligation (SOLiD sequencing), nanopore sequencing (Oxford Nanopore Technologies), or paired-end sequencing. In some embodiments, massively parallel sequencing is performed using sequencing by synthesis using reversible dye terminators.

[0394] In some embodiments, the sequence reads may be aligned to a reference genome using methods known in the art to determine alignment position information. The alignment position information may indicate the start and end positions of a region in the reference genome corresponding to the start and end nucleotide bases of a given sequence read. The alignment position information may include the length of the sequence read, which can be determined from the start and end positions. The region in the reference genome may be associated with a gene or a segment of a gene.

[0395] JPEG0007775192000021.jpg50153

[0396] Example 3 - Cell-Free Genome Atlas Study (CCGA) Cohort

[0397] Subjects from CCGA [NCT02889978] were used in the examples of this disclosure. CCGA is a prospective, multicenter, observational cfDNA-based early cancer detection study that enrolled over 15,000 demographically balanced participants at over 140 sites.

[0398] This example looks at one of the CCGA substudies. Blood was collected from subjects with newly diagnosed, treatment-naïve cancer (C, cases) and participants without a cancer diagnosis (non-cancer [NC], controls), as defined at enrollment. This pre-planned substudy included 878 cases, 580 controls, and 169 assay controls (n=1627) across 20 tumor types and all clinical stages.

[0399] All samples were analyzed using the following methods: 1) paired targeted sequencing of cfDNA and white blood cells (WBCs) (60,000X, 507-gene panel) with a joint caller to remove WBC-derived somatic mutations and residual technical noise; 2) paired whole-genome sequencing (WGS, 35X) of cfDNA and WBCs with a novel machine learning algorithm to generate cancer-related signal scores; joint analysis to identify shared events; and 3) whole-genome bisulfite sequencing (WGBS, 34X) of cfDNA with aberrant methylation fragments to generate a normalized score. Targeted assays revealed that somatic variants (SNVs / indels) in cfDNA matched to non-tumor WBCs accounted for 76% of all variants in NCs and 65% in Cs. Consistent with somatic mosaicism (e.g., clonal hematopoiesis), WBC-matched variants increased with age; some were previously unreported non-canonical loss-of-function mutations. After WBC variant removal, canonical driver somatic variants were highly specific for Cs (e.g., NCs with variants had 0 Cs, whereas for EGFR and PIK3CA, Cs were 11 and 30, respectively). Similarly, of eight NCs with somatic copy number alterations (SCNAs) detected by WGS, four were derived from WBCs. CCGA WGBS data revealed informative high- and low-fragment CpGs (1:2 ratio); a subset of these was used to calculate methylation scores. Across all assays, consistent "cancer-like" signals (representing possible undiagnosed cancer) were observed in <1% of NC participants. An increasing trend was observed in NC vs. stages I-III vs. stage IV (nonsyn. SNVs / indels per Mb [Mean±SD] NC: 1.01±0.86, stages I-III: 2.43±3.98; stage IV: 6.45±6.79; WGS score NC: 0.00±0.08, stages I-III: 0.27±0.98; stage IV: 1.95±2.33; methylation score NC: 0±0.50; stages I-III: 1.02±1.77; stage IV: 3.94±1.70).These data demonstrate that it is possible to achieve >99% specificity for invasive cancer, supporting the promise of cfDNA assays for early cancer detection.

[0400] Example 4 - Examples of cell sources

[0401] In some embodiments, the cell source of any embodiment of the present disclosure is a first cancer condition of a common primary site of origin, hi some embodiments, the first cancer condition is breast cancer, lung cancer, prostate cancer, colorectal cancer, renal cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head / neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, gastric cancer, or a combination thereof.

[0402] In some embodiments, the cell source of any embodiment of the present disclosure is a tumor of a cancer type, or a portion thereof. In some embodiments, the tumor is adrenocortical carcinoma, pediatric adrenocortical carcinoma, tumors of AIDS-related cancer, Kaposi's sarcoma, tumors related to anal cancer, tumors related to appendix cancer, astrocytoma, pediatric (brain cancer) tumors, atypical teratoid / rhabdoid tumors, central nervous system (brain cancer) tumors, basal cell carcinoma of the skin, tumors related to bile duct cancer, bladder cancer tumors, pediatric bladder cancer tumors, bone cancer (e.g., Ewing's sarcoma, osteosarcoma, malignant fibrous histiocytoma) tissue, brain tumor, breast cancer tissue, pediatric breast cancer tissue, pediatric bronchial tumor, Burkitt's lymphoma tissue, carcinoid tumor (gastrointestinal), pediatric carcinoid tumor, cell type of unknown primary origin, cell type of unknown primary origin, pediatric cardiac (heart) tumor, tumor of the central nervous system (e.g., brain tumor such as pediatric atypical teratoid / rhabdoid), pediatric embryonal tumor, pediatric germ cell tumor, cervical cancer tissue, pediatric pediatric Cervical cancer tissue, bile duct cancer tissue, pediatric chordoma tissue, chronic myeloproliferative neoplasm, colorectal cancer tumor, pediatric colorectal cancer tumor, pediatric craniopharyngioma tissue, ductal carcinoma in situ (DCIS), pediatric embryonal tumor, endometrial cancer (uterine cancer) tissue, pediatric ependymoma tissue, esophageal cancer tissue, pediatric esophageal cancer tissue, esthesioneuroblastoma (head and neck cancer) tissue, pediatric extracranial germ cell tumor, extragonadal germ cell tumor, eye cancer tissue, intraocular melanoma, retinoblastoma, fallopian tube cancer tissue, gallbladder cancer tissue, gastric (stomach) cancer tissue, pediatric gastric (stomach) cancer tissue, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), pediatric gastrointestinal stromal tumor, germ cell tumor (e.g., pediatric central nervous system germ cell tumor, pediatric extracranial germ cell tumor, extragonadal germ cell tumor, ovarian germ cell tumor, or testicular cancer tissue), head and neck cancer tissue, pediatric heart tumor, hepatocellular carcinoma (HCC) tissue. Islet cell tumors (pancreatic neuroendocrine tumors), kidney or renal cell carcinoma (RCC) tissue, laryngeal cancer tissue, leukemia, liver cancer tissue, lung cancer (non-small cell and small cell) tissue, pediatric lung cancer tissue, and male breast cancer tissue.Malignant fibrous histiocytoma and osteosarcoma of bone, melanoma, pediatric melanoma, intraocular melanoma, pediatric intraocular melanoma, Merkel cell carcinoma, malignant mesothelioma, pediatric mesothelioma, metastatic cancer tissue, metastatic squamous cell neck cancer of unknown primary, midline cell type with NUT gene alterations, oral cancer (head and neck cancer) tissue, multiple endocrine neoplasia syndrome tissue, multiple myeloma / plasma cell neoplasm, myelodysplastic syndrome tissue, myelodysplastic / myeloproliferative neoplasm, chronic myeloproliferative neoplasm, nasal cavity and paranasal sinus cancer tissue , nasopharyngeal carcinoma (NPC) tissue, neuroblastoma tissue, non-small cell lung cancer tissue, oral cavity cancer tissue, lip and oral cavity cancer and oropharynx cancer tissue, osteosarcoma and bone malignant fibrous histiocytoma tissue, ovarian cancer tissue, pediatric ovarian cancer tissue, pancreatic cancer tissue, pediatric pancreatic cancer tissue, papillomatosis (pediatric larynx) tissue, paraganglioma tissue, pediatric paraganglioma tissue, paranasal sinus and nasal cavity cancer tissue, parathyroid cancer tissue, penile cancer tissue, pharyngeal cancer tissue, pheochromocytoma tissue, pediatric pheochromocytoma tissue, Pituitary tumor, plasma cell neoplasm / multiple myeloma, pleuropulmonary blastoma, primary central nervous system (CNS) lymphoma, primary peritoneal cancer tissue, prostate cancer tissue, rectal cancer tissue, retinoblastoma, pediatric rhabdomyosarcoma, salivary gland cancer tissue, sarcoma (e.g., pediatric hemangioma, osteosarcoma, uterine sarcoma, etc.), Sézary syndrome (lymphoma) tissue, skin cancer tissue, pediatric skin cancer tissue, small cell lung cancer tissue, small intestine cancer tissue, cutaneous squamous cell carcinoma, squamous cell cervical cancer of unknown primary, cutaneous T-cell lymphoma, Testicular cancer tissue, pediatric testicular cancer tissue, throat cancer (e.g., nasopharyngeal cancer, oropharyngeal cancer, hypopharyngeal cancer) tissue, thymoma or thymic carcinoma, thyroid cancer tissue, renal pelvis and ureter transitional cell cancer tissue, cell type of unknown primary tissue, ureter or renal pelvis tissue, transitional cell carcinoma (kidney (renal cell) cancer tissue), urethral cancer tissue, endometrial uterine cancer tissue, uterine sarcoma tissue, vaginal cancer tissue, pediatric vaginal cancer tissue, vascular tumor, vulvar cancer tissue, Wilms' tumor, or other pediatric kidney tumors.

[0403] In some embodiments, the cell source of any embodiment of the present disclosure is a first cancer state. In some such embodiments, the first cancer state is a stage of breast cancer, a stage of lung cancer, a stage of prostate cancer, a stage of colon cancer, a stage of renal cancer, a stage of uterine cancer, a stage of pancreatic cancer, a stage of esophageal cancer, a stage of lymphoma, a stage of head / neck cancer, a stage of ovarian cancer, a stage of hepatobiliary cancer, a stage of melanoma, a stage of cervical cancer, a stage of multiple myeloma, a stage of leukemia, a stage of thyroid cancer, a stage of bladder cancer, or a stage of gastric cancer.

[0404] In some embodiments, the cell source of any embodiment of the present disclosure is a predetermined stage of breast cancer, a predetermined stage of lung cancer, a predetermined stage of prostate cancer, a predetermined stage of colon cancer, a predetermined stage of renal cancer, a predetermined stage of uterine cancer, a predetermined stage of pancreatic cancer, a predetermined stage of esophageal cancer, a predetermined stage of lymphoma, a predetermined stage of head / neck cancer, a predetermined stage of ovarian cancer, a predetermined stage of hepatobiliary cancer, a predetermined stage of melanoma, a predetermined stage of cervical cancer, a predetermined stage of multiple myeloma, a predetermined stage of leukemia, a predetermined stage of thyroid cancer, a predetermined stage of bladder cancer, or a predetermined stage of gastric cancer.

[0405] In some embodiments, the cell source of any embodiment of the present disclosure is from a non-cancerous tissue. In some embodiments, the cell source of any embodiment of the present disclosure is from cells derived from a healthy tissue. In some embodiments, the cell source of any embodiment of the present disclosure is from a healthy tissue such as breast, lung, prostate, colon, kidney, uterus, pancreas, esophagus, lymph, ovary, cervix, epidermis, thyroid, bladder, stomach, or a combination thereof.

[0406] In some embodiments, the cell source of any embodiment of the present disclosure is derived from one tissue type. In some embodiments, the cell source of any embodiment of the present disclosure is derived from two or more tissue types. In some embodiments, a tissue type comprises one or more cell types (e.g., a combination of healthy, non-cancerous cells and cancerous cells). In some embodiments, a tissue type comprises one cell type (e.g., either cancerous cells or healthy, non-cancerous cells).

[0407] In some embodiments, the cell source of any embodiment of the present disclosure comprises one cell type, two cell types, three cell types, four cell types, five cell types, six cell types, seven cell types, eight cell types, nine cell types, ten cell types, or more than ten cell types.

[0408] In some embodiments, the cell source of any embodiment of the present disclosure is hepatocytes. In some such embodiments, the cell source is hepatocytes, hepatic stellate adipose tissue cells (ITO cells), Kupffer cells, sinusoidal endothelial cells, or any combination thereof.

[0409] In some embodiments, the source of cells of any embodiment of the present disclosure is gastric cells, hi some such embodiments, the source of cells is parietal cells.

[0410] In some embodiments, the source of cells of any embodiment of the present disclosure is one or more types of human cells. In some such embodiments, the source of cells is adaptive NK cells, adipocytes, alveolar cells, Alzheimer's type II astrocytes, amacrine cells, ameloblasts, astrocytes, B cells, basophils, basophil-activated cells, basophilic cells, Betz cells, bilayer ganglion cells, Böttcher cells, cardiomyocytes, CD4+ T cells, cementoblasts, cerebellar granule cells, cholangiocytes, gallbladder cells, chromaffin cells, Cigar cells, club cells, or the like. alveoli, orticotropic cells, cytotoxic T cells, dendritic cells, enterochromaffin cells, enterochromaffin-like cells, eosinophils, extraglomerular mesangial cells, bassoon cells, fat pad cells, gastric long cells, goblet cells, gonadotrophic cells, hepatic stellate cells, hepatocytes, hyperdifferentiated neutrophils, intraglomerular mesangial cells, juxtaglomerular cells, keratinocytes, renal proximal tubule brush border cells, Kupffer cells, mammary gland trophic cells, Leiden cells Wich cells, macrophages, macula densa cells, mast cells, megakaryocytes, melanocytes, fold cells, monocytes, natural killer cells, natural killer T cells, glitter cells, neutrophils, osteoblasts, osteoclasts, osteocytes, eosinophilic cells (parathyroid gland), Paneth cells, parafollicular cells, parasol cells, parathyroid long cells, parietal cells, small neurosecretory cells, PEG cells, pericytes, peritubular myoid cells, platelets, podocytes, regulatory T cells, reticulocytes The cells may be retinal bipolar cells, retinal horizontal cells, retinal ganglion cells, retinal progenitor cells, sentinel cells, Sertoli cells, somatomammotrophic cells, somatotropic cells, astrocytes, supporting cells, T cells, T helper cells, telocytes, tenocytes, thyroid stimulating cells, transitional B cells, hairy cells (human), tufted cells, unipolar brush cells, leukocytes, solid alveoli, or any combination thereof. In some such embodiments, the cells of the cell source are healthy. In alternative embodiments, the cells of the cell source are cancerous.

[0411] In some embodiments, the cell source of any embodiment of the present disclosure is any combination of cell types, provided that such cell types are derived from a single organ. In some such embodiments, the single organ is breast, lung, prostate, colon / rectum, kidney, uterus, pancreas, esophagus, blood, head / neck, ovary, liver, cervix, thyroid, bladder, or stomach. In some embodiments, the single organ is healthy. In alternative embodiments, the single organ is affected by a cancer originating from a single organ. In yet another embodiment, the single organ is affected by a cancer originating from an organ other than the single organ and metastasizing to the single organ.

[0412] In some embodiments, the cell source of any embodiment of the present disclosure is any combination of cell types, provided that such cell types are derived from a predetermined set of organs. In some such embodiments, the predetermined set of organs is any two organs from the set: breast, lung, prostate, colon / rectum, kidney, uterus, pancreas, esophagus, blood, head / neck, ovary, liver, cervix, thyroid, bladder, and stomach. In some embodiments, the predetermined set of organs is healthy. In alternative embodiments, the predetermined set of organs is affected by a cancer originating from one of the organs in the predetermined set of organs. In yet further alternative embodiments, the predetermined set of organs is affected by a cancer originating from an organ other than the predetermined set of organs and that has metastasized to the predetermined set of organs.

[0413] In some embodiments, the cell source of any embodiment of the present disclosure is any combination of cell types, provided that such cell types are derived from a predetermined set of organs. In some such embodiments, the predetermined set of organs is any three organs from the set: breast, lung, prostate, colon / rectum, kidney, uterus, pancreas, esophagus, blood, head / neck, ovary, liver, cervix, thyroid, bladder, and stomach. In some embodiments, the predetermined set of organs is healthy. In alternative embodiments, the predetermined set of organs is affected by a cancer originating from one of the organs in the predetermined set of organs. In yet another alternative embodiment, the predetermined set of organs is affected by a cancer originating from an organ other than the predetermined set of organs and metastasizing to the predetermined set of organs.

[0414] In some embodiments, the cell source of any embodiment of the present disclosure is any combination of cell types, provided that such cell types are derived from a predetermined set of organs. In some such embodiments, the predetermined set of organs is any four, five, six, or seven organs from the set of breast, lung, prostate, colon / rectum, kidney, uterus, pancreas, esophagus, blood, head / neck, ovary, liver, cervix, thyroid, bladder, and stomach. In some embodiments, the predetermined set of organs is healthy. In alternative embodiments, the predetermined set of organs is affected by a cancer originating from one of the organs in the predetermined set of organs. In yet another alternative embodiment, the predetermined set of organs is affected by a cancer originating from an organ other than the predetermined set of organs and metastasizing to the predetermined set of organs.

[0415] In some specific embodiments, the cell source of any embodiment of the present disclosure is a leukocyte, hi some such embodiments, the cell source is a neutrophil, eosinophil, basophil, lymphocyte, B lymphocyte, T lymphocyte, cytotoxic T cell, monocyte, or any combination thereof.

[0416] conclusion

[0417] Multiple examples may be provided for a component, operation, or structure described herein as a single example. Finally, boundaries between various components, operations, and data stores are somewhat arbitrary, and particular operations are illustrated in the context of specific exemplary configurations. Other allocations of functionality are contemplated and may be within the scope of the implementations. In general, structures and functions shown as separate components in exemplary configurations may be implemented as a combined structure or component. Similarly, structures and functions presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements are within the scope of the implementations.

[0418] Furthermore, although terms such as "first," "second," etc. may be used herein to describe various elements, it will be understood that these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first object can be referred to as a second object, and similarly, a second object can be referred to as a first object, without departing from the scope of the present disclosure. The first object and the second object are both objects, but are not the same object.

[0419] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in the description of the invention and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. The term "and / or," as used herein, will also be understood to refer to and encompass any and all possible combinations of one or more of the associated listed items. It will be further understood that, as used herein, the terms "comprises" and / or "comprising" specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0420] As used herein, the term "if" may be interpreted to mean "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if determined" or "if [a stated condition or event] is detected" may be interpreted to mean "upon determining" or "in response to determining" or "upon detecting (a stated condition or event)" or "in response to detecting (a stated condition or event)," depending on the context.

[0421] The foregoing description has included example systems, methods, techniques, instruction sequences, and computer program products embodying example implementations. For purposes of explanation, numerous specific details have been set forth in order to provide an understanding of various implementations of the inventive subject matter. However, it will be apparent to those skilled in the art that implementations of the inventive subject matter may be practiced without these specific details. Generally, well-known example instructions, protocols, structures, and techniques have not been shown in detail.

[0422] The foregoing description has been set forth with reference to specific embodiments for purposes of explanation. However, the exemplary discussion above is not intended to be exhaustive or to limit the implementations to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The implementations were chosen and described to best explain the principles and their practical application, thereby enabling those skilled in the art to best utilize the implementations, and various modifications thereof, as suited to the particular uses contemplated.

Claims

1. 1. A method for identifying multiple features for estimating a cell origin fraction of a subject, comprising: A computer system having one or more processors and a memory storing one or more programs for execution by said one or more processors, A) obtaining a training dataset in electronic form, the training dataset comprising, for each training subject of a plurality of training subjects: a) a corresponding methylation pattern for each cell-free fragment among a plurality of corresponding training cell-free fragments, wherein the corresponding methylation pattern for each cell-free fragment (i) is determined by methylation sequencing of one or more nucleic acid samples comprising each fragment in a corresponding biological sample obtained from each of the training subjects, and (ii) the methylation pattern includes the methylation status of each CpG site among a plurality of corresponding CpG sites in each of the fragments; b) a subject cancer indication for each said training subject, wherein said subject cancer condition is one of a first cancer condition and a second cancer condition; That and; B) mapping each cell-free fragment of each of the plurality of cell-free fragments to a sequence group of a plurality of sequence groups, thereby obtaining a plurality of training sets of cell-free fragments, wherein each sequence group of the plurality of sequence groups represents a corresponding portion of the human reference genome, and each training set of cell-free fragments is mapped to a different sequence group of the plurality of sequence groups; C) determining, for each cell-free fragment in each of the training sets of cell-free fragments in the plurality of training sets of cell-free fragments, a cancer state of the cell-free fragment using a classifier trained to receive as input a methylation pattern of each cell-free fragment in each of the training sets of cell-free fragments in the plurality of training sets of cell-free fragments and to generate as output a cancer state of the cell-free fragment, wherein the cancer state of the cell-free fragment is one of the first cancer state and the second cancer state; D) assigning the determined cell-free fragment cancer state to each cell-free fragment in each training set of cell-free fragments among the plurality of training sets of cell-free fragments; E) for each sequence group of the plurality of sequence groups, determining a corresponding measure of association I between (a) a subject cancer state for each training subject of the plurality of training subjects and (b) a cell-free fragment cancer state for each cell-free fragment of a corresponding cell-free fragment training set mapped to each sequence group, wherein the corresponding measure of association I is a correlation calculation, a mutual information calculation, or a distance metric; F) identifying a plurality of features for estimating the cell origin fraction of the subject as a subset of the plurality of sequences, wherein each sequence in the subset of the plurality of sequences satisfies a selection criterion based on a corresponding measure of relatedness for each sequence, the selection criterion specifying the selection of sequences in the plurality of sequences that have the measure of relatedness that is top-ranked compared to other sequences.

2. 10. The method of claim 1, wherein the classifier is based on a formula: [Equation 1] wherein (Outside 1) 【number】 is the first model for the first cancer state, "Fragment" is the methylation pattern of each cell-free fragment, (Outside 2) 【number】 is a second model for the second cancer state, The cancer state of the cell-free fragment of each of said fragments is assigned a first cancer state when R(fragments) meets a threshold, optionally said threshold is 1 to 10, optionally said threshold is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10.

3. 2. The method of claim 1, wherein the measure of relevance I is a calculation of mutual information; [Equation 2] It is calculated as follows: i and j are independent indices into the set {first cancer state, second cancer state}; x i is the number of training subjects having cancer status i among the plurality of training subjects; y j is the number of training subjects among the plurality of training subjects that have one or more cell-free fragments mapped to each sequence group assigned to cancer status j; p(x i, y j )teeth (Outer 3) 【number】 and N(x i, y j ) is the number of training subjects among the plurality of training subjects that have cancer status i and have one or more cell-free fragments mapped to each sequence group assigned cancer status j; N T is the number of training subjects in the plurality of training subjects, p(x i ) is x i / N T and p(y j ) is y j / N T That's the method.

4. 4. The method according to claim 1, wherein (i) the plurality of sequences consists of 1000 sequences to 100,000 sequences, optionally 15,000 sequences to 80,000 sequences; and / or (ii) a method wherein each sequence group of said plurality of sequences has, on average, 10 to 1200 residues, and optionally, has, on average, 10 to 10000 residues.

5. 5. The method of any one of claims 1 to 4, wherein the selection criteria specify the selection of sequences having one of the top N relatedness measures, where N is a positive integer greater than or equal to 50, optionally N is between 500 and 5000, and optionally N is between 800 and 1500.

6. 6. The method of any one of claims 1 to 5, wherein the first cancer condition is cancer and the second cancer condition is the absence of cancer, and optionally: the first cancer condition is one of adrenal gland cancer, biliary tract cancer, bladder cancer, bone / bone marrow cancer, brain cancer, breast cancer, cervical cancer, colon cancer, esophageal cancer, gastric cancer, head / neck cancer, hepatobiliary cancer, renal cancer, liver cancer, lung cancer, ovarian cancer, pancreatic cancer, pelvic cancer, pleural cancer, prostate cancer, kidney cancer, skin cancer, stomach cancer, testicular cancer, thymic cancer, thyroid cancer, uterine cancer, lymphoma, melanoma, multiple myeloma, or leukemia, and the second cancer condition is the absence of cancer; or a stage of adrenal cancer, a stage of biliary tract cancer, a stage of bladder cancer, a stage of bone / bone marrow cancer, a stage of brain cancer, a stage of breast cancer, a stage of cervical cancer, a stage of colorectal cancer, a stage of esophageal cancer, a stage of stomach cancer, a stage of head / neck cancer, a stage of hepatobiliary cancer, a stage of kidney cancer, a stage of liver cancer, a stage of lung cancer, a stage of ovarian cancer, a stage of pancreatic cancer, a stage of pelvic cancer, a stage of pleural cancer, a stage of prostate cancer, a stage of renal cancer, a stage of skin cancer, a stage of stomach cancer, a stage of testicular cancer, a stage of thymus cancer, a stage of thyroid cancer, a stage of uterine cancer, a stage of lymphoma, a stage of melanoma, a stage of multiple myeloma, or a stage of leukemia; and a second cancer condition is the absence of cancer.

7. 2. The method of claim 1, wherein the methylation sequencing is whole genome methylation sequencing.

8. 2. The method of claim 1, wherein the methylation sequencing is targeted sequencing using a plurality of nucleic acid probes, and each sequence group of the plurality of sequence groups is associated with at least one nucleic acid probe of the plurality of nucleic acid probes.

9. 9. The method of claim 8, wherein the plurality of nucleic acid probes comprises 1,000 or more nucleic acid probes, 2,000 or more nucleic acid probes, 3,000 or more nucleic acid probes, 5,000 or more nucleic acid probes, 10,000 or more nucleic acid probes, or 1,000 to 30,000 nucleic acid probes.

10. 10. The method of claim 1, wherein each sequence group of the plurality of sequences contains 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more CpG sites.

11. 9. The method of claim 1, wherein each of the plurality of sequences comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more consecutive CpG sites.

12. 12. The method according to any one of claims 1 to 11, The methylation state of each CpG site among the corresponding plurality of CpG sites in each of the fragments is: a methylation status when each of the CpG sites is determined to be methylated by the methylation sequencing; When each of the CpG sites is determined to be unmethylated by the methylation sequencing, the CpG site is in an unmethylated state; If the methylation status of each CpG site cannot be called as methylated or unmethylated by the methylation sequencing, it is flagged as "other"; Optionally, (i) detecting one or more 5-methylcytosines (5mC) and / or 5-hydroxymethylcytosines (5hmC) in each of the fragments by methylation sequencing; or (ii) the methylation sequencing comprises converting one or more unmethylated cytosines or one or more methylated cytosines in the sequence reads of each of the fragments to one or more corresponding uracils, optionally wherein the one or more uracils are detected as one or more corresponding thymines during the methylation sequencing, or wherein the conversion of one or more unmethylated cytosines or one or more methylated cytosines comprises chemical conversion, enzymatic conversion, or a combination thereof.

13. 3. The method of claim 2, the first model is a first mixed model including a first plurality of sub-models; the second model is a second mixed model including a second plurality of sub-models; each sub-model of the first and second pluralities of sub-models represents an independent corresponding methylation model for a source of cell-free fragments in the corresponding biological sample; Optionally, (i) each of the independent corresponding methylation models is one of a binomial model, a beta-binomial model, an independent site model, or a Markov model; or (ii) two or more submodels in the first plurality of submodels are independent region models, and two or more submodels in the second plurality of submodels are independent region models.

14. 14. The method according to any one of claims 1 to 13, B) applying one or more filter conditions to the plurality of cell-free fragments prior to mapping B), and optionally (a) one of the filter conditions of the one or more filter conditions is applying a p-value threshold to a corresponding methylation pattern of each cell-free fragment among the plurality of cell-free fragments, the p-value threshold being representative of the frequency with which the methylation pattern is observed in a cohort of non-cancer subjects; Optionally, (i) the p-value threshold is between 0.001 and 0.20; or (ii) the cohort comprises at least 20 subjects and the plurality of cell-free fragments comprises at least 10,000 different corresponding methylation patterns; or (iii) the p-value threshold is met for a methylation pattern from the subject if the corresponding methylation pattern for each cell-free fragment in the plurality of cell-free fragments has a p-value of 0.10 or less, 0.05 or less, or 0.01 or less; or (b) one of the one or more filter conditions is applying a requirement that each cell-free fragment of the plurality of cell-free fragments be represented by a threshold number of sequence reads in a corresponding plurality of sequence reads measured from one or more nucleic acid samples containing each fragment in the corresponding biological sample, wherein the threshold number is 2, 3, 4, 5, 6, 7, 8, 9, 10, or an integer between 10 and 100; or (c) one of the filter conditions of the one or more filter conditions is applying a requirement that each cell-free fragment of the plurality of cell-free fragments is represented by a threshold number of cell-free nucleic acids in one or more nucleic acid samples containing the respective fragment in the corresponding biological sample, optionally the threshold number being 2, 3, 4, 5, 6, 7, 8, 9, 10, or an integer between 10 and 100; or (d) one of the filter conditions of the one or more filter conditions is applying a requirement that each cell-free fragment of the plurality of cell-free fragments has a threshold number of CpG sites, optionally the threshold number of CpG sites is at least 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 CpG sites; or (e) one of the one or more filter conditions requires that each cell-free fragment of the plurality of cell-free fragments has a length of less than a threshold number of base pairs, and optionally the threshold number of base pairs is 1,000, 2,000, 3,000, or 4,000 contiguous base pairs in length.

15. 1. A non-transitory computer-readable storage medium having stored thereon program code instructions that, when executed by a processor, cause the processor to perform a method for identifying a plurality of features for estimating a cell origin fraction of a subject, the method comprising: A) obtaining a training dataset in electronic form, the training dataset comprising, for each training subject of a plurality of training subjects: a) a corresponding methylation pattern for each cell-free fragment among a plurality of corresponding training cell-free fragments, wherein the corresponding methylation pattern for each cell-free fragment (i) is determined by methylation sequencing of one or more nucleic acid samples comprising each fragment in a corresponding biological sample obtained from each of the training subjects, and (ii) the methylation pattern includes the methylation status of each CpG site among a plurality of corresponding CpG sites in each of the fragments; b) a subject cancer indication for each said training subject, wherein said subject cancer condition is one of a first cancer condition and a second cancer condition; That and; B) mapping each cell-free fragment of each of the plurality of cell-free fragments to a sequence group of a plurality of sequence groups, thereby obtaining a plurality of training sets of cell-free fragments, wherein each sequence group of the plurality of sequence groups represents a corresponding portion of the human reference genome, and each training set of cell-free fragments is mapped to a different sequence group of the plurality of sequence groups; C) determining, for each cell-free fragment in each of the training sets of cell-free fragments in the plurality of training sets of cell-free fragments, a cancer state of the cell-free fragment using a classifier trained to receive as input a methylation pattern of each cell-free fragment in each of the training sets of cell-free fragments in the plurality of training sets of cell-free fragments and to generate as output a cancer state of the cell-free fragment, wherein the cancer state of the cell-free fragment is one of the first cancer state and the second cancer state; D) assigning the determined cell-free fragment cancer state to each cell-free fragment in each training set of cell-free fragments among the plurality of training sets of cell-free fragments; E) for each sequence group of the plurality of sequence groups, determining a corresponding measure of association I between (a) a subject cancer state for each training subject of the plurality of training subjects and (b) a cell-free fragment cancer state for each cell-free fragment of a corresponding cell-free fragment training set mapped to each sequence group, wherein the corresponding measure of association I is a correlation calculation, a mutual information calculation, or a distance metric; F) identifying a plurality of features for estimating the cell origin proportion of the subject as a subset of the plurality of sequences, wherein each sequence group in the subset of the plurality of sequences satisfies a selection criterion based on a corresponding measure of relatedness for each sequence group, the selection criterion specifying the selection of a sequence group in the plurality of sequences that has a top-ranked measure of relatedness compared to other sequence groups.

Citation Information

Patent Citations

  • Cell-free DNA methylation patterns for the analysis of diseases and conditions

    JP2019521673A

  • Methylation pattern analysis of tissues in a DNA mixture

    US20160017419A1

  • Cell-free DNA methylation patterns for disease and condition analysis

    WO2017212428A1