Multi-step screening and analysis

JP2025505395A5Pending Publication Date: 2026-02-06FLAGSHIP PIONEERING INNOVATIONS VI LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024544448
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-08-29
Filing Date
2023-01-30
Publication Date
2026-02-06

AI Technical Summary

Benefits of technology

【0093】 本発明のこれらの及び他の特徴、態様、及び利点は、以下の説明及び添付の図面に関してよりよく理解されるであろう。可能な場合は、図面に類似又は同様の参照番号が使用されることがあり、類似又は同様の機能を示し得ることに注意されたい。例えば、「第三者機関155A」などの参照番号の後の文字は、文がその特定の参照番号を有する要素を具体的に指すことを示す。後に文字が付かない文中の参照番号(「第三者機関155」など)は、その参照番号を有する図中の要素のいずれか又は全てを指す(例えば、文中の「第三者機関155」は、図中の参照番号「第三者機関155A」及び/又は「第三者機関155B」を指す)。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2023147567000001
    Figure 2023147567000001
Patent Text Reader

Abstract

Disclosed herein are methods, non-transitory computer readable media, systems, and kits for performing a multi-stage analysis to identify individuals with a health condition for monitoring, treatment, and / or enrollment in a clinical trial. Specifically, the multi-stage analysis includes a first screening in which a majority of individuals identified as not at risk for the health condition are excluded, and a subsequent second analysis that detects the presence of the health condition in the remaining individuals. Overall, the multi-stage analysis achieves improved performance and accurate identification of individuals with the health condition.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 304,536, filed January 28, 2022, U.S. Provisional Patent Application No. 63 / 312,741, filed February 22, 2022, and U.S. Provisional Patent Application No. 17 / 898,154, filed August 29, 2022, the entire disclosures of each of which are incorporated herein by reference in their entirety for all purposes. [Background technology]

[0002] Diagnostic technologies include simple point-of-care (POC) tests applied to large populations to identify relatively common diseases, and complex focused tests applied to selected populations. However, while POC tests can be applied to large populations, they cannot diagnose rare health conditions in individuals with a high enough accuracy to be feasible. Similarly, while complex focused tests can be deployed to test rare populations, such tests are often invasive, expensive, and fail when applied to detect rare health conditions in large patient populations. For example, complex focused tests perform poorly (e.g., have a high number of false positives and / or a low positive predictive value) when attempting to diagnose rare health conditions in large patient populations. Summary of the Invention [Means for solving the problem]

[0003] Disclosed herein is a method comprising multi-step analysis for identifying individuals with a certain health condition.In particular, the method disclosed herein comprising multi-step analysis is useful for identifying individuals with rare health conditions from a large population (e.g., millions of people).Multi-step analysis includes a first screening, which excludes a large proportion of individuals who are identified as not at risk for the health condition.

[0004] Disclosed herein is a step-by-step, multi-part method for detecting one or more early stage cancers in a subject, the method including performing an analysis of the subject's sequence information obtained from the subject's biological sample to identify whether the subject is at risk for having one or more early stage cancers; and, if the patient is not identified as at risk, analyzing the sequence information of the subject not identified as at risk by performing a second analysis to detect the presence of at least one particular cancer in the subject.

[0005] Further disclosed herein is a stepwise, multi-part method for detecting a candidate population of subjects having an early stage cancer among a plurality of subjects, the method comprising: for each of one or more subjects in the plurality of subjects, obtaining sequence information obtained from a first assay performed on a sample obtained from the subject; performing an analysis of the sequence information of the subject to identify whether the subject is not at risk for one or more early stage cancers; obtaining sequence information obtained from a second assay performed on the sample or an additional sample obtained from the subject to generate sequence information obtained from the second assay in response to the subject not being identified as not at risk for one or more early stage cancers; and performing an analysis of the sequence information obtained from the second assay for the subject to determine whether the subject is included in the candidate population. In various embodiments, less than 10% of the plurality of subjects are not identified as not at risk for one or more early stage cancers, and wherein performing the analysis of the sequence information obtained from the second assay is performed on less than 10% of the plurality of subjects. In various embodiments, less than 5% of the plurality of subjects are not identified as not at risk for one or more early stage cancers, and wherein performing the analysis of the sequence information obtained from the second assay is performed on less than 5% of the plurality of subjects. In various embodiments, a multi-part, stepwise method that analyzes multiple subjects across two or more stage tests achieves improved performance as a function of resource consumption compared to a single stage method. In various embodiments, a multi-part, stepwise method that analyzes multiple subjects across two or more stage tests achieves improved performance metrics compared to a single stage method. In various embodiments, a multi-part, stepwise method that analyzes multiple subjects across two or more stage tests achieves similar or reduced performance metrics compared to a single stage method.

[0006] In various embodiments, the one or more early stage cancers are 15 or more different cancers. In various embodiments, the one or more early stage cancers or preclinical stage cancers are acute lymphocytic leukemia, acute myeloid leukemia, adrenocortical carcinoma, soft tissue sarcoma, lymphoma, anal cancer, gastrointestinal cancer, brain cancer, skin cancer, bile duct cancer, bladder cancer, bone cancer, breast cancer, lung cancer, heart cancer, central nervous system cancer, cervical cancer, chronic lymphocytic leukemia, chronic myelogenous leukemia, chronic myeloproliferative neoplasms, colon cancer, uterine cancer, esophageal cancer, head and neck cancer, eye cancer, fallopian tube cancer, gallbladder cancer, gastric cancer, germ cell tumors, gestational trophoblastic cancer, leukemia ... cancer, hairy cell leukemia, liver cancer, Hodgkin's lymphoma, intraocular melanoma, pancreatic cancer, kidney cancer, leukemia, mesothelioma, metastatic cancer, oral cancer, multiple endocrine neoplasia syndrome, multiple myeloma, myelodysplastic tumor, ovarian cancer, parathyroid cancer, penile cancer, pheochromocytoma, pituitary cancer, plasma cell neoplasm, primary peritoneal cancer, prostate cancer, rectal cancer, retinoblastoma, sarcoma, small intestine cancer, testicular cancer, pharyngeal cancer, thymoma and thymic cancer, thyroid cancer, urethral cancer, uterine cancer, vaginal cancer, and vulvar cancer.

[0007] In various embodiments, one or more of the early stage or preclinical cancers is a single cancer type. In various embodiments, the single cancer type is acute lymphocytic leukemia, acute myeloid leukemia, adrenocortical carcinoma, soft tissue sarcoma, lymphoma, anal cancer, gastrointestinal cancer, brain cancer, skin cancer, bile duct cancer, bladder cancer, bone cancer, breast cancer, lung cancer, heart cancer, central nervous system cancer, cervical cancer, chronic lymphocytic leukemia, chronic myelogenous leukemia, chronic myeloproliferative neoplasms, colon cancer, uterine cancer, esophageal cancer, head and neck cancer, eye cancer, fallopian tube cancer, gallbladder cancer, gastric cancer, germ cell tumors, gestational trophoblastic carcinoma, hairy cell leukemia, ... The cancer is any one of the following: blood cancer, liver cancer, Hodgkin's lymphoma, intraocular melanoma, pancreatic cancer, kidney cancer, leukemia, mesothelioma, metastatic cancer, oral cancer, multiple endocrine neoplasia syndrome, multiple myeloma, myelodysplastic tumor, ovarian cancer, parathyroid cancer, penile cancer, pheochromocytoma, pituitary cancer, plasma cell neoplasm, primary peritoneal cancer, prostate cancer, rectal cancer, retinoblastoma, sarcoma, small intestine cancer, testicular cancer, pharyngeal cancer, thymoma and thymic cancer, thyroid cancer, urethral cancer, uterine cancer, vaginal cancer, and vulvar cancer.

[0008] In various embodiments, the early stage cancer is a preclinical cancer. In various embodiments, the preclinical cancer is a stage I or stage II cancer. In various embodiments, the method has a greater than 70% ability to detect at least one multiple early stage cancer. In various embodiments, the method has a greater than 70% ability to detect at least one multiple early stage cancer with a specificity of greater than 95%, greater than 96%, greater than 97%, greater than 98%, greater than 99%, greater than 99.5%, or greater than 99.9%. In various embodiments, the method achieves a positive predictive value of at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, or at least 85% in detecting at least one multiple early stage cancer. In various embodiments, the method achieves a negative predictive value of at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.3%, or at least 99.4% in detecting at least one multiple early stage cancer.

[0009] In various embodiments, performing an analysis of the subject's sequence information to identify whether the subject is not at risk has a negative predictive value of at least 90%, at least 95%, or at least 99%. In various embodiments, analyzing the subject's sequence information to identify whether the subject has detectable cancer or precancer has a positive predictive value of at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, or at least 85%. In various embodiments, analyzing the subject's sequence information to identify whether the subject has detectable cancer or precancer has a negative predictive value of at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, or at least 97%.

[0010] In various embodiments, the sequence information comprises methylation sequence information. In various embodiments, the methylation sequence information comprises a methylation state of a plurality of genomic sites. In various embodiments, the plurality of genomic sites comprises a plurality of CpG sites. In various embodiments, performing an analysis of the sequence information of the subject comprises applying a trained machine learning model.

[0011] In various embodiments, performing an analysis of the sequence information of the subject further comprises: calculating a window- and target region-specific metric for one or more instances of the analyte within a window of a plurality of windows on the target region of the analyte; and analyzing at least the window- and target region-specific metric using the trained machine learning model. In various embodiments, the window- and target region-specific metric comprises a ratio of a number of DNA fragments having a particular number of methylated CpGs to a number of DNA fragments in the window of the target region. In various embodiments, the window- and target region-specific metric comprises a ratio of a number of DNA fragments having a particular pattern of methylation to a number of DNA fragments in the window of the target region. In various embodiments, calculating the window- and target region-specific metric comprises performing a first function that quantifies the occurrence of methylated CpGs within the window of the target region. In various embodiments, calculating the window- and target region-specific metric further comprises performing a second function that normalizes the occurrence of methylated CpGs to the number of DNA fragments in the window of the target region. In various embodiments, the window comprises 1 to 100 CpG sites. In various embodiments, the window and target region specific metric includes an input vector that includes the proportion of DNA fragments that have a particular number of methylated CpGs among all possible CpG methylation patterns. In various embodiments, all possible methylation patterns are 2 k are the possible patterns of k, where k refers to the number of CpG sites in the window.

[0012] In various embodiments, the sequence information is obtained from an assay, where the assay includes performing one or more of: a. sequencing of nucleic acids in a sample; b. hybrid capture; c. methylation-specific PCR; d. an assay that generates methylation information; and e. sequencing of a clone library generated from a template-immortalized library.

[0013] In various embodiments, performing an assay to generate sequence information comprises obtaining bisulfite converted cell-free DNA (cfDNA); selectively amplifying a target region of the bisulfite converted cfDNA; and sequencing an amplification product comprising the amplified target region to generate methylation information. In various embodiments, the target region of the bisulfite converted cfDNA comprises a pre-identified region that is differentially methylated in cancer.

[0014] In various embodiments, the target region of the bisulfite converted cfDNA comprises one or more CpG islands or portions of one or more CpG islands set forth in Tables 1-4. In various embodiments, the target region of the bisulfite converted cfDNA comprises up to 10%, up to 20%, up to 30%, up to 40%, up to 50%, up to 55%, up to 60%, up to 65%, up to 70%, up to 75%, up to 80%, up to 85%, or up to 90% of the CpG islands or portions of CpG islands set forth in any one of Tables 1-4. In various embodiments, the target region of the bisulfite converted cfDNA comprises 100, up to 150, up to 200, up to 300, up to 400, up to 500, up to 600, up to 700, up to 800, up to 900, up to 1000, up to 1500, up to 2000, up to 2500, up to 3000, up to 3500, or up to 4000 CpG islands or portions of CpG islands selected from Tables 1-4. In various embodiments, analyzing sequence information of subjects not identified as not at risk comprises analyzing sequence information generated from a target region that comprises one or more CpG islands or portions of one or more CpG islands set forth in Tables 1-4. In various embodiments, the target region comprises at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of a CpG island or portion of a CpG island shown in any one of Tables 1-4.In various embodiments, the target region comprises at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1500, at least 2000, at least 2500, at least 3000, at least 3500, at least 4000, at least 4500, at least 5000, at least 5500, or at least 6000 CpG islands or portions of CpG islands selected from Tables 1-4. In various embodiments, performing the second analysis comprises analyzing the methylation status of a greater number of CpG islands compared to the amount of CpG islands analyzed when performing the analysis of the sequence information of the subject. In various embodiments, performing the second analysis comprises analyzing the methylation status of at least 5 times as many CpG islands compared to the amount of CpG islands analyzed when performing the analysis of the sequence information of the subject. In various embodiments, one or more of the CpG islands analyzed when performing the analysis of the sequence information of the subject represent a subset of the CpG islands analyzed when performing the second analysis. In various embodiments, all of the CpG islands analyzed when performing the analysis of the sequence information of the subject are further analyzed when performing the second analysis. In various embodiments, performing the second analysis comprises analyzing the methylation status of at least 500 CpG islands, and performing the analysis of the sequence information of the subject comprises analyzing the methylation status of at least 100 CpG islands.

[0015] In various embodiments, the target region of the bisulfite converted cfDNA comprises one or more CpG islands or portions of CpG islands as set forth in Tables 1-4. In various embodiments, the biological sample is obtained from the subject while the subject is asymptomatic. In various embodiments, the biological sample comprises any one of a blood sample, a stool sample, a urine sample, a mucus sample, and a saliva sample.

[0016] In various embodiments, the biological sample is a blood sample. In various embodiments, the biological sample does not include an invasive biopsy sample. In various embodiments, the assay performed on the biological sample processes one or more of: nucleic acid; cell-free DNA that includes selected CpGs with selected methylation status; and RNA. In various embodiments, the second analysis includes whole genome sequencing, optionally whole genome bisulfite sequencing.

[0017] In various embodiments, the method disclosed herein further comprises using the sequence information of the subject to determine the tissue of origin of at least one specific cancer in the subject. In various embodiments, the method disclosed herein further comprises: performing an analysis of additional sequence information of the subject obtained from additional biological samples of the subject obtained after the time the biological sample was obtained; determining one or more changes between the additional sequence information of the subject and the sequence information; and determining the progression of at least one specific cancer in the subject based on the determined one or more changes. In various embodiments, the method disclosed herein further comprises determining whether to provide an intervention to the subject based on the determined progression of at least one specific cancer. In various embodiments, determining one or more changes between the additional sequence information of the subject and the sequence information comprises determining one or more changes in methylation status across multiple genomic sites.

[0018] Further disclosed herein is a stepwise, multi-part method for detecting a health condition in a subject, comprising: performing an analysis of the subject's sequence information obtained from the subject's biological sample to identify whether the subject is at risk for having a health condition; and if the patient is not identified as at risk, analyzing the sequence information of the subject not identified as at risk by performing a second analysis to detect the presence of the health condition in the subject. Further disclosed herein is a stepwise, multi-part method for detecting a health condition in a subject, comprising: performing an analysis of the subject's marker information obtained from the subject's biological sample to identify whether the subject is at risk for having a health condition; and if the patient is not identified as at risk, analyzing the sequence information of the subject not identified as at risk by performing a second analysis to detect the presence of the health condition in the subject. In various embodiments, the marker information comprises a quantitative level of a protein biomarker.

[0019] Further disclosed herein is a step-wise, multi-part method for increasing the likelihood that a signal in a sample is authentic, comprising: (a) performing an analysis of sequence information of nucleic acids in the sample to determine whether the analysis produces a result that correlates with the presence or absence of a human condition; and, if a result is detected, (b) analyzing the sequence information of nucleic acids in the sample by performing a second analysis to determine whether the second analysis produces a signal, where if a signal is detected, the likelihood that the signal in the sample is authentic is increased compared to the likelihood that the signal is authentic if produced by a similar method, the similar method differing in that step (a) is omitted.

[0020] In various embodiments, the method achieves a positive predictive value of at least 20% in detecting at least one multiple early stage cancer. In various embodiments, the method achieves a positive predictive value of at least 40% in detecting at least one multiple early stage cancer. In various embodiments, the method achieves a positive predictive value of at least 60% in detecting at least one multiple early stage cancer. In various embodiments, the method achieves a positive predictive value of at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, or at least 85% in detecting a health condition. In various embodiments, the health condition is a disease risk. In various embodiments, the health condition is a rare disease or disorder. In various embodiments, the health condition has an incidence of 1 in 100, 1 in 1,000, 1 in 10,000, 1 in 100,000, 1 in 1,000,000, 1 in 10,000,000, or 1 in 100,000,000.

[0021] Further disclosed herein is a method of diagnosing a subject as having at least one of multiple early stage cancers, the method comprising obtaining sequence information from a first assay performed on a sample obtained from the subject; performing a screening by analyzing the sequence information to classify the subject as at risk for one or more multiple early stage cancers or not at risk for one or more multiple early stage cancers; responsive to the subject being classified as at risk for one or more multiple early stage cancers, obtaining sequence information from a second assay performed on the sample or an additional sample obtained from the subject to generate second assay obtained sequence information; and performing a diagnostic analysis of the sequence information obtained from the second assay on the subject to further classify the subject as at risk for one or more multiple early stage cancers as a candidate for monitoring or treatment. Further disclosed herein is a method of diagnosing a subject as being at risk for at least one of multiple early cancers, comprising obtaining sequence information from a first assay performed on a sample obtained from the subject; performing a screening by analyzing the sequence information to classify the subject as being at risk for one or more multiple early cancers or not at risk for one or more multiple early cancers; if the subject is classified as not at risk for one or more multiple early cancers, reporting that the subject is not at risk for one or more multiple early cancers; if the subject is classified as being at risk for one or more multiple early cancers, obtaining sequence information from a second assay performed on the sample or an additional sample obtained from the subject to generate second assay obtained sequence information; and performing a diagnostic analysis of the sequence information obtained from the second assay on the subject to further classify the subject as being at risk for one or more multiple early cancers as a candidate subject for monitoring.

[0022] Further disclosed herein is a method for identifying a candidate population of subjects with early cancer for enrollment in a clinical trial, the method comprising: obtaining sequence information from a first assay performed on a sample obtained from the subject for each of one or more subjects in a plurality of subjects; performing a screening by analyzing the sequence information to classify the subject as at risk for one or more multiple early cancers or not at risk for one or more multiple early cancers; obtaining sequence information from a second assay performed on the sample or an additional sample obtained from the subject in response to the subject being classified as at risk for one or more multiple early cancers to generate sequence information from a second assay; and performing a diagnostic analysis of the sequence information from the second assay on the subject to further classify the subject at risk for one or more multiple early cancers as a candidate subject for inclusion in the candidate population. In various embodiments, less than 10% of the plurality of subjects are identified as not at risk for one or more early cancers, where performing the analysis of the sequence information from the second assay is performed on less than 10% of the plurality of subjects. In various embodiments, less than 5% of the plurality of subjects are identified as not at risk for one or more early stage cancers, and wherein performing the analysis of the sequence information obtained from the second assay is performed on less than 5% of the plurality of subjects. In various embodiments, a stepwise, multi-part method analyzing a plurality of subjects across two or more stages of testing achieves improved performance as a function of resource consumption compared to a single stage method. In various embodiments, a stepwise, multi-part method analyzing a plurality of subjects across two or more stages of testing achieves improved performance metrics compared to a single stage method. In various embodiments, a stepwise, multi-part method analyzing a plurality of subjects across two or more stages of testing achieves similar or reduced performance metrics compared to a single stage method.

[0023] In various embodiments, the sequence information obtained from the first assay comprises methylation sequence information. In various embodiments, the methylation sequence information obtained from the first assay comprises a methylation state of a plurality of genomic sites. In various embodiments, the plurality of genomic sites comprises a plurality of CpG sites. In various embodiments, performing the screening by analyzing the sequence information obtained from the first assay comprises applying the trained machine learning model. In various embodiments, the sequence information obtained from the second assay comprises methylation sequence information. In various embodiments, the methylation sequence information from the second assay comprises a methylation state of a plurality of genomic sites identified as associated with the subject. In various embodiments, the plurality of genomic sites comprises a plurality of CpG sites. In various embodiments, performing the diagnostic analysis of the sequence information obtained from the second assay comprises applying the trained machine learning model.

[0024] In various embodiments, performing an analysis of the sequence information of the subject further comprises: calculating a window- and target region-specific metric for one or more instances of the analyte within a window of a plurality of windows on the target region of the analyte; and analyzing at least the window- and target region-specific metric using the trained machine learning model. In various embodiments, the window- and target region-specific metric comprises a ratio of a number of DNA fragments having a particular number of methylated CpGs to a number of DNA fragments in the window of the target region. In various embodiments, the window- and target region-specific metric comprises a ratio of a number of DNA fragments having a particular pattern of methylation to a number of DNA fragments in the window of the target region. In various embodiments, calculating the window- and target region-specific metric comprises performing a first function that quantifies the occurrence of methylated CpGs within the window of the target region. In various embodiments, calculating the window- and target region-specific metric further comprises performing a second function that normalizes the occurrence of methylated CpGs to the number of DNA fragments in the window of the target region. In various embodiments, the window comprises 1 to 100 CpG sites. In various embodiments, the window and target region specific metric comprises an input vector that comprises the proportion of DNA fragments with a particular number of methylated CpGs among all possible CpG methylation patterns. In various embodiments, all possible methylation patterns are determined by a factor of 2 k are the possible patterns of k, where k refers to the number of CpG sites in the window.

[0025] In various embodiments, obtaining sequence information obtained from a first assay comprises performing or having performed a first assay to generate sequence information obtained from a first assay. In various embodiments, performing or having performed a first assay comprises performing or having performed one or more of: a. sequencing nucleic acids in a sample; b. hybrid capture; c. methylation specific PCR; d. an assay to generate methylation information; and e. sequencing a clone library generated from a template immortalized library.

[0026] In various embodiments, performing an assay to generate sequence information comprises obtaining bisulfite converted cell-free DNA (cfDNA); selectively amplifying a target region of the bisulfite converted cfDNA; and sequencing an amplification product comprising the amplified target region to generate methylation information. In various embodiments, the target region of the bisulfite converted cfDNA comprises a pre-identified region that is differentially methylated in cancer.

[0027] In various embodiments, the target region of the bisulfite converted cfDNA comprises up to 10%, up to 20%, up to 30%, up to 40%, up to 50%, up to 55%, up to 60%, up to 65%, up to 70%, up to 75%, up to 80%, up to 85%, or up to 90% of a CpG island or a portion of a CpG island set forth in any one of Tables 1-4. In various embodiments, the target region of the bisulfite converted cfDNA comprises 100, up to 150, up to 200, up to 300, up to 400, up to 500, up to 600, up to 700, up to 800, up to 900, up to 1000, up to 1500, up to 2000, up to 2500, up to 3000, up to 3500, or up to 4000 CpG islands or portions of CpG islands selected from Tables 1-4. In various embodiments, performing a diagnostic analysis of the sequence information obtained from the second assay comprises analyzing sequence information generated from a target region that includes one or more CpG islands or portions of one or more CpG islands set forth in Tables 1-4. In various embodiments, the target region includes at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the CpG islands or portions of CpG islands set forth in any one of Tables 1-4. In various embodiments, the target region comprises at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1500, at least 2000, at least 2500, at least 3000, at least 3500, at least 4000, at least 4500, at least 5000, at least 5500, or at least 6000 CpG islands or portions of CpG islands selected from Tables 1-4.In various embodiments, performing a diagnostic analysis of the sequence information obtained from the second assay comprises analyzing the methylation status of a greater number of CpG islands compared to the amount of CpG islands analyzed when performing a screening of the subject. In various embodiments, performing a diagnostic analysis of the sequence information obtained from the second assay comprises analyzing the methylation status of at least 5 times more CpG islands compared to the amount of CpG islands analyzed when performing a screening. In various embodiments, one or more of the CpG islands analyzed when performing a screening represent a subset of the CpG islands analyzed when performing a diagnostic analysis. In various embodiments, all of the CpG islands analyzed when performing a screening are further analyzed when performing a diagnostic analysis. In various embodiments, performing a diagnostic analysis comprises analyzing the methylation status of at least 500 CpG islands and performing a screening comprises analyzing the methylation status of at least 100 CpG islands.

[0028] In various embodiments, the target region of the bisulfite converted cfDNA comprises one or more CpG islands or portions of one or more CpG islands as set forth in Tables 1-4. In various embodiments, the sample or additional sample is obtained from the subject while the subject is asymptomatic. In various embodiments, the sample or additional sample comprises any one of a blood sample, a stool sample, a urine sample, a mucus sample, and a saliva sample. In various embodiments, the sample or additional sample is a blood sample. In various embodiments, the first assay performed on the sample or the second assay performed on the sample or additional sample processes one or more of: nucleic acid; cell-free DNA comprising the selected CpG having the selected methylation state; and RNA.

[0029] In various embodiments, the cost of the second assay is greater than the cost of the first assay. In various embodiments, the second assay comprises whole genome sequencing. In various embodiments, the whole genome sequencing comprises whole genome bisulfite sequencing. In various embodiments, the diagnostic analysis achieves higher sensitivity with higher specificity compared to screening. In various embodiments, the method disclosed herein further comprises using the sequence information of the subject to determine the tissue of origin of at least one specific cancer in the subject. In various embodiments, the method disclosed herein further comprises: performing an analysis of additional sequence information of the subject obtained from additional biological samples of the subject obtained after the time the biological sample was obtained; determining one or more changes between the additional sequence information of the subject and the sequence information; and determining the progression of at least one specific cancer in the subject based on the determined one or more changes. In various embodiments, the method disclosed herein further comprises determining whether to provide an intervention to the subject based on the determined progression of at least one specific cancer.

[0030] In various embodiments, determining one or more changes between the additional sequence information and the sequence information of the subject comprises determining one or more changes in methylation status across multiple genomic sites. In various embodiments, the method disclosed herein further comprises: for each of one or more other subjects in the plurality of subjects, obtaining sequence information obtained from a first assay performed on a sample obtained from the subject; performing a screening by analyzing the sequence information to classify the subject as at risk for one or more multiple early cancers or at no risk for one or more multiple early cancers; and in response to the subject being classified as at no risk for one or more multiple early cancers, reporting the subject as at no risk for one or more multiple early cancers and removing the subject from the candidate population. In various embodiments, the method disclosed herein further comprises obtaining sequence information obtained from a third assay performed on an additional sample obtained from the subject; and performing a diagnostic analysis of the sequence information obtained from the third assay on the subject to further classify the subject.

[0031] In various embodiments, the obtained sequence information obtained from the third assay includes methylation sequence information. In various embodiments, the methylation sequence information includes the methylation state of multiple sites that are individually informative about the subject. In various embodiments, the additional sample is obtained at a time different from the time at which either the sample or the additional sample was obtained. In various embodiments, the one or more multiple early stage cancers are 15 or more different cancers. In various embodiments, the one or more multiple early stage cancers are acute lymphocytic leukemia, acute myeloid leukemia, adrenocortical carcinoma, soft tissue sarcoma, lymphoma, anal cancer, gastrointestinal cancer, brain cancer, skin cancer, bile duct cancer, bladder cancer, bone cancer, breast cancer, lung cancer, heart cancer, central nervous system cancer, cervical cancer, chronic lymphocytic leukemia, chronic myelogenous leukemia, chronic myeloproliferative neoplasms, colon cancer, uterine cancer, esophageal cancer, head and neck cancer, eye cancer, fallopian tube cancer, gallbladder cancer, gastric cancer, germ cell tumors, gestational trophoblastic carcinoma, myeloma, ovarian cancer ... The set includes hairy cell leukemia, liver cancer, Hodgkin's lymphoma, intraocular melanoma, pancreatic cancer, kidney cancer, leukemia, mesothelioma, metastatic cancer, oral cancer, multiple endocrine neoplasia syndrome, multiple myeloma, myelodysplastic tumor, ovarian cancer, parathyroid cancer, penile cancer, pheochromocytoma, pituitary cancer, plasma cell neoplasm, primary peritoneal cancer, prostate cancer, rectal cancer, retinoblastoma, sarcoma, small intestine cancer, testicular cancer, pharyngeal cancer, thymoma and thymic cancer, thyroid cancer, urethral cancer, uterine cancer, vaginal cancer, and vulvar cancer.

[0032] In various embodiments, the one or more multiple early cancers are of a single cancer type. In various embodiments, the single cancer type is acute lymphocytic leukemia, acute myeloid leukemia, adrenocortical carcinoma, soft tissue sarcoma, lymphoma, anal cancer, gastrointestinal cancer, brain cancer, skin cancer, bile duct cancer, bladder cancer, bone cancer, breast cancer, lung cancer, heart cancer, central nervous system cancer, cervical cancer, chronic lymphocytic leukemia, chronic myelogenous leukemia, chronic myeloproliferative neoplasms, colon cancer, uterine cancer, esophageal cancer, head and neck cancer, eye cancer, fallopian tube cancer, gallbladder cancer, gastric cancer, germ cell tumors, gestational trophoblastic carcinoma, hairy cell leukemia, ... The cancer is any one of blood cancer, liver cancer, Hodgkin's lymphoma, intraocular melanoma, pancreatic cancer, kidney cancer, leukemia, mesothelioma, metastatic cancer, oral cancer, multiple endocrine neoplasia syndrome, multiple myeloma, myelodysplastic tumor, ovarian cancer, parathyroid cancer, penile cancer, pheochromocytoma, pituitary cancer, plasma cell neoplasm, primary peritoneal cancer, prostate cancer, rectal cancer, retinoblastoma, sarcoma, small intestine cancer, testicular cancer, pharyngeal cancer, thymoma and thymic cancer, thyroid cancer, urethral cancer, uterine cancer, vaginal cancer, and vulvar cancer. In various embodiments, the early stage cancer is a cancer in a preclinical stage.

[0033] In various embodiments, the preclinical cancer is a stage I or stage II cancer. In various embodiments, the method has a greater than 70% ability to detect at least one multiple early stage cancer with a specificity of greater than 95%, greater than 96%, greater than 97%, greater than 98%, greater than 99%, greater than 99.5%, or greater than 99.9%. In various embodiments, the method achieves a positive predictive value of at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, or at least 85% in detecting at least one multiple early stage cancer. In various embodiments, the method achieves a negative predictive value of at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.3%, or at least 99.4% in detecting at least one multiple early stage cancer. In various embodiments, the screening has a negative predictive value of at least 90%, at least 95%, or at least 99%. In various embodiments, the diagnostic assay has a positive predictive value of at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, or at least 85%. In various embodiments, the diagnostic assay has a negative predictive value of at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, or at least 97%.

[0034] Further disclosed herein is a non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to perform an analysis of the subject's sequence information obtained from the subject's biological sample to identify whether the subject is at risk of having one or more early stage cancers; and if the patient is not identified as at risk, analyze the sequence information of the subject not identified as at risk by performing a second analysis to detect the presence of at least one specific cancer in the subject. In various embodiments, the one or more early stage cancers are 15 or more different cancers. In various embodiments, the one or more early stage or preclinical cancers are selected from the group consisting of acute lymphocytic leukemia, acute myeloid leukemia, adrenal cortical carcinoma, soft tissue sarcoma, lymphoma, anal cancer, gastrointestinal cancer, brain cancer, skin cancer, bile duct cancer, bladder cancer, bone cancer, breast cancer, lung cancer, heart cancer, central nervous system cancer, cervical cancer, chronic lymphocytic leukemia, chronic myelogenous leukemia, chronic myeloproliferative neoplasms, colon cancer, uterine cancer, esophageal cancer, head and neck cancer, eye cancer, fallopian tube cancer, gallbladder cancer, gastric cancer, germ cell tumors, gestational trophoblastic cancer, and / or ovarian cancer. cancer, hairy cell leukemia, liver cancer, Hodgkin's lymphoma, intraocular melanoma, pancreatic cancer, kidney cancer, leukemia, mesothelioma, metastatic cancer, oral cancer, multiple endocrine neoplasia syndrome, multiple myeloma, myelodysplastic tumor, ovarian cancer, parathyroid cancer, penile cancer, pheochromocytoma, pituitary cancer, plasma cell neoplasm, primary peritoneal cancer, prostate cancer, rectal cancer, retinoblastoma, sarcoma, small intestine cancer, testicular cancer, pharyngeal cancer, thymoma and thymic cancer, thyroid cancer, urethral cancer, uterine cancer, vaginal cancer, and vulvar cancer.

[0035] In various embodiments, one or more of the early stage or preclinical cancers is a single cancer type. In various embodiments, the single cancer type is acute lymphocytic leukemia, acute myeloid leukemia, adrenocortical carcinoma, soft tissue sarcoma, lymphoma, anal cancer, gastrointestinal cancer, brain cancer, skin cancer, bile duct cancer, bladder cancer, bone cancer, breast cancer, lung cancer, heart cancer, central nervous system cancer, cervical cancer, chronic lymphocytic leukemia, chronic myelogenous leukemia, chronic myeloproliferative neoplasms, colon cancer, uterine cancer, esophageal cancer, head and neck cancer, eye cancer, fallopian tube cancer, gallbladder cancer, gastric cancer, germ cell tumors, gestational trophoblastic carcinoma, hairy cell leukemia, ... The cancer is any one of hematology, liver cancer, Hodgkin's lymphoma, intraocular melanoma, pancreatic cancer, renal cancer, leukemia, mesothelioma, metastatic cancer, oral cancer, multiple endocrine neoplasia syndrome, multiple myeloma, myelodysplastic tumor, ovarian cancer, parathyroid cancer, penile cancer, pheochromocytoma, pituitary cancer, plasma cell neoplasm, primary peritoneal cancer, prostate cancer, rectal cancer, retinoblastoma, sarcoma, small intestine cancer, testicular cancer, pharyngeal cancer, thymoma and thymic cancer, thyroid cancer, urethral cancer, uterine cancer, vaginal cancer, and vulvar cancer. In various embodiments, the early cancer is a cancer in a preclinical stage. In various embodiments, the preclinical cancer is a cancer in stage I or stage II.

[0036] In various embodiments, the performance of the analysis and the analysis of the sequence information has a greater than 70% ability to detect at least one multiple early stage cancer. In various embodiments, the performance of the analysis and the analysis of the sequence information has a greater than 70% ability to detect at least one multiple early stage cancer with a specificity of greater than 95%, greater than 96%, greater than 97%, greater than 98%, greater than 99%, greater than 99.5%, or greater than 99.9%. In various embodiments, the performance of the analysis and the analysis of the sequence information achieves a positive predictive value of at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, or at least 85% in detecting at least one multiple early stage cancer. In various embodiments, the performance of the analysis and the analysis of the sequence information achieves a negative predictive value of at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.3%, or at least 99.4% in detecting at least one multiple early stage cancer. In various embodiments, the performance of the analysis has a negative predictive value of at least 90%, at least 95%, or at least 99%. In various embodiments, the analysis of the sequence information to identify whether a subject has detectable cancer or precancer has a positive predictive value of at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, or at least 85%. In various embodiments, the analysis of the sequence information to identify whether a subject has detectable cancer or precancer has a negative predictive value of at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, or at least 97%.

[0037] In various embodiments, the sequence information comprises methylation sequence information. In various embodiments, the methylation sequence information comprises a methylation state of a plurality of genomic sites. In various embodiments, the plurality of genomic sites comprises a plurality of CpG sites. In various embodiments, the instructions that cause the processor to perform an analysis of the sequence information of the subject further comprise instructions that, when executed by the processor, cause the processor to apply a trained machine learning model.

[0038] In various embodiments, the instructions for causing the processor to perform an analysis of the sequence information of the subject further include instructions that, when executed by the processor, cause the processor to calculate window and target region specific metrics for one or more instances of the analyte within a window of a plurality of windows on the target region of the analyte; and to analyze at least the window and target region specific metrics using the trained machine learning model. In various embodiments, the window and target region specific metrics include a ratio of the number of DNA fragments with a particular number of methylated CpGs to the number of DNA fragments in the window of the target region. In various embodiments, the window and target region specific metrics include a ratio of the number of DNA fragments with a particular pattern of methylation to the number of DNA fragments in the window of the target region. In various embodiments, the instructions for causing the processor to calculate the window and target region specific metrics, when executed by the processor, further include instructions that cause the processor to perform a first function of quantifying the occurrence of methylated CpGs within the window of the target region. In various embodiments, the instructions that cause the processor to calculate the window and target region specific metric further include instructions that, when executed by the processor, cause the processor to perform a second function that normalizes the occurrence of methylated CpGs to the number of DNA fragments in the window of the target region. In various embodiments, the window includes between 1 and 100 CpG sites. In various embodiments, the window and target region specific metric includes an input vector that includes a proportion of DNA fragments that have a particular number of methylated CpGs among all possible CpG methylation patterns. In various embodiments, all possible methylation patterns are normalized to a proportion of DNA fragments that have a particular number of methylated CpGs among all possible CpG methylation patterns. k are the possible patterns of k, where k refers to the number of CpG sites in the window.

[0039] In various embodiments, the sequence information is obtained from an assay, where the assay includes performing one or more of: a. sequencing the nucleic acid in the sample; b. hybrid capture; c. methylation-specific PCR; d. an assay that generates methylation information; and e. sequencing a clone library generated from a template-immortalized library. In various embodiments, performing an assay that generates sequence information includes obtaining bisulfite converted cell-free DNA (cfDNA); selectively amplifying a target region of the bisulfite converted cfDNA; and sequencing the amplification product that includes the amplified target region to generate the methylation information.

[0040] In various embodiments, the target region of the bisulfite converted cfDNA comprises a pre-identified region that is differentially methylated in cancer. In various embodiments, the target region of the bisulfite converted cfDNA comprises one or more CpG islands or portions of one or more CpG islands shown in Tables 1-4. In various embodiments, the target region of the bisulfite converted cfDNA comprises up to 10%, up to 20%, up to 30%, up to 40%, up to 50%, up to 55%, up to 60%, up to 65%, up to 70%, up to 75%, up to 80%, up to 85%, or up to 90% of the CpG islands or portions of CpG islands shown in any one of Tables 1-4. In various embodiments, the target region of the bisulfite converted cfDNA comprises 100, up to 150, up to 200, up to 300, up to 400, up to 500, up to 600, up to 700, up to 800, up to 900, up to 1000, up to 1500, up to 2000, up to 2500, up to 3000, up to 3500, or up to 4000 CpG islands or portions of CpG islands selected from Tables 1-4. In various embodiments, analyzing sequence information of subjects not identified as not at risk comprises analyzing sequence information generated from a target region that comprises one or more CpG islands or portions of one or more CpG islands set forth in Tables 1-4. In various embodiments, the target region comprises at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of a CpG island or portion of a CpG island shown in any one of Tables 1-4.In various embodiments, the target region comprises at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1500, at least 2000, at least 2500, at least 3000, at least 3500, at least 4000, at least 4500, at least 5000, at least 5500, or at least 6000 CpG islands or portions of CpG islands selected from Tables 1-4. In various embodiments, performing the second analysis comprises analyzing the methylation status of a greater number of CpG islands compared to the amount of CpG islands analyzed when performing the analysis of the sequence information of the subject. In various embodiments, performing the second analysis comprises analyzing the methylation status of at least 5 times as many CpG islands compared to the amount of CpG islands analyzed when performing the analysis of the sequence information of the subject. In various embodiments, one or more of the CpG islands analyzed when performing the analysis of the sequence information of the subject represent a subset of the CpG islands analyzed when performing the second analysis. In various embodiments, all of the CpG islands analyzed when performing the analysis of the sequence information of the subject are further analyzed when performing the second analysis. In various embodiments, performing the second analysis comprises analyzing the methylation status of at least 500 CpG islands, and performing the analysis of the sequence information of the subject comprises analyzing the methylation status of at least 100 CpG islands.

[0041] In various embodiments, the biological sample is obtained from the subject while the subject is asymptomatic. In various embodiments, the biological sample includes any one of a blood sample, a stool sample, a urine sample, a mucus sample, and a saliva sample. In various embodiments, the biological sample is a blood sample. In various embodiments, the biological sample does not include an invasive biopsy sample. In various embodiments, the assay performed on the biological sample processes one or more of: nucleic acid; cell-free DNA that includes a selected CpG having a selected methylation state; and RNA.

[0042] In various embodiments, the second analysis comprises whole genome sequencing, optionally whole genome bisulfite sequencing. In various embodiments, the non-transitory computer-readable medium, when executed by the processor, further comprises instructions that cause the processor to use the sequence information of the subject to determine the tissue of origin of at least one specific cancer in the subject. In various embodiments, the non-transitory computer-readable medium, when executed by the processor, further comprises instructions that cause the processor to perform an analysis of additional sequence information of the subject obtained from additional biological samples of the subject obtained after the biological sample was obtained; determine one or more changes between the additional sequence information of the subject and the sequence information; and determine the progression of at least one specific cancer in the subject based on the determined one or more changes. In various embodiments, the non-transitory computer-readable medium, when executed by the processor, further comprises instructions that cause the processor to determine whether to provide an intervention to the subject based on the determined progression of at least one specific cancer. In various embodiments, the instructions that cause the processor to determine one or more changes between the additional sequence information of the subject and the sequence information further include instructions that, when executed by the processor, cause the processor to determine one or more changes in methylation status across a plurality of genomic sites.

[0043] Further disclosed herein is a non-transitory computer readable medium comprising instructions, when executed by a processor, that cause the processor to perform an analysis of the subject's sequence information obtained from the subject's biological sample to identify whether the subject is at risk for having a health condition; and if the patient is not identified as at risk, analyze the sequence information of the subject not identified as at risk by performing a second analysis to detect the presence of the health condition in the subject. Further disclosed herein is a non-transitory computer readable medium comprising instructions, when executed by a processor, that cause the processor to perform an analysis of the subject's marker information obtained from the subject's biological sample to identify whether the subject is at risk for having a health condition; and if the patient is not identified as at risk, analyze the sequence information of the subject not identified as at risk by performing a second analysis to detect the presence of the health condition in the subject. In various embodiments, the marker information comprises a quantitative level of a protein biomarker.

[0044] Further disclosed herein is a non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to: (a) perform an analysis of sequence information of nucleic acids in the sample to determine whether the analysis produces a result that correlates with the presence or absence of a human condition; and if a result is detected, (b) analyze the sequence information of nucleic acids in the sample by performing a second analysis to determine whether the second analysis produces a signal, where if a signal is detected, then the likelihood that the signal in the sample is genuine is increased compared to the likelihood that the signal would be genuine if produced by a similar method, the similar method differing in that step (a) is omitted.

[0045] In various embodiments, the steps performed by the processor achieve a positive predictive value of at least 20% in detecting at least one multiple early stage cancer. In various embodiments, the steps performed by the processor achieve a positive predictive value of at least 40% in detecting at least one multiple early stage cancer. In various embodiments, the steps performed by the processor achieve a positive predictive value of at least 60% in detecting at least one multiple early stage cancer. In various embodiments, the steps performed by the processor achieve a positive predictive value of at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, or at least 85% in detecting the health condition. In various embodiments, the health condition is a disease risk. In various embodiments, the health condition is a rare disease or disorder. In various embodiments, the health condition has an incidence of 1 in 100, 1 in 1,000, 1 in 10,000, 1 in 100,000, 1 in 1,000,000, 1 in 10,000,000, or 1 in 100,000,000.

[0046] Further disclosed herein is a non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to: obtain sequence information from a first assay performed on a sample obtained from the subject; perform a screening by analyzing the sequence information to classify the subject as at risk or not at risk for a health condition; in response to the subject being classified as at risk for a health condition, obtain sequence information from a second assay performed on the sample or an additional sample obtained from the subject to generate second assay-derived sequence information; and perform a diagnostic analysis of the sequence information obtained from the second assay on the subject to further classify the subject at risk for the health condition as a candidate for monitoring. Further disclosed herein is a non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to: obtain sequence information obtained from a first assay performed on a sample obtained from the subject; perform a screening by analyzing the sequence information to classify the subject as at risk for a health condition or not at risk for a health condition; if the subject is classified as not at risk for the health condition, report that the subject is not at risk for the health condition; if the subject is classified as at risk for the health condition, obtain sequence information obtained from a second assay performed on the sample or an additional sample obtained from the subject to generate second assay obtained sequence information; and perform a diagnostic analysis of the sequence information obtained from the second assay on the subject to further classify the subject at risk for the health condition as a candidate for monitoring.

[0047] Further disclosed herein is a non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to: obtain, for each of one or more subjects in the plurality of subjects, sequence information obtained from a first assay performed on a sample obtained from the subject; perform a screening by analyzing the sequence information to classify the subject as at risk for a health condition or not at risk for a health condition; responsive to the subject being classified as at risk for a health condition, obtain sequence information obtained from a second assay performed on the sample or an additional sample obtained from the subject to generate second assay-derived sequence information; and perform a diagnostic analysis of the sequence information obtained from the second assay for the subjects to further classify the subjects at risk for the health condition as candidate subjects for inclusion in a candidate population.

[0048] Further disclosed herein is a system comprising a non-transitory computer readable medium, the non-transitory computer readable medium comprising: a processor; a data storage device comprising sequence information obtained from the subject's biological sample; instructions that, when executed by the processor, cause the processor to perform an analysis of the subject's sequence information obtained from the subject's biological sample to identify whether the subject is at risk for having one or more early stage cancers; and, if the patient is not identified as at risk, analyze the sequence information of the subject not identified as at risk by performing a second analysis to detect the presence of at least one specific cancer in the subject. In various embodiments, the one or more early stage cancers are 15 or more different cancers. In various embodiments, the one or more early stage or preclinical cancers are selected from the group consisting of acute lymphocytic leukemia, acute myeloid leukemia, adrenal cortical carcinoma, soft tissue sarcoma, lymphoma, anal cancer, gastrointestinal cancer, brain cancer, skin cancer, bile duct cancer, bladder cancer, bone cancer, breast cancer, lung cancer, heart cancer, central nervous system cancer, cervical cancer, chronic lymphocytic leukemia, chronic myelogenous leukemia, chronic myeloproliferative neoplasms, colon cancer, uterine cancer, esophageal cancer, head and neck cancer, eye cancer, fallopian tube cancer, gallbladder cancer, gastric cancer, germ cell tumors, gestational trophoblastic cancer, and / or ovarian cancer. cancer, hairy cell leukemia, liver cancer, Hodgkin's lymphoma, intraocular melanoma, pancreatic cancer, kidney cancer, leukemia, mesothelioma, metastatic cancer, oral cancer, multiple endocrine neoplasia syndrome, multiple myeloma, myelodysplastic tumor, ovarian cancer, parathyroid cancer, penile cancer, pheochromocytoma, pituitary cancer, plasma cell neoplasm, primary peritoneal cancer, prostate cancer, rectal cancer, retinoblastoma, sarcoma, small intestine cancer, testicular cancer, pharyngeal cancer, thymoma and thymic cancer, thyroid cancer, urethral cancer, uterine cancer, vaginal cancer, and vulvar cancer.

[0049] In various embodiments, one or more of the early stage or preclinical cancers is a single cancer type. In various embodiments, the single cancer type is acute lymphocytic leukemia, acute myeloid leukemia, adrenocortical carcinoma, soft tissue sarcoma, lymphoma, anal cancer, gastrointestinal cancer, brain cancer, skin cancer, bile duct cancer, bladder cancer, bone cancer, breast cancer, lung cancer, heart cancer, central nervous system cancer, cervical cancer, chronic lymphocytic leukemia, chronic myelogenous leukemia, chronic myeloproliferative neoplasms, colon cancer, uterine cancer, esophageal cancer, head and neck cancer, eye cancer, fallopian tube cancer, gallbladder cancer, gastric cancer, germ cell tumors, gestational trophoblastic carcinoma, hairy cell leukemia, ... The cancer is any one of the following: blood cancer, liver cancer, Hodgkin's lymphoma, intraocular melanoma, pancreatic cancer, kidney cancer, leukemia, mesothelioma, metastatic cancer, oral cancer, multiple endocrine neoplasia syndrome, multiple myeloma, myelodysplastic tumor, ovarian cancer, parathyroid cancer, penile cancer, pheochromocytoma, pituitary cancer, plasma cell neoplasm, primary peritoneal cancer, prostate cancer, rectal cancer, retinoblastoma, sarcoma, small intestine cancer, testicular cancer, pharyngeal cancer, thymoma and thymic cancer, thyroid cancer, urethral cancer, uterine cancer, vaginal cancer, and vulvar cancer.

[0050] In various embodiments, the early stage cancer is a preclinical cancer. In various embodiments, the preclinical cancer is a stage I or stage II cancer. In various embodiments, the performance of the analysis and the analysis of the sequence information has a greater than 70% ability to detect at least one multiple early stage cancer. In various embodiments, the performance of the analysis and the analysis of the sequence information has a greater than 70% ability to detect at least one multiple early stage cancer with a specificity of greater than 95%, greater than 96%, greater than 97%, greater than 98%, greater than 99%, greater than 99.5%, or greater than 99.9%. In various embodiments, the performance of the analysis and the analysis of the sequence information achieves a positive predictive value of at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, or at least 85% in detecting at least one multiple early stage cancer. In various embodiments, the performance of the analysis and the analysis of the sequence information achieves a negative predictive value of at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.3%, or at least 99.4% in detecting at least one multiple early stage cancer. In various embodiments, the performance of the analysis has a negative predictive value of at least 90%, at least 95%, or at least 99%. In various embodiments, the analysis of the sequence information to identify whether a subject has a detectable cancer or precancer has a positive predictive value of at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, or at least 85%. In various embodiments, the analysis of the sequence information to identify whether a subject has a detectable cancer or precancer has a negative predictive value of at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, or at least 97%.

[0051] In various embodiments, the sequence information comprises methylation sequence information. In various embodiments, the methylation sequence information comprises a methylation state of a plurality of genomic sites. In various embodiments, the plurality of genomic sites comprises a plurality of CpG sites. In various embodiments, the instructions that cause the processor to perform an analysis of the sequence information of the subject further comprise instructions that, when executed by the processor, cause the processor to apply a trained machine learning model.

[0052] In various embodiments, the instructions for causing the processor to perform an analysis of the sequence information of the subject further include instructions that, when executed by the processor, cause the processor to calculate window and target region specific metrics for one or more instances of the analyte within a window of a plurality of windows on the target region of the analyte; and to analyze at least the window and target region specific metrics using the trained machine learning model. In various embodiments, the window and target region specific metrics include a ratio of the number of DNA fragments with a particular number of methylated CpGs to the number of DNA fragments in the window of the target region. In various embodiments, the window and target region specific metrics include a ratio of the number of DNA fragments with a particular pattern of methylation to the number of DNA fragments in the window of the target region. In various embodiments, the instructions for causing the processor to calculate the window and target region specific metrics, when executed by the processor, further include instructions that cause the processor to perform a first function of quantifying the occurrence of methylated CpGs within the window of the target region. In various embodiments, the instructions causing the processor to calculate the window and target region specific metric further include instructions that, when executed by the processor, cause the processor to perform a second function that normalizes the occurrence of methylated CpGs to the number of DNA fragments in the window of the target region. In various embodiments, the window includes between 1 and 100 CpG sites. In various embodiments, the window and target region specific metric includes an input vector that includes a proportion of DNA fragments that have a particular number of methylated CpGs among all possible CpG methylation patterns. In various embodiments, all possible methylation patterns are normalized to a proportion of DNA fragments that have a particular number of methylated CpGs among all possible CpG methylation patterns. k are the possible patterns of k, where k refers to the number of CpG sites in the window.

[0053] In various embodiments, the sequence information is obtained from an assay, where the assay includes performing one or more of: a. sequencing nucleic acids in a sample; b. hybrid capture; c. methylation-specific PCR; d. an assay that generates methylation information; and e. sequencing a clone library generated from a template-immortalized library.

[0054] In various embodiments, performing the assay to generate sequence information includes obtaining bisulfite converted cell-free DNA (cfDNA); selectively amplifying a target region of the bisulfite converted cfDNA; and sequencing an amplification product that includes the amplified target region to generate methylation information. In various embodiments, the target region of the bisulfite converted cfDNA includes a pre-identified region that is differentially methylated in cancer. In various embodiments, the target region of the bisulfite converted cfDNA includes one or more CpG islands or portions of one or more CpG islands set forth in Tables 1-4.

[0055] In various embodiments, the target region of the bisulfite converted cfDNA comprises one or more CpG islands or portions of one or more CpG islands set forth in Tables 1-4. In various embodiments, the target region of the bisulfite converted cfDNA comprises up to 10%, up to 20%, up to 30%, up to 40%, up to 50%, up to 55%, up to 60%, up to 65%, up to 70%, up to 75%, up to 80%, up to 85%, or up to 90% of the CpG islands or portions of CpG islands set forth in any one of Tables 1-4. In various embodiments, the target region of the bisulfite converted cfDNA comprises 100, up to 150, up to 200, up to 300, up to 400, up to 500, up to 600, up to 700, up to 800, up to 900, up to 1000, up to 1500, up to 2000, up to 2500, up to 3000, up to 3500, or up to 4000 CpG islands or portions of CpG islands selected from Tables 1-4. In various embodiments, analyzing sequence information of subjects not identified as not at risk comprises analyzing sequence information generated from a target region that comprises one or more CpG islands or portions of one or more CpG islands set forth in Tables 1-4. In various embodiments, the target region comprises at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of a CpG island or portion of a CpG island shown in any one of Tables 1-4.In various embodiments, the target region comprises at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1500, at least 2000, at least 2500, at least 3000, at least 3500, at least 4000, at least 4500, at least 5000, at least 5500, or at least 6000 CpG islands or portions of CpG islands selected from Tables 1-4. In various embodiments, performing the second analysis comprises analyzing the methylation status of a greater number of CpG islands compared to the amount of CpG islands analyzed when performing the analysis of the sequence information of the subject. In various embodiments, performing the second analysis comprises analyzing the methylation status of at least 5 times as many CpG islands compared to the amount of CpG islands analyzed when performing the analysis of the sequence information of the subject. In various embodiments, one or more of the CpG islands analyzed when performing the analysis of the sequence information of the subject represent a subset of the CpG islands analyzed when performing the second analysis. In various embodiments, all of the CpG islands analyzed when performing the analysis of the sequence information of the subject are further analyzed when performing the second analysis. In various embodiments, performing the second analysis comprises analyzing the methylation status of at least 500 CpG islands, and performing the analysis of the sequence information of the subject comprises analyzing the methylation status of at least 100 CpG islands.

[0056] In various embodiments, the biological sample is obtained from the subject while the subject is asymptomatic. In various embodiments, the biological sample includes any one of a blood sample, a stool sample, a urine sample, a mucus sample, and a saliva sample. In various embodiments, the biological sample is a blood sample. In various embodiments, the biological sample does not include an invasive biopsy sample. In various embodiments, the assay performed on the biological sample processes one or more of: nucleic acid; cell-free DNA that includes a selected CpG having a selected methylation state; and RNA.

[0057] In various embodiments, the second analysis comprises whole genome sequencing, optionally whole genome bisulfite sequencing. In various embodiments, the non-transitory computer-readable medium further comprises instructions, when executed by the processor, that cause the processor to use the sequence information of the subject to determine the tissue of origin of at least one specific cancer in the subject. In various embodiments, the non-transitory computer-readable medium further comprises instructions, when executed by the processor, that cause the processor to perform the analysis of additional sequence information of the subject obtained from additional biological samples of the subject obtained after the biological sample is obtained; determine one or more changes between the additional sequence information of the subject and the sequence information; and determine the progression of at least one specific cancer in the subject based on the determined one or more changes.

[0058] In various embodiments, the non-transitory computer readable medium further comprises instructions that, when executed by the processor, cause the processor to determine whether to provide an intervention to the subject based on the determined progression of the at least one particular cancer. In various embodiments, the instructions that cause the processor to determine one or more changes between the additional sequence information and the sequence information of the subject, when executed by the processor, further comprise instructions that cause the processor to determine one or more changes in methylation status across a plurality of genomic sites.

[0059] Further disclosed herein is a system comprising a non-transitory computer readable medium comprising: a processor; a data storage device comprising sequence information obtained from the subject's biological sample; and instructions that, when executed by the processor, cause the processor to perform an analysis of the subject's sequence information obtained from the subject's biological sample to identify whether the subject is at risk for having a health condition; and, if the patient is not identified as at risk, analyze the sequence information of the subjects not identified as at risk by performing a second analysis to detect the presence of the health condition in the subject.

[0060] Further disclosed herein is a system comprising a non-transitory computer readable medium comprising a processor; a data storage device comprising marker information obtained from a biological sample of a subject; instructions which, when executed by the processor, cause the processor to perform an analysis of the subject's marker information obtained from the biological sample of the subject to identify whether the subject is at risk for having a health condition; and, if the patient is not identified as at risk, analyze sequence information of the subject not identified as at risk by performing a second analysis to detect the presence of the health condition in the subject. In various embodiments, the marker information comprises quantitative levels of protein biomarkers.

[0061] Further disclosed herein is a system comprising a processor; a data storage device containing marker information obtained from a biological sample of a subject; and instructions that, when executed by the processor, cause the processor to: (a) perform an analysis of sequence information of nucleic acids in the sample to determine whether the analysis produces a result that correlates with the presence or absence of a human condition; and if a result is detected, (b) analyze the sequence information of nucleic acids in the sample by performing a second analysis to determine whether the second analysis produces a signal, where if a signal is detected, the likelihood that the signal in the sample is genuine is increased compared to the likelihood that the signal is genuine if produced by a similar method, the similar method differing in that step (a) is omitted. In various embodiments, the steps performed by the processor achieve a positive predictive value of at least 20% in detecting at least one multiple early stage cancer. In various embodiments, the steps performed by the processor achieve a positive predictive value of at least 40% in detecting at least one multiple early stage cancer. In various embodiments, the steps performed by the processor achieve a positive predictive value of at least 60% in detecting at least one multiple early stage cancer. In various embodiments, the steps performed by the processor achieve a positive predictive value of at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, or at least 85% in detecting the health condition. In various embodiments, the health condition is a disease risk. In various embodiments, the health condition is a rare disease or disorder. In various embodiments, the health condition has an incidence of 1 in 100, 1 in 1,000, 1 in 10,000, 1 in 100,000, 1 in 1,000,000, 1 in 10,000,000, or 1 in 100,000,000.

[0062] Further disclosed herein is a system comprising a non-transitory computer readable medium comprising a processor; a data storage device including sequence information obtained from a first assay performed on a sample obtained from the subject; instructions that, when executed by the processor, cause the processor to perform a screening by analyzing the sequence information to classify the subject as at risk or not at risk for a health condition; in response to the subject being classified as at risk for a health condition, obtain sequence information obtained from a second assay performed on the sample or an additional sample obtained from the subject to generate second assay-derived sequence information; and perform a diagnostic analysis of the sequence information obtained from the second assay on the subject to further classify the subject at risk for the health condition as a candidate for monitoring.

[0063] Further disclosed herein is a system comprising a non-transitory computer readable medium comprising: a processor; a data storage device including sequence information obtained from a first assay performed on a sample obtained from the subject; and instructions that, when executed by the processor, cause the processor to perform a screening by analyzing the sequence information to classify the subject as at risk for a health condition or not at risk for a health condition; if the subject is classified as not at risk for the health condition, report that the subject is not at risk for the health condition; if the subject is classified as at risk for the health condition, obtain sequence information obtained from a second assay performed on the sample or an additional sample obtained from the subject to generate second assay obtained sequence information; and perform a diagnostic analysis of the sequence information obtained from the second assay on the subject to further classify the subject at risk for the health condition as a candidate for monitoring.

[0064] Further disclosed herein is a system comprising a non-transitory computer readable medium comprising: a processor; a data storage device comprising sequence information obtained from a first assay performed on a sample obtained from the subject; instructions that, when executed by the processor, cause the processor to perform a screening by analyzing the sequence information for each of one or more subjects in the plurality of subjects to classify the subject as at risk for a health condition or not at risk for a health condition; in response to the subject being classified as at risk for a health condition, obtain sequence information obtained from a second assay performed on the sample or an additional sample obtained from the subject to generate second assay-derived sequence information; and perform a diagnostic analysis of the sequence information obtained from the second assay for the subjects to further classify the subjects at risk for the health condition as candidate subjects for inclusion in a candidate population.

[0065] Further disclosed herein is a kit comprising: a. an apparatus for collecting a sample from a subject; b. a set of detection reagents that, when combined with the sample, allow for detection of biomarkers in the sample; and c. instructions for accessing computer program instructions stored on a computer storage medium that, when processed by a processor of a computer system, causes the processor to perform an analysis of the sequence information to identify whether the subject is at risk for having one or more early stage cancers; and, if the patient is not identified as at risk, analyze the sequence information of the subjects not identified as at risk from the second analysis to detect the presence of one or more early stage cancers in the subject. In various embodiments, one or more of the early stage cancers are 15 or more different cancers. In various embodiments, the one or more early stage or preclinical cancers are selected from the group consisting of acute lymphocytic leukemia, acute myeloid leukemia, adrenal cortical carcinoma, soft tissue sarcoma, lymphoma, anal cancer, gastrointestinal cancer, brain cancer, skin cancer, bile duct cancer, bladder cancer, bone cancer, breast cancer, lung cancer, heart cancer, central nervous system cancer, cervical cancer, chronic lymphocytic leukemia, chronic myelogenous leukemia, chronic myeloproliferative neoplasms, colon cancer, uterine cancer, esophageal cancer, head and neck cancer, eye cancer, fallopian tube cancer, gallbladder cancer, gastric cancer, germ cell tumors, gestational trophoblastic cancer, and / or ovarian cancer. cancer, hairy cell leukemia, liver cancer, Hodgkin's lymphoma, intraocular melanoma, pancreatic cancer, kidney cancer, leukemia, mesothelioma, metastatic cancer, oral cancer, multiple endocrine neoplasia syndrome, multiple myeloma, myelodysplastic tumor, ovarian cancer, parathyroid cancer, penile cancer, pheochromocytoma, pituitary cancer, plasma cell neoplasm, primary peritoneal cancer, prostate cancer, rectal cancer, retinoblastoma, sarcoma, small intestine cancer, testicular cancer, pharyngeal cancer, thymoma and thymic cancer, thyroid cancer, urethral cancer, uterine cancer, vaginal cancer, and vulvar cancer.

[0066] In various embodiments, one or more of the early stage or preclinical cancers is a single cancer type. In various embodiments, the single cancer type is acute lymphocytic leukemia, acute myeloid leukemia, adrenocortical carcinoma, soft tissue sarcoma, lymphoma, anal cancer, gastrointestinal cancer, brain cancer, skin cancer, bile duct cancer, bladder cancer, bone cancer, breast cancer, lung cancer, heart cancer, central nervous system cancer, cervical cancer, chronic lymphocytic leukemia, chronic myelogenous leukemia, chronic myeloproliferative neoplasms, colon cancer, uterine cancer, esophageal cancer, head and neck cancer, eye cancer, fallopian tube cancer, gallbladder cancer, gastric cancer, germ cell tumors, gestational trophoblastic carcinoma, hairy cell leukemia, ... The cancer is any one of the following: blood cancer, liver cancer, Hodgkin's lymphoma, intraocular melanoma, pancreatic cancer, kidney cancer, leukemia, mesothelioma, metastatic cancer, oral cancer, multiple endocrine neoplasia syndrome, multiple myeloma, myelodysplastic tumor, ovarian cancer, parathyroid cancer, penile cancer, pheochromocytoma, pituitary cancer, plasma cell neoplasm, primary peritoneal cancer, prostate cancer, rectal cancer, retinoblastoma, sarcoma, small intestine cancer, testicular cancer, pharyngeal cancer, thymoma and thymic cancer, thyroid cancer, urethral cancer, uterine cancer, vaginal cancer, and vulvar cancer.

[0067] In various embodiments, the early stage cancer is a preclinical cancer. In various embodiments, the preclinical cancer is a stage I or stage II cancer. In various embodiments, the performance of the analysis and the analysis of the sequence information has a greater than 70% ability to detect at least one multiple early stage cancer. In various embodiments, the performance of the analysis and the analysis of the sequence information has a greater than 70% ability to detect at least one multiple early stage cancer with a specificity of greater than 95%, greater than 96%, greater than 97%, greater than 98%, greater than 99%, greater than 99.5%, or greater than 99.9%. In various embodiments, the performance of the analysis and the analysis of the sequence information achieves a positive predictive value of at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, or at least 85% in detecting at least one multiple early stage cancer. In various embodiments, the performance of the analysis and the analysis of the sequence information achieves a negative predictive value of at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.3%, or at least 99.4% in detecting at least one multiple early stage cancer. In various embodiments, the performance of the analysis has a negative predictive value of at least 90%, at least 95%, or at least 99%. In various embodiments, the analysis of the sequence information to identify whether a subject has a detectable cancer or precancer has a positive predictive value of at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, or at least 85%. In various embodiments, the analysis of the sequence information to identify whether a subject has a detectable cancer or precancer has a negative predictive value of at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, or at least 97%.

[0068] In various embodiments, the sequence information comprises methylation sequence information. In various embodiments, the methylation sequence information comprises a methylation state of a plurality of genomic sites. In various embodiments, the plurality of genomic sites comprises a plurality of CpG sites. In various embodiments, the instructions that cause the processor to perform an analysis of the sequence information of the subject further comprise instructions that, when executed by the processor, cause the processor to apply a trained machine learning model.

[0069] In various embodiments, the sequence information is obtained from an assay, where the assay includes performing one or more of: a. sequencing the nucleic acid in the sample; b. hybrid capture; c. methylation-specific PCR; d. an assay that generates methylation information; and e. sequencing a clone library generated from a template-immortalized library. In various embodiments, performing an assay that generates sequence information includes obtaining bisulfite converted cell-free DNA (cfDNA); selectively amplifying a target region of the bisulfite converted cfDNA; and sequencing the amplification product that includes the amplified target region to generate the methylation information.

[0070] In various embodiments, the target region of the bisulfite converted cfDNA comprises a pre-identified region that is differentially methylated in cancer. In various embodiments, the target region of the bisulfite converted cfDNA comprises one or more CpG islands or portions of one or more CpG islands shown in Table 1. In various embodiments, the target region of the bisulfite converted cfDNA comprises up to 10%, up to 20%, up to 30%, up to 40%, up to 50%, up to 55%, up to 60%, up to 65%, up to 70%, up to 75%, up to 80%, up to 85%, or up to 90% of the CpG islands or portions of CpG islands shown in any one of Tables 1-4. In various embodiments, the target region of the bisulfite converted cfDNA comprises 100, up to 150, up to 200, up to 300, up to 400, up to 500, up to 600, up to 700, up to 800, up to 900, up to 1000, up to 1500, up to 2000, up to 2500, up to 3000, up to 3500, or up to 4000 CpG islands or portions of CpG islands selected from Tables 1-4. In various embodiments, analyzing sequence information of subjects not identified as not at risk comprises analyzing sequence information generated from a target region that comprises one or more CpG islands or portions of one or more CpG islands set forth in Tables 1-4. In various embodiments, the target region comprises at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of a CpG island or portion of a CpG island shown in any one of Tables 1-4.In various embodiments, the target region comprises at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1500, at least 2000, at least 2500, at least 3000, at least 3500, at least 4000, at least 4500, at least 5000, at least 5500, or at least 6000 CpG islands or portions of CpG islands selected from Tables 1-4. In various embodiments, performing the second analysis comprises analyzing the methylation status of a greater number of CpG islands compared to the amount of CpG islands analyzed when performing the analysis of the sequence information of the subject. In various embodiments, performing the second analysis comprises analyzing the methylation status of at least 5 times as many CpG islands compared to the amount of CpG islands analyzed when performing the analysis of the sequence information of the subject. In various embodiments, one or more of the CpG islands analyzed when performing the analysis of the sequence information of the subject represent a subset of the CpG islands analyzed when performing the second analysis. In various embodiments, all of the CpG islands analyzed when performing the analysis of the sequence information of the subject are further analyzed when performing the second analysis. In various embodiments, performing the second analysis comprises analyzing the methylation status of at least 500 CpG islands, and performing the analysis of the sequence information of the subject comprises analyzing the methylation status of at least 100 CpG islands.

[0071] In various embodiments, the biological sample is obtained from the subject while the subject is asymptomatic. In various embodiments, the biological sample includes any one of a blood sample, a stool sample, a urine sample, a mucus sample, and a saliva sample. In various embodiments, the biological sample is a blood sample. In various embodiments, the biological sample does not include an invasive biopsy sample.

[0072] In various embodiments, the assay performed on the biological sample processes one or more of: nucleic acid; cell-free DNA that includes selected CpGs with selected methylation status; and RNA. In various embodiments, the second analysis includes whole genome sequencing, optionally whole genome bisulfite sequencing. In various embodiments, the non-transitory computer-readable medium further includes instructions that, when executed by the processor, cause the processor to use the sequence information of the subject to determine the tissue of origin of at least one specific cancer in the subject. In various embodiments, the computer program instructions further include instructions that, when executed by the processor, cause the processor to perform an analysis of additional sequence information of the subject obtained from additional biological samples of the subject obtained after the time the biological sample was obtained; determine one or more changes between the additional sequence information of the subject and the sequence information; and determine the progression of at least one specific cancer in the subject based on the determined one or more changes.

[0073] In various embodiments, the computer program instructions, when executed by the processor, further comprise instructions that cause the processor to determine whether to provide an intervention to the subject based on the determined progression of the at least one particular cancer. In various embodiments, the computer program instructions, when executed by the processor, that cause the processor to determine one or more changes between the additional sequence information and the sequence information of the subject further comprise instructions that, when executed by the processor, cause the processor to determine one or more changes in methylation status across a plurality of genomic sites.

[0074] Further disclosed herein is a kit comprising: a. an apparatus for obtaining a sample from a subject; b. a set of detection reagents that, when combined with the sample, enable detection of biomarkers in the sample; and c. instructions for accessing computer program instructions stored on a computer storage medium which, when processed by a processor of a computer system, cause the processor to perform an analysis of the subject's sequence information obtained from the subject's sample to identify whether the subject is at risk for having a health condition; and, if the patient is not identified as at risk, analyze the sequence information of the subjects not identified as at risk by performing a second analysis to detect the presence of the health condition in the subject.

[0075] Further disclosed herein is a kit comprising: a. an apparatus for obtaining a sample from a subject; b. a set of detection reagents that, when combined with the sample, allow for detection of biomarkers in the sample; and c. instructions for accessing computer program instructions stored on a computer storage medium that, when processed by a processor of a computer system, causes the processor to perform an analysis of the subject's marker information obtained from the subject's biological sample to identify whether the subject is at risk for having a health condition; and, if the patient is not identified as at risk, analyze sequence information of the subject not identified as at risk by performing a second analysis to detect the presence of the health condition in the subject. In various embodiments, the marker information comprises quantitative levels of protein biomarkers.

[0076] Further disclosed herein is a kit comprising: a. an apparatus for collecting a sample from a subject; b. a set of detection reagents that when combined with the sample allow for detection of a biomarker in the sample; and c. instructions for accessing computer program instructions stored on a computer storage medium that, when processed by a processor of a computer system, causes the processor to: (a) perform an analysis of sequence information of nucleic acids in the sample to determine whether the analysis produces a result that correlates with the presence or absence of a human condition; and if a result is detected, (b) analyze the sequence information of nucleic acids in the sample by performing a second analysis to determine whether the second analysis produces a signal, where if a signal is detected, the likelihood that the signal in the sample is genuine is increased compared to the likelihood that the signal is genuine if produced by a similar method, the similar method differing in that step (a) is omitted. In various embodiments, the steps performed by the processor achieve a positive predictive value of at least 20% in detecting a health condition. In various embodiments, the steps performed by the processor achieve a positive predictive value of at least 40% in detecting a health condition. In various embodiments, the steps performed by the processor achieve a positive predictive value of at least 60% in detecting the health condition. In various embodiments, the steps performed by the processor achieve a positive predictive value of at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, or at least 85% in detecting the health condition. In various embodiments, the health condition is a disease risk. In various embodiments, the health condition is a rare disease or disorder. In various embodiments, the health condition has an incidence of 1 in 100, 1 in 1,000, 1 in 10,000, 1 in 100,000, 1 in 1,000,000, 1 in 10,000,000, or 1 in 100,000,000.

[0077] Further disclosed herein is a kit comprising: a. an apparatus for obtaining a sample from a subject; b. a set of primers that, when combined with the sample, enable detection of multiple sites of cell-free DNA in the sample; and c. instructions for accessing computer program instructions stored on a computer storage medium which, when processed by a processor of a computer system, causes the processor to: perform a screening by analyzing the sequence information to classify the subject as at risk or not at risk for a health condition; responsive to the subject being classified as at risk for a health condition, obtain sequence information from a second assay performed on the sample or an additional sample obtained from the subject to generate second assay derived sequence information; and perform a diagnostic analysis of the sequence information from the second assay on the subject to further classify the subject at risk for the health condition into candidate subjects for monitoring.

[0078] Further disclosed herein is a kit comprising: a. an apparatus for obtaining a sample from a subject; b. a set of primers that, when combined with the sample, enable detection of multiple sites of cell-free DNA in the sample; and c. instructions for accessing computer program instructions stored on a computer storage medium which, when processed by a processor of a computer system, causes the processor to: perform a screening by analyzing the sequence information to classify the subject as at risk for a health condition or not at risk for a health condition; if the subject is classified as not at risk for the health condition, report that the subject is not at risk for the health condition; if the subject is classified as at risk for the health condition, obtain sequence information from a second assay performed on the sample or an additional sample obtained from the subject to generate second assay obtained sequence information; and perform a diagnostic analysis of the sequence information from the second assay on the subject to further classify the subject at risk for the health condition as a candidate for monitoring.

[0079] Further disclosed herein is a kit comprising: a. an apparatus for obtaining a sample from a subject; b. a set of primers that, when combined with the sample, enable detection of multiple sites of cell-free DNA in the sample; and c. instructions for accessing computer program instructions stored on a computer storage medium which, when processed by a processor of a computer system, causes the processor to: perform screening by analyzing the sequence information for each of one or more subjects in the plurality of subjects to classify the subject as at risk for a health condition or not at risk for a health condition; responsive to the subject being classified as at risk for a health condition, obtain sequence information from a second assay performed on the sample or an additional sample obtained from the subject to generate second assay derived sequence information; and perform a diagnostic analysis of the sequence information from the second assay for the subjects to further classify the subjects at risk for the health condition as candidate subjects for inclusion in a candidate population.

[0080] In various embodiments, the multi-step analysis includes an individual-specific analysis (hereinafter referred to as an intra-individual analysis) to determine the presence or absence of an individual's health condition. In other embodiments, the multi-step analysis need not include performing an intra-individual analysis. In general, an intra-individual analysis removes an individual's baseline biological signature, which is less or not informative about the presence or absence of a health condition. By removing the baseline biological signature, the remaining signature is used to more accurately predict the presence or absence of the individual's health condition. An intra-individual analysis is useful because it accounts for a baseline biological signature that may be unique to each individual. As a result, an intra-individual analysis generates an individual's background-corrected signal that accounts for the individual's unique baseline biological signature. Specifically, an intra-individual analysis includes combining sequence information from a target nucleic acid with sequence information from a reference nucleic acid obtained from the individual. The target nucleic acid includes an informative signature for determining the presence or absence of a health condition, and the reference nucleic acid includes the individual's baseline biological signature. By combining sequence information from the target and reference nucleic acids, the resulting combined signal provides more information for determining the presence or absence of a health condition compared to the sequence information of the target nucleic acid alone.

[0081] In various embodiments, the multi-stage analysis further includes a second analysis that analyzes the background corrected signal determined by the within-individual analysis, the second analysis detecting the presence of the health condition in the remaining individuals.

[0082] Overall, multi-stage analysis (e.g., including screening, intra-individual analysis, and second analysis) can achieve improved performance (e.g., high positive predictive value, negative predictive value, sensitivity, and specificity) thereby enabling accurate identification of individuals with a health condition.

[0083] Disclosed herein is a stepwise, multi-part method for detecting circulating tumor DNA in a biological sample of a subject, the method comprising: performing a first analysis of nucleic acid sequence information obtained from a first assay performed on the biological sample to identify whether the biological sample is at risk for containing circulating tumor DNA, and if the biological sample is not identified as at risk, obtaining a target nucleic acid and a reference nucleic acid from the biological sample or an additional biological sample obtained from the subject; performing bisulfite conversion of the target nucleic acid and the reference nucleic acid; selectively amplifying target regions of the bisulfite converted target nucleic acid and / or reference nucleic acid to generate a dataset comprising methylation information from the target nucleic acid and methylation information from the reference nucleic acid; combining, using a computer processor, the methylation information from the target nucleic acid and the methylation information from the reference nucleic acid to generate background-corrected methylation information of the target nucleic acid; and performing a second analysis comprising analyzing the background-corrected methylation information to detect the presence of circulating tumor DNA in the biological sample.

[0084] In various embodiments, the biological sample or the additional biological sample is a blood sample. In various embodiments, obtaining the target nucleic acid and the reference nucleic acid comprises fractionating the biological sample or the additional sample, wherein the target nucleic acid is obtained from a first fraction of the biological sample or the additional biological sample, and the reference nucleic acid is obtained from a second fraction of the biological sample or the additional biological sample. In various embodiments, the target nucleic acid comprises cell-free DNA (cfDNA), and the reference nucleic acid comprises genomic DNA from the subject's cells. In various embodiments, the subject's cells comprise peripheral blood mononuclear cells (PBMCs) or polymorphonuclear cells.

[0085] In various embodiments, combining the methylation information from the target nucleic acid and the methylation information from the reference nucleic acid comprises aligning the methylation information from the target nucleic acid and the methylation information from the reference nucleic acid; and determining a difference between the methylation information from the target nucleic acid and the methylation information from the reference nucleic acid. In various embodiments, the methylation information of the target nucleic acid and the methylation information of the reference nucleic acid both comprise a methylation state of a plurality of genomic sites. In various embodiments, the plurality of genomic sites comprises a plurality of CpG sites set forth in any of Tables 1-4.

[0086] Further disclosed herein is a stepwise, multi-part method for detecting circulating tumor DNA in a biological sample of a subject, the method comprising: performing a first analysis of nucleic acid sequence information obtained from a first assay performed on the biological sample to identify whether the biological sample is at risk for containing circulating tumor DNA, and if the biological sample is not identified as at risk, obtaining a target nucleic acid and a reference nucleic acid from the biological sample or an additional biological sample obtained from the subject; processing the target nucleic acid and the reference nucleic acid to generate a dataset comprising methylation information from the target nucleic acid and methylation information from the reference nucleic acid, Generating the data set includes performing a second assay, the second assay including one or more of: a. sequencing the target nucleic acid and / or the reference nucleic acid by targeted sequencing, whole genome sequencing, or whole genome bisulfite sequencing, b. a nucleic acid amplification assay, and c. an assay for generating methylation information; combining the methylation information from the target nucleic acid and the methylation information from the reference nucleic acid using a computer processor to generate background-corrected methylation information for the target nucleic acid; and performing a second analysis including analyzing the background-corrected methylation information to detect the presence of circulating tumor DNA in the biological sample. In various embodiments, the biological sample or the additional biological sample is a blood sample. In various embodiments, obtaining the target nucleic acid and the reference nucleic acid includes fractionating the biological sample or the additional sample, where the target nucleic acid is obtained from a first fraction of the biological sample or the additional biological sample, and the reference nucleic acid is obtained from a second fraction of the biological sample or the additional biological sample. In various embodiments, the target nucleic acid includes cell-free DNA (cfDNA) and the reference nucleic acid includes genomic DNA from a cell of the subject. In various embodiments, the subject's cells comprise peripheral blood mononuclear cells (PBMCs) or polymorphonuclear cells.

[0087] In various embodiments, combining the methylation information from the target nucleic acid and the methylation information from the reference nucleic acid comprises aligning the methylation information from the target nucleic acid and the methylation information from the reference nucleic acid; and determining a difference between the methylation information from the target nucleic acid and the methylation information from the reference nucleic acid. In various embodiments, the methylation information of the target nucleic acid and the methylation information of the reference nucleic acid both comprise a methylation state of a plurality of genomic sites. In various embodiments, the plurality of genomic sites comprises a plurality of CpG sites set forth in any of Tables 1-4.

[0088] Further disclosed herein is a stepwise, multi-part method for detecting circulating tumor DNA in a biological sample of a subject, the method comprising: performing a first analysis of nucleic acid sequence information obtained from a first assay performed on the biological sample to identify whether the biological sample is at risk of containing circulating tumor DNA, and if the biological sample is not identified as at risk, obtaining a target nucleic acid and a reference nucleic acid from the biological sample or an additional biological sample obtained from the subject; processing the target nucleic acid and the reference nucleic acid to generate a dataset comprising methylation information from the target nucleic acid and methylation information from the reference nucleic acid; combining, using a computer processor, the methylation information from the target nucleic acid and the methylation information from the reference nucleic acid to generate background-corrected methylation information for the target nucleic acid; and performing a second analysis comprising analyzing the background-corrected methylation information to detect the presence of circulating tumor DNA in the biological sample.

[0089] In various embodiments, the biological sample or the additional biological sample is a blood sample. In various embodiments, obtaining the target nucleic acid and the reference nucleic acid comprises fractionating the biological sample or the additional sample, wherein the target nucleic acid is obtained from a first fraction of the biological sample or the additional biological sample, and the reference nucleic acid is obtained from a second fraction of the biological sample or the additional biological sample. In various embodiments, the target nucleic acid comprises cell-free DNA (cfDNA), and the reference nucleic acid comprises genomic DNA from the subject's cells. In various embodiments, the subject's cells comprise peripheral blood mononuclear cells (PBMCs) or polymorphonuclear cells.

[0090] In various embodiments, combining the methylation information from the target nucleic acid and the methylation information from the reference nucleic acid comprises aligning the methylation information from the target nucleic acid and the methylation information from the reference nucleic acid; and determining a difference between the methylation information from the target nucleic acid and the methylation information from the reference nucleic acid. In various embodiments, both the methylation information of the target nucleic acid and the methylation information of the reference nucleic acid comprise a methylation state of a plurality of genomic sites. In various embodiments, the plurality of genomic sites comprise a plurality of CpG sites set forth in any of Tables 1-4. In various embodiments, processing the target nucleic acid and the reference nucleic acid to generate the dataset further comprises performing a target enrichment assay. In various embodiments, the target enrichment assay comprises hybrid capture.

[0091] Further disclosed herein is a method of selecting informative biomarkers for inclusion in a first stage of a stepwise multi-part method for detecting circulating tumor DNA in a biological sample of a subject, the method comprising obtaining a starting set of biomarkers; determining the signals of the starting set of biomarkers across a first plurality of samples and a second plurality of samples; performing a ranking of the starting set of biomarkers using the determined signals; and selecting the top X biomarkers as informative biomarkers for inclusion in the first stage of the stepwise multi-part method. In various embodiments, the starting set of biomarkers comprises one or more of a CpG site, a set of CGIs, a gene, a protein, a nucleic acid, and a metabolite. In various embodiments, the starting set of biomarkers comprises each of a CpG site, a set of CGIs, a gene, a protein, a nucleic acid, and a metabolite. In various embodiments, the first plurality of samples comprises healthy samples or samples in the absence of a health condition. In various embodiments, the healthy samples comprise healthy normal tissue or non-cancer cell-free DNA samples. In various embodiments, the second plurality of samples comprises samples with a health condition. In various embodiments, the sample having a health condition comprises a cancer biopsy sample or a cell-free DNA sample obtained from a patient having cancer, hi various embodiments, the sample having a health condition comprises samples of different cancers.

[0092] In various embodiments, performing a ranking of the starting set of biomarkers using the determined signal comprises ranking the biomarkers based on a difference between the signals of the biomarkers of the first and second plurality of samples. In various embodiments, performing a ranking of the starting set of biomarkers using the determined signal comprises ranking the biomarkers according to their significance in distinguishing between the first and second plurality of samples. In various embodiments, performing a ranking of the starting set of biomarkers using the determined signal comprises ranking the biomarkers by determining an importance value for the biomarkers by implementing a cancer prediction algorithm. In various embodiments, the top X biomarkers comprise between 10 and 200 biomarkers.

[0093] These and other features, aspects, and advantages of the present invention will be better understood with respect to the following description and accompanying drawings. It should be noted that, where possible, similar or similar reference numbers may be used in the figures and may indicate similar or similar functionality. For example, a letter following a reference number, such as "third party agency 155A," indicates that the text specifically refers to the element having that particular reference number. A reference number in a text without a letter following it (such as "third party agency 155") refers to any or all of the elements in the figure having that reference number (e.g., "third party agency 155" in the text refers to reference numbers "third party agency 155A" and / or "third party agency 155B" in the figures). [Brief description of the drawings]

[0094] [Figure 1A] 1 illustrates an overall flow process of a multi-step process for identifying individuals with a health condition, according to an embodiment. [Figure 1B] 1 shows an overall flow process including an intra-individual analysis and a secondary analysis according to a first embodiment. [Figure 1C] 13 shows an overall flow process including an intra-individual analysis and a second analysis according to a second embodiment. [Figure 1D] 13 shows additional analysis examples (eg, a six-step analysis) according to an embodiment. [Figure 1E] 1 illustrates an overall system environment including a condition analysis system, according to an embodiment. [Figure 2A] 1 shows a block diagram of a condition analysis system according to an embodiment. [Figure 2B] 1 shows an example of methylation information useful for determining whether an individual is at risk for a health condition, according to an embodiment. [Figure 2C] 1 shows an example of a flow process for determining whether an individual is at risk for a health condition, according to an embodiment. [Figure 2D] 1 shows an exemplary process for combining sequence information of a target nucleic acid and a reference nucleic acid to generate an informative signal for determining the presence or absence of a health condition, according to an embodiment. [Figure 2E] 1 is an example of a signal providing health status information according to an embodiment. [Figure 2F] 1 shows aligned sequence reads of an analyte and corresponding kmer-sized windows according to an embodiment. [Figure 2G] 1 illustrates generation of metrics from sequence reads across 2k possible patterns, according to an embodiment. [Figure 2H] 1 illustrates an example of a data structure that contains information useful for training a machine learning model, according to an embodiment. [Figure 3A] FIG. 2 shows an interaction diagram between a third party and a condition analysis system for performing a multi-stage analysis according to a first embodiment. [Figure 3B] FIG. 13 shows an interaction diagram between a third party and a condition analysis system for performing a multi-stage analysis according to a second embodiment. [Figure 3C] 1 illustrates an interaction diagram between a first third party, a second third party, and a condition analysis system for performing a multi-stage analysis, according to an embodiment. [Figure 4A] 1 shows an example of a flow process including an in-subject analysis, according to an embodiment. [Figure 4B]FIG. 1 shows an example of a flow process for selecting informative biomarkers for inclusion in the first stage of a multi-stage analysis. [Diagram 5] 1A-1E, 2A-2C, and 3A-3C. [Figure 6A] 1 shows a first process example including a condition analysis system for performing a multi-stage analysis. [Figure 6B] 1 shows a second process example including a condition analysis system for performing a multi-stage analysis. [Figure 6C] 13 shows a third process example including a condition analysis system for performing a multi-stage analysis. [Figure 6D] 1 shows an example of the performance of different stages of a multi-stage analysis for diagnosing an individual having a health condition. [Figure 7] The performance of the single-stage and two-stage analyses for a population containing 1046 samples is shown. [Figure 8] An example of a six-step analysis is shown below. [Figure 9] 1 shows an example of a sample from which the target and reference nucleic acids are obtained. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0095] definition Terms used in the claims and specification are defined as follows, unless otherwise specified.

[0096] The terms "subject," "patient," and "individual" are used interchangeably and include cells, tissues, or organisms, human or non-human, male or female.

[0097] The term "sample" includes an aliquot of a bodily fluid, such as a single cell, multiple cells, fragments of a cell, or a blood sample, obtained from a subject by means including venipuncture, excretion, ejaculation, massage, biopsy, needle aspiration, lavage sample, scraping, surgical incision, intervention, or other means known in the art. Examples of aliquots of bodily fluids include amniotic fluid, aqueous humor, bile, lymphatic fluid, breast milk, interstitial fluid, blood, plasma, cerumen, earwax, Cowper's fluid (pre-ejaculatory fluid), chyle, chyme, female ejaculate, menses, mucus, saliva, urine, vomit, tears, vaginal lubrication, sweat, serum, semen, sebum, pus, pleural effusion, cerebrospinal fluid, synovial fluid, intracellular fluid, and vitreous humor.

[0098] The terms "obtaining information," "obtaining marker information," and "obtaining sequence information" encompass obtaining information determined from at least one sample. Obtaining information (e.g., marker information or sequence information) encompasses obtaining a sample and processing the sample to experimentally determine the information (e.g., marker information or sequence information). The phrase also encompasses receiving information from a third party, for example, who has processed the sample to experimentally determine the information.

[0099] The terms "marker", "biomarker" and "biomarker" include, but are not limited to, lipids, lipoproteins, proteins, cytokines, chemokines, growth factors, peptides, nucleic acids (e.g., DNA or RNA), genes, and oligonucleotides, as well as their associated complexes, metabolites, mutations, variants, polymorphisms, modifications, fragments, subunits, degradation products, elements, and other measurements from an analyte or sample. Markers can also include mutated proteins, mutated nucleic acids, copy number variations, and / or transcript variants in circumstances where such mutations, copy number changes, and / or transcript variants are useful for generating predictive models or are useful in predictive models developed using related markers (e.g., non-mutated versions of proteins or nucleic acids, alternative transcripts, etc.).

[0100] The term "screening" or "first analysis" refers to a step in the first stage of a multi-stage analysis. Screening achieves high specificity and excludes a large proportion of true negatives (e.g., individuals not at risk for a health condition). In various embodiments, "screening" refers to in silico screening, which involves the application of machine learning models. For example, such machine learning models analyze sequence information (e.g., methylation information) and predict whether an individual is likely to be at risk for a health condition.

[0101] The phrase "second analysis" refers to the second step of the multi-stage analysis. The second analysis is performed on individuals identified as at risk for a health condition using screening. Thus, the second analysis achieves a higher positive predictive value than screening, given that screening excludes a large proportion of true negatives. In various embodiments, the "second analysis" refers to an in silico analysis that involves the application of a machine learning model to analyze sequence information (e.g., methylation information) and predict whether an individual has a health condition.

[0102] The phrase "intra-individual analysis" refers to an analysis performed on an individual that removes baseline biological signatures that are less informative in determining whether the individual is at risk for a health condition. In various embodiments, intra-individual analysis includes combining information from the individual's target nucleic acid and reference nucleic acid to generate a signal that provides information for determining the presence or absence of one or more health conditions in the individual. By combining information from the target nucleic acid and the reference nucleic acid, the signal generated can be more informative about the presence or absence of a health condition compared to a signal obtained from the target nucleic acid alone.

[0103] The phrase "target nucleic acid" refers to an individual's nucleic acid that contains at least an informative signature for determining the presence or absence of a health condition. The target nucleic acid may further include an individual's baseline biological signature that is not informative or is less informative. In various embodiments, the target nucleic acid may be a nucleic acid derived from a diseased cell associated with a health condition. For example, the target nucleic acid may be a cell-free nucleic acid derived from a cancer cell. The target nucleic acid may be either DNA, cDNA, or RNA. In certain embodiments, the target nucleic acid includes DNA.

[0104] The phrase "reference nucleic acid" refers to an individual's nucleic acid that contains the individual's baseline biological signature. Wherein, the individual's baseline biological signature may be present when the individual is healthy, and thus, compared with the sequence information of the target nucleic acid, the baseline biological signature provides less information for determining the presence or absence of a health condition. The reference nucleic acid may be either DNA, cDNA, or RNA. In certain embodiments, the reference nucleic acid comprises DNA.

[0105] Please note that as used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.

[0106] Overview of multi-stage analysis Disclosed herein is a multi-step process for detecting the signal that indicates the health status of an individual.For example, the method disclosed herein is useful for detecting circulating tumor DNA from one or more samples obtained from an individual.By detecting circulating tumor DNA from a sample obtained from an individual, the individual can be identified as having a certain health status, such as cancer.

[0107] In various embodiments, the multi-step process is a multi-part method that includes performing a first analysis of nucleic acid sequence information obtained from a first assay performed on a biological sample obtained from an individual. This first analysis identifies whether the biological sample is at risk for containing circulating tumor DNA. In various embodiments, for biological samples determined to not be at risk for containing circulating tumor DNA, the multi-part method further includes performing an intra-individual analysis and a second analysis. In various embodiments, the intra-individual analysis includes obtaining a target nucleic acid and a reference nucleic acid from the biological sample or an additional biological sample obtained from the individual; processing the target nucleic acid and the reference nucleic acid to generate a data set including methylation information from the target nucleic acid and methylation information from the reference nucleic acid; and using a computer processor to combine the methylation information from the target nucleic acid and the methylation information from the reference nucleic acid to generate background-corrected methylation information for the target nucleic acid. Here, the background-corrected methylation information provides more information for determining the presence or absence of a health condition in the individual. In various embodiments, performing the second analysis includes analyzing the background-corrected methylation information to detect the presence of circulating tumor DNA in the biological sample. By detecting the presence of circulating tumor DNA in a biological sample, an individual can be identified as having cancer.

[0108] In general, the multi-stage testing methodologies described herein achieve significant improvements over conventional testing methodologies (e.g., single-stage testing methodologies). For example, the multi-stage testing methodologies described herein achieve improved performance metrics (e.g., sensitivity, specificity, positive predictive value (PPV), and / or negative predictive value (NPV)) over conventional methodologies. In certain embodiments, the combination of first-stage and second-stage testing achieves improved specificity (e.g., true negative rate, reported as the proportion of correctly identified negatives) over conventional methods.

[0109] In some scenarios, the multi-stage testing methodology described herein quickly and accurately excludes a majority of individuals in a first stage with a more efficient and less expensive first stage test, followed by a more stringent second stage test for the remaining subgroup of patients. Here, the multi-stage testing methodology can achieve an overall performance metric that is comparable to or not significantly inferior to the overall performance metric of the conventional methodology. Overall, by quickly and accurately screening a majority of individuals in a first stage, only a minority of individuals undergo the more stringent second stage test. This represents an improvement over conventional methodologies that attempt to apply stringent tests across an entire population, which requires significant resources. Thus, even in scenarios where the multi-stage testing methodology achieves a performance metric comparable to the conventional methodology, the multi-stage testing methodology achieves improved performance as a function of resource consumption. Examples of resource consumption include time resources, monetary resources, and consumable resources (such as consumable assay reagents). In various embodiments, the multi-stage testing methodology disclosed herein achieves at least a 10% reduction in resource consumption compared to the corresponding single-stage test. In various embodiments, the multi-stage inspection methodologies disclosed herein achieve at least a 20% reduction in resource consumption, at least a 30% reduction, at least a 40% reduction, at least a 50% reduction, at least a 60% reduction, at least a 70% reduction, at least an 80% reduction, or at least a 90% reduction, compared to a corresponding single-stage inspection. In various embodiments, the multi-stage inspection methodologies disclosed herein achieve at least a 60% reduction in resource consumption, compared to a corresponding single-stage inspection.

[0110] Further disclosed herein is a multi-step process for screening a patient population and identifying a subset of individuals in the population as having a health condition. The multi-step process includes at least a first stage of screening and excluding a majority of individuals in the population who are not at risk for the health condition. A second stage, including a second analysis, is then performed on the individuals identified as being at risk for the health condition to identify candidate subjects having the health condition. In various embodiments, the method includes performing an intra-individual analysis on the individuals identified as being at risk for the health condition before performing the second analysis. For example, the intra-individual analysis includes generating a signal by removing baseline biological signatures that are less informative for determining whether the individual is at risk for the health condition. Thus, the second analysis includes analyzing the generated signal, which provides more information for determining the presence or absence of one or more health conditions in the individual.

[0111] In various embodiments, the first stage of screening may include a simplified molecular test with high specificity to exclude the majority of true negatives. The second stage of screening includes applying a more complex molecular test that achieves a higher positive predictive value to the resulting mixed population of true positives / false positives (TP / FP). Thus, for large patient populations (e.g., millions, tens of millions, or hundreds of millions), a multi-stage process can rapidly exclude the majority of individuals (e.g., more than 80% of the patient population) that represent true negatives, and identify and diagnose a subset of the population that represents true positives with high positive predictive value (PPV). In various embodiments, individuals identified as true positives (also referred to herein as candidate subjects) can undergo subsequent monitoring and / or treatment. In some embodiments, the candidate subjects are selected for enrollment in a clinical trial (e.g., a clinical trial related to a health condition).

[0112] In certain embodiments, the multi-step process disclosed herein is useful for detecting rare or low incidence health conditions. For example, a rare or low incidence health condition may have an incidence of 1 in 100, 1 in 1,000, 1 in 10,000, 1 in 100,000, 1 in 1,000,000, 1 in 10,000,000, 1 in 100,000,000, or 1 in 1,000,000,000. Thus, the disclosed multi-step process represents a significant improvement over current methodologies that suffer from low specificity or sensitivity, which causes the rare or low incidence conditions to be unable to be detected with sufficient positive predictive value.

[0113] In various embodiments, a multi-step process can be performed to diagnose a subset of individuals in a population as having multiple health conditions. In various embodiments, a multi-step process can be performed to diagnose a subset of individuals in a population as having one of 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, or 20 or more different health conditions. In certain embodiments, the health condition is in the form of cancer. In certain embodiments, a multi-step process can be performed to diagnose a subset of individuals in a population as having one of 10 or more different cancers. In certain embodiments, a multi-step process can be performed to diagnose a subset of individuals in a population as having one of 15 or more different cancers. In certain embodiments, a multi-step process can be performed to diagnose a subset of individuals in a population as having one of 20 or more different cancers. In certain embodiments, the different cancer is an early stage cancer or a preclinical stage cancer. Further examples of health conditions are described in detail herein.

[0114] In certain embodiments, the multi-step process disclosed herein is useful for identifying a signal in a sample obtained from an individual of a patient population. For example, the signal in the sample can provide information about the presence of a health condition. In certain embodiments, the signal provides information about the presence of a rare health condition with a low incidence of 1 in 100, 1 in 1,000, 1 in 10,000, 1 in 100,000, 1 in 1,000,000, 1 in 10,000,000, 1 in 100,000,000, or 1 in 1,000,000,000. Thus, the multi-step process is useful for increasing the likelihood that a detected signal is genuine. Here, the multi-step process may include: (a) performing an analysis of sequence information of nucleic acids in the sample to determine whether the analysis produces a result that correlates with the presence of a human condition; and if a result is detected, (b) analyzing sequence information of nucleic acids in the sample by performing a second analysis to determine whether the second analysis produces a signal. In various embodiments, when a signal is detected, the likelihood that the signal in the sample is authentic is increased compared to the likelihood that the signal is authentic when generated by a similar method that omits step (a). In certain embodiments, the signal in the sample can be informative about the absence of a health condition, where the multi-step process can include: (a) performing an analysis of sequence information of nucleic acids in the sample to determine whether the analysis generates a result that correlates with the absence of a human condition; and, when a result is detected, (b) analyzing the sequence information of nucleic acids in the sample by performing a second analysis to determine whether the second analysis generates a signal. In various embodiments, when a signal is detected, the likelihood that the signal in the sample is authentic is increased compared to the likelihood that the signal is authentic when generated by a similar method that omits step (a).

[0115] 1A illustrates an overall flow process 100 of a multi-step process for identifying individuals having a health condition, according to an embodiment. Although FIG. 1A illustrates the flow process in relation to a single individual 110, in various embodiments, the flow process can be performed for more than a single individual 110 (e.g., thousands, millions, tens of millions, or hundreds of millions of individuals).

[0116] FIG. 1A illustrates a first stage (e.g., assay 120A and screening 125), an in-individual analysis 128 (optionally assay 120B), and a second stage (second analysis 130) of a multi-stage analysis. Generally, the second stage includes more complex molecular tests and analyses compared to the first stage. In various embodiments, the more complex molecular tests of the second stage are more expensive to perform than the simpler molecular tests of the first stage. By utilizing less expensive and less complex tests, the first stage can identify and eliminate individuals who are not at risk for the health condition. The more complex molecular tests and analyses of the second stage can accurately identify the remaining individuals who may have the health condition. In various embodiments, the method includes an in-individual analysis between the first and second stages that removes baseline biological signatures. For example, before performing the second stage (e.g., the more complex molecular tests compared to the first stage), an in-individual analysis can be performed to remove baseline biological signatures (hereinafter referred to as "background correction information") in the sequencing information. Therefore, a second tier of more complex molecular tests can be applied to analyze the background corrected information to achieve improved identification of individuals with health conditions.

[0117] 1A shows the first and second stages of the multi-stage analysis, in various embodiments there may be additional stages to further classify individuals, hi various embodiments the multi-stage analysis includes three or more stages, four or more stages, five or more stages, six or more stages, seven or more stages, eight or more stages, nine or more stages, or ten or more stages.

[0118] In various embodiments, the combination of the first and second stages allows for the ultimate high performance (e.g., high positive predictive value) of the multi-stage analysis. In various embodiments, the first and second stages examine different markers from a sample obtained from an individual. This can be beneficial because different markers can provide different information. In some cases, different markers can inform about different predictions (e.g., whether an individual is at risk for a health condition or whether an individual has a health condition). For example, the first stage can analyze protein markers from a sample obtained from an individual, and the second stage can analyze sequence data obtained from nucleic acids from a sample obtained from an individual.

[0119] In various embodiments, the first and second stages examine the same type of markers from samples obtained from the individual, but at different levels of detail. For example, the first stage may include analysis of the methylation status of a limited, preselected set of genomic sites. The differential methylation of the limited, preselected genomic sites is sufficient to allow identification of individuals who are not at risk for the health condition. Furthermore, the second stage may include analysis of the methylation status of a larger set of genomic sites. In one scenario, the second stage includes analysis of the methylation status of the entire genome (e.g., by whole genome bisulfite sequencing). The differential methylation of the larger set of genomic sites allows accurate identification of the remaining individuals who have the health condition. As another example, the first stage may include analysis of shallow sequencing data, where the shallow sequencing data is sufficient to identify and exclude individuals who are not at risk for the health condition. The second stage may include analysis of sequencing data obtained from deeper sequencing, which is sufficient to identify individuals who have the health condition.

[0120] As shown in FIG. 1A, one or more samples are obtained from an individual 110. In various embodiments, the sample is either a blood sample, a stool sample, a urine sample, a mucus sample, or a saliva sample. In certain embodiments, the one or more samples obtained from the individual 110 are blood samples. The samples may be obtained by the individual or a third party, such as a medical professional. Examples of medical professionals include physicians, paramedics, nurses, first responders, psychologists, phlebotomists, medical physicists, nurse practitioners, surgeons, dentists, and other obvious medical professionals known to those skilled in the art. In various embodiments, the one or more samples may be obtained from the individual 110 by a reference laboratory.

[0121] In various embodiments, the sample obtained from the individual is a liquid biopsy sample obtained at a first time point. In various embodiments, the liquid biopsy sample may include various biomarkers, examples of which include proteins, metabolites, and / or nucleic acids. In certain embodiments, the liquid biopsy sample includes cell-free DNA (cfDNA) fragments. In certain embodiments, the cfDNA fragments include genomic sequences corresponding to CpG islands whose methylation status is informative of a health status.

[0122] In various embodiments, multiple liquid biopsy samples are obtained from the individual 110 at multiple different time points. For example, a first liquid biopsy sample can be obtained at a first time point and at least a second liquid biopsy sample can be obtained from the individual 110 at a second time point. In such embodiments, the first liquid biopsy sample can be used to perform a screen (e.g., screen 125) and the second liquid biopsy can be used to perform a second analysis (e.g., second analysis 130), including an in-individual analysis. Obtaining multiple liquid biopsy samples from the individual at multiple different time points includes obtaining M liquid biopsy samples, where M is any of 2, 3, 4, ..., N-1, N, where N is a positive integer.

[0123] An assay 120A is performed on the acquired sample 115A to generate marker information. Examples of marker information may include quantitative levels of biomarkers, such as protein biomarkers, nucleic acid biomarkers, and metabolite biomarkers, present in the sample. Another example of marker information is sequence information of multiple genomic sites. In various embodiments, considering that the assay 120A may be performed on a large number of samples (e.g., millions of samples) acquired from a large patient population, the assay 120A is a simplified molecular test that generates marker information that can rapidly distinguish individuals at risk for a health condition from those at risk. For example, the marker information may include quantitative levels of biomarkers, such as protein biomarkers, nucleic acid biomarkers, and metabolite biomarkers, which can rapidly lead to the identification and exclusion of individuals at risk for a health condition. As another example, the marker information may be sequence information of a limited number of genomic sites sufficient to identify individuals at risk for a health condition (e.g., true negatives). In certain embodiments, the sequence information of the multiple genomic sites includes methylation information, such as the methylation status of the multiple genomic sites. In various embodiments, the plurality of genomic sites includes a plurality of CpG islands (CGIs) whose differential methylation status may indicate risk for a health condition. Further details regarding assay 120A are described herein.

[0124] Screen 125 is performed to analyze the marker information generated by assay 120A. For example, screen 125 may include an in silico analysis of the marker information. In various embodiments, the marker information includes a quantitative value of a biomarker. Thus, screen 125 may identify and exclude individuals whose quantitative values ​​of the biomarkers indicate that they are not at risk for a health condition. In various embodiments, the marker information is sequence information of multiple genomic sites. Thus, screen 125 includes deploying a trained machine learning model that analyzes sequence information of multiple genomic sites and predicts whether the individual is at risk for a health condition. If screen 125 identifies the individual as not at risk for a health condition (shown as a "negative case" in FIG. 1A), the individual 110 may be reported as not at risk for a health condition. Thus, the individual 110 does not need to undergo any subsequent analysis or be followed up further.

[0125] Alternatively, if the screening identifies the individual as being at risk for the health condition (denoted in FIG. 1A as a "positive case" after screening 125), the individual 110 undergoes at least another stage of testing. As shown in FIG. 1A, a second analysis 130 can be performed on the individual identified as being at risk for the health condition.

[0126] With reference to the intra-individual analysis 128, the analysis is performed on a particular individual, such as an individual identified via screening 125 as being at risk for a health condition. Thus, for a particular patient, whether the patient has a health condition or not, the intra-individual analysis is performed to remove baseline biological signatures present in the patient. These baseline biological signatures may be confounding signals when analyzed to predict whether the patient has the presence or absence of a health condition. Performing the intra-individual analysis 128 removes these confounding baseline biological signatures and retains signatures that are more informative for determining the presence or absence of a health condition. For example, when processing nucleic acid sequencing information to generate a signal that can be detected, the resulting signal may include a mixture of baseline biological signatures (e.g., the patient's germline methylation) that represent a type of background noise and signatures that are informative of a health condition (e.g., cancer). Such background noise may obscure the informative signal of the health condition. Advantageously, in certain embodiments, the methods described herein contemplate subtracting such background noise from a patient's nucleic acid sequence information, thereby improving the signal-to-noise ratio of the informative signal.

[0127] For example, it has been found that performing intra-individual analyses, as opposed to inter-individual analyses in which an average of baseline signatures from a group of normal subjects is removed from a patient's nucleic acid sequence information to determine the presence or absence of one or more health conditions in a patient, greatly improves the sensitivity or specificity of detecting signals that provide information for determining the presence or absence of a health condition.

[0128] In general, the intra-individual analysis 128 involves generating information from at least target and reference nucleic acids from one or more samples obtained from a patient. In various embodiments, the intra-individual analysis 128 is performed on sequence information. Such sequence information may be generated by assay 120A, as shown in FIG. 1A. In such a scenario, both screening 125 and the intra-individual analysis 128 can be performed using sequence information generated by assay 120A. In various embodiments, the intra-individual analysis 128 is performed on sequence information generated by an assay different from assay 120A (e.g., assay 120B). As shown in FIG. 1A, the performance of assay 120B is optional (indicated by a dotted line). In various embodiments, assay 120B is performed on sample 115A, which is the same sample 115A on which assay 120A was performed. In various embodiments, assay 120B is performed on a second sample obtained from individual 110, where the second sample is different from sample 115A. For example, the second sample may be obtained from the individual 110 at a different time than the sample 115A was obtained from the individual 110. Thus, screening 125 and intra-individual analysis 128 are performed on information generated from assays performed on the different samples. More detailed embodiments of samples and / or assays used to perform intra-individual analysis 128 are described below with reference to Figures 1B and 1C.

[0129] In various embodiments, the intra-individual analysis 128 includes combining information from the target nucleic acid and the reference nucleic acid to generate a signal that is informative for determining the presence or absence of one or more health conditions in the patient. By combining information from the target nucleic acid and the reference nucleic acid, the generated signal can be more informative about the presence or absence of a health condition compared to a signal obtained from the target nucleic acid alone. For example, the information from the reference nucleic acid can represent the baseline biology of the patient. By combining information from the target nucleic acid and the reference nucleic acid, the baseline biology of the patient that may not be informative about the presence or absence of a health condition is removed from the generated signal. Thus, the information of the target nucleic acid that is not attributable to the baseline biology of the patient remains and is included in the generated signal for determining the presence or absence of one or more health conditions in the patient.

[0130] Referring now to the second analysis 130 shown in FIG. 1A, the second analysis 130 may reveal that the individual does not have the health condition. If the individual is predicted to not have the health condition (e.g., a "negative case" as shown in FIG. 1A), the individual is reported to not have the health condition. In various embodiments, the individual reported to not have the health condition can be further monitored, given that the screening 125 identified the individual as at risk for the health condition (or not at risk). Alternatively, if the individual is predicted to have the health condition (e.g., a "positive case" as shown in FIG. 1A after the second analysis 130), the individual is reported to have the health condition. In various embodiments, the individual is monitored for the progression of the health condition. In various embodiments, the individual is provided with a treatment to control or ameliorate the health condition. In various embodiments, the individual is selected for enrollment in a clinical trial.

[0131] In general, a multi-stage analysis (e.g., a multi-stage analysis including a screening 125 and a second analysis 130, or a multi-stage analysis including each of the screening 125, the in-individual analysis 128, and the second analysis 130) can rapidly identify a large proportion of individuals (e.g., more than 80% of the patient population) that exhibit true negatives, and can also accurately identify and diagnose a subset of the population that exhibit true positives. The overall multi-stage analysis (e.g., a multi-stage analysis including a screening 125 and a second analysis 130, or a multi-stage analysis including each of the screening 125, the in-individual analysis 128, and the second analysis 130) achieves one or more performance indicators, such as the metrics of sensitivity, specificity, positive predictive value (PPV), and / or negative predictive value (NPV). Sensitivity is the true positive rate, reported as a proportion of correctly identified positives. Specificity is the true negative rate, reported as a proportion of correctly identified negatives. Positive predictive value refers to the number of true positives divided by the total number of true positives and false positives. Negative predictive value refers to the true negative rate divided by the sum of true negatives and false negatives.

[0132] In various embodiments, the overall multistage analysis (e.g., a multistage analysis including a screen 125 and a second analysis 130, or a multistage analysis including each of a screen 125, an in-individual analysis 128, and a second analysis 130) achieves a sensitivity of at least 60% in detecting the presence of a health condition. In various embodiments, the overall multistage analysis achieves a sensitivity of at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least In certain embodiments, the overall multistage analysis achieves a sensitivity of at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9%. In certain embodiments, the overall multistage analysis achieves a sensitivity of at least 70%. In certain embodiments, the overall multistage analysis achieves a sensitivity of at least 71%. In certain embodiments, the overall multistage analysis achieves a sensitivity of at least 72%. In certain embodiments, the overall multistage analysis achieves a sensitivity of at least 73%. In certain embodiments, the overall multistage analysis achieves a sensitivity of at least 74%. In certain embodiments, the overall multi-stage assay achieves a sensitivity of at least 75%.

[0133] In various embodiments, the overall multistage analysis (e.g., a multistage analysis including screen 125 and second analysis 130, or a multistage analysis including each of screen 125, in-individual analysis 128, and second analysis 130) achieves a specificity of at least 60% in excluding individuals without the health condition. In various embodiments, the overall multistage analysis achieves a specificity of at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least At least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% specificity. In certain embodiments, the overall multi-stage analysis achieves a specificity of at least 99%. In certain embodiments, the overall multi-stage analysis achieves a specificity of at least 99.5%. In certain embodiments, the overall multi-stage analysis achieves a specificity of at least 99.9%.

[0134] In various embodiments, the overall multistage analysis (e.g., a multistage analysis including a screen 125 and a second analysis 130, or a multistage analysis including each of a screen 125, an in-individual analysis 128, and a second analysis 130) achieves a particular sensitivity and a particular specificity. The combination of sensitivity and specificity limits both the number of false positives and the number of false negatives. In various embodiments, the overall multistage analysis achieves a sensitivity of 70%-90% and a specificity of 90%-100%. In various embodiments, the overall multistage analysis achieves a sensitivity of 75%-89% and a specificity of 90%-100%. In various embodiments, the overall multistage analysis achieves a sensitivity of 80%-88% and a specificity of 90%-100%. In various embodiments, the overall multistage analysis achieves a sensitivity of 83%-87% and a specificity of 90%-100%. In various embodiments, the overall multistage assay achieves a sensitivity of 84%-86% and a specificity of 90%-100%. In various embodiments, the overall multistage assay achieves a sensitivity of about 85% and a specificity of 90%-100%.

[0135] In various embodiments, the overall multistage analysis (e.g., a multistage analysis including a screen 125 and a second analysis 130, or a multistage analysis including each of a screen 125, an in-individual analysis 128, and a second analysis 130) achieves a sensitivity of 70%-90% and a specificity of 91%-99%. In various embodiments, the overall multistage analysis achieves a sensitivity of 70%-90% and a specificity of 92%-98%. In various embodiments, the overall multistage analysis achieves a sensitivity of 70%-90% and a specificity of 93%-97%. In various embodiments, the overall multistage analysis achieves a sensitivity of 70%-90% and a specificity of 97%-96%. In various embodiments, the overall multistage analysis achieves a sensitivity of 70%-90% and a specificity of about 95%.

[0136] In various embodiments, the overall multistage analysis (e.g., a multistage analysis including a screen 125 and a second analysis 130, or a multistage analysis including each of a screen 125, an in-individual analysis 128, and a second analysis 130) achieves a sensitivity of 75%-89% and a specificity of 91%-99%. In various embodiments, the overall multistage analysis achieves a sensitivity of 80%-88% and a specificity of 92%-98%. In various embodiments, the overall multistage analysis achieves a sensitivity of 83%-87% and a specificity of 93%-97%. In various embodiments, the overall multistage analysis achieves a sensitivity of 84%-86% and a specificity of 94%-96%. In various embodiments, the overall multistage analysis achieves a sensitivity of about 85% and a specificity of about 95%.

[0137] In various embodiments, the overall multistage analysis (e.g., a multistage analysis including a screen 125 and a second analysis 130, or a multistage analysis including each of a screen 125, an in-individual analysis 128, and a second analysis 130) achieves a positive predictive value of at least 60%. In various embodiments, the overall multistage analysis achieves a positive predictive value of at least 20%. In various embodiments, the overall multistage analysis achieves a positive predictive value of at least 20%, at least 21%, at least 22%, at least 23%, at least 24%, at least 25%, at least 26%, at least 27%, at least 28%, at least 29%, at least 30%, at least 31%, at least 32%, at least 33%, at least 34%, at least 35%, at least 36%, at least 37%, at least 38%, at least 39%, or at least 40%. In various embodiments, the overall multistage analysis achieves a positive predictive value of at least 40%. In various embodiments, the overall multistage analysis achieves a positive predictive value of at least 40%, at least 41%, at least 42%, at least 43%, at least 44%, at least 45%, at least 46%, at least 47%, at least 48%, at least 49%, at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, or at least 60%.In various embodiments, the overall multi-stage analysis is at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119%, at least 120%, at least 122%, at least 123%, at least 124%, at least 125%, at least 126%, at least 127%, at least 128%, at least 129%, at least 130%, at least 131%, at least 132%, at least 133%, at least 134%, at least 135%, at least 136%, at least 137%, at In certain embodiments, the overall multistage analysis achieves a positive predictive value of at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9%. In certain embodiments, the overall multistage analysis achieves a positive predictive value of at least 80%. In certain embodiments, the overall multistage analysis achieves a positive predictive value of at least 81%. In certain embodiments, the overall multistage analysis achieves a positive predictive value of at least 82%. In certain embodiments, the overall multistage analysis achieves a positive predictive value of at least 83%. In certain embodiments, the overall multistage analysis achieves a positive predictive value of at least 84%. In certain embodiments, the overall multistage analysis achieves a positive predictive value of at least 85%.

[0138] In various embodiments, the overall multistage analysis (e.g., a multistage analysis including screen 125 and second analysis 130, or a multistage analysis including each of screen 125, in-individual analysis 128, and second analysis 130) achieves a negative predictive value of at least 60%. In various embodiments, the overall multistage analysis achieves a negative predictive value of at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least In certain embodiments, the overall multistage analysis achieves a negative predictive value of at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9%. In certain embodiments, the overall multistage analysis achieves a negative predictive value of at least 98%. In certain embodiments, the overall multistage analysis achieves a negative predictive value of at least 99%. In certain embodiments, the overall multistage analysis achieves a negative predictive value of at least 99.4%.

[0139] In various embodiments, the individual identified as having a health condition can undergo additional analysis. Additional analysis may refer to classifying the individual identified as having a health condition as a candidate subject to be selected for enrollment in a clinical trial. Thus, the multi-stage analysis disclosed herein allows for accurate identification of individuals (from among a large patient population) who have a health condition and therefore meet the eligibility criteria for enrollment in a clinical trial. The multi-stage analysis allows for avoiding enrollment of individuals who do not have a health condition in a clinical trial, thereby reducing the consumption of resources that would otherwise be erroneously allocated to these individuals.

[0140] In various embodiments, the additional analysis refers to longitudinal monitoring of an individual identified as having a health condition. For example, at a later time point, an additional sample may be obtained from the individual identified as having a health condition, and an assay (e.g., assay 120A or assay 120B) may be performed to generate marker information. The marker information may be analyzed by performing one or both of the screening and second analysis. The results of the screening and / or second analysis may be compared to the results of previous screening and / or second analysis to understand longitudinal changes in the individual's health condition. In some scenarios, the longitudinal changes may guide the intervention therapy provided to the individual. Further details of longitudinal analysis are described herein.

[0141] Reference is now made to Figures 1B and 1C, which show an overall flow process including an in-individual analysis and a second analysis, respectively. In general, the in-individual analysis 128 and the second analysis 130 are performed on individual patients who have previously been determined (e.g., via screening 125 shown in Figure 1A) to be at risk for a health condition or who have not been identified as not at risk. The in-individual analysis 128 removes a baseline biological signature unique to the individual patient to generate a background-corrected signal. The second analysis 130 thus includes analyzing the background-corrected signal to determine whether the individual has a health condition. Although Figures 1B and 1C each show a flow process related to a single individual, in various embodiments, the flow process can be performed on more than a single individual (e.g., thousands, millions, tens of millions, or hundreds of millions of individuals).

[0142] Referring first to FIG. 1B, an embodiment is shown in which the intra-individual analysis 128 and the second analysis 130 are performed using a single sample 115. In various embodiments, the sample 115 is a blood sample, a stool sample, a urine sample, a mucus sample, or a saliva sample. In certain embodiments, the sample 115 obtained from the individual is a blood sample. The sample 115 may be obtained by the individual or a third party, such as a medical professional. Examples of medical professionals include physicians, paramedics, nurses, first responders, psychologists, phlebotomists, medical physicists, nurse practitioners, surgeons, dentists, and other obvious medical professionals known to those of skill in the art. In various embodiments, the sample 115 may be obtained from the individual by a reference laboratory. In various embodiments, the single sample 115 may be the same sample obtained from the individual previously used to perform the assay 120A and the screen 125 (e.g., sample 115A shown in FIG. 1A).

[0143] In various embodiments, the target nucleic acid and the reference nucleic acid may be obtained from a single sample 115. The target nucleic acid may include an informative signature for determining the presence or absence of a health condition, and may further include a baseline biological signature. Here, the target nucleic acid in a blood sample may be derived from diseased cells associated with the health condition. For example, the target nucleic acid may include cell-free DNA in blood derived from diseased cells. In certain embodiments, the target nucleic acid is cell-free DNA in blood derived from cancer cells. The reference nucleic acid in the sample 115 refers to a nucleic acid that contains the baseline biological signature of an individual. For example, the baseline biological signature of an individual may be present in the nucleic acid, regardless of whether the nucleic acid originates from a diseased or non-disease source. The baseline biological signature of the reference nucleic acid is generally less informative for determining the presence or absence of a health condition compared to the informative signature present in the target nucleic acid. In various embodiments, the reference nucleic acid refers to cellular genomic DNA obtained from healthy cells of an individual. In various embodiments, the reference nucleic acid found in the sample is derived from cells of a healthy organ of the individual. Examples of organs include brain, heart, breast, lung, abdomen, colon, cervix, pancreas, kidney, liver, muscle, lymph nodes, esophagus, intestine, spleen, stomach, and gallbladder. In certain embodiments, the reference nucleic acid refers to cellular genomic DNA found in a sample and derived from peripheral blood mononuclear cells (PBMCs) (e.g., lymphocytes or monocytes) or polymorphonuclear cells (e.g., eosinophils or neutrophils).

[0144] In various embodiments, the target nucleic acid and the reference nucleic acid are obtained separately from a single sample 115. In various embodiments, the sample is processed to separate the target nucleic acid and the reference nucleic acid. For example, the sample may be processed by any one of centrifugation, filtration, gel electrophoresis, bead capture, or matrix extraction. In certain embodiments, the target nucleic acid is cell-free nucleic acid and thus may be obtained from the supernatant of the separated sample. In certain embodiments, the reference nucleic acid is cell genomic nucleic acid and thus may be obtained from a different portion of the separated sample containing cells.

[0145] As shown in FIG. 1B, a single sample 115 can be used to perform two separate assays, such as assay 120A and assay 120B. In various embodiments, a first assay 120A is performed to generate information derived from a target nucleic acid in the sample 115. In various embodiments, a second assay 120B is performed to generate information derived from a reference nucleic acid in the sample 115. As described in more detail herein, an intra-individual analysis 128 is performed to combine the information derived from the target nucleic acid and the information derived from the reference nucleic acid. Thus, the intra-individual analysis 128 generates background-corrected information that is analyzed via a second analysis 130 to determine whether an individual has a health condition.

[0146] Reference is now made to FIG. 1C, which shows an alternative embodiment in which two samples (e.g., labeled as sample 115B and sample 115C) are used to perform the intra-individual analysis 128 and the second analysis 130. Here, samples 115B and 115C can be obtained from individuals who have been previously identified as at risk for a health condition via screening 125, or from individuals who have not been identified as not at risk for a health condition. In various embodiments, one of the samples contains a target nucleic acid and the other of the samples contains a reference nucleic acid. Thus, in such embodiments, a target nucleic acid is obtained from one of the samples and a reference nucleic acid is obtained from the other of the samples. Separate assays (e.g., assays 120A and 120B) can be performed on the target nucleic acid and the reference nucleic acid.

[0147] In a particular embodiment shown in FIG. 1C, samples 115B and 115C are obtained from the individual and used to perform the intra-individual analysis 128 and the second analysis 130. In various embodiments, one of the samples may be the sample 115A used to perform the assay 120A and the screen 125 as shown in FIG. 1A. This therefore avoids the need to obtain a new sample from the individual and allows the previous sample to be reused. In various embodiments, multiple samples 115 are obtained from the individual 110 at multiple different time points. For example, a first sample 115A is obtained at a first time point, a second sample 115B is obtained from the individual 110 at a second time point, a third sample 115C is obtained from the individual 110 at a third time point, and so on. Obtaining multiple samples 115 from the individual at multiple different time points includes obtaining M samples 115, where M is 2, 3, 4, ..., N-1, N (N is a positive integer). In such an embodiment, the target and reference nucleic acids can be obtained at different time points, thereby allowing for an intra-individual analysis 128 and a second analysis 130 across different time points. This allows for tracking the progression of a health condition across different time points.

[0148] In various embodiments, the sample 115 may be processed to extract target and reference nucleic acids. In various embodiments, the sample may undergo cell disruption methods (e.g., to obtain genomic DNA), including chemical or mechanical methods. Examples of chemical methods include osmotic shock, enzyme digestion, detergent, or alkaline treatment. Examples of mechanical methods include homogenization, sonication or cavitation, pressure cell, or ball mill. In various embodiments, the sample may undergo removal of membrane lipids or proteins, or purification of nucleic acids. Examples of chemical methods for removing membrane lipids or proteins and methods for nucleic acid purification include guanidine thiocyanate (GuSCN)-phenol-chloroform extraction, alkaline extraction, cesium chloride gradient centrifugation with ethidium bromide, Chelex® extraction, or cetyltrimethylammonium bromide extraction. Examples of physical methods for removing membrane lipids or proteins and methods for nucleic acid purification include solid phase extraction using either silica matrix, glass particles, diatomaceous earth, magnetic beads, anion exchange material, or cellulose matrix. Further details of nucleic acid extraction methods are described in Ali et al, Current Nucleic Acid Extraction Methods and Their Implications to Point-of-Care Diagnostics, Biomed Res. Int. 2017;2017:9306564, which is incorporated by reference in its entirety.

[0149] As shown in Figures 1B and 1C, one or more assays (e.g., assay 120A and / or assay 120B) are performed on the acquired sample 115A and / or sample 115B to generate sequence information. Although the methods shown in Figures 1B and 1C include performing assay 120A, which is also performed in Figure 1A, in some embodiments, the methods of Figures 1B and 1C do not need to perform assay 120A, but instead perform an assay different from assay 120A performed in Figure 1A. For intra-individual analysis 128, generally, assay 120 is performed to generate sequence information of a target nucleic acid and to generate sequence information of a reference nucleic acid. In certain embodiments, the sequence information includes the state of multiple genomic sites, such as the epigenetic state of multiple CpG sites. In various embodiments, the epigenetic state refers to a methylation state. In certain embodiments, the sequence information of the target nucleic acid and the sequence information of the reference nucleic acid include the status of 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, or 10 or more common genomic sites. In certain embodiments, the sequence information of the target nucleic acid and the sequence information of the reference nucleic acid each include the status of 15 or more, 20 or more, 25 or more, 30 or more, 40 or more, 50 or more, 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, 750 or more, 1000 or more, 2000 or more, 3000 or more, 4000 or more, 5000 or more, 6000 or more, 7000 or more, 8000 or more, 9000 or more, 10000 or more, 11000 or more, 12000 or more, 13000 or more, 14000 or more, 15000 or more, 16000 or more, 17000 or more, 18000 or more, 19000 or more, or 20000 or more genomic sites.In certain embodiments, the sequence information of the target nucleic acid and the sequence information of the reference nucleic acid each include the status of 15 or more, 20 or more, 25 or more, 30 or more, 40 or more, 50 or more, 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, 750 or more, 1000 or more, 2000 or more, 3000 or more, 4000 or more, 5000 or more, 6000 or more, 7000 or more, 8000 or more, 9000 or more, 10000 or more, 11000 or more, 12000 or more, 13000 or more, 14000 or more, 15000 or more, 16000 or more, 17000 or more, 18000 or more, 19000 or more, or 20000 or more of the same or overlapping genomic sites. In various embodiments, the plurality of genomic sites includes a plurality of CpG islands (CGIs) whose differential methylation status may be indicative of a health state. Further details regarding assay 120 are described herein.

[0150] Intra-individual analysis 128 includes combining sequence information of a target nucleic acid with sequence information of a reference nucleic acid to generate an informative signal for determining the presence or absence of a health condition, where the informative signal for determining the presence or absence of a health condition is more informative for determining the presence or absence of a health condition than the sequence information of the target nucleic acid alone. In certain embodiments, the informative signal for determining the presence or absence of a health condition includes an informative signature from the target nucleic acid (e.g., a signature derived from diseased cells) and excludes a baseline biological signature (e.g., a baseline biological signature present in a reference nucleic acid). Further details of intra-individual analysis 128, particularly the generation of an informative background-corrected signal for determining the presence or absence of a health condition, are described herein.

[0151] In various embodiments, the second analysis 130 includes analyzing the background-corrected signal from the within-individual analysis 128 to predict whether the individual has a health condition. Thus, as shown in both FIG. 1B and 1C, the output of the second analysis 130 can be a determination of whether the individual has a health condition. In various embodiments, the determination can be useful to guide a decision to treat the individual. For example, if the determination reveals that the individual has a health condition, the individual can be provided with a therapy (e.g., a prophylactic therapy, a preventative therapy) to treat the health condition.

[0152] FIG. 1D illustrates an example of an additional analysis (e.g., a six-step analysis) according to an embodiment. Specifically, FIG. 1D illustrates an example of an additional analysis including a third analysis 140, a fourth analysis 142, a fifth analysis 144, and a sixth analysis 146. Here, the additional analysis illustrated in FIG. 1D can be performed after the second analysis 130, as described in any of FIG. 1A, 1B, or 1C. For example, an individual who is reported to have a health condition after the second analysis 130 can undergo the additional analysis illustrated in FIG. 1D. In various embodiments, one or more samples are obtained from the individual when performing the third analysis 140, the fourth analysis 142, the fifth analysis 144, and the sixth analysis 146. In certain embodiments, a new sample is obtained from the individual when performing each of the third analysis 140, the fourth analysis 142, the fifth analysis 144, and the sixth analysis 146.

[0153] Generally, the third analysis 140 includes an interventional monitoring analysis of an individual reported to have a health condition. In various embodiments, the third analysis 140 may include performing one or more assays, such as one or more assays performed during the second analysis 130, as described in any of Figures 1A-1C. In various embodiments, the third analysis 140 includes a longitudinal analysis, as described in further detail herein. Thus, an individual reported to have a health condition can be continuously monitored for changes in health condition. If the third analysis 140 produces a positive result, the individual can proceed to a fourth analysis 142. In various embodiments, a positive result refers to the detection of an increasing signal, which indicates a progression or increasing severity of the health condition. If the result of the third analysis 140 is negative, the individual can continue to be monitored by repeating the third analysis 140 (e.g., a longitudinal analysis). In such a scenario, the third analysis 140 may be performed on the individual every month, every two months, every three months, every four months, every five months, every six months, every seven months, every eight months, every nine months, every ten months, every eleven months, every twelve months, every two years, every three years, every four years, every five years, every six years, every seven years, every eight years, every nine years, or every ten years. In certain embodiments, the third analysis 140 may be performed on the individual every six months.

[0154] As shown in FIG. 1D, the additional analysis may further include a fourth analysis 142 and / or a fifth analysis 144. In various embodiments, the fourth analysis 142 and the fifth analysis 144 refer to treatment analyses for the individual. In various embodiments, only one of the fourth analysis 142 and the fifth analysis 144 is performed. For example, if the fourth analysis 142 determines that the individual can be administered a standard treatment and that the standard treatment results in a successful therapeutic outcome, there is no need to perform the fifth analysis 144. Thus, as shown in FIG. 1D, a "branch A" can be selected that proceeds directly to the sixth analysis 146. As another example, if the standard treatment of the fourth analysis 142 does not result in a successful therapeutic outcome, the individual can proceed to the fifth analysis 144, which may include identifying and providing an alternative treatment. Thus, as shown in FIG. 1D, a "branch B" can be selected in which the individual who has undergone the fourth analysis 142 further undergoes the fifth analysis 144.

[0155] In various embodiments, the fourth analysis 142 refers to one or more standard analyses and treatments. For example, the fourth analysis 142 may include providing a standard of care analysis and / or treatment for the individual. Examples of standard of care analyses include imaging scans (e.g., magnetic resonance imaging (MRI) scans, computed tomography (CT) scans), tissue biopsies, body fluid tests, or companion diagnostics. In light of the standard of care analysis, a standard of care treatment may be provided for the individual.

[0156] In various embodiments, the fifth analysis 144 refers to one or more personalized analyses and treatments. An example of a personalized analysis may include sequencing (e.g., DNA or RNA sequencing) to identify the presence or absence of a particular marker in an individual. Thus, a personalized treatment may be provided to treat the individual according to the personalized analysis. One example of a personalized treatment may include a personalized vaccine that includes neoepitopes designed to induce an immune response in the individual.

[0157] By the time of the sixth analysis 146, the individual has undergone treatment to treat the health condition. The sixth analysis 146 includes performing longitudinal monitoring of the individual, for example, for recurrence of the health condition. In various embodiments, the sixth analysis 146 may include obtaining samples from the individual at set time intervals and performing one or more assays on the obtained samples. For example, the one or more assays performed for the sixth analysis 146 may be similarly assays performed for the second analysis 130 or the third analysis 140, as described above. Thus, the one or more assays performed in the sixth analysis 146 may be informative about the health condition of the individual, such as the recurrence of the health condition.

[0158] System environment overview FIG. 1E illustrates an overall system environment 150 including a condition analysis system 170 for performing a multi-stage analysis according to an embodiment. The overall system environment 150 includes a condition analysis system 170 that performs at least one or more steps illustrated in FIG. 1A, and one or more third parties 155A and 155B that communicate with each other via a network 160. FIG. 1B illustrates one embodiment of the overall system environment 150 including two third parties 155A and 155B. In other embodiments, additional or fewer third parties 155 that communicate with the condition analysis system 170 can be included. The third parties 155 can communicate with the condition analysis system 170 to enable the condition analysis system 170 to perform screening and / or second analysis.

[0159] Third Party The third party organization 155 represents a partner organization of the condition analysis system 170 that can operate upstream, downstream, or both upstream and downstream of the operation of the condition analysis system 170. As an example, the third party organization 155 operates upstream of the condition analysis system 170 and provides a sample obtained from a patient to the condition analysis system 170. Thus, the condition analysis system 170 can perform assays, screening, in-person analyses, and / or second analyses to determine whether the patient is at risk for a health condition or whether the patient has a health condition. As another example, the third party organization 155 can process the sample obtained from the patient by performing one or more assays on the sample to generate data. Thus, the third party organization 155 can provide data obtained from the assays to the condition analysis system 170, and the condition analysis system 170 can perform screening, in-person analyses, and / or second analyses.

[0160] As another example, the third party 155 operates downstream of the condition analysis system 170. In this scenario, the condition analysis system 170 may perform screening to determine whether the patient is at risk for a health condition. The condition analysis system 170 may provide instructions to the third party 155 to identify the patient at risk for the health condition. The third party 155 takes appropriate action. For example, the third party 155 notifies the patient regarding a follow-up appointment so that additional samples can be obtained from the patient at the follow-up appointment for subsequent analysis. Further explanations and examples of the interrelationship between the third party 155 and the condition analysis system 170 are described in detail herein.

[0161] network The present disclosure contemplates any suitable network 160 that allows for a connection between the condition analysis system 170 and the third party 155. The network 160 may include any combination of local area networks and / or wide area networks using both wired and / or wireless communication systems. In one embodiment, the network 160 uses standard communication technologies and / or protocols. For example, the network 160 may include communication lines using technologies such as Ethernet, 802.11, Worldwide Interoperable Microwave Access (WiMAX), 3G, 4G, Code Division Multiple Access (CDMA), Digital Subscriber Line (DSL), etc. Examples of network protocols used to communicate over the network 160 include Multiprotocol Label Switching (MPLS), Transmission Control Protocol / Internet Protocol (TCP / IP), Hypertext Transfer Protocol (HTTP), Simple Mail Transfer Protocol (SMTP), and File Transfer Protocol (FTP). Data exchanged over the network 160 may be represented using any suitable format, such as Hypertext Markup Language (HTML) or Extensible Markup Language (XML). In some embodiments, all or part of the communication links of network 160 may be encrypted using any suitable technique.

[0162] Condition Analysis System FIG. 2A shows a block diagram of a condition analysis system according to an embodiment. The block diagram of the condition analysis system 170 is introduced to show an embodiment in which the condition analysis system 170 includes one or more assay devices 205 communicatively coupled to a computing system 202. The computing system 202 may further include computing modules such as a screening module 210, a signal generation module 215, a condition analysis module 220, and optionally a longitudinal analysis module 230. The computing system 202 may further include a data storage device such as a machine learning model storage device 240 that stores one or more trained machine learning models. FIG. 2A shows an embodiment in which the condition analysis system 170 performs one or more assays (e.g., assays 120A or 120B described in FIG. 1A), performs screening (e.g., screening 125 described in FIG. 1A), performs an in-individual analysis (e.g., in-individual analysis 128 described in FIG. 1A), and performs a second analysis (e.g., second analysis 130 described in FIG. 1A).

[0163] In various embodiments, the condition analysis system 170 may be configured differently than that shown in FIG. 2A. For example, while the condition analysis system 170 shown in FIG. 2A includes three different assay devices 205, in various embodiments, the condition analysis system 170 includes fewer or additional assay devices. In certain embodiments, the condition analysis system 170 does not include an assay device. In such embodiments, the condition analysis system 170 includes only the computing system 202. In those embodiments in which the condition analysis system 170 does not include an assay device, the condition analysis system 170 may perform a screening (e.g., screening 125 as shown in FIG. 1A), an intra-individual analysis (e.g., intra-individual analysis 128 as shown in FIG. 1A), and a second analysis (e.g., second analysis 130 as shown in FIG. 1A). However, the condition analysis system 170 does not perform an assay. The assay device 205 may be operated and used by a different entity, such as a third party entity (e.g., third party entity 155 as shown in FIG. 1B). Thus, a third party can perform assays using one or more assay devices 205 and then transmit data generated from the assays to the condition analysis system 170 for performing screening and / or secondary analyses.

[0164] Assay The methods disclosed herein include performing an assay to generate marker information. The assays described in this section refer to assay 120A, assay 120B, or both assay 120A and assay 120B shown in FIGS. 1A-1C. With reference to FIG. 2A, performing an assay may include using one or more assay devices 205 to perform an assay. In various embodiments, the marker information refers to a quantitative value of a biomarker, such as a protein biomarker, a nucleic acid biomarker, or a metabolite biomarker. Thus, the quantitative value of a biomarker in a sample can be used to determine whether an individual is at risk for a health condition. In various embodiments, performing an assay to determine a quantitative value of a protein biomarker may include performing one or more of an immunoassay, a protein binding assay, an antibody-based assay, an antigen-binding protein-based assay, a protein-based array, an enzyme-linked immunosorbent assay (ELISA), or a Western blot. To determine a quantitative value of a nucleic acid biomarker, performing an assay may include performing one or more of a quantitative PCR (qPCR) or a digital PCR (dPCR). To determine the quantitative value of a metabolite, performing an assay may include performing NMR, mass spectrometry, LC-MS, or UPLC-MS / MS.

[0165] In various embodiments, the marker information refers to sequence information of a plurality of genomic sites. The sequence information can then be analyzed to generate a prediction for the individual (e.g., whether the individual is at risk for a health condition or whether the individual has a health condition). In certain embodiments, methylation sequence information is generated by performing an assay. The methylation sequence information includes the methylation status of a plurality of genomic sites. In various embodiments, the plurality of genomic sites are pre-identified and selected. For example, the plurality of genomic sites may be one or more CpG sites whose differential methylation provides information for determining whether the individual is at risk for a health condition. A CpG site is a portion of a genome that has a cytosine and a guanine separated by only one phosphate group, and is often written as "5'-C-phosphate-G-3'", or "CpG" for short. Regions with a high frequency of CpG sites are commonly referred to as "CG islands" or "CGIs". It has been found that certain CGIs and certain characteristics of certain CGIs in tumor cells tend to be different from the characteristics of the same CGIs or CGIs in healthy cells. In this specification, such CGIs and genomic features are referred to as "cancer informative CGIs."

[0166] Reference is made to Figure 2B, which shows an example of methylation information useful for determining whether an individual is at risk for a health condition, according to an embodiment. Specifically, Figure 2B shows that across various types of cancer (e.g., bladder, cervical, colon, endometrial, gastric, lung, ovarian, and prostate cancer), subregions within specific CGIs may exhibit differential methylation compared to normal plasma. Thus, Figure 2B shows an example of a cancer information CGI, where running an assay generates methylation sequence information corresponding to the cancer information CGI.

[0167] In various embodiments, performing an assay to generate sequence information of a plurality of genomic sites includes processing nucleic acid of a sample, enriching the processed nucleic acid for a preselected genomic sequence (e.g., a preselected information CGI), amplifying the genomic sequence to generate an amplification product, and quantifying the amplification product including the genomic sequence (e.g., via sequencing or via a quantitative method such as ELISA, quantitative PCR, or a DNA or RNA-based assay). In various embodiments, performing an assay to generate sequence information of a plurality of genomic sites includes a subset of the aforementioned steps. For example, enriching the processed nucleic acid may be omitted. Thus, performing an assay may include processing nucleic acid of a sample, amplifying a preselected genomic sequence, and quantifying the amplification product including the genomic sequence.

[0168] Referring again to any of FIGS. 1A-1C, in various embodiments, both assay 120A and assay 120B may include performing steps of processing nucleic acid of a sample, enriching the processed nucleic acid for a preselected genomic sequence (e.g., a preselected information CGI), amplifying the genomic sequence to generate an amplified product, and quantifying the amplified product including the genomic sequence. In some embodiments, assay 120A and assay 120B may differ. For example, assay 120A may exclude the step of enriching the nucleic acid and thus includes steps of processing the nucleic acid, amplifying the genomic sequence, and quantifying the amplified product. Assay 120B includes steps of processing the nucleic acid, enriching the genomic sequence, amplifying the genomic sequence, and quantifying the amplified product. In various embodiments, assay 120A includes quantifying the amplified product by performing an ELISA assay or performing quantitative PCR, while assay 120B includes quantifying the amplified product by performing next generation sequencing.

[0169] In various embodiments, performing an assay (e.g., assay 120A or assay 120B) includes processing nucleic acids (e.g., cfDNA fragments) from a sample (e.g., a liquid biopsy sample). In various embodiments, processing nucleic acids includes treating the nucleic acids to capture methylation modifications. In various embodiments, processing nucleic acids to capture methylation modifications includes performing deamination of cytosine residues. Other techniques include, but are not limited to, enzymatic methods. In various embodiments, processing nucleic acids to capture methylation modifications includes performing any of nucleic acid amplification, polymerase chain reaction (PCR), methylation specific PCR, bisulfite pyrosequencing, single-strand conformation polymorphism (SSCP) analysis, methylation sensitive single-strand conformation analysis limit analysis, high resolution melting analysis, methylation sensitive single nucleotide primer extension, limit analysis, microarray technology, next generation methylation sequencing, nanopore sequencing, and combinations thereof.

[0170] In various embodiments, performing deamination of cytosine residues is useful for determining the methylation state of nucleic acids from a sample. Performing deamination includes providing or exposing nucleic acids from a sample to a deaminating agent. In various embodiments, performing deamination of cytosine residues includes performing selective deamination. Selective deamination refers to a process in which cytosine residues are deaminated more selectively than 5-methylcytosine residues. Deamination of cytosine forms uracil, effectively inducing a C to T point mutation, allowing for the detection of methylated cytosine. Methods for deaminating cytosine are known in the art and include bisulfite conversion and enzymatic conversion. Bisulfite conversion allows for highly efficient conversion of unmethylated cytosine to uracil in DNA from samples such as whole blood or plasma, cultured cells, tissue samples, genomic DNA, and formalin-fixed paraffin-embedded (FFPE) tissue. Bisulfite conversion can be carried out using commercially available technologies such as Zymo Gold available from Zymo Research (Irvine, Calif.) or EpiTect Fast available from Qiagen (Germantown, Md.). In certain embodiments, enzymatic conversion involves subjecting the nucleic acid to TET2 to oxidize and thereby protect methylated cytosines, followed by exposure to APOBEC to convert unprotected (unmethylated) cytosines to uracil.

[0171] In various embodiments, performing the assay includes enriching specific genomic sequences, such as genomic sequences of preselected CGIs. In various embodiments, enrichment of preselected CGIs can be achieved by hybrid capture. Examples of such hybrid capture probe sets include the KAPA HyperPrep Kit and SeqCAP Epi Enrichment System from Roche Diagnostics (Pleasanton, Calif.). For example, hybrid capture probe sets can be designed to target (e.g., hybridize to) selected genomic sequences, thereby capturing and enriching the selected genomic sequences.

[0172] In various embodiments, performing the assay includes a step of nucleic acid amplification. Examples of such assays include, but are not limited to, performing a PCR assay, a real-time PCR assay, a quantitative real-time PCR (qPCR) assay, a digital PCR (dPCR), an allele-specific PCR assay, a reverse transcription PCR assay, and a reporter assay. For example, given a processed nucleic acid (e.g., a bisulfite-converted nucleic acid) enriched with a preselected genomic sequence, a PCR assay is performed to amplify the preselected genomic sequence to generate an amplification product. Here, a PCR primer is added to initiate the amplification. In various embodiments, the PCR primer is a whole genome primer that allows for whole genome amplification. In various embodiments, the PCR primer is a gene-specific primer that results in the amplification of a sequence of a specific gene. In various embodiments, the PCR primer is an allele-specific primer. For example, the allele-specific primer can target a genomic sequence corresponding to a preselected CGI, such that performing the nucleic acid amplification amplifies the genomic sequence of the preselected CGI.

[0173] In various embodiments, performing the assay includes quantifying a nucleic acid that comprises a preselected genomic sequence (e.g., an informative CGI). In some embodiments, quantifying the nucleic acid to generate sequence information includes performing an enzyme-linked immunosorbent assay (ELISA). In some embodiments, quantifying the nucleic acid to generate sequence information includes performing quantitative PCR (qPCR) or digital PCR (dPCR). Thus, the amount of methylated, unmethylated, or partially methylated preselected genomic sequences can be quantified.

[0174] In various embodiments, quantifying the nucleic acid includes sequencing the nucleic acid comprising the preselected genomic sequence. Thus, the sequenced reads can be aligned to a reference library to determine methylation sequence information, including the methylation state of the information CGI. Thus, the number of methylated, unmethylated, or partially methylated of the preselected genomic sequence can be quantified via the sequenced reads.

[0175] FIG. 2C shows an example of a flow process for determining whether an individual is at risk for a health condition, according to an embodiment, where a specific genomic region of an indexed library of nucleic acids (such as DNA) is targeted. For example, Locus 1 can refer to a reference genomic location, where the reference genomic location serves as a target. For example, the reference genomic location is not differentially methylated in healthy individuals compared to individuals with a health condition. Locus 2 can refer to a pre-selected genomic location, such as a pre-selected informative CGI.

[0176] Performing the assay further includes performing nucleic acid amplification (e.g., PCR) to generate marker information. In various embodiments, nucleic acid amplification includes either qPCR or dPCR, which quantifies the number of methylated, unmethylated, or partially methylated sequences at locus 1 (reference) and locus 2. In various embodiments, performing the assay includes performing an ELISA to quantitate the number of methylated, unmethylated, or partially methylated sequences at locus 1 (reference) and locus 2.

[0177] Assays for generating sequence information for performing intraindividual analyses In certain embodiments, the assays disclosed herein (e.g., assays 120A or 120B shown in Figures 1A-1C) are useful for generating sequence information for performing intra-individual analyses, e.g., assays are performed to generate sequence information for target and / or reference nucleic acids.

[0178] In various embodiments, the sequence information of the target nucleic acid and / or the sequence information of the reference nucleic acid refers to the state of multiple genomic sites. The sequence information of the target nucleic acid refers to the epigenetic state (e.g., methylation state) across multiple genomic sites in the target nucleic acid. The sequence information of the reference nucleic acid refers to the epigenetic state (e.g., methylation state) across multiple genomic sites in the reference nucleic acid. In various embodiments, the multiple genomic sites are pre-identified and selected. For example, the multiple genomic sites may be one or more CpG sites whose differential methylation provides information for determining whether an individual has a health condition. A CpG site is a portion of a genome that has a cytosine and a guanine separated by only one phosphate group, and is often written as "5'-C-phosphate-G-3'", or "CpG" for short. Regions with a high frequency of CpG sites are commonly referred to as "CG islands" or "CGIs". It has been found that certain CGIs and certain characteristics of certain CGIs in tumor cells tend to be different from the characteristics of the same CGI or CGI in healthy cells. Such CGIs and genomic features are referred to herein as "cancer informative CGIs." Cancer informative CGIs may be "CGI identifiers" or reference numbers that allow the CGIs to be referenced by their respective unique CGI identifiers during data processing. Examples of CGIs include, but are not limited to, the CGIs shown in the attached tables (hereinafter referred to as Tables 1-4) that list, for each CGI, its respective location in the human genome. Additional examples of CGIs are disclosed in WO2018209361 (see Table 1) and WO2022133315 (see Table 2 entitled "TOO methylation sites" and Table 3 entitled "Pan-cancer methylation sites"), each of which is incorporated herein by reference in its entirety. In some embodiments, the methylation status of multiple CpGs within a CGI may be analyzed. In some embodiments, at least some of the CpGs within a CGI may be analyzed. In other embodiments, all of the CpGs within a CGI may be analyzed. In some embodiments, analysis of CGIs contemplated herein may include analyzing CpGs within at least a portion of one or more regions in Tables 1-4.

[0179] In various embodiments, performing an assay to generate sequence information of a plurality of genomic sites includes processing the nucleic acid of the sample, enriching the processed nucleic acid for a preselected genomic sequence (e.g., a preselected information CGI), amplifying the genomic sequence to generate an amplified product, and quantifying the amplified product including the genomic sequence (e.g., via sequencing, such as next generation sequencing, or via a quantitative method, such as ELISA, quantitative PCR, allele-specific PCR, or a DNA or RNA-based assay). In various embodiments, performing an assay to generate sequence information of a plurality of genomic sites includes a subset of the aforementioned steps. For example, enriching the processed nucleic acid may be omitted. Thus, performing an assay may include processing the nucleic acid of the sample, amplifying the preselected genomic sequence, and quantifying the amplified product including the genomic sequence.

[0180] In various embodiments, performing an assay (e.g., assay 120A or assay 120B) includes treating the target nucleic acid and / or the reference nucleic acid. In various embodiments, treating the target nucleic acid and / or the reference nucleic acid to capture methylation modifications includes performing bisulfite conversion. Bisulfite conversion allows highly efficient conversion of unmethylated cytosines in DNA to uracils from samples such as whole blood or plasma, cultured cells, tissue samples, genomic DNA, and formalin-fixed paraffin-embedded (FFPE) tissues. Bisulfite conversion can be performed using commercially available technologies such as Zymo Gold available from Zymo Research (Irvine, CA) or EpiTect Fast available from Qiagen (Germantown, MD). Other techniques include, but are not limited to, enzymatic methods. In various embodiments, processing the target nucleic acid and / or the reference nucleic acid to capture methylation modifications includes performing any of nucleic acid amplification, polymerase chain reaction (PCR), methylation-specific PCR, bisulfite pyrosequencing, single-stranded conformation polymorphism (SSCP) analysis, methylation-sensitive single-stranded conformation analysis restriction analysis, high-resolution melting analysis, methylation-sensitive single-base primer extension, restriction analysis, microarray technology, next-generation methylation sequencing, nanopore sequencing, and combinations thereof.

[0181] In various embodiments, performing the assay includes enriching a specific sequence in the target nucleic acid and / or the reference nucleic acid. In various embodiments, the specific sequence refers to a sequence of a preselected CGI. In various embodiments, enrichment of the preselected CGI can be achieved by hybrid capture. Examples of such hybrid capture probe sets include the KAPA HyperPrep Kit and SeqCAP Epi Enrichment System from Roche Diagnostics (Pleasanton, CA). For example, a hybrid capture probe set can be designed to hybridize with a specific sequence of the target nucleic acid and / or the reference nucleic acid, thereby capturing and enriching the specific sequence.

[0182] In various embodiments, performing the assay includes performing a nucleic acid amplification to amplify a specific sequence of the target nucleic acid and / or the reference nucleic acid. Examples of such assays include, but are not limited to, performing a PCR assay, a real-time PCR assay, a quantitative real-time PCR (qPCR) assay, a digital PCR (dPCR), an allele-specific PCR assay, a reverse transcription PCR assay, and a reporter assay. For example, given a processed nucleic acid (e.g., a bisulfite converted nucleic acid) enriched for a preselected sequence, a PCR assay is performed to amplify the preselected sequence to generate an amplification product. Here, a PCR primer is added to initiate the amplification. In various embodiments, the PCR primer is a whole genome primer that allows for whole genome amplification. In various embodiments, the PCR primer is a gene specific primer that results in the amplification of a sequence of a specific gene. In various embodiments, the PCR primer is an allele specific primer. For example, the allele specific primer can target a genomic sequence corresponding to a preselected CGI, such that performing the nucleic acid amplification amplifies the sequence of the preselected CGI.

[0183] In various embodiments, performing the assay includes quantifying a nucleic acid that includes a preselected sequence (e.g., an informational CGI). In some embodiments, quantifying the nucleic acid to generate sequence information includes performing a real-time PCR assay, a quantitative real-time PCR (qPCR) assay, a digital PCR (dPCR) assay, an allele-specific PCR assay, or a reverse transcription PCR assay. Thus, the number of methylated, hypermethylated, unmethylated, or partially methylated preselected sequences is quantified.

[0184] In various embodiments, quantifying the nucleic acid comprises sequencing the nucleic acid comprising the preselected sequence. Thus, the sequenced reads are aligned to a reference library, and sequence information, including the methylation state of the information CGI of the amplification product obtained from the target nucleic acid and / or the reference nucleic acid, is determined. Thus, the number of methylated, hypermethylated, unmethylated, or partially methylated preselected sequences of the target nucleic acid and the reference nucleic acid can be quantified through the sequenced reads.

[0185] screening The description in this section relates to performing a screen, such as screen 125 of FIG. 1A, which may be performed by screening module 210 of FIG. 2A. Generally, a screen is performed on marker information generated by an assay (e.g., assay 120A). In various embodiments, a screen is performed to determine whether a biological sample is at risk for containing a signal indicative of a health condition. For example, a screen is performed to determine whether a biological sample is at risk for containing circulating tumor DNA. Circulating DNA in a biological sample may indicate that an individual (e.g., the individual from whom the biological sample was obtained) may be at risk for a health condition, such as cancer. In various embodiments, a screen is performed to classify an individual as at risk for having a health condition, or not at risk for having a health condition.

[0186] In various embodiments, the marker information represents a quantitative value of the biomarker. For example, depending on the type of biomarker, the quantitative value may be generated by one or more of an immunoassay, a protein-binding assay, an antibody-based assay, an antigen-binding protein-based assay, a protein-based array, an enzyme-linked immunosorbent assay (ELISA), a Western blot, quantitative PCR (qPCR) or digital PCR (dPCR), NMR, mass spectrometry, LC-MS, or UPLC-MS / MS.

[0187] In various embodiments, performing the screen includes comparing the quantitative value of the biomarker to one or more reference values ​​or thresholds. For example, the reference value can be a statistical measure of the quantified biomarker value corresponding to an individual known to be at risk for the health condition. Thus, if the comparison identifies that the quantified value of the biomarker for the individual is statistically significantly different from the reference value corresponding to an individual known to be at risk for the health condition, the screen can identify that the individual is not at risk for the health condition.

[0188] In various embodiments, the marker information represents sequence information of one or more genomic locations, such as one or more CpG islands. In various embodiments, performing the screening includes comparing methylation information at one or more preselected genomic locations to a quantitative value of a reference genomic location. For example, referring back to FIG. 2C, an assay may be performed that generates methylation information for locus 1, which corresponds to the reference genomic location, and locus 2, which corresponds to a preselected genomic location (e.g., a preselected information CGI). Thus, the methylation information for locus 1 is compared to the methylation information for locus 2. Based on the comparison, the screening can identify whether the individual is at risk for a health condition or not at risk for a health condition.

[0189] As an example, the methylation information of one or more preselected genomic locations and the methylation information of the reference genomic location can be a cycle threshold (Ct) value. The cycle threshold refers to the number of PCR cycles required for a sample to be amplified and exceed a threshold. In various embodiments, if the difference between the Ct value of the methylation sequence of the preselected genomic location and the Ct value of the reference genomic location is greater than the threshold, the screening identifies the individual as at risk for the health condition. If the difference between the Ct value of the methylation sequence of the preselected genomic location and the Ct value of the reference genomic location is less than the threshold, the screening identifies the individual as not at risk for the health condition.

[0190] In various embodiments, the screening is performed on sequence information generated by sequencing (e.g., next-generation sequencing) of sequences at one or more genomic locations, such as one or more CpG islands. In various embodiments, such screening is performed using a system including a computer storage and a processing system. The screening may further include implementation of a machine learning model. For example, the computer storage may store sequence information corresponding to a processed sample, the processed sample including cell-free DNA fragments derived from a liquid biopsy of the individual and processed to enrich for cancer-informative CGIs, and the sequencer information includes, for each sequenced cell-free DNA fragment corresponding to the cancer-informative CGI, the respective location on the genome of the cell-free DNA fragment and methylation information of the cell-free DNA fragment. The processing system may calculate the value of the cancer-informative CGI for the individual and apply the value as an input to the trained machine learning model. The machine learning model provides a predictive output regarding whether the individual is at risk for a health condition based on the value of the cancer-informative CGI.

[0191] In various embodiments, performing the screening includes analyzing a plurality of CGIs. For example, performing the screening includes analyzing the methylation status of a plurality of CGIs. The cancer informative CGIs may be "CGI identifiers" or reference numbers that allow the CGIs to be referenced by their respective unique CGI identifiers during data processing. The attached tables (e.g., Tables 1-4) list for each CGI its respective location in the human genome. Additional examples of CGIs are disclosed in WO2018209361 (see Table 1) and WO2022133315 (see Table 2 entitled "TOO methylation sites" and Table 3 entitled "Pan-cancer methylation sites"), each of which is incorporated herein by reference in its entirety. In some embodiments, the methylation status of a plurality of CpGs within a CGI may be analyzed. In some embodiments, at least some of the CpGs within a CGI may be analyzed. In other embodiments, all of the CpGs within a CGI may be analyzed. In some embodiments, analysis of CGIs contemplated herein may include analyzing CpGs within at least a portion of one or more regions in Tables 1-4.

[0192] In various embodiments, performing the screening includes analyzing all of the CGIs in any one of Tables 1, 2, 3, or 4. In various embodiments, performing the screening includes analyzing up to 10% of the CGIs in Table 1. In various embodiments, performing the screening includes analyzing up to 10%, up to 20%, up to 30%, up to 40%, up to 50%, up to 55%, up to 60%, up to 65%, up to 70%, up to 75%, up to 80%, up to 85%, up to 90%, up to 91%, up to 92%, up to 93%, up to 94%, up to 95%, up to 96%, up to 97%, up to 98%, or up to 99% of the CGIs in Table 1. In various embodiments, performing the screening includes analyzing up to 10% of the CGIs in Table 2. In various embodiments, performing the screening includes analyzing up to 10%, up to 20%, up to 30%, up to 40%, up to 50%, up to 55%, up to 60%, up to 65%, up to 70%, up to 75%, up to 80%, up to 85%, up to 90%, up to 91%, up to 92%, up to 93%, up to 94%, up to 95%, up to 96%, up to 97%, up to 98%, or up to 99% of the CGIs in Table 2. In various embodiments, performing the screening includes analyzing up to 10% of the CGIs in Table 3. In various embodiments, performing the screening includes analyzing up to 10%, up to 20%, up to 30%, up to 40%, up to 50%, up to 55%, up to 60%, up to 65%, up to 70%, up to 75%, up to 80%, up to 85%, up to 90%, up to 91%, up to 92%, up to 93%, up to 94%, up to 95%, up to 96%, up to 97%, up to 98%, or up to 99% of the CGIs in Table 3. In various embodiments, performing the screening includes analyzing up to 10% of the CGIs in Table 4. In various embodiments, performing the screening includes analyzing up to 10%, up to 20%, up to 30%, up to 40%, up to 50%, up to 55%, up to 60%, up to 65%, up to 70%, up to 75%, up to 80%, up to 85%, up to 90%, up to 91%, up to 92%, up to 93%, up to 94%, up to 95%, up to 96%, up to 97%, up to 98%, or up to 99% of the CGIs in Table 4.In various embodiments, performing the screening includes analyzing up to 10% of the CGIs in Tables 2 and 3. In various embodiments, performing the screening includes analyzing up to 10%, up to 20%, up to 30%, up to 40%, up to 50%, up to 55%, up to 60%, up to 65%, up to 70%, up to 75%, up to 80%, up to 85%, up to 90%, up to 91%, up to 92%, up to 93%, up to 94%, up to 95%, up to 96%, up to 97%, up to 98%, or up to 99% of the CGIs in Tables 2 and 3.

[0193] In various embodiments, performing the screening includes screening 1 CGI, 2 CGI, 3 CGI, 4 CGI, 5 CGI, 6 CGI, 7 CGI, 8 CGI, 9 CGI, 10 CGI, 11 CGI, 12 CGI, 13 CGI, 14 CGI, 15 CGI, 16 CGI, 17 CGI, 18 CGI, 19 CGI, 20 CGI, 21 CGI, 22 CGI, 23 CGI, 24 CGI, 25 CGI, 26 CGI, 27 CGI, 28 CGI, 29 CGI, 30 CGI, 31 CGI, 32 CGI, 33 CGI, 34 CGI, 35 CGI, 36 CGI, 37 CGI, 38 CGI, 39 CGI, 40 CGI, 41 CGI, 42 CGI, 43 CGI, 44 CGI, 45 CGI, 46 CGI, 47 CGI, 48 CGI, 49 CGI, 50 CGI, 51 CGI, 52 CGI, 53 CGI, 54 CGI, 55 CGI, 56 CGI, 57 CGI, 58 CGI, 59 CGI, 60 CGI, 61 CGI, 62 CGI, 63 CGI, 64 CGI, 65 CGI, 66 CGI, 67 CGI, 68 CGI, 69 CGI, 70 CGI, 71 CGI, 72 CGI, 73 CGI, 74 CGI, 75 CGI, 76 CGI, 77 CGI, 78 CGI, 79 CGI, 80 CGI, 81 CGI, 82 CGI, 83 CGI, 84 CGI, 85 CGI, The method includes analyzing 1 CGI, 29 CGI, 30 CGI, 31 CGI, 32 CGI, 33 CGI, 34 CGI, 35 CGI, 36 CGI, 37 CGI, 38 CGI, 39 CGI, 40 CGI, 41 CGI, 42 CGI, 43 CGI, 44 CGI, 45 CGI, 46 CGI, 47 CGI, 48 CGI, 49 CGI, or 50 CGI (e.g., a CGI shown in any of Tables 1-4, or a portion of a CGI shown in any of Tables 1-4). In various embodiments, performing the screening includes analyzing up to 2 CGIs, up to 5 CGIs, up to 10 CGIs, up to 15 CGIs, up to 20 CGIs, up to 25 CGIs, up to 30 CGIs, up to 35 CGIs, up to 40 CGIs, up to 45 CGIs, or up to 50 CGIs (e.g., the CGIs shown in any of Tables 1-4, or a portion of the CGIs shown in any of Tables 1-4).In various embodiments, performing the screening includes analyzing up to 50 CGIs, up to 100 CGIs, up to 150 CGIs, up to 200 CGIs, up to 300 CGIs, up to 400 CGIs, up to 500 CGIs, up to 600 CGIs, up to 700 CGIs, up to 800 CGIs, up to 900 CGIs, up to 1000 CGIs, up to 1500 CGIs, up to 2000 CGIs, up to 2500 CGIs, up to 3000 CGIs, up to 3500 CGIs, up to 4000 CGIs, up to 4500 CGIs, up to 5000 CGIs, up to 5500 CGIs, or up to 6000 CGIs (e.g., the CGIs shown in any of Tables 1-4, or a portion of the CGIs shown in any of Tables 1-4). In certain embodiments, performing the screening includes analyzing up to 500 CGIs.

[0194] In various embodiments, the screening achieves a sensitivity of at least 60% in detecting the presence of a health condition. In various embodiments, the screening achieves a sensitivity of at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least In certain embodiments, the screening achieves a sensitivity of at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9%. In certain embodiments, the screening achieves a sensitivity of at least 75%. In certain embodiments, the screening achieves a sensitivity of at least 76%. In certain embodiments, the screening achieves a sensitivity of at least 77%. In certain embodiments, the screening achieves a sensitivity of at least 78%. In certain embodiments, the screening achieves a sensitivity of at least 79%. In certain embodiments, the screening achieves a sensitivity of at least 80%.

[0195] In various embodiments, the screen achieves a specificity of at least 60% in excluding individuals without the health condition. In various embodiments, the screen achieves a specificity of at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least In certain embodiments, the screening achieves a specificity of at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9%. In certain embodiments, the screening achieves a specificity of at least 90%. In certain embodiments, the screening achieves a specificity of at least 91%. In certain embodiments, the screening achieves a specificity of at least 92%. In certain embodiments, the screening achieves a specificity of at least 93%. In certain embodiments, the screening achieves a specificity of at least 94%. In certain embodiments, the screening achieves a specificity of at least 95%.

[0196] In various embodiments, the screen achieves a positive predictive value of at least 15%. In various embodiments, the screen achieves a positive predictive value of at least 15%, at least 16%, at least 17%, at least 18%, at least 19%, at least 20%, at least 21%, at least 22%, at least 23%, at least 24%, at least 25%, at least 26%, at least 27%, at least 28%, at least 29%, at least 30%, at least 31%, at least 32%, at least 33%, at least 34%, at least 35%, at least 36%, at least 37%, at least 38%, at least 39%, or at least 40%. In certain embodiments, the screen achieves a positive predictive value of at least 20%. In certain embodiments, the screen achieves a positive predictive value of at least 21%. In certain embodiments, the screen achieves a positive predictive value of at least 22%. In certain embodiments, the screen achieves a positive predictive value of at least 23%. In certain embodiments, the screen achieves a positive predictive value of at least 24%. In certain embodiments, the screen achieves a positive predictive value of at least 25%. In certain embodiments, the screen achieves a positive predictive value of at least 26%. In certain embodiments, the screen achieves a positive predictive value of at least 27%. In certain embodiments, the screen achieves a positive predictive value of at least 28%. In certain embodiments, the screen achieves a positive predictive value of at least 29%. In certain embodiments, the screen achieves a positive predictive value of at least 30%. In certain embodiments, the screen achieves a positive predictive value of at least 31%. In certain embodiments, the screen achieves a positive predictive value of at least 32%. In certain embodiments, the screen achieves a positive predictive value of at least 33%. In certain embodiments, the screen achieves a positive predictive value of at least 34%. In certain embodiments, the screen achieves a positive predictive value of at least 35%.In certain embodiments, the screen achieves a positive predictive value of at least 36%. In certain embodiments, the screen achieves a positive predictive value of at least 37%. In certain embodiments, the screen achieves a positive predictive value of at least 38%. In certain embodiments, the screen achieves a positive predictive value of at least 39%. In certain embodiments, the screen achieves a positive predictive value of at least 40%.

[0197] In various embodiments, the screen achieves a negative predictive value of at least 60%. In various embodiments, the screen achieves a negative predictive value of at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least The screening achieves a negative predictive value of 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9%. In certain embodiments, the screening achieves a negative predictive value of at least 95%. In certain embodiments, the screening achieves a negative predictive value of at least 96%. In certain embodiments, the screening achieves a negative predictive value of at least 97%. In certain embodiments, the screening achieves a negative predictive value of at least 98%. In certain embodiments, the screening achieves a negative predictive value of at least 99%.

[0198] Intra-individual analysis The description in this section relates to performing an intra-individual analysis, such as the intra-individual analysis 128 described in FIG. 1, which may be performed by the state analysis system 1709 (more specifically, the signal generation module 215) described in FIG. 2A. In general, an intra-individual analysis is performed on sequence information of a target nucleic acid and sequence information of a reference nucleic acid. As described herein, the sequence information of the target nucleic acid and the sequence information of the reference nucleic acid are generated by performing one or more assays (e.g., assay 120A and / or assay 120B). In certain embodiments, the sequence information of the target nucleic acid includes sequence information of cell-free DNA. In certain embodiments, the sequence information of the reference nucleic acid includes sequence information of cells, such as peripheral blood mononuclear cells (PBMCs) or polymorphonuclear cells.

[0199] Intra-individual analysis involves combining sequence information of a target nucleic acid with sequence information of a reference nucleic acid to generate an informative signal for determining the presence or absence of a health condition, where the step of combining sequence information of a target nucleic acid with sequence information of a reference nucleic acid can be performed by signal generation module 210 shown in FIG.

[0200] In various embodiments, combining the sequence information of the target nucleic acid and the sequence information of the reference nucleic acid includes distinguishing between a signature that exists or does not exist in the sequence information of the target nucleic acid and a signature that exists or does not exist in the sequence information of the reference nucleic acid.For example, if a certain signature exists in the sequence information of the target nucleic acid and also exists in the sequence information of the reference nucleic acid, the signatures of both the target nucleic acid and the reference nucleic acid may represent baseline biological signatures.Therefore, these signatures may be excluded from the generated signal that provides information for determining the presence or absence of a health condition.As another example, if a certain signature exists in the sequence information of the target nucleic acid, but these signatures do not exist in the sequence information of the reference nucleic acid, the signature may not be a baseline biological signature.Therefore, these signatures may be included in the generated signal that provides information for determining the presence or absence of a health condition.

[0201] In various embodiments, combining the sequence information of the target nucleic acid with the sequence information of the reference nucleic acid includes aligning the sequence information of the target nucleic acid with the sequence information of the reference nucleic acid, for example, aligning the sequence information includes aligning the sequence of a plurality of preselected genomic sites of the target nucleic acid with the sequence of the same or overlapping plurality of preselected genomic sites of the reference nucleic acid.

[0202] In various embodiments, both the sequence information of the target nucleic acid and the sequence information of the reference nucleic acid are aligned to a reference genome library (e.g., a reference assembly) having a known sequence. Thus, the sequence information of the target nucleic acid is aligned to the sequence information of the reference nucleic acid via a reference genome library. In various embodiments, the sequence information of the target nucleic acid is directly aligned to the sequence information of the reference nucleic acid. In such an embodiment, a reference genome library may not be used.

[0203] In various embodiments, combining the sequence information of the target nucleic acid and the sequence information of the reference nucleic acid includes determining the differences between the sequence information of the target nucleic acid and the sequence information of the reference nucleic acid.

[0204] In various embodiments, the difference between the sequence information of the target nucleic acid and the sequence information of the reference nucleic acid is performed position by position. For example, at a first position of a genome site, the difference between the sequence information of the target nucleic acid at the first position and the sequence information of the reference nucleic acid at the same first position is determined. The process can then be further repeated for additional positions (e.g., additional positions across multiple genome sites). In various embodiments, if the sequence information of the target nucleic acid and the reference nucleic acid is generated using a sequencing assay (e.g., next-generation sequencing) that provides base-level resolution of the sequence, the difference is determined position by position.

[0205] In various embodiments, the difference between the sequence information of the target nucleic acid and the sequence information of the reference nucleic acid is carried out for each CGI. For example, in the first CGI of a genome site, the difference between the sequence information of the target nucleic acid in the first CGI and the sequence information of the reference nucleic acid in the same CGI or overlapping part of the first CGI is determined. Then, the process can be further repeated for additional CGIs (e.g., additional CGIs across multiple genome sites). In various embodiments, if the sequence information of the target nucleic acid and the reference nucleic acid is generated using a quantitative assay (e.g., qPCR assay), the difference is determined for each CGI.

[0206] In various embodiments, the difference between the sequence information of the target nucleic acid and the sequence information of the reference nucleic acid is performed on an allele basis. For example, at a first allele of a genome site, the difference between the sequence information of the target nucleic acid at the first allele and the sequence information of the reference nucleic acid at the same allele or overlapping portion of the first allele is determined. Then, the process can be further repeated for additional alleles (e.g., additional alleles across multiple genome sites). In various embodiments, if the sequence information of the target nucleic acid and the reference nucleic acid is generated using a quantitative assay (e.g., a qPCR assay or an allele-specific assay), the difference is determined on an allele basis.

[0207] Reference is now made to FIG. 2D, which shows an example of combining sequence information of a target nucleic acid and a reference nucleic acid to generate an informative signal of a health condition according to an embodiment. The sequence information of the target nucleic acid and the sequence information of the reference nucleic acid include methylation status across multiple genomic sites. FIG. 2D shows an example of a genomic site where a nucleotide base may be differentially methylated in the target nucleic acid and the reference nucleic acid. For example, as shown in FIG. 2D, in both the target nucleic acid and the reference nucleic acid, the nucleotide base at the second position is methylated (represented by the presence of a cytosine base resulting after bisulfite conversion). Considering that the methylation at the second position occurs in both the target nucleic acid and the reference nucleic acid, this may be a baseline biological signature. Conversely, the target nucleic acid may be further methylated at the sixth and ninth positions, while the reference nucleic acid is unmethylated at the sixth and ninth positions. Here, since the reference nucleic acid is unmethylated at the sixth and ninth positions, the presence of a methylated nucleotide base in the target nucleic acid may represent an informative signature of the presence or absence of a health condition. Furthermore, at the eleventh nucleotide position, the target nucleic acid is unmethylated, whereas the reference nucleic acid is methylated, where the methylation of the reference nucleic acid can be interpreted as a baseline biological signature.

[0208] The difference in methylation state at each position between the target nucleic acid and the reference nucleic acid can represent a cancer signal. As shown in Figure 2D, the cancer signal includes the methylation state of a genomic site, where the 6th and 9th positions are methylated. Thus, the cancer signal includes signatures from the target nucleic acid that may be informative of health status (e.g., the methylation state of the 6th and 9th nucleotide bases), and excludes baseline biological signatures (e.g., baseline biological signatures present in the reference nucleic acid, such as the methylation state of the 2nd and 11th nucleotide bases).

[0209] The intra-individual analysis may further include analyzing a signal representing a combination of sequence information of the target nucleic acid and sequence information of the reference nucleic acid to determine whether a health condition is present or absent in the individual. Here, the step of analyzing the signal to determine the presence or absence of a health condition can be performed by the signal generation module 215 shown in FIG. 2A. In various embodiments, a machine learning model is introduced to analyze the signal that provides information for determining the presence or absence of a health condition. The machine learning model analyzes the signal that represents the difference between the epigenetic state (e.g., methylation state) of the multiple genomic sites of the target nucleic acid and the epigenetic state (e.g., methylation state) of the multiple genomic sites of the reference nucleic acid. Thus, the trained machine learning model analyzes the signal across the multiple genomic sites and outputs a prediction as to whether the individual has the presence or absence of a health condition.

[0210] In certain embodiments, the machine learning model analyzes the methylation status of a plurality of genomic sites in the cell-free DNA to generate a prediction. The methylation status corresponds to a set of cancer informative CpG islands (CGIs), where the cancer informative CGIs are selected from the group consisting of the set of ranked candidate CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 50 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 100 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 150 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 200 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 250 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 300 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 400 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 500 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 600 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 700 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 800 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 900 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 1000 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 2500 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 5000 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 7500 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 10000 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 15000 CGIs.In various embodiments, the machine learning model analyzes the methylation status of at least 20,000 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 25,000 CGIs.

[0211] In various embodiments, the machine learning model analyzes the methylation status of CGIs across the entire genome. For example, the machine learning model can be implemented to analyze sequencing data generated from whole genome sequencing (e.g., whole genome bisulfite sequencing).

[0212] In certain embodiments, intra-individual analysis further reveals the tissue of origin of the health condition for individuals predicted to have the presence of the health condition. Intra-individual analysis may identify the tissue of origin of the health condition according to the methylation status of the cancer-informative CGIs. For example, a particular methylation pattern across the cancer-informative CGIs is attributed to a particular tissue, examples of which include neural tissue (e.g., brain, spinal cord, nerves), muscle tissue (cardiac muscle, smooth muscle, skeletal muscle), epithelial tissue (e.g., the lining of the digestive tract, skin), and connective tissue (e.g., fat, bone, tendons, and ligaments). As a specific example, a first set of CGIs may be frequently methylated in patients with brain cancer. Thus, if a similar methylation pattern is observed across the first set of CGIs for an individual, intra-individual analysis may identify that the individual has cancer, and further, that the cancer is localized to the brain.

[0213] Second Analysis The description in this section relates to performing a second analysis, such as second analysis 130 in FIG. 1A, which may be performed by status analysis module 220 in FIG. 2A. Generally, the second analysis is performed on sequence information generated by an assay (e.g., assay 120A or assay 120B). In various embodiments, the second analysis is performed to determine whether a biological sample obtained from an individual contains a signal indicative of a health status. For example, screening is performed to determine whether the biological sample contains circulating tumor DNA. Circulating DNA in the biological sample may indicate that an individual (e.g., the individual from whom the biological sample was obtained) has a health status, such as cancer. In various embodiments, the second analysis is performed to classify an individual as having a health status (e.g., cancer) or not having a health status (e.g., cancer).

[0214] In various embodiments, the second analysis is performed on sequence information generated by sequencing (e.g., next generation sequencing) of sequences at one or more genomic locations, such as one or more CpG islands. In various embodiments, the sequence information is generated as a result of whole genome sequencing, and thus the second analysis is performed on sequences at one or more genomic locations across the entire genome.

[0215] In various embodiments, the second analysis is performed using a system including a computer storage and a processing system. The second analysis may include the implementation of a trained machine learning model, the details of which are described in more detail herein. For example, the computer storage can store sequence information corresponding to a processed sample, the processed sample includes cell-free DNA fragments derived from an individual's liquid biopsy and processed to enrich for cancer-informative CGIs, and the sequencer information includes, for each sequenced cell-free DNA fragment corresponding to the cancer-informative CGI, the respective positions of the cell-free DNA fragments on the genome and the methylation information of the cell-free DNA fragments.

[0216] In certain embodiments, the second analysis further reveals the tissue of origin of the health condition for the individual determined to have the health condition. The second analysis may identify the tissue of origin of the health condition according to the methylation status of the cancer-informative CGIs. For example, a particular methylation pattern across the cancer-informative CGIs is attributed to a particular tissue, examples of which include neural tissue (e.g., brain, spinal cord, nerves), muscle tissue (cardiac muscle, smooth muscle, skeletal muscle), epithelial tissue (e.g., the lining of the digestive tract, skin), and connective tissue (e.g., fat, bone, tendons, and ligaments). As a specific example, the first set of CGIs may be frequently methylated in patients with brain cancer. Thus, if a similar methylation pattern is observed across the first set of CGIs for the individual under analysis, the second analysis may identify that the individual has cancer, and further, that the cancer is localized to the brain.

[0217] In various embodiments, the second analysis includes analyzing a plurality of CGIs. For example, the second analysis includes analyzing the methylation status of a plurality of CGIs. The cancer informative CGIs may be "CGI identifiers" or reference numbers that allow the CGIs to be referenced by their respective unique CGI identifiers during data processing. The attached tables (e.g., Tables 1-4) list for each CGI its respective location in the human genome. Additional examples of CGIs are disclosed in WO2018209361 (see Table 1) and WO2022133315 (see Table 2 entitled "TOO methylation sites" and Table 3 entitled "Pan-cancer methylation sites"), each of which is incorporated herein by reference in its entirety. In various embodiments, the second analysis includes analyzing all of the CGIs in any one of Tables 1, 2, 3, or 4. In various embodiments, the second analysis includes analyzing at least 10% of the CGIs in Table 1. In various embodiments, the second analysis includes analyzing at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the CGIs in Table 1. In various embodiments, the second analysis includes analyzing at least 10% of the CGIs in Table 2. In various embodiments, the second analysis includes analyzing at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the CGIs in Table 2.In various embodiments, the second analysis includes analyzing at least 10% of the CGIs in Table 3. In various embodiments, the second analysis includes analyzing at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the CGIs in Table 3. In various embodiments, the second analysis includes analyzing at least 10% of the CGIs in Table 4. In various embodiments, the second analysis includes analyzing at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the CGIs in Table 4. In various embodiments, the second analysis includes analyzing at least 10% of the CGIs in Tables 2 and 3. In various embodiments, the second analysis includes analyzing at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the CGIs in Tables 2 and 3.

[0218] In various embodiments, the second analysis includes analyzing 100 CGIs (e.g., the CGIs shown in any of Tables 1-4). In various embodiments, the second analysis includes analyzing at least 100 CGIs, at least 150 CGIs, at least 200 CGIs, at least 300 CGIs, at least 400 CGIs, at least 500 CGIs, at least 600 CGIs, at least 700 CGIs, at least 800 CGIs, at least 900 CGIs, at least 1000 CGIs, at least 1500 CGIs, at least 2000 CGIs, at least 2500 CGIs, at least 3000 CGIs, at least 3500 CGIs, at least 4000 CGIs, at least 4500 CGIs, at least 5000 CGIs, at least 5500 CGIs, or at least 6000 CGIs (e.g., the CGIs shown in any of Tables 1-4). In certain embodiments, performing the screening includes analyzing at least 500 CGIs. In some embodiments, the methylation status of multiple CpGs within a CGI may be analyzed. In some embodiments, at least a portion of the CpGs within a CGI may be analyzed. In other embodiments, all CpGs within a CGI may be analyzed. In some embodiments, the analysis of CGIs contemplated herein may include analyzing CpGs within at least a portion of one or more regions in Tables 1-4.

[0219] In various embodiments, the second analysis comprises analyzing more CGIs compared to the amount of CGIs analyzed during screening.For example, the CGIs analyzed during screening may represent a subset of the CGIs analyzed during the second analysis.In some scenarios, all CpG islands analyzed during screening are further analyzed when performing the second analysis.Therefore, the second analysis represents a more robust and rigorous analysis compared to screening, which is more rapid and cost-effective.In various embodiments, the second analysis comprises analyzing at least twice the number of CGIs analyzed during screening. In various embodiments, the second analysis includes analyzing at least 3 times, at least 4 times, at least 5 times, at least 6 times, at least 7 times, at least 8 times, at least 9 times, at least 10 times, at least 11 times, at least 12 times, at least 13 times, at least 14 times at least 15 times, at least 16 times, at least 17 times, at least 18 times, at least 19 times, at least 20 times, at least 21 times, at least 22 times, at least 23 times, at least 24 times, at least 25 times, at least 26 times, at least 27 times, at least 28 times, at least 29 times, at least 30 times, at least 31 times, at least 32 times, at least 33 times, at least 34 times, at least 35 times, at least 36 times, at least 37 times, at least 38 times, at least 39 times, or at least 40 times the number of CGIs analyzed during screening. In certain embodiments, the second analysis includes analyzing at least 5 times the number of CGIs analyzed during screening. For example, a screening may include analyzing at least 100 CGIs and a second screening may include analyzing at least 500 CGIs.

[0220] In various embodiments, the second analysis achieves a sensitivity of at least 60% in detecting the presence of the health condition. In various embodiments, the screening achieves a sensitivity of at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119%, at least 120%, at least 122%, at least 124%, at least 126%, at least 128%, at least 129%, at least 130%, at least 131%, at least 132%, at least 133%, at least 134%, at least 135%, at least In certain embodiments, the second analysis achieves a sensitivity of at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9%. In certain embodiments, the second analysis achieves a sensitivity of at least 85%. In certain embodiments, the second analysis achieves a sensitivity of at least 86%. In certain embodiments, the second analysis achieves a sensitivity of at least 87%. In certain embodiments, the second analysis achieves a sensitivity of at least 88%. In certain embodiments, the second analysis achieves a sensitivity of at least 89%. In certain embodiments, the second analysis achieves a sensitivity of at least 90%.

[0221] In various embodiments, the second analysis achieves a specificity of at least 60% in excluding individuals without the health condition. In various embodiments, the second analysis achieves a specificity of at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least The second analysis achieves a specificity of 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9%. In certain embodiments, the second analysis achieves a specificity of at least 90%. In certain embodiments, the second analysis achieves a specificity of at least 91%. In certain embodiments, the second analysis achieves a specificity of at least 92%. In certain embodiments, the second analysis achieves a specificity of at least 93%. In certain embodiments, the second analysis achieves a specificity of at least 94%. In certain embodiments, the second analysis achieves a specificity of at least 95%.

[0222] In various embodiments, the second analysis achieves a positive predictive value of at least 60%. In various embodiments, the second analysis achieves a positive predictive value of at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119%, at least 120%, at least 122%, at least 124%, at least 125%, at least 126%, at least 127%, at least 128%, at least 129%, at least 130%, at least 131%, at least 132%, at least 133%, at least 134%, at least 135%, at least In certain embodiments, the second analysis achieves a positive predictive value of at least 5%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9%. In certain embodiments, the second analysis achieves a positive predictive value of at least 80%. In certain embodiments, the second analysis achieves a positive predictive value of at least 81%. In certain embodiments, the second analysis achieves a positive predictive value of at least 82%. In certain embodiments, the second analysis achieves a positive predictive value of at least 83%. In certain embodiments, the second analysis achieves a positive predictive value of at least 84%. In certain embodiments, the second analysis achieves a positive predictive value of at least 85%.

[0223] In various embodiments, the second analysis achieves a negative predictive value of at least 60%. In various embodiments, the second analysis achieves a negative predictive value of at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119%, at least 120%, at least 122%, at least 124%, at least 125%, at least 126%, at least 127%, at least 128%, at least 129%, at least 130%, at least 131%, at least 132%, at least 133%, at least 134%, at least In certain embodiments, the second analysis achieves a negative predictive value of at least 5%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9%. In certain embodiments, the second analysis achieves a negative predictive value of at least 90%. In certain embodiments, the second analysis achieves a negative predictive value of at least 91%. In certain embodiments, the second analysis achieves a negative predictive value of at least 92%. In certain embodiments, the second analysis achieves a negative predictive value of at least 93%. In certain embodiments, the second analysis achieves a negative predictive value of at least 94%. In certain embodiments, the second analysis achieves a negative predictive value of at least 95%. In certain embodiments, the second analysis achieves a negative predictive value of at least 96%. In certain embodiments, the second analysis achieves a negative predictive value of at least 97%. In certain embodiments, the second analysis achieves a negative predictive value of at least 98%. In certain embodiments, the second analysis achieves a negative predictive value of at least 99%.

[0224] Longitudinal analysis Reference is now made to the longitudinal analysis module 230, which represents an optional module of the status analysis system 170 shown in FIG. 2A (shown in dotted lines). In various embodiments, longitudinal analysis allows monitoring an individual identified as having a health condition to determine whether the individual's health condition has progressed. In various embodiments, longitudinal analysis includes analyzing whether a first biological sample obtained from an individual at a first time point differs from a second biological sample obtained from the individual at a second time point. For example, longitudinal analysis may include determining a difference in a signal indicative of a health condition in the first biological sample and the second biological sample. The signal may be the presence or amount of circulating tumor DNA indicative of cancer. Thus, longitudinal analysis may include determining a change in circulating tumor DNA present in the first biological sample and the second biological sample, which may be indicative of a change (e.g., progression) of a health condition (e.g., cancer). In various embodiments, if the longitudinal analysis module 230 determines that an individual's health condition is progressing, an intervention may be recommended and / or provided to the individual to slow the progression of the health condition.

[0225] In various embodiments, the longitudinal analysis module 230 analyzes marker information obtained from additional samples obtained from the individual at a time point after the individual was identified as having the health condition. For example, the individual may have been previously identified as having the health condition through a screening (e.g., screening 125 of FIG. 1A) and a second analysis (e.g., second analysis 130 of FIG. 1A), where the screening and / or second analysis may have included analysis of sequence information, such as the methylation status of multiple informative CGIs.

[0226] In various embodiments, the longitudinal analysis module 230 analyzes sequence information identifying the methylation state of the plurality of information CGIs obtained from additional samples obtained at a later time point and compares it to the methylation state of the plurality of information CGIs obtained from a previous sample. In various embodiments, such sequence information may be background corrected sequence information, for example, corrected by an intra-individual analysis combining sequence information from the target nucleic acid and the reference nucleic acid. Thus, the longitudinal analysis module 230 generates a longitudinal understanding of how the methylation state of the plurality of information CGIs has changed over time. This longitudinal understanding provides information for determining the progression of the health condition. In various embodiments, if the longitudinal methylation pattern of the plurality of information CGIs indicates that the individual's health condition is progressing, the individual can be provided with an intervention to slow or stop the progression of the health condition. In various embodiments, the intervention can be a surgical intervention, a therapeutic intervention (e.g., chemotherapy, gene therapy, gene editing), or a lifestyle intervention (e.g., a change in behavior or habits).

[0227] Interaction between third-party and condition analysis systems 3A shows an interaction diagram between a third party and a condition analysis system for performing a multi-stage analysis according to a first embodiment, in which a third party 155A obtains a sample from an individual and a condition analysis system 170 performs one or more assays, screening, intra-individual analysis, and a second analysis.

[0228] Specifically, the process begins at step 305, where a third party 155A obtains a sample from an individual. The third party 155A provides the sample to a status analysis system 170 (308). The status analysis system assays (310) the sample to generate marker information. In various embodiments, the marker information includes the methylation status of multiple genomic sites, such as multiple selected CpG islands. The status analysis system 170 then performs screening 312 by analyzing the methylation status using a trained machine learning model. The screening can identify whether the individual is at risk for a health condition or not at risk for a health condition. If it is determined that the individual is not at risk for a health condition, the process ends and no further analysis is performed.

[0229] If the individual is determined to be at risk for the health condition, the status analysis system 170 provides an indication to the third party 155A that the individual is at risk for the health condition (315). In step 318, the third party 155A obtains a second sample from the individual determined to be at risk for the health condition. The third party 155A provides the second sample to the status analysis system 170 (320). The status analysis system 170 assays (322) the second sample to generate methylation information. In one embodiment, assaying the second sample includes performing whole genome bisulfite sequencing. In one embodiment, assaying the second sample includes performing hybrid capture. In various embodiments, step 322 includes assaying the second sample to generate sequence information of the target nucleic acid and sequence information of the reference nucleic acid. For example, the sequence information of the target nucleic acid can include methylation information of the target nucleic acid. The sequence information of the reference nucleic acid can include methylation information of the reference nucleic acid. In step 324, the condition analysis system 170 performs an intra-individual analysis to remove the baseline biological signature and generate background-corrected information. Thus, in step 325, the condition analysis system 170 performs a second analysis by analyzing the background-corrected information to determine the presence or absence of the individual's health condition. If the individual is determined to have a health condition, the individual may be monitored, provided treatment, and / or selected as a candidate for enrollment in a clinical trial.

[0230] FIG. 3B shows an interaction diagram between a third party and a status analysis system for performing a multi-stage analysis according to a second embodiment. Here, a multi-stage analysis can be performed using a sample collected from an individual at a single collection time point. As shown in FIG. 3B, in step 340, a third party 155A obtains a sample from an individual. The third party 155A provides (342) the sample to a status analysis system 170 for processing and analysis. For example, the status analysis system 170 assays (345) the sample to generate marker information. The status analysis system 170 further performs screening 348 by analyzing the marker information to determine whether the individual is at risk for a health condition.

[0231] If the individual is determined to be at risk for the health condition, a subsequent within-individual analysis is performed at step 354 and a second analysis is performed at step 356. Optionally, the condition analysis system 170 transmits (350) an indication back to the third party 155A that the individual is at risk for the health condition. The third party 155A may then notify the individual 352 of the indication. However, in other embodiments, steps 350 and 352 may not be performed.

[0232] In various embodiments, the condition analysis system 170 performs an intra-individual analysis in step 354 after assaying one or more samples from an individual to generate sequence information of the target nucleic acid and sequence information of the reference nucleic acid. For example, the sequence information of the target nucleic acid may include methylation information of the target nucleic acid. The sequence information of the reference nucleic acid may include methylation information of the reference nucleic acid. The condition analysis system 170 performs an intra-individual analysis to remove the baseline biological signature and generate background-corrected information. In step 356, the condition analysis system 170 performs a second analysis by analyzing the background-corrected information generated as a result of step 354 to determine the presence or absence of the individual's health condition. If the individual is determined to have a health condition, the individual may be monitored, provided with treatment, and / or selected as a candidate for enrollment in a clinical trial.

[0233] 3C shows an interaction diagram between a first third party, a second third party, and a status analysis system for performing a multi-stage analysis, according to an embodiment, where a first third party 155A obtains one or more samples from an individual, a second third party 155B performs assays on one or more samples obtained from the individual, and a status analysis system 170 performs screening and / or secondary analyses.

[0234] Specifically, in step 360, third party 155A obtains a sample from an individual. Third party 155B provides (362) the sample to third party 155B, which then assays (365) the sample to generate methylation information. Third party 155B provides (368) assay results including the generated methylation information to status analysis system 170. The status analysis system performs (370) screening to determine whether the individual is at risk for a health condition by analyzing the generated methylation information.

[0235] If the individual is determined not to be at risk for the health condition, the process ends at this point. If the individual is determined to be at risk for the health condition, the status analysis system 170 can provide (372) an indication to the third party 155A that the individual is at risk. Thus, the third party 155A can obtain a second sample 375 from the individual (e.g., at a second visit by the individual). The third party 155A provides the second sample 378 to the third party 155B, which assays (380) the second sample. In various embodiments, the third party 155B performs whole genome bisulfite sequencing. In various embodiments, the third party 155B performs hybrid capture. In various embodiments, the third party 155B generates methylation information as a result of assaying the second sample. In various embodiments, the third party 155B generates sequence information for the target nucleic acid and sequence information for the reference nucleic acid. The sequence information for the target nucleic acid and the sequence information for the reference nucleic acid can include methylation information. Thus, the third party 155B provides the results of the second assay 382, ​​which includes methylation information of the target nucleic acid and the reference nucleic acid, to the status analysis system 170.

[0236] In step 384, the condition analysis system performs an intra-individual analysis to remove the baseline biological signature and generate background-corrected information. The condition analysis system 170 performs a second analysis 385 by analyzing the background-corrected information to determine whether the individual has a health condition. If the individual is determined to have a health condition, the individual may be monitored, provided treatment, and / or selected as a candidate for enrollment in a clinical trial.

[0237] Examples of methods for performing within-individual analyses 4 shows an example of a flow process including an in-subject analysis, according to an embodiment. Step 410 includes obtaining a target nucleic acid and a reference nucleic acid from one or more samples.

[0238] Step 420 includes generating sequence information from a target nucleic acid, where the sequence information from the target nucleic acid may include a signature that provides information for determining the presence or absence of a health condition, but may also include a baseline biological signature that is present, regardless of whether the nucleic acid originates from a diseased or non-disease source. Step 430 includes generating sequence information from a reference nucleic acid, where the sequence information of the reference nucleic acid includes a baseline biological signature that provides less information for determining the presence or absence of a health condition compared to the sequence information of the target nucleic acid.

[0239] Step 440 includes combining sequence information from the target nucleic acid with sequence information from the reference nucleic acid to generate a background corrected signal that provides information for determining the presence or absence of a health condition. As shown in FIG. 4, step 440 can include both step 450 and step 460. Step 450 includes aligning sequence information from the target nucleic acid with sequence information from the reference nucleic acid. Step 460 includes determining differences between the sequence information from the target nucleic acid and the sequence information from the reference nucleic acid. In various embodiments, step 460 includes determining differences on a position-by-position basis.

[0240] Step 470 includes predicting the presence or absence of a health condition using the informative background-corrected signal. Thus, if an individual is determined to have the presence of a health condition, the individual may be provided with a treatment to prophylactically or therapeutically treat the health condition.

[0241] Examples of methods for selecting informative biomarkers Disclosed herein is a method for selecting informative biomarkers for inclusion in the first stage of a multi-stage analysis. For example, the methods described herein are useful for identifying informative biomarkers for performing a first analysis, such as screening, that achieves high specificity and excludes a large proportion of true negatives (e.g., individuals who are not at risk for a health condition). Reference is made to Figure 4B, which shows an example of a flow process for selecting informative biomarkers for inclusion in the first stage of a multi-stage analysis.

[0242] Step 480 includes obtaining a starting set of biomarkers. Exemplary biomarkers described herein include, but are not limited to, lipids, lipoproteins, proteins, cytokines, chemokines, growth factors, peptides, nucleic acids (e.g., DNA or RNA), genes, and oligonucleotides, as well as their associated complexes, metabolites, mutations, variants, polymorphisms, modifications, fragments, subunits, degradation products, elements, and other measurements from an analyte or sample. Exemplary biomarkers further include CpG sites (e.g., all CpG sites in a genome or a subset of all CpG sites), a set of CGIs (e.g., 4059 CGIs), genes (e.g., all known genes or a subset of all known genes), proteins (e.g., all known proteins or a subset of all known proteins), nucleic acids (e.g., all known coding or non-coding nucleic acids), and metabolites (e.g., all known metabolites or a subset of all known metabolites). In certain embodiments, step 480 includes obtaining a starting set of biomarkers that includes one or more of a CpG site (all CpG sites in the genome), a set of CGIs, a gene (e.g., all known genes or a subset of all known genes), and a protein (e.g., all known proteins or a subset of all known proteins). In certain embodiments, step 480 includes obtaining a starting set of biomarkers that includes each of a CpG site (all CpG sites in the genome), a set of CGIs, a gene (e.g., all known genes or a subset of all known genes), and a protein (e.g., all known proteins or a subset of all known proteins).

[0243] Step 482 includes determining the signal of the starting set of biomarkers across the first and second plurality of samples. In various embodiments, the first plurality of samples refers to healthy samples or samples in the absence of a health condition. In various embodiments, the healthy samples include healthy normal tissue. In various embodiments, the healthy samples include non-cancer cell-free DNA samples. In various embodiments, the second plurality of samples refers to samples having a health condition. For example, the samples having a health condition may be cancer biopsy samples or cell-free DNA samples obtained from patients with cancer. In various embodiments, the second plurality of samples may include samples of different health conditions. For example, the second plurality of samples may include samples of different cancers (e.g., multiple cancer samples). In other embodiments, the second plurality of samples includes samples of a common health condition. For example, the second plurality of samples may include samples of a common cancer.

[0244] Generally, determining the signal of the starting set of biomarkers across the first and second multiple samples may include obtaining biomarker signal information from one or more assays. For example, if the biomarkers are protein biomarkers, determining the signal of the protein biomarker across the first and second multiple samples may include performing an assay (e.g., a multiplex immunoassay) to determine the level of each protein biomarker in the first multiple samples (e.g., healthy samples) and the level of each protein biomarker in the second multiple samples (e.g., one or more health state samples). For example, if the biomarkers refer to CpG sites, determining the signal of the protein biomarker across the first and second multiple samples may include performing a bisulfite conversion assay and nucleic acid sequencing to determine the methylation status of the CpG sites in the first multiple samples (e.g., healthy samples) and the methylation status of the CpG sites in the second multiple samples (e.g., one or more health state samples).

[0245] Step 484 includes performing a ranking of the biomarkers using the determined signals. As shown in FIG. 4B, step 484 may include one or more of steps 486, 488, and 490. In various embodiments, step 484 includes performing only one of steps 486, 488, and 490. In various embodiments, step 484 includes performing two of steps 486, 488, and 490. In various embodiments, step 484 includes performing each of steps 486, 488, and 490.

[0246] Step 486 includes ranking the biomarkers based on the difference between the signals of the first and second plurality of samples. For example, a biomarker that is highly differentially expressed or present across the first and second plurality of samples may be ranked higher than another biomarker that is similarly expressed or present across both the first and second plurality of samples.

[0247] Step 488 includes ranking the biomarkers according to their significance in distinguishing between the first or second plurality of samples. For example, significance can be represented as a p-value determined by a statistical test. Thus, by performing a statistical test comparing signals from the first plurality of samples to signals from the second plurality of samples, the resulting p-value can represent a significance value for distinguishing between samples of the first or second plurality of samples. Thus, biomarkers with smaller p-values ​​(indicating statistical significance) can be ranked higher than biomarkers with larger p-values.

[0248] Step 490 includes ranking the biomarkers by determining their importance values ​​by running a cancer prediction algorithm. Thus, biomarkers associated with higher importance values ​​may be ranked higher than biomarkers associated with lower importance values.

[0249] The significance (e.g., statistical test + p-value) in distinguishing samples (e.g., distinguishing between healthy and cancer, or the difference between a particular cancer sample and a particular non-cancerous sample (e.g., indicative of health and / or other types of cancer)).

[0250] Step 492 includes selecting the top X biomarkers according to the ranking for inclusion in the first stage test of the multi-stage analysis. In various embodiments, X refers to between 1 and 1000 biomarkers. In various embodiments, X refers to between 2 and 900 biomarkers, between 3 and 800 biomarkers, between 4 and 700 biomarkers, between 5 and 600 biomarkers, between 6 and 500 biomarkers, between 7 and 400 biomarkers, between 8 and 300 biomarkers, between 9 and 200 biomarkers, or between 10 and 200 biomarkers. In certain embodiments, X refers to between 10 and 200 biomarkers.

[0251] Machine learning models for analyzing sequence information As disclosed herein, a trained machine learning model can be deployed to analyze sequence information and predict whether an individual is at risk for or has a health condition. In various embodiments, the sequence information includes the methylation status of multiple genomic sites. Thus, the trained machine learning model analyzes differential methylation of multiple genomic sites and outputs a prediction.

[0252] In various embodiments, the trained machine learning model is introduced as part of a screen (e.g., screen 125 shown in FIG. 1A). Thus, the trained machine learning model can analyze sequence information generated via an assay (e.g., assay 120A shown in FIG. 1A) to determine whether an individual is at risk for a health condition. In various embodiments, the trained machine learning model is introduced as part of a second analysis (e.g., second analysis 130 shown in FIG. 1A). Thus, the trained machine learning model can analyze sequence information. In some embodiments, the sequence information includes the methylation status of a plurality of genomic sites, such as a plurality of CpG sites disclosed herein, to determine whether an individual has a health condition. In various embodiments, the sequence information includes background-corrected sequence information generated by an intra-individual analysis (e.g., intra-individual analysis 128 shown in FIG. 1A) to determine whether an individual has a health condition. In some embodiments, the sequence information need not be background-corrected sequence information, but instead includes methylation sequence information from a target nucleic acid (not corrected using sequence information from a reference nucleic acid).

[0253] In various embodiments, the machine learning model is one of a regression model (e.g., linear regression, logistic regression, or polynomial regression), a decision tree, a random forest, a support vector machine, a naive Bayes model, a K-means cluster, or a neural network (e.g., a feedforward network, a convolutional neural network (CNN), a deep neural network (DNN), an autoencoder neural network, a generative adversarial network, or a recurrent network (e.g., a long short-term memory network (LSTM), a bidirectional recurrent network, a deep bidirectional recurrent network)).

[0254] The machine learning model can be trained using a method that implements machine learning, such as a linear regression algorithm, a logistic regression algorithm, a decision tree algorithm, a support vector machine classification, a naive Bayes classification, a K-nearest neighbor classification, a random forest algorithm, a deep learning algorithm, a gradient boosting algorithm, and any one or combination of dimensionality reduction techniques such as manifold learning, principal component analysis, factor analysis, autoencoder regularization, and independent component analysis. In various embodiments, the machine learning model is trained using a supervised learning algorithm, an unsupervised learning algorithm, a semi-supervised learning algorithm (e.g., partially supervised), weakly supervised, transfer, multi-task learning, or any combination thereof.

[0255] In various embodiments, the machine learning model has one or more parameters, such as hyperparameters or model parameters. The hyperparameters are typically set prior to training. Examples of hyperparameters include learning rates, depths or leaves of decision trees, the number of hidden layers in a deep neural network, the number of clusters in a K-means cluster, penalties in regression models, and regularization parameters associated with a cost function. The model parameters are typically adjusted during training. Examples of model parameters include weights associated with nodes in a layer of a neural network, support vectors in a support vector machine, and coefficients in a regression model. The model parameters of the machine learning model are trained (e.g., adjusted) using training data to improve the predictive power of the machine learning model.

[0256] In certain embodiments, the machine learning model analyzes the methylation status of a plurality of genomic sites in the cell-free DNA to generate a prediction. The methylation status corresponds to a set of cancer informative CpG islands (CGIs), where the cancer informative CGIs are selected from the group consisting of the set of ranked candidate CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 50 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 100 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 150 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 200 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 250 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 300 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 400 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 500 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 600 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 700 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 800 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 900 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 1000 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 2500 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 5000 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 7500 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 10000 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 15000 CGIs.In various embodiments, the machine learning model analyzes the methylation status of at least 20,000 CGIs. In various embodiments, the machine learning model analyzes the methylation status of at least 25,000 CGIs.

[0257] In various embodiments, the machine learning model analyzes the methylation status of CGIs across the entire genome. For example, the machine learning model can be implemented to analyze sequencing data generated from whole genome sequencing (e.g., whole genome bisulfite sequencing).

[0258] Additionally, disclosed herein are specific genomic sites, such as CpG islands (CGIs), whose methylation status can provide information for determining whether an individual is at risk for or has a health condition. These informative CGIs can represent signals in a sample. In some embodiments, the methylation status of an informative CGI representing a signal in a sample can indicate the presence of a health condition. In some embodiments, the methylation status of an informative CGI representing a signal in a sample can indicate the absence of a health condition. In various embodiments, the methods disclosed herein, such as those involving multi-stage analysis, are useful for detecting or identifying signals in a sample (e.g., the methylation status of an informative CGI). In various embodiments, the methods disclosed herein, such as those involving multi-stage analysis, are useful for increasing the likelihood that a detected signal in a sample (e.g., the methylation status of an informative CGI) is authentic. Thus, there is confidence that a signal detected by multi-stage analysis (e.g., the methylation status of an informative CGI) is present in a sample.

[0259] The methylation status of the cancer informative CGI can be useful in predicting whether an individual has a health condition. In various embodiments, the methylation status of the cancer informative CGI is a background-corrected methylation status of the cancer informative CGI. For example, the background-corrected methylation status of the cancer informative CGI can be determined by an intra-individual analysis. For example, the background-corrected methylation status of the cancer informative CGI can be determined by combining the methylation information of the cancer informative CGI of the target nucleic acid and the methylation information of the cancer informative CGI of the reference nucleic acid.

[0260] In various embodiments, each cancer informative CGI may be a "CGI identifier" or reference number that allows the CGI to be referenced by its respective unique CGI identifier during data processing. The attached tables (e.g., Tables 1-4) list for each CGI its respective location in the human genome. Additional examples of CGIs are disclosed in WO2018209361 (see Table 1) and WO2022133315 (see Table 2 entitled "TOO methylation sites" and Table 3 entitled "Pan-cancer methylation sites"), each of which is incorporated herein by reference in its entirety. In some embodiments, the methylation status of multiple CpGs within a CGI may be analyzed. In some embodiments, at least some of the CpGs within a CGI may be analyzed. In other embodiments, all of the CpGs within a CGI may be analyzed. In some embodiments, the analysis of CGIs contemplated herein may include analyzing CpGs within at least a portion of one or more regions in Tables 1-4.

[0261] Reference is now made to FIG. 2E, which is an example of an informative signal. In various embodiments, the informative signal shown in FIG. 2E can be generated as a result of an in-individual analysis. The informative signal thus represents background-corrected sequence information, for example, corrected by an in-individual analysis combining sequence information from a target nucleic acid and a reference nucleic acid. In various embodiments, the informative signal shown in FIG. 2E can represent sequence information of a target nucleic acid. In such embodiments, the signal is not obtained from an in-individual analysis.

[0262] As shown in FIG. 2E, for each instance of an analyte, e.g., a cell-free DNA fragment, for multiple positions along the instance of the analyte, e.g., for each individual CpG site along the DNA fragment, there is information about the marker at that position, e.g., data indicating whether that CpG is methylated or unmethylated. An instance of an analyte can be a single sequenced DNA fragment, or a portion of a single sequenced DNA fragment. In various embodiments, the DNA fragment can be a bisulfite converted DNA fragment. Thus, an instance of an analyte can refer to a sequenced bisulfite converted DNA fragment or a portion thereof.

[0263] Conceptually, using CpG methylation of cell-free DNA as an example, the signal shown in FIG. 2E includes a row (e.g., row 240) for each instance of an analyte, such as a single sequenced DNA fragment. Thus, data for 16 instances of an analyte, e.g., 16 DNA fragments, is shown in FIG. 2E. In FIG. 2, each circle corresponds to a position along an analyte, such as a CpG site. In this example, whether the circle is shown as black or white in FIG. 2E indicates whether the CpG site is methylated (black) or unmethylated (white). In some cases, the information about the marker at a position in the nucleic acid may not be binary.

[0264] Information about the markers for each instance of an analyte in a sample can amount to a large amount of data. As an example, in practice, if deep sequencing is used to obtain the methylation status of CpGs in cell-free DNA from blood samples, using a DNA sequencer that outputs such data in a FASTQ format data file, the signal generated by processing a single blood sample can amount to several gigabytes of data, e.g., 20-30 gigabytes.

[0265] FIG. 2E also shows the relative alignment between individual instances of the analyte. For example, in the DNA example, DNA fragments can be located in the genome of the individual from whom the sample originated, and each location in the genome can have a respective set of coordinates that identify it. Thus, DNA fragments can be assigned coordinates based on their respective locations in the genome, and aligned or grouped by those coordinates. Thus, in FIG. 2E, column 242 shows locations on the analyte, such as single CpG sites in the genome, and individual instances of the analyte are shown aligned by their locations on the analyte.

[0266] Using the location information of each instance of the analyte, the individual instances of the analyte can be grouped into regions within the analyte. Typically, markers associated with health conditions are localized within identifiable regions of the analyte, such as regions within a particular gene or genome. Thus, the signals generated for each instance of the analyte are grouped and processed by health condition information region. In certain embodiments, the information region is a CGI (or at least a portion thereof) disclosed in any of Tables 1-4. The example in FIG. 2E may show data regarding methylation of CpG sites within one information region of the genome for multiple DNA fragments obtained from a biological sample. There may be multiple health condition information regions.

[0267] The different metrics and health information fields, when used, can be useful in detecting various diseases or conditions, such as cancer, autoimmune diseases, metabolic disorders, neurological disorders, aging, and trauma. Further examples of health conditions and diseases are described herein.

[0268] As disclosed herein, the trained machine learning model is deployed to generate predictions that are informative about the presence or absence of a health condition. In this context, the use of the trained machine learning model presents several technical challenges related to coding the signals obtained by processing the biological sample into features. Several challenges arise because the signals contain a large amount of information. One of the challenges includes reducing the amount of data to a set of informative features. However, as the number of features increases, the complexity of the computational model increases. However, as the number of features decreases, information relevant to the detection of a health condition may be lost. Several challenges arise due to the uncertainty of which metrics and analyte regions are truly informative of a health condition. Omitting some metrics or some regions from the set of features may negatively impact the performance of the trained computational model.

[0269] To address such issues, in various embodiments, very specific engineered features are generated from the biological sample. Such engineered features may depend on one or more health information regions (e.g., CGIs) and / or one or more individual windows within the health information region (e.g., CGIs). Each window may have a specific location range within the health information region and a specific size. The size is specified by the number of contiguous sites of interest within the analyte. Thus, metrics are calculated for multiple windows within the health information region. Thus, in certain embodiments, engineered features (e.g., CGIs) that represent metrics within a particular window within the health information region provide health information.

[0270] To train the machine learning model, in some embodiments, a first set of features is computed for a training set, which may include multiple candidate features. The candidate features may include one or more candidate metrics, one or more candidate health information domains, or a combination of both. A computational model is trained using the candidate features and then analyzed to determine which candidate features have a greater impact on the output of the trained computational model. Such analysis can be used to identify features that are more influential to the model, whether by metric or by health information domain. A second set of features can be defined by pruning the first set of features based on the more influential identified features, and a trained machine learning model can be constructed using the second set of features.

[0271] In various embodiments, to generate data for a machine learning model (e.g., for training or deployment), the methodology includes calculating a window- and target region-specific metric for one or more instances of the analyte within a window of a plurality of windows on the target region of the analyte. The particular metric used and the health status information region selected may depend on various factors and may be determined empirically. The machine learning model may be implemented to analyze at least the window- and target region-specific metric. In various embodiments, the window- and target region-specific metric includes a ratio of a number of DNA fragments having a particular number of methylated CpGs to a number of DNA fragments in the window of the target region. In various embodiments, the window- and target region-specific metric includes a ratio of a number of DNA fragments having a particular pattern of methylation to a number of DNA fragments in the window of the target region. As described in more detail below, calculating the metric may include applying two or more functions. For example, calculating the window- and target region-specific metric may include performing a first function that quantifies the occurrence of a methylated CpG within the window of the target region. As another example, calculating a window and target region specific metric may include performing a second function that normalizes the occurrence of methylated CpGs to the number of DNA fragments in the window of the target region.

[0272] In various embodiments, each instance of an analyte (e.g., cell-free DNA) is processed to generate a feature. For each instance of an analyte in the biological sample and for each window of the plurality of windows on the health information area of ​​the analyte, a respective value is generated. After processing the analyte instances, the feature calculation module calculates, for each window of the plurality of windows on the health information area, one or more respective metrics of the window based on the first function and / or the second function of the analyte insta...

Claims

1. 1. A stepwise, multi-part method for analyzing sequence information as an indicator of one or more early stage cancers in a subject, comprising: performing an analysis of said subject's sequence information obtained from said subject's biological sample, wherein the results of said analysis indicate whether said subject is not at risk for having one or more of said early stage cancers; and if the results of said analysis do not indicate said subject is not at risk, analyzing the sequence information of the subjects not indicated as being at risk by performing a second analysis to detect the presence of at least one particular cancer in the subjects.

2. wherein the one or more of the early stage cancers are 15 or more different cancers, and optionally the one or more of the early stage or preclinical cancers are selected from the group consisting of acute lymphocytic leukemia, acute myeloid leukemia, adrenocortical carcinoma, soft tissue sarcoma, lymphoma, anal cancer, gastrointestinal cancer, brain cancer, skin cancer, bile duct cancer, bladder cancer, bone cancer, breast cancer, lung cancer, heart cancer, central nervous system cancer, cervical cancer, chronic lymphocytic leukemia, chronic myelogenous leukemia, chronic myeloproliferative neoplasm, colon cancer, uterine cancer, esophageal cancer, head and neck cancer, eye cancer, fallopian tube cancer, gallbladder cancer, stomach cancer, germ cell tumor 2. The method of claim 1, wherein the cancer is a set of tumors selected from the group consisting of ovarian cancer, gestational choriocarcinoma, hairy cell leukemia, liver cancer, Hodgkin's lymphoma, intraocular melanoma, pancreatic cancer, kidney cancer, leukemia, mesothelioma, metastatic cancer, oral cancer, multiple endocrine neoplasia syndrome, multiple myeloma, myelodysplastic tumor, ovarian cancer, parathyroid cancer, penile cancer, pheochromocytoma, pituitary cancer, plasma cell neoplasm, primary peritoneal cancer, prostate cancer, rectal cancer, retinoblastoma, sarcoma, small intestine cancer, testicular cancer, pharyngeal cancer, thymoma and thymic cancer, thyroid cancer, urethral cancer, uterine cancer, vaginal cancer, and vulvar cancer.

3. 10. The method of claim 1, wherein the early stage cancer is a preclinical stage cancer, optionally wherein the preclinical stage cancer is a stage I or stage II cancer.

4. 2. The method of claim 1, wherein the sequence information comprises methylation sequence information, optionally wherein the methylation sequence information comprises the methylation status of a plurality of genomic sites, optionally wherein the plurality of genomic sites comprises a plurality of CpG sites.

5. 10. The method of claim 1, wherein performing an analysis of the subject's sequence information comprises applying a trained machine learning model.

6. performing an analysis of the sequence information of the subject, calculating, for one or more instances of the analyte within a window of a plurality of windows over a target area of ​​the analyte, a metric specific to the window and the target area; and The method of claim 5 , further comprising using the trained machine learning model to analyze the metrics specific to at least the window and the target region.

7. The sequence information is obtained from an assay, the assay comprising: a. sequencing the nucleic acid in said sample; b. Hybrid capture; c. Methylation-specific PCR; d. performing an assay that generates methylation information, and optionally, that generates sequence information; Obtaining bisulfite converted cell-free DNA (cfDNA); Selectively amplifying a target region of the bisulfite converted cfDNA; and an assay comprising sequencing an amplification product comprising said amplified target region to generate said sequence information; and e. Sequencing of clone libraries generated from template-immortalized libraries The method of claim 1 , comprising performing one or more of:

8. the target region of the bisulfite converted cfDNA comprises a pre-identified region that is differentially methylated in cancer, and optionally the target region of the bisulfite converted cfDNA comprises one or more CpG islands or portions of one or more CpG islands set forth in Tables 1-4, and optionally the target region of the bisulfite converted cfDNA comprises: comprising up to 10%, up to 20%, up to 30%, up to 40%, up to 50%, up to 55%, up to 60%, up to 65%, up to 70%, up to 75%, up to 80%, up to 85%, or up to 90% of a CpG island or portion of a CpG island as set out in any one of Tables 1 to 4; or up to 100, up to 150, up to 200, up to 300, up to 400, up to 500, up to 600, up to 700, up to 800, up to 900, up to 1000, up to 1500, up to 2000, up to 2500, up to 3000, up to 3500, or up to 4000 CpG islands or portions of CpG islands selected from Tables 1 to 4, The method of claim 7.

9. Analyzing the sequence information of the subject not indicated to be at risk includes analyzing sequence information generated from a target region that includes one or more CpG islands or portions of one or more CpG islands set forth in Tables 1-4, and optionally, the target region includes: at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the CpG islands or portions of CpG islands shown in any one of Tables 1 to 4, or at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1500, at least 2000, at least 2500, at least 3000, at least 3500, at least 4000, at least 4500, at least 5000, at least 5500, or at least 6000 CpG islands or portions of CpG islands selected from Tables 1-4 The method of claim 1 , comprising:

10. 2. The method of claim 1, wherein performing the second analysis comprises analyzing the methylation status of a greater number of CpG islands compared to the amount of CpG islands analyzed when performing the analysis of the subject's sequence information, and optionally, performing the second analysis comprises analyzing the methylation status of at least five times as many CpG islands compared to the amount of CpG islands analyzed when performing the analysis of the subject's sequence information, and optionally, performing the second analysis comprises analyzing the methylation status of at least 500 CpG islands, and performing the analysis of the subject's sequence information comprises analyzing the methylation status of at least 100 CpG islands.

11. The biological sample is obtained from the subject while the subject is asymptomatic, and optionally the biological sample comprises: blood samples, stool samples, Urine sample, Mucus samples, Saliva sample The method of claim 1 , comprising any one of:

12. 10. The method of claim 1, wherein the second analysis comprises whole genome sequencing, optionally whole genome bisulfite sequencing.

13. 10. The method of claim 1, wherein the sequence information of the subject determines the tissue of origin of the at least one particular cancer in the subject.

14. performing an analysis of additional sequence information for the subject obtained from additional biological samples for the subject obtained after the time the biological sample was obtained; determining one or more variations between the additional sequence information and the sequence information of the subject; further comprising The method of any one of claims 1 to 13, wherein the determined one or more changes determine the progression of the at least one particular cancer in the subject.

15. 15. The method of claim 14, wherein determining one or more changes between the additional sequence information and the sequence information of the subject comprises determining one or more changes in methylation status across a plurality of genomic sites.