Classification of breast tumors using DNA methylation from liquid biopsy

By detecting the methylation levels at multiple sites in breast cancer samples, a classification model was constructed, which solved the inaccuracy problem of monitoring changes in breast cancer subtypes and treatment response in existing technologies, and realized accurate classification of breast cancer subtypes and personalized treatment decisions.

CN121532829APending Publication Date: 2026-02-13GUARDANT HEALTH INC
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202480046975.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-07-14
Filing Date
2024-07-11
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Current technologies struggle to accurately monitor changes in breast cancer subtypes and individual responses to treatment during the treatment process, leading to uncertainty in treatment decisions.

Method used

By detecting methylation levels at multiple sites, generating multiple metrics, and constructing a classification model, this method utilizes cell-free DNA samples to classify breast cancer subtypes and predict treatment responses. This includes training models using one- and multi-class classification and cross-validation, and processing methylation data to characterize samples.

Benefits of technology

It enables accurate classification of breast cancer subtypes and prediction of individual treatment responses, improving the accuracy of treatment decisions and the effectiveness of personalized treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121532829A_ABST
    Figure CN121532829A_ABST
Patent Text Reader

Abstract

Described herein are gene features for providing prognosis, diagnosis, treatment and molecular subtype classification of cancer by genomic and epigenomic profiling, methods and compositions for determining cancer and subtypes, including breast cancer, and provide specific and sensitive detection of biomarkers of interest. Such biomarkers indicate disease pathogenesis, which provides opportunities to select treatment, including treatment regimens intended to identify responsive candidates and overcome resistance mechanisms.
Need to check novelty before this filing date? Find Prior Art

Description

Cross Reference to Related Applications

[0001] This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 513,760, filed July 14, 2023, which is incorporated by reference herein in its entirety for all purposes. BACKGROUND

[0002] Cancer is a leading cause of disease worldwide. Each year, tens of millions of people are diagnosed with cancer worldwide, and more than half of patients eventually die from it. In many countries, cancer is listed as the second most common cause of death after cardiovascular disease. Early detection is associated with improved outcomes for many cancers.

[0003] To detect cancer, several screening tests are available. Physical examination and history take a general health history, including checking for signs of disease, such as lumps or other unusual physical symptoms. The patient’s health habits and history of past ailments and treatments will also be collected. Laboratory tests are another type of screening test, and can include medical procedures to take a sample of tissue, blood, urine, or other substances in the body before the laboratory tests are performed. Imaging procedures screen for cancer by generating a visual representation of internal areas of the body. Genetic tests detect harmful mutations in certain genes associated with some types of cancer. Genetic tests are particularly useful for many diagnostic methods.

[0004] Breast cancer clinical subtypes are determined by expression of hormone receptors (ER / PR) and human epidermal growth factor receptor 2 (HER2). Different subtypes have different incidence, prognosis, risk of recurrence, and overall survival. Accurate subtype classification is of critical importance to provide treatment decision information. Identification of breast cancer subtypes using immunohistochemistry (IHC) is currently considered the standard of care in clinical practice. Additional confirmatory tests such as FISH can be performed. The PAM50 (Prosigna) test, which is less frequently used, provides more information about molecular subtypes (e.g., risk of metastasis, survival, etc.) based on gene expression of 50 genes. Monitoring using cfDNA can address the challenge, where IHC tests or other tests are typically performed at the beginning of treatment, however breast cancer subtypes can change and evolve over time (phenomenon of subtype switching or receptor transformation). Subtype switching can occur as a result of genetic mutations and other molecular changes within tumor cells and also as a result of therapeutic intervention. There is a great need in the art to detect cfDNA for monitoring subjects before, during, and after treatment. SUMMARY

[0005] This document describes a method comprising: detecting methylation at at least one of more than one loci; generating more than one or more measures for each of the more than one loci; and processing one or more measures to characterize a sample. In other embodiments, one or more measures are obtained from methylation determinations from each of the more than one loci. In other embodiments, the method includes obtaining a sample. In other embodiments, the method includes having the obtained sample. In other embodiments, the method includes constructing a classification model from methylation data of a set of training samples, said set of training samples including one or more of the following: HR+ / HER2-, HR+ / HER2+, HR- / HER2+, and HR- / HER2-. In other embodiments, the classification model is trained using one vs. rest classification. In other embodiments, training includes ranking. In other embodiments, regions are selected based on ranking. In other embodiments, ranking is performed based on the saliency and / or weight of each of the more than one loci. In other embodiments, cross-validation is used to train the classification model. In other embodiments, the method includes cross-validation using 1-fold, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, or more fold cross-validation. In other embodiments, generating more than one or more measures for each of more than one site includes Bernoulli probability distributions and / or mixture models. In other embodiments, processing one or more measures to characterize the sample includes calculating the probability of tumor molecules based on the amount of input DNA at one or more sites. In other embodiments, the sites include custom groups.

[0006] In other embodiments, the custom group is configured as a computer simulation group. In other embodiments, the custom group is configured within a physical group. In other embodiments, the custom group includes a set of oncogenes, promoter regions of a set of oncogenes, HRR genes, immuno-oncology (IO) genes, cancer pathways, methylation peaks found in cancer, or methylation peaks found in clinical samples. In other embodiments, the custom group is refined at least based on literature annotations, common methylation peak locations, and / or public datasets. In other embodiments, more than one site includes one or more genes selected from the group consisting of: TLCD3B, PPP1R1B, TMEM200B, GLP1R, FJX1, ZBED4, PFKFB3, MDFI, ZNF280B, CIZ1, POU4F3, MED1, ARHGEF4, SERTM1, FRMD3, NOTUM, MBP, CLDN10, FABP3, ANXA6, PM20D2, BHLHE23, LRFN5, HOXD10, ECHDC1, GSTP1, CCK BR, MIR4500HG, CALD1, PDE4B, INSC, HLA-E, NRG1, CYYR1, MIR3185, CDH12, ANP32E, SLC7A4, LTBP4, CCDC28B, DLX4, ADCY8, SP8, ASPHD1, HOXB8, AGPS, HLF, NFIX, KCNA2, ROBO3, CCDC92B, KCNH8, PPP4R4, FOXC2, PPP2R2C, CAPG, ZNF329, IL12RB2, HOXD-AS2, S ORCS1, HOXD13, SAMM50, UNC13A, CACNG8, CADPS, HIF3A, CSMD3, MIR137, DLX4, BEGAIN, PHACTR2, ZNF667, FENDRR, SRP68, TUBB2 A. SLC16A12, MARCKS, C1QL1, FOXL2, ERBB2, LAYN, PDGFRA, HRK, FAM20A, LOC105377879, WNT3A, MYH11, NMNAT2, KCNIP4, TRANK1 , SYT2, ZNF536, NEUROG3, LIN28A, DMRT3, MCOLN2, ITGA4, PNKY, CHRNB1, GLIS3, HAGLR, PIK3CD, DIO3OS, DUOX2, CRHR1, FAM20A, RFX4, PLXNC1, P2RX6, TLCD3B, PPP1R1B, TMEM200B, GLP1R, FJX1, ZBED4, PFKFB3, MDFI, ZNF280B, CIZ1, POU4F3, MED1, ARHGEF4,SERTM1、FRMD3、NOTUM、MBP、CLDN10、FABP3、ANXA6、PM20D2、BHLHE23、LRFN5、HOXD10、ECHDC1、GSTP1、CCKBR、MIR4500HG、CALD1、PDE4B、INSC、HLA-E、NR G1、CYYR1、MIR3185、CDH12、ANP32E、SLC7A4、LTBP4、CCDC28B、DLX4、ADCY8、SP8、ASPHD1、HOXB8、AGPS、HLF、NFIX、KCNA2、ROBO3、CCDC92B、KCNH8、PPP4R 4、FOXC2、PPP2R2C、CAPG、ZNF329、IL12RB2、HOXD-AS2、SORCS1、HOXD13、SAMM50、UNC13A、CACNG8、CADPS、HIF3A、CSMD3、MIR137、DLX4、BEGAIN、PHACTR2 、ZNF667、FENDRR、SRP68、TUBB2A、SLC16A12、MARCKS、C1QL1、FOXL2、ERBB2、LAYN、PDGFRA、HRK、FAM20A、LOC105377879、WNT3A、MYH11、NMNAT2、KCNIP4、T RANK1、SYT2、ZNF536、NEUROG3、LIN28A、DMRT3、MCOLN2、ITGA4、PNKY、CHRNB1、GLIS3、HAGLR、PIK3CD、DIO3OS、DUOX2、CRHR1、FAM20A、RFX4、PLXNC1、P2R X6、TLCD3B、PPP1R1B、TMEM200B、GLP1R、FJX1、ZBED4、PFKFB3、MDFI、ZNF280B、CIZ1、POU4F3、MED1、ARHGEF4、SERTM1、FRMD3、NOTUM、MBP、CLDN10、FABP3 ,ANXA6, PM20D2, BHLHE23, LRFN5, HOXD10, ECHDC1, GSTP1, CCKBR, MIR4500HG, CALD1, PDE4B, INSC, HLA-E, NRG1, CYYR1, MIR3185, CDH12, ANP32E, SLC7A4, LTBP4, CCDC28B, DLX4, ADCY8, SP8, ASPHD1, HOXB8, AGPS, HLF, NFIX, KCNA2, ROBO3, CCDC92B, KCNH8, PPP4R4, FOXC2, PPP2R2C, CAPG, ZNF329, IL12RB2,HOXD-AS2, SORCS1, HOXD13, SAMM50, UNC13A, CACNG8, CADPS, HIF3A, CSMD3, MIR137, DLX4, BEGAIN, PHAC TR2, ZNF667, FENDRR, SRP68, TUBB2A, SLC16A12, MARCKS, C1QL1, FOXL2, ERBB2, LAYN, PDGFRA, HRK, FAM2 0A, LOC105377879, WNT3A, MYH11, NMNAT2, KCNIP4, TRANK1, SYT2, ZNF536, NEUROG3, LIN28A, DMRT3, MCOLLN2, ITGA4, PNKY, CHRNB1, GLIS3, HAGLR, PIK3CD, DIO3OS, DUOX2, CRHR1, FAM20A, RFX4, PLXNC1, P2RX60. In other embodiments, more than one site includes one or more genes selected from the group consisting of: TLCD3B, MIR137, ROBO3, UNC13A, LIN28A, ASPHD1, KCNH8, CADPS, NMNAT2, and POU4F3. In other embodiments, characterizing the sample includes determining the gene expression of one or more biomarkers. In other embodiments, the method includes diagnosing a patient with cancer. In other embodiments, the method includes prognosing a patient as susceptible to cancer. In other embodiments, the sample includes cell-free DNA.

[0007] This document describes a method comprising: detecting methylation at at least one of more than one sites; generating more than one methylation determination for each of the more than one sites; obtaining one or more measures from the methylation determination; and processing the one or more measures to generate a probability that a patient has cancer. In other embodiments, generating the probability that a patient has cancer includes determining a cancer subtype. In other embodiments, the cancer subtype is characterized as one or more of HR, HER2, and TNBC. In other embodiments, the method includes diagnosing a patient with cancer. In other embodiments, the method includes prognosing a patient as susceptible to cancer. In other embodiments, the sample comprises cell-free DNA. In other embodiments, more than one site includes one or more genes selected from the group consisting of: TLCD3B, PPP1R1B, TMEM200B, GLP1R, FJX1, ZBED4, PFKFB3, MDFI, ZNF280B, CIZ1, POU4F3, MED1, ARHGEF4, SERTM1, FRMD3, NOTUM, MBP, CLDN10, FABP3, ANXA6, PM20D2, BHLHE23, LRFN5, HOX D10, ECHDC1, GSTP1, CCKBR, MIR4500HG, CALD1, PDE4B, INSC, HLA-E, NRG1, CYYR1, MIR3185, CDH12, ANP32E, SLC7A4 , LTBP4, CCDC28B, DLX4, ADCY8, SP8, ASPHD1, HOXB8, AGPS, HLF, NFIX, KCNA2, ROBO3, CCDC92B, KCNH8, PPP4R4, FOXC 2. PPP2R2C, CAPG, ZNF329, IL12RB2, HOXD-AS2, SORCS1, HOXD13, SAMM50, UNC13A, CACNG8, CADPS, HIF3A, CSMD3, MI R137, DLX4, BEGAIN, PHACTR2, ZNF667, FENDRR, SRP68, TUBB2A, SLC16A12, MARCKS, C1QL1, FOXL2, ERBB2, LAYN, PDG FRA, HRK, FAM20A, LOC105377879, WNT3A, MYH11, NMNAT2, KCNIP4, TRANK1, SYT2, ZNF536, NEUROG3, LIN28A, DMRT3, MCOLN2, ITGA4, PNKY, CHRNB1, GLIS3, HAGLR, PIK3CD, DIO3OS, DUOX2, CRHR1, FAM20A, RFX4, PLXNC1, P2RX6, TLCD3B,PPP1R1B、TMEM200B、GLP1R、FJX1、ZBED4、PFKFB3、MDFI、ZNF280B、CIZ1、POU 4F3、MED1、ARHGEF4、SERTM1、FRMD3、NOTUM、MBP、CLDN10、FABP3、ANXA6、PM20 D2, BHLHE23, LRFN5, HOXD10, ECHDC1, GSTP1, CCKBR, MIR4500HG, CALD1, PDE4B, INSC, HLA-E, NRG1, CYYR1, MIR3185, CDH12, ANP32E, SLC7A4, LTBP4, CCDC 28B、DLX4、ADCY8、SP8、ASPHD1、HOXB8、AGPS、HLF、NFIX、KCNA2、ROBO3、CCDC 92B、KCNH8、PPP4R4、FOXC2、PPP2R2C、CAPG、ZNF329、IL12RB2、HOXD-AS2、SOR CS1、HOXD13、SAMM50、UNC13A、CACNG8、CADPS、HIF3A、CSMD3、MIR137、DLX4、 BEGAIN、PHACTR2、ZNF667、FENDRR、SRP68、TUBB2A、SLC16A12、MARKKS、C1QL1 、FOXL2、ERBB2、LAYN、PDGFRA、HRK、FAM20A、LOC105377879、WNT3A、MYH11、N MNAT2、KCNIP4、TRANK1、SYT2、ZNF536、NEUROG3、LIN28A、DMRT3、MCOLN2、ITG A4、PNKY、CHRNB1、GLIS3、HAGLR、PIK3CD、DIO3OS、DUOX2、CRHR1、FAM20A、RF X4, PLXNC1, P2RX6, TLCD3B, PPP1R1B, TMEM200B, GLP1R, FJX1, ZBED4, PFKFB3 、MDFI、ZNF280B、CIZ1、POU4F3、MED1、ARHGEF4、SERTM1、FRMD3、NOTUM、MBP、 CLDN10、FABP3、ANXA6、PM20D2、BHLHE23、LRFN5、HOXD10、ECHDC1、GSTP1、CCK BR、MIR4500HG、CALD1、PDE4B、INSC、HLA-E、NRG1、CYYR1、MIR3185、CDH12、AN P32E、SLC7A4、LTBP4、CCDC28B、DLX4、ADCY8、SP8、ASPHD1、HOXB8、AGPS、HLF、NFIX, KCNA2, ROBO3, CCDC92B, KCNH8, PPP4R4, FOXC2, PPP2R2C, CAPG, ZNF329, IL12RB2, HOXD-AS2, SORCS1, HOXD13, SAMM50, U NC13A, CACNG8, CADPS, HIF3A, CSMD3, MIR137, DLX4, BEGAIN, PHACTR2, ZNF667, FENDRR, SRP68, TUBB2A, SLC16A12, MARCKS, C1 QL1, FOXL2, ERBB2, LAYN, PDGFRA, HRK, FAM20A, LOC105377879, WNT3A, MYH11, NMNAT2, KCNIP4, TRANK1, SYT2, ZNF536, NEUROG3, LIN28A, DMRT3, MCOLLN2, ITGA4, PNKY, CHRNB1, GLIS3, HAGLR, PIK3CD, DIO3OS, DUOX2, CRHR1, FAM20A, RFX4, PLXNC1, P2RX60. In other embodiments, more than one site includes one or more genes selected from the group consisting of: MIR137, ROBO3, UNC13A, LIN28A, ASPHD1, KCNH8, CADPS, NMNAT2, and POU4F3. In other embodiments, the method includes selecting a treatment chosen from trastuzumab, pertuzumab, trastuzumab emtansine, neratinib, lapatinib, tucatinib, margetuximab, and trastuzumab deruxtecan. In other embodiments, the method includes administering treatment to a patient, wherein the treatment is trastuzumab, pertuzumab, trastuzumab emtansine, neratinib, lapatinib, tucatinib, margetuximab, and / or trastuzumab deruxtecan. In other embodiments, at least one of more than one sites is located in one or more genes selected from the group consisting of: TLCD3B, PPP1R1B, TMEM200B, GLP1R, FJX1, ZBED4, PFKFB3, MDFI, ZNF280B, CIZ1, POU4F3, MED1, ARHGEF4, SERTM1, FRMD3, NOTUM, MBP, CLDN10, FABP3, ANXA6, PM20D2, BHLHE23, LRFN5, HOXD10, ECHDC1, GSTP1, CCKBR, MIR4500HG, CALD1, PDE4B, INSC, HLA-E, NRG1, CYYR1, MIR3185, CDH12.ANP32E、SLC7A4、LTBP4、CCDC28B、DLX4、ADCY8、SP8、ASPHD1、HOXB8、AGPS、H LF、NFIX、KCNA2、ROBO3、CCDC92B、KCNH8、PPP4R4、FOXC2、PPP2R2C、CAPG、ZN F329、IL12RB2、HOXD-AS2、SORCS1、HOXD13、SAMM50、UNC13A、CACNG8、CADPS 、HIF3A、CSMD3、MIR137、DLX4、BEGAIN、PHACTR2、ZNF667、FENDRR、SRP68、TUB B2A, SLC16A12, MARCKS, C1QL1, FOXL2, ERBB2, LAYN, PDGFRA, HRK, FAM20A, LOC105377879, WNT3A, MYH11, NMNAT2, KCNIP4, TRANK1, SYT2, ZNF536, NEURO G3、LIN28A、DMRT3、MCOLN2、ITGA4、PNKY、CHRNB1、GLIS3、HAGLR、PIK3CD、DI O3OS、DUOX2、CRHR1、FAM20A、RFX4、PLXNC1、P2RX6、TLCD3B、PPP1R1B、TMEM20 0B、GLP1R、FJX1、ZBED4、PFKFB3、MDFI、ZNF280B、CIZ1、POU4F3、MED1、ARHGE F4、SERTM1、FRMD3、NOTUM、MBP、CLDN10、FABP3、ANXA6、PM20D2、BHLHE23、LR FN5、HOXD10、ECHDC1、GSTP1、CCKBR、MIR4500HG、CALD1、PDE4B、INSC、HLA-E 、NRG1、CYYR1、MIR3185、CDH12、ANP32E、SLC7A4、LTBP4、CCDC28B、DLX4、ADCY 8、SP8、ASPHD1、HOXB8、AGPS、HLF、NFIX、KCNA2、ROBO3、CCDC92B、KCNH8、PPP 4R4、FOXC2、PPP2R2C、CAPG、ZNF329、IL12RB2、HOXD-AS2、SORCS1、HOXD13、SA MM50、UNC13A、CACNG8、CADPS、HIF3A、CSMD3、MIR137、DLX4、BEGAIN、PHACTR 2、ZNF667、FENDRR、SRP68、TUBB2A、SLC16A12、MARCKS、C1QL1、FOXL2、ERBB2、LAYN, PDGFRA, HRK, FAM20A, LOC105377879, WNT3A, MYH11, NMNAT2, KCNIP4, TRANK1, SYT2, ZNF536, NEUROG3, LIN28A, DMRT3, MCOLN2, ITG A4, PNKY, CHRNB1, GLIS3, HAGLR, PIK3CD, DIO3OS, DUOX2, CRHR1, FAM20A, RFX4, PLXNC1, P2RX6, TLCD3B, PPP1R1B, TMEM200B, GLP1R, FJX1, ZBED4, PFKFB3, MDFI, ZNF280B, CIZ1, POU4F3, MED1, ARHGEF4, SERTM1, FRMD3, NOTUM, MBP, CLDN10, FABP3, ANXA6, PM20D2, BHLHE23, LRFN 5. HOXD10, ECHDC1, GSTP1, CCKBR, MIR4500HG, CALD1, PDE4B, INSC, HLA-E, NRG1, CYYR1, MIR3185, CDH12, ANP32E, SLC7A4, LTBP4, CCDC28B , DLX4, ADCY8, SP8, ASPHD1, HOXB8, AGPS, HLF, NFIX, KCNA2, ROBO3, CCDC92B, KCNH8, PPP4R4, FOXC2, PPP2R2C, CAPG, ZNF329, IL12RB2, HO XD-AS2, SORCS1, HOXD13, SAMM50, UNC13A, CACNG8, CADPS, HIF3A, CSMD3, MIR137, DLX4, BEGAIN, PHACTR2, ZNF667, FENDRR, SRP68, TUBB2A SLC16A12, MARCKS, C1QL1, FOXL2, ERBB2, LAYN, PDGFRA, HRK, FAM20A, LOC105377879, WNT3A, MYH11, NMNAT2, KCNIP4, TRANK1, SYT2, ZNF536, NEUROG3, LIN28A, DMRT3, MCOLLN2, ITGA4, PNKY, CHRNB1, GLIS3, HAGLR, PIK3CD, DIO3OS, DUOX2, CRHR1, FAM20A, RFX4, PLXNC1, P2RX60. In other embodiments, at least one of more than one site is located in one or more genes selected from the group consisting of: MIR137, ROBO3, UNC13A, LIN28A, ASPHD1, KCNH8, CADPS,NMNAT2 and POU4F3. In other embodiments, the method includes diagnosing a patient as having cancer. In other embodiments, the method includes assessing a patient's susceptibility to cancer. In other embodiments, the sample includes cell-free DNA.

[0008] This document describes a method for determining the diagnosis, prognosis, and susceptibility to cancer in an individual, including: determining whether there are lower or higher levels of gene expression relative to a normal baseline standard in the individual. In various embodiments, the genes include RE1 silencing transcription factor (REST) ​​pathway targets. In various embodiments, the genes include one or more of the following: MIR137, ROBO3, UNC13A, LIN28A, ASPHD1, KCNH8, CADPS, NMNAT2, and / or POU4F3. In other embodiments, the method includes cancer subtyping, such as breast cancer subtypes HER2, HR, and TNBC. In other embodiments, the method includes selecting a treatment for the subject. In other embodiments, the method includes administering a treatment to the subject.

[0009] For example, a method for selecting a therapeutic treatment for a subject is disclosed, comprising: a) obtaining a biological sample from the subject; b) determining the expression levels of MIR137, ROBO3, UNC13A, LIN28A, ASPHD1, KCNH8, CADPS, NMNAT2, and / or POU4F3 in the biological sample; c) comparing the expression levels with a control; and d) selecting a therapeutic treatment based on the comparison. In some embodiments, the control is derived from one or more subjects without cancer. In some embodiments, the expression levels are inferred from analysis of cfDNA. In some embodiments, abnormal expression of MIR137, ROBO3, UNC13A, LIN28A, ASPHD1, KCNH8, CADPS, NMNAT2, and / or POU4F3 indicates resistance to the therapeutic treatment in the subject, and thus an alternative therapeutic treatment is selected. In some embodiments, the method involves administering an alternative therapy. In some implementations, the method involves determining the expression of MIR137, ROBO3, UNC13A, LIN28A, ASPHD1, KCNH8, CADPS, NMNAT2, and / or POU4F3. In other implementations, the method includes cancer subtyping, such as breast cancer subtypes HER2, HR, and TNBC.

[0010] A method for selecting a therapeutic treatment for a subject is also disclosed, comprising: a) obtaining a biological sample from the subject; b) determining the expression level of one or more components of the RE1 silencing transcription factor (REST) ​​pathway in the biological sample; c) comparing the expression level with a control; and d) selecting a therapeutic treatment based on the comparison. In some embodiments, the control is derived from one or more subjects without cancer. In some embodiments, the expression level is inferred from analysis of cfDNA. In some embodiments, abnormal expression of one or more components of the RE1 silencing transcription factor (REST) ​​pathway indicates resistance to the therapeutic treatment in the subject, and thus alternative therapeutic treatment is selected. In some embodiments, the method involves administering an alternative therapy.

[0011] A method for identifying whether a subject is resistant to a therapeutic treatment is also disclosed, wherein the method includes: a) obtaining a biological sample from the subject; b) determining the expression level of one or more components of the RE1 silencing transcription factor (REST) ​​pathway in the biological sample; c) comparing the expression level with a control; and d) classifying the subject as resistant to the therapeutic treatment based on the comparison. In some embodiments, the control is derived from one or more subjects without cancer. In some embodiments, the expression level is inferred from analysis of cfDNA. In some embodiments, abnormal expression of one or more components of the RE1 silencing transcription factor (REST) ​​pathway indicates resistance to the therapeutic treatment in the subject. In some embodiments, the method involves administering an alternative therapy.

[0012] The use of one or more components of the RE1 silencing transcription factor (REST) ​​pathway as biomarkers for cancer is also disclosed. In some embodiments, the biomarker is resistance to therapeutic treatments for cancer. In other embodiments, the method includes cancer subtyping, such as breast cancer subtypes HER2, HR, and TNBC.

[0013] In some embodiments, the results of the systems and methods disclosed herein are used as input to generate a report. The report may be in paper or electronic format. For example, genetic results determined by the methods and systems disclosed herein, such as the detection of the presence of nucleic acid variants in a sample, may be directly displayed in such a report. In some embodiments, such a report may only show the presence or absence of a disease such as cancer.

[0014] The various steps of the methods disclosed herein, or the steps implemented through the systems disclosed herein, may be performed at the same or different times, in the same or different geographical locations (e.g., countries), and / or by the same or different people. In some embodiments, reports will be communicated to subjects, such as those with cancer who have undergone testing using the methods and systems described herein, or to healthcare professionals, such as physicians treating subjects with cancer. Brief description of the attached diagram

[0015] Figure 1 The Cancer Genome Atlas (TCGTA) classifies samples using a one-vs-rest classification of breast cancer subtypes.

[0016] Figure 2 The number of methylation determinations across BRCA subtypes.

[0017] Figure 3 A comparison of methylation detection platforms with TCGA, including single and multi-classification using three breast cancer subtypes.

[0018] Figure 4 Tumor molecules were modeled using a probabilistic generative model. The methylation data exhibited sparsity, with regions sporadically turning on / off. The degree of sparsity of the signal correlated with DNA input and TF (tumor mitochondrial response). Region scores (molecule counts within a region) depended on TF. The observed molecules were a mixture of normal molecules (background) and tumor molecules.

[0019] Figure 5 The probability of tumor molecules being present in genomic regions. The molecule counts in 15,229 promoter regions were then fitted to the background distribution of region scores from 1,585 cancer-free samples; then, a mixture model could be fitted to the distribution of region scores from 1,014 BRCA samples.

[0020] Figure 6 Applications of probabilistic modeling include normalized methylation scores for cancer subtyping and predicting the likelihood of different cancer subtypes using the described methods.

[0021] Figure 7 BRCA subtyping methodology. Subtype prediction models were applied across different techniques, including training three models (HR / HER2 / TNBC) using 450K methylation determination data from raw tissue samples from TCGA, and then applying these subtype prediction models to methylation determination in our Infinity data for subtype prediction.

[0022] Figure 8. Epigenomic breast cancer subtype classification results.Figure 8A HR & TNBC classifiers trained on methylation determination from the Cancer Genome Atlas (TCGA) of primary tumors, including logistic regression using LASSO, and also Figure 8B The classifier for HER2 was developed using 5-fold CV on samples with methyl binding domain partitioning methylation detection in samples with TF>0.01.

[0023] Figure 9 Subtype predictions are performed across time points using HR and TNBC classifiers. Samples with TF > 0.001 are considered. White corresponds to samples with missing data or below the TF threshold.

[0024] Figure 10 The HER2 predictor region is a target of RE1 silencing transcription factors (REST). Using ENCODE_TF_ChIP-seq_2015 from the library in gp.enrichr, we found that genes associated with the HER2 predictor region are targets of RE1 silencing transcription factors (REST) ​​across a variety of cell lines, including MCF-7, a well-characterized human breast cancer cell line. Detailed Explanation

[0025] While various embodiments of this disclosure have been shown and described herein, those skilled in the art will understand that such embodiments are provided by way of example only. Many variations, modifications, and substitutions will occur to those skilled in the art without departing from this disclosure. It should be understood that various alternatives may be adopted to the embodiments of this disclosure described herein.

[0026] The term “about” and its grammatical equivalents associated with a reference value can include a range of values ​​that are plus or minus 10% of that value. For example, the quantity “about 10” can include quantities from 9 to 11. The term “about” associated with a reference value can include a range of values ​​that are plus or minus 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1%.

[0027] The term "at least" and its grammatical equivalents associated with a reference value can include the reference value and be greater than that value. For example, a quantity "at least 10" can include the value 10 and any value higher than 10, such as 11, 100, and 1,000.

[0028] The term “at most” and its syntactic equivalents associated with a reference value can include the reference value and be less than that value. For example, a quantity “at most 10” can include any value of 10 and below 10, such as 9, 8, 5, 1, 0.5, and 0.1.

[0029] Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” as used herein can include plural referents. Thus, for example, reference to “cell” can include more than one such cell, and reference to “culture” can include reference to one or more cultures and their equivalents known to those skilled in the art, and so on. Unless otherwise expressly indicated, all technical and scientific terms used herein may have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0030] Cancer can be indicated by epigenetic variations such as methylation. Examples of methylation changes in cancer include localized increases in DNA methylation at CpG islands at transcription start sites (TSS) of genes involved in normal growth control, DNA repair, cell cycle regulation, and / or cell differentiation. This hypermethylation may be associated with an abnormal loss of transcriptional capacity of the genes involved and occurs at least as frequently as point mutations and deletions that cause altered gene expression. DNA methylation profiling can be used to detect regions in the genome with different levels of methylation (“differentially methylated regions” or “DMRs”) that have changed during development or been perturbed by disease (e.g., cancer or any cancer-related disease). The genome of cancer cells has an imbalance in the aforementioned DNA methylation patterns, and therefore an imbalance in the functional packaging of DNA. Thus, the combination of chromatin organization abnormalities with methylation changes, when analyzed together, may help enhance cancer profiling. Combining MBD partitioning with fragmentomics data (such as the start and end positions of fragment mappings (related to nucleosome position), fragment length, and associated nucleosome occupancy) can be used for chromatin structure analysis in hypermethylation studies, with the aim of improving biomarker detection rates.

[0031] Methylation profiling can include identifying methylation patterns across different regions of the genome. For example, after partitioning and sequencing molecules based on their methylation levels (e.g., the relative number of methylation sites per molecule), the sequences of molecules in different partitions can be mapped to a reference genome. This can reveal regions of the genome that are more or less methylated compared to other regions. In this way, genomic regions can differ in their methylation levels in contrast to individual molecules.

[0032] Nucleic acid molecules can be characterized by modifications, which can include various chemical or protein modifications (i.e., epigenetic modifications). Non-limiting examples of chemical modifications may include, but are not limited to, covalent DNA modifications, including DNA methylation. In some embodiments, DNA methylation includes adding a methyl group to cytosine (the cytosine followed by guanine in a nucleic acid sequence) at a CpG site. In some embodiments, DNA methylation includes adding a methyl group to adenine, such as N6-methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the fifth carbon of the 6-carbon ring of cytosine). In some embodiments, 5-methylation includes adding a methyl group to the 5C position of cytosine to produce 5-methylcytosine (m5c). In some embodiments, methylation includes derivatives of m5c. Derivatives of m5c include, but are not limited to, 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5-carboxycytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the third carbon of the 6-carbon ring of cytosine). In some embodiments, 3C methylation involves adding a methyl group to the 3C position of cytosine to generate 3-methylcytosine (3mC). Other examples include N6-methyladenine or glycosylation. DNA methylation involves adding a methyl group to DNA (e.g., CpG) and can alter the expression of methylated DNA regions. Methylation can also occur at non-CpG sites; for example, methylation can occur at CpA, CpT, or CpC sites. DNA methylation can alter the activity of methylated DNA regions. For example, when DNA in a promoter region is methylated, gene transcription can be repressed. DNA methylation is essential for normal development, and abnormalities in methylation can disrupt epigenetic regulation. Disruptions in epigenetic regulation, such as repression, can lead to diseases such as cancer. DNA methylation of promoters may indicate cancer.

[0033] The CpG binary is a dinucleotide CpG (cytosine-phosphate-guanine, i.e., at the 5' end of the nucleic acid sequence) on the sense strand of a double-stranded DNA molecule. In the 3' direction, cytosine is followed by guanine and its complementary CpG on the antisense chain. The CpG binary can be fully methylated or hemimethylated (methylated only on one chain).

[0034] CpG dinucleotides are underrepresented in the normal human genome, where most CpG dinucleotide sequences are transcriptionally inert (e.g., near-centromere regions of chromosomes and heterochromatin regions of DNA in repetitive elements) and are methylated. However, many CpG islands are protected from such methylation, especially around transcription start sites (TSS).

[0035] Specifically, as an example, different breast cancer subtypes have different incidences, prognoses, recurrence risks, and overall survival rates. Current descriptions of breast cancer subtypes include HR+ / HER2-, HR+ / HER2+, HR- / HER2+ (Her2-positive), and HR- / HER2- (triple-negative). Accurate subtype classification is crucial for providing information for treatment decisions. It has been reported that ER, PR, and HER2 status changes between primary and metastatic tumors, with a 27.0% rate of change in HR or HER2. Inconsistencies across different biopsy sites may appear similar, with loss of HR status significantly associated with mortality risk (adjusted HR = 1.51, p = 0.002), while gain of HR and HER2 inconsistency are not significantly associated with mortality risk. Given this understanding, sampling metastatic sites (including methylation testing in liquid biopsies, which can reflect systemic cancer status) is necessary for treatment adjustment, which is impractical given current technologies using IHC and other forms of gene expression testing.

[0036] Protein modifications include components that bind chromatin, particularly histones (including their modified forms), as well as components that bind other proteins, such as those involved in replication or transcription. This disclosure provides methods for processing and analyzing nucleic acids with varying degrees of modification, such that the nature of their original modifications is correlated with nucleic acid tags, which can be decoded during nucleic acid analysis by sequencing. The genetic variation in the nucleic acid modifications of a sample can then be correlated with the degree of modification (epigenetic variation) of that nucleic acid in the original sample, said nucleic acids including single-stranded (e.g., ssDNA or RNA) or double-stranded molecules (e.g., dsDNA).

[0037] DNA loss can reduce the presence of one or more types of DNA, making them difficult to detect (e.g., cfDNA). In one or more other cases, existing methods for measuring DNA methylation, such as enrichment or depletion methods, can have relatively high levels of resolution, such as from about 100 base pairs (bp) to about 200 bp, which can make it difficult to accurately determine the amount of DNA methylation. The accuracy of DNA methylation determination can affect the accuracy of tumor score estimation in a sample. Since tumor scores are used to determine whether a sample is derived from a subject with or without a tumor, the accuracy of tumor score estimation can influence individual diagnostic and / or treatment decisions.

[0038] Sample The sample can be any biological sample isolated from the subject. The sample can be a bodily sample. Samples can include body tissues such as known or suspected solid tumors, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells or leukocytes, endothelial cells, tissue biopsies, cerebrospinal fluid, synovial fluid, lymph, ascites, interstitial fluid or extracellular fluid, fluids in the intercellular spaces (including gingival crevicular fluid), bone marrow, pleural effusion, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. Samples are preferably bodily fluids, particularly blood and its fractions, and urine. Samples can be in the form initially isolated from the subject, or can be further processed to remove or add components, such as cells, or to enrich one component relative to other components. Therefore, preferred bodily fluids for analysis are plasma or serum containing cell-free nucleic acids. Samples can be isolated or obtained from the subject and transported to the sample analysis site. Samples can be stored and transported at desirable temperatures, such as room temperature, 4°C, -20°C, and / or -80°C. Samples may be isolated from or obtained from the subject at the sample analysis site. Subjects may be humans, mammals, animals, companion animals, service animals, or pets. Subjects may have cancer. Subjects may not have cancer or may have detectable symptoms of cancer. Subjects may have been treated with one or more cancer therapies, such as chemotherapy, antibodies, vaccines, or biologics. Subjects may be in remission. Subjects may be diagnosed or may not be diagnosed with a susceptibility to cancer or any cancer-related genetic mutations / disorders.

[0039] The volume of plasma can depend on the desired read depth for the sequenced region. Exemplary volumes are 0.4–40 ml, 5–20 ml, and 10–20 ml. For example, the volume can be 0.5 mL, 1 mL, 5 mL, 10 mL, 20 mL, 30 mL, or 40 mL. The volume of plasma sampled can be from 5 ml to 20 ml.

[0040] Samples can contain varying amounts of nucleic acids containing genomic equivalents. For example, a sample of approximately 30 ng DNA can contain approximately 10,000 (10^3) ​​nucleotides. 4 The genome equivalent of 200 billion haploid human genomes, and in the case of cfDNA, it can contain approximately 200 billion (2 × 10⁻⁶) haploid human genomes. 11 ) individual polynucleotide molecules. Similarly, a sample of about 100 ng DNA can contain about 30,000 haploid human genome equivalents, and in the case of cfDNA, about 600 billion individual molecules.

[0041] The sample may contain nucleic acids from different sources, such as nucleic acids from cells and cell-free samples from the same subject, or from cells and cell-free samples from different subjects. The sample may contain nucleic acids carrying mutations. For example, the sample may contain DNA carrying germline mutations and / or somatic mutations. Germline mutations refer to mutations present in the germline DNA of the subject. Somatic mutations refer to mutations originating from somatic cells of the subject, such as cancer cells. The sample may contain DNA carrying cancer-related mutations (e.g., cancer-related somatic mutations). The sample may contain epigenetic variations (i.e., chemical or protein modifications) that are associated with the presence of genetic variations (such as cancer-related mutations). In some embodiments, the sample contains epigenetic variations associated with the presence of genetic variations, wherein the sample does not contain said genetic variations.

[0042] Exemplary amounts of cell-free nucleic acid in the pre-amplification sample range from about 1 fg to about 1 µg, such as 1 pg to 200 ng, 1 ng to 100 ng, 10 ng to 1000 ng. For example, amounts can be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. Amounts can be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng of cell-free nucleic acid molecules. The quantity can be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or 200 ng of cell-free nucleic acid molecules. The method may include obtaining 1 femtogram (fg) to 200 ng.

[0043] Cell-free nucleic acids are nucleic acids that are not contained within cells or otherwise bound to cells, or in other words, nucleic acids that remain in a sample after the removal of intact cells. Cell-free nucleic acids include DNA, RNA, and their hybrids, including genomic DNA, mitochondrial DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), and / or fragments of any of these. Cell-free nucleic acids can be double-stranded, single-stranded, or hybrids thereof. Cell-free nucleic acids can be released into body fluids through secretion or cell death procedures such as cell necrosis and apoptosis. Some cell-free nucleic acids are released into body fluids from cancer cells, such as circulating tumor DNA (ctDNA). Others are released from healthy cells. In some embodiments, cfDNA is cell-free fetal DNA (cffDNA). In some embodiments, cell-free nucleic acids are produced by tumor cells. In some embodiments, cell-free nucleic acids are produced by a mixture of tumor cells and non-tumor cells.

[0044] Cell-free nucleic acids have an example size distribution of approximately 100-500 nucleotides, with molecules of 110 to 230 nucleotides representing approximately 90% of the molecules, a mode of approximately 168 nucleotides, and a second small peak in the range of 240 to 440 nucleotides. Cell-free nucleic acids can be separated from body fluids via a fractionation or partitioning step, in which the cell-free nucleic acids present in solution are separated from intact cells and other insoluble components in the body fluid. Partitioning can include techniques such as centrifugation or filtration. Optionally, cells in the body fluid can be lysed, and cell-free and cellular nucleic acids are processed together. Typically, nucleic acids can be precipitated with alcohol after adding buffer and washing steps. Further cleaning steps, such as silica-based columns, can be used to remove contaminants or salts. Nonspecific bulk carrier nucleic acids, such as Cot-1 DNA, DNA or proteins for bisulfite sequencing, hybridization, and / or ligation, can be added throughout the reaction to optimize certain aspects of the procedure, such as yield.

[0045] Following such processing, the sample can include various forms of nucleic acids, including double-stranded DNA, single-stranded DNA, and single-stranded RNA. In some embodiments, single-stranded DNA and RNA can be converted into double-stranded forms, and therefore included in subsequent processing and analytical steps.

[0046] Analyte Analytes may include nucleic acid analytes and non-nucleic acid analytes. This disclosure provides a method for detecting genetic variations in biological samples from a subject. Biological samples may include polynucleotides from cancer cells. Polynucleotides may be DNA (e.g., genomic DNA, cDNA), RNA (e.g., mRNA, small RNA), or any combination thereof. Biological samples may include, for example, tumor tissue from a biopsy. In some cases, biological samples may include blood or saliva. In specific cases, biological samples may contain cell-free DNA (“cfDNA”) or circulating tumor DNA (“ctDNA”). Cell-free DNA may be present, for example, in blood.

[0047] Examples of non-nucleic acid analytes include, but are not limited to, lipids, carbohydrates, peptides, proteins, glycoproteins (N-linked or O-linked), lipoproteins, phosphoproteins, specific phosphorylated or acetylated variants of proteins, amidated variants of proteins, hydroxylated variants of proteins, methylated variants of proteins, ubiquitinated variants of proteins, sulfated variants of proteins, viral proteins (e.g., viral capsid, viral envelope, viral outer shell, viral appendages, viral glycoproteins, viral spikes, etc.), extracellular and intracellular proteins, antibodies, and antigen-binding fragments. This also includes receptors, antigens, surface proteins, transmembrane proteins, differentiation protein clusters, protein channels, protein pumps, carrier proteins, phospholipids, glycoproteins, glycolipids, cell-cell interaction protein complexes, antigen-presenting complexes, major histocompatibility complexes, engineered T-cell receptors, T-cell receptors, B-cell receptors, chimeric antigen receptors, extracellular matrix proteins, post-translational modifications (e.g., phosphorylation, glycosylation, ubiquitination, nitrosation, methylation, acetylation, or lipidation) of cell surface proteins, gap junctions, and adhesion junctions.

[0048] Typically, systems, apparatus, methods, and compositions can be used to analyze any number of analytes, further including both nucleic acid analytes and non-nucleic acid analytes. For example, the number of analytes being analyzed can be at least about 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 100, at least 1,000, at least 10,000, at least 100,000, or more different analytes present in a region of the sample or a single feature of the substrate.

[0049] One or more nucleic acid analytes and / or non-nucleic acid analytes constitute a set of molecular interactions in the biological system under study (e.g., a cell), which can be considered as an “interaction set”—molecular interactions occurring between molecules belonging to different biochemical families (proteins, nucleic acids, lipids, carbohydrates, etc.) and also within a given family. In various embodiments, the interaction set is a protein-DNA interaction set (a network formed by transcription factors (and DNA or chromatin regulatory proteins) and their target genes). In other embodiments, the interaction set refers to a protein-protein interaction network (PPI) or a protein-protein interaction network (PIN). The methods described herein allow for the study and analysis of interaction sets. Techniques such as proteogenomics (whole genome sequencing, whole exome sequencing, and RNA-seq, as well as mass spectrometry, as examples) can support the study of interaction sets.

[0050] Analysis The methods of this invention can be used to diagnose the presence of a condition, particularly cancer, in a subject, to characterize the condition (e.g., to stage the cancer or determine its heterogeneity), monitor the condition's response to treatment, and achieve prognostic assessment of the risk of condition progression or subsequent disease development. This disclosure can also be used to determine the efficacy of a particular treatment option. If treatment is successful, a successful treatment option may increase the amount of copy number variations or rare mutations detected in the subject's blood as more cancer cells may die and shed DNA. In other instances, this may not occur. In yet another instance, some treatment options may be correlated with the genetic profile of the cancer over time. This correlation can be used to select a therapy. Additionally, if cancer is observed to be in remission after treatment, the methods of this invention can be used to monitor residual disease or disease recurrence.

[0051] The types and numbers of cancers that can be detected include leukemia, brain cancer, lung cancer, skin cancer, nasal cancer, laryngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, colorectal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, solid tumors, heterogeneous tumors, and homogeneous tumors. Cancer type and / or stage can be detected based on genetic variations, including mutations, rare mutations, insertions / deletions, copy number variations, transversions, translocations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structural changes, gene fusions, gene truncation, gene amplification, gene duplication, chromosomal damage, DNA damage, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in the 5-methylcytosine content of nucleic acids.

[0052] Genetic and other analyte data can also be used to characterize specific forms of cancer. Cancers are often heterogeneous in both composition and stage. Genetic profiling data can allow for the characterization of specific subtypes of cancer, which may be important in the diagnosis or treatment of that specific subtype. This information can also provide subjects or practitioners with clues about the prognosis of a specific type of cancer and allow them to adjust treatment options based on disease progression. Some cancers can progress and become more aggressive and genetically unstable. Other cancers may remain benign, inactive, or dormant. The systems and methods disclosed herein can be used to determine disease progression.

[0053] The analyses of this invention can also be used to determine the efficacy of a particular treatment option. If the treatment is successful, a successful treatment option may increase the amount of copy number variations or rare mutations detected in the subject's blood as more cancer cells may die and shed DNA. In other instances, this may not occur. In yet another instance, some treatment options may be correlated with the genetic profile of the cancer over time. This correlation can be used to select a therapy. Additionally, if cancer is observed to be in remission after treatment, the method of this invention can be used to monitor residual disease or recurrence of disease.

[0054] The methods of this invention can also be used to detect genetic variations in conditions other than cancer. Following the onset of certain diseases, immune cells, such as B cells, can undergo rapid clonal expansion. Copy number variation detection can be used to monitor clonal expansion and certain immune states. In this instance, copy number variation analysis can be performed over time to generate a spectrum of how a particular disease may progress. Copy number variation or even rare mutation detection can be used to determine how a pathogen population changes during the course of infection. This can be particularly important during chronic infections (such as HIV / AID or hepatitis infections), where the virus can alter its life cycle state and / or mutate into a more virulent form during the course of infection. When immune cells attempt to destroy transplanted tissue, the methods of this invention can be used to determine or analyze the host body's rejection activity to monitor the state of the transplanted tissue and to modify the process of rejection treatment or prevention.

[0055] Furthermore, the methods of this disclosure can be used to characterize the heterogeneity of anomalies in a subject. Such methods may include, for example, generating a genetic profile of extracellular polynucleotides derived from the subject, wherein the genetic profile includes more than one set of data obtained from copy number variation and rare mutation analysis. In some embodiments, the anomaly is cancer. In some embodiments, the anomaly may be a condition leading to a heterogeneous genomic population. In the example of cancer, some tumors are known to contain tumor cells at different stages of cancer. In other examples, heterogeneity may include multiple lesions of the disease. Again, in the example of cancer, multiple tumor lesions may be present, perhaps one or more of which are the result of metastases that have spread from the primary site.

[0056] The method of this invention can be used to generate or analyze a fingerprint or dataset that represents the sum of genetic information derived from different cells in heterogeneous diseases. This dataset can include individual or combined copy number variation and mutation analyses.

[0057] The methods of this invention can be used for the diagnosis, prognosis, monitoring, or observation of cancer or other diseases. In some embodiments, the methods herein do not involve the diagnosis, prognosis, or monitoring of the fetus, and therefore do not involve noninvasive prenatal testing. In other embodiments, these methods can be used in pregnant subjects to diagnose, prognose, monitor, or observe cancer or other diseases in unborn subjects whose DNA and other polynucleotides may co-circulate with maternal molecules.

[0058] Determination of 5-methylcytosine patterns of nucleic acids Bisulfite-based sequencing and its variations provide a means of determining the methylation patterns of nucleic acids. In some embodiments, determining the methylation pattern includes distinguishing between 5-methylcytosine (5mC) and unmethylated cytosine. In some embodiments, determining the methylation pattern includes distinguishing between N6-methyladenine and unmethylated adenine. In some embodiments, determining the methylation pattern includes distinguishing between 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), and 5-carboxycytosine (5caC) and unmethylated cytosine. Examples of bisulfite sequencing include, but are not limited to, oxidized bisulfite sequencing (OX-BS-seq), Tet-assisted bisulfite sequencing (TAB-seq), and reduced bisulfite sequencing (redBS-seq).

[0059] Oxidized bisulfite sequencing (OX-BS-seq) is used to distinguish between 5mC and 5hmC by first converting 5hmC to 5fC, followed by bisulfite sequencing as described above. Tet-assisted bisulfite sequencing (TAB-seq) can also be used to distinguish between 5mC and 5hmC. In TAB-seq, 5hmC is protected by glycosylation. As mentioned earlier, 5mC is converted to 5caC using the Tet enzyme before bisulfite sequencing. Reduced bisulfite sequencing is used to distinguish 5fC from modified cytosine.

[0060] Typically, in bisulfite sequencing, nucleic acid samples are split into two aliquots, and one aliquot is treated with bisulfite. Bisulfite converts native cytosine and certain modified cytosine nucleotides (e.g., 5-formylcytosine or 5-carboxycytosine) to uracil, while other modified cytosines (e.g., 5-methylcytosine, 5-hydroxymethylcytosine) are not converted. Comparison of the nucleic acid sequences of molecules from the two aliquots indicates which cytosines were converted to uracil and which were not. Therefore, modified and unmodified cytosines can be identified. Initially splitting the sample into two aliquots is disadvantageous for samples containing only small amounts of nucleic acids and / or including heterogeneous cellular / tissue-derived samples such as body fluids containing cell-free DNA.

[0061] This disclosure provides methods for enabling bisulfite sequencing and its variations. These methods function by linking nucleic acids in a population to a capture motif (i.e., a tag that can be captured or immobilized). Capture motifs include, but are not limited to, biotin, avidin, streptavidin, nucleic acids containing a specific nucleotide sequence, haptens recognized by antibodies, and magnetically attractive particles. Extraction motifs can be members of binding pairs such as biotin / streptavidin or haptens / antibodies. In some embodiments, the capture motif attached to the analyte is captured by its binding pair, which is attached to a separable motif, such as magnetically attractive particles or large particles that can be precipitated by centrifugation. The capture motif can be any type of molecule that allows affinity separation of nucleic acids with the capture motif from nucleic acids lacking the capture motif. Example capture motifs are biotin or oligonucleotides, where biotin allows affinity separation by binding to streptavidin linked to or capable of being linked to a solid phase, and oligonucleotides allow affinity separation by binding to complementary oligonucleotides linked to or capable of being linked to a solid phase. After the capture motif is linked to the sample nucleic acid, the sample nucleic acid is used as an amplification template. After amplification, the original template remains connected to the captured portion, but the amplicon does not connect to the captured portion.

[0062] The capture portion can be ligated to the sample nucleic acid as a component of an adaptor, which can also provide binding sites for amplification and / or sequencing primers. In some methods, the sample nucleic acid is ligated to the adaptor at both ends, with both adaptors containing the capture portion. Preferably, any cytosine residues in the adaptor are modified, such as by 5-methylcytosine, to protect against bisulfite. In some cases, the capture portion is ligated to the original template via a cleavable linker (e.g., photocleavable dethiobiotin-TEG or uracil residues cleavable by the USER™ enzyme, Chem. Commun. (Camb). 21 Feb 2015; 51(15): 3266-3269), in which case the capture portion can be removed if necessary.

[0063] The amplicon is denatured and then contacted with an affinity reagent used to capture the tag. The original template binds to the affinity reagent, while the amplified nucleic acid molecules do not. Therefore, the original template can be separated from the amplified nucleic acid molecules.

[0064] After isolation or partitioning, the corresponding populations of nucleic acids (i.e., the original template and amplification products) can be subjected to bisulfite treatment, with the original template population receiving bisulfite treatment while the amplification products do not. Optionally, the amplification products can undergo bisulfite treatment while the original template population does not. After such treatment, the corresponding populations can be amplified (in the case of the original template population, this converts uracil to thymine). The populations can also undergo biotinylated probe hybridization for enrichment. The corresponding populations are then analyzed and sequences are compared to determine which cytosines are 5-methylated (or 5-hydroxymethylated) in the original sample. Detection of T nucleotides (corresponding to unmethylated cytosine converted to uracil) in the template population and C nucleotides at corresponding positions in the amplification population indicates unmodified C. The presence of C at corresponding positions in the original template and amplification population indicates the presence of modified C in the original sample.

[0065] In some embodiments, the method uses sequential DNA-seq and bisulfite-seq (BIS-seq) NGS library preparation with molecularly tagged DNA libraries. The process involves tagging of an adaptor (e.g., biotin), DNA-seq amplification of the entire library, parental molecule recovery (e.g., streptavidin bead pull-down), bisulfite conversion, and BIS-seq. In some embodiments, the method identifies 5-methylcytosine at single-base resolution through sequential NGS preparative amplification of parental library molecules with and without bisulfite treatment. This can be achieved by modifying the 5-methylated NGS adaptor used in BIS-seq (oriented adaptor; Y-shaped / forked, replaced with 5-methylcytosine) with a marker (e.g., biotin) on one of the two adaptor strands. Sample DNA molecules are linked adaptors and are amplified (e.g., by PCR). Since only parental molecules will have tagged adaptor ends, they can be selectively recovered from their amplified progeny using tag-specific capture methods (e.g., streptavidin magnetic beads). Because the parental molecules retain the 5-methylation marker, bisulfite conversion on the captured library will produce a 5-methylation state at single-base resolution during BIS-seq, thus preserving molecular information in the corresponding DNA-seq. In some embodiments, the bisulfite-treated library can be combined with the untreated library prior to enrichment / NGS by adding a sample-tagged DNA sequence in a standard multiplex NGS workflow. As with the BIS-seq workflow, bioinformatics analysis can be performed for genome alignment and 5-methylation base recognition. In summary, this method provides the ability to selectively recover parental, ligated molecules carrying the 5-methylcytosine marker after library amplification, allowing for parallel processing of bisulfite-converted DNA. This overcomes the detrimental nature of bisulfite treatment to the quality / sensitivity of DNA-seq information extracted from the workflow. With this method, the recovered ligated, parental DNA molecules (via tagged adaptors) allow for the amplification of a complete DNA library and the parallel application of treatments that induce epigenetic DNA modifications. This disclosure discusses the use of BIS-seq methods to identify 5-methylated cytosine (5-methylcytosine), but this should not be limiting. Variations of BIS-seq have been developed to identify hydroxymethylated cytosine (5hmC; OX-BS-seq, TAB-seq), formylcytosine (5fC; redBS-seq), and carboxycytosine. These methods can be implemented using sequential / parallel library preparation as described herein.

[0066] Alternative methods of analyzing modified nucleic acids This disclosure provides alternative methods for analyzing modified nucleic acids (e.g., methylated, histone-linked, and other modifications discussed above). In some such methods, a population of nucleic acids with varying degrees of modification (e.g., each nucleic acid molecule has 0, 1, 2, 3, 4, 5, or more methyl groups) is contacted with an adaptor, and the population is then stratified according to the degree of modification. The adaptor is attached to one or both ends of the nucleic acid molecules in the population. Preferably, the adaptor contains a sufficient number of different tags such that the number of tag combinations results in a high probability, e.g., 95%, 99%, or 99.9%, that two nucleic acids with the same start and end points receive different tag combinations. After attaching the adaptor, the nucleic acid is amplified from a primer that binds to a primer binding site within the adaptor. Adaptors with the same or different tags may contain the same or different primer binding sites, but preferably the adaptor contains the same primer binding site. After amplification, the nucleic acid is contacted with an agent that preferably binds the modified nucleic acid (such as those previously described). The nucleic acid is divided into at least two partitions, the difference between the at least two partitions being the degree to which the modified nucleic acid binds to the agent. For example, if the reagent has an affinity for the modified nucleic acid, the overrepresented modified nucleic acid (compared to the median representation in the population) preferentially binds to the reagent, while the underrepresented modified nucleic acid does not bind to the reagent or is more easily eluted from it. After separation, the different partitions can then undergo additional processing steps, which typically include parallel but separate additional amplification and sequence analysis. The sequence data from the different partitions can then be compared.

[0067] Nucleic acid molecules can be coupled to Y-adaptors containing primer binding sites and tags. The molecule is then amplified. The amplified molecule is then partitioned by contacting an antibody that preferentially binds to 5-methylcytosine to produce two partitions. One partition contains the unmethylated original molecule and the amplified copy that has lost methylation. The other partition contains the original DNA molecule with methylation. The two partitions are then processed and sequenced separately, and the methylated partition is further amplified. The sequence data of the two partitions can then be compared. In this example, the tag is not used to distinguish between methylated and unmethylated DNA, but rather to distinguish the different molecules within these partitions, allowing one to determine whether reads with the same start and end points are based on the same or different molecules.

[0068] This disclosure also provides methods for analyzing nucleic acid populations, wherein at least some nucleic acids contain one or more modified cytosine residues, such as 5-methylcytosine and any other modifications previously described. In these methods, the nucleic acid population is contacted with an adaptor comprising one or more cytosine residues, such as 5-methylcytosine, modified at the 5C position. Preferably, all cytosine residues in such an adaptor are also modified, or all such cytosine residues in the primer-binding region of the adaptor are modified. The adaptor is attached to both ends of the nucleic acid molecules in the population. Preferably, the adaptor contains a sufficient number of different tags such that the number of tag combinations results in a high probability, for example, 95%, 99%, or 99.9%, that two nucleic acids with the same start and end points receive different tag combinations. The primer-binding sites in such an adaptor may be the same or different, but are preferably the same. After the adaptor is attached, the nucleic acid is amplified by primers that bind to the primer-binding sites of the adaptor. The amplified nucleic acid is divided into a first aliquot and a second aliquot. The first aliquot is sequenced, with or without further processing. This determines the sequence data of molecules in the first aliquot regardless of the initial methylation state of the nucleic acid molecules. Nucleic acid molecules in the second aliquot are treated with bisulfite. This treatment converts unmodified cytosine to uracil. The bisulfite-treated nucleic acids then undergo amplification, initiated by primers targeting the original primer binding sites of the adaptor attached to the nucleic acid. Now only the nucleic acid molecules initially attached to the adaptor (unlike their amplification products) are amplified because these nucleic acids retain cytosine at the primer binding sites of the adaptor, while the amplification products have lost the methylation of these cytosine residues, which were converted to uracil during the bisulfite treatment. Therefore, only the original molecules in the population (at least some of which are methylated) undergo amplification. After amplification, these nucleic acids are sequenced. Comparison of the sequences determined from the first and second aliquots can particularly indicate which cytosine residues in the nucleic acid population have undergone methylation.

[0069] Partitioning a sample into more than one subsample; aspects of the sample; analysis of epigenetic features In some embodiments described herein, different forms of nucleic acid populations (e.g., hypermethylated and hypomethylated DNA in a sample, such as capture groups of cfDNA as described herein) can be physically partitioned based on one or more characteristics of the nucleic acids, followed by further analysis, such as differential modification or isolation of nucleotides, tagging, and / or sequencing. This approach can be used to determine, for example, whether certain sequences are hypermethylated or hypomethylated. In some embodiments, analysis is performed on hypermethylated variable epigenetic target regions to determine whether they exhibit the hypermethylation signature of tumor cells, and / or analysis is performed on hypomethylated variable epigenetic target regions to determine whether they exhibit the hypomethylation signature of tumor cells. Additionally, by partitioning heterogeneous nucleic acid populations, one can increase rare signals, for example, by enriching rare nucleic acid molecules that are more prevalent in one fraction (or partition) of the population. For example, by partitioning a sample into hypermethylated and hypomethylated nucleic acid molecules, it is easier to detect genetic variations that are present in hypermethylated DNA but less so (or absent) in hypomethylated DNA. By analyzing more than one fraction of a sample, multidimensional analysis of individual loci or nucleic acid species in the genome can be performed, thus enabling greater sensitivity.

[0070] In some cases, heterogeneous nucleic acid samples are partitioned into two or more partitions (e.g., at least three, four, five, six, or seven partitions). In some implementations, each partition is differentially tagged. The tagged partitions can then be pooled together for collective sample preparation and / or sequencing. The partitioning-tag-pooling step can occur more than once, with each round of partitioning occurring based on different characteristics (in the examples provided herein) and tagged using differential tags that distinguish them from other partitioning and partitioning methods.

[0071] Examples of features that can be used for partitioning include sequence length, methylation level, nucleosome binding, sequence mismatch, immunoprecipitation, and / or proteins that bind to DNA. The resulting partitions can include one or more of the following nucleic acid forms: single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), shorter DNA fragments, and longer DNA fragments. In some embodiments, partitioning is typically performed based on cytosine modification (e.g., cytosine methylation) or methylation, and optionally in combination with at least one additional partitioning step, which can be based on any of the aforementioned features or forms of DNA. In some embodiments, a heterogeneous population of nucleic acids is partitioned into nucleic acids having one or more epigenetic modifications and those not having said one or more epigenetic modifications. Examples of epigenetic modifications include the presence or absence of methylation, methylation level, methylation type (e.g., 5-methylcytosine with other types of methylation, such as adenine methylation and / or cytosine hydroxymethylation), and association with one or more proteins (such as histones) and the level of association. Optionally or additionally, the heterogeneous nucleic acid population can be partitioned into nucleosome-associated nucleic acid molecules and nucleosome-free nucleic acid molecules. Optionally or additionally, the heterogeneous nucleic acid population can be partitioned into single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA). Optionally or additionally, the heterogeneous nucleic acid population can be partitioned based on nucleic acid length (e.g., molecules with a maximum length of 160 bp and molecules with a length greater than 160 bp).

[0072] In some cases, each partition (representing a different nucleic acid form) is differentially labeled, and the partitions are pooled together and then sequenced. In other cases, the different forms are sequenced separately. In some implementations, different nucleic acid populations are partitioned into two or more distinct partitions. Each partition represents a different nucleic acid form, and the first partition (also called a subsample) contains DNA with a larger proportion of cytosine modifications than the second subsample. Each partition is tagged differently. The first subsample undergoes a procedure that differently affects the first and second nucleotides in the DNA of the first subsample, where the first nucleotide is modified or unmodified, and the second nucleotide is modified or unmodified, different from the first nucleotide, and the first and second nucleotides have the same base-pairing specificity. The tagged nucleic acids are pooled together and then sequenced. Sequence reads are obtained and analyzed, including computer simulation (in silico) to distinguish the first and second nucleotides in the DNA of the first subsample. Tags are used to sort reads from different partitions. Analysis can be performed at the level of individual partitions and at the level of the entire nucleic acid population to detect genetic variation. For example, the analysis may include computer simulation analysis to determine genetic variations, such as CNVs, SNVs, insertions / deletions, and fusions in the nucleic acids of each partition. In some cases, computer simulation analysis may include determining chromatin structure. For example, the coverage of sequence reads can be used to determine the location of nucleosomes in chromatin. Higher coverage may be associated with higher nucleosome occupancy in a genomic region, while lower coverage may be associated with lower nucleosome occupancy or nucleosome depleted regions (NDRs).

[0073] Samples may include nucleic acids with various modifications, including post-replication modifications of nucleotides and binding to one or more proteins (typically non-covalent).

[0074] In the implementation scheme, the nucleic acid population is a population of nucleic acids obtained from serum, plasma, or blood samples of subjects suspected of having vegetations, tumors, or cancer, or previously diagnosed with vegetations, tumors, or cancer. The nucleic acid population includes nucleic acids with different levels of methylation. Methylation can occur by any one or more post-replication or post-transcriptional modifications. Post-replication modifications include modifications to nucleotide cytosine, particularly at the 5-position of the nucleobase, such as 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, and 5-carboxycytosine. The affinity agent can be an antibody with desired specificity, a natural binding partner or a variant thereof (Bock et al., Nat Biotech 28: 1106-1114 (2010); Song et al., Nat Biotech 29: 68-72 (2011)), or, for example, an artificial peptide selected for specificity to a given target via phage display.

[0075] Examples of capture fractions envisioned herein include methyl-binding domains (MBDs) and methyl-binding proteins (MBPs) as described herein, including proteins such as MeCP2 and antibodies that preferentially bind to 5-methylcytosine. Similarly, partitioning of different forms of nucleic acids can be performed using histone-binding proteins that can separate histone-bound nucleic acids from free or unbound nucleic acids. Examples of histone-binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48, and SANT domain peptides. For some affinity agents and modifications, although binding to the agent may occur substantially all-or-nothing depending on whether the nucleic acid is modified, separation may be to a certain extent. In such cases, nucleic acids overrepresented in a modification bind to the agent to a greater extent than nucleic acids underrepresented in the modification. Optionally, modified nucleic acids may bind in an all-or-nothing manner. However, various levels of modification can then be eluted sequentially from the binding agent.

[0076] For example, in some implementations, partitioning can be binary or based on the degree / level of modification. For instance, a methyl-binding domain protein (e.g., the MethylMiner methylated DNA enrichment kit (ThermoFisherScientific)) can be used to partition all methylated fragments with unmethylated fragments. Subsequently, additional partitioning can include eluting fragments with different methylation levels by adjusting the salt concentration of the solution containing the methyl-binding domain and the binding fragment. As the salt concentration increases, fragments with higher methylation levels are eluted. In some cases, the final partitioning represents nucleic acids with different degrees of modification (overrepresentation or underrepresentation). Overrepresentation and underrepresentation can be defined by the number of modifications a nucleic acid carries relative to the median number of modifications per strand in the population. For example, if the median number of 5-methylcytosine residues in the nucleic acids in a sample is 2, then nucleic acids containing more than two 5-methylcytosine residues are overrepresented, while nucleic acids with one or zero 5-methylcytosine residues are underrepresented. The purpose of affinity separation is to enrich over-represented nucleic acids in the binding phase and under-represented nucleic acids in the non-binding phase (i.e., in solution). The nucleic acids in the binding phase can be eluted before subsequent processing.

[0077] When using the MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific), sequential elution can be used to separate methylated regions with different levels of methylation. For example, low-methylated regions (e.g., unmethylated) can be separated from methylated regions by contacting a group of nucleic acids with MBDs attached to magnetic beads from the kit. The beads are used to isolate methylated nucleic acids from unmethylated nucleic acids. Subsequently, one or more elution steps are performed sequentially to elute nucleic acids with different methylation levels. For example, the first group of methylated nucleic acids can be eluted at a salt concentration of 160 mM or higher, such as at least 150 mM, at least 200 mM, at least 300 mM, at least 400 mM, at least 500 mM, at least 600 mM, at least 700 mM, at least 800 mM, at least 900 mM, at least 1000 mM, or at least 2000 mM. After such methylated nucleic acids have been eluted, magnetic separation is again used to separate nucleic acids with higher levels of methylation from those with lower levels of methylation. The elution and magnetic separation steps can be repeated to produce various partitions, such as hypomethylated partitions (representing no methylation), methylated partitions (representing low methylation levels), and hypermethylated partitions (representing high methylation levels).

[0078] In some methods, nucleic acids bound to an affinity separator undergo a washing step. This washing step removes nucleic acids that are weakly bound to the affinity agent. Such nucleic acids can be enriched with a degree of modification close to the mean or median (i.e., an intermediate value between nucleic acids that remain bound to the solid and those that do not upon initial contact with the agent). Affinity separation results in at least two, and sometimes three or more, partitions of nucleic acids with different degrees of modification. While the partitions remain separate, nucleic acids from at least one partition, and usually two or three (or more) partitions, are attached to nucleic acid tags, which are typically provided as part of an adaptor, and the nucleic acids in different partitions receive different tags that distinguish members of one partition from members of another. Tags attached to nucleic acid molecules in the same partition can be the same or different from each other. However, if they are different, the tags can have a partially shared encoding to identify the molecules to which they are attached as belonging to a particular partition. For more details on partitioning nucleic acid samples based on characteristics such as methylation, see WO2018 / 119452, which is incorporated herein by reference. In some implementations, nucleic acid molecules can be hierarchically separated into different partitions based on whether they bind to a specific protein or a fragment thereof and whether they do not bind to that specific protein or a fragment thereof.

[0079] Nucleic acid molecules can be fractionated based on DNA-protein binding. Protein-DNA complexes can be fractionated based on specific properties of the protein. Examples of such properties include various epitopes, modifications (e.g., histone methylation or acetylation), or enzymatic activities. Examples of proteins that can bind DNA and be used as the basis for fractionation can include, but are not limited to, protein A and protein G. Any suitable method can be used for fractionation of nucleic acid molecules based on protein-binding regions. Examples of methods for fractionation of nucleic acid molecules based on protein-binding regions include, but are not limited to, SDS-PAGE, chromatin immunoprecipitation (ChIP), heparin chromatography, and asymmetric field flow fractionation (AF4).

[0080] In some implementations, nucleic acid fractionation is performed by contacting the nucleic acid with the methylation-binding domain (“MBD”) of a methylation-binding protein (“MBP”). The MBD binds to 5-methylcytosine (5mC). The MBD is coupled to paramagnetic beads (such as Dynabeads® M-280 streptavidin) via a biotin linker. Fractionation into fractions with different degrees of methylation can be performed by eluting the fractions with increasing NaCl concentrations.

[0081] An exemplary method for molecular tag identification of MBD bead partition libraries using NGS is as follows: Extracted DNA samples (e.g., plasma DNA extracted from human samples) are physically partitioned using a methyl-binding domain protein-bead purification kit, retaining all eluents from the process for downstream processing.

[0082] Differential molecular tags and NGS-feasible linker sequences were applied in parallel to each partition. For example, hypermethylated, residual methylated ('washed'), and hypomethylated partitions were linked to NGS linkers with molecular tags.

[0083] All molecularly tagged partitions were reassembled and subsequently amplified using adaptor-specific DNA primer sequences.

[0084] Enrich / hybridize the recombined and amplified total library to target genomic regions of interest (e.g., cancer-specific genetic variations and differentially methylated regions).

[0085] The enriched total DNA library was re-amplified and tagged with samples. Different samples were pooled and subjected to multiplex assays on an NGS instrument.

[0086] Bioinformatics analysis of NGS data was performed, using molecular tags to identify unique molecules and deconvolving samples into molecules with differentially expressed molecular dividers (MBDs). This analysis can simultaneously produce information on the relative 5-methylcytosine content of genomic regions with standard gene sequencing / variation detection.

[0087] Examples of MBPs envisioned in this article include, but are not limited to: (a) MeCP2, a protein that preferentially binds to 5-methyl-cytosine compared to binding to unmodified cytosine.

[0088] (b) RPL26, PRP8 and DNA mismatch repair protein MHS6 preferentially bind to 5-hydroxymethyl-cytosine compared to binding to unmodified cytosine.

[0089] (c) FOXK1, FOXK2, FOXP1, FOXP4 and FOXI3 preferentially bind 5-formyl-cytosine compared to binding unmodified cytosine (Iurlaro et al., Genome Biol. 14: R119 (2013)).

[0090] (d) Antibodies specific to one or more methylated nucleotide bases.

[0091] Typically, elution varies with the number of methylation sites per molecule, with molecules having more methylation eluting at increasing salt concentrations. To elute DNA into different populations based on the degree of methylation, a series of elution buffers with increasing NaCl concentrations can be used. Salt concentrations can range from about 100 nM to about 2500 mM NaCl. In one embodiment, the process produces three (3) partitions. Molecules are contacted with a solution of a first salt concentration, and this solution contains molecules containing methyl-binding domains that can attach to capture moieties such as streptavidin. At the first salt concentration, one population of molecules will bind to MBD, and one population will remain unbound. The unbound population can be separated into a “hypomethylated” population. For example, the first partition representing hypomethylated DNA is the partition that remains unbound at low salt concentrations (e.g., 100 mM or 160 mM). The second partition representing moderately methylated DNA is eluted using a moderate salt concentration (e.g., between 100 mM and 2000 mM). This is also separated from the sample. The third partition, representing the highly methylated form of DNA, was eluted with a high salt concentration (e.g., at least about 2000 mM).

[0092] This disclosure also provides methods for analyzing nucleic acid populations, wherein at least some nucleic acids contain one or more modified cytosine residues, such as 5-methylcytosine and any other modifications previously described. In these methods, after partitioning, a sample of nucleic acid subsamples is contacted with an adaptor containing one or more cytosine residues modified at the 5C position (such as 5-methylcytosine). Preferably, all cytosine residues in such an adaptor are also modified, or all such cytosine residues in the primer-binding region of the adaptor are modified. The adaptor is attached to both ends of nucleic acid molecules in the population. Preferably, the adaptor contains a sufficient number of different tags such that the number of tag combinations results in a high probability, for example, 95%, 99%, or 99.9%, that two nucleic acids with the same start and end points receive different tag combinations. The primer-binding sites in such an adaptor can be the same or different, but are preferably the same. After the adaptor is attached, the nucleic acid is amplified by primers that bind to the primer-binding sites of the adaptor. The amplified nucleic acid is divided into a first aliquot and a second aliquot. With or without further processing, the first aliquot is sequenced. This determines the sequence data of molecules in the first aliquot regardless of the initial methylation state of the nucleic acid molecules. Nucleic acid molecules in the second aliquot undergo a procedure that differently affects the first and second nucleobases in the DNA, where the first nucleobase includes cytosine modified at position 5, and the second nucleobase includes unmodified cytosine. This procedure can be bisulfite treatment or another procedure that converts unmodified cytosine to uracil. The nucleic acids that have undergone this procedure are then amplified using primers targeting the original primer binding sites of the adaptor attached to the nucleic acid. Now only the nucleic acid molecules initially attached to the adaptor (unlike their amplification products) are amplified because these nucleic acids retain cytosine at the primer binding sites of the adaptor, while the amplification products have lost the methylation of these cytosine residues, which have been converted to uracil during bisulfite treatment. Therefore, only the original molecules in the population (at least some of which are methylated) undergo amplification. After amplification, these nucleic acids are sequenced. Comparison of sequences determined from the first and second aliquots can particularly indicate which cytosines in the nucleic acid population have undergone methylation.

[0093] Such analysis can be performed using the following exemplary procedure. After partitioning, the two ends of the methylated DNA are ligated to a Y-shaped adaptor containing primer binding sites and a tag. The cytosine in the adaptor is modified at position 5 (e.g., 5-methylated). This modification of the adaptor serves to protect the primer binding sites in subsequent transformation steps (e.g., bisulfite treatment, TAP transformation, or any other transformation that does not affect the modified cytosine but affects the unmodified cytosine). After adaptor attachment, the DNA molecule is amplified. The amplified product is divided into two aliquots for sequencing with and without transformation. The untransformed aliquot may undergo sequence analysis with or without further treatment. The other aliquot undergoes a procedure that differently affects the first and second nucleotides in the DNA, where the first nucleotide includes the cytosine modified at position 5, and the second nucleotide includes the unmodified cytosine. This procedure may be bisulfite treatment or another procedure to convert the unmodified cytosine to uracil. When contacted with primers specific to the original primer binding site, only primer binding sites protected by cytosine modification can support amplification. Therefore, only the original molecule, not copies from the first amplification, undergoes further amplification. The further amplified molecules then undergo sequence analysis. The sequences from the two aliquots can then be compared. In the isolation scheme discussed above, the nucleic acid tag in the adaptor is not used to distinguish between methylated and unmethylated DNA, but rather to distinguish nucleic acid molecules within the same partition.

[0094] Subjecting a first subsample to a procedure that differentially affects a first nucleobase in DNA and a second nucleobase in DNA of the first subsample Enrichment / capture step; amplification; adapters; barcodes The method disclosed herein includes steps of subjecting a first subsample to procedures that differently affect a first nucleotide and a second nucleotide in the DNA of the first subsample, wherein the first nucleotide is a modified or unmodified nucleotide, the second nucleotide is a modified or unmodified nucleotide different from the first nucleotide, and the first and second nucleotides have the same base-pairing specificity. In some embodiments, if the first nucleotide is modified or unmodified adenine, then the second nucleotide is modified or unmodified adenine; if the first nucleotide is modified or unmodified cytosine, then the second nucleotide is modified or unmodified cytosine; if the first nucleotide is modified or unmodified guanine, then the second nucleotide is modified or unmodified guanine; if the first nucleotide is modified or unmodified thymine, then the second nucleotide is modified or unmodified thymine (wherein, for the purposes of this step, modified and unmodified uracil are included in modified thymine).

[0095] In some embodiments, the first nucleobase is a modified or unmodified cytosine, and then the second nucleobase is a modified or unmodified cytosine. For example, the first nucleobase may comprise unmodified cytosine (C), and the second nucleobase may comprise one or more of 5-methylcytosine (mC) and 5-hydroxymethylcytosine (hmC). Alternatively, the second nucleobase may comprise C, and the first nucleobase may comprise one or more of mC and hmC. Other combinations are also possible, such as those indicated, for example, in the overview above and the discussion below, such as where one of the first and second nucleobases comprises mC and the other comprises hmC.

[0096] In some implementations, the procedure that differently affects the first and second nucleobases in the DNA of the first subsample includes bisulfite conversion. Bisulfite treatment converts unmodified cytosine and certain modified cytosine nucleotides (e.g., 5-formylcytosine (fC) or 5-carboxycytosine (caC)) into uracil, while other modified cytosines (e.g., 5-methylcytosine and 5-hydroxymethylcytosine) are not converted. Therefore, in the case of bisulfite conversion, the first nucleobase comprises one or more of unmodified cytosine, 5-formylcytosine, 5-carboxycytosine, or other bisulfite-affected cytosine forms, and the second nucleobase may comprise one or more of mC and hmC, such as mC and optionally hmC. Sequencing of the bisulfite-treated DNA identifies the position read as a cytosine as the mC position or the hmC position. Simultaneously, positions read as T are identified as T or bisulfite-susceptible forms of C, such as unmodified cytosine, 5-formylcytosine, or 5-carboxycytosine. Therefore, bisulfite conversion of the first subsample as described herein facilitates the identification of positions containing mC or hmC using sequence reads obtained from the first subsample. For an exemplary description of bisulfite conversion, see, for example, Moss et al., Nat Commun. 2018; 9: 5068.

[0097] In some embodiments, the procedures that differently affect the first and second nucleotides of the DNA in the first subsample include oxidized bisulfite (Ox-BS) conversion. In some embodiments, the procedures that differently affect the first and second nucleotides of the DNA in the first subsample include Tet-assisted substituted borane reducing agent conversion, optionally wherein the substituted borane reducing agent is 2-methylpyridineborane, pyridineborane, tert-butylamineborane, or aminoborane. In some embodiments, the procedures that differently affect the first and second nucleotides of the DNA in the first subsample include chemically assisted substituted borane reducing agent conversion, optionally wherein the substituted borane reducing agent is 2-methylpyridineborane, pyridineborane, tert-butylamineborane, or aminoborane. In some implementations, procedures that differently affect the first and second nucleobases in the DNA of the first subsample include APOBEC-coupled epigenetic (ACE) transformation.

[0098] In some implementations, procedures that differently affect the first and second nucleotides in the DNA of the first subsample include enzymatic conversion of the first nucleotide, for example, as in EM-Seq. See, for example, Vaisvila R, et al. (2019) EM-seq: Detection of DNA methylation at single base resolution from picograms of DNA. bioRxiv; DOI: 10.1101 / 2019.12.20.884692, available at www.biorxiv.org / content / 10.1101 / 2019.12.20.884692v1. For example, TET2 and T4-βGT can be used to convert 5mC and 5hmC into substrates that cannot be deaminated by deaminases (e.g., APOBEC3A), and then the deaminase (e.g., APOBEC3A) can be used to deaminate unmodified cytosine, converting it to uracil.

[0099] In some implementations, the procedure that differently affects the first nucleobase in the DNA of the first subsample and the second nucleobase in the DNA includes separating the DNA that initially contains the first nucleobase from the DNA that initially does not contain the first nucleobase.

[0100] In some embodiments, the first nucleobase is a modified or unmodified adenine, and the second nucleobase is a modified or unmodified adenine. In some embodiments, the modified adenine is N6-methyladenine (mA). In some embodiments, the modified adenine is one or more of N6-methyladenine (mA), N6-hydroxymethyladenine (hmA), or N6-formyladenine (fA).

[0101] Techniques including methylated DNA immunoprecipitation (MeDIP) can be used to separate DNA containing modified bases (such as mA) from other DNA. See, for example, Kumar et al., Frontiers Genet. 2018; 9: 640; Greer et al., Cell 2015; 161: 868-878. Antibodies specific to mA are described in Sun et al., Bioessays 2015; 37:1155-62. Antibodies against various modified nucleobases (such as thymine / uracil forms, including halogenated forms such as 5-bromouracil) are commercially available. Various modified bases can also be detected based on changes in their base pairing specificity. For example, hypoxanthine is a modified form of adenine that can be produced by deamination and is read as G in sequencing. See, for example, U.S. Patent 8,486,630; Brown, Genomes, 2nd Edition, John Wiley & Sons, Inc., New York, NY, 2002, Chapter 14, “Mutation, Repair, and Recombination”.

[0102] Capture panel In some embodiments, the methods disclosed herein include the step of capturing one or more target regions of DNA, such as cfDNA. Capture can be performed using any suitable method known in the art. In some embodiments, capture includes contacting the DNA to be captured with a set of target-specific probes. The target-specific probe set may have any of the characteristics of the target-specific probe set described herein, including but not limited to the embodiments set forth above and the characteristics described below in the probe-related section. One or more subsamples prepared during the methods disclosed herein may be captured. In some embodiments, DNA is captured from at least a first subsample or a second subsample, e.g., at least a first subsample and a second subsample. If the first subsample undergoes a separation step (e.g., separating DNA initially containing a first nucleobase (e.g., hmC) from DNA initially not containing a first nucleobase, such as hmC-seal), capture can be performed on any one, any two, or all of the DNA initially containing a first nucleobase (e.g., hmC), the DNA initially not containing a first nucleobase, and the second subsample. In some embodiments, the subsamples are differentially tagged (e.g., as described herein) and then pooled prior to capture.

[0103] The capture step can be performed using conditions suitable for a specific nucleic acid hybridization, which are typically dependent to some extent on the characteristics of the probe, such as length, base composition, etc. Given the general knowledge of nucleic acid hybridization in the art, those skilled in the art will be familiar with appropriate conditions. In some embodiments, a complex of the target-specific probe and DNA is formed.

[0104] In some embodiments, the method described herein includes capturing more than one set of target regions of cfDNA obtained from a test subject. Target regions include epigenetic target regions that may exhibit differences in methylation levels and / or fragmentation patterns, depending on whether they originate from tumor cells or healthy cells. Target regions also include sequence-variable target regions that may exhibit sequence differences, depending on whether they originate from tumor cells or healthy cells. The capture step produces a capture set of cfDNA molecules, and within the capture set of cfDNA molecules, cfDNA molecules corresponding to the sequence-variable target region set are captured with a greater capture yield than cfDNA molecules corresponding to the epigenetic target region set. For further discussion of the capture steps, capture yield, and related aspects, see WO2020 / 160414, which is incorporated herein by reference for all purposes.

[0105] In some implementations, the method described herein includes contacting cfDNA obtained from a test subject with a target-specific probe set, wherein the target-specific probe set is configured to capture cfDNA corresponding to a sequence-variable target region set with a greater capture yield than cfDNA corresponding to an epigenetic target region set.

[0106] Capturing cfDNA corresponding to sequence-variable target regions at a higher capture yield than that corresponding to epigenetic target regions is beneficial because analyzing sequence-variable target regions with sufficient confidence or accuracy may require a greater sequencing depth than analyzing epigenetic target regions. The amount of data required to determine fragmentation patterns (e.g., testing for perturbations of transcription start sites or CTCF binding sites) or fragment abundance (e.g., in hypermethylated and hypomethylated regions) is generally less than the amount of data required to determine the presence or absence of cancer-related sequence mutations. Capturing target regions at different yields can facilitate sequencing target regions to different sequencing depths within the same sequencing run (e.g., using pooled mixtures and / or within the same sequencing pool).

[0107] In various embodiments, the method further includes sequencing the captured cfDNA to, for example, different sequencing depths for epigenetic target groups and sequence-variable target groups, consistent with those discussed herein. In some embodiments, the complex of the target-specific probe and DNA is separated from DNA not bound to the target-specific probe. For example, in cases where the target-specific probe is covalently or nonvalently bound to a solid support, washing or aspiration steps can be used to separate the unbound material. Alternatively, chromatography can be used where the complex has different chromatographic properties than the unbound material (e.g., where the probe contains ligands that bind to chromatographic resins).

[0108] As discussed in detail elsewhere herein, a target-specific probe set may include more than one set, such as probes for a sequence-variable target set and probes for an epigenetic target set. In some such embodiments, the capture step is performed simultaneously using probes for sequence-variable targets and probes for epigenetic targets in the same container; for example, probes for both sequence-variable and epigenetic target sets are in the same composition. This method provides a relatively more efficient workflow. In some embodiments, the concentration of probes for sequence-variable targets is greater than the concentration of probes for epigenetic target sets.

[0109] Optionally, a capture step is performed in a first container with a sequence-variable target region probe set and in a second container with an epigenetic target region probe set, or a contact step is performed at a first time and in the first container with a sequence-variable target region probe set and at a second time before or after the first time with an epigenetic target region probe set. This method allows for the preparation of separate first and second compositions comprising captured DNA corresponding to the sequence-variable target region set and captured DNA corresponding to the epigenetic target region set. The compositions can be processed individually as desired (e.g., graded based on methylation, as described elsewhere herein) and recombine in appropriate proportions to provide material for further processing and analysis, such as sequencing.

[0110] In some implementations, DNA is amplified. In some implementations, amplification occurs before the capture step. In some implementations, amplification occurs after the capture step.

[0111] In some implementations, the DNA contains an adaptor. This can be performed simultaneously with the amplification process, for example, by providing the adaptor in the 5' portion of the primer, as described above. Alternatively, the adaptor can be added by other methods such as ligation.

[0112] In some implementations, the DNA contains a tag, which may be a barcode or contain a barcode. The tag can aid in identifying the origin of the nucleic acid. For example, a barcode can be used to allow identification of the DNA's origin, such as a subject, after pooling more than one sample for parallel sequencing. This can be performed concurrently with the amplification procedure, for example, by providing a barcode in the 5' portion of the primer, as described above. In some implementations, the adaptor and the tag / barcode are provided by the same primer or primer set. For example, the barcode may be located at the 3' of the adaptor and the 5' of the target hybridization portion of the primer. Alternatively, the barcode can be added by other methods, such as ligation, optionally along with the adaptor in the same ligation substrate.

[0113] Further details regarding amplification, labeling, and barcodes are discussed in the following “General Characteristics of the Method” section, and these details can be combined, to a degree of feasibility, with any of the aforementioned implementation schemes and the implementation schemes described in the “Introduction and Overview” section.

[0114] Epigenetic target region panel In some embodiments, a capture set of DNA (e.g., cfDNA) is provided. For the disclosed methods, the capture set of DNA may be provided, for example, by performing a capture step after a partitioning step as described herein. The capture set may include DNA corresponding to a sequence-variable target region group, DNA corresponding to an epigenetic target region group, or a combination thereof. In some embodiments, the amount of captured sequence-variable target region DNA is greater than the amount of captured epigenetic target region DNA when normalized for differences in target region size (footprint size).

[0115] Optionally, a first capture group and a second capture group may be provided, comprising DNA corresponding to a sequence-variable target region group and DNA corresponding to an epigenetic target region group, respectively. The first capture group and the second capture group may be combined to provide a combined capture group.

[0116] In some embodiments, which include capture groups (including capture groups of combinations as discussed above) that include DNA corresponding to sequence-variable target groups and epigenetic target groups, the DNA corresponding to the sequence-variable target groups may be present at a higher concentration than the DNA corresponding to the epigenetic target groups, for example, 1.1 to 1.2 times higher, 1.2 to 1.4 times higher, 1.4 to 1.6 times higher, 1.6 to 1.8 times higher, 1.8 to 2.0 times higher, or 2.0 to 2.2 times higher. Large concentration, 2.2 to 2.4 times larger concentration, 2.4 to 2.6 times larger concentration, 2.6 to 2.8 times larger concentration, 2.8 to 3.0 times larger concentration, 3.0 to 3.5 times larger concentration, 3.5 to 4.0 times larger concentration, 4.0 to 4.5 times larger concentration, 4.5 to 5.0 times larger concentration, 5.0 to 5.5 times larger concentration, 5.5 to 6.0 times larger concentration, 6.0 to 6.5 times larger concentration, 6.5 to 7.0 times larger concentration, 7.0 to 7 0.5 times larger concentration, 7.5 to 8.0 times larger concentration, 8.0 to 8.5 times larger concentration, 8.5 to 9.0 times larger concentration, 9.0 to 9.5 times larger concentration, 9.5 to 10.0 times larger concentration, 10 to 11 times larger concentration, 11 to 12 times larger concentration, 12 to 13 times larger concentration, 13 to 14 times larger concentration, 14 to 15 times larger concentration, 15 to 16 times larger concentration, 16 to 17 times larger concentration, 17 to 18 times larger concentration, 18 times Concentrations ranging from 19 to 20 times larger, 20 to 30 times larger, 30 to 40 times larger, 40 to 50 times larger, 50 to 60 times larger, 60 to 70 times larger, 70 to 80 times larger, 80 to 90 times larger, 90 to 100 times larger, 10 to 20 times larger, 10 to 40 times larger, 10 to 50 times larger, 10 to 70 times larger, or 10 to 100 times larger. The degree of concentration difference is calculated based on normalization for the target footprint size, as discussed in the definitions section.

[0117] Hypermethylated variable target regions An epigenetic target set may include one or more types of target regions that can distinguish DNA from DNA from vegetative (e.g., tumor or cancer) cells from DNA from healthy cells (e.g., non-vegetative circulating cells). Example types of such regions are discussed in detail herein. An epigenetic target set may also include one or more control regions, such as those described herein. In some embodiments, the epigenetic target set has a footprint of at least 100 kb, for example, at least 200 kb, at least 300 kb, or at least 400 kb. In some embodiments, the epigenetic target set has a footprint in the range of 100-1000 kb, for example, 100-200 kb, 200-300 kb, 300-400 kb, 400-500 kb, 500-600 kb, 600-700 kb, 700-800 kb, 800-900 kb, and 900-1000 kb.

[0118] Hypomethylated variable target regions In some implementations, the epigenetic target set includes one or more hypermethylated variable target regions. Typically, a hypermethylated variable target region refers to a region in which, for example in a cfDNA sample, an observed increase in methylation levels indicates an increased likelihood that the sample (e.g., cfDNA) contains DNA produced by invasive cells (such as tumor cells or cancer cells). For example, hypermethylation of tumor suppressor gene promoters has been repeatedly observed. See, for example, Kang et al., Genome Biol. 18:53 (2017) and the references cited therein. In examples, a hypermethylated variable target region may include a region in which the methylation is not necessarily different from that of DNA from the same type of healthy tissue, but is indeed different from that of typical cfDNA in healthy subjects (e.g., having more methylation). For example, such a hypermethylated variable target region can be used, at least in part, to detect cancer when the presence of cancer leads to increased cell death (such as apoptosis corresponding to the tissue type of cancer). In some implementations, the hypermethylated variable target region includes one or more genomic regions in which the methylation status of cfDNA molecules in these regions is not different from that of cfDNA from healthy subjects, but the presence / increased amount of hypermethylated cfDNA in these regions indicates a specific tissue type (e.g., cancer origin) and is presented as cfDNA entering circulation due to increased apoptosis (e.g., tumor shedding).

[0119] Hypermethylated target regions can be obtained, for example, from the Cancer Genome Atlas. Kang et al., Genome Biology 18:53 (2017) describe a probabilistic approach to constructing a cancer locator using hypermethylated target regions from the breast, colon, kidney, liver, and lung. In some embodiments, the hypermethylated target regions can be specific to one or more types of cancer. Thus, in some embodiments, the hypermethylated target regions comprise one, two, three, four, or five subgroups of hypermethylated target regions that collectively exhibit hypermethylation in one, two, three, four, or five of the following cancers: breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer.

[0120] In some embodiments, probes targeting an epigenetic target set include probes specific to one or more hypermethylated variable target regions. Hypermethylated variable target regions can be any of the hypermethylated variable target regions listed above. For example, in some embodiments, probes specific to hypermethylated variable target regions include probes specific to more than one locus. In some embodiments, for each locus included as a target region, there may be one or more probes having hybridization sites that bind between the transcription start site and the stop codon (or the final stop codon for genes undergoing alternative splicing). In some embodiments, one or more probes bind within 300 bp (e.g., within 200 bp or 100 bp) at the listed locations. In some embodiments, probes have hybridization sites that overlap with the locations listed above. In some implementations, probes specific to hypermethylated target regions include probes specific to one, two, three, four, or five subgroups of hypermethylated target regions that collectively exhibit hypermethylation in one, two, three, four, or five of the following cancers: breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer.

[0121] Composition comprising captured DNA Universal hypomethylation is a phenomenon commonly observed in a variety of cancers. See, for example, Hon et al., GenomeRes. 22:246-258 (2012) (breast cancer); Ehrlich, Epigenomics 1:239-259 (2009) (a review article noting observations of hypomethylation in colon cancer, ovarian cancer, prostate cancer, leukemia, hepatocellular carcinoma, and cervical cancer). For example, regions that are normally methylated in healthy cells (such as repetitive elements (e.g., LINE1 elements, Alu elements, centromere tandem repeats, pericentromere tandem repeats, and satellite DNA) and intergenic regions) may show reduced methylation in tumor cells. Therefore, in some embodiments, the epigenetic target set includes hypomethylated variable target regions, where the observed reduction in methylation levels indicates an increased likelihood that the sample (e.g., cfDNA) contains DNA generated by cytoplasmic cells (such as tumor cells or cancer cells). In examples, a hypomethylated variable target region may include a region whose methylation state is not necessarily different in cancerous tissue relative to DNA from the same type of healthy tissue, but is indeed different in methylation (e.g., less methylated) relative to cfDNA typical of healthy subjects. For example, such a hypomethylated variable target region can be used at least in part to detect cancer when the presence of cancer leads to increased cell death (such as apoptosis corresponding to the tissue type of cancer). In some embodiments, the hypomethylated variable target region comprises one or more genomic regions in which the methylation state of cfDNA molecules in these regions is not different from cfDNA from healthy subjects, but the amount of hypomethylated cfDNA present / increased in these regions indicates a specific tissue type (e.g., cancer origin) and is presented as cfDNA entering circulation accompanied by increased apoptosis (e.g., tumor shedding).

[0122] In some embodiments, the hypomethylated variable target region includes repeating elements and / or intergenic regions. In some embodiments, repeating elements include one, two, three, four, or five of the following: LINE1 elements, Alu elements, centromere tandem repeat sequences, pericentromere tandem repeat sequences, and / or satellite DNA.

[0123] Example specific genomic regions exhibiting cancer-related hypomethylation include nucleotides 8403565-8953708 and 151104701-151106035 on human chromosome 1. In some embodiments, the hypomethylated variable target region overlaps with or includes one or both of these regions.

[0124] In some implementations, probes targeting a set of epigenetic target regions include probes specific to one or more hypomethylated variable target regions. Hypomethylated variable target regions can be any of the hypomethylated target regions listed above. For example, probes specific to one or more hypomethylated variable target regions can include probes targeting regions such as repetitive elements (e.g., LINE1 elements, Alu elements, centromere tandem repeats, pericentromere tandem repeats, and satellite DNA) and intergenic regions that are normally methylated in healthy cells but may exhibit reduced methylation in tumor cells.

[0125] In some embodiments, probes specific to hypomethylated variable target regions include probes specific to repetitive elements and / or intergenic regions. In some embodiments, probes specific to repetitive elements include probes specific to one, two, three, four, or five of the following: LINE1 elements, Alu elements, centromere tandem repeat sequences, pericentromere tandem repeat sequences, and / or satellite DNA.

[0126] Example probes specific to genomic regions exhibiting cancer-related hypomethylation include probes specific to nucleotides 8403565-8953708 and / or 151104701-151106035 of human chromosome 1. In some embodiments, probes specific to hypomethylated variable target regions include probes specific to regions that overlap with or contain nucleotides 8403565-8953708 and / or 151104701-151106035 of human chromosome 1.

[0127] Probes used to detect this set of regions may include probes for detecting genomic regions of interest (hotspot regions) and nucleosome-sensing probes (e.g., KRAS codons 12 and 13), and can be designed to optimize capture based on analysis of cfDNA coverage and fragment size variation and GC sequence composition affected by nucleosome binding patterns. The regions used in this paper may also include non-hotspot regions optimized based on nucleosome location and GC models.

[0128] Subjects In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject with cancer. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject suspected of having cancer. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject with a tumor. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject suspected of having a tumor. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject with a growth. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject suspected of having a growth. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject in remission from a tumor, cancer, or growth (e.g., after chemotherapy, surgical resection, radiation, or a combination thereof). In any of the foregoing embodiments, the cancer, tumor, or growth, or suspected cancer, tumor, or growth, may be of the lung, colon, rectum, kidney, breast, prostate, or liver. In some embodiments, the cancer, tumor, or growth, or suspected cancer, tumor, or growth, is of the lung. In some embodiments, the cancer, tumor, or growth, or suspected cancer, tumor, or growth, is of the colon or rectum. In some embodiments, the cancer, tumor, or growth, or suspected cancer, tumor, or growth, is of the breast. In some embodiments, the cancer, tumor, or growth, or suspected cancer, tumor, or growth, is of the prostate. In any of the foregoing embodiments, the subject may be a human subject.

[0129] In some embodiments, the sequence-variable target probe set has a footprint of at least 0.5 kb (e.g., at least 1 kb, at least 2 kb, at least 5 kb, at least 10 kb, at least 20 kb, at least 30 kb, or at least 40 kb). In some embodiments, the epigenetic target probe set has a footprint in the range of 0.5–100 kb (e.g., 0.5–2 kb, 2–10 kb, 10–20 kb, 20–30 kb, 30–40 kb, 40–50 kb, 50–60 kb, 60–70 kb, 70–80 kb, 80–90 kb, and 90–100 kb).

[0130] Computer system, processing of real world evidence (RWE) This document provides a combination of a first group and a second group comprising captured DNA. The first group may comprise or be derived from DNA with a greater proportion of cytosine modifications than the second group. The first group may comprise a form of a first nucleobase initially present in the DNA with altered base pairing specificity and a second nucleobase without altered base pairing specificity, wherein the form of the first nucleobase initially present in the DNA before the alteration of base pairing specificity is a modified or unmodified nucleobase, and the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the form of the first nucleobase initially present in the DNA before the alteration of base pairing specificity and the second nucleobase have the same base pairing specificity. The second group does not comprise a form of the first nucleobase initially present in the DNA with altered base pairing specificity. In some embodiments, the cytosine modification is cytosine methylation. In some embodiments, the first nucleobase is modified or unmodified cytosine, and the second nucleobase is modified or unmodified cytosine. The first and second nucleobases can be any nucleobases discussed in the overview or regarding procedures that subject the first subsample to different effects on the first and second nucleobases in the DNA of the first subsample.

[0131] In some implementations, the first group includes sequence tags selected from one or more sequence tags in the first group, and the second group includes sequence tags selected from one or more sequence tags in the second group, wherein the sequence tags in the second group are different from the sequence tags in the first group. The sequence tags may include barcodes.

[0132] In some embodiments, the first population comprises protected hmC, such as glucosylated hmC. In some embodiments, the first population undergoes any of the transformation procedures discussed herein, such as bisulfite transformation, Ox-BS transformation, TAB transformation, ACE transformation, TAP transformation, TAPSβ transformation, or CAP transformation. In some embodiments, the first population is protected by hmC followed by deamination of mC and / or C. In some combined embodiments, the first population comprises or is derived from DNA having a larger proportion of cytosine modification than the second population, and the first population comprises both first and second subpopulations, and the first nucleotide is a modified or unmodified nucleotide, the second nucleotide is a modified or unmodified nucleotide different from the first nucleotide, and the first and second nucleotides have the same base-pairing specificity. In some embodiments, the second population does not contain the first nucleotide. In some embodiments, the first nucleotide is a modified or unmodified cytosine, and the second nucleotide is a modified or unmodified cytosine, optionally wherein the modified cytosine is mC or hmC. In some embodiments, the first nucleobase is a modified or unmodified adenine, and the second nucleobase is a modified or unmodified adenine, optionally wherein the modified adenine is mA.

[0133] In some embodiments, the first nucleobase (e.g., modified cytosine) is biotinylated. In some embodiments, the first nucleobase (e.g., modified cytosine) is a product of Huisgen cycloaddition of β-6-azido-glucosyl-5-hydroxymethylcytosine, the product containing an affinity tag (e.g., biotin).

[0134] In any of the combinations described herein, the captured DNA may include cfDNA. The captured DNA may have any of the characteristics described herein regarding the capture set, including, for example, a higher concentration of DNA corresponding to a sequence-variable target region set than the concentration of DNA corresponding to an epigenetic target region set (as normalized for footprint size as discussed above). In some embodiments, the DNA of the capture set contains a sequence tag, which may be added to the DNA as described herein. Typically, the inclusion of a sequence tag results in DNA molecules that differ from their naturally occurring, untagged form.

[0135] This combination may also include the probe set or sequencing primers described herein, each of which may differ from naturally occurring nucleic acid molecules. For example, the probe set described herein may include a capture portion, and the sequencing primers may contain a non-naturally occurring marker.

[0136] Therapeutics and related administration The methods of this disclosure can be implemented using or by means of a computer system. For example, such a method may include: partitioning a sample into more than one subsample, said more than one subsample including a first subsample and a second subsample, wherein the first subsample contains DNA with a larger proportion of cytosine modification than the second subsample; subjecting the first subsample to procedures that differently affect a first nucleobase and a second nucleobase in the DNA of the first subsample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first and second nucleobases have the same base-pairing specificity; and sequencing the DNA in the first subsample and the DNA in the second subsample in a manner that distinguishes the first and second nucleobases in the DNA of the first subsample.

[0137] In one aspect, this disclosure provides a non-transitory computer-readable medium including computer-executable instructions that, when executed by at least one electronic processor, perform at least a portion of a method comprising: collecting cfDNA from a test subject; capturing more than one set of target regions from the cfDNA, wherein the more than one set of target regions includes a sequence-variable target region set and an epigenetic target region set, thereby generating a set of captured cfDNA molecules; sequencing the captured cfDNA molecules, wherein the captured cfDNA molecules of the sequence-variable target region set are sequenced to a deeper sequencing depth than the captured cfDNA molecules of the epigenetic target region set; obtaining more than one sequence read generated by a nucleic acid sequencer through sequencing the captured cfDNA molecules; mapping the more than one sequence read to one or more reference sequences to generate mapped sequence reads; and processing the mapped sequence reads corresponding to the sequence-variable target region set and the epigenetic target region set to determine the likelihood that the subject has cancer.

[0138] The code can be pre-compiled and configured for use with machines having processors suitable for executing the code, or it can be compiled during runtime. The code can be provided in a programming language, which can be selected to enable the code to be executed either pre-compiled or as-compiled.

[0139] Additional details relating to computer systems and networks, databases, and computer program products are provided, for example, in the following: Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Ed. (2011); Kurose, Computer Networking: A Top-Down Approach, Pearson, 7th Ed. (2016); Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Ed. (2010); Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11th Ed. (2014); Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd Ed. (2006); and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011), all of which are hereby incorporated in their entirety by reference.

[0140] Kit In some embodiments, the methods disclosed herein involve identifying and administering a customized therapy to a patient based on the state of nucleic acid variation being of somatic or germline origin. In some embodiments, virtually any cancer therapy (e.g., surgical therapy, radiation therapy, chemotherapy therapy, and / or similar therapies) can be included as part of these methods. Typically, the customized therapy includes at least one immunotherapy (or immunotherapeutic agent). Immunotherapy generally refers to a method of enhancing the immune response against a given cancer type. In some embodiments, immunotherapy refers to a method of enhancing the T-cell response against a tumor or cancer.

[0141] In some implementations, the status of nucleic acid variants in a sample from a subject as somatic or germline can be compared with a database of comparator results from a reference population to identify a custom or targeted therapy for that subject. Typically, the reference population includes patients with the same type of cancer or disease as the subject being tested and / or those receiving or having received the same therapies as the subject being tested. When nucleic acid variants and comparator results meet certain classification criteria (e.g., a near-perfect or approximate match), a custom or targeted therapy (or more than one therapy) can be identified.

[0142] In some embodiments, the custom-made therapies described herein are typically administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions containing immunotherapeutic agents are typically administered intravenously. Some therapeutic agents are administered orally. However, custom-made therapies (e.g., immunotherapeutic agents, etc.) can also be administered by methods such as, for example, sublingual, sublingual, rectal, vaginal, urethral, ​​topical, intraocular, intranasal, and / or intraauricular administration, which may include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, salves, ointments, etc.

[0143] Biomarker Kits comprising compositions as described herein are also provided. The kits can be used to perform the methods described herein. In some embodiments, the kit contains a first reagent for partitioning a sample into more than one subsample as described herein, such as any partitioning reagent described elsewhere herein. In some embodiments, the kit contains a second reagent (e.g., any reagent described elsewhere herein for converting nucleosides such as cytosine or methylated cytosine into different nucleosides), the second reagent being used to subject the first subsample to procedures that differently affect a first nucleobase and a second nucleobase in the DNA of the first subsample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first and second nucleosides have the same base-pairing specificity. The kit may contain the first and second reagents, as well as additional elements discussed below and / or elsewhere herein.

[0144] The kit may also contain more than one oligonucleotide probe, which selectively hybridizes to at least 5, 6, 7, 8, 9, 10, 20, 30, 40, or all genes selected from the following groups: TLCD3B, PPP1R1B, TMEM200B, GLP1R, FJX1, ZBED4, PFKFB3, MDFI, ZNF280B, CIZ1, POU4F3, MED1, ARHGEF4, SERTM1, FRMD3, NOTUM, MBP, CLDN10, FABP3, ANXA6, PM20D2, BHLHE23, LRFN5, HOXD10, ECHDC1, GSTP1. , CCKBR, MIR4500HG, CALD1, PDE4B, INSC, HLA-E, NRG1, CYYR1, MIR3185, CDH12, ANP32E, SLC7A4, LTBP4, CCDC28B, DLX4, ADCY8, SP8, ASPHD1, HOXB8, AGPS , HLF, NFIX, KCNA2, ROBO3, CCDC92B, KCNH8, PPP4R4, FOXC2, PPP2R2C, CAPG, ZNF329, IL12RB2, HOXD-AS2, SORCS1, HOXD13, SAMM50, UNC13A, CACNG8, CADP S, HIF3A, CSMD3, MIR137, DLX4, BEGAIN, PHACTR2, ZNF667, FENDRR, SRP68, TUBB2A, SLC16A12, MARCKS, C1QL1, FOXL2, ERBB2, LAYN, PDGFRA, HRK, FAM20A, LOC105377879, WNT3A, MYH11, NMNAT2, KCNIP4, TRANK1, SYT2, ZNF536, NEUROG3, LIN28A, DMRT3, MCOLN2, ITGA4, PNKY, CHRNB1, GLIS3, HAGLR, PIK3CD, DI O3OS, DUOX2, CRHR1, FAM20A, RFX4, PLXNC1, P2RX6, TLCD3B, PPP1R1B, TMEM200B, GLP1R, FJX1, ZBED4, PFKFB3, MDFI, ZNF280B, CIZ1, POU4F3, MED1, ARHGE F4, SERTM1, FRMD3, NOTUM, MBP, CLDN10, FABP3, ANXA6, PM20D2, BHLHE23, LRFN5, HOXD10, ECHDC1, GSTP1, CCKBR, MIR4500HG, CALD1, PDE4B, INSC, HLA-E,NRG1, CYYR1, MIR3185, CDH12, ANP32E, SLC7A4, LTBP4, CCDC28B, DLX4, ADCY8, SP8, ASPHD1, HOXB8, AGPS, HLF, NFIX, KCNA2, ROBO3, CCDC92B, KCNH8, PPP4R4, FOXC2, PPP2R2C, CAPG, ZNF329, IL12RB2, HOXD-AS2, SORCS1, HOXD13, SAMM50, UNC13A, CACNG8, CADPS, HIF3A, CSMD3, MIR137, DLX4, BEGAIN, PHACTR 2、ZNF667、FENDRR、SRP68、TUBB2A、SLC16A12、MARCKS、C1QL1、FOXL2、ERBB2、LAYN、PDGFRA、HRK、FAM20A、LOC105377879、WNT3A、MYH11、NMNAT2、KCNIP4 、TRANK1、SYT2、ZNF536、NEUROG3、LIN28A、DMRT3、MCOLN2、ITGA4、PNKY、CHRNB1、GLIS3、HAGLR、PIK3CD、DIO3OS、DUOX2、CRHR1、FAM20A、RFX4、PLXNC1、P2 RX6、TLCD3B、PPP1R1B、TMEM200B、GLP1R、FJX1、ZBED4、PFKFB3、MDFI、ZNF280B、CIZ1、POU4F3、MED1、ARHGEF4、SERTM1、FRMD3、NOTUM、MBP、CLDN10、FABP 3、ANXA6、PM20D2、BHLHE23、LRFN5、HOXD10、ECHDC1、GSTP1、CCKBR、MIR4500HG、CALD1、PDE4B、INSC、HLA-E、NRG1、CYYR1、MIR3185、CDH12、ANP32E、SLC7A 4、LTBP4、CCDC28B、DLX4、ADCY8、SP8、ASPHD1、HOXB8、AGPS、HLF、NFIX、KCNA2、ROBO3、CCDC92B、KCNH8、PPP4R4、FOXC2、PPP2R2C、CAPG、ZNF329、IL12RB2、HOXD-AS2、SORCS1、HOXD13、SAMM50、UNC13A、CACNG8、CADPS、HIF3A、CSMD3、MIR137、DLX4、BEGAIN、PHACTR2、ZNF667、FENDRR、SRP68、TUBB2A、SLC16A12、The oligonucleotide probes can selectively hybridize to different numbers of genes, including MARCKS, C1QL1, FOXL2, ERBB2, LAYN, PDGFRA, HRK, FAM20A, LOC105377879, WNT3A, MYH11, NMNAT2, KCNIP4, TRANK1, SYT2, ZNF536, NEUROG3, LIN28A, DMRT3, MCOLLN2, ITGA4, PNKY, CHRNB1, GLIS3, HAGLR, PIK3CD, DIO3OS, DUOX2, CRHR1, FAM20A, RFX4, PLXNC1, and P2RX6. For example, the number of genes may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, or 54. The kit may contain a container containing more than one oligonucleotide probe for performing any of the methods described herein and instructions.

[0145] Oligonucleotide probes can selectively hybridize with exon regions of genes (e.g., at least 5 genes). In some cases, oligonucleotide probes can selectively hybridize with at least 30 exons of genes (e.g., at least 5 genes). In some cases, more than one probe can selectively hybridize with each of the at least 30 exons. The probe hybridizing with each exon may have a sequence overlapping with at least one other probe. In some embodiments, oligonucleotide probes can selectively hybridize with non-coding regions of genes disclosed herein (e.g., intronic regions of genes). Oligonucleotide probes can also selectively hybridize with regions of genes that contain both exon and intronic regions of genes disclosed herein.

[0146] Oligonucleotide probes can target any number of exons. For example, they can target at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, and 14 exons. 5, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, 230, 235, 240, 245, 250, 255, 260, 265, 270, 275, 280, 285, 290, 295, 300, 400, 500, 600, 700, 800, 900, 1,000 or more exons.

[0147] The kit may contain at least 4, 5, 6, 7, or 8 different library adaptors with different molecular barcodes and the same sample barcode. Library adaptors may not be sequencing adaptors. For example, library adaptors may not contain flow cell sequences or sequences that allow the formation of hairpin loops for sequencing. Different variations and combinations of molecular and sample barcodes are described throughout the kit and apply to the kit. Furthermore, in some cases, the adaptor is not a sequencing adaptor. Additionally, the adaptors provided with the kit may also include sequencing adaptors. Sequencing adaptors may contain sequences that hybridize with one or more sequencing primers. Sequencing adaptors may also contain sequences that hybridize with a solid support, such as a flow cell sequence. For example, a sequencing adaptor may be a flow cell adaptor. Sequencing adaptors may be attached to one or both ends of a polynucleotide fragment. In some cases, the kit may contain at least 8 different library adaptors with different molecular barcodes and the same sample barcode. Library adaptors may not be sequencing adaptors. The kit may also contain a sequencing adaptor having a first sequence that selectively hybridizes with a library adaptor and a second sequence that selectively hybridizes with a flow-through sequence. In another example, the sequencing adaptor may be hairpin-shaped. For example, a hairpin-shaped adaptor may contain complementary double-stranded portions and circular portions, wherein the double-stranded portions may be attached (e.g., linked) to a double-stranded polynucleotide. The hairpin-shaped sequencing adaptor may be attached to both ends of a polynucleotide fragment to produce a circular molecule that can be sequenced multiple times. Sequencing adaptors can range from end-to-end in the number of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, and 54 adaptors. One, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more bases. Sequencing adaptors can contain 20-30, 20-40, 30-50, 30-60, 40-60, 40-70, 50-60, or 50-70 bases end-to-end. In specific instances, sequencing adaptors can contain 20-30 bases end-to-end.In another example, a sequencing adaptor may contain 50-60 bases end-to-end. A sequencing adaptor may contain one or more barcodes. For example, a sequencing adaptor may contain a sample barcode. A sample barcode may contain a predetermined sequence. A sample barcode can be used to identify the source of a polynucleotide. A sample barcode may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 21, 22, 23, 24, 25, or more nucleic acid bases (or any length as described throughout), for example, at least 8 bases. The barcode may be a continuous or non-continuous sequence, as described above.

[0148] Library adapters can be blunt-ended or Y-shaped, and their length can be less than or equal to 40 nucleic acid bases. Other variations of library adapters can be found in the full text and are applicable to this kit.

[0149] Cancer treatment, therapy This disclosure provides methods for diagnosing, prognosticating, and selecting treatments for subjects with diseases such as cancer using biomarkers. A biomarker can be any gene or gene variant whose presence, mutation, deletion, substitution, copy number, or translation (i.e., translation into a protein) is an indicator of a disease state. Biomarkers in this disclosure may include the presence, mutation, deletion, substitution, copy number, or translation of any one or more of EGFR, KRAS, MET, BRAF, MYC, NRAS, ERBB2, ALK, Notch, PIK3CA, APC, and SMO.

[0150] A biomarker is a genetic variant associated with one or more types of cancer. Biomarkers can be identified using any of several resources or methods. Biomarkers can be previously discovered or discovered de novo using experimental or epidemiological techniques. Detection of a biomarker can indicate cancer when it is highly associated with the cancer. Detection of a biomarker can also indicate cancer when it occurs in a region or gene at a frequency greater than that of a given background population or dataset.

[0151] Publicly available resources such as scientific literature and databases can detail the discovery of genetic variants associated with cancer. Scientific literature can describe experimental or genome-wide association studies (GWAS) that associate one or more genetic variants with cancer. Databases can aggregate information gathered from sources such as scientific literature to provide a more comprehensive resource for identifying one or more biomarkers. Non-limiting examples of databases include FANTOM, GTex, GEO, BodyAtlas, INSiGHT, OMIM (Online Human Mendelian Genetics, omim.org), cBioPortal (cbioportal.org), CIVIC (Clinical Interpretation of Cancer Variants, civic.genome.wustl.edu), DOCM (Database of Selected Mutations, docm.genome.wustl.edu), and the ICGC Data Portal (dcc.icgc.org). In another example, the COSMIC (Catalogue of Somatic Mutations in Cancer) database allows searching for biomarkers by cancer, gene, or mutation type. Biomarkers can also be identified de novo by conducting experimental studies such as case-control or association studies (e.g., genome-wide association studies).

[0152] As an example, a significance test can be used to identify biomarkers. Here, a significance test is a statistical method used to determine whether an effect or relationship observed in data is likely due to chance. For example, setting a significance level (α) of 0.05 indicates a 5% risk of Type I error (i.e., concluding an effect exists when it is not actually present). Typical steps include 1) formulating a null (no effect) hypothesis and an alternative (effect exists) hypothesis, calculating test statistics (e.g., Z-score, t-score) based on sample data, and determining a p-value, which shows the probability of obtaining an outcome at least as extreme as that observed under the null hypothesis. If the p-value is less than 0.05, the null hypothesis is rejected, indicating that the effect is statistically significant and unlikely to be due to chance; if it is greater, then there is insufficient evidence to reject the null hypothesis.

[0153] In another instance, weights in a machine learning model can be used to determine biomarkers.

[0154] In machine learning models, weights (or coefficients) determine the influence of each input feature or assign importance to each input feature to the output. They are initially set to small random values ​​and adjusted during training to minimize prediction error. The training process involves making predictions (forward pass), calculating the error using a loss function, and calculating the gradient (backward pass) that shows how each weight affects the error. Using an optimization algorithm like gradient descent, the weights are iteratively updated in the direction that reduces the error. This process is repeated until the model converges, at which point the loss function reaches its minimum, producing the optimal weights that improve the model's prediction accuracy.

[0155] One or more biomarkers can be detected in a sequencing panel. Biomarkers can be one or more genetic variants associated with cancer. Biomarkers can be selected from single nucleotide variants (SNVs), copy number variants (CNVs), insertions or deletions (e.g., insertions / deletions), gene fusions, and inversions. Biomarkers can affect protein levels. Biomarkers can be in promoters or enhancers and can alter gene transcription. Biomarkers can affect gene transcription and / or translation efficiency. Biomarkers can affect the stability of transcribed mRNA. Biomarkers can lead to changes in the amino acid sequence of translated proteins. Biomarkers can affect splicing, can change amino acids encoded by specific codons, can cause frameshifts, or can cause premature stop codons. Biomarkers can lead to conserved substitutions of amino acids. One or more biomarkers can lead to conserved substitutions of amino acids. One or more biomarkers can lead to non-conserved substitutions of amino acids.

[0156] One or more biomarkers can be driver mutations. Driver mutations are mutations that give tumor cells a selective advantage in their microenvironment by increasing their survival or proliferation. Biomarkers may not all be driver mutations. One or more biomarkers can be passenger mutations. Passenger mutations are mutations that do not affect the adaptation of tumor cells but may be associated with clonal expansion because they occur in the same genome as driver mutations.

[0157] The frequency of biomarkers can be as low as 0.001%. The frequency of biomarkers can be as low as 0.005%. The frequency of biomarkers can be as low as 0.01%. The frequency of biomarkers can be as low as 0.02%. The frequency of biomarkers can be as low as 0.03%. The frequency of biomarkers can be as low as 0.05%. The frequency of biomarkers can be as low as 0.1%. The frequency of biomarkers can be as low as 1%.

[0158] A single biomarker may be absent in more than 50% of subjects with cancer. A single biomarker may be absent in more than 40% of subjects with cancer. A single biomarker may be absent in more than 30% of subjects with cancer. A single biomarker may be absent in more than 20% of subjects with cancer. A single biomarker may be absent in more than 10% of subjects with cancer. A single biomarker may be absent in more than 5% of subjects with cancer. A single biomarker may be present in 0.001% to 50% of subjects with cancer. A single biomarker may be present in 0.01% to 30% of subjects with cancer. A single biomarker may be present in 0.01% to 20% of subjects with cancer. A single biomarker may be present in 0.01% to 10% of subjects with cancer. A single biomarker may be present in 0.1% to 10% of subjects with cancer. A single biomarker can be present in 0.1% to 5% of subjects with cancer.

[0159] The detection of biomarkers can indicate the presence of one or more types of cancer. The detection can indicate the presence of cancer selected from the group comprising: ovarian cancer, pancreatic cancer, breast cancer, colorectal cancer, non-small cell lung cancer (e.g., squamous cell carcinoma or adenocarcinoma), or any other cancer. The detection can indicate the presence of any cancer selected from the group comprising: ovarian cancer, pancreatic cancer, breast cancer, colorectal cancer, non-small cell lung cancer (squamous cell or adenocarcinoma), or any other cancer. The detection can indicate the presence of any of more than one cancer selected from the group comprising: ovarian cancer, pancreatic cancer, breast cancer, colorectal cancer, non-small cell lung cancer (squamous cell or adenocarcinoma), or any other cancer. The detection can indicate the presence of one or more of any cancers mentioned in this application.

[0160] One or more cancers may express a biomarker in at least one exon of the group. One or more cancers selected from the group including ovarian cancer, pancreatic cancer, breast cancer, colorectal cancer, non-small cell lung cancer (squamous cell or adenocarcinoma), or any other cancer may each express a biomarker in at least one exon of the group. Each of at least three cancers may express a biomarker in at least one exon of the group. Each of at least four cancers may express a biomarker in at least one exon of the group. Each of at least five cancers may express a biomarker in at least one exon of the group. Each of at least eight cancers may express a biomarker in at least one exon of the group. Each of at least ten cancers may express a biomarker in at least one exon of the group. All cancers may express a biomarker in at least one exon of the group.

[0161] If a participant has cancer, the participant must express a biomarker in at least one exon or gene in the group. At least 85% of participants with cancer must express a biomarker in at least one exon or gene in the group. At least 90% of participants with cancer must express a biomarker in at least one exon or gene in the group. At least 92% of participants with cancer must express a biomarker in at least one exon or gene in the group. At least 95% of participants with cancer must express a biomarker in at least one exon or gene in the group. At least 96% of participants with cancer must express a biomarker in at least one exon or gene in the group. At least 97% of participants with cancer must express a biomarker in at least one exon or gene in the group. At least 98% of participants with cancer must express a biomarker in at least one exon or gene in the group. At least 99% of participants with cancer must express a biomarker in at least one exon or gene in the group. At least 99.5% of subjects with cancer can express a biomarker in at least one exon or gene in the group.

[0162] If a subject has cancer, the subject must exhibit the biomarker in at least one region of the group. At least 85% of subjects with cancer must exhibit the biomarker in at least one region of the group. At least 90% of subjects with cancer must exhibit the biomarker in at least one region of the group. At least 92% of subjects with cancer must exhibit the biomarker in at least one region of the group. At least 95% of subjects with cancer must exhibit the biomarker in at least one region of the group. At least 96% of subjects with cancer must exhibit the biomarker in at least one region of the group. At least 97% of subjects with cancer must exhibit the biomarker in at least one region of the group. At least 98% of subjects with cancer must exhibit the biomarker in at least one region of the group. At least 99% of subjects with cancer must exhibit the biomarker in at least one region of the group. At least 99.5% of subjects with cancer must exhibit the biomarker in at least one region of the group.

[0163] Detection can be performed with high sensitivity and / or high specificity. Sensitivity can be a measure of the proportion of positive results correctly identified as positive. In some cases, sensitivity refers to the percentage of all detected biomarkers present. In some cases, sensitivity refers to the percentage of individuals correctly identified as having a certain disease. Specificity can be a measure of the proportion of negative results correctly identified as negative. In some cases, specificity refers to the proportion of individuals correctly identified as having no altered bases. In some cases, specificity refers to the percentage of individuals correctly identified as being in healthy individuals without a certain disease. The previously described non-unique labeling method significantly increases the specificity of detection by reducing noise generated by amplification and sequencing errors, which reduces the frequency of false positives. Detection can be performed with a sensitivity of at least 95%, 97%, 98%, 99%, 99.5%, or 99.9% and / or a specificity of at least 80%, 90%, 95%, 97%, 98%, or 99%. Detection can be performed with a sensitivity of at least 90%, 95%, 97%, 98%, 99%, 99.5%, 99.6%, 99.98%, 99.9%, or 99.95%. Detection can be performed with a specificity of at least 90%, 95%, 97%, 98%, 99%, 99.5%, 99.6%, 99.98%, 99.9%, or 99.95%. The method can be used to detect biomarkers with at least 70% specificity and at least 70% sensitivity, at least 75% specificity and at least 75% sensitivity, at least 80% specificity and at least 80% sensitivity, at least 85% specificity and at least 85% sensitivity, at least 90% specificity and at least 90% sensitivity, at least 95% specificity and at least 95% sensitivity, at least 96% specificity and at least 96% sensitivity, at least 97% specificity and at least 97% sensitivity, at least 98% specificity and at least 98% sensitivity, at least 99% specificity and at least 99% sensitivity, or 100% specificity and 100% sensitivity. In some cases, the method can detect biomarkers with about 80% or higher sensitivity or specificity or about 95% or higher sensitivity or specificity.

[0164] The detection can be highly accurate. This accuracy can be applied to the identification of biomarkers in cell-free DNA and / or the diagnosis of cancer. Statistical tools such as covariate analysis described above can be used to increase and / or measure accuracy. The method can detect biomarkers with an accuracy of at least 80%, 90%, 95%, 97%, 98%, or 99%, 99.5%, 99.6%, 99.98%, 99.9%, or 99.95%. In some cases, the method can detect biomarkers with an accuracy of at least 95% or higher.

[0165] Cancers In various implementation schemes, cancer treatment may include trastuzumab, pertuzumab, Ado-emtansine trastuzumab, neratinib, lapatinib, tucatinib, magitoximab, and / or detrastuzumab.

[0166] Treatment of colon cancer cells and tumor organoids with another derivative of the hypomethylating agent (5-aza-2'-deoxycytidine) has been reported to be sufficient to induce a growth-inhibiting immune response by triggering retrotransposon expression. The combination of DNMTi and HDACi more effectively and selectively induces LTR retrotransposons than either drug alone. (Baxevanis, “The Regulation and Immune Signature of Retrotransposons in Cancer”) Sequencing panel(Basel). Sep 2023; 15(17): 4340. In other embodiments, the described methods and compositions can be applied to cancer therapies with known mechanisms of action that significantly affect methylation. Such therapies include HDAC inhibitors, such as vorinostat, romidesin, pabistal, chidamide, and belinostat; HAT inhibitors, such as inhibitors of the GCN5-associated N-acetyltransferase (GNAT) family, including GCN5 and p300 / CBP-associated factor (PCAF), and the MYST superfamily, such as mangosteen extract (Garcinol), PU141, C646, Tip60 inhibitors TH1834, NU9056, and 6-alkyl salicylates; and BRD inhibitors, such as I-BET 151, I-BET... 762, OTX015, MK-8628, birabresib, and Yervoy, a CTLA4 inhibitor based on PD-L1 protein expression, also include tesimumab (Imjuno); nivolumab (Opdivo) is a PD-1 inhibitor that can be used in combination with ipilimumab (optionally including platinum); other PD-1 inhibitors include pembrolizumab (Keytruda), cimipril-rwlc (Libtayo), and durvalumab (Imfinzi), which are used for unresectable NSCLC. Further information can be found in Basudan Clin Pract. Feb. 2023; 13(1): 22–40 and Meng et al., Cell Death Dis 15, 3 (2024), each of which is incorporated herein by reference in its entirety.

[0167] In various implementation schemes, cancer treatment may include abecilibi, ADO-trastuzumab emtansine, apeliximab, anastrozole, elacestrant dihydrochloride, capivasertib, everolimus, exemestane, FAM-trastuzumab-NXKI, fulvestrant, lapatinib besylate, letrozole / reboxil succinate, magitoximab-cmkb, nellatinib maleate, olaparib, palbociclib, pembrolizumab, pertuzumab, pertuzumab, trastuzumab and hyaluronidase-zzxf, reboxil, sacituzumab govitecan-hziy, taprazole tosylate, tamoxifen citrate, toremifene, trastuzumab and tucatinib. In various implementation schemes, cancer treatment may include abraxane (paclitaxel albumin-stabilized nanoparticles), adapalene (pamidronate disodium), capecitabine, cyclophosphamide, doxorubicin hydrochloride, ellence (epirarubicin hydrochloride), 5-fu (fluorouracil injection), gemcitabine hydrochloride, halaven (eribulin mesylate), infugem (gemcitabine hydrochloride), ixempra (ixaprone), paclitaxel, taxotere (docetaxel), tecentriq (atezolizumab), tepadina (thiotepa), thiotepa, trexall (methoprene sodium), truqap (capivasertib), vinblastine sulfate, and / or zoladex (goserelin acetate). In various implementation schemes, cancer treatment may include olaparib, taprazoprib, rucapaprib, and nirapaprib + abiraterone acetate.

[0168] In various implementation schemes, cancer treatment may include imatinib, gefitinib, afatinib, dacomitinib, sunitinib, sorafenib, vandetanib, brivanib, cabozantinib, neratinib, tevantinib, bevacizumab, cixutumumab, dalotuzumab, figitumumab, rilotumumab, onartuzumab, ganitumab, ramucirumab, ridaforolimus, tesimolimus, everolimus, BMS-690514, BMS-754807, EMD 525797, GDC-0973, GDC-0941, MK-2206, AZD6244, GSK1120212, PX-866, XL821, IMC-A12, MM-121, PF-02341066, RG7160, and Sym004. Antibodies suitable for use as anti-EGFR therapy include cetuximab (trade name: Erbitux) and panitumumab (trade name: Vectibex). In some cases, cancer treatment includes EGFR tyrosine kinase inhibitors, such as gefitinib (trade name: Iressa), erlotinib (trade name: Tarceva), lapatinib, canertinib, and cetuximab.

[0169] In other implementation schemes, cancer treatment may include atezolizumab (Tecentriq), imatinib, gefitinib, afatinib, dacomitinib, sunitinib, sorafenib, vandetanib, brinib, cabozantinib, neratinib, tevantinib, bevacizumab, cetuximab, darotozumab, fentuximab, rituximab, onatuxumab, ganituxumab, ramucirumab, levofloxacin, tesimolimus, everolimus, levofloxacin, osimertinib, BMS-690514, BMS-754807, EMD 525797, GDC-0973, GDC-0941, MK-2206, AZD6244, GSK1120212, PX-866, XL821, IMC-A12, MM-121, PF-02341066, RG7160, and Sym004. Antibodies suitable for use as anti-EGFR therapy include cetuximab (trade name: Erbitux) and panitumumab (trade name: Vectibex). In some cases, cancer treatment includes EGFR tyrosine kinase inhibitors, such as gefitinib (trade name: Iressa), erlotinib (trade name: Tarceva), lapatinib, canenatinib, and cetuximab.

[0170] In some cases, therapies may be used in combination, such as anti-EGFR therapy and anti-EGFR treatment. Anti-EGFR therapy can be used in any combination with a chemotherapy agent or chemotherapy regimen, such as FOLFOX (fluorouracil [5-FU] / leucovorin / oxaliplatin), FOLFIRI (5-FU / leucovorin / irinotecan), etc. In some aspects, cancer treatment is administered to the subject. In some cases, cancer treatment is administered in combination with another therapy (such as non-anti-EGFR therapy with anti-EGFR therapy). In other embodiments, treatment may include monoclonal antibodies (mAbs), multispecific antibodies, antibody-drug conjugates (ADCs), cell therapy, fusion proteins, drug combinations, antibody fragments (Ab fragments), vaccines, and / or nanoparticles. Further information can be found in US20060018899A1, US7449184B2, US7560111B2, WO2015115091A1, US20060275305A1, US20090202546A1, US20090202536A1, and US20120107302A1, each of which is incorporated herein by reference in its entirety.

[0171] Example 1 - cfDNA methylation detection To increase the likelihood of detecting tumor indicator mutations, the sequenced DNA region can include a set of genes or genomic regions. Selecting a limited number of regions for sequencing (e.g., a limited set) can reduce the total sequencing effort required (e.g., the total number of nucleotides sequenced). A sequencing set can target more than one different gene or region to detect a single cancer, a collection of cancers, or all cancers.

[0172] In some respects, groups targeting more than one distinct gene or genomic region are selected such that a defined proportion of cancer-affected subjects exhibit genetic variants or biomarkers in one or more distinct gene or genomic regions within that group. A group can be selected to limit the region used for sequencing to a fixed number of base pairs. This group can be selected to sequence a desired amount of DNA. A group can also be selected to achieve a desired read depth. A group can be selected to achieve a desired read depth or read coverage for a given number of sequenced base pairs. A group can be selected to achieve theoretical sensitivity, theoretical specificity, and / or theoretical accuracy for detecting one or more genetic variants in a sample.

[0173] Probes used to detect this set of regions may include probes for detecting hotspot regions as well as nucleosome-sensing probes (e.g., KRAS codons 12 and 13), and can be designed to optimize capture based on analysis of cfDNA coverage and fragment size variation and GC sequence composition affected by nucleosome binding patterns. The regions used in this paper may also include non-hotspot regions optimized based on nucleosome location and GC models. This group may include more than one subpanel, including subpanels for identifying: source tissues (e.g., using published literature to define 50-100 decoys representing genes (not necessarily promoters) with the most diverse transcriptional profiles across tissues), whole-genome scaffolds (e.g., for identifying highly conserved genomic content and sparsely tiling it across chromosomes with a small number of probes for copy number base alignment purposes), and transcription start sites (TSS) / CpG islands (e.g., for capturing differentially methylated regions (DMRs) in promoters of tumor suppressor genes (e.g., SEPT9 / VIM in colorectal cancer). In some embodiments, the markers of source tissues are tissue-specific epigenetic markers.

[0174] One or more regions within a group may contain one or more loci from one or more genes. More than one gene may be selected for sequencing and biomarker detection. Genes included in the regions to be sequenced may be selected from genes known to be involved in cancer or from genes not involved in cancer. For example, more than one gene in a group may be an oncogene, tumor suppressor gene, growth factor gene, DNA repair gene, signal transduction gene, transcription factor gene, receptor gene, or metabolic gene. Examples of genes that may be included in the group include, but are not limited to: TLCD3B, PPP1R1B, TMEM200B, GLP1R, FJX1, ZBED4, PFKFB3, MDFI, ZNF280B, CIZ1, POU4F3, MED1, ARHGEF4, SERTM1, FRMD3, NOTUM, MBP, CLDN10, FABP3, ANXA6, PM20D2, BHLHE23, LRFN5, HOXD10, ECHDC1, GSTP1, CCKBR, MIR4500HG, CALD1, PDE4B, INSC, HLA-E, NRG1, CYYR1, MIR3185, CDH12, ANP32E, SLC7A4, LTBP4, CCDC28B, DLX4, ADCY8, SP8, ASPHD1, HOXB8, AGPS, H LF, NFIX, KCNA2, ROBO3, CCDC92B, KCNH8, PPP4R4, FOXC2, PPP2R2C, CAPG, ZNF329, IL12RB2, HOXD-AS2, SORCS1, HOXD13, SAMM50, UNC13A, CACNG8, CADPS, HIF3A, CSMD3, MIR137, DLX4, BEGAIN, PHACTR2, ZNF667, FENDRR, SRP68, TUBB2A, SLC16A12, MARCKS, C1 QL1, FOXL2, ERBB2, LAYN, PDGFRA, HRK, FAM20A, LOC105377879, WNT3A, MYH11, NMNAT2, KCNIP4, TRANK1, SYT2, ZNF536, NEUROG3, LIN28A, DMRT3, MCOLN2, ITGA4, PNKY, CHRNB1, GLIS3, HAGLR, PIK3CD, DIO3OS, DUOX2, CRHR1, FAM20A, RFX4, PLXNC1, P2RX6, TLCD 3B, PPP1R1B, TMEM200B, GLP1R, FJX1, ZBED4, PFKFB3, MDFI, ZNF280B, CIZ1, POU4F3, MED1, ARHGEF4, SERTM1, FRMD3, NOTUM, MBP,CLDN10, FABP3, ANXA6, PM20D2, BHLHE23, LRFN5, HOXD10, ECHDC1, GSTP1, CCKBR, MIR4500HG, CALD1, PDE4B, INSC, HLA-E, NRG1, CYYR1, MIR3185, CDH12, ANP32E, SLC7A4, LTBP4, CCDC28B, DLX4, ADCY8, SP8, ASPHD1, HOXB8, AGPS, HLF, NFIX, KCNA2, ROBO3, CCDC92B, KCNH8, PPP4R4, FOXC2, PPP2R2C, CAPG, ZNF3 29、IL12RB2、HOXD-AS2、SORCS1、HOXD13、SAMM50、UNC13A、CACNG8、CADPS、HIF3A、CSMD3、MIR137、DLX4、BEGAIN、PHACTR2、ZNF667、FENDRR、SRP68、TUBB2 A、SLC16A12、MARCKS、C1QL1、FOXL2、ERBB2、LAYN、PDGFRA、HRK、FAM20A、LOC105377879、WNT3A、MYH11、NMNAT2、KCNIP4、TRANK1、SYT2、ZNF536、NEUROG3、 LIN28A、DMRT3、MCOLN2、ITGA4、PNKY、CHRNB1、GLIS3、HAGLR、PIK3CD、DIO3OS、DUOX2、CRHR1、FAM20A、RFX4、PLXNC1、P2RX6、TLCD3B、PPP1R1B、TMEM200B GLP1R、FJX1、ZBED4、PFKFB3、MDFI、ZNF280B、CIZ1、POU4F3、MED1、ARHGEF4、SERTM1、FRMD3、NOTUM、MBP、CLDN10、FABP3、ANXA6、PM20D2、BHLHE23、LRFN5、 HOXD10, ECHDC1, GSTP1, CCKBR, MIR4500HG, CALD1, PDE4B, INSC, HLA-E, NRG1, CYYR1, MIR3185, CDH12, ANP32E, SLC7A4, LTBP4, CCDC28B, DLX4, ADCY8, SP8, ASPHD1, HOXB8, AGPS, HLF, NFIX, KCNA2, ROBO3, CCDC92B, KCNH8, PPP4R4, FOXC2, PPP2R2C, CAPG, ZNF329, IL12RB2, HOXD-AS2, SORCS1, HOXD13, SAMM50,UNC13A、CACNG8、CADPS、HIF3A、CSMD3、MIR137、DLX4、BEGAIN、PHACTR2、ZNF667、FENDRR、SRP68 、TUBB2A、SLC16A12、MARCKS、C1QL1、FOXL2、ERBB2、LAYN、PDGFRA、HRK、FAM20A、LOC105377879、 WNT3A、MYH11、NMNAT2、KCNIP4、TRANK1、SYT2、ZNF536、NEUROG3、LIN28A、DMRT3、MCOLN2、ITGA4 、PNKY、CHRNB1、GLIS3、HAGLR、PIK3CD、DIO3OS、DUOX2、CRHR1、FAM20A、RFX4、PLXNC1、P2RX60。、

[0175] In some cases, one or more regions of a group may include one or more loci from one or more genes, including one or more of the following: TLCD3B, PPP1R1B, TMEM200B, GLP1R, FJX1, ZBED4, PFKFB3, MDFI, ZNF280B, CIZ1, POU4F3, MED1, ARHGEF4, SERTM1, FRMD3, NOTUM, MBP, CLDN10, FABP3, ANXA6, PM20D2, BHLHE23, LRFN5, HOXD10, ECHDC1, GSTP1, CCKBR, MIR4500HG, CALD1, PD E4B, INSC, HLA-E, NRG1, CYYR1, MIR3185, CDH12, ANP32E, SLC7A4, LTBP4, CCDC28B, DLX4, ADCY8, SP8, ASPHD1, HOXB8, AGPS, HLF, NFIX, KCNA2, ROBO3, CCD C92B, KCNH8, PPP4R4, FOXC2, PPP2R2C, CAPG, ZNF329, IL12RB2, HOXD-AS2, SORCS1, HOXD13, SAMM50, UNC13A, CACNG8, CADPS, HIF3A, CSMD3, MIR137, DLX4 , BEGAIN, PHACTR2, ZNF667, FENDRR, SRP68, TUBB2A, SLC16A12, MARCKS, C1QL1, FOXL2, ERBB2, LAYN, PDGFRA, HRK, FAM20A, LOC105377879, WNT3A, MYH11, NMNAT2, KCNIP4, TRANK1, SYT2, ZNF536, NEUROG3, LIN28A, DMRT3, MCOLN2, ITGA4, PNKY, CHRNB1, GLIS3, HAGLR, PIK3CD, DIO3OS, DUOX2, CRHR1, FAM20A, R FX4, PLXNC1, P2RX6, TLCD3B, PPP1R1B, TMEM200B, GLP1R, FJX1, ZBED4, PFKFB3, MDFI, ZNF280B, CIZ1, POU4F3, MED1, ARHGEF4, SERTM1, FRMD3, NOTUM, MBP , CLDN10, FABP3, ANXA6, PM20D2, BHLHE23, LRFN5, HOXD10, ECHDC1, GSTP1, CCKBR, MIR4500HG, CALD1, PDE4B, INSC, HLA-E, NRG1, CYYR1, MIR3185, CDH12,ANP32E、SLC7A4、LTBP4、CCDC28B、DLX4、ADCY8、SP8、ASPHD1、HOXB8、AGPS、H LF、NFIX、KCNA2、ROBO3、CCDC92B、KCNH8、PPP4R4、FOXC2、PPP2R2C、CAPG、ZN F329、IL12RB2、HOXD-AS2、SORCS1、HOXD13、SAMM50、UNC13A、CACNG8、CADPS 、HIF3A、CSMD3、MIR137、DLX4、BEGAIN、PHACTR2、ZNF667、FENDRR、SRP68、TUB B2A, SLC16A12, MARCKS, C1QL1, FOXL2, ERBB2, LAYN, PDGFRA, HRK, FAM20A, LOC105377879, WNT3A, MYH11, NMNAT2, KCNIP4, TRANK1, SYT2, ZNF536, NEURO G3、LIN28A、DMRT3、MCOLN2、ITGA4、PNKY、CHRNB1、GLIS3、HAGLR、PIK3CD、DI O3OS、DUOX2、CRHR1、FAM20A、RFX4、PLXNC1、P2RX6、TLCD3B、PPP1R1B、TMEM20 0B、GLP1R、FJX1、ZBED4、PFKFB3、MDFI、ZNF280B、CIZ1、POU4F3、MED1、ARHGE F4、SERTM1、FRMD3、NOTUM、MBP、CLDN10、FABP3、ANXA6、PM20D2、BHLHE23、LR FN5、HOXD10、ECHDC1、GSTP1、CCKBR、MIR4500HG、CALD1、PDE4B、INSC、HLA-E 、NRG1、CYYR1、MIR3185、CDH12、ANP32E、SLC7A4、LTBP4、CCDC28B、DLX4、ADCY 8、SP8、ASPHD1、HOXB8、AGPS、HLF、NFIX、KCNA2、ROBO3、CCDC92B、KCNH8、PPP 4R4、FOXC2、PPP2R2C、CAPG、ZNF329、IL12RB2、HOXD-AS2、SORCS1、HOXD13、SA MM50、UNC13A、CACNG8、CADPS、HIF3A、CSMD3、MIR137、DLX4、BEGAIN、PHACTR 2、ZNF667、FENDRR、SRP68、TUBB2A、SLC16A12、MARCKS、C1QL1、FOXL2、ERBB2、LAYN, PDGFRA, HRK, FAM20A, LOC105377879, WNT3A, MYH11, NMNAT2, KCNIP4, TRANK1, SYT2, ZNF536, NEUROG3, LIN28A , DMRT3, MCOLN2, ITGA4, PNKY, CHRNB1, GLIS3, HAGLR, PIK3CD, DIO3OS, DUOX2, CRHR1, FAM20A, RFX4, PLXNC1, P2RX6. ,

[0176] In some implementations, one or more regions of the group include one or more loci from one or more genes for detecting residual cancer after surgery. This detection can be performed earlier than existing cancer detection methods. In some implementations, one or more regions of the group include one or more loci from one or more genes for detecting cancer in high-risk patient populations. For example, smokers have a much higher incidence of lung cancer than the general population. Furthermore, smokers may develop other lung conditions that make cancer detection more difficult, such as the development of irregular lung nodules. In some implementations, the methods described herein may detect cancer in high-risk patients earlier than existing cancer detection methods.

[0177] Regions can be selected for inclusion in the sequencing group based on the number of subjects with cancer who possess biomarkers in that gene or region. Alternatively, regions can be selected for inclusion based on the prevalence of cancer among subjects and the presence of biomarkers in that gene. The presence of biomarkers in a region can indicate that a subject has cancer.

[0178] In some cases, information from one or more databases can be used to select groups. Information about cancer can be derived from cancer tumor biopsies or cfDNA assays. Databases can include information describing the population of sequenced tumor samples. Databases can include information about mRNA expression in tumor samples. Databases can include information about regulatory elements in tumor samples. Information associated with the sequenced tumor samples can include the frequencies of various genetic variants and describe the genes or regions in which the genetic variants are present. Genetic variants can be biomarkers. A non-limiting example of such a database is COSMIC. COSMIC is a catalog of somatic mutations found in various cancers. For a specific cancer, COSMIC sorts genes according to mutation frequency. By having mutations with a high frequency in a given gene, genes can be selected to be included in a group. For example, COSMIC shows that 33% of the sequenced breast cancer sample population has a mutation in TP53, and 22% of the sampled breast cancer population has a mutation in KRAS. Other sequenced genes, including APC, have mutations found in only about 4% of the sequenced breast cancer sample population. Based on the relatively high frequencies of TP53 and KRAS in the sampled breast cancers (e.g., APC occurs at a frequency of approximately 4% compared to APC), TP53 and KRAS can be included in the sequencing group. COSMIC is provided as a non-limiting example; however, any database or information set that associates cancer with biomarkers located in genes or genetic regions can be used. In another example provided by COSMIC, 380 out of 1156 biliary tract cancer samples (33%) carried TP53 mutations. Several other genes, such as APC, were mutated in 4%–8% of all samples. Therefore, TP53 can be selected for inclusion in the group based on its relatively high frequency in the biliary tract cancer sample population.

[0179] A group can be selected from genes or regions in which the frequency of the biomarker in sampled tumor tissue or circulating tumor DNA is significantly higher than its frequency found in a given background population. Combinations of regions can be selected to include a group such that at least a majority of subjects with cancer will have a biomarker present in at least one region or gene within that group. Combinations of regions can be selected based on data indicating that, for a specific cancer or set of cancers, a majority of subjects have one or more biomarkers in one or more selected regions. For example, to detect cancer 1, a group including regions A, B, C, and / or D can be selected based on data indicating that 90% of subjects with cancer 1 have biomarkers in regions A, B, C, and / or D within that group. Optionally, a biomarker can manifest independently in two or more regions of subjects with cancer, such that, when combined, the biomarker in two or more regions is present in a majority of the population of subjects with cancer. For example, to detect cancer2, groups including regions X, Y, and Z can be selected based on data indicating that 90% of subjects have biomarkers in one or more regions, and that in 30% of such subjects, the biomarker is detected only in region X, while for the remaining subjects where the biomarker is detected, it is detected only in regions Y and / or Z. If the biomarker is detected in one or more of these regions 50% or more of the time, then the presence of a biomarker in one or more regions previously shown to be associated with one or more cancers can indicate or predict that a subject has cancer. Computational methods, such as models of the conditional probability of detecting cancer given a known cancer frequency for a set of biomarkers in one or more regions, can be used to predict which regions, alone or in combination, can predict cancer. Other methods for group selection include using databases describing information from studies employing comprehensive genomic mapping analyses and / or whole-genome sequencing (WGS, RNA-seq, Chip-seq, bisulfate sequencing, ATAC-seq, etc.) of tumors with large panels. Information gathered from the literature can also describe pathways that are commonly affected and mutated in certain cancers. Group selection can also be communicated by using an ontology that describes genetic information.

[0180] The genes included in the sequencing set may include fully transcribed regions, promoter regions, enhancer regions, regulatory elements, and / or downstream sequences. To further increase the likelihood of detecting tumor indicator mutations, only exons may be included in the set. The set may contain all exons of the selected genes, or only one or more exons of the selected genes. The set may include exons from each of more than one different gene. The set may contain at least one exon from each of more than one different gene.

[0181] In some respects, a set of exons from each of more than one different gene is selected such that a certain proportion of subjects with cancer exhibit genetic variation in at least one exon of that set of exons.

[0182] At least one complete exon from each distinct gene in a set of genes can be sequenced. The sequenced set can contain exons from more than one gene. The set can contain exons from 2 to 100 distinct genes, 2 to 70 genes, 2 to 50 genes, 2 to 30 genes, 2 to 15 genes, or 2 to 10 genes.

[0183] The selected group can contain varying numbers of exons. This group can contain 2 to 3000 exons. This group can contain 2 to 1000 exons. This group can contain 2 to 500 exons. This group can contain 2 to 100 exons. This group can contain 2 to 50 exons. This group can contain no more than 300 exons. This group can contain no more than 200 exons. This group can contain no more than 100 exons. This group can contain no more than 50 exons. This group can contain no more than 40 exons. This group can contain no more than 30 exons. This group can contain no more than 25 exons. This group can contain no more than 20 exons. This group can contain no more than 15 exons. This group can contain no more than 10 exons. This group can contain no more than 9 exons. This group can contain no more than 8 exons. This group can contain no more than 7 exons.

[0184] This group may contain one or more exons from more than one different gene. This group may contain one or more exons from each of the more than one different gene in a certain proportion. This group may contain at least two exons from each of at least 25%, 50%, 75%, or 90% of the different genes. This group may contain at least three exons from each of at least 25%, 50%, 75%, or 90% of the different genes. This group may contain at least four exons from each of at least 25%, 50%, 75%, or 90% of the different genes.

[0185] The size of a sequencing group can vary. A sequencing group can be large or small (in terms of nucleotide size) depending on several factors, including, for example, the total amount of nucleotides sequenced or the number of unique molecules sequenced for a specific region within the group. Sequencing group sizes can range from 5 kb to 50 kb. Sequencing group sizes can range from 10 kb to 30 kb. Sequencing group sizes can range from 12 kb to 20 kb. Sequencing group sizes can range from 12 kb to 60 kb. Sequencing group sizes can be at least 10 kb, 12 kb, 15 kb, 20 kb, 25 kb, 30 kb, 35 kb, 40 kb, 45 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb. The size of the sequencing group can be less than 100 kb, 90 kb, 80 kb, 70 kb, 60 kb or 50 kb.

[0186] The groups selected for sequencing can contain at least 1, 5, 10, 15, 20, 25, 30, 40, 50, 60, 80, or 100 regions. In some cases, the regions within a group are selected such that the region size is relatively small. In some cases, the regions within a group have a size of approximately 10 kb or less, approximately 8 kb or less, approximately 6 kb or less, approximately 5 kb or less, approximately 4 kb or less, approximately 3 kb or less, approximately 2.5 kb or less, approximately 2 kb or less, approximately 1.5 kb or less, or approximately 1 kb or less. In some cases, the regions within a group have a size of approximately 0.5 kb to approximately 10 kb, approximately 0.5 kb to approximately 6 kb, approximately 1 kb to approximately 11 kb, approximately 1 kb to approximately 15 kb, approximately 1 kb to approximately 20 kb, approximately 0.1 kb to approximately 10 kb, or approximately 0.2 kb to approximately 1 kb. For example, regions within a group can range in size from approximately 0.1 kb to approximately 5 kb.

[0187] The group selected in this paper allows for deep sequencing sufficient to detect low-frequency genetic variants (e.g., in cell-free nucleic acid molecules obtained from a sample). The amount of a genetic variant in a sample can be referred to by the minor allele frequency of a given genetic variant. Minor allele frequency can refer to the frequency of a minor allele (e.g., not the most common allele) in a given population of nucleic acids, such as in a sample. Genetic variants with low minor allele frequencies may have a relatively low frequency of presence in the sample. In some cases, this group allows the detection of genetic variants with minor allele frequencies of at least 0.0001%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, or 0.5%. This group can also allow the detection of genetic variants with minor allele frequencies of 0.001% or higher. This group allows the detection of genetic variants present in a sample at frequencies as low as 0.0001%, 0.001%, 0.005%, 0.01%, 0.025%, 0.05%, 0.075%, 0.1%, 0.25%, 0.5%, 0.75%, or 1.0%. This group allows the detection of biomarkers present in a sample at frequencies of at least 0.0001%, 0.001%, 0.005%, 0.01%, 0.025%, 0.05%, 0.075%, 0.1%, 0.25%, 0.5%, 0.75%, or 1.0%. This group allows the detection of biomarkers at frequencies as low as 1.0% in a sample. This group allows the detection of biomarkers at frequencies as low as 0.75% in a sample. This group allows the detection of biomarkers at frequencies as low as 0.5% in a sample. This group allows the detection of biomarkers at frequencies as low as 0.25% in a sample. This group allows the detection of biomarkers at frequencies as low as 0.1% in a sample. This group allows the detection of biomarkers at frequencies as low as 0.075% in a sample. This group allows the detection of biomarkers at frequencies as low as 0.05% in a sample. This group allows the detection of biomarkers at frequencies as low as 0.025% in a sample. This group allows the detection of biomarkers at frequencies as low as 0.01% in a sample. This group allows the detection of biomarkers at frequencies as low as 0.005% in a sample. This group allows the detection of biomarkers at frequencies as low as 0.001% in a sample. This group allows the detection of biomarkers at frequencies as low as 0.0001% in a sample. This group allows the detection of biomarkers in sequenced cfDNA at frequencies as low as 1.0% to 0.0001% in a sample. This group allows the detection of biomarkers in sequenced cfDNA at frequencies as low as 0.01% to 0.0001% in a sample.

[0188] In a population of subjects with a disease (e.g., cancer), a certain proportion of genetic variants may be expressed. In some cases, at least 1%, 2%, 3%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the population with cancer may express one or more genetic variants in at least one region of the group. For example, at least 80% of the population with cancer may express one or more genetic variants in at least one region of the group.

[0189] This group may contain one or more regions from each of one or more genes. In some cases, the group may contain one or more regions from each of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, or 80 genes. In some cases, the group may contain one or more regions from each of at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, or 80 genes. In some cases, the group may include one or more regions from each of about 1 to about 80, 1 to about 50, about 3 to about 40, 5 to about 30, or 10 to about 20 different genes.

[0190] Regions within the group can be selected to detect one or more epigenetic modifications. These epigenetic modifications can be acetylated, methylated, ubiquitinated, phosphorylated, ubiquitinated like other substances, ribosylated, and / or citrullinated. For example, regions within the group can be selected to detect one or more methylated regions.

[0191] Regions within a group can be selected such that they contain sequences differentially transcribed across one or more tissues. In some cases, regions may contain sequences transcribed at higher levels in certain tissues compared to other tissues. For example, a region may contain sequences transcribed in some tissues but not in others.

[0192] Regions within a group may contain coding and / or non-coding sequences. For example, regions within a group may contain one or more sequences of exons, introns, promoters, 3' untranslated regions, 5' untranslated regions, regulatory elements, transcription start sites, and / or splicing sites. In some cases, regions within a group may contain other non-coding sequences, including pseudogenes, repetitive sequences, transposons, viral elements, and telomeres. In some cases, regions within a group may contain sequences from non-coding RNAs, such as ribosomal RNA, transfer RNA, Piwi-interacting RNA, and microRNAs.

[0193] Regions within a group can be selected to detect (diagnose) cancer at a desired level of sensitivity (e.g., by detecting one or more genetic variants). For example, regions within a group can be selected to detect cancer with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% (e.g., by detecting one or more genetic variants). Regions within a group can also be selected to detect cancer with 100% sensitivity.

[0194] Regions within a group can be selected to detect (diagnose) cancer (e.g., by detecting one or more genetic variants) at a desired level of specificity. For example, regions within a group can be selected to detect cancer (e.g., by detecting one or more genetic variants) with at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% specificity. Regions within a group can also be selected to detect one or more genetic variants with 100% specificity.

[0195] Regions within a group can be selected to detect (diagnose) cancer with the desired positive predictive value. The positive predictive value can be increased by increasing sensitivity (e.g., the chance of detecting an actual positive) and / or specificity (e.g., the chance of not mistaking an actual negative for a positive). As a non-limiting example, regions within a group can be selected to detect one or more genetic variants with a positive predictive value of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. Regions within a group can be selected to detect one or more genetic variants with a 100% positive predictive value.

[0196] Regions within a group can be selected for detecting (diagnosing) cancer with the desired accuracy. As used herein, the term "accuracy" can refer to the ability of a test to distinguish between a disease condition (e.g., cancer) and health. Accuracy can be quantified using measures such as sensitivity and specificity, predictive value, likelihood ratio, area under the ROC curve, Youden index, and / or diagnostic odds ratio.

[0197] Accuracy can be expressed as a percentage, which is the ratio between the number of tests that gave correct results and the total number of tests performed. Regions within a group can be selected to detect cancer with at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy. Regions within a group can also be selected to detect cancer with 100% accuracy.

[0198] Groups can be selected such that removing one or more regions or genes from that group results in a significant decrease in specificity. Removing a region from that group can lead to a decrease in specificity of at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, or more.

[0199] You can select groups such that adding one or more regions or genes to a group does not significantly increase the specificity of that group, for example, it does not increase the specificity by more than 1%, 2%, 5%, 10%, 15%, or 20%.

[0200] The size of the group can be such that when one or more regions or genes in the group are removed, its sensitivity is significantly reduced, for example, by at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50% or more.

[0201] You can select groups such that adding one or more regions or genes to a group does not significantly increase the sensitivity of that group, for example, it does not increase the sensitivity by more than 1%, 2%, 5%, 10%, 15%, or 20%.

[0202] The size of the group can be such that when one or more regions or genes in the group are removed, the accuracy is significantly reduced, for example, by at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50% or more.

[0203] You can select groups such that adding one or more regions or genes to a group does not significantly increase the accuracy of that group, for example, it does not increase the accuracy by more than 1%, 2%, 5%, 10%, 15%, or 20%.

[0204] The size of the group can be such that when one or more regions or genes in the group are removed, the positive predictive value is significantly reduced, for example, by at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50% or more.

[0205] You can select groups such that adding one or more regions or genes to a group does not significantly increase the positive predictive value of that group, for example, it does not increase the positive predictive value by more than 1%, 2%, 5%, 10%, 15%, or 20%.

[0206] Groups can be selected for high sensitivity and detection of low-frequency genetic variants. For example, groups can be selected such that genetic variants or biomarkers present in a sample at frequencies as low as 0.01%, 0.05%, or 0.001% can be detected with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. Regions within a group can be selected for detection of biomarkers present in a sample at frequencies of 1% or lower with a sensitivity of 70% or higher. Groups can be selected for detection of biomarkers in a sample at frequencies as low as 0.1% with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. Groups can be selected to detect biomarkers at frequencies as low as 0.01% in a sample with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. Groups can also be selected to detect biomarkers at frequencies as low as 0.001% in a sample with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.

[0207] Groups can be selected for high specificity and detection of low-frequency genetic variants. For example, groups can be selected such that genetic variants or biomarkers present in a sample at frequencies as low as 0.01%, 0.05%, or 0.001% can be detected with at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% specificity. Regions within a group can be selected for detection of biomarkers present in a sample at frequencies as low as 1% with 70% or higher specificity. Groups can be selected for detection of biomarkers in a sample at frequencies as low as 0.1% with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% specificity. Groups can be selected to detect biomarkers with a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% at a frequency as low as 0.001% in the sample.

[0208] Groups can be selected for high accuracy and detection of low-frequency genetic variants. Groups can be selected such that genetic variants or biomarkers present in the sample at frequencies as low as 0.01%, 0.05%, or 0.001% can be detected with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy. Regions within a group can be selected for detection of biomarkers present in the sample at frequencies as low as 1% with 70% or higher accuracy. Groups can be selected for detection of biomarkers in the sample at frequencies as low as 0.1% with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy. Groups can be selected to detect biomarkers in samples with an accuracy of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% at a frequency as low as 0.001%.

[0209] Groups can be selected as highly predictive and for detecting low-frequency genetic variants. Groups can be selected such that genetic variants or biomarkers present in the sample at frequencies as low as 0.01%, 0.05%, or 0.001% can have a positive predictive value of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.

[0210] The concentration of the probe or decoy used in the group can be increased (2 to 6 ng / µL) to capture more nucleic acid molecules in the sample. The concentration of the probe or decoy used in the group can be at least 2 ng / µL, 3 ng / µL, 4 ng / µL, 5 ng / µL, 6 ng / µL, or higher. The probe concentration can be about 2 ng / µL to about 3 ng / µL, about 2 ng / µL to about 4 ng / µL, about 2 ng / µL to about 5 ng / µL, or about 2 ng / µL to about 6 ng / µL. The concentration of the probe or decoy used in the group can be 2 ng / µL or higher up to 6 ng / µL or lower. In some cases, this can allow the analysis of more molecules in biological samples, thus enabling the detection of lower frequency alleles.

[0211] Although preferred embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that these embodiments are provided by way of example only. The invention is not intended to be limited to the specific instances provided herein. While the invention has been described with reference to the foregoing description, the description and illustration of the embodiments herein are not intended to be construed in a limiting sense. Various variations, modifications, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it should be understood that all aspects of the invention are not limited to the specific descriptions, configurations, or relative proportions set forth herein under various conditions and variables. It should be understood that various alternatives to the embodiments of this disclosure described herein may be employed in the practice of the invention. Therefore, it is contemplated that this disclosure should also cover any such alternatives, modifications, variations, or equivalents. The following claims are intended to define the scope of the invention and thereby cover the methods and structures within the scope of these claims and their equivalents.

[0212] While the foregoing disclosure has been described in detail with reference to illustrations and examples for clarity and understanding, it will be apparent to those skilled in the art, upon reading this disclosure, that various changes in form and detail may be made without departing from the true scope of this disclosure, and that it may be practiced within the scope of the appended claims. For example, all methods, systems, computer-readable media and / or component features, steps, elements, or other aspects may be used in various combinations. Example

[0213] Example 2 - Methods As described, current breast cancer detection involves an IHC test, typically performed at the start of treatment. This paradigm overlooks the fact that breast cancer subtypes can change and evolve over time, a result of subtype switching or receptor conversion phenomena.

[0214] In particular, subtype switching can occur due to genetic mutations and other molecular changes within tumor cells. Specifically, HR+ / HER2- may acquire mutations that lead to resistance to hormone therapy; the tumor can transform into a different subtype. Subtype switching can also occur as a result of therapeutic interventions: certain treatments, such as targeted therapy or chemotherapy, can exert selective pressure on tumor cells, leading to the emergence of resistant clones with altered molecular characteristics. Mixtures of clinical subtypes can also occur across or coexist within metastatic lesions in the same patient, presenting significant clinical challenges.

[0215] Monitoring with cfDNA can address these challenges by providing a non-invasive, routine monitoring platform that can collect and analyze samples at a range of time points, something currently untapped in IHC. Similarly, liquid biopsy allows for the detection of metastatic lesions across the entire body in ways impossible with tissue biopsy, given the infeasibility of repeatedly obtaining tissue samples across different locations.

[0216] Example 3 - TCGA A training set of approximately 250 breast plasma samples was used to generate methylation scores from the methylation-based liquid biopsy group, with the majority being HR+ / HER2- (e.g., 75%). Factors included in the analysis included tumor score (higher TF was preferred), time interval between IHC testing and blood collection (for all samples), and the use of pretreated samples with a high degree of homogeneity.

[0217] Example 4 - Results The methylation detection platform described in this paper can be compared with TCGA by using one-to-many classification of four breast cancer subtypes (e.g., HR+ / HER2-, HR+ / HER2+, HR- / HER2+, HR- / HER2-). The methylation detection platform described in this paper includes methylation-based groups that report somatic changes in >750 genes. Tumor fraction-independent methylation regions with significant signals can be ranked according to their variability across the training set. Methylation β values ​​generated by the Illumina 450K methylation microarray from TCGA include probes corresponding to the methylation detection platform, with data augmentation using Gaussian noise ~N(0, σ=0.01) to balance the categories and 5-fold cross-validation.

[0218] Example 5 - Preliminary findings Using a machine learning model with binary methylation data, a set of 200 most important epigenomic methylation regions was extracted from an ML model trained on binary TCGA methylation data. Promoter methylation determination of breast cancer samples was then extracted for these 200 selected epigenomic methylation regions. Here, single and multi-class classification of three types of breast cancer was applied.

[0219] Example 6 - Generating probabilistic modeling The results achieved in this paper demonstrate that methylation-based liquid biopsy NGS assays can characterize subtype profiles previously observed in tissue-based profiling analyses. This allows for the characterization of a patient's breast cancer subtype profile using a convenient and non-invasive liquid biopsy method for methylation detection.

[0220] Example 7 - Exemplary method for BRCA subtype typing Probabilistic modeling methods can be generated that allow determining the likelihood of tumor molecules being present in genomic regions of a sample given a TF and DNA input. Methylation data is sparsity-dependent, meaning regions can be sporadically turned on / off. The degree of sparsity of the signal is related to both the DNA input and TF. The region score (the number of molecules in a region) depends on the TF. The observed molecules are a mixture of normal molecules (background) and tumor molecules. Applications of probabilistic modeling include normalized methylation scores for cancer subtyping and predicting the likelihood of different cancer subtypes.

[0221] Table 1. TCGA samples used to train breast cancer subtype classifiers. Subtype prediction models can be applied across different technologies. Here, three models (HR / HER2 / TNBC) can be trained using 450K methylation determination data from TCGA raw tissue samples. Optionally, the subtype prediction model can be applied to methyl-binding domain partitioning methylation detection or subtype prediction. In some implementations, the model can be trained directly on methylation data (LR+Lasso), including samples with known subtypes, and performance can be evaluated using cross-validation.

[0222] The model was fully trained using the TCGA data described in Table 1. After training, the model was used to classify infinity samples.

[0223] Example 8 - Performance characteristics

[0224]

[0225] The model was fully trained using the TCGA data described in Table 1. After training, the model was used to classify samples with methyl binding domain partitioning for methylation detection.

[0226] Figure 5 Molecular counts in 15,229 promoter regions were then fitted to a background distribution of region scores from 1,585 cancer-free samples; subsequently, a mixture model was fitted to the distribution of region scores from 1,014 BRCA samples, as shown below. Figure 7 As shown in the diagram. Here, the following is used: Figure 9 The 5-fold cross-validation shown in Figure 8 and Example 9 - Enriched gene sets The subtype predictions shown across time points were trained on HER2 classifiers on samples with methyl binding domain partition methylation detection (TF>0.01).

[0227] Figure 1 The top-ranked genes near genomic regions that predict breast cancer subtypes were used for gene set enrichment analysis using the gp.enrichr a module within the Enrichr suite. The analysis revealed that the input gene set was enriched for genes in the ERBB2 signaling pathway (see [link to Enrichr module]). Example 10 - REST pathway Furthermore, using ENCODE_TF_ChIP-seq_2015 from the library in gp.enrichr, we found that genes associated with the HER2 predictor region are targets of RE1 silencing transcription factors (REST) ​​across multiple cell lines, including MCF-7, a well-characterized human breast cancer cell line derived from metastatic sites (pleural effusion) of breast adenocarcinoma. Genes included MIR137, ROBO3, UNC13A, LIN28A, ASPHD1, KCNH8, CADPS, NMNAT2, and POU4F3.

[0228] Example 11 - Targeting the REST pathway As described, the HER2 predictor region is a target of the RE1 silencing transcription factor (REST). REST is a transcriptional repressor that recruits various chromatin-modifying enzymes and co-repressor complexes that modify chromatin structure and function. One mechanism by which REST represses transcription is through the recruitment of the CoREST complex, which includes histone deacetylases and other chromatin-modifying enzymes. This complex modifies the chromatin structure around the REST target gene, making it more condensed and harder for transcriptional mechanisms to access. Without being bound by any particular theory, it is possible that by recruiting complexes that alter histone modifications, REST may create a chromatin environment conducive to DNA methylation by DNA methyltransferases. This can lead to stable silencing of its target genes through the combined action of histone modification and DNA methylation. Given the above, the HER2+ subtype predictor genomic region (the promoter region of the REST regulatory network) is methylated in HER2+ via a REST-mediated mechanism.

[0229] Importantly, the genes are present only in the HER2 predictor region, suggesting that they together represent the REST regulatory network and could be used as a predictive biomarker for HER2-positive breast cancer.

[0230] Example 12 - HER2-specific signature Due to its role in proliferation, REST can be a therapeutic target for cancers with high REST levels. REST levels are primarily regulated through phosphorylation-triggered degradation, such as via small molecule inhibitory compounds, molecules targeting small C-terminal domain phosphatase 1 (SCP1), or antibodies and mimics of SCP1, which dephosphorylates REST at the site. Since targeting transcription factors like REST via small molecules is often difficult, alternative approaches include using epigenetic regulators such as HDAC and HAT inhibitors, considering that the CoREST complex comprises histone deacetylases and other chromatin-modifying enzymes.

[0231] Example 13 - Use of gene signatures in selecting drug treatments Many genes enriched in HER2-positive samples have been identified, including IR137, ROBO3, UNC13A, LIN28A, ASPHD1, KCNH8, CADPS, NMNAT2, and POU4F3. It is of interest to assess their performance in breast cancer subtyping. Here, evaluation of the HER2, HR, and TNBC datasets reveals compelling results, in which almost all of the aforementioned genes are able to identify HER2-positive status, excluding HR and TNBC subtypes.

[0232]

[0233] +Indicates the presence of promoter methylation Table 2. FDA-approved HER2 drugs In light of the foregoing, the enrichment of genes associated with HER2 status, including IR137, ROBO3, UNC13A, LIN28A, ASPHD1, KCNH8, CADPS, NMNAT2, and POU4F3, indicates that subjects enriched with one or more of these genes are good candidates for receiving HER2 therapeutic agents (including those listed in Table 2). In one instance, liquid biopsy can determine the expression of the aforementioned enriched genes, and therefore, subjects can be identified as good responsive candidates for HER2 drug use based on cancer subtype.

[0234]

Claims

1. A method comprising: Detect methylation at at least one of more than one sites; Generate more than one or more measures for each of the more than one site; and Process one or more metrics to characterize the sample.

2. The method of claim 1, wherein the one or more measures are obtained from methylation determinations from each of the more than one site.

3. The method according to claim 1, comprising obtaining a sample.

4. The method of claim 1, comprising having the obtained sample.

5. The method of claim 1, further comprising constructing a classification model from methylation data of a set of training samples, said training samples comprising one or more of the following: HR+ / HER2-, HR+ / HER2+, HR- / HER2+, and HR- / HER2-.

6. The method of claim 5, wherein the classification model is trained using one-to-many classification.

7. The method of claim 6, wherein the training includes sorting.

8. The method of claim 7, wherein the region is selected based on sorting.

9. The method of claim 7, wherein the ranking is based on the significance and / or weight of each of the more than one site.

10. The method of claim 6, wherein cross-validation is used to train the classification model.

11. The method of claim 10, wherein cross-validation includes cross-validation using 1-fold, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, or more folds.

12. The method of claim 1, wherein generating more than one or more measures for each of the more than one site includes a Bernoulli probability distribution and / or a mixture model.

13. The method of claim 1, wherein processing the one or more measures to characterize the sample includes calculating the probability of tumor molecules based on the amount of input DNA at one or more sites of the more than one site.

14. The method of claim 1, wherein the site comprises a customized group.

15. The method of claim 14, wherein the customized group is configured as a computer simulation group.

16. The method of claim 14, wherein the customized group is configured in a physical group.

17. The method of claim 14, wherein the customized set comprises a set of oncogenes, promoter regions of a set of oncogenes, HRR genes, immuno-oncology (IO) genes, cancer pathways, methylation peaks found in cancer or methylation peaks found in clinical samples.

18. The method of claim 14, wherein the customized group is refined based at least on literature annotations, common methylation peak positions, and / or public datasets.

19. The method of claim 1, wherein the more than one site comprises one or more genes selected from the group consisting of: TLCD3B, PPP1R1B, TMEM200B, GLP1R, FJX1, ZBED4, PFKFB3, MDFI, ZNF280B, CIZ1, POU4F3, MED1, ARHGEF4, SERTM1, FRMD3, NOTUM, MBP, CLDN10, FABP3, ANXA6, PM20D2, BHLHE23, LRFN5, HOXD10, ECHDC1, GSTP1, CCKBR, MIR4500HG, CALD1, PDE4B, INSC, HLA-E, NRG1, CYYR1, MIR3185, CDH12, ANP32E, SLC7A4, LTBP4, CCDC28B, DLX4, ADCY8, SP8, ASPHD1, HOXB8, AGPS, HLF, NFIX, KCNA2, ROBO3, CCDC92B, KCN H8, PPP4R4, FOXC2, PPP2R2C, CAPG, ZNF329, IL12RB2, HOXD-AS2, SORCS1, HOXD13, SAMM50, UNC13A, CACNG8, CADPS, HIF3A, CSMD3, MIR137, DLX4, BEGAIN, PHACTR2, ZNF667, FENDRR, SRP68, TUBB2A, SLC16A12, MARCKS, C1QL1, FOXL2, ERBB2, LAYN, PDGFRA, HRK, FAM20A, LOC105377879, WNT3A, MYH11, NMNAT2, KCNIP4, TRANK1, SYT2, ZNF536, NEUROG3, LIN28A, DMRT3, MCOLN2, ITGA4, PNKY, CHRNB1, GLIS3, HAGLR, PIK3CD, DIO3OS, DUOX2, CRHR1, FAM20A, RFX4, PLX NC1, P2RX6, TLCD3B, PPP1R1B, TMEM200B, GLP1R, FJX1, ZBED4, PFKFB3, MDFI, ZNF280B, CIZ1, POU4F3, MED1, ARHGEF4, SERTM1, FRMD3, NOTUM, MBP, CLDN10 , FABP3, ANXA6, PM20D2, BHLHE23, LRFN5, HOXD10, ECHDC1, GSTP1, CCKBR, MIR4500HG, CALD1, PDE4B, INSC, HLA-E, NRG1, CYYR1, MIR3185, CDH12, ANP32E,SLC7A4、LTBP4、CCDC28B、DLX4、ADCY8、SP8、ASPHD1、HOXB8、AGPS、HLF、NFIX 、KCNA2、ROBO3、CCDC92B、KCNH8、PPP4R4、FOXC2、PPP2R2C、CAPG、ZNF329、IL 12RB2、HOXD-AS2、SORCS1、HOXD13、SAMM50、UNC13A、CACNG8、CADPS、HIF3A、 CSMD3、MIR137、DLX4、BEGAIN、PHACTR2、ZNF667、FENDRR、SRP68、TUBB2A、SL C16A12、MARKS、C1QL1、FOXL2、ERBB2、LAYN、PDGFRA、HRK、FAM20A、LOC1053 77879、WNT3A、MYH11、NMNAT2、KCNIP4、TRANK1、SYT2、ZNF536、NEUROG3、LIN 28A、DMRT3、MCOLN2、ITGA4、PNKY、CHRNB1、GLIS3、HAGLR、PIK3CD、DIO3OS、D UOX2、CRHR1、FAM20A、RFX4、PLXNC1、P2RX6、TLCD3B、PPP1R1B、TMEM200B、GLP 1R、FJX1、ZBED4、PFKFB3、MDFI、ZNF280B、CIZ1、POU4F3、MED1、ARHGEF4、SER TM1、FRMD3、NOTUM、MBP、CLDN10、FABP3、ANXA6、PM20D2、BHLHE23、LRFN5、HO ECHDC1 CYYR1、MIR3185、CDH12、ANP32E、SLC7A4、LTBP4、CCDC28B、DLX4、ADCY8、SP8、 ASPHD1、HOXB8、AGPS、HLF、NFIX、KCNA2、ROBO3、CCDC92B、KCNH8、PPP4R4、FO XC2、PPP2R2C、CAPG、ZNF329、IL12RB2、HOXD-AS2、SORCS1、HOXD13、SAMM50、 UNC13A、CACNG8、CADPS、HIF3A、CSMD3、MIR137、DLX4、BEGAIN、PHACTR2、ZNF 667、FENDRR、SRP68、TUBB2A、SLC16A12、MARKS、C1QL1、FOXL2、ERBB2、LAYN、PDGFRA、HRK、FAM20A、LOC105377879、WNT3A、MYH11、NMNAT2、KCNIP4、TRANK1、SYT2、ZNF536、NEUROG3、LIN28A、DMRT3、MCOLN2、ITGA4、PNKY、CHRNB1、GLIS3、HAGLR、PIK3CD、DIO3OS、DUOX2、CRHR1、FAM20A、RFX4、PLXNC1、P2RX60。、 20. The method of claim 1, wherein the more than one site comprises one or more genes selected from the group consisting of: TLCD3B, MIR137, ROBO3, UNC13A, LIN28A, ASPHD1, KCNH8, CADPS, NMNAT2, and POU4F3.

21. The method of claim 1, wherein characterizing the sample comprises determining the gene expression of one or more biomarkers.

22. A method comprising: Detect methylation at at least one of more than one sites; For each of the more than one site, generate more than one methylation determination; One or more measures are obtained from the methylation determination; and The one or more metrics are processed to generate the probability that a patient has cancer.

23. The method of claim 22, wherein generating the probability that a patient has cancer includes determining the cancer subtype.

24. The method of claim 23, wherein the cancer subtype is characterized as one or more of HR, HER2, and TNBC.

25. The method of claim 22, wherein the more than one site comprises one or more genes selected from the group consisting of: TLCD3B, PPP1R1B, TMEM200B, GLP1R, FJX1, ZBED4, PFKFB3, MDFI, ZNF280B, CIZ1, POU4F3, MED1, ARHGEF4, SERTM1, FRMD3, NOTUM, MBP, CLDN10, FABP3, ANXA6, PM20D2, BHLHE23, LRFN5, HOXD10, ECHDC1, GSTP1, CCKBR, MIR4500HG, CALD1, PDE4B, INSC , HLA-E, NRG1, CYYR1, MIR3185, CDH12, ANP32E, SLC7A4, LTBP4, CCDC28B, DLX4, ADCY8, SP8, ASPHD1, HOXB8, AGPS, HLF, NFIX, KCNA2, ROBO3, CCDC92B, KCN H8, PPP4R4, FOXC2, PPP2R2C, CAPG, ZNF329, IL12RB2, HOXD-AS2, SORCS1, HOXD13, SAMM50, UNC13A, CACNG8, CADPS, HIF3A, CSMD3, MIR137, DLX4, BEGAIN, PHACTR2, ZNF667, FENDRR, SRP68, TUBB2A, SLC16A12, MARCKS, C1QL1, FOXL2, ERBB2, LAYN, PDGFRA, HRK, FAM20A, LOC105377879, WNT3A, MYH11, NMNAT2, KCNIP4, TRANK1, SYT2, ZNF536, NEUROG3, LIN28A, DMRT3, MCOLN2, ITGA4, PNKY, CHRNB1, GLIS3, HAGLR, PIK3CD, DIO3OS, DUOX2, CRHR1, FAM20A, RFX4, PLX NC1, P2RX6, TLCD3B, PPP1R1B, TMEM200B, GLP1R, FJX1, ZBED4, PFKFB3, MDFI, ZNF280B, CIZ1, POU4F3, MED1, ARHGEF4, SERTM1, FRMD3, NOTUM, MBP, CLDN10 , FABP3, ANXA6, PM20D2, BHLHE23, LRFN5, HOXD10, ECHDC1, GSTP1, CCKBR, MIR4500HG, CALD1, PDE4B, INSC, HLA-E, NRG1, CYYR1, MIR3185, CDH12, ANP32E,SLC7A4、LTBP4、CCDC28B、DLX4、ADCY8、SP8、ASPHD1、HOXB8、AGPS、HLF、NFIX 、KCNA2、ROBO3、CCDC92B、KCNH8、PPP4R4、FOXC2、PPP2R2C、CAPG、ZNF329、IL 12RB2、HOXD-AS2、SORCS1、HOXD13、SAMM50、UNC13A、CACNG8、CADPS、HIF3A、 CSMD3、MIR137、DLX4、BEGAIN、PHACTR2、ZNF667、FENDRR、SRP68、TUBB2A、SL C16A12、MARKS、C1QL1、FOXL2、ERBB2、LAYN、PDGFRA、HRK、FAM20A、LOC1053 77879、WNT3A、MYH11、NMNAT2、KCNIP4、TRANK1、SYT2、ZNF536、NEUROG3、LIN 28A、DMRT3、MCOLN2、ITGA4、PNKY、CHRNB1、GLIS3、HAGLR、PIK3CD、DIO3OS、D UOX2、CRHR1、FAM20A、RFX4、PLXNC1、P2RX6、TLCD3B、PPP1R1B、TMEM200B、GLP 1R、FJX1、ZBED4、PFKFB3、MDFI、ZNF280B、CIZ1、POU4F3、MED1、ARHGEF4、SER TM1、FRMD3、NOTUM、MBP、CLDN10、FABP3、ANXA6、PM20D2、BHLHE23、LRFN5、HO ECHDC1 CYYR1、MIR3185、CDH12、ANP32E、SLC7A4、LTBP4、CCDC28B、DLX4、ADCY8、SP8、 ASPHD1、HOXB8、AGPS、HLF、NFIX、KCNA2、ROBO3、CCDC92B、KCNH8、PPP4R4、FO XC2、PPP2R2C、CAPG、ZNF329、IL12RB2、HOXD-AS2、SORCS1、HOXD13、SAMM50、 UNC13A、CACNG8、CADPS、HIF3A、CSMD3、MIR137、DLX4、BEGAIN、PHACTR2、ZNF 667、FENDRR、SRP68、TUBB2A、SLC16A12、MARKS、C1QL1、FOXL2、ERBB2、LAYN、PDGFRA、HRK、FAM20A、LOC105377879、WNT3A、MYH11、NMNAT2、KCNIP4、TRANK1、SYT2、ZNF536、NEUROG3、LIN28A、DMRT3、MCOLN2、ITGA4、PNKY、CHRNB1、GLIS3、HAGLR、PIK3CD、DIO3OS、DUOX2、CRHR1、FAM20A、RFX4、PLXNC1、P2RX60。、 26. The method of claim 22, wherein the more than one site comprises one or more genes selected from the group consisting of: MIR137, ROBO3, UNC13A, LIN28A, ASPHD1, KCNH8, CADPS, NMNAT2, and POU4F3.

27. The method of claim 22, comprising selecting a treatment from trastuzumab, pertuzumab, trastuzumab Ado-emtansine, neratinib, lapatinib, tucatinib, magitoximab, and detrastuzumab.

28. The method of claim 22, comprising administering treatment to the patient, wherein the treatment is trastuzumab, pertuzumab, Ado-emtansine trastuzumab, neratinib, lapatinib, tucatinib, magitoximab, and / or detrastuzumab.

29. The method according to any one of the preceding claims, comprising: The patient was diagnosed with cancer.

30. The method according to any one of the preceding claims, comprising: The prognosis is that the patient is susceptible to cancer.

31. The method according to any one of the preceding claims, wherein the sample comprises cell-free DNA.

Citation Information

Patent Citations

  • HER2 antibody composition

    US20060018899A1

  • HERCEPTIN adjuvant therapy

    US20060275305A1

  • Antibody-drug conjugates and methods

    US20090202536A1

  • Composition comprising antibody that binds to domain ii of her2 and acidic variants thereof

    US20090202546A1

  • Combinations of an Anti-her2 antibody-drug conjugate and chemotherapeutic agents, and methods of use

    US20120107302A1