Methods for the early detection of cancer

A binary classification model using methylation data from an epigenome panel addresses the limitations of current PD-L1 detection methods, enabling precise treatment recommendations for cancer patients.

JP2026525206APending Publication Date: 2026-07-29GUARDANT HEALTH INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
GUARDANT HEALTH INC
Filing Date
2024-06-28
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

Current methods for detecting PD-L1 expression in cancer, such as immunohistochemical staining, are laborious and time-consuming, limiting their usefulness in determining effective treatment options.

Method used

A method utilizing methylation data from an epigenome panel to construct a binary classification model for PD-L1 expression, allowing for the prediction of PD-L1 levels in bodily fluids, which informs immunotherapy and chemotherapy recommendations.

Benefits of technology

Enables rapid and accurate determination of PD-L1 levels for personalized treatment decisions, improving treatment efficacy by leveraging methylation patterns in genomic and epigenomic profiling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026525206000001_ABST
    Figure 2026525206000001_ABST
Patent Text Reader

Abstract

Genetic signatures that lead to prognosis, diagnosis, treatment, and molecular subtype classification of cancer by genomic and epigenomic profiling, including immune checkpoint modulators such as programmed cell death ligand 1 (PDL-1), are described herein. The methods and compositions described herein provide specific and highly sensitive detection of biomarkers of interest. Such biomarkers serve as indicators of disease pathogenicity and provide opportunities for treatment selection, including treatment regimens targeting resistance mechanisms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 511,493, filed Jun. 30, 2023, which is hereby incorporated by reference in its entirety.

[0002] Field of the Invention Methods and compositions for cancer prognosis determination, diagnosis, treatment, and molecular subtype classification by genomic, epigenomic, and transcriptomic profiling are described herein.

Background Art

[0003] Background Cancer is a major global cause of disease. Every year, tens of millions of people worldwide are diagnosed with cancer, and more than half of patients ultimately die from cancer. In many countries, cancer is the second leading cause of death after cardiovascular disease. For many cancers, early detection leads to improved outcomes.

[0004] To detect cancer, several screening tests are available. General health signs are investigated from physical examinations and medical histories, including confirmation of disease symptoms, such as lumps or other unusual physical symptoms. Information on the patient's health habits and history of past illnesses and treatments is also obtained. Clinical tests are another type of screening test, which may include medical procedures to obtain samples of tissues, blood, urine, or other substances in the body before performing the clinical test. Imaging procedures screen for cancer by generating a visual representation of areas within the body. Genetic tests detect harmful mutations in certain genes associated with some types of cancer. Genetic tests are particularly useful in several diagnostic methods. There is a strong need in the art for genetic test methods that include biomarkers having information value for diagnostic methods.

[0005] PD-L1 has attracted attention, and the PD-1 / PD-L1 pathway plays a major role in immunomodulation by delivering inhibitory signals to maintain a balance between T cell activation, tolerance, and immune-mediated tissue damage. Generally speaking, current methods using PD-L1 testing measure the percentage of cells in tumors that "express PD-L1 (programmed cell death ligand-1 protein)," encoded by the CD274 gene. This can be informative because PD-L1 levels can influence treatment options. Nevertheless, such techniques are limited in usefulness, laborious, and time-consuming because they rely on visual inspection of tumor tissue and immunohistochemical staining. [Overview of the project] [Means for solving the problem]

[0006] Summary of the Invention This specification describes a method for constructing a binary classification model from methylation data by using normalized molecular counts in an epigenome panel at a hyper partition as a measure of methylation. In one application, lung cancer patients can have their PD-L1 levels tested via a bodily fluid sample such as blood. Based on the expression level, immunotherapy may be recommended and / or administered as first-line treatment, or in other cases, immunotherapy and / or chemotherapy may be recommended and / or administered.

[0007] A method is described herein that includes the steps of detecting methylation at at least one of a plurality of sites, generating one or more metrics for each of the plurality of sites, and processing one or more metrics to characterize a sample. In other embodiments, the one or more metrics are obtained from methylation calls from each of the plurality of sites. In other embodiments, the method includes obtaining a sample. In other embodiments, the method includes the sample being obtained. In other embodiments, the step of characterizing a sample includes determining the gene expression of one or more biomarkers. In other embodiments, the one or more biomarkers include PD-L1, MSI, and / or BRAF. In other embodiments, the method includes constructing a binary classification model from methylation data of a set of training samples including PD-L1 high and PD-L1 low status. In other embodiments, the classification model is trained using cross-validation. In other embodiments, the cross-validation includes using 3-fold or 10-fold cross-validation. In other embodiments, regions are selected using penalized logistic regression and Least Absolute Shrinkage Selection Operator (LASSO) regularization. In other embodiments, the penalized logistic regression model includes PD-L1 as the response variable and methylation calls as the predictor variable for each of multiple sites. In other embodiments, the sites include a custom panel. In other embodiments, the custom panel is comprised of an in silico panel. In other embodiments, the custom panel is comprised of a physical panel. In other embodiments, the custom panel includes a set of oncogenes, promoter regions for the set of oncogenes, HRR genes, immuno-oncology (IO) genes, cancer pathways, methylation peaks found in cancer, or methylation peaks found in clinical samples. In other embodiments, the custom panel is refined based on at least literature annotations, typical methylation peak locations, and / or public datasets.In other embodiments, the PD-L1 status is determined based on gene expression data, PD-L1 promoter region nucleosome location, or histological data. In other embodiments, the PD-L1 status is a prediction of the treatment response. In other embodiments, the treatment comprises one or more of the following: immune checkpoint inhibitors (ICIs), poly(ADP-ribose) polymerase (PARP) inhibitors, kinase inhibitors, aromatase inhibitors, CTLA4 inhibitors, PD-L1 inhibitors, and PD-1 inhibitors, either alone or in combination with fluoropyrimidine-containing chemotherapy and platinum-containing chemotherapy. In other embodiments, the immune checkpoint inhibitor is pembrolizumab. In other embodiments, the poly(ADP-ribose) polymerase (PARP) inhibitor is olaparib or talazoparib. In other embodiments, the method includes the step of diagnosing that a subject has cancer. In other embodiments, the method includes the step of determining the prognosis that a subject is susceptible to cancer. In other embodiments, the method includes the step of selecting a treatment for the subject. In other embodiments, the method includes the step of performing a treatment on a subject.

[0008] A method is described herein that includes the steps of detecting methylation at at least one of several sites, generating several methylation calls for each of the multiple sites, obtaining one or more metrics from the methylation calls, and processing one or more metrics to generate a probability that a patient exhibits PD-L1 expression. In other embodiments, the patient is a lung cancer patient, and the PD-L1 level corresponds to high PD-L1 expression as measured by proteomics, histological examination, or immunohistochemical examination. In other embodiments, high PD-L1 expression includes PD-L1 expression in 1% or more of tumor cells. In other embodiments, high PD-L1 expression includes 50% or more of tumor cells being stained with PD-L1 [TC≧50%] or PD-L1 stained tumor-infiltrating immune cells [IC] accounting for 10% or more of the tumor area [IC≧10%]. In other embodiments, the patient does not exhibit either EGFR or ALK genomic abnormalities. In other embodiments, the patient does not exhibit EGFR genomic abnormalities, ALK genomic abnormalities, or ROS genomic abnormalities. In other embodiments, the patient is administered a PD-L1 inhibitor or a CTLA4 inhibitor alone or in combination with platinum-containing chemotherapy. In other embodiments, the method includes the step of diagnosing that the subject has cancer. In other embodiments, the method includes the step of determining the prognosis that the subject is susceptible to cancer. In other embodiments, the method includes the step of selecting a treatment for the subject. In other embodiments, the method includes the step of administering a treatment to the subject.

[0009] A method is described herein that includes the steps of detecting methylation at at least one of several sites; generating several methylation calls for each of the several sites; obtaining one or more metrics from the methylation calls; processing one or more metrics to generate a probability that a patient exhibits PD-L1 expression; and determining that the patient is a candidate for treatment using PARPi. In other embodiments, the method includes the step of diagnosing that a subject has cancer. In other embodiments, the method includes the step of determining that a subject is prone to cancer. In other embodiments, the method includes the step of selecting a treatment for a subject. In other embodiments, the method includes the step of administering a treatment to a subject.

[0010] A method is described herein that includes the steps of detecting methylation at at least one of several sites; generating several methylation calls for each of the several sites; obtaining one or more metrics from the methylation calls; processing one or more metrics to generate a probability that a patient exhibits PD-L1 expression; and determining that the patient is a candidate for treatment with jedatricib and talazoparib. In other embodiments, jedatricib sensitizes advanced TNBC or BRCA1 / 2 mutant breast cancer to talazoparib-mediated PARP inhibition.

[0011] In other embodiments, cancer is breast cancer, bladder cancer, cervical cancer, colon cancer, head and neck cancer, Hodgkin lymphoma, liver cancer, lung cancer, renal cell carcinoma, skin cancer including melanoma, gastric cancer, rectal cancer, and any solid tumor in which errors in DNA that occur during copying cannot be repaired. In other embodiments, the sample comprises cell-free DNA. In other embodiments, the method comprises the step of diagnosing that a subject has cancer. In other embodiments, the method comprises the step of determining the prognosis of a subject to cancer. In other embodiments, the method comprises the step of selecting a treatment for the subject. In other embodiments, the method comprises the step of performing a treatment on the subject.

[0012] This specification describes a method comprising the steps of: detecting nucleosomal positioning in at least one of several genomic regions to generate a nucleosome occupancy profile of the genomic region; obtaining one or more metrics from the nucleosome occupancy profile; and processing one or more metrics to generate the probability that a patient exhibits PD-L1 expression.

[0013] In various embodiments, the method is SKI, THEMIS2, RPA2, TEKT2, STK40, GJA9-MYCBP, LOC105378663, HEYL, CNN3, JTB, FAM78B, ARV1, ADSS2, ZNF672, MBOAT2, ASXL2, SERTAD2, TMEM131, CLASP1, SATB2, ABHD14B, NISCH, TMEM45A, LGI2, KLHL5, NEUROG2-AS1, ABHD18, MFSD8, ELF2, T RIM2, AHRR, PDCD6-AHRR, SEMA5A, IQGAP2, TSLP, SLC25A48, RELL2, ARHGAP26, SLC36A1, CNPY3, FAM229B, MAN1A1, ADCYAP1R1, KIA A0895, TRAPPC14, LINC01004, FAM131B, GIMAP4, SLC4A2, CD274, TOX, GDAP1, ZNF623, GNA14, S1PR3, C9orf47, ROR2, ERCC6L2, LINC 00476, ECPAS, ASTN2, PHF19, PTGES2-AS1, RALGDS, HACD1, ABLIM1, LOC101927692, GFRA1, C11orf21, TRIM44, CHST1, TMX2-CTNND 1, LOC101928069, PDE2A, DLG2, ENDOD1, DDX6, TULP3, PTPRO, ZCRB1, TMPO-AS1, HSP90B1, SIRT4, SRSF9, SLITRK1, MMP14, BCL2L2- In other embodiments, the method includes at least one of a plurality of sites located in one or more genes selected from the group consisting of PABPN1, KCNH5, TRAF3, IDH2, CIB1, MAN2A2, KDM8, ZFHX3, HSBP1, TOP3A, RETREG3, ADAM11, KPNB1, GRIN2C, GALR2, ZBTB14, EPB41L3, PDE4A, KLF1, SIX5, DM1-AS, ZNF114, CLEC11A, and LINC01530. In other embodiments, the plurality of sites are located in one or more genes that are targets of hsa-miR-6132, hsa-miR-6836-5p, hsa-miR-1909-3p, and / or hsa-miR-6722-3p. In other embodiments, the method includes the step of diagnosing that a subject has cancer.In other embodiments, the method includes the step of determining the prognosis of the subject as being susceptible to cancer. In other embodiments, the method includes the step of selecting a treatment for the subject. In other embodiments, the method includes the step of performing a treatment on the subject.

[0014] Methods for diagnosing cancer in an individual, determining the prognosis of cancer, and determining susceptibility to cancer, comprising the step of determining the presence or absence of high levels of expression of a microRNA compared to a normal baseline standard in the individual, are described herein. In various embodiments, the microRNA comprises one or more of hsa-miR-6132, hsa-miR-6836-5p, hsa-miR-1909-3p, and hsa-miR-6722-3p. In other embodiments, the method comprises the step of selecting a treatment for a subject. In other embodiments, the method comprises the step of performing the treatment on the subject.

[0015] For example, a method for selecting a therapeutic treatment for a subject is disclosed, comprising the steps of: a) obtaining a biological sample from the subject; b) determining the expression levels of hsa-miR-6132, hsa-miR-6836-5p, hsa-miR-1909-3p, and / or hsa-miR-6722-3p in the biological sample; c) comparing the expression levels with a control; and d) selecting a therapeutic treatment based on the comparison. In some embodiments, the control is derived from one or more subjects without cancer. In some embodiments, the expression levels are inferred from cfDNA analysis, for example, by nucleosome positioning analysis. In some embodiments, the cancer is non-small cell lung cancer (NSCLC). In some embodiments, abnormal expression of hsa-miR-6132, hsa-miR-6836-5p, hsa-miR-1909-3p, and / or hsa-miR-6722-3p indicates that the subject is resistant to therapeutic treatment (e.g., tyrosine kinase inhibitors, e.g., EGFR inhibitors, e.g., osimertinib), and therefore, an alternative therapeutic treatment is selected. In some embodiments, the method involves administering the alternative treatment. In some embodiments, the method involves determining the expression of hsa-miR-6132, hsa-miR-1909-3p, and / or hsa-miR-6722-3p.

[0016] Also disclosed is a method for determining whether a subject is resistant to therapeutic treatment, comprising the steps of: a) obtaining a biological sample from the subject; b) determining the expression levels of hsa-miR-6132, hsa-miR-6836-5p, hsa-miR-1909-3p, and / or hsa-miR-6722-3p in the biological sample; c) comparing the expression levels with a control; and d) classifying the subject as resistant to therapeutic treatment based on the comparison. In some embodiments, the control is derived from one or more subjects without cancer. In some embodiments, the expression levels are inferred from cfDNA analysis, for example, by nucleosome positioning analysis. In some embodiments, the cancer is non-small cell lung cancer (NSCLC). In some embodiments, abnormal expression of hsa-miR-6132, hsa-miR-6836-5p, hsa-miR-1909-3p, and / or hsa-miR-6722-3p indicates that the subject is resistant to therapeutic treatment (e.g., tyrosine kinase inhibitors, e.g., EGFR inhibitors, e.g., osimertinib). In some embodiments, the method involves administering an alternative treatment. In some embodiments, the method involves determining the expression of hsa-miR-6132, hsa-miR-1909-3p, and / or hsa-miR-6722-3p.

[0017] The use of hsa-miR-6132, hsa-miR-6836-5p, hsa-miR-1909-3p, and / or hsa-miR-6722-3p as cancer biomarkers is also disclosed. In some embodiments, the biomarker relates to resistance to therapeutic treatment for cancer. In some embodiments, the cancer is non-small cell lung cancer. In some embodiments, the therapeutic treatment is a tyrosine kinase inhibitor, e.g., an EGFR inhibitor, e.g., osimertinib.

[0018] A method for selecting a therapeutic treatment for a subject, comprising the steps of a) obtaining a biological sample from the subject, and b) obtaining SKI, THEMIS2, RPA2, TEKT2, STK40, GJA9-MYCBP, LOC105378663, HEYL, CNN3, JTB, FAM78B, ARV1, ADSS2, ZNF672, MBOAT2, ASXL2, SERTAD2, TMEM131, CLASP1, SATB2, ABHD14B, NISCH, TMEM45A, LGI2, KLHL5, NEUROG 2-AS1, ABHD18, MFSD8, ELF2, TRIM2, AHRR, PDCD6-AHRR, SEMA5A, IQGAP2, TSLP, SLC25A48, RELL2, ARHGAP26, SLC36A1, CNPY3, FAM229B, MAN 1A1, ADCYAP1R1, KIAA0895, TRAPPC14, LINC01004, FAM131B, GIMAP4, SLC4A2, CD274, TOX, GDAP1, ZNF623, GNA14, S1PR3, C9orf47, ROR2, ER CC6L2, LINC00476, ECPAS, ASTN2, PHF19, PTGES2-AS1, RALGDS, HACD1, ABLIM1, LOC101927692, GFRA1, C11orf21, TRIM44, CHST1, TMX2-CT NND1, LOC101928069, PDE2A, DLG2, ENDOD1, DDX6, TULP3, PTPRO, ZCRB1, TMPO-AS1, HSP90B1, SIRT4, SRSF9, SLITRK1, MMP14, BCL2L2-PABPN A method is also disclosed that includes the steps of: 1) determining the expression levels of KCNH5, TRAF3, IDH2, CIB1, MAN2A2, KDM8, ZFHX3, HSBP1, TOP3A, RETREG3, ADAM11, KPNB1, GRIN2C, GALR2, ZBTB14, EPB41L3, PDE4A, KLF1, SIX5, DM1-AS, ZNF114, CLEC11A, and LINC01530; c) comparing the expression levels with a control; and d) selecting a therapeutic treatment based on the comparison.Also disclosed is a method for selecting a therapeutic treatment for a subject, comprising the steps of: a) obtaining a biological sample from the subject; b) determining the expression levels of SKI, SEMA5A, FAM131B, SLC4A2, CLASP1, and / or HSP90B1 in the biological sample; c) comparing the expression levels with a control; and d) selecting a therapeutic treatment based on the comparison. In some embodiments, the control is derived from one or more subjects without cancer. In some embodiments, the expression levels are inferred from cfDNA analysis, for example, by nucleosome positioning analysis. In some embodiments, the cancer is non-small cell lung cancer (NSCLC). In some embodiments, abnormal expression of SKI, SEMA5A, FAM131B, SLC4A2, CLASP1, and / or HSP90B1 indicates that the subject is resistant to a therapeutic treatment (e.g., a tyrosine kinase inhibitor, e.g., an EGFR inhibitor, e.g., osimertinib), and therefore an alternative therapeutic treatment is selected. In some embodiments, the method involves administering an alternative treatment.

[0019] A method for determining whether a subject is resistant to therapeutic treatment, comprising the steps of a) obtaining a biological sample from the subject, and b) identifying SKI, THEMIS2, RPA2, TEKT2, STK40, GJA9-MYCBP, LOC105378663, HEYL, CNN3, JTB, FAM78B, ARV1, ADSS2, ZNF672, MBOAT2, ASXL2, SERTAD2, TMEM131, CLASP1, SATB2, ABHD14B, NISCH, TMEM45A, LGI2, KLHL5, N EUROG2-AS1, ABHD18, MFSD8, ELF2, TRIM2, AHRR, PDCD6-AHRR, SEMA5A, IQGAP2, TSLP, SLC25A48, RELL2, ARHGAP26, SLC36A1, CNPY3, FAM229B, M AN1A1, ADCYAP1R1, KIAA0895, TRAPPC14, LINC01004, FAM131B, GIMAP4, SLC4A2, CD274, TOX, GDAP1, ZNF623, GNA14, S1PR3, C9orf47, ROR2, ERCC 6L2, LINC00476, ECPAS, ASTN2, PHF19, PTGES2-AS1, RALGDS, HACD1, ABLIM1, LOC101927692, GFRA1, C11orf21, TRIM44, CHST1, TMX2-CTNND1, L OC101928069, PDE2A, DLG2, ENDOD1, DDX6, TULP3, PTPRO, ZCRB1, TMPO-AS1, HSP90B1, SIRT4, SRSF9, SLITRK1, MMP14, BCL2L2-PABPN1, KCNH5, T A method is also disclosed that includes the steps of: a) determining the expression levels of RAF3, IDH2, CIB1, MAN2A2, KDM8, ZFHX3, HSBP1, TOP3A, RETREG3, ADAM11, KPNB1, GRIN2C, GALR2, ZBTB14, EPB41L3, PDE4A, KLF1, SIX5, DM1-AS, ZNF114, CLEC11A, and LINC01530; c) comparing the expression levels with a control; and d) classifying the subject as resistant to therapeutic treatment based on the comparison.Also disclosed is a method for determining whether a subject is resistant to therapeutic treatment, comprising the steps of: a) obtaining a biological sample from the subject; b) determining the expression levels of SKI, SEMA5A, FAM131B, SLC4A2, CLASP1, and / or HSP90B1 in the biological sample; c) comparing the expression levels with a control; and d) classifying the subject as resistant to therapeutic treatment based on the comparison. In some embodiments, the control is derived from one or more subjects without cancer. In some embodiments, the expression levels are inferred from cfDNA analysis, for example, by nucleosome positioning analysis. In some embodiments, the cancer is non-small cell lung cancer (NSCLC). In some embodiments, abnormal expression of SKI, SEMA5A, FAM131B, SLC4A2, CLASP1, and / or HSP90B1 indicates that the subject is resistant to therapeutic treatment (e.g., tyrosine kinase inhibitors, e.g., EGFR inhibitors, e.g., osimertinib). In some embodiments, the method involves administering an alternative treatment.

[0020] SKI, THEMIS2, RPA2, TEKT2, STK40, GJA9-MYCBP, LOC105378663, HEYL, CNN3, JTB, FAM78B, ARV1, ADSS2, ZNF672, MBOAT2, A SXL2, SERTAD2, TMEM131, CLASP1, SATB2, ABHD14B, NISCH, TMEM45A, LGI2, KLHL5, NEUROG2-AS1, ABHD18, MFSD8, ELF2, TRIM 2, AHRR, PDCD6-AHRR, SEMA5A, IQGAP2, TSLP, SLC25A48, RELL2, ARHGAP26, SLC36A1, CNPY3, FAM229B, MAN1A1, ADCYAP1R1, K IAA0895, TRAPPC14, LINC01004, FAM131B, GIMAP4, SLC4A2, CD274, TOX, GDAP1, ZNF623, GNA14, S1PR3, C9orf47, ROR2, ERCC 6L2, LINC00476, ECPAS, ASTN2, PHF19, PTGES2-AS1, RALGDS, HACD1, ABLIM1, LOC101927692, GFRA1, C11orf21, TRIM44, CH ST1, TMX2-CTNND1, LOC101928069, PDE2A, DLG2, ENDOD1, DDX6, TULP3, PTPRO, ZCRB1, TMPO-AS1, HSP90B1, SIRT4, SRSF9, SL The use of ITRK1, MMP14, BCL2L2-PABPN1, KCNH5, TRAF3, IDH2, CIB1, MAN2A2, KDM8, ZFHX3, HSBP1, TOP3A, RETREG3, ADAM11, KPNB1, GRIN2C, GALR2, ZBTB14, EPB41L3, PDE4A, KLF1, SIX5, DM1-AS, ZNF114, CLEC11A, and LINC01530 as cancer biomarkers is also disclosed. The use of SKI, SEMA5A, FAM131B, SLC4A2, CLASP1, and / or HSP90B1 as cancer biomarkers is also disclosed. In some embodiments, the biomarkers relate to resistance to therapeutic treatment for cancer. In some embodiments, the cancer is non-small cell lung cancer.In some embodiments, the therapeutic treatment is a tyrosine kinase inhibitor, such as an EGFR inhibitor, such as osimertinib.

[0021] A method for selecting a therapeutic treatment for a subject is also disclosed, comprising the steps of: a) obtaining a biological sample from the subject; b) determining the expression level of one or more components of the ER-associated degradation (ERAD) pathway in the biological sample; c) comparing the expression level with a control; and d) selecting a therapeutic treatment based on the comparison. In some embodiments, the control is derived from one or more subjects without cancer. In some embodiments, the expression level is inferred from cfDNA analysis, for example, by nucleosome positioning analysis. In some embodiments, the cancer is non-small cell lung cancer (NSCLC). In some embodiments, abnormal expression of one or more components of the ER-associated degradation (ERAD) pathway indicates that the subject is resistant to a therapeutic treatment (e.g., a tyrosine kinase inhibitor, e.g., an EGFR inhibitor, e.g., osimertinib), and therefore an alternative therapeutic treatment is selected. In some embodiments, the method involves administering the alternative treatment. In some embodiments, one or more components of the ERAD path are MAN2A2 and / or MAN1A1.

[0022] Also disclosed is a method for determining whether a subject is resistant to therapeutic treatment, comprising the steps of: a) obtaining a biological sample from the subject; b) determining the expression level of one or more components of the ER-associated degradation (ERAD) pathway in the biological sample; c) comparing the expression level to a control; and d) classifying the subject as resistant to therapeutic treatment based on the comparison. In some embodiments, the control is derived from one or more subjects without cancer. In some embodiments, the expression level is inferred from cfDNA analysis, for example, by nucleosome positioning analysis. In some embodiments, the cancer is non-small cell lung cancer (NSCLC). In some embodiments, abnormal expression of one or more components of the ER-associated degradation (ERAD) pathway indicates that the subject is resistant to therapeutic treatment (e.g., tyrosine kinase inhibitors, e.g., EGFR inhibitors, e.g., osimertinib). In some embodiments, the method involves administering an alternative treatment. In some embodiments, one or more components of the ERAD path are MAN2A2 and / or MAN1A1.

[0023] The use of one or more components of the ER-associated degradation (ERAD) pathway as a cancer biomarker is also disclosed. In some embodiments, the biomarker relates to resistance to therapeutic treatment for cancer. In some embodiments, the cancer is non-small cell lung cancer. In some embodiments, the therapeutic treatment is a tyrosine kinase inhibitor, e.g., an EGFR inhibitor, e.g., osimertinib. In some embodiments, one or more components of the ERAD pathway are MAN2A2 and / or MAN1A1.

[0024] In some embodiments, reports are created using the results of the systems and methods disclosed herein as input. The reports may be on paper or in electronic form. For example, such reports can directly display genetic results determined by the methods and systems disclosed herein, such as the presence of nucleic acid variants detected in a sample. In some embodiments, such reports display only the presence or absence of a disease such as cancer.

[0025] The various steps of the methods disclosed herein, or the steps performed by the systems disclosed herein, can be performed at the same or different times, in the same or different geographical locations, e.g., countries, and / or by the same or different persons. In some embodiments, the report is notified to a subject, e.g., a subject having cancer who has been examined by the methods and systems described herein, or a medical professional such as a physician treating a subject having cancer.

Brief Description of the Drawings

[0026] [Figure 1] PD-L1 PD-1 / PD-L1 Signaling: Reduced CD8+ T cell proliferation, survival, and cytokine production. Abbreviations: DC, dendritic cell; Treg, regulatory T cell; ICOS, inducible co-stimulator; ICOS-L, inducible co-stimulator-ligand; CD28, surface antigen classification 28; CTLA-4, cytotoxic T lymphocyte-associated antigen-4; PD-L1, programmed death-ligand 1; PD-1, programmed death-1; MHC, major histocompatibility complex; TCR, T cell receptor; IFN-γ, interferon-γ; IFN-γR, interferon-γ receptor.

[0027] [Figure 2] Liquid biopsy genomic data. When using existing genomic panels, the predictive power is limited because the number of probes in the PD-L1 promoter region is limited.

[0028] [Figure 3]PD-L1 prediction model. For model generation, 384 CRC samples were analyzed using methylation data from a 450k Illumina microarray (single-site bisulfite sequencing) and gene expression from RNASeq (normalized). Beta values ​​can be used as a comparison with other epigenome detection platforms. Subsequently, each sample was labeled as PD-L1-low (bottom 50%), PD-L1-medium (25%), or PD-L1-high (top 25%) based on CD274 gene expression. After 10-fold cross-validation, the inventors used a penalized logistic regression model (LASSO) with sample ID PD-L1 high or low as the response variable and the methylation score (beta) of all Infinity-targeted regions as the predictor variable.

[0029] [Figure 4] A PD-L1 prediction model based on methylation data. Approximately 40 target regions were selected as predictor variables using LASSO.

[0030] [Figure 5] PD-L1 gene expression from methylation data. The inventors established a prediction of PD-L1 gene expression using a 10-fold cross-validation, penalized logistic regression model (LASSO), with sample PD-L1 gene expression as the response variable and the methylation score (beta) of all targeted regions as the predictor variable. Here, 52 regions were selected by LASSO.

[0031] [Figure 6] PD-L1 status correlated with BRAF V600E and MSI-H status. Here, the inventors identified the association between PD-L1 promoter methylation status and MSI-H and BRCA. Of the samples tested, 39 were MSI-H (mainly CRC and breast) and 311 were BRAF V600E positive (mainly CRC).

[0032] [Figure 7]Prediction of MSI-H and BRCA status from methylation data in lung cancer.

[0033] [Figure 8] Nucleosome location of the PD-L1 promoter region. Here, we generated molecular center point (a possible indicator of nucleosome location) coverage from low-level reads. Average coverage profile from a subset of samples.

[0034] [Figure 9] Nucleosome locations. Coverage profiles for individual PD-L1 samples are shown.

[0035] [Figure 10] Gene set enrichment analysis using miRNA target interactions was performed using the gp.enrichr function from the gseapy library. The axis of the image represents the negative logarithm (base 10) of the corrected p-value, shown as -log10(corrected p-value). A larger value on this axis indicates a higher statistical significance of the result, and a lower p-value (higher -log10 value) indicates stronger evidence against the null hypothesis that miRNAs do not regulate genes in the input set. 2 represents a corrected p-value of 0.01. MicroRNAs (miRNAs) are small non-coding RNA molecules that regulate gene expression by binding to complementary sequences on target mRNAs, typically resulting in the silencing of the target mRNA through translational repression or targeted degradation.

[0036] [Figure 11] The microRNAs hsa-miR-6132, hsa-miR-6836-5p, hsa-miR-1909-3p, and hsa-miR-6722-3p are significant regulators of genes within our set (corrected P-value < 0.05).

[0037] Blue indicates an regulatory relationship between the gene and the miR. Yellow indicates no regulatory relationship between the gene and the miR. Rows correspond to the genes in the set, and columns correspond to the microRNAs. [Modes for carrying out the invention]

[0038] Detailed explanation While various embodiments of the Disclosure are shown and described herein, it will be understood by those skilled in the art that such embodiments are presented merely as examples. Those skilled in the art can conceive of numerous variations, alterations, and substitutions without departing from the Disclosure. It should be understood that various alternatives to the embodiments of the Disclosure described herein are available.

[0039] With respect to a reference number, the term "approximately" and its grammatical equivalents can encompass a range of values ​​up to plus or minus 10% from that value. For example, the quantity "approximately 10" can encompass quantities from 9 to 11. With respect to a reference number, the term "approximately" can encompass a range of values ​​up to plus or minus 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% from that value.

[0040] With respect to a reference number, the term "at least" and its grammatical equivalents can encompass the reference number and any value greater than that. For example, the quantity "at least 10" can encompass the value of 10 and any number greater than 10, such as 11, 100, and 1,000.

[0041] With respect to a reference number, the term "at most" and its grammatical equivalent can encompass the reference number and any values ​​less than that value. For example, the quantity "at most 10" can encompass the value 10 and any numbers less than 10, such as 9, 8, 5, 1, 0.5, and 0.1.

[0042] As used herein, the singular forms “a,” “an,” and “the” may encompass multiple references unless the context explicitly indicates otherwise. For example, a reference to “a cell” may encompass multiple such cells, and a reference to “the culture” may encompass one or more cultures and their equivalents known to those skilled in the art. All scientific and technical terms used herein may have the same meaning as commonly understood by those skilled in the art in which this disclosure is made, unless explicitly indicated otherwise.

[0043] As described, in cancer diagnosis, blood-based assessment of PD-L1 status in tumor types such as NSCLC or other distant cancers is of great benefit. This specification describes a method for constructing a binary classification model from methylation data by using normalized molecular counts from an epigenome panel in high-segmentation areas as a measure of methylation. It should be noted that since PD-L1 is expressed in both normal and cancer cells, a limited number of hypermolecules are present in the promoter region, and the PD-L1 TSS region may be hypomethylated.

[0044] Current methods either omit testing both genomic and epigenomic attributes of patient samples, or perform multiple tests separately. Omitting genomic or epigenomic information may lead to the prescription of cancer treatments that would have been known to be ineffective, or to the withholding of cancer treatments that would have been known to be effective, if both genomic and epigenomic information were available. Cancer can be indicated by epigenetic variations, such as methylation. An example of methylation changes in cancer is the localized acquisition of DNA methylation in CpG islands at transcription start sites (TSSs) of genes involved in normal growth regulation, DNA repair, cell cycle regulation, and / or cell differentiation. This hypermethylation may be associated with an abnormal loss of transcriptional ability of the genes involved and occurs at least as frequently as point mutations and deletions as a cause of altered gene expression. DNA methylation profiling can be used to detect regions of the genome with different degrees of methylation ("differential methylation regions" or "DMRs") that are altered during development or perturbed by disease, such as cancer or any cancer-related disease. The genomes of cancer cells exhibit imbalances in the DNA methylation patterns described above, and therefore in the functional packaging of DNA. Thus, abnormalities in chromatin organization, when analyzed together with methylation changes, can contribute to the enhancement of cancer profiling. MBD distribution and fragment mix data, such as fragments mapped to start and stop positions (correlated with nucleosome locations), fragment length, and associated nucleosome occupancy, can be used for chromatin structure analysis in hypermethylation studies with the aim of improving biomarker detection rates.

[0045] Methylation profiling can involve determining methylation patterns across different regions of the genome. For example, molecules can be partitioned based on their degree of methylation (e.g., the relative number of methylation sites per molecule), sequenced, and then the sequences of molecules in different partitions can be mapped to a reference genome. This can reveal regions of the genome that are more or less methylated than other regions. Thus, genomic regions can have different degrees of methylation, in contrast to individual molecules.

[0046] A defining characteristic of nucleic acid molecules is modification, which can include various chemical modifications or protein modifications (i.e., epigenetic modifications). Non-exclusive examples of chemical modifications include, but are not limited to, covalent DNA modifications, including DNA methylation. In some embodiments, DNA methylation involves the addition of a methyl group to cytosine at a CpG site (where guanine follows cytosine in the nucleic acid sequence). In some embodiments, DNA methylation involves the addition of a methyl group to adenine, such as N6-methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the fifth carbon in the six-membered carbon ring of cytosine). In some embodiments, 5-methylation involves the addition of a methyl group to the 5C position of cytosine, thereby creating 5-methylcytosine (m5c). In some embodiments, methylation includes derivatives of m5c. Derivatives of m5c include, but are not limited to, 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5-carboxylcytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the third carbon of the 6-membered carbon ring of cytosine). In some embodiments, 3C methylation involves the addition of a methyl group to the 3C position of cytosine, thereby producing 3-methylcytosine (3mC). Other examples include N6-methyladenine or glycosylation. DNA methylation involves the addition of a methyl group to DNA (e.g., CpG), which can alter the expression of methylated DNA regions. Methylation can also occur at non-CpG sites; for example, methylation can occur at CpA, CpT, or CpC sites. DNA methylation can alter the activity of methylated DNA regions. For example, methylation of DNA within a promoter region can suppress gene transcription. DNA methylation is crucial for normal development, and abnormalities in methylation can disrupt epigenetic regulation. This disruption, or suppression, of epigenetic regulation may lead to diseases such as cancer. Promoter methylation in DNA can be an indicator of cancer.

[0047] A CpG dyad is a dinucleotide CpG (cytosine-phosphate-guanine, i.e., cytosine followed by guanine in the 5'→3' direction of the nucleic acid sequence) on the sense strand of a double-stranded DNA molecule and its complementary CpG on the antisense strand. CpG dyads can be fully methylated or semi-methylated (methylated on only one strand).

[0048] CpG dinucleotides are relatively rare in the normal human genome, and the vast majority of CpG dinucleotide sequences are transcriptionally inactive (e.g., in DNA heterochromatin regions and repeat elements around the centromere of chromosomes) and are also methylated. However, many CpG islands are protected from such methylation, particularly around transcription start sites (TSSs).

[0049] Protein modifications include binding to chromatin components, particularly histones, including their modified forms, and binding to other proteins, such as proteins involved in replication or transcription. This disclosure provides a method for processing and analyzing nucleic acids with varying degrees of modification, wherein the nature of the original modification of the nucleic acid is correlated with a nucleic acid tag, which can be deciphered by sequencing during nucleic acid analysis. The genetic variation of the sample nucleic acid modification can then be associated with the degree of modification of that nucleic acid in the original sample (epigenetic variation). This includes single-stranded (e.g., ssDNA or RNA) or double-stranded (e.g., dsDNA) molecules.

[0050] DNA loss can reduce the presence of one or more types of DNA, making it difficult to detect the presence of one or more types of DNA, such as cfDNA. In one or more additional scenarios, existing methods for measuring DNA methylation, such as enrichment or depletion methods, may have relatively high resolutions, such as about 100 base pairs (bp) to about 200 bp, but it may be difficult to accurately determine the amount of DNA methylation. The accuracy of determining DNA methylation may affect the accuracy of the tumor fraction estimate of a sample. Since tumor fraction can be used to determine whether a sample originates from an object in which a tumor is present, the accuracy of determining the tumor fraction estimate may affect diagnostic and / or treatment decisions for an individual. Fee

[0051] The sample may be any biological sample isolated from the subject. The sample may be a body sample. The sample may include body tissues, such as known or suspected solid tumors, whole blood, platelets, serum, plasma, stool, red blood cells, white blood cells or leucocytes, endothelial cells, tissue biopsy material, cerebrospinal fluid, synovial fluid, lymph, ascites, interstitial fluid or extracellular fluid, intercellular fluid including gingival crevicular exudate, bone marrow, pleural fluid, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. Preferably, the sample is a body fluid, particularly blood and its fractions, as well as urine. The sample may be in the form originally isolated from the subject, or may have been subjected to further processing to remove or add components such as cells, or to enrich one component with another. Therefore, preferred body fluids for analysis are plasma or serum containing cell-free nucleic acids. The sample can be isolated or obtained from the subject and transported to the sample analysis site. Samples can be stored and transported at a desired temperature, e.g., room temperature, 4°C, -20°C, and / or -80°C. Samples can be isolated or obtained from subjects at the sample analysis site. Subjects may be humans, mammals, animals, companion animals, service animals, or pets. Subjects may have cancer. Subjects may not have cancer or detectable cancer symptoms. Subjects may have been treated with one or more cancer treatments, e.g., chemotherapy, antibodies, vaccines, or biological agents. Subjects may be in remission. Subjects may or may not have been diagnosed as susceptible to cancer or any cancer-related genetic mutation / disorder.

[0052] The volume of plasma may depend on the desired read depth of the region to be sequenced. Exemplary volumes are 0.4–40 ml, 5–20 ml, and 10–20 ml. For example, the volume could be 0.5 mL, 1 mL, 5 mL, 10 mL, 20 mL, 30 mL, or 40 mL. The volume of sampled plasma may be 5–20 mL.

[0053] The sample may contain varying amounts of nucleic acids, including genome equivalents. For example, a sample of approximately 30 ng of DNA may contain approximately 10,000 (10⁴) haploid human genome equivalents, or approximately 200 billion (2 × 10¹¹) individual polynucleotide molecules in the case of cfDNA. Similarly, a sample of approximately 100 ng of DNA may contain approximately 30,000 haploid human genome equivalents, or approximately 600 billion individual molecules in the case of cfDNA.

[0054] The sample may include nucleic acids from different sources, e.g., nucleic acids and cell-free nucleic acids from the same target cells, or nucleic acids and cell-free nucleic acids from different target cells. The sample may include nucleic acids with mutations. For example, the sample may include DNA with germline mutations and / or somatic mutations. Germline mutations refer to mutations present in the germline DNA of the target. Somatic mutations refer to mutations originating from somatic cells of the target, e.g., cancer cells. The sample may include DNA with cancer-associated mutations (e.g., cancer-associated somatic mutations). The sample may include epigenetic variants (i.e., chemical or protein modifications), where the epigenetic variants are associated with the presence of genetic variants such as cancer-associated mutations. In some embodiments, the sample includes epigenetic variants associated with the presence of genetic variants, where the sample does not include genetic variants.

[0055] Exemplary amounts of cell-free nucleic acids in the sample before amplification range from approximately 1 fg to approximately 1 μg, for example, 1 pg to 200 ng, 1 ng to 100 ng, and 10 ng to 1000 ng. For example, the amount may be up to approximately 600 ng, up to approximately 500 ng, up to approximately 400 ng, up to approximately 300 ng, up to approximately 200 ng, up to approximately 100 ng, up to approximately 50 ng, or up to approximately 20 ng of cell-free nucleic acid molecules. The amount may be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng of cell-free nucleic acid molecules. The quantity may be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or 200 ng of cell-free nucleic acid molecules. The method may include obtaining 1 femtogram (fg) to 200 ng.

[0056] Cell-free nucleic acids are nucleic acids that are not contained within cells and are not bound to cells in any other form, or in other words, nucleic acids that remain in a sample after intact cells have been removed. Cell-free nucleic acids include DNA, RNA, and hybrids thereof, including genomic DNA, mitochondrial DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA chains (long ncRNA), or fragments of any of these. Cell-free nucleic acids may be double-stranded, single-stranded, or hybrids thereof. Cell-free nucleic acids may be released into body fluids by secretion or cell death processes, such as cell necrosis and apoptosis. Some cell-free nucleic acids are released into body fluids from cancer cells, such as circulating tumor DNA (ctDNA). Other cell-free nucleic acids are released from healthy cells. In some embodiments, cfDNA is cell-free embryonic DNA (cffDNA). In some embodiments, cell-free nucleic acids are produced by tumor cells. In some embodiments, cell-free nucleic acids are produced by a mixture of tumor cells and non-tumor cells.

[0057] Cell-free nucleic acids have an exemplary size distribution of approximately 100–500 nucleotides, with molecules of 110–approximately 230 nucleotides accounting for about 90% of the molecules, the mode being approximately 168 nucleotides, and a second minor peak in the range of 240–440 nucleotides. Cell-free nucleic acids can be isolated from body fluids by a fractionation or partitioning step that separates the cell-free nucleic acids found in solution from intact cells and other insoluble components of the body fluid. Partitioning may involve techniques such as centrifugation or filtration. Alternatively, cells in the body fluid can be lysed and the cell-free and cellular nucleic acids processed together. Generally, after buffer addition and washing steps, nucleic acids can be precipitated using alcohol. Furthermore, contaminants or salts can be removed using a cleanup step, such as a silica-based column. In certain aspects of the procedure, for example to optimize yield, nonspecific bulk carrier nucleic acids, such as Cot-1 DNA, DNA, or proteins for bisulfite sequencing, hybridization, and / or ligation, can be added throughout the reaction.

[0058] After such processing, the sample may contain various forms of nucleic acids, including double-stranded DNA, single-stranded DNA, and single-stranded RNA. In some embodiments, single-stranded DNA and RNA can be converted to double-stranded form for subsequent processing and analysis steps.

[0059] specimen The specimen may include nucleic acid specimens and non-nucleic acid specimens. This disclosure provides for the detection of genetic variations in biological specimens derived from the subject. The biological specimen may include polynucleotides derived from cancer cells. Polynucleotides may be DNA (e.g., genomic DNA, cDNA), RNA (e.g., mRNA, small RNA), or any combination thereof. The biological specimen may include tumor tissue derived from a biopsy, for example. In some cases, the biological specimen may include blood or saliva. In certain cases, the biological specimen may include cell-free DNA ("cfDNA") or circulating tumor DNA ("ctDNA"). Cell-free DNA may be present in blood, for example.

[0060] Examples of non-nucleic acid samples include, but are not limited to, lipids, carbohydrates, peptides, proteins, glycoproteins (N-linked or O-linked), lipoproteins, phosphorylated proteins, specific phosphorylated or acetylated variants of proteins, amidated variants of proteins, hydroxylated variants of proteins, methylated variants of proteins, ubiquitinated variants of proteins, sulfated variants of proteins, viral proteins (e.g., viral capsids, viral envelopes, viral coats, viral accessories, viral glycoproteins, viral spikes, etc.), extracellular and intracellular proteins, antibodies, and antigen-binding fragments. Non-nucleic acid samples include receptors, antigens, surface proteins, transmembrane proteins, surface antigen classification proteins, protein channels, protein pumps, carrier proteins, phospholipids, glycoproteins, glycolipids, intercellular interaction protein complexes, antigen presentation complexes, major histocompatibility complexes, engineered T cell receptors, T cell receptors, B cell receptors, chimeric antigen receptors, extracellular matrix proteins, post-translational modification states of cell surface proteins (e.g., phosphorylation, glycosylation, ubiquitination, nitrosylation, methylation, acetylation, or lipid addition), gap junctions, and adhesion junctions.

[0061] In general, any number of samples, including both nucleic acid and non-nucleic acid samples, can be analyzed using the systems, apparatus, methods, and compositions. For example, the number of samples analyzed may be at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 20, at least about 25, at least about 30, at least about 40, at least about 50, at least about 100, at least about 1,000, at least about 10,000, at least about 100,000 or more different samples present within a region of the sample or within individual features of the substrate. Methods for performing multiplexed assays to analyze two or more different samples will be discussed in later sections of this disclosure.

[0062] One or more nucleic acid and / or non-nucleic acid samples constitute a set of intermolecular interactions in the biological system under test (e.g., cells), which can be thought of as an "interactome"—intermolecular interactions that occur between molecules belonging to different biochemical families (proteins, nucleic acids, lipids, carbohydrates, etc.) and also between molecules within a given family. In various embodiments, the interactome is a protein-DNA interactome (a network formed by transcription factors (and DNA or chromatin regulatory proteins) and their target genes). In other embodiments, the interactome refers to a protein-protein interaction network (PPI) or protein-protein interaction network (PIN). The methods described herein enable the testing and analysis of interactomes. Techniques such as proteogenomics (e.g., whole-genome sequencing, whole-exome sequencing and RNA-seq, and mass spectrometry) can support testing of interactomes.

[0063] analysis This method can be used to diagnose a condition in a subject, particularly the presence of cancer; to characterize a condition (e.g., to stage cancer or determine cancer heterogeneity); to monitor the response to treatment of a condition; and to determine the risk of developing a condition or the prognosis of the subsequent course of the condition. This disclosure may also be useful in determining the efficacy of a particular treatment option. A successful treatment option may increase the amount of copy number variation or rare mutations detected in the subject's blood, as more cancer cells may be killed and DNA may be shed if the treatment is successful. In other cases, this may not occur. In another case, a particular treatment option may correlate over time with the genetic profile of the cancer. This correlation may be useful in selecting a treatment. Furthermore, if the cancer is observed to be in remission after treatment, this method can be used to monitor residual disease or disease recurrence.

[0064] The types and number of cancers that can be detected include blood cancers, brain cancers, lung cancers, skin cancers, nasal cancers, throat cancers, liver cancers, bone cancers, lymphomas, pancreatic cancers, intestinal cancers, rectal cancers, thyroid cancers, bladder cancers, kidney cancers, oral cancers, stomach cancers, solid tumors, heterogeneous tumors, and homogeneous tumors. The type and / or stage of cancer can be detected from genetic variations, including mutations, rare mutations, indels, copy number variations, base transpositions, translocations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structural changes, gene fusions, chromosome fusions, gene shortening, gene amplification, gene duplication, chromosomal damage, DNA damage, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5-methylcytosine.

[0065] Genetic and other specimen data can also be used to characterize specific forms of cancer. Cancers are often heterogeneous in both composition and staging. Genetic profiling data may enable the characterization of specific subtypes of cancer, which may be important in the diagnosis or treatment of that particular subtype. This information may also provide subjects or practitioners with clues regarding prognosis for specific types of cancer, enabling either the subject or practitioner to adapt treatment options as the disease progresses. Some cancers may progress and become more invasive and genetically unstable. Other cancers may remain benign, inactive, or quiescent. The systems and methods of this disclosure may be useful in determining disease progression.

[0066] This analysis is also useful in determining the efficacy of specific treatment options. A successful treatment option may increase the amount of copy number variations or rare mutations detected in the subject's blood, as more cancers may be killed and DNA may be released if the treatment is successful. In other cases, this may not occur. In other cases, a particular treatment option may correlate with the cancer's genetic profile over time. This correlation may be useful in selecting a treatment. Furthermore, if the cancer is observed to be in remission after treatment, this method can be used to monitor residual disease or disease recurrence.

[0067] This method can also be used to detect genetic variations in non-cancerous conditions. Immune cells, such as B cells, can undergo rapid clonal expansion in the presence of certain diseases. Clonal expansion can be monitored using copy number variation detection, and a particular immune state can be monitored. In this example, copy number variation analysis can be performed over time to generate a profile of how a particular disease may progress. Using the detection of copy number variations, or even rare mutations, it is possible to determine how a population of pathogens changes during the course of infection. This can be particularly important during chronic infections, such as HIV / AIDS or hepatitis infections, where the virus may change its life cycle and / or mutate into a more virulent form during the course of infection. This method can be used to determine or profile the rejection activity of the host body as immune cells attempt to destroy transplanted tissue, thereby monitoring the status of the transplanted tissue and modifying strategies for treating or preventing rejection.

[0068] Furthermore, the methods of this disclosure can be used to characterize heterogeneity of an abnormal condition in a subject. Such a method may include, for example, a step of generating a gene profile of extracellular polynucleotides derived from the subject, wherein the gene profile includes multiple data obtained from the analysis of copy number variations and rare mutations. In some embodiments, the abnormal condition is cancer. In some embodiments, the abnormal condition may result in a heterogeneous genomic population. In the example of cancer, it is known that several tumors may contain tumor cells at different stages of cancer. In other examples, heterogeneity may include multiple lesions of the disease. Again, in the example of cancer, multiple tumor lesions may be present, and perhaps one or more lesions are the result of metastases that have spread from the primary site.

[0069] This method can be used to generate or profile fingerprints or sets of data, which are summaries of genetic information derived from different cells in heterogeneous diseases. These data sets may include, alone or in combination with, analysis of copy number variations and mutations.

[0070] This method can be used to diagnose, prognose, monitor, or observe cancer or other diseases. In some embodiments, the methods described herein do not involve diagnosing, prognosing, or monitoring a fetus, and therefore do not apply to non-invasive prenatal testing. In other embodiments, these methodologies can be used to diagnose, prognose, monitor, or observe cancer or other diseases in unborn subjects, in pregnant subjects, whose DNA and other polynucleotides may co-circulate with maternal molecules.

[0071] Determination of the 5-methylcytosine pattern of nucleic acids Bisulfite-based sequencing and its variations provide means for determining the methylation patterns of nucleic acids. In some embodiments, determining the methylation pattern includes distinguishing 5-methylcytosine (5mC) from unmethylated cytosine. In some embodiments, determining the methylation pattern includes distinguishing N6-methyladenine from unmethylated adenine. In some embodiments, determining the methylation pattern includes distinguishing 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), and 5-carboxylcytosine (5caC) from unmethylated cytosine. Examples of bisulfite sequencing, but not limited to these, include oxidative bisulfite sequencing (OX-BS-seq), Tet-assisted bisulfite sequencing (TAB-seq), and reductive bisulfite sequencing (redBS-seq).

[0072] Oxidative bisulfite sequencing (OX-BS-seq) is used to distinguish between 5mC and 5hmC by first converting 5hmC to 5fC and then proceeding with bisulfite sequencing as previously described. Tet-assisted bisulfite sequencing (TAB-seq) can also be used to distinguish between 5mC and 5hmC. In TAB-seq, 5hmC is protected by glucosylation. Then, 5mC is converted to 5caC using the Tet enzyme, and then bisulfite sequencing is proceeded as previously described. Reductive bisulfite sequencing is used to distinguish between 5fC and modified cytosine.

[0073] Generally, bisulfite sequencing involves dividing a nucleic acid sample into two aliquots and treating one aliquot with bisulfite. Bisulfite converts native cytosines and certain modified cytosine nucleotides (e.g., 5-formylcytosine or 5-carboxylcytosine) to uracil, while other modified cytosines (e.g., 5-methylcytosine, 5-hydroxymethylcytosine) are not. Comparison of the nucleic acid sequences of molecules from the two aliquots reveals which cytosines were converted to uracil and which were not. Consequently, modified and unmodified cytosines can be determined. Initially dividing the sample into two aliquots is inconvenient for samples containing only small amounts of nucleic acids and / or composed of heterogeneous cell / tissue origins, such as body fluids containing cell-free DNA.

[0074] This disclosure provides bisulfite sequencing and methods therefor, in which variations thereof are permitted. These methods function by linking nucleic acids in a population to a capture moiety, i.e., a label that can be captured or immobilized. Examples of capture moieties include, but are not limited to, biotin, avidin, streptavidin, nucleic acids containing specific nucleotide sequences, haptens recognized by antibodies, and magnetically attractable particles. The extraction moiety may be a member of a binding pair, such as biotin / streptavidin or hapten / antibody. In some embodiments, the capture moiety attached to the sample is captured by its binding pair attached to an isolateable portion, such as a magnetically attractable particle or a large particle that can be settled by centrifugation. The capture moiety may be any type of molecule that enables affinity separation of nucleic acids having the capture moiety from nucleic acids lacking the capture moiety. Exemplary capture moieties are biotin that enables affinity separation by binding to streptavidin linked to or linkable to a solid phase, or oligonucleotides that enable affinity separation by binding to complementary oligonucleotides linked to or linkable to a solid phase. After the capture portion is attached to the sample nucleic acid, the sample nucleic acid functions as an amplification template. After amplification, the original template remains attached to the capture portion, but the amplicon is not attached to the capture portion.

[0075] The capture portion can be ligated to the sample nucleic acid as a component of the adapter, thereby also providing amplification and / or sequencing primer binding sites. In some methods, both ends of the sample nucleic acid are ligated to the adapter, and both adapters have capture portions. Preferably, any cytosine residues in the adapter are modified with, for example, 5-methylcytosine to protect them from the action of bisulfites. In some cases, the capture portion is ligated to the original template by a cleavable linkage (e.g., a photocleavable desthiobiotin-TEG or a uracil residue cleavable with USER® enzyme, Chem. Commun. (Camb). 2015 Feb 21; 51(15): 3266-3269), in which case the capture portion can be removed as desired.

[0076] The amplicon is denatured and brought into contact with an affinity reagent for the capture tag. The original template binds to the affinity reagent, but the nucleic acid molecule produced by amplification does not. Therefore, the original template can be separated from the nucleic acid molecule produced by amplification.

[0077] After separation or segregation, each population of nucleic acids (i.e., the original template and the amplified product) can be subjected to bisulfite treatment, with the original template population being treated and the amplified product not. Alternatively, the amplified product can be subjected to bisulfite treatment, while the original template population is not. After such treatment, each population can be amplified (in the case of the original template population, uracil is converted to thymine). The populations can also be subjected to biotin probe hybridization for enrichment. Then, each population is analyzed and its sequences are compared to determine which cytosines were 5-methylated (or 5-hydroxymethylated) in the original population. Unmodified C is indicated by the detection of T nucleotides (corresponding to unmethylated cytosines converted to uracil) in the template population and C nucleotides at the corresponding positions in the amplified population. Modified C is indicated in the original sample by the presence of C at the corresponding positions in the original template and the amplified population.

[0078] In some embodiments, the method utilizes sequential DNA-seq and bisulfite-seq (BIS-seq) NGS library preparation of molecularly tagged DNA libraries. This process is carried out by labeling an adapter (e.g., biotin), DNA-seq amplification of the entire library, parent molecule recovery (e.g., streptavidin bead pulldown), bisulfite conversion, and BIS-seq. In some embodiments, the method identifies 5-methylcytosine at single-nucleotide resolution by sequential NGS preparative amplification of parent library molecules with and without bisulfite treatment. This can be achieved by modifying one of the two adapter strands of the 5-methylated NGS adapter (directional adapter; Y-shaped / fork-shaped due to the substitution of 5-methylcytosine) used in BIS-seq with a label (e.g., biotin). The adapter is then ligated to the sample DNA molecule and amplified (e.g., by PCR). Since only the parent molecule has a labeled adapter end, the parent molecule can be selectively recovered from its amplified offspring by a label-specific capture method (e.g., streptavidin-magnetic beads). Because the parent molecule retains a 5-methylation mark, bisulfite conversion of the captured library provides 5-methylation status at single-nucleotide resolution during BIS-seq, and the molecular information is preserved in the corresponding DNA-seq. In some embodiments, enrichment / NGS can be performed in a standard multiplexed NGS workflow by combining the bisulfite-treated library with an untreated library and then adding a sample-tagged DNA sequence. Bioinformatics analysis can be performed for genomic alignment and 5-methylated base identification, as in the BIS-seq workflow. In short, this method provides the ability to selectively recover ligated molecules with the parent 5-methylcytosine mark after library amplification, thereby enabling parallel processing of bisulfite-converted DNA. This overcomes the destructive nature of bisulfite treatment on the quality / sensitivity of the DNA-seq information extracted by the workflow.This method allows for the parallel application of complete DNA library amplification and processing to extract epigenetic DNA modifications using recovered, ligated parental DNA molecules (via labeled adapters). While this disclosure discusses, but is not limited to, the use of the BIS-seq method for identifying cytosine 5-methylation (5-methylcytosine), variations of BIS-seq have been developed to identify hydroxymethylated cytosine (5hmC; OX-BS-seq, TAB-seq), formylcytosine (5fC; redBS-seq), and carboxylcytosine. These methodologies can be implemented in conjunction with the sequential / parallel library preparation described herein.

[0079] Alternative methods for modified nucleic acid analysis This disclosure provides alternative methods for analyzing modified nucleic acids (e.g., methylation, histone linkage, and other modifications described above). Some such methods involve contacting a population of nucleic acids having different degrees of modification (e.g., 0, 1, 2, 3, 4, 5, or more methyl groups per nucleic acid molecule) with an adapter, and then fractionating the population according to the degree of modification. The adapter is attached to either one or both ends of the nucleic acid molecules in the population. Preferably, the adapter contains a sufficient number of different tags such that the probability of two nucleic acids having the same start and end points receiving the same tag combination is low (e.g., 95%, 99%, or 99.9%). After the adapter is attached, the nucleic acid is amplified from a primer that binds to a primer-binding site in the adapter. The adapter may contain the same or different primer-binding sites, whether they have the same or different tags, but preferably the adapter contains the same primer-binding site. After amplification, the nucleic acids are brought into contact with an activator that preferably binds to modified nucleic acids (e.g., such activators previously described). The nucleic acids are separated from binding to the activator into at least two sets of nucleic acids with different degrees of modification. For example, if the activator has affinity for modified nucleic acids, nucleic acids with a high degree of modification (compared to the median degree of modification in the population) will preferentially bind to the activator, while nucleic acids with a low degree of modification will not bind to the activator or will elute more easily from the activator. After separation, the different sets can then be subjected to further processing steps, which typically include further amplification and sequence analysis, performed in parallel but separately. The sequence data from the different sets can then be compared.

[0080] Both ends of a nucleic acid can be ligated to a Y-shaped adapter containing a primer-binding site and a tag. The molecule is amplified. The amplified molecule is then fractionated by contacting it with an antibody that preferentially binds to 5-methylcytosine to produce two fractions. One fraction contains the original molecule lacking methylation and the amplified copy with lost methylation. The other fraction contains the original DNA molecule with methylation. The two fractions are then processed and sequenced separately, with further amplification of the methylated fraction. The sequence data of the two fractions can then be compared. In this example, the tag is used not to distinguish between methylated and unmethylated DNA, but to distinguish between different molecules within these fractions, and thus it can be determined whether reads with the same start and end points are based on the same molecule or different molecules.

[0081] This disclosure provides further methods for analyzing a population of nucleic acids in which at least a portion of the nucleic acids contain one or more modified cytosine residues, e.g., 5-methylcytosine and any of the other modifications described above. These methods involve contacting the population of nucleic acids with an adapter containing one or more cytosine residues modified at the 5C position, e.g., 5-methylcytosine. Preferably, all cytosine residues of such an adapter are also modified, or all such cytosines within the primer-binding region of the adapter are modified. The adapter is attached to both ends of the nucleic acid molecules in the population. Preferably, the adapter contains a sufficient number of different tags such that the probability of two nucleic acids having the same start and end points receiving the same tag combination is low (e.g., 95%, 99%, or 99.9%), depending on the number of tag combinations. The primer-binding sites of such an adapter may be the same or different, but are preferably the same. After the adapter is attached, the nucleic acids are amplified from a primer that binds to the primer-binding site of the adapter. The amplified nucleic acids are split into a first aliquot and a second aliquot. The first aliquot is assayed for sequence data with or without further processing. Therefore, the sequence data for the molecule in the first aliquot is determined independently of the initial methylation state of the nucleic acid molecule. The nucleic acid molecule in the second aliquot is treated with bisulfite. This treatment converts unmodified cytosine to uracil. The bisulfite-treated nucleic acid is then subjected to amplification, primed with a primer targeting the original primer-binding site of the adapter ligated to the nucleic acid. At this point, only the nucleic acid molecule originally ligated to the adapter (separate from its amplified product) is amplified because these nucleic acids retain cytosine at the primer-binding site of the adapter, while the amplified product loses methylation of these cytosine residues that were converted to uracil by bisulfite treatment. Therefore, only the original molecule in the population, which is at least partially methylated, undergoes amplification. After amplification, these nucleic acids are subjected to sequence analysis.By comparing the sequences determined from the first and second aliquots, it may be possible to indicate, in particular, which cytosines within the nucleic acid population were subjected to methylation.

[0082] Distributing the sample into multiple subsamples; sample characteristics; analysis of epigenetic features. In certain embodiments described herein, a population of different forms of nucleic acids (e.g., highly methylated and hypomethylated DNA in a sample, e.g., a set of captured cfDNA described herein) can be physically distributed based on one or more characteristics of the nucleic acids before further analysis, e.g., differential modification or isolation, tagging, and / or sequencing of nucleic acid bases. This technique can be used, for example, to determine whether a particular sequence is highly methylated or hypomethylated. In some embodiments, highly methylated variable epigenetic target regions are analyzed to determine whether they exhibit the highly methylated characteristics of tumor cells, and / or hypomethylated variable epigenetic target regions are analyzed to determine whether they exhibit the hypomethylated characteristics of tumor cells. Furthermore, by distributing heterogeneous nucleic acid populations, rare signals can be increased, for example, by enriching rare nucleic acid molecules that are more dominant in one fraction (or section) of the population. For example, genetic variations that are present in highly methylated DNA but less so (or absent) in less methylated DNA can be more easily detected by distributing the sample into highly methylated and less methylated nucleic acid molecules. By analyzing multiple fractions of the sample, multidimensional analysis of a single locus or nucleic acid species in the genome can be performed, thus achieving higher sensitivity.

[0083] In some cases, heterogeneous nucleic acid samples are distributed into two or more sections (e.g., at least three, four, five, six, or seven sections). In some embodiments, each section is differentially tagged. The tagged sections can then be pooled together for collective sample preparation and / or sequencing. The distribution-tagging-pooling step may be performed more than once, in which case each round of distribution is based on different characteristics (examples are presented herein) and tagged using differential tags that distinguish them from other sections and distribution means.

[0084] Examples of features that can be used for partitioning include sequence length, methylation level, nucleosome binding, sequence mismatch, immunoprecipitation, and / or proteins that bind to DNA. The resulting partitions may include one or more of the following nucleic acid forms: single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), shorter DNA fragments, and longer DNA fragments. In some embodiments, partitioning based on cytosine modification (e.g., cytosine methylation) or methylation is commonly performed and, if necessary, combined with at least one additional partitioning step which may be based on any of the aforementioned features or DNA forms. In some embodiments, a heterogeneous population of nucleic acids is partitioned into nucleic acids having one or more epigenetic modifications and nucleic acids without one or more epigenetic modifications. Examples of epigenetic modifications include the presence or absence of methylation; the level of methylation; the type of methylation (e.g., 5-methylcytosine or other types of methylation, e.g., adenine methylation and / or cytosine hydroxymethylation); and association with one or more proteins, such as histones, and the level of association. Alternatively, or in addition to the above, a heterogeneous population of nucleic acids can be distributed between nucleic acid molecules associated with nucleosomes and nucleic acid molecules lacking nucleosomes. Alternatively, or in addition to the above, a heterogeneous population of nucleic acids can be distributed between single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA). Alternatively, or in addition to the above, a heterogeneous population of nucleic acids can be distributed based on the length of the nucleic acid (e.g., molecules up to 160 bp and molecules with a length greater than 160 bp).

[0085] In some cases, each segment (representing a different nucleic acid morphology) is differentially labeled before sequencing, and the segments are pooled together. In other cases, the different morphologies are sequenced separately. In some embodiments, a population of different nucleic acids is divided into two or more distinct segments. Each segment represents a different nucleic acid morphology, and the first segment (also called a subsample) contains a higher proportion of cytosine-modified DNA than the second subsample. Each segment is tagged separately. The first subsample is subjected to a procedure that affects the first nucleic acid base of the DNA in the first subsample differently from the second nucleic acid base of the DNA, where the first nucleic acid base is modified or unmodified, and the second nucleic acid base is a different modified or unmodified nucleic acid base from the first nucleic acid base, and the first and second nucleic acid bases have the same base-pairing specificity. The tagged nucleic acids are pooled together before sequencing. Sequence reads are obtained and analyzed in silico, including distinguishing between the first and second nucleic acid bases of the DNA in the first partial sample. Tags are used to sort reads from different segmentes. Analysis to detect genetic variants can be performed at the segmental level and at the whole nucleic acid population level. For example, the analysis may include in silico analysis to determine genetic variants in the nucleic acids within each segment, such as CNVs, SNVs, indels, and fusions. In some cases, in silico analysis may include determining chromatin structure. For example, the coverage of sequence reads can be used to determine nucleosome positioning in chromatin. Higher coverage may correlate with higher nucleosome occupancy in genomic regions, while lower coverage may correlate with lower nucleosome occupancy or nucleosome-depleted regions (NDRs).

[0086] The samples may contain nucleic acids with various modifications, including post-replication modifications to nucleotides, and typically non-covalent bonding to one or more proteins.

[0087] In one embodiment, the nucleic acid population is obtained from serum, plasma, or blood samples from subjects suspected of having neoplasia, tumor, or cancer, or who have been previously diagnosed with neoplasia, tumor, or cancer. The nucleic acid population includes nucleic acids with varying levels of methylation. Methylation can result from any one or more post-replication or post-transcriptional modifications. Post-replication modifications include modifications of nucleotide cytosines, particularly those at the 5-position of the nucleic acid base, e.g., 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, and 5-carboxylcytosine. The affinity agonist may be an antibody with desired specificity, its natural binding partner or variant (Bock et al., Nat Biotech 28: 1106-1114 (2010); Song et al., Nat Biotech 29: 68-72 (2011)), or an artificial peptide selected, for example, by phage display to have specificity for a given target.

[0088] Examples of capture portions intended herein include methyl-binding domains (MBDs) and methyl-binding proteins (MBPs) described herein, including proteins such as antibodies that preferentially bind to MeCP2 and 5-methylcytosine. Similarly, the distribution of different forms of nucleic acids can be carried out using histone-binding proteins that can separate histone-bound nucleic acids from free or unbound nucleic acids. Examples of histone-binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48, and SANT domain peptides. With respect to some affinity agonists and modifications, depending on whether the nucleic acid is modified, binding to the agonist may occur in an all-or-nothing manner, while separation may be degree-dependent. In such cases, nucleic acids with a high degree of modification will bind to the agonist to a greater degree than nucleic acids with a low degree of modification. Alternatively, modified nucleic acids may bind in an all-or-nothing manner. Nevertheless, various levels of modification can be sequentially eluted from the binding agonist.

[0089] For example, in some embodiments, partitioning may be binary or based on the degree / level of modification. For instance, all methylated fragments can be partitioned from unmethylated fragments using a methyl-binding domain protein (e.g., MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific)). Subsequent partitioning may involve eluting fragments with different levels of methylation by adjusting the salt concentration in the solution containing the methyl-binding domain and the bound fragments. As the salt concentration increases, fragments with higher levels of methylation are eluted. In some cases, the final partitioning may represent nucleic acids with different degrees of modification (high or low abundance of modification). High and low abundance can be defined by the number of modifications a nucleic acid has compared to the median number of modifications per strand in the population. For example, if the median number of 5-methylcytosine residues in nucleic acids in a sample is 2, then nucleic acids containing more than 2 5-methylcytosine residues have a high abundance of this modification, while nucleic acids with 1 or zero 5-methylcytosine residues have a low abundance. The effect of affinity separation is to enrich the bound phase with nucleic acids that have a high degree of modification, and enrich the unbound phase (i.e., in solution) with nucleic acids that have a low degree of modification. After eluting the nucleic acids in the bound phase, further processing can be performed.

[0090] When using the MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific), various levels of methylation can be partitioned using sequential elution. For example, a low-methylation fraction (e.g., no methylation) can be separated from a methylated fraction by contacting the nucleic acid population with MBD from the kit, which has been attached to magnetic beads. The beads are used to separate methylated nucleic acids from unmethylated nucleic acids. Then, one or more elution steps are performed sequentially to elute nucleic acids with different levels of methylation. For example, a first set of methylated nucleic acids can be eluted at a salt concentration of 160 mM or higher, e.g., at least 150 mM, at least 200 mM, at least 300 mM, at least 400 mM, at least 500 mM, at least 600 mM, at least 700 mM, at least 800 mM, at least 900 mM, at least 1000 mM, or at least 2000 mM. After eluting such methylated nucleic acids, magnetic separation is used again to separate the nucleic acids with higher levels of methylation from those with lower levels of methylation. The elution and magnetic separation steps themselves can be repeated to create various segments, such as a low-methylation segment (representing no methylation), a methylated segment (representing low levels of methylation), and a high-methylation segment (representing high levels of methylation).

[0091] In some methods, nucleic acids bound to the activator used for affinity separation are subjected to a washing step. The washing step washes away nucleic acids that are weakly bound to the affinity activator. Such nucleic acids may be enriched with nucleic acids having a degree of modification close to the mean or median (i.e., an intermediate value between nucleic acids that remained bound to the solid phase and nucleic acids that did not bind to the solid phase when the sample was first brought into contact with the activator). Affinity separation yields at least two, and sometimes three or more, partitions of nucleic acids with different degrees of modification. Although the partitions remain separated, nucleic acids from at least one, and usually two or three (or more), partitions are ligated to nucleic acid tags, usually provided as components of an adapter, so that nucleic acids in different partitions receive different tags that distinguish members of one partition from members of another. Tags ligated to nucleic acid molecules of the same partition may be the same or different. However, if they are different, the tags may share a common part of their code to identify the molecule to which they are attached as belonging to a particular partition. For further details regarding the portioning of nucleic acid samples based on characteristics such as methylation, see WO2018 / 119452, which is incorporated herein by reference. In some embodiments, nucleic acid molecules can be fractionated into different parts based on nucleic acid molecules bound to a particular protein or fragment thereof and nucleic acid molecules not bound to that particular protein or fragment thereof.

[0092] Nucleic acid molecules can be fractionated based on DNA-protein binding. Protein-DNA complexes can be fractionated based on specific properties of the protein. Examples of such properties include various epitopes, modifications (e.g., histone methylation or acetylation), or enzymatic activity. Examples of proteins that can bind to DNA and serve as a basis for fractionation include, but are not limited to, protein A and protein G. Nucleic acid molecules can be fractionated based on the region bound to the protein using any suitable method. Examples of methods used to fractionate nucleic acid molecules based on the region bound to the protein include, but are not limited to, SDS-PAGE, chromatin immunoprecipitation (ChIP), heparin chromatography, and asymmetric field flow fractionation (AF4).

[0093] In some embodiments, nucleic acid distribution is carried out by contacting the nucleic acid with the methylation-binding domain ("MBD") of a methylation-binding protein ("MBP"). The MBD binds to 5-methylcytosine (5mC). The MBD is coupled to paramagnetic beads such as Dynabeads® M-280 streptavidin via a biotin linker. Distribution to fractions with different degrees of methylation can be carried out by eluting the fractions by increasing the NaCl concentration.

[0094] An example method for identifying molecular tags in a library distributed with MBD beads using NGS is as follows:

[0095] Physical distribution of extracted DNA samples (e.g., plasma DNA extracted from human samples) using a methyl-binding domain protein-bead purification kit. All eluates from the process are stored for downstream processing.

[0096] Parallel application of differential molecular tags and adapter sequences enabling NGS for each segment. For example, ligating a highly methylated segment, a residually methylated ("washed") segment, and a lowly methylated segment with an NGS adapter having a molecular tag.

[0097] All molecularly tagged segments are recombined and then amplified using adapter-specific DNA primer sequences.

[0098] Enrichment / hybridization of the combined and amplified total library. Targeting the desired genomic region (e.g., cancer-specific genetic variants and differential methylation regions).

[0099] Re-amplification of the enriched total DNA library, and addition of sample tags. Different samples are pooled and assayed in multiplex in an NGS instrument.

[0100] Bioinformatics analysis of NGS data. Molecular tags are used to identify unique molecules and to deconvolve samples into differentially MBD-distributed molecules. This analysis allows for obtaining relative information about 5-methylcytosine in genomic regions, in parallel with standard gene sequencing / variant detection.

[0101] Examples of MBPs intended in this specification include, but are not limited to, the following:

[0102] (a) MeCP2 is a protein that preferentially binds to 5-methylcytosine rather than unmodified cytosine.

[0103] (b) RPL26, PRP8, and the DNA mismatch repair protein MHS6 preferentially bind to 5-hydroxymethylcytosine rather than unmodified cytosine.

[0104] (c)FOXK1, FOXK2, FOXP1, FOXP4, and FOXI3 preferentially bind to 5-formyl-cytosine rather than unmodified cytosine (Iurlaro et al., Genome Biol. 14: R119 (2013)).

[0105] (d) An antibody specific to one or more methylated nucleotide bases.

[0106] Generally, elution is dependent on the number of methylation sites per molecule, with higher salt concentrations eluting molecules that have more methylation. A series of elution buffers with increasing NaCl concentrations can be used to elute DNA into separate populations based on the degree of methylation. Salt concentrations can range from approximately 100 nM to approximately 2500 mM NaCl. In one embodiment, the process yields three divisions. Molecules are brought into contact with a solution of a first salt concentration containing molecules that have methyl-binding domains and can be attached to a capture moiety such as streptavidin. At the first salt concentration, some populations of molecules bind to the MBD, while others remain unbound. The unbound populations can be separated as a "low-methylated" population. For example, the first division, representing the low-methylated form of DNA, remains unbound at low salt concentrations, e.g., 100 mM or 160 mM. The second segment, representing intermediate methylated DNA, is eluted using an intermediate salt concentration, e.g., between 100 mM and 2000 mM. This segment is also separated from the sample. The third segment, representing the highly methylated form of DNA, is eluted using a high salt concentration, e.g., at least about 2000 mM.

[0107] This disclosure provides further methods for analyzing a population of nucleic acids in which at least a portion of the nucleic acids contains one or more modified cytosine residues, e.g., 5-methylcytosine and any of the other modifications described above. In these methods, after distribution, a portion of the nucleic acid sample is brought into contact with an adapter containing one or more cytosine residues modified at the 5C position, e.g., 5-methylcytosine. Preferably, all cytosine residues of such an adapter are also modified, or all such cytosines within the primer-binding region of the adapter are modified. The adapter is attached to both ends of the nucleic acid molecule in the population. Preferably, the adapter contains a sufficient number of different tags such that the probability of two nucleic acids having the same start and end points receiving the same tag combination is low (e.g., 95%, 99%, or 99.9%), depending on the number of tag combinations. The primer-binding sites of such an adapter may be the same or different, but are preferably the same. After the adapter is attached, the nucleic acid is amplified from a primer that binds to the primer-binding site of the adapter. The amplified nucleic acid is split into a first aliquot and a second aliquot. The first aliquot is assayed for sequence data with or without further processing. Thus, the sequence data for the molecule in the first aliquot is determined independently of the initial methylation state of the nucleic acid molecule. The nucleic acid molecule in the second aliquot is subjected to a procedure that affects the first nucleic acid base of the DNA differently than the second nucleic acid base of the DNA, where the first nucleic acid base contains cytosine modified at position 5 and the second nucleic acid base contains unmodified cytosine. This procedure may be bisulfite treatment or another procedure that converts unmodified cytosine to uracil. The nucleic acid subjected to the procedure is then amplified using a primer to the original primer-binding site of the adapter ligated to the nucleic acid. At this point, only the nucleic acid molecule originally ligated to the adapter (which is separate from its amplified product) is amplified because these nucleic acids retain cytosine at the primer-binding site of the adapter, while in the amplified product, the methylation of these cytosine residues that were converted to uracil by bisulfite treatment is lost.Therefore, only the original molecules in the population that are at least partially methylated undergo amplification. After amplification, these nucleic acids are subjected to sequence analysis. By comparing the sequences determined from the first and second aliquots, it may be possible to indicate, in particular, which cytosines in the nucleic acid population were subjected to methylation.

[0108] Such analysis can be performed using the following exemplary procedure: After distribution, both ends of the methylated DNA are ligated to a Y-shaped adapter containing a primer-binding site and a tag. The cytosine on the adapter is modified at position 5 (e.g., 5-methylation). The modification of the adapter serves to protect the primer-binding site during subsequent conversion steps (e.g., bisulfite treatment, TAP conversion, or any other conversion that does not affect the modified cytosine but does affect the unmodified cytosine). After the adapter is attached, the DNA molecule is amplified. The amplified product is split into two aliquots for sequencing with and without conversion. The aliquot not subjected to conversion can be subjected to sequencing analysis with or without further processing. The other aliquot is subjected to a procedure that affects the first nucleic acid base of the DNA differently than the second nucleic acid base of the DNA, where the first nucleic acid base contains cytosine modified at position 5 and the second nucleic acid base contains unmodified cytosine. This procedure may be bisulfite treatment or another procedure to convert unmodified cytosine to uracil. When contacted with a primer specific to the original primer-binding site, only the primer-binding site protected by cytosine modification can support amplification. Therefore, only the original molecule is subjected to further amplification, and the copy from the first amplification is not subjected to further amplification. The further amplified molecule is then subjected to sequence analysis. The sequences from the two aliquots can then be compared. Similar to the separation scheme described above, the nucleic acid tag on the adapter is used to distinguish nucleic acid molecules within the same segment, rather than to distinguish between methylated and unmethylated DNA.

[0109] A step of subjecting a first partial sample to a procedure that affects the first nucleic acid base of the DNA of the first partial sample in a way that is different from that of the second nucleic acid base of the DNA. A method disclosed herein is a step of subjecting a first partial sample to a procedure that affects a first nucleic acid base of the DNA of the first partial sample differently from a second nucleic acid base of the DNA, wherein the first nucleic acid base is a modified or unmodified nucleic acid base, the second nucleic acid base is a modified or unmodified nucleic acid base different from the first nucleic acid base, and the first and second nucleic acid bases have the same base-pairing specificity. In some embodiments, if the first nucleic acid base is a modified or unmodified adenine, then the second nucleic acid base is a modified or unmodified adenine; if the first nucleic acid base is a modified or unmodified cytosine, then the second nucleic acid base is a modified or unmodified cytosine; if the first nucleic acid base is a modified or unmodified guanine, then the second nucleic acid base is a modified or unmodified guanine; and if the first nucleic acid base is a modified or unmodified thymine, then the second nucleic acid base is a modified or unmodified thymine (for the purposes of this step, modified and unmodified uracil are included in modified thymine).

[0110] In some embodiments, the first nucleic acid base is modified or unmodified cytosine, and the second nucleic acid base is modified or unmodified cytosine. For example, the first nucleic acid base may contain unmodified cytosine (C), and the second nucleic acid base may contain one or more of 5-methylcytosine (mC) and 5-hydroxymethylcytosine (hmC). Alternatively, the second nucleic acid base may contain C, and the first nucleic acid base may contain one or more of mC and hmC. Other combinations are also possible, for example, as shown in the above summary and the following discussion, such as when one of the first and second nucleic acid bases contains mC and the other contains hmC.

[0111] In some embodiments, a procedure that affects the first nucleic acid base of the DNA in a first partial sample differently from that affecting the second nucleic acid base of the DNA involves bisulfite conversion. Treatment with bisulfite converts unmodified cytosine and certain modified cytosine nucleotides (e.g., 5-formylcytosine (fC) or 5-carboxylcytosine (caC)) to uracil, while other modified cytosines (e.g., 5-methylcytosine, 5-hydroxymethylcytosine) are not converted. Therefore, when using bisulfite conversion, the first nucleic acid base includes one or more of unmodified cytosine, 5-formylcytosine, 5-carboxylcytosine, or other cytosine forms affected by bisulfite, and the second nucleic acid base may include one or more of mC and hmC, e.g., mC and optionally hmC. Sequencing of the bisulfite-treated DNA identifies the positions read as cytosine as mC or hmC positions. On the other hand, positions read as T are identified as T or bisulfite-sensitive forms of C, such as unmodified cytosine, 5-formylcytosine, or 5-carboxylcytosine. Therefore, by performing bisulfite conversion on the first partial sample described herein, it becomes easier to identify mC or hmC-containing positions using sequence reads obtained from the first partial sample. For an illustrative description of bisulfite conversion, see, for example, Moss et al., Nat Commun. 2018; 9: 5068.

[0112] In some embodiments, a procedure that affects the first nucleic acid base of the DNA in the first sample differently from the second nucleic acid base of the DNA includes oxidative bisulfite (Ox-BS) conversion. In some embodiments, a procedure that affects the first nucleic acid base of the DNA in the first sample differently from the second nucleic acid base of the DNA includes Tet-assisted bisulfite (TAB) conversion. In some embodiments, a procedure that affects the first nucleic acid base of the DNA in the first sample differently from the second nucleic acid base of the DNA includes Tet-assisted conversion using a substituted borane reducing agent, which may optionally be 2-picoline borane, borampyridine, tert-butylamine borane, or ammonia borane. In some embodiments, a procedure that affects the first nucleic acid base of the DNA in the first sample differently from the second nucleic acid base of the DNA includes chemical-assisted conversion using a substituted borane reducing agent, which may optionally be 2-picoline borane, borampyridine, tert-butylamine borane, or ammonia borane. In some embodiments, the procedure that affects the first nucleic acid base of the DNA in the first partial sample differently from the second nucleic acid base of the DNA includes APOBEC coupling epigenetic (ACE) conversion.

[0113] In some embodiments, the procedure for affecting the first nucleic acid base of the DNA in a first partial sample differently from the second nucleic acid base of the DNA includes, for example, enzymatic conversion of the first nucleic acid base, similar to EM-Seq. See, for example, Vaisvila R, et al. (2019) EM-seq: Detection of DNA methylation at single base resolution from picograms of DNA. bioRxiv; DOI: 10.1101 / 2019.12.20.884692v1, available at www.biorxiv.org / content / 10.1101 / 2019.12.20.884692v1. For example, 5mC and 5hmC can be converted to substrates that cannot be deaminated by deaminase (e.g., APOBEC3A) using TET2 and T4-βGT, and then the unmodified cytosines can be deaminated using deaminase (e.g., APOBEC3A) to convert them to uracil.

[0114] In some embodiments, the procedure for affecting the first nucleic acid base of the DNA of a first partial sample differently from that affecting the second nucleic acid base of the DNA includes separating the DNA that originally contains the first nucleic acid base from the DNA that originally does not contain the first nucleic acid base.

[0115] In some embodiments, the first nucleic acid base is a modified or unmodified adenine, and the second nucleic acid base is a modified or unmodified adenine. In some embodiments, the modified adenine is N6-methyladenine (mA). In some embodiments, the modified adenine is one or more of N6-methyladenine (mA), N6-hydroxymethyladenine (hmA), or N6-formyladenine (fA).

[0116] Techniques including methylated DNA immunoprecipitation (MeDIP) can be used to isolate DNA containing modified bases such as mA from other DNA. See, for example, Kumar et al., Frontiers Genet. 2018; 9: 640; Greer et al., Cell 2015; 161: 868-878. Antibodies specific to mA are described in Sun et al., Bioessays 2015; 37: 1155-62. Antibodies against various modified nucleic acid bases, such as thymine / uracil forms including halogenated forms like 5-bromouracil, are commercially available. Various modified bases can also be detected based on changes in their base pairing specificity. For example, hypoxanthine is a modified form of adenine that can result from deamination and is read as G in sequencing. For example, see U.S. Patent 8,486,630; Brown, Genomes, 2nd Ed., John Wiley & Sons, Inc., New York, NY, 2002, chapter 14, "Mutation, Repair, and Recombination."

[0117] Enrichment / capture steps, amplification, adapter, barcode In some embodiments, the methods disclosed herein include the step of capturing one or more sets of target regions of DNA, such as cfDNA. Capture can be carried out using any suitable method known in the art. In some embodiments, the capture step includes contacting the DNA to be captured with a set of target-specific probes. The set of target-specific probes may have any of the features described herein with respect to a set of target-specific probes, including, but not limited to, those in the embodiments described above and in the section on probes below. The capture step can be carried out on one or more partial samples prepared during the methods disclosed herein. In some embodiments, DNA is captured from at least a first or second partial sample, for example, at least a first and a second partial sample. If the first partial sample undergoes a separation step (for example, separating DNA that originally contains a first nucleic acid base (e.g., hmC) from DNA that originally does not contain a first nucleic acid base (e.g., hmC-seal)), the capture step can be performed on any, any two, or all of the DNA that originally contains a first nucleic acid base (e.g., hmC), the DNA that originally does not contain a first nucleic acid base, and the second partial sample. In some embodiments, the partial samples are differentially tagged (e.g., as described herein), then pooled, and then captured.

[0118] The capture step can be carried out using conditions suitable for a particular nucleic acid hybridization, and these conditions generally depend to some extent on the characteristics of the probe, such as its length and base composition. Those skilled in the art will be able to determine the appropriate conditions based on general knowledge in the art regarding nucleic acid hybridization. In some embodiments, a complex is formed between the target-specific probe and DNA.

[0119] In some embodiments, the methods described herein include a step of capturing cfDNA obtained from a test subject for multiple sets of target regions. The target regions include epigenetic target regions, which may exhibit differences in methylation levels and / or fragmentation patterns depending on whether they originate from tumors or healthy cells. The target regions also include sequence-variable target regions, which may exhibit differences in sequence depending on whether they originate from tumors or healthy cells. The capture step generates a set of captured cfDNA molecules, where cfDNA molecules corresponding to the sequence-variable target region set are captured in a higher capture yield than cfDNA molecules corresponding to the epigenetic target region set in the set of captured cfDNA molecules. For further consideration of the capture step, capture yields, and related embodiments, see WO2020 / 160414, incorporated herein by reference for all purposes.

[0120] In some embodiments, the method described herein includes the step of contacting cfDNA obtained from a test subject with a set of target-specific probes, wherein the set of target-specific probes is configured to capture cfDNA corresponding to a sequence-variable target region set with a higher capture yield than cfDNA corresponding to an epigenetic target region set.

[0121] To analyze sequence-variable target regions with sufficient confidence or accuracy, higher sequencing depths may be required than those necessary for analyzing epigenetic target regions. Therefore, capturing cfDNA corresponding to a set of sequence-variable target regions at a higher capture yield than cfDNA corresponding to a set of epigenetic target regions may be beneficial. The amount of data required to determine fragmentation patterns (e.g., for examining perturbations at transcription start sites or CTCF-binding sites) or fragment abundances (e.g., in highly methylated and hypomethylated segments) is generally less than the amount of data required to determine the presence or absence of cancer-associated sequence mutations. By capturing target region sets at different yields, it may be easier to sequence target regions to different sequencing depths in the same sequencing run (e.g., using a pooled mixture and / or in the same sequencing cell).

[0122] In various embodiments, the method further includes the step of sequencing the captured cfDNA to varying degrees of sequencing depth with respect to, for example, an epigenetic target region set and a sequence variable target region set, consistent with the discussion herein. In some embodiments, the target-specific probe-DNA complex is separated from DNA not bound to the target-specific probe. For example, if the target-specific probe is covalently or noncovalently bound to a solid support, washing or aspiration steps can be used to separate the unbound material. Alternatively, if the complex has chromatographic properties distinct from the unbound material (for example, if the probe contains a ligand that binds to a chromatography resin), chromatography can be used.

[0123] As discussed in detail elsewhere in this specification, a set of target-specific probes may include multiple sets, such as probes for sequence-variable target regions and probes for epigenetic target regions. In some such embodiments, the capture step is performed simultaneously in the same container using probes for sequence-variable target regions and probes for epigenetic target regions, for example, the probes for sequence-variable target regions and probes for epigenetic target regions are in the same composition. This method provides a relatively streamlined workflow. In some embodiments, the concentration of probes for sequence-variable target regions is higher than the concentration of probes for epigenetic target regions.

[0124] Alternatively, the capture step may be performed using the sequence-variable target region probe set in a first container and the epigenetic target region probe set in a second container, or the contact step may be performed using the sequence-variable target region probe set at a first time point and in the first container, and the epigenetic target region probe set at a second time point before or after the first time point. This method allows for the preparation of separate first and second compositions, each containing captured DNA corresponding to the sequence-variable target region set and captured DNA corresponding to the epigenetic target region set. The compositions can be processed separately as desired (e.g., to fractionate based on methylation as described elsewhere in this specification) and then recombined in proportions suitable for further processing and analysis, such as sequencing.

[0125] In some embodiments, DNA is amplified. In some embodiments, amplification is performed before the capture step. In some embodiments, amplification is performed after the capture step.

[0126] In some embodiments, the adapter is included in the DNA. This can be done in parallel with the amplification procedure, for example, by attaching the adapter to the 5' portion of the primer as described above. Alternatively, the adapter can be added by other methods such as ligation.

[0127] In some embodiments, the DNA is accompanied by a tag, which may include a barcode. The tag can facilitate the identification of the nucleic acid's origin. For example, a barcode can be used to identify the origin (e.g., target) from which the DNA originates after pooling multiple samples for parallel sequencing. This can be done in parallel with the amplification procedure, for example, by affixing a barcode to the 5' portion of the primer, as described above. In some embodiments, the adapter and tag / barcode are provided by the same primer or primer set. For example, the barcode may be located on the 3' side of the adapter and on the 5' side of the portion that hybridizes with the primer's target. Alternatively, the barcode may be added together with the adapter in the same ligation substrate as needed, by other methods such as ligation.

[0128] Further details regarding amplification, tagging, and barcodes are discussed in the following section, “General Features of the Method,” and these can be combined to an operational degree with any of the embodiments described above, as well as those described in the Introduction and Summary sections.

[0129] Captured set In some embodiments, a set of captured DNA (e.g., cfDNA) is provided. With respect to the disclosed method, the set of captured DNA can be obtained, for example, by performing a capture step after a distribution step, as described herein. The captured set may include DNA corresponding to a set of sequence variable target regions, a set of epigenetic target regions, or a combination thereof. In some embodiments, when normalized for differences in the size (footprint size) of the targeting regions, the quantity of captured sequence variable target region DNA is greater than the quantity of captured epigenetic target region DNA.

[0130] Alternatively, this can result in a first captured set and a second captured set, each containing DNA corresponding to a sequence-variable target region set and DNA corresponding to an epigenetic target region set, respectively. The first and second captured sets can be combined to result in a combined captured set.

[0131] In some embodiments, including the combined captured sets discussed above, the captured sets containing DNA corresponding to the sequence variable target region set and the epigenetic target region set are present at higher concentrations than the DNA corresponding to the epigenetic target region set, e.g., 1.1 to 1.2 times, 1.2 to 1.4 times, or 1.4 to 1.6 times. Concentration of 0, 1.6 to 1.8 times, 1.8 to 2.0 times, 2.0 to 2.2 times, 2.2 to 2.4 times, 2.4 to 2.6 times, 2.6 to 2.8 times, 2.8 to 3.0 times, 3.0 to 3.5 times, 3.5 to 4.0 times, 4.0 to 4.5 times, 4.5 to 5.0 times, 5.0 to 5.5 times, 5.5 to 6.0 times, 6.0 to 6.5 2x concentration, 6.5x to 7.0x concentration, 7.0x to 7.5x concentration, 7.5x to 8.0x concentration, 8.0x to 8.5x concentration, 8.5x to 9.0x concentration, 9.0x to 9.5x concentration, 9.5x to 10.0x concentration, 10x to 11x concentration, 11x to 12x concentration, 12x to 13x concentration, 13x to 14x concentration, 14x to 15x concentration, 15x to 16x concentration, 16x to 17x concentration, 17x to 18x The concentrations can be twice as high, 18 to 19 times higher, 19 to 20 times higher, 20 to 30 times higher, 30 to 40 times higher, 40 to 50 times higher, 50 to 60 times higher, 60 to 70 times higher, 70 to 80 times higher, 80 to 90 times higher, 90 to 100 times higher, 10 to 20 times higher, 10 to 40 times higher, 10 to 50 times higher, 10 to 70 times higher, or 10 to 100 times higher. The degree of concentration difference is the main cause of normalization with respect to the footprint size of the target region, as discussed in the definition section.

[0132] Epigenetic target region set An epigenetic target region set may include one or more types of target regions that can discriminate DNA from neoplastic (e.g., tumor or cancer) cells from DNA from healthy cells, e.g., non-neoplastic circulating cells. Exemplary types of such regions are discussed in detail herein. An epigenetic target region set may also include, for example, one or more control regions described herein. In some embodiments, an epigenetic target region set has a footprint of at least 100 kb, e.g., at least 200 kb, at least 300 kb, or at least 400 kb. In some embodiments, the epigenetic target region set has footprints ranging from 100 to 1000 kb, for example, 100 to 200 kb, 200 to 300 kb, 300 to 400 kb, 400 to 500 kb, 500 to 600 kb, 600 to 700 kb, 700 to 800 kb, 800 to 900 kb, and 900 to 1,000 kb.

[0133] Highly methylated variable target region In some embodiments, the epigenetic target region set includes one or more hypermethylated variable target regions. Generally, a hypermethylated variable target region refers to a region, for example in a cfDNA sample, where the observed elevated level of methylation indicates an increased likelihood that the sample (e.g., a cfDNA sample) contains DNA produced by neoplastic cells such as tumor or cancer cells. For example, hypermethylation of tumor suppressor gene promoters has been repeatedly observed. See, for example, Kang et al., Genome Biol. 18:53 (2017) and the references cited therein. In one example, a hypermethylated variable target region may include a region in cancerous tissue that does not necessarily have different methylation compared to DNA from the same type of healthy tissue, but has different methylation (e.g., more methylation) compared to cfDNA that is typical in healthy subjects. For example, if the presence of cancer leads to increased cell death, e.g., apoptosis, of cells of the tissue type corresponding to cancer, such cancer can be detected, at least in part, using such hypermethylated variable target regions. In some embodiments, the hypermethylated variable target regions include one or more genomic regions in which cfDNA molecules in those regions have no different methylation status in cancer subjects compared to cfDNA from healthy subjects, but the increased presence / quantity of hypermethylated cfDNA in those regions serves as an indicator of a specific tissue type (e.g., cancer origin) and is represented as cfDNA with increased apoptosis (e.g., tumor efflux) into circulation.

[0134] Highly methylated target regions can be obtained, for example, from the Cancer Genome Atlas. Kang et al., Genome Biology 18:53 (2017) describe the construction of a stochastic method called CancerLocator using highly methylated target regions derived from breast, colon, kidney, liver, and lung. In some embodiments, highly methylated target regions may be specific to one or more types of cancer. Thus, in some embodiments, the highly methylated target regions include one, two, three, four, or five subsets of highly methylated target regions that collectively represent highly methylated in one, two, three, four, or five of the following: breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer.

[0135] In some embodiments, the probes for the epigenetic target region set include probes specific to one or more highly methylated variable target regions. The highly methylated variable target regions may be any of the above. For example, in some embodiments, the probes specific to highly methylated variable target regions include probes specific to a plurality of loci listed in Table 1, e.g., probes specific to at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1. In some embodiments, the probes specific to highly methylated variable target regions include probes specific to a plurality of loci listed in Table 2, e.g., probes specific to at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 2. In some embodiments, probes specific to highly methylated variable target regions include probes specific to a plurality of loci listed in Table 1 or Table 2, for example, probes specific to at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1 or Table 2. In some embodiments, for each locus included as a target region, there may be one or more probes having a hybridization site that binds between the transcription start site and the stop codon (the last stop codon of a gene that undergoes alternative splicing). In some embodiments, one or more probes bind within 300 bp, for example, within 200 bp or 100 bp, of the listed locations. In some embodiments, the probes have hybridization sites that overlap with the locations listed above. In some embodiments, the hypermethylation target region-specific probes include probes specific to one, two, three, four, or five subsets of hypermethylation target regions that collectively represent hypermethylation in one, two, three, four, or five of the following cancers: breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer.

[0136] Low-methylation variable target region Overall hypomethylation is a phenomenon commonly observed in various cancers. See, for example, Hon et al., Genome Res. 22:246-258 (2012) (breast cancer); Ehrlich, Epigenomics 1:239-259 (2009) (a review article mentioning the observation of hypomethylation in colon cancer, ovarian cancer, prostate cancer, leukemia, hepatocellular carcinoma, and cervical cancer). For example, regions such as repeat elements, e.g., LINE1 elements, Alu elements, centromere tandem repeats, periconomere tandem repeats, and satellite DNA, as well as intergenetic regions that are normally methylated in healthy cells, may show reduced methylation in tumor cells. Therefore, in some embodiments, a set of epigenetic target regions includes hypomethylated variable target regions in which the observed decrease in methylation levels indicates an increased likelihood that the sample (e.g., a cfDNA sample) contains DNA produced by neoplastic cells such as tumor or cancer cells. For example, a hypomethylated variable target region may include regions in cancerous tissue that are not necessarily methylated differently from DNA from healthy tissue of the same type, but are methylated differently (e.g., less methylated) than cfDNA that is typical in healthy subjects. For example, if the presence of cancer leads to increased cell death, e.g., apoptosis, of cells of the tissue type corresponding to cancer, such cancer can be detected, at least in part, using such a hypomethylated variable target region. In some embodiments, a hypomethylated variable target region includes one or more genomic regions in which cfDNA molecules in those regions are not methylated differently from cfDNA from healthy subjects in cancer subjects, but the increased presence / quantity of hypomethylated cfDNA in those regions serves as an indicator of a particular tissue type (e.g., cancer origin) and is indicated as cfDNA with increased apoptosis (e.g., tumor efflux) into circulation.

[0137] In some embodiments, the low-methylation variable target region includes repeat elements and / or intergenetic regions. In some embodiments, the repeat elements include one, two, three, four, or five of the following: LINE1 elements, Alu elements, centromere tandem repeats, peristromere tandem repeats, and / or satellite DNA.

[0138] Exemplary specific genomic regions exhibiting cancer-related hypomethylation include nucleotides 8403565–8953708 and 151104701–151106035 on human chromosome 1. In some embodiments, the hypomethylation variable target region overlaps with or includes one or both of these regions.

[0139] In some embodiments, a probe for a set of epigenetic target regions includes a probe specific to one or more hypomethylated variable target regions. The hypomethylated variable target regions may be any of the above. For example, a probe specific to one or more hypomethylated variable target regions may include probes for repeating elements, such as LINE1 elements, Alu elements, centromere tandem repeats, peristromereal tandem repeats, and satellite DNA, where intergenetic regions that are normally methylated in healthy cells may show reduced methylation in tumor cells.

[0140] In some embodiments, probes specific to low-methylation variable target regions include probes specific to repeat elements and / or intergenetic regions. In some embodiments, probes specific to repeat elements include probes specific to one, two, three, four, or five of the following: LINE1 elements, Alu elements, centromere tandem repeats, pericentromere tandem repeats, and / or satellite DNA.

[0141] Exemplary probes specific to genomic regions exhibiting cancer-related hypomethylation include probes specific to human chromosome 1 nucleotides 8403565-8953708 and / or 151104701-151106035. In some embodiments, probes specific to hypomethylation variable target regions include probes specific to regions that overlap with or contain human chromosome nucleotides 8403565-8953708 and / or 151104701-151106035.

[0142] Probes for detecting a panel of regions include probes for detecting target genomic regions (hotspot regions), as well as nucleosome recognition probes (e.g., KRAS codons 12 and 13), which can be designed to optimize capture based on analysis of cfDNA coverage and fragment size variations influenced by nucleosome binding patterns and GC sequence composition. Regions used herein may also include non-hotspot regions optimized based on nucleosome location and GC model.

[0143] Gene-specific probes include SKI, THEMIS2, RPA2, TEKT2, STK40, GJA9-MYCBP, LOC105378663, HEYL, CNN3, JTB, FAM78B, ARV1, ADSS2, ZNF672, MBOAT2, ASXL2, SERTAD2, TMEM131, CLASP1, SATB2, ABHD14B, NISCH, TMEM45A, LGI2, KLHL5, NEUROG2-AS1, ABHD18, MFSD8, ELF2, TRIM2, AHRR, PDCD6-AHRR, SEMA5A, IQGAP2, TSLP, SLC25A48, RELL2, ARHGAP26, SLC36A1, CNPY3, FAM229B, MA N1A1, ADCYAP1R1, KIAA0895, TRAPPC14, LINC01004, FAM131B, GIMAP4, SLC4A2, CD274, TOX, GDAP1, ZNF623, GNA14, S1PR3, C 9orf47, ROR2, ERCC6L2, LINC00476, ECPAS, ASTN2, PHF19, PTGES2-AS1, RALGDS, HACD1, ABLIM1, LOC101927692, GFRA1, C1 1orf21, TRIM44, CHST1, TMX2-CTNND1, LOC101928069, PDE2A, DLG2, ENDOD1, DDX6, TULP3, PTPRO, ZCRB1, TMPO-AS1, HSP90B 1. SIRT4, SRSF9, SLITRK1, MMP14, BCL2L2-PABPN1, KCNH5, TRAF3, IDH2, CIB1, MAN2A2, KDM8, ZFHX3, HSBP1, TOP3A, RETREG3, ADAM11, KPNB1, GRIN2C, GALR2, ZBTB14, EPB41L3, PDE4A, KLF1, SIX5, DM1-AS, ZNF114, CLEC11A, and LINC01530 can also be listed.

[0144] In some embodiments, DNA (e.g., cfDNA) is obtained from subjects having cancer. In some embodiments, DNA (e.g., cfDNA) is obtained from subjects suspected of having cancer. In some embodiments, DNA (e.g., cfDNA) is obtained from subjects having a tumor. In some embodiments, DNA (e.g., cfDNA) is obtained from subjects suspected of having a tumor. In some embodiments, DNA (e.g., cfDNA) is obtained from subjects having a neoplasia. In some embodiments, DNA (e.g., cfDNA) is obtained from subjects in remission of tumor, cancer, or neoplasia (e.g., after chemotherapy, surgical resection, radiation therapy, or a combination thereof). In any of the embodiments described above, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia may be of the lung, colon, rectum, kidney, breast, prostate, or liver. In some embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia may be of the lung. In some embodiments, the cancer, tumor, neoplasia, or suspected cancer, tumor, or neoplasia is of the colon or rectum. In some embodiments, the cancer, tumor, neoplasia, or suspected cancer, tumor, or neoplasia is of the breast. In some embodiments, the cancer, tumor, neoplasia, or suspected cancer, tumor, or neoplasia is of the prostate. In any of the embodiments described above, the subject may be a human subject.

[0145] In some embodiments, the sequence-variable target region probe set has a footprint of at least 0.5kb, for example, at least 1kb, at least 2kb, at least 5kb, at least 10kb, at least 20kb, at least 30kb, or at least 40kb. In some embodiments, the epigenetic target region probe set has a footprint ranging from 0.5 to 100kb, for example, in the range of 0.5 to 2kb, 2 to 10kb, 10 to 20kb, 20 to 30kb, 30 to 40kb, 40 to 50kb, 50 to 60kb, 60 to 70kb, 70 to 80kb, 80 to 90kb, and 90 to 100kb.

[0146] In some embodiments, probes specific to a sequence-variable target region set include at least 10, 20, 30, or 35 cancer-related genes, such as SKI, THEMIS2, RPA2, TEKT2, STK40, GJA9-MYCBP, LOC105378663, HEYL, CNN3, JTB, FAM78B, ARV1, ADSS2, ZNF672, MBOAT2, ASXL2, SERTAD2, TMEM131, CLASP1, SATB2, ABHD14B, NISCH, T MEM45A, LGI2, KLHL5, NEUROG2-AS1, ABHD18, MFSD8, ELF2, TRIM2, AHRR, PDCD6-AHRR, SEMA5A, IQGAP2, TSLP, SLC25A48, RELL2, ARHG AP26, SLC36A1, CNPY3, FAM229B, MAN1A1, ADCYAP1R1, KIAA0895, TRAPPC14, LINC01004, FAM131B, GIMAP4, SLC4A2, CD274, TOX, GDAP1 , ZNF623, GNA14, S1PR3, C9orf47, ROR2, ERCC6L2, LINC00476, ECPAS, ASTN2, PHF19, PTGES2-AS1, RALGDS, HACD1, ABLIM1, LOC10192 7692, GFRA1, C11orf21, TRIM44, CHST1, TMX2-CTNND1, LOC101928069, PDE2A, DLG2, ENDOD1, DDX6, TULP3, PTPRO, ZCRB1, TMPO-AS1, H Includes target region-specific probes derived from SP90B1, SIRT4, SRSF9, SLITRK1, MMP14, BCL2L2-PABPN1, KCNH5, TRAF3, IDH2, CIB1, MAN2A2, KDM8, ZFHX3, HSBP1, TOP3A, RETREG3, ADAM11, KPNB1, GRIN2C, GALR2, ZBTB14, EPB41L3, PDE4A, KLF1, SIX5, DM1-AS, ZNF114, CLEC11A, and LINC01530.

[0147] composition containing captured DNA A combination comprising a first population and a second population of captured DNA is provided herein. The first population may contain, or may be derived from, DNA having a higher proportion of cytosine modifications than the second population. The first population may include a first nucleic acid base form in which the base-pairing specificity originally present in the DNA has been altered, and a second nucleic acid base in which the base-pairing specificity has not been altered, where the first nucleic acid base form before alteration of base-pairing specificity originally present in the DNA is a modified or unmodified nucleic acid base, and the second nucleic acid base is a modified or unmodified nucleic acid base different from the first nucleic acid base, and the first nucleic acid base form before alteration of base-pairing specificity originally present in the DNA and the second nucleic acid base have the same base-pairing specificity. The second population does not include a first nucleic acid base form in which the base-pairing specificity originally present in the DNA has been altered. In some embodiments, the cytosine modification is cytosine methylation. In some embodiments, the first nucleic acid base is a modified or unmodified cytosine, and the second nucleic acid base is a modified or unmodified cytosine. The first and second nucleic acid bases may be any of those discussed herein, in summary, or in relation to the step of subjecting a first partial sample to a procedure that affects the first nucleic acid base of the DNA of the first partial sample differently from the second nucleic acid base of the DNA.

[0148] In some embodiments, the first group includes array tags selected from a first set of one or more array tags, and the second group includes array tags selected from a second set of one or more array tags, wherein the second set of array tags is different from the first set of array tags. The array tags may include barcodes.

[0149] In some embodiments, the first population includes protected hmC, e.g., glucosylated hmC. In some embodiments, the first population has been subjected to one of the conversion procedures discussed herein, e.g., bisulfite conversion, Ox-BS conversion, TAB conversion, ACE conversion, TAP conversion, TAPSβ conversion, or CAP conversion. In some embodiments, the first population has been subjected to protection of hmC, followed by deamination of mC and / or C. In some embodiments of the combination, the first population includes or is derived from DNA having a higher proportion of cytosine modifications than the second population, and the first population includes a first subpopulation and a second subpopulation, where the first nucleic acid bases are modified or unmodified nucleic acid bases, the second nucleic acid bases are modified or unmodified nucleic acid bases different from the first nucleic acid bases, and the first and second nucleic acid bases have the same base pairing specificity. In some embodiments, the second population does not include the first nucleic acid bases. In some embodiments, the first nucleic acid base is modified or unmodified cytosine, and the second nucleic acid base is modified or unmodified cytosine, and the modified cytosine is optionally mC or hmC. In some embodiments, the first nucleic acid base is modified or unmodified adenine, and the second nucleic acid base is modified or unmodified adenine, and the modified adenine is optionally mA.

[0150] In some embodiments, the first nucleic acid base (e.g., modified cytosine) is biotinylated. In some embodiments, the first nucleic acid base (e.g., modified cytosine) is the product of hysgen cycloaddition to β-6-azido-glucosyl-5-hydroxymethylcytosine containing an affinity label (e.g., biotin).

[0151] In any of the combinations described herein, the captured DNA may include cfDNA. The captured DNA may have any of the features described herein with respect to the captured set, including, for example, a higher concentration of DNA corresponding to a sequence variable target region set (normalized with respect to footprint size as discussed above) than the concentration of DNA corresponding to an epigenetic target region set. In some embodiments, the DNA of the captured set includes a sequence tag, which can be attached to the DNA as described herein. Generally, including a sequence tag results in a DNA molecule that differs from the naturally occurring, untagged form.

[0152] The combination may further include the probe sets or sequencing primers described herein, each of which may differ from naturally occurring nucleic acid molecules. For example, the probe sets described herein may include a capture moiety, and the sequencing primers may include labels that do not exist in nature.

[0153] Computer systems, processing of real-world evidence (RWE) The methods of the present disclosure can be implemented using a computer system or with the assistance of a computer system. For example, such a method may include the steps of: distributing a sample into a plurality of subsamples, including a first subsample and a second subsample, wherein the first subsample contains a higher proportion of cytosine-modified DNA than the second subsample; subjecting the first subsample to a procedure that affects the first nucleic acid base of the DNA of the first subsample differently from the second nucleic acid base of the DNA, wherein the first nucleic acid base is a modified or unmodified nucleic acid base, the second nucleic acid base is a modified or unmodified nucleic acid base different from the first nucleic acid base, and the first and second nucleic acid bases have the same base-pairing specificity; and sequencing the DNA in the first subsample and the DNA in the second subsample in such a manner that the first and second nucleic acid bases of the DNA of the first subsample are distinguishable.

[0154] In one embodiment, the Disclosure provides a non-temporary computer-readable medium containing computer-executable instructions that, when executed by at least one electronic processor, carries out at least part of a method comprising: collecting cfDNA from a test subject; capturing a plurality of sets of target regions from the cfDNA, wherein the plurality of target region sets include a sequence-variable target region set and an epigenetic target region set, thereby creating a set of captured cfDNA molecules; sequencing the captured cfDNA molecules, wherein the captured cfDNA molecules of the sequence-variable target region set are sequenced to a higher sequencing depth than the captured cfDNA molecules of the epigenetic target region set; obtaining a plurality of sequence reads generated by sequencing the captured cfDNA molecules using a nucleic acid sequencer; mapping the plurality of sequence reads to one or more reference sequences to generate mapped sequence reads; and processing the mapped sequence reads corresponding to the sequence-variable target region set and the mapped sequence reads corresponding to the epigenetic target region set to determine the likelihood that the subject has cancer.

[0155] The code may be pre-compiled and configured for use in a machine having a processor adapted to run the code, or it may be compiled during execution time. The code may be supplied in a programming language that can be selected to allow the code to be executed in a pre-compiled manner or in an as-compiled manner.

[0156] Further details regarding computer systems and networks, databases, and computer program products are also provided, for example, in Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Ed. (2011), Kurose, Computer Networking: A Top-Down Approach, Pearson, 7th Ed. (2016), each of which is thus incorporated herein by reference in its entirety, in Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Ed. (2010), in Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11th Ed. (2014), in Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd Ed. (2006), and in Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011).

[0157] Treatment and related administrations In certain embodiments, the methods disclosed herein relate to identifying and administering customized treatments to a patient, taking into account the status of whether a nucleic acid variant is of somatic or germline origin. In some embodiments, essentially any cancer treatment (e.g., surgery, radiotherapy, chemotherapy, and / or similar) can be included as part of these methods. Typically, a customized treatment comprises at least one immunotherapy (or immunotherapy agent). Immunotherapy generally refers to a method of enhancing the immune response to a given type of cancer. In certain embodiments, immunotherapy refers to a method of enhancing the T-cell response to a tumor or cancer.

[0158] In certain embodiments, the status of nucleic acid variants derived from a sample of a subject—whether somatic or germline origin—can be compared to a database of comparative results obtained from a reference population to identify customized or targeted therapies for that subject. Typically, the reference population includes patients with the same type of cancer or disease as the subject of study, and / or patients who are receiving or have received the same treatment as the subject of study. If the nucleic acid variants and comparative results meet certain classification criteria (e.g., substantially or substantially identical), then a customized or targeted therapy(s) may be identified.

[0159] In certain embodiments, the customized treatments described herein are typically administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions containing immunotherapeutic agents are typically administered intravenously. Certain therapeutic agents are administered orally. However, customized treatments (e.g., immunotherapeutic agents) may also be administered by means such as buccal, sublingual, rectal, vaginal, urethral, ​​topical, intraocular, intranasal, and / or intraauricular, and the administration may include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, salves, ointments, etc.

[0160] kit Kits containing the compositions described herein are also provided. The kits may be useful for carrying out the methods described herein. In some embodiments, the kit includes a first reagent for the step of distributing a sample into a plurality of partial samples as described herein, e.g., any of the distribution reagents described elsewhere herein. In some embodiments, the kit includes a second reagent for the step of subjecting the first partial sample to a procedure that affects a first nucleic acid base of the DNA of the first partial sample differently from a second nucleic acid base of the DNA, wherein the first nucleic acid base is a modified or unmodified nucleic acid base, the second nucleic acid base is a modified or unmodified nucleic acid base different from the first nucleic acid base, and the first and second nucleic acid bases have the same base-pairing specificity (e.g., any of the reagents described elsewhere herein for converting a nucleic acid base such as cytosine or methylated cytosine to a different nucleic acid base). The kit may include the first and second reagents and additional elements discussed below and / or elsewhere herein.

[0161] The kit includes SKI, THEMIS2, RPA2, TEKT2, STK40, GJA9-MYCBP, LOC105378663, HEYL, CNN3, JTB, FAM78B, ARV1, ADSS2, ZNF672, MBOAT2, ASXL2, SERTAD2, TMEM131, CLASP1, SATB2, ABHD14B, NISCH, TMEM45A, LGI2, KLHL5, NEUROG2-AS1, ABHD18, MFSD8, ELF2, TRIM2, AHRR, and PDCD. 6-AHRR, SEMA5A, IQGAP2, TSLP, SLC25A48, RELL2, ARHGAP26, SLC36A1, CNPY3, FAM229B, MAN1A1, ADCYAP1R1, KIAA0895, TRAPPC14, L INC01004, FAM131B, GIMAP4, SLC4A2, CD274, TOX, GDAP1, ZNF623, GNA14, S1PR3, C9orf47, ROR2, ERCC6L2, LINC00476, ECPAS, ASTN2, PHF19, PTGES2-AS1, RALGDS, HACD1, ABLIM1, LOC101927692, GFRA1, C11orf21, TRIM44, CHST1, TMX2-CTNND1, LOC101928069, PDE2A , DLG2, ENDOD1, DDX6, TULP3, PTPRO, ZCRB1, TMPO-AS1, HSP90B1, SIRT4, SRSF9, SLITRK1, MMP14, BCL2L2-PABPN1, KCNH5, TRAF3, IDH2 The present invention may further include multiple oligonucleotide probes that selectively hybridize with at least 5, 6, 7, 8, 9, 10, 20, 30, 40, or all of the genes selected from the group consisting of CIB1, MAN2A2, KDM8, ZFHX3, HSBP1, TOP3A, RETREG3, ADAM11, KPNB1, GRIN2C, GALR2, ZBTB14, EPB41L3, PDE4A, KLF1, SIX5, DM1-AS, ZNF114, CLEC11A, and LINC01530. The number of genes that an oligonucleotide probe can selectively hybridize with may vary.For example, the number of genes may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, or 54. The kit may include a container containing multiple oligonucleotide probes and instructions for carrying out any of the methods described herein.

[0162] Oligonucleotide probes can selectively hybridize with the exon regions of genes, for example, at least five genes. In some cases, oligonucleotide probes can selectively hybridize with at least 30 exons of genes, for example, at least five genes. In some cases, multiple probes can selectively hybridize with each of the at least 30 exons. Each probe hybridizing with an exon may have a sequence that overlaps with at least one other probe. In some embodiments, oligo probes can selectively hybridize with non-coding regions of genes disclosed herein, for example, intron regions of genes. Oligo probes can also selectively hybridize with regions of genes that include both exon and intron regions of genes disclosed herein.

[0163] Any number of exons can be targeted by oligonucleotide probes. For example, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 1 It is possible to target 65, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, 230, 235, 240, 245, 250, 255, 260, 265, 270, 275, 280, 285, 290, 295, 300, 400, 500, 600, 700, 800, 900, 1,000, or more exons.

[0164] The kit may include at least four, five, six, seven, or eight different library adapters, each having a distinct molecular barcode and the same sample barcode. The library adapters do not have to be sequencing adapters. For example, a library adapter may not include a flow cell sequence or a sequence that allows for the formation of a hairpin loop for sequencing. Various variations and combinations of molecular and sample barcodes are described throughout and applicable to the kit. Furthermore, in some cases, the adapters are not sequencing adapters. Furthermore, the adapters provided in the kit may also include sequencing adapters. A sequencing adapter may include a sequence that hybridizes with one or more sequencing primers. A sequencing adapter may further include a sequence that hybridizes with a solid support, such as a flow cell sequence. For example, a sequencing adapter may be a flow cell adapter. A sequencing adapter may be attached to one or both ends of a polynucleotide fragment. In some cases, the kit may include at least eight different library adapters, each having a distinct molecular barcode and the same sample barcode. The library adapters do not have to be sequencing adapters. The kit may further include a sequencing adapter having a first sequence that selectively hybridizes with a library adapter and a second sequence that selectively hybridizes with a flow cell sequence. In another example, the sequencing adapter may be hairpin-shaped. For example, a hairpin-shaped adapter may include complementary double-stranded portions and loop portions, the double-stranded portion of which can be attached to a double-stranded polynucleotide (e.g., ligation). By attaching the hairpin-shaped sequencing adapter to both ends of a polynucleotide fragment, a cyclic molecule that can be sequenced multiple times can be generated.The sequencing adapter can span from end to end up to 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, and 55 bases. The bases may be 1, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, or fewer than that. The sequencing adapter may contain 20-30 bases, 20-40 bases, 30-50 bases, 30-60 bases, 40-60 bases, 40-70 bases, 50-60 bases, and 50-70 bases from end to end. In a specific example, the sequencing adapter may contain 20-30 bases from end to end. In another example, the sequencing adapter may contain 50-60 bases from end to end. The sequencing adapter may contain one or more barcodes. For example, the sequencing adapter may contain a sample barcode. The sample barcode may contain a predetermined sequence. The sample barcode can be used to identify the source of the polynucleotide.The sample barcode may consist of at least 1 nucleic acid base, 2 nucleic acid bases, 3 nucleic acid bases, 4 nucleic acid bases, 5 nucleic acid bases, 6 nucleic acid bases, 7 nucleic acid bases, 8 nucleic acid bases, 9 nucleic acid bases, 10 nucleic acid bases, 11 nucleic acid bases, 12 nucleic acid bases, 13 nucleic acid bases, 14 nucleic acid bases, 15 nucleic acid bases, 16 nucleic acid bases, 17 nucleic acid bases, 18 nucleic acid bases, 19 nucleic acid bases, 21 nucleic acid bases, 22 nucleic acid bases, 23 nucleic acid bases, 24 nucleic acid bases, 25 nucleic acid bases, or more nucleic acid bases (or any length listed throughout), for example, at least 8 bases. The barcode may be a continuous or discontinuous sequence as described above.

[0165] The library adapter may be blunt-ended and Y-shaped, and may be shorter than or equal to 40 nucleic acid bases in length. Other variations may be found throughout and are applicable to the kit.

[0166] biomarkers This disclosure provides, for example, a method for using biomarkers for diagnosis, prognosis, and treatment selection for subjects with cancer. The biomarker may be any gene or gene variant whose presence, mutation, deletion, substitution, copy number, or translation (i.e., to a protein) is an indicator of a pathological condition. The biomarkers of this disclosure may include presence, mutation, deletion, substitution, copy number, or translation in one or more of EGFR, KRAS, MET, BRAF, MYC, NRAS, ERBB2, ALK, Notch, PIK3CA, APC, and SMO.

[0167] A biomarker is a genetic variant associated with one or more types of cancer. Biomarkers can be determined using one of several resources or methods. Biomarkers may be previously discovered or newly discovered using experimental or epidemiological techniques. Detection of a biomarker can be an indicator of cancer if the biomarker is highly correlated with cancer. Detection of a biomarker can be an indicator of cancer if the biomarker is present in a region or gene at a higher frequency than its frequency in a given background population or dataset.

[0168] Publicly available resources such as scientific literature and databases may contain detailed information on genetic variants found to be associated with cancer. Scientific literature may describe experiments or genome-wide association studies (GWAS) that associate one or more genetic variants with cancer. Databases may contain information gathered from sources such as scientific literature to provide a more comprehensive resource for determining one or more biomarkers. Non-exclusive examples of databases include FANTOM, GTex, GEO, Body Atlas, INSiGHT, OMIM (Online Mendelian Inheritance in Man, omim.org), cBioPortal (cbioportal.org), CIViC (Clinical Interpretations of Variants in Cancer, civic.genome.wustl.edu), DOCM (Database of Curated Mutations, docm.genome.wustl.edu), and ICGC Data Portal (dcc.icgc.org). In a further example, the COSMIC (Catalogue of Somatic Mutations in Cancer) database allows for the searching of biomarkers by cancer, gene, or mutation type. Novel biomarkers can also be identified through experiments such as case-control studies or association studies (e.g., genome-wide association studies).

[0169] One or more biomarkers can be detected in a sequencing panel. A biomarker may be one or more genetic variants associated with cancer. Biomarkers can be selected from single nucleotide variants (SNVs), copy number variants (CNVs), insertions or deletions (e.g., indels), gene fusions, and inversions. Biomarkers may affect protein levels. Biomarkers may be present in promoters or enhancers and may alter gene transcription. Biomarkers may affect gene transcription and / or translational efficacy. Biomarkers may affect the stability of transcribed mRNA. Biomarkers may alter the amino acid sequence of translated proteins. Biomarkers may affect splicing, alter amino acids encoded by specific codons, induce frameshifts, or induce stop codons. Biomarkers may induce conserved amino acid substitutions. One or more biomarkers may induce conserved amino acid substitutions. One or more biomarkers may induce non-conserved amino acid substitutions.

[0170] One or more of the biomarkers may be driver mutations. Driver mutations are mutations that provide a selective advantage to tumor cells by increasing either their survival or regeneration in their microenvironment. None of the biomarkers have to be driver mutations. One or more of the biomarkers may be passenger mutations. Passenger mutations are mutations that do not affect the fitness of tumor cells but may be associated with clonal expansion due to their location within the same genome as the driver mutation.

[0171] The frequency of the biomarker can be as low as 0.001%. The frequency of the biomarker can be as low as 0.005%. The frequency of the biomarker can be as low as 0.01%. The frequency of the biomarker can be as low as 0.02%. The frequency of the biomarker can be as low as 0.03%. The frequency of the biomarker can be as low as 0.05%. The frequency of the biomarker can be as low as 0.1%. The frequency of the biomarker can be as low as 1%.

[0172] A single biomarker may be absent in more than 50% of subjects with cancer. A single biomarker may be absent in more than 40% of subjects with cancer. A single biomarker may be absent in more than 30% of subjects with cancer. A single biomarker may be absent in more than 20% of subjects with cancer. A single biomarker may be absent in more than 10% of subjects with cancer. A single biomarker may be absent in more than 5% of subjects with cancer. A single biomarker may be present in 0.001% to 50% of subjects with cancer. A single biomarker may be present in 0.01% to 50% of subjects with cancer. A single biomarker may be present in 0.01% to 30% of subjects with cancer. A single biomarker may be present in 0.01% to 20% of subjects with cancer. A single biomarker may be present in 0.01% to 10% of subjects with cancer. A single biomarker may be present in 0.1% to 10% of subjects with cancer. A single biomarker may be present in 0.1% to 5% of individuals with cancer.

[0173] Detection of a biomarker may indicate the presence of one or more cancers. Detection may indicate the presence of a cancer selected from the group including ovarian cancer, pancreatic cancer, breast cancer, colorectal cancer, non-small cell lung cancer (e.g., squamous cell carcinoma or adenocarcinoma) or any other cancer. Detection may indicate the presence of any cancer selected from the group including ovarian cancer, pancreatic cancer, breast cancer, colorectal cancer, non-small cell lung cancer (squamous cell or adenocarcinoma) or any other cancer. Detection may indicate the presence of any of the cancers selected from the group including ovarian cancer, pancreatic cancer, breast cancer, colorectal cancer and non-small cell lung cancer (squamous cell or adenocarcinoma), or any other cancer. Detection may indicate the presence of one or more of the cancers referred to in this application.

[0174] One or more cancers may exhibit a biomarker in at least one exon within the panel. One or more cancers selected from the group including ovarian cancer, pancreatic cancer, breast cancer, colorectal cancer, non-small cell lung cancer (squamous cell or adenocarcinoma), or any other cancer, each exhibit a biomarker in at least one exon within the panel. Each of at least three cancers may exhibit a biomarker in at least one exon within the panel. Each of at least four cancers may exhibit a biomarker in at least one exon within the panel. Each of at least five cancers may exhibit a biomarker in at least one exon within the panel. Each of at least eight cancers may exhibit a biomarker in at least one exon within the panel. Each of at least ten cancers may exhibit a biomarker in at least one exon within the panel. All cancers may exhibit a biomarker in at least one exon within the panel.

[0175] If a subject has cancer, the subject may exhibit a biomarker in at least one exon or gene within the panel. At least 85% of subjects with cancer may exhibit a biomarker in at least one exon or gene within the panel. At least 90% of subjects with cancer may exhibit a biomarker in at least one exon or gene within the panel. At least 92% of subjects with cancer may exhibit a biomarker in at least one exon or gene within the panel. At least 95% of subjects with cancer may exhibit a biomarker in at least one exon or gene within the panel. At least 96% of subjects with cancer may exhibit a biomarker in at least one exon or gene within the panel. At least 97% of subjects with cancer may exhibit a biomarker in at least one exon or gene within the panel. At least 98% of subjects with cancer may exhibit a biomarker in at least one exon or gene within the panel. At least 99% of subjects with cancer may exhibit a biomarker in at least one exon or gene within the panel. At least 99.5% of subjects with cancer may exhibit a biomarker in at least one exon or gene within the panel.

[0176] If a subject has cancer, the subject may exhibit a biomarker in at least one region within the panel. At least 85% of subjects with cancer may exhibit a biomarker in at least one region within the panel. At least 90% of subjects with cancer may exhibit a biomarker in at least one region within the panel. At least 92% of subjects with cancer may exhibit a biomarker in at least one region within the panel. At least 95% of subjects with cancer may exhibit a biomarker in at least one region within the panel. At least 96% of subjects with cancer may exhibit a biomarker in at least one region within the panel. At least 97% of subjects with cancer may exhibit a biomarker in at least one region within the panel. At least 98% of subjects with cancer may exhibit a biomarker in at least one region within the panel. At least 99% of subjects with cancer may exhibit a biomarker in at least one region within the panel. At least 99.5% of subjects with cancer may exhibit a biomarker in at least one region within the panel.

[0177] Detection can be performed with high sensitivity and / or high specificity. Sensitivity can refer to a measure of the proportion of positives that are correctly identified as positive. In some cases, sensitivity refers to the percentage of all existing biomarkers detected. In some cases, sensitivity refers to the percentage of people with a disease that are correctly identified as having a particular disease. Specificity can refer to a measure of the proportion of negatives that are correctly identified as negative. In some cases, specificity refers to the proportion of unaltered bases that are correctly identified. In some cases, specificity refers to the percentage of healthy people that are correctly identified as not having a particular disease. The non-unique tagging method described above significantly increases the specificity of detection by reducing noise and sequencing errors generated by amplification, thereby reducing the frequency of false positives. Detection can be performed with a sensitivity of at least 95%, 97%, 98%, 99%, 99.5%, or 99.9% and / or a specificity of at least 80%, 90%, 95%, 97%, 98%, or 99%. Detection can be performed with a sensitivity of at least 90%, 95%, 97%, 98%, 99%, 99.5%, 99.6%, 99.98%, 99.9%, or 99.95%. Detection can be performed with a specificity of at least 90%, 95%, 97%, 98%, 99%, 99.5%, 99.6%, 99.98%, 99.9%, or 99.95%. Detection may be performed with at least 70% specificity and at least 70% sensitivity, at least 75% specificity and at least 75% sensitivity, at least 80% specificity and at least 80% sensitivity, at least 85% specificity and at least 85% sensitivity, at least 90% specificity and at least 90% sensitivity, at least 95% specificity and at least 95% sensitivity, at least 96% specificity and at least 96% sensitivity, at least 97% specificity and at least 97% sensitivity, at least 98% specificity and at least 98% sensitivity, at least 99% specificity and at least 99% sensitivity, or 100% specificity and 100% sensitivity. In some cases, the method may be able to detect the biomarker with a sensitivity of about 80% or higher. In some cases, the method may be able to detect the biomarker with a sensitivity of about 95% or higher.In some cases, the method can detect biomarkers with a sensitivity of approximately 80% or higher, and approximately 95% or higher.

[0178] Detection can be highly accurate. Accuracy can be applied to the identification of biomarkers in cell-free DNA and / or the diagnosis of cancer. Accuracy can be increased and / or measured using statistical tools such as the covariance analysis described above. The method may detect biomarkers with an accuracy of at least 80%, 90%, 95%, 97%, 98%, or 99%, 99.5%, 99.6%, 99.98%, 99.9%, or 99.95%. In some cases, the method may detect biomarkers with an accuracy of at least 95% or higher.

[0179] Cancer treatment and management In various embodiments, cancer treatments may include: ipilimumab (Yervoy), a CTLA4 inhibitor applied based on PD-L1 protein expression; tremelimumab (Imjuno); and the PD-1 inhibitor nivolumab (Opdivo). These can be used in combination with ipilimumab and may include platinum-based drugs as needed. Other PD-1 inhibitors include pembrolizumab (Keytruda), semiprimab-rwlc (Libtayo), and durvalumab (Imfinzi), which are used for unresectable NSCLC. Further information can be found in Basudan Clin Pract. 2023 Feb; 13(1): 22-40 and Meng et al., Cell Death Dis 15, 3 (2024), respectively, which are fully incorporated herein by reference.

[0180] In other embodiments, cancer treatments include atezolizumab (Tecentriq), imatinib, gefitinib, afatinib, dacomitinib, sunitinib, sorafenib, vandetanib, brivanib, cabozantib, neratinib, tivantinib, bevacizumab, cictumumab, darotuzumab, figtumumab, rilotumumab, onarutuzumab, ganitumumab, ramucirumab, ridafololimus, tesirolimus, everolimus, relatrimab, osimertinib, BMS-690514, BMS-754807, and EMD. Examples of antibodies suitable for use as anti-EGFR therapy include 525797, GDC-0973, GDC-0941, MK-2206, AZD6244, GSK1120212, PX-866, XL821, IMC-A12, MM-121, PF-02341066, RG7160, and Sym004. In some cases, EGFR tyrosine kinase inhibitors, such as gefitinib (Iressa), erlotinib (Tarceva), lapatinib, canertinib, and cetuximab, are used for cancer treatment.

[0181] In some cases, treatments can be used in combination, such as anti-EGFR therapy with other anti-EGFR therapy. Anti-EGFR therapy can be used in combination with any combination of chemotherapy agents or chemotherapy regimens, such as FOLFOX (fluorouracil [5-FU] / leucovorin / oxaliplatin) and FOLFIRI (5-FU / leucovorin / irinotecan).

[0182] In some cases, one type of cancer treatment is administered to a single patient. In other cases, cancer treatment is administered in combination with other treatments, such as non-anti-EGFR therapy and anti-EGFR therapy.

[0183] Sequence determination panel To improve the likelihood of detecting tumors exhibiting mutations, the DNA regions to be sequenced may include a panel of genes or genomic regions. Selecting limited regions for sequencing (e.g., a limited panel) can reduce the total sequencing required (e.g., the total amount of nucleotides sequenced). The sequencing panel may target multiple different genes or regions to detect a single cancer, a set of cancers, or all cancers.

[0184] In some embodiments, a panel targeting multiple different genes or genomic regions is selected such that a determined proportion of subjects with cancer exhibit a genetic variant or biomarker in one or more different genes or genomic regions within the panel. The panel may be selected so that the region for sequencing is limited to a fixed number of base pairs. The panel may be selected so that a desired amount of DNA is sequenced. The panel may be further selected to achieve a desired sequence read depth. The panel may be selected so that a desired sequence read depth or sequence read coverage is achieved for a given amount of sequenced base pairs. The panel may be selected so that theoretical sensitivity, theoretical specificity, and / or theoretical accuracy are achieved for detecting one or more genetic variants in a sample.

[0185] Probes for detecting panels of regions include probes for detecting hotspot regions, as well as nucleosome recognition probes (e.g., KRAS codons 12 and 13), which can be designed to optimize capture based on analysis of cfDNA coverage and fragment size variations influenced by nucleosome binding patterns and GC sequence composition. Regions used herein may also include non-hotspot regions optimized based on nucleosome location and GC model. The panel may include multiple subpanels, including subpanels for identifying origin tissue (e.g., using published literature to define 50-100 baits representing genes with the most diverse transcriptional profiles across tissues (not necessarily promoters)), whole-genome scaffolds (e.g., for identifying hyperconservative genomic content and tiling to low density with a small number of probes for copy number-based lining across chromosomes), and transcription start site (TSS) / CpG islands (e.g., for capturing differential methylation regions (e.g., differential methylation regions (DMRs)) in promoters of tumor suppressor genes (e.g., SEPT9 / VIM in colorectal cancer). In some embodiments, the origin tissue markers are tissue-specific epigenetic markers.

[0186] One or more regions within a panel may contain one or more loci from one or more genes. Multiple genes can be selected for sequencing and biomarker detection. The genes included in the regions to be sequenced can be selected from genes known to be involved in cancer or genes not involved in cancer. For example, the multiple genes in a panel may be oncogenes, tumor suppressors, growth factors, DNA repair genes, signaling genes, transcription factors, receptors, or genes involved in metabolism.Examples of genes that may be present in the panel include, but are not limited to, SKI, THEMIS2, RPA2, TEKT2, STK40, GJA9-MYCBP, LOC105378663, HEYL, CNN3, JTB, FAM78B, ARV1, ADSS2, ZNF672, MBOAT2, ASXL2, SERTAD2, TMEM131, CLASP1, SATB2, ABHD14B, NISCH, TMEM45A, LGI2, KLHL5, and NEUROG. 2-AS1, ABHD18, MFSD8, ELF2, TRIM2, AHRR, PDCD6-AHRR, SEMA5A, IQGAP2, TSLP, SLC25A48, RELL2, ARHGAP26, SLC36A1, CNPY3 , FAM229B, MAN1A1, ADCYAP1R1, KIAA0895, TRAPPC14, LINC01004, FAM131B, GIMAP4, SLC4A2, CD274, TOX, GDAP1, ZNF623, GNA 14, S1PR3,C9orf47, ROR2, ERCC6L2, LINC00476, ECPAS, ASTN2, PHF19, PTGES2-AS1, RALGDS, HACD1, ABLIM1, LOC101927692, GFRA1, C11orf21, TRIM44, CHST1, TMX2-CTNND1, LOC101928069, PDE2A, DLG2, ENDOD1, DDX6, TULP3, PTPRO, ZCRB1, TMPO-AS1 Examples include HSP90B1, SIRT4, SRSF9, SLITRK1, MMP14, BCL2L2-PABPN1, KCNH5, TRAF3, IDH2, CIB1, MAN2A2, KDM8, ZFHX3, HSBP1, TOP3A, RETREG3, ADAM11, KPNB1, GRIN2C, GALR2, ZBTB14, EPB41L3, PDE4A, KLF1, SIX5, DM1-AS, ZNF114, CLEC11A, and LINC01530.

[0187] In some cases, one or more regions within the panel are SKI, THEMIS2, RPA2, TEKT2, STK40, GJA9-MYCBP, LOC105378663, HEYL, CNN3, JTB, FAM78B, ARV1, ADSS2, ZNF672, MBOAT2, ASXL2, SERTAD2, TMEM131, CLASP1, SATB2, ABHD14B, NISCH, TMEM45A, LGI2, KLHL5, NEUROG2-AS1, ABHD18, MFSD8, ELF2, TRIM2, AHRR, PDCD6-AHRR, SEMA5A, IQGAP2, TSLP, SLC25A48, RELL2, ARHGAP26, SLC36A1, CNPY3, FAM229B, MAN1A1, AD CYAP1R1, KIAA0895, TRAPPC14, LINC01004, FAM131B, GIMAP4, SLC4A2, CD274, TOX, GDAP1, ZNF623, GNA14, S1PR3, C9orf47, ROR2, E RCC6L2, LINC00476, ECPAS, ASTN2, PHF19, PTGES2-AS1, RALGDS, HACD1, ABLIM1, LOC101927692, GFRA1, C11orf21, TRIM44, CHST1, TMX2-CTNND1, LOC101928069, PDE2A, DLG2, ENDOD1, DDX6, TULP3, PTPRO, ZCRB1, TMPO-AS1, HSP90B1, SIRT4, SRSF9, SLITRK1, MMP1 4, may include one or more loci from one or more genes, including one or more of the following: BCL2L2-PABPN1, KCNH5, TRAF3, IDH2, CIB1, MAN2A2, KDM8, ZFHX3, HSBP1, TOP3A, RETREG3, ADAM11, KPNB1, GRIN2C, GALR2, ZBTB14, EPB41L3, PDE4A, KLF1, SIX5, DM1-AS, ZNF114, CLEC11A, LINC01530.

[0188] In some embodiments, one or more regions within the panel include one or more loci from one or more genes for detecting residual cancer after surgery. This detection may be faster than what is possible with existing cancer detection methods. In some embodiments, one or more regions within the panel include one or more loci from one or more genes for detecting cancer in high-risk patient populations. For example, smokers have a much higher rate of lung cancer than the general population. Furthermore, smokers may develop other lung conditions that make cancer detection more difficult, such as the development of irregular nodules in the lungs. In some embodiments, the methods described herein detect cancer in high-risk patients faster than what is possible with existing cancer detection methods.

[0189] Regions can be selected for inclusion in the sequencing panel based on the number of subjects with cancer who have a biomarker in that gene or region. Regions can also be selected for inclusion in the sequencing panel based on the prevalence of subjects with cancer and the biomarkers present in that gene. The presence of a biomarker in a region can serve as an indicator of subjects with cancer.

[0190] In some cases, panels can be selected using information from one or more databases. Information about cancer can be obtained from cancer tumor biopsies or cfDNA assays. Databases may contain information describing populations of sequenced tumor samples. Databases may contain information about mRNA expression in tumor samples. Databases may contain information about regulatory elements in tumor samples. Information about sequenced tumor samples may include the frequencies of various genetic variants and describe the genes or regions in which the genetic variants exist. Genetic variants can be biomarkers. A non-limiting example of such a database is COSMIC, a list of somatic mutations found in various cancers. For a particular cancer, COSMIC ranks genes based on their mutation frequency. Genes can be selected for inclusion in a panel due to their high frequency of mutations within a given gene. For example, COSMIC has shown that 33% of a population of sequenced breast cancer samples have a mutation in TP53, and 22% of a population of sampled breast cancers have a mutation in KRAS. Other ranked genes, including APC, have mutations found in only about 4% of the population of breast cancer samples being sequenced. TP53 and KRAS can be included in the sequencing panel based on their relatively high frequency across sampled breast cancers (e.g., compared to APC, which is present in about 4% of samples). While COSMIC is provided as a non-limiting example, any database or set of information that associates cancer with biomarkers located in genes or gene regions can be used. In another example, as provided by COSMIC, 380 out of 1156 biliary tract cancer samples (33%) had mutations in TP53. Several other genes, such as APC, have mutations in 4–8% of all samples. Therefore, TP53 can be selected for inclusion in the panel based on its relatively high frequency in the population of biliary tract cancer samples.

[0191] Genes or regions in the sampled tumor tissue or circulating tumor DNA that have a significantly higher frequency of biomarkers than those found in a given background population can be selected for the panel. Combinations of regions can be selected to include in the panel such that at least the majority of subjects with cancer have a biomarker present in at least one of the regions or genes in the panel. Combinations of regions can be selected based on data showing that for a particular cancer or set of cancers, the majority of subjects have one or more biomarkers in one or more of the selected regions. For example, to detect cancer 1, a panel containing regions A, B, C, and / or D can be selected based on data showing that 90% of subjects with cancer 1 have biomarkers in regions A, B, C, and / or D in the panel. Alternatively, a biomarker may be shown to be independently present in two or more regions of subjects with cancer, and therefore, when combined, the majority of the population of subjects with cancer will have the biomarker in two or more regions. For example, to detect cancer 2, a panel including regions X, Y, and Z may be selected based on data showing that 90% of subjects have biomarkers in one or more regions, that in 30% of such subjects the biomarker is detected only in region X, and for the remaining subjects where the biomarker is detected, the biomarker is detected only in regions Y and / or Z. Biomarkers present in one or more regions previously shown to be associated with one or more cancers can be an indicator or predictor of subjects having cancer if the biomarker is detected 50% or more frequently in one or more of those regions. Computer techniques, such as models that use conditional probabilities of detecting cancer considering known cancer frequencies for sets of biomarkers within one or more regions, can be used to predict which regions, individually or in combination, may be predictors of cancer.Other methods for panel selection involve the use of databases containing information from studies using comprehensive genomic profiling and / or whole-genome sequencing (WGS, RNA-seq, Chip-seq, bisulfite sequencing, ATAC-seq, and others) of tumors with large panels. Information gathered from the literature may also generally describe pathways that are influenced and mutated in certain cancers. The use of ontologs containing genetic information can provide further information for panel selection.

[0192] The genes included in the panel for sequencing may include fully transcribed regions, promoter regions, enhancer regions, regulatory elements, and / or downstream sequences. To further increase the likelihood of detecting tumors exhibiting mutations, only exons may be included in the panel. The panel may include all exons of a selected gene, or it may include only one or more exons of a selected gene. The panel may include exons from each of several different genes. The panel may include at least one exon from each of several different genes.

[0193] In some embodiments, a panel of exons from each of several different genes is selected such that a determined proportion of subjects with cancer exhibits a genetic variant in at least one exon within the panel of exons.

[0194] At least one complete exon can be sequenced from each of the different genes within a panel of genes. The panel to be sequenced may contain exons from multiple genes. The panel may contain exons from 2 to 100 different genes, 2 to 70 genes, 2 to 50 genes, 2 to 30 genes, 2 to 15 genes, or 2 to 10 genes.

[0195] The selected panel may contain a variety of exons. The panel may contain 2 to 3000 exons. The panel may contain 2 to 1000 exons. The panel may contain 2 to 500 exons. The panel may contain 2 to 100 exons. The panel may contain 2 to 50 exons. The panel may contain 300 or fewer exons. The panel may contain 200 or fewer exons. The panel may contain 100 or fewer exons. The panel may contain 50 or fewer exons. The panel may contain 40 or fewer exons. The panel may contain 30 or fewer exons. The panel may contain 25 or fewer exons. The panel may contain 20 or fewer exons. The panel may contain 15 or fewer exons. The panel may contain 10 or fewer exons. The panel may contain 9 or fewer exons. The panel may contain 8 or fewer exons. The panel may contain 7 or fewer exons.

[0196] The panel may contain one or more exons from multiple different genes. The panel may contain one or more exons from each of a certain proportion of multiple different genes. The panel may contain at least two exons from at least 25%, 50%, 75%, or 90% of each of the different genes. The panel may contain at least three exons from at least 25%, 50%, 75%, or 90% of each of the different genes. The panel may contain at least four exons from at least 25%, 50%, 75%, or 90% of each of the different genes.

[0197] The size of the sequencing panel can vary. The sequencing panel can be larger or smaller (in terms of nucleotide size) depending on several factors, including, for example, the total amount of nucleotides to be sequenced or the number of unique molecules sequenced for a particular region within the panel. A sequencing panel can be between 5kb and 50kb in size. A sequencing panel may be between 10kb and 30kb in size. A sequencing panel may be between 12kb and 20kb in size. A sequencing panel may be between 12kb and 60kb in size. A sequencing panel may be at least 10kb, 12kb, 15kb, 20kb, 25kb, 30kb, 35kb, 40kb, 45kb, 50kb, 60kb, 70kb, 80kb, 90kb, 100kb, 110kb, 120kb, 130kb, 140kb, or 150kb in size. The sequencing panel can be less than 100kb, less than 90kb, less than 80kb, less than 70kb, less than 60kb, or less than 50kb in size.

[0198] The panels selected for sequencing may contain at least 1, 5, 10, 15, 20, 25, 30, 40, 50, 60, 80, or 100 regions. In some cases, the regions within the panel are selected so that the size of the regions is relatively small. In some cases, the regions within the panel have a size of approximately 10kb or less, approximately 8kb or less, approximately 6kb or less, approximately 5kb or less, approximately 4kb or less, approximately 3kb or less, approximately 2.5kb or less, approximately 2kb or less, approximately 1.5kb or less, or approximately 1kb or less, or less. In some cases, the areas within the panel have sizes ranging from approximately 0.5kb to 10kb, 0.5kb to 6kb, 1kb to 11kb, 1kb to 15kb, 1kb to 20kb, 0.1kb to 10kb, or 0.2kb to 1kb. For example, the size of the areas within the panel may range from approximately 0.1kb to 5kb.

[0199] The panels selected herein may enable sufficient deep sequencing to detect low-frequency genetic variants (e.g., in cell-free nucleic acid molecules obtained from a sample). The amount of a genetic variant in a sample may be referred to in terms of the minor allele frequency of a given genetic variant. Minor allele frequency may refer to the frequency of a small number of alleles (e.g., alleles that are not the most common) present in a given population of nucleic acids, such as a sample. Genetic variants with low minor allele frequencies may have a relatively low presence in a sample. In some cases, the panel may enable the detection of genetic variants with minor allele frequencies of at least 0.0001%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, or 0.5%. The panel may enable the detection of genetic variants with minor allele frequencies of 0.001% or higher. The panel may enable the detection of genetic variants with minor allele frequencies of 0.01% or higher. The panel may enable the detection of genetic variants present in a sample at frequencies as low as 0.0001%, 0.001%, 0.005%, 0.01%, 0.025%, 0.05%, 0.075%, 0.1%, 0.25%, 0.5%, 0.75%, or 1.0%. The panel may enable the detection of biomarkers present in a sample at frequencies of at least 0.0001%, 0.001%, 0.005%, 0.01%, 0.025%, 0.05%, 0.075%, 0.1%, 0.25%, 0.5%, 0.75%, or 1.0%. The panel may enable the detection of biomarkers present in a sample at frequencies as low as 1.0%. The panel may enable the detection of biomarkers present in a sample at frequencies as low as 0.75%. The panel may enable the detection of biomarkers with a frequency of 0.5% in the sample. The panel may enable the detection of biomarkers with a frequency of 0.25% in the sample. The panel may enable the detection of biomarkers with a frequency of 0.1% in the sample. The panel may enable the detection of biomarkers with a frequency of 0.075% in the sample. The panel may enable the detection of biomarkers with a frequency of 0.05% in the sample.The panel may enable the detection of biomarkers with a frequency of 0.025% in a sample. The panel may enable the detection of biomarkers with a frequency of 0.01% in a sample. The panel may enable the detection of biomarkers with a frequency of 0.005% in a sample. The panel may enable the detection of biomarkers with a frequency of 0.001% in a sample. The panel may enable the detection of biomarkers with a frequency of 0.0001% in a sample. The panel may enable the detection of biomarkers in sequenced cfDNA with a frequency of 1.0% to 0.0001% in a sample. The panel may enable the detection of biomarkers in sequenced cfDNA with a frequency of 0.01% to 0.0001% in a sample.

[0200] Genetic variants can be expressed as a percentage of the target population having a disease (e.g., cancer). In some cases, at least 1%, 2%, 3%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the population with cancer exhibit one or more genetic variants in at least one region within the panel. For example, at least 80% of the population with cancer may exhibit one or more genetic variants in at least one region within the panel.

[0201] A panel may include one or more regions from each of one or more genes. In some cases, a panel may include one or more regions from each of at least one, two, three, four, five, six, seven, eight, nine, ten, fifteen, twenty, twenty-five, thirty, forty, fifty, or eighty genes. In some cases, a panel may include one or more regions from each of up to one, two, three, four, five, six, seven, eight, nine, ten, fifteen, twenty, twenty-five, thirty, forty, fifty, or eighty genes. In some cases, a panel may include one or more regions from each of about 1 to about 80, 1 to about 50, about 3 to about 40, 5 to about 30, or 10 to about 20 different genes.

[0202] Regions within the panel can be selected to detect one or more epigenetic modification regions. These epigenetic modification regions may be acetylated, methylated, ubiquitinated, phosphorylated, SUMOlated, ribosylated, and / or citrullinated. For example, regions within the panel can be selected to detect one or more methylated regions.

[0203] Regions within a panel can be selected to contain sequences that are differentially transcribed across one or more tissues. In some cases, regions may contain sequences that are transcribed at a higher level in a particular tissue compared to other tissues. For example, regions may contain sequences that are transcribed in a particular tissue but not in other tissues.

[0204] Regions within a panel may contain coding and / or non-coding sequences. For example, a region within a panel may contain one or more sequences within exons, introns, promoters, 3' untranslated regions, 5' untranslated regions, regulatory elements, transcription start sites, and / or splice sites. In some cases, regions within a panel may contain other non-coding sequences, including pseudogenes, repetitive sequences, transposons, viral elements, and telomeres. In some cases, regions within a panel may contain sequences within non-coding RNA, such as ribosomal RNA, transfer RNA, Piwi-interacting RNA, and microRNA.

[0205] Regions within the panel can be selected so that cancer is detected (diagnosed) with a desired level of sensitivity (e.g., by detecting one or more genetic variants). For example, regions within the panel can be selected so that cancer is detected with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% (e.g., by detecting one or more genetic variants). Regions within the panel can also be selected so that cancer is detected with 100% sensitivity.

[0206] Regions within the panel can be selected so that cancer is detected (diagnosed) with a desired level of specificity (e.g., by detecting one or more genetic variants). For example, regions within the panel can be selected so that cancer is detected with a specificity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% (e.g., by detecting one or more genetic variants). Regions within the panel can be selected so that one or more genetic variants are detected with 100% specificity.

[0207] Regions within the panel can be selected so that cancer is detected (diagnosed) with a desired positive predictive value. Positive predictive value can be increased by increasing sensitivity (e.g., the likelihood of an actual positive being detected) and / or specificity (e.g., the likelihood of an actual negative not being falsely identified as positive). As a non-limiting example, regions within the panel can be selected so that one or more genetic variants are detected with a positive predictive value of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. Regions within the panel can also be selected so that one or more genetic variants are detected with a 100% positive predictive value.

[0208] Areas within the panel can be selected to allow cancer to be detected (diagnosed) with the desired accuracy. As used herein, the term “accuracy” may refer to the test’s ability to distinguish between a diseased state (e.g., cancer) and health. Accuracy can be quantified using measures such as sensitivity and specificity, predictive value, likelihood ratio, area under the ROC curve, Joden index, and / or diagnostic odds ratio.

[0209] Accuracy can be presented as a percentage, referring to the ratio of the number of tests that yielded correct results to the total number of tests performed. Areas within the panel can be selected to detect cancer with at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy. Areas within the panel can also be selected to detect cancer with 100% accuracy.

[0210] The panel can be selected such that the specificity decreases noticeably when one or more regions or genes within the panel are deleted. Deleting a single region from the panel may result in a specificity reduction of at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, or greater.

[0211] The panel can be selected so that the addition of one or more regions or genes to the panel does not cause a noticeable increase in the panel's specificity, for example, so that the specificity does not increase by more than 1%, more than 2%, more than 5%, more than 10%, more than 15%, or more than 20%.

[0212] The panel may be of a size such that the sensitivity decreases noticeably when one or more regions or genes within the panel are deleted, for example, by at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, or more.

[0213] The panel can be selected so that the addition of one or more regions or genes to the panel does not cause a noticeable increase in the panel's sensitivity, for example, so that the sensitivity does not increase by more than 1%, more than 2%, more than 5%, more than 10%, more than 15%, or more than 20%.

[0214] The panel may be of a size such that the accuracy decreases noticeably when one or more regions or genes within the panel are deleted, for example, by at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, or more.

[0215] The panel can be selected such that adding one or more regions or genes to the panel does not increase the panel's accuracy to a noticeable degree, for example, not by more than 1%, more than 2%, more than 5%, more than 10%, more than 15%, or more than 20%.

[0216] The panel may be of a size such that deleting one or more regions or genes within the panel results in a noticeable decrease in the positive predictive value, for example, by at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, or more.

[0217] The panel can be selected so that the addition of one or more regions or genes to the panel does not cause a noticeable increase in the panel's positive predictive value, for example, so that the positive predictive value does not increase by more than 1%, more than 2%, more than 5%, more than 10%, more than 15%, or more than 20%.

[0218] The panel can be selected to detect low-frequency genetic variants with high sensitivity. For example, the panel can be selected to detect genetic variants or biomarkers present at frequencies as low as 0.01%, 0.05%, or 0.001% in the sample with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. Regions within the panel can be selected to detect biomarkers present at frequencies of 1% or less in the sample with a sensitivity of 70% or higher. The panel can be selected to detect biomarkers with a frequency of 0.1% or lower in the sample with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. The panel can be selected to detect biomarkers with a frequency of 0.01% or lower in the sample with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. The panel can be selected to detect biomarkers with a frequency of 0.001% or lower in the sample with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.

[0219] The panel can be selected to detect low-frequency genetic variants with high specificity. For example, the panel can be selected to detect genetic variants or biomarkers present at frequencies as low as 0.01%, 0.05%, or 0.001% in the sample with specificity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. Regions within the panel can be selected to detect biomarkers present at frequencies of 1% or less in the sample with specificity of 70% or higher. The panel can be selected so that biomarkers with a frequency of 0.1% or lower in the sample can be detected with a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. The panel can be selected so that biomarkers with a frequency of 0.01% or lower in the sample can be detected with a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. The panel can be selected so that biomarkers with a frequency of 0.001% or lower in the sample are detected with a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.

[0220] The panel can be selected to detect low-frequency genetic variants with high accuracy. The panel can be selected to detect genetic variants or biomarkers present at frequencies as low as 0.01%, 0.05%, or 0.001% in the sample with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy. Regions within the panel can be selected to detect biomarkers present at frequencies of 1% or less in the sample with 70% or higher accuracy. The panel can be selected to detect biomarkers present at frequencies as low as 0.1% in the sample with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy. The panel can be selected so that biomarkers with a frequency of 0.01% or lower in the sample can be detected with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy. The panel can be selected so that biomarkers with a frequency of 0.001% or lower in the sample can be detected with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy.

[0221] The panel can be selected to provide high predictive power and detect low-frequency genetic variants. The panel can be selected so that genetic variants or biomarkers present at a frequency of 0.01%, 0.05%, or 0.001% in the sample have a positive predictive value of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.

[0222] The concentration of the probe or bait used in the panel can be increased (2–6 ng / μL) to capture more nucleic acid molecules in the sample. The concentration of the probe or bait used in the panel may be at least 2 ng / μL, 3 ng / μL, 4 ng / μL, 5 ng / μL, 6 ng / μL, or higher. The probe concentration may be approximately 2 ng / μL to 3 ng / μL, approximately 2 ng / μL to 4 ng / μL, approximately 2 ng / μL to 5 ng / μL, or approximately 2 ng / μL to 6 ng / μL. The concentration of the probe or bait used in the panel may be 2 ng / μL or higher, up to 6 ng / μL, or less. In some cases, this may allow for the analysis of more molecules in the biological sample, thereby potentially enabling the detection of less frequent alleles.

[0223] While preferred embodiments of the present invention are shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided merely as examples. The present invention is not intended to be limited by any specific examples provided herein. Although the present invention is described in relation to the above specifications, the descriptions and illustrations of embodiments herein are not intended to be construed as limiting. Those skilled in the art will readily conceive of numerous variations, alterations, and substitutions without departing from the present invention. Furthermore, it should be understood that all aspects of the present invention are not limited to any specific descriptions, configurations, or relative proportions described herein, which depend on various conditions and variables. It should be understood that various alternatives to the embodiments disclosed herein can be used in the practice of the present invention. Therefore, this disclosure is intended to encompass all such alternatives, alterations, variations, or equivalents. The scope of the present invention is defined by the following claims, and the methods and structures within these claims and their equivalents are intended to be encompassed thereby.

[0224] While the foregoing disclosure is described in some detail as explanations and examples for clarity and understanding, it will be apparent to those skilled in the art that various variations in form and detail can be made without departing from the true scope of this disclosure and can be implemented within the scope of the appended claims. For example, all methods, systems, computer-readable media, and / or component features, steps, elements, or other embodiments thereof can be used in various combinations. [Examples]

[0225] (Example 1) PD-L1 as a biomarker for general and therapeutic response. Blood-based assessment of PD-L1 status has significant value for NSCLC and other distant cancers. In this context, PD-L1 is included in methods for testing patients for PD-L1 expression, often accompanied by immunohistochemical testing, as a biomarker of treatment response. These methods require separate protocols, samples, and can be time-consuming.

[0226] The methods described herein support the measurement of PD-L1 expression from methylation data and, in some embodiments, in samples containing cell-free DNA. Therefore, no additional tests or samples are required, the workflow is simplified, and costs and time are saved by increased informative capacity from a single test.

[0227] Those skilled in the art will readily understand that the techniques described herein can be extended to methylation data for measuring MSI or BRAF status from sample MSI or BRAF status and the location of nucleosomes within the promoter region.

[0228] (Example 2) Build a predictive model Here, to construct a predictive model in one example, a PD-L1 expression predictive model can be deployed from tissue bisulfite TCGA data and overlaid on an epigenome panel. The TCGA data (COADREAD tissue cohort) includes 384 CRC samples, with methylation data obtained from a 450k Illumina microarray (single-site bisulfite sequencing) and gene expression measured by normalized RNASeq. This methylation data can be converted to other measured epigenome panels by averaging the beta values ​​of probes overlapping with the targeted Infinity region (Infinity 23,936 regions are represented on the 450k Illumina array).

[0229] Here, samples are labeled as PDL1-low (bottom 50%), PD-L1-medium (25%), or PD-L1-high (top 25%) based on CD274 gene expression. A penalized logistic regression model (LASSO) is applied, with the response variable being sample id PD-L1 high or low, and the predictor variable being the methylation score (beta) of all Infinity targeting regions, and 10-fold cross-validation is used. In one example, approximately 50 regions were selected by LASSO.

[0230] (Example 3) Relationship between PD-L1 promoter region methylation status and sample MSI or BRAF status Without being bound by any particular theory, BRAFV600E may upregulate PD-L1 expression transcriptionally, which has been shown to enhance chemotherapy-induced apoptosis. Such ability may reflect the endogenous non-immune function of PD-L1, suggesting its potential as a predictive biomarker.

[0231] Here, the promoter region of PD-L1, as measured by the epigenome panel using MBD segmentation, was measured as molecular count = 0 for all samples except 15 high-segment samples and a significant number of molecules for low-segment samples. Approximately 40 samples were identified as MSI-H (mainly CRC and breast), and 300 samples were identified as BRAF V600E positive (mainly CRC).

[0232] Similarly, it is possible to predict sample MSI or BRAF status from genome-wide methylation.

[0233] (Example 4) Prediction of sample MSI or BRAF status from nucleosome location in the PD-L1 promoter region. Without being bound by any particular theory, cell-free DNA may possess a nucleosome footprint that is potentially informative regarding the tissue of origin. Nucleosomes positioned in highly favorable locations adjacent to nucleosome-depleted regions are likely generated in a transcriptionally independent manner by the nucleosome-remodeling complex and regulated by the pre-initiation complex (PIC) and related factors. Transcriptional elongation, and the recruitment of nucleosome-remodeling active histone chaperones by elongation mechanisms, may lead to further downstream positioning.

[0234] (Example 5) Treatment in NSCLC For example, PD-L1 status can be used to determine the treatment of cancer, such as non-small cell lung cancer (NSCLC). One example is ipilimumab (Yervoy), which is a CTLA4 inhibitor and is applied based on PD-L1 protein expression. Nivolumab (Opdivo) is a PD-1 inhibitor and can be used in combination with ipilimumab, and platinum-based drugs may also be included as needed.

[0235] Other PD-1 inhibitors used for unresectable NSCLC include pembrolizumab (Keytruda), semiprimab-rwlc (Libtayo), and durvalumab (Imfinzi). After resection, atezolizumab (Tecentriq) is used. Each of the above can be used as an adjunct treatment and is included in combination with other therapies (e.g., carboplatin, platinum-based drugs).

[0236] Using the methods described herein, it is possible to monitor the target PD-L1 to determine whether to perform these or other procedures, which is not possible when using current tissue sectioning techniques.

[0237] (Example 6) Genetic enrichment Gene set enrichment analyses were performed using the PD-L1 predictor variable region. Various databases, including miRNA target interactions, BioCarta, and Gene Ontology (GO) molecular function analyses, were used for the gene sets, and the gp.enrichr function in the gseapy library was employed. The inventors found that the microRNAs (miRNAs) hsa-miR-6132, hsa-miR-6836-5p, hsa-miR-1909-3p, and hsa-miR-6722-3p were significant regulators of the inventors' gene sets (corrected P-value < 0.05). These results are summarized in Figure 10. In particular, the inventors found that the microRNA mhsa-miR-6836-5p modulates genes in the input set, including SKI, SEMA5A, FAM131B, SLC4A2, CLASP1, and HSP90B1 (corrected P-value = 0.02). hsa-miR-6836-5p has previously been suggested to promote osimertinib (Tagrisso) resistance in non-small cell lung cancer (NSCLC) through its role in the MSTRG.292666.16 / miR-6836-5p / MAPK8IP3 axis. Specifically, hsa-miR-6836-5p is downregulated in the presence of M2 tumor-associated macrophage-derived exosomes, which leads to upregulation of the long non-coding RNA (lncRNA) MSTRG.292666.16 and MAPK8IP3, thereby contributing to resistance to osimertinib treatment. Therefore, the genes SKI, SEMA5A, FAM131B, SLC4A2, CLASP1, and HSP90B1 within the gene set may contribute to osimertinib resistance.

[0238] (Example 7) Treatment, resistance mechanism Osimertinib is an EGFR (epidermal growth factor receptor) tyrosine kinase inhibitor specifically designed for the treatment of non-small cell lung cancer (NSCLC) with certain EGFR mutations. The indications are as follows:

[0239] Adjunctive therapy for EGFR mutation-positive non-small cell lung cancer (NSCLC)

[0240] TAGRISSO is indicated as adjuvant therapy after tumor resection in adult patients with non-small cell lung cancer (NSCLC) whose tumors have epidermal growth factor receptor (EGFR) exon 19 deletion or exon 21 L858R mutation detected by an FDA-approved test.

[0241] First-line treatment for EGFR mutation-positive metastatic NSCLC

[0242] TAGRISSO is indicated as a first-line treatment for adult patients with metastatic NSCLC whose tumors have an EGFR exon 19 deletion or exon 21 L858R mutation detected by an FDA-approved test.

[0243] First-line treatment for EGFR mutation-positive locally advanced or metastatic NSCLC.

[0244] TAGRISSO, in combination with pemetrexed and platinum-based chemotherapy, is indicated as a first-line treatment for adult patients with locally advanced or metastatic NSCLC whose tumors have an EGFR exon 19 deletion or exon 21 L858R mutation detected by an FDA-approved test.

[0245] Previously treated EGFR T790M mutation-positive metastatic NSCLC

[0246] TAGRISSO is indicated for the treatment of adult patients with metastatic EGFR T790M mutation-positive NSCLC detected by FDA-approved testing, whose disease has progressed during or after treatment with an EGFR tyrosine kinase inhibitor (TKI).

[0247] (Example 8) ER-associated degradation (ERAD) pathway The PDL1 predictor region was further enriched with information related to the function of the ERAD pathway. The corrected P-value for the ER-Associated Degradation (ERAD) Pathway in Homo sapiens (BioCarta_2016) was 0.055. ERAD is crucial for maintaining cellular protein homeostasis by identifying and degrading misfolded proteins in the endoplasmic reticulum. In HER2-positive breast cancer, the ERAD pathway plays a vital role in mitigating ER stress induced by increased protein toxicity due to increased HER2 / mTOR activity. This stress management is essential for the survival and resistance of HER2-positive cancer cells. Furthermore, genetic and pharmacological inhibition of the ERAD pathway leads to irreversible ER stress and selective death of HER2-positive cancer cells, including those resistant to conventional HER2-targeted therapies.

[0248] Without being bound by any particular theory, these results indicate that the ERAD pathway is also associated with osimertinib resistance, at least via the MAN2A2 and MAN1A1 genes. Furthermore, these results further support our previous observation that the genes in the set are regulated by hsa-miR-6836-5p, a key regulator of resistance to osimertinib.

[0249] Finally, Ontology (GO) molecular function analysis predicted that the gene set possesses protein-binding molecular function (GO:0005515), with a p-value of 0.0094, indicating that the protein selectively interacts with one or more specific proteins via non-covalent bonds. These interactions may include enzyme-substrate interactions, receptor-ligand binding, and protein complex formation.

Claims

1. A step of detecting methylation at at least one of multiple sites, A step of generating one or more measurement criteria for each of the aforementioned multiple parts, The steps include processing one or more of the aforementioned metrics to characterize the sample, and Methods that include...

2. The method according to claim 1, wherein the one or more measurement criteria are obtained from methylated chols from each of the plurality of sites.

3. The method according to claim 1, comprising the step of obtaining a sample.

4. The method according to claim 1, further comprising the condition that a sample has been obtained.

5. The method according to claim 1, wherein the step of characterizing the sample includes determining the gene expression of one or more biomarkers.

6. The method according to claim 4, wherein the one or more biomarkers include PD-L1, MSI and / or BRAF.

7. The method according to claim 1, further comprising the step of constructing a binary classification model from methylation data of a set of training samples containing PDL-1 high and PDL-1 low status.

8. The method according to claim 6, wherein the classification model is trained using cross-validation.

9. The method according to claim 7, wherein the cross-validation includes using 3-segment or 10-segment cross-validation.

10. The method according to claim 6, wherein a region is selected using penalized logistic regression and Least Absolute Contraction Selection Operator (LASSO) regularization.

11. The method according to claim 9, wherein the penalized logistic regression model includes, for each of the plurality of sites, a response variable PD-L1 and a predictor variable methylation call.

12. The method according to claim 1, wherein the portion includes a custom panel.

13. The method according to claim 11, wherein the custom panel is constructed in an in silico panel.

14. The method according to claim 11, wherein the custom panel is made up of a physical panel.

15. The method according to claim 11, wherein the custom panel includes a set of oncogenes, promoter regions for the set of oncogenes, HRR genes, immuno-oncology (IO) genes, cancer pathways, methylation peaks found in cancer, or methylation peaks found in clinical samples.

16. The method according to claim 11, wherein the custom panel is refined based on at least literature annotations, typical methylation peak locations, and / or public datasets.

17. The method according to claim 1, wherein the PDL-1 status is determined based on gene expression data, PD-L1 promoter region nucleosome position, or histological data.

18. The method according to claim 16, wherein the PD-L1 status is used to predict the treatment response.

19. The method according to claim 17, wherein the treatment comprises one or more of the following: an immune checkpoint inhibitor (ICI), a poly(ADP-ribose) polymerase (PARP) inhibitor, a kinase inhibitor, or an aromatase inhibitor, a CTLA4 inhibitor, a PD-L1 inhibitor, or a PD-1 inhibitor, either alone or in combination with fluoropyrimidine-containing chemotherapy and platinum-containing chemotherapy.

20. The method according to claim 18, wherein the immune checkpoint inhibitor is pembrolizumab.

21. The method according to claim 18, wherein the poly(ADP-ribose) polymerase (PARP) inhibitor is olaparib or talazoparib.

22. A step of detecting methylation at at least one of multiple sites, The steps include generating multiple methylation calls for each of the aforementioned multiple sites, A step of obtaining one or more metrics from the methylated coal, The steps include processing one or more of the aforementioned metrics to generate the probability that a patient exhibits PD-L1 expression, and Methods that include...

23. The method according to claim 21, wherein the patient is a lung cancer patient and the PD-L1 level corresponds to high PD-L1 expression as measured by proteomics, histological examination or immunohistochemical examination.

24. The method according to claim 22, wherein the high PD-L1 expression includes PD-L1 expression in 1% or more of tumor cells.

25. The method according to claim 22, wherein high PD-L1 expression includes that 50% or more of tumor cells are stained with PD-L1 [TC ≥ 50%] or that PD-L1 stained tumor-infiltrating immune cells [IC] account for 10% or more of the tumor area [IC ≥ 10%].

26. The method according to claim 22, wherein the patient does not exhibit either an EGFR genomic abnormality or an ALK genomic abnormality.

27. The method according to claim 22, wherein the patient does not exhibit EGFR genomic abnormalities, ALK genomic abnormalities, or ROS genomic abnormalities.

28. The method according to any one of claims 21 to 26, wherein the patient is administered a PD-L1 inhibitor or a CTLA4 inhibitor alone or in combination with platinum-containing chemotherapy.

29. A step of detecting methylation at at least one of multiple sites, The steps include generating multiple methylation calls for each of the aforementioned multiple sites, A step of obtaining one or more metrics from the methylated coal, A step of processing one or more of the aforementioned metrics to generate the probability that a patient exhibits PD-L1 expression, The step of determining that the aforementioned patient is a candidate for treatment using PARPi, Methods that include...

30. A step of detecting methylation at at least one of multiple sites, The steps include generating multiple methylation calls for each of the aforementioned multiple sites, A step of obtaining one or more metrics from the methylated coal, A step of processing one or more of the aforementioned metrics to generate the probability that a patient exhibits PD-L1 expression, The steps include determining that the patient is a candidate for treatment with jedatricib and talazoparib, and Methods that include...

31. The method according to claim 29, wherein jedatricib sensitizes advanced TNBC or BRCA1 / 2 mutant breast cancer to talazoparib-mediated PARP inhibition.

32. The method according to any one of the preceding claims, wherein the cancer is breast cancer, bladder cancer, cervical cancer, colon cancer, head and neck cancer, Hodgkin lymphoma, liver cancer, lung cancer, renal cell carcinoma, skin cancer including melanoma, gastric cancer, rectal cancer, and any solid tumor in which errors in the DNA that occur when the DNA is copied cannot be repaired.

33. The method according to any one of the preceding claims, wherein the sample comprises cell-free DNA.

34. The steps include detecting nucleosome positioning in at least one of several genomic regions and generating a nucleosome occupancy profile of the genomic region, A step of obtaining one or more metrics from the nucleosome occupancy profile, The steps include processing one or more of the aforementioned metrics to generate the probability that a patient exhibits PD-L1 expression, and Methods that include...

35. At least one of the aforementioned multiple parts is SKI, THEMIS2, RPA2, TEKT2, STK40, GJA9-MYCBP, LOC105378663, HEYL, CNN3, JTB, FAM78B, ARV1, ADSS2, ZNF672, MBOAT 2, ASXL2, SERTAD2, TMEM131, CLASP1, SATB2, ABHD14B, NISCH, TMEM45A, LGI2, KLHL5, NEUROG2-AS1, ABHD18, MFSD8, ELF 2, TRIM2, AHRR, PDCD6-AHRR, SEMA5A, IQGAP2, TSLP, SLC25A48, RELL2, ARHGAP26, SLC36A1, CNPY3, FAM229B, MAN1A1, A DCYAP1R1, KIAA0895, TRAPPC14, LINC01004, FAM131B, GIMAP4, SLC4A2, CD274, TOX, GDAP1, ZNF623, GNA14, S1PR3, C9or f47, ROR2, ERCC6L2, LINC00476, ECPAS, ASTN2, PHF19, PTGES2-AS1, RALGDS, HACD1, ABLIM1, LOC101927692, GFRA1, C1 1orf21, TRIM44, CHST1, TMX2-CTNND1, LOC101928069, PDE2A, DLG2, ENDOD1, DDX6, TULP3, PTPRO, ZCRB1, TMPO-AS1, HSP 90B1, SIRT4, SRSF9, SLITRK1, MMP14, BCL2L2-PABPN1, KCNH5, TRAF3, IDH2, CIB1, MAN2A2, KDM8, ZFHX3, HSBP1, TOP3A, RETREG3, ADAM11, KPNB1, GRIN2C, GALR2, ZBTB14, EPB41L3, PDE4A, KLF1, SIX5, DM1-AS, ZNF114, CLEC11A, and LINC01530 The method according to any of the preceding claims, wherein the gene is one or more genes selected from the group consisting of the following.

36. Steps to diagnose that the subject has cancer A method according to any of the prior claims, including the method described above.

37. Steps to determine the prognosis of a subject who is likely to develop cancer. A method according to any of the prior claims, including the method described above.

38. Steps to select the treatment for the target A method according to any of the prior claims, including the method described above.

39. Steps to implement treatment on the target A method according to any of the prior claims, including the method described above.