Methods to monitor patients treated with a cancer vaccine
Integrating genomic and epigenetic assays for personalized cancer vaccines addresses the limitations of current monitoring methods by providing real-time MRD detection and enabling tailored treatment strategies, improving vaccine efficacy.
Patent Information
- Application Number
- PCT/US2025/042054
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-21
- Filing Date
- 2025-08-14
- Publication Date
- 2026-02-26
AI Technical Summary
Current methods for monitoring patients treated with personalized cancer vaccines (PCV) are inadequate as they primarily focus on genetic alterations, neglecting epigenetic changes that occur in conjunction with treatment response, resistance, and tumor evolution, leading to potential disease recurrence due to minimal residual disease (MRD).
Developing highly sensitive assays that integrate genomic and epigenetic information, including DNA methylation status and differentially methylated regions (DMRs), to monitor neoantigen targets and MRD, using sequencing and computational models to assess therapeutic response and adjust treatment regimens.
Provides real-time monitoring of therapeutic response and MRD, enabling personalized treatment adjustments based on comprehensive molecular analysis, enhancing the effectiveness of cancer vaccines.
Smart Images

Figure US2025042054_26022026_PF_FP_ABST
Abstract
Description
Attorney Docket No. : GH0224WOMETHODS TO MONITOR PATIENTS TREATED WITH A CANCER VACCINECROSS-REFERENCE
[0001] This application claims the benefit of, and relies on the filing date of, U.S. provisional patent application number 63 / 685,393, which was filed August 21, 2024, the entire disclosure of which is incorporated herein by reference.FIELD OF THE INVENTION
[0002] The present disclosure relates generally to therapeutic cancer vaccines, specifically methods and assays designed to monitor neoantigen targets and manage minimal residual disease (MRD) post-vaccination.BACKGROUND
[0003] Recent studies in oncology have underscored the potential of personalized cancer vaccines (PCV) in treating various diseases. These vaccines are designed to trigger de novo immune responses against target neoantigens, which are unique peptides generated from tumor-specific mutations. By stimulating an immune response against these neoantigens, personalized cancer vaccines aim to enhance the immune system’s ability to fight cancer. Even after a successful treatment response with personalized cancer vaccines, a patient may still harbor minimal residual disease (MRD), which could lead to disease recurrence if not effectively monitored and managed. MRD refers to the small fraction of cancer cells that may persist even after a patient exhibits a complete remission of the disease following treatment.
[0004] Current methods to monitor patients treated with the personalized therapeutic cancer vaccine have largely focused on tumor-specific genomic alterations such as nucleic acid mutations and fusions that can be traced in the blood after treatment. These assays first profile genetic mutations on the primary tumor tissue to identify somatic clonal variants which are then followed using targeted resequencing in plasma.
[0005] However, these methods ignore epigenetic changes that occur in conjunction with genetic alterations. Such epigenetic changes occur with disease state, in response to treatment (both PCV and standard of care cancer therapies), resistance to treatment, tumor evolution, and metastasis. Therefore, patients undergoing treatment with a PCV would benefit from highlyAttorney Docket No. : GH0224WO sensitive assays that combine both genomic and epigenetic information to monitor neoantigen targets and molecular residual disease.SUMMARY
[0006] The present disclosure provides methods and systems to develop highly sensitive assays that combine both genomic and epigenetic information to monitor neoantigen targets and molecular residual disease (MRD) in patients undergoing treatment with a PCV alone or in combination with one or more therapies. The epigenetic data used to characterize changes in response to cancer vaccine treatment in individuals undergoing PCV therapy may include DNA methylation status at single CpG sites, regional methylation patterns, methylation-based signatures, and / or methylation levels measured at both single-site and / or fragment-level resolution.
[0007] In one aspect, the disclosure provides a method for monitoring therapeutic response in a subject treated with a personalized cancer vaccine (PCV), comprising: (a) obtaining a tumor and match normal tissue sample from the subject prior to initiation of treatment with the PCV, extracting a plurality of molecules from the tumor tissue sample, and sequencing the plurality of molecules to identify one or more biomarkers comprising tumor-specific neoantigens and differentially methylated regions (DMRs); (b) obtaining a plasma sample from the subject after the initiation of treatment with the PCV and selecting the biomarkers based on the sequencing data from the tumor tissue sample in (a) and detecting the biomarkers in cell free DNA from the plasma sample to determine the therapeutic response by assessing changes in neoantigen levels and DMRs, wherein a decrease in neoantigen levels indicates a response to treatment, and an increase or unchanged level indicates potential resistance or disease progression, and wherein the presence of at least one of the DMRs correlate with MRD in the subject, thereby monitoring therapeutic response in the subject treated with the PCV.
[0008] In another aspect, the disclosure provides a method for monitoring therapeutic response in a subject treated with a personalized cancer vaccine (PCV), comprising, obtaining a tumor and match normal tissue sample from the subject prior to initiation of treatment with the PCV, extracting a plurality of molecules from the tumor tissue sample, and sequencing the plurality of molecules to identify one or more tumor-specific neoantigens and differentially methylated regions (DMRs); (b) selecting a personalized panel incorporating one or more of the neoantigens and the DMRs identified (a); (c) obtaining a plasma sample from the subject after the initiation of treatment with the PCV; and (d) determining the therapeutic response by assessing changes in neoantigen levels, wherein a decrease in neoantigen levels indicates a response to treatment, and an increase or unchanged level indicates potential resistance orAttorney Docket No. : GH0224WO disease progression, and wherein the presence of at least one of the DMRs in the personalized panel correlate with MRD in the subject, thereby monitoring therapeutic response in the subject treated with the PCV.
[0009] In additional aspects, the disclosure provides a method for monitoring therapeutic response in a subject treated with a personalized cancer vaccine (PCV), comprising, (a) obtaining a tumor and match normal tissue sample from the subject prior to initiation of treatment with the PCV, extracting a plurality of molecules from the tumor tissue sample, and sequencing the plurality of molecules to identify one or more tumor-specific neoantigens and differentially methylated regions (DMRs); (b) selecting a personalized panel incorporating one or more of the neoantigens and the DMRs identified (a); (c) obtaining a plasma sample from the subject after the initiation of treatment with the PCV and determining the therapeutic response by assessing changes in neoantigen levels and DMRs, wherein a decrease in neoantigen levels indicates a response to treatment, and an increase or unchanged level indicates potential resistance or disease progression, and wherein the presence of at least one of the DMRs in the personalized panel correlate with MRD in the subject, thereby monitoring therapeutic response in the subject treated with the PCV.
[0010] Moreover, the method may include (a) obtaining additional plasma samples from the subject at a plurality of subsequent time points following the initiation of treatment; (b) analyzing the additional plasma samples to monitor therapeutic response using the method described in claim 1(b); and (c) monitoring temporal changes in the levels of the detected neoantigens and / or DMRs across the plurality of time points to assess the therapeutic response over time.
[0011] Furthermore, the treatment may include a combination of the PCV and one or more therapies. In additional embodiments, the one or more therapies comprise molecules that enhance the immune system’s ability to recognize and attack cancer cells. In additional embodiments, the one or more therapies are selected from the group consisting of: immune checkpoint inhibitors (ICIs), monoclonal antibodies, and immune modulators. In some embodiments, the one or more ICIs comprise an anti -programmed cell death protein 1 (Anti- PD-1) molecule, monoclonal antibody, bispecific antibody (BsAbs) or a multispecific antibody. In some embodiments, the one or more ICIs comprise an anti-Cytotoxic T lymphocyte-associated protein-4 (Anti-CTLA-4) molecule, monoclonal antibody, bispecific antibody (BsAbs) or a multispecific antibody. In some embodiments, the one or more ICIs are selected from the group consisting of: Pembrolizumab (Keytruda), Nivolumab (Opdivo), Atezolizumab (Tecentriq), Durvalumab (Imfinzi), Ipilimumab (Yervoy), CAR T-CellAttorney Docket No. : GH0224WOTherapies, Tisagenlecleucel (Kymriah), Axicabtagene ciloleucel (Yescarta), and Brexucabtagene autoleucel (Tecartus). In some embodiments, the one or more cytokines are selected from the group consisting of: Interleukin-2 (IL-2), Interferon-alpha (IFN-a), Oncolytic Virus Therapy, and Talimogene laherparepvec (T-VEC, Imlygic). In some embodiments, the one or more monoclonal antibodies are selected from the group consisting of: Rituximab (Rituxan), Trastuzumab (Herceptin), and Bevacizumab (Avastin). In some embodiments, the one or more immune modulators are selected from the group consisting of: Thalidomide (Thalomid), Lenalidomide (Revlimid), and Pomalidomide (Pomalyst).
[0012] In some embodiments, the tumor-specific neoantigens corresponds to tumor-associated mutations in the subject. In some embodiments, the mutations are selected from the group consisting of: point mutation, insertions, deletions, frameshift mutation, duplication, inversion, translocation, copy number variation (CNV), repeat expansion, substitution, gene fusion, chromosome fusion, amplification, loss of heterozygosity, structural variations, large-scale deletions, large-scale duplications, and complex rearrangements.
[0013] In additional embodiments, the differentially methylated regions(DMRs) comprise promoter regions. In some embodiments, the DMRs comprise promoter regions, intronic regions, coding regions. In some embodiments, the DMRs are selected based at least on one functional element. In some embodiments, the functional element comprises transcription factor binding sites (TFBS), on or more overlapping variants, one or more histone modifications, and or one or more chromatin interaction sites.
[0014] Moreover, the biomarkers may include clinically actionable genomic loci. In additional embodiments, the biomarkers further comprise resistance sequence variants and / or epigenetic variants emerging under therapeutic pressure.
[0015] Furthermore, the method may include (a) generating a real-time dynamic monitoring report of the patient’s immune response to the PCV based on the therapeutic response assessment, and (b) adjusting the treatment regimen for the patient based on the information provided in the report.
[0016] In some embodiments, the PCV is designed based on whole exome alone or in combination with transcriptome sequencing data from the tumor.
[0017] In some embodiments the PCV is administered in combination with one or more plasmid-encoded IL-12.
[0018] In yet additional embodiments the PCV is administered in combination with pembrolizumab. In additional embodiments, the neoantigens comprise tumor-specific genetic mutations that result in the production of altered proteins.Attorney Docket No. : GH0224WO
[0019] Moreover, the method may include determining neoantigen levels and / or a plurality of DMRs to predict clinical outcomes in the subject undergoing treatment with the PCV. In some embodiments, the DMRs are selected based on whole epigenome, whole genome, and / or whole transcriptome sequencing of the tumor and the normal tissue samples from the patient, the DMRs overlapping at least one functional element in the whole genome and / or whole transcriptome data. In some embodiments, the DMRs are selected based on targeted epigenome, targeted genome, and / or targeted transcriptome sequencing of the tumor and the normal tissue samples from the patient, the DMRs overlapping at least one functional element in the targeted genome and / or targeted transcriptome data.
[0020] In some embodiments, the neoantigens are selected based on whole genome and / or whole transcriptome sequencing of the tumor and the normal tissue samples from the patient, the neoantigens overlapping at least one functional element in the whole genome and / or whole transcriptome data. In additional embodiments, the neoantigens are selected based on targeted- genome and / or targeted-transcriptome sequencing of the tumor and the normal tissue samples from the patient, the neoantigens overlapping at least one functional element in the targeted genome and / or targeted transcriptome data.
[0021] In additional embodiments, the PCV selected from the group consisting of: neoantigen vaccines, dendritic cell vaccines, oncolytic virus vaccines, peptide vaccines, and DNA / RNA vaccines.
[0022] In some embodiments, the plurality of molecules extracted from the tissue sample and / or the biomarkers of interest are selected from the group consisting of: DNA, RNA, proteins, DNA methylation, histone modifications, epigenetic marks indicative of chromatin accessibility, non-coding RNAs, DNA sequences encoding transcription factor binding sites (TFBS) and / or bound by transcription factors, chromatin remodeling complexes, crossed linked-molecules or interacting molecules indicative of 3D chromatin architecture, histone protein and / or histone tail variants, DNA sequences encoding nucleosome positioning information, and DNA sequences comprising fragmentomic patters. The fragmentomic patterns may comprise fragment end density, fragment end motif, fragment length distribution, fragment size ratio, jagged ends, coverage oscillation, nucleosome positioning, mono- and dinucleosome ratios, and GC-content bias.
[0023] In some embodiments, the plurality of molecules extracted from the tissue sample is DNA.
[0024] In some embodiments, the sequencing comprises Illumina Sequencing, Nanopore Sequencing, Pacbio, Ion torrent, Sanger Sequencing, lOx genomics, QIAGEN, OxfordAttorney Docket No. : GH0224WONanopore, Complete Genomics, Ultima Genomics, Element Biosciences, or Singular Genomics. In additional embodiments, the sequencing is selected from the group consisting of:, RNA-seq, bisulfite sequencing, ATAC-seq, ChlP-seq, Hi-C, CUT&RUN, Cut&Tag, direct RNA sequencing, Methylated DNA Immunoprecipitation Sequencing (MeDIP-Seq), Methylation Cap Analysis (MethylCap-Seq), Methylation Enriched Sequencing (Methyl-Seq), Methylation-specific Enrichment Sequencing (Methyl-Seq), 5-hmC Enrichment Sequencing (5-hmC-Seq), hybrid capture sequencing, targeted enrichment sequencing, long-read sequencing, single-molecule real-time (SMRT) sequencing, exome sequencing, whole-genome sequencing (WGS), transcriptome sequencing, amplicon sequencing, low-pass sequencing, paired-end sequencing, and / or multiplex sequencing.
[0025] In additional aspects, the disclosure provides a computer-implemented method configured to use a previously generated computational model, trained to determine cancer recurrence in a subject treated with a personalized cancer vaccine (PCV), comprising, (a) extracting a plurality of molecules from a tumor and match normal samples obtained from the subject prior to initiation of treatment with the PCV and sequencing a portion of the plurality of molecules to acquire tumor-specific neoantigen and differentially methylated regions (DMRs) data; (b) extracting a plurality of cell-free DNA(cfDNA) from a blood sample obtained from the same subject, wherein the blood sample is collected from the subject at a later timepoint following treatment with the PCV; (c) sequencing a portion of the cfDNA molecules, to acquire neoantigens and DMRs data and selecting the neoantigens and DMRs based on the sequencing data obtained from the tumor and match normal tissue samples in (a) to obtain a plurality of features; (d) inputting the plurality of features to the previously trained computational model; and (e) outputting from the computational model a determination of whether or not the cancer reoccurred in the subj ect treated with the PCV.
[0026] In some embodiments, the tumor and match normal samples obtained from the subject are genetically assayed to further obtain tumor specific somatic genetic variation. In additional embodiments, the input features for the computational model comprise at least a portion of the somatic genetic variation obtained from the tumor. In some embodiments, the tumor and match normal samples obtained from the subject epigenetically assayed to further obtain tumor specific epigenetic patterns. In some embodiments, the input features for the computational model comprise at least a portion of the tumor specific epigenetic patterns.
[0027] In some embodiments, the tumor specific epigenetic patterns are selected from the group consisting of: cytosine methylation, transcription factor binding sites (TFBS), fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylationAttorney Docket No. : GH0224WO or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, H3S10ph.
[0028] In additional embodiments, prior to sequencing, the cfDNA molecules are amplified and subsequently selectively enriched to select target regions. In some embodiments, prior to sequencing, the cfDNA molecules are selectively enriched to select target regions and the target regions are subsequently amplified.
[0029] In some embodiments, the target regions comprise a panel including transcription factors and / or transcription factor binding sites(TFBS). In some embodiments, the target regions comprise a panel including DMRs in cancer. In some embodiments, the target regions comprise a panel including clinically actionable genetic and epigenetic variants. In some embodiments, the target regions comprise a panel comprising cancer pathways. In additional embodiments, the target regions comprise resistance loci to detect variants emerging under therapeutic pressure. In some embodiments, the target regions comprise proteomic signatures of disease, treatment response, and or resistance to treatment. In some embodiments, the target regions comprise a plurality of panels including transcription factors, TFBS, DMRs in cancer, clinically actionable genetic and epigenetic variants, genes with a function in cancer pathways, resistance loci, and / or proteomic signatures of disease, treatment response, or resistance to treatment. In some embodiments, the method further comprises partitioning the sample to selectively enrich each panel from the different partitions.
[0030] In additional embodiments, the method further comprises performing a retrospective biomarker analysis to correlate biomarker changes with radiographic imaging, clinical data, and / or histology findings, to validate the method’s efficacy in detecting MRD and / or monitoring disease progression and treatment response.
[0031] In some embodiments, the cancer is selected from the group consisting of adenocarcinoma, basal cell carcinoma, bladder cancer, bone cancer, brain cancer, breast cancer, cervical cancer, cholangiocarcinoma, colorectal cancer, endometrial cancer, esophageal cancer, gallbladder cancer, gastric cancer, germ cell tumors, glioma, head and neck cancer, hepatocellular carcinoma, Kaposi sarcoma, kidney cancer, lip and oral cavity cancer, liver cancer, lung cancer, melanoma, mesothelioma, neuroendocrine tumors, ovarian cancer, pancreatic cancer, penile cancer, prostate cancer, sarcoma, skin cancer, small cell lung cancer, squamous cell carcinoma, stomach cancer, testicular cancer, thyroid cancer, uterine cancer, vaginal cancer, vulvar cancer.Attorney Docket No. : GH0224WO
[0032] In one aspect, the disclosure provides a method for monitoring therapeutic response in a subject treated with a personalized cancer vaccine (PCV), comprising: (a) obtaining a tumor tissue sample from the subject prior to initiation of treatment with the PCV, extracting a plurality of molecules from the tumor tissue sample, and sequencing the plurality of molecules to identify one or more tumor-derived biomarkers; (b) determining from the tumor-derived biomarkers obtained in (a) a panel of biomarkers characteristic of the tumor sample; and (c) obtaining a blood sample from the subject after initiation of treatment with the PCV, and detecting the selected biomarkers in the blood sample to monitor therapeutic response, wherein a decrease in biomarkers levels indicates a response to treatment, and an increase or unchanged level indicates resistance to treatment and / or disease progression, thereby monitoring therapeutic response in the subject treated with the PCV.
[0033] In another aspect, the disclosure provides a method for monitoring therapeutic response in a subject treated with a PCV, comprising: (a) obtaining a first blood sample from the subject prior to initiation of treatment with the PCV, extracting a plurality of molecules from the first blood sample, and sequencing the plurality of molecules to determine one or more biomarkers characteristic of a tumor in the subject; and (b) obtaining a second blood sample from the subject after initiation of treatment with the PCV, and determining in the second sample the biomarkers determined in (a) in the first sample to determine changes or levels of the biomarkers; and (c) applying a trained classifier to the biomarker data obtained from the second sample to classify the subject as responsive or resistant to treatment, thereby monitoring therapeutic response in the subject treated with the PCV.
[0034] In yet another aspect, the disclosure provides a method for monitoring therapeutic response in a subject treated with PCV, the PCV comprising a panel of neoantigens, the method comprising: (a) obtaining a blood sample from the subject after initiation of treatment with the PCV, and determining the panel of neoantigens comprising the PCV to monitor therapeutic response in the subject, wherein a decrease in neoantigen levels indicates a response to treatment, and an increase or unchanged level indicates resistance to treatment and / or disease progression, thereby monitoring therapeutic response in the subject treated with the PCV.
[0035] In some embodiments, the detecting comprises direct detection of the neoantigens using immunoassays. In other embodiments, the immunoassays comprise ELISA, western blot, lateral flow immunoassay (LFIA), chemiluminescent immunoassay (CLIA), radioimmunoassay (RIA), fluorescent immunoassay (FIA), multiplex bead-based immunoassay, immunohistochemistry (IHC). In yet other embodiments, the detectingAttorney Docket No. : GH0224WO comprises indirect detection of the neoantigens through detection of an epigenetic signature characteristic of the neoantigen levels in the subject.
[0036] In some embodiments, the epigenetic signature comprises a methylation signature. In additional embodiments, the detecting comprises detecting a metabolomic signature characteristic of the neoantigen levels in the subject. In some embodiments, the detecting comprises detecting a transcriptomic signature characteristic of the neoantigen levels in the subject.
[0037] In an additional aspect, the disclosure provides, a method for monitoring therapeutic response in a subject treated with a PCV, the PCV comprising a panel of neoantigens, the method comprising: obtaining a tissue sample from the subject after initiation of treatment with the PCV, and detecting the panel of neoantigens comprising the PCV to monitor therapeutic response in the subject, wherein a decrease in neoantigen levels indicates a response to treatment, and an increase or unchanged level indicates resistance to treatment and / or disease progression, thereby monitoring therapeutic response in the subject treated with the PCV.
[0038] In some embodiments, the detecting comprises detecting a histological signature characteristic of the neoantigen levels in the subject. In other embodiments, the detecting comprises detecting a radiomic signature characteristic of the neoantigen levels in the subject. In yet other embodiments, the radiomic signature comprises tissue volume, irregularity, entropy, contrast, mean pixel values, histogram features, gray-level co-occurrence matrix. In additional embodiments, the detecting comprises detecting a radiogenomic signature characteristic of the neoantigen levels in the subject.
[0039] In another aspect, the disclosure provides a method for monitoring therapeutic response in a subject treated with a PCV, the PCV comprising a panel of neoantigens, the method comprising: (a) obtaining a blood sample from the subject after initiation of treatment with the PCV, (b) determining an epigenetic profile from the blood sample, the epigenetic profile comprising methylation data, (c) applying a trained classifier to the epigenetic data to detect an epigenetic signature indicative of neoantigen levels in the subject after treatment with the PCV, wherein detection of the epigenetic signature classifies the subject as responsive or resistant to treatment, thereby monitoring therapeutic response in the subject treated with the PCV.
[0040] In certain aspects, the disclosure provides a method for monitoring therapeutic response in a subject treated with a PCV, the PCV comprising a panel of neoantigens, the method comprising: (a) obtaining a blood sample from the subject after initiation of treatment with the PCV, (b) determining a neoantigen profile from the blood sample, the neoantigen profile comprising post treatment levels the neoantigen panel in the PCV administered to the subject;Attorney Docket No. : GH0224WO(c) applying a trained classifier to the neoantigen profile to detect a neoantigen signature indicative of neoantigen levels, wherein detection of the neoantigen signature enables classification of the subject as responsive or resistant to treatment, thereby monitoring therapeutic response in the subject treated with the PCV.
[0041] In some embodiments, the neoantigen signature further comprise genetic and / or epigenetic lesions in the tumor from the subject. In other embodiments, determining the neoantigen profile comprises extracting a plurality of molecules from the blood sample to obtain genetic, epigenetic, and / or transcriptomic data to determine sequences encoding the neoantigens.
[0042] Furthermore, in a group of embodiments, in response to determining resistance to treatment and / or disease progression in the subject, administering a second dose of the PCV to the subject.
[0043] In an additional group of embodiments, in response to determining resistance to treatment in the subject, (a) obtaining a second and / or third blood sample from the subject; (b) extracting a plurality of molecules from the sample to determine a new set of neoantigens in the sample obtained from the subject to determine an updated PCV for the subject, and (c) administering the updated PCV to the subject.
[0044] In another aspect, the disclosure provides a method for monitoring therapeutic response in a subject treated with an off-the-shelf cancer vaccine, the method comprising: (a) obtaining a sample from the subject after initiation of treatment with the off-the-shelf cancer vaccine; (b) extracting a plurality of molecules from the sample and sequencing the molecules to obtain sequencing reads; (c) determining one or more features in the sequencing reads; and (d) applying a classifier trained to recognize the features determined in (c) to classify the sample obtained from the subject as responsive or resistant to the off-the-shelf cancer vaccine, thereby monitoring therapeutic response in the subject treated with the off-the-shelf cancer vaccine.
[0045] In some embodiments, the one or more features comprise methylation signals indicative of immune activation. In other embodiments, the sample comprises blood. In yet other embodiments, the one or more features comprise methylation signals indicative of treatment response and / or resistance. In additional embodiments, methylation signals comprise differentially methylated regions (DMRs) in immune-regulatory or cancer-associated loci.
[0046] In some embodiments, the off-the-shelf cancer vaccine is selected from the group consisting of Bacillus Calmette-Guerin (BCG), Sipuleucel-T (Provenge), Talimogene laherparepvec (T-VEC), CIMAvax-EGF, GV AX, and New York-ESO-1 vaccine.Attorney Docket No. : GH0224WO
[0047] In yet other embodiments, the classifier is trained using a supervised learning algorithm based on labeled datasets from vaccinated and non -vaccinated subjects.
[0048] In another aspect, the disclosure provides a method for monitoring therapeutic response in a subject treated with a cancer vaccine, the method comprising: (a) obtaining a sample from the subject after initiation of treatment with the cancer vaccine; (b) extracting a plurality of molecules from the sample and sequencing the molecules to obtain sequencing reads; (c) determining one or more features in the sequencing reads; and (d) applying a classifier trained to recognize the features determined in (c) to classify the sample obtained from the subject as responsive or resistant to the cancer vaccine, thereby monitoring therapeutic response in the subject treated with the cancer vaccine.
[0049] The various steps of the methods disclosed herein, or steps carried out by the systems disclosed herein, may be carried out at the same or different times, in the same or different geographical locations, e.g., countries, and / or by the same or different people. In some embodiments, the report is communicated to a subject, for example, a subject who has cancer and has undergone testing and / or treatment by the methods and systems described herein, or to a healthcare professional, such as a physician treating the subject that has cancer and / or is undergoing treatment with a PC V.
[0050] This summary is provided for purposes of illustrating some exemplary embodiments, so as to provide a basic understanding of some aspects of the subject matter described herein. Accordingly, it will be appreciated that the above-described features are examples and should not be construed to narrow the scope or spirit of the subject matter described herein in any way. Other features, aspects, and advantages of the subject matter described herein will become apparent from the following Detailed Description, Figures, and Claims.BRIEF DESCRIPTON OF THE DRAWINGS
[0051] Figure 1. The figure shows an exemplary workflow for testing a subject undergoing treatment with a personalized cancer vaccine (PCV). A tumor sample is collected from the subject, and a personalized therapeutic cancer vaccine is designed based on whole exome / transcriptome tumor sequencing data. The personalized vaccine is then synthesized and administered to the subject. The figure also shows the subsequent steps of immune response monitoring, including the use a tumor-informed ctDNA assay to (1) monitor efficacy by tracking tumor neoantigen target levels (2) monitor molecular residual disease (MRD) by tracking variants and methylation biomarkers.Attorney Docket No. : GH0224WODETAILED DESCRIPTION
[0052] The methods described herein are directed to the integration of a plurality of sequencing datasets comprising a plurality of genetic and epigenetic states to monitor neoantigen targets and / or MRD in patients undergoing treatment with a PC V. Data integration is achieved through a variety of artificial intelligence (Ai), machine learning, and deep learning methods. These methods can elucidate the genetic and epigenetic changes that occur as a result of treatment, response to treatment, resistance to treatment, and in MRD.
[0053] In certain embodiments, machine learning or deep learning models may be trained using high-dimensional molecular data derived from samples obtained from subjects treated with a cancer vaccine. The training data may include a combination of tumor-specific mutations, differentially methylated regions, fragmentomic features including fragment end density, size distribution, and / or end motifs, transcriptomic profiles, epigenetic modifications including histone marks, chromatin accessibility, and / or TFBS, and neoantigen abundance or dynamics. Each sample may be annotated with clinical outcome data such as therapeutic response, recurrence, or MRD status. Classifiers are trained to learn associations between these molecular features and clinically meaningful endpoints. This enables the trained models to recognize molecular signatures associated with MRD and / or therapeutic response to cancer vaccines, thereby supporting real-time disease monitoring. Additionally, the models facilitate stratification of subjects who may require additional vaccine doses or who did not respond to treatment and may benefit from a modified cancer vaccine composition tailored to emerging biomarkers or treatment-resistant features. In some embodiments, data are collected at multiple timepoints from the same subject to support longitudinal model training and vaccine adaptation strategies.
[0054] In some embodiments, of the disclosure epigenetic data used to detect MRD and / or therapeutic response in subjects treated with a PCV can comprise methylation data, histone modification data, chromatin conformation capture data, nucleosome positioning, histone variants for example, replacement of the canonical histone H2A with the variant H2A.Z, RNA methylation, chromatin accessibility, DNA hydromethylation, DNA phosphorylation, acetylation, transcription factor binding sites, and / or chromatin looping or DNA-DNA interactions data. In some embodiments, of the disclosure the TFBS are directly assayed using methods such as ChlP-seq or CUT&Tag, in alternative embodiments the TFBS are inferred based on epigenetic data sets and / or machine learning algorithms. These datasets may be analyzed individually or in combination to detect MRD and / or therapeutic response in subjectsAttorney Docket No. : GH0224WO treated with a cancer vaccine. The cancer vaccine may comprise a PCV or an off-the-shelf cancer vaccine.
[0055] Thus, the disclosed methods provide a comprehensive molecular view of treatment- induced changes across the tumor and the corresponding changes that occur in the immune system in response to disease, resistance, MRD, and / or response to treatment with the cancer vaccine. Additionally, the disclosure provides methods to screen, diagnose, stage, and / or subtype a disease by monitoring changes in disease specific genetic and / or epigenetic biomarkers.
[0056] Reference will now be made in detail to certain embodiments of the invention. While the invention will be described in conjunction with such embodiments, it will be understood that they are not intended to limit the invention to those embodiments. On the contrary, the invention is intended to cover all alternatives, modifications, and equivalents, which may be included within the invention as defined by the appended claims.
[0057] Before describing the present teachings in detail, it is to be understood that the disclosure is not limited to specific compositions or process steps, as such may vary. It should be noted that, as used in this specification and the appended claims, the singular form “a”, “an” and “the” include plural references unless the context clearly dictates otherwise. Thus, for example, reference to “a nucleic acid” includes a plurality of nucleic acids, reference to “a cell” includes a plurality of cells, and the like.
[0058] Numeric ranges are inclusive of the numbers defining the range. Measured and measurable values are understood to be approximate, taking into account significant digits and the error associated with the measurement. Also, the use of “comprise”, “comprises”, “comprising”, “contain”, “contains”, “containing”, “include”, “includes”, and “including” are not intended to be limiting. It is to be understood that both the foregoing general description and detailed description are exemplary and explanatory only and are not restrictive of the teachings.
[0059] Unless specifically noted in the above specification, embodiments in the specification that recite “comprising” various components are also contemplated as “consisting of’ or “consisting essentially of’ the recited components; embodiments in the specification that recite “consisting of’ various components are also contemplated as “comprising” or “consisting essentially of’ the recited components; and embodiments in the specification that recite “consisting essentially of’ various components are also contemplated as “consisting of’ or “comprising” the recited components (this interchangeability does not apply to the use of these terms in the claims).Attorney Docket No. : GH0224WO
[0060] The section headings used herein are for organizational purposes and are not to be construed as limiting the disclosed subject matter in any way. In the event that any document or other material incorporated by reference contradicts any explicit content of this specification, including definitions, this specification controls.
[0061] Definitions
[0062] “Cell-free DNA,” “cfDNA molecules,” or simply “cfDNA” include DNA molecules that naturally occur in a subject in extracellular form (e.g., in blood, serum, plasma, or other bodily fluids such as lymph, cerebrospinal fluid, urine, or sputum). While the cfDNA originally existed in a cell or cells in a large complex biological organism, e.g., a mammal, it has undergone release from the cell(s) into a fluid found in the organism, and may be obtained by obtaining a sample of the fluid without the need to perform an in vitro cell lysis step.
[0063] As used herein, a modification or other feature is present in “a greater proportion” in a first subsample or population of nucleic acid than in a second subsample or population when the fraction of nucleotides with the modification or other feature is higher in the first subsample or population than in the second population. For example, if in a first subsample, one tenth of the nucleotides are mC, and in a second subsample, one twentieth of the nucleotides are mC, then the first subsample comprises the cytosine modification of 5-methylation in a greater proportion than the second subsample.
[0064] As used herein, “machine learning model” (or “model”) refers to a collection of parameters and functions, where the parameters are trained on a set of training samples or individual data points or instances used to train a machine learning model. These samples are part of the dataset that provides the model with examples of input data along with the corresponding output (for supervised learning), or just input data (for unsupervised learning). The parameters and functions may be a collection of linear algebra operations, non-linear algebra operations, and tensor algebra operations. The parameters and functions may include statistical functions, tests, and probability models. The training samples can correspond to samples having measured properties of the sample (e.g., genomic, epigenomic, transcriptomic, metabolites etc. data and other subject data, such as histology, imaging data and / or electronic medical health records, or insurance claim data), as well as known patient / sample metadata including classifications or labels for example molecular phenotypes or specific cancer or disease therapies. Other phenotypes can include patient biomedical information including “cardiovascular phenotypes” or “cardiovascular risk factors” such as weight, height, Body Mass Index(BMI) , and other physical characteristics. Yet other phenotypes can include cancer risk factors including smoking, excessive alcohol consumption, poor diet, physical inactivity,Attorney Docket No. : GH0224WO obesity, genetic predispositions, exposure to harmful chemicals and radiation, chronic inflammation, certain infections (such as human papillomavirus, hepatitis B and C), hormonal imbalances, and advanced age. The model can learn from the training samples in a training process that optimizes the parameters (and potentially the functions) to provide an optimal quality metric (e.g., accuracy) for classifying new samples. A variety of advanced statistical and computational methods that can be employed as training functions including Expectation Maximization(EM) to find maximum likelihood estimates of parameters in probabilistic models, especially for models with latent variables, Maximum Likelihood Estimation (MLE) to estimate the parameters of a statistical model. MLE methods select the set of parameters that maximize the likelihood function i.e., the parameters under which the observed data is most probable. Bayesian Parameter Estimation Methods which incorporate prior knowledge in addition to the data at hand through the use of probability distributions. These include Markov Chain Monte Carlo (MCMC), Gibbs Sampling, Hamiltonian Monte Carlo (HMC), and Variational Inference (VI), or Gradient-Based Methods including Stochastic Gradient Descent (SGD) and the Broyden-Fletcher-Goldfarb-Shanno (BFGS) algorithm.
[0065] As used herein, “Deep learning models” refers to a collection of “architectures” or “algorithms” useful in scenarios where machine learning approaches may fall short, for example due to the complexity of the data (e.g., high dimensional genomics, epigenomics datasets used alone or in one or more combinations). Additional high dimensional data can include medical images such as MRIs, or histology reports. Deep learning models, especially CNNs, can automatically extract relevant features without manual feature engineering.
[0066] Exemplary applications for deep learning models include complex pattern recognition tasks for example, recognizing specific functional elements in genomics data, functional elements can include transcription factor binding sites (TFBS) or chromatin interaction sites for example promoter-enhancer interactions. Additionally, recognition of disease specific features in histology images including patters or specific features of PDL-1 expression.
[0067] As used herein, “without substantially altering base-pairing specificity” of a given nucleobase means that a majority of molecules comprising that nucleobase that can be sequenced do not have alterations of the base pairing specificity of the second nucleobase relative to its base pairing specificity as it was in the originally isolated sample. In some embodiments, 75%, 90%, 95%, or 99% of molecules comprising that nucleobase that can be sequenced do not have alterations of the base pairing specificity of the second nucleobase relative to its base pairing specificity as it was in the originally isolated sample.Attorney Docket No. : GH0224WO
[0068] As used herein, “base pairing specificity” refers to the standard DNA base (A, C, G, or T) for which a given base most preferentially pairs. Thus, for example, unmodified cytosine and 5-methylcytosine have the same base pairing specificity (i.e., specificity for G) whereas uracil and cytosine have different base pairing specificity because uracil has base pairing specificity for A while cytosine has base pairing specificity for G. The ability of uracil to form a wobble pair with G is irrelevant because uracil nonetheless most preferentially pairs with A among the four standard DNA bases.
[0069] As used herein, a “combination” comprising a plurality of members refers to either of a single composition comprising the members or a set of compositions in proximity, e.g., in separate containers or compartments within a larger container, such as a multiwell plate, tube rack, refrigerator, freezer, incubator, water bath, ice bucket, machine, or other form of storage.
[0070] The “capture yield” of a collection of probes for a given target set refers to the amount (e.g., amount relative to another target set or an absolute amount) of nucleic acid corresponding to the target set that the collection of probes captures under typical conditions. Exemplary typical capture conditions are an incubation of the sample nucleic acid and probes at 65°C for 10-18 hours in a small reaction volume (about 20 pL) containing stringent hybridization buffer. The capture yield may be expressed in absolute terms or, for a plurality of collections of probes, relative terms. When capture yields for a plurality of sets of target regions are compared, they are normalized for the footprint size of the target region set (e.g., on a per-kilobase basis). Thus, for example, if the footprint sizes of first and second target regions are 50 kb and 500 kb, respectively (giving a normalization factor of 0.1), then the DNA corresponding to the first target region set is captured with a higher yield than DNA corresponding to the second target region set when the mass per volume concentration of the captured DNA corresponding to the first target region set is more than 0.1 times the mass per volume concentration of the captured DNA corresponding to the second target region set. As a further example, using the same footprint sizes, if the captured DNA corresponding to the first target region set has a mass per volume concentration of 0.2 times the mass per volume concentration of the captured DNA corresponding to the second target region set, then the DNA corresponding to the first target region set was captured with a two-fold greater capture yield than the DNA corresponding to the second target region set.
[0071] “Capturing” one or more target nucleic acids refers to preferentially isolating or separating the one or more target nucleic acids from non-target nucleic acids.
[0072] A “captured set” of nucleic acids refers to nucleic acids that have undergone capture.Attorney Docket No. : GH0224WO
[0073] A “target-region set” or “set of target regions” refers to a plurality of genomic loci targeted for capture and / or targeted by a set of probes (e.g., through sequence compl ementarity ) .
[0074] “Corresponding to a target region set” means that a nucleic acid, such as cfDNA, originated from a locus in the target region set or specifically binds one or more probes for the target-region set.
[0075] “Specifically binds” in the context of an probe or other oligonucleotide and a target sequence means that under appropriate hybridization conditions, the oligonucleotide or probe hybridizes to its target sequence, or replicates thereof, to form a stable probe:target hybrid, while at the same time formation of stable probemon-target hybrids is minimized. Thus, a probe hybridizes to a target sequence or replicate thereof to a sufficiently greater extent than to a nontarget sequence, to enable capture or detection of the target sequence. Appropriate hybridization conditions are well-known in the art, may be predicted based on sequence composition, or can be determined by using routine testing methods (see, e.g., Sambrook et al., Molecular Cloning, A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989) at §§ 1.90-1.91, 7.37-7.57, 9.47-9.51 and 11.47-11.57, particularly §§ 9.50-9.51, 11.12-11.13, 11.45-11.47 and 11.55-11.57, incorporated by reference herein).
[0076] “Sequence-variable target region set” refers to a set of target regions that may exhibit changes in sequence such as nucleotide substitutions (i.e., single nucleotide variations), insertions, deletions, or gene fusions or transpositions in neoplastic cells (e.g., tumor cells and cancer cells).
[0077] “Epigenetic target region set” refers to a set of target regions that may show sequenceindependent changes in neoplastic cells (e.g., tumor cells and cancer cells) or that may show sequence-independent changes in cfDNA from subjects having cancer relative to cfDNA from healthy subjects. Examples of sequence-independent changes include, but not limited to, changes in methylation (increases or decreases), nucleosome distribution, CTCF binding, transcription start sites, and regulatory protein binding regions. For present purposes, loci susceptible to neoplasia-, tumor-, or cancer-associated focal amplifications and / or gene fusions may also be included in an epigenetic target region set because detection of a change in copy number by sequencing or a fused sequence that maps to more than one locus in a reference genome tends to be more similar to detection of exemplary epigenetic changes discussed above than detection of nucleotide substitutions, insertions, or deletions, e.g., in that the focal amplifications and / or gene fusions can be detected at a relatively shallow depth of sequencing because their detection does not depend on the accuracy of base calls at one or a few individualAttorney Docket No. : GH0224WO positions. In some embodiments, the epigenetic target region set includes one or more genomic regions, where the epigenetic state (e.g., methylation state) of cfDNA molecules in these regions is unchanged in cancer, but their presence / quantity in blood indicates increased, aberrant presentation of cfDNA from certain tissue (e.g. cancer origin) into circulation.
[0078] A nucleic acid is “produced by a tumor” or ctDNA or circulating tumor DNA, if it originated from a tumor cell. Tumor cells are neoplastic cells that originated from a tumor, regardless of whether they remain in the tumor or become separated from the tumor (as in the cases, e.g., of metastatic cancer cells and circulating tumor cells).
[0079] The term “methylation” or “DNA methylation” refers to addition of a methyl group to a nucleotide base in a nucleic acid molecule. In some embodiments, methylation refers to addition of a methyl group to a cytosine at a CpG site (cytosine-phosphate-guanine site (i.e., a cytosine followed by a guanine in a 5’3’ direction of the nucleic acid sequence). In some embodiments, DNA methylation refers to addition of a methyl group to adenine, such as in N6- methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the 5th carbon of cytosine). In some embodiments, 5-methylation refers to addition of a methyl group to the 5C position of the cytosine to create 5-methylcytosine (5mC). In some embodiments, methylation comprises a derivative of 5mC. Derivatives of 5mC include, but are not limited to, 5 -hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5- caryboxylcytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the 3rd carbon of cytosine). In some embodiments, 3C methylation comprises addition of a methyl group to the 3C position of the cytosine to generate 3 -methylcytosine (3mC). Methylation can also occur at non CpG sites, for example, methylation can occur at a CpA, CpT, or CpC site. DNA methylation can change the activity of methylated DNA region. For example, when DNA in a promoter region is methylated, transcription of the gene may be repressed. DNA methylation is critical for normal development and abnormality in methylation may disrupt epigenetic regulation. The disruption, e.g., repression, in epigenetic regulation may cause diseases, such as cancer. Promoter methylation in DNA may be indicative of cancer.
[0080] The term “hypermethylation” refers to an increased level or degree of methylation of nucleic acid molecule(s) relative to the other nucleic acid molecules within a population (e.g., sample) of nucleic acid molecules. In some embodiments, hypermethylated DNA can include DNA molecules comprising at least 1 methylated residue, at least 2 methylated residues, at least 3 methylated residues, at least 5 methylated residues, or at least 10 methylated residues.
[0081] The term “hypomethylation” refers to a decreased level or degree of methylation of nucleic acid molecule(s) relative to the other nucleic acid molecules within a population (e.g.,Attorney Docket No. : GH0224WO sample) of nucleic acid molecules. In some embodiments, hypomethylated DNA includes unmethylated DNA molecules. In some embodiments, hypomethylated DNA can include DNA molecules comprising 0 methylated residues, at most 1 methylated residue, at most 2 methylated residues, at most 3 methylated residues, at most 4 methylated residues, or at most 5 methylated residues.
[0082] The terms “or a combination thereof’ and “or combinations thereof’ as used herein refers to any and all permutations and combinations of the listed terms preceding the term. For example, “A, B, C, or combinations thereof’ is intended to include at least one of: A, B, C, AB, AC, BC, or ABC, and if order is important in a particular context, also BA, CA, CB, ACB, CBA, BCA, BAC, or CAB. Continuing with this example, expressly included are combinations that contain repeats of one or more item or term, such as BB, AAA, AAB, BBC, AAABCCCC, CBBAAA, CABABB, and so forth. The skilled artisan will understand that typically there is no limit on the number of items or terms in any combination, unless otherwise apparent from the context.
[0083] “ Or” is used in the inclusive sense, i.e., equivalent to “and / or,” unless the context requires otherwise.
[0084] Exemplary methods
[0085] In accordance with an embodiment of the present disclosure, Figure 1 is illustrative of an exemplary workflow for testing a subject undergoing treatment with a personalized cancer vaccine (PCV). A tumor sample is collected from the subject, and a personalized therapeutic cancer vaccine is designed based on whole exome and / or transcriptome tumor sequencing data. The personalized vaccine is then formulated and administered to the subject. The figure also shows the subsequent steps of immune response monitoring, including the use a tumor- informed ctDNA assay to (1) monitor efficacy by tracking tumor neoantigen target levels (2) monitor molecular residual disease (MRD) by tracking variants and / or methylation biomarkers. The methylation biomarkers may comprise single site methylation sites, hypermethylated and / or hypomethylated fragments. Additionally, the methylation biomarkers may comprise a group of methylation sites, a methylation pattern and / or a methylation signature.
[0086] Subjects
[0087] The methods described herein may be applied to samples from subjects undergoing treatment with a cancer vaccine comprising a PCV or an off-the-shelf cancer vaccine, either alone or in combination with other therapies. Samples obtained from subjects may include tumor tissue, blood, and / or matched normal tissue or fluid samples. Blood and other biological fluids including plasma, serum, urine, saliva, or cerebrospinal fluid may be obtained andAttorney Docket No. : GH0224WO analyzed according to the methods of the disclosure to identify biomarkers indicative of disease status, therapeutic response to the cancer vaccine, minimal residual disease (MRD), and / or resistance. Additionally, samples may be collected before or after treatment with the cancer vaccine. In some embodiments, samples may be collected at multiple timepoints to enable longitudinal monitoring. DNA may be obtained from such subjects to assess treatment-induced molecular changes, including methylation alterations, fragmentomic signatures, or neoantigen levels, as disclosed in the present disclosure. The subject may have an existing cancer diagnosis, be suspected of having cancer, be undergoing initial diagnosis or staging, or be in remission following prior therapy.
[0088] A matched normal sample which may comprise non-tumor tissue such as obtained from subjects undergoing treatment with a cancer vaccine, may be incorporated into the various workflows described herein, to support a range of applications. One key use is in distinguishing somatic variants from germline variants. In this case, match normal samples may be used to isolate and obtain a reference inherited genomic background. This enables identification of tumor-specific alterations by subtraction of germline variants that may otherwise be indistinguishable from somatic changes that occur in the tumor. Matched normal samples may also be used to detect and filter out variants arising from clonal hematopoiesis of indeterminate potential (CHIP). Subtraction of CHIP variants may be applied in any sample described herein, and may be of particular interest when using samples from older individuals including those over the age of 45 years old. Matched normal samples may further be used to detect technical artifacts, such as sequencing errors or formalin-induced changes, by providing a control for comparison. Furthermore, matched normal data may be used to train computational models of mutational signatures, structural variation, and background mutation rates. Matched normal tissue suitable for use in developing a cancer vaccine, or in assays used to monitor individuals receiving such vaccines, may include peripheral blood mononuclear cells (PBMCs), whole blood, saliva, skin biopsies, or normal tissue adjacent to the tumor site.
[0089] In additional embodiments of the disclosure, a matched normal sample may be excluded from the analysis for broader clinical and research utility. Excluding the need for a matched normal sample reduces assay turnaround time and lowers overall assay cost. In such cases, CHIP variants and other non-tumor-derived alterations may still be detected and filtered using computational approaches, including statistical methods, population-based variant filters, and variant classification models. These models may include statistical classifiers, machine learning algorithms, deep learning architectures, or artificial intelligence systems trained to distinguish somatic variants from germline variants and / or technical noise.Attorney Docket No. : GH0224WO
[0090] In some embodiments, the DNA (e.g., cfDNA or genomic DNA from tissue) is obtained from a subject having a cancer. In some embodiments, the DNA (e.g., cfDNA or genomic DNA from tissue) is obtained from a subject suspected of having a cancer. In some embodiments, the DNA (e.g., cfDNA or genomic DNA from tissue) is obtained from a subj ect having a tumor. In some embodiments, the DNA (e.g., cfDNA or genomic DNA from tissue) is obtained from a subject suspected of having a tumor. In some embodiments, the DNA (e.g., cfDNA or genomic DNA from tissue) is obtained from a subject having neoplasia. In some embodiments, the DNA (e.g., cfDNA, or genomic DNA from tissue) is obtained from a subject suspected of having neoplasia. In some embodiments, the DNA (e.g., cfDNA, or genomic DNA from tissue) is obtained from a subject in remission from a tumor, cancer, or neoplasia (e.g., following chemotherapy, surgical resection, radiation, or a combination thereof). In any of the foregoing embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia may be of the lung, colon, rectum, kidney, breast, prostate, or liver. In some embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia is of the lung. In some embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia is of the colon or rectum. In some embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia is of the breast. In some embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia is of the prostate. In any of the foregoing embodiments, the subject may be a human subject.
[0091] In certain embodiments, training samples are used to develop, validate, or optimize classifiers, deep learning and / or artificial intelligence algorithms applied to the biological data, including sequencing data obtained from the various samples from subjects treated with a cancer vaccine. Training samples may include biological specimens such as blood, plasma, serum, tumor tissue, or matched normal samples obtained from subjects with known clinical outcomes. These may include samples from subjects who responded to vaccine treatment, exhibited resistance, or experienced recurrence or progression following vaccination. Samples may be collected at defined clinical timepoints including pre-treatment baseline, on-treatment, and / or post-treatment, and may be processed to generate molecular profiles, including methylation data, fragmentomic patterns, neoantigen abundance, transcriptomic signatures, and protein expression.
[0092] In some embodiments, labels corresponding to clinical outcomes for example responder vs. non-responder, MRD-positive vs. MRD-negative are associated with the training samples used in supervised learning. The resulting models may be trained to recognize molecular features or combinations thereof that are predictive of therapeutic response to the cancerAttorney Docket No. : GH0224WO vaccine, immune activation, or disease progression. Training data may also include longitudinal samples from the same individual to account for temporal dynamics in biomarker patterns. Temporal molecular data obtained from the same subject across multiple timepoints may be used to inform the development of an updated or boosted cancer vaccine in response to observed treatment resistance or lack of clinical response. In additional embodiments, biomarkers such as persistent or re-emerging neoantigens, stable or increasing methylation patterns at DMRs, or evolving fragmentomic signatures may indicate vaccine escape or incomplete immune clearance. These data can guide the design of a modified vaccine formulation that incorporates additional or newly emerging neoantigens, adjusted antigen dosages, or altered delivery strategies to enhance therapeutic efficacy. This integration of realtime longitudinal data to upgrade or redesign cancer vaccines for the subject during treatment results in a personalized, adaptive vaccination approach, as described throughout the present disclosure.
[0093] Hi-C technique to capture chromatin conformation
[0094] Three-dimensional chromatin organization varies among different cell types and plays a crucial role in gene regulation by bringing distant functional elements into close spatial proximity. These functional elements contribute to maintaining homeostasis in health and play a pivotal role in disease regulation. The methods, systems, and methods described herein may be applied to detect and monitor therapeutic response in a subject treated with a PCV by assessing changes in chromatin architecture that correlate with treatment-induced modulation of tumor-derived biomarkers, such as neoantigens and changes in the epigenome. Epigenetic changes that may occur in response to PCV treatment and / or response as well as tumor presence and / or tumor evolution include methylation status at specific loci in the genome. Such loci may include promoter regions, functional elements such as enhancers, insulators, transcription factor binding sites and regulatory elements found across the genome. In certain embodiments, the chromatin landscape may be altered in response to vaccine-induced immune pressure, leading to modulation of chromatin interactions, enhancer-promoter looping, or chromatin accessibility at loci encoding neoantigens or immune response genes.
[0095] Such chromatin-level changes may be captured using chromatin conformation capture (3C) techniques including 4C, 5C, Hi-C, ChlA-PET, Capture-C, Capture Hi-C, HiChIP, or Micro-C, and may be integrated as part of the molecular signature used to classify the subject as responsive or resistant to PCV treatment. In some embodiments, classifiers trained on 3D chromatin features or alterations in topologically associating domains (TADs) or loopAttorney Docket No. : GH0224WO structures are applied to chromatin interaction data to infer epigenetic or structural biomarkers associated with treatment outcome.
[0096] The implementations of techniques, architectures, frameworks, systems, processes, and computer-readable instructions described herein are directed to the analysis of three- dimensional chromatin organization as captured by chromatin conformation capture (3C) techniques including, 4C (Circular Chromosome Conformation Capture), 5C (Chromosome Conformation Capture Carbon Copy), Hi-C, ChlA-PET (Chromatin Interaction Analysis by Paired-End Tag Sequencing), Capture-C, Capture Hi-C, HiChIP, Micro-C.
[0097] Following is an exemplary workflow that outlines the steps in preparing a Hi-C library, from cell culture to final library preparation and quality control.
[0098] In exemplary workflows for the Hi-C technique cells are cultured and chromatin is crosslinked e.g., fix cells in 1% formaldehyde, quench with glycine, harvest by centrifugation, and store at -80°C for future use. Crosslinking conditions are typically standardized to ensure consistency across experiments. Cells are then lysed, and chromatin is digested. For example, cells can be lysed with a Dounce homogenizer in the presence of cold hypotonic buffer supplemented with protease inhibitors and IGEPAL CA-630. Lysates are wash and resuspended in restriction enzyme buffer, chromatin is then solubilized with SDS and incubated at 65°C, quench SDS with Triton X-100, and digested with a restriction enzyme (e.g., Hindlll). Biotin marking of DNA ends and blunt end ligation: DNA ends are labeled with biotin-14- dCTP using the KI enow fragment of DNA polymerase I. Then, DNA fragments are ligated in a diluted condition at 16°C to favor intra-molecular ligation. DNA is then purified with the following steps, degradation of proteins with Proteinase K, extraction with phenol: chloroform, DNA precipitation, resuspension, and RNase A treatment to yield high-quality DNA. Following DNA purification several quality control steps can be performed to ensure the library meets quality metrics. For example, by profiling fragment size distribution and quantify DNA using Agilent Bioanalyzer or Agilent TapeStation systems. Biotin removal from un-ligated ends: T4 DNA polymerase is used to remove biotin-labeled ends that have not been ligated. DNA is then precipitated and washed to prepare for downstream applications. DNA fragmentation and size fractionation: DNA is fractionated based on size using the Covaris 8700 and AMPure XP beads to achieve the desired size distribution for sequencing. End repair and “A” tailing: DNA molecules are repaired for asymmetric breaks and prepare for Illumina adapter ligation by filling in overhangs, and adenylating the 3’ ends. Streptavidin pull-down of biotinylated Hi-C ligation products: Mix biotinylated Hi-C ligation products with streptavidin beads, wash, and prepare for Illumina adapter ligation. Paired-end adapter ligation and libraryAttorney Docket No. : GH0224WO amplification: Illumina paired-end adapters are then ligated while the DNA is bound to streptavidin beads, libraries are PCR amplified with minimal cycles to avoid PCR artifacts the resulting amplified DNA is purified. Final quality control and library quantification: Assess the quality of the final amplified library and quantify it, ideally using a Bioanalyzer, before sequencing.
[0099] Moreover, the chromatin conformation capture experiment (Hi-C, 3C, 4C, capture Hi- C etc.) might utilize chromatin fragmentation methods that do not depend on sequence specificity. Exemplary protocols include TopoLink™ (Catalog #: 21010) from Dovetail Genomics (part of Cantata Bio) . In certain embodiments, a targeted approach might be more desirable, for example, when the research question is focused on promoter-enhancer interactions. Exemplary protocols include a capture Hi-C protocol using Dovetail Genomics Dovetail® Targeted Enrichment Panels (Catalog # 25013). In some embodiments, experimental workflows can combine one or more experiments in a single workflow using a variety of methods. For example, the ChlP-seq and the Hi-C workflows can be combined in one workflow using Dovetail® HiChIP MNase Kit (Catalog #: 21007) see, for example, Yang, Jae-Hyun et al. “Loss of epigenetic information as a cause of mammalian aging.” Cell vol. 186,2 (2023): 305-326. e27. doi: 10.1016 / j. cell.2022.12.027, which is incorporated herein by reference. This approach can investigate chromatin interactions mediated by specific proteins of interest. In additional embodiments, the methods disclosed herein can be used alone or in combination with Dovetail® Micro-C Kit (Catalog # 21006) or similar methods, to generate uniform fragments that capture nucleosome positioning information, and maintain even coverage across the genome. This approaches obtain ultra-high-resolution topology mapping down to the mono-nucleosome level (150-200 bp conformation), see, for example, Bayanjargal, Ariunaa et al. “The DBD-a4 helix of EWS::FLI is required for GGAA microsatellite binding that underlies genome regulation in Ewing sarcoma.” bioRxiv : the preprint server for biology 2024.01.31.578127. 31 Jan. 2024, doi: 10.1101 / 2024.01.31.578127. Preprint which is incorporated herein by reference. Additionally, Hi-C experiments can improve and / or obtain haplotype phasing, genome assembly, and / or variant detection. Protocols and reagents to obtain such information include for example, Dovetail® Omni-C® Kit (Catalog # 21006). See, for example, Milevskiy, Michael J G et al. “Three-dimensional genome architecture coordinates key regulators of lineage specification in mammary epithelial cells.” Cell genomics vol. 3,11 100424. 16 Oct. 2023, doi: 10.1016 / j .xgen.2023.100424, which is incorporated herein by reference.
[0100] Exemplary workflows for the capture Hi-C techniqueAttorney Docket No. : GH0224WO
[0101] Cross-linking of Chromatin: Cells are treated with formaldehyde or a combination of cross-linkers to preserve physical interactions between chromosomal regions. Digestion of DNA: The cross-linked chromatin is then digested using a restriction enzyme. This step is crucial for creating ends that can be ligated later. Some protocols might use a combination of enzymes for more efficient digestion. Ligation under Dilute Conditions: The digested chromatin is diluted and ligated, allowing for the ligation of interacting DNA ends that are in close proximity due to chromatin folding. Purification and Shearing: Cross-links are reversed, and the DNA is purified. The DNA may then be sheared into smaller fragments to prepare for library preparation. Capture Step: Biotinylated probes, designed to hybridize to regions of interest, are used to selectively capture specific fragments from the ligated DNA pool. This step enhances the resolution and specificity of interactions being analyzed. Library Preparation: Captured DNA fragments are processed into a sequencing library, including end repair, A- tailing, adapter ligation, and enrichment of targeted fragments through biotin-streptavidin pulldown. Sequencing: The prepared library is sequenced using high-throughput sequencing technology. Capture-Hi-C data analysis involves processing sequencing data is to identify chromatin interactions, including mapping reads to a reference genome, filtering, and identifying significant interactions within the captured regions.
[0102] Bioinformatic pipeline for Hi-C data
[0103] Various bioinformatic pipelines exists for the analysis of Hi-C datasets. Exemplary pipelines include, snHiC (Gregoricchio & Zwart, 2023). Briefly the analysis steps include (1) Generation of Contact Matrices: snHiC facilitates the creation of contact matrices at multiple resolutions in a single run, streamlining the initial analysis of Hi-C data. (2) Aggregation of Individual Samples: The pipeline allows for the aggregation of individual samples into user- specified groups, enabling comparative analyses across different conditions or time points. (3)Detection of Chromatin Features: It includes steps for the detection of domains, compartments, loops, and stripes, which are critical structural features of the genome organization revealed by Hi-C data. (4)Differential Analysis: snHiC supports differential compartment and chromatin interaction analyses, allowing users to identify changes in genome organization under different experimental conditions. The snHiC workflow can be automated using snakemake, which, makes it less prone to errors and more reproducible. To setup a yaml- formatted file is available to build a compatible conda environment, simplifying the setup and ensuring that users have all the necessary software and dependencies.
[0104] Similar analysis of Hi-C and Capture Hi-C data can be accomplished using publically available algorithms including HiCUP: A pipeline for mapping and processing Hi-C and CHi-Attorney Docket No. : GH0224WOC data, removing artefacts and producing quality control reports. CHICAGO: Capture Hi-C Analysis of Genomic Organization. HiC-bench, a comprehensive and reproducible Hi-C data analysis platform. HiC-Pro optimized pipeline for processing Hi-C data from raw reads to normalized contact maps. ChiCMaxima, a pipeline for detection and visualization of chromatin looping in CHi-C, which also allows integrating information from biological replicates.
[0105] DNA methylation plays a crucial role in regulating gene expression and maintaining genome stability. Disruption of DNA methylation control mechanisms can lead to various diseases, including cancer. A body of evidence supports that cancer cells exhibit largely different DNA methylation patterns compared to normal cells. In general, cancer cells are characterized by genome-wide hypomethylation, as well as hypermethylation of CpG islands associated with tumor suppressor genes and developmental regulators. Hypermethylation in the promoter regions of tumor suppressor genes leads to their silencing, while hypomethylation in the promoter regions of oncogenes can activate them, both mechanisms play a significant role in the development of cancer.
[0106] DNA methylation sequencing workflows
[0107] In some embodiments of the present disclosure, methylation data is generated from biological samples including blood and / or tumor tissue, matched normal tissue, pre and / or posttreatment blood samples to monitor therapeutic response in a subject treated with a PCV. The methylation data may be derived from methods such as whole-genome bisulfite sequencing (WGBS), methyl-CpG-binding domain enrichment (MBD-seq), reduced representation bisulfite sequencing (RRBS), targeted bisulfite sequencing, methylation arrays (e.g., EPIC), methylated DNA immunoprecipitation sequencing (MeDIP-seq), enzymatic methylation detection (EM-seq), or direct detection via nanopore or PacBio long-read sequencing. These methods may be applied to detect tumor-specific or immune-responsive DMRs or to assess broader epigenetic changes modulated by neoantigen and / or broader systemic immune activity following PCV administration.
[0108] Methylation features used for response monitoring may include single-CpG methylation levels, methylation haplotypes, regional methylation density, entropy across CpG islands, partially methylated domains (PMDs), and changes in methylation at immune regulatory loci. In some embodiments, hypermethylated or hypom ethylated fragments are selectively enriched using MBD proteins, and abundance is quantified across predefined regions relevant to cancer and / or immune response. Classifiers trained on these methylation- derived features may be applied to determine whether the epigenetic profile reflects a therapeutic response, resistance to treatment, or MRD. The resulting data provide a nonAttorney Docket No. : GH0224WO invasive, genome-wide view of dynamic epigenetic changes in the subject during or after vaccination, enabling longitudinal monitoring of PCV efficacy.
[0109] In specific embodiments, panels comprising immune-related genes and / or regulatory regions may be used to capture methylation data at regions of interest. These panels may include promoter regions, enhancers, transcription factor binding sites, and gene bodies associated with immune activation, antigen presentation, T cell signaling, cytokine expression, and immune checkpoint regulation. Methylation profiling of these regions enables detection of immune system modulation following PCV administration. For example, hypomethylation at interferon response gene promoters or changes in methylation at loci encoding immune checkpoint molecules (e.g., PDCD1, CTLA4) may reflect vaccine-induced immune activation or resistance mechanisms. In some embodiments, methylation changes at these immune loci are quantified and used as features for classifying subjects as responders or non-responders using trained computational models.
[0110] In certain embodiments, methylation panels targeting immune-related pathways may be used to monitor therapeutic response following administration of a personalized cancer vaccine (PCV). These panels may include promoter, enhancer, and / or gene body regions of key immune genes whose expression is modulated during vaccine-induced immune activation. Exemplary loci include interferon signaling genes (IFNG, IFNB1, IRF1, IRF7, STAT1, STAT2), antigen processing and presentation genes (HLA-A, HLA-B, HLA-C, B2M, TAPI, TAP2), immune checkpoint regulators (PDCD1 [PD-1], CTLA4, LAG3, TIGIT), and / or T-cell activation markers (CD3E, CD8A, CD69, GZMB). Methylation changes at loci encoding key cytokines (IL2, IL6, TNF, CXCL9, CXCL10) and / or dendritic cell activation regulators (e.g., CD80, CD86, CCR7) may also be assessed. In certain embodiments, CpG-rich regions upstream of miR-155 and miR-146a may be included to detect epigenetic regulation of microRNA-mediated immune programming. Methylation data may be generated from posttreatment samples collected at various timepoints, including blood, plasma, serum, saliva, urine, cerebrospinal fluid, pleural fluid, ascitic fluid, synovial fluid, breast milk, semen, vaginal fluid, amniotic fluid, bile, sweat, and / or tears. Genomic DNA, RNA, cell-free DNA, and / or cell-free RNA may be extracted from these samples and used as input for bisulfite-based or affinity-based methylation profiling methods. The resulting data may be analyzed to detect immune activation signatures, PCV treatment-induced epigenetic remodeling, or molecular signatures associated with therapeutic resistance.
[0111] Single site methylation sequencing workflowsAttorney Docket No. : GH0224WO
[0112] Various single-site methylation (SSM) sequencing methods exist, these are broadly categorized into RRBS-based and WGBS-based approaches. The following section describes SSM workflows to detect epigenetic changes in subjects undergoing cancer treatment, including treatment with a PCV and / or an off-the-shelf cancer vaccine. The vaccine may be administered to the subject alone or in combination with other cancer therapies, including pembrolizumab, as previously described in the specification. In some embodiments the workflows comprise targeted methylation profiling using panel of genes of interest. In additional embodiments the workflows comprise genome-wide methylation profiling. The workflows may additionally comprise CpG methylation profiling at base-pair resolution or fragment level methylation profiling, including hypermethylated fragment enrichment and / or hypomethylated fragment enrichment. The workflows may be compatible with low-input or single-cell DNA samples. The resulting data may be used to track tumor or immune-related methylation changes associated with PCV treatment, therapeutic response, resistance, or disease progression. RRBS-based methods focus on GC-rich regions, optimizing cost and efficiency for single-cell research. WGBS-based methods offer genome-wide coverage, with variations to enhance DNA preservation and reduce loss. Both methodologies include adaptations for single-cell analysis, integrating technologies like microfluidics and unique molecular identifiers to improve accuracy and minimize DNA loss. Workflows for single-site methylation sequencing has seen significant recent advances, with methodologies focusing on high-throughput, cost-efficiency, and enhanced sensitivity. A typical outline of the workflow based on the most recent research includes the following steps: (1) Sample Preparation and DNA Isolation: Collection of target samples, followed by DNA extraction and purification. This initial step is crucial for ensuring the quality of DNA for subsequent methylation analysis. Bisulfite Treatment (Optional for Some Methods): Traditional workflows often involve bisulfite conversion, where unmethylated cytosines are converted to uracil, while methylated cytosines remain unchanged. However, newer methods like MLAD-seq and EAC-seq offer bi sulfite-free alternatives, providing single-base resolution and quantitative detection of 5mC without the DNA degradation associated with bisulfite treatment details of such protocols can be found in the literature for example Xiong et al., 2022 and Wang et al., 2022. Library Preparation and Sequencing: library preparation may involve enrichment of regions of interest or tagging DNA fragments with unique molecular identifiers (UMIs) before sequencing. Data Analysis and Interpretation: Post-sequencing, the data undergoes processing to identify methylated sites. Computational tools and algorithms, such as those incorporated in Methyl Score, can accurately identify differentially methylated regions (DMRs) and predictAttorney Docket No. : GH0224WO phenotypes or disease states based on methylation profiles (Hiither et al., 2022). Methods like EAC-seq be applied for direct and bi sulfite-free detection of DNA methylation.
[0113] Enzymes used in single site methylation profiling
[0114] Bisulfite sequencing does not directly use enzymes for the conversion but relies on chemical treatment to differentiate between methylated and unmethylated cytosines. Methods that use enzymes include Ten-eleven translocation (TET)-assisted pyridine borane sequencing (TAPS): this method utilizes TET enzymes to oxidize 5mC to 5caC, which is then converted to thymine via pyridine borane reduction, allowing for sequencing without bisulfite conversion. Enzymatic methyl-seq (EM-seq) this method involves the use of enzymes for conversion, this approach is a less damaging alternative to bisulfite treatment for identifying methylated sites.
[0115] Subjecting the first subsample to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample
[0116] In some embodiments, methods disclosed herein comprise a step of subjecting DNA, or a subsample thereof, to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity. In some embodiments, the procedure chemically converts the first or second nucleobase such that the base pairing specificity of the converted nucleobase is altered. In some embodiments, DNA is subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA before library preparation using the DNA, before a first amplification of the DNA, before dividing the DNA into a plurality of subsamples, or any combination thereof. In certain embodiments, the DNA is subjected to the procedure before or after contacting the DNA with a methylation-sensitive nuclease.
[0117] In some embodiments, the procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA is performed prior to the sequencing and / or (a) prior to or after the selectively depleting the target nucleic acid comprising the wild-type sequence, the target nucleic acid comprising the converted nucleotide, or the target nucleic acid that does not comprise the converted nucleotide; (b) prior to the amplifying the selectively digested population of target nucleic acids; (c) prior to or after the partitioning the population of target nucleic acids into a plurality of subsamples; and / or (d) prior to or after a step of enriching for one or more sets of target regions of DNA.
[0118] In some embodiments, if the first nucleobase is a modified or unmodified adenine, then the second nucleobase is a modified or unmodified adenine; if the first nucleobase is a modifiedAttorney Docket No. : GH0224WO or unmodified cytosine, then the second nucleobase is a modified or unmodified cytosine; if the first nucleobase is a modified or unmodified guanine, then the second nucleobase is a modified or unmodified guanine; and if the first nucleobase is a modified or unmodified thymine, then the second nucleobase is a modified or unmodified thymine (where modified and unmodified uracil are encompassed within modified thymine for the purpose of this step).
[0119] In some embodiments, the first nucleobase is a modified or unmodified cytosine, then the second nucleobase is a modified or unmodified cytosine. For example, first nucleobase may comprise unmodified cytosine (C) and the second nucleobase may comprise one or more of 5- methylcytosine (mC) and 5-hydroxymethylcytosine (hmC). Alternatively, the second nucleobase may comprise C and the first nucleobase may comprise one or more of mC and hmC. Other combinations are also possible, such as where one of the first and second nucleobases comprises mC and the other comprises hmC.
[0120] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises bisulfite conversion. Treatment with bisulfite converts unmodified cytosine and certain modified cytosine nucleotides (e.g. 5- formyl cytosine (fC) or 5-carboxylcytosine (caC)) to uracil whereas other modified cytosines (e.g., 5-methylcytosine, 5-hydroxylmethylcystosine) are not converted. Thus, where bisulfite conversion is used, the first nucleobase comprises one or more of unmodified cytosine, 5- formyl cytosine, 5-carboxylcytosine, or other cytosine forms affected by bisulfite, and the second nucleobase may comprise one or more of mC and hmC, such as mC and optionally hmC. Sequencing of bisulfite-treated DNA identifies positions that are read as cytosine as being mC or hmC positions. Meanwhile, positions that are read as T are identified as being T or a bisulfite-susceptible form of C, such as unmodified cytosine, 5-formyl cytosine, or 5- carboxylcytosine. Performing bisulfite conversion, such as on a DNA sample as described herein, facilitates identifying positions containing mC or hmC using the sequence reads obtained from the exemplary sample. For an exemplary description of bisulfite conversion, see, e.g., Moss et al., Nat Commun. 2018; 9: 5068.
[0121] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises oxidative bisulfite (Ox-BS) conversion. This procedure first converts hmC to fC, which is bisulfite susceptible, followed by bisulfite conversion. Thus, when oxidative bisulfite conversion is used, the first nucleobase comprises one or more of unmodified cytosine, fC, caC, hmC, or other cytosine forms affected by bisulfite, and the second nucleobase comprises mC. Sequencing of Ox-BS converted DNA identifies positions that are read as cytosine as being mC positions. Meanwhile, positions thatAttorney Docket No. : GH0224WO are read as T are identified as being T, hmC, or a bisulfite-susceptible form of C, such as unmodified cytosine, fC, or hmC. Performing Ox-BS conversion, such as on a DNA sample as described herein, thus facilitates identifying positions containing mC using the sequence reads obtained from the sample. For an exemplary description of oxidative bisulfite conversion, see, e.g., Booth et al., Science 2012; 336: 934-937.
[0122] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises Tet-assisted bisulfite (TAB) conversion. In TAB conversion, hmC is protected from conversion and mC is oxidized in advance of bisulfite treatment, so that positions originally occupied by mC are converted to U while positions originally occupied by hmC remain as a protected form of cytosine. For example, as described in Yu et al., Cell 2012; 149: 1368-80, P-glucosyl transferase can be used to protect hmC (forming 5-glucosylhydroxymethylcytosine (ghmC)), then a TET protein such as mTetl can be used to convert mC to caC, and then bisulfite treatment can be used to convert C and caC to U while ghmC remains unaffected.
[0123] Alternatively, a carbamoyltransferase enzyme, such as 5 -hydroxymethylcytosine carbamoyltransferase as described in Yang et al., Bio-protocol, 2023; 12(17): e4496, can be used to protect hmC (by converting hmC to 5-carbamoyloxymethylcytosine (5cmC)), then a TET protein such as mTetl or a TET2 comprising a T1372S mutation, can be used to convert mC to caC, and then bisulfite treatment can be used to convert C and caC to U while 5cmC remains unaffected. Thus, when TAB conversion is used, the first nucleobase comprises one or more of unmodified cytosine, fC, caC, mC, or other cytosine forms affected by bisulfite, and the second nucleobase comprises hmC. Sequencing of TAB-converted DNA identifies positions that are read as cytosine as being hmC positions. Meanwhile, positions that are read as T are identified as being T, mC, or a bisulfite-susceptible form of C, such as unmodified cytosine, fC, or caC. Performing TAB conversion, such as on a DNA sample as described herein, thus facilitates identifying positions containing hmC using the sequence reads obtained from the sample.
[0124] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises Tet-assisted conversion with a substituted borane reducing agent, optionally wherein the substituted borane reducing agent is 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane. In Tet-assisted pic-borane conversion with a substituted borane reducing agent conversion, a TET protein is used to convert mC and hmC to caC, without affecting unmodified C. caC, and fC if present, are then converted to dihydrouracil (DHU) by treatment with 2-picoline borane (pic-borane) orAttorney Docket No. : GH0224WO another substituted borane reducing agent such as borane pyridine, tert-butylamine borane, or ammonia borane, also without affecting unmodified C. See, e.g., Liu et al., Nature Biotechnology 2019; 37:424-429 (e.g., at Supplementary Fig. 1 and Supplementary Note 7). Thus, when this type of conversion is used, the first nucleobase comprises one or more of 5mC, 5fC, 5caC, or 5hmC, and the second nucleobase comprises unmodified cytosine. DHU is read as a T in sequencing. Thus, when this type of conversion is used, the first nucleobase comprises one or more of mC, fC, caC, or hmC, and the second nucleobase comprises unmodified cytosine. Sequencing of the converted DNA identifies positions that are read as cytosine as being unmodified C positions. Meanwhile, positions that are read as T are identified as being T, mC, fC, caC, or hmC. Performing TAP conversion, such as on a DNA sample as described herein, thus facilitates identifying positions containing unmodified C using the sequence reads obtained from the sample. This procedure encompasses Tet-assisted pyridine borane sequencing (TAPS), described in further detail in Liu et al. 2019, supra.
[0125] Alternatively, protection of hmC (e.g., using PGT or 5-hydroxymethylcytosine carbamoyltransferase) can be combined with Tet-assisted conversion with a substituted borane reducing agent, e.g. as described above. In this method (TAPS-P), 5hmC can be protected from conversion, for example through glucosylation using P-glucosyl transferase (PGT), forming 5- glucosylhydroxymethylcytosine (5ghmC), or through carbamoylation using 5- hydroxymethylcytosine carbamoyltransferase, forming 5cmC. This is described in Yu et al., Cell 2012; 149: 1368-80. Treatment with a TET protein, such as mTetl or a TET2 comprising a T1372S mutation, then converts mC to caC but does not convert C, 5ghmC, or 5cmC. 5caC is then converted to DHU by treatment with pic-borane or another substituted borane reducing agent such as borane pyridine, tert-butylamine borane, or ammonia borane, also without affecting ghmC, 5cmC, or unmodified C. Thus, when Tet-assisted conversion with a substituted borane reducing agent is used, the first nucleobase comprises mC, and the second nucleobase comprises one or more of unmodified cytosine or hmC, such as unmodified cytosine and optionally hmC, fC, and / or caC. Sequencing of the converted DNA identifies positions that are read as cytosine as being either hmC or unmodified C positions. Meanwhile, positions that are read as T are identified as being T, fC, caC, or mC. Performing TAPSP conversion, such as on a DNA sample as described herein, thus facilitates distinguishing positions containing unmodified C or hmC on the one hand from positions containing mC using the sequence reads obtained from the sample. For an exemplary description of this type of conversion, see, e.g., Liu et al., Nature Biotechnology 2019; 37:424-429. 5-Attorney Docket No. : GH0224WO hydroxymethylcytosine carbamoyltransferase is described in Yang et al., Bio-protocol, 2023; 12(17): e4496.
[0126] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises chemi cal -assisted conversion with a substituted borane reducing agent, optionally wherein the substituted borane reducing agent is 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane. In chemical- assisted conversion with a substituted borane reducing agent, an oxidizing agent such as potassium perruthenate (KRuO4) (also suitable for use in ox-BS conversion) is used to specifically oxidize hmC to fC. Treatment with pic-borane or another substituted borane reducing agent such as borane pyridine, tert-butylamine borane, or ammonia borane converts fC and caC to DHU but does not affect mC or unmodified C. Thus, when this type of conversion is used, the first nucleobase comprises one or more of hmC, fC, and caC, and the second nucleobase comprises one or more of unmodified cytosine or mC, such as unmodified cytosine and optionally mC. Sequencing of the converted DNA identifies positions that are read as cytosine as being either mC or unmodified C positions. Meanwhile, positions that are read as T are identified as being T, fC, caC, or hmC. Performing this type of conversion, such as on a DNA sample as described herein, thus facilitates distinguishing positions containing unmodified C or mC on the one hand from positions containing hmC using the sequence reads obtained from the sample. For an exemplary description of this type of conversion, see, e.g., Liu et al., Nature Biotechnology 2019; 37:424-429.
[0127] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises APOBEC-coupled epigenetic (ACE) conversion. In ACE conversion, an AID / APOBEC family DNA deaminase enzyme such as APOBEC3A (A3A) is used to deaminate unmodified cytosine and mC without deaminating hmC, fC, or caC. Thus, when ACE conversion is used, the first nucleobase comprises unmodified C and / or mC (e.g., unmodified C and optionally mC), and the second nucleobase comprises hmC. Sequencing of ACE-converted DNA identifies positions that are read as cytosine as being hmC, fC, or caC positions. Meanwhile, positions that are read as T are identified as being T, unmodified C, or mC. Performing ACE conversion on a DNA sample as described herein thus facilitates distinguishing positions containing hmC from positions containing mC or unmodified C using the sequence reads obtained from the sample. For an exemplary description of ACE conversion, see, e.g., Schutsky et al., Nature Biotechnology 2018; 36: 1083-1090.Attorney Docket No. : GH0224WO
[0128] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises enzymatic conversion of the first nucleobase, e.g., as in EM-Seq. See, e.g., Vaisvila R, et al. (2019) EM-seq: Detection of DNA methylation at single base resolution from picograms of DNA. bioRxiv; DOI: 10.1101 / 2019.12.20.884692, available at www.biorxiv.org / content / 10.1101 / 2019.12.20.884692vl. For example, TET2 and T4-PGT or 5-hydroxymethylcytosine carbamoyltransferase (described in Yang et al., Bio-protocol, 2023; 12(17): e4496) can be used to convert 5mC and 5hmC into substrates that cannot be deaminated by a deaminase (e.g., APOBEC3A), and then a deaminase (e.g., APOBEC3A) can be used to deaminate unmodified cytosines converting them to uracils.
[0129] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises enzymatic conversion of the first nucleobase using a non-specific, modification-sensitive double-stranded DNA deaminase, e.g., as in SEM-seq. See, e.g., Vaisvila et al. (2023) Discovery of novel DNA cytosine deaminase activities enables a nondestructive single-enzyme methylation sequencing method for base resolution high-coverage methylome mapping of cell-free and ultra-low input DNA. bioRxiv; DOI: 10.1101 / 2023.06.29.547047, available at https: / / www.biorxiv.org / content / 10.1101 / 2023.06.29.547047vl. SEM-Seq employs a nonspecific, modification-sensitive double-stranded DNA deaminase (MsddA) in a nondestructive single-enzyme 5-methylctyosine sequencing (SEM-seq) method that deaminates unmodified cytosines. Accordingly, SEM-seq does not require the TET2 and T4-PGT or 5- hydroxymethylcytosine carbamoyltransferase protection and denaturing steps that are of use, e.g., in APOEC3A-based protocols. Additionally, MsddA does not deaminate 5-formylated cytosines (5fC) or 5-carboxylated cytosines (5caC). In SEM-seq, unmodified cytosines in the DNA are deaminated to uracil and is read as “T” during sequencing. Modified cytosines (e.g., 5mC) are not converted and are read as “C” during sequencing. Cytosines that are read as thymines are identified as unmodified (e.g., unmethylated) cytosines or as thymines in the DNA. Performing SEM-seq conversion thus facilitates identifying positions containing 5mC using the sequence reads obtained. In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises enzymatic conversion of the first nucleobase using MsddA.
[0130] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample converts a modified nucleoside. In some embodiments, the conversion procedure which converts a modifiedAttorney Docket No. : GH0224WO nucleosides comprises enzymatic conversion, such as DM-seq, for example, as described in WO2023 / 288222A1. In DM-seq, unmodified cytosines in the DNA are enzymatically protected from a subsequent deamination step wherein 5mC in 5mCpG is converted to T. The enzymatically protected unmodified (e.g., unmethylated) cytosines are not converted and are read as “C” during sequencing. Cytosines that are read as thymines (in a CpG context) are identified as methylated cytosines in the DNA. Thus, when this type of conversion is used, the first nucleobase comprises unmodified (such as unmethylated) cytosine, and the second nucleobase comprises modified (such as methylated) cytosine. Sequencing of the converted DNA identifies positions that are read as cytosine as being unmodified C positions. Meanwhile, positions that are read as T are identified as being T or 5mC. Performing DM-seq conversion thus facilitates identifying positions containing 5mC using the sequence reads obtained.
[0131] Exemplary cytosine deaminases for use herein include APOBEC enzymes, for example, APOBEC3A. Generally, AID / APOBEC family DNA deaminase enzymes such as APOBEC3 A (A3 A) are used to deaminate (unprotected) unmodified cytosine and 5mC. For an exemplary description of APOBEC conversion, see, e.g., Schutsky et al., Nature Biotechnology 2018; 36: 1083-1090.
[0132] The enzymatic protection of unmodified cytosines in the DNA comprises addition of a protective group to the unmodified cytosines. Such protective groups can comprise an alkyl group, an alkyne group, a carboxyl group, a carboxyalkyl group, an amino group, a hydroxymethyl group, a glucosyl group, a glucosylhydroxymethyl group, an isopropyl group, or a dye. For example, DNA can be treated with a methyltransferase, such as a CpG-specific methyltransferase, which adds the protective group to unmodified cytosines. The term methyltransferase is used broadly herein to refer to enzymes capable of transferring a methyl or substituted methyl (e.g., carboxymethyl) to a substrate (e.g., a cytosine in a nucleic acid). In some embodiments, the DNA is contacted with a CpG-specific DNA methyltransferase (MTase), such as a CpG-specific carboxymethyltransferase (CxMTase), and a substituted methyl donor, such as a carboxymethyl donor (e.g., carboxymethyl-S-adenosyl-L-methionine). See, e.g., WO2021 / 236778A2. In particular embodiments, the CxMTase can facilitate the addition of a protective carboxymethyl group to an unmethylated cytosine. In some embodiments, the unmethylated cytosine is unmodified cytosine. The carboxymethyl group can prevent deamination of the cytosine during a deamination step (such as a deamination step using an APOBEC enzyme, such as A3 A). Substituted methyl or carboxymethyl donors useful in the disclosed methods include but are not limited to, S-adenosyl-L-methionine (SAM) analogs, optionally wherein the SAM analog is carboxy-S-adenosyl-L-methionine (CxSAM).Attorney Docket No. : GH0224WOSAM analogs are described, for example, in WO2022 / 197593A1. The MTase may be, for example, a CpG methyltransferase from Spiroplasma sp. strain MQ1 (M.SssI), DNA- methyltransferase 1 (DNMT1), DNA-methyltransferase 3 alpha (DNMT3A), DNA- methyltransferase 3 beta (DNMT3B), or DNA adenine methyltransferase (Dam). The CxMTase may be a CpG methyltransferase from Mycoplasma penetrans (M.Mpel). In a particular embodiment, the methyltransferase enzyme is a variant of M.Mpel having SEQ ID NO: 1 or SEQ ID NO: 2, or a sequence at least 90%, at least 92%, at least 94%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto, optionally wherein the amino acid corresponding to position 374 is R or K.
[0133] In one embodiment, the methyltransferase enzyme is a variant of M.Mpel having an N374R substitution or an N374K substitution. The methyltransferase of SEQ ID NO: 1 or SEQ ID NO: 2 can further comprise one or more amino acid substitutions selected from a) substitution of one or both residues T300 and E305 with S, A, G, Q, D, orN; b) substitution of one or more residues A323, N306, and Y299 with a positively charged amino acid selected from K, R or H; and / or c) substitution of S323 with A, G, K, R or H, which may enhance the activity of the enzyme.
[0134] Optionally, the conversion procedure further includes enzymatic protection of 5hmCs, such as by glucosylation of the 5hmCs (e.g., using PGT) or by carbamoylation of the 5hmCs (e.g., using 5-hydroxymethylcytosine carbamoyltransferase), in the DNA prior to the deamination of unprotected modified cytosines. In this method, 5hmC can be protected from conversion, for example through glucosylation using P-glucosyl transferase (PGT), forming (5- glucosylhydroxymethylcytosine) 5ghmC, or through carbamoylation using 5- hydroxymethylcytosine carbamoyltransferase, forming 5cmC. This is described, for example, in Yu et al., Cell 2012; 149: 1368-80, and in Yang et al., Bio-protocol, 2023; 12(17): e4496. Glucosylation or carbamoylation of 5hmC can reduce or eliminate deamination of 5hmC by a deaminase such as APOBEC3 A. Treatment with an MTase or CxMTase then adds a protecting group to unmodified (unmethylated) cytosines in the DNA. 5mC (but not protected, unmodified cytosine and not 5ghmC or 5cmC) is then deaminated (converted to T in the case of 5mC) by treatment with a deaminase, for example, an APOBEC enzyme (such as APOBEC3 A). Sequencing of the converted DNA identifies positions that are read as cytosine as being either 5hmC or unmodified C positions. Meanwhile, positions that are read as T are identified as being T or 5mC. Performing DM-seq conversion with glucosylation of 5hmC on a sample as described herein thus facilitates distinguishing positions containing unmodified C or 5hmC on the one hand from positions containing 5mC using the sequence reads obtained.Attorney Docket No. : GH0224WO
[0135] Also provided herein are methods in which alternative base conversion schemes are used. For example, unmethylated cytosines can be left intact while methylated cytosines and hydroxymethylcytosines are converted to a base read as a thymine (e.g., uracil, thymine, or dihydrouracil).
[0136] In some embodiments, methylating a cytosine in at least one first complementary strand or second complementary strand comprises contacting the cytosine with a methyltransferase such as DNMT1 or DNMT5. In such embodiments, the step of oxidizing a 5- hydroxymethylated cytosine to 5-formylcytosine (such as by contacting the 5 -hydroxymethyl cytosine in a first strand and a second strand with KRuO4) can be optional.
[0137] In some embodiments, converting the modified cytosine in at least one first or second strand to a thymine or a base read as thymine comprises oxidizing a hydroxymethyl cytosine, e.g., the hydroxymethyl cytosine is oxidized to formylcytosine. In some embodiments, oxidizing the hydroxymethyl cytosine to formylcytosine comprises contacting the hydroxymethyl cytosine with a ruthenate, such as potassium ruthenate (KRuO4).
[0138] In some embodiments, the modified cytosine is converted to thymine, uracil, or dihydrouracil. In any such embodiments, amplification methods may comprise uracil- and / or dihydrouracil-tolerant amplification methods, such as PCR using a uracil- and / or dihydrouracil-tolerant DNA polymerase.
[0139] In some embodiments, the method comprises converting a formylcytosine and / or a methylcytosine to carboxylcytosine as part of converting the modified cytosine in at least one first or second strand to a thymine or a base read as thymine. For example, converting the formylcytosine and / or the methylcytosine to carboxylcytosine can comprise contacting the formylcytosine and / or the methylcytosine with a TET enzyme, such as TET1, TET2, TET3, or a TET2 comprising a T1372S mutation. In some embodiments, the method comprises reducing the carboxylcytosine as part of converting the modified cytosine in at least one first or second strand to a thymine or a base read as thymine, and / or the carboxylcytosine is reduced to dihydrouracil. In some embodiments, reducing the carboxylcytosine comprises contacting the carboxylcytosine with a borane or borohydride reducing agent.
[0140] In some embodiments, the borane or borohydride reducing agent comprises pyridine borane, 2-picoline borane, borane, tert-butylamine borane, ammonia borane, sodium borohydride, sodium cyanoborohydride (NaBH3CN), lithium borohydride (LiBH4), ethylenediamine borane, dimethylamine borane, sodium triacetoxyborohydride, morpholine borane, 4-methylmorpholine borane, trimethylamine borane, dicyclohexylamine borane, or a salt thereof. In other embodiments, the reducing agent comprises lithium aluminum hydride,Attorney Docket No. : GH0224WO sodium amalgam, amalgam, sulfur dioxide, dithionate, thiosulfate, iodide, hydrogen peroxide, hydrazine, diisobutylaluminum hydride, oxalic acid, carbon monoxide, cyanide, ascorbic acid, formic acid, dithiothreitol, beta-mercaptoethanol, or any combination thereof.
[0141] Various TET enzymes may be used in the disclosed methods as appropriate. In some embodiments, the one or more TET enzymes comprise TETv. TETv is described in US Patent 10,260,088 and its sequence is SEQ ID NO: 1 therein (SEQ ID NO: 3 in the present application). In some embodiments, the one or more TET enzymes comprise TETcd. TETcd is described in US Patent 10,260,088 and its sequence is SEQ ID NO: 3 therein (SEQ ID NO: 4 in the present application). In some embodiments, the one or more TET enzymes comprise TET1. In some embodiments, the one or more TET enzymes comprise TET2. TET2 may be expressed and used as a fragment comprising TET2 residues 1129-1480 joined to TET2 residues 1844-1936 by a linker (SEQ ID NO: 5 of the present application) as described, e.g., in US Patent 10,961,525. In some embodiments, the one or more TET enzymes comprise TET1 and TET2. In some embodiments, the one or more TET enzymes comprise a VI 900 TET mutant, such as a VI 900 A, V1900C, V1900G, VI 9001, or V1900P TET mutant. In some embodiments, the one or more TET enzymes comprise a VI 900 TET2 mutant, such as a V1900A, V1900C, V1900G, VI 9001, or V1900P TET2 mutant. Examples of V1900A, V1900C, V1900G, VI 9001, and V1900P TET2 mutants are provided as SEQ ID NOs: 6-10. In some embodiments, the VI 900 TET mutant has at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 6, 7, 8, 9, or 10. Position 1900 of the wild-type TET2 sequence corresponds to position 438 in each of SEQ ID NOs: 5-10. It can be beneficial to use a TET enzyme that maximizes formation of 5-carboxylcytosine (5-caC) relative to less oxidized modified cytosines, particularly 5-formylcytosine, because 5-caC is not a substrate for enzymatic deamination, e.g., by APOBEC enzymes such as APOBEC3A. Maximizing formation of 5-caC thus reduces the risk of false calls in which a base is identified as unmethylated because it underwent deamination even though it was methylated (or hydroxymethylated) in the original sample. Accordingly, in some embodiments, the TET enzyme comprises a mutation that increases formation of 5-caC. Exemplary mutations are set forth above. “A mutation that increases formation of 5-caC” means that the TET enzyme having the mutation produces more 5-caC than a TET enzyme that lacks the mutation but is otherwise identical. 5-caC production can be measured as described, e.g., in Liu et al., Nat Chem Biol 13: 181-187 (2017) (see Online Methods section, TET reactions in vitro subsection, “driving” conditions). Any variants and / or mutants described in Liu et al. (2017) can be used in the disclosed methods as appropriate.Attorney Docket No. : GH0224WO
[0142] In some embodiments, the one or more TET enzymes comprise a TET2 enzyme comprising a T1372S mutation, such as TET2-CS-T1372S and TET2-CD-T1372S. Examples of TET2-CS-T1372S and TET2-CD-T 1372 S are provided as SEQ ID NOs: 11 and 12. A TET2 comprising a T1372S mutation is described in US Patent 10,961,525 and may be expressed and used as a fragment comprising TET2 residues 1129-1480 joined to TET2 residues 1844- 1936 by a linker. Position 1372 of TET2 corresponds to position 258 of SEQ ID NO: 21 (wild type TET2 catalytic domain) of US Patent 10,961,525. Thus, the sequence of a T1372S TET2 catalytic domain may be obtained by changing the threonine at position 258 of SEQ ID NO: 21 of US Patent 10,961,525 to serine. TET2 comprising a T1372S mutation is also described in Liu et al., Nat Chem Biol. 2017 February; 13(2): 181-187. As demonstrated in Liu et al., TET2 comprising a T1372S mutation can more efficiently oxidize 5mC to produce 5- carboxylcytosine (5caC) than other versions of TET2 such as TET2 lacking a T1372S mutation. In some embodiments, the TET2 enzyme comprises SEQ ID NO: 14 or optionally a variant of SEQ ID NO: 14 in which at least 5, 6, 7, or 8 positions match SEQ ID NO: 14 including position 5 of SEQ ID NO: 14. In some embodiments, the TET2 enzyme is a human TET2 enzyme comprising a T1372S mutation. In some embodiments, the TET2 enzyme comprises the sequence of SEQ ID NO: 11. In some embodiments, the TET2 enzyme comprises a sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 11. In some embodiments, the TET2 enzyme comprises a sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 12. In some embodiments, the TET2 enzyme comprises the sequence of SEQ ID NO: 12. The sequences of SEQ ID NOs: 11 and 12 are shown in the Table of Sequences herein.
[0143] Provided herein is a method comprising contacting DNA contacting DNA with a TET2 enzyme comprising a T1372S mutation to oxidize 5-methylcytosine (5mC) and / or 5- hydroxymethylcytosine (5hmC) present in the DNA to 5-carboxycytosine (5caC), subsequently contacting at least a portion of the DNA with a substituted borane reducing agent, thereby converting 5-caC in the DNA to dihydrouracil (DHU), thereby producing treated DNA, and sequencing at least a portion of the treated DNA.
[0144] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises separating DNA originally comprising the first nucleobase from DNA not originally comprising the first nucleobase. In some such embodiments, the first nucleobase is hmC. DNA originally comprising the first nucleobase may be separated from other DNA using a labeling procedure comprising biotinylating positions that originally comprised the first nucleobase. In some embodiments,Attorney Docket No. : GH0224WO the first nucleobase is first derivatized with an azide-containing moiety, such as a glucosylazide containing moiety. The azide-containing moiety then may serve as a reagent for attaching biotin, e.g., through Huisgen cycloaddition chemistry. Then, the DNA originally comprising the first nucleobase, now biotinylated, can be separated from DNA not originally comprising the first nucleobase using a biotin-binding agent, such as avidin, neutravidin (deglycosylated avidin with an isoelectric point of about 6.3), or streptavidin. An example of a procedure for separating DNA originally comprising the first nucleobase from DNA not originally comprising the first nucleobase is hmC-seal, which labels hmC to form P-6-azide-glucosyl-5- hydroxymethylcytosine and then attaches a biotin moiety through Huisgen cycloaddition, followed by separation of the biotinylated DNA from other DNA using a biotin-binding agent. For an exemplary description of hmC-seal, see, e.g., Han et al., Mol. Cell 2016; 63: 711-719. This approach is useful for identifying fragments that include one or more hmC nucleobases.
[0145] In some embodiments, following such a separation, the method further comprises differentially tagging each of the DNA originally comprising the first nucleobase, the DNA not originally comprising the first nucleobase. The method may further comprise pooling the DNA originally comprising the first nucleobase and the DNA not originally comprising the first nucleobase following differential tagging. The DNA originally comprising the first nucleobase and the DNA not originally comprising the first nucleobase may then be used in downstream analyses. For example, the pooled DNA originally comprising the first nucleobase and the DNA not originally comprising the first nucleobase may be sequenced in the same sequencing cell (such as after being subjected to further treatments, such as those described herein) while retaining the ability to resolve whether a given read came from a molecule of DNA originally comprising the first nucleobase or DNA not originally comprising the first nucleobase using the differential tags.
[0146] In some embodiments, the first nucleobase is a modified or unmodified adenine, and the second nucleobase is a modified or unmodified adenine. In some embodiments, the modified adenine is N6-methyladenine (mA). In some embodiments, the modified adenine is one or more of N6-methyladenine (mA), N6-hydroxymethyladenine (hmA), or N6- formyladenine (fA).
[0147] Techniques comprising partitioning based on methylation status or methylated DNA immunoprecipitation (MeDIP) can be used to separate DNA containing modified bases such as mC, mA, caC (which may be generated by oxidation of mC or hmC with Tet2, e.g., before enzymatic conversion of unmodified C to U, e.g., using a deaminase such as APOBEC3 A), or dihydrouracil from other DNA. See, e.g., Kumar et al., Frontiers Genet. 2018; 9: 640; Greer etAttorney Docket No. : GH0224WO al., Cell 2015; 161 : 868-878. An antibody specific for mA is described in Sun et al., Bioessays 2015; 37: 1155-62. Antibodies for various modified nucleobases, such as mC, caC, and forms of thymine / uracil including dihydrouracil or halogenated forms such as 5 -bromouracil, are commercially available. Various modified bases can also be detected based on alterations in their base pairing specificity. For example, hypoxanthine is a modified form of adenine that can result from deamination and is read in sequencing as a G. See, e.g., US Patent 8,486,630; Brown, Genomes, 2nd Ed., John Wiley & Sons, Inc., New York, N.Y., 2002, chapter 14, “Mutation, Repair, and Recombination.”
[0148] Partitioning the sample
[0149] In certain exemplary embodiments, described herein, physical partitioning of nucleic acid populations may be employed to separate and distinguish hypermethylated from hypomethylated DNA fragments. Subsequent analysis of these partitioned nucleic acid populations can be used to identify differentially methylated regions (DMRs), including DMRs that contain sequences encoding neoantigens incorporated into the personalized cancer vaccine (PCV) or off-the-shelf vaccine administered to the subject. Joint assessment of epigenetic status and neoantigen-associated sequences may improve sensitivity for detecting therapeutic response to the PCV and / or off-the-shelf vaccine. In some embodiments, sequence analysis may be performed on the hypermethylated partition to identify tumor-specific variants, such as somatic mutations encoding neoantigens.
[0150] In some embodiments, a population of different forms of nucleic acids(e.g., hypermethylated and hypomethylated DNA in a sample, such as a captured set of cfDNA as described herein) can be physically partitioned based on one or more characteristics of the nucleic acids prior to further analysis, e.g., differentially modifying or isolating a nucleobase, tagging, and / or sequencing. Additionally the population of different forms of nucleic acids may include transcription factor binding sites (TFBS), mRNA expression, fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, H3S10ph. This approach can be used to determine, for example, whether certain sequences are hypermethylated or hypomethylated or contain a specific histone modification or functional element such as a TFBS. In some embodiments, hypermethylation variable epigenetic target regions are analyzed to determine whether they show hypermethylation characteristic of tumor cells and / or hypomethylation variable epigenetic target regions are analyzed to determineAttorney Docket No. : GH0224WO whether they show hypomethylation characteristic of tumor cells. Additionally, by partitioning a heterogeneous nucleic acid population, one may increase rare signals, e.g., by enriching rare nucleic acid molecules that are more prevalent in one fraction (or partition) of the population. For example, a genetic variation present in hyper-methylated DNA but less (or not) in hypomethylated DNA can be more easily detected by partitioning a sample into hypermethylated and hypo-methylated nucleic acid molecules. By analyzing multiple fractions of a sample, a multi-dimensional analysis of a single locus of a genome or species of nucleic acid can be performed and hence, greater sensitivity can be achieved. In some embodiments of the disclosure each partition may comprise different forms of nucleic acids including transcription factor binding sites (TFBS), mRNA expression, fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, H3S10ph. This approach can be used to determine, for example, whether certain sequences are hypermethylated or hypomethylated or contain a specific histone modification or functional element such as a TFBS.
[0151] In some instances, a heterogeneous nucleic acid sample is partitioned into two or more partitions. For instance, a minimum of three partitions, extending to four, five, six, seven, up to any number of partitions, where the total number can be any non-negative real number). In some embodiments, each partition is differentially tagged. Tagged partitions can then be pooled together for collective sample prep and / or sequencing. The partitioning-tagging-pooling steps can occur more than once, with each round of partitioning occurring based on a different characteristics (examples provided herein), and tagged using differential tags that are distinguished from other partitions and partitioning means.
[0152] Examples of characteristics that can be used for partitioning include sequence length, methylation level, nucleosome binding, sequence mismatch, immunoprecipitation, and / or proteins that bind to DNA. Additional characteristics that can be used for partitioning include, transcription factor binding sites (TFBS), fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, H3S10ph, type of nucleic acid for example RNA, mRNA, cDNA, or DNA. Resulting partitions can include one or moreAttorney Docket No. : GH0224WO of the following nucleic acid forms: single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), shorter DNA fragments and longer DNA fragments. In some embodiments, partitioning based on a cytosine modification (e.g., cytosine methylation) or methylation generally is performed and is optionally combined with at least one additional partitioning step, which may be based on any of the foregoing characteristics or forms of DNA. In some embodiments, a heterogeneous population of nucleic acids is partitioned into nucleic acids with one or more epigenetic modifications and without the one or more epigenetic modifications. Examples of epigenetic modifications include presence or absence of methylation; level of methylation; type of methylation (e.g., 5-methylcytosine versus other types of methylation, such as adenine methylation and / or cytosine hydroxymethylation); and association and level of association with one or more proteins, such as histones. Alternatively or additionally, a heterogeneous population of nucleic acids can be partitioned into nucleic acid molecules associated with nucleosomes and nucleic acid molecules devoid of nucleosomes. Alternatively or additionally, a heterogeneous population of nucleic acids may be partitioned into singlestranded DNA (ssDNA) and double-stranded DNA (dsDNA). Alternatively, or additionally, a heterogeneous population of nucleic acids may be partitioned based on nucleic acid length (e.g., molecules of up to 160 bp and molecules having a length of greater than 160 bp).
[0153] In some instances, each partition (representative of a different nucleic acid form) is differentially labelled, and the partitions are pooled together prior to sequencing. In other instances, the different forms are separately sequenced.
[0154] In some embodiments, a population of different nucleic acids is partitioned into two or more different partitions. Each partition is representative of a different nucleic acid form, and a first partition (also referred to as a subsample) comprises DNA with a cytosine modification in a greater proportion than a second subsample. Each partition is distinctly tagged. The first subsample is subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity. The tagged nucleic acids are pooled together prior to sequencing. Sequence reads are obtained and analyzed, including to distinguish the first nucleobase from the second nucleobase in the DNA of the first subsample, in silico. Tags are used to sort reads from different partitions. Analysis to detect genetic variants, a variety of epigenetic marks or functional elements for example TFBS, CTCF binding sites can be performed on a partition-by-partition level, as well as whole nucleic acid population level. ForAttorney Docket No. : GH0224WO example, analysis can include in silico analysis to determine genetic variants, such as CNV, SNV, indel, fusion in nucleic acids in each partition. Additionally, analysis can include in silico analysis to determine transcription factor binding sites (TFBS), mRNA expression, fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, H3S10ph. In some instances, in silico analysis can include determining chromatin structure. For example, coverage of sequence reads can be used to determine nucleosome positioning in chromatin. Higher coverage can correlate with higher nucleosome occupancy in genomic region while lower coverage can correlate with lower nucleosome occupancy or nucleosome depleted region (NDR).
[0155] Samples can include nucleic acids varying in modifications including post-replication modifications to nucleotides and binding, usually noncovalently, to one or more proteins.
[0156] In an embodiment, the population of nucleic acids is one obtained from a serum, plasma or blood sample from a subject suspected of having neoplasia, a tumor, or cancer or previously diagnosed with neoplasia, a tumor, or cancer. The population of nucleic acids includes nucleic acids having varying levels of methylation. Methylation can occur from any one or more postreplication or transcriptional modifications. Post-replication modifications include modifications of the nucleotide cytosine, particularly at the 5-position of the nucleobase, e.g., 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine and 5-carboxylcytosine.
[0157] The affinity agents can be antibodies with the desired specificity, natural binding partners or variants thereof (Bock et al., Nat Biotech 28: 1106-1114 (2010); Song et al., Nat Biotech 29: 68-72 (2011)), or artificial peptides selected e.g., by phage display to have specificity to a given target.
[0158] Examples of capture moieties contemplated herein include methyl binding domain (MBDs) and methyl binding proteins (MBPs) as described herein, including proteins such as MeCP2 and antibodies preferentially binding to 5-methylcytosine.
[0159] Likewise, partitioning of different forms of nucleic acids can be performed using histone binding proteins which can separate nucleic acids bound to histones from free or unbound nucleic acids. Examples of histone binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48 and SANT domain peptides.
[0160] Although for some affinity agents and modifications, binding to the agent may occur in an essentially all or none manner depending on whether a nucleic acid bears a modification,Attorney Docket No. : GH0224WO the separation may be one of degree. In such instances, nucleic acids overrepresented in a modification bind to the agent at a greater extent that nucleic acids underrepresented in the modification. Alternatively, nucleic acids having modifications may bind in an all or nothing manner. But then, various levels of modifications may be sequentially eluted from the binding agent.
[0161] For example, in some embodiments, partitioning can be binary or based on degree / level of modifications. For example, all methylated fragments can be partitioned from unmethylated fragments using methyl -binding domain proteins (e.g., MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific)). Subsequently, additional partitioning may involve eluting fragments having different levels of methylation by adjusting the salt concentration in a solution with the methyl-binding domain and bound fragments. As salt concentration increases, fragments having greater methylation levels are eluted.
[0162] In some instances, the final partitions are representative of nucleic acids having different extents of modifications (over representative or under representative of modifications). Overrepresentation and underrepresentation can be defined by the number of modifications born by a nucleic acid relative to the median number of modifications per strand in a population. For example, if the median number of 5-methylcytosine residues in nucleic acid in a sample is 2, a nucleic acid including more than two 5-methylcytosine residues is overrepresented in this modification and a nucleic acid with 1 or zero 5-methylcytosine residues is underrepresented. The effect of the affinity separation is to enrich for nucleic acids overrepresented in a modification in a bound phase and for nucleic acids underrepresented in a modification in an unbound phase (i.e. in solution). The nucleic acids in the bound phase can be eluted before subsequent processing.
[0163] When using MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific) various levels of methylation can be partitioned using sequential elutions. For example, a hypom ethylated partition (e.g., no methylation) can be separated from a methylated partition by contacting the nucleic acid population with the MBD from the kit, which is attached to magnetic beads. Similarly, histone marks, or TFBS can be can be separated using appropriate antibodies specific to the biochemical properties of each molecule, for example Anti-CTCF, Anti-Pol II (RNA Polymerase II), Anti-H3K4me3 (Histone H3 tri-methylated at lysine 4), Anti-H3K27me3 (Histone H3 tri-methylated at lysine 27), Anti-H3K9me3 (Histone H3 trimethylated at lysine 9), Anti-H3K27ac (Histone H3 acetylated at lysine 27), Anti-H3K4mel (Histone H3 mono-methylated at lysine 4), Anti-H3K36me3 (Histone H3 tri-methylated at lysine 36). The beads are used to separate out the methylated nucleic acids from the nonAttorney Docket No. : GH0224WO methylated nucleic acids. Subsequently, one or more elution steps are performed sequentially to elute nucleic acids having different levels of methylation. For example, a first set of methylated nucleic acids can be eluted at a salt concentration of 160 mM or higher, e.g., at least 150 mM, at least 200 mM, at least 300 mM, at least 400 mM, at least 500 mM, at least 600 mM, at least 700 mM, at least 800 mM, at least 900 mM, at least 1000 mM, or at least 2000 mM. After such methylated nucleic acids are eluted, magnetic separation is once again used to separate higher level of methylated nucleic acids from those with lower level of methylation. The elution and magnetic separation steps can repeat themselves to create various partitions such as a hypomethylated partition (representative of no methylation), a methylated partition (representative of low level of methylation), and a hyper methylated partition (representative of high level of methylation).
[0164] In some methods, nucleic acids bound to an agent used for affinity separation are subjected to a wash step. The wash step washes off nucleic acids weakly bound to the affinity agent. Such nucleic acids can be enriched in nucleic acids having the modification to an extent close to the mean or median (i.e., intermediate between nucleic acids remaining bound to the solid phase and nucleic acids not binding to the solid phase on initial contacting of the sample with the agent).
[0165] The affinity separation results in at least two, and sometimes three or more partitions of nucleic acids with different extents of a modification. While the partitions are still separate, the nucleic acids of at least one partition, and usually two or three (or more) partitions are linked to nucleic acid tags, usually provided as components of adapters, with the nucleic acids in different partitions receiving different tags that distinguish members of one partition from another. The tags linked to nucleic acid molecules of the same partition can be the same or different from one another. But if different from one another, the tags may have part of their code in common so as to identify the molecules to which they are attached as being of a particular partition.
[0166] For further details regarding portioning nucleic acid samples based on characteristics such as methylation, see WO2018 / 119452, which is incorporated herein by reference.
[0167] In some embodiments, the nucleic acid molecules can be fractionated into different partitions based on the nucleic acid molecules that are bound to a specific protein or a fragment thereof and those that are not bound to that specific protein or fragment thereof.
[0168] Nucleic acid molecules can be fractionated based on DNA-protein binding. Protein- DNA complexes can be fractionated based on a specific property of a protein. Examples of such properties include various epitopes, modifications (e.g., histone methylation orAttorney Docket No. : GH0224WO acetylation) or enzymatic activity. Examples of proteins which may bind to DNA and serve as a basis for fractionation may include, but are not limited to, protein A and protein G. Any suitable method can be used to fractionate the nucleic acid molecules based on protein bound regions. Examples of methods used to fractionate nucleic acid molecules based on protein bound regions include, but are not limited to, SDS-PAGE, chromatin-immuno-precipitation (ChIP), heparin chromatography, and asymmetrical field flow fractionation (AF4).
[0169] In certain exemplary embodiments, partitioning of the nucleic acids is performed by contacting the nucleic acids with a methylation binding domain (“MBD”) of a methylation binding protein (“MBP”). MBD binds to 5-methylcytosine (5mC). MBD is coupled to paramagnetic beads, such as Dynabeads® M-280 Streptavidin via a biotin linker. Partitioning into fractions with different extents of methylation can be performed by eluting fractions by increasing the NaCl concentration.
[0170] Examples of MBPs contemplated herein include, but are not limited to: (a) MeCP2 is a protein preferentially binding to 5 -methyl -cytosine over unmodified cytosine, (b) RPL26, PRP8 and the DNA mismatch repair protein MHS6 preferentially bind to 5- hydroxymethyl - cytosine over unmodified cytosine, (c) FOXK1, FOXK2, FOXP1, FOXP4 and FOXI3 preferably bind to 5-formyl-cytosine over unmodified cytosine (lurlaro et al., Genome Biol. 14: R119 (2013)), and (d) Antibodies specific to one or more methylated nucleotide bases.
[0171] In general, elution is a function of number of methylated sites per molecule, with molecules having more methylation eluting under increased salt concentrations. To elute the DNA into distinct populations based on the extent of methylation, one can use a series of elution buffers of increasing NaCl concentration. Salt concentration can range from about 100 nM to about 2500 mM NaCl. In one embodiment, the process results in three (3) partitions. Molecules are contacted with a solution at a first salt concentration and comprising a molecule comprising a methyl binding domain, which molecule can be attached to a capture moiety, such as streptavidin. At the first salt concentration a population of molecules will bind to the MBD and a population will remain unbound. The unbound population can be separated as a “hypomethylated” population. For example, a first partition representative of the hypomethylated form of DNA is that which remains unbound at a low salt concentration, e.g., 100 mM or 160 mM. A second partition representative of intermediate methylated DNA is eluted using an intermediate salt concentration, e.g., between 100 mM and 2000 mM concentration. This is also separated from the sample. A third partition representative of hypermethylated form of DNA is eluted using a high salt concentration, e.g., at least about 2000 mM.Attorney Docket No. : GH0224WO
[0172] Single cell methylation profiling
[0173] Additionally, single cell methylation profiling techniques can be applied to one or more of the methods in the disclosure, including for example, MLAD-seq a technique for single-base resolution and quantitative detection of 5mC in DNA, EAC-seq utilizes engineered proteins for bi sulfite-free, quantitative mapping of 5mC at single-base resolution, Digital-scRRBS a microfluidics-based platform for single-cell methylation sequencing, msRRBS a scalable single-cell reduced representation bisulfite sequencing technology that allows pooling of cellspecific barcoded DNA fragments before bisulfite conversion which improves efficiency and reduces cost.
[0174] DNA methylation data analysis tools
[0175] The following section outlines tools and frameworks suitable for processing bisulfite sequencing or affinity-enriched methylation data obtained from blood, tumor, and / or normal match samples obtained from subjects undergoing cancer treatment. The cancer treatment may comprise a PCV or an off-the-shelf cancer vaccine. The vaccine may be administered to the subject alone or in combination with additional cancer therapies including pembrolizumab as previously described in the disclosure. These tools enable detection of biological changes that occur as a result of cancer, resistance, tumor evolution, and / or treatment, including treatment with a PCV.
[0176] RnBeads is a software tool for large-scale analysis and interpretation of DNA methylation data Msuite: is an analysis toolkit for DNA methylation profiling, specifically optimized for emerging bi sulfite-free methods. methylKit: An R package for the analysis of genome-wide DNA methylation profiles which also supports epigenome-wide association studies and biomarker discovery. BSmooth: Provides alignment, quality control, and analysis pipeline for whole-genome bisulfite sequencing. Similar open-source methods include MethLAB, MethCy and Methylation plotter. Additional methods include BEAT (BS-Seq Epimutation Analysis Toolkit), an R / B ioconductor package for quantitative analysis of DNA methylation from bisulfite sequencing data, utilizing a binomial mixture model. SINBAD, designed for pre-processing, quality assessment, and analysis of single-cell methylation data, starting from multiplexed sequencing reads. CpGtools, a Python package for analyzing DNA methylation data, offering a comprehensive suite for analyzing, annotating, QC, and visualizing the data and MethTools a toolbox for visualizing and analyzing DNA methylation data generated by the bisulfite sequencing.
[0177] Methods to error correct DNA methylation dataAttorney Docket No. : GH0224WO
[0178] Various methods exist to profile methylation status across the genome, including sodium bisulfite conversion and sequencing, differential enzymatic cleavage of DNA, and affinity-mediated capture of methylated DNA; however, these methods have limitations that may lead to misclassification, as the presence of unmethylated cytosines or DNA fragments may be erroneously recognized as methylated, and conversely, methylated cytosines may be detected as unmethylated. Additionally, the resolution of some techniques may not discern methylation variations in specific regions accurately. Therefore, there is a need for methods that can correct DNA methylation data, enhancing both specificity and sensitivity in analysis. The methods described herein may be applied to analyze data from subjects undergoing treatment with a PCV or an off-the-shelf therapeutic cancer vaccine alone or in combination with other cancer therapies.
[0179] In additional aspects, the present disclosure provides methods to filter background noise in methylation data and to error correct DNA methylation data using chromatin interaction data to boost detection of disease associated signals. These methods work by leveraging the congruent methylation states of chromatin interaction partners Xi, Xii, Xiii...Xn, to identify and rectify inaccuracies in the methylation data. In essence, when interaction partners Xi and Xii are observed, they should consistently exhibit the same methylation status.
[0180] To address the noise in the methylation dataset, we construct a probabilistic model that captures the relationship between Xi and Xii and facilitates the removal of inaccuracies by emphasizing concordance and treating discordant instances as indicative of assay introduced errors not true biochemical states. An example of such model can comprise a linear model with an interaction or co-occurrence term. For example, given that Xi and Xii are interaction partners, methylation__status=P0+ l ><Xi +p2*Xii +p3*(Xi xXii)+e Where P3 represents the interaction or co-occurrence effect, indicating that the status of Xi depends on the status of Xii. In this model methylation status is the dependent variable and Xi status and Xii status are independent variables along with an interaction term. Individually, the terms in this linear model represent:
[0181] Methylation status: Is the dependent variable predicted by the model.
[0182] pO:This is the intercept term representing the predicted value of methylation_ status when both Xi and Xii are zero.
[0183] pi*Xi: This term represents the contribution of Xi to the predicted methylation_ status, pi is the coefficient associated the Xi determining the impact of a one-unit change on Xi on the predicted methylation status.Attorney Docket No. : GH0224WO
[0184] P 2*Xii: Similarly, this term represents the contribution of Xii to the predicted methylation_ status. P2 is the coefficient associated with Xii determining the impact of a one- unit change on Xii on the predicted methylation status.
[0185] P 3*(Xi *Xii): is the interaction term, capturing the combined effect of Xi and Xii on the predicted methylation_ status when they interact.
[0186] e: Represents the error or residual in the model, capturing the difference between the predicted methylation status and the actual observed values.
[0187] In essence, this model describes how Xi is influenced by its own value, the value of Xii, and the interaction between them. The coefficients pi, P2, P3 determine the strength and direction of these influences.
[0188] Alternatively, the relationship between the status of Xi, Xii, Xiii. . .Xn and their impact on the target variable methylation status can be represented in a way that is not confined to a specific mathematical form. Exemplary models can include the function Methylation status = f(Xi, Xii)- c Instead of using a predefined linear model, this is a more flexible approach using machine learning. By feeding a diverse set of data points into the model, we enable it to autonomously learn the relationship and dependencies between the features (Xi and Xii) and the target variable(methylation status).
[0189] In this scenario, the model is configured to learn about the congruence between the statuses of Xi and Xii and use this learned relationship to identify and correct discordant methylation states between interaction partners in test data. In this context, training data can comprise a table listing all possible combinations of Xi and Xii such as in Table 1, below:Table 1. Possible combination of Xi and Xii.Attorney Docket No. : GH0224WO
[0190] Additionally, a training data can comprise chromatin interaction data comprising multiple interaction partners and the corrected status. Furthermore, the training data can comprise data with labeled examples of Xi and Xii and their corrected status, specifically tailored to the relationship between the variables.
[0191] Methylation_status = f(Xi, Xii)- c, where:
[0192] Methylation status: represents the underlying relationship between Xi and Xii, it’s the dependent variable to be predicted or output of the model.
[0193] f(Xi, Xii) : represents the function that the machine learning model learns.
[0194] c: is the error term, accounting for any unexplained variability in the target variable that the model does not capture. The error term may also indicate that the observed value of the outcome variable may deviate from the predicted value due to random or unaccounted factors.
[0195] Such general function captures the relationship and interactions between Xi and Xii. Furthermore, this function can accommodate linear relationships, interactions, or more complex patterns based on the characteristics of the data. The target variable Methylation status is determined by a relationship involving independent variables Xi and Xii. In this model, the actual form of the function f is not explicitly defined; it is determined by the learning algorithm based on the patterns in the training data. This allows the model to adapt to the specific characteristics of the data without imposing rigid assumptions about the underlying relationships.
[0196] In some cases, a concordance filter may be sufficient to filter out discordant Xi, Xii, Xiii . . . Xn interaction partners. A concordance filter can be a rule-based approach that leverages the knowledge that Xi and Xii should always have the same status (methylated or unmethylated) to filter out data points where their statuses between the pairs are discordant. A data structure comprising columns Xi status, Xii status, and filtered where:
[0197] Xi status: is a binary value (0 or 1) representing the status of Xi.
[0198] Xii status: is a binary value (0 or 1) representing the status of Xii.
[0199] filtered: is a Boolean flag (True / False) indicating whether the data point is filtered or not (this flag can be initially set to False for all).
[0200] Additionally, a filtering rule can be applied to the chromatin interaction data whereby we iterate through each data point and apply the following rule: If Xi status is equal to Xii status, set filtered to False (keep the data point). If Xi status is not equal to Xii status, set filtered to True (mark the data point as noise). This filtering rule approach assumes a perfect concordance between Xi and Xii, which might not always be true in chromatin interaction data and does not explicitly model the noise in the data collection process. However, assuming aAttorney Docket No. : GH0224WO perfect concordance between Xi and Xii, setting a filtering rule can improve the reliability of methylation assays based on the inherent relationship between one Xi, Xii or multiple interaction partners Xi, Xii, Xiii. . .Xn.
[0201] As described earlier in this disclosure, to capture more complex relationships between Xi, Xii, Xiii. . . .Xn, and other assay driven noise characteristics, machine learning models that learn the patterns in chromatin interaction data may be better suited. Other suitable options include Logistic Regression to predict the methylation status (methylated or unmethylated) of Xi based on both Xi status and Xii status. Alternately, a Decision Tree can be employed to learn a set of rules based on the methylation status or levels of methylation to classify data points as concordant or discordant. In additional embodiments, a Gaussian Mixture Models can be employed to model the joint distribution of Xi status and Xii status (methylation status or methylation level) to identify regions of high concordance and filter outliers.
[0202] RNA extraction and isolation
[0203] In some cases, RNA expression changes may occur in individuals undergoing treatment with a personalized or off-the-shelf cancer vaccine, and these changes may serve as biomarkers of disease status, therapeutic response, or immune activation. RNA-derived biomarkers may be used alone or in combination with other molecular features such as DNA methylation, fragmentomic signatures, or neoantigen levels to assess treatment efficacy or disease progression. In certain embodiments, RNA expression may also act as an indirect measure of neoantigen levels or immune activity targeting neoantigen-expressing cells. In certain embodiments, neoantigen levels may be tracked at the DNA, RNA, or protein level in subjects undergoing treatment, to obtain direct or surrogate measures of antigen burden for assessing or monitoring therapeutic response and disease progression. These data may be combined or analyzed together to confirm or cross-validate neoantigen presence and activity. For example, assessing whether a neoantigen-encoding DNA variant is transcribed into RNA and further translated into protein to obtain more accurate neoantigen levels and biological relevance. In some embodiments of the disclosed methods, the population of target nucleic acids comprises RNA and the method further comprises a cDNA synthesis step. RNA for use in the methods disclosed herein may be isolated from a blood sample or a sample comprising cells (such as a sample that includes immune and / or cancer-derived cells (e.g., a blood sample such as a whole blood sample, a buffy coat sample, a leukapheresis sample, or a peripheral blood PBMC sample)). General methods for RNA extraction and isolation (such as mRNA extraction and isolation) are known in the art and are disclosed in standard textbooks of molecular biology, including Ausubel et al., Current Protocols of Molecular Biology, John Wiley and Sons (1997).Attorney Docket No. : GH0224WOMethods for RNA extraction from paraffin embedded tissues are disclosed, for example, in Rupp and Locker, Lab Invest. 56:A67 (1987), and De Andres et al., BioTechniques 18:42044 (1995). In particular, RNA isolation can be performed using a purification kit, buffer set, and protease(s) from commercial manufacturers, such as PreAnalytix GmbH or Qiagen, according to the manufacturer’s instructions. For example, RNA can be extracted from whole blood samples using the PAXgene® Blood RNA Kit (PreAnalytix GmbH). Other commercially available RNA isolation kits include MasterPure Complete DNA and RNA Purification Kit (EPICENTRE, Madison, WI), and Paraffin Block RNA Isolation Kit (Ambion, Inc.). Total RNA from tissue samples can be isolated using RNA Stat-60 (Tel-Test). RNA prepared from tumor tissue can be isolated, for example, by cesium chloride density gradient centrifugation.
[0204] cDNA library preparation
[0205] Following RNA extraction from a sample (such as a blood sample), a cDNA library is typically prepared in preparation for sequencing, e.g., as in RNA-Seq. In some embodiments, the cDNAs in a library, such as an RNA-Seq library, can comprise a cDNA insert flanked by adapter sequences, such as adapter sequences used for amplification and sequencing on a particular platform. Exemplary cDNA library preparation methods are discussed below; however, cDNA library preparation methods can vary depending on the RNA species under investigation, which can differ in size, sequence, structural features and abundance. One of ordinary skill in the art will be able to select cDNA library preparation methods suitable for cDNA library preparation using an RNA species of interest.
[0206] Some embodiments of the disclosed methods comprise preparing cDNA from RNA (such as RNA extracted from a blood sample), such as by reverse transcription of the RNA template into cDNA. Reverse transcription is generally followed by exponential amplification of the cDNA, e.g., in a PCR reaction. Two commonly used reverse transcriptases are avian myeloblastosis vims reverse transcriptase (AMV-RT) and Moloney murine leukemia virus reverse transcriptase (MMLV-RT). The reverse transcription step is typically primed using specific primers, random hexamers, or oligo-dT primers, depending on the circumstances and the goal of expression profiling. For example, extracted RNA can be reverse transcribed using a Gene Amp RNA PCR kit (Perkin Elmer, Calif., USA), following the manufacturer's instructions. The derived cDNA can then be used as a template in the subsequent amplification (e.g., PCR) reaction. In some embodiments, RNA is converted to cDNA using random priming, followed by second strand synthesis, end repair, and optional A-tailing. Adapters comprising barcodes can then be ligated to the cDNA, which is then amplified.Attorney Docket No. : GH0224WO
[0207] Amplification is typically primed by primers that anneal or bind to primer binding sites in adapters flanking a cDNA molecule to be amplified. Amplification methods can involve cycles of denaturation, annealing and extension, resulting from thermocycling or can be isothermal as in transcription-mediated amplification. Other amplification methods include the ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and self-sustained sequence-based replication.
[0208] Although a PCR step can use a variety of thermostable DNA-dependent DNA polymerases, it typically employs the Taq DNA polymerase. TaqMan® PCR typically utilizes the 5'-nuclease activity of Taq or Tth polymerase to hydrolyze a hybridization probe bound to its target amplicon, but any enzyme with equivalent 5’ nuclease activity can be used. Two oligonucleotide primers are used to generate an amplicon typical of a PCR reaction. A third oligonucleotide, or probe, is designed to detect nucleotide sequence located between the two PCR primers. The probe is non-extendible by Taq DNA polymerase enzyme, and is labeled with a reporter fluorescent dye and a quencher fluorescent dye. Any laser-induced emission from the reporter dye is quenched by the quenching dye when the two dyes are located close together as they are on the probe. During the amplification reaction, the Taq DNA polymerase enzyme cleaves the probe in a template-dependent manner. The resultant probe fragments disassociate in solution, and signal from the released reporter dye is free from the quenching effect of the second fluorophore. One molecule of reporter dye is liberated for each new molecule synthesized, and detection of the unquenched reporter dye provides the basis for quantitative interpretation of the data.
[0209] The primers used for the amplification are selected so as to amplify a unique segment of the gene of interest, such as RNA (such as mRNA) encoding a gene of a target gene set described herein. In some embodiments, expression of other genes is also detected, such as other known disease markers (such as known cancer markers) or housekeeping genes. Primers that can be used to amplify disease-related molecules are commercially available or can be designed and synthesized. In some examples, the primers specifically hybridize to a promoter or promoter region of a disease-related molecule. An alternative quantitative nucleic acid amplification procedure is described in U.S. Pat. No. 5,219,727. In this procedure, the amount of a target sequence in a sample is determined by simultaneously amplifying the target sequence and an internal standard nucleic acid segment. The amount of amplified cDNA from each segment is determined and compared to a standard curve to determine the amount of the target nucleic acid segment that was present in the sample prior to amplification. In some embodiments, the expression of a “housekeeping” gene or “internal control” can also beAttorney Docket No. : GH0224WO evaluated. These terms include any constitutively or globally expressed gene whose presence enables an assessment of mRNA levels provided herein. Such an assessment includes a determination of the overall constitutive level of gene transcription and a control for variations in RNA recovery. Exemplary housekeeping genes include tubulin, glyceraldehyde-3- phosphate-dehydrogenase (GAPDH), beta-actin, and 18S ribosomal RNA.
[0210] rRNA and / or globin mRNA depletion; polv(A) selection
[0211] Ribosomal RNAs (rRNAs) are the most abundant RNA species in most cells. Globin mRNA is also abundant in certain cell types found in the blood. Thus, some embodiments of the present disclosure comprise a step of ribosomal RNA (rRNA) depletion and / or a step of globin mRNA depletion. Such steps can be performed, e.g., following RNA extraction from a sample, and prior to a step of RNA fragmentation or cDNA fragmentation, prior to a step preparing cDNA from the RNA, prior to a step of ligating adapters to the cDNA, and prior to a sequencing step. In some embodiments, the methods include a step of rRNA depletion. In other embodiments, the methods include a step of globin mRNA depletion. In yet other embodiments, the methods disclosed herein include both a step of rRNA depletion and a step of globin mRNA depletion.
[0212] Any suitable rRNA depletion and / or globin mRNA depletion methods are of use in the present disclosure. One approach is to eliminate rRNAs uses sequence-specific probes that can hybridize to rRNAs (Hrdlickova et al., Wiley Interdiscip Rev RNA. 2017;8(l): 10.1002 / wrna.1364). Unwanted rRNAs or their cDNAs are hybridized with biotinylated DNA or locked nucleic acid (LNA) probes, followed by depletion with streptavidin beads. Alternatively, rRNAs can be targeted by anti-sense DNA oligos and digested by RNase H, a method also known as probe-directed degradation (PDD). Another approach for rRNA reduction uses specific, not-so-random (NSR) primers that bind to the RNA molecules of interest during reverse transcription, thus avoiding reverse transcription of the rRNAs. For example, a method known as Ovation RNA-Seq (NuGen) uses hexamer or heptamer primers whose sequences are not present in rRNAs. In addition to sequence-based approaches, some methods take advantage of certain features of rRNAs for their elimination. The COT-hybridization method is based on heat denaturation, re-annealing, and selective degradation by a duplex-specific nuclease (DSN). Double-stranded cDNAs from abundant sequences are preferentially degraded because of their more rapid annealing kinetics compared to less abundant ones. Selective degradation has also been achieved using the enzyme terminator 5 ’-phosphate-dependent exonuclease (TEX), which recognizes RNA molecules with 5 ’-monophosphate, as with rRNAs and tRNAs. Further, commercial kits are available forAttorney Docket No. : GH0224WO rRNA and globin mRNA depletion, including, e.g., the Watchmaker Genomics RNA Library Prep Kit with Polaris Depletion.
[0213] Other embodiments of the present disclosure comprise a step of poly(A) selection. Such a step can be performed, e.g., following RNA extraction from a sample, and prior to a step of RNA fragmentation or cDNA fragmentation, prior to a step preparing cDNA from the RNA, prior to a step of ligating adapters to the cDNA, and prior to a sequencing step. In eukaryotic organisms, most protein coding RNAs (mRNAs) and many long noncoding RNAs (IncRNAs) (>200 nt) comprise a poly(A) tail (“polyadenylated RNAs”). The poly(A) tail may be used to enrich for polyadenylated RNAs from total cellular RNA, in which polyadenylated RNAs may account for approximately 1-5% of total cellular RNA (Hrdlickova et al., Wiley Interdiscip Rev RNA. 2017;8(l): 10.1002 / wrna.l364). Exemplary poly(A) selection methods include, but are not limited to, use of magnetic or cellulose beads coated with oligo-dT molecules. Alternatively, polyadenylated RNAs can be selected using oligo-dT priming for reverse transcription (RT). Poly(A) selection may be combined with globin mRNA depletion.
[0214] Fragmentation
[0215] In some embodiments, methods disclosed herein comprise fragmenting RNA isolated from a sample (such as RNA isolated from a sample comprising cells, such as a whole blood sample, a buffy coat sample, a leukapheresis sample, or a PBMC sample), such as following poly(A) selection or rRNA and / or globin mRNA depletion. RNA fragmentation methods can include physical fragmentation, chemical fragmentation, and / or enzymatic fragmentation. Physical fragmentation methods include, but are not limited to, acoustic or hydrodynamic shearing (such as sonication or point-sink shearing), needle shearing, and nebulization. Enzymatic fragmentation methods can include use of a ribonuclease (such as RNase III). RNA may also be fragmented using chemical shearing methods. Chemical fragmentation methods can include, but are not limited to, heat treatment of RNA in the presence of a divalent metal cation (such as magnesium or zinc). In some embodiments, the fragmenting provides RNA (such as mRNA) fragments of 25-400, 25-300, 25-200, 50-400, 50-300, 50-250, 50-200, 100- 400, 100-300, 100-200, 125-400, 125-300, 125-200, 125-175, 150-400, 150-300, 200-400, 250-400, 300-400, 200-350, 200-300, 225-375, 250-350, or 275-325 base pairs in length.
[0216] Alternatively, non-fragmented RNAs can be reverse transcribed, and the resultant cDNA can be fragmented. cDNA fragmentation methods can include physical fragmentation, chemical fragmentation, and / or enzymatic fragmentation. Physical fragmentation methods include, but are not limited to, acoustic or hydrodynamic shearing (such as sonication or pointsink shearing), needle shearing, and nebulization. Enzymatic fragmentation methods canAttorney Docket No. : GH0224WO include use of a restriction endonuclease (such as a 4-cutter or 5 -cutter restriction endonuclease, e.g., Alul, Dpnl, Eco47I, Haelll, Hpall, Mbo I, Msel, MspI, PspGI, Rsal, Sse9I, or TaqI), a non-specific nuclease (e.g., micrococcal nuclease), or a transposase (for example, when insertion of an adapter into a fragmented double-stranded cDNA molecule is desired). cDNA may also be fragmented using chemical shearing methods. Chemical fragmentation methods can include, but are not limited to, heat digestion of cDNA in the presence of a divalent metal cation (such as magnesium or zinc). In some embodiments, the fragmenting provides cDNA fragments of 25-400, 25-300, 25-200, 50-400, 50-300, 50-250, 50-200, 100-400, 100- 300, 100-200, 125-400, 125-300, 125-200, 125-175, 150-400, 150-300, 200-400, 250-400, 300-400, 200-350, 200-300, 225-375, 250-350, or 275-325 base pairs in length.
[0217] Adapter ligation or addition; tagging
[0218] In certain exemplary embodiments, the disclosed methods comprise adding adapters to DNA (such as cDNA, cell-free DNA, or fragmented genomic DNA). In some embodiments, adapters are added to the DNA before or after subjecting the DNA to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, such as after the subjecting. When adapters are added to the DNA before subjecting the DNA to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, they may comprise nucleotides that are resistant to the procedure. For example, where the procedure comprises contacting the DNA with a deaminase, the adapters may comprise deaminase-resistant cytosines such as 5-ghmC or 5-propynyl cytosine. Similarly, where the procedure comprises contacting the DNA with bisulfite, the adapters may comprise bi sulfite- resistant cytosines such as 5-mC or 5-hmC. In some embodiments, adapters may be added or to DNA concurrently with an amplification procedure, e.g., by providing the adapters in a 5’ portion of a primer (where PCR is used, this can be referred to as library prep-PCR or LP- PCR), before or after an amplification step. In some embodiments, adapters are added by other approaches. In some such methods, first adapters are added to the nucleic acids by ligation to the 3’ ends thereof, which may include ligation to single-stranded DNA. The adapter can be used as a priming site for second-strand synthesis, e.g., using a universal primer and a DNA polymerase. A second adapter can then be ligated to at least the 3’ end of the second strand of the now double-stranded molecule. In some embodiments, the first adapter comprises an affinity tag, such as biotin, and nucleic acid ligated to the first adapter is bound to a solid support (e.g., bead), which may comprise a binding partner for the affinity tag such as streptavidin. For further discussion of a related procedure, see Gansauge et al., Nature Protocols 8:737-748 (2013). Commercial kits for sequencing library preparation compatible with singleAttorney Docket No. : GH0224WO stranded nucleic acids are available, e.g., the Accel-NGS® Methyl-Seq DNA Library Kit from Swift Biosciences. In some embodiments, after adapter ligation, nucleic acids are amplified. In some embodiments, end repair of the DNA is performed prior to addition of adapters.
[0219] In some embodiments, the single-stranded DNA library preparation is performed in a one-step combined phosphorylation / ligation reaction, e.g., as described in Troll et al., BMC Genomics, 20: 1023 (2019), available at https: / / doi.org / 10.1186 / sl2864-019-6355-0. This method, called Single Reaction Single-stranded LibrarY (“SRSLY,”) can be performed without end-polishing. SRSLY may be useful for converting short and fragmented DNA molecules, e.g., cfDNA fragments, into sequencing libraries while retaining native lengths and ends. The SRSLY method can create sequencing libraries (e.g., Illumina sequencing libraries) from fragmented or degraded template (input) DNA. In particular embodiments, template DNA is first heat denatured and then immediately cold shocked to render the template DNA molecules single-stranded. The DNA can be maintained as single-stranded throughout the ligation reaction by the inclusion of a thermostable single-stranded binding protein (SSB). Next, the template DNA, which at this point can be single-stranded and coated with SSB, is placed in a phosphorylation / ligation dual reaction with directional dsDNA NGS adapters that contain single-stranded overhangs. Both the forward and reverse sequencing adapters can share similar structures but differ in which termini is unblocked in order to facilitate proper ligations. Both sequencing adapters can comprise a dsDNA portion and a single-stranded splint overhang of random nucleotides that occurs on the 3 -prime terminus of the bottom strand of the forward adapter and the 5-prime terminus of the bottom strand of the reverse adapter. In this way, the forward adapter (e.g., (P5) Illumina adapter) can delivered to the 5-prime end of template molecules and the reverse adapter (e.g., (P7) Illumina adapter) is delivered to the 3-prime end of template molecules. Thus, the native polarity of input DNA molecules can be retained.
[0220] During the dual phosphorylation / ligation reaction, T4 Polynucleotide Kinase (PNK) can be used to prepare template DNA termini for ligation by phosphorylating 5-prime termini and dephosphorylating 3-prime termini. T4 PNK works on both ssDNA and dsDNA molecules and has no activity on the phosphorylation state of proteins. Simultaneously, the random nucleotides of the splint adapter can be annealed to the single-stranded template molecule. This creates a short, localized dsDNA molecule, enabling ligation of template to adapter with a ligase such as T4 DNA ligase, which has high ligation efficiency on dsDNA templates but low efficiency on ssDNA. After the single phosphorylation / ligation reaction is complete, the library DNA can be, e.g., purified and placed directly into standard NGS indexing PCR, compatible with both traditional single or dual index primers.Attorney Docket No. : GH0224WO
[0221] In some embodiments, following attachment of adapters, the nucleic acids are subject to amplification. The amplification can use, e.g., universal primers that recognize primer binding sites in the adapters.
[0222] In some embodiments, the DNA is linked at both ends to Y-shaped adapters including primer binding sites and tags. In some such embodiments, the DNA is amplified.
[0223] In embodiments of the disclosed methods, a target nucleic acid comprises a 5’ adapter, a 3’ adapter, or both a 5’ adapter and a 3’ adapter. In such embodiments, the 5’ adapter, or both the 5’ adapter and the 3’ adapter comprise at least one sequence that is recognized by at least one restriction enzyme, such as a restriction enzyme described elsewhere herein. In particular embodiments, the 5’ adapter is downstream of an oligonucleotide probe binding site within the target nucleic acid.
[0224] Tagging DNA molecules is a procedure in which a tag is attached to or associated with the DNA molecules. Such tags can be molecules, such as nucleic acids, containing information that indicates a feature of the molecule with which the tag is associated. Tags can allow one to differentiate molecules from which sequence reads originated. For example, molecules can bear a sample tag (which distinguishes molecules in one sample from those in a different sample) or a molecular tag / molecular barcode / barcode (which distinguishes different molecules from one another in both unique and non-unique tagging scenarios). For methods that involve a partitioning step, a partition tag (which distinguishes molecules in one partition from those in a different partition) may be included. In some embodiments, adapters added to DNA molecules comprise tags. In some such embodiments, the tag comprises one or a combination of barcodes. As used herein, the term “barcode” refers to a nucleic acid molecule having a particular nucleotide sequence, or to the nucleotide sequence, itself, depending on context. A barcode can have, for example, between 10 and 100 nucleotides. A collection of barcodes can have degenerate sequences or can have sequences having a certain hamming distance, as desired for the specific purpose. So, for example, a molecular barcode can be comprised of one barcode or a combination of two barcodes, each attached to different ends of a molecule. Additionally or alternatively, for different partitions and / or samples, different sets of molecular barcodes, or molecular tags can be used such that the barcodes serve as a molecular tag through their individual sequences and also serve to identify the partition and / or sample to which they correspond based the set of which they are a member. Tags comprising barcodes can be incorporated into or otherwise joined to adapters. Tags can be incorporated by ligation, overlap extension PCR among other methods.Attorney Docket No. : GH0224WO
[0225] Tagging strategies can be divided into unique tagging and non-unique tagging strategies. In unique tagging, all or substantially all of the molecules in a sample bear a different tag, so that reads can be assigned to original molecules based on tag information alone. Tags used in such methods are sometimes referred to as “unique tags”. In non-unique tagging, different molecules in the same sample can bear the same tag, so that other information in addition to tag information is used to assign a sequence read to an original molecule. Such information may include start and stop coordinate, coordinate to which the molecule maps, start or stop coordinate alone, etc. Tags used in such methods are sometimes referred to as “non-unique tags”. Accordingly, it is not necessary to uniquely tag every molecule in a sample. It suffices to uniquely tag molecules falling within an identifiable class within a sample. Thus, molecules in different identifiable families can bear the same tag without loss of information about the identity of the tagged molecule.
[0226] In some embodiments, the adapters include different tags of sufficient numbers that the number of combinations of tags results in a low probability e.g., 95, 99 or 99.9% of two nucleic acids with the same start and stop points receiving the same combination of tags. Adapters, whether bearing the same or different tags, can include the same or different primer binding sites. In some embodiments, adapters include the same primer binding site.
[0227] In certain embodiments of non-unique tagging, the number of different tags used can be sufficient that there is a very high likelihood (e.g., at least 99%, at least 99.9%, at least 99.99% or at least 99.999% that all molecules of a particular group bear a different tag. In some embodiments comprising barcode attachment, e.g., randomly, to both ends of a molecule, the combination of barcodes, together, constitutes a tag. This number, in term, is a function of the number of molecules falling into the calls. For example, the class may be all molecules mapping to the same start-stop position on a reference genome. The class may be all molecules mapping across a particular genetic locus, e.g., a particular base or a particular region (e.g., up to 100 bases or a gene or an exon of a gene). In certain embodiments, the number of different tags used to uniquely identify a number of molecules, z, in a class can be between any of 2*z, 3*z, 4*z, 5*z, 6*z, 7*z, 8*z, 9*z, 10*z, 11 *z, 12*z, 13*z, 14*z, 15*z, 16*z, 17*z, 18*z, 19*z, 20*z or 100*z (e.g., lower limit) and any of 100,000*z, 10,000*z, 1000*z or 100*z (e.g., upper limit).
[0228] For example, in a sample of about 5 ng to 30 ng of DNA, one expects around 3000 molecules to map to a particular nucleotide coordinate, and between about 3 and 10 molecules having any start coordinate to share the same stop coordinate. Accordingly, about 50 to about 50,000 different tags (e.g., between about 6 and 220 barcode combinations) can suffice toAttorney Docket No. : GH0224WO uniquely tag all such molecules. To uniquely tag all 3000 molecules mapping across a nucleotide coordinate, about 1 million to about 20 million different tags would be required.
[0229] Generally, assignment of unique or non-unique tags barcodes in reactions follows methods and systems described by US patent applications 20010053519, 20030152490, 20110160078, and U.S. Pat. No. 6,582,908 and U.S. Pat. No. 7,537,898 and US Pat. No. 9,598,731. Tags can be linked to sample nucleic acids randomly or non-randomly.
[0230] In some embodiments, the tagged nucleic acids are sequenced after loading into a microwell plate. The microwell plate can have 96, 384, or 1536 microwells. In some cases, they are introduced at an expected ratio of unique tags to microwells. For example, the unique tags may be loaded so that more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 unique tags are loaded per genome sample. In some cases, the unique tags may be loaded so that less than about 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 unique tags are loaded per genome sample. In some cases, the average number of unique tags loaded per sample genome is less than, or greater than, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 unique tags per genome sample
[0231] In some embodiments, 20-50 different tags (e.g., barcodes) are ligated to both ends of target nucleic acids. For example, 35 different tags (e.g., barcodes) ligated to both ends of target molecules creating 35 x 35 permutations, which equals 1225 for 35 tags. Such numbers of tags are sufficient so that different molecules having the same start and stop points have a high probability (e.g., at least 94%, 99.5%, 99.99%, 99.999%) of receiving different combinations of tags. Other barcode combinations include any number between 10 and 500, e.g., about 15x15, about 35x35, about 75x75, about 100x100, about 250x250, about 500x500.
[0232] In some cases, unique tags may be predetermined or random or semi-random sequence oligonucleotides. In other cases, a plurality of barcodes may be used such that barcodes are not necessarily unique to one another in the plurality. In this example, barcodes may be ligated to individual molecules such that the combination of the barcode and the sequence it may be ligated to creates a unique sequence that may be individually tracked. As described herein, detection of non-unique barcodes in combination with sequence data of beginning (start) and end (stop) portions of sequence reads may allow assignment of a unique identity to a particular molecule. The length or number of base pairs, of an individual sequence read may also be used to assign a unique identity to such a molecule. As described herein, fragments from a singleAttorney Docket No. : GH0224WO strand of nucleic acid having been assigned a unique identity, may thereby permit subsequent identification of fragments from the parent strand.
[0233] Tagging of partitions
[0234] In certain exemplary embodiments, two or more partitions, e.g., each partition, is / are differentially tagged. Tags or indexes can be molecules, such as nucleic acids, containing information that indicates a feature of the molecule with which the tag is associated. For example, molecules can bear a sample tag or sample index (which distinguishes molecules in one sample from those in a different sample), a partition tag (which distinguishes molecules in one partition from those in a different partition) and / or a molecular tag / molecular barcode / barcode (which distinguishes different molecules from one another (in both unique and non-unique tagging scenarios). In certain embodiments, a tag can comprise one or a combination of barcodes. As used herein, the term “barcode” refers to a nucleic acid molecule having a particular nucleotide sequence, or to the nucleotide sequence, itself, depending on context. A barcode can have, for example, between 10 and 100 nucleotides. A collection of barcodes can have degenerate sequences or can have sequences having a certain Hamming distance, as desired for the specific purpose. So, for example, a molecular barcode can be comprised of one barcode or a combination of two barcodes, each attached to different ends of a molecule. Additionally or alternatively, for different partitions and / or samples, different sets of molecular barcodes, molecular tags, or molecular indexes can be used such that the barcodes serve as a molecular tag through their individual sequences and also serve to identify the partition and / or sample to which they correspond based the set of which they are a member.
[0235] Tags can be used to label the individual polynucleotide population partitions so as to correlate the tag (or tags) with a specific partition. Alternatively, tags can be used in embodiments of the invention that do not employ a partitioning step. In some embodiments, a single tag can be used to label a specific partition. In some embodiments, multiple different tags can be used to label a specific partition. In embodiments employing multiple different tags to label a specific partition, the set of tags used to label one partition can be readily differentiated for the set of tags used to label other partitions. In some embodiments, the tags may have additional functions, for example the tags can be used to index sample sources or used as unique molecular identifiers (which can be used to improve the quality of sequencing data by differentiating sequencing errors from mutations, for example as in Kinde et al., Proc Nat’l Acad Sci USA 108: 9530-9535 (2011), Kou et al., PLoS 0NE, e0146638 (2016)) or used as non-unique molecule identifiers, for example as described in US Pat. No. 9,598,731. Similarly, in some embodiments, the tags may have additional functions, for example the tagsAttorney Docket No. : GH0224WO can be used to index sample sources or used as non-unique molecular identifiers (which can be used to improve the quality of sequencing data by differentiating sequencing errors from mutations).
[0236] In one embodiment, partition tagging comprises tagging molecules in each partition with a partition tag. After re-combining partitions (e.g., to reduce the number of sequencing runs needed and avoid unnecessary cost) and sequencing molecules, the partition tags identify the source partition. In another embodiment, different partitions are tagged with different sets of molecular tags, e.g., comprised of a pair of barcodes. In this way, each molecular barcode indicates the source partition as well as being useful to distinguish molecules within a partition. For example, a first set of 35 barcodes can be used to tag molecules in a first partition, while a second set of 35 barcodes can be used tag molecules in a second partition.
[0237] In some embodiments, after partitioning and tagging with partition tags, the molecules may be pooled for sequencing in a single run. In some embodiments, a sample tag is added to the molecules, e.g., in a step subsequent to addition of partition tags and pooling. Sample tags can facilitate pooling material generated from multiple samples for sequencing in a single sequencing run.
[0238] Alternatively, in some embodiments, partition tags may be correlated to the sample as well as the partition. As a simple example, a first tag can indicate a first partition of a first sample; a second tag can indicate a second partition of the first sample; a third tag can indicate a first partition of a second sample; and a fourth tag can indicate a second partition of the second sample.
[0239] While tags may be attached to molecules already partitioned based on one or more characteristics for example transcription factor binding sites (TFBS), mRNA expression, fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, H3S10ph, the final tagged molecules in the library may no longer possess that characteristic. In additional examples, while single stranded DNA molecules may be partitioned and tagged, the final tagged molecules in the library are likely to be double stranded. Similarly, while DNA may be subject to partition based on different levels of methylation, in the final library, tagged molecules derived from these molecules are likely to be unmethylated. Accordingly, the tag attached to molecule in the library typically indicatesAttorney Docket No. : GH0224WO the characteristic of the “parent molecule” from which the ultimate tagged molecule is derived, not necessarily to characteristic of the tagged molecule, itself.
[0240] As an example, barcodes 1, 2, 3, 4...n, etc. are used to tag and label molecules in the first partition; barcodes A, B, C, D, etc. are used to tag and label molecules in the second partition; and barcodes a, b, c, d, etc. are used to tag and label molecules in the third partition. Differentially tagged partitions can be pooled prior to sequencing. Differentially tagged partitions can be separately sequenced or sequenced together concurrently, e.g., in the same flow cell of an Illumina sequencer.
[0241] After sequencing, analysis of reads to detect genetic variants can be performed on a partition-by-partition level, as well as a whole nucleic acid population level. Tags are used to sort reads from different partitions. Analysis can include in silico analysis to determine genetic and epigenetic variation (one or more of methylation, chromatin structure, etc.) using sequence information, genomic coordinates length, coverage, and / or copy number. In some embodiments, higher coverage can correlate with higher nucleosome occupancy in genomic region while lower coverage can correlate with lower nucleosome occupancy or a nucleosome depleted region (NDR).
[0242] Alternative Methods of Modified Nucleic Acid Analysis
[0243] In some embodiments the adapters are added to the nucleic acids after partitioning the nucleic acids, in other embodiments the adapters may be added to the nucleic acids prior to partitioning the nucleic acids based on molecular characteristics for example transcription factor binding sites (TFBS), mRNA expression, fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, H3S10ph. In some such methods, a population of nucleic acids bearing the modification to different extents (e.g., 0, 1, 2, 3, 4, or more methyl groups per nucleic acid molecule) is contacted with adapters before fractionation of the population depending on the extent of the modification. Adapters attach to either one end or both ends of nucleic acid molecules in the population. Preferably, the adapters include different tags of sufficient numbers that the number of combinations of tags results in a low probability e.g., 95, 99 or 99.9% of two nucleic acids with the same start and stop points receiving the same combination of tags. Adapters, whether bearing the same or different tags, can include the same or different primer binding sites, but preferably adapters include the same primer binding site. Following attachment of adapters, the nucleic acids are contacted with anAttorney Docket No. : GH0224WO agent that preferentially binds to nucleic acids bearing the modification (such as the previously described such agents). The nucleic acids are partitioned into at least two subsamples differing in the extent to which the nucleic acids bear the modification from binding to the agents. For example, if the agent has affinity for nucleic acids bearing the modification, nucleic acids overrepresented in the modification (compared with median representation in the population) preferentially bind to the agent, whereas nucleic acids underrepresented for the modification do not bind or are more easily eluted from the agent. Following partitioning, the first subsample is subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity. The nucleic acids are then amplified from primers binding to the primer binding sites within the adapters. Following amplification, the different partitions can then be subject to further processing steps, which typically include further (e.g., clonal) amplification, and sequence analysis, in parallel but separately. Sequence data from the different partitions can then be compared.
[0244] In another embodiment, a partitioning scheme can be performed using the following exemplary procedure. Nucleic acids are linked at both ends to Y-shaped adapters including primer binding sites and tags. The molecules are amplified. The amplified molecules are then fractionated by contact with an antibody preferentially binding to 5-methylcytosine to produce two partitions. One partition includes original molecules lacking methylation and amplification copies having lost methylation. The other partition includes original DNA molecules with methylation. The partition including original DNA molecules with methylation is subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity. The two partitions are then processed and sequenced separately with further amplification of the methylated partition. The sequence data of the two partitions can then be compared. In this example, tags are not used to distinguish between methylated and unmethylated DNA but rather to distinguish between different molecules within these partitions so that one can determine whether reads with the same start and stop points are based on the same or different molecules.Attorney Docket No. : GH0224WO
[0245] The disclosure provides further methods for analyzing a population of nucleic acids in which at least some of the nucleic acids include one or more modified cytosine residues, such as 5-methylcytosine and any of the other modifications described previously. In these methods, after partitioning, the sub samples of nucleic acids are contacted with adapters including one or more cytosine residues modified at the 5C position, such as 5-methylcytosine. Preferably all cytosine residues in such adapters are also modified, or all such cytosines in a primer binding region of the adapters are modified. Adapters attach to both ends of nucleic acid molecules in the population. Preferably, the adapters include different tags of sufficient numbers that the number of combinations of tags results in a low probability e.g., 95, 99 or 99.9% of two nucleic acids with the same start and stop points receiving the same combination of tags. The primer binding sites in such adapters can be the same or different, but are preferably the same. After attachment of adapters, the nucleic acids are amplified from primers binding to the primer binding sites of the adapters. The amplified nucleic acids are split into first and second aliquots. The first aliquot is assayed for sequence data with or without further processing. The sequence data on molecules in the first aliquot is thus determined irrespective of the initial methylation state of the nucleic acid molecules. The nucleic acid molecules in the second aliquot are subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, wherein the first nucleobase comprises a cytosine modified at the 5 position, and the second nucleobase comprises unmodified cytosine. This procedure may be bisulfite treatment or another procedure that converts unmodified cytosines to uracils. The nucleic acids subjected to the procedure are then amplified with primers to the original primer binding sites of the adapters linked to nucleic acid. Only the nucleic acid molecules originally linked to adapters (as distinct from amplification products thereof) are now amplifiable because these nucleic acids retain cytosines in the primer binding sites of the adapters, whereas amplification products have lost the methylation of these cytosine residues, which have undergone conversion to uracils in the bisulfite treatment. Thus, only original molecules in the populations, at least some of which are methylated, undergo amplification. After amplification, these nucleic acids are subject to sequence analysis. Comparison of sequences determined from the first and second aliquots can indicate among other things, which cytosines in the nucleic acid population were subject to methylation.
[0246] Such an analysis can be performed using the following exemplary procedure. After partitioning, methylated DNA is linked to Y-shaped adapters at both ends including primer binding sites and tags. The cytosines in the adapters are modified at the 5 position (e.g., 5- methylated). The modification of the adapters serves to protect the primer binding sites in aAttorney Docket No. : GH0224WO subsequent conversion step (e.g., bisulfite treatment, TAP conversion, or any other conversion that does not affect the modified cytosine but affects unmodified cytosine). After attachment of adapters, the DNA molecules are amplified. The amplification product is split into two aliquots for sequencing with and without conversion. The aliquot not subjected to conversion can be subjected to sequence analysis with or without further processing. The other aliquot is subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, wherein the first nucleobase comprises a cytosine modified at the 5 position, and the second nucleobase comprises unmodified cytosine. This procedure may be bisulfite treatment or another procedure that converts unmodified cytosines to uracils. Only primer binding sites protected by modification of cytosines can support amplification when contacted with primers specific for original primer binding sites. Thus, only original molecules and not copies from the first amplification are subjected to further amplification. The further amplified molecules are then subjected to sequence analysis. Sequences can then be compared from the two aliquots. As in the separation scheme discussed above, nucleic acid tags in adapters are not used to distinguish between methylated and unmethylated DNA but to distinguish nucleic acid molecules within the same partition.
[0247] Enriching / Capturing
[0248] Enrichment and capture methods may also be applied to process samples obtained from subjects undergoing treatment with a personalized cancer vaccine (PCV) or an off-the-shelf therapeutic cancer vaccine. Panels may be selected based on their biological relevance, for example, panels may comprise regions of the genome known to harbor tumor-specific mutations, encode neoantigens included in the vaccine, exhibit differential methylation patterns associated with immune activation or resistance, or regulate immune-related pathways. Panels may also comprise biomarkers used to assess therapeutic response, minimal residual disease (MRD), or vaccine-induced epigenetic remodeling, as described throughout the present disclosure. In additional embodiments, panels may comprise regions of the genome harboring one or more of tumor-specific mutations, differentially methylated sites or patterns, and sequences encoding neoantigens used in vaccine design.
[0249] In some embodiments, capture panels can comprise biomarkers such as persistent or re- emerging neoantigens, stable or increasing methylation patterns at DMRs, or evolving fragmentomic signatures may indicate vaccine escape or incomplete immune clearance. These data can guide the design of a modified vaccine formulation that incorporates additional or newly emerging neoantigens, adjusted antigen dosages, or altered delivery strategies to enhance therapeutic efficacy. This integration of real-time longitudinal data to upgrade orAttorney Docket No. : GH0224WO redesign cancer vaccines for the subject during treatment results in a personalized, adaptive vaccination approach, as described throughout the present disclosure.
[0250] In some embodiments, methods disclosed herein comprise a step of capturing one or more sets of target regions of DNA, such as cfDNA. Capture may be performed using any suitable approach known in the art.
[0251] In some embodiments, capturing comprises contacting the DNA to be captured with a set of target-specific probes. The set of target-specific probes may have any of the features described herein for sets of target-specific probes, including but not limited to in the embodiments set forth above and the sections relating to probes below. Capturing may be performed on one or more subsamples prepared during methods disclosed herein. In some embodiments, DNA is captured from at least the first subsample or the second subsample, e.g., at least the first subsample and the second subsample. Where the first subsample undergoes a separation step (e.g., separating DNA originally comprising the first nucleobase (e.g., hmC) from DNA not originally comprising the first nucleobase, such as hmC-seal), capturing may be performed on any, any two, or all of the DNA originally comprising the first nucleobase (e.g., hmC), the DNA not originally comprising the first nucleobase, and the second subsample. In some embodiments, the subsamples are differentially tagged (e.g., as described herein) and then pooled before undergoing capture.
[0252] The capturing step may be performed using conditions suitable for specific nucleic acid hybridization, which generally depend to some extent on features of the probes such as length, base composition, etc. Those skilled in the art will be familiar with appropriate conditions given general knowledge in the art regarding nucleic acid hybridization. In some embodiments, complexes of target-specific probes and DNA are formed.
[0253] In some embodiments, a method described herein comprises capturing cfDNA obtained from a test subject for a plurality of sets of target regions. The target regions comprise epigenetic target regions, which may show differences in methylation levels and / or fragmentation patterns depending on whether they originated from a tumor or from healthy cells. The target regions also comprise sequence-variable target regions, which may show differences in sequence depending on whether they originated from a tumor or from healthy cells. The capturing step produces a captured set of cfDNA molecules, and the cfDNA molecules corresponding to the sequence-variable target region set are captured at a greater capture yield in the captured set of cfDNA molecules than cfDNA molecules corresponding to the epigenetic target region set. For additional discussion of capturing steps, capture yields,Attorney Docket No. : GH0224WO and related aspects, see W02020 / 160414, which is incorporated herein by reference for all purposes.
[0254] In some embodiments, a method described herein comprises contacting cfDNA obtained from a test subject with a set of target-specific probes, wherein the set of targetspecific probes is configured to capture cfDNA corresponding to the sequence-variable target region set at a greater capture yield than cfDNA corresponding to the epigenetic target region set.
[0255] It can be beneficial to capture cfDNA corresponding to the sequence-variable target region set at a greater capture yield than cfDNA corresponding to the epigenetic target region set because a greater depth of sequencing may be necessary to analyze the sequence-variable target regions with sufficient confidence or accuracy than may be necessary to analyze the epigenetic target regions. The volume of data needed to determine fragmentation patterns (e.g., to test fsor perturbation of transcription start sites or CTCF binding sites) or fragment abundance (e.g., in hypermethylated and hypomethylated partitions) is generally less than the volume of data needed to determine the presence or absence of cancer-related sequence mutations. Capturing the target region sets at different yields can facilitate sequencing the target regions to different depths of sequencing in the same sequencing run (e.g., using a pooled mixture and / or in the same sequencing cell).
[0256] In various embodiments, the methods further comprise sequencing the captured cfDNA, e.g., to different degrees of sequencing depth for the epigenetic and sequence-variable target region sets, consistent with the discussion herein.
[0257] In some embodiments, complexes of target-specific probes and DNA are separated from DNA not bound to target-specific probes. For example, where target-specific probes are bound covalently or noncovalently to a solid support, a washing or aspiration step can be used to separate unbound material. Alternatively, where the complexes have chromatographic properties distinct from unbound material (e.g., where the probes comprise a ligand that binds a chromatographic resin), chromatography can be used.
[0258] As discussed in detail elsewhere herein, the set of target-specific probes may comprise a plurality of sets such as probes for a sequence-variable target region set and probes for an epigenetic target region set. In some such embodiments, the capturing step is performed with the probes for the sequence-variable target region set and the probes for the epigenetic target region set in the same vessel at the same time, e.g., the probes for the sequence-variable and epigenetic target region sets are in the same composition. This approach provides a relatively streamlined workflow. In some embodiments, the concentration of the probes for the Sequence-Attorney Docket No. : GH0224WO variable target region set is greater that the concentration of the probes for the epigenetic target region set.
[0259] Alternatively, the capturing step is performed with the sequence-variable target region probe set in a first vessel and with the epigenetic target region probe set in a second vessel, or the contacting step is performed with the sequence-variable target region probe set at a first time and a first vessel and the epigenetic target region probe set at a second time before or after the first time. This approach allows for preparation of separate first and second compositions comprising captured DNA corresponding to the sequence-variable target region set and captured DNA corresponding to the epigenetic target region set. The compositions can be processed separately as desired (e.g., to fractionate based on methylation as described elsewhere herein) and recombined in appropriate proportions to provide material for further processing and analysis such as sequencing.
[0260] In some embodiments, the DNA is amplified. In some embodiments, amplification is performed before the capturing step. In some embodiments, amplification is performed after the capturing step.
[0261] In some embodiments, adapters are included in the DNA. This may be done concurrently with an amplification procedure, e.g., by providing the adapters in a 5’ portion of a primer, e.g., as described above. Alternatively, adapters can be added by other approaches, such as ligation.
[0262] In some embodiments, tags, which may be or include barcodes, are included in the DNA. Tags can facilitate identification of the origin of a nucleic acid. For example, barcodes can be used to allow the origin (e.g., subject) whence the DNA came to be identified following pooling of a plurality of samples for parallel sequencing. This may be done concurrently with an amplification procedure, e.g., by providing the barcodes in a 5’ portion of a primer, e.g., as described above. In some embodiments, adapters and tags / barcodes are provided by the same primer or primer set. For example, the barcode may be located 3’ of the adapter and 5’ of the target-hybridizing portion of the primer. Alternatively, barcodes can be added by other approaches, such as ligation, optionally together with adapters in the same ligation substrate.
[0263] Additional details regarding amplification, tags, and barcodes are discussed in the “General Features of the Methods” section below, which can be combined to the extent practicable with any of the foregoing embodiments and the embodiments set forth in the introduction and summary section.
[0264] Captured setAttorney Docket No. : GH0224WO
[0265] In some embodiments, a captured set of DNA (e.g., cfDNA) is provided. With respect to the disclosed methods, the captured set of DNA may be provided, e.g., by performing a capturing step after a partitioning step as described herein. The capture set may comprise target including loci known to undergo remodeling in response to cancer vaccine treatment. These regions may serve as biomarkers to detect treatment-induced immune activation, resistance mechanisms, or minimal residual disease. Furthermore, the regions that undergo remodeling may include methylation sites exhibiting one or more methylation patterns indicative of disease state, therapeutic response, or resistance to the cancer vaccine. The captured set may comprise DNA corresponding to a sequence-variable target region set, an epigenetic target region set, or a combination thereof. In some embodiments the quantity of captured sequence-variable target region DNA is greater than the quantity of the captured epigenetic target region DNA, when normalized for the difference in the size of the targeted regions (footprint size).
[0266] Alternatively, first and second captured sets may be provided, comprising, respectively, DNA corresponding to a sequence-variable target region set and DNA corresponding to an epigenetic target region set. The first and second captured sets may be combined to provide a combined captured set.
[0267] In some embodiments in which a captured set comprising DNA corresponding to the sequence-variable target region set and the epigenetic target region set includes a combined captured set as discussed above, the DNA corresponding to the sequence-variable target region set may be present at a greater concentration than the DNA corresponding to the epigenetic target region set, e.g., a 1.1 to 1.2-fold greater concentration, a 1.2- to 1.4-fold greater concentration, a 1.4- to 1.6-fold greater concentration, a 1.6- to 1.8-fold greater concentration, a 1.8- to 2.0-fold greater concentration, a 2.0- to 2.2-fold greater concentration, a 2.2- to 2.4- fold greater concentration a 2.4- to 2.6-fold greater concentration, a 2.6- to 2.8-fold greater concentration, a 2.8- to 3.0-fold greater concentration, a 3.0- to 3.5-fold greater concentration, a 3.5- to 4.0, a 4.0- to 4.5-fold greater concentration, a 4.5- to 5.0-fold greater concentration, a 5.0- to 5.5-fold greater concentration, a 5.5- to 6.0-fold greater concentration, a 6.0- to 6.5-fold greater concentration, a 6.5- to 7.0-fold greater, a 7.0- to 7.5-fold greater concentration, a 7.5- to 8.0-fold greater concentration, an 8.0- to 8.5-fold greater concentration, an 8.5- to 9.0-fold greater concentration, a 9.0- to 9.5-fold greater concentration, 9.5- to 10.0-fold greater concentration, a 10- to 11-fold greater concentration, an 11- to 12-fold greater concentration a 12- to 13-fold greater concentration, a 13- to 14-fold greater concentration, a 14- to 15-fold greater concentration, a 15- to 16-fold greater concentration, a 16- to 17-fold greater concentration, a 17- to 18-fold greater concentration, an 18- to 19-fold greater concentration, aAttorney Docket No. : GH0224WO19- to 20-fold greater concentration, a 20- to 30-fold greater concentration, a 30- to 40-fold greater concentration, a 40- to 50-fold greater concentration, a 50- to 60-fold greater concentration, a 60- to 70-fold greater concentration, a 70- to 80-fold greater concentration, a 80- to 90-fold greater concentration, a 90- to 100-fold greater concentration, a 10- to 20-fold greater concentration, a 10- to 40-fold greater concentration, a 10- to 50-fold greater concentration, a 10- to 70-fold greater concentration, or a 10- to 100-fold greater concentration. The degree of difference in concentrations accounts for normalization for the footprint sizes of the target regions, as discussed in the definition section.
[0268] Epigenetic target region set
[0269] The epigenetic target region set may comprise one or more types of target regions likely to differentiate DNA from neoplastic (e.g., tumor or cancer) cells and from healthy cells, e.g., non-neoplastic circulating cells. In additional embodiments the epigenetic target regions may also include loci known to undergo remodeling in response to cancer vaccine treatment. These regions may serve as biomarkers to detect treatment-induced immune activation, resistance mechanisms, or minimal residual disease. Furthermore, the regions that undergo remodeling may include methylation sites exhibiting one or more methylation patterns indicative of disease state, therapeutic response, or resistance to the cancer vaccine. Exemplary types of such regions are discussed in detail herein. The epigenetic target region set may also comprise one or more control regions, e.g., as described herein. In additional embodiments the epigenetic target region set may comprise different forms of nucleic acids may include transcription factor binding sites (TFBS), mRNA expression, fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, H3S10ph.
[0270] In some embodiments, the epigenetic target region set has a footprint of at least 100 kb, e.g., at least 200 kb, at least 300 kb, or at least 400 kb. In some embodiments, the epigenetic target region set has a footprint in the range of 100-1000 kb, e.g., 100-200 kb, 200-300 kb, 300- 400 kb, 400-500 kb, 500-600 kb, 600-700 kb, 700-800 kb, 800-900 kb, and 900-1,000 kb.
[0271] HvDermethylation variable target regions
[0272] In some embodiments, the epigenetic target region set comprises one or more hypermethylation variable target regions. In general, hypermethylation variable target regions refer to regions where an increase in the level of observed methylation, e.g., in a cfDNA sample, indicates an increased likelihood that a sample (e.g., of cfDNA) contains DNAAttorney Docket No. : GH0224WO produced by neoplastic cells, such as tumor or cancer cells. For example, hypermethylation of promoters of tumor suppressor genes has been observed repeatedly. See, e.g., Kang et al., Genome Biol. 18:53 (2017) and references cited therein. In an example, hypermethylation variable target regions can include regions that do not necessarily differ in methylation in cancerous tissue relative to DNA from healthy tissue of the same type, but do differ in methylation (e.g., have more methylation) relative to cfDNA that is typical in healthy subjects. Where, for example, the presence of a cancer results in increased cell death such as apoptosis of cells of the tissue type corresponding to the cancer, such a cancer can be detected at least in part using such hypermethylation variable target regions. In some embodiments, hypermethylation variable target regions include one or more genomic regions, where the cfDNA molecules in those regions do not differ in methylation state in cancer subjects relative to cfDNA from healthy subjects, but the presence / increased quantity of hypermethylated cfDNA in those regions is indicative of a particular tissue type (e.g., cancer origin) and is presented as cfDNA with increased apoptosis (e.g. tumor shedding) into circulation.
[0273] An extensive discussion of methylation variable target regions in colorectal cancer is provided in Lam et al., Biochim Biophys Acta. 1866: 106-20 (2016). These include VIM, SEPT9, ITGA4, OSM4, GATA4 and NDRG4. An exemplary set of hypermethylation variable target regions based on colorectal cancer (CRC) studies is provided in Table 2. Many of these genes likely have relevance to cancers beyond colorectal cancer; for example, TP53 is widely recognized as a critically important tumor suppressor and hypermethylation-based inactivation of this gene may be a common oncogenic mechanism.Table 2. Exemplary Hypermethylation Target Regions based on CRC studies.Attorney Docket No. : GH0224WO
[0274] In some embodiments, the hypermethylation variable target regions comprise a plurality of loci listed in Table 2, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 2. For example, for each locus included as a target region, there may be one or more probes with a hybridization site that binds between the transcription start site and the stop codon (the last stop codon for genes that are alternatively spliced) of the gene, or in the promoter region of the gene. In some embodiments, the one or more probes bind within 300 bp of the transcription start site of a gene in Table 2, e.g., within 200 or 100 bp.
[0275] Methylation variable target regions in various types of lung cancer are discussed in detail, e.g., in Ooki et al., Clin. Cancer Res. 23:7141-52 (2017); Belinksy, Annu. Rev. Physiol. 77:453-74 (2015); Hulbert et al., Clin. Cancer Res. 23: 1998-2005 (2017); Shi et al., BMC Genomics 18:901 (2017); Schneider et al., BMC Cancer. 11 : 102 (2011); Lissa et al., Transl Lung Cancer Res 5(5):492-504 (2016); Skvortsova et al., Br. J. Cancer. 94(10): 1492-1495 (2006); Kim et al., Cancer Res. 61 :3419-3424 (2001); Furonaka et al., Pathology International 55:303-309 (2005); Gomes et al., Rev. Port. Pneumol. 20:20-30 (2014); Kim et al., Oncogene. 20: 1765-70 (2001); Hopkins-Donaldson et al., Cell Death Differ. 10:356-64 (2003); Kikuchi et al., Clin. Cancer Res. 11 :2954-61 (2005); Heller et al., Oncogene 25:959-968 (2006); Licchesi et al., Carcinogenesis. 29:895-904 (2008); Guo et al., Clin. Cancer Res. 10:7917-24 (2004); Palmisano et al., Cancer Res. 63:4620-4625 (2003); and Toyooka et al., Cancer Res. 61 :4556-4560, (2001).
[0276] An exemplary set of hypermethylation variable target regions based on lung cancer studies is provided in Table 3. Many of these genes likely have relevance to cancers beyond lung cancer; for example, Casp8 (Caspase 8) is a key enzyme in programmed cell death and hypermethylation-based inactivation of this gene may be a common oncogenic mechanism notAttorney Docket No. : GH0224WO limited to lung cancer. Additionally, a number of genes appear in both Tables 2 and 3, indicating generality.Table 3. Exemplary Hypermethylation Target Regions based on Lung Cancer studies
[0277] Any of the foregoing embodiments concerning target regions identified in Table 3 may be combined with any of the embodiments described above concerning target regions identified in Table 2. In some embodiments, the hypermethylation variable target regions comprise a plurality of loci listed in Table 2 or Table 3, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 2 or Table 3.
[0278] Additional hypermethylation target regions may be obtained, e.g., from the Cancer Genome Atlas. Kang et al., Genome Biology 18:53 (2017), describe construction of aAttorney Docket No. : GH0224WO probabilistic method called CancerLocator using hypermethylation target regions from breast, colon, kidney, liver, and lung. In some embodiments, the hypermethylation target regions can be specific to one or more types of cancer. Accordingly, in some embodiments, the hypermethylation target regions include one, two, three, four, or five subsets of hypermethylation target regions that collectively show hypermethylation in one, two, three, four, or five of breast, colon, kidney, liver, and lung cancers.
[0279] In yet additional embodiments the, epigenetic target region set comprises transcription factor binding sites (TFBS), mRNA expression, fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, H3S10ph. This approach can be used to determine, for example, whether certain sequences are hypermethylated or hypomethylated or contain a specific histone modification or functional element such as a TFBS.
[0280] Hypomethylation variable target regions
[0281] Global hypomethylation is a commonly observed phenomenon in various cancers. See, e.g., Hon et al., Genome Res. 22:246-258 (2012) (breast cancer); Ehrlich, Epigenomics 1:239- 259 (2009) (review article noting observations of hypomethylation in colon, ovarian, prostate, leukemia, hepatocellular, and cervical cancers). For example, regions such as repeated elements, e.g., LINE1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats, and satellite DNA, and intergenic regions that are ordinarily methylated in healthy cells may show reduced methylation in tumor cells. Accordingly, in some embodiments, the epigenetic target region set includes hypomethylation variable target regions, where a decrease in the level of observed methylation indicates an increased likelihood that a sample (e.g., of cfDNA) contains DNA produced by neoplastic cells, such as tumor or cancer cells. In additional embodiments, the hypomethylated regions comprise hypomethylationvariable regions, where a decrease in methylation level is indicative of the presence of DNA derived from neoplastic cells, such as tumor or cancer cells, in a sample (e.g., cfDNA). In certain embodiments, changes in methylation at such regions may also correlate with treatment response, resistance, or minimal residual disease in subjects treated with a cancer vaccine. In an example, hypomethylation variable target regions can include regions that do not necessarily differ in methylation state in cancerous tissue relative to DNA from healthy tissue of the same type, but do differ in methylation (e.g., are less methylated) relative to cfDNA that is typical inAttorney Docket No. : GH0224WO healthy subjects. Where, for example, the presence of a cancer results in increased cell death such as apoptosis of cells of the tissue type corresponding to the cancer, such a cancer can be detected at least in part using such hypomethylation variable target regions. In some embodiments, hypomethylation variable target regions include one or more genomic regions, where the cfDNA molecules in those regions do not differ in methylation state in cancer subjects relative to cfDNA from healthy subjects, but the presence / increased quantity of hypomethylated cfDNA in those regions is indicative of a particular tissue type (e.g., cancer origin) and is presented as cfDNA with increased apoptosis (e.g. tumor shedding) into circulation.
[0282] In some embodiments, hypomethylation variable target regions include repeated elements and / or intergenic regions. In some embodiments, repeated elements include one, two, three, four, or five of LINE1 elements, Alu elements, centromeric tandem repeats, peri centromeric tandem repeats, and / or satellite DNA.
[0283] Exemplary specific genomic regions that show cancer-associated hypomethylation include nucleotides 8403565-8953708 and 151104701-151106035 of human chromosome 1. In some embodiments, the hypomethylation variable target regions overlap or comprise one or both of these regions.
[0284] CTCF binding sites
[0285] CTCF is a DNA-binding protein that contributes to chromatin organization and often colocalizes with cohesin. Perturbation of CTCF binding sites has been reported in a variety of different cancers. See, e.g., Katainen et al., Nature Genetics, doi: 10.1038 / ng.3335, published online 8 June 2015; Guo et al., Nat. Commun. 9: 1520 (2018). CTCF binding results in recognizable patterns in cfDNA that can be detected by sequencing, e.g., through fragment length analysis. Details regarding sequencing-based fragment length analysis are provided in Snyder et al., Cell 164:57-68 (2016); WO 2018 / 009723; and US20170211143A1, each of which are incorporated herein by reference.
[0286] Thus, perturbations of CTCF binding result in variation in the fragmentation patterns of cfDNA. As such, CTCF binding sites represent a type of fragmentation variable target regions.
[0287] There are many known CTCF binding sites. See, e.g., the CTCFBSDB (CTCF Binding Site Database), available on the Internet at insulatordb.uthsc.edu / ; Cuddapah et al., Genome Res. 19:24-32 (2009); Martin et al., Nat. Struct. Mol. Biol. 18:708-14 (2011); Rhee et al., Cell. 147: 1408-19 (2011), each of which are incorporated by reference. Exemplary CTCF bindingAttorney Docket No. : GH0224WO sites are at nucleotides 56014955-56016161 on chromosome 8 and nucleotides 95359169- 95360473 on chromosome 13.
[0288] Accordingly, in some embodiments, the epigenetic target region set includes CTCF binding regions. In some embodiments, the CTCF binding regions comprise at least 10, 20, 50, 100, 200, or 500 CTCF binding regions, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500- 1000 CTCF binding regions, e.g., such as CTCF binding regions described above or in one or more of CTCFBSDB or the Cuddapah et al., Martin et al., or Rhee et al. articles cited above.
[0289] In some embodiments, at least some of the CTCF sites can be methylated or unmethylated, wherein the methylation state is correlated with the whether or not the cell is a cancer cell. In some embodiments, the epigenetic target region set comprises at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, at least 1000 bp upstream and downstream regions of the CTCF binding sites.
[0290] Transcriptional start sites
[0291] Transcriptional start sites may also show perturbations in neoplastic cells and / or in cells or cell free DNA obtained from subjects treated with a cancer vaccine. For example, nucleosome organization at various transcription start sites in healthy cells of the hematopoietic lineage — which contributes substantially to cfDNA in healthy individuals — may differ from nucleosome organization at those transcription start sites in neoplastic cells. This results in different cfDNA patterns that can be detected by sequencing, as discussed generally in Snyder et al., Cell 164:57-68 (2016); WO 2018 / 009723; and US20170211143A1. In another example, transcription start sites that do not necessarily differ epigenetically in cancerous tissue relative to DNA from healthy tissue of the same type, but do differ epigenetically (e.g., with respect to nucleosome organization) relative to cfDNA that is typical in healthy subjects. Where, for example, the presence of a cancer results in increased cell death such as apoptosis of cells of the tissue type corresponding to the cancer, such a cancer can be detected at least in part using such transcription start sites.
[0292] Thus, perturbations of transcription start sites in response to cancer vaccine treatment, resistance, and / or tumor evolutions may also result in variation in the fragmentation patterns of cfDNA. As such, transcription start sites also represent a type of fragmentation variable target regions.
[0293] Human transcriptional start sites are available from DBTSS (DataBase of Human Transcription Start Sites), available on the Internet at dbtss.hgc.jp and described in Yamashita et al., Nucleic Acids Res. 34(Database issue): D86-D89 (2006), which is incorporated herein by reference.Attorney Docket No. : GH0224WO
[0294] Accordingly, in some embodiments, the epigenetic target region set and / or regions associated with cancer vaccine treatment and / or treatment response includes transcriptional start sites. In some embodiments, the transcriptional start sites comprise at least 10, 20, 50, 100, 200, or 500 transcriptional start sites, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 transcriptional start sites, e.g., such as transcriptional start sites listed in DBTSS. In some embodiments, at least some of the transcription start sites can be methylated or unmethylated, wherein the methylation state is correlated with disease, therapeutic response, resistance and / or MRD in subjects treated with a cancer vaccine. In some embodiments, the epigenetic target region set comprises at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, at least 1000 bp upstream and downstream regions of the transcription start sites of genes. These genes may comprise genes involved in immune regulation, tumor suppression, oncogenic signaling, antigen presentation, cytokine response, and genes encoding neoantigens or their regulatory elements, such as IFNG, IL2, TNF, PDCD1, CTLA4, HL A- A, HLA-B, TAPI, B2M, CD8A, GZMB, and CD274 (PD-L1).
[0295] Focal amplifications
[0296] Although focal amplifications are somatic mutations, they can be detected by sequencing based on read frequency in a manner analogous to approaches for detecting certain epigenetic changes such as changes in methylation. As such, regions that may show focal amplifications in cancer can be included in the epigenetic target region set and may comprise one or more of AR, BRAF, CCND1, CCND2, CCNE1, CDK4, CDK6, EGFR, ERBB2, FGFR1, FGFR2, KIT, KRAS, MET, MYC, PDGFRA, PIK3CA, and RAFI. For example, in some embodiments, the epigenetic target region set comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 of the foregoing targets. In yet other embodiments, regions that may show focal amplifications in response to cancer vaccine treatment may comprise genes involved in immune regulation, tumor suppression, oncogenic signaling, antigen presentation, cytokine response, and genes encoding neoantigens or their regulatory elements, such as IFNG, IL2, TNF, PDCD1, CTLA4, HLA-A, HLA-B, TAPI, B2M, CD8A, GZMB, and CD274 (PD- Ll).
[0297] Methylation control regions
[0298] It can be useful to include control regions to facilitate data validation. In some embodiments, the epigenetic target region set includes control regions that are expected to be methylated or unmethylated in essentially all samples, regardless of whether the DNA is derived from a cancer cell or a normal cell. In some embodiments, the epigenetic target region set includes control hypomethylated regions that are expected to be hypomethylated inAttorney Docket No. : GH0224WO essentially all samples. In some embodiments, the epigenetic target region set includes control hypermethylated regions that are expected to be hypermethylated in essentially all samples.
[0299] Sequence-variable target region set
[0300] In some embodiments, the sequence-variable target region set comprises a plurality of regions known to undergo somatic mutations in cancer.
[0301] In some aspects, the sequence-variable target region set targets a plurality of different genes or genomic regions (“panel”) selected such that a determined proportion of subjects having a cancer exhibits a genetic variant or tumor marker in one or more different genes or genomic regions in the panel. The panel may be selected to limit a region for sequencing to a fixed number of base pairs. The panel may be selected to sequence a desired amount of DNA, e.g., by adjusting the affinity and / or amount of the probes as described elsewhere herein. The panel may be further selected to achieve a desired sequence read depth. The panel may be selected to achieve a desired sequence read depth or sequence read coverage for an amount of sequenced base pairs. The panel may be selected to achieve a theoretical sensitivity, a theoretical specificity, and / or a theoretical accuracy for detecting one or more genetic variants in a sample.
[0302] Probes for detecting the panel of regions can include those for detecting genomic regions of interest (hotspot regions) as well as nucleosome-aware probes (e.g., KRAS codons 12 and 13) and may be designed to optimize capture based on analysis of cfDNA coverage and fragment size variation impacted by nucleosome binding patterns and GC sequence composition. Regions used herein can also include non-hotspot regions optimized based on nucleosome positions and GC models.
[0303] Examples of listings of genomic locations of interest may be found in Table 4 and Table 5. In some embodiments, a sequence-variable target region set used in the methods of the present disclosure comprises at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the genes of Table 4. In some embodiments, a sequence-variable target region set used in the methods of the present disclosure comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the SNVs of Table 4. In some embodiments, a sequence-variable target region set used in the methods of the present disclosure comprises at least 1, at least 2, at least 3, at least 4, at least 5, or 6 of the fusions of Table 4. In some embodiments, a sequence-variable target region set used in the methods of the present disclosure comprise at least a portion of at least 1, at least 2, or 3 of the indels of Table 4. In some embodiments, a sequence-variableAttorney Docket No. : GH0224WO target region set used in the methods of the present disclosure comprises at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the genes of Table 5. In some embodiments, a sequence-variable target region set used in the methods of the present disclosure comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the SNVs of Table 5. In some embodiments, a sequence-variable target region set used in the methods of the present disclosure comprises at least 1, at least 2, at least 3, at least 4, at least 5, or 6 of the fusions of Table 5. In some embodiments, a sequence-variable target region set used in the methods of the present disclosure comprises at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 of the indels of Table 5. Each of these genomic locations of interest may be identified as a backbone region or hot-spot region for a given panel. An example of a listing of hot-spot genomic locations of interest may be found in Table 6. In some embodiments, a sequence-variable target region set used in the methods of the present disclosure comprises at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 of the genes of Table 6. Each hot-spot genomic region is listed with several characteristics, including the associated gene, chromosome on which it resides, the start and stop position of the genome representing the gene’s locus, the length of the gene’s locus in base pairs, the exons covered by the gene, and the critical feature (e.g., type of mutation) that a given genomic region of interest may seek to capture.Table 4. Certain Genomic Regions of InterestAttorney Docket No. : GH0224WOTable 5. Certain Genomic Regions of InterestTable 6. Certain Genomic Regions of InterestAttorney Docket No. : GH0224WOAttorney Docket No. : GH0224WO
[0304] Additionally or alternatively, suitable target region sets are available from the literature. For example, Gale et al., PLoS One 13: e0194630 (2018), which is incorporated herein by reference, describes a panel of 35 cancer-related gene targets that can be used as part or all of a sequence-variable target region set. These 35 targets are AKT1, ALK, BRAF, CCND1, CDK2A, CTNNB1, EGFR, ERBB2, ESRI, FGFR1, FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11, TP53, and U2AF1.
[0305] In some embodiments, the sequence-variable target region set comprises target regions from at least 10, 20, 30, or 35 cancer-related genes, such as the cancer-related genes listed above.
[0306] Sequencing
[0307] In general, sample nucleic acids flanked by adapters with or without prior amplification can be subject to sequencing. Sequencing methods include, for example, Sanger sequencing, high-throughput sequencing, pyrosequencing, sequencing-by-synthesis, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by-hybridization, Digital Gene Expression (Helicos), Next generation sequencing (NGS), Single Molecule Sequencing by Synthesis (SMSS) (Helicos), massively-parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking, and sequencing using PacBio, SOLiD, Ion Torrent, or Nanopore platforms. Sequencing reactions can be performed in a variety of sample processing units, which may multiple lanes, multiple channels, multiple wells, or other mean of processing multiple sample sets substantially simultaneously. Sample processing unit can also include multiple sample chambers to enable processing of multiple runs simultaneously. Additionally, sequencing chemistries can include Illumina sequencing (MiSeq, HiSeq, NextSeq, NovaSeq, MiniSeq, iSeq 100), Oxford Nanopore sequencingAttorney Docket No. : GH0224WO(MinlON, GridlON, PromethlON, Flongle), Sanger sequencing (ABI 3730x1), Ion Torrent sequencing (Ion PGM, Ion S5, Ion GeneStudio S5), PacBio SMRT sequencing (Sequel, Sequel lie), Illumina NovaSeq 6000, PacBio HiFi sequencing (Sequel II and Sequel lie Systems using HiFi chemistry), Element Biosciences (AVITI System), Ultima Genomics (Ultima Sequencing), and Singular Genomics (G4 Sequencing System).
[0308] The sequencing reactions can be performed on one or more forms of nucleic acids at least one of which is known to contain markers of cancer or of other disease. The sequencing reactions can also be performed on any nucleic acid fragments present in the sample. In some embodiments, sequence coverage of the genome may be less than 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100%. In some embodiments, the sequence reactions may provide for sequence coverage of at least 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, or 80% of the genome. Sequence coverage can performed on at least 5, 10, 20, 70, 100, 200 or 500 different genes, or at most 5000, 2500, 1000, 500 or 100 different genes.
[0309] Simultaneous sequencing reactions may be performed using multiplex sequencing. In some cases, cell-free nucleic acids may be sequenced with at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. In other cases cell-free nucleic acids may be sequenced with less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. Sequencing reactions may be performed sequentially or simultaneously. Subsequent data analysis may be performed on all or part of the sequencing reactions. In some cases, data analysis may be performed on at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. In other cases, data analysis may be performed on less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. An exemplary read depth is 1000-50000 reads per locus (base).
[0310] Differential depth of sequencing
[0311] In some embodiments, sequencing depth may intentionally varied across different target region sets to optimize detection sensitivity while saving sequencing cost. For example, regions such as those harboring tumor-specific mutations and / or encoding neoantigens may be sequenced at greater depth than epigenetic target regions to enable more accurate disease classification variant calling, allele frequency estimation, and / or neoantigen quantification.
[0312] Sequencing coverage may be adjusted across different target regions depending on molecular features, the biological relevance, and the required sensitivity. For example, higher coverage may be applied to certain regions including those containing low-frequencyAttorney Docket No. : GH0224WO mutations, subclonal variants, or subtle methylation changes to improve detection sensitivity. Conversely, other regions such as broadly altered methylation domains or stable fragmentomic patterns may be adequately characterized with lower coverage. This flexibility in sequencing depth may be particularly important when redesigning or boosting a cancer vaccine, as more sensitive detection of emerging mutations, evolving neoantigens, or epigenetic shifts may inform the selection of additional targets or modifications to the vaccine formulation.
[0313] Higher sequencing coverage at these loci improves the ability to detect low-frequency mutations or subclonal variants that may influence vaccine design, therapeutic response, and / or resistance. In contrast, epigenetic regions, including DMRs or chromatin state markers, may require lower coverage to generate reliable methylation or accessibility profiles. Allocation of sequencing depth across region types helps balance sensitivity, cost, and data yield, particularly in workflows designed to monitor subjects treated with personalized or off-the-shelf cancer vaccines.
[0314] In some embodiments, nucleic acids corresponding to the sequence-variable target region set are sequenced to a greater depth of sequencing than nucleic acids corresponding to the epigenetic target region set. For example, the depth of sequencing for nucleic acids corresponding to the sequence variant target region set may be at least 1.25-, 1.5-, 1.75-, 2-, 2.25-, 2.5-, 2.75-, 3-, 3.5-, 4-, 4.5-, 5-, 6-, 7-, 8-, 9-, 10-, 11-, 12-, 13-, 14-, or 15-fold greater, or 1.25- to 1.5-, 1.5- to 1.75-, 1.75- to 2-, 2- to 2.25-, 2.25- to 2.5-, 2.5- to 2.75-, 2.75- to 3-, 3- to 3.5-, 3.5- to 4-, 4- to 4.5-, 4.5- to 5-, 5- to 5.5-, 5.5- to 6-, 6- to 7-, 7- to 8-, 8- to 9-, 9- to 10-, 10- to 11-, 11- to 12-, 13- to 14-, 14- to 15-fold, or 15- to 100-fold greater, than the depth of sequencing for nucleic acids corresponding to the epigenetic target region set. In some embodiments, said depth of sequencing is at least 2-fold greater. In some embodiments, said depth of sequencing is at least 5-fold greater. In some embodiments, said depth of sequencing is at least 10-fold greater. In some embodiments, said depth of sequencing is 4- to 10-fold greater. In some embodiments, said depth of sequencing is 4- to 100-fold greater. Each of these embodiments refer to the extent to which nucleic acids corresponding to the sequence-variable target region set are sequenced to a greater depth of sequencing than nucleic acids corresponding to the epigenetic target region set.
[0315] In some embodiments, the captured cfDNA corresponding to the sequence-variable target region set and the captured cfDNA corresponding to the epigenetic target region set are sequenced concurrently, e.g., in the same sequencing cell (such as the flow cell of an Illumina sequencer) and / or in the same composition, which may be a pooled composition resulting from recombining separately captured sets or a composition obtained by capturing the cfDNAAttorney Docket No. : GH0224WO corresponding to the sequence-variable target region set and the captured cfDNA corresponding to the epigenetic target region set in the same vessel.
[0316] Analysis
[0317] In some embodiments, a method described herein comprises identifying the presence or absence of DNA produced by a tumor (or neoplastic cells, or cancer cells).
[0318] The present methods can be used to diagnose presence or absence of conditions, particularly cancer, in a subject, to characterize conditions (e.g., staging cancer or determining heterogeneity of a cancer), monitor response to treatment of a condition, effect prognosis risk of developing a condition or subsequent course of a condition. The present disclosure can also be useful in determining the efficacy of a particular treatment option. Successful treatment options may increase the amount of copy number variation or rare mutations detected in subject's blood if the treatment is successful as more cancers may die and shed DNA. In other examples, this may not occur. In another example, perhaps certain treatment options may be correlated with genetic profiles of cancers over time. This correlation may be useful in selecting a therapy.
[0319] Additionally, if a cancer is observed to be in remission after treatment, the present methods can be used to monitor residual disease or recurrence of disease.
[0320] The types and number of cancers that may be detected may include blood cancers, brain cancers, lung cancers, skin cancers, nose cancers, throat cancers, liver cancers, bone cancers, lymphomas, pancreatic cancers, skin cancers, bowel cancers, rectal cancers, thyroid cancers, bladder cancers, kidney cancers, mouth cancers, stomach cancers, solid state tumors, heterogeneous tumors, homogenous tumors and the like. Type and / or stage of cancer can be detected from genetic variations including mutations, rare mutations, indels, copy number variations, transversions, translocations, inversion, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structure alterations, gene fusions, chromosome fusions, gene truncations, gene amplification, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5 -methylcytosine.
[0321] Genetic data can also be used for characterizing a specific form of cancer. Cancers are often heterogeneous in both composition and staging. Genetic profile data may allow characterization of specific sub-types of cancer that may be important in the diagnosis or treatment of that specific sub-type. This information may also provide a subject or practitioner clues regarding the prognosis of a specific type of cancer and allow either a subject or practitioner to adapt treatment options in accord with the progress of the disease. Some cancersAttorney Docket No. : GH0224WO can progress to become more aggressive and genetically unstable. Other cancers may remain benign, inactive or dormant. The system and methods of this disclosure may be useful in determining disease progression.
[0322] Further, the methods of the disclosure may be used to characterize the heterogeneity of an abnormal condition in a subject. Such methods can include, e.g., generating a genetic profile of extracellular polynucleotides derived from the subject, wherein the genetic profile comprises a plurality of data resulting from copy number variation and rare mutation analyses. In some embodiments, an abnormal condition is cancer. In some embodiments, the abnormal condition may be one resulting in a heterogeneous genomic population. In the example of cancer, some tumors are known to comprise tumor cells in different stages of the cancer. In other examples, heterogeneity may comprise multiple foci of disease. Again, in the example of cancer, there may be multiple tumor foci, perhaps where one or more foci are the result of metastases that have spread from a primary site.
[0323] The present methods can be used to generate or profile, fingerprint or set of data that is a summation of genetic information derived from different cells in a heterogeneous disease. This set of data may comprise copy number variation, epigenetic variation, and mutation analyses alone or in combination.
[0324] The present methods can be used to diagnose, prognose, monitor or observe cancers, or other diseases. In some embodiments, the methods herein do not involve the diagnosing, prognosing or monitoring a fetus and as such are not directed to non-invasive prenatal testing. In other embodiments, these methodologies may be employed in a pregnant subject to diagnose, prognose, monitor or observe cancers or other diseases in an unborn subject whose DNA and other polynucleotides may co-circulate with maternal molecules.
[0325] An exemplary method for molecular tag identification of MBD-bead partitioned libraries through NGS which includes a step of subjecting the first subsample to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample is as follows:
[0326] 1. Physical partitioning of an extracted DNA sample (e.g., extracted blood plasma DNA from a human sample, which has optionally been subjected to target capture as described herein) using a methyl -binding domain protein-bead purification kit, saving all elutions from process for downstream processing.
[0327] 2. Parallel application of differential molecular tags and NGS-enabling adapter sequences to each partition. For example, the hypermethylated, residual methylation ('wash'), and hypomethylated partitions are ligated with NGS- adapters with molecular tags.Attorney Docket No. : GH0224WO
[0328] 3. Subject hypermethylated partition to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, such as any of those described herein.
[0329] 4. Re-combining all molecular tagged partitions, and subsequent amplification using adapter-specific DNA primer sequences.
[0330] 5. Capture / hybridization of re-combined and amplified total library, targeting genomic regions of interest (e.g., cancer-specific genetic variants and differentially methylated regions).
[0331] 6. Re-amplification of the captured DNA library, appending a sample tag. Different samples are pooled, and assayed in multiplex on an NGS instrument.
[0332] 7. Bioinformatics analysis of NGS data, with the molecular tags being used to identify unique molecules, as well deconvolution of the sample into molecules that were differentially MBD-partitioned. This analysis can yield information on relative 5-methylcytosine for genomic regions, concurrent with standard genetic sequencing / variant detection.
[0333] In some embodiments of methods described herein, including but not limited to the method shown above, the molecular tags consist of nucleotides that are not altered by the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, such as any of those described herein (e.g., mC along with A, T, and G where the procedure is bisulfite conversion or any other conversion that does not affect mC; hmC along with A, T, and G where the procedure is a conversion that does not affect hmC; etc.). In some embodiments of methods described herein, including but not limited to the method shown above, the molecular tags do not comprise nucleotides that are altered by the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, such as any of those described herein (e.g., the tags do not comprise unmodified C where the procedure is bisulfite conversion or any other conversion that affects C; the tags do not comprise mC where the procedure is a conversion that affects mC; the tags do not comprise hmC where the procedure is a conversion that affects hmC; etc.).
[0334] In general, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA may instead be performed before the step of parallel application of differential molecular tags and NGS-enabling adapter sequences to each partition. For example, this may be done where the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA is a separation, such as hmC-seal, and in such a case the separated populations may themselves be differentially tagged relative to each other. Such an exemplary method is as follows:
[0335] 1. Physical partitioning of an extracted DNA sample (e.g., extracted blood plasma DNA from a human sample, which has optionally been subjected to target capture as describedAttorney Docket No. : GH0224WO herein) using a methyl -binding domain protein-bead purification kit, saving all elutions from process for downstream processing.
[0336] 2. Subject hypermethylated partition to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, such as any of those described herein.
[0337] 3. Parallel application of differential molecular tags and NGS-enabling adapter sequences to each partition. For example, the hypermethylated partition (or where applicable, two or more sub-partitions of the hypermethylated partition), residual methylation ('wash') partition, and hypomethylated partition are ligated with NGS- adapters with molecular tags.
[0338] 4. Re-combining all molecular tagged partitions, and subsequent amplification using adapter-specific DNA primer sequences.
[0339] 5. Capture / hybridization of re-combined and amplified total library, targeting genomic regions of interest (e.g., cancer-specific genetic variants and differentially methylated regions).
[0340] 6. Re-amplification of the captured DNA library, appending a sample tag. Different samples are pooled, and assayed in multiplex on an NGS instrument.
[0341] 7 Bioinformatics analysis of NGS data, with the molecular tags being used to identify unique molecules, as well deconvolution of the sample into molecules that were differentially MBD-partitioned. This analysis can yield information on relative 5-methylcytosine for genomic regions, concurrent with standard genetic sequencing / variant detection.
[0342] Methods to integrating sets of molecular data
[0343] Multiple methods can be employed to integrate multiple layers of epigenome information such as methylation patterns, transcription factor binding sites (TFBS), and histone modifications described earlier in the disclosure to predict regional features like the presence of a TFBS or chromatin interaction sites (e.g., enhancer promoter interaction) at a specific genomic location. For example, a neural network approach can integrate multiple layers of epigenomic information to predict regional features, using deep learning to handle highdimensional data. Such model can undergo continuous iteration and validation against known biological insights can refine the model specificity and sensitivity. In some embodiments, this integrative modeling approach may be used to infer regulatory or structural features relevant to disease detection, immune response, or therapeutic resistance by learning patterns across multiple genomic, epigenomic, and / or transcriptomic layers, to more accurately classify molecular states in subjects undergoing cancer vaccine treatment.
[0344] Neural Network Architecture: Multi-Modal Deep Learning Framework
[0345] The following describes neural network architectures suitable for implementing the classification methods disclosed herein. The methods described herein may be applied toAttorney Docket No. : GH0224WO analyze data from subjects undergoing treatment with a PCV or an off-the-shelf therapeutic cancer vaccine.
[0346] Input Layer: the methods in the present disclosure provide datasets involving various types of epigenomic information, hence the input layer is designed to handle multiple data modalities. We achieve this by using separate input channels or sub-networks for each data type (methylation, TFBS, and histone modifications). Each channel can preprocess its respective data type, normalizing and encoding it in a form suitable for deep learning, that is including several steps to convert raw data into a format that neural networks can effectively process and learn from. These steps include data normalization / standardization, encoding, reshaping, handling missing values, feature engineering / selection, and data augmentation. Feature extraction layers: For each data modality, a convolutional neural network (CNNs) or recurrent neural network (RNN) can be used to capture spatial dependencies and patterns within the genomic sequences. CNNs are particularly useful for identifying patterns in histone modifications and TFBS, while RNNs or transformer-based models can effectively process sequential data, capturing long-range dependencies in methylation patterns for example. Integration Layer: After feature extraction, the outputs of the separate channels are integrated. We achieve this through concatenation, followed by dense layers, or by using more sophisticated integration techniques like attention mechanisms, which allow the model to weigh the importance of information from different epigenomic layers dynamically. Prediction Layer: We then feed the integrated features into one or more dense layers with nonlinear activation functions to enable the prediction of regional features, for example the presence of a TFBS at a specific genomic location( this holds true for other feature like histone marks and interaction sites). The output layer is designed according to the specific prediction task, for example binary classification for predicting the presence / absence of a TFBS. Training: The model should be trained on labeled datasets where the ground truth (e.g., presence or absence of TFBS in specific regions or other functional element or genomic feature) is known. We use a cross-entropy loss function for classification tasks, and optimize the model using gradient descent algorithms like Adam or SGD. Regularization and Dropout: To prevent overfitting, given the complexity of epigenomic data and the deep architecture, we incorporate regularization techniques (L1 / L2 regularization) and dropout layers, particularly after dense layers in the network. Evaluation and Fine-tuning: We evaluate the model’ s performance using standard metrics like accuracy, precision, recall, and Fl score. Depending on the results, finetune the model by adjusting the architecture, hyperparameters, or training procedure.Attorney Docket No. : GH0224WO
[0347] In some embodiments we use techniques like cross-validation for a more robust evaluation. Similarly, for enhanced per Performance methods such as data augmentation techniques specific to genomic data to increase the diversity of the training set. Other techniques include multi-task learning when for example predicting multiple regional features simultaneously, under this framework the network is designed to make several predictions at once, sharing representations between tasks to improve learning efficiency and prediction accuracy. Lastly techniques such as transfer learning can be employed to leverage pre-trained models on related tasks to improve performance, this method is especially useful when labeled data are limited.
[0348] In certain embodiments, the methods described herein may utilize a unified deep learning model to integrate multiple biomarker modalities including methylation profiles (Fi), neoantigen abundance (F2), and / or fragmentomic or transcriptomic features (F3) for therapeutic response prediction and MRD assessment in subjects undergoing cancer vaccine treatment. Each data modality may be processed through a modality-specific encoder, such as a selfattention layer, to produce a learned representation.
[0349] These representations may then be integrated using a modality-weighted fusion mechanism, where the final feature embedding is computed as:
[0350] Ffinai = WiFi + W2F2+ W3F3, with weights w , w2, w3learned during training to reflect the relative importance of each biomarker type. This fused embedding is used by a classification module to assign the sample to outcome categories such as “responder,” “nonresponder,” or “MRD-positive,” thereby providing data-driven, multi-modal assessment of vaccine efficacy and / or decease status in the subject.
[0351] Each individual biomarker feature such as a methylation site, transcript expression value, or neoantigen-linked variant may be mapped to its corresponding protein-coding gene, and these genes are used to define nodes within a protein-protein interaction (PPI) or gene regulatory network graph. Features mapped to the same gene may be aggregated to form a node-level feature vector. The resulting graph, with biomarker-derived node attributes, is processed using a graph attention network (GAT) to learn context-aware representations for downstream classification tasks. Node embeddings are updated through attention-weighted message passing, where the representation of a node / + 1 is given by:Attorney Docket No. : GH0224WOHere a® is the learned attention coefficient between nodes i and j, allowing the model to focus on the most informative relationships. These attention weights are computed using: a® = softmaXj LeakyReLU(where II denotes feature concatenation, and a is a trainable attention vector. The final graphlevel representation may then be passed through a classification layer to assign categories such as “responder,” “non-responder,” and / or “MRD-positive,” or “MRD-positive” for sample classification in cancer vaccine treatment monitoring.
[0352] Integrating informative features across multiple datasets offers several advantages by enabling the capture of subtle and / or non-linear patterns that may be missed by traditional statistical or rule-based analyses. The combination of diverse biomarker modalities including fragmentomic signatures, methylation profiles, neoantigen levels or signatures, and transcriptomic data allows models to learn higher-order relationships and integrate signals that span biological pathways, immune dynamics, and tumor evolution. This integration enhances assay sensitivity and specificity, particularly in samples obtained from tumors with low ctDNA shedding and / or those with biologically heterogeneous tumor contexts. Moreover, Al-driven methods can adaptively refine predictions based on prior training data, leading to more robust generalization across diverse patient populations, disease types, and clinical scenarios.
[0353] Integrating molecular data with patient outcomes data, electronic health records and / or health insurance claims data
[0354] In various examples, the implementations described herein can integrate molecular data comprising transcription factor binding sites (TFBS), fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, and / or H3S10ph with patient outcomes data, electronic health records and / or health insurance claims data to: (1) Identify common methylation changes associated with cancer types or subtypes and stages. The goal of this analysis is to reveal potential methylation biomarkers for cancer diagnosis (2) Correlate methylation changes with clinical outcomes. The goal of this analysis is to understand the impact of these methylation changes on cancer progression, resistance, and treatment response. (3) Integrate the identified methylation changes with existing biological pathways and networks. The goal of this analysis is to uncover how methylation changes disrupt cellular processes and contribute to cancer development, resistance, and response. (4) Compare methylation data across different cancer types or subtypes to identify unique andAttorney Docket No. : GH0224WO shared mechanisms. The goal of this analysis is to understand cancer heterogeneity and similarities across different cancers. (5) Develop a predictive model using methylation data to predict treatment outcomes, recurrence, or drug resistance. The goal of this analysis is to help guide personalized treatment strategies. Such analysis can be achieved using a variety of statistical and machine learning models including Linear Regression and Logistic Regression: For continuous and binary outcomes, respectively, to model the relationship between methylation levels at specific sites and the presence or severity of cancer. Cox Proportional Hazards Model, can be useful for survival analysis to correlate methylation levels with the time to event data, such as time to cancer recurrence or progression. Mixed Models are helpful when the data comprises multiple measurements or hierarchical structures, mixed models can account for the correlation within subjects or groups. Multivariate Additionally, analysis like principal component analysis (PCA) or partial least squares regression (PLSR) can reduce dimensionality and identify patterns in methylation data that correlate with cancer types.
[0355] Other methods employed to explore relationships or correlations between methylation data and cancer type / subtype include machine learning models. For example, to handle complex, non-linear relationships we employ Decision Trees and Random Forests these are particularly helpful in situations where the association between the methylation status of certain genes (or CpG sites) and cancer characteristics does not follow a straight-line pattern. For example, to model interaction effects i.e., methylation at one site might affect the impact of methylation at another site. Alone or in one or more combinations these models can identify specific methylation sites that are important for classifying cancer types. In some embodiments, Support Vector Machines (SVM) can be employed for classification tasks, including distinguishing between different types of cancer based on methylation patterns. In yet other embodiments, for example when the data is sufficiently large, deep learning approaches (e.g., convolutional neural networks for structured data like methylation arrays) can capture complex patterns including interactions in the data. Methods such as Gradient Boosting Machines (GBM) including models like XGBoost, LightGBM, and CatBoost can provide robust predictive models for cancer classification based on methylation data. These models include sequential addition of weak learners (e.g., decision trees) in such a way that each new tree corrects the errors made by the previous ones. GBMs can handle various types of data, including categorical and continuous variables. In additional embodiments, Cluster Analysis are employed to identify subgroups within cancer types that share similar methylation patterns. These include unsupervised learning models, including K-means clustering or hierarchical clustering.Attorney Docket No. : GH0224WO
[0356] In some embodiments, the cell-free DNA is from a subject having or suspected of having cancer and / or the cell-free DNA includes DNA from cancer cells. In some embodiments, the DNA is partitioned into a first subsample and a second subsample, wherein the first subsample comprises DNA with a nucleotide modification (e.g., a cytosine modification) in a greater proportion than the second subsample, and the first subsample is subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample, and the DNA is sequenced in a manner that distinguishes the first nucleobase from the second nucleobase in the DNA of the first sub sample.
[0357] Computer Systems
[0358] All methods of the present disclosure can be implemented using, or with the aid of, computer systems. For example, such methods may comprise: partitioning the sample into a plurality of subsamples, including a first subsample and a second subsample, wherein the first subsample comprises DNA with a cytosine modification in a greater proportion than the second subsample; subjecting the first subsample to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first sub sample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity; and sequencing DNA in the first subsample and DNA in the second subsample in a manner that distinguishes the first nucleobase from the second nucleobase in the DNA of the first subsample. The computer system can also include Hardware, for example the hardware components of the system include a central processing unit (CPU), random access memory (RAM), storage devices (such as hard disk drives or solid-state drives), and input / output devices (such as keyboards, mice, and display monitors). The specific configuration of the hardware components may vary based on the System model and user requirements. The System can also include software including operating system software that manages the hardware resources and provides a platform for running application software. In addition to the operating system, the System may come with pre-installed application software designed to meet the needs of specific tasks or industries. Users may also install additional applications as required. Additional features in the computer system can include security systems. For example the system can incorporate multiple layers of security measures, including firewalls, antivirus software, and encryption protocols, to protect against unauthorized access and ensure the confidentiality, integrity, and availability of data. The system can also include connectivity systems comprising various connectivityAttorney Docket No. : GH0224WO options, including wired and wireless network connections, to enable communication and data exchange with other systems and devices. Compatibility with standard networking protocols ensures the System can integrate seamlessly into existing network environments. The computer system can also include support and Maintenance systems. For example, a comprehensive support and maintenance services are provided to ensure the system operates efficiently and effectively. This includes technical support, software updates, and hardware repair or replacement services. Other features of the computer system can include compliance protocols. For example, the system can be designed to comply with relevant industry standards and regulatory requirements, ensuring reliability and safety in its operation.
[0359] Additional details relating to computer systems and networks, databases, and computer program products are also provided in, for example, Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Ed. (2011), Kurose, Computer Networking: A Top-Down Approach, Pearson, 7thEd. (2016), Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Ed. (2010), Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11thEd. (2014), Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd Ed. (2006), and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011), each of which is hereby incorporated by reference in its entirety.
[0360] Practical applications of the methods
[0361] The present methods can be used to diagnose presence of conditions, particularly cancer, in a subject, to characterize conditions (e.g., staging cancer or determining heterogeneity of a cancer), monitor response to treatment of a condition, effect prognosis risk of developing a condition or subsequent course of a condition. The present disclosure can also be useful in determining the efficacy of a particular treatment option. Successful treatment options may increase the amount of copy number variation or rare mutations detected in subject's blood if the treatment is successful as more cancers may die and shed DNA. In other examples, this may not occur. In another example, perhaps certain treatment options may be correlated with genetic profiles of cancers over time. This correlation may be useful in selecting a therapy.
[0362] Additionally, if a cancer is observed to be in remission after treatment, the present methods can be used to monitor residual disease or recurrence of disease.
[0363] In some embodiments, the methods and systems disclosed herein may be used to identify customized or targeted therapies to treat a given disease or condition in patients based on the classification of a nucleic acid variant as being of somatic or germline origin. Typically,Attorney Docket No. : GH0224WO the disease under consideration is a type of cancer. Non-limiting examples of such cancers include biliary tract cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, gliomas, astrocytomas, breast carcinoma, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal carcinoma, colon cancer, hereditary nonpolyposis colorectal cancer, colorectal adenocarcinomas, gastrointestinal stromal tumors (GISTs), endometrial carcinoma, endometrial stromal sarcomas, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder carcinomas, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinomas, Wilms tumor, leukemia, acute lymphocytic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), chronic myelomonocytic leukemia (CMML), liver cancer, liver carcinoma, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, Lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphomas, non-Hodgkin lymphoma, diffuse large B-cell lymphoma, Mantle cell lymphoma, T cell lymphomas, non-Hodgkin lymphoma, precursor T- lymphoblastic lymphoma / leukemia, peripheral T cell lymphomas, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal cancer, oral cavity squamous cell carcinomas, osteosarcoma, ovarian carcinoma, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasms, acinar cell carcinomas. Prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine carcinomas, stomach cancer, gastric carcinoma, gastrointestinal stromal tumor (GIST), uterine cancer, or uterine sarcoma. Type and / or stage of cancer can be detected from genetic variations including mutations, rare mutations, indels, copy number variations, transversions, translocations, inversion, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structure alterations, gene fusions, chromosome fusions, gene truncations, gene amplification, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5-methylcytosine.
[0364] Genetic data can also be used for characterizing a specific form of cancer. Cancers are often heterogeneous in both composition and staging. Genetic profile data may allow characterization of specific sub-types of cancer that may be important in the diagnosis or treatment of that specific sub-type. This information may also provide a subject or practitioner clues regarding the prognosis of a specific type of cancer and allow either a subject or practitioner to adapt treatment options in accord with the progress of the disease. Some cancersAttorney Docket No. : GH0224WO can progress to become more aggressive and genetically unstable. Other cancers may remain benign, inactive, or dormant. The system and methods of this disclosure may be useful in determining disease progression.
[0365] Further, the methods of the disclosure may be used to characterize the heterogeneity of an abnormal condition in a subject. Such methods can include, e.g., generating a genetic profile of extracellular polynucleotides derived from the subject, wherein the genetic profile comprises a plurality of data resulting from copy number variation and rare mutation analyses. In some embodiments, an abnormal condition is cancer. In some embodiments, the abnormal condition may be one resulting in a heterogeneous genomic population. In the example of cancer, some tumors are known to comprise tumor cells in different stages of the cancer. In other examples, heterogeneity may comprise multiple foci of disease. Again, in the example of cancer, there may be multiple tumor foci, perhaps where one or more foci are the result of metastases that have spread from a primary site.
[0366] The present methods can be used to generate or profile, fingerprint, or set of data that is a summation of genetic information derived from different cells in a heterogeneous disease. This set of data may comprise copy number variation, epigenetic variation, and mutation analyses alone or in combination.
[0367] The present methods can be used to diagnose, prognose, monitor or observe cancers, or other diseases. In some embodiments, the methods herein do not involve the diagnosing, prognosing or monitoring a fetus and as such are not directed to non-invasive prenatal testing. In other embodiments, these methodologies may be employed in a pregnant subject to diagnose, prognose, monitor or observe cancers or other diseases in an unborn subject whose DNA and other polynucleotides may co-circulate with maternal molecules.
[0368] Non-limiting examples of other genetic-based diseases, disorders, or conditions that are optionally evaluated using the methods and systems disclosed herein include achondroplasia, alpha- 1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-Tooth (CMT), cri du chat, Crohn's disease, cystic fibrosis, Dercum disease, down syndrome, Duane syndrome, Duchenne muscular dystrophy, Factor V Leiden thrombophilia, familial hypercholesterolemia, familial Mediterranean fever, fragile X syndrome, Gaucher disease, hemochromatosis, hemophilia, holoprosencephaly, Huntington's disease, Klinefelter syndrome, Marfan syndrome, myotonic dystrophy, neurofibromatosis, Noonan syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland anomaly, porphyria, progeria, retinitis pigmentosa, severe combined immunodeficiency (SCID), sickle cell disease, spinal muscular atrophy, Tay-Sachs,Attorney Docket No. : GH0224WO thalassemia, trimethylaminuria, Turner syndrome, velocardiofacial syndrome, WAGR syndrome, Wilson disease, or the like.
[0369] In some embodiments, a method described herein comprises detecting a presence or absence of DNA originating or derived from a tumor cell at a preselected timepoint following a previous cancer treatment of a subject previously diagnosed with cancer using a set of sequence information obtained as described herein. The method may further comprise determining a cancer recurrence score that is indicative of the presence or absence of the DNA originating or derived from the tumor cell for the test subject.
[0370] Where a cancer recurrence score is determined, it may further be used to determine a cancer recurrence status. The cancer recurrence status may be at risk for cancer recurrence, e.g., when the cancer recurrence score is above a predetermined threshold. The cancer recurrence status may be at low or lower risk for cancer recurrence, e.g., when the cancer recurrence score is above a predetermined threshold. In particular embodiments, a cancer recurrence score equal to the predetermined threshold may result in a cancer recurrence status of either at risk for cancer recurrence or at low or lower risk for cancer recurrence.
[0371] In some embodiments, a cancer recurrence score is compared with a predetermined cancer recurrence threshold, and the test subject is classified as a candidate for a subsequent cancer treatment when the cancer recurrence score is above the cancer recurrence threshold or not a candidate for therapy when the cancer recurrence score is below the cancer recurrence threshold. In particular embodiments, a cancer recurrence score equal to the cancer recurrence threshold may result in classification as either a candidate for a subsequent cancer treatment or not a candidate for therapy.
[0372] The methods discussed above may further comprise any compatible feature or features set forth elsewhere herein, including in the section regarding methods of determining a risk of cancer recurrence in a test subject and / or classifying a test subject as being a candidate for a subsequent cancer treatment.
[0373] Methods of determining a risk of cancer recurrence in a test subject and / or classifying a test subject as being a candidate for a subsequent cancer treatment
[0374] In some embodiments, the methods disclosed herein are used to determine a subject’s risk of cancer recurrence. . In other embodiments, the methods disclosed herein are used to determine subject’s risk of cancer recurrence. In certain embodiments, the disclosed methods are applied to assess MRD and predict risk of cancer recurrence in subjects treated with a cancer vaccine. MRD assessment may be based on the detection and interpretation of molecular biomarkers such as tumor-specific mutations, regions exhibiting treatment-induced epigeneticAttorney Docket No. : GH0224WO changes including methylation changes, neoantigen abundance, and / or fragmentomic alterations identified in biological samples collected from subjects undergoing treatment with a cancer vaccine.
[0375] The molecular data including tumor-specific mutations, regions exhibiting treatment- induced epigenetic changes including methylation changes, neoantigen abundance, and / or fragmentomic alterations identified in biological samples collected from subjects undergoing treatment with a cancer vaccine, may be further analyzed using trained classifiers to determine whether a subject remains at elevated risk for recurrence and / or may benefit from a subsequent or modified treatment regimen, such as a vaccine boost, immune-based therapy, and / or targeted intervention.
[0376] The methods thus support clinical decision-making post-vaccination as they may be used to stratify subjects for continued surveillance and / or additional therapeutic intervention.
[0377] Samples may be collected longitudinally before, during, and / or after treatment for temporal comparison. In some embodiments, classifiers trained on molecular signatures of MRD are applied to interpret the presence or absence of disease-related signals. Hence, the methods provide a non-invasive approach to detect subclinical relapse, evaluate treatment efficacy, and inform follow-up interventions such as vaccine boosting and / or adjunctive therapy.
[0378] Any of such methods may comprise collecting DNA (e.g., originating or derived from a tumor cell) from the test subject diagnosed with the cancer at one or more preselected timepoints following one or more previous cancer treatments to the test subject. The subject may be any of the subjects described herein. The DNA may be cfDNA. The DNA may be obtained from a tissue sample.
[0379] Any of such methods may comprise capturing a plurality of sets of target regions from DNA from the subject, wherein the plurality of target region sets comprises a sequence-variable target region set and an epigenetic target region set, whereby a captured set of DNA molecules is produced. The capturing step may be performed according to any of the embodiments described elsewhere herein.
[0380] In any of such methods, the previous cancer treatment may comprise surgery, administration of a therapeutic composition, and / or chemotherapy.
[0381] Any of such methods may comprise sequencing the captured DNA molecules, whereby a set of sequence information is produced. The captured DNA molecules of the sequencevariable target region set may be sequenced to a greater depth of sequencing than the captured DNA molecules of the epigenetic target region set.Attorney Docket No. : GH0224WO
[0382] Any of such methods may comprise detecting a presence or absence of DNA originating or derived from a tumor cell at a preselected timepoint using the set of sequence information. The detection of the presence or absence of DNA originating or derived from a tumor cell may be performed according to any of the embodiments thereof described elsewhere herein.
[0383] Methods of determining a risk of cancer recurrence in a test subject may comprise determining a cancer recurrence score that is indicative of the presence or absence, or amount, of the DNA originating or derived from the tumor cell for the test subject. The cancer recurrence score may further be used to determine a cancer recurrence status. The cancer recurrence status may be at risk for cancer recurrence, e.g., when the cancer recurrence score is above a predetermined threshold. The cancer recurrence status may be at low or lower risk for cancer recurrence, e.g., when the cancer recurrence score is above a predetermined threshold. In particular embodiments, a cancer recurrence score equal to the predetermined threshold may result in a cancer recurrence status of either at risk for cancer recurrence or at low or lower risk for cancer recurrence.
[0384] Methods of classifying a test subject as being a candidate for a subsequent cancer treatment may comprise comparing the cancer recurrence score of the test subject with a predetermined cancer recurrence threshold, thereby classifying the test subject as a candidate for the subsequent cancer treatment when the cancer recurrence score is above the cancer recurrence threshold or not a candidate for therapy when the cancer recurrence score is below the cancer recurrence threshold. In particular embodiments, a cancer recurrence score equal to the cancer recurrence threshold may result in classification as either a candidate for a subsequent cancer treatment or not a candidate for therapy. In some embodiments, the subsequent cancer treatment comprises chemotherapy or administration of a therapeutic composition.
[0385] Any of such methods may comprise determining a disease-free survival (DFS) period for the test subject based on the cancer recurrence score; for example, the DFS period may be 1 year, 2 years, 3, years, 4 years, 5 years, or 10 years.
[0386] In some embodiments, the set of sequence information comprises sequence-variable target region sequences, and determining the cancer recurrence score may comprise determining at least a first subscore indicative of the amount of SNVs, insertions / deletions, CNVs and / or fusions present in sequence-variable target region sequences.
[0387] In some embodiments, a number of mutations in the sequence-variable target regions chosen from 1, 2, 3, 4, or 5 is sufficient for the first subscore to result in a cancer recurrenceAttorney Docket No. : GH0224WO score classified as positive for cancer recurrence. In some embodiments, the number of mutations is chosen from 1, 2, or 3.
[0388] In some embodiments, the set of sequence information comprises epigenetic target region sequences, and determining the cancer recurrence score comprises determining a second subscore indicative of the amount of molecules (obtained from the epigenetic target region sequences) that represent an epigenetic state different from DNA found in a corresponding sample from a healthy subject (e.g., cfDNA found in a blood sample from a healthy subject, or DNA found in a tissue sample from a healthy subject where the tissue sample is of the same type of tissue as was obtained from the test subject). These abnormal molecules (i.e., molecules with an epigenetic state different from DNA found in a corresponding sample from a healthy subject) may be consistent with epigenetic changes associated with cancer, e.g., methylation of hypermethylation variable target regions and / or perturbed fragmentation of fragmentation variable target regions, where “perturbed” means different from DNA found in a corresponding sample from a healthy subject.
[0389] In some embodiments, a proportion of molecules corresponding to the hypermethylation variable target region set and / or fragmentation variable target region set that indicate hypermethylation in the hypermethylation variable target region set and / or abnormal fragmentation in the fragmentation variable target region set greater than or equal to a value in the range of 0.001%-10% is sufficient for the second subscore to be classified as positive for cancer recurrence. The range may be 0.001%-l%, 0.005%-l%, 0.01%-5%, 0.01%-2%, or O.OI%-1%.
[0390] In some embodiments, any of such methods may comprise determining a fraction of tumor DNA from the fraction of molecules in the set of sequence information that indicate one or more features indicative of origination from a tumor cell. This may be done for molecules corresponding to some or all of the epigenetic target regions, e.g., including one or both of hypermethylation variable target regions and fragmentation variable target regions (hypermethylation of a hypermethylation variable target region and / or abnormal fragmentation of a fragmentation variable target region may be considered indicative of origination from a tumor cell). This may be done for molecules corresponding to sequence variable target regions, e.g., molecules comprising alterations consistent with cancer, such as SNVs, indels, CNVs, and / or fusions. The fraction of tumor DNA may be determined based on a combination of molecules corresponding to epigenetic target regions and molecules corresponding to sequence variable target regions.Attorney Docket No. : GH0224WO
[0391] Determination of a cancer recurrence score may be based at least in part on the fraction of tumor DNA, wherein a fraction of tumor DNA greater than a threshold in the range of 10'11to 1 or IO'10to 1 is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence. In some embodiments, a fraction of tumor DNA greater than or equal to a threshold in the range of 1010to 109, 109to 108, 108to 107, 107to 106, 106to 105, 105to 10^, 10^ to I O3, I O3to I O2, or 102to 101is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence. In some embodiments, the fraction of tumor DNA greater than a threshold of at least 10'7is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence. A determination that a fraction of tumor DNA is greater than a threshold, such as a threshold corresponding to any of the foregoing embodiments, may be made based on a cumulative probability. For example, the sample was considered positive if the cumulative probability that the tumor fraction was greater than a threshold in any of the foregoing ranges exceeds a probability threshold of at least 0.5, 0.75, 0.9, 0.95, 0.98, 0.99, 0.995, or 0.999. In some embodiments, the probability threshold is at least 0.95, such as 0.99.
[0392] In some embodiments, the set of sequence information comprises sequence-variable target region sequences and epigenetic target region sequences, and determining the cancer recurrence score comprises determining a first subscore indicative of the amount of SNVs, insertions / deletions, CNVs and / or fusions present in sequence-variable target region sequences and a second subscore indicative of the amount of abnormal molecules in epigenetic target region sequences, and combining the first and second subscores to provide the cancer recurrence score. Where the first and second subscores are combined, they may be combined by applying a threshold to each subscore independently (e.g., greater than a predetermined number of mutations (e.g., > 1) in sequence-variable target regions, and greater than a predetermined fraction of abnormal molecules (i.e., molecules with an epigenetic state different from the DNA found in a corresponding sample from a healthy subject; e.g., tumor) in epigenetic target regions), or training a machine learning classifier to determine status based on a plurality of positive and negative training samples.
[0393] In some embodiments, a value for the combined score in the range of -4 to 2 or -3 to 1 is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence.
[0394] In any embodiment where a cancer recurrence score is classified as positive for cancer recurrence, the cancer recurrence status of the subject may be at risk for cancer recurrence and / or the subject may be classified as a candidate for a subsequent cancer treatment.Attorney Docket No. : GH0224WO
[0395] In some embodiments, the cancer is any one of the types of cancer described elsewhere herein, e.g., colorectal cancer.
[0396] Therapies and Related Administration
[0397] In certain embodiments, the methods disclosed herein relate to identifying and administering customized therapies to patients given the status of a nucleic acid variant as being of somatic or germline origin. In some embodiments, essentially any cancer therapy (e.g., surgical therapy, radiation therapy, chemotherapy, and / or the like) may be included as part of these methods. Typically, customized therapies include at least one immunotherapy (or an immunotherapeutic agent). Immunotherapy refers generally to methods of enhancing an immune response against a given cancer type. In certain embodiments, immunotherapy refers to methods of enhancing a T cell response against a tumor or cancer.
[0398] In certain embodiments, the status of a nucleic acid variant from a sample from a subject as being of somatic or germline origin may be compared with a database of comparator results from a reference population to identify customized or targeted therapies for that subject. Typically, the reference population includes patients with the same cancer or disease type as the test subject and / or patients who are receiving, or who have received, the same therapy as the test subject. A customized or targeted therapy (or therapies) may be identified when the nucleic variant and the comparator results satisfy certain classification criteria (e.g., are a substantial or an approximate match).
[0399] In certain embodiments, the customized therapies described herein are typically administered parenterally (e.g., intravenously, or subcutaneously). Pharmaceutical compositions containing an immunotherapeutic agent are typically administered intravenously. Certain therapeutic agents are administered orally. However, customized therapies (e.g., immunotherapeutic agents, etc.) may also be administered by methods such as, for example, buccal, sublingual, rectal, vaginal, intraurethral, topical, intraocular, intranasal, and / or intraauricular, which administration may include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, salves, ointments, or the like.
[0400] In addition to monitoring patients undergoing treatment with a PCV the methods can also be applied to monitor patients treated with off-the-shelf vaccines alone or in combination with other therapies. For example, to monitor patients with early-stage bladder cancer treated with Bacillus Calmette-Guerin (BCG) or prostate cancer patients treated with Sipuleucel-T (Provenge®). Similarly, the methods can be applied to patients treated with off-the-shelf vaccines having personalized elements tailored to the patient’s immune cells.Attorney Docket No. : GH0224WO
[0401] In some embodiments, the cancer vaccines include: Sipuleucel-T (Provenge) for prostate cancer, Talimogene laherparepvec (T-VEC, Imlygic) for melanoma, BCG (Bacillus Calmette-Guerin) for bladder cancer, MAGE-A3 antigen-specific cancer immunotherapeutic (ASCI) for various cancers (investigational), CIMAvax-EGF for non-small cell lung cancer (NSCLC), GV AX pancreas for pancreatic cancer (investigational), Rindopepimut (CDX-110) for glioblastoma (investigational), ROVA-T for small cell lung cancer (investigational), OncoVAX for colon cancer (investigational), and New York-ESO-1 Vaccine for various cancers (investigational).
[0402] In some embodiments, the methods use enveloped virus-like particles (eVLPs). eVLPs are protein-based vaccines engineered to closely resemble the natural structure of viruses as recognized by our immune system. For instance, VBI-1901, developed to treat aggressive brain tumors, incorporates two highly immunogenic antigens from cytomegalovirus (CMV). These antigens activate CD4+ and CD8+ T cells, key components of the immune response. CMV is commonly expressed in various solid tumors, including glioblastoma multiforme. Patients can receive VBI-1901 through two intradermal injections administered monthly.
[0403] In some embodiments, antibodies for boosting vaccine effectiveness are used, for example, NEO-201. NEO-201 can be used for treating advanced lung, head and neck, endometrial, and cervical cancers. This antibody can counteract immune suppression linked to vaccines and other immunotherapies. NEO-201 specifically targets and eliminates cells that express truncated core 1 O-glycan. These glycans are associated with immunosuppressive cells, which can weaken immune responses to immunotherapy and vaccines. By binding to and destroying these cells, NEO-201 can enhance the effectiveness of treatments.
[0404] In some embodiments, the cancer vaccine is SCIB1 (SCIB 1-001), which is used for the treatment of patients with metastatic melanoma. SCIB1 incorporates specific epitopes from the proteins gplOO and TRP-2, which play key roles in melanin production in the skin. These epitopes were identified from the cloning of T cells from patients who experienced spontaneous recovery from melanoma. In some embodiments, SCIB1 is used in combination with immune checkpoint inhibitors (CPIs), such as Keytruda® (pembrolizumab) and Opdivo® (nivolumab) for the treatment of late-stage melanoma. The rationale is that while CPIs relieve the immunosuppression in the tumor microenvironment, SCIB1 primes the immune response against the tumor, potentially enhancing the efficacy of CPIs. Preclinical studies have shown that SCIB1 and anti-PDl antibody CPIs have a strong synergistic effect when administeredAttorney Docket No. : GH0224WO together, and this combination could increase the proportion of patients who respond to the treatment. In some embodiments, iSCIBl+, a modified version of SCIB1, is used. SCIB1 expresses a select number of melanoma-specific T cell epitopes, restricting its efficacy to about 40% of patients. iSCIBl+ expands the number of epitopes and includes an improved Fc region, enhanced with AvidiMab® technology, to better target dendritic cells and induce higher frequency T cell responses.
[0405] Off- The-Shelf Therapeutic Cancer Vaccines
[0406] The methods described throughout the disclosure may also be applied to process and analyze samples from subjects undergoing treatment with an off-the-shelf cancer vaccine. Analysis of such samples will yield information related to treatment-induced immune activation, epigenetic remodeling, minimal residual disease (MRD), tumor evolution, and therapeutic resistance, analogous to the applications described for personalized cancer vaccines (PCVs). These methods enable detection of molecular signatures such as changes in DNA methylation, cfDNA fragmentation patterns, or transcriptomic alterations that reflect biological responses to off-the-shelf vaccination, either alone or in combination with other therapies. In additional embodiments the levels of neoantigens may also be tracked as indicators of response to treatment, and may be used alone or in combination with other biomarkers, such as DNA methylation or fragmentomic patterns, to assess therapeutic efficacy and disease progression. In addition to monitoring patients undergoing treatment with a PCV the methods can also be applied to monitor patients treated with off-the-shelf vaccines alone or in combination with other therapies. For example, to monitor patients with early-stage bladder cancer treated with Bacillus Calmette-Guerin (BCG) or prostate cancer patients treated with Sipuleucel-T (Provenge®). Similarly, the methods can be applied to patients treated with off-the-shelf vaccines having personalized elements tailored to the patient’s immune cells.
[0407] Additional off-the-self cancer vaccines include: Sipuleucel-T (Provenge) for prostate cancer, Talimogene laherparepvec (T-VEC, Imlygic) for melanoma, BCG (Bacillus Calmette- Guerin) for bladder cancer, MAGE -A3 antigen-specific cancer immunotherapeutic (ASCI) for various cancers (investigational), CIMAvax-EGF for non-small cell lung cancer (NSCLC), GV AX pancreas for pancreatic cancer (investigational), Rindopepimut (CDX-110) for glioblastoma (investigational), ROVA-T for small cell lung cancer (investigational), OncoVAX for colon cancer (investigational), and New York-ESO-1 Vaccine for various cancers (investigational).Attorney Docket No. : GH0224WO
[0408] Unlike personalized cancer vaccines, these off-the-shelf products are not designed based on individual tumor sequencing, but may nonetheless induce systemic or local immune activation in subjects.
[0409] In some embodiments, samples are collected from subjects undergoing treatment with an off-the-self vaccine to treat a cancer. Samples including pre and / or post-treatment biological samples including blood, plasma, serum, urine, saliva, cerebrospinal fluid, pleural fluid, ascitic fluid, synovial fluid, breast milk, semen, vaginal fluid, amniotic fluid, bile, sweat, tears, stool, sputum, nasal fluid, and / or bronchoalveolar lavage fluid may be collected, sequenced to generate sequencing data. Raw and / or pre processed data may then be used to extract signals of disease and / or response to off-the-self cancer vaccine. Signals of disease including MRD and / or response to off-the-shelf vaccines include methylation changes. In some embodiments, the methylation changes occur within genes involved in immune activation, antigen presentation, cytokine signaling, T cell receptor signaling, and immune checkpoint regulation, including but not limited to IFNG, IL2, IL6, TNF, CXCL9, CXCL10, CD3E, CD8A, PDCD1, CTLA4, LAG3, and HL A class I and II genes, as previously described.
[0410] Signals detectable in subjects undergoing treatment with an off-the-shelf vaccine may reflect immune activation, tumor cell death and / or immune targeting, or therapeutic resistance. The methods enable real-time monitoring of treatment efficacy even when the cancer vaccine is not personalized to the subject.
[0411] Epigenetic therapeutic drugs
[0412] The intricate interplay of chemical modifications to DNA and histone proteins, known as the epigenome, plays a significant role in modulating gene expression in health and disease including cancer. Epigenetic changes in conjunction with genetic alterations contribute to the acquisition of cancer hallmarks such as sustaining proliferative signaling, evading growth suppressors, resisting cell death, enabling replicative immortality, inducing angiogenesis, and activating invasion and metastasis (Hanahan, D. 2022). Given the reversible nature of epigenetic modifications, understanding these mechanisms offers promising avenues for therapeutic intervention, with several epigenetic drugs already approved or in clinical trials for the treatment of cancer.
[0413] FDA approved epigenetic drugs
[0414] Methods of the disclosure can be applied to monitor patients treated with a PCV and / or an off-the-self vaccine, alone or in combination with other therapies including epigenetic drugs such as DNA methyltransferase (DNMT) inhibitors, Histone deacetylase (HDAC) inhibitors,Attorney Docket No. : GH0224WOLysine methyltransferase inhibitors, Lysine demethylase inhibitors, and Bromodomain inhibitors.
[0415] DNA Methyltransferase Inhibitors (DNMTi)- inhibit DNA methyltransferases, enzymes that add methyl groups to DNA, typically silencing gene expression. DNMTi drugs include Azacitidine (Vidaza) which was approved for the treatment of myelodysplastic syndromes (MDS) and Decitabine (Dacogen) also approved for MDS. Histone Deacetylase Inhibitors (HDACi) inhibit histone deacetylases, enzymes that remove acetyl groups from histone proteins, typically leading to a closed chromatin structure and gene silencing. Inhibiting these enzymes can reactivate silenced genes beneficial in cancer treatment. HDACi drugs include Vorinostat (Zolinza) approved for the treatment of cutaneous T cell lymphoma (CTCL), Romidepsin (Istodax) approved for CTCL and peripheral T-cell lymphoma (PTCL), Belino...
Claims
1. Attorney Docket No. : GH0224WOCLAIMSWhat is claimed is:
1. A method for monitoring therapeutic response in a subject treated with a personalized cancer vaccine (PCV), comprising:(a) obtaining a tumor and match normal tissue sample from the subject prior to initiation of treatment with the PCV, extracting a plurality of molecules from the tumor tissue sample, and sequencing the plurality of molecules to identify one or more biomarkers comprising tumor-specific neoantigens and differentially methylated regions (DMRs); and(b) obtaining a plasma sample from the subject after the initiation of treatment with the PCV and selecting the biomarkers based on the sequencing data from the tumor tissue sample in (a) and detecting the biomarkers in cell-free DNA from the plasma sample to determine the therapeutic response by assessing changes in neoantigen levels and DMRs, wherein a decrease in neoantigen levels indicates a response to treatment, and an increase or unchanged level indicates potential resistance or disease progression, and wherein the presence of at least one of the DMRs correlate with MRD in the subject, thereby monitoring therapeutic response in the subject treated with the PCV.
2. A method for monitoring therapeutic response in a subject treated with a personalized cancer vaccine (PCV), comprising:(a) obtaining a tumor and match normal tissue sample from the subject prior to initiation of treatment with the PCV, extracting a plurality of molecules from the tumor tissue sample, and sequencing the plurality of molecules to identify one or more tumor-specific neoantigens and differentially methylated regions (DMRs);(b) selecting a personalized panel incorporating one or more of the neoantigens and the DMRs identified (a);(c) obtaining a plasma sample from the subject after the initiation of treatment with the PCV; and(d) determining the therapeutic response by assessing changes in neoantigen levels, wherein a decrease in neoantigen levels indicates a response to treatment, and an increase or unchanged level indicates potential resistance or disease progression, and wherein the presence of at least one of the DMRs in the personalized panelAttorney Docket No. : GH0224WO correlate with MRD in the subject, thereby monitoring therapeutic response in the subject treated with the PCV.
3. A method for monitoring therapeutic response in a subject treated with a personalized cancer vaccine (PCV), comprising:(a) obtaining a tumor and match normal tissue sample from the subject prior to initiation of treatment with the PCV, extracting a plurality of molecules from the tumor tissue sample, and sequencing the plurality of molecules to identify one or more tumor-specific neoantigens and differentially methylated regions (DMRs);(b) selecting a personalized panel incorporating one or more of the neoantigens and the DMRs identified (a); and(c) obtaining a plasma sample from the subject after the initiation of treatment with the PCV and determining the therapeutic response by assessing changes in neoantigen levels and DMRs, wherein a decrease in neoantigen levels indicates a response to treatment, and an increase or unchanged level indicates potential resistance or disease progression, and wherein the presence of at least one of the DMRs in the personalized panel correlate with MRD in the subject, thereby monitoring therapeutic response in the subject treated with the PCV.
4. The method of claim 1, further comprising:(a) obtaining additional plasma samples from the subject at a plurality of subsequent time points following the initiation of treatment;(b) analyzing the additional plasma samples to monitor therapeutic response using the method described in claim 1(b); and(c) monitoring temporal changes in the levels of the detected neoantigens and / or DMRs across the plurality of time points to assess the therapeutic response over time.
5. The method of claim 1 or 4, wherein the treatment further comprises a combination of the PCV and one or more therapies.
6. The method of claim 5, wherein the one or more therapies comprise molecules that enhance the immune system’s ability to recognize and attack cancer cells.
7. The method of claims 5 to 6, wherein the one or more therapies are selected from the group consisting of: immune checkpoint inhibitors (ICIs), monoclonal antibodies, and immune modulators.
8. The method of claim 7, wherein the one or more ICIs comprise an anti-programmed cell death protein 1 (Anti-PD-1) molecule, monoclonal antibody, bispecific antibody (BsAbs) or a multispecific antibody.Attorney Docket No. : GH0224WO9. The method of claim 7, wherein the one or more ICIs comprise an anti-Cytotoxic T lymphocyte-associated protein-4 (Anti-CTLA-4) molecule, monoclonal antibody, bispecific antibody (BsAbs) or a multispecific antibody.
10. The method of claim 7, wherein the one or more ICIs are selected from the group consisting of: Pembrolizumab (Keytruda), Nivolumab (Opdivo), Atezolizumab (Tecentriq), Durvalumab (Imfinzi), Ipilimumab (Yervoy), CAR T-Cell Therapies, Tisagenlecleucel (Kymriah), Axicabtagene ciloleucel (Yescarta), and Brexucabtagene autoleucel (Tecartus).
11. The method of claim 7, wherein the one or more cytokines are selected from the group consisting of: Interleukin-2 (IL-2), Interferon-alpha (IFN-a), Oncolytic Virus Therapy, and Talimogene laherparepvec (T-VEC, Imlygic).
12. The method of claim 7, wherein the one or more monoclonal antibodies are selected from the group consisting of: Rituximab (Rituxan), Trastuzumab (Herceptin), and Bevacizumab (Avastin).
13. The method of claim 7, wherein the one or more immune modulators are selected from the group consisting of: Thalidomide (Thalomid), Lenalidomide (Revlimid), and Pomalidomide (Pomalyst).
14. The method of claim 1 or 4, wherein the tumor-specific neoantigens corresponds to tumor- associated mutations in the subject.
15. The method of claim 1 or 4, wherein the mutations are selected from the group consisting of: point mutation, insertions, deletions, frameshift mutation, duplication, inversion, translocation, copy number variation (CNV), repeat expansion, substitution, gene fusion, chromosome fusion, amplification, loss of heterozygosity, structural variations, large-scale deletions, large-scale duplications, and complex rearrangements.
16. The method of claims 1 or 4, wherein the differentially methylated regions(DMRs) comprise promoter regions.
17. The method of claims 1 or 4, wherein the DMRs comprise promoter regions, intronic regions, coding regions.
18. The method of claims 1 or 4, wherein the DMRs are selected based at least on one functional element.
19. The method of claim 18, wherein the functional element comprises transcription factor binding sites (TFBS), on or more overlapping variants, one or more histone modifications, and or one or more chromatin interaction sites.Attorney Docket No. : GH0224WO20. The method of claims 1 or 4, wherein the biomarkers further comprise clinically actionable genomic loci.
21. The method of claim 1, wherein the biomarkers further comprise resistance sequence variants and / or epigenetic variants emerging under therapeutic pressure.
22. The method of claims 1 or 4, wherein the method further comprises (a) generating a realtime dynamic monitoring report of the patient’s immune response to the PCV based on the therapeutic response assessment, and (b) adjusting the treatment regimen for the patient based on the information provided in the report.
23. The method of claim 1, wherein the PCV is designed based on whole exome / transcriptome tumor sequencing, and is administered in combination with one or more plasmid-encoded IL- 12 and / or pembrolizumab.
24. The method of claim 1, wherein the neoantigens comprise tumor-specific genetic mutations that result in the production of altered proteins.
25. The method of claim 1, further comprising determining neoantigen levels and / or a plurality of DMRs to predict clinical outcomes in the subject undergoing treatment with the PCV.
26. The method of claim 1, wherein the DMRs are selected based on whole epigenome, whole genome, and / or whole transcriptome sequencing of the tumor and the normal tissue samples from the patient, the DMRs overlapping at least one functional element in the whole genome and / or whole transcriptome data.
27. The method of claim 1, wherein the DMRs are selected based on targeted epigenome, targeted genome, and / or targeted transcriptome sequencing of the tumor and the normal tissue samples from the patient, the DMRs overlapping at least one functional element in the targeted genome and / or targeted transcriptome data.
28. The method of claim 1, wherein the neoantigens are selected based on whole genome and / or whole transcriptome sequencing of the tumor and the normal tissue samples from the patient, the neoantigens overlapping at least one functional element in the whole genome and / or whole transcriptome data.
29. The method of claim 1, wherein the neoantigens are selected based on targeted-genome and / or targeted-transcriptome sequencing of the tumor and the normal tissue samples from the patient, the neoantigens overlapping at least one functional element in the targeted genome and / or targeted transcriptome data.
30. The method of any one of the preceding claims, wherein the PCV selected from the group consisting of: neoantigen vaccines, dendritic cell vaccines, oncolytic virus vaccines, peptide vaccines, and DNA / RNA vaccines.Attorney Docket No. : GH0224WO31. The method of any one of the preceding claims, wherein the plurality of molecules extracted from the tissue sample are selected from the group consisting of: DNA, RNA, proteins, DNA methylation, histone modifications, chromatin accessibility, non-coding RNAs, transcription factor binding sites (TFBS), chromatin remodeling complexes, 3D genome architecture, histone variants, nucleosome positioning, and DNA sequences comprising fragmentomic patters.
32. The method of any one of the preceding claims, wherein the sequencing comprises Illumina Sequencing, Nanopore Sequencing, Pacbio, Ion torrent, Sanger Sequencing, lOx genomics, QIAGEN, Oxford Nanopore, Complete Genomics, Ultima Genomics, Element Biosciences, or Singular Genomics.
33. The method of any one of the preceding claims, wherein the sequencing is selected from the group consisting of:, RNA-seq, bisulfite sequencing, ATAC-seq, ChlP-seq, Hi-C, CUT&RUN, Cut&Tag, direct RNA sequencing, Methylated DNA Immunoprecipitation Sequencing (MeDIP-Seq), Methylation Cap Analysis (MethylCap-Seq), Methylation Enriched Sequencing (Methyl-Seq), Methylation-specific Enrichment Sequencing (Methyl-Seq), 5-hmC Enrichment Sequencing (5-hmC-Seq), hybrid capture sequencing, targeted enrichment sequencing, long-read sequencing, single-molecule real-time (SMRT) sequencing, exome sequencing, whole-genome sequencing (WGS), transcriptome sequencing, amplicon sequencing, low-pass sequencing, paired-end sequencing, and / or multiplex sequencing.
34. A computer-implemented method configured to use a previously generated computational model, trained to determine cancer recurrence in a subject treated with a personalized cancer vaccine (PCV), comprising:(a) extracting a plurality of molecules from a tumor and match normal samples obtained from the subject prior to initiation of treatment with the PCV and sequencing a portion of the plurality of molecules to acquire tumor-specific neoantigen and differentially methylated regions (DMRs) data;(b) extracting a plurality of cell-free DNA(cfDNA) from a blood sample obtained from the same subject, wherein the blood sample is collected from the subject at a later timepoint following treatment with the PCV;(c) sequencing a portion of the cfDNA molecules, to acquire neoantigens and DMRs data and selecting the neoantigens and DMRs based on the sequencing data obtained from the tumor and match normal tissue samples in (a) to obtain a plurality of features;Attorney Docket No. : GH0224WO(d) inputting the plurality of features to the previously trained computational model; and(e) outputting from the computational model a determination of whether or not the cancer reoccurred in the subject treated with the PCV.
35. The method of claim 34, wherein the tumor and match normal samples obtained from the subject are genetically assayed to further obtain tumor specific somatic genetic variation.
36. The method of claim 35, wherein the input features for the computational model comprise at least a portion of the somatic genetic variation obtained from the tumor.
37. The method of claims 34 to 46, wherein the tumor and match normal samples obtained from the subject epigenetically assayed to further obtain tumor specific epigenetic patterns.
38. The method of claim 34 to 37, wherein the input features for the computational model comprise at least a portion of the tumor specific epigenetic patterns.
39. The method of claims 38, wherein the tumor specific epigenetic patterns are selected from the group consisting of: cytosine methylation, transcription factor binding sites (TFBS), fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, H3S10ph.
40. The method of any one of the preceding claims, wherein prior to sequencing, the cfDNA molecules are amplified and subsequently selectively enriched to select target regions.
41. The method of any one of the preceding claims, wherein prior to sequencing, the cfDNA molecules are selectively enriched to select target regions and the target regions are subsequently amplified.
42. The method of claims 40 or 41, wherein the target regions comprise a panel including transcription factors and / or transcription factor binding sites(TFBS).
43. The method of claims 40 or 41, wherein the target regions comprise a panel including DMRs in cancer.
44. The method of claims 40 or 41, wherein the target regions comprise a panel including clinically actionable genetic and epigenetic variants.
45. The method of claims 40 or 41, wherein the target regions comprise a panel comprising cancer pathways.Attorney Docket No. : GH0224WO46. The method of claims 40 or 41, wherein the target regions comprise resistance loci to detect variants emerging under therapeutic pressure.
47. The method of claims 40 or 41, wherein the target regions comprise proteomic signatures of disease, treatment response, and or resistance to treatment.
48. The method of claims 40 or 41, wherein the target regions comprise a plurality of panels including transcription factors, TFBS, DMRs in cancer, clinically actionable genetic and epigenetic variants, genes with a function in cancer pathways, resistance loci, and / or proteomic signatures of disease, treatment response, or resistance to treatment.
49. The method of claim 48, further comprising partitioning the sample to selectively enrich each panel from the different partitions.
50. The method of any one of the preceding claims, further comprising retrospective biomarker analysis to correlate biomarker changes with radiographic imaging, clinical data, and / or histology findings, to validate the method’s efficacy in detecting MRD and / or monitoring disease progression and treatment response.
51. The method of any one of the preceding claims, wherein the cancer is selected from the group consisting of: adenocarcinoma, basal cell carcinoma, bladder cancer, bone cancer, brain cancer, breast cancer, cervical cancer, cholangiocarcinoma, colorectal cancer, endometrial cancer, esophageal cancer, gallbladder cancer, gastric cancer, germ cell tumors, glioma, head and neck cancer, hepatocellular carcinoma, Kaposi sarcoma, kidney cancer, lip and oral cavity cancer, liver cancer, lung cancer, melanoma, mesothelioma, neuroendocrine tumors, ovarian cancer, pancreatic cancer, penile cancer, prostate cancer, sarcoma, skin cancer, small cell lung cancer, squamous cell carcinoma, stomach cancer, testicular cancer, thyroid cancer, uterine cancer, vaginal cancer, vulvar cancer.
52. A method for monitoring therapeutic response in a subject treated with a PCV, comprising:(a) obtaining a tumor tissue sample from the subject prior to initiation of treatment with the PCV, extracting a plurality of molecules from the tumor tissue sample, and sequencing the plurality of molecules to identify one or more tumor-derived biomarkers;(b) determining from the tumor-derived biomarkers obtained in (a) a panel of biomarkers characteristic of the tumor sample; and(c) obtaining a blood sample from the subject after initiation of treatment with the PCV, and detecting the selected biomarkers in the blood sample to monitor therapeutic response, wherein a decrease in biomarkers levels indicates a response to treatment, and an increase or unchanged level indicates resistance to treatmentAttorney Docket No. : GH0224WO and / or disease progression, thereby monitoring therapeutic response in the subject treated with the PCV.
53. A method for monitoring therapeutic response in a subj ect treated with a PCV, comprising:(a) obtaining a first blood sample from the subject prior to initiation of treatment with the PCV, extracting a plurality of molecules from the first blood sample, and sequencing the plurality of molecules to determine one or more biomarkers characteristic of a tumor in the subject; and(b) obtaining a second blood sample from the subject after initiation of treatment with the PCV, and determining in the second sample the biomarkers determined in (a) in the first sample to determine changes or levels of the biomarkers; and(c) applying a trained classifier to the biomarker data obtained from the second sample to classify the subject as responsive or resistant to treatment, thereby monitoring therapeutic response in the subject treated with the PCV.
54. A method for monitoring therapeutic response in a subject treated with PCV, the PCV comprising a panel of neoantigens, the method comprising:(a) obtaining a blood sample from the subj ect after initiation of treatment with the PCV, and determining the panel of neoantigens comprising the PCV to monitor therapeutic response in the subject, wherein a decrease in neoantigen levels indicates a response to treatment, and an increase or unchanged level indicates resistance to treatment and / or disease progression, thereby monitoring therapeutic response in the subject treated with the PCV.
55. The method of claim 54, wherein the detecting comprises direct detection of the neoantigens using immunoassays.
56. The method of claim 55, wherein the immunoassays comprise ELISA, western blot, lateral flow immunoassay (LFIA), chemiluminescent immunoassay (CLIA), radioimmunoassay (RIA), fluorescent immunoassay (FIA), multiplex bead-based immunoassay, immunohistochemistry (H4C).
57. The method of claim 54, wherein the detecting comprises indirect detection of the neoantigens through detection of an epigenetic signature characteristic of the neoantigen levels in the subject.
58. The method of claim 57, wherein the epigenetic signature comprises a methylation signature.
59. The method of claim 54, wherein the detecting comprises detecting a metabolomic signature characteristic of the neoantigen levels in the subject.Attorney Docket No. : GH0224WO60. The method of claim 54, wherein the detecting comprises detecting a transcriptomic signature characteristic of the neoantigen levels in the subject.
61. A method for monitoring therapeutic response in a subject treated with a PCV, the PCV comprising a panel of neoantigens, the method comprising: obtaining a tissue sample from the subject after initiation of treatment with the PCV, and detecting the panel of neoantigens comprising the PCV to monitor therapeutic response in the subject, wherein a decrease in neoantigen levels indicates a response to treatment, and an increase or unchanged level indicates resistance to treatment and / or disease progression, thereby monitoring therapeutic response in the subject treated with the PCV.
62. The method of claim 61, wherein the detecting comprises detecting a histological signature characteristic of the neoantigen levels in the subject.
63. The method of claim 61, wherein the detecting comprises detecting a radiomic signature characteristic of the neoantigen levels in the subject.
64. The method of claim 63, wherein the radiomic signature comprises tissue volume, irregularity, entropy, contrast, mean pixel values, histogram features, gray-level cooccurrence matrix.
65. The method of claim 61, wherein the detecting comprises detecting a radigenomic signature characteristic of the neoantigen levels in the subject.
66. A method for monitoring therapeutic response in a subject treated with a PCV, the PCV comprising a panel of neoantigens, the method comprising:(a) obtaining a blood sample from the subject after initiation of treatment with the PCV,(b) determining an epigenetic profile from the blood sample, the epigenetic profile comprising methylation data,(c) applying a trained classifier to the epigenetic data to detect an epigenetic signature indicative of neoantigen levels in the subject after treatment with the PCV, wherein detection of the epigenetic signature classifies the subject as responsive or resistant to treatment, thereby monitoring therapeutic response in the subject treated with the PCV.
67. A method for monitoring therapeutic response in a subject treated with a PCV, the PCV comprising a panel of neoantigens, the method comprising:(a) obtaining a blood sample from the subject after initiation of treatment with the PCV,Attorney Docket No. : GH0224WO(b) determining a neoantigen profile from the blood sample, the neoantigen profile comprising post treatment levels the neoantigen panel in the PCV administered to the subject;(c) applying a trained classifier to the neoantigen profile to detect a neoantigen signature indicative of neoantigen levels, wherein detection of the neoantigen signature enables classification of the subject as responsive or resistant to treatment, thereby monitoring therapeutic response in the subject treated with the PCV.
68. The method of claim 67, wherein the neoantigen signature further comprise genetic and / or epigenetic lesions in the tumor from the subject.
69. The method of claim 68, wherein determining the neoantigen profile comprises extracting a plurality of molecules from the blood sample to obtain genetic, epigenetic, and / or transcriptomic data to determine sequences encoding the neoantigens.
70. The method of any one of claims 52 to 69, further comprising, in response to determining resistance to treatment and / or disease progression in the subject, administering a second dose of the PCV to the subject.
71. The method of any one of claims 52 to 69, further comprising, in response to determining resistance to treatment in the subject,(a) obtaining a second and / or third blood sample from the subject;(b) extracting a plurality of molecules from the sample to determine a new set of neoantigens in the sample obtained from the subject to determine an updated PCV for the subject, and(c) administering the updated PCV to the subject.
72. A method for monitoring therapeutic response in a subject treated with an off-the-shelf cancer vaccine, the method comprising:(a) obtaining a sample from the subject after initiation of treatment with the off-the- shelf cancer vaccine;(b) extracting a plurality of molecules from the sample and sequencing the molecules to obtain sequencing reads;(c) determining one or more features in the sequencing reads; and(d) applying a classifier trained to recognize the features determined in (c) to classify the sample obtained from the subject as responsive or resistant to the off-the-shelf cancer vaccine, thereby monitoring therapeutic response in the subject treated with the off-the-shelf cancer vaccine.Attorney Docket No. : GH0224WO73. The method of claim 71, wherein the one or more features comprise methylation signals indicative of immune activation.
74. The method of claim 71, wherein the sample comprises blood.
75. The method of claim 71, wherein the one or more features comprise methylation signals indicative of treatment response and / or resistance.
76. The method of claim 74, wherein methylation signals comprise differentially methylated regions (DMRs) in immune-regulatory or cancer-associated loci.
77. The method of claim 71, wherein the off-the-shelf cancer vaccine is selected from the group consisting of: Bacillus Calmette-Guerin (BCG), Sipuleucel-T (Provenge), Talimogene laherparepvec (T-VEC), CIMAvax-EGF, GV AX, and New York-ESO-1 vaccine.
78. The method of claim 71, wherein the classifier is trained using a supervised learning algorithm based on labeled datasets from vaccinated and non-vaccinated subjects.
79. A method for monitoring therapeutic response in a subject treated with a cancer vaccine, the method comprising:(a) obtaining a sample from the subject after initiation of treatment with the cancer vaccine;(b) extracting a plurality of molecules from the sample and sequencing the molecules to obtain sequencing reads;(c) determining one or more features in the sequencing reads; and(d) applying a classifier trained to recognize the features determined in (c) to classify the sample obtained from the subject as responsive or resistant to the cancer vaccine, thereby monitoring therapeutic response in the subject treated with the cancer vaccine.
Citation Information
Patent Citations
Compositions and methods for analyzing modified nucleotides
US10260088B2
Hyperactive AID / APOBEC and hmC dominant TET enzymes
US10961525B2
Oligonucleotides
US20010053519A1
Method and apparatus for imaging a sample on a device
US20030152490A1
Digital Counting of Individual Molecules by Stochastic Attachment of Diverse Labels
US20110160078A1