Combination group cell-free DNA monitoring
By enriching cfDNA with polynucleotide probe sets and performing cell-free DNA sequencing with high sequencing depth and extensive target coverage, the accuracy and reliability of cancer status monitoring and vaccine efficacy assessment in the prior art were solved, and efficient and non-invasive monitoring and evaluation were achieved.
Patent Information
- Application Number
- CN202380080069.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-31
- Filing Date
- 2023-10-31
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to accurately and reliably monitor cancer status and vaccine efficacy, especially under conditions of high sequencing read depth and extensive target coverage.
The polynucleotide probe set, including tumor-informed and tumor-uninformed probes, was used to enrich circulating tumor DNA (cfDNA), and monitor the effects of cancer-related mutations and vaccines on tumors through cell-free DNA sequencing methods with high sequencing depth and extensive target coverage.
A significant improvement in non-invasiveness and efficiency has been achieved, enabling accurate monitoring of cancer status and evaluating vaccine efficacy, providing a wider monitoring capability, including monitoring tumor evasion mutations.
Smart Images

Figure CN120225259A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 381,747, filed October 31, 2022, which is hereby incorporated by reference in its entirety. Background Art
[0003] Therapeutic vaccines based on tumor - specific antigens hold great promise as the next generation of personalized cancer immunotherapies. For example, cancers with high mutational burdens, such as non - small cell lung cancer (NSCLC) and melanoma, are particularly attractive targets for such therapies due to the relatively high likelihood of generating neoantigens. Early evidence suggests that neoantigen - based vaccination can trigger T - cell responses, and neoantigen - targeted cell therapies can cause tumor regression in selected patients in some cases.
[0004] One problem in neoantigen vaccine design is which of the many coding mutations present in a subject's tumor can produce the "best" therapeutic neoantigen, e.g., an antigen that can trigger anti - tumor immunity and cause tumor regression. Targeted antigens common to patients with cancer hold great promise as vaccine strategies, including targeting neoantigens with mutations and tumor antigens without mutations (e.g., incorrectly expressed tumor antigens).
[0005] Challenges of shared - antigen vaccine strategies include at least monitoring the cancer status and / or vaccine efficacy before or after administering a cancer vaccine to a subject. For example, many standard methods of monitoring the disease are invasive or burdensome, such as radiological assessments (e.g., CT scans) or tumor biopsies. In addition, certain existing cell - free DNA monitoring methods suffer from reduced ability to monitor cancer status and burden, e.g., reduced monitoring sensitivity, because they only monitor a small subset of mutations (e.g., fewer than 50) related to the tumor exome. Similarly, certain existing cell - free DNA monitoring methods (e.g., Wan et al.; Science Translational Medicine, June 17, 2020: Vol. 12, No. 548) suffer from reduced accuracy and reliability because they only monitor a large number of mutations at low sequencing depths.
[0006] To monitor cfDNA, assays are generally divided into tumor - unaware (tumor - )Determination and tumor-informed determination. Tumor-agnostic monitoring utilizes a panel approach where DNA targets are fixed and tend to capture only a few variants from many patients. Tumor-informed methods rely on sequencing biopsies and longitudinally tracking a defined set of individual variants over time, but typically monitor a smaller footprint. To overcome the smaller footprint of fixed genomic and individual panels, whole exome sequencing (WES) and whole genome sequencing (WGS) of liquid biopsy samples provide an expanded breadth suitable for de novo variant discovery or detection without the need for tissue that may be available for early detection or recurrence. However, the increased breadth of coverage typically increases cost, making these techniques impractical in a clinical setting and / or requiring a lower overall sequencing depth to maintain the cost and use of different bioinformatics strategies.
[0007] Accordingly, there is a need in the art for accurate, reliable, and less invasive cancer monitoring methods, such as cell-free DNA sequencing methods that provide broad target coverage (e.g., at least 95% of the mutations present in the cancer exome) at a high sequencing read depth (e.g., at least 1000-fold). There is also a need for compositions and methods that can effectively monitor subject-specific efficacy as well as broader monitoring capabilities, including monitoring for tumor escape mutations. SUMMARY OF THE INVENTION
[0008] Provided herein is a polynucleotide probe set for enriching cfDNA, the set comprising: (A) one or more tumor-informed polynucleotide probes; and (B) one or more tumor-agnostic polynucleotide probes.
[0009] In some aspects, the one or more tumor-informed polynucleotide probes are configured to capture target sequences comprising epitope sequences encoded by a cancer vaccine administered to a subject, wherein the subject has been determined to have a tumor expressing the epitope sequences. In some aspects, the KRAS mutations are selected from the group consisting of: KRAS_G12C mutation, KRAS_G12D mutation, KRAS_G12V mutation, and KRAS_Q61H mutation. In some aspects, the epitope sequences comprise mutations selected from the group consisting of: KRAS_G13D, KRAS_Q61K, TP53_R249M, CTNNB1_S45P, CTNNB1_S45F, ERBB2_Y772_A775dup, KRAS_G12D, KRAS_Q61R, CTNNB1_T41A, TP53_K132N, KRAS_G12A, KRAS_Q61L, TP53_R213L, BRAF_G466V, KRAS_G12V, KRAS_Q61H, CTNNB1_S37F, TP53_S127Y, TP53_K132E, and KRAS_G12C. In some aspects, the epitope sequences comprise EGFR mutations. In some aspects, the EGFR mutations include the EGFR_L858R mutation.
[0010] In some aspects, the epitope sequences comprise one or more subject-specific epitopes, wherein the subject's tumor has been sequenced to identify the subject-specific epitopes to be encoded by the cancer vaccine. In some aspects, the one or more subject-specific epitopes include at least 2 subject-specific epitopes, at least 10 subject-specific epitopes, at least 20 subject-specific epitopes, or from 2 to 20 subject-specific epitopes. In some aspects, the one or more subject-specific epitopes comprise from 2 to 20 subject-specific epitopes.
[0011] In some aspects, the set further comprises additional tumor-informed polynucleotide probes that capture additional target sequences, wherein the tumor has been determined to express the additional target sequences, and wherein the additional target sequences are not encoded by the cancer vaccine. In some aspects, the additional target sequences include at least 10 target sequences, at least 20 target sequences, at least 30 target sequences, at least 100 target sequences, from 10 to 500 target sequences, from 30 to 500 target sequences, from 100 to 500 target sequences, from 10 to 100 target sequences, from 30 to 100 target sequences, or from 100 to 100 target sequences. In some aspects, the additional target sequences have been predicted to be presented by at least one HLA of the subject.
[0012] In some aspects, the one or more tumor-naïve polynucleotide probes are configured to capture a target sequence that comprises a sequence of interest selected from the group consisting of: cancer-related genes, oncogenes, tumor suppressor genes, interferon-gamma signaling pathway genes, antigen processing pathway genes, and combinations thereof.
[0013] In some aspects, the one or more tumor-naïve polynucleotide probes are configured to capture a target sequence that comprises a sequence of interest selected from each of the following: cancer-related genes, oncogenes, tumor suppressor genes, interferon-gamma signaling pathway genes, and antigen processing pathway genes.
[0014] In some aspects, cancer-related genes are selected from the group consisting of: ABCA12, ACVR2A, AKAP9, BMPR2, COL12A1, CSMD3, DNAH5, DOCK3, FAT2, FAT3, FAT4, FGF10, FGF6, FLG, MAGI1, MDN1, MMAB, NBEA, OBSCN, PCBP1, PCLO, PLEKHA6, PROC, RAD54L, RELN, RPL22, RYR2, TCERG1, WRN, and ZDBF2. In some aspects, cancer-related genes include each of the following: ABCA12, ACVR2A, AKAP9, BMPR2, COL12A1, CSMD3, DNAH5, DOCK3, FAT2, FAT3, FAT4, FGF10, FGF6, FLG, MAGI1, MDN1, MMAB, NBEA, OBSCN, PCBP1, PCLO, PLEKHA6, PROC, RAD54L, RELN, RPL22, RYR2, TCERG1, WRN, and ZDBF2.
[0015] In some aspects, the oncogenes are selected from the group consisting of: ABL1, AKT2, ALK, AR, BCL6, BCL9L, BRAF, BTK, CARD11, CCND1, CCND3, CTNNB1, DDR2, EGFR, ERBB2, ERBB3, FGFR1, FGFR3, FHOD3, FLT1, FLT3, GNAS, HRAS, KDR, KIT, KRAS, MAP2K1, MAP2K2, MECOM, MED12, MET, MTOR, NOTCH1, NOTCH2, NOTCH3, NRAS, NRG1, NTRK1, NTRK3, PDGFRA, PDGFRB, PIK3CA, PIK3CG, PSMB2, RET, ROS1, SF3B1, SMO, SYNE1 and ZBTB20. In some aspects, the oncogenes include each of the following: ABL1, AKT2, ALK, AR, BCL6, BCL9L, BRAF, BTK, CARD11, CCND1, CCND3, CTNNB1, DDR2, EGFR, ERBB2, ERBB3, FGFR1, FGFR3, FHOD3, FLT1, FLT3, GNAS, HRAS, KDR, KIT, KRAS, MAP2K1, MAP2K2, MECOM, MED12, MET, MTOR, NOTCH1, NOTCH2, NOTCH3, NRAS, NRG1, NTRK1, NTRK3, PDGFRA, PDGFRB, PIK3CA, PIK3CG, PSMB2, RET, ROS1, SF3B1, SMO, SYNE1 and ZBTB20.
[0016] In some aspects, the tumor suppressor genes are selected from the group consisting of: TP53, PTEN, ARID1A, APC, AMER1, ASXL1, ATM, ATR, ATRX, AXIN2, BARD1, BRCA1, BRCA2, CASP8, CFH, CREBBP, DNMT3A, EP300, ERCC2, ERCC5, EXT1, FANCA, FANCD2, FANCI, FANCM, FAT1, FBXW7, HNF1A, MAX, MLH1, MSH3, MSH6, NF1, PIK3R1, PTCH1, PTPRT, RECQL4, RNF43, ROBO1, SLX4, SMAD2, SMAD3, SMAD4, SOX9, TCF7L2, TERT promoter, TET2, TGFBR2, TP53BP1, TSC1, TSC2, WNT16, XPC, ZFP36L2, and ZNRF3. In some aspects, the tumor suppressor genes include each of the following: TP53, PTEN, ARID1A, APC, AMER1, ASXL1, ATM, ATR, ATRX, AXIN2, BARD1, BRCA1, BRCA2, CASP8, CFH, CREBBP, DNMT3A, EP300, ERCC2, ERCC5, EXT1, FANCA, FANCD2, FANCI, FANCM, FAT1, FBXW7, HNF1A, MAX, MLH1, MSH3, MSH6, NF1, PIK3R1, PTCH1, PTPRT, RECQL4, RNF43, ROBO1, SLX4, SMAD2, SMAD3, SMAD4, SOX9, TCF7L2, TERT promoter, TET2, TGFBR2, TP53BP1, TSC1, TSC2, WNT16, XPC, ZFP36L2, and ZNRF3.
[0017] In some aspects, the interferon-γ signaling pathway genes are selected from the group consisting of: IFNGR1, INFGR2, JAK1, JAK2, and STAT1. In some aspects, the interferon-γ signaling pathway genes include each of the following: IFNGR1, INFGR2, JAK1, JAK2, and STAT1.
[0018] In some aspects, the antigen processing pathway genes are selected from the group consisting of: B2M, HLA-A, HLA-B, HLA-C, HLA-E, TAP1, TAP2, NLRC5, CALR, CANX, PSMB2, and TAPBP. In some aspects, the antigen processing pathway genes include each of the following: B2M, HLA-A, HLA-B, HLA-C, HLA-E, TAP1, TAP2, NLRC5, CALR, CANX, PSMB2, and TAPBP.
[0019] In some aspects, one or more tumor-agnostic polynucleotide probes are configured to capture target sequences that include sequences of interest selected from the group consisting of: ABCA12, ACVR2A, AKAP9, BMPR2, COL12A1, CSMD3, DNAH5, DOCK3, FAT2, FAT3, FAT4, FGF10, FGF6, FLG, MAGI1, MDN1, MMAB, NBEA, OBSCN, PCBP1, PCLO, PLEKHA6, PROC, RAD54L, RELN, RPL22, RYR2, TCERG1, WRN, ZDBF2, ABL1, AKT2, ALK, AR, BCL6, BCL9L, BRAF, BTK, CARD11, CCND1, CCND3, CTNNB1, DDR2, EGFR, ERBB2, ERBB3, FGFR1, FGFR3, FHOD3, FLT1, FLT3, GNAS, HRAS, KDR, KIT, KRAS, MAP2K1, MAP2K2, MECOM, MED12, MET, MTOR, NOTCH1, NOTCH2, NOTCH3, NRAS, NRG1, NTRK1, NTRK3, PDGFRA, PDGFRB, PIK3CA, PIK3CG, RET, ROS1, SF3B1, SMO, SYNE1, ZBTB20, TP53, PTEN, ARID1A, APC, AMER1, ASXL1, ATM, ATR, ATRX, AXIN2, BARD1, BRCA1, BRCA2, CASP8, CFH, CREBBP, DNMT3A, EP300, ERCC2, ERCC5, EXT1, FANCA, FANCD2, FANCI, FANCM, FAT1, FBXW7, HNF1A, MAX, MLH1, MSH3, MSH6, NF1, PIK3R1, PTCH1, PTPRT, RECQL4, RNF43, ROBO1, SLX4, SMAD2, SMAD3, SMAD4, SOX9, TCF7L2, TERT promoter, TET2, TGFBR2, TP53BP1, TSC1, TSC2, WNT16, XPC, ZFP36L2, ZNRF3, IFNGR1, INFGR2, JAK1, JAK2, STAT1, B2M, HLA-A, HLA-B, HLA-C, HLA-E, TAP1, TAP2, NLRC5, CALR, CANX, PSMB2, TAPBP, and combinations thereof.
[0020] In some aspects, one or more tumor-agnostic polynucleotide probes are configured to capture a target sequence that includes a sequence of interest selected from each of the following: ABCA12, ACVR2A, AKAP9, BMPR2, COL12A1, CSMD3, DNAH5, DOCK3, FAT2, FAT3, FAT4, FGF10, FGF6, FLG, MAGI1, MDN1, MMAB, NBEA, OBSCN, PCBP1, PCLO, PLEKHA6, PROC, RAD54L, RELN, RPL22, RYR2, TCERG1, WRN, ZDBF2, ABL1, AKT2, ALK, AR, BCL6, BCL9L, BRAF, BTK, CARD11, CCND1, CCND3, CTNNB1, DDR2, EGFR, ERBB2, ERBB3, FGFR1, FGFR3, FHOD3, FLT1, FLT3, GNAS, HRAS, KDR, KIT, KRAS, MAP2K1, MAP2K2, MECOM, MED12, MET, MTOR, NOTCH1, NOTCH2, NOTCH3, NRAS, NRG1, NTRK1, NTRK3, PDGFRA, PDGFRB, PIK3CA, PIK3CG, PSMB2, RET, ROS1, SF3B1, SMO, SYNE1, ZBTB20, TP53, PTEN, ARID1A, APC, AMER1, ASXL1, ATM, ATR, ATRX, AXIN2, BARD1, BRCA1, BRCA2, CASP8, CFH, CREBBP, DNMT3A, EP300, ERCC2, ERCC5, EXT1, FANCA, FANCD2, FANCI, FANCM, FAT1, FBXW7, HNF1A, MAX, MLH1, MSH3, MSH6, NF1, PIK3R1, PTCH1, PTPRT, RECQL4, RNF43, ROBO1, SLX4, SMAD2, SMAD3, SMAD4, SOX9, TCF7L2, TERT promoter, TET2, TGFBR2, TP53BP1, TSC1, TSC2, WNT16, XPC, ZFP36L2, ZNRF3, IFNGR1, INFGR2, JAK1, JAK2, STAT1, B2M, HLA-A, HLA-B, HLA-C, HLA-E, TAP1, TAP2, NLRC5, CALR, CANX, PSMB2, and TAPBP.
[0021] In some aspects, one or more tumor-naïve polynucleotide probes are configured to capture a target sequence that includes a sequence of interest selected from each of the following: ABL1, AKT2, ALK, APC, AR, ATR, ATRX, BARD1, BCL6, BMPR1A, BRAF, BRCA1, BRCA2, BTK, CARD11, CCND1, CCND3, CDK12, CFH, CREBBP, CTNNB1, DDR2, DNMT3A, EGFR, EP300, ERBB2, ERBB3, ERCC2, ERCC5, EXT1, FANCA, FANCD2, FANCI, FANCM, FBXW7, FGF10, FGF6, FGFR1, FGFR3, FLI1, FLT1, FLT3, GNAS, HNF1A, HRAS, KDR, KIT, KRAS, MAGI1, MAP2K1, MAP2K2, MAX, MED12, MET, MLH1, MMAB, MSH3, MSH6, MTOR, NF1, NFE2L2, NOTCH1, NOTCH2, NOTCH3, NRAS, NRG1, NTRK1, NTRK3, PDGFRA, PDGFRB, PIK3CA, PIK3CG, PIK3R1, PMS2, PPARG, PROC, PTCH1, RAD54L, RAF1, RECQL4, RET, ROS1, SF3B1, SF3B2, SLX4, SMO, TERT promoter, TET2, TP53BP1, TSC1, TSC2, WRN, XPA, XPC, ZNF395, B2M, HLA-A, HLA-B, HLA-C, TAP1, TAP2, NLRC5, IFNGR1, INFGR2, JAK1, JAK2, TP53, PTEN, and ARID1A.
[0022] In some aspects, one or more tumor-naïve polynucleotide probes include two or more probes configured to capture all of the coding exon sequences of a given gene. In some aspects, one or more tumor-naïve polynucleotide probes include two or more probes configured to capture genomic regions of interest associated with cancer.
[0023] In some aspects, tumor-informed polynucleotide probes and / or tumor-naïve polynucleotide probes include probes that contain overlapping sequences.
[0024] In some aspects, the set comprises at least 20 probes, at least 30 probes, at least 40 probes, at least 50 probes, at least 60 probes, at least 70 probes, at least 80 probes, at least 90 probes, at least 100 probes, at least 200 probes, at least 300 probes, at least 400 probes, or at least 500 probes.
[0025] In some aspects, the set is configured to cover at least 100 kb, at least 300 kb, at least 300 kb, at least 400 kb, 100 to 400 kb, 200 to 400 kb, 300 to 400 kb, 100 to 500 kb, 200 to 500 kb, 300 to 500 kb, or 340 to 400 kb of the subject's genome.
[0026] In some aspects, one or more tumor-naïve polynucleotide probes include polynucleotide probes configured to capture sequences associated with a given cancer that the subject is known or suspected to have, optionally wherein the cancer is CRC or NSCLC.
[0027] In some aspects, the set further comprises additional polynucleotide probes configured to capture sequences containing polymorphisms in the human population, wherein the sequences containing polymorphisms can be combined to uniquely identify the subject.
[0028] Also provided herein is a method for enriching cfDNA, the method comprising: (a) providing a sample comprising cfDNA; (b) providing a set of polynucleotide probes comprising any one of the tumor-informed / tumor-naïve combination sets provided herein; (c) contacting the sample comprising cfDNA with the set of polynucleotide probes under conditions sufficient to hybridize the cfDNA containing the target sequence of interest to its corresponding polynucleotide probe; and (d) capturing the hybridized cfDNA and polynucleotide probe pairs to enrich cfDNA.
[0029] The present disclosure also provides a method for monitoring the cancer status of a subject having, having had, or suspected of having cancer, the method comprising the steps of: a. obtaining or having obtained sequencing data of cell-free DNA (cfDNA) from a sample from the subject, and wherein the sequencing data comprises a target coverage of at least 50% of all polynucleotide regions of interest corresponding to mutations present in the cancer exome, and wherein the polynucleotide regions of interest being sequenced have a read depth of at least 1000-fold, optionally wherein the polynucleotide regions of interest contain at least 50 mutations, optionally wherein the average read depth is the average duplex read depth, and optionally wherein obtaining the sequencing data comprises collecting or having collected a sample from the subject, isolating or having isolated cfDNA, enriching or having enriched cfDNA, and / or sequencing or having sequenced cfDNA; and b. determining or having determined the mutation frequency present in the exome to assess the cancer status, optionally wherein the assessment of the status comprises an assessment of the presence and / or cancer burden, wherein the cfDNA has been enriched prior to sequencing using: (1) a subject-specific polynucleotide probe set; (2) a tumor-naive polynucleotide probe set; and / or (3) a combined set of a tumor-informed polynucleotide probe set and a tumor-naive polynucleotide probe set, wherein the polynucleotide probes are configured to capture the polynucleotide regions of interest.
[0030] The present disclosure also provides a method for monitoring the cancer status of a subject having, having had, or suspected of having cancer, the method comprising the steps of: a. obtaining or having obtained sequencing data of cell-free DNA (cfDNA) from a sample from the subject, and wherein the sequencing data comprises a target coverage of at least 95% of all polynucleotide regions of interest corresponding to mutations present in the cancer exome, wherein the polynucleotide regions of interest contain at least 50 mutations, and wherein the polynucleotide regions of interest being sequenced have a duplex read depth of at least 1000-fold, and optionally wherein obtaining the sequencing data comprises collecting or having collected a sample from the subject, isolating or having isolated cfDNA, enriching or having enriched cfDNA, and / or sequencing or having sequenced cfDNA; and b. determining or having determined the frequency of at least 50 mutations present in the exome to assess the cancer status, optionally wherein the assessment of the status comprises an assessment of the presence and / or cancer burden, wherein the cfDNA has been enriched prior to sequencing using: (1) a tumor-informed polynucleotide probe set; (2) a tumor-naive polynucleotide probe set; and / or (3) a combined set of a tumor-informed polynucleotide probe set and a tumor-naive polynucleotide probe set, wherein the polynucleotide probes are configured to capture the polynucleotide regions of interest.
[0031] The present disclosure also provides a method for evaluating the efficacy of a therapy in a subject having, having had, or suspected of having cancer, the method comprising the steps of: a. obtaining or having obtained sequencing data of cell-free DNA (cfDNA) from a pre-therapy sample from the subject, wherein the sequencing data comprises target coverage of at least 50% of all polynucleotide regions of interest corresponding to mutations present in the cancer exome, and wherein the polynucleotide regions of interest being sequenced have a read depth of at least 1000-fold, optionally wherein the polynucleotide regions of interest contain at least 50 mutations, optionally wherein the average read coverage is the average duplex read coverage, and optionally wherein obtaining the sequencing data comprises collecting or having collected a pre-therapy sample from the subject, isolating or having isolated pre-therapy cfDNA, enriching or having enriched pre-therapy cfDNA and / or sequencing or having sequenced pre-therapy cfDNA; b. obtaining or having obtained sequencing data of cell-free DNA (cfDNA) from a post-therapy sample from the subject, optionally wherein the therapy comprises a cancer vaccine comprising a neoantigen or an expression system encoding the neoantigen, and wherein the sequencing data comprises target coverage of at least 50% of all polynucleotide regions of interest corresponding to mutations present in the cancer exome, and wherein the polynucleotide regions of interest being sequenced have a read depth of at least 1000-fold, optionally wherein the polynucleotide regions of interest contain at least 50 mutations, optionally wherein the average read coverage is the average duplex read coverage, and optionally wherein obtaining the sequencing data comprises collecting or having collected a post-therapy sample from the subject, isolating or having isolated post-therapy cfDNA, enriching or having enriched post-therapy cfDNA and / or sequencing or having sequenced post-therapy cfDNA; and c. determining or having determined the frequency of mutations present in the exome of the pre-therapy cfDNA relative to the post-therapy cfDNA to evaluate the efficacy of the therapy, optionally wherein an increase in the frequency of mutations in the post-therapy cfDNA relative to the pre-therapy cfDNA indicates an increased likelihood of an increase in the tumor burden of the subject, and optionally wherein a decrease or maintenance in the frequency of mutations in the post-therapy cfDNA relative to the pre-therapy cfDNA indicates an increased likelihood of a decrease or stability in the tumor burden of the subject, wherein the cfDNA has been enriched prior to sequencing using: (1) a tumor-informed polynucleotide probe set; (2) a tumor-uninformed polynucleotide probe set; and / or (3) a combined set of a tumor-informed polynucleotide probe set and a tumor-uninformed polynucleotide probe set, wherein the polynucleotide probes are configured to capture polynucleotide regions of interest.
[0032] The present disclosure also provides a method for assessing the efficacy of a therapy in a subject having, having had, or suspected of having cancer, the method comprising the steps of: a. obtaining or having obtained sequencing data of tumor-derived DNA from cancer-affected tissue from the subject, optionally wherein obtaining the sequencing data comprises collecting or having collected cancer-affected tissue, isolating or having isolated tumor-derived DNA, and sequencing or having sequenced the tumor-derived DNA; b. determining or having determined from the tumor-derived DNA sequencing data one or more tumor-associated mutations relative to the wild-type germline nucleic acid sequence of the subject, optionally wherein one or more of the one or more tumor-associated mutations are associated with a neoantigen comprising at least one alteration such that the peptide sequence encoded by the tumor-derived DNA is different from the corresponding peptide sequence encoded by the wild-type germline nucleic acid sequence of the subject; c. designing and / or selecting or having designed and / or having selected (1) a tumor-informed polynucleotide probe set; (2) a tumor-uninformed polynucleotide probe set; and / or (3) a combined set of a tumor-informed polynucleotide probe set and a tumor-uninformed polynucleotide probe set, wherein the polynucleotide probes are configured to capture at least the tumor-associated mutations, optionally wherein the polynucleotide regions of interest comprise at least 50 tumor-associated mutations; d. obtaining or having obtained sequencing data of cell-free DNA (cfDNA) from a pre-therapy sample from the subject, wherein the pre-therapy cfDNA is enriched using polynucleotide probes prior to sequencing, and wherein the sequencing data comprises a target coverage of at least 50% of all polynucleotide regions of interest corresponding to the tumor-associated mutations, and wherein the polynucleotide regions of interest sequenced have a read depth of at least 1000-fold, optionally wherein the average read coverage is the average duplex read coverage, and optionally wherein obtaining the sequencing data comprises collecting or having collected the pre-therapy sample from the subject, isolating or having isolated the pre-therapy cfDNA, enriching or having enriched the pre-therapy cfDNA, and / or sequencing or having sequenced the pre-therapy cfDNA; e. obtaining or having obtained sequencing data of cell-free DNA (cfDNA) from a post-therapy sample from the subject, optionally wherein the therapy comprises a cancer vaccine comprising a neoantigen or an expression system encoding the neoantigen, wherein the post-therapy cfDNA is enriched using polynucleotide probes prior to sequencing, and wherein the sequencing data comprises a target coverage of at least 50% of all polynucleotide regions of interest corresponding to the tumor-associated mutations, and wherein the polynucleotide regions of interest sequenced have a read depth of at least 1000-fold, optionally wherein the average read coverage is the average duplex read coverage, and optionally wherein obtaining the sequencing data comprises collecting or having collected the post-therapy sample from the subject, isolating or having isolated the post-therapy cfDNA, enriching or having enriched the post-therapy cfDNA, and / or sequencing or having sequenced the post-therapy cfDNA; and f.Determine or have determined the frequency of tumor - associated mutations in cfDNA before therapy relative to cfDNA after therapy to evaluate the efficacy of the therapy, optionally wherein at least one or more tumor - associated mutations related to neoantigens are determined, optionally wherein an increase in the frequency of mutations in cfDNA after therapy relative to cfDNA before therapy indicates an increased likelihood of an increase in the tumor burden of the subject, and optionally wherein a decrease or maintenance in the frequency of mutations in cfDNA after therapy relative to cfDNA before therapy indicates an increased likelihood of a decrease or stability in the tumor burden of the subject.
[0033] In some aspects, the method includes designing and / or selecting or having designed and / or having selected a combined set of tumor - informed polynucleotide probes and tumor - uninformed polynucleotide probes. In some aspects, the combined set that is designed and / or selected includes any of the tumor - informed / tumor - uninformed combined sets provided herein.
[0034] Also provided herein is a method for enriching cfDNA, the method comprising: (a) providing a sample comprising cfDNA; (b) providing a polynucleotide probe set, wherein the set comprises: (i) one or more tumor - informed polynucleotide probes; and (ii) one or more tumor - uninformed polynucleotide probes; (c) contacting the sample comprising cfDNA with the polynucleotide probe set under conditions sufficient to hybridize cfDNA comprising a target sequence of interest to its corresponding polynucleotide probe; and (d) capturing the hybridized cfDNA and polynucleotide probe pairs to enrich cfDNA. In some aspects, the set includes any of the tumor - informed / tumor - uninformed combined sets provided herein.
[0035] In some aspects, the method includes one or more of the following steps: a. collecting or having collected a sample from a subject; b. isolating or having isolated cfDNA; c. enriching or having enriched cfDNA; or d. sequencing or having sequenced cfDNA.
[0036] In some aspects, the method includes each of the following steps: a. collecting or having collected a sample from a subject; b. isolating or having isolated the cfDNA; c. enriching or having enriched the cfDNA; and d. sequencing or having sequenced cfDNA.
[0037] In some aspects, the average read depth includes an average read coverage of at least 1500-fold, at least 2000-fold, at least 2500-fold, 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. In some aspects, the average read depth includes an average read coverage in the range of 1000-fold to 5000-fold. In some aspects, the average read depth includes an average read coverage in the ranges of 1000-fold to 4000-fold, 1000-fold to 3000-fold, 1000-fold to 2000-fold, 2000-fold to 5000-fold, 2000-fold to 4000-fold, 2000-fold to 3000-fold, 3000-fold to 5000-fold, 3000-fold to 4000-fold, or 4000-fold to 5000-fold. In some aspects, the average read depth includes an average read duplex depth.
[0038] In some aspects, each polynucleotide region of interest corresponding to a mutation present in the exome has a read depth of at least 1000-fold. In some aspects, each polynucleotide region of interest corresponding to a mutation present in the exome has a read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. In some aspects, the target coverage includes at least 60%, at least 70%, at least 80%, or at least 90% of the polynucleotide regions of interest corresponding to mutations present in the cancer exome. In some aspects, the target coverage includes at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9%, or 100% of the polynucleotide regions of interest corresponding to mutations present in the cancer exome.
[0039] In some aspects, the target coverage includes at least 95% of the polynucleotide regions of interest corresponding to mutations present in the cancer exome. In some aspects, the polynucleotide region of interest contains at least 50, at least 60, at least 70, at least 80, or at least 90 mutations.
[0040] In some aspects, the polynucleotide region of interest contains at least 50 mutations. In some aspects, the polynucleotide region of interest contains at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1000 mutations.
[0041] In some aspects, the method comprises the steps of: a. obtaining or having obtained sequencing data of tumor-derived DNA from cancerous tissue of a subject, optionally wherein obtaining the sequencing data comprises collecting or having collected cancerous tissue, isolating or having isolated tumor-derived DNA, and sequencing or having sequenced the tumor-derived DNA; b. determining or having determined from the tumor-derived DNA sequencing data one or more tumor-associated mutations relative to the wild-type germline nucleic acid sequence of the subject, optionally wherein one or more of the one or more tumor-associated mutations are associated with a neoantigen comprising at least one alteration that causes the peptide sequence encoded by the tumor-derived DNA to be different from the corresponding peptide sequence encoded by the wild-type germline nucleic acid sequence of the subject; c. designing and / or selecting or having designed and / or having selected (1) a tumor-informed polynucleotide probe set; (2) a tumor-naive polynucleotide probe set; and / or (3) a combined set of a tumor-informed polynucleotide probe set and a tumor-naive polynucleotide probe set, wherein the polynucleotide probes are configured to capture regions of polynucleotides of interest corresponding to the tumor-associated mutations, optionally wherein the regions of polynucleotides of interest comprise at least 50 tumor-associated mutations; and d. enriching or having enriched cfDNA using the polynucleotide probes prior to sequencing.
[0042] In some aspects, the cancer is selected from the group consisting of: lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer.
[0043] In some aspects, the subject has been administered a therapy. In some aspects, the therapy comprises a cancer vaccine. In some aspects, the cancer vaccine comprises an epitope-encoding nucleic acid sequence encoding at least one mutation present in the cancer exome. In some aspects, the cancer vaccine comprises a self-amplifying alphavirus-based expression system. In some aspects, the cancer vaccine comprises a chimpanzee adenovirus (ChAdV)-based expression system.
[0044] In some aspects, the method includes obtaining sequencing data of cfDNA from two or more samples from a subject. In some aspects, two or more samples are collected at different time points. In some aspects, the two or more samples are collected at different time points relative to the administration of a therapy. In some aspects, a pre-therapy sample is collected before the administration of the therapy, and post-therapy cfDNA is collected after the administration of the therapy. In some aspects, the determining step includes determining or having determined the mutation frequency of pre-therapy cfDNA relative to post-therapy cfDNA to evaluate the efficacy of the therapy, optionally wherein at least one or more tumor-associated mutations related to neoantigens are determined, optionally wherein an increase in the frequency of mutations in post-therapy cfDNA relative to pre-therapy cfDNA indicates an increased likelihood of an increase in the tumor burden of the subject, and optionally wherein a decrease or maintenance in the frequency of mutations in post-therapy cfDNA relative to pre-therapy cfDNA indicates an increased likelihood of a decrease or stability in the tumor burden of the subject.
[0045] In some aspects, an increase in the frequency of one or more mutations in post-therapy cfDNA relative to pre-therapy cfDNA in a tumor-naive group indicates the likelihood of tumor mutations in an immune evasion mechanism. In some aspects, an increase in the frequency of mutations in post-therapy cfDNA relative to pre-therapy cfDNA indicates an increased likelihood of an increase in the tumor burden of the subject. In some aspects, a decrease or maintenance in the frequency of mutations in post-therapy cfDNA relative to pre-therapy cfDNA indicates an increased likelihood of a decrease or stability in the tumor burden of the subject. In some aspects, the decrease includes a complete response (CR) or a partial response (PR).
[0046] In some aspects, the method further includes administering a therapy to the subject after assessing the cancer status. In some aspects, the assessment of the mutation frequency in cfDNA indicates the likelihood that the subject has or still has cancer.
[0047] In some aspects, the therapy includes a cancer vaccine. In some aspects, the cancer vaccine contains an epitope-encoding nucleic acid sequence that encodes at least one mutation present in the exome. In some aspects, the cancer vaccine contains a self-amplifying alphavirus-based expression system. In some aspects, the cancer vaccine contains a chimpanzee adenovirus (ChAdV)-based expression system.
[0048] In some aspects, the collecting step includes collecting a blood sample.
[0049] In some aspects, the separating step includes centrifuging to separate cfDNA from cells and / or cell debris. In some aspects, the separating step includes separating cfDNA from whole blood. In some aspects, separating cfDNA from whole blood includes separating the plasma layer, the buffy coat, and the red blood cells. In some aspects, cfDNA is separated from the plasma layer.
[0050] In some aspects, the sequencing step includes next-generation sequencing (NGS) or Sanger sequencing. In some aspects, NGS includes duplex sequencing, whole exome sequencing, whole genome sequencing, de novo sequencing, phased sequencing, targeted amplicon sequencing, or shotgun sequencing.
[0051] In some aspects, the enrichment step includes enriching a region of a polynucleotide of interest corresponding to a mutation present in the exome in cfDNA prior to sequencing. In some aspects, the enrichment includes a combination of a tumor-informed polynucleotide probe set and a tumor-uninformed polynucleotide probe set. In some aspects, each of the tumor-informed polynucleotide probe set and the tumor-uninformed polynucleotide probe set in a separate sample is enriched separately.
[0052] In some aspects, the tumor-informed polynucleotide probe contains each region of a polynucleotide of interest corresponding to a mutation present in the exome. In some aspects, the tumor-informed polynucleotide probe contains at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9%, or 100% of the regions of a polynucleotide of interest corresponding to a mutation present in a cancer exome. In some aspects, the tumor-informed polynucleotide probe contains at least 50, at least 60, at least 70, at least 80, at least 90 mutations, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1000 mutations, optionally, these mutations are present in the exome of a cancer.
[0053] In some aspects, the enrichment step includes hybridizing one or more polynucleotide probes to one or more regions of a polynucleotide of interest.
[0054] In some aspects, the polynucleotide probe has a length of 80 to 150 base pairs (bp). In some aspects, the length of the polynucleotide probe is 50-100, 50-150, 80 to 140, 80 to 130, 80 to 120, 80 to 110, 80 to 100, 80 to 90, 90 to 150, 90 to 140, 90 to 130, 90 to 120, 90 to 110, 90 to 100, 100 to 150, 100 to 140, 100 to 130, 100 to 120, 100 to 110, 110 to 150, 110 to 140, 110 to 130, 110 to 120, 120 to 150, 120 to 140, 120 to 130, 130 to 150, 130 to 140, 140 to 150, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140 or 150 bp.
[0055] In some aspects, one or more polynucleotide probes are biotinylated.
[0056] In some aspects, tumor-informed polynucleotide probes are designed or selected after sequencing the tumor of a subject. In some aspects, tumor-informed polynucleotide probes are designed or selected after exome sequencing of the tumor of a subject. In some aspects, the tumor-informed polynucleotide probes are designed or selected to target all mutations of the sequenced tumor.
[0057] In some aspects, the sequencing step includes ligating sequencing adapters to cfDNA. In some aspects, the sequencing adapters are configured for dual sequencing.
[0058] In some aspects, one or more mutations include point mutations, frameshift mutations, non-frameshift mutations, deletion mutations, insertion mutations, splicing variants, genomic rearrangements, proteasome-generated spliced antigens, or combinations thereof. In some aspects, one or more mutations include at least one alteration that changes such that the peptide sequence encoded by the cfDNA is different from the corresponding peptide sequence encoded by the wild-type germline nucleic acid sequence of the subject. In some aspects, one or more mutations consist of coding mutations comprising at least one alteration that changes such that the peptide sequence encoded by the cfDNA is different from the corresponding peptide sequence encoded by the wild-type germline nucleic acid sequence of the subject. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] These and other features, aspects, and advantages of the present invention will be better understood with reference to the following description and drawings, in which:
[0060] Figure 1Shows a detailed process for separating and processing ctDNA from a patient. Briefly, tumor-specific DNA variant alleles are identified from biopsied tumor tissue (point 1). Blood is drawn from the patient at specific time points of the dosing regimen, ctDNA is isolated and used to generate a UMI library (points 2 and 4). Baits designed based on the variants identified in the patient's tumor DNA (point 3) are used to purify the ctDNA containing the identified variants (point 5).
[0061] Figure 2 Shows after separation and processing for analysis of ctDNA isolated from a patient after separation and processing as outlined in Figure 1 . The purified ctDNA is sequenced (point 6) to quantify the prevalence of specific identified variants. Repeated testing of ctDNA during treatment allows monitoring of tumor progression or response to therapy.
[0062] Figure 3A Illustrates the isolation and sequencing of circulating tumor DNA (ctDNA) in two patients receiving GRANITE therapy, and is a graph showing the absolute duplex read coverage of DNA variants identified in ctDNA isolated from patient #1 (identified as pt0009).
[0063] Figure 3B Illustrates the isolation and sequencing of circulating tumor DNA (ctDNA) in two patients receiving GRANITE therapy, and is a graph showing the normalized duplex read coverage of DNA variants identified in ctDNA isolated from patient #1 (identified as pt0009).
[0064] Figure 3C Illustrates the isolation and sequencing of circulating tumor DNA (ctDNA) in two patients receiving GRANITE therapy, and is a graph showing the monitoring of tumor-specific DNA variant alleles in patient #1 during treatment, highlighting TP52 R175H, APC T1556fs, and CDKN2A W110*.
[0065] Figure 3D Illustrates the isolation and sequencing of circulating tumor DNA (ctDNA) in two patients receiving GRANITE therapy, and is a graph showing the absolute duplex read coverage of DNA variants identified in ctDNA isolated from patient #2 (identified as pt0005).
[0066] Figure 3EIllustrates the isolation and sequencing of circulating tumor DNA (ctDNA) in two patients receiving GRANITE therapy, and is a graph showing the normalized duplex read coverage of DNA variants identified in ctDNA isolated from Patient #2 (identified as pt0005).
[0067] Figure 3F Illustrates the isolation and sequencing of circulating tumor DNA (ctDNA) in two patients receiving GRANITE therapy, and is a graph showing the monitoring of tumor-specific DNA variant alleles in Patient #2 during treatment, said alleles including TRABD2B A385T, ADAR G751R, VILL L273fs, SURF2 P146L, TP53 P153fs, CSH2 A156V, and MAP2K2 E66K.
[0068] Figure 4A Is a graph showing the monitoring of variant allele frequencies (VAFs) in Patient #1 (pt0009) during GRANITE therapy, and shows the frequencies of 11 identified tumor-specific variant alleles during treatment.
[0069] Figure 4B Is a graph showing the monitoring of variant allele frequencies (VAFs) in Patient #1 (pt0009) during GRANITE therapy, and shows the VAF trends of all variant alleles in ctDNA isolated during treatment.
[0070] Figure 4C Is a graph showing the monitoring of variant allele frequencies (VAFs) in Patient #1 (pt0009) during GRANITE therapy, and shows the average percentage change in VAF between consecutive doses during treatment.
[0071] Figure 5A Is a graph illustrating the monitoring of ctDNA in additional patients receiving GRANITE therapy, and shows the monitoring of ctDNA in patients with non-small cell lung cancer (NSCLC) receiving GRANITE therapy.
[0072] Figure 5B Is a graph illustrating the monitoring of ctDNA in additional patients receiving GRANITE therapy, and shows the tracking of ctDNA in patients with microsatellite stable colorectal cancer (MSS-CRC).
[0073] Figure 6AIt is a graph showing the monitoring of ctDNA in a patient (identified as pt0101) receiving SLATE therapy, and shows the absolute duplex read coverage of a specific KRAS allele variant in ctDNA isolated from the patient's plasma.
[0074] Figure 6B It is a graph showing the monitoring of ctDNA in a patient (identified as pt0101) receiving SLATE therapy, and shows the normalized duplex read coverage of a specific KRAS allele variant in ctDNA isolated from the patient's plasma.
[0075] Figure 6C It is a graph showing the monitoring of ctDNA in a patient (identified as pt0101) receiving SLATE therapy, and shows the change in KRAS variant allele duplexes between consecutive doses.
[0076] Figure 7 It is a graph showing the monitoring of ctDNA associated with the KRAS G12C mutation in patients with NSCLC.
[0077] Figure 8 Shows the detailed process for isolating and processing ctDNA from a patient for patient-specific vaccine screening and manufacturing.
[0078] Figure 9 Shows a schematic diagram of the ctDNA monitoring assay. A shotgun library of cfDNA, biopsy DNA, or gDNA from whole blood is prepared with duplex UMI. Duplex sequencing reduces noise by requiring variants to be observed on both strands of a duplex molecule. Multiple patient-specific pools are combined to create a superset containing probes for 6 - 9 patients. The universal set captures a common set of targets in all patient samples.
[0079] Figure 10A Shows the number of potential variants covered by the indicated NGS panel. Figure 10B Shows the percentage of WES variants potentially covered by various NGS panels.
[0080] Figure 11 Shows the blood collection protocol for patients enrolled in SLATE ("off-the-shelf" vaccine program) and GRANITE ("personalized cancer vaccine" program).
[0081] Figure 12 Shows that for variant calls with >1000-fold duplex consensus coverage, patient assays monitored an average of approximately 140 variants per patient at high sequencing depth.
[0082] Figure 13AShows the cassette mutations observed in the ctDNA and biopsies of the indicated patients. * Indicates patients without a biopsy or with too low tumor content to detect variants in the assay. Figure 13B Shows that when comparing variants in cfDNA and corresponding biopsies using the GRANITE assay, significant overlap was found, particularly the ability to call variants at lower frequencies in high-quality (RNALater or fresh frozen) biopsies.
[0083] Figure 14A Shows the presence of de novo variants in cfDNA samples from the indicated patients and tumor tissue types (GEA, CRC, or NSCLC). Additional variants were present in the cfDNA of many patients that were not present in the initial biopsies. New variants often occurred at positions where another patient had a targeted variant. Figure 14B Shows that using matched normal gDNA from whole blood or PMBC, CHIP mutations were identified and excluded as somatic tumor variants. Figure 14C Shows representative patient G08 with two NLRC5 mutations and two TAP1 mutations, one of the NLRC5 mutations tracked with the average VAF of all variants, and the two TAP1 mutations appearing nearly a year after therapy. Figure 14D Shows a summary of additional analysis of variants observed in cfDNA and found outside of patient-specific variants. Figure 14E Shows the G09 cfDNA kinetics of new variants including multiple KRAS variants.
[0084] Figure 15 Shows that all patient-specific variants captured in the WES of the biopsy were also captured using the patient-specific assay (100% concordance).
[0085] Figure 16A Shows that variants in the baseline biopsy were at low frequency in the archived biopsy. Figure 16B Shows that biopsy variants during treatment were more representative of those present in the archived biopsy. Although all biopsies were from the primary site, only 12 / 135 of the targeted variants were common among the three, indicating tumor heterogeneity.
[0086] Figure 17A Shows the variant dynamics of cfDNA over time in patient G01. Figure 17B Shows targeted low-frequency variants of the indicated variants (SSH3, GRIA4, ZNF541, TMEM217, ZNF697, AHNAK2, SCHIP1, and CNR1) in ctDNA. Figure 17C Shows the targeted variants over time in the WES of the ctDNA of patient G01. Figure 17DShows brain met biopsy variants over time in the WES of ctDNA of patient G01.
[0087] Figure 18A - 18E Shows ctDNA monitoring of tumor variants in SLATE patients. Figure 18A Shows ctDNA %VAF of patient S2. Figure 18B Shows ctDNA %VAF of patient S5. Figure 18C Shows ctDNA %VAF of patient S10. Figure 18D Shows ctDNA %VAF of patient S13. Figure 18E Provides a representative patient without MR, showing loss of B2M start codon. SD = stable disease, PD = progressive disease, indicating the best overall response.
[0088] Figure 19A Shows fold change in HLA allele read fraction from molecular responders (MR). Figure 19B Shows fold change in HLA allele read fraction from non - molecular responders (non - MR).
[0089] Figure 20 Provides a figure outlining considerations including subject - specific tumor - informed probes.
[0090] Figure 21A Provides the percentage of CRC and NSCLC samples covered by target probes in the general group. Data is based on the analysis of 10,586 samples from cbioportal.org. Figure 21B Shows a retrospective analysis of variants identified by general group version 1 (v1) or version 2 (v2) in patients of a previous study (GO - 004).
[0091] Figure 22 Shows a general strategy for monitoring loss of heterozygosity of HLA genes on chromosome 6. Detailed Description
[0092] Definitions
[0093] Generally, the terms used in the claims and the specification are intended to be interpreted to have the ordinary meaning understood by a person of ordinary skill in the art. Certain terms are defined below to provide additional clarity. If the ordinary meaning conflicts with the provided definition, the provided definition shall be used. As used herein, the term "antigen" is a substance that induces an immune response. The antigen can be a neoantigen. The antigen can be a "shared antigen", i.e., an antigen found in a specific population (e.g., a specific group of cancer patients). As used herein, the term "neoantigen" is an antigen that has at least one alteration that makes it different from the corresponding wild-type antigen, e.g., via a mutation or a post-translational modification in a tumor cell. The neoantigen can include a polypeptide sequence or a nucleotide sequence. The mutation can include a frameshift or non-frameshift indel, a missense or nonsense substitution, a splice site alteration, a genomic rearrangement or gene fusion, or any genomic alteration or expression alteration that results in a neoORF. The mutation can also include a splice variant. The post-translational modification in a tumor cell can include abnormal phosphorylation. The post-translational modification in a tumor cell can also include a spliced antigen generated by a proteasome. See Liepe et al., A large fraction of HLA classI ligands are proteasome-generated spliced peptides; Science. October 21, 2016; 354(6310):354-358. Such shared neoantigens can be used to induce an immune response in a subject via administration. The subject for administration can be determined by using various diagnostic methods (e.g., the patient selection methods further described below).
[0094] As used herein, the term "tumor antigen" is an antigen that is present in the tumor cells or tissues of a subject but not in the corresponding normal cells or tissues of the subject, or an antigen derived from a polypeptide whose expression is altered in tumor cells or cancer tissues compared to normal cells or tissues, as known or discovered.
[0095] As used herein, the term "antigen-based vaccine" is a vaccine composition based on one or more antigens (e.g., multiple antigens). The vaccine can be a nucleotide-based (e.g., virus-based, RNA-based or DNA-based), protein-based (e.g., peptide-based) or a combination thereof vaccine.
[0096] As used herein, the term "candidate antigen" is a mutation or other aberration that generates a sequence that may represent an antigen.
[0097] As used herein, the term "coding region" is the part of a gene that encodes a protein.
[0098] As used herein, the term "coding mutation" is a mutation that occurs in the coding region.
[0099] As used herein, the term "ORF" means open reading frame.
[0100] As used herein, the term "NEO-ORF" is a tumor-specific ORF created by a mutation or other aberration (such as splicing).
[0101] As used herein, the term "missense mutation" is a mutation that results in the substitution of one amino acid for another.
[0102] As used herein, the term "nonsense mutation" is a mutation that results in the substitution of an amino acid for a stop codon or the removal of a canonical start codon.
[0103] As used herein, the term "frameshift mutation" is a mutation that results in a change in the protein frame.
[0104] As used herein, the term "indel" is an insertion or deletion of one or more nucleic acids.
[0105] As used herein, in the context of two or more nucleic acid or polypeptide sequences, the term "percent identity" refers to two or more sequences or subsequences that have a specified percentage of identical nucleotide or amino acid residues when compared and aligned for maximum correspondence, as measured using one of the sequence comparison algorithms described below (e.g., BLASTP and BLASTN or other algorithms available to those of skill in the art) or by visual inspection. Depending on the application, the "percent identity" may exist over regions of the sequences being compared, e.g., over functional domains, or alternatively may exist over the full length of the two sequences being compared.
[0106] For sequence comparison, typically one sequence acts as a reference sequence to which the test sequence is compared. When using a sequence comparison algorithm, the test and reference sequences are input into a computer, subsequence coordinates are specified (if necessary), and sequence algorithm program parameters are designated. The sequence comparison algorithm then calculates the percent sequence identity of the test sequence relative to the reference sequence based on the designated program parameters. Alternatively, sequence similarity or dissimilarity can be determined by the presence or absence of a particular combination of nucleotides, or for translated sequences, by the presence or absence of a combination of amino acids at selected sequence positions (e.g., sequence motifs).
[0107] The optimal alignment of sequences for comparison can be determined, for example, by the local homology algorithm of Smith and Waterman, Adv. Appl. Math. 2:482 (1981), by the homology alignment algorithm of Needleman & Wunsch, J. Mol. Biol. 48:443 (1970), by the similarity search method of Pearson & Lipman, Proc. Nat’l. Acad. Sci. USA 85:2444 (1988), by computer implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, Wis.), or by visual inspection (generally see Ausubel et al., infra) to achieve comparison.
[0108] An example of an algorithm suitable for determining percent sequence identity and percent sequence similarity is the BLAST algorithm described in Altschul et al., J. Mol. Biol. 215:403-410 (1990). Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information.
[0109] As used herein, the term “no stop or readthrough” is a mutation that results in the removal of a natural stop codon.
[0110] As used herein, the term “epitope” is a specific portion of an antigen that is typically bound by an antibody or a T cell receptor.
[0111] As used herein, the term “immunogenicity” is the ability to elicit an immune response, for example, by T cells, B cells, or both.
[0112] As used herein, the terms “HLA binding affinity” “MHC binding affinity” mean the binding affinity between a specific antigen and a specific MHC allele.
[0113] As used herein, the term “bait” is a nucleic acid probe used to enrich a specific sequence of DNA or RNA from a sample.
[0114] As used herein, the term “variant” is a difference between the nucleic acid of a subject and a reference human genome used as a control.
[0115] As used herein, the term “variant calling” is the algorithmic determination of the presence of a variant, typically by sequencing.
[0116] As used herein, the term "polymorphism" refers to a germline variant, i.e., a variant found in all cells of an individual that carry DNA.
[0117] As used herein, the term "somatic variant" refers to a variant that arises in non-germline cells of an individual.
[0118] As used herein, the term "allele" refers to a form of a gene, a form of a gene sequence, or a form of a protein.
[0119] As used herein, the term "HLA type" refers to the complementary sequence of HLA gene alleles.
[0120] As used herein, the term "nonsense-mediated decay" or "NMD" refers to the cellular degradation of mRNA caused by premature termination codons.
[0121] As used herein, the term "truncal mutation" refers to a mutation that originates early in tumor development and is present in a substantial fraction of tumor cells.
[0122] As used herein, the term "subclonal mutation" refers to a mutation that originates late in tumor development and is present only in a subset of tumor cells.
[0123] As used herein, the term "exome" refers to a subset of the genome that encodes proteins. The exome can be the collection of exons of the genome.
[0124] As used herein, the term "logistic regression" refers to a regression model for statistically derived binary data, where the logit of the probability that the dependent variable equals 1 is modeled as a linear function of the dependent variable.
[0125] As used herein, the term "neural network" refers to a machine learning model for classification or regression, consisting of multiple layers of linear transformations, followed by element-wise non-linearities typically trained by stochastic gradient descent and backpropagation.
[0126] As used herein, the term "proteome" refers to the collection of all proteins expressed and / or translated by a cell, cell population, or individual.
[0127] As used herein, the term "peptidome" refers to the collection of all peptides presented on the cell surface by MHC-I or MHC-II. The peptidome can refer to the property of a cell or cell population (e.g., tumor peptidome, meaning the union of the peptidomes of all cells that make up a tumor).
[0128] As used herein, the term "ELISpot" refers to an enzyme-linked immunosorbent spot assay, which is a commonly used method for monitoring the immune response in humans and animals.
[0129] As used herein, the term "tolerance or immunotolerance" is a state of immune non-responsiveness to one or more antigens (e.g., autoantigens).
[0130] As used herein, the term "central tolerance" is tolerance that is affected in the thymus, either by deletion of autoreactive T cell clones or by promoting the differentiation of autoreactive T cell clones into immunosuppressive regulatory T cells (Tregs).
[0131] As used herein, the term "peripheral tolerance" affects peripheral tolerance by downregulating or anergizing autoreactive T cells that survive central tolerance or by promoting the differentiation of these T cells into Tregs.
[0132] The term "sample" can include an aliquot of single cells or multiple cells or cell fragments or body fluid taken from a subject by including venipuncture, excretion, ejaculation, massage, biopsy, needle aspiration, lavage sample, scraping, surgical incision or intervention or other means known in the art.
[0133] The term "subject" encompasses cells, tissues or organisms, human or non-human, whether in vivo, ex vivo or in vitro, male or female. The term subject includes mammals, including humans.
[0134] The term "mammal" encompasses both human and non-human, and includes but is not limited to humans, non-human primates, canines, felines, murine, bovines, equines and porcines.
[0135] The term "clinical factor" refers to a measure of a subject's condition, such as disease activity or severity. "Clinical factor" encompasses all markers of a subject's health status, including non-sample markers, and / or other characteristics of the subject, such as but not limited to age and gender. A clinical factor can be a score, value or set of values that can be obtained by evaluating a sample (or population of samples) from a subject or the subject under defined conditions. A clinical factor can also be predicted by markers and / or other parameters such as gene expression surrogates. Clinical factors can include tumor type, tumor subtype and smoking history.
[0136] The term "alphavirus" refers to members of the Togaviridae family and is a positive-sense single-stranded RNA virus. Alphaviruses are generally divided into Old World viruses, such as Sindbis virus, Ross River virus, Mayaro virus, Chikungunya virus and Semliki Forest virus, or New World viruses, such as Eastern equine encephalitis virus, Oropouche virus, Fort Morgan virus or Venezuelan equine encephalitis virus and its derivative strain TC-83. Alphaviruses are generally self-replicating RNA viruses.
[0137] The term "alphavirus backbone" refers to the minimal sequence of an alphavirus that permits self-replication of the viral genome. The minimal sequence can include conserved sequences for nonstructural protein-mediated amplification, the nonstructural protein 1 (nsP1) gene, the nsP2 gene, the nsP3 gene, the nsP4 gene, and the polyA sequence, as well as a subgenomic viral RNA expression sequence that includes the 26S promoter element.
[0138] The term "sequence for nonstructural protein-mediated amplification" includes alphavirus conserved sequence elements (CSEs) well known to those skilled in the art. CSEs include, but are not limited to, the alphavirus 5’ UTR, the 51-nt CSE, the 24-nt CSE, or other 26S subgenomic promoter sequences, the 19-nt CSE, and the alphavirus 3’ UTR.
[0139] The term "RNA polymerase" includes polymerases that catalyze the production of RNA polynucleotides from a DNA template. RNA polymerases include, but are not limited to, phage-derived polymerases, including T3, T7, and SP6.
[0140] The term "lipid" includes hydrophobic and / or amphiphilic molecules. Lipids can be cationic, anionic, or neutral. Lipids can be synthetic or naturally derived and, in some cases, biodegradable. Lipids can include cholesterol, phospholipids, lipid conjugates (including, but not limited to, polyethylene glycol (PEG) conjugates (PEGylated lipids)), waxes, oils, glycerides, fats, and fat-soluble vitamins. Lipids can also include dilinoleyl methyl-4-dimethylaminobutyrate (MC3) and MC3-like molecules.
[0141] The term "lipid nanoparticle" or "LNP" includes vesicle-like structures formed by surrounding an aqueous interior with a lipid-containing membrane, also known as liposomes. Lipid nanoparticles include lipid-based compositions having a solid lipid core stabilized by surfactants. The core lipids can be fatty acids, acylglycerols, waxes, and mixtures of these surfactants. Biomembrane lipids such as phospholipids, sphingomyelins, bile salts (sodium taurocholate), and sterols (cholesterol) can be used as stabilizers. Lipid nanoparticles can be formed using specific ratios of different lipid molecules, including but not limited to specific ratios of one or more cationic, anionic, or neutral lipids. Lipid nanoparticles can encapsulate molecules within an outer membrane shell and can subsequently contact target cells to deliver the encapsulated molecules to the host cell cytosol. Lipid nanoparticles can be modified or functionalized with non-lipid molecules, including on their surface. Lipid nanoparticles can be single-layered / unilamellar or multi-layered / multilamellar. Lipid nanoparticles can be complexed with nucleic acids. Single-layer lipid nanoparticles can be complexed with nucleic acids, where the nucleic acids are located within the aqueous interior. Multi-layer lipid nanoparticles can be complexed with nucleic acids, where the nucleic acids are located within the aqueous interior, or formed or sandwiched in between.
[0142] The term "medically effective amount" is the amount of a vaccine component (e.g., a peptide, an engineered vector, and / or an adjuvant) that effectively provides a cell with a sufficient level of protein, protein expression, and / or cell signaling activity (e.g., adjuvant-mediated activation) in a route of administration to provide a vaccine benefit, i.e., some measurable level of immunity.
[0143] As used herein, terms such as "obtaining", "isolating", "enriching", "sequencing", "acquiring", "collecting", and "determining" refer to directly performing a process (e.g., directly performing a method) to obtain a result, such as directly obtaining a product, including but not limited to directly sequencing cfDNA to obtain cfDNA sequencing data, directly isolating cfDNA to obtain isolated cfDNA, directly enriching cfDNA to obtain an enriched cfDNA sample including cfDNA, etc. As used herein, terms such as "have obtained", "have isolated", "have enriched", "have sequenced", "have acquired", "have collected", and "have determined" refer to indirectly receiving information or receiving a product without directly performing a process (e.g., not directly performing a method), such as by receiving knowledge or a product from another party or source (e.g., a third-party laboratory that directly obtains cfDNA sequencing data, isolates cfDNA, enriches cfDNA, and / or collects a sample including cfDNA, etc.). In some cases, the other party or source is directed to directly perform the process (e.g., the third-party laboratory is directed to obtain cfDNA sequencing data, isolate cfDNA, enrich cfDNA, and / or collect a sample including cfDNA, etc.). In some cases, the knowledge or product is purchased from another party or source that directly performs the process (e.g., purchasing cfDNA sequencing data, isolated cfDNA, enriched cfDNA, and / or a collected sample including cfDNA, etc.).
[0144] Abbreviations: MHC: major histocompatibility complex; HLA: human leukocyte antigen or human MHC locus; NGS: next-generation sequencing; PPV: positive predictive value; TSNA: tumor-specific neoantigen; FFPE: formalin-fixed, paraffin-embedded; NMD: nonsense-mediated decay; NSCLC: non-small cell lung cancer; DC: dendritic cell.
[0145] It should be noted that, unless the context clearly dictates otherwise, as used in the specification and the appended claims, the singular forms "a", "an", and "the" include plural referents.
[0146] Unless explicitly stated or otherwise apparent from the context, as used herein, the term "about" is to be understood as within the normal tolerance in the art, e.g., within 2 standard deviations of the mean. About can be understood as within 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5%, 0.1%, 0.05%, or 0.01% of the stated value. Unless the context clearly dictates otherwise, all numerical values provided herein are modified by the term about.
[0147] Any terms not directly defined herein should be understood to have the meanings customarily associated with them within the art of the present invention. Certain terms are discussed herein to provide additional guidance to the practitioner in describing the compositions, devices, methods, etc. of the present invention and how to make or use them. It should be understood that the same thing can be described in more than one way. Thus, alternative languages and synonyms may be used for any one or more of the terms discussed herein. It is not important whether a term is elaborated or discussed herein. Some synonyms or alternative methods, materials, etc. are provided. The elaboration of one or several synonyms or equivalents does not exclude the use of other synonyms or equivalents unless expressly stated. The use of examples, including examples of terms, is for illustrative purposes only and does not limit the scope and meaning of the aspects of the present invention herein.
[0148] For all purposes, all references, published patents, and patent applications cited in the body of the specification are hereby incorporated by reference in their entirety herein.
[0149] Monitoring disease status and therapeutic efficacy
[0150] Methods for monitoring the disease status of a subject are provided herein, which are performed by analyzing cell-free DNA (cfDNA), particularly by monitoring mutation frequencies (e.g., tumor-associated mutations associated with cancer). For example, cfDNA can be used to monitor the disease progression of patients receiving therapy. The cfDNA analysis methods described herein provide a non-invasive way to assess and / or monitor diseases, particularly as compared to more invasive procedures such as tumor biopsies. The cfDNA analysis methods described herein are particularly useful for analyzing a large number of mutations, such as analyzing all or most of the tumor exome. Generally, monitoring is performed by sequencing cfDNA that has broad target coverage (e.g., at least 50% of all polynucleotide regions of interest corresponding to mutations present in the subject's cancer exome) and has a high sequencing read depth ("deep sequencing", e.g., at least 1000-fold average read depth).
[0151] In one aspect, a method for monitoring the cancer status of a subject comprises the steps of: a. obtaining or having obtained sequencing data of cfDNA from a sample of the subject, and wherein the sequencing data includes a target coverage of at least 50% of all polynucleotide regions of interest corresponding to mutations present in the cancer exome, and wherein the polynucleotide regions of interest sequenced have an average read depth of at least 1000-fold; and b. determining or having determined the mutation frequency present in the exome to evaluate the cancer status.
[0152] More than one sample can be analyzed to assess the disease state of a subject. Thus, in one aspect, a method for monitoring the cancer state of a subject comprises the steps of: a. obtaining or having obtained sequencing data of cfDNA from a first sample from the subject, and wherein the sequencing data comprises a target coverage of at least 50% of all polynucleotide regions of interest corresponding to mutations present in the cancer exome, and wherein the polynucleotide regions of interest sequenced have an average read depth of at least 1000-fold; b. obtaining or having obtained sequencing data of cfDNA from a second sample from the subject, wherein the sequencing data comprises a target coverage of at least 50% of all polynucleotide regions of interest corresponding to mutations present in the cancer exome, and wherein the polynucleotide regions of interest sequenced have an average read depth of at least 1000-fold; and c. determining or having determined the mutation frequency in the exome of the first cfDNA relative to the second cfDNA to assess the cancer state.
[0153] Multiple samples containing cfDNA can be collected from a subject at different time points and used to monitor the disease, such as monitoring the disease burden and / or response to therapy during the course of treatment. The time points can be selected at specific intervals to monitor the disease state. For example, the time points can be selected based on the therapy dosing schedule. The time points based on the dosing schedule can include the same day as the therapy administration. The time points based on the dosing schedule can include, but are not limited to, one, two, three, four, five, six days after dosing. The time points based on the dosing schedule can include, but are not limited to, one, two, three, four, five, six, eight, ten, twelve weeks after dosing. The time points based on the dosing schedule can include, but are not limited to, one, two, three, six, twelve months after dosing.
[0154] The time points can be regular time intervals, such as regular time intervals during the course of therapy, including but not limited to daily, every two days, every three days, every four days, every five days, every six days. The time points based on regular time intervals can include, but are not limited to, once a week, once every two weeks, once every three weeks, once every four weeks, once every five weeks, once every six weeks, once every eight weeks, once every ten weeks, once every twelve weeks. The time points can also be selected based on regular time intervals, including but not limited to once a month, once every two months, once every three months, once every six months, and once every twelve months. Combinations of one or more of the above time intervals can also be used.
[0155] Analysis of cfDNA can be used to monitor disease progression in patients undergoing therapy. For example, longitudinal samples can be collected during therapy to monitor cancer status (e.g., tumor burden over time). An increase in the mutation frequency monitored in longitudinal samples can indicate an increased likelihood of an increase in the tumor burden of the subject. A decrease or maintenance of the mutation frequency of the mutations monitored in longitudinal samples can indicate an increased likelihood of a decrease or stability of the tumor burden of the subject.
[0156] Analysis of cfDNA can be used to evaluate the efficacy of a therapy administered to a subject. Thus, in one aspect, a method for evaluating the efficacy of a therapy in a subject with cancer comprises the steps of: a. obtaining or having obtained sequencing data of cfDNA from a pre-therapy sample of the subject, and wherein the sequencing data comprises target coverage of at least 50% of all polynucleotide regions of interest corresponding to mutations present in the cancer exome, and wherein the polynucleotide regions of interest sequenced have an average read depth of at least 1000-fold; b. obtaining or having obtained sequencing data of cfDNA from a post-therapy sample from the subject, wherein the sequencing data comprises target coverage of at least 50% of all polynucleotide regions of interest corresponding to mutations present in the cancer exome, and wherein the polynucleotide regions of interest sequenced have an average read depth of at least 1000-fold; and c. determining or having determined the mutation frequency present in the exome of the pre-therapy cfDNA relative to the post-therapy cfDNA to evaluate the efficacy of the therapy.
[0157] Multiple samples with cfDNA can be collected at different time points relative to the administration of the therapy. A sample with cfDNA can be collected before the administration of the therapy. A sample with cfDNA can be collected after the administration of the therapy. A sample with cfDNA can be collected while the therapy is being administered. Samples with cfDNA can be collected before and after the administration of the therapy. For example, a first sample with cfDNA can be collected before administering the therapy to the subject, and a second sample with cfDNA can be collected after the administration of the therapy. Samples with cfDNA can be collected while and after the administration of the therapy. For example, a first sample with cfDNA can be collected while administering the therapy to the subject, and a second sample with cfDNA can be collected after the administration of the therapy. After the administration of the therapy, multiple samples with cfDNA (e.g., longitudinal samples) can be collected.
[0158] Obtaining sequencing data can include one or more of the following steps: collecting or having collected a sample from a subject; isolating or having isolated cfDNA; enriching or having enriched cfDNA, and / or sequencing or having sequenced cfDNA. Obtaining sequencing data can include each of the following steps: collecting or having collected a sample from a subject; isolating or having isolated cfDNA; enriching or having enriched cfDNA, and / or sequencing or having sequenced cfDNA. Intermediates can be obtained for performing any of the above steps. For example, isolated cfDNA can be obtained from a third-party source and used to perform one or more of the remaining steps, such as enrichment and sequencing. Intermediates can be generated and directed to a third party to perform any of the above steps. For example, enriched cfDNA can be generated and provided to a third-party source for performing one or more of the remaining steps, such as sequencing.
[0159] Cancer monitoring
[0160] The methods described herein can be used to monitor cancer status, such as tumor burden.
[0161] The diseases of the subject can include cancer. Cancer cells can release their genomic DNA into the circulation upon cell death, which is called circulating tumor DNA (ctDNA) or cfDNA from cancer cells. Multiple cancers can be monitored. For example, the cancers that can be monitored include, but are not limited to, carcinoma, sarcoma, lymphoma or leukemia, germ cell tumor, blastoma or other cancers. Carcinomas include, but are not limited to, epithelial tumors, squamous cell tumors squamous cell carcinoma, basal cell tumors basal cell carcinoma, transitional cell papillomas and carcinomas, adenomas and adenocarcinomas (glandular), adenomas, adenocarcinomas, linitis plastica insulinoma, glucagonoma, gastrinoma, vasoactive intestinal peptide tumor (vipoma), cholangiocarcinoma, hepatocellular carcinoma, adenoid cystic carcinoma, appendiceal carcinoid tumor, prolactinoma, oxyphil cell adenoma, Hurthle cell adenoma, renal cell carcinoma, Grawitz tumor, multiple endocrine adenoma, endometrioid adenoma, adnexal and skin appendage tumors, mucoepidermoid tumors, cystic, mucinous and serous tumors, cystadenomas, pseudomyxoma peritonei, ductal, lobular and medullary tumors, acinar cell tumors, complex epithelial tumors, Warthin’s tumor, thymoma, specialized glandular tumors, sex cord-stromal tumors, theca cell tumors, granulosa cell tumors, androblastomas, Sertoli-Leydig cell tumor, glomus tumor, paraganglioma, pheochromocytoma, glomus tumor, nevus and melanoma, melanocytic nevus, malignant melanoma, melanoma, nodular melanoma, dysplastic nevus, lentigo maligna melanoma, superficial spreading melanoma and acral lentiginous melanoma. Sarcomas include, but are not limited to, Askin tumor, botryodies, chondrosarcoma, Ewing’s sarcoma, malignant hemangioendothelioma, malignant schwannoma, osteosarcoma, soft tissue sarcomas, including: alveolar soft part sarcoma, angiosarcoma, cystosarcoma phyllodes, dermatofibrosarcoma, desmoid tumor, desmoplastic small round cell tumor, epithelioid sarcoma, extraskeletal chondrosarcoma, extraskeletal osteosarcoma, fibrosarcoma, hemangiopericytoma, angiosarcoma, kaposi’s sarcoma, leiomyosarcoma, liposarcoma, lymphangiosarcoma, lymphosarcoma, malignant fibrous histiocytoma, neurofibrosarcoma, rhabdomyosarcoma and synovial sarcoma.Lymphomas and leukemias include, but are not limited to, chronic lymphocytic leukemia / small lymphocytic lymphoma, B-cell prolymphocytic leukemia, lymphoplasmacytic lymphoma (such as Waldenström macroglobulinemia), splenic marginal zone lymphoma, plasma cell myeloma, plasmacytoma, monoclonal immunoglobulin deposition disease, heavy chain disease, extranodal marginal zone B-cell lymphoma (also known as MALT lymphoma), nodal marginal zone B-cell lymphoma, follicular lymphoma, mantle cell lymphoma, diffuse large B-cell lymphoma, mediastinal (thymic) large B-cell lymphoma, intravascular large B-cell lymphoma, primary effusion lymphoma, Burkitt lymphoma / leukemia, T-cell prolymphocytic leukemia, T-cell large granular lymphocytic leukemia, aggressive NK-cell leukemia, adult T-cell leukemia / lymphoma, extranodal NK / T-cell lymphoma, nasal type, enteropathy-type T-cell lymphoma, hepatosplenic T-cell lymphoma, blastic NK-cell lymphoma, mycosis fungoides / Sezary syndrome, primary cutaneous CD30-positive T-cell lymphoproliferative disorders, primary cutaneous anaplastic large cell lymphoma, lymphomatoid papulosis, angioimmunoblastic T-cell lymphoma, peripheral T-cell lymphoma, anaplastic large cell lymphoma, not otherwise specified, classical Hodgkin lymphoma (nodular sclerosis, mixed cellularity, lymphocyte-rich, lymphocyte-depleted or non-depleted), and nodular lymphocyte-predominant Hodgkin lymphoma. Germ cell tumors include, but are not limited to, germinoma, dysgerminoma, seminoma, non-germinomatous germ cell tumor, embryonal carcinoma, endodermal sinus tumor, choriocarcinoma, teratoma, polyembryoma, and gonadoblastoma. Blastomas include, but are not limited to, nephroblastoma, medulloblastoma, and retinoblastoma. Other cancers include, but are not limited to, lip cancer, laryngeal cancer, hypopharyngeal cancer, tongue cancer, salivary gland cancer, gastric cancer, adenocarcinoma, thyroid cancer (medullary thyroid cancer and papillary thyroid cancer), kidney cancer, renal parenchymal cancer, cervical cancer, corpus cancer, endometrial cancer, choriocarcinoma, testicular cancer, urinary cancer, melanoma, brain tumors such as glioblastoma, astrocytoma, meningioma, medulloblastoma, and peripheral neuroectodermal tumor, gallbladder cancer, bronchial cancer, multiple myeloma, basal cell carcinoma, teratoma, retinoblastoma, choroidal melanoma, seminoma, rhabdomyosarcoma, craniopharyngioma, osteosarcoma, chondrosarcoma, myosarcoma, liposarcoma, fibrosarcoma, Ewing sarcoma, and plasmacytoma. Cancers that can be monitored include, but are not limited to, lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer.
[0162] Cancer monitoring can also include monitoring cancer escape mutations (also referred to as secondary mutations or escape mutants). For example, cancer monitoring can include monitoring de novo mutations relative to an earlier sequencing dataset, such as an initial biopsy, longitudinal sample, pre-therapy sample, or any other archived sample. Monitoring of cancer escape mutations can inform whether additional therapy is needed (e.g., a therapy effective against a cancer with a specific de novo mutation) and / or whether the efficacy of current or proposed therapy will be affected. Cancer escape mutations include those genes targeted by the tumor agnostic probe sets described herein, such as genes and mutations generally considered to be oncogenic (e.g., "driver" mutations considered to contribute to cancer and commonly recognized gain-of-function mutations), tumor suppressor genes (e.g., genes generally considered to monitor and / or control tumor-related properties such as cell division, where mutations can interfere with the control of such properties and are generally considered loss-of-function mutations), interferon-gamma signaling pathway genes (including JAK / STAT signaling pathway genes), antigen processing pathway genes (including monitoring loss of HLA heterozygosity), and additional mutations generally associated with cancer but that may not be annotated (e.g., mutations not yet annotated as oncogenes or tumor suppressor genes).
[0163] The tumor agnostic / agnostic combination sets described herein can be used to simultaneously monitor cancer status, such as tumor burden, and cancer escape mutations using a single set.
[0164] Tumor-specific mutations
[0165] The methods described herein are applicable to tracking the presence of tumor-specific mutations associated with cancer cells present in cfDNA ("ctDNA"). Tumor-specific mutations can include previously identified tumor-specific mutations, such as those found in the Catalogue of Somatic Mutations in Cancer (COSMIC) database.
[0166] Methods for identifying certain mutations (e.g., variants or alleles present in cancer cells) are also disclosed herein. Specifically, these mutations can be present in the genome, transcriptome, proteome, or exome of cancer cells of a subject having cancer, but not in the normal tissue of the subject. Specific methods for identifying neoantigens (including shared neoantigens) that are specific to a tumor are known to those of skill in the art, such as the methods described in more detail in International Patent Application Publications WO / 2017 / 106638, WO / 2018 / 195357, and WO / 2018 / 208856, each of which is incorporated herein by reference in its entirety for all purposes.
[0167] If genetic mutations in a tumor result in changes in the amino acid sequence of a protein that is unique to the tumor, they can be considered useful for immunotargeting the tumor and / or monitoring tumor burden (e.g., disease state). Useful mutations include: (1) nonsynonymous mutations, which result in a different amino acid in the protein; (2) readthrough mutations, in which a stop codon is modified or deleted, resulting in the translation of a longer protein with a novel tumor-specific sequence at the C-terminus; (3) splice site mutations, which result in the inclusion of an intron in the mature mRNA and thus the production of a unique tumor-specific protein sequence; (4) chromosomal rearrangements, which result in the production of a chimeric protein with a tumor-specific sequence at the junction of two proteins (i.e., gene fusion); (5) frameshift mutations or deletions, which result in the production of a new open reading frame with a novel tumor-specific protein sequence. Mutations can also include non-frameshift indels, missense or nonsense substitutions, splice site alterations, genomic rearrangements or gene fusions, or any one or more of genomic alterations or expression alterations that result in a neoORF.
[0168] By sequencing DNA, RNA, or protein in tumor cells and normal cells, mutant peptides or mutant polypeptides resulting from, for example, splice site, frameshift, readthrough, or gene fusion mutations in tumor cells can be identified.
[0169] A variety of methods can be used to detect the presence of specific mutations or alleles in an individual's DNA or RNA. Any of the sequencing methods described herein can be used to determine tumor-specific mutations. Advances in the field have provided accurate, simple, and inexpensive large-scale SNP genotyping. For example, several techniques have been described, including dynamic allele-specific hybridization (DASH), microplate array diagonal gel electrophoresis (MADGE), pyrosequencing, oligonucleotide-specific ligation, the TaqMan system, and various DNA "chip" technologies such as the Affymetrix SNP chip. These methods utilize the amplification of the target gene region, typically by PCR. There are also some other methods based on the generation of small signaling molecules by invasive cleavage, followed by mass spectrometry or immobilized padlock probes and rolling circle amplification. Several methods known in the art for detecting specific mutations are summarized below.
[0170] PCR-based detection modalities can include multiplex amplification of multiple markers simultaneously. For example, it is well known in the art to select PCR primers to generate PCR products that do not overlap in size and can be analyzed simultaneously. Alternatively, different markers can be amplified with differentially labeled primers and thus differentially detected separately. Of course, hybridization-based detection modalities allow for the differential detection of multiple PCR products in a sample. Other techniques known in the art allow for the multiplex analysis of multiple markers.
[0171] Several methods have been developed to facilitate the analysis of single nucleotide polymorphisms in genomic DNA or cellular RNA. For example, single base polymorphisms can be detected by using specialized exonuclease-resistant nucleotides, as disclosed, for example, in Mundy, C.R. (U.S. Patent No. 4,656,127). According to the method, a primer complementary to the allelic sequence at the 3' of the polymorphic site is allowed to hybridize to a target molecule obtained from a specific animal or human. If the polymorphic site on the target molecule contains a nucleotide complementary to the specific exonuclease-resistant nucleotide derivative present, then the derivative will be incorporated into the end of the hybridized primer. This incorporation renders the primer resistant to exonuclease and thus allows its detection. Since the identity of the exonuclease-resistant derivative of the sample is known, the finding that the primer has become resistant to exonuclease reveals that the nucleotide present in the polymorphic site of the target molecule is complementary to the nucleotide of the nucleotide derivative used in the reaction. The advantage of this method is that it does not require the determination of large amounts of irrelevant sequence data.
[0172] Solution-based methods can be used to determine the identity of the nucleotide at the polymorphic site. Cohen, D. et al. (French Patent 2,650,840; PCT Application No. WO91 / 02087). As in the Mundy method of U.S. Patent No. 4,656,127, a primer complementary to the allelic sequence at the 3' of the polymorphic site is employed. The method uses labeled dideoxynucleotide derivatives to determine the identity of the nucleotide at the site, which will be incorporated into the primer end if it is complementary to the nucleotide at the polymorphic site.
[0173] Goelet, P. et al. (PCT Application No. 92 / 15712) describe an alternative method called genetic bit analysis or GBA. The method of Goelet, P. et al. uses a mixture of labeled terminators and a primer complementary to the 3' sequence of the polymorphic site. Thus, the incorporated labeled terminator is determined by and complementary to the nucleotide present in the polymorphic site of the target molecule being evaluated. Compared to the method of Cohen et al. (French Patent 2,650,840; PCT Application No. WO91 / 02087), the method of Goelet, P. et al. can be a heterogeneous analysis, in which the primer or the target molecule is immobilized on a solid phase.
[0174] Several primer-directed nucleotide incorporation procedures for determining polymorphic sites in DNA have been described (Komher, J.S. et al., Nucl. Acids Res. 17:7779-7784 (1989); Sokolov, B.P., Nucl. Acids Res. 18:3671 (1990); Syvanen, A.-C., et al., Genomics 8:684-692 (1990); Kuppuswamy, M.N. et al., Proc. Natl. Acad. Sci. (U.S.A.) 88:1143-1147 (1991); Prezant, T.R. et al., Hum. Mutat. 1:159-164 (1992); Ugozzoli, L. et al., GATA 9:107-112 (1992); Nyren, P. et al., Anal. Biochem. 208:171-175 (1993)). These methods differ from GBA in that they utilize the incorporation of labeled deoxynucleotides to discriminate between bases at polymorphic sites. In this format, since the signal is proportional to the number of incorporated deoxynucleotides, polymorphisms occurring during runs of the same nucleotide can produce signals proportional to the run length (Syvanen, A.-C., et al., Amer. J. Hum. Genet. 52:46-59 (1993)).
[0175] Many initiatives directly obtain sequence information from millions of parallel individual DNA or RNA molecules. Real-time single molecule synthesis sequencing techniques rely on the detection of fluorescent nucleotides as they are incorporated into nascent DNA strands complementary to the template being sequenced. In one method, oligonucleotides 30-50 bases in length are covalently anchored to a glass coverslip at the 5' end. These anchored strands serve two functions. First, they act as capture sites for target template strands if the template is configured with a capture tail complementary to the surface-bound oligonucleotide. They also serve as primers for template-directed primer extension, which forms the basis of sequence reads. The capture primer serves as a fixed position site for sequence determination, using multiple cycles of synthesis, detection, and chemical cleavage of the dye linker to remove the dye for determination. Each cycle adds a polymerase / labeled nucleotide mixture, rinsing, imaging, and dye cleavage. In an alternative method, the polymerase is modified with a fluorescent donor molecule and immobilized on a glass slide, while each nucleotide is color-coded with an acceptor fluorescent moiety attached to the γ-phosphate. When the nucleotide is incorporated into the nascent strand, the system detects the interaction between the fluorescently labeled polymerase and the fluorescently modified nucleotide. There are also other synthetic sequencing techniques.
[0176] Any suitable synthetic sequencing platform can be used to identify mutations. As described above, the four main synthetic sequencing platforms currently available are: the Genome Sequencer from Roche / 454 Life Sciences, the 1G Analyzer from Illumina / Solexa, the SOLiD System from Applied BioSystems, and the Heliscope System from Helicos Biosciences. PacificBioSciences and VisiGen Biotechnologies have also described synthetic sequencing platforms. In some embodiments, the plurality of nucleic acid molecules to be sequenced are bound to a support (e.g., a solid support). To immobilize the nucleic acid on the support, capture sequences / universal priming sites can be added at the 3' and / or 5' ends of the template. The nucleic acid can be bound to the support by hybridizing the capture sequence to a complementary sequence covalently attached to the support. The capture sequence (also referred to as a universal capture sequence) is a nucleic acid sequence complementary to the sequence attached to the support, which can also serve as a universal primer.
[0177] As an alternative to the capture sequence, members of a coupling pair (such as, for example, an antibody / antigen, a receptor / ligand, or an avidin-biotin pair as described in, for example, U.S. Patent Application No. 2006 / 0252077) can be linked to each fragment to be captured, which is located on a surface coated with the respective second member of the coupling pair.
[0178] After capture, the sequence can be analyzed, for example, by single molecule detection / sequencing, for example, as described in U.S. Patent No. 7,283,337, including template-dependent synthetic sequencing. In synthetic sequencing, surface-bound molecules are exposed to a plurality of labeled nucleotide triphosphates in the presence of a polymerase. The sequence of the template is determined by the order of the labeled nucleotides incorporated into the 3' end of the growing strand. This can be done in real time or in a stepwise repeating mode. For real-time analysis, different optical labels can be incorporated into each nucleotide, and multiple lasers can be used to stimulate the incorporated nucleotides.
[0179] Sequencing can also include other massively parallel sequencing or next-generation sequencing (NGS) technologies and platforms. Additional examples of massively parallel sequencing technologies and platforms are Illumina HiSeq or MiSeq, Thermo PGM or Proton, Pac BioRS II or Sequel, Qiagen's Gene Reader, and Oxford Nanopore MinION. Additional similar currently available massively parallel sequencing technologies, as well as future generations of these technologies, can be used.
[0180] Any cell type or tissue can be utilized to isolate nucleic acid samples for methods of identifying tumor-specific mutations described herein. For example, DNA or RNA samples can be isolated from tumors or body fluids (e.g., blood collected by known techniques such as venipuncture) or saliva. Alternatively, nucleic acid testing can be performed on dried samples (e.g., hair or skin). Additionally, samples can be collected from tumors for sequencing, and another sample can be collected from normal tissue for sequencing, where the normal tissue is of the same tissue type as the tumor. Samples can be collected from tumors for sequencing, and another sample can be obtained from normal tissue for sequencing, where the normal tissue is of a different tissue type relative to the tumor. Tumors from which tumor-specific mutations can be identified include, but are not limited to, any of the tumors described herein, such as lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, and T cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer. Alternatively, protein mass spectrometry can be used to identify or verify the presence of mutant peptides that bind to MHC proteins on tumor cells. Peptides can be eluted by acid washing from tumor cells or from HLA molecules immunoprecipitated from tumors, and then identified using mass spectrometry.
[0181] Processing of cfDNA
[0182] Methods for processing cfDNA (e.g., isolation and purification of cfDNA) are generally known to those skilled in the art. For example, general methods for isolating cfDNA are described in US-2020 / 0277667-A1, which is incorporated herein by reference for all purposes. See also, e.g., Current Protocols in Molecular Biology, latest version.Exemplary methods for isolating cfDNA are also described in US-10,385,369-B2 and US-2020 / 0277667-A1,Cell-Free Plasma DNA as a Predictor of Outcome in Severe Sepsis and Septic Shock.Clin.Chem.2008,Vol. 54, p. 1000-Diagnostics.Clin.Chem 1007; Prediction of MYCN Amplification in Neuroblastoma Using Serum DNA and Real-Time Quantitative Polymerase Chain Reaction.JCO 2005,Vol. 23, pp. 5205-5210; Circulating Nucleic Acids in Blood of Healthy Male and Female Donors.Clin.Chem.2005,Vol. 51, pp. 1317-1319; Use of Magnetic Beads for Plasma Cell-free DNA Extraction: Toward Automation of Plasma DNA Analysis for Molecular.2003,Vol. 49, pp. 1953-1955; Chiu R W K,Poon L M,Lau T K,Leung T N,Wong E M C,Lo Y M D.effects of blood-processing protocols on fetal and total DNA quantification in maternal plasma.Clin Chem 2001;47:1607-1613; and Swinkels et al.Effects of Blood-Processing Protocols on Cell-free DNA Quantification in Plasma.Clinical Chemistry,2003,Vol. 49, No. 3, 525-526, each of which is incorporated herein by reference for all purposes.
[0183] Commercially available kits for the isolation and purification of cfDNA are known to those skilled in the art, including but not limited to the QIAamp Circulating Nucleic Acid Kit and the Apostle MiniMax cfDNA Isolation Kit (Beckman Coulter; Indianapolis, IN).
[0184] Blood / plasma samples can be collected from a subject, and cfDNA can be isolated from the blood / plasma samples. Samples containing cfDNA other than blood (e.g., feces, mucus) can be collected for cfDNA isolation and purification. Isolation of cfDNA can be achieved, for example, by centrifuging to separate cfDNA from cells or cell debris, or by separating the plasma layer that may contain cfDNA from the buffy coat and red blood cells in whole blood. Whole blood can be collected in a cell-free DNA BCT tube and centrifuged at an appropriate speed to separate the plasma layer, buffy coat, and red blood cells. The plasma layer can then be removed and spun again to remove any residual cellular material. The supernatant can then be collected and stored at -80°C until extraction. As an exemplary non-limiting example, whole blood can be collected in a 10 mL Streck cell-free DNA BCT tube (Streck; La Vista, NE, USA), spun at 1600 x g for 10 minutes at ambient temperature to separate the plasma layer, buffy coat, and red blood cells. The plasma layer can then be removed and spun again at 5000 x g for 10 minutes to remove any residual cellular material. The supernatant can then be collected and stored at -80°C until extraction. Those of ordinary skill in the art will recognize that the above non-limiting exemplary protocol can be optimized based on specific experimental conditions.
[0185] To prepare a cfDNA library for sequencing, cfDNA is typically fragmented, e.g., by shearing or enzymatic preparation (e.g., using the NEBNext Ultra II FS DNA Module; NEB, Ipswich, MA) to generate a library of polynucleotide regions of interest. The isolated nucleic acid (e.g., isolated cfDNA) can be fragmented or sheared by practicing conventional techniques. For example, DNA can be fragmented by physical shearing methods, enzymatic cleavage methods, chemical cleavage methods, and other methods well known to those skilled in the art. One of ordinary skill in the art will recognize that the above non-limiting illustrative protocols can be optimized to generate a library of desired fragment lengths, e.g., optimized for whole exome sequencing, depending on the desired sequencing application. For example, the time of enzymatic digestion can be optimized (e.g., as an illustrative example, using the NEBNext Ultra II FS DNA module for 25 minutes). The fragment length can be at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1000 bp. The fragment length can be 100 - 250, 150 - 350, 200 - 450, 300 - 700, or 500 - 1000 bp. The fragment length can average at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1000 bp. The fragment length can average 100 - 250, 150 - 350, 200 - 450, 300 - 700, or 500 - 1000 bp.
[0186] cfDNA Enrichment
[0187] cfDNA can be enriched to improve the detection and measurement of polynucleotide regions of interest. Typically, enrichment is performed on a fragmented cfDNA library (e.g., a library of polynucleotide regions of interest). The region of interest can contain a polynucleotide known or suspected of encoding one or more mutations. The region of interest can also contain gene translocations (e.g., Bcr-Abl fusion). The region of interest can contain a polynucleotide encoding a gene coding region or a fragment of a gene coding region, which can include tumor exome polynucleotides, such as tumor exome polynucleotides known or suspected of having subject- and / or tumor-specific mutations. Enrichment of the polynucleotide region of interest can generally improve the targeted measurement (e.g., increase sensitivity) of the DNA region of interest by subtracting noise from the sequencing results. The term "enrich / enrichment" refers to the partial purification of an analyte having a certain characteristic (e.g., a nucleic acid known or suspected of having a tumor-specific mutation) from an analyte that does not have that characteristic (e.g., a nucleic acid that does not contain a tumor-specific mutation). Relative to the analyte that does not have the characteristic, enrichment generally increases the concentration of the analyte having the characteristic (e.g., a nucleic acid containing a tumor-specific mutation) by at least 2-fold, at least 5-fold, or at least 10-fold. After enrichment, at least 10%, at least 20%, at least 50%, at least 80%, or at least 90% of the analytes in the sample can have the characteristic for which enrichment was performed. For example, at least 10%, at least 20%, at least 50%, at least 80%, or at least 90% of the nucleic acid molecules in the enriched composition can contain a strand having one or more tumor-specific mutations, which has been modified to contain a capture tag.
[0188] Enrichment of cfDNA can include hybridizing one or more polynucleotide probes (also referred to herein as "baits") to one or more polynucleotide regions of interest. The bait sequences can be based on tumor-specific mutations derived from genomic sequencing (e.g., sequencing of the tumor exome of a biopsy). The bait can contain a single polynucleotide sequence or a library of polynucleotide sequences derived from tumor sequencing. The bait sequences derived from tumor sequencing can be subject-specific. For example, a biopsy of a subject's tumor can be performed and sequenced to determine mutations associated with the subject's tumor, and subsequently subject-specific baits can be designed using the subject- and tumor-specific sequences to enrich regions of interest in the tumor exome, including baits capable of enriching all regions of interest having patient-specific tumor variants.
[0189] The bait can include a group (also referred to as a "combination group") containing a combination of tumor-informed polynucleotide probes and tumor-uninformed polynucleotide probes.
[0190] Tumor-informed polynucleotide probes include probes configured to capture target sequences (e.g., by hybridization and other modifications such as biotinylation as described elsewhere herein). The target sequences can include epitope sequences encoded by a cancer vaccine administered to a subject, where the subject has been determined to have a tumor expressing the epitope sequence. For example, when a cancer vaccine is administered to a subject, such a probe for an epitope sequence can be considered tumor-informed when: (a) the vaccine is an individualized vaccine such that prior cancer / tumor sequencing informed the selection of the epitopes included in the cancer vaccine itself, or (b) the vaccine is an "off-the-shelf" vaccine where the vaccine contains commonly occurring epitopes, but prior cancer / tumor sequencing is required to determine if the subject meets the eligibility requirements for receiving the vaccine.
[0191] Exemplary epitope sequences that can be encoded by a cancer vaccine include epitopes having mutations including, but not limited to, KRAS, G13D, KRAS_Q61K, TP53_R249M, CTNNB1_S45P, CTNNB1_S45F, ERBB2_Y772_A775dup, KRAS_G12D, KRAS_Q61R, CTNNB1_T41A, TP53_K132N, KRAS_G12A, KRAS_Q61L, TP53_R213L, BRAF_G466V, KRAS_G12V, KRAS_Q61H, CTNNB1_S37F, TP53_S127Y, TP53_K132E, and KRAS_G12C.
[0192] Exemplary epitope sequences that can be encoded by a cancer vaccine include KRAS mutations such as KRAS_G12C mutation, KRAS_G12D mutation, KRAS_G12V mutation, and KRAS_Q61H mutation. Exemplary epitope sequences that can be encoded by a cancer vaccine include KRAS_G12C mutation. Exemplary epitope sequences that can be encoded by a cancer vaccine include KRAS_G12D mutation. Exemplary epitope sequences that can be encoded by a cancer vaccine include KRAS_G12V mutation. Exemplary epitope sequences that can be encoded by a cancer vaccine include KRAS_Q61H mutation.
[0193] Exemplary epitope sequences that can be encoded by a cancer vaccine include EGFR mutations such as EGFR_L858R mutation. Exemplary epitope sequences that can be encoded by a cancer vaccine include EGFR_L858R mutation.
[0194] The epitope sequence can comprise one or more subject-specific epitopes. The epitope sequence can comprise one or more subject-specific epitopes, wherein the subject's tumor has been sequenced to determine the subject-specific epitopes to be encoded by the cancer vaccine. The subject-specific epitopes can comprise at least 2 subject-specific epitopes, at least 10 subject-specific epitopes, at least 20 subject-specific epitopes, or from 2 to 20 subject-specific epitopes. The subject-specific epitopes can comprise at least 2 subject-specific epitopes. The subject-specific epitopes can comprise at least 10 subject-specific epitopes. The subject-specific epitopes can comprise at least 20 subject-specific epitopes. The subject-specific epitopes can comprise from 2 to 20 subject-specific epitopes.
[0195] The group can also comprise additional tumor-informed polynucleotide probes that capture additional target sequences not encoded by the cancer vaccine, such as additional target sequences determined by cancer / tumor sequencing. For example, the group can comprise additional target sequences that have been predicted to be presented by the subject's HLA alleles but not selected to be included in the cancer vaccine. The group can comprise additional target sequences determined by cancer / tumor sequencing and known or believed to be cancer-related. The additional target sequences can comprise at least 10 target sequences, at least 20 target sequences, at least 30 target sequences, at least 100 target sequences, from 10 to 500 target sequences, from 30 to 500 target sequences, from 100 to 500 target sequences, from 10 to 100 target sequences, from 30 to 100 target sequences, or from 100 to 100 target sequences. The additional target sequences can comprise at least 10 target sequences. The additional target sequences can comprise at least 20 target sequences. The additional target sequences can comprise at least 30 target sequences. The additional target sequences can comprise at least 100 target sequences. The additional target sequences can comprise from 10 to 500 target sequences. The additional target sequences can comprise from 30 to 500 target sequences. The additional target sequences can comprise from 100 to 500 target sequences. The additional target sequences can comprise from 10 to 100 target sequences. The additional target sequences can comprise from 30 to 100 target sequences. The additional target sequences can comprise from 100 to 100 target sequences.
[0196] The tumor-uninformed polynucleotide probes can be configured to capture target sequences comprising sequences of interest, including but not limited to cancer-related genes, oncogenes, tumor suppressor genes, interferon-γ signaling pathway genes, antigen processing pathway genes, and combinations thereof.
[0197] Oncogenes (also known as "driver" mutations) are genes that are generally considered or predicted to contribute to cancer and are typically considered gain-of-function mutations (e.g., KRAS mutations). The set can include probes specific for oncogenes that contain "hotspots" within the genes, which include but are not limited to ABL1, AKT2, ALK, AR, BCL6, BCL9L, BRAF, BTK, CARD11, CCND1, CCND3, CTNNB1, DDR2, EGFR, ERBB2, ERBB3, FGFR1, FGFR3, FHOD3, FLT1, FLT3, GNAS, HRAS, KDR, KIT, KRAS, MAP2K1, MAP2K2, MECOM, MED12, MET, MTOR, NOTCH1, NOTCH2, NOTCH3, NRAS, NRG1, NTRK1, NTRK3, PDGFRA, PDGFRB, PIK3CA, PIK3CG, PSMB2, RET, ROS1, SF3B1, SMO, SYNE1, and ZBTB20. The set can include probes specific for oncogenes that contain "hotspots" within the genes, which include each of the following: ABL1, AKT2, ALK, AR, BCL6, BCL9L, BRAF, BTK, CARD11, CCND1, CCND3, CTNNB1, DDR2, EGFR, ERBB2, ERBB3, FGFR1, FGFR3, FHOD3, FLT1, FLT3, GNAS, HRAS, KDR, KIT, KRAS, MAP2K1, MAP2K2, MECOM, MED12, MET, MTOR, NOTCH1, NOTCH2, NOTCH3, NRAS, NRG1, NTRK1, NTRK3, PDGFRA, PDGFRB, PIK3CA, PIK3CG, PSMB2, RET, ROS1, SF3B1, SMO, SYNE1, and ZBTB20.
[0198] Tumor suppressor genes are genes that are generally thought to monitor and / or control tumor-related properties such as cell division, where mutations interfere with the control of such properties and are generally considered loss-of-function mutations. The group can contain probes specific for tumor suppressor genes, containing "hot spots" within the genes, which genes include but are not limited to TP53, PTEN, ARID1A, APC, AMER1, ASXL1, ATM, ATR, ATRX, AXIN2, BARD1, BRCA1, BRCA2, CASP8, CFH, CREBBP, DNMT3A, EP300, ERCC2, ERCC5, EXT1, FANCA, FANCD2, FANCI, FANCM, FAT1, FBXW7, HNF1A, MAX, MLH1, MSH3, MSH6, NF1, PIK3R1, PTCH1, PTPRT, RECQL4, RNF43, ROBO1, SLX4, SMAD2, SMAD3, SMAD4, SOX9, TCF7L2, TERT promoter, TET2, TGFBR2, TP53BP1, TSC1, TSC2, WNT16, XPC, ZFP36L2, and ZNRF3. The group can contain probes specific for tumor suppressor genes, containing "hot spots" within the genes, which genes include each of the following: TP53, PTEN, ARID1A, APC, AMER1, ASXL1, ATM, ATR, ATRX, AXIN2, BARD1, BRCA1, BRCA2, CASP8, CFH, CREBBP, DNMT3A, EP300, ERCC2, ERCC5, EXT1, FANCA, FANCD2, FANCI, FANCM, FAT1, FBXW7, HNF1A, MAX, MLH1, MSH3, MSH6, NF1, PIK3R1, PTCH1, PTPRT, RECQL4, RNF43, ROBO1, SLX4, SMAD2, SMAD3, SMAD4, SOX9, TCF7L2, TERT promoter, TET2, TGFBR2, TP53BP1, TSC1, TSC2, WNT16, XPC, ZFP36L2, and ZNRF3.
[0199] Interferon-γ signaling pathway genes are genes involved in interferon-γ signaling, such as JAK / STAT signaling pathway genes. The set can include probes specific for interferon-γ signaling pathway genes, containing "hot spots" within the genes, which genes include but are not limited to IFNGR1, INFGR2, JAK1, JAK2, and STAT1. The set can include probes specific for interferon-γ signaling pathway genes, containing "hot spots" within the genes, which genes include each of IFNGR1, INFGR2, JAK1, JAK2, and STAT1.
[0200] Antigen processing pathway genes are genes involved in antigen processing and / or presentation (e.g., presentation via MHC), and can include monitoring for loss of HLA heterozygosity. The set can include probes specific for antigen processing pathway genes, containing "hot spots" within the genes, which genes include but are not limited to B2M, HLA-A, HLA-B, HLA-C, HLA-E, TAP1, TAP2, NLRC5, CALR, CANX, PSMB2, and TAPBP. The set can include probes specific for antigen processing pathway genes, containing "hot spots" within the genes, which genes include each of the following: B2M, HLA-A, HLA-B, HLA-C, HLA-E, TAP1, TAP2, NLRC5, CALR, CANX, PSMB2, and TAPBP.
[0201] Other mutations that are also commonly associated with cancer can also be monitored, although not otherwise annotated, e.g., not yet annotated as oncogenes or tumor suppressors. The group can contain probes specific for cancer-related genes that contain "hot spots" within the genes, which include but are not limited to ABCA12, ACVR2A, AKAP9, BMPR2, COL12A1, CSMD3, DNAH5, DOCK3, FAT2, FAT3, FAT4, FGF10, FGF6, FLG, MAGI1, MDN1, MMAB, NBEA, OBSCN, PCBP1, PCLO, PLEKHA6, PROC, RAD54L, RELN, RPL22, RYR2, TCERG1, WRN, and ZDBF2. The group can contain probes specific for cancer-related genes that contain "hot spots" within the genes, which include each of the following: ABCA12, ACVR2A, AKAP9, BMPR2, COL12A1, CSMD3, DNAH5, DOCK3, FAT2, FAT3, FAT4, FGF10, FGF6, FLG, MAGI1, MDN1, MMAB, NBEA, OBSCN, PCBP1, PCLO, PLEKHA6, PROC, RAD54L, RELN, RPL22, RYR2, TCERG1, WRN, and ZDBF2.
[0202] A group having tumor-naive polynucleotide probes can be designed to capture each of cancer-related genes, oncogenes, tumor suppressor genes, interferon-γ signaling pathway genes, and antigen processing pathway genes.
[0203] An exemplary set of non-limiting tumor-agnostic polynucleotide probes includes probes specific for genes that contain "hotspots" within the genes, including but not limited to ABCA12, ACVR2A, AKAP9, BMPR2, COL12A1, CSMD3, DNAH5, DOCK3, FAT2, FAT3, FAT4, FGF10, FGF6, FLG, MAGI1, MDN1, MMAB, NBEA, OBSCN, PCBP1, PCLO, PLEKHA6, PROC, RAD54L, RELN, RPL22, RYR2, TCERG1, WRN, ZDBF2, ABL1, AKT2, ALK, AR, BCL6, BCL9L, BRAF, BTK, CARD11, CCND1, CCND3, CTNNB1, DDR2, EGFR, ERBB2, ERBB3, FGFR1, FGFR3, FHOD3, FLT1, FLT3, GNAS, HRAS, KDR, KIT, KRAS, MAP2K1, MAP2K2, MECOM, MED12, MET, MTOR, NOTCH1, NOTCH2, NOTCH3, NRAS, NRG1, NTRK1, NTRK3, PDGFRA, PDGFRB, PIK3CA, PIK3CG, RET, ROS1, SF3B1, SMO, SYNE1, ZBTB20, TP53, PTEN, ARID1A, APC, AMER1, ASXL1, ATM, ATR, ATRX, AXIN2, BARD1, BRCA1, BRCA2, CASP8, CFH, CREBBP, DNMT3A, EP300, ERCC2, ERCC5, EXT1, FANCA, FANCD2, FANCI, FANCM, FAT1, FBXW7, HNF1A, MAX, MLH1, MSH3, MSH6, NF1, PIK3R1, PTCH1, PTPRT, RECQL4, RNF43, ROBO1, SLX4, SMAD2, SMAD3, SMAD4, SOX9, TCF7L2, TERT promoter, TET2, TGFBR2, TP53BP1, TSC1, TSC2, WNT16, XPC, ZFP36L2, ZNRF3, IFNGR1, INFGR2, JAK1, JAK2, STAT1, B2M, HLA-A, HLA-B, HLA-C, HLA-E, TAP1, TAP2, NLRC5, CALR, CANX, PSMB2, TAPBP, and combinations thereof.
[0204] A set of exemplary non-limiting tumor-agnostic polynucleotide probes includes probes specific for each of the following: ABCA12, ACVR2A, AKAP9, BMPR2, COL12A1, CSMD3, DNAH5, DOCK3, FAT2, FAT3, FAT4, FGF10, FGF6, FLG, MAGI1, MDN1, MMAB, NBEA, OBSCN, PCBP1, PCLO, PLEKHA6, PROC, RAD54L, RELN, RPL22, RYR2, TCERG1, WRN, ZDBF2, ABL1, AKT2, ALK, AR, BCL6, BCL9L, BRAF, BTK, CARD11, CCND1, CCND3, CTNNB1, DDR2, EGFR, ERBB2, ERBB3, FGFR1, FGFR3, FHOD3, FLT1, FLT3, GNAS, HRAS, KDR, KIT, KRAS, MAP2K1, MAP2K2, MECOM, MED12, MET, MTOR, NOTCH1, NOTCH2, NOTCH3, NRAS, NRG1, NTRK1, NTRK3, PDGFRA, PDGFRB, PIK3CA, PIK3CG, PSMB2, RET, ROS1, SF3B1, SMO, SYNE1, ZBTB20, TP53, PTEN, ARID1A, APC, AMER1, ASXL1, ATM, ATR, ATRX, AXIN2, BARD1, BRCA1, BRCA2, CASP8, CFH, CREBBP, DNMT3A, EP300, ERCC2, ERCC5, EXT1, FANCA, FANCD2, FANCI, FANCM, FAT1, FBXW7, HNF1A, MAX, MLH1, MSH3, MSH6, NF1, PIK3R1, PTCH1, PTPRT, RECQL4, RNF43, ROBO1, SLX4, SMAD2, SMAD3, SMAD4, SOX9, TCF7L2, TERT promoter, TET2, TGFBR2, TP53BP1, TSC1, TSC2, WNT16, XPC, ZFP36L2, ZNRF3, IFNGR1, INFGR2, JAK1, JAK2, STAT1, B2M, HLA-A, HLA-B, HLA-C, HLA-E, TAP1, TAP2, NLRC5, CALR, CANX, PSMB2.
[0205] A set of exemplary non - limiting tumor - agnostic polynucleotide probes includes probes specific for each of the following: ABL1, AKT2, ALK, APC, AR, ATR, ATRX, BARD1, BCL6, BMPR1A, BRAF, BRCA1, BRCA2, BTK, CARD11, CCND1, CCND3, CDK12, CFH, CREBBP, CTNNB1, DDR2, DNMT3A, EGFR, EP300, ERBB2, ERBB3, ERCC2, ERCC5, EXT1, FANCA, FANCD2, FANCI, FANCM, FBXW7, FGF10, FGF6, FGFR1, FGFR3, FLI1, FLT1, FLT3, GNAS, HNF1A, HRAS, KDR, KIT, KRAS, MAGI1, MAP2K1, MAP2K2, MAX, MED12, MET, MLH1, MMAB, MSH3, MSH6, MTOR, NF1, NFE2L2, NOTCH1, NOTCH2, NOTCH3, NRAS, NRG1, NTRK1, NTRK3, PDGFRA, PDGFRB, PIK3CA, PIK3CG, PIK3R1, PMS2, PPARG, PROC, PTCH1, RAD54L, RAF1, RECQL4, RET, ROS1, SF3B1, SF3B2, SLX4, SMO, TERT promoter, TET2, TP53BP1, TSC1, TSC2, WRN, XPA, XPC, ZNF395, B2M, HLA - A, HLA - B, HLA - C, TAP1, TAP2, NLRC5, IFNGR1, INFGR2, JAK1, JAK2, TP53, PTEN, and ARID1A.
[0206] The probes in the set can be designed to monitor the entire coding region of a selected gene, for example, by a series of overlapping probes that span all exons of a particular gene. The probes in the set can be designed to monitor a specific region or mutation ("hotspot") within a gene. The probes in the set can be designed to include two or more probes that are configured to capture genomic regions of interest related to cancer (e.g., the entire coding region or "hotspot"). The probes in the set can be designed to include probes that contain sequences that overlap with each other. As an illustrative non - limiting example, an overlapping probe design can include two probes, each 90 nucleotides in length and offset from each other by 20 bases, which can cover 110 bp of each target.
[0207] The probe set can include at least 20 probes, at least 30 probes, at least 40 probes, at least 50 probes, at least 60 probes, at least 70 probes, at least 80 probes, at least 90 probes, at least 100 probes, at least 200 probes, at least 300 probes, at least 400 probes, or at least 500 probes. The probe set can include at least 20 probes. The probe set can include at least 30 probes. The probe set can include at least 40 probes. The probe set can include at least 50 probes. The probe set can include at least 60 probes. The probe set can include at least 70 probes. The probe set can include at least 80 probes. The probe set can include at least 90 probes. The probe set can include at least 100 probes. The probe set can include at least 200 probes. The probe set can include at least 300 probes. The probe set can include at least 400 probes. The probe set can include at least 500 probes.
[0208] The probe set can be configured to cover at least 100 kb, at least 300 kb, at least 300 kb, at least 400 kb, 100 to 400 kb, 200 to 400 kb, 300 to 400 kb, 100 to 500 kb, 200 to 500 kb, 300 to 500 kb, or 340 to 400 kb of the subject's genome. The probe set can be configured to cover at least 100 kb of the subject's genome. The probe set can be configured to cover at least 300 kb of the subject's genome. The probe set can be configured to cover at least 300 kb of the subject's genome. The probe set can be configured to cover at least 400 kb of the subject's genome. The probe set can be configured to cover 100 to 400 kb of the subject's genome. The probe set can be configured to cover 200 to 400 kb of the subject's genome. The probe set can be configured to cover 300 to 400 kb of the subject's genome. The probe set can be configured to cover 100 to 500 kb of the subject's genome. The probe set can be configured to cover 200 to 500 kb of the subject's genome. The probe set can be configured to cover 300 to 500 kb of the subject's genome. The probe set can be configured to cover 340 to 400 kb of the subject's genome.
[0209] The tumor-agnostic polynucleotide probes can include polynucleotide probes configured to capture sequences associated with a given cancer (such as CRC or NSCLC) that the subject is known or suspected to have. The tumor-agnostic polynucleotide probes can include polynucleotide probes configured to capture sequences associated with CRC. The tumor-agnostic polynucleotide probes can include polynucleotide probes configured to capture sequences associated with NSCLC. The tumor-agnostic polynucleotide probes can include polynucleotide probes configured to capture sequences associated with GEA.
[0210] The set can further include additional polynucleotide probes that are configured to capture sequences that contain polymorphisms (e.g., single nucleotide polymorphisms “SNPs”) in a human population, where the sequences containing polymorphisms can be combined to uniquely identify (“fingerprint”) a subject. Such sequences can be used, for example, if multiple subject samples are multiplex sequenced.
[0211] Hybridization generally refers to the process of joining nucleic acid strands to complementary strands through base pairing known in the art. Nucleic acids are generally considered to selectively hybridize to a reference nucleic acid sequence if the two sequences specifically hybridize to each other under moderately to highly stringent hybridization and wash conditions. Moderate and high stringency hybridization conditions are known (see, e.g., Ausubel et al., Short Protocols in Molecular Biology, 3rd ed., Wiley & Sons 1995 and Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd ed., 2001 Cold Spring Harbor, N.Y.). The hybridization protocol can be carried out at about 42° C. The hybridization buffer can include, but is not limited to, formamide, SSC, Denhardt's solution, SDS, and / or denatured carrier DNA. The hybridization protocol can include a washing step in a buffer that can include SSC and SDS at 42° C. Illustrative non-limiting hybridization protocols include hybridization at about 42° C. in 50% formamide, 5×SSC, 5×Denhardt's solution, 0.5% SDS, and 100 μg / ml denatured carrier DNA, followed by two washes at room temperature in 2×SSC and 0.5% SDS and two additional washes at 42° C. in 0.1×SSC and 0.5% SDS. Another illustrative non-limiting example of high stringency conditions includes overnight hybridization using custom designed xGen locked probes and xGen hybridization and wash kits (IDT), which includes hybridization in xGen hybridization buffer, plus a hybridization buffer enhancer in a thermal cycler at 95° C. for 30 seconds and then at 65° C. for 4-16 hours; then one wash at room temperature in xGen wash buffer; then two washes at 65° C. in xGen stringent wash buffer; and finally three washes at room temperature in wash buffer 1, wash buffer 2, and wash buffer 3 (according to the manufacturer's instructions). One of ordinary skill in the art can recognize that the above non-limiting illustrative protocols can be optimized based on a particular hybridization reaction.
[0212] The length of the bait can be from 80 to 150 base pairs (bp), including 80 to 140, 80 to 130, 80 to 120, 80 to 110, 80 to 100, 80 to 90, 90 to 150, 90 to 140, 90 to 130, 90 to 120, 90 to 110, 90 to 100, 100 to 150, 100 to 140, 100 to 130, 100 to 120, 100 to 110, 110 to 150, 110 to 140, 110 to 130, 110 to 120, 120 to 150, 120 to 140, 120 to 130, 130 to 150, 130 to 140, 140 to 150 bp. The length of the bait can be from 80 to 150 bp. The length of the bait can be from 80 to 140 bp. The length of the bait can be from 80 to 130 bp. The length of the bait can be from 80 to 120 bp. The length of the bait can be from 80 to 110 bp. The length of the bait can be from 80 to 100 bp. The length of the bait can be from 80 to 90 bp. The length of the bait can be from 90 to 150 bp. The length of the bait can be from 90 to 140 bp. The length of the bait can be from 90 to 130 bp. The length of the bait can be from 90 to 120 bp. The length of the bait can be from 90 to 110 bp. The length of the bait can be from 90 to 100 bp. The length of the bait can be from 100 to 150 bp. The length of the bait can be from 100 to 140 bp. The length of the bait can be from 100 to 130 bp. The length of the bait can be from 100 to 120 bp. The length of the bait can be from 100 to 110 bp. The length of the bait can be from 110 to 150 bp. The length of the bait can be from 110 to 140 bp. The length of the bait can be from 110 to 130 bp. The length of the bait can be from 110 to 120 bp. The length of the bait can be from 120 to 150 bp. The length of the bait can be from 120 to 140 bp. The length of the bait can be from 120 to 130 bp. The length of the bait can be from 130 to 150 bp. The length of the bait can be from 130 to 140 bp. The length of the bait can be from 140 to 150 bp.
[0213] The polynucleotide probe can include an affinity tag. An affinity tag is generally a molecule that can be covalently linked to a substrate molecule (such as a hybridization probe) and is used for subsequent purification by binding the tag to another surface or material (such as binding a biotin tag to streptavidin resin). Enrichment of the polynucleotide can be performed by affinity purification or any other suitable method based on the affinity tag used. In some embodiments, an affinity tag is added to the polynucleotide probe, DNA molecules hybridized to the probe labeled with the affinity tag are enriched; and the enriched DNA molecules are sequenced.
[0214] A polynucleotide probe (“bait”) can be biotinylated. Biotinylation refers to the covalent addition of a biotin moiety to the polynucleotide probe. The biotin moiety can include biotin or a biotin analogue, such as desthiobiotin, oxybiotin, 2-iminobiotin, diaminobiotin, biotin sulfoxide, biocytin, etc. The biotin moiety typically binds to streptavidin with an affinity of at least 10-8 M. The enrichment step using the biotinylated polynucleotide probe can be accomplished using magnetic streptavidin beads, although other supports can be used, including but not limited to microparticles, fibers, beads, and supports.
[0215] In an illustrative non-limiting example, enrichment can include the steps of: (a) attaching a biotin moiety to an oligonucleotide probe; (b) hybridizing the biotinylated probe to cfDNA; (c) enriching the biotinylated DNA molecules by binding to a support that binds biotin (such as streptavidin beads); (d) amplifying the enriched DNA using polymerase chain reaction; and (f) sequencing the amplified DNA to generate a plurality of sequence reads.
[0216] Based on the particular disease or therapy being monitored, multiple polynucleotide regions of interest can be selected for enrichment. For example, in cancer patients, sequence analysis of tumor genomic DNA can be used to identify tumor-specific mutations, which can be used to select regions of interest for disease monitoring.
[0217] Prior to sequencing, regions of interest can be enriched from cfDNA. The regions of interest can also contain polynucleotides encoding coding regions, which can include tumor exome polynucleotides.
[0218] cfDNA sequencing
[0219] Methods for cfDNA sequencing are generally known to those skilled in the art. For example, general methods for cfDNA sequencing are described in US-2020 / 0277667-A1, which is incorporated herein by reference for all purposes. Generally, any sequencing method described herein can be used.
[0220] Sequencing of the isolated cfDNA can include next-generation sequencing (NGS) or Sanger sequencing. As used herein, the term "next-generation sequencing" or "high-throughput sequencing" refers to so-called parallel synthesis sequencing or ligation sequencing platforms. NGS methods can also include nanopore sequencing methods or methods based on electronic detection. NGS can include duplex sequencing, whole exome sequencing, whole genome sequencing, de novo sequencing, phased sequencing, targeted amplicon sequencing, or shotgun sequencing. NGS can be performed on a platform such as NovaSeq using 2x151bp and 8bp index reads. Other NGS platforms include, but are not limited to, Illumina HiSeq or MiSeq, Thermo PGM or Proton, Pac Bio RS II or Sequel, Qiagen’s Gene Reader, and Oxford Nanopore MinION, or any other suitable platform. Examples of such methods are described in Margulies et al. (Nature 2005 437:376-80); Ronaghi et al. (Analytical Biochemistry 1996 242:84-9); Shendure (Science 2005 309:1728); Imelfort et al. (Brief Bioinform. 2009 10:609-18); Fox et al. (Methods Mol Biol. 2009; 553:79-108); Appleby et al. (Methods Mol Biol. 2009; 513:19-39); English (PloS One. 2012 7:e47768), and Morozova (Genomics. 2008 92:255-64), which are incorporated herein by reference for the general description of the methods and the specific steps of the methods, including all starting materials, reagents, and final products of each step.
[0221] NGS can generate at least 10,000, at least 50,000, at least 100,000, at least 500,000, at least 1M, at least 10M, at least 100M, or at least 1B sequence reads. NGS can generate at least 10,000 sequence reads. NGS can generate at least 50,000 sequence reads. NGS can generate at least 100,000 sequence reads. NGS can generate at least 500,000 sequence reads. NGS can generate at least 1M sequence reads. NGS can generate at least 10M sequence reads. NGS can generate at least 100M sequence reads. NGS can generate at least 1B sequence reads. The sequence reads can be analyzed by a computer, and thus the instructions for performing the steps can be set forth as a program that can be recorded on a suitable physical computer-readable storage medium.
[0222] Kits such as KAPA HiFi HotStart ReadyMix and NEBNext Multiple Oligos for Illumina can be used to perform whole library amplification on cfDNA (including enriched cfDNA).
[0223] As an illustrative non - limiting example of the methods described herein, whole blood can be collected from a given subject, or from a cancer subject undergoing therapy, and cfDNA can be isolated from the whole blood. Sequencing of DNA from diseased tissue (e.g., cancerous diseased tissue such as from a tumor biopsy) can be used to identify subject - specific mutations and / or tumor - specific mutations. Subject - specific mutations and / or tumor - specific mutations can be used to design a library of biotinylated polynucleotide probes and / or to guide the selection of biotinylated polynucleotide probes to enrich regions of polynucleotides of interest from the subject's cfDNA that are unique to the subject's cancer / tumor. Dual - sequencing adapters can be ligated to the cfDNA, and then the cfDNA can be analyzed by dual - sequencing to measure the frequency of all detected variant alleles.
[0224] Sequencing adapters and dual - sequencing
[0225] Generally, for methods involving next - generation sequencing, adapters are ligated to cfDNA to facilitate sequencing. The term "sequencing adapter" or "adapter" refers to an oligonucleotide that is ligated to the ends of polynucleotides from a prepared library (e.g., a fragmented cfDNA library of regions of polynucleotides of interest) prior to sequencing. Adapter ligation can be performed on end - repaired DNA with 5 - mer non - random unique molecular identifiers (IDT, Coralville, Iowa).
[0226] Sequencing adapters can be configured for dual - sequencing. Generally, dual - sequencing allows for independent tracking during the sequencing of both strands of a single DNA molecule. Paired sequences can be compared to reduce sequencing errors by excluding variants that are not present on both DNA strands. Adapters configured for dual - sequencing can include xGen UMI adapters (IDT). A general description of sequencing adapters for dual - sequencing and their uses is described in US 2017 / 0211140 A1, which is incorporated herein by reference for all purposes.
[0227] Read depth
[0228] As used herein, the sequencing read depth (expressed as X-fold, e.g., a read depth of 1000-fold, and in some cases referred to as sequencing read coverage) refers to the read coverage level (e.g., the number of unique reads) after detection and removal of duplicate reads (e.g., PCR duplicates). Generally, a greater sequencing read depth is associated with greater variant detection reliability. For example, reliable detection of variants (e.g., point mutations) occurring at frequencies greater than 5% and up to 10%, 15%, or 20% typically requires a sequencing depth of >200-fold to ensure high detection reliability.
[0229] The sequencing read depth can be the read depth of a single mutation. The sequencing read depth of a single mutation can be at least 1000-fold. The sequencing read depth of a single mutation can be at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. The sequencing read depth of a single mutation can be at least 1500-fold. The sequencing read depth of a single mutation can be at least 2000-fold. The sequencing read depth of a single mutation can be at least 2500-fold. The sequencing read depth of a single mutation can be at least 3000-fold. The sequencing read depth of a single mutation can be at least 3500-fold. The sequencing read depth of a single mutation can be at least 4000-fold. The sequencing read depth of a single mutation can be at least 4500-fold. The sequencing read depth of a single mutation can be at least 5000-fold. The range of the sequencing read depth of a single mutation can be from 1000-fold to 5000-fold, including from 1000-fold to 4000-fold, from 1000-fold to 3000-fold, from 1000-fold to 2000-fold, from 2000-fold to 5000-fold, from 2000-fold to 4000-fold, from 2000-fold to 3000-fold, from 3000-fold to 5000-fold, from 3000-fold to 4000-fold, and from 4000-fold to 5000-fold. The range of the sequencing read depth of a single mutation can be from 1000-fold to 5000-fold. The range of the sequencing read depth of a single mutation can be from 1000-fold to 4000-fold. The range of the sequencing read depth of a single mutation can be from 1000-fold to 3000-fold. The range of the sequencing read depth of a single mutation can be from 1000-fold to 2000-fold. The range of the sequencing read depth of a single mutation can be from 2000-fold to 5000-fold. The range of the sequencing read depth of a single mutation can be from 2000-fold to 4000-fold. The range of the sequencing read depth of a single mutation can be from 2000-fold to 3000-fold. The range of the sequencing read depth of a single mutation can be from 3000-fold to 5000-fold. The range of the sequencing read depth of a single mutation can be from 3000-fold to 4000-fold. The range of the sequencing read depth of a single mutation can be from 4000-fold to 5000-fold. The range of the sequencing read depth of a single mutation can be from at least 100-fold to 1000-fold.
[0230] The sequencing read depth can be the duplex read depth. The sequencing read depth can be the duplex read depth of a single mutation. The duplex read depth of a single mutation can be at least 1000-fold. The duplex read depth of a single mutation can be at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. The duplex read depth of a single mutation can be at least 1500-fold. The duplex read depth of a single mutation can be at least 2000-fold. The duplex read depth of a single mutation can be at least 2500-fold. The duplex read depth of a single mutation can be at least 3000-fold. The duplex read depth of a single mutation can be at least 3500-fold. The duplex read depth of a single mutation can be at least 4000-fold. The duplex read depth of a single mutation can be at least 4500-fold. The duplex read depth of a single mutation can be at least 5000-fold. The range of the duplex read depth of a single mutation can be from 1000-fold to 5000-fold, including from 1000-fold to 4000-fold, from 1000-fold to 3000-fold, from 1000-fold to 2000-fold, from 2000-fold to 5000-fold, from 2000-fold to 4000-fold, from 2000-fold to 3000-fold, from 3000-fold to 5000-fold, from 3000-fold to 4000-fold, and from 4000-fold to 5000-fold. The range of the duplex read depth of a single mutation can be from 1000-fold to 5000-fold. The range of the duplex read depth of a single mutation can be from 1000-fold to 4000-fold. The range of the duplex read depth of a single mutation can be from 1000-fold to 3000-fold. The range of the duplex read depth of a single mutation can be from 1000-fold to 2000-fold. The range of the duplex read depth of a single mutation can be from 2000-fold to 5000-fold. The range of the duplex read depth of a single mutation can be from 2000-fold to 4000-fold. The range of the duplex read depth of a single mutation can be from 2000-fold to 3000-fold. The range of the duplex read depth of a single mutation can be from 3000-fold to 5000-fold. The range of the duplex read depth of a single mutation can be from 3000-fold to 4000-fold. The range of the duplex read depth of a single mutation can be from 4000-fold to 5000-fold. The range of the duplex read depth of a single mutation can be from at least 100-fold to 1000-fold.
[0231] The sequencing read depth can be the average read depth. The average read depth refers to the average sequencing depth of multiple polynucleotide regions of interest (e.g., a cancer exome and / or a region of interest that is targeted for enrichment, such as by any of the tumor-informed / tumor-uninformed panel groups described herein, and / or by baits for regions with subject-specific variants and tumor-specific variants). The average read depth can be the average read depth of a cancer exome. The average read depth can be the average read depth of a region of interest targeted for enrichment by any of the tumor-informed / tumor-uninformed panel groups described herein. The average read depth can be the average read depth of a previously identified region of interest with subject-specific mutations and / or tumor-specific mutations. The average read depth can be the average read depth of enriched cfDNA. The average read depth can be the average read depth of cfDNA enriched by baits for regions with subject-specific variants and tumor-specific variants.
[0232] The average read depth can be at least 1000-fold. The average read depth can be at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. The average read depth can be at least 1500-fold. The average read depth can be at least 2000-fold. The average read depth can be at least 2500-fold. The average read depth can be at least 3000-fold. The average read depth can be at least 3500-fold. The average read depth can be at least 4000-fold. The average read depth can be at least 4500-fold. The average read depth can be at least 5000-fold. The range of the average read depth can be from 1000-fold to 5000-fold, including from 1000-fold to 4000-fold, from 1000-fold to 3000-fold, from 1000-fold to 2000-fold, from 2000-fold to 5000-fold, from 2000-fold to 4000-fold, from 2000-fold to 3000-fold, from 3000-fold to 5000-fold, from 3000-fold to 4000-fold, and from 4000-fold to 5000-fold. The range of the average read depth can be from 1000-fold to 5000-fold. The range of the average read depth can be from 1000-fold to 4000-fold. The range of the average read depth can be from 1000-fold to 3000-fold. The range of the average read depth can be from 1000-fold to 2000-fold. The range of the average read depth can be from 2000-fold to 5000-fold. The range of the average read depth can be from 2000-fold to 4000-fold. The range of the average read depth can be from 2000-fold to 3000-fold. The range of the average read depth can be from 3000-fold to 5000-fold. The range of the average read depth can be from 3000-fold to 4000-fold. The range of the average read depth can be from 4000-fold to 5000-fold. The range of the average read depth can be from at least 100-fold to 1000-fold.
[0233] The average read depth can be the average duplex read depth. The average duplex read depth can be at least 1000-fold. The average duplex read depth can be at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. The average duplex read depth can be at least 1500-fold. The average duplex read depth can be at least 2000-fold. The average duplex read depth can be at least 2500-fold. The average duplex read depth can be at least 3000-fold. The average duplex read depth can be at least 3500-fold. The average duplex read depth can be at least 4000-fold. The average duplex read depth can be at least 4500-fold. The average duplex read depth can be at least 5000-fold. The range of the average duplex read depth can be from 1000-fold to 5000-fold, including from 1000-fold to 4000-fold, from 1000-fold to 3000-fold, from 1000-fold to 2000-fold, from 2000-fold to 5000-fold, from 2000-fold to 4000-fold, from 2000-fold to 3000-fold, from 3000-fold to 5000-fold, from 3000-fold to 4000-fold, and from 4000-fold to 5000-fold. The range of the average duplex read depth can be from 1000-fold to 5000-fold. The range of the average duplex read depth can be from 1000-fold to 4000-fold. The range of the average duplex read depth can be from 1000-fold to 3000-fold. The range of the average duplex read depth can be from 1000-fold to 2000-fold. The range of the average duplex read depth can be from 2000-fold to 5000-fold. The range of the average duplex read depth can be from 2000-fold to 4000-fold. The range of the average duplex read depth can be from 2000-fold to 3000-fold. The range of the average duplex read depth can be from 3000-fold to 5000-fold. The range of the average duplex read depth can be from 3000-fold to 4000-fold. The range of the average duplex read depth can be from 4000-fold to 5000-fold. The range of the average duplex read depth can be from at least 100-fold to 1000-fold.
[0234] Multiplex analysis
[0235] The methods described herein include multiplex arrays that can sequence (“detect”) multiple polynucleotide regions of interest in a cfDNA sample. The cfDNA sample can contain ctDNA containing one or more mutant alleles that encode genes in the tumor exome. One or more polynucleotide regions of interest can be selectively enriched by designing baits that target one or more polynucleotide regions of interest. One or more polynucleotide regions of interest can be selectively enriched by designing baits that target one or more polynucleotide regions of interest from the tumor exome. One or more polynucleotide regions of interest can be selectively enriched by designing baits that target one or more polynucleotide regions of interest from the tumor exome that are known or suspected to have subject-specific mutations and tumor-specific mutations.
[0236] One or more regions of interest of polynucleotides may include 10 or more regions of interest of polynucleotides. One or more regions of interest of polynucleotides may include 20 regions of interest of polynucleotides. One or more regions of interest of polynucleotides may include 30 or more regions of interest of polynucleotides. One or more regions of interest of polynucleotides may include 40 or more regions of interest of polynucleotides. One or more regions of interest of polynucleotides may include 50 or more regions of interest of polynucleotides. One or more regions of interest of polynucleotides may include 60 or more regions of interest of polynucleotides. One or more regions of interest of polynucleotides may include 70 or more regions of interest of polynucleotides. One or more regions of interest of polynucleotides may include 80 or more regions of interest of polynucleotides. One or more regions of interest of polynucleotides may include 90 or more regions of interest of polynucleotides. One or more regions of interest of polynucleotides may include 100 or more regions of interest of polynucleotides. One or more regions of interest of polynucleotides may include 150 or more regions of interest of polynucleotides. One or more regions of interest of polynucleotides may include 200 or more regions of interest of polynucleotides. One or more regions of interest of polynucleotides may include 250 or more regions of interest of polynucleotides. One or more regions of interest of polynucleotides may include 300 or more regions of interest of polynucleotides. One or more regions of interest of polynucleotides may include 400 or more regions of interest of polynucleotides. One or more regions of interest of polynucleotides may include 500 or more regions of interest of polynucleotides. One or more regions of interest of polynucleotides may include 600 or more regions of interest of polynucleotides. One or more regions of interest of polynucleotides may include 700 or more regions of interest of polynucleotides. One or more regions of interest of polynucleotides may include 800 or more regions of interest of polynucleotides. Or one or more regions of interest of polynucleotides may include 900 or more regions of interest of polynucleotides.
[0237] One or more regions of interest of polynucleotides may comprise at least 10% of the regions of interest of polynucleotides corresponding to mutations present in the tumor exome of the subject (in other words, at least 10% of all subject-specific and tumor-specific mutations are associated with the tumor exome). One or more regions of interest of polynucleotides may comprise at least 20% of the regions of interest of polynucleotides corresponding to mutations present in the tumor exome of the subject. One or more regions of interest of polynucleotides may comprise at least 30% of the regions of interest of polynucleotides corresponding to mutations present in the tumor exome of the subject. One or more regions of interest of polynucleotides may comprise at least 40% of the regions of interest of polynucleotides corresponding to mutations present in the tumor exome of the subject. One or more regions of interest of polynucleotides may comprise at least 50% of the regions of interest of polynucleotides corresponding to mutations present in the tumor exome of the subject. One or more regions of interest of polynucleotides may comprise at least 60% of the regions of interest of polynucleotides corresponding to mutations present in the tumor exome of the subject. One or more regions of interest of polynucleotides may comprise at least 70% of the regions of interest of polynucleotides corresponding to mutations present in the tumor exome of the subject. One or more regions of interest of polynucleotides may comprise at least 80% of the regions of interest of polynucleotides corresponding to mutations present in the tumor exome of the subject. One or more regions of interest of polynucleotides may comprise at least 90% of the regions of interest of polynucleotides corresponding to mutations present in the tumor exome of the subject. One or more regions of interest of polynucleotides may comprise at least 95% of the regions of interest of polynucleotides corresponding to mutations present in the tumor exome of the subject. One or more regions of interest of polynucleotides may comprise at least 96% of the regions of interest of polynucleotides corresponding to mutations present in the tumor exome of the subject. One or more regions of interest of polynucleotides may comprise at least 97% of the regions of interest of polynucleotides corresponding to mutations present in the tumor exome of the subject. One or more regions of interest of polynucleotides may comprise at least 98% of the regions of interest of polynucleotides corresponding to mutations present in the tumor exome of the subject. One or more regions of interest of polynucleotides may comprise at least 99% of the regions of interest of polynucleotides corresponding to mutations present in the tumor exome of the subject. One or more regions of interest of polynucleotides may comprise at least 99.5% of the regions of interest of polynucleotides corresponding to mutations present in the tumor exome of the subject. One or more regions of interest of polynucleotides may comprise at least 99.9% of the regions of interest of polynucleotides corresponding to mutations present in the tumor exome of the subject. One or more regions of interest of polynucleotides may comprise 100% of the regions of interest of polynucleotides corresponding to mutations present in the tumor exome of the subject.
[0238] Mutations can include, but are not limited to, point mutations, frameshift mutations, non-frameshift mutations, deletion mutations, insertion mutations, splicing variants, genomic rearrangements, proteasome-generated spliced antigens, or combinations thereof. A mutation can include at least one alteration that changes the peptide sequence encoded by the cfDNA to be different from the corresponding peptide sequence encoded by the wild-type germline nucleic acid sequence of the subject. A mutation can consist of a coding mutation comprising at least one alteration that changes the peptide sequence encoded by the cfDNA to be different from the corresponding peptide sequence encoded by the wild-type germline nucleic acid sequence of the subject. One or more mutations can include 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, or 90 or more mutations. One or more mutations can include 100 or more, 150 or more, 200 or more, 250 or more, 300 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, or 900 or more mutations. Mutations can be associated with the tumor exome. One or more mutations can include at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the mutations present in the tumor exome of the subject. One or more mutations can include at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9%, 100% of the mutations present in the tumor exome of the subject.
[0239] Target coverage
[0240] As used herein, target coverage (usually expressed as a percentage) refers to the proportion of one or more polynucleotide regions that are sequenced (e.g., regions represented in a sequencing dataset at least some read depth). Generally, target coverage is described as the proportion of one or more desired regions to be covered (e.g., multiple polynucleotide regions of interest). For example, target coverage can be the proportion of the whole genome, exome, cancer genome, cancer exome, and / or enriched regions (e.g., cancer exome and / or regions of interest that are targeted for enrichment, such as by any of the tumor-informed / tumor-uninformed combination groups described herein, and / or by baits for regions with subject-specific variants and tumor-specific variants).
[0241] Target coverage can be the proportion of the tumor and / or cancer exome of a subject that is sequenced. Target coverage can be at least 10% of the tumor and / or cancer exome. Target coverage can be at least 20% of the tumor and / or cancer exome. Target coverage can be at least 30% of the tumor and / or cancer exome. Target coverage can be at least 40% of the tumor and / or cancer exome. Target coverage can be at least 50% of the tumor and / or cancer exome. Target coverage can be at least 60% of the tumor and / or cancer exome. Target coverage can be at least 70% of the tumor and / or cancer exome. Target coverage can be at least 80% of the tumor and / or cancer exome. Target coverage can be at least 90% of the tumor and / or cancer exome. Target coverage can be at least 95% of the tumor and / or cancer exome. Target coverage can be at least 96% of the tumor and / or cancer exome. Target coverage can be at least 97% of the tumor and / or cancer exome. Target coverage can be at least 98% of the tumor and / or cancer exome. Target coverage can be at least 99% of the tumor and / or cancer exome. Target coverage can be at least 99.5% of the tumor and / or cancer exome. Target coverage can be at least 99.9% of the tumor and / or cancer exome. Target coverage can be 100% of the tumor and / or cancer exome.
[0242] Target coverage can be the proportion of a polynucleotide region of interest that is sequenced. Target coverage can be at least 10% of the polynucleotide region of interest. Target coverage can be at least 20% of the polynucleotide region of interest. Target coverage can be at least 30% of the polynucleotide region of interest. Target coverage can be at least 40% of the polynucleotide region of interest. Target coverage can be at least 50% of the polynucleotide region of interest. Target coverage can be at least 60% of the polynucleotide region of interest. Target coverage can be at least 70% of the polynucleotide region of interest. Target coverage can be at least 80% of the polynucleotide region of interest. Target coverage can be at least 90% of the polynucleotide region of interest. Target coverage can be at least 95% of the polynucleotide region of interest. Target coverage can be at least 96% of the polynucleotide region of interest. Target coverage can be at least 97% of the polynucleotide region of interest. Target coverage can be at least 98% of the polynucleotide region of interest. Target coverage can be at least 99% of the polynucleotide region of interest. Target coverage can be at least 99.5% of the polynucleotide region of interest. Target coverage can be at least 99.9% of the polynucleotide region of interest. Target coverage can be 100% of the polynucleotide region of interest.
[0243] Target coverage can be the proportion of the polynucleotide regions of interest targeted for enrichment that are sequenced (e.g., cancer exomes and / or regions of interest, which are targeted for enrichment, e.g., by any of the tumor-informed / tumor-uninformed panel groups described herein, and / or by baits for regions with subject-specific variants and tumor-specific variants). Target coverage can be at least 10% of the polynucleotide regions of interest targeted for enrichment. Target coverage can be at least 20% of the polynucleotide regions of interest targeted for enrichment. Target coverage can be at least 30% of the polynucleotide regions of interest targeted for enrichment. Target coverage can be at least 40% of the polynucleotide regions of interest targeted for enrichment. Target coverage can be at least 50% of the polynucleotide regions of interest targeted for enrichment. Target coverage can be at least 60% of the polynucleotide regions of interest targeted for enrichment. Target coverage can be at least 70% of the polynucleotide regions of interest targeted for enrichment. Target coverage can be at least 80% of the polynucleotide regions of interest targeted for enrichment. Target coverage can be at least 90% of the polynucleotide regions of interest targeted for enrichment. Target coverage can be at least 95% of the polynucleotide regions of interest targeted for enrichment. Target coverage can be at least 96% of the polynucleotide regions of interest targeted for enrichment. Target coverage can be at least 97% of the polynucleotide regions of interest targeted for enrichment. Target coverage can be at least 98% of the polynucleotide regions of interest targeted for enrichment. Target coverage can be at least 99% of the polynucleotide regions of interest targeted for enrichment. Target coverage can be at least 99.5% of the polynucleotide regions of interest targeted for enrichment. Target coverage can be at least 99.9% of the polynucleotide regions of interest targeted for enrichment. Target coverage can be 100% of the polynucleotide regions of interest targeted for enrichment.
[0244] Target coverage can be the proportion of polynucleotide regions targeted for enrichment by any of the tumor-informed / tumor-uninformed panel groups described herein.
[0245] Target coverage can be the proportion of the polynucleotide regions of interest that are sequenced corresponding to mutations present in the tumor and / or cancer exome of a subject. Target coverage can be at least 10% of all polynucleotide regions of interest corresponding to mutations present in the tumor and / or cancer exome of a subject (e.g., the coverage is at least 10% of all subject-specific and tumor-specific mutations associated with the tumor and / or cancer exome). Target coverage can be at least 20% of all polynucleotide regions of interest corresponding to mutations present in the tumor and / or cancer exome of a subject. Target coverage can be at least 30% of all polynucleotide regions of interest corresponding to mutations present in the tumor and / or cancer exome of a subject. Target coverage can be at least 40% of all polynucleotide regions of interest corresponding to mutations present in the tumor and / or cancer exome of a subject. Target coverage can be at least 50% of all polynucleotide regions of interest corresponding to mutations present in the tumor and / or cancer exome of a subject. Target coverage can be at least 60% of all polynucleotide regions of interest corresponding to mutations present in the tumor and / or cancer exome of a subject. Target coverage can be at least 70% of all polynucleotide regions of interest corresponding to mutations present in the tumor and / or cancer exome of a subject. Target coverage can be at least 80% of all polynucleotide regions of interest corresponding to mutations present in the tumor and / or cancer exome of a subject. Target coverage can be at least 90% of all polynucleotide regions of interest corresponding to mutations present in the tumor and / or cancer exome of a subject. Target coverage can be at least 95% of all polynucleotide regions of interest corresponding to mutations present in the tumor and / or cancer exome of a subject. Target coverage can be at least 96% of all polynucleotide regions of interest corresponding to mutations present in the tumor and / or cancer exome of a subject. Target coverage can be at least 97% of all polynucleotide regions of interest corresponding to mutations present in the tumor and / or cancer exome of a subject. Target coverage can be at least 98% of all polynucleotide regions of interest corresponding to mutations present in the tumor and / or cancer exome of a subject. Target coverage can be at least 99% of all polynucleotide regions of interest corresponding to mutations present in the tumor and / or cancer exome of a subject. Target coverage can be at least 99.5% of all polynucleotide regions of interest corresponding to mutations present in the tumor and / or cancer exome of a subject. Target coverage can be at least 99.9% of all polynucleotide regions of interest corresponding to mutations present in the tumor and / or cancer exome of a subject. Target coverage can be 100% of all polynucleotide regions of interest corresponding to mutations present in the tumor and / or cancer exome of a subject.
[0246] Target coverage can be the proportion of the tumor and / or cancer genome of a subject for which sequencing is obtained. Target coverage can be at least 10% of the tumor and / or cancer genome. Target coverage can be at least 20% of the tumor and / or cancer genome. Target coverage can be at least 30% of the tumor and / or cancer genome. Target coverage can be at least 40% of the tumor and / or cancer genome. Target coverage can be at least 50% of the tumor and / or cancer genome. Target coverage can be at least 60% of the tumor and / or cancer genome. Target coverage can be at least 70% of the tumor and / or cancer genome. Target coverage can be at least 80% of the tumor and / or cancer genome. Target coverage can be at least 90% of the tumor and / or cancer genome. Target coverage can be at least 95% of the tumor and / or cancer genome. Target coverage can be at least 96% of the tumor and / or cancer genome. Target coverage can be at least 97% of the tumor and / or cancer genome. Target coverage can be at least 98% of the tumor and / or cancer genome. Target coverage can be at least 99% of the tumor and / or cancer genome. Target coverage can be at least 99.5% of the tumor and / or cancer genome. Target coverage can be at least 99.9% of the tumor and / or cancer genome. Target coverage can be 100% of the tumor and / or cancer genome.
[0247] Target coverage can be the percentage of the region of interest that is sequenced, and where the region of interest that is sequenced has a read depth greater than a particular value. Target coverage can be the percentage of the region of interest that is sequenced, and where the region of interest that is sequenced has a read depth of at least 1000-fold. Target coverage can be the percentage of the region of interest that is sequenced, and where the region of interest that is sequenced has a read depth of at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. Target coverage can be the percentage of the region of interest that is sequenced, and where the region of interest that is sequenced has a read depth of at least 1500-fold. Target coverage can be the percentage of the region of interest that is sequenced, and where the region of interest that is sequenced has a read depth of at least 2000-fold. Target coverage can be the percentage of the region of interest that is sequenced, and where the region of interest that is sequenced has a read depth of at least 2500-fold. Target coverage can be the percentage of the region of interest that is sequenced, and where the region of interest that is sequenced has a read depth of at least 3000-fold. Target coverage can be the percentage of the region of interest that is sequenced, and where the region of interest that is sequenced has a read depth of at least 3500-fold. Target coverage can be the percentage of the region of interest that is sequenced, and where the region of interest that is sequenced has a read depth of at least 4000-fold. Target coverage can be the percentage of the region of interest that is sequenced, and where the region of interest that is sequenced has a read depth of at least 4500-fold. Target coverage can be the percentage of the region of interest that is sequenced, and where the region of interest that is sequenced has a read depth of at least 5000-fold.
[0248] Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has more than a particular average read depth. Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has an average read depth of at least 1000-fold. Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has an average read depth of at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has an average read depth of at least 1500-fold. Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has an average read depth of at least 2000-fold. Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has an average read depth of at least 2500-fold. Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has an average read depth of at least 3000-fold. Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has an average read depth of at least 3500-fold. Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has an average read depth of at least 4000-fold. Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has an average read depth of at least 4500-fold. Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has an average read depth of at least 5000-fold.
[0249] Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has a depth of duplex reads above a particular level. Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has a depth of duplex reads of at least 1000-fold. Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has a depth of duplex reads of at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has a depth of duplex reads of at least 1500-fold. Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has a depth of duplex reads of at least 2000-fold. Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has a depth of duplex reads of at least 2500-fold. Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has a depth of duplex reads of at least 3000-fold. Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has a depth of duplex reads of at least 3500-fold. Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has a depth of duplex reads of at least 4000-fold. Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has a depth of duplex reads of at least 4500-fold. Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has a depth of duplex reads of at least 5000-fold.
[0250] Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has an average depth of duplex reads above a particular level. Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has an average depth of duplex reads of at least 1000-fold. Target coverage can be the percentage of the region of interest that is sequenced and where the region of interest that is sequenced has an average depth of duplex reads of at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold.
[0251] The target coverage can be at least 10% of the region of interest, and the sequenced region of interest has a read depth or average read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. The target coverage can be at least 20% of the region of interest, and the sequenced region of interest has a read depth or average read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. The target coverage can be at least 30% of the region of interest, and the sequenced region of interest has a read depth or average read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. The target coverage can be at least 40% of the region of interest, and the sequenced region of interest has a read depth or average read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. The target coverage can be at least 50% of the region of interest, and the sequenced region of interest has a read depth or average read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. The target coverage can be at least 60% of the region of interest, and the sequenced region of interest has a read depth or average read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. The target coverage can be at least 70% of the region of interest, and the sequenced region of interest has a read depth or average read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. The target coverage can be at least 80% of the region of interest, and the sequenced region of interest has a read depth or average read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold.The target coverage can be at least 90% of the region of interest, and the region of interest being sequenced has a read depth or average read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold or at least 5000-fold. The target coverage can be at least 95% of the region of interest, and the region of interest being sequenced has a read depth or average read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold or at least 5000-fold. The target coverage can be at least 96% of the region of interest, and the region of interest being sequenced has a read depth or average read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold or at least 5000-fold. The target coverage can be at least 97% of the region of interest, and the region of interest being sequenced has a read depth or average read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold or at least 5000-fold. The target coverage can be at least 98% of the region of interest, and the region of interest being sequenced has a read depth or average read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold or at least 5000-fold. The target coverage can be at least 99% of the region of interest, and the region of interest being sequenced has a read depth or average read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold or at least 5000-fold. The target coverage can be at least 99.5% of the region of interest, and the region of interest being sequenced has a read depth or average read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold or at least 5000-fold. The target coverage can be at least 99.9% of the region of interest, and the region of interest being sequenced has a read depth or average read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold or at least 5000-fold.The target coverage can be 100% of the region of interest, and the region of interest being sequenced has a read depth or average read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold.
[0252] The target coverage can be at least 10% of the region of interest, and the sequenced region of interest has a duplex read depth or average duplex read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold or at least 5000-fold. The target coverage can be at least 20% of the region of interest, and the sequenced region of interest has a duplex read depth or average duplex read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold or at least 5000-fold. The target coverage can be at least 30% of the region of interest, and the sequenced region of interest has a duplex read depth or average duplex read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold or at least 5000-fold. The target coverage can be at least 40% of the region of interest, and the sequenced region of interest has a duplex read depth or average duplex read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold or at least 5000-fold. The target coverage can be at least 50% of the region of interest, and the sequenced region of interest has a duplex read depth or average duplex read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold or at least 5000-fold. The target coverage can be at least 60% of the region of interest, and the sequenced region of interest has a duplex read depth or average duplex read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold or at least 5000-fold. The target coverage can be at least 70% of the region of interest, and the sequenced region of interest has a duplex read depth or average duplex read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold or at least 5000-fold. The target coverage can be at least 80% of the region of interest, and the sequenced region of interest has a duplex read depth or average duplex read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold or at least 5000-fold.The target coverage can be at least 90% of the region of interest, and the region of interest being sequenced has a duplex read depth or average duplex read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. The target coverage can be at least 95% of the region of interest, and the region of interest being sequenced has a duplex read depth or average duplex read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. The target coverage can be at least 96% of the region of interest, and the region of interest being sequenced has a duplex read depth or average duplex read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. The target coverage can be at least 97% of the region of interest, and the region of interest being sequenced has a duplex read depth or average duplex read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. The target coverage can be at least 98% of the region of interest, and the region of interest being sequenced has a duplex read depth or average duplex read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. The target coverage can be at least 99% of the region of interest, and the region of interest being sequenced has a duplex read depth or average duplex read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. The target coverage can be at least 99.5% of the region of interest, and the region of interest being sequenced has a duplex read depth or average duplex read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold. The target coverage can be at least 99.9% of the region of interest, and the region of interest being sequenced has a duplex read depth or average duplex read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold, or at least 5000-fold.The target coverage can be 100% of the region of interest, and the sequenced region of interest has a duplex read depth or average duplex read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold or at least 5000-fold.
[0253] The target coverage can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% of the region of interest, and the sequenced region of interest has a read depth or average read depth of at least 1000-fold. The target coverage can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% of the region of interest, and the sequenced region of interest has a read depth or average read depth of at least 1500-fold. The target coverage can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% of the region of interest, and the sequenced region of interest has a read depth or average read depth of at least 2000-fold. The target coverage can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% of the region of interest, and the sequenced region of interest has a read depth or average read depth of at least 2500-fold. The target coverage can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% of the region of interest, and the sequenced region of interest has a read depth or average read depth of at least 3000-fold. The target coverage can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% of the region of interest, and the sequenced region of interest has a read depth or average read depth of at least 3500-fold. The target coverage can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% of the region of interest, and the sequenced region of interest has a read depth or average read depth of at least 4000-fold. The target coverage can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% of the region of interest, and the sequenced region of interest has a read depth or average read depth of at least 4500-fold. The target coverage can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% of the region of interest, and the sequenced region of interest has a read depth or average read depth of at least 5000-fold.
[0254] The target coverage can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% of the region of interest, and the sequenced region of interest has a duplex read depth or average duplex read depth of at least 1000-fold. The target coverage can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% of the region of interest, and the sequenced region of interest has a duplex read depth or average duplex read depth of at least 1500-fold. The target coverage can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% of the region of interest, and the sequenced region of interest has a duplex read depth or average duplex read depth of at least 2000-fold. The target coverage can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% of the region of interest, and the sequenced region of interest has a duplex read depth or average duplex read depth of at least 2500-fold. The target coverage can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% of the region of interest, and the sequenced region of interest has a duplex read depth or average duplex read depth of at least 3000-fold. The target coverage can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% of the region of interest, and the sequenced region of interest has a duplex read depth or average duplex read depth of at least 3500-fold. The target coverage can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% of the region of interest, and the sequenced region of interest has a duplex read depth or average duplex read depth of at least 4000-fold. The target coverage can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% of the region of interest, and the sequenced region of interest has a duplex read depth or average duplex read depth of at least 4500-fold.The target coverage can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% of the region of interest, and the sequenced region of interest has a double-stranded read depth or average double-stranded read depth of at least 5000-fold.
[0255] Evaluate
[0256] After sequencing, the sequence reads can be analyzed to provide a quantitative determination of the variant allele frequency (also referred to as the mutant allele frequency) within the cfDNA of the subject. Methods for quantifying the sequence reads and variant allele frequency (VAF) are known to those of skill in the art. Computational programs for sequencing analysis and VAF calculation include, but are not limited to, BWA-MEM (Durbin et al., Bioinformatics, 2010), the fgbio toolkit (Fulcrum Genomics), and freebayes (Marth et al., arXiv 2012), the disclosures of which are hereby incorporated by reference in their entireties for all purposes. Generally, the frequency of one or more mutations (e.g., VAF) in the subject's cfDNA is expressed as the percentage of mutant-specific sequencing reads relative to the wild-type germline nucleic acid sequence reads of the subject. For example, the mutation frequency can be determined by counting the specific variant allele reads relative to the total cfDNA count of a sample taken from the subject. Additionally, VAF evaluation can be combined with the cfDNA concentration (e.g., ng / ml) in the plasma to estimate the tumor genomic concentration in the plasma (see Bos et al., Molecular Oncology (2020) doi:10.1002 / 1878-0261.12827 and Reinert et al., JAMA Oncol. 2019;5(8):1124-1131. Doi:10.1001 / jamaoncol.2019.0528, the disclosures of which are hereby incorporated by reference in their entireties for all purposes).
[0257] After determining the frequency of one or more mutations (or the estimated tumor genome per milliliter of plasma) (e.g., VAF) in the subject's cfDNA, the mutation frequency or estimated tumor genome content can then be evaluated to characterize various diseases or subject attributes, such as the disease state of the subject, the efficacy of a therapy, or a combination thereof. Mutation frequency
[0258] For example, an assessment can be made to evaluate the disease state of a subject, such as evaluating the tumor burden of the subject. The assessment of tumor burden can be used for various applications, such as part of disease diagnosis, disease prognosis, disease prediction, and / or monitoring of disease progression. The assessment of disease progression can be performed by comparing the mutation frequencies in samples taken from the subject at different time points. The change in mutation frequency can be relative to a fixed time point, such as a baseline mutation frequency, such as the mutation frequency determined on the first day of a treatment regimen.
[0259] An increase in the mutation frequency from cfDNA mutation analysis of a first collected sample (e.g., an early longitudinal sample) relative to the mutation frequency from cfDNA mutation analysis of a second sample (e.g., a later longitudinal sample) can be evaluated as disease progression, non - response to therapy, and / or disease recurrence. A decrease in the mutation frequency from cfDNA mutation analysis of a first collected sample (e.g., an early longitudinal sample) relative to the mutation frequency from cfDNA mutation analysis of a second sample (e.g., a later longitudinal sample) can be evaluated as a response. The response can be a complete response (CR) or a partial response (PR). An increase in the mutation frequency in cfDNA after therapy relative to cfDNA before therapy can indicate an increased likelihood of an increase in the tumor burden of the subject. A decrease or maintenance of the mutation frequency in cfDNA after therapy relative to cfDNA before therapy can indicate an increased likelihood of a decrease or stability of the tumor burden of the subject.
[0260] An increase in mutation frequency (or the estimated tumor genome per milliliter of plasma) over time can be evaluated as disease progression and / or recurrence. The increase in mutation frequency can be a relative increase of at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 100% of the mutation frequency between the time points to be evaluated for progression and / or recurrence. The increase in mutation frequency can be a relative increase of at least 2 - fold, at least 3 - fold, at least 4 - fold, at least 5 - fold, at least 6 - fold, at least 7 - fold, at least 8 - fold, at least 9 - fold, or at least 10 - fold of the mutation frequency between the time points to be evaluated for progression and / or recurrence. The increase in mutation frequency can be a relative increase of at least 20 - fold, at least 30 - fold, at least 40 - fold, at least 50 - fold, at least 60 - fold, at least 70 - fold, at least 80 - fold, at least 90 - fold, or at least 100 - fold of the mutation frequency between the time points to be evaluated for progression and / or recurrence.
[0261] A decrease in the mutation frequency (or the estimated tumor genome per milliliter of plasma) over time can be evaluated as disease remission. The decrease in the mutation frequency can be a relative increase of at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 100% in the mutation frequency between the time points to be evaluated for remission. The decrease in the mutation frequency can be a relative increase of at least 2-fold, at least 3-fold, at least 4-fold, at least 5-fold, at least 6-fold, at least 7-fold, at least 8-fold, at least 9-fold, or at least 10-fold in the mutation frequency between the time points to be evaluated for remission. The decrease in the mutation frequency can be a relative increase of at least 20-fold, at least 30-fold, at least 40-fold, at least 50-fold, at least 60-fold, at least 70-fold, at least 80-fold, at least 90-fold, or at least 100-fold in the mutation frequency between the time points to be evaluated for remission. The decrease in the mutation frequency can reach a mutation level undetectable in cfDNA, thus being evaluated as remission, such as being evaluated as complete remission.
[0262] An assessment can be made to evaluate the de novo mutation status of a subject, such as evaluating whether the subject's cancer or tumor has developed into a tumor escape mutant (e.g., see "Cancer Monitoring"). The increase in the mutation frequency of the cfDNA mutation analysis from the first collected sample (e.g., an early longitudinal sample) relative to the mutation frequency of the cfDNA mutation analysis from the second sample (e.g., a later longitudinal sample) can be evaluated based on the mutation frequency (including de novo occurrences) related to the regions targeted by the tumor-naïve probe sets described herein.
[0263] To evaluate the effect of a therapy on a disease, the mutation frequency (or the estimated tumor genome per milliliter of plasma) in cfDNA can be compared between a sample collected before the therapy and a sample collected after the therapy. An increase in the mutation frequency of the cfDNA mutation analysis of the sample collected before the therapy relative to the mutation frequency of the cfDNA mutation analysis of the sample collected after the therapy can be evaluated as disease progression, non-response to the therapy, and / or disease recurrence. A decrease in the mutation frequency of the cfDNA mutation analysis of the sample collected before the therapy relative to the mutation frequency of the cfDNA mutation analysis of the sample collected after the therapy can be evaluated as a response. The response can be a complete response (CR) or a partial response (PR). An increase in the mutation frequency in cfDNA after the therapy relative to cfDNA before the therapy can indicate an increased likelihood of an increase in the tumor burden of the subject. A decrease or maintenance of the mutation frequency in cfDNA after the therapy relative to cfDNA before the therapy can indicate an increased likelihood of a decrease or stability of the tumor burden of the subject.
[0264] After the assessment step, further therapy can be administered to the subject. For example, initial measurements can be obtained from a patient prior to initiating a multi-dose anti-cancer therapy regimen. Follow-up measurements can be performed prior to administering each dose. Analysis of the variant allele frequency in cfDNA at each stage can allow assessment of the patient's response to each dose of the therapy regimen. The assessment can further guide clinical decisions, including dose, therapy selection, etc. For example, clinical decisions (including introducing a new therapy or stopping the current therapy) can be informed by assessing the tumor escape mutation status associated with the regions targeted by the tumor-naive probe sets described herein.
[0265] Therapeutic treatment
[0266] The methods described herein can be performed after administering therapy to a patient. The therapy can include a cancer vaccine. The therapy can include targeted radiotherapy (e.g., external beam radiation, brachytherapy). The therapy can include immune checkpoint inhibitors, including but not limited to PD-1 inhibitors (e.g., nivolumab, pembrolizumab), PD-L1 inhibitors (e.g., avelumab, durvalumab) or CTLA-4 inhibitors (e.g., ipilimumab). The therapy can include targeted therapy techniques such as monoclonal antibody therapy (e.g., trastuzumab, bevacizumab), retinoids (e.g., ATRA, bexarotene), selective steroid hormone receptor modulators (e.g., tamoxifen, toremifene), or oncoprotein inhibitors such as tyrosine kinase (TK) (e.g., imatinib, erlotinib), mammalian target of rapamycin (mTOR) (e.g., everolimus, temsirolimus), or histone deacetylase (HDAC) (e.g., valproate, vorinostat). The therapy can include cytotoxic chemotherapy. Examples of cytotoxic chemotherapeutic agents include cisplatin, carboplatin, oxaliplatin, nedaplatin, azacitidine, capecitabine, carmofur, cladribine, clofarabine, cytarabine, decitabine, fluorouracil, floxuridine, fludarabine, mercaptopurine, nelarabine, pentostatin, tegafur, thioguanine, methotrexate, pemetrexed, raltitrexed, hydroxyurea, irinotecan, topotecan, danorubicin, doxorubicin, epirubicin, idarubicin, mitoxantrone, valrubicin, etoposide, teniposide, docetaxel, paclitaxel, vinblastine, vincristine, vindesine, vinflunine, vinorelbine, bendamustine, busulfan, carmustine, chlorambucil, mechlorethamine, dacarbazine, fotemustine, ifosfamide, lomustine, melphalan, streptozocin, gemcitabine, cyclophosphamide, temozolomide, dacarbazine, altretamine, bleomycin, bortezomib, actinomycin D, estramustine, ixabepilone, mitomycin and procarbazine.
[0267] Also provided is a method of inducing a tumor - specific immune response, vaccinating against a tumor, treating, and / or alleviating cancer symptoms in a subject by administering to the subject one or more antigens, such as a plurality of antigens identified using the methods disclosed herein.
[0268] In some aspects, the subject has been diagnosed with cancer or is at risk of developing cancer. The subject can be a human, dog, cat, horse, or any animal in which a tumor - specific immune response is desired. The tumor can be any solid tumor, such as breast cancer, ovarian cancer, prostate cancer, lung cancer, kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, melanoma, and other tissue and organ tumors, as well as hematological tumors, such as lymphoma and leukemia, including acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, T - cell lymphocytic leukemia, and B - cell lymphoma.
[0269] An amount of antigen sufficient to induce a CTL response can be administered. An amount of antigen sufficient to induce a T - cell response can be administered. An amount of antigen sufficient to induce a B - cell response can be administered.
[0270] The antigen can be administered alone or in combination with other therapeutic agents, such as chemotherapy, immune checkpoint blockade, and / or other immunotherapies.
[0271] The optimal amount and optimal dosing regimen of each antigen included in the vaccine composition can be determined. For example, the antigen or its variant can be prepared for intravenous (i.v.) injection, subcutaneous (s.c.) injection, intradermal (i.d.) injection, intraperitoneal (i.p.) injection, intramuscular (i.m.) injection. Methods of injection include subcutaneous injection, intradermal injection, intraperitoneal injection, intramuscular injection, and intravenous injection. Methods of DNA or RNA injection include intradermal injection, intramuscular injection, subcutaneous injection, intraperitoneal injection, and intravenous injection. Other methods of administering the vaccine composition are known to those of skill in the art.
[0272] The vaccine can be compiled such that the selection, number, and / or amount of antigen present in the composition is tissue, cancer, and / or subject - specific. For example, the exact selection of peptides can be guided by the expression pattern of the parental protein in a given tissue, or by the patient's mutation or disease state. The selection can depend on the specific type of cancer, the state of the disease, the goal of vaccination (e.g., prevention or against an ongoing disease), the previous treatment regimen, the patient's immune status, and of course the patient's HLA haplotype. Additionally, the vaccine can contain individualized components according to the individual needs of a particular patient. Examples include changing the selection of antigen based on the expression of antigen in a particular patient or adjusting secondary treatment after the first round or regimen of treatment.
[0273] Patients suitable for antigen vaccine administration can be determined by using various diagnostic methods (such as the patient selection methods further described below). Patient selection can involve determining mutations or expression patterns of one or more genes. In some cases, patient selection involves determining the patient's haplotype. Various patient selection methods can be performed in parallel. For example, sequencing diagnostics can simultaneously determine a patient's mutations and haplotype. Various patient selection methods can be performed sequentially. For example, one diagnostic test determines the mutations and a separate diagnostic test determines the patient's haplotype, and each test can be the same (e.g., two high-throughput sequencing) or different (e.g., one high-throughput sequencing and another Sanger sequencing) diagnostic methods.
[0274] For compositions used as vaccines against cancer, antigens with similar normal self-peptides that are highly expressed in normal tissues can be avoided or present in small amounts in the compositions described herein. On the other hand, if a patient's tumor is known to express a large amount of a certain antigen, the corresponding pharmaceutical composition for treating this cancer can be present in large amounts and / or can include more than one antigen specific to this particular antigen or its pathway.
[0275] Compositions containing antigens can be administered to individuals who have already developed cancer. In a therapeutic application, the composition is administered to the patient in an amount sufficient to elicit a therapeutically effective response, for example, in an amount sufficient to stimulate an effective CTL response against tumor antigens and cure or at least partially inhibit symptoms and / or complications. The amount sufficient to achieve this goal is defined as a "therapeutically effective dose". The amount effective for this use will depend on, for example, the composition, the mode of administration, the stage and severity of the disease being treated, the patient's weight and overall health status, and the judgment of the prescribing physician. It should be remembered that the compositions are generally used in serious conditions, i.e., life-threatening or potentially life-threatening situations, especially when the cancer has metastasized. In such cases, given the minimization of foreign substances and the relatively non-toxic nature of the antigens, it is possible and felt desirable for the treating physician to administer significantly excessive amounts of these compositions.
[0276] For therapeutic use, administration can be initiated after detection or surgical resection of the tumor. Booster doses can be administered thereafter until at least the symptoms are substantially eliminated and for a period of time thereafter.
[0277] A pharmaceutical composition (e.g., a vaccine composition) for therapeutic treatment is intended for parenteral, topical, nasal, oral, or local administration. The pharmaceutical composition can be administered parenterally, e.g., intravenously, subcutaneously, intradermally, or intramuscularly. The composition can be administered at the surgical resection site to induce a local immune response against the tumor. The composition can be administered to target specific diseased tissues and / or cells of a subject. Compositions for parenteral administration are disclosed herein, which compositions comprise a solution of an antigen, and the vaccine composition is dissolved or suspended in an acceptable carrier (e.g., an aqueous carrier). A variety of aqueous carriers can be used, such as water, buffered water, 0.9% saline, 0.3% glycine, hyaluronic acid, etc. These compositions can be sterilized by conventional, well-known sterilization techniques, or they can be filter sterilized. The resulting aqueous solution can be packaged for use as is or in lyophilized form, and the lyophilized preparation is combined with a sterile solution prior to administration. The composition can contain pharmaceutically acceptable auxiliary substances as required to approximate physiological conditions, such as pH regulators and buffers, tonicity regulators, wetting agents, etc., such as sodium acetate, sodium lactate, sodium chloride, potassium chloride, calcium chloride, sorbitan monolaurate, triethanolamine oleate, etc.
[0278] The antigen can also be administered via liposomes, which target specific cell tissues, such as lymphoid tissues. Liposomes can also be used to extend the half-life. Liposomes include emulsions, foams, micelles, insoluble monolayers, liquid crystals, phospholipid dispersions, lamellar layers, and the like. In these formulations, the antigen to be delivered is incorporated as part of the liposome, either alone or in combination with a molecule that binds to a receptor prevalent in, for example, lymphoid cells, such as a monoclonal antibody that binds to the CD45 antigen, or in combination with other therapeutic or immunogenic compositions. Thus, liposomes filled with the desired antigen can be directed to the site of lymphoid cells, where the liposomes then deliver the selected therapeutic / immunogenic composition. Liposomes can be formed from standard vesicle-forming lipids, which lipids generally include neutral or negatively charged phospholipids and sterols, such as cholesterol. The choice of lipid is generally guided by considering, for example, liposome size, acid lability, and stability of the liposome in the bloodstream. There are a variety of methods for preparing liposomes, as described, for example, in Szoka et al., Ann. Rev. Biophys. Bioeng. 9; 467 (1980), U.S. Patents 4,235,871, 4,501,728, 4,501,728, 4,837,028, and 5,019,369.
[0279] To target immune cells, ligands incorporated into liposomes can include, for example, antibodies or fragments thereof that are specific for cell surface determinants of the desired immune system cells. The liposome suspension can be administered intravenously, locally / topically, etc., at a dose that varies particularly according to the mode of administration, the peptide being delivered, and the stage of the disease being treated.
[0280] For therapeutic or immunization purposes, nucleic acids encoding peptides and optionally one or more of the peptides described herein can also be administered to a patient. Many methods are conveniently used to deliver these nucleic acids to the patient. For example, the nucleic acid can be delivered directly as "naked DNA". For example, such methods are described in Wolff et al., Science 247:1465-1468 (1990) and U.S. Patents 5,580,859 and 5,589,466. Ballistic delivery can also be used to administer nucleic acids, as described in, for example, U.S. Patent No. 5,204,253. Particles consisting solely of DNA can be administered. Alternatively, the DNA can be adhered to particles such as gold particles. Methods for delivering nucleic acid sequences can include viral vectors, mRNA vectors, and DNA vectors, with or without electroporation.
[0281] Nucleic acids can also be delivered by complexing with cationic compounds such as cationic lipids. Lipid-mediated gene delivery methods have been described in, for example, WO 96 / 18372; WO 93 / 24640; Mannino and Gould-Fogerite, BioTechniques 6(7):682-691 (1988); U.S. Patent No. 5,279,833 to Rose; WO 91 / 06309; and Felgner et al., Proc. Natl. Acad. Sci. USA 84:7413-7414 (1987).
[0282] Antigens can also be included in virus vector-based vaccine platforms such as vaccinia, fowlpox, self-replicating alphaviruses, Maraba virus, adenoviruses (see, e.g., Tatsis et al., Adenoviruses, Molecular Therapy (2004) 10, 616-629) or lentiviruses, including but not limited to second-generation, third-generation or mixed second / third-generation lentiviruses and any generation of recombinant lentiviruses designed to target specific cell types or receptors (see, e.g., Hu et al., Immunization Delivered by Lentiviral Vectors for Cancer and Infectious Diseases, Immunol Rev. (2011) 239(1):45-61, Sakuma et al., Lentiviral vectors: basic to translational, Biochem J. (2012) 443(3):603-18, Cooper et al., Rescue of splicing-mediated intron loss maximizes expression in lentiviral vectors containing the human ubiquitin C promoter, Nucl. Acids Res. (2015) 43(1):682-690, Zufferey et al., Self-Inactivating Lentivirus Vector for Safe and Efficient In Vivo Gene Delivery, J. Virol. (1998) 72(12):9873-9880). Depending on the packaging capacity of the virus vector-based vaccine platforms described above, the method can deliver one or more nucleotide sequences encoding one or more antigenic peptides.The sequences can be flanked by non-mutated sequences, can be separated by linkers, or can have one or more sequences targeting subcellular compartments in front (see, e.g., Gros et al., Prospective identification of neoantigen-specific lymphocytes in the peripheral blood of melanoma patients, Nat Med. (2016) 22(4):433-8, Stronen et al., Targeting of cancer neoantigens with donor-derived T cell receptor repertoires, Science. (2016) 352(6291):1337-41, Lu et al., Efficient identification of mutated cancer antigens recognized by T cells associated with durable tumor regressions, Clin Cancer Res. (2014) 20(13):3401-10). After being introduced into the host, the vector-infected cells express the antigen and thus trigger a host immune (e.g., CTL) response against the peptide. Vaccinia vectors and methods for immunization protocols are described, for example, in U.S. Patent No. 4,722,848. Another vector is BCG (Bacillus Calmette-Guérin). The BCG vector is described in Stover et al. (Nature 351:456-460 (1991)). Based on the description herein, those skilled in the art will be clear about a variety of other vaccine vectors that can be used for therapeutic administration or antigen immunization, such as Salmonella typhi vectors, etc.
[0283] The vaccine may comprise an epitope-encoding nucleic acid, the sequence of which encodes one or more tumor-specific mutations and / or subject-specific mutations, such as one or more mutations whose frequencies in cfDNA have been determined. The vaccine system may comprise an alphavirus-based self-replicating expression system that encodes the epitope-encoding nucleic acid, the sequence of which encodes one or more tumor-specific mutations and / or subject-specific mutations. The alphavirus-based self-replicating expression system used as a cancer vaccine is described in International Patent Application Publication WO / 2018 / 208856, which is hereby incorporated by reference in its entirety for all purposes. The vaccine system may comprise a chimpanzee adenovirus (ChAdV)-based expression system that encodes the epitope-encoding nucleic acid, the sequence of which encodes one or more tumor-specific mutations and / or subject-specific mutations. The ChAdV-based expression system used as a cancer vaccine is described in International Patent Application Publication WO / 2018 / 098362, which is hereby incorporated by reference in its entirety for all purposes.
[0284] The method of administering the nucleic acid is to use a small gene construct encoding one or more epitopes. To generate the DNA sequence (mini-gene) encoding the selected CTL epitopes for expression in human cells, the amino acid sequences of these epitopes are reverse-translated. The human codon usage table is used to guide the codon selection for each amino acid. These epitope-encoding DNA sequences are directly adjacent to create a continuous polypeptide sequence. To optimize expression and / or immunogenicity, additional elements can be incorporated into the mini-gene design. Examples of amino acid sequences that can be reverse-translated and included in the mini-gene sequence include: helper T lymphocytes, epitopes, leader (signal) sequences, and endoplasmic reticulum retention signals. In addition, MHC presentation of CTL epitopes can be improved by incorporating synthetic (e.g., polyalanine) or naturally occurring flanking sequences adjacent to the CTL epitopes. The mini-gene sequence can be converted into DNA by assembling oligonucleotides encoding the positive and negative strands of the mini-gene. Using well-known techniques, overlapping oligonucleotides (30-100 bases long) are synthesized, phosphorylated, purified, and annealed under appropriate conditions. The ends of the oligonucleotides can be ligated using T4 DNA ligase. Subsequently, this synthetic mini-gene encoding the CTL epitope polypeptide can be cloned into the desired expression vector.
[0285] Purified plasmid DNA for injection can be prepared using a variety of formulations. The simplest method among them is to reconstitute the lyophilized DNA in sterile phosphate-buffered saline (PBS). A variety of methods have been described, and new techniques may become available. As described above, the nucleic acid can be conveniently formulated with cationic lipids. In addition, glycolipids, fusogenic liposomes, peptides, and compounds collectively referred to as protective, interactive, non-condensing (PINC) can also be complexed with the purified plasmid DNA to affect variables such as stability, intramuscular dispersion, or transport to specific organs or cell types.
[0286] Also disclosed is a method of making a vaccine, comprising performing the steps of the method disclosed herein; and producing a vaccine comprising a plurality of antigens or a subset of a plurality of antigens.
[0287] Antigens disclosed herein can be manufactured using methods known in the art. For example, methods for producing antigens disclosed herein or vectors (e.g., vectors comprising at least one sequence encoding one or more antigens) may include culturing host cells under conditions suitable for expressing the antigen or vector, wherein the host cell comprises at least one polynucleotide encoding the antigen or vector, and purifying the antigen or vector. Standard purification methods include chromatographic techniques, electrophoresis, immunology, precipitation, dialysis, filtration, concentration, and chromatofocusing techniques.
[0288] The host cell may include Chinese hamster ovary (CHO) cells, NS0 cells, yeast or HEK293 cells. The host cell may be transformed with one or more polynucleotides, the one or more polynucleotides comprising at least one nucleic acid sequence encoding an antigen or a vector disclosed herein, optionally wherein the isolated polynucleotide further comprises a promoter sequence operably connected to at least one nucleic acid sequence encoding an antigen or a vector. In certain embodiments, the isolated polynucleotide may be a cDNA.
[0289] antigen
[0290] Antigens may include nucleotides or polypeptides. For example, an antigen may be an RNA sequence encoding a polypeptide sequence. Therefore, antigens useful in vaccines may include nucleotide sequences or polypeptide sequences. Antigens that can be used for cancer vaccines are described in International Patent Application Publication WO / 2019 / 226941, which is incorporated herein by reference in its entirety for all purposes.
[0291] Disclosed herein are isolated peptides comprising tumor-specific mutations identified by the methods disclosed herein, peptides comprising known tumor-specific mutations, and mutant polypeptides or fragments thereof identified by the methods disclosed herein. Neoantigenic peptides can be described in the context of their coding sequences, wherein the neoantigens include nucleotide sequences (e.g., DNA or RNA) encoding related polypeptide sequences.
[0292] Also disclosed herein are peptides derived from any polypeptide known or found to be expressed in tumor cells or cancerous tissues altered compared to normal cells or tissues, for example, any polypeptide known or found to be expressed abnormally in tumor cells or cancerous tissues compared to normal cells or tissues. For example, suitable polypeptides from which antigenic peptides can be derived can be found in the COSMIC database. COSMIC has collected comprehensive information on somatic mutations in human cancers. The peptides contain tumor-specific mutations.
[0293] One or more polypeptides encoded by an antigen nucleotide sequence may comprise at least one of the following: a binding affinity for MHC with an IC50 value of less than 1000 nM, a length of 8 - 15, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids for MHC class I peptides, the presence of a sequence motif promoting proteasomal cleavage within or near the peptide, and the presence of a sequence motif promoting TAP transport. For MHC class II peptides of length 6 - 30, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acids, the presence of a sequence motif promoting cleavage by extracellular or lysosomal proteases (e.g., cathepsin) or HLA binding catalyzed by HLA - DM within or near the peptide.
[0294] One or more antigens may be presented on the surface of a tumor.
[0295] One or more antigens may be immunogenic in a subject having a tumor, e.g., capable of eliciting a T - cell response or a B - cell response in the subject.
[0296] In the context of vaccine generation for a subject having a tumor, one or more antigens that induce an autoimmune response in the subject may be excluded from consideration.
[0297] The size of at least one antigen peptide molecule may include, but is not limited to, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 60, about 70, about 80, about 90, about 100, about 110, about 120 or more amino acid residues, and any range derivable therefrom. In a specific embodiment, the antigen peptide molecule is equal to or less than 50 amino acids.
[0298] Antigen peptides and polypeptides may: for MHC class I, have a length of 15 residues or less and typically consist of about 8 to about 11 residues, particularly 9 or 10 residues; for MHC class II, be 6 - 30 residues, including the end values.
[0299] If desired, longer peptides can be designed in several ways. In one case, when the presentation potential of the peptide on the HLA allele is predicted or known, the longer peptide can be composed of one of the following: (1) a single presented peptide extending 2-5 amino acids toward the N-terminus and C-terminus of each corresponding gene product; (2) a concatenation of some or all of the presented peptides with their respective extension sequences. In another case, when sequencing reveals the presence of longer (>10 residues) neoepitope sequences in the tumor (e.g., due to frameshift, read-through, or intron inclusion resulting in novel peptide sequences), the longer peptide is composed of: (3) an entire stretch of novel tumor-specific amino acids - thereby bypassing the need to select the strongest HLA-presenting shorter peptides by computational or in vitro testing. In both cases, the use of longer peptides allows for endogenous processing by patient cells and can result in more efficient antigen presentation and induction of T cell responses.
[0300] Antigenic peptides and polypeptides can be presented on HLA proteins. In some aspects, antigenic peptides and polypeptides are presented on HLA proteins with greater affinity than wild-type peptides. In some aspects, antigenic peptides or polypeptides can have an IC50 of at least less than 5000nM, at least less than 1000nM, at least less than 500nM, at least less than 250nM, at least less than 200nM, at least less than 150nM, at least less than 100nM, at least less than 50nM or less.
[0301] In some aspects, the antigenic peptides and polypeptides do not induce an autoimmune response and / or provoke immune tolerance when administered to a subject.
[0302] Compositions comprising at least two or more antigenic peptides are also provided. In some embodiments, the composition contains at least two different peptides. At least two different peptides can be derived from the same polypeptide. Different polypeptides mean that the peptides change in length, amino acid sequence or both. Peptides are derived from any polypeptide known or found to contain tumor-specific mutations, or peptides are derived from any polypeptide known or found to express changes in tumor cells or cancerous tissues compared to normal cells or tissues, such as any polypeptide known or found to express abnormalities in tumor cells or cancerous tissues compared to normal cells or tissues. For example, suitable polypeptides from which antigenic peptides can be derived can be found in the COSMIC database or the AACR Genomic Evidence Tumor Information Exchange (GENIE) database. COSMIC brings together comprehensive information on somatic mutations in human cancers. AACR GENIE gathers clinical-grade cancer genome data and the clinical outcomes of tens of thousands of cancer patients and links them. The peptides contain tumor-specific mutations. In some aspects, tumor-specific mutations are driver mutations for specific cancer types.
[0303] Antigenic peptides and polypeptides having desired activity or properties can be modified to provide certain desired attributes, e.g., improved pharmacological characteristics, while increasing or at least maintaining substantially all of the biological activity of the unmodified peptide to bind to the desired MHC molecules and activate appropriate T cells. For example, antigenic peptides and polypeptides can be subjected to various changes, such as conservative or non-conservative substitutions, where such changes may provide certain advantages in their use, such as improved MHC binding, stability, or presentation. Conservative substitution means replacing one biologically and / or chemically similar amino acid residue with another amino acid residue, e.g., replacing one hydrophobic residue with another hydrophobic residue, or replacing one polar residue with another polar residue. Substitutions include combinations such as Gly, Ala; Val, Ile, Leu, Met; Asp, Glu; Asn, Gln; Ser, Thr; Lys, Arg; and Phe, Tyr. The effect of single amino acid substitutions can also be probed using D-amino acids. Such modifications can be carried out using well-known peptide synthesis procedures, e.g., as described in Merrifield, Science 232:341-347 (1986), Barany & Merrifield, The Peptides, Gross & Meienhofer eds. (NY, Academic Press), pp. 1-284 (1979); and Stewart & Young, Solid Phase Peptide Synthesis, (Rockford, Ill., Pierce), 2nd ed. (1984).
[0304] Modification of peptides and polypeptides with various amino acid mimics or unnatural amino acids can be particularly useful for increasing the in vivo stability of the peptides and polypeptides. Stability can be determined in a variety of ways. For example, peptidases and various biological media, such as human plasma and serum, have been used to test stability. See, e.g., Verhoef et al., Eur. J. Drug Metab Pharmacokin. 11:291-302 (1986). The half-life of a peptide can conveniently be determined using a 25% human serum (v / v) assay. The protocol is generally as follows. Pooled human serum (type AB, non-heat inactivated) is defatted by centrifugation prior to use. The serum is then diluted to 25% with RPMI tissue culture medium and used to test peptide stability. At predetermined time intervals, a small aliquot of the reaction solution is removed and added to 6% aqueous trichloroacetic acid or ethanol. The turbid reaction sample is cooled (4 °C) for 15 minutes and then spun to pellet the precipitated serum proteins. The presence of the peptide is then determined by reverse-phase HPLC using stability-specific chromatographic conditions.
[0305] Peptides and polypeptides can be modified to provide desired properties other than improved serum half-life. For example, the ability of a peptide to induce CTL activity can be enhanced by linking it to a sequence containing at least one epitope capable of inducing a T helper cell response. Immunogenic peptide / T helper cell conjugates can be linked through a spacer molecule. The spacer generally consists of relatively small neutral molecules that are substantially uncharged under physiological conditions such as amino acids or amino acid mimics. The spacer is generally selected from, for example, Ala, Gly, or other neutral spacers of nonpolar amino acids or neutral polar amino acids. It should be understood that the optionally present spacer need not consist of the same residues and can thus be a hetero-oligomer or a homo-oligomer. When present, the spacer will generally be at least one or two residues, more usually three to six residues. Alternatively, the peptide can be linked to the T helper peptide without a spacer.
[0306] The antigenic peptide can be linked to the T helper peptide directly or through a spacer at the amino or carboxyl terminus of the peptide. The amino terminus of the antigenic peptide or the T helper peptide can be acylated. Exemplary T helper peptides include tetanus toxoid 830-843, influenza 307-319, malaria circumsporozoite 382-398 and 378-389.
[0307] A protein or peptide can be prepared by any technique known to those of skill in the art, including expressing a protein, polypeptide, or peptide by standard molecular biology techniques, isolating a protein or peptide from a natural source, or chemically synthesizing a protein or peptide. Nucleotide and protein, polypeptide, and peptide sequences corresponding to various genes have been previously disclosed and can be found in computerized databases known to those of ordinary skill in the art. One such database is the Genbank and GenPept databases of the National Center for Biotechnology Information located on the website of the National Institutes of Health. The coding regions of known genes can be amplified and / or expressed using techniques disclosed herein or known to those of ordinary skill in the art. Alternatively, various commercial preparations of proteins, polypeptides, and peptides are known to those of skill in the art.
[0308] On the other hand, the antigen includes nucleic acids (such as polynucleotides) encoding antigenic peptides or portions thereof. The polynucleotide can be, for example, DNA, cDNA, PNA, CNA, RNA (such as mRNA), single-stranded and / or double-stranded, or polynucleotides in natural or stabilized forms, such as polynucleotides having a phosphorothioate backbone, or combinations thereof, and it may or may not contain introns. On yet another hand, an expression vector capable of expressing a polypeptide or a portion thereof is provided. Expression vectors for different cell types are well known in the art and can be selected without undue experimentation. Generally, DNA is inserted into an expression vector, such as a plasmid, in an appropriate orientation and in the correct expression reading frame. If necessary, the DNA can be ligated to appropriate transcriptional and translational regulatory control nucleotide sequences recognizable by the desired host, although such controls are generally available in the expression vector. The vector is then introduced into the host by standard techniques. Guidance can be found, for example, in Sambrook et al. (1989) Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, N.Y.
[0309] Examples
[0310] The following are examples for carrying out specific embodiments of the present invention. The examples are provided for illustrative purposes only and are not intended to limit the scope of the present invention in any way. Efforts have been made to ensure the accuracy of the values used (such as amounts, temperatures, etc.), but of course some experimental errors and deviations should be taken into account.
[0311] Unless otherwise indicated, the practice of the present invention will employ conventional protein chemistry, biochemistry, recombinant DNA techniques, and pharmacological methods within the skill of the art. Such techniques are well explained in the literature. See, for example, T.E. Creighton, Proteins: Structures and Molecular Properties (W.H. Freeman and Company, 1993); A.L. Lehninger, Biochemistry (Worth Publishers, Inc., this edition newly added); Sambrook et al., Molecular Cloning: A Laboratory Manual (2nd Edition, 1989); Methods In Enzymology (edited by S. Colowick and N. Kaplan, Academic Press, Inc.); Reming’on's Pharmaceutical Sciences, 18th Edition (Easton, Pennsylvania: Mack Publishing Company, 1990); Carey and Sundberg Advanced Organic Chemistry 3rd Edition. (Plenum Press) Volumes A and B (1992).
[0312] This example outlines a cell-free DNA (cfDNA) assay for monitoring mutation frequency. In addition, the provided protocol (see below) is used to process and analyze the monitoring data of cfDNA mutation frequency in the treatment from patient plasma. Notably, for GRANITE patients, more than 200 mutations were monitored, representing all or most of the high-quality mutations associated with the tumor exome of each patient. The results indicate that the method provides a reliable method for monitoring mutation frequency.
[0313] Example 1 - Cell-Free DNA Monitoring
[0314] Method
[0315] The following is a protocol describing the method of the cell-free DNA monitoring assay.
[0316] Plasma Sample Collection
[0317] At regular, scheduled clinical visits (separated by approximately 1 month) that coincide with the dosing times, whole blood is drawn from the patients. The whole blood is collected in 10 mL Streck cell-free DNA BCT tubes (Streck; La Vista, NE, USA), and spun at 1600 x g for 10 minutes at ambient temperature to separate the plasma layer, buffy coat, and red blood cells. The plasma layer is removed and spun again at 5000 x g for 10 minutes to remove any residual cellular material. The supernatant is collected and stored at -80 °C until extraction.
[0318] cfDNA Extraction and Quantification
[0319] After thawing the separated plasma at ambient temperature, the plasma is spun at 5,000 x g for 5 minutes to remove any cold precipitates that formed during storage. cfDNA is extracted using the Apostle MiniMax cfDNA Isolation Kit (Beckman Coulter; Indianapolis, IN). The extracted cfDNA is quantified using the Qubit 1x High Sensitivity dsDNA Assay on a Qubit Fluorometer 4.0 (Thermo Fisher Scientific). For selected samples, 1 uL will be used to visualize the samples on an Agilent TapeStation using the HSD1000 kit.
[0320] gDNA Isolation
[0321] For genomic DNA from each sample, 50,000 PMBCs are isolated and extracted using the Qiagen Tissue AllPrep kit. For RNAlater samples, genomic DNA is isolated from the tissue stored in RNAlater using the Qiagen DNA / RNA Mini AllPrep kit.
[0322] Library Preparation and Hybridization Capture of the Duplex Library
[0323] The library was prepared with up to 20 ng of cfDNA using the KAPA Hyper Prep Kit (KAPA Biosystems; Wilmington, MA) according to the manufacturer's instructions. For libraries from gDNA, 30 ng of gDNA was first fragmented using the NEBNext Ultra II FS DNA Module (NEB, Ipswich, MA) under the following conditions: 25 minutes at 37 °C followed by 30 minutes at 65 °C. After end repair, adapter ligation was performed for 30 minutes with a double-stranded adapter library containing 5-mer non-random unique molecular identifiers (IDT, Coralville, Iowa). Whole library amplification was performed using KAPA HiFi HotStart ReadyMix and NEBNext Multiple Oligos for Illumina (96 unique dual-index primer pairs).
[0324] After preparing the duplex library, the selected regions of interest were hybridized overnight with 750 ng of the duplex library using custom-designed xGen locked probes and the xGen hybridization and wash kit according to the manufacturer's instructions (IDT). The final library was quantified and normalized using the Qubit 1x High Sensitivity dsDNA assay.
[0325] Sequencing and Analysis
[0326] The normalized samples were pooled in equimolar amounts and sequenced on the NovaSeq using 2x151 bp and 8 bp index reads.
[0327] Figure 1 and Figure 2 The figures and Table 1 show the specifications of the method for isolating and monitoring mutant alleles in ctDNA of individual patients.
[0328] Table 1 - Assay Specifications for ctDNA Monitoring
[0329] cfDNA input 20 ng (as low as 5 ng - limited variant sensitivity) Group footprint 296 Kb Bait number 5,460 Variants per patient 90 - 461 (average: 283) Duplex depth 2000 - fold - 4000 - fold duplex consensus
[0330] Tumor-specific DNA variant alleles were identified from biopsied tumor tissue in patients and used to create baits to isolate tumor-specific DNA from all circulating cell-free DNA (cfDNA) in patient blood samples. The isolated ctDNA was sequenced in duplicate and analyzed for duplex concordance. Sequencing multiple blood draws during treatment enables less invasive monitoring of patient response.
[0331] After sequencing is completed and FASTQ files are generated, UMIs are extracted and assigned to each read tag before alignment with BWA-MEM (Durbin et al., Bioinformatics, 2010). Reads are grouped by UMIs on the aligned bam file using the fgbio toolkit (Fulcrum Genomics), and duplex consensus reads are called. Before data analysis, the unaligned bam file from fgbio is aligned using BWA-MEM, and then freebayes (Marth et al., arXiv2012) is used to obtain the variant allele frequency (VAF) of each somatic variant of interest.
[0332] Cancer vaccine administration
[0333] An open-label, multi-center, multi-dose phase 1 / 2 study was conducted to evaluate the dose, safety and tolerability, immunogenicity and early clinical activity of a heterologous prime / boost vaccination strategy. Two vaccine regimens, GRANITE and SLATE, were evaluated. The clinical trial design is described in international patent application publication WO / 2019 / 226941, which is hereby incorporated by reference in its entirety for all purposes.
[0334] In patients with advanced cancer, a combination of an individualized neoantigen cancer vaccine (“GRANITE”) and immune checkpoint blockade was administered. The GRANITE heterologous prime / boost vaccine regimen included (1) ChAdV[GRT-C901] used as the prime vaccination and (2) SAM[GRT-R902] formulated in LNP and used to boost vaccination after GRT-C901. The ChAdV vector is based on a modified ChAdV68 sequence. The SAM vector is based on an RNA alphavirus backbone. GRT-C901 and GRT-R902 express the same 20 individualized neoantigens as well as two universal CD4 T cell epitopes (PADRE and tetanus toxoid). Tumors were used for whole exome and transcriptome sequencing to detect somatic mutations, and blood was used for HLA typing and detection / deduction of germline exome variants to generate individualized neoantigen cassettes for 10 subjects (patients 1-10, herein referred to as patients G1-G10) using the EDGE algorithm.
[0335] In patients with advanced cancer, a combination of a shared neoantigen cancer vaccine (“SLATE”) and immune checkpoint blockade is administered. The SLATE heterologous prime / boost vaccine regimen consists of (1) ChAdV[GRT-C903] used as the prime vaccination and (2) SAM[GRT-R904] formulated in LNP and used to boost vaccination after GRT-C903. GRT-C903 and GRT-R904 express the same 20 shared neoantigens, which are derived from a panel of specific oncogenic mutations as well as two universal CD4 T cell epitopes (PADRE and tetanus toxoid). For subject inclusion, tumors were used for whole exome and transcriptome sequencing to detect somatic mutations, and blood was used for HLA typing. Enrolled SLATE subjects were identified as having HLA A02:01 and a KRAS mutation G12C predicted to be presented by HLA A02:01 (patients S1, S2, and S3), HLA A01:01 and a KRAS mutation Q61H predicted to be presented by HLA A01:01 (patients S4 and S7), or HLA A03:01 or A11:01 and a KRAS mutation G12V predicted to be presented by HLA A03:01 or A11:01 (A03:01 for patient S9; A11:01 for patients S11 and S15).
[0336] Both treatment studies (i.e., the GRANITE and SLATE vaccine regimens) administered the vaccine via bilateral (e.g., in each deltoid) IM injection, in combination with immune checkpoint blockade (specifically SC ipilimumab and IV nivolumab). These studies followed two consecutive phases.
[0337] GRT-C901 and GRT-C903 are replication-deficient, E1- and E3-deleted adenovirus vectors based on chimpanzee adenovirus 68. The vectors contain expression cassettes encoding 20 neoantigens as well as two universal CD4 T cell epitopes (PADRE and tetanus toxoid). GRT-C901 and GRT-C903 were formulated as a solution of 5×10 11 vp / mL and IM injected at 1.0 mL at each of 2 bilateral vaccine injection sites in opposite deltoids. The GRT-C901 and GRT-C903 vectors differ only in the neoantigens encoded within the cassette.
[0338] GRT-R902 and GRT-R904 are SAM vectors derived from alphaviruses. The GRT-R902 and GRT-R904 vectors encode the 5’ and 3’ RNA sequences required for viral protein and RNA amplification, but do not encode structural proteins. The SAM vectors are formulated in LNPs, which include four lipids: an ionizable amino lipid, phosphatidylcholine, cholesterol, and a PEG-based coating lipid, to encapsulate the SAM and form the LNP. For each patient, the GRT-R902 vector contains the same neoantigen expression cassette used in GRT-C901. The GRT-R904 vector contains the same neoantigen expression cassette used in GRT-C903. The GRT-R902 and GRT-R904 are formulated as a 1 mg / mL solution and given IM injections at each of two bilateral vaccine injection sites in the opposing deltoids (preferably the deltoids, but the gluteal [dorsal or abdominal muscles] or rectus femoris on each side can be used). The booster vaccine injection sites are as close as possible to the initial vaccine injection sites. The injection volume is based on the dose to be administered. The dose level specifically refers to the amount of the SAM vector, i.e., it does not refer to other components such as the LNP. The LNP:SAM ratio is approximately 24:1. Thus, for each corresponding GRT-R902 / GRT-R904 dose level, the doses of LNP are 720 μg, 2400 μg, and 7200 μg, respectively (see below).
[0339] Ipilimumab is a human monoclonal IgG1 antibody that binds to cytotoxic T lymphocyte-associated antigen 4 (CTLA-4). Ipilimumab is formulated as a 5 mg / mL solution and given SC injections proximal (within approximately 2 cm) to each bilateral vaccine injection site. Ipilimumab is administered at a dose of 30 mg of antibody in four 1.5 mL (7.5 mg) injections proximal to the draining LNs at each bilateral vaccine injection site (i.e., at each bilateral on each deltoid, rectus abdominis, dorsal tongue muscle, or rectus femoris [preferably the deltoids, but depending on the clinical site and patient preference] 1.5 mL below and 1.5 mL above the vaccine injection site).
[0340] Nivolumab is a human monoclonal IgG4 antibody that blocks the interaction of PD-1 with its ligands PD-L1 and PD-L2. Nivolumab is formulated as a 10 mg / mL solution and administered by IV infusion (480 mg) at the dose specified in the protocol through a low-protein-binding in-line filter with a pore size of 0.2 to 1.2 microns. Nivolumab is not administered by IV bolus or push injection. After nivolumab infusion, the tubing is immediately flushed with diluent to clear it. Nivolumab is administered on the same day, with or without ipilimumab, after each vaccination (i.e., each of GRT-C901, GRT-R902, GRT-C903, or GRT-R904). The dose and route of nivolumab are based on the dose and route approved by the Food and Drug Administration.
[0341] Results
[0342] Monitoring of ctDNA in cfDNA-containing samples was used to track patients' responses to therapy. Specifically, patients receiving neoantigen-based cancer vaccines (GRANITE and SLATE) were monitored during treatment. Sequencing of cancer exome-related mutations was performed at high target coverage and high read depth.
[0343] CtDNA from two independent patients (G1 and G2) receiving GRANITE therapy was monitored to examine responses. Details of all ctDNA isolations from each patient are given in Table 2.
[0344] Table 2 - Details of ctDNA Isolation from Patients Receiving GRANITE Therapy
[0345]
[0346]
[0347] Figure 3A and Figure 3B Shows the duplex read coverage of patient G1 during treatment. In cfDNA samples, the average sequencing read depth of targets (average target duplex read coverage [x]) ranged from 2817x - 5017x, where >87% of the targets (monitoring over 330 variants) had ≥2000-fold duplex reads, and >68% of the targets had ≥4000-fold duplex reads (excluding D5D1 and D6D1). The sequencing profiles showed high target coverage at high read depth.
[0348] Monitoring of the mutant allele frequency of cfDNA in GRANITE patient G1 during treatment. As Figure 3CAs shown in Table 3, 117 mutant alleles out of more than 330 subject-specific and tumor-specific variants were detected in the ctDNA of G1. Figure 4A - 4C Also shown is the frequency of mutant alleles in ctDNA isolated from G1 during the disease process. Figure 4A Shown is the mutant allele frequency of 11 out of 20 mutations detected at baseline. Figure 4B Shown is the average mutant allele frequency. Figure 4C Shown is the percentage change in the average mutant allele frequency. After the first and second doses, the tumor-specific variant allele frequency (VAF) (also known as mutant allele frequency (MAF)) initially peaked and then decreased after the third dose, indicating a response to treatment, and then moderately increased within the first 168 days, associated with stable disease. Then, the mutant allele frequency increased significantly after 168 days (week 24), associated with disease progression. Thus, monitoring the mutant allele frequency in cfDNA serves as an effective non-invasive alternative for monitoring disease status, including assessing disease progression and the efficacy of treatment regimens.
[0349]
[0350]
[0351]
[0352]
[0353]
[0354]
[0355] In Figure 3D and Figure 3E Shown is the duplex read coverage of patient G2 during treatment. After consistent results in cfDNA samples, the average read coverage of the target ranged from 3877x - 4534x, where >93% of the targets (more than 240 variants were monitored) had a duplex read of ≥2000-fold, and >76% of the targets had a duplex read of ≥3000-fold. The sequencing profiles showed high target coverage at high read depths.
[0356] Monitoring the mutant allele frequency of cfDNA of GRANITE patient G2 during treatment. As Figure 3FAs shown, during the treatment regimen of patient G2, no ctDNA above the minimum call threshold was detected, which was associated with an extended disease-free period (no signs of disease at any time point in the post-operative study). Thus, monitoring the mutant allele frequency in cfDNA serves as an effective non-invasive alternative for monitoring disease, including assessing the presence and disease burden of the disease.
[0357] Table 4 - VAF values of ctDNA from patient G2
[0358]
[0359] The mutant allele frequency of cfDNA was monitored during the treatment of GRANITE patients G3 and G8. Figure 5A - 5B Tracking of multiple variant alleles in the ctDNA of each patient is shown respectively. After an initial peak around one month after initial treatment, the VAF of both patients showed a stable decrease. In both patients, this decrease was associated with an overall reduction in tumor volume. Patient G3 showed a maximum 6-fold decrease in VAF (mean VAF of 0.69% at week 4 vs. mean VAF of 0.12% at week 20), with all 20 monitored variants detected. The cfDNA profile of patient G3 was associated with disease progression at week 8, then stabilized at week 16, and then minimally progressed at week 24 (T cell decline). Patient G8 showed a continuous decrease in mutant allele frequency, including loss of detection of some variants (16 out of 20 variants detected), associated with stable disease. Thus, monitoring the mutant allele frequency in cfDNA serves as an effective non-invasive alternative for monitoring disease status, including assessing disease progression.
[0360] During the treatment of SLATE patient S1, the mutant allele frequency in cfDNA was also monitored. Details of all ctDNA isolations are detailed in Table 5.
[0361] Table 5 - Details of ctDNA isolation from patient S1 receiving SLATE therapy
[0362] Sample Yield ng / μL Total ng Plasma mL ng / mL Day 1 Dose 1 1.43 143 8.00 17.88 Day 1 Dose 2 0.988 98.8 8.00 12.35 Day 1 Dose 3 1.06 106 8.00 13.25
[0363] Figure 6A and Figure 6B shows the duplex read coverage of patient S1 during treatment. After agreement in cfDNA samples, the average read coverage of the target ranged from 2728x - 3660x, where >98% of the targets had ≥1000-fold duplex reads and >78% of the targets had ≥2000-fold duplex reads.
[0364] The mutant allele frequency of cfDNA was monitored during treatment. As Figure 6CAs shown, a steady increase in ctDNA tumor content was observed, indicating progressive tumors. All ctDNA analysis results for Patient S1 are given in Table 4. Thus, monitoring the mutant allele frequency in cfDNA serves as an effective non-invasive alternative for monitoring disease status, including assessing disease progression.
[0365] The tumor of SLATE Patient S2 was determined to have a KRAS G12C mutation and was monitored using variant-specific tracking of the KRAS G12C mutation. As Figure 7 shown, an overall decrease in the VAF of the KRAS mutant was observed, which was associated with a 20% reduction in tumor volume at week 8. Thus, monitoring the mutant allele frequency in cfDNA serves as an effective non-invasive alternative for monitoring disease status, including assessing disease progression.
[0366] The results indicate that the mutant allele frequency in cfDNA can be monitored during the treatment of a large number of tumor-specific and subject-specific mutations. The results also indicate that monitoring the mutant allele frequency in cfDNA serves as an effective non-invasive alternative for monitoring disease status, including assessing disease progression, assessing the presence and disease burden of the disease, and the efficacy of treatment regimens.
[0367] Table 6 - ctDNA analysis results of patients receiving SLATE therapy
[0368]
[0369]
[0370] Example 2 - Cell-free DNA monitoring using a combination group in GRANITE and SLATE vaccine programs
[0371] Circulating tumor DNA (ctDNA) is an emerging minimally invasive diagnostic and prognostic biomarker for patients receiving immunotherapy. Example 2 describes the assessment of ctDNA dynamics and tumor evolution over time by a combined approach determined by tumor-informed and tumor-uninformed ctDNA monitoring.
[0372] The method for describing the cell-free DNA monitoring assay was generally described in Example 1 and is briefly described as follows: Patient biopsies were collected for GRANITE vaccine generation screening. Individualized vaccines were manufactured for patients with sufficient neoantigens. After the patient was enrolled in the study, baseline biopsies were collected. Then the patient was administered the vaccine for treatment ("GRANITE" program). ( Figure 8 ). In contrast, "off-the-shelf" common neoantigen vaccines were administered to certain subjects ("SLATE" program).
[0373] A combinatorial probe set for enriching targets was designed, which includes tumor-informed probes and tumor-uninformed probes. Neoantigens were predicted from whole exome sequencing (WES) of tumor DNA and whole transcriptome sequencing of tumor RNA. For all coding mutations detected in the whole exome sequencing (WES) of archival tissue, a tumor-informed patient-specific set (median: 123; range: 67 - 402) was designed, including patients selected for the Granularity of Recurrence Assessment in Neoplasia by Individualized Testing (GRANITE) program. A tumor-uninformed set (also referred to as the "universal" set) was designed and incorporated into the patient set to monitor tumor hotspots of recurrent mutations and genes involved in immunotherapy resistance. Specifically, the set monitors genes and mutations generally considered to be oncogenic (e.g., "driver" mutations considered to contribute to cancer and commonly recognized gain-of-function mutations), tumor suppressor genes (e.g., genes generally considered to monitor and / or control tumor-related properties such as cell division, where mutations can interfere with the control of such properties and are generally considered loss-of-function mutations), interferon-γ signaling pathway genes (including JAK / STAT signaling pathway genes), antigen processing pathway genes (including monitoring loss of HLA heterozygosity), and additional mutations generally associated with cancer but possibly not annotated (e.g., mutations not yet annotated as oncogenes or tumor suppressor genes). The probes in the set were designed to monitor the entire exome coding regions of the selected genes or to monitor specific regions or mutations ("hotspots") within the genes.
[0374] Monthly cell-free DNA (cfDNA) samples (mean 7; range: 1 - 18) were collected during treatment. Libraries with duplex unique molecular identifiers (UMIs) were prepared from cfDNA, matched normal DNA, and biopsy DNA and captured using a combination of the individualized set, universal set, or WES (9). Shotgun libraries of cfDNA, biopsy DNA, or gDNA from whole blood or PBMCs were prepared with duplex UMIs. Duplex sequencing reduces noise by requiring variants to be observed on both strands of the duplex molecule. The enriched duplex libraries were sequenced to an average target depth >65,000-fold before consensus duplicate removal.
[0375] Patient-specific variants and universal regions were enriched for de novo variant calling. The patient-specific regions were tumor-informed, while the universal set was tumor-uninformed. Multiple patient-specific collections were combined to create a superset containing probes for 6 - 9 patients. The universal set captured a common set of targets in all patient samples (e.g., see Table 7). The combination of sets was not only patient-specific but also provided the flexibility to observe variants not present in the original tumor.
[0376] A summary of the GRANITE cfDNA monitoring assay is also provided in Table 6.
[0377] Table 6 GRANITE cfDNA Monitoring Assay
[0378]
[0379]
[0380] Table 7 provides a summary of cancer-related genes and hotspots monitored in the general tumor-naïve group. The group also includes approximately 290 probes that are configured to capture positions with common polymorphisms (SNPs) in the human population for fingerprinting to uniquely identify a subject, e.g., in the case of multiplexing multiple subject sequences.
[0381] Table 7 - Exemplary Genes in the General Group
[0382]
[0383] Figure 10A and 10B Shows the variant coverage of manufacturing variants in the individualized GRANITE group. Figure 10A Shows the number of potential variants covered by various commercially available NGS groups compared to the coverage of the designed GRANITE group and whole exome sequencing, including individual tumor-naïve and tumor-informed. Figure 10B Shows the percentage of WES variants potentially covered by various NGS groups. Neoantigen coverage is determined by variants present in the manufacturing biopsies of GRANITE patients receiving treatment. In the absence of an archived biopsy, SLATE patients have only 1 mutation intentionally monitored (i.e., the tumor-informed mutation determined by sequencing that makes the SLATE patient eligible for the SLATE vaccine). Using the static group, an average of 5 - 10 variants were monitored in GRANITE patients. However, with the use of the tumor-informed group, this range increases to 16 - 50 variants. Regardless of the group type, currently available assays do not overlap with most tumor-specific neoantigens in the GRANITE vaccine.
[0384] Figure 11 Shows the blood collection protocol for SLATE and GRANITE vaccine patients. Whole blood is collected at the time of vaccine administration and centrifuged to collect plasma serum and buffy coat. Most cfDNA comes from the turnover of white blood cells. The collection protocol includes collecting buffy coat and / or whole blood for sequencing of patient-matched normal gDNA. Buffy coat collection allows for the exclusion of clonal hematopoiesis of matched normal cfDNA via CHIP. Patients with advanced disease have a higher cfDNA concentration. The median yield of whole blood plasma of patient samples is 15 ng / ml. This is used to calculate hGE at 3000 he / 10 ng. The number of molecules limits the sensitivity. Figure 11Shows cfDNA yields in ng / ml plasma collected from GEA, CRC, NSCLC, or other tumor tissues or healthy donors in GRANITE or SLATE patients.
[0385] Figure 12 Shows that GRANITE patient assays monitored an average of approximately 140 variants per patient at high sequencing depths for variant calling at >1000-fold duplex consensus coverage. Patient samples were sequenced to obtain an average target depth of approximately 100,000-fold (paired-end) depth, which decreased to approximately 3900-fold after duplex consensus.
[0386] Table 8 provides a summary of patient samples, including tissue type, number of cfDNA samples, biopsies, number of targeted variants, input cfDNA amount (ng), and average VAF range for longitudinal samples.
[0387] Table 8
[0388]
[0389]
[0390] *Samples not collected
[0391] Figure 13A and 13B Shows that most neoantigens were found in cfDNA and patient biopsies using the GRANITE assay. After multiple lines of therapy, most neoantigens were found in the patient's cfDNA or tumor biopsies, indicating that many neoantigens are trunk variants suitable for targeting with personalized neoantigen vaccines. Figure 13A Shows cassette mutations observed in ctDNA and biopsies of the indicated patients (* indicates patients for whom biopsies were not available or in whom tumor content was too low to detect variants in the assay). When comparing variants in cfDNA and corresponding biopsies using the GRANITE assay, significant overlap was found, particularly the ability to call variants at lower frequencies in high-quality (RNALater or fresh frozen) biopsies ( Figure 13B ).
[0392] The common set based on tumor-naive regions captures de novo variants after GRANITE vaccination. Figure 14AShows the presence of de novo variants in cfDNA samples from the indicated patients and tumor tissue types (GEA, CRC, or NSCLC). Importantly, many patients have additional variants in their cfDNA that were not present in the original biopsy. New variants often occur at sites where another patient has a targeted variant. Using matched normal gDNA from whole blood or PMBC, CHIP mutations were identified and excluded as somatic tumor variants ( Figure 14B ). For example, patient G08 had two NLRC5 mutations and two TAP1 mutations, one of the NLRC5 mutations tracked with the average VAF of all variants, and the two TAP1 mutations appeared nearly a year after therapy ( Figure 14C ). Figure 14D Shows a summary of additional analysis of variants observed in cfDNA and found outside of patient-specific variants, indicating that 75% of the patients evaluated had newly detected variants in their cfDNA, including driver variants (KRAS and BRAF) and resistance variants (TAP1). Figure 14E Shows that in patient G09, multiple complex KRAS variants were detected that were not detected in the archived biopsy or follow-up biopsies. The KRAS Q61H variant followed the same trajectory as the archived tumor variant, while three KRAS G12 variants occurred at lower VAFs with different kinetics, illustrating how multiple KRAS G12 hotspot variants can be captured with cfDNA to observe metastatic disease. Thus, the tumor-naïve group effectively monitors for additional mutations, including those that may be involved in immune evasion.
[0393] Duplex sequencing enables higher sensitivity of biopsies. Targeted variants also provide insights into tumor heterogeneity. All patient-specific variants captured in the biopsy WES were also captured using patient-specific assays (100% concordance), see Figure 15 and Table 9. A certain number of de novo variants were observed in both assays, including new variants that could not be captured by WES. Most de novo variants would not be captured if unbiased WES was not performed. Despite the differences, the GRANITE monitoring assay captured 219 / 353 (62%) of the variant sets observed between the two methods.
[0394] Table 9: Comparison of patient-specific variants captured in patient-specific cfDNA assays and WES
[0395] GRANITE WES Both Total variants Cartridge 1 0 19 20 Targeted coding 37 0 148 185 Synonymous 0 45 0 45 New 7 89 7 103 Total 45 134 174 353
[0396] Using the monitoring assay on patient biopsies, variants in the baseline biopsy were at low frequencies in the archived biopsy ( Figure 16A)。The biopsy variants during treatment are more representative of the variants present in the archival biopsies. Although all biopsies were from the primary site, only 12 / 135 of the targeted variants were common among the three, indicating tumor heterogeneity( Figure 16B )。In cfDNA, 115 of the 135 variants were observed. Among the variants not observed, all but one were unique to the baseline biopsy( Figure 16B )。
[0397] Exploratory analysis identified variants for brain metastases in longitudinal cfDNA samples. A subset of variants was found to occur only near the end of treatment in the individualized assays, some of which were found only in the final biopsy (brain met). Figure 17A Shows the dynamics of cfDNA over time in patient G01. Figure 17B Shows targeted low-frequency variants of the indicated variants (SSH3, GRIA4, ZNF541, TMEM217, ZNF697, AHNAK2, SCHIP1, and CNR1) in ctDNA. Figure 17C Shows the targeted variants over time of the WES of ctDNA of patient G01. Figure 17D Shows the brain met biopsy variants of the WES of ctDNA of patient G01 over time. The brain met biopsy contained variants not found in the early treatment biopsies. Using WES of cfDNA, the kinetics of 32 variants could be tracked. Figure 17D The dashed lines in [Figure number] indicate variants also found in the patient-specific assays.
[0398] Thus, in 24 patients, a median of 92.5% (range: 45 - 100%) of neoantigens and a median of 84% (range: 24 - 99%) of all targeted variants were found in cfDNA. Signs of heterogeneity were found in both cfDNA and biopsies, and duplex sequencing improved the detection of target variants in biopsies with low tumor content. Combining the tumor-informed and tumor-naive groups, de novo variants were found in the cfDNA of 19 patients. De novo variants were found in regions targeted by the tumor-naive group or regions where another patient had targeted variants. For example, evidence of acquired immune escape was observed in a patient with colorectal cancer through a biallelic loss-of-function mutation in TAP1. Using WES of longitudinal cfDNA in patients with gastroesophageal adenocarcinoma, changes in copy number (including loss of HLA heterozygosity) and newly emerging subclonal variants were confirmed between treatment biopsies and cfDNA.
[0399] Based on the results in Example 2, for patients treated with neoantigen vaccines, longitudinal monitoring of cfDNA provided early insights into patients who responded to treatment. Using a combination cohort approach of tumor-informed and tumor-uninformed monitoring, ctDNA kinetics shown by targeting numerous mutations also tracked tumor burden, evolution, and emerging resistance. Specifically: (1) Comprehensive tumor-informed ctDNA monitoring provided improved breadth and maintained sensitivity to monitor the longitudinal dynamics of tumor burden and neoantigens delivered by the vaccine in patients treated in an individualized cancer vaccine regimen (GRANITE); significant concordance was observed between plasma and tumor biopsy collections, and the ability of the ctDNA monitoring assay to detect and monitor variants in metastatic sites was demonstrated; and (3) By combining a rationally designed tumor-uninformed cohort, important tumor-intrinsic immune escape events were detected and monitored, thus providing insights into the mechanism of action of individualized neoantigen immunotherapy.
[0400] Example 3: Cell-free DNA Monitoring with Cohorts in SLATE Vaccination
[0401] Circulating tumor (ct) DNA sequencing and analysis methods
[0402] A panel of universal capture probes was designed to capture mutations targeted by an "off-the-shelf" common neoantigen vaccine cassette ("SLATE"), cancer hotspots, and SNPs for fingerprinting. Vaccine cassette design and manufacture were performed as previously described (Palmer et al. "Individualized, heterologous chimpanzee adenovirus and self-amplifying mRNA neoantigen vaccine for advanced metastatic solid tumors: phase 1 trial interim results." Nat Med. August 15, 2022; incorporated herein by reference for all purposes). Certain genes included probes designed to capture the entire coding region, such as probes for TP53, PTEN, ARID1A, and genes involved in the antigen presentation machinery (B2M, TAP1 / 2, and HLA-A, HLA-B, and HLA-C) were also designed to be captured. Figure 22 A general strategy for monitoring loss of heterozygosity of HLA genes on chromosome 6 is shown.
[0403] The probes were designed and synthesized by Integrated DNA Technologies (IDT). Genomic DNA fragments from whole blood or PMBCs of patient matches were fragmented prior to library preparation using the NEB FS module (NEB, Ipswich, MA). Using the KAPA HyperPrep (KAPA Biosystems, Wilmington, MA) kit, shotgun libraries of cfDNA (up to 30 ng) and fragmented, patient-matched genomic DNA (20 - 30 ng) were prepared using a custom library with duplex adapters containing unique molecular identifiers (UMIs) (IDT). The shotgun libraries were captured overnight using the IDT xGen hybridization and wash kit. The enriched libraries were sequenced on an Illumina NovaSeq to a minimum average raw depth of 65,000x. Briefly, UMIs were clipped from the raw sequencing reads prior to alignment with hg38 using BWA-MEM. Paired reads were grouped using fgbio by position and duplex identity. Consensus reads were generated using 3x duplexes (three supporting reads from each strand) and realigned to hg38. Variant calling was performed using FreeBayes and VarDictJava. The percent change in ctDNA was calculated as the change in the VAF of SLATE variants selected from the baseline sample.
[0404] Effective control of tumor growth was observed in a subset of patients treated with SLATEv1
[0405] The clinical activity of the SLATE vaccine regimen was evaluated using the RECIST v1.1 criteria, by the secondary endpoints of the phase 1 study, ORR, and PFS, as well as OS. Eight of 19 patients (42%) had a best overall response (BOR) of stable disease (SD), including 4 patients with NSCLC, 2 patients with PDA, 1 patient with CRC, and 1 patient with pancreatobiliary adenocarcinoma, and all other patients had a BOR of PD (Supplementary Table 2). The median PFS for all phase 1 patients was 1.9 months (95% confidence interval (CI) = [1.7, 3.9 months]). The median OS across all tumor types was 7.9 months (95% CI = [4.7, 10.9 months]) (data not shown). At the time when neoantigen-specific T cell responses were still being generated (11 / 15 within 2 months after the first vaccination and 4 / 15 within 4 months), 79% of patients (15 / 19) in the phase 1 treatment showed an OR of PD, and all patients progressed early in the course of treatment. Despite the rapid progression of many patients, the target lesions decreased in 2 / 4 patients with NSCLC and SD (whose tumors had previously progressed on ICB), indicating that vaccine-induced T cells lysed tumor cells.
[0406] Although CT scans can provide insights into the antitumor effects of cytotoxic therapies, the effects of immunotherapies that induce robust T cell responses may be misclassified due to T cell infiltration into tumors and subsequent antigen-induced expansion, which can lead to an increase in lesion size. Monitoring of ctDNA in the blood has been shown to be associated with clinical outcomes such as PFS and OS17-19, providing an alternative to CT scans for longitudinal assessment of antitumor effects in patients treated with immunotherapy. In fact, the latest evidence suggests that, compared with imaging, a decrease in ctDNA may be a more sensitive marker of early treatment efficacy of immunotherapy and is more associated with improved survival outcomes. Therefore, ctDNA levels corresponding to targeted neoantigens within the vaccine cassette were evaluated as an exploratory endpoint using tumor-informed probes, such as monitoring of neoepitopes encoded by the SLATE cassette that qualified subjects for the trial. Molecular responses (MR) were observed in 23% of evaluable patients (4 / 17), characterized by a ≥30% decrease in neoantigen-specific ctDNA compared with baseline levels, and two of these patients also showed SD despite prior progression on ICB, indicating effective immune control of vaccine-induced T cell-driven tumor growth in a subset of patients (data not shown).
[0407] In addition to the decrease in ctDNA levels of vaccine-targeted neoantigens, a decrease in additional variants in ctDNA that were not encoded by the vaccine was observed in all four patients with molecular responses (MR), such as by monitoring with additional tumor-informed probes targeting mutations predicted to be presented by the subject's HLA and / or additional mutations commonly associated with cancer (e.g., driver mutations). Figure 18A - 18E ) thus providing further evidence of effective tumor targeting by vaccine-induced T cells specific for one of the tumor neoantigens. Figure 18A - 18E ctDNA monitoring of tumor variants in SLATE patients is shown. The levels of ctDNA over time of vaccine-encoded tumor mutations (solid lines) and additionally detected somatic mutations (dashed lines) after the first vaccination for each patient are shown. ctDNA levels are reported as the percentage of variant allele frequency (VAF) of the total reads. Figure 18A The ctDNA %VAF of patient S2 is shown. Figure 18B The ctDNA %VAF of patient S5 is shown. Figure 18C The ctDNA %VAF of patient S10 is shown. Figure 18D The ctDNA %VAF of patient S13 is shown. Figure 18E A representative patient without MR is provided, showing loss of the B2M start codon. SD = stable disease, PD = progressive disease, indicating the best overall response.
[0408] Although the immune system can control tumor growth, tumors can evade immune control of targeted immune responses. One mechanism by which tumors have been shown to evade immune responses is by disrupting the antigen presentation pathway. The ctDNA panel includes a general tumor-agnostic panel that includes probes that monitor antigen presentation genes to assess whether patients treated with SLATEv1 exhibit antigen presentation defects at the time of progression. Two patients, one a molecular responder and the other a molecular non-responder, exhibited loss of heterozygosity (LOH) of relevant neoantigen-matched HLA alleles during treatment ( Figure 19A - 19B ), indicating that vaccine-induced T cells exerted immune pressure on the tumor, triggering this immune escape mechanism. Figure 19A Shows the fold change in HLA allele read fractions from a molecular responder (MR). Figure 19B Shows the fold change in HLA allele read fractions from a non-molecular responder (non-MR). One of the molecular responders, S13, showed some evidence of HLA LOH (p = 0.015) in the baseline sample, which was not evident in subsequent post-treatment samples and corresponded to a decrease in ctDNA. At all subsequent time points, evidence of HLA LOH (p < 0.01) was observed, which was associated with PD and indicated the growth of tumor cells resistant to T cell control. Another patient, S4, had a variant that resulted in the loss of the B2M start codon, which has been shown to reduce antigen presentation, was detected at low frequency at baseline, and increased in frequency during treatment with the vaccine regimen ( Figure 18E ). HLA LOH was also observed in HLA-A (the neoantigen-matched HLA allele) and HLA-C in samples collected 9 weeks after the first vaccination (last collection). The B2M mutation and complete LOH could explain the lack of clinical activity observed in this patient. Collectively, these data demonstrate preliminary signs of clinical activity in a subset of patients with advanced / metastatic solid tumors treated with the SLATEv1 neoantigen vaccine.
[0409] Example 4: Selection of Group Probe Targets
[0410] Figure 20 A figure is provided that outlines considerations for including subject-specific tumor-agnostic probes.
[0411] In addition, the common group targets in version 1 were further updated to target specific tissue types (e.g., CRC and NSCLC), resulting in common group version 2. The common group was updated by reviewing the literature, hotspots in the TCGA, COSMIC, and MyCancerGenome datasets, and included monitoring genes and mutations that are generally considered carcinogenic (e.g., "driver" mutations thought to contribute to cancer and commonly recognized gain-of-function mutations), tumor suppressor genes (e.g., genes generally considered to monitor and / or control tumor-related properties such as cell division, where mutations can interfere with the control of such properties and are generally considered loss-of-function mutations), interferon-γ signaling pathway genes, JAK / STAT signaling pathway genes, antigen processing pathway genes (including monitoring loss of HLA heterozygosity), and mutations that are generally associated with cancer but may not be annotated (e.g., mutations that have not been annotated as oncogenes or tumor suppressor genes). The initial results were prioritized based on prevalence, proximity, and oncogene / suppressor gene status, resulting in 133 prioritized regions. Examples of common group version 1 targets are provided in Table 10.
[0412] Table 10: Common group version 1
[0413]
[0414]
[0415] Approximately 290 SNPs were used for fingerprinting; 115 kb footprint Table 11 provides the updated common group version 2.
[0416] Table 11: Common group version 2 (focused on CRC / NSCLC)
[0417]
[0418]
[0419] Approximately 290 SNPs were used for fingerprinting; 135 kb footprint
[0420] Figure 21A The percentage of CRC and NSCLC samples covered by target probes in the common group is provided. Up to 86.4% of all samples with more than or equal to one mutation are covered by the common group. Up to 49.8% of all samples with more than or equal to two mutations are covered by the common group. The data is based on the analysis of 10,586 samples from cbioportal.org. Figure 21B A retrospective analysis of variants identified by common group version 1 (v1) or version 2 (v2) in patients from a previous study (GO-004) is shown. A data summary is provided in Table 12.
[0421] Table 12: Variants covered by general groups v1 and v2.
[0422]
[0423] Although the present invention has been specifically shown and described with reference to preferred embodiments and various alternative embodiments, those skilled in the relevant art will understand that various forms and details may be changed therein without departing from the spirit and scope of the present invention.
[0424] All references, issued patents, and patent applications cited within the body of this specification are hereby incorporated by reference in their entirety for all purposes.
Claims
1. A polynucleotide probe set for enriching cfDNA, the set comprising: (A) one or more tumor-informed polynucleotide probes; and (B) one or more tumor-uninformed polynucleotide probes.
2. The set according to claim 1, wherein the one or more tumor-informed polynucleotide probes are configured to capture a target sequence comprising an epitope sequence encoded by a cancer vaccine administered to a subject, wherein the subject has been determined to have a tumor expressing the epitope sequence.
3. The set according to claim 2, wherein the epitope sequence comprises a KRAS mutation.
4. The set according to claim 3, wherein the KRAS mutation is selected from the group consisting of: KRAS_G12C mutation, KRAS_G12D mutation, KRAS_G12V mutation, and KRAS_Q61H mutation.
5. The set according to claim 2, wherein the epitope sequence comprises a mutation selected from the group consisting of: KRAS_G13D, KRAS_Q61K, TP53_R249M, CTNNB1_S45P, CTNNB1_S45F, ERBB2_Y772_A775dup, KRAS_G12D, KRAS_Q61R, CTNNB1_T41A, TP53_K132N, KRAS_G12A, KRAS_Q61L, TP53_R213L, BRAF_G466V, KRAS_G12V, KRAS_Q61H, CTNNB1_S37F, TP53_S127Y, TP53_K132E, and KRAS_G12C.
6. The set according to claim 2, wherein the epitope sequence comprises an EGFR mutation.
7. The set according to claim 6, wherein the EGFR mutation includes the EGFR_L858R mutation.
8. The set according to claim 2, wherein the epitope sequence comprises one or more subject-specific epitopes, wherein the subject's tumor has been sequenced to determine the subject-specific epitopes to be encoded by the cancer vaccine.
9. The set according to claim 8, wherein the one or more subject-specific epitopes comprise at least 2 subject-specific epitopes, at least 10 subject-specific epitopes, at least 20 subject-specific epitopes, or 2 to 20 subject-specific epitopes.
10. The set according to claim 8, wherein the one or more subject-specific epitopes comprise 2 to 20 subject-specific epitopes.
11. The set according to any one of claims 2-10, wherein the set further comprises additional tumor-informed polynucleotide probes that capture additional target sequences, wherein the tumor has been determined to express the additional target sequences, and wherein the additional target sequences are not encoded by the cancer vaccine.
12. The group according to claim 11, wherein the additional target sequences comprise at least 10 target sequences, at least 20 target sequences, at least 30 target sequences, at least 100 target sequences, 10 to 500 target sequences, 30 to 500 target sequences, 100 to 500 target sequences, 10 to 100 target sequences, 30 to 100 target sequences, or 100 to 100 target sequences.
13. The group according to claim 11 or 12, wherein the additional target sequences have been predicted to be presented by at least one HLA of the subject.
14. The group according to any one of claims 1-13, wherein the one or more tumor-naïve polynucleotide probes are configured to capture target sequences that comprise sequences of interest selected from the group consisting of: cancer-related genes, oncogenes, tumor suppressor genes, interferon-γ signaling pathway genes, antigen processing pathway genes, and combinations thereof.
15. The group according to any one of claims 1-13, wherein the one or more tumor-naïve polynucleotide probes are configured to capture target sequences that comprise sequences of interest selected from each of the following: cancer-related genes, oncogenes, tumor suppressor genes, interferon-γ signaling pathway genes, and antigen processing pathway genes.
16. The group according to claim 14 or 15, wherein the cancer-related genes are selected from the group consisting of: ABCA12, ACVR2A, AKAP9, BMPR2, COL12A1, CSMD3, DNAH5, DOCK3, FAT2, FAT3, FAT4, FGF10, FGF6, FLG, MAGI1, MDN1, MMAB, NBEA, OBSCN, PCBP1, PCLO, PLEKHA6, PROC, RAD54L, RELN, RPL22, RYR2, TCERG1, WRN, and ZDBF2.
17. The group according to claim 14 or 15, wherein the cancer-related genes include each of the following: ABCA12, ACVR2A, AKAP9, BMPR2, COL12A1, CSMD3, DNAH5, DOCK3, FAT2, FAT3, FAT4, FGF10, FGF6, FLG, MAGI1, MDN1, MMAB, NBEA, OBSCN, PCBP1, PCLO, PLEKHA6, PROC, RAD54L, RELN, RPL22, RYR2, TCERG1, WRN, and ZDBF2.
18. The group according to any one of claims 14 - 17, wherein the oncogene is selected from the group consisting of: ABL1, AKT2, ALK, AR, BCL6, BCL9L, BRAF, BTK, CARD11, CCND1, CCND3, CTNNB1, DDR2, EGFR, ERBB2, ERBB3, FGFR1, FGFR3, FHOD3, FLT1, FLT3, GNAS, HRAS, KDR, KIT, KRAS, MAP2K1, MAP2K2, MECOM, MED12, MET, MTOR, NOTCH1, NOTCH2, NOTCH3, NRAS, NRG1, NTRK1, NTRK3, PDGFRA, PDGFRB, PIK3CA, PIK3CG, PSMB2, RET, ROS1, SF3B1, SMO, SYNE1, and ZBTB20.
19. The group according to any one of claims 14 - 17, wherein the oncogene comprises each of the following: ABL1, AKT2, ALK, AR, BCL6, BCL9L, BRAF, BTK, CARD11, CCND1, CCND3, CTNNB1, DDR2, EGFR, ERBB2, ERBB3, FGFR1, FGFR3, FHOD3, FLT1, FLT3, GNAS, HRAS, KDR, KIT, KRAS, MAP2K1, MAP2K2, MECOM, MED12, MET, MTOR, NOTCH1, NOTCH2, NOTCH3, NRAS, NRG1, NTRK1, NTRK3, PDGFRA, PDGFRB, PIK3CA, PIK3CG, PSMB2, RET, ROS1, SF3B1, SMO, SYNE1, and ZBTB20.
20. The group according to any one of claims 14 - 19, wherein the tumor suppressor gene is selected from the group consisting of: TP53, PTEN, ARID1A, APC, AMER1, ASXL1, ATM, ATR, ATRX, AXIN2, BARD1, BRCA1, BRCA2, CASP8, CFH, CREBBP, DNMT3A, EP300, ERCC2, ERCC5, EXT1, FANCA, FANCD2, FANCI, FANCM, FAT1, FBXW7, HNF1A, MAX, MLH1, MSH3, MSH6, NF1, PIK3R1, PTCH1, PTPRT, RECQL4, RNF43, ROBO1, SLX4, SMAD2, SMAD3, SMAD4, SOX9, TCF7L2, TERT promoter, TET2, TGFBR2, TP53BP1, TSC1, TSC2, WNT16, XPC, ZFP36L2, and ZNRF3.
21. The group according to any one of claims 14-19, wherein the tumor suppressor genes include each of the following: TP53, PTEN, ARID1A, APC, AMER1, ASXL1, ATM, ATR, ATRX, AXIN2, BARD1, BRCA1, BRCA2, CASP8, CFH, CREBBP, DNMT3A, EP300, ERCC2, ERCC5, EXT1, FANCA, FANCD2, FANCI, FANCM, FAT1, FBXW7, HNF1A, MAX, MLH1, MSH3, MSH6, NF1, PIK3R1, PTCH1, PTPRT, RECQL4, RNF43, ROBO1, SLX4, SMAD2, SMAD3, SMAD4, SOX9, TCF7L2, TERT promoter, TET2, TGFBR2, TP53BP1, TSC1, TSC2, WNT16, XPC, ZFP36L2, and ZNRF3.
22. The group according to any one of claims 14-21, wherein the interferon-γ signaling pathway genes are selected from the group consisting of: IFNGR1, INFGR2, JAK1, JAK2, and STAT1.
23. The group according to any one of claims 14-21, wherein the interferon-γ signaling pathway genes include each of the following: IFNGR1, INFGR2, JAK1, JAK2, and STAT1.
24. The group according to any one of claims 14-23, wherein the antigen processing pathway genes are selected from the group consisting of: B2M, HLA-A, HLA-B, HLA-C, HLA-E, TAP1, TAP2, NLRC5, CALR, CANX, PSMB2, and TAPBP.
25. The group according to any one of claims 14-23, wherein the antigen processing pathway genes include each of the following: B2M, HLA-A, HLA-B, HLA-C, HLA-E, TAP1, TAP2, NLRC5, CALR, CANX, PSMB2, and TAPBP.
26. The group according to any one of claims 1-13, wherein the one or more tumor-agnostic polynucleotide probes are configured to capture a target sequence, the target sequence comprising an interesting sequence selected from the group consisting of: ABCA12, ACVR2A, AKAP9, BMPR2, COL12A1, CSMD3, DNAH5, DOCK3, FAT2, FAT3, FAT4, FGF10, FGF6, FLG, MAGI1, MDN1, MMAB, NBEA, OBSCN, PCBP1, PCLO, PLEKHA6, PROC, RAD54L, RELN, RPL22, RYR2, TCERG1, WRN, ZDBF2, ABL1, AKT2, ALK, AR, BCL6, BCL9L, BRAF, BTK, CARD11, CCND1, CCND3, CTNNB1, DDR2, EGFR, ERBB2, ERBB3, FGFR1, FGFR3, FHOD3, FLT1, FLT3, GNAS, HRAS, KDR, KIT, KRAS, MAP2K1, MAP2K2, MECOM, MED12, MET, MTOR, NOTCH1, NOTCH2, NOTCH3, NRAS, NRG1, NTRK1, NTRK3, PDGFRA, PDGFRB, PIK3CA, PIK3CG, RET, ROS1, SF3B1, SMO, SYNE1, ZBTB20, TP53, PTEN, ARID1A, APC, AMER1, ASXL1, ATM, ATR, ATRX, AXIN2, BARD1, BRCA1, BRCA2, CASP8, CFH, CREBBP, DNMT3A, EP300, ERCC2, ERCC5, EXT1, FANCA, FANCD2, FANCI, FANCM, FAT1, FBXW7, HNF1A, MAX, MLH1, MSH3, MSH6, NF1, PIK3R1, PTCH1, PTPRT, RECQL4, RNF43, ROBO1, SLX4, SMAD2, SMAD3, SMAD4, SOX9, TCF7L2, TERT promoter, TET2, TGFBR2, TP53BP1, TSC1, TSC2, WNT16, XPC, ZFP36L2, ZNRF3, IFNGR1, INFGR2, JAK1, JAK2, STAT1, B2M, HLA-A, HLA-B, HLA-C, HLA-E, TAP1, TAP2, NLRC5, CALR, CANX, PSMB2, TAPBP, and combinations thereof.
27. The group according to any one of claims 1-13, wherein the one or more tumor-agnostic polynucleotide probes are configured to capture a target sequence that comprises a sequence of interest selected from each of the following: ABCA12, ACVR2A, AKAP9, BMPR2, COL12A1, CSMD3, DNAH5, DOCK3, FAT2, FAT3, FAT4, FGF10, FGF6, FLG, MAGI1, MDN1, MMAB, NBEA, OBSCN, PCBP1, PCLO, PLEKHA6, PROC, RAD54L, RELN, RPL22, RYR2, TCERG1, WRN, ZDBF2, ABL1, AKT2, ALK, AR, BCL6, BCL9L, BRAF, BTK, CARD11, CCND1, CCND3, CTNNB1, DDR2, EGFR, ERBB2, ERBB3, FGFR1, FGFR3, FHOD3, FLT1, FLT3, GNAS, HRAS, KDR, KIT, KRAS, MAP2K1, MAP2K2, MECOM, MED12, MET, MTOR, NOTCH1, NOTCH2, NOTCH3, NRAS, NRG1, NTRK1, NTRK3, PDGFRA, PDGFRB, PIK3CA, PIK3CG, PSMB2, RET, ROS1, SF3B1, SMO, SYNE1, ZBTB20, TP53, PTEN, ARID1A, APC, AMER1, ASXL1, ATM, ATR, ATRX, AXIN2, BARD1, BRCA1, BRCA2, CASP8, CFH, CREBBP, DNMT3A, EP300, ERCC2, ERCC5, EXT1, FANCA, FANCD2, FANCI, FANCM, FAT1, FBXW7, HNF1A, MAX, MLH1, MSH3, MSH6, NF1, PIK3R1, PTCH1, PTPRT, RECQL4, RNF43, ROBO1, SLX4, SMAD2, SMAD3, SMAD4, SOX9, TCF7L2, TERT promoter, TET2, TGFBR2, TP53BP1, TSC1, TSC2, WNT16, XPC, ZFP36L2, ZNRF3, IFNGR1, INFGR2, JAK1, JAK2, STAT1, B2M, HLA-A, HLA-B, HLA-C, HLA-E, TAP1, TAP2, NLRC5, CALR, CANX, PSMB2, and TAPBP.
28. The group according to any one of claims 1 - 13, wherein the one or more tumor-naive polynucleotide probes are configured to capture a target sequence, the target sequence comprising a sequence of interest selected from each of the following: ABL1, AKT2, ALK, APC, AR, ATR, ATRX, BARD1, BCL6, BMPR1A, BRAF, BRCA1, BRCA2, BTK, CARD11, CCND1, CCND3, CDK12, CFH, CREBBP, CTNNB1, DDR2, DNMT3A, EGFR, EP300, ERBB2, ERBB3, ERCC2, ERCC5, EXT1, FANCA, FANCD2, FANCI, FANCM, FBXW7, FGF10, FGF6, FGFR1, FGFR3, FLI1, FLT1, FLT3, GNAS, HNF1A, HRAS, KDR, KIT, KRAS, MAGI1, MAP2K1, MAP2K2, MAX, MED12, MET, MLH1, MMAB, MSH3, MSH6, MTOR, NF1, NFE2L2, NOTCH1, NOTCH2, NOTCH3, NRAS, NRG1, NTRK1, NTRK3, PDGFRA, PDGFRB, PIK3CA, PIK3CG, PIK3R1, PMS2, PPARG, PROC, PTCH1, RAD54L, RAF1, RECQL4, RET, ROS1, SF3B1, SF3B2, SLX4, SMO, TERT promoter, TET2, TP53BP1, TSC1, TSC2, WRN, XPA, XPC, ZNF395, B2M, HLA-A, HLA-B, HLA-C, TAP1, TAP2, NLRC5, IFNGR1, INFGR2, JAK1, JAK2, TP53, PTEN, and ARID1A.
29. The group according to any one of claims 1 - 28, wherein the one or more tumor-naive polynucleotide probes comprise two or more probes configured to capture all the coding exon sequences of a given gene.
30. The group according to any one of claims 1 - 29, wherein the one or more tumor-naive polynucleotide probes comprise two or more probes configured to capture genomic regions of interest associated with cancer.
31. The group according to any one of claims 1 - 30, wherein the tumor-informed polynucleotide probes and / or the tumor-naive polynucleotide probes comprise probes having overlapping sequences.
32. The set according to any one of claims 1-31, wherein the set comprises at least 20 probes, at least 30 probes, at least 40 probes, at least 50 probes, at least 60 probes, at least 70 probes, at least 80 probes, at least 90 probes, at least 100 probes, at least 200 probes, at least 300 probes, at least 400 probes or at least 500 probes.
33. The set according to any one of claims 1-32, wherein the set is configured to cover at least 100 kb, at least 300 kb, at least 300 kb, at least 400 kb, 100 to 400 kb, 200 to 400 kb, 300 to 400 kb, 100 to 500 kb, 200 to 500 kb, 300 to 500 kb or 340 to 400 kb of the subject's genome.
34. The set according to any one of claims 1-33, wherein the one or more tumor-naïve polynucleotide probes comprise polynucleotide probes configured to capture sequences associated with a given cancer that the subject is known or suspected to have, optionally wherein the cancer is CRC or NSCLC.
35. The set according to any one of claims 1-34, wherein the set further comprises additional polynucleotide probes configured to capture sequences containing polymorphisms in the human population, wherein the sequences containing polymorphisms can be combined to uniquely identify the subject.
36. A method for enriching cfDNA, the method comprising: (a) providing a sample comprising cfDNA; (b) providing a set of polynucleotide probes comprising the set according to any one of claims 1-35; (c) contacting the sample comprising cfDNA with the set of polynucleotide probes under conditions sufficient to hybridize the cfDNA containing the target sequence of interest to its corresponding polynucleotide probe; and (d) capturing the hybridized cfDNA and polynucleotide probe pairs to enrich the cfDNA.
37. A method for monitoring the cancer status of a subject having, having had, or suspected of having cancer, wherein the method comprises the steps of: a. obtaining or having obtained sequencing data of cell-free DNA (cfDNA) from a sample from the subject, and wherein the sequencing data comprises target coverage of at least 50% of all polynucleotide regions of interest corresponding to mutations present in the cancer exome, and wherein the polynucleotide regions of interest sequenced have a read depth of at least 1000-fold, optionally wherein the polynucleotide regions of interest comprise at least 50 mutations, optionally wherein the average read depth is the average duplex read depth, and optionally wherein obtaining the sequencing data comprises collecting or having collected the sample from the subject, isolating or having isolated the cfDNA, enriching or having enriched the cfDNA and / or sequencing or having sequenced the cfDNA; and b. Determine or have determined the frequency of the mutations present in the exome to assess the status of the cancer, optionally wherein the assessment of the status includes an assessment of the presence and / or cancer burden, wherein the cfDNA has been enriched prior to sequencing using: (1) a subject-specific polynucleotide probe set; (2) a tumor-naive polynucleotide probe set; and / or (3) a combined set of a tumor-informed polynucleotide probe set and a tumor-naive polynucleotide probe set, wherein the polynucleotide probes are configured to capture the polynucleotide regions of interest.
38. A method for monitoring the cancer status of a subject having, having had, or suspected of having cancer, wherein the method comprises the steps of: a. Obtain or have obtained sequencing data of cell-free DNA (cfDNA) from a sample from the subject, and wherein the sequencing data includes target coverage of at least 95% of all polynucleotide regions of interest corresponding to mutations present in the cancer exome, wherein the polynucleotide regions of interest contain at least 50 mutations, and wherein the sequenced polynucleotide regions of interest have a duplex read depth of at least 1000-fold, and optionally wherein obtaining the sequencing data includes collecting or having collected the sample from the subject, isolating or having isolated the cfDNA, enriching or having enriched the cfDNA and / or sequencing or having sequenced the cfDNA; and b. Determine or have determined the frequency of at least 50 mutations present in the exome to assess the status of the cancer, optionally wherein the assessment of the status includes an assessment of the presence and / or cancer burden, wherein the cfDNA has been enriched prior to sequencing using: (1) a tumor-informed polynucleotide probe set; (2) a tumor-naive polynucleotide probe set; and / or (3) a combined set of a tumor-informed polynucleotide probe set and a tumor-naive polynucleotide probe set, wherein the polynucleotide probes are configured to capture the polynucleotide regions of interest.
39. A method for assessing the efficacy of a therapy in a subject having, having had, or suspected of having cancer, wherein the method comprises the steps of: a. Sequencing data of cell-free DNA (cfDNA) is obtained or has been obtained from a pre-therapy sample from the subject, and wherein the sequencing data includes target coverage of at least 50% of all polynucleotide regions of interest corresponding to mutations present in the cancer exome, and wherein the polynucleotide regions of interest sequenced have a read depth of at least 1000-fold, optionally wherein the polynucleotide regions of interest contain at least 50 mutations, optionally wherein the average read coverage is the average duplex read coverage, and optionally wherein obtaining the sequencing data includes collecting or having collected the pre-therapy sample from the subject, isolating or having isolated the pre-therapy cfDNA, enriching or having enriched the pre-therapy cfDNA and / or sequencing or having sequenced the pre-therapy cfDNA; b. Sequencing data of cell-free DNA (cfDNA) is obtained or has been obtained from a post-therapy sample from the subject, optionally wherein the therapy includes a cancer vaccine comprising the neoantigen or an expression system encoding the neoantigen, and wherein the sequencing data includes target coverage of at least 50% of all polynucleotide regions of interest corresponding to mutations present in the cancer exome, and wherein the polynucleotide regions of interest sequenced have a read depth of at least 1000-fold, optionally wherein the polynucleotide regions of interest contain at least 50 mutations, optionally wherein the average read coverage is the average duplex read coverage, and optionally wherein obtaining the sequencing data includes collecting or having collected the post-therapy sample from the subject, isolating or having isolated the post-therapy cfDNA, enriching or having enriched the post-therapy cfDNA and / or sequencing or having sequenced the post-therapy cfDNA; and c. Determining or having determined the frequency of the mutations present in the exome of the pre-therapy cfDNA relative to the post-therapy cfDNA to evaluate the efficacy of the therapy, optionally wherein an increase in the frequency of the mutations in the post-therapy cfDNA relative to the pre-therapy cfDNA indicates an increased likelihood of an increase in the tumor burden of the subject, and optionally wherein a decrease or maintenance of the frequency of the mutations in the post-therapy cfDNA relative to the pre-therapy cfDNA indicates an increased likelihood of a decrease or stability of the tumor burden of the subject, wherein the cfDNA has been enriched prior to sequencing using: (1) a tumor-informed polynucleotide probe set; (2) a tumor-uninformed polynucleotide probe set; and / or (3) a combined set of a tumor-informed polynucleotide probe set and a tumor-uninformed polynucleotide probe set, wherein the polynucleotide probes are configured to capture the polynucleotide regions of interest.
40. A method for evaluating the efficacy of a therapy in a subject having, having had, or suspected of having cancer, wherein the method comprises the steps of: a. Obtain or have obtained sequencing data of tumor-derived DNA from cancer-affected tissue from the subject, optionally wherein obtaining the sequencing data comprises collecting or having collected the cancer-affected tissue, isolating or having isolated the tumor-derived DNA, and sequencing or having sequenced the tumor-derived DNA; b. Determine or have determined from the tumor-derived DNA sequencing data one or more tumor-associated mutations relative to the wild-type germline nucleic acid sequence of the subject, optionally wherein one or more of the one or more tumor-associated mutations are associated with a neoantigen comprising at least one alteration that causes the peptide sequence encoded by the tumor-derived DNA to be different from the corresponding peptide sequence encoded by the wild-type germline nucleic acid sequence of the subject; c. Design and / or select or have designed and / or have selected (1) a tumor-informed polynucleotide probe set; (2) a tumor-uninformed polynucleotide probe set; and / or (3) a combined set of a tumor-informed polynucleotide probe set and a tumor-uninformed polynucleotide probe set, wherein the polynucleotide probes are configured to capture at least the tumor-associated mutations, optionally wherein the polynucleotide regions of interest comprise at least 50 tumor-associated mutations; d. Obtain or have obtained sequencing data of cell-free DNA (cfDNA) from a pre-treatment sample from the subject, wherein the pre-treatment cfDNA has been enriched using polynucleotide probes prior to sequencing, and wherein the sequencing data comprises a target coverage of at least 50% of all polynucleotide regions of interest corresponding to the tumor-associated mutations, and wherein the polynucleotide regions of interest sequenced have a read depth of at least 1000-fold, optionally wherein the average read coverage is the average duplex read coverage, and optionally wherein obtaining the sequencing data comprises collecting or having collected the pre-treatment sample from the subject, isolating or having isolated the pre-treatment cfDNA, enriching or having enriched the pre-treatment cfDNA, and / or sequencing or having sequenced the pre-treatment cfDNA; e. Obtaining or having obtained sequencing data of cell-free DNA (cfDNA) from a post-therapy sample from the subject, optionally wherein the therapy comprises a cancer vaccine comprising the neoantigen or an expression system encoding the neoantigen, wherein the post-therapy cfDNA has been enriched using polynucleotide probes prior to sequencing, and wherein the sequencing data comprises a target coverage of at least 50% of all polynucleotide regions of interest corresponding to the tumor-associated mutations, and wherein the polynucleotide regions of interest sequenced have a read depth of at least 1000-fold, optionally wherein the average read coverage is the average duplex read coverage, and optionally wherein obtaining the sequencing data comprises collecting or having collected the post-therapy sample from the subject, isolating or having isolated the post-therapy cfDNA, enriching or having enriched the post-therapy cfDNA, and / or sequencing or having sequenced the post-therapy cfDNA; and f. Determining or having determined the frequency of the tumor-associated mutations in the pre-therapy cfDNA relative to the post-therapy cfDNA to evaluate the efficacy of the therapy, optionally wherein determining at least the one or more tumor-associated mutations associated with the neoantigen, optionally wherein an increase in the frequency of the mutations in the post-therapy cfDNA relative to the pre-therapy cfDNA indicates an increased likelihood of an increase in the tumor burden of the subject, and optionally wherein a decrease or maintenance of the frequency of the mutations in the post-therapy cfDNA relative to the pre-therapy cfDNA indicates an increased likelihood of a decrease or stability of the tumor burden of the subject.
41. The method according to any one of the preceding method claims, wherein the method comprises designing and / or selecting or having designed and / or having selected a combined set of tumor-informed polynucleotide probe sets and tumor-uninformed polynucleotide probe sets.
42. The method according to claim 41, wherein the combined set designed and / or selected comprises the set according to any one of claims 1-35.
43. A method for enriching cfDNA, the method comprising: (a) Providing a sample comprising cfDNA; (b) Providing a polynucleotide probe set, wherein the set comprises: (i) One or more tumor-informed polynucleotide probes; and (ii) One or more tumor-uninformed polynucleotide probes; (c) Contacting the sample comprising cfDNA with the polynucleotide probe set under conditions sufficient to hybridize the cfDNA comprising the target sequence of interest with its corresponding polynucleotide probe; and (d) Capturing the hybridized cfDNA and polynucleotide probe pairs to enrich the cfDNA.
44. The method according to claim 43, wherein the set comprises the set according to any one of claims 1-35.
45. The method according to any one of the preceding claims, wherein the method comprises one or more of the following steps: a. Collecting or having collected the sample from the subject; b. Isolating or having isolated the cfDNA; c. enrich or have enriched the cfDNA; or d. sequence or have sequenced the cfDNA.
46. The method according to any one of the preceding claims, wherein the method comprises each of the following steps: a. collect or have collected the sample from the subject; b. isolate or have isolated the cfDNA; c. enrich or have enriched the cfDNA; and d. sequence or have sequenced the cfDNA.
47. The method according to any one of the preceding claims, wherein the average read depth comprises an average read coverage of at least 1500-fold, at least 2000-fold, at least 2500-fold, 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold or at least 5000-fold.
48. The method according to any one of the preceding claims, wherein the average read depth comprises an average read coverage range of 1000-fold to 5000-fold.
49. The method according to any one of the preceding claims, wherein the average read depth comprises an average read coverage range of 1000-fold to 4000-fold, 1000-fold to 3000-fold, 1000-fold to 2000-fold, 2000-fold to 5000-fold, 2000-fold to 4000-fold, 2000-fold to 3000-fold, 3000-fold to 5000-fold, 3000-fold to 4000-fold or 4000-fold to 5000-fold.
50. The method according to any one of the preceding claims, wherein the average read depth comprises an average read duplex depth.
51. The method according to any one of the preceding claims, wherein each of the polynucleotide regions of interest corresponding to the mutations present in the exome has a read depth of at least 1000-fold.
52. The method according to any one of the preceding claims, wherein each of the polynucleotide regions of interest corresponding to the mutations present in the exome has a read depth of at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, 3000-fold, at least 3500-fold, at least 4000-fold, at least 4500-fold or at least 5000-fold.
53. The method according to any one of the preceding claims, wherein the target coverage comprises at least 60%, at least 70%, at least 80% or at least 90% of the polynucleotide regions of interest corresponding to the mutations present in the cancer exome.
54. The method according to any one of the preceding claims, wherein the target coverage comprises at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9% or 100% of the polynucleotide regions of interest corresponding to the mutations present in the cancer exome.
55. The method according to any one of the preceding claims, wherein the target coverage comprises at least 95% of the polynucleotide regions of interest corresponding to the mutations present in the cancer exome.
56. The method according to any one of the preceding claims, wherein the polynucleotide region of interest comprises at least 50, at least 60, at least 70, at least 80, or at least 90 mutations.
57. The method according to any one of the preceding claims, wherein the polynucleotide region of interest comprises at least 50 mutations.
58. The method according to any one of the preceding claims, wherein the polynucleotide region of interest comprises at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1000 mutations.
59. The method according to any one of the preceding claims, wherein the method comprises the following steps: a. Obtaining or having obtained sequencing data of tumor-derived DNA from cancerous tissue of the subject, optionally wherein obtaining the sequencing data comprises collecting or having collected the cancerous tissue, isolating or having isolated the tumor-derived DNA, and sequencing or having sequenced the tumor-derived DNA; b. Determining or having determined from the tumor-derived DNA sequencing data one or more tumor-related mutations relative to the wild-type germline nucleic acid sequence of the subject, optionally wherein one or more of the one or more tumor-related mutations are associated with a neoantigen comprising at least one alteration such that the peptide sequence encoded by the tumor-derived DNA is different from the corresponding peptide sequence encoded by the wild-type germline nucleic acid sequence of the subject; c. Designing and / or selecting or having designed and / or having selected (1) a tumor-informed polynucleotide probe set; (2) a tumor-uninformed polynucleotide probe set; and / or (3) a combined set of a tumor-informed polynucleotide probe set and a tumor-uninformed polynucleotide probe set, wherein the polynucleotide probes are configured to capture the polynucleotide region of interest corresponding to the tumor-related mutations, optionally wherein the polynucleotide region of interest comprises at least 50 tumor-related mutations; and d. Enriching or having enriched the cfDNA using the polynucleotide probes prior to sequencing.
60. The method according to any one of the preceding claims, wherein the cancer is selected from the group consisting of: lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer.
61. The method according to any one of claims 1 to 60, wherein the subject has been administered a therapy.
62. The method according to claim 61, wherein the therapy comprises a cancer vaccine.
63. The method according to claim 62, wherein the cancer vaccine comprises an epitope-encoding nucleic acid sequence encoding at least one of the mutations present in the cancer exome.
64. The method according to claim 62 or 63, wherein the cancer vaccine comprises a self - amplifying expression system based on alphavirus.
65. The method according to claim 62 or 63, wherein the cancer vaccine comprises an expression system based on chimpanzee adenovirus (ChAdV).
66. The method according to any one of the preceding claims, wherein the method comprises obtaining sequencing data of cfDNA from two or more samples from the subject.
67. The method according to claim 66, wherein the two or more samples are collected at different time points.
68. The method according to claim 67, wherein the two or more samples are collected at different time points relative to the administration of the therapy.
69. The method according to claim 68, wherein a pre - therapy sample is collected before the administration of the therapy, and post - therapy cfDNA is collected after the administration of the therapy.
70. The method according to claim 69, wherein the determining step comprises determining or having determined the frequency of the mutations of the pre - therapy cfDNA relative to the post - therapy cfDNA to evaluate the efficacy of the therapy, optionally wherein at least one or more of the tumor - associated mutations related to the neoantigen are determined, optionally wherein an increase in the frequency of the mutations in the post - therapy cfDNA relative to the pre - therapy cfDNA indicates an increased likelihood of an increase in the tumor burden of the subject, and optionally wherein a decrease or maintenance in the frequency of the mutations in the post - therapy cfDNA relative to the pre - therapy cfDNA indicates an increased likelihood of a decrease or stability in the tumor burden of the subject.
71. The method according to claim 70, wherein an increase in the frequency of one or more of the mutations in the post - therapy cfDNA relative to the pre - therapy cfDNA in the tumor - naive group indicates the possibility of tumor mutations of an immune escape mechanism.
72. The method according to claim 70, wherein an increase in the frequency of the mutations in the post - therapy cfDNA relative to the pre - therapy cfDNA indicates an increased likelihood of an increase in the tumor burden of the subject.
73. The method according to claim 70, wherein a decrease or maintenance in the frequency of the mutations in the post - therapy cfDNA relative to the pre - therapy cfDNA indicates an increased likelihood of a decrease or stability in the tumor burden of the subject.
74. The method according to claim 71, wherein the decrease comprises a complete response (CR) or a partial response (PR).
75. The method according to any one of the preceding claims, wherein the method further comprises administering a therapy to the subject after evaluating the status of the cancer.
76. The method according to claim 75, wherein the evaluation of the frequency of the mutations in the cfDNA indicates the likelihood that the subject has or still has cancer.
77. The method according to claim 75 or 76, wherein the therapy comprises a cancer vaccine.
78. The method according to claim 77, wherein the cancer vaccine comprises an epitope-encoding nucleic acid sequence encoding at least one of the mutations present in the exome.
79. The method according to claim 77 or 78, wherein the cancer vaccine comprises a self-amplifying expression system based on alphavirus.
80. The method according to claim 77 or 78, wherein the cancer vaccine comprises an expression system based on chimpanzee adenovirus (ChAdV).
81. The method according to any one of the preceding claims, wherein the collecting step comprises collecting a blood sample.
82. The method according to any one of the preceding claims, wherein the separating step comprises centrifuging to separate cfDNA from cells and / or cell debris.
83. The method according to any one of the preceding claims, wherein the separating step comprises separating cfDNA from whole blood.
84. The method according to claim 83, wherein separating cfDNA from whole blood comprises separating the plasma layer, the buffy coat, and the red blood cells.
85. The method according to claim 84, wherein the cfDNA is separated from the plasma layer.
86. The method according to any one of the preceding claims, wherein the sequencing step comprises next-generation sequencing (NGS) or Sanger sequencing.
87. The method according to claim 86, wherein NGS comprises duplex sequencing, whole exome sequencing, whole genome sequencing, de novo sequencing, phased sequencing, targeted amplicon sequencing, or shotgun sequencing.
88. The method according to any one of the preceding claims, wherein the enriching step comprises enriching the region of interest of the polynucleotide corresponding to the mutation present in the exome in the cfDNA prior to sequencing.
89. The method according to claim 88, wherein the enrichment comprises the combination of the tumor-informed polynucleotide probe set and the tumor-uninformed polynucleotide probe set, optionally wherein the tumor-informed polynucleotide probe set and the tumor-uninformed polynucleotide probe set are enriched separately in their respective individual samples.
90. The method according to claim 89, wherein the tumor-informed polynucleotide probe comprises each region of interest of the polynucleotide corresponding to the mutation present in the exome.
91. The method according to any one of claims 88-90, wherein the tumor-informed polynucleotide probe comprises at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9%, or 100% of the region of interest of the polynucleotide corresponding to the mutation present in the cancer exome.
92. The method according to any one of claims 88-90, wherein the tumor-informed polynucleotide probe comprises at least 50, at least 60, at least 70, at least 80, at least 90 mutations, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900 or at least 1000 mutations, optionally wherein the mutations are present in the cancer exons.
93. The method according to any one of the preceding claims, wherein the enrichment step comprises hybridizing one or more polynucleotide probes to the one or more regions of polynucleotides of interest.
94. The method according to any one of the preceding claims, wherein the polynucleotide probe has a length of 80 to 150 base pairs (bp).
95. The kit or method according to claim 94, wherein the polynucleotide probe has a length of 50-100, 50-150, 80 to 140, 80 to 130, 80 to 120, 80 to 110, 80 to 100, 80 to 90, 90 to 150, 90 to 140, 90 to 130, 90 to 120, 90 to 110, 90 to 100, 100 to 150, 100 to 140, 100 to 130, 100 to 120, 100 to 110, 110 to 150, 110 to 140, 110 to 130, 110 to 120, 120 to 150, 120 to 140, 120 to 130, 130 to 150, 130 to 140, 140 to 150, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140 or 150 bp.
96. The kit or method according to any one of claims 88-95, wherein the one or more polynucleotide probes are biotinylated.
97. The method according to any one of the preceding claims, wherein the tumor-informed polynucleotide probe is designed or selected after sequencing the tumor of the subject.
98. The kit or method according to claim 97, wherein the tumor-informed polynucleotide probe is designed or selected after exome sequencing of the tumor of the subject.
99. The kit or method according to claim 98, wherein the tumor-informed polynucleotide probe is designed or selected to target all mutations of the sequenced tumor.
100. The method according to any one of the preceding claims, wherein the sequencing step comprises ligating sequencing adapters to the cfDNA.
101. The method according to claim 100, wherein the sequencing adapter is configured for dual sequencing.
102. The kit or method according to any one of the preceding claims, wherein one or more of the mutations comprise point mutations, frameshift mutations, non-frameshift mutations, deletion mutations, insertion mutations, splicing variants, genomic rearrangements, proteasome-generated spliced antigens or combinations thereof.
103. The group or method according to any one of the preceding claims, wherein one or more of said mutations comprise at least one alteration that alters such that the peptide sequence encoded by said cfDNA is different from the corresponding peptide sequence encoded by the wild-type germline nucleic acid sequence of said subject.
104. The group or method according to any one of the preceding claims, wherein said one or more mutations consist of coding mutations comprising at least one alteration that alters such that the peptide sequence encoded by said cfDNA is different from the corresponding peptide sequence encoded by the wild-type germline nucleic acid sequence of said subject.
Citation Information
Patent Citations
Fast method for detection and / or identification of a single base in a nucleic acid sequence and its applications.
FR2650840A1
Stabilizing a nucleic acid for nucleic acid sequencing
US20060252077A1
Adapters, methods, and compositions for duplex sequencing
US20170211140A1
Noninvasive diagnostics by sequencing 5-hydroxymethylated cell-free DNA
US20200277667A1
Method of encapsulating biologically active materials in lipid vesicles
US4235871A