Dual assay to boost accuracy of detected actionable variants in liquid biopsy

The dual assay method enhances the detection of actionable variants in liquid biopsies by combining a sensitive discovery assay with a specific confirmation assay, ensuring accurate and continuous monitoring of cancer-related genetic changes.

WO2026073127A1PCT designated stage Publication Date: 2026-04-02AMAZON TECH INC
View PDF 16 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Current methods for identifying actionable variants in liquid biopsies face challenges in achieving high sensitivity and specificity, leading to high false positive rates and missed detections of novel variants, which hinders effective cancer treatment decisions and monitoring.

Method used

A dual assay approach comprising a discovery assay with high sensitivity and a confirmation assay with high specificity, utilizing a personalized sequencing panel to identify actionable variants through sequencing at varying depths and employing bioinformatics tools for variant scoring.

Benefits of technology

The method achieves high sensitivity and specificity in detecting actionable variants, enabling accurate monitoring and treatment decisions for cancer patients, including during and after treatment, while allowing for longitudinal tracking of variant landscapes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025048478_02042026_PF_FP_ABST
    Figure US2025048478_02042026_PF_FP_ABST
Patent Text Reader

Abstract

This disclosure relates to a method of detecting actionable variants in liquid biopsies collected from patients. The method includes a dual assay system, wherein the first sequencing protocol (i.e. discovery assay) sequences a large number of variants at a low sequencing depth to exhibit high sensitivity and, the second sequencing protocol (i.e. confirmation assay) sequences prioritized variants from the first sequencing round based on quality and confidence of variant called at high sequencing depth to exhibit high specificity. The method can further include repeating the first sequencing protocol and / or second sequencing protocol to monitor response to treatment.
Need to check novelty before this filing date? Find Prior Art

Description

DUAL ASSAY TO BOOST ACCURACY OF DETECTED ACTIONABLE VARIANTS TN LIQUID BIOPSY CROSS-REFERENCE TO RELATED APPLICATIONS

[0000] The present application claims the benefit of United States Provisional Application No.63 / 701,033, filed on September 30, 2024, the entire contents of which are incorporated herein by reference.BACKGROUND

[0001] Cancer immunotherapy (e.g., cancer vaccine) has emerged as a promising cancer treatment modality. The goal of cancer immunotherapy is to harness the immune system for selective destruction of cancer while leaving normal tissues unharmed. Traditional cancer vaccines typically target tumor-associated antigens. Tumor-associated antigens are typically present in normal tissues, but overexpressed in cancer. However, because these antigens are often present in normal tissues immune tolerance can prevent immune activation. Several clinical trials targeting tumor-associated antigens have failed to demonstrate a durable beneficial effect compared to standard of care treatment. Li et al., Ann Oncol., 28 (Suppl 12): xii 11- xii 17 (2017).

[0002] One of the hallmarks of cancer is the accumulation of somatic alterations, acquired during a cell’s life cycle and development. Shen, J. Mol Cell Biol. 3(1): 1-3 (2011). Throughout cancer progression, some of these variants may dictate response to therapy. Dancey et al., Cell, 148(3):409-420 (2012).

[0003] As such, neoantigens represent an attractive target for cancer immunotherapies.Neoantigens are non-autologous proteins with individual specificity. Neoantigens are derived from random somatic mutations (e.g., somatic single nucleotide variants and small insertions and deletions (indels)) in the tumor cell genome and are not expressed on the surface of normal cells. Id. Because neoantigens are expressed exclusively on tumor cells, and thus do not induce central immune tolerance, cancer vaccines targeting cancer neoantigens have potential advantages, including decreased central immune tolerance and an improved safety profile.

[0004] The reliable detection of neoantigens (i.e., mutations) in cancer genomes is important for developing effective therapies, as well as guiding treatment choices for cancer, such as immunogenic compositions. Identification of somatic mutations is challenged by theheterogeneous composition of tumors. Evaluating heterogeneity to guide choice and sequence of therapy could be achieved by tumor biopsies, but this is impractical due to the associated risk of complications and costs. Alternatively, circulating tumor (ct)DNA can be used to monitor cancer dynamics noninvasively. ctDNAis DNA originating from tumor cells (e.g., from primary tumors, micrometastases, and overt metastases) that is released into the circulation. Fiala et al., (2018), The Journal of Applied Laboratory Medicine, 3(2):300-313. However, as the abundance and quality of ctDNA is low, variant calling is prone to false positives due to noisy sequencing readouts.

[0005] Currently, two approaches to identifying actionable variants in liquid biopsies are practiced, tumor-naive sequencing or tumor-informed sequencing. Tumor-naive sequencing protocols have high sensitivity, as assays are designed to have lower sequencing depth to maximize variant calling resulting in a decreased number of false negatives. However, as a result, tumor-naive sequencing tends to identify a high number of false positives (i.e. low specificity). Conversely, tumor-informed sequencing exhibits high specificity due to prioritizing a select number of known tumor variants and sequencing at high sequencing depth. This results in a decreased number of false positives variants identified but an increased number of false negatives (i.e. low sensitivity). This compromises the tumor-informed sequencing protocol, preventing the discovery of actionable variants outside the predetermined panel of variants.

[0006] Accordingly, there is a significant unmet need for identifying actionable variants in liquid biopsies that exhibits high sensitivity and high specificity, allowing for reduction in false identification of variants (or neoantigens) while permitting discovery of actionable variants and, further, providing patient longitudinal tracking of the variants present during or after treatment.SUMMARY OF THE INVENTION

[0007] This disclosure relates to methods of detecting actionable variants in a liquid biopsy. To detect actionable variants in liquid biopsies, a test sample can be obtained from a liquid biopsy. A discovery assay can be run, wherein the discovery assay can include sequencing the test sample. A list of prioritized variants can be identified to design a personalized sequencing panel based on the list of prioritized variants. A confirmation assay can be run, wherein the confirmation assaycan include sequencing the test sample using the personalized sequencing panel to detect one or more actionable variants in the liquid biopsy.

[0008] The liquid biopsy of the disclosure can be from an amniotic fluid sample, ascitic fluid sample, bile sample, blood sample, buccal sample, cerebral spinal fluid sample, fecal sample, hair sample, peritoneal fluid sample, plasma sample, pleural effusion sample, saliva sample, semen sample, serum sample, skin sample, synovial fluid sample, urine sample, non-solid tumor sample, or a combination of any of the foregoing.

[0009] In some aspects, the test sample is cell-free DNA, cell-free RNA or a combination thereof.

[0010] In one aspect, the cell-free DNA comprises circulating tumor DNA.

[0011] The methods of this disclosure can relate to the discovery assay, wherein the discovery assay can have high sensitivity to detect variants.

[0012] In one aspect, the high sensitivity can be about 95% or more.

[0013] In some aspects, the discovery assay has a variant panel that can comprise a high number of variant sites.

[0014] The methods of this disclosure can include a variant panel that can be used for whole genome sequencing or whole exome sequencing.

[0015] In one aspect, the high number of variant sites can be about 50000000 variant sites.

[0016] In some aspects, the discovery assay has a low sequencing depth.

[0017] In some aspects, the low sequencing depth is between about 200X to about 5000X.

[0018] The methods of this disclosure can relate to a variant caller that can be a bioinformatics tool capable of quantifying a score based on the confidence or quality of a variant called.

[0019] In some aspects, the bioinformatics tool can be capable of using a machine learning model.

[0020] In some aspects, the variant caller can be DRAGEN or Mutect2.

[0021] In some aspects, the score can be based on Somatic Quality (SQ) quantified by DRAGEN or tumor-likelihood (TLOD) quantified by Mutect2.

[0022] The methods of this disclosure can relate to a list of prioritized variants that can include a select number of variants that can be chosen based on the confidence or quality of a variant called.

[0023] In some aspects, the select number of variants can be between about 100 to about 10000 variants.

[0024] As disclosed above, the methods can relate to a personalized sequencing panel that can include a select number of variants.

[0025] The confirmation assay can use the personalized sequencing panel and the personalized sequencing panel can have high specificity to detect the actionable variants.

[0026] In some aspects, the confirmation assay can have high sequencing depth to detect the actionable variants.

[0027] In some aspects, the high sequencing depth is about 10000X to about 100000X.

[0028] The confirmation assay can detect the actionable variants in the liquid biopsy using a variant caller.

[0029] This disclosure also relates to methods of detecting actionable variants in a liquid biopsy collected from a patient having or suspected to have cancer. A test sample can be obtained that can be cell-free DNA or cell-free RNAfrom a liquid biopsy. A discovery assay on cell-free DNA can be run and the discovery assay can sequence about 50000000 variant sites at a low sequencing depth between about 200X to about 5000X sequencing depth. This disclosure can further relate to methods of identifying a list of prioritized variants using a bioinformatic tool that can be capable of quantifying a score based on confidence or quality of variants called. A personalized sequencing panel can be based on the list of prioritized variants. The methods of this disclosure can further relate to a confirmation assay that can sequence the cell-free DNA and / or the cell-free RNA using the personalized sequencing panel at 100000X sequencing depth. The variants can be called using a variant caller to identify actional variants from the detected variants.

[0030] The test sample can include two blood samples collected simultaneously or sequentially from a patient having or suspected to have cancer.

[0031] In some aspects, the personalized sequencing panel can be used to monitor a patient having or suspected to have cancer.

[0032] In some aspects, the personalized sequencing panel can be used to monitor a patient having or suspected to have cancer during and / or after treatment.

[0033] In some aspects, the personalized sequencing panel can be re-designed to add or remove variants not of interest.BRIEF DESCRIPTION OF THE DRAWINGS

[0034] FIG. 1 is a schematic depicting the dual assay approach to discover and identify actionable variants with high sensitivity and high specificity.

[0035] FIG. 2 shows the quantified receiver operating characteristic (ROC) curve demonstrating distinctive ability of a trained random forest classifier to identify variants in a shallow sequencing run that would be confirmed using a targeted deep duplex sequencing. The classifier was trained using a downsampled fragmentomics training set to simulate shallow sequencing and then out-of-bag prediction was used to evaluate the predictive performance of the classifier. The X shows a cutoff point that maximizes the Fs score (where recall is 5x as important as precision).

[0036] FIG 3. shows a dual-axis plot evaluating the predictive performance of a trained random forest classifier, where N beyond cutoff represents how many variants exceeds a probability threshold and p(Confirmed in Tissue) represents how likely those variants are to be confirmed true positives. The vertical line indicates the Fs cutoff shown in FIG. 2.DETAILED DESCRIPTION

[0037] The disclosure relates to a novel approach to identify actionable variants in liquid biopsies obtained from patients. This information can be used to support decisions on patient treatment and monitor actionable variants presence in circulating tumor (ctDNA) before, during and / or after treatment. Identification of actionable variants employing this novel sequencing protocol, tailored to achieve high sensitivity and high specificity, can be used to inform drug development (e.g. immunotherapy, vaccine) and treatment management in patients. Without wishing to be bound by any particular mechanism or theory, current methodologies for identifying variants can have high sensitivity but low specificity. For example, sequencing protocols can employ low sequencing depth to detect a high number of variants, however this approach leads to low specificity (i.e. identifying a number of false positives). Without wishing to be bound by any particular mechanism or theory, current methodologies for identifying variants can have high specificity but low sensitivity. For example, sequencing protocols can usehigh sequencing depth employing a panel of known variants, however this does not allow for the discovery of novel variants nor does this approach have the compacity to monitor changes in the variant landscape over time.

[0038] Genetic material for use in the method can be obtained (e.g., extracted) by processing the biological sample with any technique. The processing can include adding an anticoagulant (e.g., ethylenediaminetetraacetic acid (EDTA), EDTA salt (e.g., K3EDTA, Na2EDTA), citrate salt, citrate buffer, heparin, or heparin salt (e.g., sodium heparin)), centrifugation, separation of layers (e.g., separation of plasma, buffy coat, erythrocytes), filtration (e.g., to remove cells and / or debris), or a combination of any of the foregoing. The processing can include any technique to lyse a cell in the biological sample, including, but not limited to, physical lysis (e.g., grinding under liquid nitrogen, bead beating, French press, a grinder, sonication), enzymatic lysis (e.g., treatment with lysozyme, zymolase, lyticase, proteinase K, collagenase, lipase, and the like), chemical lysis (e.g., treatment with a detergent or surfactant (e.g., sodium dodecyl sulfate), chaotrope (e.g., guanidine salt, alkaline solution) chemical solvents), or a combination of any of the foregoing. The biological sample can be processed (e.g., lysed) in bulk. For example, the biological sample (e.g., tissue or cells of the biological sample) can be lysed in bulk via sonication to release the genetic material from the bulk biological sample (e.g., from all cells of the biological sample). As another example, the biological sample can be processed by bead beating of the sample. The biological sample (e.g., cells or tissue of the biological sample) can be processed on a single cell basis (e.g., by lysis of single cells). For example, a biological sample can be separated by flow cytometry and each cell of the biological sample can be isolated and lysed individually. Continuing this example, the genetic material from each cell can be barcoded to identify the originating cell and sequenced. Genetic material can be purified before sequencing by any purification technique including gel purification, centrifugal spin column purification, magnetic bead-based purification, extraction (e.g., phenol-chloroform extraction), and the like. The quantity and / or quality of the genetic material can be assessed prior to sequencing by any technique, such as optical density, absorbance (e. ., NanoDrop™ (THERMO FISHER SCIENTIFIC INC.) analysis), agarose gel electrophoresis, high-performance liquid chromatography (HPLC), fluorescent nucleic acid-binding dyes, microfluidics measurement(e.g., 2100 Bioanalyzer (AGILENT TECHNOLOGIES, INC.)), real-time quantitative PCR (RT-qPCR), reverse transcriptase qPCR, and the like.

[0039] Any biological sample (or test sample) collected from a patient can be used in the methods described herein. Suitable biological samples include, but are not limited to, an amniotic fluid sample, ascitic fluid sample, bile sample, blood sample, buccal sample, cerebral spinal fluid sample, fecal sample, hair sample, peritoneal fluid sample, plasma sample, pleural effusion sample, saliva sample, semen sample, serum sample, skin sample, synovial fluid sample, urine sample, tissue sample, tumor sample, or a combination of any of the foregoing. Blood samples can be separated to isolate a buffy coat (e.g., comprising leukocytes and / or platelets), a plasma, a serum, erythrocytes, or a combination of any of the foregoing. The biological sample can be a plasma sample, a serum sample, a buffy coat, an erythrocyte sample, or a combination of any of the foregoing. In some embodiments, the biological sample is a blood sample.

[0040] A biological sample used in the methods described herein can contain cells, tissue, or genetic material from any type of cancer, including hematological malignancies, solid tumors, sarcomas, carcinomas, and other solid and non-solid tumors. A patient from whom the biological sample is obtained can have any type of cancer. Illustrative suitable cancerous tumors include, for example, adrenocortical carcinoma, anal cancer, appendiceal cancer, astrocytoma, basal cell carcinoma, brain tumor, bile duct cancer, bladder cancer, bone cancer, breast cancer, bronchial tumor, carcinoma of unknown primary origin, cardiac tumor, cervical cancer, chordoma, colon cancer, colorectal cancer, craniopharyngioma, ductal carcinoma, embryonal tumor, endometrial cancer, ependymoma, esophageal cancer, esthesioneuroblastoma, fibrous histiocytoma, Ewing sarcoma, eye cancer, germ cell tumor, gallbladder cancer, gastric cancer, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor, gestational trophoblastic disease, glioma, head and neck cancer, hepatocellular cancer, histiocytosis, Hodgkin lymphoma, hypopharyngeal cancer, intraocular melanoma, islet cell tumor, Kaposi sarcoma, kidney cancer, Langerhans cell histiocytosis, laryngeal cancer, leukemias (e.g., acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myelogenous leukemia, hairy cell leukemia, myelodysplastic syndrome, prolymphocytic leukemia, large granular lymphocytic leukemia, adult T-cell leukemia, clonal eosinophilias), lip and oral cavity cancer, liver cancer, lobular carcinoma in situ, lung cancer, macroglobulinemia, malignant fibrous histiocytoma, melanoma,Merkel cell carcinoma, mesothelioma, metastatic squamous neck cancer with occult primary, midline tract carcinoma involving NUT gene, mouth cancer, multiple endocrine neoplasia syndrome, multiple myeloma, mycosis fungoides, myelodysplastic syndrome, myelodysplastic / myeloproliferative neoplasm, nasal cavity and par nasal sinus cancer, nasopharyngeal cancer, neuroblastoma, non-Hodgkin lymphoma, non-small cell lung cancer, oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, papillomatosis, paraganglioma, parathyroid cancer, penile cancer, pharyngeal cancer, pheochromocytomas, pituitary tumor, pleuropulmonary blastoma, primary central nervous system lymphoma, prostate cancer, rectal cancer, renal cell cancer, renal pelvis and ureter cancer, retinoblastoma, rhabdoid tumor, salivary gland cancer, Sezary syndrome, skin cancer, small cell lung cancer, small intestine cancer, soft tissue sarcoma, spinal cord tumor, stomach cancer, T-cell lymphoma, teratoid tumor, testicular cancer, throat cancer, thymoma and thymic carcinoma, thyroid cancer, urethral cancer, uterine cancer, vaginal cancer, vulvar cancer, and Wilms tumor. In some embodiments, the cancerous tumor is selected from a melanoma, leukemia, lymphoma, or breast cancer. The biological sample can contain cells, tissue, and / or genetic material from an adenoma, Barrett’s esophagus, a benign cancer, a cervical intraepithelial neoplasia (CIN), chronic lymphocytic leukemia (CLL), a clonal hematopoiesis of indeterminate potential (CHIP), a colorectal polyp, a monoclonal gammopathy of undetermined significance (MGUS), a precancer, a prostate cancer, smoldering multiple myeloma (SMM), or a combination of any of the foregoing. The biological sample can contain cells, tissue, and / or genetic material from an adenocarcinoma, esophageal cancer, cervical cancer, myelodysplastic syndrome, acute myeloid leukemia, colorectal cancer, myeloma (e.g., multiple myeloma, active myeloma, light chain myeloma, non-secretory myeloma, solitary plasmacytoma, immunoglobulin G myeloma, immunoglobulin A myeloma, immunoglobulin M myeloma, immunoglobulin E myeloma, immunoglobulin D myeloma), breast cancer, or a combination of any of the foregoing. The biological sample can contain cells, tissue, and / or genetic material from a primary tumor or secondary (e.g., metastatic) tumor of a patient.

[0041] Any sequencing technique can be used in the methods described herein. Suitable sequencing techniques include, but are not limited to, whole genome sequencing (WGS), shotgun metagenomic sequencing, whole exome sequencing (WES), transcriptome sequencing,next-generation sequencing (NGS), cancer personalized profiling by deep sequencing (CAPP-Seq), tagged-amplicon deep sequencing (Tam-Seq), single nucleotide polymorphism (SNP) array, single cell sequencing, or combinations of any of the foregoing. The sequencing technique can be long-read sequencing (e.g. PacBio, Oxford Nanopore) or short-read sequencing (e.g. Illumina).

[0042] A patient can have (or a biological sample can contain genetic material from) any type, grade, or stage of cancer. The cancer can be any stage of cancer, including, but not limited to, stage 0, stage I, stage II, stage III, or stage IV The patient can have, or the biological sample can contain genetic material from, a carcinoma in-situ. For example, the biological sample can contain tissue from a ductal carcinoma in-situ (DCIS). The cancer can be in remission (e.g., partial remission). The cancer can be a relapsed tumor (e.g., a tumor of a relapsed cancer). The cancer can have any grade, including, but not limited to, X, 1, 2, 3, or 4. The cancer can be a recalcitrant tumor. The cancerous tumor can be resistant to therapy (e.g., resistant to chemotherapy, resistant to immunotherapy). The cancer can be susceptible to therapy (e.g., susceptible to chemotherapy, susceptible to immunotherapy, susceptible to radiotherapy). The patient can have, or the biological sample can contain tissue, cells, and / or genetic material from a pre-cancerous tumor or a benign tumor (e.g., a benign tumor at risk for progressing to a cancerous tumor).

[0043] This disclosure relates to a dual assay setup composed of a discovery assay and a confirmation assay to reduce the false positive rate (i.e. increase specificity), while providing compacity for longitudinal tracking of variants before, during and / or after treatment. Without wishing to be bound by any particular mechanism or theory, the discovery assay relates to identifying variants through maximizing sensitivity by increasing the number of variant sites while decreasing the sequencing depth. Without wishing to be bound by any particular mechanism or theory, the confirmation assay relates to prioritizing a select number of variants, identified using the sequencing results of the discovery assay, on a personalized panel and sequencing at high sequencing depth.

[0044] Disclosed herein are methods for identifying actionable variants in liquid biopsies. The methods disclosed herein comprise sequencing cell-free (cf)DNA, cell-free (cf)RNA or a combination thereof from a liquid biopsy sample from a subject. The cfDNA and / or cfRNA issequenced to obtain sequence data. The sequence data comprises sequence reads of a plurality of polynucleotides from the subject. The cfDNA sequences and / or cfRNA sequences are used to analyzed actionable variants. In some embodiments, the cfDNA comprises circulating tumor (ct)DNA.

[0045] The sequence data can be obtained from whole genome sequencing, whole exome sequencing, targeted sequencing, DNA hybridization methods or combinations thereof. The genome sequencing data can be sequence data derived from high-depth whole genome sequencing data.

[0046] In one aspect, the method of detecting actionable variants in a liquid biopsy can comprise the steps of: (a) obtaining a test sample from a liquid biopsy, (b) running a discovery assay, wherein the discovery assay comprises sequencing the test sample, (c) identifying a list of prioritized variants, (d) designing a personalized sequencing panel based on the list of prioritized variants, (e) running a confirmation assay, wherein the confirmation assay comprises sequencing the test sample using the personalized sequencing panel and (f) detecting one or more actionable variants in the liquid biopsy.

[0047] Methods described herein can repeat one or more steps of the method. Any step of the method can be repeated. For example, the step of sequencing genetic material from a biological sample collected from a patient to obtain a sequencing result can be repeated for multiple (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10) samples. Repeated steps of a method can be practiced in any order (c. ., repeated steps can be practiced in succession, repeated steps can be practiced not in succession (e. ., disconnected, with different steps in between the repeated steps)).

[0048] Methods described herein and steps of methods described herein can be repeated any number of times. Methods can be repeated any number of times during the course of treatment of a patient. The number of repetitions of the method can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14. The number of repetitions of the method can be at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, or at least 14. The number of repetitions of the method can be at most 1, at most 2, at most 3, at most 4, at most 5, at most 6, at most 7, at most 8, at most 9, at most 10, at most 11, at most 12, at most 13, or at most 14. The number of repetitions of the method can be between about 1 to about 10, about 1 to about 9, about 1 to about 8, about 1 to about 7, about 1 to about 6, about 1 to about5, about 1 to about 4, about 1 to about 3, or about 1 to about 2. For example, the method of predicting a sequencing depth, batch size, and / or amount of biological sample needed to obtain a desired number of neoantigens can be repeated twice: once for each of two biological samples.

[0049] The number of repetitions of a step can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14. The number of repetitions of a step can be at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, or at least 14. The number of repetitions of a step can be at most 1, at most 2, at most 3, at most 4, at most 5, at most 6, at most 7, at most 8, at most 9, at most 10, at most 11, at most 12, at most 13, or at most 14. The number of repetitions of a step can be between about 1 to about 10, about 1 to about 9, about 1 to about 8, about 1 to about 7, about 1 to about 6, about 1 to about 5, about 1 to about 4, about 1 to about 3, or about 1 to about 2.

[0050] Steps of the method, repetitions of the method, and / or repetitions of steps of the method can be separated by any time interval. The time interval can be about 30 seconds, about 1 minute, about 2 minutes, about 3 minutes, about 4 minutes, about 5 minutes, about 6 minutes, about 7 minutes, about 8 minutes, about 9 minutes, about 10 minutes, about 11 minutes, about 12 minutes, about 13 minutes, about 14 minutes, about 15 minutes, about 16 minutes, about 17 minutes, about 18 minutes, about 19 minutes, about 20 minutes, about 22 minutes, about 24 minutes, about 26 minutes, about 28 minutes, about 30 minutes, about 33 minutes, about 36 minutes, about 39 minutes, about 40 minutes, about 44 minutes, about 48 minutes, about 50 minutes, about 55 minutes, about 1 hour, about 1.5 hours, about 2 hours, about 2.5 hours, about 3 hours, about 3.5 hours, about 4 hours, about 4.5 hours, about 5 hours, about 5.5 hours, about 6 hours, about 6.5 hours, about 7 hours, about 7.5 hours, about 8 hours, about 8.5 hours, about 9 hours, about 9.5 hours, about 10 hours, about 10.5 hours, about 11 hours, about 11.5 hours, about 12 hours, about 12.5 hours, about 13 hours, about 13.5 hours, about 14 hours, about 14.5 hours, about 15 hours, about 15.5 hours, about 16 hours, about 16.5 hours, about 17 hours, about 17.5 hours, about 18 hours, about 18.5 hours, about 19 hours, about 19.5 hours, about 20 hours, about 20.5 hours, about 21 hours, about 21.5 hours, about 22 hours, about 22.5 hours, about 23 hours, about 23.5 hours, about 1 day, about 1.5 days, about 2 days, about 3 days, about 4 days, about 5 days, about 6 days, about 1 week, about 2 weeks, about 3 weeks, about 4 weeks, about 1 month, about 1.5 months, about 2 months, about 2.5 months, about 3 months, about 3.5 months, about 4months, about 4.5 months, about 5 months, about 5.5 months, about 6 months, about 6.5 months, about 7 months, about 7.5 months, about 8 months, about 8.5 months, about 9 months, about 9.5 months, about 10 months, about 10.5 months, about 11 months, about 11.5 months, about 1 year, about 1.5 years, about 2 years, about 2.5 years, about 3 years, about 3.5 years, about 4 years, about 4.5 years, or. about 5 years. The time interval can be at least 30 seconds, at least 1 minute, at least 2 minutes, at least 3 minutes, at least 4 minutes, at least 5 minutes, at least 6 minutes, at least 7 minutes, at least 8 minutes, at least 9 minutes, at least 10 minutes, at least 11 minutes, at least 12 minutes, at least 13 minutes, at least 14 minutes, at least 15 minutes, at least 16 minutes, at least 17 minutes, at least 18 minutes, at least 19 minutes, at least 20 minutes, at least 22 minutes, at least 24 minutes, at least 26 minutes, at least 28 minutes, at least 30 minutes, at least 33 minutes, at least 36 minutes, at least 39 minutes, at least 40 minutes, at least 44 minutes, at least 48 minutes, at least 50 minutes, at least 55 minutes, at least 1 hour, at least 1.5 hours, at least 2 hours, at least 2.5 hours, at least 3 hours, at least 3.5 hours, at least 4 hours, at least 4.5 hours, at least 5 hours, at least 5.5 hours, at least 6 hours, at least 6.5 hours, at least 7 hours, at least 7.5 hours, at least 8 hours, at least 8.5 hours, at least 9 hours, at least 9.5 hours, at least 10 hours, at least 10.5 hours, at least 11 hours, at least 11.5 hours, at least 12 hours, at least 12.5 hours, at least 13 hours, at least 13.5 hours, at least 14 hours, at least 14.5 hours, at least 15 hours, at least 15.5 hours, at least 16 hours, at least 16.5 hours, at least 17 hours, at least 17.5 hours, at least 18 hours, at least 18.5 hours, at least 19 hours, at least 19.5 hours, at least 20 hours, at least 20.5 hours, at least 21 hours, at least 21.5 hours, at least 22 hours, at least 22.5 hours, at least 23 hours, at least 23.5 hours, at least 1 day, at least 1.5 days, at least 2 days, at least 3 days, at least 4 days, at least 5 days, at least 6 days, at least 1 week, at least 2 weeks, at least 3 weeks, at least 4 weeks, at least 1 month, at least 1.5 months, at least 2 months, at least 2.5 months, at least 3 months, at least 3.5 months, at least 4 months, at least 4.5 months, at least 5 months, at least 5.5 months, at least 6 months, at least 6.5 months, at least 7 months, at least 7.5 months, at least 8 months, at least 8.5 months, at least 9 months, at least 9.5 months, at least 10 months, at least 10.5 months, at least 11 months, at least 11.5 months, at least 1 year, at least 1.5 years, at least 2 years, at least 2.5 years, at least 3 years, at least 3.5 years, at least 4 years, at least 4.5 years, or at least 5 years. The time interval can be at most 30 seconds, at most 1 minute, at most 2 minutes, at most 3 minutes, at most 4 minutes, at most 5 minutes, at most 6 minutes, atmost 7 minutes, at most 8 minutes, at most 9 minutes, at most 10 minutes, at most 11 minutes, at most 12 minutes, at most 13 minutes, at most 14 minutes, at most 15 minutes, at most 16 minutes, at most 17 minutes, at most 18 minutes, at most 19 minutes, at most 20 minutes, at most 22 minutes, at most 24 minutes, at most 26 minutes, at most 28 minutes, at most 30 minutes, at most 33 minutes, at most 36 minutes, at most 39 minutes, at most 40 minutes, at most 44 minutes, at most 48 minutes, at most 50 minutes, at most 55 minutes, at most 1 hour, at most 1.5 hours, at most 2 hours, at most 2.5 hours, at most 3 hours, at most 3.5 hours, at most 4 hours, at most 4.5 hours, at most 5 hours, at most 5.5 hours, at most 6 hours, at most 6.5 hours, at most 7 hours, at most 7.5 hours, at most 8 hours, at most 8.5 hours, at most 9 hours, at most 9.5 hours, at most 10 hours, at most 10.5 hours, at most 11 hours, at most 11.5 hours, at most 12 hours, at most 12.5 hours, at most 13 hours, at most 13.5 hours, at most 14 hours, at most 14.5 hours, at most 15 hours, at most 15.5 hours, at most 16 hours, at most 16.5 hours, at most 17 hours, at most 17.5 hours, at most 18 hours, at most 18.5 hours, at most 19 hours, at most 19.5 hours, at most 20 hours, at most 20.5 hours, at most 21 hours, at most 21.5 hours, at most 22 hours, at most 22.5 hours, at most 23 hours, at most 23.5 hours, at most 1 day, at most 1.5 days, at most 2 days, at most 3 days, at most 4 days, at most 5 days, at most 6 days, at most 1 week, at most 2 weeks, at most 3 weeks, at most 4 weeks, at most 1 month, at most 1.5 months, at most 2 months, at most 2.5 months, at most 3 months, at most 3.5 months, at most 4 months, at most 4.5 months, at most 5 months, at most 5.5 months, at most 6 months, at most 6.5 months, at most 7 months, at most 7.5 months, at most 8 months, at most 8.5 months, at most 9 months, at most 9.5 months, at most 10 months, at most 10.5 months, at most 11 months, at most 11.5 months, at most 1 year, at most 1.5 years, at most 2 years, at most 2.5 years, at most 3 years, at most 3.5 years, at most 4 years, at most 4.5 years, or at most 5 years. The time interval can be between about 1 minute and about 5 years, about 1 minute and about 4 years, about 1 minute and about 3 years, about 1 minute and about 2 years, about 1 minute and about 1 year, about 1 minute and about 6 months, about 1 minute and about 3 months, about 1 minute and about 1 month, about 1 minute and about 1 week, about 1 minute and about 1 day, about 1 minute and about 1 hour, about 1 hour and about 5 years, about 1 hour and about 4 years, about 1 hour and about 3 years, about 1 hour and about 2 years, about 1 hour and about 1 year, about 1 hour and about 6 months, about 1 hour and about 3 months, about 1 hour and about 1 month, about 1 hour and about 1week, about 1 hour and about 1 day, about 1 day and about 5 years, about 1 day and about 4 years, about 1 day and about 3 years, about 1 day and about 2 years, about 1 day and about 1 year, about 1 day and about 6 months, about 1 day and about 3 months, about 1 day and about 1 month, about 1 day and about 1 week, about 1 month and about 5 years, about 1 month and about 4 years, about 1 month and about 3 years, about 1 month and about 2 years, about 1 month and about 1 year, about 1 month and about 6 months, about 1 month and about 3 months, about 1 year and about 5 years, about 1 year and about 4 years, about 1 year and about 3 years, or about 1 year and about 2 years.

[0051] Methods described herein can include a step of obtaining a test sample from a liquid biopsy. The liquid biopsy can be a blood sample, fecal sample, saliva sample or the like. For example, the liquid biopsy can be a blood sample. The test sample can be genetic material obtained from the blood sample. For example, the test sample can obtain genetic material from plasma isolated from blood. For example, the genetic material can be deoxyribonucleic acid (DNA, e.g., an oligonucleotide containing 2’ -deoxyribonucleotides). The DNA can be cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), DNA from circulating tumor cells (CTCs), mitochondrial DNA, nuclear DNA (i.e., DNA from the nucleus of a cell), complementary DNA (cDNA), or a combination of any of the foregoing. The genetic material can contain ribonucleic acid (RNA, e.g., an oligonucleotide containing ribonucleotides). The RNA enetic material can be messenger RNA (mRNA), short-interfering RNA (siRNA), microRNA (miRNA), circular RNA (circRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), small nucleolar RNA (snRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), sub-genomic RNA (sgRNA), RNA from integrating or non-integrating viruses, a combination of any of the foregoing, or any other RNA. In some embodiments, the genetic material contains a nucleic acid selected from DNA, RNA, or a combination thereof. In some embodiments, the genetic material comprises a nucleic acid selected from cell free DNA (cfDNA), circulating tumor DNA (ctDNA), cell free RNA (cfRNA), circulating tumor RNA (ctRNA), or a combination of any of the foregoing isolated from a plasma sample obtained from a blood sample.

[0052] Methods described herein can include a step of running a discovery assay. The discovery assay relates to running a sequencing protocol on a test sample from a liquid biopsy. The test sample can be genetic material. For example, the genetic material can be cfDNA, cfRNA, or acombination thereof. The discovery assay sequencing protocol institutes parameters (e.g. sequencing depth, batch size), which can result in high sensitivity to detect variants. As disclosed herein high sensitivity can be a sensitivity of about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or about 100%. High sensitivity can be a sensitivity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%. High sensitivity can be a sensitivity of at most 50%, at most 55%, at most 60%, at most 65%, at most 70%, at most 75%, at most 80%, at most 85%, at most 90%, at most 95%, or at most 100%. High sensitivity can be a sensitivity between about 50% to about 100%, between about 50% to about 95%, between about 50% to about 90%, between about 50% to about 85%, between about 50% to about 80%, between about 50% to about 75%, between about 50% to about 70%, between about 50% to about 65%, between about 50% to about 60%, between about 50% to about 55%, between about 55% to about 100%, between about 60% to about 100%, between about 65% to about 100%, between about 70% to about 100%, between about 75% to about 100%, between about 80% to about 100%, between about 85% to about 100%, between about 90% to about 100%, or between about 95% to about 100% between about 96% and 100%, between about 97% to about 100%, between about 98% to about 100%, or between about 99% to about 100%.

[0053] Methods described herein can include running a discovery assay wherein the sequencing protocol comprises a variant panel. For example, the variant panel can be a whole genome panel, a whole exome panel, or a custom panel (e.g. targeted gene panel). As disclosed herein the variant panel for the discovery assay can comprise a high number of variant sites. The discovery assay variant panel can comprise a high number of variant sites of about 10000, about 20000, about 30000, about 40000, about 50000, about 60000, about 70000, about 80000, about 90000, about 100000, about 150000, about 200000, about 250000, about 300000, about 350000, about 400000, about 450000, about 500000, about 550000, about 600000, about 6500000, about 700000, about 750000, about 800000, about 850000, about 900000, about 950000, about 1000000, about 2000000, about 3000000, about 4000000, about 5000000, about 6000000, about 7000000, about 8000000, about 9000000, about 10000000, about 20000000, about 30000000, about 40000000, about 50000000, about 60000000, about 70000000, about 80000000, about 90000000, or about 100000000. The discovery assay variant panel can comprise a high numberof variant sites of at least 10000, at least 20000, at least 30000, at least 40000, at least 50000, at least 60000, at least 70000, at least 80000, at least 90000, at least 100000, at least 150000, at least 200000, at least 250000, at least 300000, at least 350000, at least 400000, at least 450000, at least 500000, at least 550000, at least 600000, at least 6500000, at least 700000, at least 750000, at least 800000, at least 850000, at least 900000, at least 950000, at least 1000000, at least 2000000, at least 3000000, at least 4000000, at least 5000000, at least 6000000, at least 7000000, at least 8000000, at least 9000000, at least 10000000, at least 20000000, at least 30000000, at least 40000000, at least 50000000, at least 60000000, at least 70000000, at least 80000000, at least 90000000, or at least 100000000. The discovery assay variant panel can comprise a high number of variant sites of at most 10000, at most 20000, at most 30000, at most 40000, at most 50000, at most 60000, at most 70000, at most 80000, at most 90000, at most 100000, at most 150000, at most 200000, at most 250000, at most 300000, at most 350000, at most 400000, at most 450000, at most 500000, at most 550000, at most 600000, at most 6500000, at most 700000, at most 750000, at most 800000, at most 850000, at most 900000, at most 950000, at most 1000000, at most 2000000, at most 3000000, at most 4000000, at most 5000000, at most 6000000, at most 7000000, at most 8000000, at most 9000000, at most 10000000, at most 20000000, at most 30000000, at most 40000000, at most 50000000, at most 60000000, at most 70000000, at most 80000000, at most 90000000, or at most 100000000.

[0054] Methods described herein can include running a discovery assay wherein the sequencing protocol has a low sequencing depth. Without wishing to be bound by a particular mechanism or theory, a low sequencing depth allows the discovery assay to identify a large number of variants. The sequencing depth of the discovery assay can be about 10X, about 20X, about 30X, about 40X, about 50X, about 60X, about 70X, about 80X, about 90X, about 100X, about 11 OX, about 120X, about 130X, about 140X, about 150X, about 160X, about 170X, about 180X, about 190X, about 200X, about 210X, about 220X, about 230X, about 240X, about 250X, about 260X, about 270X, about 280X, about 290X, about 300X, about 350X, about 400X, about 450X, about 500X, about 550X, about 600X, about 650X, about 700X, about 750X, about 800X, about 850X, about 900X, about 950X, about 1000X, about 1500X, about 2000X, about 2500X, about 3000X, about 3500X, about 4000X, about 4500X, about 5000X, about 5500X, about 6000X, about 6500X, about 7000X, about 7500X, about 8000X, about 8500X, about 9000X, about 9500X, or about10000X. The sequencing depth of the discovery assay can be at least 10X, at least 20X, at least 3OX, at least 40X, at least 5OX, at least 60X, at least 70X, at least 80X, at least 90X, at least 1OOX, at least 11OX, at least 120X, at least 13OX, at least 140X, at least 15OX, at least 160X, at least 170X, at least 180X, at least 190X, at least 200X, at least 210X, at least 220X, at least 230X, at least 240X, at least 250X, at least 260X, at least 270X, at least 280X, at least 290X, at least 3OOX, at least 35OX, at least 400X, at least 450X, at least 5OOX, at least 55OX, at least 600X, at least 65OX, at least 700X, at least 750X, at least 8OOX, at least 850X, at least 900X, at least 95OX, at least 1OOOX, at least 15OOX, at least 2000X, at least 2500X, at least 3OOOX, at least 35OOX, at least 4000X, at least 45OOX, at least 5OOOX, at least 55OOX, at least 6000X, at least 6500X, at least 7000X, at least 75OOX, at least 8OOOX, at least 85OOX, at least 9000X, at least 9500X, or at least 1OOOOX. The sequencing depth of the discovery assay can be at most 1OX, at most 20X, at most 3OX, at most 40X, at most 5OX, at most 60X, at most 70X, at most 80X, at most 90X, at most 1OOX, at most 11OX, at most 120X, at most I3OX, at most 140X, at most 15OX, at most 160X, at most 170X, at most 18OX, at most 190X, at most 200X, at most 21 OX, at most 220X, at most 23OX, at most 240X, at most 25OX, at most 260X, at most 270X, at most 280X, at most 290X, at most 3OOX, at most 35OX, at most 400X, at most 450X, at most 500X, at most 55OX, at most 600X, at most 65OX, at most 700X, at most 750X, at most 800X, at most 85OX, at most 900X, at most 95OX, at most 1OOOX, at most 15OOX, at most 2000X, at most 2500X, at most 3000X, at most 35OOX, at most 4000X, at most 4500X, at most 5000X, at most 5500X, at most 6000X, at most 6500X, at most 7000X, at most 7500X, at most 8000X, at most 85OOX, at most 9000X, at most 9500X, or at most 1OOOOX. The sequencing depth of the discovery assay can be between about 200X to about 10000X, about 3OOX to about 10000X, about 400X to about 1OOOOX, about 5OOX to about 10000X, about 600X to about 10000X, about 700X to about 10000X, about 800X to about 10000X, about 900X to about 10000X, about 1000X to about 10000X, about 200X to about 9000X, about 200X to about 8000X, about 200X to about 7000X, about 200X to about 6000X, about 200X to about 5000X, about 200X to about 4000X, about 200X to about 3000X, about 200X to about 2000X, or about 200X to about 1000X. For example, the sequencing depth of the discovery assay can be between about 200X to about 5000X.

[0055] Methods described herein include identification of variants from genetic material isolated from liquid biopsy samples collected from a patient. Without wishing to be bound by any particular mechanism of operation or theory, sequencing variation (e.g. genetic variants) can be identified using a variant caller. The variant caller can be used to identify variants on germline, somatic and population-level sequencing. The variant caller can be used to identify variants using short-read sequencing data. Assembly of short-read sequencing data sets can use any de novo assembly approach known in the art. For example, de Bruijn graph-based assembly (e.g. DRAGEN, Mutect2, Velvet, ABySS, Clover), overlap layout consensus (OLC) (e.g. Edena), string graph (e.g. SGA), greedy (e.g. SSAKE, Perga), and / or hybrid assembler algorithms (e.g. Ray). Short-read sequencing data can be paired-end or single-end data sets. The structural variant caller can be used to identify SVs using long-read sequencing data (e.g. Sv AB A, PBSV, Sniffles2). Long-read sequencing can be continuous long read (CLR), high-fidelity (HiFi), and / or Oxford Nanopore Technologies (ONT) sequencing. Assembly of long-read sequencing data sets can use any de novo assembly approach known in the art. For example, Canu, Flye, Miniasm, Raven and / or wtdbg2 for ONT and CLR reads, or HiCanu, Hifiasm, LJA and / or MBG for HiFi reads.

[0056] Methods described herein include use of a variant caller (i.e. a bioinformatics tool) to identify variants and quantify a score based on the confidence or quality of the variant called (e g. confident score, quality score). Without wishing to be bound by any particular mechanism of operation or theory, the confidence or quality score can be a scaled probability score that quantifies the probability for a given variant call that the tumor genotype and normal genotype are different. For example, a low probability score can indicate no difference between the tumor and normal genotype, while a higher probability score indicates the genotypes are different. Any variant caller coded to include a confidence or quality score filter can be used to prioritize variants called. In some embodiments, the variant caller can be DRAGEN and the score can be a Somatic Quality score. In some embodiments, the variant caller can be Mutect2 and the score can be tumor-likelihood (TLOD).

[0057] Methods described herein include use of DRAGEN, a variant caller bioinformatic tool capable of quantifying a score based on the confidence or quality of the variant called. For example, DRAGEN quantifies a Somatic Quality (SQ) score. Without wishing to be bound byany particular mechanism of operation or theory, the SQ filter can be adjusted to filter low confidence variant calls, thereby decreasing false positive calls. The SQ filter threshold may be adjusted in DRAGEN, wherein the threshold is increased or decreased to filter out variants based on the desired quality and / or quantity of variants called. In some embodiments, sorted and aligned human-readable text files (e.g. BAM, SAM, CRAM) comprising tumor sequencing data or tumor sequencing data and normal sequencing data can be 1) read to establish active regions, 2) assembled locally using a De Brujin graph to generate candidate haplotypes, 3) align haplotypes against a human reference, 4) quantify number of reads supporting a variant, 5) quantify a likelihood for each haplotype identified based on sample specific noise levels, and 6) calculate a posterior probability of the variant being present. Based on the posterior probability the variant will either be filtered out (i.e. SQ is below threshold filter set) or outputted as a variant called (i.e. SQ is above threshold filter set) as described above.

[0058] Methods described herein include use of Mutect2 a variant caller bioinformatic tool capable of quantifying a score based on the confidence or quality of the variant called. For example, Mutect2 quantifies tumor-likelihood (TLOD). Without wishing to be bound by any particular mechanism of operation or theory, Mutect2 can use a predetermined filter to adjusted variant calls based on confidence and quality, thereby decreasing false positive calls. The TLOD threshold may be adjusted in Mutect2, wherein the threshold is increased or decreased to filter out variants based on the desired quality and / or quantity of variants called. In some embodiments, sorted and aligned human-readable text files (e.g. BAM, SAM, CRAM) comprising tumor sequencing data or tumor sequencing data and normal sequencing data are inputted, allele sets identified and matched, and the log evidence ratio of an allele set containing all alleles compared to an allele set excluding that allele (i.e. log odds that an allele exists). The TLOD score can be adjusted based on number of variants present in a normal allele set as described above. For example, the more the called variant appears in the normal alleles the lower the confidence score gets as a result.

[0059] Methods described herein can include creating a list of prioritized variants based on a variant callers confidence or quality score. As described above, variants called can be filtered out based on confidence or quality that the variant called is a true positive (i.e. a tumor-derived variant). The list of prioritized variants comprise a select number of variants. The select numberof variants can be about 1 , about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 150, about 200, about 250, about 300, about 350, about 400, about 450, about 500, about 550, about 600, about 650, about 700, about 750, about 800, about 850, about 900, about 950, about 1000, about 1500, about 2000, about 2500, about 3000, about 3500, about 4000, about 4500, about 5000, about 5500, about 6000, about 6500, about 7000, about 7500, about 8000, about 8500, about 9000, about 9500, or about 10000. The select number of variants can be at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1500, at least 2000, at least 2500, at least 3000, at least 3500, at least 4000, at least 4500, at least 5000, at least 5500, at least 6000, at least 6500, at least 7000, at least 7500, at least 8000, at least 8500, at least 9000, at least 9500, or at least 10000. The select number of variants can be at most 1, at most 2, at most 3, at most 4, at most 5, at most 6, at most 7, at most 8, at most 9, at most 10, at most 15, at most 20, at most 25, at most 30, at most 35, at most 40, at most 45, at most 50, at most 55, at most 60, at most 65, at most 70, at most 75, at most 80, at most 85, at most 90, at most 95, at most 100, at most 150, at most 200, at most 250, at most 300, at most 350, at most 400, at most 450, at most 500, at most 550, at most 600, at most 650, at most 700, at most 750, at most 800, at most 850, at most 900, at most 950, at most 1000, at most 1500, at most 2000, at most 2500, at most 3000, at most 3500, at most 4000, at most 4500, at most 5000, at most 5500, at most 6000, at most 6500, at most 7000, at most 7500, at most 8000, at most 8500, at most 9000, at most 9500, or at most 10000. The select number of variants can be between about 100 to about 10000, about 200 to about 10000, about 300 to about 10000, about 400 to about 10000, about 500 to about 10000, about 600 to about 10000, about 700 to about 10000, about 800 to about 10000, about 900 to about 10000, about 1000 to about 10000, about 2000 to about 10000, about 3000 to about 10000, about 4000 to about 10000, about 5000 to about 10000, about 6000 to about 10000, about 7000to about 10000, about 8000 to about 10000, about 9000 to about 10000, about 100 to about 9000, about 100 to about 8000, about 100 to about 7000, about 100 to about 6000, about 100 to about 5000, about 100 to about 4000, about 100 to about 3000, about 100 to about 2000, about 100 to about 1000, about 100 to about 900, about 100 to about 800, about 100 to about 700, about 100 to about 600, about 100 to about 500, about 100 to about 400, about 100 to about 300, or about 100 to about 200. In some embodiments, the select number of variants can be between about 100 to about 10000.

[0060] The method disclosed herein can include identification of a variant that can be further characterized using proteomic and / or immunopeptidomic approaches. Without wishing to be bound by any particular mechanism or theory, the variant can be translated to a peptide variant, wherein the peptide variant can bind to a major histocompatibility complex (MHC) molecule to form a complex or induce T cells to cross react. Peptide variants can be analyzed using thin layer chromatography, electrophoresis (e.g. capillary electrophoresis, solid phase extraction, reverse phase high performance liquid chromatography, amino acid analysis after acid hydrolysis, and fast atom bombardment mass spectroscopy), MALDI-TOF, and / or ESI-Q-TOF mass spectroscopy.

[0061] The methods disclosed herein can include a confirmation assay. The confirmation assay as described herein follows the discovery assay, using the list of prioritized variants to design a personalized sequencing panel. The personalized panel can include variants up to a pre-specified number, as prioritized according to their confidence or score in the discovery assay. For DNA or RNA variants, the personalized panel can be manufactured by synthesizing and pooling a set of oligonucleotide probes (single stranded or double stranded) to allow target enrichment prior to sequencing. The set of oligonucleotide probes are then optimized to ensure uniform capture and minimal off-target effects. For peptide variants, the panel can consist of affinity-based (e.g. immunoprecipitation) or enrichment-based (e.g. chromatography, proteomics). The test sample can be run on the personalized sequencing panel at high sequencing depth in order to increase assay specificity (i.e. decrease false positive) variants called. For example, the personalized sequencing panel may comprise about 100 to about 5000 variants. For example, the personalized sequencing panel may comprise about 5000 variants.

[0062] As described above, the methods can include a step of running a confirmation assay. The confirmation assay relates to running a sequencing protocol on a test sample from a liquid biopsy. The test sample can be genetic material. For example, the genetic material can be cfDNA, cfRNA, or a combination thereof. The confirmation assay sequencing protocol institutes parameters (e.g. sequencing depth, batch size), which can result in high specificity to detect actionable variants. As disclosed herein high specificity can be a specificity of about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%o, or about 100%. High specificity can be a specificity of at least 50%, at least 55%o, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%. High specificity can be a specificity of at most 50%, at most 55%, at most 60%, at most 65%, at most 70%, at most 75%, at most 80%, at most 85%, at most 90%, at most 95%, or at most 100%. High specificity can be a specificity between about 50% to about 100%, between about 50% to about 95%, between about 50% to about 90%, between about 50% to about 85%, between about 50% to about 80%, between about 50% to about 75%, between about 50%o to about 70%, between about 50% to about 65%, between about 50% to about 60%o, between about 50% to about 55%, between about 55% to about 100%, between about 60% to about 100%, between about 65% to about 100%, between about 70% to about 100%, between about 75% to about 100%, between about 80% to about 100%, between about 85% to about 100%, between about 90% to about 100%, between about 95%o to about 100%, between about 96% and 100%, between about 97% to about 100%, between about 98% to about 100%, or between about 99% to about 100%.

[0063] This disclosure relates to calculated sensitivity and specificity. Calculations of the sensitivity value uses formula (I): ( -T7ue Positlve- x 100; calculations of the \True Positive+False Negative / 100.

[0064] The methods disclosed herein can include a confirmation assay to decrease the rate of false positive identification by increase the raw sequencing depth. High sequencing depth, as described herein for the confirmation assay, increases the specificity as increasing the number of reads on a specific base or position leads to decreased identification of false positives. The sequencing depth of the confirmation assays can have a high sequencing depth of about 10000X,20000X, about 30000X, about 40000X, about 50000X, about 60000X, about 70000X, about 80000X, about 90000X, about 100000X, about 200000X, about 300000X, about 400000X, about 500000X, about 600000X, about 700000X, about 800000X, about 900000X, or about 1000000X. The sequencing depth of the confirmation assays can have a high sequencing depth of at least 10000X, at least 20000X, at least 30000X, at least 40000X, at least 50000X, at least 60000X, at least 70000X, at least 80000X, at least 90000X, at least 100000X, at least 200000X, at least 300000X, at least 400000X, at least 500000X, at least 600000X, at least 700000X, at least 800000X, at least 900000X, or at least 1000000X. The sequencing depth of the confirmation assays can have a high sequencing depth of at most 10000X, at most 20000X, at most 30000X, at most 40000X, at most 50000X, at most 60000X, at most 70000X, at most 80000X, at most 90000X, at most 100000X, at most 200000X, at most 300000X, at most 400000X, at most 500000X, at most 600000X, at most 700000X, at most 800000X, at most 900000X, or at most 1000000X. The sequencing depth of the confirmation assays can have a high sequencing depth between about 10000X to about 100000X, about 20000X to about 100000X, about 30000X to about 100000X, about 40000X to about 100000, about 50000X to about 100000X, about 60000 to about 100000X, about 70000X to about 100000X, about 80000X to about 100000X, about 90000X to about 100000X, about 10000X to about 90000X, about 10000X to about 80000X, about 10000X to about 70000X, about 10000X to about 60000X, about 10000X to about 50000X, about 10000X to about 40000X, about 10000X to about 30000X, or about 10000X to about 20000X. In some embodiments, the high sequencing depth chosen for the confirmation assay is 1000000X.

[0065] The method as disclosed herein can input the sequencing results of the confirmation assay into a variant caller. The variant caller, as described above, can call variants based on threshold set of a confidence or quality filter. As the high sequencing depth used in the confirmation assay allows for high specificity, it is to be expected that variants called are true positives. Therefore, variants discovered and identified from the confirmation assay can be actionable variants. Without wishing to be bound by a particular mechanism or theory, actionable variants (e.g. DNA mutations) can affect a patient’s response to treatment. For example, the presence of an actionable variant may support use of a specific course of treatment and / or, alternatively, remove specific treatment options. Further, as disclosed herein actionable variantscan be used as a diagnosis. For example, the identification of an actionable variant can identify a type of cancer. For example, the identification of an actionable variant can relate to evaluation of hereditary predisposition.

[0066] The method disclosed herein describes use of a test sample for a discovery assay and for a confirmation assay. The test sample is genetic material isolated from a liquid biopsy collected from a patient. The patient can have or can be suspected to have cancer. The test samples can be collected simultaneously or sequentially. The test sample can be collected at any time. For example, the test sample can be collected prior to diagnosis, prior to treatment, during treatment, during progression, after treatment, while in partial remission or while in complete remission.

[0067] In one aspect, the method of detecting actionable variants in a liquid biopsy collected from a patient having or suspected to have cancer, can comprise the steps of (a) obtaining a test sample comprising cell-free DNA or cell-free RNA from a liquid biopsy, (b) running a discovery assay on cell-free DNA, wherein the discovery assay comprises sequencing about 50000000 variant sites at a low sequencing depth between about 200X to about 5000X sequencing depth, (c) identifying a list of prioritized variants using a bioinformatic tool capable of quantifying a score based on confidence or quality of variants called, (d) designing a personalized sequencing panel based on the list of prioritized variants, (e) running a confirmation assay, wherein the confirmation assay comprises sequencing the cell-free DNA and / or the cell-free RNA using the personalized sequencing panel at 100000X sequencing depth (f) detecting variants using a variant caller and (g) identifying actionable variants from the detected variants.

[0068] Methods described herein can repeat one or more steps of the method. Any step of the method can be repeated. For example, the step of sequencing genetic material from a biological sample collected from a patient to obtain a sequencing result can be repeated for multiple (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10) samples. Repeated steps of a method can be practiced in any order (e. ., repeated steps can be practiced in succession, repeated steps can be practiced not in succession (e.g., disconnected, with different steps in between the repeated steps)).

[0069] Methods described herein and steps of methods described herein can be repeated any number of times. Methods can be repeated any number of times during the course of treatment of a patient. The number of repetitions of the method can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14. The number of repetitions of the method can be at least 1, at least 2, at least 3, at least 4, atleast 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, or at least 14. The number of repetitions of the method can be at most 1, at most 2, at most 3, at most 4, at most 5, at most 6, at most 7, at most 8, at most 9, at most 10, at most 11, at most 12, at most 13, or at most 14. The number of repetitions of the method can be between about 1 to about 10, about 1 to about 9, about 1 to about 8, about 1 to about 7, about 1 to about 6, about 1 to about 5, about 1 to about 4, about 1 to about 3, or about 1 to about 2. For example, the method of predicting a sequencing depth, batch size, and / or amount of biological sample needed to obtain a desired number of neoantigens can be repeated twice: once for each of two biological samples.

[0070] The number of repetitions of a step can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14. The number of repetitions of a step can be at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, or at least 14. The number of repetitions of a step can be at most 1, at most 2, at most 3, at most 4, at most 5, at most 6, at most 7, at most 8, at most 9, at most 10, at most 11, at most 12, at most 13, or at most 14. The number of repetitions of a step can be between about 1 to about 10, about 1 to about 9, about 1 to about 8, about 1 to about 7, about 1 to about 6, about 1 to about 5, about 1 to about 4, about 1 to about 3, or about 1 to about 2.

[0071] Steps of the method, repetitions of the method, and / or repetitions of steps of the method can be separated by any time interval. The time interval can be about 30 seconds, about 1 minute, about 2 minutes, about 3 minutes, about 4 minutes, about 5 minutes, about 6 minutes, about 7 minutes, about 8 minutes, about 9 minutes, about 10 minutes, about 11 minutes, about 12 minutes, about 13 minutes, about 14 minutes, about 15 minutes, about 16 minutes, about 17 minutes, about 18 minutes, about 19 minutes, about 20 minutes, about 22 minutes, about 24 minutes, about 26 minutes, about 28 minutes, about 30 minutes, about 33 minutes, about 36 minutes, about 39 minutes, about 40 minutes, about 44 minutes, about 48 minutes, about 50 minutes, about 55 minutes, about 1 hour, about 1.5 hours, about 2 hours, about 2.5 hours, about 3 hours, about 3.5 hours, about 4 hours, about 4.5 hours, about 5 hours, about 5.5 hours, about 6 hours, about 6.5 hours, about 7 hours, about 7.5 hours, about 8 hours, about 8.5 hours, about 9 hours, about 9.5 hours, about 10 hours, about 10.5 hours, about 11 hours, about 11.5 hours, about 12 hours, about 12.5 hours, about 13 hours, about 13.5 hours, about 14 hours, about 14.5 hours, about 15 hours, about 15.5 hours, about 16 hours, about 16.5 hours, about 17 hours, about 17.5hours, about 18 hours, about 18.5 hours, about 19 hours, about 19.5 hours, about 20 hours, about 20.5 hours, about 21 hours, about 21.5 hours, about 22 hours, about 22.5 hours, about 23 hours, about 23.5 hours, about 1 day, about 1.5 days, about 2 days, about 3 days, about 4 days, about 5 days, about 6 days, about 1 week, about 2 weeks, about 3 weeks, about 4 weeks, about 1 month, about 1.5 months, about 2 months, about 2.5 months, about 3 months, about 3.5 months, about 4 months, about 4.5 months, about 5 months, about 5.5 months, about 6 months, about 6.5 months, about 7 months, about 7.5 months, about 8 months, about 8.5 months, about 9 months, about 9.5 months, about 10 months, about 10.5 months, about 11 months, about 11.5 months, about 1 year, about 1.5 years, about 2 years, about 2.5 years, about 3 years, about 3.5 years, about 4 years, about 4.5 years, or. about 5 years. The time interval can be at least 30 seconds, at least 1 minute, at least 2 minutes, at least 3 minutes, at least 4 minutes, at least 5 minutes, at least 6 minutes, at least 7 minutes, at least 8 minutes, at least 9 minutes, at least 10 minutes, at least 11 minutes, at least 12 minutes, at least 13 minutes, at least 14 minutes, at least 15 minutes, at least 16 minutes, at least 17 minutes, at least 18 minutes, at least 19 minutes, at least 20 minutes, at least 22 minutes, at least 24 minutes, at least 26 minutes, at least 28 minutes, at least 30 minutes, at least 33 minutes, at least 36 minutes, at least 39 minutes, at least 40 minutes, at least 44 minutes, at least 48 minutes, at least 50 minutes, at least 55 minutes, at least 1 hour, at least 1.5 hours, at least 2 hours, at least 2.5 hours, at least 3 hours, at least 3.5 hours, at least 4 hours, at least 4.5 hours, at least 5 hours, at least 5.5 hours, at least 6 hours, at least 6.5 hours, at least 7 hours, at least 7.5 hours, at least 8 hours, at least 8.5 hours, at least 9 hours, at least 9.5 hours, at least 10 hours, at least 10.5 hours, at least 11 hours, at least 11.5 hours, at least 12 hours, at least 12.5 hours, at least 13 hours, at least 13.5 hours, at least 14 hours, at least 14.5 hours, at least 15 hours, at least 15.5 hours, at least 16 hours, at least 16.5 hours, at least 17 hours, at least 17.5 hours, at least 18 hours, at least 18.5 hours, at least 19 hours, at least 19.5 hours, at least 20 hours, at least 20.5 hours, at least 21 hours, at least 21.5 hours, at least 22 hours, at least 22.5 hours, at least 23 hours, at least 23.5 hours, at least 1 day, at least 1.5 days, at least 2 days, at least 3 days, at least 4 days, at least 5 days, at least 6 days, at least 1 week, at least 2 weeks, at least 3 weeks, at least 4 weeks, at least 1 month, at least 1.5 months, at least 2 months, at least 2.5 months, at least 3 months, at least 3.5 months, at least 4 months, at least 4.5 months, at least 5 months, at least 5.5 months, at least 6 months, at least 6.5 months, at least 7 months, at least 7.5months, at least 8 months, at least 8.5 months, at least 9 months, at least 9.5 months, at least 10 months, at least 10.5 months, at least 11 months, at least 11.5 months, at least 1 year, at least 1.5 years, at least 2 years, at least 2.5 years, at least 3 years, at least 3.5 years, at least 4 years, at least 4.5 years, or at least 5 years. The time interval can be at most 30 seconds, at most 1 minute, at most 2 minutes, at most 3 minutes, at most 4 minutes, at most 5 minutes, at most 6 minutes, at most 7 minutes, at most 8 minutes, at most 9 minutes, at most 10 minutes, at most 11 minutes, at most 12 minutes, at most 13 minutes, at most 14 minutes, at most 15 minutes, at most 16 minutes, at most 17 minutes, at most 18 minutes, at most 19 minutes, at most 20 minutes, at most 22 minutes, at most 24 minutes, at most 26 minutes, at most 28 minutes, at most 30 minutes, at most 33 minutes, at most 36 minutes, at most 39 minutes, at most 40 minutes, at most 44 minutes, at most 48 minutes, at most 50 minutes, at most 55 minutes, at most 1 hour, at most 1.5 hours, at most 2 hours, at most 2.5 hours, at most 3 hours, at most 3.5 hours, at most 4 hours, at most 4.5 hours, at most 5 hours, at most 5.5 hours, at most 6 hours, at most 6.5 hours, at most 7 hours, at most 7.5 hours, at most 8 hours, at most 8.5 hours, at most 9 hours, at most 9.5 hours, at most 10 hours, at most 10.5 hours, at most 11 hours, at most 11.5 hours, at most 12 hours, at most 12.5 hours, at most 13 hours, at most 13.5 hours, at most 14 hours, at most 14.5 hours, at most 15 hours, at most 15.5 hours, at most 16 hours, at most 16.5 hours, at most 17 hours, at most 17.5 hours, at most 18 hours, at most 18.5 hours, at most 19 hours, at most 19.5 hours, at most 20 hours, at most 20.5 hours, at most 21 hours, at most 21.5 hours, at most 22 hours, at most 22.5 hours, at most 23 hours, at most 23.5 hours, at most 1 day, at most 1.5 days, at most 2 days, at most 3 days, at most 4 days, at most 5 days, at most 6 days, at most 1 week, at most 2 weeks, at most 3 weeks, at most 4 weeks, at most 1 month, at most 1.5 months, at most 2 months, at most 2.5 months, at most 3 months, at most 3.5 months, at most 4 months, at most 4.5 months, at most 5 months, at most 5.5 months, at most 6 months, at most 6.5 months, at most 7 months, at most 7.5 months, at most 8 months, at most 8.5 months, at most 9 months, at most 9.5 months, at most 10 months, at most 10.5 months, at most 11 months, at most 11.5 months, at most 1 year, at most 1.5 years, at most 2 years, at most 2.5 years, at most 3 years, at most 3.5 years, at most 4 years, at most 4.5 years, or at most 5 years. The time interval can be between about 1 minute and about 5 years, about 1 minute and about 4 years, about 1 minute and about 3 years, about 1 minute and about 2 years, about 1 minute and about 1 year, about 1 minute andabout 6 months, about 1 minute and about 3 months, about 1 minute and about 1 month, about 1 minute and about 1 week, about 1 minute and about 1 day, about 1 minute and about 1 hour, about 1 hour and about 5 years, about 1 hour and about 4 years, about 1 hour and about 3 years, about 1 hour and about 2 years, about 1 hour and about 1 year, about 1 hour and about 6 months, about 1 hour and about 3 months, about 1 hour and about 1 month, about 1 hour and about 1 week, about 1 hour and about 1 day, about 1 day and about 5 years, about 1 day and about 4 years, about 1 day and about 3 years, about 1 day and about 2 years, about 1 day and about 1 year, about 1 day and about 6 months, about 1 day and about 3 months, about 1 day and about 1 month, about 1 day and about 1 week, about 1 month and about 5 years, about 1 month and about 4 years, about 1 month and about 3 years, about 1 month and about 2 years, about 1 month and about 1 year, about 1 month and about 6 months, about 1 month and about 3 months, about 1 year and about 5 years, about 1 year and about 4 years, about 1 year and about 3 years, or about 1 year and about 2 years.

[0072] The methods described herein can include identifying actionable variants from a list of detected variants. The detected variants can be identified from a prioritized list of variants and / or the variants identified from the confirmation assay. As described herein, the actionable variant can cause a change in response to treatment. The actionable variant can be expressed in the tumor. The actionable variant can modify the translated protein, referred to herein as a neoantigen. The neoantigen can be a candidate for a vaccine. In some embodiments the neoantigen is an immunogenic compound that can be used to generate a treatment. In some embodiments, more than one immunogenic compound can be used to generate a treatment.

[0073] The methods described herein, can relate to a patient having or suspected to have cancer. The patient can have any type of cancer, at any stage, as described above. The personalized sequencing panel can be used to monitor presence of actionable variants. Monitoring can be done before treatment, during treatment or after treatment. Without wishing to be bound by any particular mechanism or theory, minimal residual disease (MRD) monitoring is a diagnostic tool that can identify a small number of cancer cells that remain after treatment to assess response to therapy. MRD monitoring can also be employed to detect early relapse and / or guide treatment. The personalized sequencing panel can be used to assess MRD similarly. For example, thepersonalized sequencing panel can be used to monitor after treatment to evaluate response to treatment.

[0074] As disclosed herein, the personalized sequencing panel can be re-designed to include variants of interest or remove variants. Variants of interest can be identified by running subsequent discovery assays before, during or after treatment. For example, if a second discovery assay is run and new variants identified these can be included in the personalized sequencing panel. As disclosed above, the discovery assay can be run repeatedly, thus allowing for discovery of novel variants. Variants can be removed if the confirmation assay does not identify them as an actionable variant (i.e. not present in the test sample), as described above.Machine Learning Models

[0075] The methods described herein can include using a bioinformatic tool capable of quantifying a score based on the confidence or quality of variants called. As described herein, the bioinformatics tool can be a machine learning model trained to quantify the score based on the confidence or quality of the variant called. The machine learning model can operate by any type of machine learning methodology, including, but not limited to, supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, and combinations thereof.

[0076] Supervised learning models can involve training a machine learning model or its algorithm using labeled data. Labeled data can comprise a training data set labeled with input and output values. The labeled data can teach the machine what elements it needs to recognize and how to identify labeled elements from raw data. The labeled data can be fed and refed into the machine learning model to train the machine and increase accuracy in arriving at an output for a new data set. The user can provide feedback on accuracy of the machine algorithm.Supervised learning algorithms can be any supervised learning algorithm, including, but not limited to, classification algorithms (e.g., k-nearest neighbor (KNN), naive Bayes classifier algorithms, support vector machine (SVM) algorithms, decision trees, random forest models, logistic regression, Proaftn, stochastic gradient descent algorithms, linear classifiers, and combinations thereof) and regression algorithms (e.g., simple linear regression algorithms, multivariate regression algorithms, decision tree algorithms, lasso regression algorithms, logisticregression algorithms, ridge regression algorithms, and combinations thereof). In some embodiments, the machine learning model includes a classification algorithm. Any classification algorithm described herein can be used in the machine learning model. In some embodiments, the classification algorithm is selected from logistic regression, naive Bayes, K-nearest neighbor (KNN), decision tree, support vector machine, K-means clustering, random forest, artificial neural network (ANN), or a combination of any of the foregoing.

[0077] Unsupervised learning models can be trained on raw data without labels or tags.Unsupervised learning models can be used for clustering large, unstructured data sets.Unsupervised learning models can cluster data or perform dimensionality reduction of the data. Clustering examines similarities in raw data and groups the data by similarity, providing structure to unstructured raw data. Dimensionality reduction reduces the number of features in a data set, thereby reducing processing time, storage space, complexity, and overfitting in the machine learning model. Dimensionality reduction can include feature selection, wherein a subset of relevant features from the total original features present in the raw data are selected for use as an input. Dimensionality reduction can also include feature extraction, wherein a new set of features is generated from the existing data set. These new features can be used as further inputs for the machine learning model. Unsupervised learning models can operate under any unsupervised learning algorithm, including, but not limited to, clustering algorithms (e.g., hierarchical clustering, k-means clustering algorithm, density-based clustering, graph-based clustering, Cophenetic correlation, Bayesian information criterion, t-distributed stochastic neighbor embedding (t-SNE), uniform manifold approximation and projection (UMAP), and combinations thereof) and dimensionality reduction algorithms (e.g., principle component analysis (PCA), non-negative matrix factorization (NMF), linear discriminant analysis (LDA), generalized discriminant analysis (GDA), and combinations thereof).

[0078] Semi-supervised learning models can be a hybrid of supervised and unsupervised learning models. Semi-supervised learning models can utilize small amounts of labeled data processed alongside larger raw data sets. Any algorithm for semi-supervised learning can be used, including any unsupervised algorithm, any supervised algorithm, any modified unsupervised algorithm, and any modified supervised learning algorithm. Modified unsupervised and supervised algorithms include, but are not limited to, self-training algorithms (e.g., pseudo-labeler algorithms using pre-existing supervised classifier model trained on a small portion of labeled data, followed by predictions on remainder of the data set (e.g., the unlabeled portion)), label propagation algorithms (e.g., algorithms that assign labels to unlabeled observations by propagating or allocating labels through the data set over time based on the labeled data set (e.g., a graph neural network)), and combinations thereof.

[0079] Reinforcement learning models can include any type of reinforcement learning model. Reinforcement learning models can be sensor models or neuron models. Sensor models can use a break algorithm or a sigmoid algorithm. The input of sensor models can be a real number vector. The input of neuron models can be a binary vector. Reinforcement learning models can use any reinforcement learning algorithm, including, but not limited to, temporal difference (TD) algorithms e.g., tabular TD(lamb da), replacing traces; Tabular TD(0); tabular TD(lambda) with function approximation; gradient temporal difference 2 (GTD2); least squares temporal difference (LSTD)), temporal distribution characterization (TDC), Monte Carlo methods (e.g., every-visit Monte-Carlo, first-visit Monte-Carlo), value iteration, policy iteration, Q-leaming with function approximation, Explicit Explore or Exploit (E3) algorithm, State-aggregation based Q-learning, fitted Q-iteration, generalized policy iteration, tabular q-leaming algorithm, upper confidence bound 1 algorithm, delayed-Q algorithm, MoRmax algorithm, upper confidence reinforcement learning 2 (UCRL2) algorithm, state-action-reward-state-action algorithm, R-Max algorithm, REINFORCE algorithm, model-based interval estimation (MB IE), Boltzmann exploration, action elimination, generalized policy iteration, natural actor-critic, soft-state aggregation based Q-learning, epsilon-greedy, and combinations thereof. Reinforcement learning models can be any deep reinforcement learning model. Deep reinforcement learning can be based on an artificial neural network.

[0080] The machine learning model can comprise an artificial neural network (ANN). An artificial neural network can comprise neurons. Neurons in an artificial neural network can receive signals, process them, and send signals to neurons connected to them. Neurons can be connected via links.

[0081] The connected pattern or network of neurons can be a directed, weighted graph. The neural network can comprise multiple layers. The layers can comprise input layers, hidden layers, and output layers. Input layers can receive raw data or labelled data. The raw data orlabelled data can be a training data set or a non-training data set. Input layers can also receive outputs from preceding machine learning models. Hidden layers are intermediate layers between output and input layers. Hidden layers can process data by applying functions to the data. Output layers can receive processed data from preceding layers and produce an output from the machine learning model. The output can be a final result. The output can be an input for a subsequent machine learning model.

[0082] The machine learning model can comprise any connection pattern between each pair of layers. The layers can be fully connected. Fully connected layers comprise layers in which every neuron in one layer is connected to every neuron in the subsequent layer. The layers can comprise pooling layers. Pooling layers comprise layers in which a group of neurons in one layer connects to a single neuron in the next layer. The machine learning model can contain any feedforward network, comprising pooling layers. The machine learning model can comprise any quantum neural network (QNN). Feedforward networks include, but are not limited to, convolutional neural network (CNN) (c. ., including, but not limited to, LeNet, AlexNet, ZF Net, GoogLeNet, VGGNet, ResNet, MobileNets, GoogLeNet_DeepDream, or combinations thereof), radial basis function network, linear neural network, perceptron, multilayer perceptron, and combinations thereof. The machine learning model can comprise any recurrent neural network (RNN), including, but not limited to, long-term memory (LSTM network), gated recurrent units (GRU), echo state network (ESN), fully recurrent neural network (FRNN), Elman network, Jordan network, Hopfield network, bidirectional associative memory, independently RNN (IndRNN), recursive neural network, neural history compressor, second order RNN, transformer (e.g., transformer architecture), continuous-time RNN (CTRNN), hierarchical RNN (HRNN), recurrent multilayer perceptron network (RMLP), multiple timescales RNN (MTRNN), neural Turing machine (NTM), differentiable neural computer (DNC), neural network pushdown automata (NNPDA), Memristive network, and combinations thereof. Recurrent neural networks can comprise connections between neurons in the same or prior layers. Recurrent neural networks can be bi-directional in the transmission of signal between neurons or layers.

[0083] The machine learning model can include a dropout layer. Dropout layers can be used to reduce statistical noise in a data set and overfitting. A dropout layer can randomly set input units to 0 with a frequency of rate at each step during training time. Many types and architectures ofdropout layers are known (e.g., U.S. Pat. No. 9,406,017), and any can be used in the methods described herein. A machine learning model can comprise any number of dropout layers. A machine learning model can comprise at least one dropout layer, e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 14, at least 16, at least 18, at least 20, at least 24, at least 28, at least 30, at least 33, at least 36, at least 39, at least 40 at least 44, at least 48, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 150, at least 200, at least 250, at least 300, and at least 350 dropout layers. Amachine learning model can comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, 230, 235, 240, 245, 250, 255, 260, 265, 270, 275, 280, 285, 290, 295, 300, 305, 310, 315, 320, 325, 330, 335, 340, 345, and 350 dropout layers.

[0084] The machine learning model can comprise applying an activation function. Hidden layers and output layers can comprise activation functions. Activation functions can have different properties (e.g., linearity, range, differentiability). Activation functions can be linear. Activation functions can be non-linear. Activation functions can have a finite range. Activation functions can have an infinite range. Activation functions can be continuously differentiable. Any activation function can be applied in the methods described herein including, but not limited to, nonlinear, ridge activation functions (e.g., linear activation, rectified linear unit (ReLU), Heaviside activation, logistic activation), radial activation functions (e.g., Gaussian,multi quadratic, inverse multi quadratic, polyharmonic spine), folding activation function, and combinations thereof. Activation functions include, but are not limited to, identity, binary step, exponential linear unit (ELU), Gaussian, Gaussian error linear unit (GELU), hyperbolic tangent, logistic, leaky rectified linear unit, Maxout, parametric rectified linear unit (PreLU), ReLU (e.g., S-shaped ReLU, Sine ReLU), quantum activation functions, scaled exponential linear unit (SELU), ), sigmoid linear unit (e.g., SiLU, sigmoid shrinkage, SiL, Swish-1), Softplus, and combinations thereof. In the methods described herein, a machine learning model can comprise applying any number of activation functions. The number of activation functions can be at least one, e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, atleast 10, at least 12, at least 14, at least 16, at least 18, at least 20, at least 24, at least 28, at least 30, at least 33, at least 36, at least 39, at least 40, at least 44, at least 48, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 150, at least 200, at least 250, at least 300, or at least 350 activation functions. The machine learning model can comprise applying 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, 230, 235, 240, 245, 250, 255, 260, 265, 270, 275, 280, 285, 290, 295, 300, 305, 310, 315, 320, 325, 330, 335, 340, 345, or 350 activation functions.

[0085] Learning in a machine learning model can comprise adjusting the weights of the machine learning model to improve accuracy of the task. Learning in a machine learning model can comprise adjusting the optional thresholds of the machine learning model to improve accuracy of the task. For example, the weights and optional thresholds of the machine learning model can be adjusted to improve the accuracy of predicting the sequencing depth, batch size, and / or amount of biological sample needed to obtain the desired sequencing output (e.g., the desired number of neoantigens or sequence variants). For example, the weights and optional thresholds of the machine learning model can be adjusted to improve the accuracy of predicting the likelihood that a neoantigen or sequence variant is a true positive.

[0086] Machine learning models in the methods described herein can be trained by any data set (e.g., any training data set). A training data set can be data of sequence variants identified from a biological sample (e.g, liquid biopsy). A training data set can comprise sequence variants identified from a biological sample labeled to indicate whether each identified variant is a true positive, false positive, true negative, and / or false negative. The method can include a step of training the machine learning model on training set data.

[0087] Training data sets can be a set of labeled SVs, wherein the sequence variants are known true positive SVs. For example, the labeled SVs can be SVs identified from liquid biopsy sequencing data paired with high-purity tumor tissue sequencing data, collected around the same time to minimize biological difference. The ctDNA-derived SVs obtained from the liquid biopsy that also show up in the paired tumor tissue sequencing data are labeled as true positiveexamples, whereas the variants absent from paired tumor tissue sequencing data are labeled false positives. Additionally, the training data sets can include SVs that are confirmed as true positive SVs by multiple independent sequencing methods or carefully curated by a group of experts. Liquid biopsy sequencing data from known cancer patients can be used to train true positive variants, ideally from patients with a wide range of cancer types (e.g. breast cancer, lung cancer, melanoma etc.) and stages (stage I to IV, from early to late stage) to maximize generality of the machine learning model trained. Liquid biopsy sequencing data from healthy subjects, wherein the subjects have no cancer diagnoses, can be used to train true negative SVs.

[0088] Training data sets e.g., negative training data sets, positive training data sets) can contain any number of data points. The number of data points in a training data set can be at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 14, at least 16, at least 18, at least 20, at least 22, at least 24, at least 26, at least 28, at least 30, at least 33, at least 36, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1800, at least 1900, at least 2000, at least 2200, at least 2400, at least 2600, at least 2800, at least 3000, at least 3300, at least 3600, at least 3900, at least 4000, at least 4400, at least 4800, at least 5000, at least 5500, at least 6000, at least 6500, at least 7000, at least 7500, at least 8000, at least 8500, at least 9000, at least 9500, at least 10000, at least 11000, at least 12000, at least 13000, at least 14000, at least 15000, at least 16000, at least 17000, at least 18000, at least 19000, or at least 20000. The number of data points can be between about 1 to about 20000, about 5 to about 20000, about 10 to about 20000, about 15 to about 20000, about 20 to about 20000, about 25 to about 20000, about 30 to about 20000, about 35 to about 20000, about 40 to about 20000, about 45 to about 20000, about 50 to about 20000, about 55 to about 20000, about 60 to about 20000, about 65 to about 20000, about 70 to about 20000, about 75 to about 20000, about 80 to about 20000, about 85 to about 20000, about 90 to about 20000, about 95 to about 20000, about 100 to about 20000, about 200 to about 20000,about 300 to about 20000, about 400 to about 20000, about 500 to about 20000, about 600 to about 20000, about 700 to about 20000, about 800 to about 20000, about 900 to about 20000, about 1000 to about 20000, about 2000 to about 20000, about 3000 to about 20000, about 4000 to about 20000, about 5000 to about 20000, about 6000 to about 20000, about 7000 to about 20000, about 8000 to about 20000, about 9000 to about 20000, about 10000 to about 20000, about 11000 to about 20000, about 12000 to about 20000, about 13000 to about 20000, about 14000 to about 20000, about 15000 to about 20000, about 16000 to about 20000, about 17000 to about 20000, about 18000 to about 20000, or about 19000 to about 20000. The number of data points in a training data set can be about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 12, about 14, about 16, about 18, about 20, about 22, about 24, about 26, about 28, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 200, about 300, about 400, about 500, about 600, about 700, about 800, about 900, about 1000, about 1100, about 1200, about 1300, about 1400, about 1500, about 1600, about 1700, about 1800, about 1900, about 2000, about 2200, about 2400, about 2600, about 2800, about 3000, about 3300, about 3600, about 3900, about 4000, about 4400, about 4800, about 5000, about 5500, about 6000, about 6500, about 7000, about 7500, about 8000, about 8500, about 9000, about 9500, about 10000, about 11000, about 12000, about 13000, about 14000, about 15000, about 16000, about 17000, about 18000, about 19000, or about 20000.

[0089] The machine learning model can comprise one or more hyperparameters. The machine learning model can comprise any number of hyperparameters, such as at least 1 hyperparameter, e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 hyperparameters. Hyperparameters include, but are not limited to, learning rate, number of total layers, number of hidden layers, machine learning model batch size, the number of neurons per layer, the number of sensors per layer, and combinations thereof.

[0090]

[0120] The machine learning model can comprise any number of total layers. The total number of layers can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 24, 26, 28, 30, 33, 36, 39, 40, 44, 48, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 220, 240, 260, 280, 300, 330, 360, 390, 400, 440, 480, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, or 1000. The total number of layers can be at least 1, atleast 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 22, at least 24, at least 26, at least 28, at least 30, at least 33, at least 36, at least 39, at least 40, at least 44, at least 48, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 220, at least 240, at least 260, at least 280, at least 300, at least 330, at least 360, at least 390, at least 400, at least 440, at least 480, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, or at least 1000. The machine learning model can be a single layer or an unlayered network.

[0091] The machine learning model can comprise any number of hidden layers. The total number of hidden layers can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 24, 26, 28, 30, 33, 36, 39, 40, 44, 48, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 220, 240, 260, 280, 300, 330, 360, 390, 400, 440, 480, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, and 1000. The total number of hidden layers can be at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 22, at least 24, at least 26, at least 28, at least 30, at least 33, at least 36, at least 39, at least 40, at least 44, at least 48, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 220, at least 240, at least 260, at least 280, at least 300, at least 330, at least 360, at least 390, at least 400, at least 440, at least 480, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, or at least 1000.

[0092] The machine learning model batch size is the number of samples that will be propagated through the network in one machine learning model batch (e.g., the number of training examples in one forward and backward pass). Any machine learning model batch size can be used in the methods described herein. The machine learning model batch size can be at least 1, e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12,at least 14, at least 16, at least 18, at least 20, at least 22, at least 24, at least 26, at least 28, at least 30, at least 33, at least 36, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, or at least 1000. The machine learning model batch size can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000.

[0093] A layer (e.g., an input layer, a hidden layer, an output layer) can comprise any number of sensors or neurons. The number of sensors or neurons in a layer can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 24, 26, 28, 30, 33, 36, 39, 40, 44, 48, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 220, 240, 260, 280, 300, 330, 360, 390, 400, 440, 480, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, or 1000. The number of sensors or neurons in a layer can be at least 1, e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 14, at least 16, at least 18, at least 20, at least 22, at least 24, at least 26, at least 28, at least 30, at least 33, at least 36, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, or at least 1000.Immunogenic Composition

[0094] The methods described herein can include a step of selecting one or more actionable variant that results in the expression of a neoantigen (e.g., one or more neoantigens) for inclusion in an immunogenic composition. The methods described herein can include a step of generating an immunogenic composition (e.g., a cancer vaccine). The methods described herein can include a step of administering an immunogenic composition to a patient (e.g., a patient in need thereof).

[0095] The methods described herein can determine, score, and / or select neoantigens for generating an immunogenic composition (e.g., a cancer vaccine). Any method of determining, scoring, and selecting neoantigens can be used. Examples of suitable methods include those described by WO2022 / 159176A1, US 20230197192A1, US 20230173045A1, and US 20220093209A1, the entire contents of each of which are incorporated herein.

[0096] Neoantigens are self-antigens generated by tumor cells due to genomic mutations or dysregulated RNA splicing. Neoantigens (e.g., neoantigens predicted as true positives by methods described herein) identified by the methods described herein can be in the form of any sequencing results, including, but not limited to, sequence reads (DNA sequence reads, RNA sequence reads), sequence variants (e.g., peptide-modifying sequence variants), encoded peptides, or combinations of the foregoing. The neoantigens can be any type of peptide, including, but not limited to, long peptides, short peptides, or combinations thereof. Neoantigens of the methods described herein can be tumor-specific (e.g., neoantigens that are only present in a cancer or tumor and not in healthy or germ-line cells of a subject).

[0097] As used herein, the terms “short identified neoantigen peptide” and “short peptide” can refer to a peptide with an amino acid length that is between about 3 to about 15, about 3 to about 14, about 3 to about 13, about 3 to about 12, about 3 to about 11, about 3 to about 10, about 3 to about 9, about 3 to about 8, about 3 to about 7, about 3 to about 6, about 3 to about 5, about 3 to about 4, about 4 to about 15, about 5 to about 15, about 6 to about 15, about 7 to about 15, about 8 to about 15, about 9 to about 15, about 10 to about 15, about 11 to about 15, about 12 to about 15, about 13 to about 15, about 14 to about 15, about 8 to about 11, about 8 to about 10, about 8 to about 9, about 9 to about 11, about 10 to about 11, about 7 to about 11, about 6 to about 11, about 5 to about 11, about 4 to about 11, about 3 to about 11, about 8 to about 12, about 8 to about 13, or about 8 to about 14. “Short identified neoantigen peptide” and “short peptide” can refer to a peptide with an amino acid length that is about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, or about 15. “Short identified neoantigen peptide” and “short peptide” can refer to a peptide with an amino acid length that is at most about 15, at most about 14, at most about 13, at most about 12, at most about 11, at most about 10, at most about 9, at most about 8, at most about 7, at most about 6, at most about 5, at most about 4, at most about 3, or at most about 2. As used herein, the terms “long identifiedneoantigen peptide” and “long peptide” can refer to a peptide with an amino acid length that is between about 13 and about 30, about 13 and about 29, about 13 and about 28, about 13 and about 27, about 13 and about 26, about 13 and about 25, about 13 and about 24, about 13 and about 23, about 13 and about 22, about 13 and about 21, about 13 and about 20, about 13 and about 19, about 13 and about 18, about 13 and about 17, about 13 and about 16, about 13 and about 15, about 13 and about 14, about 14 and about 30, about 15 and about 30, about 16 and about 30, about 17 and about 30, about 18 and about 30, about 19 and about 30, about 20 and about 21, about 22 and about 30, about 23 and about 30, about 24 and about 30, about 25 and about 30, about 26 and about 30, about 27 and about 30, about 28 and about 30, about 29 and about 30, or about 13 and about 25. “Long identified neoantigen peptide” and “long peptide” can refer to a peptide with an amino acid length that is about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, or about 30. “Long identified neoantigen peptide” and “long peptide” can refer to a peptide with an amino acid length that is at least about 13, at least about 14, at least about 15, at least about 16, at least about 17, at least about 18, at least about 19, at least about 20, at least about 21, at least about 22, at least about 23, at least about 24, at least about 25, at least about 26, at least about 27, at least about 28, at least about 29, or at least about 30.

[0098] The methods described herein can include steps to determine neoantigens (e.g., neoantigens predicted to be true positives by methods described herein, neoantigens identified from a sequencing result by methods described herein) appropriate for inclusion in an immunogenic composition. The methods can further include a step of determining the human leukocyte antigen (HLA) or major histocompatibility complex (MHC) type for a patient. The HLAtype can be determined through any method, including, but not limited to, nucleic acid sequencing (e.g., RNA sequencing, DNA sequencing, complementary (cDNA) sequencing), immunoassay (e.g, enzyme linked immunosorbent assay (ELISA)), mass spectrometry, flow cytometry, real-time PCR, or combinations of any of the foregoing. The methods can include a step of scoring neoantigens of a patient (e.g, neoantigens predicted to be true positives by methods described herein, neoantigens identified from a sequencing result by methods described herein). The methods can include a step of selecting one or more neoantigens (e.g., of abiological sample, of a sequencing result) for inclusion in an immunogenic composition. Any number of neoantigens can be selected for inclusion in the immunogenic composition. The number of neoantigens selected can be at most 20, at most 19, at most 18, at most 17, at most 16, at most 15, at most 14, at most 13, at most 12, at most 11, at most 10, at most 9, at most 8, at most 7, at most 6, at most 5, at most 4, at most 3, at most 2, or at most 1. The number of neoantigens selected can be at least 20, at least 19, at least 18, at least 17, at least 16, at least 15, at least 14, at least 13, at least 12, at least 11, at least 10, at least 9, at least 8, at least 7, at least 6, at least 5, at least 4, at least 3, at least 2, or at least 1. The number of neoantigens selected can be between about 1 to about 20, about 2 to about 20, about 3 to about 20, about 4 to about 20, about 5 to about 20, about 6 to about 20, about 7 to about 20, about 8 to about 20, about 9 to about 20, about 10 to about 20, about 11 to about 20, about 12 to about 20, about 13 to about 20, about 14 to about 20, about 15 to about 20, about 16 to about 20, about 17 to about 20, about 18 to about 20, about 19 to about 20, about 1 to about 19, about 1 to about 18, about 1 to about 17, about 1 to about 16, about 1 to about 15, about 1 to about 14, about 1 to about 13, about 1 to about 12, about 1 to about 11, about 1 to about 10, about 1 to about 9, about 1 to about 8, about 1 to about 7, about 1 to about 6, about 1 to about 5, about 1 to about 4, about 1 to about 3, or about 1 to about 2.

[0099] Neoantigens suitable for use in generating an immunogenic composition can be determined by sequencing nucleic acids from a biological sample of a patient. Sequencing in the methods described herein can be any type of sequencing, including, but not limited to, whole genome sequencing, shotgun metagenomic sequencing, whole exome sequencing, nextgeneration sequencing (NGS), cancer personalized profiling by deep sequencing (CAPP-Seq), tagged-amplicon deep sequencing (Tam-Seq), or combinations of any of the foregoing.

[0100] Neoantigen peptides can be predicted by any method. One suitable method for predicting neoantigen peptides is analysis (via an algorithm) of protein variants expressed and isoforms translated from population-level statistics. Another suitable method for predicting neoantigen peptides is by translation of RNA sequencing data (e.g, RNA sequencing data that overlaps each somatic variant). Another suitable method for predicting neoantigen peptides is by transcription of DNA sequences (e.g., cfDNA sequences) into RNA and translation of the resulting RNA sequencing data.

[0101] Any strategy or method can be used to reduce the background noise of nucleic acids or proteins originating from healthy non-cancerous cells, and to increase the signal to noise ratio of nucleic acids or proteins from cancerous cells (e.g., sequencing variants that encode for neoantigens). One suitable method for increasing the signal to noise ratio of sequencing variants is by sequencing a biological sample to yield separate non-cancer sequencing results (e.g., sequencing results from non-cancerous cells, sequencing results lacking oncogenic somatic mutations) and cancer sequencing results (e.g., sequencing results from cancer cells, sequencing results containing somatic mutations). In such a method, the non-cancer sequencing results can be subtracted (e.g., eliminated, deleted) from the cancer sequencing results to improve the signal to noise ratio for variant calling of neoantigens. The non-cancer sequencing results and the cancer sequencing results can be obtained from the same biological sample. For example, the cancer sequencing results, and the non-cancer sequencing results can both be obtained from the same liquid biopsy sample (e.g., the same blood sample). Continuing this example, the liquid biopsy sample can be a blood sample and the non-cancer sequencing results are obtained from sequencing the genetic material of white blood cells, while the cancer sequencing results are obtained from sequencing cell free DNA (cfDNA). The non-cancer sequencing results and the cancer sequencing results can be obtained from two different biological samples. For example, the non-cancer sequencing results can be obtained from a tissue biopsy (e.g., a biopsy of healthy, non-cancerous tissue) and the cancer sequencing results can be obtained from tissue resection of a tumor. Another method to increase the signal to noise is to use unique molecular identifiers (UMI) and oversample during sequencing to facilitate error correction. In some embodiments, the method comprises a step of sequencing nucleic acids (e.g., genetic material) of a biological sample to yield non-cancer sequencing results and cancer sequencing results, wherein the non-cancer sequencing results are subtracted from the cancer sequencing results to improve the signal to noise ratio for variant calling of neoantigens.

[0102] Any number of biological samples can be sequenced and analyzed to identify, determine, score, and / or select neoantigens. The number of biological samples can be about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20. The number of biological samples can be at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7,at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20. The number of biological samples can be between about 1 to about 20, about 1 to about 19, about 1 to about 18, about 1 to about 17, about 1 to about 16, about 1 to about 15, about 1 to about 14, about 1 to about 13, about 1 to about 12, about 1 to about 11, about 1 to about 10, about 1 to about 9, about 1 to about 8, about 1 to about 7, about 1 to about 6, about 1 to about 5, about 1 to about 4, about 1 to about 3, about 1 to about 2, about 2 to about 20, about 3 to about 20, about 4 to about 20, about 5 to about 20, about 6 to about 20, about 7 to about 20, about 8 to about 20, about 9 to about 20, about 10 to about 20, about 11 to about 20, about 12 to about 20, about 13 to about 20, about 14 to about 20, about 15 to about 20, about 16 to about 20, about 17 to about 20, about 18 to about 20, about 19 to about 20, about 2 to about 19, about 2 to about 18, about 2 to about 17, about 2 to about 16, about 2 to about 15, about 2 to about 14, about 2 to about 13, about 2 to about 12, about 2 to about 11, about 2 to about 10, about 2 to about 9, about 2 to about 8, about 2 to about 7, about 2 to about 6, about 2 to about 5, about 2 to about 4, about 2 to about 3, about 3 to about 19, about 3 to about 18, about 3 to about 17, about 3 to about 16, about 3 to about 15, about 3 to about 14, about 3 to about 13, about 3 to about 12, about 3 to about 11, about 3 to about 10, about 3 to about 9, about 3 to about 8, about 3 to about 7, about 3 to about 6, about 3 to about 5, about 3 to about 4, about 4 to about 19, about 4 to about 18, about 4 to about 17, about 4 to about 16, about 4 to about 15, about 4 to about 14, about 4 to about 13, about 4 to about 12, about 4 to about 11, about 4 to about 10, about 4 to about 9, about 4 to about 8, about 4 to about 7, about 4 to about 6, or about 4 to about 5.

[0103] Sequencing results obtained by the methods described herein (e.g., DNA sequencing result, RNA sequencing result) can be used to determine peptide sequences that are encoded by nucleic acids (e.g., encoded by DNA, encoded by RNA). The nucleic acid that is sequenced to obtain a sequencing result can be any type of nucleic acid, including, but not limited to, RNA, DNA, cfDNA, cfRNA (cell-free RNA), circulating tumor DNA (ctDNA), circulating tumor RNA (ctRNA), or combinations of any of the foregoing. The sequencing result can be in any format, including, but not limited to, CRAM format, General Feature Format (e.g., GFF3), FASTA format, FASTQ format, NeXML format, Nexus format, Pileup format, Sequence Alignment Map (SAM) format, Stockholm format, Variant Call Format (VCF) format, genomic VCF (gVCF)format, or a combination of any of the foregoing. In some embodiments, the sequencing result is circulating tumor DNA and the format of the sequencing result is VCF format, wherein the VCF file contains tumor-specific somatic variants found in the biological sample. In some embodiments, the sequencing result is obtained from sequencing healthy and tumor tissue and the format is separate FASTQ files of the sequencing results from healthy and tumor tissue. The resulting peptide sequences can be analyzed to determine whether the encoded peptide is a neoantigen based on the presence of sequence variants (e.g., sequence variants in the peptide amino acid sequence compared to the healthy cells or healthy tissue of a subject, sequence variants in the peptide amino acid sequence compared to a reference genome). Neoantigens (e.g., neoantigen peptides) identified by the methods described herein can be analyzed to determine a predicted immunogenicity based on any factor, including, but not limited to, whether the neoantigen is immunogenic (e.g., whether a neoantigen can elicit an immune response in a subject, prediction of major histocompatibility complex (MHC) binding affinity, whether peptide is predicted to be presented on a cell surface by an MHC molecule), whether the tumor expresses an amount of neoantigen sufficient to elicit an immune response, whether the neoantigen is expressed on a sufficient fraction of the tumor cells, relative or absolute amount of nucleic acids encoding for the neoantigen present in the sample (e.g. , present in the total cfDNA) or a combination thereof. Any methodology for predicting immunogenicity of a neoantigen can be used in the methods described herein. For example, the number of sequenced variants (e.g., DNA sequence variants, RNA sequence variants, encoded peptide variants) can be used to predict the expression of a neoantigen in a biological sample and infer the expression in the originating tissue of the subject (e.g., in a tumor of the subject). For several examples of predicting immunogenicity of a neoantigen, see WO2022 / 159176A1, US 20230197192A1, US 20230173045A1, and US 20220093209A1, all of which are hereby incorporated by reference in their entireties. The predicted immunogenicity can be used to score a neoantigen. Neoantigens can be scored by any parameter of the neoantigen. For example, the neoantigens can be scored by the counts of nucleic acids (e.g., DNA, RNA) encoding the neoantigen (e.g., RNA transcripts per million (RNA TPM), RNA expression values). For example, neoantigens can be scored by cellular prevalence of the neoantigen (e.g., cells expressing the neoantigen as determined by single cell sequencing of circulating tumor cells or ctc-FISH staining, cells expressing theneoantigen protein as determined by cell cytometry of circulating tumor cells). Neoantigen peptides can be scored by binding affinity (e.g., experimental binding affinity, predicted binding affinity) to a major histocompatibility complex (MHC), frequency of mutant alleles, gene expression, C-terminal cleavage affinity, transporter associated with antigen processing (TAP) transport of the neoantigen peptide, or combinations of any of the foregoing. The scores of predicted immunogenicity can be ranked (e.g., neoantigens can be ranked in order of predicted immunogenicity). The scored neoantigens can be further selected for inclusion (e.g., inclusion of a neoantigen peptide, inclusion of a nucleic acid encoding for a neoantigen peptide) in an immunogenic composition (e.g., a cancer vaccine). The scored neoantigens can be further selected for exclusion (e.g., exclusion of a neoantigen peptide, exclusion of a nucleic acid encoding for a neoantigen peptide) from an immunogenic composition (e.g., a cancer vaccine).

[0104] Sequence variants (e.g., mutations) identified from biological samples (e.g., liquid biopsy samples, blood samples) by variant calling can be from non-cancer cells or cancer cells (e.g., ctDNA, ctRNA). Any method to distinguish non-cancer cell and cancer cell sequence variants can be used in the methods described herein. For example, multiple sequence variants can be screened simultaneously to improve the probability of detecting circulating tumor nucleic acids (e.g., ctDNA, ctRNA). As another example, the probability of detecting circulating tumor nucleic acids can be increased based on detection of epigenetic modifications (e.g., methylation profile, including 5-methylcytosine (5mC) and / or 5 -hydroxy methylcytosine (5hmC)) of the nucleic acids (e.g., cfDNA). As another example, the probability of detecting circulating tumor nucleic acids can be increased based on the fragmentation pattern (e.g., fragment lengths, fragment start positions, fragment end nucleotide motifs) of nucleic acids (e.g., cfDNA). As another example, the probability of detecting circulating tumor nucleic acids can be increased based on the mutational signature profiles of known human carcinogens, existence of co-mutations, and / or evolutionary cancer signatures.

[0105] Methods described herein can include a step of generating an immunogenic composition based on the neoantigen identified by methods described herein (e.g., neoantigens predicted as true positives). The immunogenic composition (e.g., cancer vaccine) can be any type of immunogenic composition, such as those described in W02022 / 170067 Al, US2023 / 0173045 Al, WO2022 / 159176 Al, WO2022 / 251034 Al, or US2023 / 0173046 Al, the entirety of each ofwhich are incorporated by reference herein. The immunogenic composition can comprise or encode neoantigens (e.g., one or more neoantigens) scored for predicted immunogenicity by any method including by the methods described herein. The immunogenic composition can contain neoantigens in any form, including, but not limited to, peptides (e.g., long identified peptides, short identified peptides, synthesized peptides, isolated peptides), RNA sequences encoding for peptides (e.g., mRNA, mRNA containing unnatural nucleotides (e.g., pseudouridine, Nl-methylpseudouridine, 7-methylguanosine, N6-methyladenosine, 2’-O-methyl nucleotide), mRNA containing inverted nucleotides, or combinations thereof), DNA sequences encoding for peptides (e.g., plasmid DNA, viral DNA), or combinations thereof. Immunogenic compositions containing nucleic acids (e.g., RNA, DNA) encoding for peptides can be in the form of a virus (e.g., an adenovirus, lentivirus, fowl pox, vaccinia, self-replicating alphavirus, Maraba virus) or free nucleic acids in a drug delivery particle (e.g., a lipid nano particle (LNP)).

[0106] A nucleic acid (e.g., RNA, DNA) of an immunogenic composition can encode for any number of peptide antigens. The number of peptide antigens encoded by a nucleic acid can be about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20. The number of peptide antigens encoded by a nucleic acid can be at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 11, at most about 12, at most about 13, at most about 14, at most about 15, at most about 16, at most about 17, at most about 18, at most about 19, or at most about 20. In some embodiments, the method includes a step of generating an immunogenic composition, wherein the immunogenic composition comprises or encodes for one or more neoantigens scored for the predicted immunogenicity by the method. The immunogenic composition can comprise neoantigens predicted as true positives by methods described herein. The immunogenic composition can comprise neoantigens identified from the sequencing results in the methods described herein.

[0107] An immunogenic composition can contain or encode for any number of neoantigens (e.g., neoantigen proteins, DNA sequences encoding for neoantigen proteins, RNA sequences encoding for neoantigen proteins). An immunogenic composition can contain or encode for a number of neoantigens that is about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9,about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, or about 30. The number of neoantigens contained or encoded by an immunogenic composition can be one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, 20 or more, 21 or more, 22 or more, 23 or more, 24 or more, 25 or more, 26 or more, 27 or more, 28 or more, or 29 or more, or 30 or more. The number of neoantigens contained or encoded by an immunogenic composition can be at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 16, at least about 17, at least about 18, at least about 19, at least about 20, at least about 21, at least about 22, at least about 23, at least about 24, at least about 25, at least about 26, at least about 27, at least about 28, at least about 29, or at least about 30. The number of neoantigens contained or encoded by an immunogenic composition can be between about 1 to about 50, about 2 to about 50, about 3 to about 50, about 4 to about 50, about 5 to about 50, about 6 to about 50, about 7 to about 50, about 8 to about 50, about 9 to about 50, about 10 to about 50, about 11 to about 50, about 12 to about 50, about 13 to about 50, about 14 to about 50, about 15 to about 50, about 16 to about 50, about 17 to about 50, about 18 to about 50, about 19 to about 50, about 20 to about 50, about 22 to about 50, about 24 to about 50, about 26 to about 50, about 28 to about 50, about 30 to about 50, about 33 to about 50, about 36 to about 50, about 40 to about 50, about 45 to about 50, about 1 to about 45, about 1 to about 40, about 1 to about 36, about 1 to about 33, about 1 to about 30, about 1 to about 28, about 1 to about 26, about 1 to about 24, about 1 to about 22, about 1 to about 20, about 1 to about 18, about 1 to about 16, about 1 to about 14, about 1 to about 12, about 1 to about 10, about 1 to about 9, about 1 to about 8, about 1 to about 7, about 1 to about 6, about 1 to about 5, about 1 to about 4, about 1 to about 3, about 1 to about 2, about 2 to about 14, about 2 to about 12, about 2 to about 10, about 2 to about 8, about 2 to about 6, about 2 to about 4, about 3 to about 14, about 3 to about 12, about 3 to about 10, about 3 to about 8, about 3 to about 6, about 3 to about 4, about 4 to about 14, about 4 to about 12, about 4 to about 10, about 4 to about 8, about 4 to about 6, about 5 to about 14,about 5 to about 12, about 5 to about 10, about 5 to about 8, about 5 to about 6, about 6 to about 14, about 6 to about 12, about 6 to about 10, about 6 to about 8, about 7 to about 14, about 7 to about 12, about 7 to about 10, about 7 to about 8, about 8 to about 14, about 8 to about 12, about 8 to about 10, about 9 to about 14, about 9 to about 12, about 9 to about 10, about 10 to about 14, about 10 to about 12, about 11 to about 14, about 11 to about 12, about 12 to about 14, or about 13 to about 14.

[0108] Immunogenic compositions described herein may comprise up to about 50 neoantigen long peptides and / or short peptides. The immunogenic composition may comprise about 10 to about 20 neoantigen long peptides and / or short peptides. In some embodiments, the immunogenic composition comprises about 19 neoantigen long peptides and / or short peptides.

[0109] The immunogenic composition may comprise at least about 2 or more neoantigen long peptides. The immunogenic composition may comprise about 2 to about 18 neoantigen long peptides. The immunogenic composition can comprise at least about 10 to about 15 neoantigen long peptides. The immunogenic composition may comprise at least about 2 or more neoantigen short peptides. The immunogenic composition may comprise at least about 2 to about 10 neoantigen short peptides.

[0110] The methods described herein can include a step of administering the immunogenic composition to a patient in need thereof. The selected neoantigens can be divided into pools for separate administration of immunogenic compositions to a patient in need thereof. The pools of neoantigens can contain any number of neoantigens. The number of neoantigens in a pool can be about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, or about 50. The number of neoantigens in a pool can be at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, or at least 50. The number of neoantigens in a pool can be at most 2, at most 3, at most 4, at most 5, at most 6, at most 7, at most 8, at most 9, at most 10, at most 11, at most 12, at most 13, at most 14, at most 15, at most 16, at most 17, at most 18, at most 19, at most 20, at most 21, at most 22, at most 23, at most 24, at most 25, at most 26, at most 27, at most 28, at most 29, at most 30, at most 31, at most 32, at most 33, at most 34, at most 35, at most 36, at most 37, at most 38, at most 39, at most 40, at most 41, at most 42, at most 43, at most 44, at most 45, at most 46, at most 47, at most 48, at most 49, or at most 50. The number of neoantigens in a pool can be about 2 to about 30, 2 to about 28, 2 to about 26, 1 to about 24, 2 to about 22, 2 to about 20, about 2 to about 18, about 2 to about 16, about 2 to about 14, about 2 to about 12, about 2 to about 10, about 2 to about 9, about 2 to about 8, about 2 to about 7, about 2 to about 6, about 2 to about 5, about 2 to about 4, or about 2 to about 3. As an example, three peptide pools of the immunogenic composition may comprise about 5 neoantigen long peptides and / or short peptides and one peptide pool may comprise 4 neoantigen long peptides and / or short peptides and a helper peptide. Each peptide pool may comprise different neoantigen long peptides and / or short peptides. Neoantigen long peptides may be about 15 to about 30 amino acids in length.Neoantigen short peptides may be about 5 to about 15 amino acids in length.[0U1] The selected neoantigens for the immunogenic composition can be split into any number of pools. The number of pools can be about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20. The number of pools can be at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20. The number of pools can be at most 1, at most 2, at most 3, at most 4, at most 5, at most 6, at most 7, at most 8, at most 9, at most 10, at most 11, at most 12, at most 13, at most 14, at most 15, at most 16, at most 17, at most 18, at most 19, or at most 20. The number of pools can be between about 1 to about 2, about 1 to about 3, about 1 to about 4, about 1 to about 5, about 1 to about 6, about 1 to about 7, about 1 to about 8, about 1 to about 9, about 1 to about 10, about 2 to about 10, about 3 to about 10, about 4 to about 10, about 5 to about 10, about 6 to about 10, about 7 to about 10, about 8 to about 10, about 9 to about 10, about 2 to about 3, about 2 to about 4, about 2to about 5, about 2 to about 6, about 2 to about 7, about 2 to about 8, about 2 to about 9, about 2 to about 10, about 3 to about 4, about 3 to about 5, about 3 to about 6, about 3 to about 7, about 3 to about 8, about 3 to about 9, about 3 to about 10, about 4 to about 5, about 4 to about 6, about 4 to about 7, about 4 to about 8, about 4 to about 9, or about 4 to about 10.

[0112] The immunogenic composition can comprise at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 16, at least about 17, at least about 18, at least about 19, at least about 20, at least about 21, at least about 22, at least about 23, at least about 24, at least about 25, at least about 26, at least about 27, at least about 28, at least about 29, at least about 30, at least about 31, at least about 32, at least about 33, at least about 34, at least about 35, at least about 36, at least about 37, at least about 38, at least about 39, at least about 40, at least about 41, at least about 42, at least about 43, at least about 44, at least about 45, at least about 46, at least about 47, at least about 48, at least about 49, at least about 50 or more neoantigen peptides (e.g., neoantigen long peptides and / or short peptides). The immunogenic composition can comprise up to about 100 neoantigen peptides. The immunogenic composition can contain about 1-5 neoantigens, 1-10 neoantigens, about 1-15 neoantigens, about 4-10 neoantigens, about 4-15 neoantigens, about 10-20 neoantigens, about 10-30 neoantigens, about 10-40 neoantigens, about 10-50 neoantigens, about 10-60 neoantigens, about 10-70 neoantigens, about 10-80 neoantigens, about 10-90 neoantigens, or about 10-100 neoantigens. For example, the immunogenic composition can comprise about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 neoantigens. In some embodiments, the immunogenic composition can comprise about 19 neoantigens. Each of the neoantigens in the immunogenic composition can be different.

[0113] The immunogenic composition described herein can further comprise an adjuvant.Adjuvants are any substance whose admixture into an immunogenic composition increase, or otherwise enhances and / or boosts, the immune response to a tumor-specific neoantigen, but when the substance is administered alone does not generate an immune response to a tumorspecific neoantigen. The adjuvant can generate an immune response to the neoantigen and does not produce an allergy or other adverse reaction. It is contemplated herein that adjuvant can beadministered before, together, concomitantly with, or after administration of the immunogenic composition.

[0114] Adjuvants can enhance an immune response by several mechanisms including, e.g., lymphocyte recruitment, stimulation of B and / or T cells, and stimulation of macrophages. When an immunogenic composition described herein comprises adjuvants or is administered together with one or more adjuvants, the adjuvants that can be used include, but are not limited to, mineral salt adjuvants or mineral salt gel adjuvants, particulate adjuvants, microparticulate adjuvants, mucosal adjuvants, immunostimulatory adjuvants, or combinations thereof. Examples of adjuvants include, but are not limited to, aluminum salts (alum) (such as aluminum hydroxide, aluminum phosphate, and aluminum sulfate), 3 De-O-acylated monophosphoryl lipid A (MPL) (see, GB 2220211), MF59 (Novartis), AS03 (Glaxo SmithKline), AS04 (Glaxo SmithKline), polysorbate 80 (Tween 80; ICL Americas, Inc ), imidazopyridine compounds (see, International Application No. PCT / US2007 / 064857, published as International Publication No.W02007 / 109812, published as U.S. Pat. App. publication No. US20090232844A1 and corresponding to U.S. Pat. No. 8,063,063), imidazoquinoxaline compounds (see, International Application No. PCT / US2007 / 064858, published as International Publication No.W02007 / 109813, published as U.S. Pat. App. publication No. US20090311288A1 and corresponding to U.S. Pat. No. 8,173,657) and saponins, such as QS21 (see, Kensil et al, in Vaccine Design: The Subunit and Adjuvant Approach (eds. Powell & Newman, Plenum Press, NY, 1995); U.S. Pat. No. 5,057,540). In some embodiments, the adjuvant is Freund's adjuvant (complete or incomplete). Other adjuvants are oil in water emulsions (such as squalene or peanut oil), optionally in combination with immune stimulants, such as monophosphoryl lipid A (see, Stoute et al, N. Engl. J. Med. 336, 86-91 (1997)). CpG immunostimulatory oligonucleotides have also been reported to enhance the effects of adjuvants in a vaccine setting. Other TLR binding molecules such as RNA binding TLR 7, TLR 8 and / or TLR 9 may also be used.

[0115] Other examples of useful adjuvants include, but are not limited to, chemically modified CpGs (e. ., CpR, Idera), Poly(I:C)(e.g., polyi:CI2U), poly ICLC, non-CpG bacterial DNAor RNA as well as immunoactive small molecules and antibodies such as cyclophosphamide, sunitmib, bevacizumab, Celebrex (celecoxib), NCX-4016, sildenafil, tadalafil, vardenafil,sorafenib, XL-999, CP-547632, pazopanib, ZD2171, AZD2171, ipilimumab, tremelimumab, and SC58175, which may act therapeutically and / or as an adjuvant.

[0116] An immunogenic composition described herein can be administered to a subject that has not been diagnosed with cancer, has been diagnosed with cancer, is already suffering from cancer, has recurrent cancer (e.g., relapse), or is at risk of developing cancer. An immunogenic composition can be administered to a subject that is resistant to other forms of cancer treatment (e.g., chemotherapy, immunotherapy, or radiation). Immunogenic composition can be administered to the subject prior to other standard of care cancer therapies (e.g., chemotherapy, immunotherapy, or radiation). An immunogenic composition can be administered to the subject concurrently, after, or in combination to other standard of care cancer therapies (e.g., chemotherapy, immunotherapy, or radiation). In some embodiments, the method includes administering an immunogenic composition to a patient in need thereof.

[0117] Immunogenic compositions described herein can be administered to a subject in an amount sufficient to elicit an immune response to the tumor-specific neoantigen and to destroy, or at least partially arrest, symptoms and / or complications. In embodiments, the immunogenic composition can provide a long-lasting immune response. A long-lasting immune response can be established by administering a boosting dose of the immunogenic composition to the subject. The immune response to the immunogenic composition can be extended by administering to the subject a boosting dose. In embodiments, at least one, at least two, at least three or more boosting doses can be administered to abate the cancer. A first boosting dose may increase the immune response by at least 50%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 1000%. A second boosting dose may increase the immune response by at least 50%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 1000%. A third boosting dose may increase the immune response by at least 50%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 1000%.

[0118] An amount adequate to elicit an immune response is defined as a “therapeutically effective dose.” Amounts effective for this use will depend on, e.g., the composition, the manner of administration, the stage and severity of the disease being treated, the weight and general state of health of the individual, and the judgment of the prescribing physician. It should be kept in mind that an immunogenic composition can generally be employed in serious disease states, thatis, life-threatening or potentially life-threatening situations, especially when the cancer has metastasized. In such cases, in view of the minimization of extraneous substances and the relative nontoxic nature of a neoantigen, it is possible and can be considered desirable by the treating physician to administer substantial excesses of immunogenic compositions (e.g., cancer vaccines).Definitions

[0119] All publications and patents cited in this disclosure are incorporated by reference in their entirety. To the extent the material incorporated by reference contradicts or is inconsistent with this specification, the specification will supersede any such material. The citation of any references herein is not an admission that such references are prior art to the present disclosure. When a range of values is expressed, it includes embodiments using any particular value within the range. Further, reference to values stated in ranges includes each and every value within that range. All ranges are inclusive of their endpoints and combinable. When values are expressed as approximations, by use of the antecedent “about,” it will be understood that the particular value forms another embodiment. Reference to a particular numerical value includes at least that particular value, unless the context clearly dictates otherwise. The use of “or” will mean “and / or” unless the specific context of its use dictates otherwise.

[0120] Various terms relating to aspects of the description are used throughout the specification and claims. Such terms are to be given their ordinary meaning in the art unless otherwise indicated. Other specifically defined terms are to be construed in a manner consistent with the definitions provided herein. The techniques and procedures described or referenced herein are generally well understood and commonly employed using conventional methodologies by those skilled in the art, such as, for example, the widely utilized molecular cloning methodologies described in Sambrook et al., Molecular Cloning: A Laboratory Manual 4th ed. (2012) Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY. As appropriate, procedures involving the use of commercially available kits and reagents are generally carried out in accordance with manufacturer-defined protocols and conditions unless otherwise noted.

[0121] As used herein, the singular forms “a,” “an,” and “the” include plural forms unless the context clearly indicates otherwise. The terms “include,” “such as,” and the like are intended to convey inclusion without limitation, unless otherwise specifically indicated.

[0122] Unless otherwise indicated, the terms “at least,” “less than,” and “about,” or similar terms preceding a series of elements or a range are to be understood to refer to every element in the series or range. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described herein. Such equivalents are intended to be encompassed by the following claims.

[0123] As used herein, the term “about” is used to refer to an amount that is approximately, nearly, almost, or in the vicinity of being equal to or is equal to a stated amount, e.g., the stated amount plus / minus about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, about 2% or about 1%.

[0124] The terms “patient” or “individual” (e.g., individual of a cohort of interest) as used herein refer to any animal, such as any mammal, including, but not limited to, humans, non-human primates, rodents, mammals commonly kept as pets (e.g., dogs and cats, among others), livestock (e.g., cattle, sheep, goats, pigs, horses, and camels, among others) and the like. In some embodiments, the mammal is a mouse. In some embodiments, the mammal is a human.EQUIVALENTS

[0125] It will be readily apparent to those skilled in the art that other suitable modifications and adaptions of the methods of the invention described herein are obvious and may be made using suitable equivalents without departing from the scope of the disclosure or the embodiments. Having now described certain compositions and methods in detail, the same will be more clearly understood by reference to the following example, which is introduced for illustration only and not intended to be limiting.EXAMPLESExample 1- Actionable Variant Identification

[0126] Four patients undergoing immunotherapy were followed over the course of their treatment, including initial tumor biopsy (“TumorBx”), baseline biopsies pre-treatment(“Baseline”), as well as follow-up during treatment every 4 weeks (“Week 0 to 24”). At each follow-up timepoint, at least 10ml of peripheral blood samples were collected from the patients, accessioned and stored in Streck tubes to stabilize cell-free DNA (cfDNA). The tumor tissue was initially profiled to identify sequencing variants in order to create a personalized sequencing panel. Of the prioritized list of variants from each individual tissue tumor profile, 400 variants were chosen for the personalized sequencing panel, creating 4 unique panels for each patient. The personalized sequencing panels were then used to enrich the prioritized variants prior to sequencing. Each variant was sequenced to ~100,000x raw depth and -3,000 collapsed depth (after removing duplicate reads due to amplification). Each variant was then assessed to be either “confirmed” or “not confirmed” in the confirmation assay. An estimated tumor fraction (% of molecules with tumor origin in blood) was available for each time point for each patient. It was demonstrated that across a wide range of tumor fractions, even tumor fractions considered very low (<0.1%) a large percentage of variants were confirmed (Table 1). Detailed analysis with manual reviews of the missing variants in the initial tumor biopsy sequencing data revealed that tissue variants not found in plasma show these variants are likely false positives from the initial discovery assay (e g. low variant fraction analyzed with low sequencing depth, noisy sequencing data output, poor-quality variant calls, discordant variant calls in different assays) further validating the use of the personalized sequencing panel use in the confirmation assay.Table 1. Confirmation of Actionable Variants*Patient B entered tumor remission in week 12 - “Partial Response” (PR) according to RECIST criteria.Example 2: Classifier Predictive Performance of a Dual Assay

[0127] To evaluate the dual assay workflow’s predictive performance, subject sequencing data was constructed using a deep duplex sequencing run. The deep duplex sequencing run used a sequencing depth of 22,000X. As duplex sequencing tags each strand of double-stranded DNA, there is a significant reduction in noise and mutations are predicted with higher accuracy, making it an adequate proxy for quantifying predictive performance of a classifier. The sequencing data was constructed to simulate a normal sequence run where the sequencing depth is 5000X. Theconstructed sequencing data retained only fragment length and did not differentiate between simplex and duplex reads. The number of fragments were downsampled so that the total number of reference (REF) fragments and alternative (ALT) fragments (i.e., variant containing fragments) were at most 5000. The distribution of the downsampled ALT fragments were compared to a reference distribution of fragment lengths constructed from an independent sample’s buffy coat sample sequencing. These features were then used to train a machine learning model, that is, a random forest classifier.

[0128] The random forest classifier was trained to predict which of the variants detected in the normal sequencing run (n = 99,264 detected variants) would be identified to be found in the deep duplex sequencing read (n = 1,527 detected variants) with at least weak evidence. The predictive performance of the classifier was then evaluated using out-of-bag (OOB) predictions. The OOB samples, that is about a third of the data left out of the training set, was input into the classifier to predict actionable variants. Recall was set as five times more important than precision, meaning missing an actionable variant was five times worse than identifying a variant that would not be confirmed using deep sequencing methodologies. Area under the curve (AUC) was then quantified to measure how well the classifier distinguished between an actionable variants (i.e., true positive) and a false positive variant.

[0129] The classifier identified 807 variants of the 1527 variants of the deep duplex sequencing confirmed variants and predicted 11,583 variants total demonstrating good predictive performance. This would significantly reduce the list of prioritized variants by about tenfold. The AUC value equaled 0.766 with p-value<le-6, further indicating the discriminative ability of the classifier (FIG. 2). The classifier was further assessed to quantify how many variants exceed a probability threshold and how likely those variants are to be confirmed, along with predictive outcome measures to demonstrate predictive performance of the classifier (FIG. 3 and Table 2). As this classifier only considered downsampled length, inclusion of other descriptors of variants will further increase predictive performance.Table 2. Classifier Predictive Performance

Claims

CLAIMS1. A method of detecting actionable variants in a liquid biopsy comprising the steps of:(a) obtaining a test sample from a liquid biopsy;(b) running a discovery assay, wherein the discovery assay comprises sequencing the test sample; (c) identifying a list of prioritized variants;(d) designing a personalized sequencing panel based on the list of prioritized variants;(e) running a confirmation assay, wherein the confirmation assay comprises sequencing the test sample using the personalized sequencing panel; and(f) detecting one or more actionable variants in the liquid biopsy.

2. The method of claim 1, wherein the liquid biopsy is an amniotic fluid sample, ascitic fluid sample, bile sample, blood sample, buccal sample, cerebral spinal fluid sample, fecal sample, hair sample, peritoneal fluid sample, plasma sample, pleural effusion sample, saliva sample, semen sample, serum sample, skin sample, synovial fluid sample, urine sample, non-solid tumor sample, or a combination of any of the foregoing.

3. The method of claim 1 or 2, wherein the test sample is cell -free DNA, cell-free RNA or a combination thereof.

4. The method of any one of claims 1-3, wherein the cell-free DNA comprises circulating tumor DNA.

5. The method of any one of claims 1-4, wherein the discovery assay has high sensitivity to detect variants.

6. The method of any one of claims 1-5, wherein the high sensitivity is about 95% or more.

7. The method of any one of claims 1-6, wherein the discovery assay has a variant panel, comprising a high number of variant sites.

8. The method of claim 7, wherein the variant panel is used for whole genome sequencing or whole exome sequencing.

9. The method of claim 7 or 8, wherein the high number of variant sites is about 50000000 variant sites.

10. The method of any one of claims 1-9, wherein the discovery assay has a low sequencing depth.

11. The method of claim 10, wherein the low sequencing depth is between about 200X to about 5000X.

12. The method of any one of claims 1-11, wherein the variant caller is a bioinformatics tool capable of quantifying a score based on the confidence or quality of a variant called.

13. The method of any one of claims 1-12, wherein the bioinformatics tool capable is a machine learning model.

14. The method of any one of claims 1-13, wherein the variant caller is DRAGEN or Mutect2.

15. The method of claim 14, wherein the score is based on Somatic Qualify (SQ) quantified by DRAGEN or tumor-likelihood (TLOD) quantified by Mutect2.

16. The method of any one of claims 1-15, wherein the list of prioritized variants comprises a select number of variants chosen based on the confidence or qualify of a variant called.

17. The method of claim 16, wherein the select number of variants is between about 100 to about 10000 variants.

18. The method of any one of claims 1-17, wherein the personalized sequencing panel comprises the select number of variants.

19. The method of any one of claims 1-18, wherein the confirmation assay using the personalized sequencing panel has high specificity to detect the actionable variants.

20. The method of any one of claims 1-19, wherein the confirmation assay has a high sequencing depth to detect the actionable variants.

21. The method of claim 20, wherein the high sequencing depth is about 10000X to about 100000X.

22. The method of any one of claims 1-21, wherein the confirmation assay detects the actionable variants in the liquid biopsy using the variant caller.

23. A method of detecting actionable variants in a liquid biopsy collected from a patient having or suspected to have cancer, comprising the steps of:(a) obtaining a test sample comprising cell-free DNA or cell-free RNA from a liquid biopsy; (b) running a discovery assay on cell-free DNA, wherein the discovery assay comprises sequencing about 50000000 variant sites at a low sequencing depth between about 200X to about 5000X sequencing depth;(c) identifying a list of prioritized variants using a bioinformatic tool capable of quantifying a score based on confidence or quality of variants called;(d) designing a personalized sequencing panel based on the list of prioritized variants;(e) running a confirmation assay, wherein the confirmation assay comprises sequencing the cell-free DNA and / or the cell-free RNA using the personalized sequencing panel at 100000X sequencing depth;(f) detecting variants using a variant caller; and(g) identifying actionable variants from the detected variants.

24. The method of claim 23, wherein the test sample comprises two blood samples collected simultaneously or sequentially from the patient having or suspected to have cancer.

25. The method of claim 23 or 24, wherein the personalized sequencing panel can be used to monitor the patient having or suspected to have cancer.

26. The method of any one of claims 23-25, wherein the personalized sequencing panel can be used to monitor the patient having or suspected to have cancer during and / or after treatment.

1. The method of any one of claims 23-26, wherein the personalized sequencing panel can be redesigned to add or remove variants not of interest.

Citation Information

Patent Citations

  • Modified lipopolysaccharides

    GB2220211A

  • Immunopotentiating compounds

    US20090232844A1

  • Imidazoquinoxaline compounds as immunomodulators

    US20090311288A1

  • Predicting immunogenicity of t cell epitopes

    US20220093209A1

  • Ranking neoantigens for personalized cancer vaccine

    US20230173045A1