Structural variant detection in circulating tumor DNA

A machine learning model accurately predicts the presence of ctDNA-derived SVs in tumor tissue, addressing the challenge of identifying these variants in liquid biopsies and enhancing cancer treatment options.

WO2026072916A1PCT designated stage Publication Date: 2026-04-02AMAZON TECH INC
View PDF 16 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing sequencing methodologies struggle to accurately identify structural variants (SVs) in circulating tumor DNA (ctDNA) isolated from liquid biopsies, as these variants are rarely found in tumor tissue, hindering the development of effective cancer immunotherapies and vaccines.

Method used

A machine learning model is trained using paired tumor tissue and ctDNA sequencing results to predict the probability of ctDNA-derived SVs being present in tumor tissue, employing short-read or long-read sequencing techniques and a structural variant caller to identify true positive SVs.

Benefits of technology

The model achieves favorable sensitivity and positive predictive values in identifying ctDNA-derived SVs, enabling the selection of appropriate treatment targets and expanding the pool of immunotherapy candidates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025048129_02042026_PF_FP_ABST
    Figure US2025048129_02042026_PF_FP_ABST
Patent Text Reader

Abstract

This disclosure relates to methods of predicting the probability that a circulating tumor DNA (ctDNA) structural variant (SV) is present in tumor tissue using a machine learning model. The methods can further include training and validating the machine learning model using paired ctDNA-derived SVs and tumor tissue-derived SVs. The training set data can include multiple passes of quantified sensitivity values and positive predictive values calculated from known true positive SVs, false positive SVs or false negative SVs until a favorable sensitivity values and / or favorable positive predictive value is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.: 146401.094703STRUCTURAL VARIANT DETECTION IN CIRCULATING TUMOR DNACROSS-REFERENCE TO RELATED APPLICATIONS[0000J The present application claims the benefit of United States Provisional Application No.63 / 701,027, filed on September 30, 2024, the entire contents of which are incorporated herein by reference.BACKGROUND

[0001] Advances in sequencing methodologies, such as high-throughput sequencing, automation, massively parallel sequencing, and nucleotide barcoding, have increased the precision and utility of sequencing for biological samples. This has decreased the cost of sequencing and increased the quantity and quality of sequencing reads. See G. Dorado, et al. Analyzing Modern Biomolecules: The Re volution of Nucleic-Acid Sequencing - Re view . Biomolecules (2021) 11 (8), 1111, 1-18. Thus, as sequencing methodologies continue to advance they can be employed to identify structural variations, such as gene deletions, duplications, inversions, translocations, or gene fusion, which play a pivotal role in cancer development and progression.

[0002] Cancer immunotherapy (e.g., cancer vaccine) has emerged as a promising cancer treatment modality. The goal of cancer immunotherapy is to harness the immune system for selective destruction of cancer while leaving normal tissues unharmed. Traditional cancer vaccines typically target tumor-associated antigens. Tumor-associated antigens are typically present in normal tissues, but overexpressed in cancer. However, because these antigens are often present in normal tissues immune tolerance can prevent immune activation. Several clinical trials targeting tumor-associated antigens have failed to demonstrate a durable beneficial effect compared to standard of care treatment. Li et al., Ann Oncol., 28 (Suppl 12): xii 11— xii 17 (2017). Thus, neoantigens derived from structural variants (SV) are of particular interest, as SVs are a unique biological feature of tumor tissue as compared to normal tissue. iAmazon Privileged and ConfidentialAttorney Docket No.: 146401.094703

[0003] Existing sequencing methodologies allow for the identification of SVs in isolated cell- free DNA collected from liquid biopsies. Advantageously, this approach does not require surgical resection of tumor material, and the biological sample can be collected quickly and painlessly, from sources such as a blood sample, fecal sample, saliva sample, or the like. Further, collecting liquid biopsies can allow for identification of SVs in early-stage cancers and pre-cancers, where a tumor has not been identified for tumor tissue collection.

[0004] However, detection of reliable SVs (i.e., SVs present in a tumor), in ctDNA isolated from liquid biopsies presents a challenge, as ctDNA-derived SVs are rarely found in tumor tissue. Thus, there is an unmet need to accurately identify ctDNA-derived SVs likely to be present in tumor tissue. Identification of reliable ctDNA-derived SVs would be vital for the selection of appropriate treatment and would expand the pool of potential targets for immunotherapy and vaccines.SUMMARY OF THE INVENTION

[0005] This disclosure relates to methods of predicting a probability score that a circulating tumor DNA (ctDNA)-derived structural variant (SV) is present in tumor tissue. Genetic material can be isolated from a liquid biopsy and the genetic material can be ctDNA. The ctDNA can be sequenced to obtain sequencing results. The sequencing results can be input into a structural variant caller to identify ctDNA-derived SVs. The identified ctDNA-derived SVs can be input into a machine learning model. The machine learning model can have an output of the probability score that the ctDNA-derived SV is a true positive, wherein the true positive is a ctDNA-derived SV present in the tumor tissue.

[0006] This disclosure also relates to methods of training a machine learning model to output a probability score that a structural variant (SV) in circulating tumor DNA (ctDNA) is present in tumor tissue. Genetic material from a paired tumor tissue sample and liquid biopsy sample collected from a patient in which the genetic material from the liquid biopsy is ctDNA, can be sequenced to obtain tumor tissue sequencing results and ctDNA sequencing results. The tumor tissue sequencing result and the ctDNA sequencing result can be input into a structural variant caller to obtain an output of tumor tissue-derived SVs and ctDNA-derived SVs. The output can be received identifying true positive SVs, false positive SVs, and false negative SVs from the 2Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 tumor tissue-derived SVs and ctDNA-derived SVs sequencing results. Training set data can be input into the machine learning model, wherein the training set data can be the true positive SVs, false positive SVs, false negative SVs, a calculated sensitivity value, a calculated positive predictive value, and / or identified predictors of true positive SVs based on the calculated sensitivity value and the calculated positive predictive value. The methods of this disclosure can repeat steps until a favorable sensitivity value and / or a favorable positive predictive value is achieved.

[0007] In some aspects, the liquid biopsy can be collected from an amniotic fluid sample, ascitic fluid sample, bile sample, blood sample, buccal sample, cerebral spinal fluid sample, fecal sample, hair sample, peritoneal fluid sample, plasma sample, pleural effusion sample, saliva sample, semen sample, serum sample, skin sample, synovial fluid sample, urine sample, or a combination thereof.

[0008] The methods of this disclosure relate to genetic material that can be a nucleic acid selected from DNA, RNA, or a combination thereof.

[0009] In some aspects, the genetic material can be a nucleic acid selected from cell free DNA (cfDNA), circulating tumor DNA (ctDNA), cell free RNA (cfRNA), circulating tumor RNA (ctRNA), or a combination of any of the foregoing.

[0010] The sequencing results can be DNA sequencing results, RNA sequencing results, or a combination thereof.

[0011] The sequencing results can be obtained using either short-read sequencing or long-read sequencing.

[0012] In some aspects, the structural variant caller is DRAGEN WES, DRAGEN RNA, DRAGEN WGS or similar.

[0013] The methods of this disclosure relate to tumor tissue-derived SVs that can be a deletion, duplication, inversion, translocation, gene fusion, or any combination thereof.

[0014] The methods of this disclosure relate to ctDNA-derived SVs that can be a deletion, duplication, inversion, translocation, gene fusion, or any combination thereof.

[0015] The methods of this disclosure relate to true positive SVs that can be ctDNA-derived SVs present in the tumor tissue.3Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703

[0016] The methods of this disclosure relate to false negative SVs that can be tumor tissue- derived SVs not present in the ctDNA.

[0017] The methods of this disclosure relate to the false positive SVs that can be ctDNA-derived SVs not present in the tumor tissue.

[0018] The methods of this disclosure relate to a sensitivity value that can be calculated using formula (I): -T ue Positlve- ' x 100. As described above, the methods of the\True Positive + False Negative disclosure can relate to a true positive that can be a number of ctDNA-derived SVs present in tumor tissue and a false negative that can be a number of tumor tissue-derived SVs present in the tumor tissue sample but not present in the ctDNA.

[0019] The methods of this disclosure relate to a positive predictive value that can be calculated using formula (II): f -rrMe Positlve- \Xdescribed above, the disclosure can\True Positive+False Positive relate to a true positive that can be a number of ctDNA-derived SVs present in tumor tissue and a false positive that can be a number of ctDNA-derived SVs present in the ctDNA but not present in the tumor tissue.

[0020] In some aspects, the favorable sensitivity value can be between about 50% and about 100%.

[0021] In some aspects, the favorable positive predictive value can be between about 50% and about 100%.

[0022] The methods of this disclosure can further relate to predictors that can be recurrent true positive SVs within a gene type and / or a SV class.

[0023] In some aspects, the gene type can be a kinase, a tumor suppressor gene, a phosphatase, an oncogene, an essential gene, a spliceosome, or a protease.

[0024] In some aspects, the SV class can be a deletion, translocation, insertion, duplication, or a gene fusion.

[0025] In some aspects, the machine learning model can be a classification algorithm.

[0026] The methods of this disclosure can relate to the classification algorithm that can be selected from logistic regression, naive Bayes, K-nearest neighbor (KNN), decision tree, support vector machine, K-means clustering, random forest, artificial neural network (ANN), or a combination of any of the foregoing.4Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703

[0027] The methods of this disclosure can further relate to validating the machine learning model by inputting a known true positive SV into the machine learning model and receiving an output of the probability score indicating the ctDNA-derived SV is a true positive SV.

[0028] The methods of this disclosure relate to identifying one or more true positive SV, wherein the true positive SV can be a neoantigen.

[0029] In some aspects, the methods can include selecting a neoantigen for inclusion in an immunogenic composition.

[0030] In some aspects, the methods can include generating an immunogenic composition.

[0031] In some aspects, the methods can include administering the immunogenic composition to a patient in need thereof.BRIEF DESCRIPTION OF THE DRAWINGS

[0032] FIGs. 1A -ID illustrate plasma / tissue concordance, showing the number of ctDNA- derived SV, tissue-derived SVs and corresponding true positive ctDNA-derived SVs. FIG 1A. shows the number of SVs identified when ctDNA sequencing data was run through DRAGEN PASS whole exome sequencing (WES) coding compared to the number of tumor tissue-derived SVs run through DRAGEN PASS WES coding and overlapping SVs. The calculated sensitivity was 3.89% and the positive predictive value was 7.10%. FIG. IB shows the number of SVs identified when ctDNA sequencing data was run through DRAGEN PASS WES coding compared to the number of SVs identified when tumor tissue RNA was validated and overlapping SVs. The calculated sensitivity was 4.82% and the positive predictive value was 4.01%. FIG. 1C shows the number of SVs identified from ctDNA sequencing data when run through DRAGEN Non-PASS WES coding compared to the number of tumor tissue-derived SVs was run through DRAGEN Non-PASS WES coding and overlapping SVs. The calculated sensitivity was 8.53% and the positive predictive value was 4.81%. FIG. ID shows the number of SVs identified from the ctDNA sequencing data when run through DRAGEN Non-PASS WES coding compared to the number of SVs identified when tumor tissue RNA was validated and overlapping SVs. The calculated sensitivity was 7.59% and the positive predictive value was 1.83%.Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703

[0033] FIG. 2 is a Box and Whisker plot that shows the number of SVs called in tumor tissue using DRAGEN in each cancer indication for the Cohort- 1.DETAILED DESCRIPTION

[0034] Comprehensive genomic characterization stands to uncover novel targets for the development of immunotherapy or vaccines. Solid tumors shed cell-free DNA (cfDNA) and cell- free RNA (cfRNA) into the bloodstream presenting the opportunity for genomic characterization, especially in settings where tumor biopsies are not accessible. Characterization of structural variations, such as gene deletions, duplications, inversions, translocations and gene fusion have previously demonstrated a crucial role in cancer development and progression through impacting gene function and regulation. Further, being able to identify SVs in circulating tumor DNA (ctDNA) would support noninvasive analysis of tumor-derived SVs and overcome obstacles faced in early-stage cancer or inoperable tumors. However, ctDNA-derived SVs are rarely found on tumor tissue, thus, there is an unmet need to discern ctDNA-derived SVs that are present on a tumor.

[0035] This disclosure relates to methods of predicting the probability that a ctDNA-derived SV is present in tumor tissue. As described herein, machine learning models can be trained to output the probability that a ctDNA-derived SV is present in tumor tissue. The trained machine learning model can be trained and validated using paired tumor tissue and ctDNA isolated from liquid biopsies. SVs can be discovered and / or identified using short-read sequencing (e.g., Illumina) or long-read sequencing e.g., PacBio, Oxford Nanopore). Sequencing performed on paired tumor tissue and ctDNA samples can be used for training a machine learning model and, subsequently, using the machine learning model to identify true positive ctDNA-derived SVs, wherein a “true positive” is defined as a ctDNA-derived SV that is present in tumor tissue.

[0036] Methods described herein can include sequencing genetic material from a biological sample collected from a patient to obtain a sequencing result. Any type of genetic material can be used in the methods described herein. The genetic material can contain a deoxyribonucleic acid (DNA, e.g., an oligonucleotide containing 2’ -deoxyribonucleotides). The DNA can be cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), DNA from circulating tumor cells (CTCs), mitochondrial DNA, nuclear DNA (i.e., DNA from the nucleus of a cell), complementary DNA 6Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703(cDNA), or a combination of any of the foregoing. The DNA can be from a single cell. The DNA can be from a plurality of single cells. The DNA can be from a bulk lysis of tissue (e.g., bulk lysis of normal tissue, bulk lysis of neoplastic tissue). The genetic material can contain a ribonucleic acid (RNA, e.g., an oligonucleotide containing ribonucleotides). The RNA genetic material can be messenger RNA (mRNA), short-interfering RNA (siRNA), microRNA (miRNA), circular RNA (circRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), small nucleolar RNA (snRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), sub-genomic RNA (sgRNA), RNA from integrating or non-integrating viruses, a combination of any of the foregoing, or any other RNA. The RNA can be from a single cell. The RNA can be from a plurality of single cells. The RNA can be from a bulk lysis of tissue. The genetic material can contain RNA, DNA, cfDNA, cfRNA (cell-free RNA), circulating tumor DNA (ctDNA), circulating tumor RNA (ctRNA), or combinations of any of the foregoing. The sequencing result can contain DNA sequencing results, RNA sequencing results, or a combination thereof. In some embodiments, the genetic material contains a nucleic acid selected from DNA, RNA, or a combination thereof. In some embodiments, the genetic material comprises a nucleic acid selected from cell free DNA (cfDNA), circulating tumor DNA (ctDNA), cell free RNA (cfRNA), circulating tumor RNA (ctRNA), or a combination of any of the foregoing.

[0037] Genetic material for use in the method can be obtained (e.g., extracted) by processing the biological sample with any technique. The processing can include adding an anticoagulant (e.g., ethylenediaminetetraacetic acid (EDTA), EDTA salt (e.g., KiEDTA, Na2EDTA), citrate salt, citrate buffer, heparin, or heparin salt (e.g., sodium heparin)), centrifugation, separation of layers (e.g., separation of plasma, buffy coat, erythrocytes), filtration (e.g., to remove cells and / or debris), or a combination of any of the foregoing. The processing can include any technique to lyse a cell in the biological sample, including, but not limited to, physical lysis (e.g., grinding under liquid nitrogen, bead beating, French press, a grinder, sonication), enzymatic lysis (e.g., treatment with lysozyme, zymolase, lyticase, proteinase K, collagenase, lipase, and the like), chemical lysis (e.g, treatment with a detergent or surfactant (e.g, sodium dodecyl sulfate), chaotrope (e.g., guanidine salt, alkaline solution) chemical solvents), or a combination of any of the foregoing. The biological sample can be processed (e.g, lysed) in bulk. For example, the biological sample (e.g, tissue or cells of the biological sample) can be lysed in bulk via7Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 sonication to release the genetic material from the bulk biological sample (e.g., from all cells of the biological sample). As another example, the biological sample can be processed by bead beating of the sample. The biological sample (e.g., cells or tissue of the biological sample) can be processed on a single cell basis (e.g., by lysis of single cells). For example, a biological sample can be separated by flow cytometry and each cell of the biological sample can be isolated and lysed individually. Continuing this example, the genetic material from each cell can be barcoded to identify the originating cell and sequenced. Genetic material can be purified before sequencing by any purification technique including gel purification, centrifugal spin column purification, magnetic bead-based purification, extraction (e.g., phenol-chloroform extraction), and the like. The quantity and / or quality of the genetic material can be assessed prior to sequencing by any technique, such as optical density, absorbance (e. ., NanoDrop™ (THERMO FISHER SCIENTIFIC INC.) analysis), agarose gel electrophoresis, high-performance liquid chromatography (HPLC), fluorescent nucleic acid-binding dyes, microfluidics measurement (e.g., 2100 Bioanalyzer (AGILENT TECHNOLOGIES, INC.)), real-time quantitative PCR (RT- qPCR), reverse transcriptase qPCR, and the like.

[0038] Any biological sample collected from a patient can be used in the methods described herein. Biological samples (e.g., liquid biopsies) include, but are not limited to, an amniotic fluid sample, ascitic fluid sample, bile sample, blood sample, buccal sample, cerebral spinal fluid sample, fecal sample, hair sample, peritoneal fluid sample, plasma sample, pleural effusion sample, saliva sample, semen sample, serum sample, skin sample, synovial fluid sample, urine sample, tissue sample, tumor sample, or a combination of any of the foregoing. Blood samples can be separated to isolate a buffy coat (e.g., comprising leukocytes and / or platelets), a plasma, a serum, erythrocytes, or a combination of any of the foregoing. The biological sample can be a plasma sample, a serum sample, a buffy coat, an erythrocyte sample, or a combination of any of the foregoing. In some embodiments, the biological sample is a blood sample.

[0039] A biological sample used in the methods described herein can contain cells, tissue, or genetic material from any type of cancer, including hematological malignancies, solid tumors, sarcomas, carcinomas, and other solid and non-solid tumors. A patient from whom the biological sample is obtained can have any type of cancer. Illustrative cancerous tumors include, for example, adrenocortical carcinoma, anal cancer, appendiceal cancer, astrocytoma, basal cell 8Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 carcinoma, brain tumor, bile duct cancer, bladder cancer, bone cancer, breast cancer, bronchial tumor, carcinoma of unknown primary origin, cardiac tumor, cervical cancer, chordoma, colon cancer, colorectal cancer, craniopharyngioma, ductal carcinoma, embryonal tumor, endometrial cancer, ependymoma, esophageal cancer, esthesioneuroblastoma, fibrous histiocytoma, Ewing sarcoma, eye cancer, germ cell tumor, gallbladder cancer, gastric cancer, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor, gestational trophoblastic disease, glioma, head and neck cancer, hepatocellular cancer, histiocytosis, Hodgkin lymphoma, hypopharyngeal cancer, intraocular melanoma, islet cell tumor, Kaposi sarcoma, kidney cancer, Langerhans cell histiocytosis, laryngeal cancer, leukemias (e.g., acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myelogenous leukemia, hairy cell leukemia, myelodysplastic syndrome, prolymphocytic leukemia, large granular lymphocytic leukemia, adult T-cell leukemia, clonal eosinophilias), lip and oral cavity cancer, liver cancer, lobular carcinoma in situ, lung cancer, macroglobulinemia, malignant fibrous histiocytoma, melanoma, Merkel cell carcinoma, mesothelioma, metastatic squamous neck cancer with occult primary, midline tract carcinoma involving NUT gene, mouth cancer, multiple endocrine neoplasia syndrome, multiple myeloma, mycosis fungoides, myelodysplastic syndrome, myelodysplastic / myeloproliferative neoplasm, nasal cavity and par nasal sinus cancer, nasopharyngeal cancer, neuroblastoma, non-Hodgkin lymphoma, non-small cell lung cancer, oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, papillomatosis, paraganglioma, parathyroid cancer, penile cancer, pharyngeal cancer, pheochromocytomas, pituitary tumor, pleuropulmonary blastoma, primary central nervous system lymphoma, prostate cancer, rectal cancer, renal cell cancer, renal pelvis and ureter cancer, retinoblastoma, rhabdoid tumor, salivary gland cancer, Sezary syndrome, skin cancer, small cell lung cancer, small intestine cancer, soft tissue sarcoma, spinal cord tumor, stomach cancer, T-cell lymphoma, teratoid tumor, testicular cancer, throat cancer, thymoma and thymic carcinoma, thyroid cancer, urethral cancer, uterine cancer, vaginal cancer, vulvar cancer, and Wilms tumor. In some embodiments, the cancerous tumor is selected from a melanoma, leukemia, lymphoma, or breast cancer. The biological sample can contain cells, tissue, and / or genetic material from an adenoma, Barrett’s esophagus, a benign cancer, a cervical intraepithelial neoplasia (CIN), chronic lymphocytic leukemia (CLL), a clonal hematopoiesis of indeterminate potential (CHIP), a9Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 colorectal polyp, a monoclonal gammopathy of undetermined significance (MGUS), a precancer, a prostate cancer, smoldering multiple myeloma (SMM), or a combination of any of the foregoing. The biological sample can contain cells, tissue, and / or genetic material from an adenocarcinoma, esophageal cancer, cervical cancer, myelodysplastic syndrome, acute myeloid leukemia, colorectal cancer, myeloma (e. ., multiple myeloma, active myeloma, light chain myeloma, non-secretory myeloma, solitary plasmacytoma, immunoglobulin G myeloma, immunoglobulin A myeloma, immunoglobulin M myeloma, immunoglobulin E myeloma, immunoglobulin D myeloma), breast cancer, or a combination of any of the foregoing. The biological sample can contain cells, tissue, and / or genetic material from a primary tumor or secondary (e.g., metastatic) tumor of a patient.

[0001] Any sequencing technique can be used in the methods described herein. Sequencing techniques include, but are not limited to, whole genome sequencing (WGS), shotgun metagenomic sequencing, whole exome sequencing (WES), transcriptome sequencing, nextgeneration sequencing (NGS), cancer personalized profiling by deep sequencing (CAPP-Seq), tagged-amplicon deep sequencing (Tam-Seq), single nucleotide polymorphism (SNP) array, single cell sequencing, or combinations of any of the foregoing. Sequencing techniques can include short-read sequencing techniques (e.g., Illuminia) or long-read sequencing techniques (e.g., PacBio, Oxford Nanopore). In some instances, the foregoing techniques and procedures can be performed according to the methods described in e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual 4th ed. (2012) Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY. See also, Austell et al., Current Protocols in Molecular Biology, ed., Greene Publishing and Wiley-Interscience New York (1992) (with periodic updates). The methods as described herein can result in DNA sequencing, RNA sequencing or a combination thereof.

[0040] A patient can have (or a biological sample can contain genetic material from) any type, grade, or stage of cancer. The cancer can be any stage of cancer, including, but not limited to, stage 0, stage I, stage II, stage III, or stage IV. The patient can have, or the biological sample can contain genetic material from, a carcinoma in-situ. For example, the biological sample can contain tissue from a ductal carcinoma in-situ (DCIS). The cancer can be in remission (e.g., partial remission). The cancer can be a relapsed tumor (e.g., a tumor of a relapsed cancer). The cancer can have any grade, including, but not limited to, X, 1, 2, 3, or 4. The cancer can be a ioAmazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 recalcitrant tumor. The cancerous tumor can be resistant to therapy (e.g., resistant to chemotherapy, resistant to immunotherapy). The cancer can be susceptible to therapy e.g., susceptible to chemotherapy, susceptible to immunotherapy, susceptible to radiotherapy). The patient can have, or the biological sample can contain tissue, cells, and / or genetic material from a pre-cancerous tumor or a benign tumor (e.g., a benign tumor at risk for progressing to a cancerous tumor).

[0041] In one aspect, a method of predicting probability that a circulating tumor DNA (ctDNA)- derived structural variant (SV) is present in tumor tissue, can comprise the steps of (a) Isolating genetic material from a liquid biopsy, wherein the genetic material is ctDNA, (b) sequencing the ctDNA to obtain a sequencing result, (c) inputting the sequencing result into a structural variant caller to identify ctDNA-derived SVs, (d) inputting a ctDNA-derived SV into a machine learning model, (e) receiving an output from the machine learning model of the probability that the ctDNA-derived SV is a true positive, wherein the true positive is a ctDNA-derived SV present in the tumor tissue and (f) repeating steps (d)-(e) for all identified ctDNA-derived SVs.

[0042] Any one or more steps of the methods described herein can be repeated. For example, the step of sequencing genetic material from a biological sample collected from a patient to obtain a sequencing result can be repeated for multiple (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10) samples. Repeated steps of a method can be practiced in any order (e.g., repeated steps can be practiced in succession, repeated steps can be practiced not in succession (e.g., disconnected, with different steps in between the repeated steps)).

[0043] Steps of the method, repetitions of the method, and / or repetitions of steps of the method can be separated by any time interval. The time intervals can correlative with any known medical procedure and / or treatment. For example, the liquid biopsy may be taken before, during or after administration of a treatment. The time interval can be about 30 seconds, about 1 minute, about 2 minutes, about 3 minutes, about 4 minutes, about 5 minutes, about 6 minutes, about 7 minutes, about 8 minutes, about 9 minutes, about 10 minutes, about 11 minutes, about 12 minutes, about 13 minutes, about 14 minutes, about 15 minutes, about 16 minutes, about 17 minutes, about 18 minutes, about 19 minutes, about 20 minutes, about 22 minutes, about 24 minutes, about 26 minutes, about 28 minutes, about 30 minutes, about 33 minutes, about 36 minutes, about 39 minutes, about 40 minutes, about 44 minutes, about 48 minutes, about 50 minutes, about 55 nAmazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 minutes, about 1 hour, about 1 .5 hours, about 2 hours, about 2.5 hours, about 3 hours, about 3.5 hours, about 4 hours, about 4.5 hours, about 5 hours, about 5.5 hours, about 6 hours, about 6.5 hours, about 7 hours, about 7.5 hours, about 8 hours, about 8.5 hours, about 9 hours, about 9.5 hours, about 10 hours, about 10.5 hours, about 11 hours, about 11.5 hours, about 12 hours, about12.5 hours, about 13 hours, about 13.5 hours, about 14 hours, about 14.5 hours, about 15 hours, about 15.5 hours, about 16 hours, about 16.5 hours, about 17 hours, about 17.5 hours, about 18 hours, about 18.5 hours, about 19 hours, about 19.5 hours, about 20 hours, about 20.5 hours, about 21 hours, about 21.5 hours, about 22 hours, about 22.5 hours, about 23 hours, about 23.5 hours, about 1 day, about 1.5 days, about 2 days, about 3 days, about 4 days, about 5 days, about 6 days, about 1 week, about 2 weeks, about 3 weeks, about 4 weeks, about 1 month, about 1.5 months, about 2 months, about 2.5 months, about 3 months, about 3.5 months, about 4 months, about 4.5 months, about 5 months, about 5.5 months, about 6 months, about 6.5 months, about 7 months, about 7.5 months, about 8 months, about 8.5 months, about 9 months, about 9.5 months, about 10 months, about 10.5 months, about 11 months, about 11.5 months, about 1 year, about1.5 years, about 2 years, about 2.5 years, about 3 years, about 3.5 years, about 4 years, about 4.5 years, or about 5 years. The time interval can be at least 30 seconds, at least 1 minute, at least 2 minutes, at least 3 minutes, at least 4 minutes, at least 5 minutes, at least 6 minutes, at least 7 minutes, at least 8 minutes, at least 9 minutes, at least 10 minutes, at least 11 minutes, at least 12 minutes, at least 13 minutes, at least 14 minutes, at least 15 minutes, at least 16 minutes, at least 17 minutes, at least 18 minutes, at least 19 minutes, at least 20 minutes, at least 22 minutes, at least 24 minutes, at least 26 minutes, at least 28 minutes, at least 30 minutes, at least 33 minutes, at least 36 minutes, at least 39 minutes, at least 40 minutes, at least 44 minutes, at least 48 minutes, at least 50 minutes, at least 55 minutes, at least 1 hour, at least 1.5 hours, at least 2 hours, at least 2.5 hours, at least 3 hours, at least 3.5 hours, at least 4 hours, at least 4.5 hours, at least 5 hours, at least 5.5 hours, at least 6 hours, at least 6.5 hours, at least 7 hours, at least 7.5 hours, at least 8 hours, at least 8.5 hours, at least 9 hours, at least 9.5 hours, at least 10 hours, at least 10.5 hours, at least 11 hours, at least 11.5 hours, at least 12 hours, at least 12.5 hours, at least 13 hours, at least 13.5 hours, at least 14 hours, at least 14.5 hours, at least 15 hours, at least15.5 hours, at least 16 hours, at least 16.5 hours, at least 17 hours, at least 17.5 hours, at least 18 hours, at least 18.5 hours, at least 19 hours, at least 19.5 hours, at least 20 hours, at least 20.512Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 hours, at least 21 hours, at least 21.5 hours, at least 22 hours, at least 22.5 hours, at least 23 hours, at least 23.5 hours, at least 1 day, at least 1.5 days, at least 2 days, at least 3 days, at least 4 days, at least 5 days, at least 6 days, at least 1 week, at least 2 weeks, at least 3 weeks, at least 4 weeks, at least 1 month, at least 1.5 months, at least 2 months, at least 2.5 months, at least 3 months, at least 3.5 months, at least 4 months, at least 4.5 months, at least 5 months, at least 5.5 months, at least 6 months, at least 6.5 months, at least 7 months, at least 7.5 months, at least 8 months, at least 8.5 months, at least 9 months, at least 9.5 months, at least 10 months, at least 10.5 months, at least 11 months, at least 11.5 months, at least 1 year, at least 1.5 years, at least 2 years, at least 2.5 years, at least 3 years, at least 3.5 years, at least 4 years, at least 4.5 years, or at least 5 years. The time interval can be at most 30 seconds, at most 1 minute, at most 2 minutes, at most 3 minutes, at most 4 minutes, at most 5 minutes, at most 6 minutes, at most 7 minutes, at most 8 minutes, at most 9 minutes, at most 10 minutes, at most 11 minutes, at most 12 minutes, at most 13 minutes, at most 14 minutes, at most 15 minutes, at most 16 minutes, at most 17 minutes, at most 18 minutes, at most 19 minutes, at most 20 minutes, at most 22 minutes, at most 24 minutes, at most 26 minutes, at most 28 minutes, at most 30 minutes, at most 33 minutes, at most 36 minutes, at most 39 minutes, at most 40 minutes, at most 44 minutes, at most 48 minutes, at most 50 minutes, at most 55 minutes, at most 1 hour, at most 1.5 hours, at most 2 hours, at most 2.5 hours, at most 3 hours, at most 3.5 hours, at most 4 hours, at most 4.5 hours, at most 5 hours, at most 5.5 hours, at most 6 hours, at most 6.5 hours, at most 7 hours, at most 7.5 hours, at most 8 hours, at most 8.5 hours, at most 9 hours, at most 9.5 hours, at most 10 hours, at most 10.5 hours, at most 11 hours, at most 11.5 hours, at most 12 hours, at most 12.5 hours, at most 13 hours, at most 13.5 hours, at most 14 hours, at most 14.5 hours, at most 15 hours, at most 15.5 hours, at most 16 hours, at most 16.5 hours, at most 17 hours, at most 17.5 hours, at most 18 hours, at most 18.5 hours, at most 19 hours, at most 19.5 hours, at most 20 hours, at most 20.5 hours, at most 21 hours, at most 21.5 hours, at most 22 hours, at most 22.5 hours, at most 23 hours, at most 23.5 hours, at most 1 day, at most 1.5 days, at most 2 days, at most 3 days, at most 4 days, at most 5 days, at most 6 days, at most 1 week, at most 2 weeks, at most 3 weeks, at most 4 weeks, at most 1 month, at most 1.5 months, at most 2 months, at most 2.5 months, at most 3 months, at most 3.5 months, at most 4 months, at most 4.5 months, at most 5 months, at most 5.5 months, at most 6 months, at most 6.5 months, at most 7 months, at most 7.5 13Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 months, at most 8 months, at most 8.5 months, at most 9 months, at most 9.5 months, at most 10 months, at most 10.5 months, at most 11 months, at most 11.5 months, at most 1 year, at most 1.5 years, at most 2 years, at most 2.5 years, at most 3 years, at most 3.5 years, at most 4 years, at most 4.5 years, or at most 5 years. The time interval can be between about 1 minute and about 5 years, about 1 minute and about 4 years, about 1 minute and about 3 years, about 1 minute and about 2 years, about 1 minute and about 1 year, about 1 minute and about 6 months, about 1 minute and about 3 months, about 1 minute and about 1 month, about 1 minute and about 1 week, about 1 minute and about 1 day, about 1 minute and about 1 hour, about 1 hour and about 5 years, about 1 hour and about 4 years, about 1 hour and about 3 years, about 1 hour and about 2 years, about 1 hour and about 1 year, about 1 hour and about 6 months, about 1 hour and about 3 months, about 1 hour and about 1 month, about 1 hour and about 1 week, about 1 hour and about 1 day, about 1 day and about 5 years, about 1 day and about 4 years, about 1 day and about 3 years, about 1 day and about 2 years, about 1 day and about 1 year, about 1 day and about 6 months, about 1 day and about 3 months, about 1 day and about 1 month, about 1 day and about 1 week, about 1 month and about 5 years, about 1 month and about 4 years, about 1 month and about 3 years, about 1 month and about 2 years, about 1 month and about 1 year, about 1 month and about 6 months, about 1 month and about 3 months, about 1 year and about 5 years, about 1 year and about 4 years, about 1 year and about 3 years, or about 1 year and about 2 years.

[0044] Exemplary amounts of circulating tumor DNAin a biological sample (e.g., plasma or serum) can range from about 1 femtogram (fg) to about 1000 nanograms (ng), e.g., 1 picogram (pg) to 200 ng, 1 nanogram (ng) to 100 ng, 10 ng to 1000 ng. For example, the amount can be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of circulating tumor DNA acid molecules. The amount can be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng of cell-free nucleic acid molecules. The amount can be up to 1 fg, 10 fg, 100 fg, 1 pg, 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or 200 ng of circulating tumor DNA molecules. The method can comprise obtaining 1 fg to 200 ng circulating tumor DNA.

[0045] The circulating tumor DNA can have an exemplary size distribution of about 100-500 nucleotides. The circulating tumor DNA can be about 100, about 105, about 110, about 115, 14Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 about 120, about 125, about 130, about 135, about 140, about 145, about 150, about 155, about 160, about 165, about 170, about 175, about 180, about 185, about 190, about 195, about 200, about 210, about 215, about 220, about 225, about 230, about 235, about 240, about 245, about 250, about 255, about 260, about 265, about 270, about 275, about 280, about 285, about 290, about 295, about 300, about 305, about 310, about 315, about 320, about 325, about 330, about 335, about 340, about 345, about 350, about 355, about 360, about 365, about 370, about 375, about 380, about 385, about 390, about 395, about 400, about 405, about 410, about 415, about 420, about 425, about 430, about 435, about 440, about 445, about 450, about 455, about 460, about 465, about 470, about 475, about 480, about 485, about 490, about 495, or about 500 nucleotides.

[0046] Methods described herein can include identification of SVs from tumor tissue samples and / or ctDNA samples collected from a patient. SVs can be identified using a structural variant caller. The structural variant caller can be used to identify SVs on germline, somatic and population-level sequencing. The structural variant caller can be used to identify SVs using short-read sequencing data (e.g., DRAGEN WES, DRAGEN RNA, DRAGEN WGS). Assembly of short-read sequencing data sets can use any de novo assembly approach known in the art. For example, Bruijn graph-based assembly (e.g., Velvet, ABySS, Clover), overlap layout consensus (OLC) (e.g., Edena), string graph (e.g., SGA), greedy (e.g., SSAKE, Perga), and / or hybrid assembler algorithms (e.g., Ray). Short-read sequencing data can be paired-end or single-end data sets. The structural variant caller can be used to identify SVs using long-read sequencing data (e.g., Sv ABA, PBSV, Sniffles2). Long-read sequencing can be continuous long read (CLR), high-fidelity (HiFi), and / or Oxford Nanopore Technologies (ONT) sequencing. Assembly of long-read sequencing data sets can use any de novo assembly approach known in the art. For example, Canu, Flye, Miniasm, Raven and / or wtdbg2 for ONT and CLR reads, or HiCanu, Hifiasm, LJA and / or MBG for HiFi reads. The methods described herein can identify novel DNA adjacencies through joining multiple genomic parts of chimeric reads. Methods described herein can include inputting tumor tissue sequencing results or ctDNA sequencing results into any of the aforementioned structural variant callers and receiving an output of tumor tissue-derived SVs and / or ctDNA-derived SVs.Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703

[0047] Methods described herein can include a step of identifying one or more SVs from the sequencing results. For example, methods described herein can include a step of identifying one or more SVs from the tumor tissue sequencing result. For example, methods described herein can include a step of identifying one or more SVs from the ctDNA sequencing result. Any number of SVs can be identified from the tumor tissue sequencing results or the ctDNA sequencing results. The number of SVs can be about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about47, about 48, about 49, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 150, about 200, about 250, about 300, about 350, about 400, about 450, about 500, about 550, about 600, about 650, about 700, about 750, about 800, about 850, about 900, about 950, about 1000, about 1500, about 2000, about 2500, about 3000, about 3500, about 4000, about 4500, about 5000, about 5500, about 6000, about 6500, about 7000, about 7500, about 8000, about 8500, about 9000, about 10000, about 20000, about 30000, about 40000, about 50000, about 60000, about 70000, about 80000, about 90000, or about 100000. The number of SVs can be at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1500, at least 2000, at least 2500, at least 3000, at least 3500, at least 4000, at least 4500, at least 5000, at least 5500, at least 6000, at least 6500, at least 7000, at least 7500, at least 8000, at least 8500, at least 9000, at least 9500, at least 10000, at least 20000, at least 30000, at least 40000, at least 50000, 16Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 at least 60000, at least 70000, at least 80000, at least 90000, or at least 100000. The number of SVs can be at most 1, at most 2, at most 3, at most 4, at most 5, at most 6, at most 7, at most 8, at most 9, at most 10, at most 11, at most 12, at most 13, at most 14, at most 15, at most 16, at most 17, at most 18, at most 19, at most 20, at most 21, at most 22, at most 23, at most 24, at most 25, at most 26, at most 27, at most 28, at most 29, at most 30, at most 31, at most 32, at most 33, at most 34, at most 35, at most 36, at most 37, at most 38, at most 39, at most 40, at most 41, at most 42, at most 43, at most 44, at most 45, at most 46, at most 47, at most 48, at most 49, at most 50, at most 55, at most 60, at most 65, at most 70, at most 75, at most 80, at most 85, at most 90, at most 95, at most 100, at most 150, at most 200, at most 250, at most 300, at most 350, at most 400, at most 450, at most 500, at most 550, at most 600, at most 650, at most 700, at most 750, at most 800, at most 850, at most 900, at most 950, at most 1000, at most 1500, at most 2000, at most 2500, a most 3000, at most 3500, at most 4000, at most 4500, at most 5000, at most 5500, at most 6000, at most 6500, at most 7000, at most 7500, at most 8000, at most 8500, at most 9000, at most 9500, at most 10000, at most 20000, at most 30000, at most 40000, at most 50000, at most 60000, at most 70000, at most 80000, at most 90000, or at most 100000.

[0048] Methods described herein can include identification of SVs, wherein the SV is a gene mutation (e.g., deletion, duplication, inversion, translocation, gene function, any combination thereof). The gene mutation can result in amino acid changes of about 1, about, 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 22, about 24, about 26, about 28, about 30, about 33, about 36, about 40, about 44, about 48, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 110, about 120, about 130, about 140, about 150, about 160, about 170, about 180, about 190, about 200, about 220, about 240, about 260, about 280, about 300, about 330, about 360, about 390, about 400, about 440, about 480, about 500, about 550, about 600, about 650, about 700, about 750, about 800, about 850, about 900, about 950, about 1000, about 1100, about 1200, about 1300, about 1400, about 1500, about 1600, about 1700, about 1800, about 1900, about 2000, about 2200, about 2400, about 2600, about 2800, about 3000, about 4000, about 5000, about 6000, about 7000, about 8000, about 9000, about 10000, about 100000, about 1000000, about 10000000, about 100000000, or about 1000000000. The gene mutation can result in amino acid 17Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 changes of at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 22, at least 24, at least 26, at least 28, at least 30, at least 33, at least 36, at least 40, at least 44, at least 48, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 220, at least 240, at least 260, at least 280, at least 300, at least 330, at least 360, at least 390, at least 400, at least 440, at least 480, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1800, at least 1900, at least 2000, at least 2200, at least 2400, at least 2600, at least 2800, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10000, at least 100000, at least 1000000, at least 10000000, at least100000000, or at least 1000000000. The gene mutation can result in amino acid changes of at most 1, at most 2, at most 3, at most 4, at most 5, at most 6, at most 7, at most 8, at most 9, at most 10, at most 11, at most 12, at most 13, at most 14, at most 15, at most 16, at most 17, at most 18, at most 19, at most 20, at most 22, at most 24, at most 26, at most 28, at most 30, at most 33, at most 36, at most 40, at most 44, at most 48, at most 50, at most 55, at most 60, at most 65, at most 70, at most 75, at most 80, at most 85, at most 90, at most 95, at most 100, at most 110, at most 120, at most 130, at most 140, at most 150, at most 160, at most 170, at most 180, at most 190, at most 200, at most 220, at most 240, at most 260, at most 280, at most 300, at most 330, at most 360, at most 390, at most 400, at most 440, at most 480, at most 500, at most 550, at most 600, at most 650, at most 700, at most 750, at most 800, at most 850, at most 900, at most 950, at most 1000, at most 1100, at most 1200, at most 1300, at most 1400, at most 1500, at most 1600, at most 1700, at most 1800, at most 1900, at most 2000, at most 2200, at most 2400, at most 2600, at most 2800, at most 3000, at most 4000, at most 5000, at most 6000, at most 7000, at most 8000, at most 9000, at most 10000, at most 100000, at most 1000000, at most 10000000, at most 100000000, or at most 1000000000. The amino acid changes can be the length of any given gene within the genome.18Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703

[0049] Methods described herein can include a step of receiving as an output a probability score that a ctDNA-derived SV is a true positive SV (i.e., a ctDNA-derived SV is present in the tumor of a patient). The probability score can be for each ctDNA-derived SV identified by the method. The probability score can be for a plurality of ctDNA-derived SVs (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20). The probability score can be any statistical value or probability, such as a p-value, q-value, percentage, odds ratio, decimal, fraction, a combination of any of the foregoing, and the like. Any number of probability scores can be received as an output from the method. The probability score can be between about 0 to about 1.0, wherein a probability score of 1.0 means the ctDNA-derived SV is most likely a true positive and a probability score of 0 means the ctDNA-derived SV is most likely a false positive. The number of probability scores can be about 0, about 0.1, about 0.2, about 0.3, about 0.4, about 0.5, about 0.6, about 0.7, about 0.8, about 0.9, or about 1.0. The number of probability scores can be at least 0, at least 0.1, at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, at least 0.8, at least 0.9, or at least 1.0. The number of probability scores can be at most 1.0, at most 0.9, at most 0.8, at most 0.7, at most 0.6, at most 0.5, at most 0.4, at most 0.3, at most 0.2, at most 0.1, or at most 0. The number of probability scores can be between about 0 to about 1.0, about 0 to about 0.9, about 0 to about 0.8, about 0 to about 0.7, about 0 to about 0.6, about 0 to about 0.5, about 0 to about 0.4, about 0 to about 0.3, about 0 to about 0.2, about 0 to about 0.1, about 0.1 to about 1.0, about 0.2 to about 1.0, about 0.3 to about 1.0, about 0.4 to about 1.0, about 0.5 to about 1.0, about 0.6 to about 1.0, about 0.7 to about 1.0, about 0.8 to about 1.0, or about 0.9 to about 1.0.

[0050] In one aspect, a method of training a machine learning model to output a probability score that a structural variants (SV) in circulating tumor DNA (ctDNA) is present in tumor tissue, can comprise the steps of (a) sequencing genetic material from a paired tumor tissue sample and liquid biopsy sample collected from a patient, wherein the genetic material from the liquid biopsy is ctDNA, to obtain a tumor tissue sequencing result and a ctDNA sequencing result, (b) inputting the tumor tissue sequencing result and the ctDNA sequencing result into a structural variant caller to obtain an output of tumor tissue-derived SVs and ctDNA-derived SVs, (c) identifying true positive SVs, false positive SVs and false negative SVs from the output of tumor tissue-derived SVs and ctDNA-derived SVs, (d) inputting training set data into the 19Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 machine learning model, wherein the training set data comprises the true positive SVs, false positive SVs, false negative SVs, a calculated sensitivity value, a calculated positive predictive value and / or identified predictors of true positive SVs based on the calculated sensitivity value and the calculated positive predictive value and (e) repeating steps (a)-(d) until a favorable sensitivity value and / or a favorable positive predictive value is achieved.

[0051] This disclosure relates to methods of calculating a sensitivity value and / or a positive predictive value. Calculations of the sensitivity value uses formula (I): f -True Positlve- x 100; calculations of the positive predictive value (PPV) uses\True Positives False Negative formula (II): ( -True Positlve- ) x 100. The method disclosed herein relates to training xTrue Positive+False Positive set data that can include calculated sensitivity, specificity, positive predictive value, and / or negative predictive value. For example, the training set data can use sensitivity to train a machine learning model to predict if a ctDNA-derived SV is a true positive. For example, the training set data can use the positive predictive value to train a machine learning model to predict if a ctDNA-derived SV is a true positive.

[0052] The methods described herein refer to true positive SVs. True positive SVs as used herein refer to ctDNA-derived SVs or ctRNA-derived SVs that are present on tumor tissue. The methods described herein refer to false positive SV. False positive SVs as used herein refers to ctDNA-derived SV or ctRNA-derived SV that are not present on tumor tissue. The methods described herein refer to false negative SV. False negative SVs as used herein refers to tumor tissue-derived SVs that are not present in ctDNA or ctRNA samples. The methods as described herein can comprise a step of calculating a sensitivity value or a positive predictive value using formula (I) and formula (II). To calculate sensitivity the total number of true positive SVs and false negative SVs identified in a paired tumor tissue sample and ctDNA sample can be inputted into formula (I). To calculate PPV the total number of true positive SVs and false positives SVs identified in a paired tumor tissue sample and ctDNA sample can be inputted into formula (II). This step can be repeated based on collected paired tumor tissue sample and ctDNA sample. For example, the step can be repeated about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 times.

[0053] The sensitivity value can be about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 12%, about 14%, about 16%, about 18%, 20Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 about 20%, about 25%, about 30%, about 3 %, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or about 100%. The sensitivity value can be at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, at least 10%, at least 12%, at least 14%, at least 16%, at least 18%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%. The sensitivity value can be at most 1%, at most 2%, at most 3%, at most 4%, at most 5%, at most 6%, at most 7%, at most 8%, at most 9%, at most 10%, at most 12%, at most 14%, at most 16%, at most 18%, at most 20%, at most 25%, at most 30%, at most 35%, at most 40%, at most 45%, at most 50%, at most 55%, at most 60%, at most 65%, at most 70%, at most 75%, at most 80%, at most 85%, at most 90%, at most 95%, or at most 100%. The sensitivity value can be between about 50% to about 100%, about 55% to about 100%, about 60% to about 100%, about 65% to about 100%, about 70% to about 100%, about 75% to about 100%, about 80% to about 100%, about 85% to about 100%, about 90% to about 100%, or about 95% to about 100%.

[0054] A favorable sensitivity value can be about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 100%. A favorable sensitivity value can be at least at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%. A favorable sensitivity value can be at most at most 50%, at most 55%, at most 60%, at most 65%, at most 70%, at most 75%, at most 80%, at most 85%, at most 90%, at most 95%, or at most 100%. A favorable sensitivity value can be between about 50% to about 100%, about 55% to about 100%, about 60% to about 100%, about 65% to about 100%, about 70% to about 100%, about 75% to about 100%, about 80% to about 100%, about 85% to about 100%, about 90% to about 100%, or about 95% to about 100%.

[0055] The PPV can be about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 12%, about 14%, about 16%, about 18%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 100%. The PPV can be at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, 21Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 at least 7%, at least 8%, at least 9%, at least 10%, at least 12%, at least 14%, at least 16%, at least 18%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 100%. The PPV can be at most 1%, at most 2%, at most 3%, at most 4%, at most 5%, at most 6%, at most 7%, at most 8%, at most 9%, at most10%, at most 12%, at most 14%, at most 16%, at most 18%, at most 20%, at most 25%, at most30%, at most 35%, at most 40%, at most 45%, at most 50%, at most 55%, at most 60%, at most65%, at most 70%, at most 75%, at most 80%, at most 85%, at most 90%, at most 95%, at most100%. The sensitivity value can be between about 50% to about 100%, about 55% to about 100%, about 60% to about 100%, about 65% to about 100%, about 70% to about 100%, about 75% to about 100%, about 80% to about 100%, about 85% to about 100%, about 90% to about 100%, or about 95% to about 100%.

[0056] A favorable PPV value can be about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 100%. A favorable PPV value can be at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 100%. A favorable sensitivity value can be at most 50%, at most 55%, at most 60%, at most 65%, at most 70%, at most 75%, at most 80%, at most 85%, at most 90%, at most 95%, at most 100%. A favorable PPV value can be between about 50% to about 100%, about 55% to about 100%, about 60% to about 100%, about 65% to about 100%, about 70% to about 100%, about 75% to about 100%, about 80% to about 100%, about 85% to about 100%, about 90% to about 100%, about 95% to about 100%.

[0057] Predictors of true positive SVs can be any biological feature, including any key feature that can increase the probability that a ctDNA-derived SV is present in tumor tissue. For example, the predictor could be any gene type. Gene types can be a kinase, tumor suppressor, phosphatase, oncogene, essential gene, spliceosome, protease. For example, the predictor can be a SV class. The SV class can be a deletion, duplication, inversion, translocation, or gene fusion.

[0058] The methods described herein can include a step of validating the machine learning model. The steps as described herein can include obtaining paired tumor tissue and ctDNA sequencing results and identifying tumor tissue-derived SVs and ctDNA-derived SVs to identify true positive SVs. Inputting a known true positive ctDNA-derived SV into the machine learning 22Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 model and receiving a probability score indicating the ctDNA-derived SV is likely a true positive. The methods described herein can also validate a machine learning model through input of known true negative SVs and receiving an output of a probability score indicating the ctDNA- derived SV is likely not a true positive.Machine Learning Models

[0059] The methods described herein can input a ctDNA-derived SV and receive an output of the probability that the ctDNA-derived SV is a true positive SV (i.e., the ctDNA-derived SV is present in the tumor tissue). The machine learning model can operate by any type of machine learning methodology, including, but not limited to, supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, and combinations thereof.

[0060] In one aspect, the method uses a machine learning model to predict the probability that a ctDNA-derived SV is present in the tumor tissue. The machine learning model can predict the probability that a ctDNA-derived SV is present in the tumor tissue based on (e.g., based in part) the input of true positive SVs, false positive SVs, false negative SVs, a calculated sensitivity value, a calculated positive predictive value and / or identified predictors of true positive SVs based on the calculated sensitivity value and the calculated positive predictive value obtained from paired tumor tissue and ctDNA samples collected from a patient.

[0061] Supervised learning models can involve training a machine learning model or its algorithm using labeled data. Labeled data can comprise a training data set labeled with input and output values. The labeled data can teach the machine what elements it needs to recognize and how to identify labeled elements from raw data. The labeled data can be fed and refed into the machine learning model to train the machine and increase accuracy in arriving at an output for a new data set. The user can provide feedback on accuracy of the machine algorithm. Supervised learning algorithms can be any supervised learning algorithm, including, but not limited to, classification algorithms (e.g., k-nearest neighbor (KNN), naive Bayes classifier algorithms, support vector machine (SVM) algorithms, decision trees, random forest models, logistic regression, Proaftn, stochastic gradient descent algorithms, linear classifiers, and combinations thereof) and regression algorithms (e.g., simple linear regression algorithms, multivariate regression algorithms, decision tree algorithms, lasso regression algorithms, logistic 23Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 regression algorithms, ridge regression algorithms, and combinations thereof). In some embodiments, the machine learning model includes a classification algorithm. Any classification algorithm described herein can be used in the machine learning model. In some embodiments, the classification algorithm is selected from logistic regression, naive Bayes, K-nearest neighbor (KNN), decision tree, support vector machine, K-means clustering, random forest, artificial neural network (ANN), or a combination of any of the foregoing.

[0062] Unsupervised learning models can be trained on raw data without labels or tags. Unsupervised learning models can be used for clustering large, unstructured data sets. Unsupervised learning models can cluster data or perform dimensionality reduction of the data. Clustering examines similarities in raw data and groups the data by similarity, providing structure to unstructured raw data. Dimensionality reduction reduces the number of features in a data set, thereby reducing processing time, storage space, complexity, and overfitting in the machine learning model. Dimensionality reduction can include feature selection, wherein a subset of relevant features from the total original features present in the raw data are selected for use as an input. Dimensionality reduction can also include feature extraction, wherein a new set of features is generated from the existing data set. These new features can be used as further inputs for the machine learning model. Unsupervised learning models can operate under any unsupervised learning algorithm, including, but not limited to, clustering algorithms (e.g., hierarchical clustering, k-means clustering algorithm, density-based clustering, graph-based clustering, Cophenetic correlation, Bayesian information criterion, t-distributed stochastic neighbor embedding (t-SNE), uniform manifold approximation and projection (UMAP), and combinations thereof) and dimensionality reduction algorithms (e.g., principle component analysis (PCA), non-negative matrix factorization (NMF), linear discriminant analysis (LDA), generalized discriminant analysis (GDA), and combinations thereof).

[0063] Semi-supervised learning models can be a hybrid of supervised and unsupervised learning models. Semi-supervised learning models can utilize small amounts of labeled data processed alongside larger raw data sets. Any algorithm for semi-supervised learning can be used, including any unsupervised algorithm, any supervised algorithm, any modified unsupervised algorithm, and any modified supervised learning algorithm. Modified unsupervised and supervised algorithms include, but are not limited to, self-training algorithms (e.g., pseudo- 24Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 labeler algorithms using pre-existing supervised classifier model trained on a small portion of labeled data, followed by predictions on remainder of the data set (e.g., the unlabeled portion)), label propagation algorithms (e.g., algorithms that assign labels to unlabeled observations by propagating or allocating labels through the data set over time based on the labeled data set (e.g., a graph neural network)), and combinations thereof.

[0064] Reinforcement learning models can include any type of reinforcement learning model. Reinforcement learning models can be sensor models or neuron models. Sensor models can use a break algorithm or a sigmoid algorithm. The input of sensor models can be a real number vector. The input of neuron models can be a binary vector. Reinforcement learning models can use any reinforcement learning algorithm, including, but not limited to, temporal difference (TD) algorithms e.g., tabular TD(lamb da), replacing traces; Tabular TD(0); tabular TD(lambda) with function approximation; gradient temporal difference 2 (GTD2); least squares temporal difference (LSTD)), temporal distribution characterization (TDC), Monte Carlo methods (e.g., every-visit Monte-Carlo, first-visit Monte-Carlo), value iteration, policy iteration, Q-leaming with function approximation, Explicit Explore or Exploit (E3) algorithm, State-aggregation based Q-learning, fitted Q-iteration, generalized policy iteration, tabular q-leaming algorithm, upper confidence bound 1 algorithm, delayed-Q algorithm, MoRmax algorithm, upper confidence reinforcement learning 2 (UCRL2) algorithm, state-action-reward-state-action algorithm, R-Max algorithm, REINFORCE algorithm, model-based interval estimation (MB IE), Boltzmann exploration, action elimination, generalized policy iteration, natural actor-critic, soft- state aggregation based Q-learning, epsilon-greedy, and combinations thereof. Reinforcement learning models can be any deep reinforcement learning model. Deep reinforcement learning can be based on an artificial neural network.

[0065] The machine learning model can comprise an artificial neural network (ANN). An artificial neural network can comprise neurons. Neurons in an artificial neural network can receive signals, process them, and send signals to neurons connected to them. Neurons can be connected via links.

[0066] The connected pattern or network of neurons can be a directed, weighted graph. The neural network can comprise multiple layers. The layers can comprise input layers, hidden layers, and output layers. Input layers can receive raw data or labelled data. The raw data or 25Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 labelled data can be a training data set or a non-training data set. Input layers can also receive outputs from preceding machine learning models. Hidden layers are intermediate layers between output and input layers. Hidden layers can process data by applying functions to the data. Output layers can receive processed data from preceding layers and produce an output from the machine learning model. The output can be a final result. The output can be an input for a subsequent machine learning model.

[0067] The machine learning model can comprise any connection pattern between each pair of layers. The layers can be fully connected. Fully connected layers comprise layers in which every neuron in one layer is connected to every neuron in the subsequent layer. The layers can comprise pooling layers. Pooling layers comprise layers in which a group of neurons in one layer connects to a single neuron in the next layer. The machine learning model can contain any feedforward network, comprising pooling layers. The machine learning model can comprise any quantum neural network (QNN). Feedforward networks include, but are not limited to, convolutional neural network (CNN) (c. ., including, but not limited to, LeNet, AlexNet, ZF Net, GoogLeNet, VGGNet, ResNet, MobileNets, GoogLeNet_DeepDream, or combinations thereof), radial basis function network, linear neural network, perceptron, multilayer perceptron, and combinations thereof. The machine learning model can comprise any recurrent neural network (RNN), including, but not limited to, long-term memory (LSTM network), gated recurrent units (GRU), echo state network (ESN), fully recurrent neural network (FRNN), Elman network, Jordan network, Hopfield network, bidirectional associative memory, independently RNN (IndRNN), recursive neural network, neural history compressor, second order RNN, transformer (e.g., transformer architecture), continuous-time RNN (CTRNN), hierarchical RNN (HRNN), recurrent multilayer perceptron network (RMLP), multiple timescales RNN (MTRNN), neural Turing machine (NTM), differentiable neural computer (DNC), neural network pushdown automata (NNPDA), Memristive network, and combinations thereof. Recurrent neural networks can comprise connections between neurons in the same or prior layers. Recurrent neural networks can be bi-directional in the transmission of signal between neurons or layers.

[0068] The machine learning model can include a dropout layer. Dropout layers can be used to reduce statistical noise in a data set and overfitting. A dropout layer can randomly set input units to 0 with a frequency of rate at each step during training time. Many types and architectures of 26Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 dropout layers are known (e.g., U.S. Pat. No. 9,406,017), and any can be used in the methods described herein. A machine learning model can comprise any number of dropout layers. A machine learning model can comprise at least one dropout layer, e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 14, at least 16, at least 18, at least 20, at least 24, at least 28, at least 30, at least 33, at least 36, at least 39, at least 40, at least 44, at least 48, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 150, at least 200, at least 250, at least 300, or at least 350 dropout layers. A machine learning model can comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, 230, 235, 240, 245, 250, 255, 260, 265, 270, 275, 280, 285, 290, 295, 300, 305, 310, 315, 320, 325, 330, 335, 340, 345, or 350 dropout layers.

[0069] The machine learning model can comprise applying an activation function. Hidden layers and output layers can comprise activation functions. Activation functions can have different properties (e.g., linearity, range, differentiability). Activation functions can be linear. Activation functions can be non-linear. Activation functions can have a finite range. Activation functions can have an infinite range. Activation functions can be continuously differentiable. Any activation function can be applied in the methods described herein including, but not limited to, nonlinear, ridge activation functions (e.g., linear activation, rectified linear unit (ReLU), Heaviside activation, logistic activation), radial activation functions (e.g., Gaussian, multi quadratic, inverse multi quadratic, polyharmonic spine), folding activation function, and combinations thereof. Activation functions include, but are not limited to, identity, binary step, exponential linear unit (ELU), Gaussian, Gaussian error linear unit (GELU), hyperbolic tangent, logistic, leaky rectified linear unit, Maxout, parametric rectified linear unit (PreLU), ReLU (e.g., S-shaped ReLU, Sine ReLU), quantum activation functions, scaled exponential linear unit (SELU), ), sigmoid linear unit (e.g., SiLU, sigmoid shrinkage, SiL, Swish-1), Softplus, and combinations thereof. In the methods described herein, a machine learning model can comprise applying any number of activation functions. The number of activation functions can be at least one, e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at 27Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 least 10, at least 12, at least 14, at least 16, at least 18, at least 20, at least 24, at least 28, at least 30, at least 33, at least 36, at least 39, at least 40, at least 44, at least 48, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 150, at least 200, at least 250, at least 300, or at least 350 activation functions. The machine learning model can comprise applying 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, 230, 235, 240, 245, 250, 255, 260, 265, 270, 275, 280, 285, 290, 295, 300, 305, 310, 315, 320, 325, 330, 335, 340, 345, or 350 activation functions.

[0070] Learning in a machine learning model can comprise adjusting the weights of the machine learning model to improve accuracy of the task. Learning in a machine learning model can comprise adjusting the optional thresholds of the machine learning model to improve accuracy of the task. For example, the weights and optional thresholds of the machine learning model can be adjusted to improve the accuracy of predicting the sequencing depth, batch size, and / or amount of biological sample needed to obtain the desired sequencing output (e.g., the desired number of neoantigens or sequence variants). For example, the weights and optional thresholds of the machine learning model can be adjusted to improve the accuracy of predicting the likelihood that a neoantigen or sequence variant is a true positive.

[0071] Machine learning models in the methods described herein can be trained by any data set (e.g., any training data set). A training data set can be data of neoantigens and / or sequence variants identified from a biological sample (e.g., that is not a tumor tissue resection or tumor tissue biopsy, distal from a tumor of patient) and matched neoantigens and / or specific variants from a tumor biopsy (or tissue from a surgical resection of a tumor). A training data set can comprise neoantigens and / or sequence variants identified from a biological sample labeled to indicate whether each identified neoantigen and / or variant is a true positive, false positive, true negative, and / or false negative. A true positive can be a neoantigen or sequence variant identified in a distal biological sample (e.g, plasma sample, non-tumor tissue sample) that is found (e.g., expressed, mutated, identified) in a tumor of the patient. A false positive can be a neoantigen or sequence variant identified in a distal biological sample (e.g., plasma sample, non-tumor tissue 28Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 sample) that is not found (e.g., expressed, mutated, identified) in a tumor of the patient. A true negative can be a neoantigen or sequence variant that is not identified in a distal biological sample (e.g., plasma sample, non-tumor tissue sample) that is not found (e.g., expressed, mutated, identified) in a tumor of the patient. A false negative can be a neoantigen or sequence variant that is not identified in a distal biological sample (e.g., plasma sample, non-tumor tissue sample) that is found (e.g., expressed, mutated, identified) in a tumor of the patient. The method can include a step of training the machine learning model on training set data. In some embodiments, the training set data comprises matched neoantigens identified from the biological sample and neoantigens identified from a tumor biopsy.

[0072] Training data sets (e.g., negative training data sets, positive training data sets) can contain any number of data points. The number of data points in a training data set can be at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 14, at least 16, at least 18, at least 20, at least 22, at least 24, at least 26, at least 28, at least 30, at least 33, at least 36, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1800, at least 1900, at least 2000, at least 2200, at least2400, at least 2600, at least 2800, at least 3000, at least 3300, at least 3600, at least 3900, at least4000, at least 4400, at least 4800, at least 5000, at least 5500, at least 6000, at least 6500, at least7000, at least 7500, at least 8000, at least 8500, at least 9000, at least 9500, at least 10000, at least 11000, at least 12000, at least 13000, at least 14000, at least 15000, at least 16000, at least 17000, at least 18000, at least 19000, or at least 20000. The number of data points can be between about 1 to about 20000, about 5 to about 20000, about 10 to about 20000, about 15 to about 20000, about 20 to about 20000, about 25 to about 20000, about 30 to about 20000, about 35 to about 20000, about 40 to about 20000, about 45 to about 20000, about 50 to about 20000, about 55 to about 20000, about 60 to about 20000, about 65 to about 20000, about 70 to about 20000, about 75 to about 20000, about 80 to about 20000, about 85 to about 20000, about 90 to29Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 about 20000, about 95 to about 20000, about 100 to about 20000, about 200 to about 20000, about 300 to about 20000, about 400 to about 20000, about 500 to about 20000, about 600 to about 20000, about 700 to about 20000, about 800 to about 20000, about 900 to about 20000, about 1000 to about 20000, about 2000 to about 20000, about 3000 to about 20000, about 4000 to about 20000, about 5000 to about 20000, about 6000 to about 20000, about 7000 to about 20000, about 8000 to about 20000, about 9000 to about 20000, about 10000 to about 20000, about 11000 to about 20000, about 12000 to about 20000, about 13000 to about 20000, about 14000 to about 20000, about 15000 to about 20000, about 16000 to about 20000, about 17000 to about 20000, about 18000 to about 20000, or about 19000 to about 20000. The number of data points in a training data set can be about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 12, about 14, about 16, about 18, about 20, about 22, about 24, about 26, about 28, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 200, about 300, about 400, about 500, about 600, about 700, about 800, about 900, about 1000, about 1100, about 1200, about 1300, about 1400, about 1500, about 1600, about 1700, about 1800, about 1900, about 2000, about 2200, about 2400, about 2600, about 2800, about 3000, about 3300, about 3600, about 3900, about 4000, about 4400, about 4800, about 5000, about 5500, about 6000, about 6500, about 7000, about 7500, about 8000, about 8500, about 9000, about 9500, about 10000, about 11000, about 12000, about 13000, about 14000, about 15000, about 16000, about 17000, about 18000, about 19000, or about 20000.

[0073] The machine learning model can comprise one or more hyperparameters. The machine learning model can comprise any number of hyperparameters, such as at least 1 hyperparameter, e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 hyperparameters. Hyperparameters include, but are not limited to, learning rate, number of total layers, number of hidden layers, machine learning model batch size, the number of neurons per layer, the number of sensors per layer, and combinations thereof.

[0074]

[0120] The machine learning model can comprise any number of total layers. The total number of layers can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 24, 26, 28, 30, 33, 36, 39, 40, 44, 48, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 220, 240, 260, 280, 300, 330, 360, 390, 400, 440, 480, 500, 550, 30Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703600, 650, 700, 750, 800, 850, 900, 950, or 1000. The total number of layers can be at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 22, at least 24, at least 26, at least 28, at least 30, at least 33, at least 36, at least 39, at least 40, at least 44, at least 48, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 220, at least 240, at least 260, at least 280, at least 300, at least 330, at least 360, at least 390, at least 400, at least 440, at least 480, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, and at least 1000. The machine learning model can be a single layer or an unlayered network.

[0075] The machine learning model can comprise any number of hidden layers. The total number of hidden layers can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 24, 26, 28, 30, 33, 36, 39, 40, 44, 48, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 220, 240, 260, 280, 300, 330, 360, 390, 400, 440, 480, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, or 1000. The total number of hidden layers can be at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 22, at least 24, at least 26, at least 28, at least 30, at least 33, at least 36, at least 39, at least 40, at least 44, at least 48, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 220, at least 240, at least 260, at least 280, at least 300, at least 330, at least 360, at least 390, at least 400, at least 440, at least 480, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, or at least 1000.

[0076] The machine learning model batch size is the number of samples that will be propagated through the network in one machine learning model batch (e.g., the number of training examples in one forward and backward pass). Any machine learning model batch size can be used in the methods described herein. The machine learning model batch size can be at least 1, e.g., at least 31Amazon Privileged and ConfidentialAttorney Docket No.: 146401.0947032, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 14, at least 16, at least 18, at least 20, at least 22, at least 24, at least 26, at least 28, at least 30, at least 33, at least 36, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, or at least 1000. The machine learning model batch size can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000.

[0077] A layer e.g., an input layer, a hidden layer, an output layer) can comprise any number of sensors or neurons. The number of sensors or neurons in a layer can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 24, 26, 28, 30, 33, 36, 39, 40, 44, 48, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 220, 240, 260, 280, 300, 330, 360, 390, 400, 440, 480, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, or 1000.The number of sensors or neurons in a layer can be at least 1, e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 14, at least 16, at least 18, at least 20, at least 22, at least 24, at least 26, at least 28, at least 30, at least 33, at least 36, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, or at least 1000.Immunogenic Composition

[0078] The methods described herein can include a step of selecting one or more true positive SV that results in the expression of a neoantigen {e.g., one or more neoantigens) for inclusion in an immunogenic composition. The methods described herein can include a step of generating an32Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 immunogenic composition (e.g, a cancer vaccine). The methods described herein can include a step of administering an immunogenic composition to a patient (e.g., a patient in need thereof).

[0079] The methods described herein can determine, score, and / or select neoantigens for generating an immunogenic composition (e.g., a cancer vaccine). Any method of determining, scoring, and selecting neoantigens can be used. Examples of suitable methods include those described by WO2022 / 159176A1, US 20230197192A1, US 20230173045A1, and US 20220093209A1, the entire contents of each of which are incorporated herein.

[0080] Neoantigens are self-antigens generated by tumor cells due to genomic mutations or dysregulated RNA splicing. Neoantigens (e.g, neoantigens predicted as true positives by methods described herein) identified by the methods described herein can be in the form of any sequencing results, including, but not limited to, sequence reads (DNA sequence reads, RNA sequence reads), sequence variants (e.g, peptide-modifying sequence variants), encoded peptides, or combinations of the foregoing. The neoantigens can be any type of peptide, including, but not limited to, long peptides, short peptides, or combinations thereof. Neoantigens of the methods described herein can be tumor-specific (e.g, neoantigens that are only present in a cancer or tumor and not in healthy or germ-line cells of a subject).

[0081] As used herein, the terms “short identified neoantigen peptide” and “short peptide” can refer to a peptide with an amino acid length that is between about 3 to about 15, about 3 to about14, about 3 to about 13, about 3 to about 12, about 3 to about 11, about 3 to about 10, about 3 to about 9, about 3 to about 8, about 3 to about 7, about 3 to about 6, about 3 to about 5, about 3 to about 4, about 4 to about 15, about 5 to about 15, about 6 to about 15, about 7 to about 15, about 8 to about 15, about 9 to about 15, about 10 to about 15, about 11 to about 15, about 12 to about15, about 13 to about 15, about 14 to about 15, about 8 to about 11, about 8 to about 10, about 8 to about 9, about 9 to about 11, about 10 to about 11, about 7 to about 11, about 6 to about 11, about 5 to about 11, about 4 to about 11, about 3 to about 11, about 8 to about 12, about 8 to about 13, or about 8 to about 14. “Short identified neoantigen peptide” and “short peptide” can refer to a peptide with an amino acid length that is about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, or about 15. “Short identified neoantigen peptide” and “short peptide” can refer to a peptide with an amino acid length that is at most about 15, at most about 14, at most about 13, at most about 12, at most about 11, at most33Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 about 10, at most about 9, at most about 8, at most about 7, at most about 6, at most about 5, at most about 4, at most about 3, or at most about 2. As used herein, the terms “long identified neoantigen peptide” and “long peptide” can refer to a peptide with an amino acid length that is between about 13 and about 30, about 13 and about 29, about 13 and about 28, about 13 and about 27, about 13 and about 26, about 13 and about 25, about 13 and about 24, about 13 and about 23, about 13 and about 22, about 13 and about 21, about 13 and about 20, about 13 and about 19, about 13 and about 18, about 13 and about 17, about 13 and about 16, about 13 and about 15, about 13 and about 14, about 14 and about 30, about 15 and about 30, about 16 and about 30, about 17 and about 30, about 18 and about 30, about 19 and about 30, about 20 and about 21, about 22 and about 30, about 23 and about 30, about 24 and about 30, about 25 and about 30, about 26 and about 30, about 27 and about 30, about 28 and about 30, about 29 and about 30, or about 13 and about 25. “Long identified neoantigen peptide” and “long peptide” can refer to a peptide with an amino acid length that is about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, or about 30. “Long identified neoantigen peptide” and “long peptide” can refer to a peptide with an amino acid length that is at least about 13, at least about 14, at least about 15, at least about 16, at least about 17, at least about 18, at least about 19, at least about 20, at least about 21, at least about 22, at least about 23, at least about 24, at least about 25, at least about 26, at least about 27, at least about 28, at least about 29, or at least about 30.

[0082] The methods described herein can include steps to determine neoantigens (e.g., neoantigens predicted to be true positives by methods described herein, neoantigens identified from a sequencing result by methods described herein) appropriate for inclusion in an immunogenic composition. The methods can further include a step of determining the human leukocyte antigen (HLA) or major histocompatibility complex (MHC) type for a patient. The HLAtype can be determined through any method, including, but not limited to, nucleic acid sequencing (e.g., RNA sequencing, DNA sequencing, complementary (cDNA) sequencing), immunoassay e.g., enzyme linked immunosorbent assay (ELISA)), mass spectrometry, flow cytometry, real-time PCR, or combinations of any of the foregoing. The methods can include a step of scoring neoantigens of a patient (e.g., neoantigens predicted to be true positives by 34Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 methods described herein, neoantigens identified from a sequencing result by methods described herein). The methods can include a step of selecting one or more neoantigens (e.g., of a biological sample, of a sequencing result) for inclusion in an immunogenic composition. Any number of neoantigens can be selected for inclusion in the immunogenic composition. The number of neoantigens selected can be at most 20, at most 19, at most 18, at most 17, at most 16, at most 15, at most 14, at most 13, at most 12, at most 11, at most 10, at most 9, at most 8, at most 7, at most 6, at most 5, at most 4, at most 3, at most 2, or at most 1. The number of neoantigens selected can be at least 20, at least 19, at least 18, at least 17, at least 16, at least 15, at least 14, at least 13, at least 12, at least 11, at least 10, at least 9, at least 8, at least 7, at least 6, at least 5, at least 4, at least 3, at least 2, or at least 1. The number of neoantigens selected can be between about 1 to about 20, about 2 to about 20, about 3 to about 20, about 4 to about 20, about 5 to about 20, about 6 to about 20, about 7 to about 20, about 8 to about 20, about 9 to about 20, about 10 to about 20, about 11 to about 20, about 12 to about 20, about 13 to about 20, about 14 to about 20, about 15 to about 20, about 16 to about 20, about 17 to about 20, about 18 to about 20, about 19 to about 20, about 1 to about 19, about 1 to about 18, about 1 to about 17, about 1 to about 16, about 1 to about 15, about 1 to about 14, about 1 to about 13, about 1 to about 12, about 1 to about 11, about 1 to about 10, about 1 to about 9, about 1 to about 8, about 1 to about 7, about 1 to about 6, about 1 to about 5, about 1 to about 4, about 1 to about 3, or about 1 to about 2.

[0083] Neoantigens suitable for use in generating an immunogenic composition can be determined by sequencing nucleic acids from a biological sample of a patient. Sequencing in the methods described herein can be any type of sequencing, including, but not limited to, whole genome sequencing, shotgun metagenomic sequencing, whole exome sequencing, nextgeneration sequencing (NGS), cancer personalized profiling by deep sequencing (CAPP-Seq), tagged-amplicon deep sequencing (Tam-Seq), or combinations of any of the foregoing.

[0084] Neoantigen peptides can be predicted by any method. One suitable method for predicting neoantigen peptides is analysis (via an algorithm) of protein variants expressed and isoforms translated from population-level statistics. Another suitable method for predicting neoantigen peptides is by translation of RNA sequencing data (e.g., RNA sequencing data that overlaps each somatic variant). Another suitable method for predicting neoantigen peptides is by transcription 35Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 of DNA sequences (e.g., cfDNA sequences) into RNA and translation of the resulting RNA sequencing data.

[0085] Any strategy or method can be used to reduce the background noise of nucleic acids or proteins originating from healthy non-cancerous cells, and to increase the signal to noise ratio of nucleic acids or proteins from cancerous cells (e.g., sequencing variants that encode for neoantigens). One suitable method for increasing the signal to noise ratio of sequencing variants is by sequencing a biological sample to yield separate non-cancer sequencing results (e.g., sequencing results from non-cancerous cells, sequencing results lacking oncogenic somatic mutations) and cancer sequencing results (e.g., sequencing results from cancer cells, sequencing results containing somatic mutations). In such a method, the non-cancer sequencing results can be subtracted (e.g., eliminated, deleted) from the cancer sequencing results to improve the signal to noise ratio for variant calling of neoantigens. The non-cancer sequencing results and the cancer sequencing results can be obtained from the same biological sample. For example, the cancer sequencing results, and the non-cancer sequencing results can both be obtained from the same liquid biopsy sample (e.g., the same blood sample). Continuing this example, the liquid biopsy sample can be a blood sample and the non-cancer sequencing results are obtained from sequencing the genetic material of white blood cells, while the cancer sequencing results are obtained from sequencing cell free DNA (cfDNA). The non-cancer sequencing results and the cancer sequencing results can be obtained from two different biological samples. For example, the non-cancer sequencing results can be obtained from a tissue biopsy (e.g., a biopsy of healthy, non-cancerous tissue) and the cancer sequencing results can be obtained from tissue resection of a tumor. Another method to increase the signal to noise is to use unique molecular identifiers (UMI) and oversample during sequencing to facilitate error correction. In some embodiments, the method comprises a step of sequencing nucleic acids (e.g., genetic material) of a biological sample to yield non-cancer sequencing results and cancer sequencing results, wherein the non- cancer sequencing results are subtracted from the cancer sequencing results to improve the signal to noise ratio for variant calling of neoantigens.

[0086] Any number of biological samples can be sequenced and analyzed to identify, determine, score, and / or select neoantigens. The number of biological samples can be about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 36Amazon Privileged and ConfidentialAttorney Docket No.: 146401.09470313, about 14, about 15, about 16, about 17, about 18, about 19, or about 20. The number of biological samples can be at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20. The number of biological samples can be between about 1 to about 20, about 1 to about 19, about 1 to about 18, about 1 to about 17, about 1 to about 16, about 1 to about 15, about 1 to about 14, about 1 to about 13, about 1 to about 12, about 1 to about 11, about 1 to about 10, about 1 to about 9, about 1 to about 8, about 1 to about 7, about 1 to about 6, about 1 to about 5, about 1 to about 4, about 1 to about 3, about 1 to about 2, about 2 to about 20, about 3 to about 20, about 4 to about 20, about 5 to about 20, about 6 to about 20, about 7 to about 20, about 8 to about 20, about 9 to about 20, about 10 to about 20, about 11 to about 20, about 12 to about 20, about 13 to about 20, about 14 to about 20, about 15 to about 20, about 16 to about 20, about 17 to about 20, about 18 to about 20, about 19 to about 20, about 2 to about 19, about 2 to about 18, about 2 to about 17, about 2 to about 16, about 2 to about 15, about 2 to about 14, about 2 to about 13, about 2 to about 12, about 2 to about 11, about 2 to about 10, about 2 to about 9, about 2 to about 8, about 2 to about 7, about 2 to about 6, about 2 to about 5, about 2 to about 4, about 2 to about 3, about 3 to about 19, about 3 to about 18, about 3 to about 17, about 3 to about 16, about 3 to about 15, about 3 to about 14, about 3 to about 13, about 3 to about 12, about 3 to about 11, about 3 to about 10, about 3 to about 9, about3 to about 8, about 3 to about 7, about 3 to about 6, about 3 to about 5, about 3 to about 4, about4 to about 19, about 4 to about 18, about 4 to about 17, about 4 to about 16, about 4 to about 15, about 4 to about 14, about 4 to about 13, about 4 to about 12, about 4 to about 11, about 4 to about 10, about 4 to about 9, about 4 to about 8, about 4 to about 7, about 4 to about 6, or about 4 to about 5.

[0087] Sequencing results obtained by the methods described herein (e.g., DNA sequencing result, RNA sequencing result) can be used to determine peptide sequences that are encoded by nucleic acids (e.g., encoded by DNA, encoded by RNA). The nucleic acid that is sequenced to obtain a sequencing result can be any type of nucleic acid, including, but not limited to, RNA, DNA, cfDNA, cfRNA (cell-free RNA), circulating tumor DNA (ctDNA), circulating tumor RNA (ctRNA), or combinations of any of the foregoing. The sequencing result can be in any format, including, but not limited to, CRAM format, General Feature Format (e.g., GFF3), FASTA 37Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 format, FASTQ format, NeXML format, Nexus format, Pileup format, Sequence Alignment Map (SAM) format, Stockholm format, Variant Call Format (VCF) format, genomic VCF (gVCF) format, or a combination of any of the foregoing. In some embodiments, the sequencing result is circulating tumor DNA and the format of the sequencing result is VCF format, wherein the VCF file contains tumor-specific somatic variants found in the biological sample. In some embodiments, the sequencing result is obtained from sequencing healthy and tumor tissue and the format is separate FASTQ files of the sequencing results from healthy and tumor tissue. The resulting peptide sequences can be analyzed to determine whether the encoded peptide is a neoantigen based on the presence of sequence variants (e.g., sequence variants in the peptide amino acid sequence compared to the healthy cells or healthy tissue of a subject, sequence variants in the peptide amino acid sequence compared to a reference genome). Neoantigens (e.g., neoantigen peptides) identified by the methods described herein can be analyzed to determine a predicted immunogenicity based on any factor, including, but not limited to, whether the neoantigen is immunogenic (e.g., whether a neoantigen can elicit an immune response in a subject, prediction of major histocompatibility complex (MHC) binding affinity, whether peptide is predicted to be presented on a cell surface by an MHC molecule), whether the tumor expresses an amount of neoantigen sufficient to elicit an immune response, whether the neoantigen is expressed on a sufficient fraction of the tumor cells, relative or absolute amount of nucleic acids encoding for the neoantigen present in the sample (e.g. , present in the total cfDNA) or a combination thereof. Any methodology for predicting immunogenicity of a neoantigen can be used in the methods described herein. For example, the number of sequenced variants (e.g., DNA sequence variants, RNA sequence variants, encoded peptide variants) can be used to predict the expression of a neoantigen in a biological sample and infer the expression in the originating tissue of the subject (e.g., in a tumor of the subject). For several examples of predicting immunogenicity of a neoantigen, see WO2022 / 159176A1, US 20230197192A1, US 20230173045A1, and US 20220093209A1, all of which are hereby incorporated by reference in their entireties. The predicted immunogenicity can be used to score a neoantigen. Neoantigens can be scored by any parameter of the neoantigen. For example, the neoantigens can be scored by the counts of nucleic acids (e.g., DNA, RNA) encoding the neoantigen (e.g., RNA transcripts per million (RNA TPM), RNA expression values). For example, neoantigens can be scored by 38Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 cellular prevalence of the neoantigen (e.g., cells expressing the neoantigen as determined by single cell sequencing of circulating tumor cells or ctc-FISH staining, cells expressing the neoantigen protein as determined by cell cytometry of circulating tumor cells). Neoantigen peptides can be scored by binding affinity (e.g., experimental binding affinity, predicted binding affinity) to a major histocompatibility complex (MHC), frequency of mutant alleles, gene expression, C-terminal cleavage affinity, transporter associated with antigen processing (TAP) transport of the neoantigen peptide, or combinations of any of the foregoing. The scores of predicted immunogenicity can be ranked (e.g., neoantigens can be ranked in order of predicted immunogenicity). The scored neoantigens can be further selected for inclusion (e.g., inclusion of a neoantigen peptide, inclusion of a nucleic acid encoding for a neoantigen peptide) in an immunogenic composition (e.g., a cancer vaccine). The scored neoantigens can be further selected for exclusion (e.g., exclusion of a neoantigen peptide, exclusion of a nucleic acid encoding for a neoantigen peptide) from an immunogenic composition (e.g., a cancer vaccine).

[0088] Sequence variants (e.g., mutations) identified from biological samples (e.g., liquid biopsy samples, blood samples) by variant calling can be from non-cancer cells or cancer cells (e.g., ctDNA, ctRNA). Any method to distinguish non-cancer cell and cancer cell sequence variants can be used in the methods described herein. For example, multiple sequence variants can be screened simultaneously to improve the probability of detecting circulating tumor nucleic acids (e.g., ctDNA, ctRNA). As another example, the probability of detecting circulating tumor nucleic acids can be increased based on detection of epigenetic modifications (e.g., methylation profile, including 5-methylcytosine (5mC) and / or 5-hydroxymethylcytosine (5hmC)) of the nucleic acids (e.g., cfDNA). As another example, the probability of detecting circulating tumor nucleic acids can be increased based on the fragmentation pattern (e.g., fragment lengths, fragment start positions, fragment end nucleotide motifs) of nucleic acids (e.g., cfDNA). As another example, the probability of detecting circulating tumor nucleic acids can be increased based on the mutational signature profiles of known human carcinogens, existence of co-mutations, and / or evolutionary cancer signatures.

[0089] Methods described herein can include a step of generating an immunogenic composition based on the neoantigen identified by methods described herein (e.g., neoantigens predicted as true positives). The immunogenic composition (e.g., cancer vaccine) can be any type of 39Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 immunogenic composition, such as those described in WO2022 / 170067 Al, US2023 / 0173045 Al, WO2022 / 159176 Al, WO2022 / 251034 Al, or US2023 / 0173046 Al, the entirety of each of which are incorporated by reference herein. The immunogenic composition can comprise or encode neoantigens (e.g., one or more neoantigens) scored for predicted immunogenicity by any method including by the methods described herein. The immunogenic composition can contain neoantigens in any form, including, but not limited to, peptides (e.g., long identified peptides, short identified peptides, synthesized peptides, isolated peptides), RNA sequences encoding for peptides (e.g., mRNA, mRNA containing unnatural nucleotides (e.g., pseudouridine, Nl- methylpseudouridine, 7-methylguanosine, N6-methyladenosine, 2’-O-methyl nucleotide), mRNA containing inverted nucleotides, or combinations thereof), DNA sequences encoding for peptides (e.g., plasmid DNA, viral DNA), or combinations thereof. Immunogenic compositions containing nucleic acids (e.g., RNA, DNA) encoding for peptides can be in the form of a virus (e.g., an adenovirus, lentivirus, fowl pox, vaccinia, self-replicating alphavirus, Maraba virus) or free nucleic acids in a drug delivery particle (e.g., a lipid nano particle (LNP)).

[0090] A nucleic acid e.g., RNA, DNA) of an immunogenic composition can encode for any number of peptide antigens. The number of peptide antigens encoded by a nucleic acid can be about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20. The number of peptide antigens encoded by a nucleic acid can be at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 11, at most about 12, at most about 13, at most about 14, at most about 15, at most about 16, at most about 17, at most about 18, at most about 19, or at most about 20. In some embodiments, the method includes a step of generating an immunogenic composition, wherein the immunogenic composition comprises or encodes for one or more neoantigens scored for the predicted immunogenicity by the method. The immunogenic composition can comprise neoantigens predicted as true positives by methods described herein. The immunogenic composition can comprise neoantigens identified from the sequencing results in the methods described herein.

[0091] An immunogenic composition can contain or encode for any number of neoantigens (e.g., neoantigen proteins, DNA sequences encoding for neoantigen proteins, RNA sequences encoding 40Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 for neoantigen proteins). An immunogenic composition can contain or encode for a number of neoantigens that is about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, or about 30. The number of neoantigens contained or encoded by an immunogenic composition can be one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, 20 or more, 21 or more, 22 or more, 23 or more, 24 or more, 25 or more, 26 or more, 27 or more, 28 or more, or 29 or more, or 30 or more. The number of neoantigens contained or encoded by an immunogenic composition can be at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 16, at least about 17, at least about 18, at least about 19, at least about 20, at least about 21, at least about 22, at least about 23, at least about 24, at least about 25, at least about 26, at least about 27, at least about 28, at least about 29, or at least about 30. The number of neoantigens contained or encoded by an immunogenic composition can be between about 1 to about 50, about 2 to about 50, about 3 to about 50, about 4 to about 50, about 5 to about 50, about 6 to about 50, about 7 to about 50, about 8 to about 50, about 9 to about 50, about 10 to about 50, about 11 to about 50, about 12 to about 50, about 13 to about 50, about 14 to about 50, about 15 to about 50, about 16 to about 50, about 17 to about 50, about 18 to about 50, about 19 to about 50, about 20 to about 50, about 22 to about 50, about 24 to about 50, about 26 to about 50, about 28 to about 50, about 30 to about 50, about 33 to about 50, about 36 to about 50, about 40 to about 50, about 45 to about 50, about 1 to about 45, about 1 to about 40, about 1 to about 36, about 1 to about 33, about 1 to about 30, about 1 to about 28, about 1 to about 26, about 1 to about 24, about 1 to about 22, about 1 to about 20, about 1 to about 18, about 1 to about 16, about 1 to about 14, about 1 to about 12, about 1 to about 10, about 1 to about 9, about 1 to about 8, about 1 to about 7, about 1 to about 6, about 1 to about 5, about 1 to about 4, about 1 to about 3, about 1 to about 2, about 2 to about 14, about 2 to about 12, about 2 to about 10, about 2 to about 8, about 2 to about 6, about 2 to about 4, about 3 to about 14, about 3 to about 12, about 341Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 to about 10, about 3 to about 8, about 3 to about 6, about 3 to about 4, about 4 to about 14, about 4 to about 12, about 4 to about 10, about 4 to about 8, about 4 to about 6, about 5 to about 14, about 5 to about 12, about 5 to about 10, about 5 to about 8, about 5 to about 6, about 6 to about 14, about 6 to about 12, about 6 to about 10, about 6 to about 8, about 7 to about 14, about 7 to about 12, about 7 to about 10, about 7 to about 8, about 8 to about 14, about 8 to about 12, about 8 to about 10, about 9 to about 14, about 9 to about 12, about 9 to about 10, about 10 to about 14, about 10 to about 12, about 11 to about 14, about 11 to about 12, about 12 to about 14, or about 13 to about 14.

[0092] Immunogenic compositions described herein may comprise up to about 50 neoantigen long peptides and / or short peptides. The immunogenic composition may comprise about 10 to about 20 neoantigen long peptides and / or short peptides. In some embodiments, the immunogenic composition comprises about 19 neoantigen long peptides and / or short peptides.

[0093] The immunogenic composition may comprise at least about 2 or more neoantigen long peptides. The immunogenic composition may comprise about 2 to about 18 neoantigen long peptides. The immunogenic composition can comprise at least about 10 to about 15 neoantigen long peptides. The immunogenic composition may comprise at least about 2 or more neoantigen short peptides. The immunogenic composition may comprise at least about 2 to about 10 neoantigen short peptides.

[0094] The methods described herein can include a step of administering an immunogenic composition to a patient in need thereof. The selected neoantigens can be divided into pools for separate administration of immunogenic compositions to a patient in need thereof. The pools of neoantigens can contain any number of neoantigens. The number of neoantigens in a pool can be about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, or about 50. The number of neoantigens in a pool can be at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 42Amazon Privileged and ConfidentialAttorney Docket No.: 146401.09470323, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31 , at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, or at least 50. The number of neoantigens in a pool can be at most 2, at most 3, at most 4, at most 5, at most 6, at most 7, at most 8, at most 9, at most 10, at most 11, at most 12, at most 13, at most 14, at most 15, at most 16, at most 17, at most 18, at most 19, at most 20, at most 21, at most 22, at most 23, at most 24, at most 25, at most 26, at most 27, at most 28, at most 29, at most 30, at most 31, at most 32, at most 33, at most 34, at most 35, at most 36, at most 37, at most 38, at most 39, at most 40, at most 41, at most 42, at most 43, at most 44, at most 45, at most 46, at most 47, at most 48, at most 49, or at most 50. The number of neoantigens in a pool can be about 2 to about 30, 2 to about 28, 2 to about 26, 1 to about 24, 2 to about 22, 2 to about 20, about 2 to about 18, about 2 to about 16, about 2 to about 14, about 2 to about 12, about 2 to about 10, about 2 to about 9, about 2 to about 8, about 2 to about 7, about 2 to about 6, about 2 to about 5, about 2 to about 4, or about 2 to about 3. As an example, three peptide pools of the immunogenic composition may comprise about 5 neoantigen long peptides and / or short peptides and one peptide pool may comprise 4 neoantigen long peptides and / or short peptides and a helper peptide. Each peptide pool may comprise different neoantigen long peptides and / or short peptides. Neoantigen long peptides may be about 15 to about 30 amino acids in length. Neoantigen short peptides may be about 5 to about 15 amino acids in length.

[0095] The selected neoantigens for the immunogenic compositions can be split into any number of pools. The number of pools can be about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20. The number of pools can be at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20. The number of pools can be at most 1, at most 2, at most 3, at most 4, at most 5, at most 6, at most 7, at most 8, at most 9, at most 10, at most 11, at most 12, at most 13, at most 14, at most 15, at most 16, at most 17, at most 18, at most 19, or at most 20. The number of pools can be between about 1 to about 2, about 1 to about 3, about 1 to about 4, about 1 to about 5, about 1 to about 6, about 1 to about 7, about 1 to about 8, about 1 to about 9, about 1 to about 10, about 2 to about43Amazon Privileged and ConfidentialAttorney Docket No.: 146401.09470310, about 3 to about 10, about 4 to about 10, about 5 to about 10, about 6 to about 10, about 7 to about 10, about 8 to about 10, about 9 to about 10, about 2 to about 3, about 2 to about 4, about 2 to about 5, about 2 to about 6, about 2 to about 7, about 2 to about 8, about 2 to about 9, about 2 to about 10, about 3 to about 4, about 3 to about 5, about 3 to about 6, about 3 to about 7, about 3 to about 8, about 3 to about 9, about 3 to about 10, about 4 to about 5, about 4 to about 6, about 4 to about 7, about 4 to about 8, about 4 to about 9, or about 4 to about 10.

[0096] The immunogenic compositions can comprise at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 16, at least about 17, at least about 18, at least about 19, at least about 20, at least about 21, at least about 22, at least about 23, at least about 24, at least about 25, at least about 26, at least about 27, at least about 28, at least about 29, at least about 30, at least about 31, at least about 32, at least about 33, at least about 34, at least about 35, at least about 36, at least about 37, at least about 38, at least about 39, at least about 40, at least about 41, at least about 42, at least about 43, at least about 44, at least about 45, at least about 46, at least about 47, at least about 48, at least about 49, at least about 50 or more neoantigen peptides (e.g., neoantigen long peptides and / or short peptides). The immunogenic compositions can comprise up to about 100 neoantigen peptides. The immunogenic compositions can contain about 1-5 neoantigens, 1-10 neoantigens, about 1-15 neoantigens, about 4-10 neoantigens, about 4-15 neoantigens, about 10-20 neoantigens, about 10-30 neoantigens, about 10-40 neoantigens, about 10-50 neoantigens, about 10-60 neoantigens, about 10-70 neoantigens, about 10-80 neoantigens, about 10-90 neoantigens, or about 10-100 neoantigens. For example, the immunogenic compositions can comprise about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 neoantigens. In some embodiments, the immunogenic compositions can comprise about 19 neoantigens. Each of the neoantigens in the immunogenic compositions can be different.

[0097] The immunogenic compositions described herein can further comprise an adjuvant. Adjuvants are any substance whose admixture into an immunogenic composition increase, or otherwise enhances and / or boosts, the immune response to a tumor-specific neoantigen, but when the substance is administered alone does not generate an immune response to a tumor- 44Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 specific neoantigen. The adjuvant can generate an immune response to the neoantigen and does not produce an allergy or other adverse reaction. It is contemplated herein that adjuvant can be administered before, together, concomitantly with, or after administration of the immunogenic composition.

[0098] Adjuvants can enhance an immune response by several mechanisms including, e.g., lymphocyte recruitment, stimulation of B and / or T cells, and stimulation of macrophages. When an immunogenic composition described herein comprises adjuvants or is administered together with one or more adjuvants, the adjuvants that can be used include, but are not limited to, mineral salt adjuvants or mineral salt gel adjuvants, particulate adjuvants, microparticulate adjuvants, mucosal adjuvants, immunostimulatory adjuvants, or combinations thereof. Examples of adjuvants include, but are not limited to, aluminum salts (alum) (such as aluminum hydroxide, aluminum phosphate, and aluminum sulfate), 3 De-O-acylated monophosphoryl lipid A (MPL) (see, GB 2220211), MF59 (Novartis), AS03 (Glaxo SmithKline), AS04 (Glaxo SmithKline), polysorbate 80 (Tween 80; ICL Americas, Inc.), imidazopyridine compounds (see, International Application No. PCT / US2007 / 064857, published as International Publication No.W02007 / 109812, published as U.S. Pat. App. publication No. US20090232844A1 and corresponding to U.S. Pat. No. 8,063,063), imidazoquinoxaline compounds (see, International Application No. PCT / US2007 / 064858, published as International Publication No.W02007 / 109813, published as U.S. Pat. App. publication No. US20090311288A1 and corresponding to U.S. Pat. No. 8,173,657) and saponins, such as QS21 (see, Kensil et al, in Vaccine Design: The Subunit and Adjuvant Approach (eds. Powell & Newman, Plenum Press, NY, 1995); U.S. Pat. No. 5,057,540). In some embodiments, the adjuvant is Freund's adjuvant (complete or incomplete). Other adjuvants are oil in water emulsions (such as squalene or peanut oil), optionally in combination with immune stimulants, such as monophosphoryl lipid A (see, Stoute et al, N. Engl. J. Med. 336, 86-91 (1997)). CpG immunostimulatory oligonucleotides have also been reported to enhance the effects of adjuvants in a vaccine setting. Other TLR binding molecules such as RNA binding TLR 7, TLR 8 and / or TLR 9 may also be used.

[0099] Other examples of useful adjuvants include, but are not limited to, chemically modified CpGs (e. ., CpR, Idera), Poly(I:C)(e.g., polyi:CI2U), poly ICLC, non-CpG bacterial DNA or RNA as well as immunoactive small molecules and antibodies such as cyclophosphamide, 45Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 sunitmib, bevacizumab, Celebrex (celecoxib), NCX-4016, sildenafil, tadalafil, vardenafil, sorafenib, XL-999, CP-547632, pazopanib, ZD2171, AZD2171, ipilimumab, tremelimumab, and SC58175, which may act therapeutically and / or as an adjuvant.

[0100] An immunogenic composition described herein can be administered to a subject that has not been diagnosed with cancer, has been diagnosed with cancer, is already suffering from cancer, has recurrent cancer (e.g., relapse), or is at risk of developing cancer. An immunogenic composition can be administered to a subject that is resistant to other forms of cancer treatment (e. ., chemotherapy, immunotherapy, or radiation). Immunogenic composition can be administered to the subject prior to other standard of care cancer therapies (e.g., chemotherapy, immunotherapy, or radiation). An immunogenic composition can be administered to the subject concurrently, after, or in combination to other standard of care cancer therapies (e.g, chemotherapy, immunotherapy, or radiation). In some embodiments, the method includes administering an immunogenic composition to a patient in need thereof.

[0101] Immunogenic compositions described herein can be administered to a subject in an amount sufficient to elicit an immune response to the tumor-specific neoantigen and to destroy, or at least partially arrest, symptoms and / or complications. In embodiments, the immunogenic composition can provide a long-lasting immune response. A long-lasting immune response can be established by administering a boosting dose of the immunogenic composition to the subject. The immune response to the immunogenic composition can be extended by administering to the subject a boosting dose. In embodiments, at least one, at least two, at least three or more boosting doses can be administered to abate the cancer. A first boosting dose may increase the immune response by at least 50%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 1000%. A second boosting dose may increase the immune response by at least 50%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 1000%. A third boosting dose may increase the immune response by at least 50%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 1000%.

[0102] An amount adequate to elicit an immune response is defined as a “therapeutically effective dose.” Amounts effective for this use will depend on, e.g., the composition, the manner of administration, the stage and severity of the disease being treated, the weight and general state of health of the individual, and the judgment of the prescribing physician. It should be kept in 46Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 mind that an immunogenic composition can generally be employed in serious disease states, that is, life-threatening or potentially life-threatening situations, especially when the cancer has metastasized. In such cases, in view of the minimization of extraneous substances and the relative nontoxic nature of a neoantigen, it is possible and can be considered desirable by the treating physician to administer substantial excesses of immunogenic compositions (e.g., cancer vaccines).Definitions

[0103] All publications and patents cited in this disclosure are incorporated by reference in their entirety. To the extent the material incorporated by reference contradicts or is inconsistent with this specification, the specification will supersede any such material. The citation of any references herein is not an admission that such references are prior art to the present disclosure. When a range of values is expressed, it includes embodiments using any particular value within the range. Further, reference to values stated in ranges includes each and every value within that range. All ranges are inclusive of their endpoints and combinable. When values are expressed as approximations, by use of the antecedent “about,” it will be understood that the particular value forms another embodiment. Reference to a particular numerical value includes at least that particular value, unless the context clearly dictates otherwise. The use of “or” will mean “and / or” unless the specific context of its use dictates otherwise.

[0104] Various terms relating to aspects of the description are used throughout the specification and claims. Such terms are to be given their ordinary meaning in the art unless otherwise indicated. Other specifically defined terms are to be construed in a manner consistent with the definitions provided herein. The techniques and procedures described or referenced herein are generally well understood and commonly employed using conventional methodologies by those skilled in the art, such as, for example, the widely utilized molecular cloning methodologies described in Sambrook et al., Molecular Cloning: A Laboratory Manual 4th ed. (2012) Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY. As appropriate, procedures involving the use of commercially available kits and reagents are generally carried out in accordance with manufacturer-defined protocols and conditions unless otherwise noted.Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703

[0105] As used herein, the singular forms “a,” “an,” and “the” include plural forms unless the context clearly indicates otherwise. The terms “include,” “such as,” and the like are intended to convey inclusion without limitation, unless otherwise specifically indicated.

[0106] Unless otherwise indicated, the terms “at least,” “less than,” and “about,” or similar terms preceding a series of elements or a range are to be understood to refer to every element in the series or range. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described herein. Such equivalents are intended to be encompassed by the following claims.

[0107] As used herein, the term “about” is used to refer to an amount that is approximately, nearly, almost, or in the vicinity of being equal to or is equal to a stated amount, e.g., the stated amount plus / minus about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, about 2% or about 1%.

[0108] The terms “patient” or “individual” (e.g., individual of a cohort of interest) as used herein refer to any animal, such as any mammal, including, but not limited to, humans, non-human primates, rodents, mammals commonly kept as pets (e.g., dogs and cats, among others), livestock (e.g., cattle, sheep, goats, pigs, horses, and camels, among others) and the like. In some embodiments, the mammal is a mouse. In some embodiments, the mammal is a human.

[0109] The term “true positive” (e.g., true positive SV) as used herein refers to an identified ctDNA-derived SV or ctRNA-derived SV that is present on tumor tissue.

[0110] The term “false positive” (e.g., false positive SV) as used herein refers to an identified ctDNA-derived SV or ctRNA-derived SV that is not present on tumor tissue.

[0111] The term “false negative” (e.g., false negative SV) as used herein refers to an identified tumor tissue-derived SV that is not present in ctDNA or ctRNA samples.EQUIVALENTS

[0112] It will be readily apparent to those skilled in the art that other suitable modifications and adaptions of the methods of the invention described herein are obvious and may be made using suitable equivalents without departing from the scope of the disclosure or the embodiments. Having now described certain compositions and methods in detail, the same will be more clearlyAmazon Privileged and ConfidentialAtorney Docket No.: 146401.094703 understood by reference to the following example, which is introduced for illustration only and not intended to be limiting.EXAMPLESExample 1- Cohort-Level Somatic Structural Variant Discovery

[0113] Paired tissue and blood samples were collected from patients in two separate cohorts, Cohort- 1 and Cohort-2. Both DNA and RNA were extracted from FFPE (Cohort- 1) or fresh frozen (Cohort-2) tumor tissues. Cell-free tumor DNA (cfDNA) was extracted from the plasma portion of the blood samples, while DNA from buffy coat (mostly white blood cells) was extracted as normal controls. All extracted DNA from tumor tissue, plasma and white blood cells were captured with WES probes to limit to mostly protein-coding regions, and sequenced at a depth of at least 125x with paired-end Illumina. RNA was sequenced with strand-specific protocols following reverse transcription. To discover somatic structural variants, the sequencing results were run through DRAGEN WES tumor-normal structural variant caller and a custom RNA confirmation algorithm loosely based on isovar. The sequencing results were input through both the PASS and NON-PASS SV filters. Additionally, DRAGEN WGS structural variant caller was run on Cohort-1. DRAGEN RNA gene fusion caller can further be used to discover additional targets. Concordance metrics between coding plasma and tissue SVs demonstrated variance in sensitivity and PPV between conditions (Table 1 and FIG.l)Table 1. Plasma / tissue concordanceCohort- 1

[0114] Cohort- 1 had three individual patient groups separated by level of tumor mutational burden (TMB) (i.e., high TMB, medium TMB, and low TMB) with 16 patients total. The median49Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 number of SVs were called in each of the TMB groups. It was observed that more SVs were called in tissue and plasma for the medium and low TMB groups compared to the high TMB groups (Table 2). When sorting Cohort- 1 by indication, the number of structural variants called varied, demonstrating the impact type of cancer has on presence of structural variants (FIG. 2).Table 2. Structural variants called in Cohort- 1

[0115] Calculation of sensitivity and positive predictive value (PPV) showed low TMB groups had no observable actionable variants (true positive plasma SVs) (Table 3).Table 3. Calculation of sensitivity and PPV

[0116] Further, in the high TMB group, SVs discovered in DNA and confirmed in RNA increased the number of actionable variants by 12%, while in medium and low TMB groups the number of actional variants increased by an average of 47%.Cohort-2

[0117] The Cohort-2 contains patients with melanoma and breast cancer. It was observed that distribution of SV classes, specifically , translocations (TRA), large deletions (DEL), largeAmazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 insertions (INS), and duplications (DUP) among paired tissue and plasma varied (Table 4). The vast majority of SVs called in tissue DNA are deletions, while SVs observed in plasma are relatively equally distributed among duplications, deletions, and insertions. Calculation of sensitivity and PPV between variant class revealed translocations have a high percentage of true positive plasma-derived SVs resulting in an average sensitivity of 20.0% and a PPV of 62.5% (Table 5). Genomic origin distribution of SVs also varied between tissue and plasma samples. Within tissue samples SVs were largely observed in kinase genes, while plasma-derived SVs were observed among tumor suppressor genes, oncogene and kinase genes (Table 6). Calculation of sensitivity and PPV between gene type revealed SVs located in kinase gene regions have the highest PPV, while SVs located in phosphatase and oncogenes have the highest sensitivity (Table 7).Table 4. Distribution of structural variants by variant classTable 5. Calculations of sensitivity and PPV by variant class51Amazon Privileged and ConfidentialAtorney Docket No.: 146401.094703Table 6. Distribution of structural variants by gene typeTable 7. Calculations of sensitivity and PPV by gene typeExample 2. Training a machine learning model to predict true positive ctDNA-derived SVs52Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703

[0118] Paired tumor tissue and blood biopsies will be collected from a patient suspected to have or having cancer. DNA will be extracted from the samples and libraries will be prepared for sequencing. The DNA samples will be run on Illumina, receiving an output of sequencing data files. Data files will be prepared and processed using a variant caller (e.g. DRAGEN, Mutect2) to call structural variants. Output data may use a PASS or Non-PASS filter to either filter out or keep potential artifact SVs, respectively. DNA may further be validated using a custom RNA confirmation algorithm loosely based on isovar. Once datasets for the paired tumor tissue- derived SVs and ctDNA-derived SVs isolated from the blood biopsy are prepared, overlapping SVs will be identified. These data will then be used as training set data. Training set data will evaluate various characteristics of SVs identified in true positive ctDNA-derived SVs (i.e., SVs present on tumor tissue) and be used to calculate sensitivity and positive predictive values. This will be repeated until sensitivity and PPV increase to confidently predict a true positive ctDNA- derived SV.Example 3. Prediction of a ctDNA-derived SV being present on a tumor

[0119] A liquid biopsy will be collected from a patient suspected to have or having cancer. DNA will be extracted and library samples will be prepared. The prepared DNA will be run on a sequencing capable of outputting sequencing data (e.g. Illumina). The sequencing data will then be run through a variant caller capable of identifying SV in the isolated genetic material. Each identified ctDNA-derived SV will be passed through the machine learning model to receive an output of the probability that the ctDNA-derived SV is present on the tumor.Amazon Privileged and Confidential

Claims

Attorney Docket No.: 146401.094703CLAIMSWHAT IS CLAIMED IS:

1. A method of predicting a probability score that a circulating tumor DNA (ctDNA)- derived structural variant (SV) is present in tumor tissue, comprising the steps of:(a) Isolating genetic material from a liquid biopsy, wherein the genetic material is ctDNA;(b) sequencing the ctDNA to obtain a sequencing result;(C) inputting the sequencing result into a structural variant caller to identify ctDNA- derived SVs;(d) inputting a ctDNA-derived SV into a machine learning model; and(e) receiving an output from the machine learning model of the probability score that the ctDNA-derived SV is a true positive, wherein the true positive is a ctDNA-derived SV present in the tumor tissue.

2. A method of training a machine learning model to output a probability score that a structural variant (SV) in circulating tumor DNA (ctDNA) is present in tumor tissue, comprising the steps of:(a) sequencing genetic material from a paired tumor tissue sample and liquid biopsy sample collected from a patient, wherein the genetic material from the liquid biopsy is ctDNA, to obtain a tumor tissue sequencing result and a ctDNA sequencing result;(b) inputting the tumor tissue sequencing result and the ctDNA sequencing result into a structural variant caller to obtain an output of tumor tissue-derived SVs and ctDNA-derived SVs;(c) identifying true positive SVs, false positive SVs and false negative SVs from the output of tumor tissue-derived SVs and ctDNA-derived SVs;(d) inputting training set data into the machine learning model, wherein the training set data comprises the true positive SVs, false positive SVs, false negative SVs, a calculated sensitivity value, a calculated positive predictive value and / or identified predictors of true54Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 positive SVs based on the calculated sensitivity value and the calculated positive predictive value;(e) repeating steps (a)-(d) until a favorable sensitivity value and / or a favorable positive predictive value is achieved.

3. The method of claim 1 or 2, wherein the liquid biopsy is collected from an amniotic fluid sample, ascitic fluid sample, bile sample, blood sample, buccal sample, cerebral spinal fluid sample, fecal sample, hair sample, peritoneal fluid sample, plasma sample, pleural effusion sample, saliva sample, semen sample, serum sample, skin sample, synovial fluid sample, urine sample, or a combination thereof.

4. The method of any one of claims 1-3, wherein the genetic material comprises a nucleic acid selected from DNA, RNA, or a combination thereof.

5. The method of any one of claims 1-4, wherein the genetic material comprises a nucleic acid selected from cell free DNA (cfDNA), circulating tumor DNA (ctDNA), cell free RNA (cfRNA), circulating tumor RNA (ctRNA), or a combination of any of the foregoing.

6. The method of any one of claims 1-5, wherein the sequencing result comprises DNA sequencing results, RNA sequencing results, or a combination thereof.

7. The method of any one of claims 1-6, wherein the sequencing results can be obtained using either short-read sequencing or long-read sequencing.

8. The method of any one of claims 1-7, wherein the structural variant caller is DRAGEN WES, DRAGEN RNA, DRAGEN WGS or similar.

9. The method of any one of claims 2-8, wherein the tumor tissue-derived SVs comprise a deletion, duplication, inversion, translocation, gene fusion, or any combination thereof.55Amazon Privileged and ConfidentialAttorney Docket No.: 146401.09470310. The method of any one of claims 1-8, wherein the ctDNA-derived SVs comprise a deletion, duplication, inversion, translocation, gene fusion, or any combination thereof.

11. The method of any one of claims 2-10, wherein the true positive SVs are ctDNA-derived SVs present in the tumor tissue.

12. The method of any one of claims 2-10, wherein the false negative SVs are tumor tissue- derived SVs not present in the ctDNA.

13. The method of any one of claims 2-10, wherein the false positive SVs are ctDNA-derived SVs not present in the tumor tissue.

14. The method of any one of claims 2-13, wherein the sensitivity value is calculated using formula (I): 100wherein:(i) a true positive is a number of ctDNA-derived SVs present in tumor tissue; and(ii) a false negative is a number of tumor tissue-derived SVs present in the tumor tissue sample but not present in the ctDNA.

15. The method of any one of claims 2-13, wherein the positive predictive value is calculated using formula (II): 100wherein:(i) a true positive is a number of ctDNA-derived SVs present in tumor tissue; and(ii) a false positive is a number of ctDNA-derived SVs present in the ctDNA but not present in the tumor tissue.Amazon Privileged and ConfidentialAttorney Docket No.: 146401.09470316. The method of any one of claims 2-15, wherein the favorable sensitivity value is between about 50% and about 100%.

17. The method of any one of claims 2-16, wherein the favorable positive predictive value is between about 50% and about 100%.

18. The method of any one of claims 2-17, wherein the predictors comprise recurrent true positive SVs within a gene type and / or a SV class.

19. The method of claim 18, wherein the gene type comprises a kinase, a tumor suppressor gene, a phosphatase, an oncogene, an essential gene, a spliceosome, or a protease.

20. The method of claim 18, wherein the SV class is a deletion, translocation, insertion, duplication, or a gene fusion.

21. The method of any one of claims 2-20, wherein the machine learning model comprises a classification algorithm.

22. The method of claim 21, wherein the classification algorithm is selected from logistic regression, naive Bayes, K-nearest neighbor (KNN), decision tree, support vector machine, K- means clustering, random forest, artificial neural network (ANN), or a combination of any of the foregoing.

23. The method of any one of claims 2-22 further comprising the step of:(i) validating the machine learning model by inputting a known true positive SV into the machine learning model; and(ii) receiving an output of the probability score indicating the ctDNA-derived SV is a true positive SV.

24. The method of any one of claims 1-23, further comprising a step of:57Amazon Privileged and ConfidentialAttorney Docket No.: 146401.094703 identifying one or more true positive SV, wherein the true positive SV results in a neoantigen.

25. The method of claim 24, further comprising a step of: selecting the neoantigen for inclusion in an immunogenic composition.

26. The method of claim 24 or 25, further comprising a step of: generating the immunogenic composition.

27. The method of any one of claims 24-26, further comprising a step of: administering the immunogenic composition to a patient in need thereof.Amazon Privileged and Confidential

Citation Information

Patent Citations

  • Modified lipopolysaccharides

    GB2220211A

  • Immunopotentiating compounds

    US20090232844A1

  • Imidazoquinoxaline compounds as immunomodulators

    US20090311288A1

  • Predicting immunogenicity of t cell epitopes

    US20220093209A1

  • Ranking neoantigens for personalized cancer vaccine

    US20230173045A1