Rule-based methods for prioritizing tumor neoantigens, viral antigens, and tumor-associated antigens or high-confidence tumor-specific antigens
A rule-based method prioritizes tumor neoantigens and associated antigens using binding affinity and gene expression criteria, enhancing the identification and effectiveness of cancer immunotherapies by focusing on high-confidence tumor-specific antigens.
Patent Information
- Application Number
- PCT/US2025/041077
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-12
- Filing Date
- 2025-08-07
- Publication Date
- 2026-02-19
AI Technical Summary
Current clinical pipelines for identifying tumor neoantigens, viral antigens, and tumor-associated antigens are laborious and restricted, failing to support immunopeptidomics and are limited to predicted neoantigens, while immune editing due to immune pressure can hinder effective immune recognition.
A rule-based method for prioritizing tumor neoantigens, viral antigens, and tumor-associated antigens by selecting amino acid sequences based on criteria such as binding affinity, likelihood of presentation, gene expression, and mutation prevalence, followed by mass spectrometry-based data acquisition and analysis.
Enhances the identification and prioritization of clinically relevant antigens, improving the effectiveness of cancer immunotherapies by focusing on high-confidence tumor-specific antigens capable of inducing robust immune responses.
Smart Images

Figure US2025041077_19022026_PF_FP_ABST
Abstract
Description
[0001] Docket No.: 084276.00417 / LUD 6239
[0002] RULE-BASED METHODS EOR PRIORITIZING TUMOR NEOANTIGENS, VIRAL ANTIGENS, AND TUMOR-ASSOCIATED ANTIGENS OR HIGH-CONFIDENCE TUMOR-SPECIFIC ANTIGENS
[0003] REFERENCE TO SEQUENCE LISTING
[0004] The present application is being filed along with a Sequence Listing in electronic format. The Sequence Listing is provided as a file entitled SeqListing_084276.00417.xml, created on August 5, 2025, which is 113,659 bytes in size. The information in the electronic format of the sequence listing is incorporated herein by reference in its entirety.
[0005] CROSS-REFERENCE TO RELATED APPLICATIONS
[0006] This application claims priority under 35 U.S.C. §119(e) to U.S. Provisional Patent Application No. 63 / 681,995 filed August 12, 2024. The foregoing application is incorporated by reference herein in its entirety.
[0007] FIELD OF THE INVENTION
[0008] This invention relates to methods for rule-based ranking for prioritizing tumor neoantigens, viral antigens, and tumor-associated antigens (TAAs) or high-confidence tumor-specific antigens (HC-TSAs).
[0009] BACKGROUND OF THE INVENTION
[0010] Personalized antigen discovery, enabling the identification of peptides presented on a patient’s cancer cells by class-I and -II human leukocyte antigen (HLA-I and HLA-II) and recognized by autologous T cells, is crucial for the development of cancer vaccines, adoptive transfer of T cells, and various T-cell targeting molecules. Whole genome (WGS) or whole exome (WES) and RNA sequencing (RNAseq) based methods for mutated antigen (neoantigens) predictions are commonly used for translational research and in clinical trials. Recently, the application of mass spectrometry (MS) to identify HLA-bound peptides, in combination with proteogenomics, facilitated the exploration of novel targets from a variety of antigens naturally processed and presented in cancer, including neoantigens, tumor-associated and tumor-specific antigens (TAAs and TSAs), oncoviruses, and peptides translated from putative non-protein-coding Docket No.: 084276.00417 / LUD 6239 transcripts. However, their identification is laborious, and current clinical pipelines do not support immunopeptidomics and are restricted to predicted neoantigens.
[0011] Immunotherapies are remarkably effective against some tumor indications. However, robust immune pressure may lead to immune editing, resulting in the selection of tumor cells displaying diminished antigenicity. Mutations, like somatic single nucleotide variants (SNVs), insertions, deletions, and copy number variations (CNVs), can disable parts of the antigen processing and presentation machinery (APPM). These alterations affect components like P2- microglobulin (B2M), transporters associated with presentation (TAPI and TAP2), and the HLA locus, hindering immune recognition. It is thus essential to gain a comprehensive understanding of the heterogenous antigenic landscape and of the tumor’s capacity to present antigens.
[0012] SUMMARY OF THE INVENTION
[0013] This disclosure addresses the need mentioned above in a number of aspects. In one aspect, this disclosure provides a method of ranking neoantigens capable of inducing a tumor-specific immune response. In some embodiments, the method comprises: (a) selecting, from candidate amino acid sequences containing one or more mutations, a first set of amino acid sequences, each of which is free of a false-positive mutation; (b) identifying, from the first set of amino acid sequences, a second set of amino acid sequences, each of which contains a tumor-specific mutant amino acid sequence; (c) identifying, from the second set of amino acid sequences, a third set of amino acid sequences, each of which contains a mutation within a binding core region interacting with a binding groove of a human leukocyte antigen (HLA) molecule; (d) selecting, from the third set of amino acid sequences, a fourth set of amino acid sequences, each of which exhibits high binding affinity to HLA-I and / or HLA-II, and identifying remaining unselected third set of amino acid sequences as a fifth set of amino acid sequences; (e) selecting, from the fourth set of amino acid sequences, a sixth set of amino acid sequences, each of which has a high likelihood of presentation, and identifying remaining unselected fourth set of amino acid sequences as a seventh set of amino acid sequences; (f) selecting, from the sixth set of amino acid sequences, an eighth set of amino acid sequences, each of which is from a gene having a transcripts-per-million expression value greater than 15, and adding remaining unselected sixth set of amino acid sequences as a seventh set of amino acid sequences; (g) selecting, from the eighth set of amino acid sequences, a ninth set of amino acid sequences, each of which has a RNAseq mutation Docket No.: 084276.00417 / LUD 6239 coverage between 0% and 30%; and (h) selecting, from the eighth set of amino acid sequences, a tenth set of amino acid sequences, each of which has a RNAseq mutation coverage greater than 30%, as having a highest likelihood of inducing a tumor-specific immune response, and adding the remaining unselected eighth set of amino acid sequences to the seventh set of amino acid sequences.
[0014] In some embodiments, the method comprises ranking a likelihood of inducing a tumorspecific immune response of the neoantigens in an order from high to low as: the tenth set of amino acid sequences, the ninth set of amino acid sequences, the seventh set of amino acid sequences, and the fifth set of amino acid sequences.
[0015] In some embodiments, the method comprises sequentially sorting the fifth set of amino acid sequences, the seventh set of amino acid sequences, the ninth set of amino acid sequences, or the tenth set of amino acid sequences based on: a likelihood that a mutation within an amino acid sequence is oncogenic, binding affinity to HLA-I and / or HLA-II, a likelihood of presentation, a cancer cell fraction (CCF) value measuring a proportion of cancer cells within a tumor sample that contains a specific genetic mutation, number of predicted overlapping HLA-II binding amino acid sequences, number of predicted included HLA-I binding amino acid sequences, or a combination thereof.
[0016] In some embodiments, the candidate amino acid sequences are in form of peptides.
[0017] In some embodiments, the likelihood that a mutation within an amino acid sequence is oncogenic is determined by CScape.
[0018] In some embodiments, the binding affinity to HLA-I is determined by mixMHCpred, and the binding affinity to HLA-II is determined by mixMHC2pred.
[0019] In some embodiments, the likelihood of presentation is represented as an ipMSDB presentation score.
[0020] In some embodiments, the cancer cell fraction (CCF) value equal to or greater than 1 indicates that a mutation is clonal and affects nearly all cancer cells, and the cancer cell fraction (CCF) value smaller than 1 indicates that the mutation is sub-clonal and affects only a fraction of the tumor cells. Docket No.: 084276.00417 / LUD 6239
[0021] In some embodiments, the number of predicted overlapping HLA-II binding amino acid sequences indicates a potential for broader immune recognition and is determined based on number of HLA-II restricted peptides that cover a given HLA class-I peptide.
[0022] In some embodiments, the number of predicted included HLA-I binding amino acid sequences indicates a potential of a HLA-II restricted peptide to simultaneously present multiple HLA-I peptides.
[0023] In some embodiments, the false-positive mutation is repeatedly identified across samples and cancer types in highly polymorphic genes. In some embodiments, the tumor-specific mutant amino acid sequence does not exist in a reference proteome of a healthy subject.
[0024] In some embodiments, the binding core region is determined by mixMHC2pred. In some embodiments, for a HLA-I binding amino acid sequence, the HLA-I binding amino acid sequence in whole is treated as the binding core region. In some embodiments, for a HLA-II binding amino acid sequence, the binding core region comprises about nine amino acids within the HLA-II binding amino acid sequence.
[0025] In some embodiments, the fourth set of amino acid sequences that exhibit high binding affinity to HLA-I and / or HLA-II have a binding score equal to or less than 1 as determined by mixMHCpred for the HLA-I and mixMHC2pred for the HLA-II.
[0026] In some embodiments, the sixth set of amino acid sequences having the high likelihood of presentation have an ipMSDB presentation score higher than 5 as determined by bestWTPeptideCount l for the HLA-I and bestWTPeptideCount ll for the HLA-II.
[0027] In some embodiments, the transcripts-per-million value indicates gene expression levels and is determined based on RNA sequencing data.
[0028] In some embodiments, the RNAseq mutation coverage indicates prevalence of a specific genetic mutation within RNAseq data.
[0029] In another aspect, this disclosure provides a method for identifying and prioritizing neoantigens in a subject based on mass spectrometry data. In some embodiments, the method comprises: (a) identifying and ranking the neoantigens according to the method as described herein; (b) predicting chromatographic retention times for the neoantigens; (c) acquiring immunopepti domic data using a mass spectrometry-based data acquisition system configured to Docket No.: 084276.00417 / LUD 6239 perform real-time spectral acquisition based on the predicted retention times; and (d) searching raw mass spectrometry data fdes generated in step (c) against personalized peptide reference sequences to prioritize the neoantigens based on search results.
[0030] In some embodiments, the mass spectrometry-based data acquisition system comprises NeoDiscMS.
[0031] In some embodiments, the personalized peptide reference sequences are derived from somatic mutations identified from next generation sequencing data of the subject.
[0032] In some embodiments, the prioritizing of the neoantigens is based at least in part on predicted immunogenicity and expression levels.
[0033] In some embodiments, the predicted chromatographic retention times are used to guide real-time spectral acquisition during mass spectrometry analysis.
[0034] In some embodiments, the mass spectrometry-based data acquisition system configured to perform a data acquisition cycle comprising: (a) performing a full MSI scan to detect precursor ions; (b) for each precursor ion detected in the MSI scan that matches a precursor mass in an inclusion list: (i) performing a first-stage scouting MS2 scan (sMS2); (ii) calculating a peptide- spectrum match (PSM) score in real time by comparing fragment ions from the sMS2 scan to predicted fragment ion spectra from a peptide database; and (iii) if the PSM score meets a predefined quality threshold, performing a second-stage high-sensitivity MS2 scan (hMS2) of the same precursor ion using a higher automatic gain control (AGC) target, an extended maximum injection time, and stepped collision energies; (c) allocating any remaining time to a discovery data-dependent acquisition (DDA) scan branch using wide MS2 isolation windows; wherein the sMS2 and hMS2 scans are exempted from dynamic exclusion and MSI precursor intensity threshold restrictions.
[0035] In some embodiments, the PSM score is calculated using a cross-correlation algorithm between observed and predicted fragment ion spectra.
[0036] In some embodiments, the method further comprises processing raw mass spectrometry data using chimeric spectra deconvolution with a deconvolution algorithm configured to assign peptide-spectrum matches to precursor ions having actual masses deviating from the nominal Docket No.: 084276.00417 / LUD 6239 isolation mass. In some embodiments, the deconvolution algorithm comprises MSFragger operating in DDA+ mode.
[0037] In some embodiments, the data acquisition cycle is repeated to enable repeated MS2 acquisition of the same precursor ion to improve identification confidence.
[0038] In another aspect, this disclosure provides a method of ranking viral antigens. In some embodiments, the method comprises: (i) identifying, from candidate amino acid sequences containing one or more mutations, a first set of amino acid sequences, each of which exhibits high binding affinity to HLA-I and / or HLA-II, and identifying remaining unselected candidate amino acid sequences as a second set of amino acid sequences; (ii) selecting, from the first set of amino acid sequences, a third set of amino acid sequences, each of which is immunogenic in Immune Epitope Database (IEDB) on an HLA allele which a patient expresses; (iii) selecting, from the first set of amino acid sequences, a fourth set of amino acid sequences, each of which is immunogenic in IEDB but not on an HLA allele which the patient expresses; (iv) identifying remaining unselected first set of amino acid sequences as a fifth set of amino acid sequences; and (v) identifying the third set of amino acid sequences as having a highest likelihood as a viral antigen.
[0039] In some embodiments, the method comprises ranking a likelihood as the viral antigen in an order from high to low as: the third set of amino acid sequences, the fourth set of amino acid sequences, the fifth set of amino acid sequences, and the second set of amino acid sequences.
[0040] In some embodiments, the method comprises sequentially sorting the second set of amino acid sequences, the third set of amino acid sequences, the fourth set of amino acid sequences, or the fifth set of amino acid sequences based on: IEDB immunogenicity, number of IEDB validations, viral gene expression, binding affinity to HLA-I and / or HLA-II, or a combination thereof.
[0041] In some embodiments, IEDB immunogenicity is determined based on whether an amino acid sequence has been reported as immunogenic in IEDB. In some embodiments, the number of IEDB validations is determined based on number of times an amino acid sequence has been reported as immunogenic in IEDB.
[0042] In some embodiments, the viral gene expression is determined based on expression level of a viral gene. Docket No.: 084276.00417 / LUD 6239
[0043] In some embodiments, the first set of amino acid sequences that exhibit high binding affinity to HLA-I and / or HLA-II have a binding score equal to or less than 1 as determined by mixMHCpred for the HLA-I and mixMHC2pred for the HLA-II.
[0044] In yet another aspect, this disclosure provides a method of ranking tumor-associated antigens (TAAs) or high-confidence tumor-specific antigens (HC-TSAs) capable of inducing a tumor-specific immune response. In some embodiments, the method comprises: (a) selecting, from candidate amino acid sequences containing one or more mutations, a first set of amino acid sequences, each of which exhibits high binding affinity to HLA-I and / or HLA-II, and identifying remaining unselected third set of amino acid sequences as a second set of amino acid sequences; (b) selecting, from the first set of amino acid sequences, a third set of amino acid sequences, each of which has a high likelihood of presentation, and identifying remaining unselected first set of amino acid sequences as a fourth set of amino acid sequences; (c) selecting, from the third set of amino acid sequences, a fifth set of amino acid sequences, each of which is from a gene having a transcripts-per-million expression value greater than 15, and adding remaining unselected third set of amino acid sequences to the fourth set of amino acid sequences; (d) selecting, from the fifth set of amino acid sequences, a sixth set of amino acid sequences, each of which is immunogenic in Immune Epitope Database (IEDB) on an HLA allele which a patient expresses; (e) selecting, from the fifth set of amino acid sequences, a seventh set of amino acid sequences, each of which is immunogenic in IEDB but not on an HLA allele which the patient expresses; (f) adding remaining unselected fifth set of amino acid sequences to the fourth set of amino acid sequences; and (g) identifying the sixth set of amino acid sequences as having a highest likelihood as a tumor- associated antigen (TAA) or high-confidence tumor-specific antigen (HC-TSA).
[0045] In some embodiments, the method comprises ranking a likelihood as the tumor-associated antigen (TAA) or high-confidence tumor-specific antigen (HC-TSA) in an order from high to low as: the sixth set of amino acid sequences, the seventh set of amino acid sequences, the fourth set of amino acid sequences, and the second set of amino acid sequences.
[0046] In some embodiments, the method comprises sequentially sorting the second set of amino acid sequences, the fourth set of amino acid sequences, the sixth set of amino acid sequences, or the seventh set of amino acid sequences based on: binding affinity to HLA-I and / or HLA-II, number of HLA restrictions, likelihood of presentation, gene expression, or a combination thereof. Docket No.: 084276.00417 / LUD 6239
[0047] In some embodiments, the binding affinity to HLA-I is determined by mixMHCpred, and the binding affinity to HLA-II is determined by mixMHC2pred.
[0048] In some embodiments, the number of HLA restrictions is determined based on number of HLA-I and HLA-II alleles in a patient that presents an amino acid sequence. In some embodiments, the number of HLA restrictions is determined by mixMHCpred for HLA-I and mixMHC2pred for HLA-II.
[0049] In some embodiments, the gene expression is determined based on sample-specific RNAseq data.
[0050] In some embodiments, the first set of amino acid sequences that exhibit high binding affinity to HLA-I and / or HLA-II have a binding score equal to or less than 1 as determined by mixMHCpred for the HLA-I and mixMHC2pred for the HLA-II.
[0051] In some embodiments, the likelihood of presentation is represented as an ipMSDB presentation score. In some embodiments, the third set of amino acid sequences having the ipMSDB presentation score higher than 20 as determined by bestWTPeptideCount l for the HLA- I and bestWTPeptideCount ll for the HLA-II.
[0052] In some embodiments, immunogenicity in IEDB is determined based on whether an amino acid sequence has been reported as immunogenic in IEDB.
[0053] The foregoing summary is not intended to define every aspect of the disclosure, and additional aspects are described in other sections, such as the following detailed description. The entire document is intended to be related as a unified disclosure, and it should be understood that all combinations of features described herein are contemplated, even if the combination of features are not found together in the same sentence, or paragraph, or section of this document. Other features and advantages of the invention will become apparent from the following detailed description. It should be understood, however, that the detailed description and the specific examples, while indicating specific embodiments of the disclosure, are given by way of illustration only, because various changes and modifications within the spirit and scope of the disclosure will become apparent to those skilled in the art from this detailed description. Docket No.: 084276.00417 / LUD 6239
[0054] BRIEF DESCRIPTION OF THE DRAWINGS
[0055] FIG. 1 shows a schematic overview of NeoDisc pipeline. Input data are shown in the top white boxes, while the different modules of NeoDisc are represented in gray boxes and their output is shown as white boxes. The data types used by the modules are indicated. Arrows display the flow of data between modules. Squares below module output boxes highlight which data are used in combination for multiple sample analysis.
[0056] FIG. 2 shows a detailed overview of the NeoDisc pipeline. Input data are shown in the top white boxes, while the different modules of NeoDisc are represented in boxes, with steps in the analysis shown as boxes, and their output is shown as white boxes. The data types used by the modules are indicated. Arrows display the flow of data between modules. Squares below module output boxes highlight which data are used in combination for multiple sample analysis. Tools and databases are annotated next to the boxes.
[0057] FIGS. 3A, 3B, and 3C show NeoDisc rule-based prioritizations. FIG. 3A shows an example flowchart of NeoDisc rule-based prioritization for HLA-I and HLA-II neoantigens. FIG. 3B shows an example flowchart of NeoDisc rule-based prioritization for HLA-I and HLA-II viral antigens. FIG. 3C shows an example flowchart of NeoDisc rule-based prioritization for HLA-I and HLA-II HC-TSAs.
[0058] FIGS. 4A, 4B, and 4C show a schematic overview illustrating the NeoDiscMS workflow. FIG. 4A shows the initial steps in which matched tumor and germline genome sequencing data, together with tumor transcriptome data, are processed by the NeoDisc pipeline to generate a personalized proteome reference specific to the tumor sample. This reference is annotated with single-nucleotide polymorphisms (SNPs) and somatic mutations. The annotated proteome is then utilized for HLA-binding affinity prediction and immunogenicity scoring of candidate antigenic peptides. From the resulting prioritized list of peptides, a personalized inclusion list is generated for downstream use in NeoDiscMS mass spectrometry (MS)-based immunopeptidome acquisition. The MS data acquired from the matched tumor sample is subsequently searched against the personalized proteome by NeoDisc to identify naturally presented peptides, thereby refining and prioritizing clinically relevant and potentially immunogenic targets. FIG. 4B shows a comparative illustration of three MS acquisition strategies: data-dependent acquisition (DDA), DDA with an inclusion list (ilDDA), and the NeoDiscMS method. In NeoDiscMS, each three-second acquisition Docket No.: 084276.00417 / LUD 6239 cycle is divided into three hierarchical levels of priority: an MSI scan, followed by targeted branch MS2 scans, and lastly, discovery branch MS2 scans. FIG. 4C shows the targeted branch of NeoDiscMS, which includes scheduled scouting MS2 scans (sMS2) that are triggered by the detection of a precursor ion in the MSI scan that matches a mass and retention time in the personalized inclusion list. Each sMS2 scan is evaluated in real-time using a database search to determine whether a high-sensitivity MS2 scan (hMS2) should be acquired. The decision is based on real-time metrics including the cross-correlation (Xcorr) between predicted and observed fragment ions, and the mass deviation between predicted and measured precursor ions (expressed in parts per million, PPM).
[0059] DETAILED DESCRIPTION OF THE INVENTION
[0060] Accurate identification and prioritization of antigenic peptides presented by class-I and -II human leukocyte antigens (HLA-I and -II) recognized by autologous T cells is crucial for the development of cancer immunotherapies. To address this challenge, this disclosure provides rulebased models for the prioritization and selection of clinically relevant antigens and demonstrates superiority of the disclosed models in prioritizing neoantigens over existing neoantigen prioritization pipelines and highlighted different features for personalized antigen discovery in the context of heterogenic antigenic landscape and defective cellular antigen presentation machineries.
[0061] FIGS. 3A-C illustrate the exemplary rule-based models for prioritizing tumor neoantigens, viral antigens, and tumor-associated antigens (TAAs) or high-confidence tumor-specific antigens (HC-TSAs). The starting input is amino acid sequences covering sample-specific non-synonymous somatic mutations. The general workflow involves: (a) sorting all peptides: the initial sorting of the peptides is based on specified criteria. This step eventually ensures the consistent sorting of peptides within each of the peptide groups (defined in the next step); (b) peptide groups: assign peptides to groups based on the thresholds defined in the flow charts; and (c) order peptide groups: ordering of the peptide groups.
[0062] Neoantigen Prioritization
[0063] Sorting all peptides involves sequential sorting of the peptides with the following six criteria) (FIG. 3A): Docket No.: 084276.00417 / LUD 6239
[0064] CScape: The CScape score is a computational metric used to predict the likelihood that a particular genetic mutation is oncogenic (Rogers, M.F., et al. Sci Rep 7, 11597 (2017)). A higher CScape score indicates a more oncogenic mutation.
[0065] Binding Rank %: The binding rank is a measure used to evaluate the binding affinity of peptides to human leukocyte antigen (HLA) molecules, specifically for HLA-I and HLA-II types. This measure is crucial for predicting which peptides are most likely to be presented by cancer cells and antigen presenting cells, and potentially elicit an immune response. The binding rank is typically expressed as a percentile, indicating how a peptide’s predicted binding affinity compares to other peptides. Low Percentile Rank', indicates stronger predicted binding affinity (e.g., top 1% of peptides), while high percentile rank indicates weaker predicted binding affinity. The binding rank is derived from predictions made by mixMHCpred (HLA-I) and mixMHC2pred (HLA-II) computational tools (Daniel M. Tadros, et al. bioRxiv 2024.05.08.593183; Julien Racle et al. Immunity, volume 56, issue 6, pl359-1375.e!3 (2023)). A Binding rank % is calculated for all sequences (range 0-100 %) and the peptides are sorted based on the % rank. ipMSDB Presentation: ipMSDB (Muller M, et al. Front Immunol. 2017 Oct 20;8:1367) is a comprehensive database derived from immunopeptidomic analyses of hundreds of samples (cell lines, tissues, etc.), where mass spectrometry is used to identify experimentally eluted nonmutated HLA bound peptides naturally processed and presented peptides on HLA molecules. Some proteins tend to be more presented than others, and that within a protein, some regions are also more presented than others (Muller M, et al. Front Immunol. 2017 Oct 20;8:1367; Kraemer Al, et al. Nat Cancer. 2023 May; 4(5):608-628). These regions were defined as ‘hotspots’ of presentation. For example, hotspots of presentation for both HLA-I and HLA-II were found in both MLANA and MAGEC2 proteins (ipMSDB; (top) part of the plots). The bottom part of the plot shows all the peptides predicted to bind the HLA alleles of patient Mel- 1 (binding rank of 0.5%). In MAGEC2, some regions are predicted to contain many predicted peptides (bottom), but ipMSDB (top) shows that these regions are not covered by peptides that could be detected by mass spectrometry (MS).
[0066] The ipMSDB metric helps identify peptides having a higher potential for natural presentation by HLA molecules. Experimentally validated immunogenic peptides (obtained from the IEDB database; Immune Epitope Database, www.iedb.org) derived from High-Confidence Docket No.: 084276.00417 / LUD 6239
[0067] Tumor Specific Antigen proteins (HC-TSAs) have higher ipMSDB presentation score (ipMSDB coverage) compared to all other predicted peptides from these HC-TSAs.
[0068] The ipMSDB score used is determined with bestWTMatchOverlap I and bestWTMatchOverlap II peptides prioritization. That score reflects the percentage of the peptide sequence overlapping by the best matched peptide sequence in ipMSDB. A higher ipMSDB score indicates a more likely presented peptide.
[0069] Cancer Cell Fraction: The cancer cell fraction (CCF) is a measure used in cancer genomics to estimate the proportion of cancer cells within a tumor sample that contains a specific genetic mutation. This metric provides insight into the clonal architecture of the tumor and helps in understanding the prevalence and significance of mutations within the cancer cell population. CCF ~ 1 or > 1: Indicates that the mutation is clonal and affects nearly all cancer cells. CCF < 1 : indicates that the mutation is sub-clonal and affects only a fraction of the tumor cells. A higher CCF value defines a neoantigen that is more likely to induce a potent response targeting a greater number of cancer cells.
[0070] Number of Predicted Overlapping HLA-II Peptides: determines how many HLA-II restricted peptides (to the patient HLAs) cover a given HLA class-I peptide, indicating the potential for broader immune recognition.
[0071] Number of Predicted Included HLA-I Peptides: evaluates how many distinct HLA-I peptides are included within a given HLA -II peptide sequence. This measure is used to assess the potential of a HLA-II restricted peptide to simultaneously present multiple HLA-I peptides, potentially enhancing immune recognition and response.
[0072] Peptide Groups:
[0073] 1. False positive (FP) mutations o Yes / No: Indicates whether the peptide derives from a false-positive mutation call (mutations being repeatedly identified across samples and cancer types in highly polymorphic genes), considering mutations in the following genes: HLA-A, HLA- B, HLA-C, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5,HLA-DPA1, HLA-DPB1, HLA-DQA1, HLA-DQB1,HLA-DMA,TRBV3, TRBV5, TRBV6, TRBV6-1, TRBV10, TRBV10-1, TRBV11,TRAV12, KRT1, and PRSS3. Docket No.: 084276.00417 / LUD 6239 Impact: False positive mutations calls are de-prioritized. -specific Yes / No: Defines whether the mutant peptide sequence exists or does not exist in a healthy reference proteome. Yes: the peptide is uniquely found in the tumor, No: the peptide can also be found in the healthy proteome. Non-tumor-specific peptides could lead to autoimmune responses. Impact: Non-tumor-specific peptides are removed. on within binding core Yes / No: Determines if the mutation resides within the binding core region of the peptide that directly interacts with the HLA molecule in the binding groove, forming a stable HLA-peptide complex. For HLA-I peptides, the entire peptide sequence is considered as the binding core. For HLA-II peptides, the binding core is usually around 9 amino acids long contained within the sequence of the HLA-II peptide. The binding core is predicted with mixMHC2pred. Impact: Mutations within the binding core may have direct impact on binding to the HLA and on TCR recognition, and therefore they are prioritized over mutations in the flanking regions. g rank: Binding Rank < 1 : High affinity peptides are prioritized over others. mixMHCpred and mixMHC2pred are used for HLA-I and HLA-II binding predictions, respectively. Importance: Represents the strength of the peptide’ s binding affinity and, therefore, likelihood of presentation. Immunogenicity of neoantigens is positively associated with high HLA binding affinity. B presentation score The scores used here are bestWTPeptideCount l and bestWTPeptideCount ll, which reflect the number of HLA-I and HLA-II peptide sequences that cover the position of the mutation in ipMSDB, respectively. Docket No.: 084276.00417 / LUD 6239 >5: Predicted neoantigen peptides with a presentation score greater than 5 (5 or more peptides in ipMSDB cover the mutation site) are more likely to be presented than others. Relevance: Higher scores indicate an increased likelihood of presentation. xpression (TPM) TPM (Transcripts Per Million) is a normalization method used in RNA sequencing (RNA-seq) data analysis to quantify gene expression levels. It provides a way to compare the expression of genes across different samples and conditions. If available, sample-specific (derived from RNAseq data) gene expression is used. If not, the median gene expression derived from the matched cancer type in TCGA is used instead. >15: Peptides derived from medium / highly expressed genes (>15TPM) are prioritized over others. Significance: High gene expression indicates an increased likelihood of protein expression, processing, and presentation. Peptide presentation is positively associated with a higher source gene expression, and neoantigen immunogenicity is positively associated with a higher gene expression. q Mutation Coverage (%) Quantification of the prevalence of a specific genetic mutation within RNAseq data. It indicates how frequently the mutation appears in the sequenced reads covering the mutation site. > 30: High support of mutation expression > 0: (> 0 and <30): Support of mutation expression 0: No support of mutation expression Impact: High mutation coverage supports the accuracy of the mutation calling and its expression level. Peptide presentation is positively associated with a higher source gene expression, and neoantigen immunogenicity is positively associated with a higher gene expression. Docket No.: 084276.00417 / LUD 6239
[0074] Order peptide groups:
[0075] Peptide groups are ordered from the best (A) to the worst (Reject) group as A>B>OD>Rej ect
[0076] Performance: The performance of NeoDisc rule-based neoantigen prioritization was benchmarked for HLA-I restricted neoantigens on the NCI cohort. Superior performance of the prioritization approach compared to pTuneos and pVACseq was observed. This prioritization method was also applied to HLA-II restricted neoantigens.
[0077] In one aspect, this disclosure provides a method of ranking neoantigens capable of inducing a tumor-specific immune response. In some embodiments, the method comprises: (a) selecting, from candidate amino acid sequences containing one or more mutations, a first set of amino acid sequences, each of which is free of a false-positive mutation; (b) identifying, from the first set of amino acid sequences, a second set of amino acid sequences, each of which contains a tumorspecific mutant amino acid sequence; (c) identifying, from the second set of amino acid sequences, a third set of amino acid sequences, each of which contains a mutation within a binding core region interacting with a binding groove of a human leukocyte antigen (HLA) molecule; (d) selecting, from the third set of amino acid sequences, a fourth set of amino acid sequences, each of which exhibits high binding affinity to HLA-I and / or HLA-II, and identifying remaining unselected third set of amino acid sequences as a fifth set of amino acid sequences; (e) selecting, from the fourth set of amino acid sequences, a sixth set of amino acid sequences, each of which has a high likelihood of presentation, and identifying remaining unselected fourth set of amino acid sequences as a seventh set of amino acid sequences; (f) selecting, from the sixth set of amino acid sequences, an eighth set of amino acid sequences, each of which is from a gene having a transcripts-per- million expression value greater than 15 (e.g., 15 to 106), and adding remaining unselected sixth set of amino acid sequences as a seventh set of amino acid sequences; (g) selecting, from the eighth set of amino acid sequences, a ninth set of amino acid sequences, each of which has a RNAseq mutation coverage, for example, between 0% and 30% (e.g., 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, or any intermediate values therebetween); and (h) selecting, from the eighth set of amino acid sequences, a tenth set of amino acid sequences, each of which has a RNAseq mutation coverage greater than 30%, as having a highest likelihood of inducing a tumor- Docket No.: 084276.00417 / LUD 6239 specific immune response, and adding the remaining unselected eighth set of amino acid sequences to the seventh set of amino acid sequences.
[0078] In some embodiments, the candidate amino acid sequences are in the form of peptides (or polypeptides).
[0079] In some embodiments, the method comprises ranking a likelihood of inducing a tumorspecific immune response of the neoantigens in an order from high to low as: the tenth set of amino acid sequences, the ninth set of amino acid sequences, the seventh set of amino acid sequences, and the fifth set of amino acid sequences.
[0080] In some embodiments, the method comprises sequentially sorting the fifth set of amino acid sequences, the seventh set of amino acid sequences, the ninth set of amino acid sequences, or the tenth set of amino acid sequences based on: a likelihood that a mutation within an amino acid sequence is oncogenic, binding affinity to HLA-I and / or HLA-II, a likelihood of presentation, a cancer cell fraction (CCF) value measuring a proportion of cancer cells within a tumor sample that contains a specific genetic mutation, number of predicted overlapping HLA-II binding amino acid sequences, number of predicted included HLA-I binding amino acid sequences, or a combination thereof.
[0081] In some embodiments, the likelihood that a mutation within an amino acid sequence is oncogenic is determined by CScape. CScape refers to a computational tool or algorithm configured to predict the oncogenic potential of somatic mutations, particularly single-nucleotide variants (SNVs), in cancer genomes. CScape may be implemented to classify mutations as likely cancer drivers or passengers based on machine learning models trained on large-scale cancer genomic datasets. Features used in CScape may include, but are not limited to, evolutionary conservation, gene context, sequence properties, and functional impact scores. CScape may be employed in silico to support the prioritization of candidate neoantigens or therapeutic targets in oncology- related applications.
[0082] In some embodiments, the binding affinity to HLA-I is determined by mixMHCpred, and the binding affinity to HLA-II is determined by mixMHC2pred. mixMHCpred refers to a computational tool for predicting peptide binding affinity to major histocompatibility complex (MHC) class I molecules. The mixMHCpred algorithm is based on position weight matrices derived from large-scale monoallelic mass spectrometry datasets, and it accounts for peptide length Docket No.: 084276.00417 / LUD 6239 distribution and motifs specific to different HLA alleles. The tool is used to score peptides based on their predicted likelihood of being naturally presented on the cell surface in complex with MHC class I molecules, thereby aiding in the identification and prioritization of potential T cell epitopes and neoantigens.
[0083] In some embodiments, the likelihood of presentation is represented as an ipMSDB presentation score. As used herein, the term “ipMSDB presentation score” refers to a numerical or probabilistic value that quantifies the likelihood that a given peptide is presented on the surface of a cell in complex with a major histocompatibility complex (MHC) molecule, based on evidence from the ipMSDB (immunopeptidomics mass spectrometry database). The score may be derived from experimental mass spectrometry -based immunopeptidomic data, computational models trained on such data, or a combination thereof. The ipMSDB presentation score may incorporate factors including, but not limited to, peptide abundance, binding affinity, peptide length, HLA allele context, and frequency of detection across datasets. A higher score indicates a greater probability that the peptide is naturally processed and presented in vivo.
[0084] In some embodiments, the cancer cell fraction (CCF) value equal to or greater than 1 indicates that a mutation is clonal and affects nearly all cancer cells, and the cancer cell fraction (CCF) value smaller than 1 (e. ., 0.001, 0.01, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or any intermediate values therebetween) indicates that the mutation is sub-clonal and affects only a fraction of the tumor cells. As used herein, the term “cancer cell fraction (CCF) value” refers to an estimated measure of the proportion of cancer cells within a biological sample (e.g., a tumor biopsy) that harbor a particular genetic alteration, such as a somatic mutation. The CCF value may be expressed as a percentage or fractional value ranging from 0 to 1 (or 0% to 100%), and reflects the clonality of the mutation within the cancer cell population. In some embodiments, the CCF value is derived from variant allele frequency (VAF) data, corrected for local copy number variations, tumor purity, and ploidy. A CCF value approaching 1 (or 100%) may indicate that the mutation is clonal and present in nearly all cancer cells in the sample, whereas a lower CCF value may indicate that the mutation is subclonal and present in only a subset of cancer cells. CCF values can be used to prioritize mutations for therapeutic targeting, neoantigen prediction, or clonality analysis in precision oncology applications. Docket No.: 084276.00417 / LUD 6239
[0085] In some embodiments, the number of predicted overlapping HLA-II binding amino acid sequences indicates a potential for broader immune recognition and is determined based on number of HLA-II restricted peptides that cover a given HLA class-I peptide.
[0086] In some embodiments, the number of predicted included HLA-I binding amino acid sequences indicates a potential of a HLA-II restricted peptide to simultaneously present multiple HLA-I peptides.
[0087] In some embodiments, the false-positive mutation is repeatedly identified across samples and cancer types in highly polymorphic genes. In some embodiments, the tumor-specific mutant amino acid sequence does not exist in a reference proteome of a healthy subject.
[0088] In some embodiments, the binding core region is determined by mixMHC2pred. In some embodiments, for a HLA-I binding amino acid sequence, the HLA-I binding amino acid sequence in whole is treated as the binding core region. In some embodiments, for a HLA-II binding amino acid sequence, the binding core region comprises about nine amino acids within the HLA-II binding amino acid sequence.
[0089] In some embodiments, the fourth set of amino acid sequences that exhibit high binding affinity to HLA-I and / or HLA-II have a binding score equal to or less than 1 (e.g., 0.001, 0.01, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, or any intermediate values therebetween) as determined by mixMHCpred for the HLA-I and mixMHC2pred for the HLA-II. In some embodiments, the fifth set of amino acid sequences have a binding score greater than 1 (e.g., 1 to 100).
[0090] In some embodiments, the sixth set of amino acid sequences having the high likelihood of presentation have an ipMSDB presentation score higher than 5 as determined by bestWTPeptideCount l for the HLA-I and bestWTPeptideCount ll for the HLA-II. In some embodiments, the remaining unselected fourth set of amino acid sequences as a seventh set of amino acid sequences have an ipMSDB presentation score lower than or equal to 5 (e.g., 1, 2, 3, 4, 5, or any intermediate values therebetween) as determined by bestWTPeptideCount l for the HLA-I and bestWTPeptideCount ll for the HLA-II.
[0091] In some embodiments, the transcripts-per-million value indicates gene expression levels and is determined based on RNA sequencing data. Docket No.: 084276.00417 / LUD 6239
[0092] In some embodiments, the RNAseq mutation coverage indicates the prevalence (or frequency) of a specific genetic mutation within RNAseq data.
[0093] NeoDiscMS
[0094] In another aspect, this disclosure provides a method for identifying and prioritizing neoantigens in a subject based on mass spectrometry data. In some embodiments, the method comprises: (a) identifying and ranking the neoantigens according to the method as described herein; (b) predicting chromatographic retention times for the neoantigens; (c) acquiring immunopeptidomic data using a mass spectrometry-based data acquisition system configured to perform real-time spectral acquisition based on the predicted retention times; and (d) searching raw mass spectrometry data files generated in step (c) against personalized peptide reference sequences to prioritize the neoantigens based on search results.
[0095] As used herein, the term “immunopeptidomic data” refers to information obtained through the analysis and characterization of peptides that are naturally processed and presented on cell surface major histocompatibility complex (MHC) molecules, also known as human leukocyte antigen (HLA) molecules. Such data typically includes the identity, abundance, sequence, and post-translational modifications of HLA-bound peptides, and is generally generated using mass spectrometry-based techniques. Immunopeptidomic data may further comprise metadata associated with the sample, experimental conditions, retention times, spectral quality, peptide- MHC binding affinities, and source protein annotations. This data is used, for example, to identify tumor-associated or neoantigenic peptides for applications in immunotherapy, vaccine development, or diagnostic profiling.
[0096] As used herein, the term “personalized peptide reference sequences” refers to a set of peptide sequences that are specifically derived from an individual subject’s genomic and / or transcriptomic data, such as data obtained through next-generation sequencing (NGS) of tumor and / or normal tissue. These sequences may include peptides corresponding to somatic mutations, gene fusions, alternative splicing events, or other subject-specific genetic alterations that result in the expression of neoantigens. The personalized peptide reference sequences are generated in silico and are used as a custom reference database for the identification and analysis of peptides in mass spectrometry -based immunopeptidomics workflows. Such reference sequences enable subjectspecific identification and prioritization of tumor-associated or tumor-specific antigens. Docket No.: 084276.00417 / LUD 6239
[0097] In some embodiments, the mass spectrometry -based data acquisition system comprises NeoDiscMS.
[0098] As used herein, “NeoDiscMS” refers to a mass spectrometry-based immunopeptidomics data acquisition platform that operates as an extension of the NeoDisc workflow. NeoDiscMS is configured to perform targeted, real-time spectral acquisition guided by subject-specific genomic data, including predicted chromatographic retention times of tumor-specific peptides. The platform enables enhanced detection of immunogenic tumor-associated antigens by integrating next-generation sequencing (NGS) data with mass spectrometry analysis, thereby facilitating the personalized identification and prioritization of neoantigens. NeoDiscMS is designed to improve sensitivity and confidence in antigen discovery while maintaining broad coverage of the immunopeptidome, even in clinical settings with limited sample input and time constraints.
[0099] In some embodiments, the personalized peptide reference sequences are derived from somatic mutations identified from next generation sequencing data of the subject.
[0100] In some embodiments, the prioritizing of the neoantigens is based at least in part on predicted immunogenicity and expression levels.
[0101] In some embodiments, the predicted chromatographic retention times are used to guide real-time spectral acquisition during mass spectrometry analysis.
[0102] In some embodiments, the mass spectrometry-based data acquisition system configured to perform a data acquisition cycle comprising: (a) performing a full MSI scan to detect precursor ions; (b) for each precursor ion detected in the MSI scan that matches a precursor mass in an inclusion list: (i) performing a first-stage scouting MS2 scan (sMS2); (ii) calculating a peptide- spectrum match (PSM) score in real time by comparing fragment ions from the sMS2 scan to predicted fragment ion spectra from a peptide database; and (iii) if the PSM score meets a predefined quality threshold, performing a second-stage high-sensitivity MS2 scan (hMS2) of the same precursor ion using a higher automatic gain control (AGC) target, an extended maximum injection time, and stepped collision energies; (c) allocating any remaining time to a discovery data-dependent acquisition (DDA) scan branch using wide MS2 isolation windows; wherein the sMS2 and hMS2 scans are exempted from dynamic exclusion and MSI precursor intensity threshold restrictions. Docket No.: 084276.00417 / LUD 6239
[0103] As used herein, the term “peptide-spectrum match (PSM) score” refers to a numerical or probabilistic value that quantifies the confidence of a match between an observed mass spectrometry (MS / MS) fragmentation spectrum and a candidate peptide sequence. The PSM score is typically generated by a search algorithm that compares the experimental spectrum to theoretical spectra derived from one or more peptide sequences in a reference database. A higher PSM score indicates a better match and greater confidence that the peptide sequence corresponds to the observed spectrum. In some implementations, PSM scores may be derived using scoring algorithms such as SEQUEST, Mascot, Andromeda, or other suitable computational methods, and may incorporate factors such as fragment ion intensity, mass accuracy, and the number of matching fragment ions.
[0104] As used herein, the term “automatic gain control (AGC)” refers to a signal processing technique or system configured to automatically adjust the gain of an input signal to maintain an output signal within a desired amplitude range. AGC systems typically detect the amplitude level of the input or output signal and apply a feedback mechanism to dynamically increase or decrease gain in response to variations in signal strength. This allows for consistent signal levels despite fluctuations in input signal amplitude and may be implemented in analog, digital, or mixed-signal circuits.
[0105] As used herein, the term “discovery data-dependent acquisition (DDA)” refers to a mass spectrometry (MS)-based data acquisition strategy in which precursor ions are selected for fragmentation in real time based on their intensity in a survey scan. In a typical DDA workflow, a full MS scan (MSI) is first performed to detect all ions present in a sample at a given time point. A predefined number of the most intense ions are then automatically selected for subsequent fragmentation and tandem MS analysis (MS2). This approach enables broad and unbiased detection of peptides or analytes in complex mixtures and is commonly employed in discovery proteomics and immunopeptidomics to identify novel or low-abundance targets. The DDA method is referred to as “data-dependent” because the selection of precursor ions for fragmentation depends on the data obtained from the preceding MSI scan.
[0106] In some embodiments, the PSM score is calculated using a cross-correlation algorithm between observed and predicted fragment ion spectra. As used herein, the term “cross-correlation algorithm” refers to a computational method or process that measures the similarity or degree of Docket No.: 084276.00417 / LUD 6239 correlation between two signals, datasets, or functions as a function of the displacement (or time lag) of one relative to the other. In certain embodiments, the cross-correlation algorithm may be used to identify temporal or spatial alignment between a reference signal and a target signal, to detect repeating patterns or features, or to quantify the degree of similarity between observed and predicted data. The algorithm may be implemented using discrete or continuous mathematical formulations and may be executed in hardware, software, or a combination thereof.
[0107] In some embodiments, the method further comprises processing raw mass spectrometry data using chimeric spectra deconvolution with a deconvolution algorithm configured to assign peptide-spectrum matches to precursor ions having actual masses deviating from the nominal isolation mass. In some embodiments, the deconvolution algorithm comprises MSFragger operating in DDA+ mode.
[0108] As used herein, “MSFragger” refers to an open-source, ultrafast database search tool designed for mass spectrometry-based proteomics. MSFragger performs peptide-spectrum matching by rapidly comparing acquired tandem mass spectrometry (MS / MS) spectra against theoretical spectra generated from a protein or peptide database. MSFragger supports both closed and open searches, including variable modifications and mass shifts, and is capable of identifying post-translational modifications, sequence variants, and non-canonical peptides. The algorithm utilizes fragment ion indexing to achieve high-speed performance while maintaining high sensitivity and accuracy in peptide identification.
[0109] As used herein, the term “DDA+ mode” refers to an enhanced data-dependent acquisition (DDA) mode used in mass spectrometry-based analysis. In DDA+ mode, precursor ions are selected for fragmentation based on predefined criteria, such as intensity or charge state, with additional real-time or retrospective prioritization features that improve sampling of low- abundance or user-defined target peptides. DDA+ mode may incorporate real-time spectral acquisition guidance, such as retention time prediction, dynamic exclusion, inclusion lists, or machine learning-based prioritization, to increase the sensitivity, specificity, and reproducibility of peptide detection. DDA+ mode may be implemented in conjunction with personalized peptide reference libraries and / or informed by genomic or transcriptomic data.
[0110] In some embodiments, the data acquisition cycle is repeated to enable repeated MS2 acquisition of the same precursor ion to improve identification confidence. Docket No.: 084276.00417 / LUD 6239
[0111] Viral Antigen Prioritization
[0112] Sorting all peptides (FIG. 3B):
[0113] IEDB Immunogenicity: The IEDB (Immune Epitope Database, www.iedb.org) was parsed to retrieve reported viral immunogenic peptide sequences. Predicted peptides are sorted based on whether they were reported immunogenic (ranking better) or not (ranking lower) in IEDB. Here, this is a binary ranking of reported immunogenicity as “true” or “false.”
[0114] Number of IEDB Validations: Peptides are sorted based on the number of times they were reported immunogenic in IEDB. More validations indicate higher chances of immunogenicity.
[0115] Viral Gene Expression: Sample-specific RNAseq data (RNAseq data (transcriptome) of the patient’s tumor) is used to quantify viral gene expression, which is used to prioritize peptides derived from highly expressed viral genes over others.
[0116] Binding Rank %: same as neoantigens “Sorting all peptides” description above.
[0117] Peptide Groups:
[0118] 1. Binding Rank: same as neoantigens “Group peptides” description above
[0119] 2. IEDB Immunogenicity o No: Peptide was not reported immunogenic in IEDB o Other HLA: Peptide was reported immunogenic in IEDB but on different HLA restriction(s) (i.e., the patient does not express the HLA allele(s) reported in IEDB to bind the immunogenic peptide) o Patient HLA: Peptide was reported immunogenic in IEDB on an HLA allele which the patient express. o Significance: Evidence of immunogenicity in IEDB is a strong indicator of peptide immunogenicity.
[0120] Order peptide groups:
[0121] Peptide groups are ordered from the best (A) to the worst (D) group as A>B>OD Docket No.: 084276.00417 / LUD 6239
[0122] Performance: The performance of NeoDisc rule-based viral antigen prioritization was not benchmarked yet. NeoDisc ranked an immunogenic HPV18 derived peptide in second position (patient CESC1) and allowed the identification of 34 EBV derived immunogenic peptides among a total of 108 peptides tested (patient NPC1).
[0123] In another aspect, this disclosure provides a method of ranking viral antigens. In some embodiments, the method comprises: (i) identifying, from candidate amino acid sequences containing one or more mutations, a first set of amino acid sequences, each of which exhibits high binding affinity to HLA-I and / or HLA-II, and identifying remaining unselected candidate amino acid sequences as a second set of amino acid sequences; (ii) selecting, from the first set of amino acid sequences, a third set of amino acid sequences, each of which is immunogenic in Immune Epitope Database (IEDB) on an HLA allele which a patient expresses; (iii) selecting, from the first set of amino acid sequences, a fourth set of amino acid sequences, each of which is immunogenic in IEDB but not on an HLA allele which the patient expresses; (iv) identifying remaining unselected first set of amino acid sequences as a fifth set of amino acid sequences; and (v) identifying the third set of amino acid sequences as having a highest likelihood as a viral antigen.
[0124] In some embodiments, the method comprises ranking a likelihood as the viral antigen in an order from high to low as: the third set of amino acid sequences, the fourth set of amino acid sequences, the fifth set of amino acid sequences, and the second set of amino acid sequences.
[0125] In some embodiments, the method comprises sequentially sorting the second set of amino acid sequences, the third set of amino acid sequences, the fourth set of amino acid sequences, or the fifth set of amino acid sequences based on: IEDB immunogenicity, number of IEDB validations, viral gene expression, binding affinity to HLA-I and / or HLA-II, or a combination thereof.
[0126] In some embodiments, IEDB immunogenicity is determined based on whether an amino acid sequence has been reported as immunogenic in IEDB. In some embodiments, the number of IEDB validations is determined based on number of times an amino acid sequence has been reported as immunogenic in IEDB.
[0127] In some embodiments, the viral gene expression is determined based on expression level of a viral gene. Docket No.: 084276.00417 / LUD 6239
[0128] In some embodiments, the first set of amino acid sequences that exhibit high binding affinity to HLA-I and / or HLA-II have a binding score equal to or less than 1 (e.g., 0.001, 0.01, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, or any intermediate values therebetween) as determined by mixMHCpred for the HLA-I and mixMHC2pred for the HLA-II. In some embodiments, the second set of amino acid sequences have a binding score greater than 1 (e.g., 1 tolOO).
[0129] TAAs (Tumor-associated Antigens) and HC-TSAs (High-Confidence Tumor-Specific Antigens) Prioritization
[0130] Sorting all peptides (FIG. 3C):
[0131] Binding Rank %: same as neoantigens “Sorting all peptides” description above.
[0132] Number of HLA Restrictions: is metric used to evaluate the diversity and breadth of peptide binding to the HLA alleles of the patient. It indicates how many different HLA-I and HLA- II alleles in a patient can potentially present a given peptide, as predicted by computational tools (mixMHCpred for HLA-I and mixMHC2pred for HLA-II). A higher number of HLA restrictions indicates that a peptide can be presented by multiple HLA alleles, increasing the likelihood presentation and recognition by various T cell populations. ipMSDB Score: the bestWTMatchScore l and bestWTMatchScore ll metrics are used for HLA-I and HLA-II peptides prioritization, respectively. These scores represent the sum of the HLA-I / II density profile height over the position of the peptide (i.e., the sum of the number of peptides in ipMSDB which cover each amino acid in the predicted peptide sequence). The match score is more suitable for TAAs-derived peptides as it considers the entire peptide sequence, while for Neoantigens, the position of the mutation is more important.
[0133] Gene Expression: quantification of the source gene expression from sample-specific RNAseq data. Peptides derived from genes with higher expression are prioritized.
[0134] Group peptides:
[0135] 1. Binding Rank %: same as neoantigen “Peptide Groups” description above
[0136] 2. ipMSDB presentation score: same as neoantigen “Peptide Groups” description above. However, here the bestWTMatchScore I and bestWTMatch Score II metrics are used.
[0137] 3. Gene Expression (TPM): same as neoantigen “Peptide Groups” description above Docket No.: 084276.00417 / LUD 6239
[0138] 4. IEDB Immunogenicity: same as viral antigen “Peptide Groups” description above. The difference is that IEDB is restricted to HC-TSAs-derived peptides.
[0139] Order peptide groups:
[0140] Peptide groups are ordered from the best (A) to the worst (D) group as A>B>OD
[0141] Performance: HC-TSAs ranking is demonstrated in Table 4. Ranks of immunogenic HC- TSAs (that were identified by MS) are provided for patients MEL-1, MEL-2, and MEL-4.
[0142] In another aspect, this disclosure provides a method of ranking tumor-associated antigens (TAAs) or high-confidence tumor-specific antigens (HC-TSAs) capable of inducing a tumorspecific immune response. In some embodiments, the method comprises: (a) selecting, from candidate amino acid sequences containing one or more mutations, a first set of amino acid sequences, each of which exhibits high binding affinity to HLA-I and / or HLA-II, and identifying remaining unselected third set of amino acid sequences as a second set of amino acid sequences; (b) selecting, from the first set of amino acid sequences, a third set of amino acid sequences, each of which has a high likelihood of presentation, and identifying remaining unselected first set of amino acid sequences as a fourth set of amino acid sequences; (c) selecting, from the third set of amino acid sequences, a fifth set of amino acid sequences, each of which is from a gene having a transcripts-per-million expression value greater than 15 (c. ., 15 to 106), and adding remaining unselected third set of amino acid sequences (e.g., from a gene having a transcripts-per-million expression value less than or equal to 15 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or any intermediate values therebetween) to the fourth set of amino acid sequences; (d) selecting, from the fifth set of amino acid sequences, a sixth set of amino acid sequences, each of which is immunogenic in Immune Epitope Database (IEDB) on an HLA allele which a patient expresses; (e) selecting, from the fifth set of amino acid sequences, a seventh set of amino acid sequences, each of which is immunogenic in IEDB but not on an HLA allele which the patient expresses; (f) adding remaining unselected fifth set of amino acid sequences to the fourth set of amino acid sequences; and (g) identifying the sixth set of amino acid sequences as having a highest likelihood as a tumor-associated antigen (TAA) or high-confidence tumor-specific antigen (HC-TSA).
[0143] In some embodiments, the method comprises ranking a likelihood as the tumor-associated antigen (TAA) or high-confidence tumor-specific antigen (HC-TSA) in an order from high to low Docket No.: 084276.00417 / LUD 6239 as: the sixth set of amino acid sequences, the seventh set of amino acid sequences, the fourth set of amino acid sequences, and the second set of amino acid sequences.
[0144] In some embodiments, the method comprises sequentially sorting the second set of amino acid sequences, the fourth set of amino acid sequences, the sixth set of amino acid sequences, or the seventh set of amino acid sequences based on: binding affinity to HLA-I and / or HLA-II, number of HLA restrictions, likelihood of presentation, gene expression, or a combination thereof.
[0145] In some embodiments, the binding affinity to HLA-I is determined by mixMHCpred, and the binding affinity to HLA-II is determined by mixMHC2pred.
[0146] In some embodiments, the number of HLA restrictions is determined based on number of HLA-I and HLA-II alleles in a patient that presents an amino acid sequence. In some embodiments, the number of HLA restrictions is determined by mixMHCpred for HLA-I and mixMHC2pred for HLA-II.
[0147] In some embodiments, the gene expression is determined based on sample-specific RNAseq data.
[0148] In some embodiments, the first set of amino acid sequences that exhibit high binding affinity to HLA-I and / or HLA-II have a binding score equal to or less than 1 (e.g., 0.001, 0.01, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, or any intermediate values therebetween) as determined by mixMHCpred for the HLA-I and mixMHC2pred for the HLA-II. In some embodiments, the second set of amino acid sequences have a binding score greater than 1 (e.g., 1 tolOO).
[0149] In some embodiments, the likelihood of presentation is represented as an ipMSDB presentation score. In some embodiments, the third set of amino acid sequences having the ipMSDB presentation score higher than 20 as determined by bestWTPeptideCount l for the HLA- I and bestWTPeptideCount ll for the HLA-II. In some embodiments, the remaining unselected first set of amino acid sequences have the ipMSDB presentation score less than or equal to 20 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or any intermediate values therebetween) as determined by bestWTPeptideCount l for the HLA-I and bestWTPeptideCount ll for the HLA-II.
[0150] In some embodiments, immunogenicity in IEDB is determined based on whether an amino acid sequence has been reported as immunogenic in IEDB. As used herein, the term “IEDB” refers Docket No.: 084276.00417 / LUD 6239 to the Immune Epitope Database, a publicly available bioinformatics resource maintained by the National Institute of Allergy and Infectious Diseases (NIAID). The IEDB contains curated experimental data and computational tools related to immune epitopes, including T cell and B cell epitopes, as well as major histocompatibility complex (MHC) binding data. The IEDB may be utilized to predict or evaluate the immunogenicity of peptides based on their potential to bind specific HLA alleles and to elicit immune responses. In some embodiments, IEDB prediction tools may be employed to assess antigenic peptide candidates identified through in silico analysis, nextgeneration sequencing data, or mass spectrometry-based immunopeptidomics.
[0151] Additional Definitions
[0152] As used herein, the terms “first,” “second,” “third,” and the like are employed solely for clarity of distinction between different elements and do not necessarily imply any order or hierarchy unless expressly stated. Such terms are intended to denote that the referenced entities are different from one another.
[0153] The term “ / / / vitro,” as used herein, refers to processes or events that occur outside of a living organism, such as in a test tube, reaction vessel, or cell culture system.
[0154] The term “zn vivo,” as used herein, refers to processes or events that occur within a living, multicellular organism, including, for example, a non-human animal model.
[0155] It is noted that, as used in the present specification and in the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise.
[0156] The terms “including,” “comprising,” “containing,” “having,” and grammatical variations thereof, are intended to be inclusive and open-ended, such that the recitation of items following such terms is not intended to exclude other, unrecited items or equivalents thereof, unless explicitly stated otherwise.
[0157] The phrases “in one embodiment,” “in various embodiments,” “in some embodiments,” and similar expressions, as used herein, are not intended to refer to the same embodiment unless expressly stated, and may refer to different or alternative embodiments.
[0158] The terms “and / or” and “ / ” as used herein are intended to mean any one of the listed items, any combination of the listed items, or all of the listed items. Docket No.: 084276.00417 / LUD 6239
[0159] The term “substantially,” as used herein, is not intended to exclude the possibility of complete correspondence or absence. For example, a composition that is “substantially free” of a component may be entirely free of that component. Where appropriate, the term “substantially” may be omitted from the present disclosure without loss of generality.
[0160] The terms “approximately” and “about,” as used herein in connection with numerical values, are intended to mean a value that is within an acceptable range of deviation from a stated reference value. In some embodiments, the range may be within ±25%, ±20%, ±15%, ±10%, ±5%, ±2%, ±1%, or less, of the stated value, unless context dictates otherwise or unless such deviation would exceed the maximum possible value (e. ., 100%). Unless indicated otherwise, the term “about” is intended to encompass values proximate to the recited value that do not materially affect the functionality or intended outcome of the invention.
[0161] It is to be understood that wherever a value or range of values is stated herein, all values and sub-ranges encompassed within that value or range are intended to be explicitly disclosed and encompassed within the scope of the present disclosure. This includes all values within the stated range as well as the endpoints of the range.
[0162] As used herein, the term “each,” when referring to a collection of elements or components, is intended to identify individual elements within the collection but does not necessarily imply that all such elements are included unless the context clearly indicates otherwise.
[0163] The use of examples or exemplary language (e.g., “such as”) is intended for illustrative purposes only and is not to be construed as limiting the scope of the invention unless expressly recited in the claims. No language in the specification should be interpreted as designating any unclaimed feature as essential to the invention.
[0164] Unless otherwise indicated or clearly contradicted by context, the steps of any method disclosed herein may be performed in any order, and steps may be performed simultaneously or sequentially. All permutations, combinations, and sub-combinations of method steps are intended to be encompassed by the present disclosure unless expressly excluded.
[0165] Each publication, patent application, patent, and other reference cited herein is hereby incorporated by reference in its entirety to the extent that it is not inconsistent with the present disclosure. Such references are provided solely for their disclosure as of the filing date of the Docket No.: 084276.00417 / LUD 6239 present application. No admission is made that any cited reference constitutes prior art against the present invention. The actual publication dates of references may differ from those cited herein and may require independent verification.
[0166] It is to be understood that the examples and embodiments described herein are provided for illustrative purposes only. Variations, modifications, and alternative embodiments that would be apparent to those of ordinary skill in the art in view of the present disclosure are considered to fall within the scope of the appended claims and the spirit of the invention as disclosed.
[0167] Examples
[0168] EXAMPLE 1
[0169] Installation and configuration of NeoDisc
[0170] NeoDisc is available upon registration. Access is available at https: / / neodisc.unil. ch / A manual file, describing NeoDisc requirements and outputs, is provided with NeoDisc to guide the user through the installation and configuration process.
[0171] DNA reads ali nment
[0172] FastQ reads were aligned and processed following GATK best practice workflow for data pre-processing for variant discovery (Geraldine A Van der Auwera et a , Curr Protoc Bioinformatics. 2013;43(l 110): 11.10.1-11.10.33). Briefly, FastQ reads were aligned to the reference human genome (GRCh37) (https: / / www.ncbi.nlm.nih.gov / assembly / GCF_ 000001405.13 / ) with bwa (0.7.17) (Li, H. & Durbin, R, Bioinformatics 25, 1754-1760 (2009)). Duplicate reads were marked, and base quality scores were recalibrated with GATK 4.4.0.0 tools (MarkDuplicatesSpark, BaseRecalibrator, ApplyBQSR) (Van der Auwera, G.A. etal. Curr Protoc Bioinformatics 43, 11 10 11-11 10 33 (2013)). Alignment and reads quality were assessed with picard tools v3.0.0 (https: / / broadinstitute.github.io / picard / ) (CollectAlignmentSummaryMetrics, CollectHsMetrics) and fastQC 0.11.9 (https: / / www.bioinformatics.babraham.ac.uk / projects / fastqc / ).
[0173] CNV. tumor content and variant calling
[0174] CNV and tumor content were estimated from the recalibrated BAM files with Sequenza 3.0.0 workflow (Favero, F. etal., Annals of oncology: official journal of the European Society for Medical Oncology / ESMO 26, 64-70 (2015)), using matched-cancer type CNV frequencies Docket No.: 084276.00417 / LUD 6239 derived from PCAWG (Consortium, I.T.P.-C.A.o.W.G., Nature 578, 82-93 (2020)) as weights to improve estimations. Four variant calling algorithms (Mutect 1.1.7 (Cibulskis, K. et al. Nature biotechnology 31, 213-219 (2013)), Varscan 2.4.6 (Koboldt, D.C. etal. Genome research 22, 568- 576 (2012)), and GATK 4.4.0.0 (Van der Auwera, G.A. et al. Curr Protoc Bioinformatics 43, 11 10 11-11 10 33 (2013)) tools HaplotypeCaller (Cibulskis, K. et al. Nature biotechnology 31, 213- 219 (2013)) and Mutect2) were applied to the recalibrated BAM files.
[0175] The GATK HaplotypeCaller algorithm reduces the global false-positive variant call rate by relying on haplotypes de-novo assembly in variable regions (McKenna, A. et al. Genome research 20, 1297-1303 (2010); DePristo, M.A. et al. Nature genetics 43, 491-498 (2011)). HaplotypeCaller was run in GVCF mode on tumor and germline recalibrated BAM file separately to detect SNV and Indel variants. Tumor and germline gVCFs were combined using GATK GenotypeGVCF to produce a single VCF. GATK best practices were followed for subsequent variant quality score recalibration and variants quality assessment. Patient-specific SNPs were defined as variants present in both tumor and germline, while tumor-specific variants were defined as somatic mutations.
[0176] The MuTect 1.1.7 variant calling algorithm predicts somatic mutations based on log odds scores of two Bayesian classifiers. One classifier identifies non-reference variants in the tumor sample and the second classifier determines tumor-specificity of those variants. Read support is used for the filtering of candidate somatic mutations, reducing NGS artifacts. Mutect 1.1.7 ran with default parameters and selected SMs were exported in VCF format.
[0177] The Mutect V2 variant calling algorithm also relies on a Bayesian somatic genotyping model, but this model is different from Mutect VI and it uses de-novo assembly of haplotypes in variable regions. Mutect V2 ran with default parameters and identified variants (SNPs and SMs) were exported in VCF.
[0178] The VarScan2 algorithm relies on hard filtering of calls, making it less sensitive to bias such as extreme read coverage or sample contamination. Variants were filtered on parameters which include read quality, strand bias, coverage, and frequency. VarScan2 ran in somatic mode with custom parameters (-normal-purity 1.00, -min-normal-coverage 6, -min-tumor-coverage 8, — min-tumor-coverage 8, — somatic-p-value 0.05, -strand-filter 0), and the -tumor-purity and - min-var-freq were set based on the tumor content estimation, as the tumor-content estimation and Docket No.: 084276.00417 / LUD 6239 min (0.4 x tumor _content, 0.2) , respectively. Resulting variant (SMs and SNPs) calls were further filtered using Varscan2 fpfilter (— dream3-settings — min-var-basequal 25 — min-ref-avgrl 70 — min-var-avgrl 70 -keep-failures) and exported in VCF format. Finally, only variants passing fpfilter were retained.
[0179] Variant calls from GATK, MuTect vl, MuTect v2, and VarScan 2 were combined into a single VCF that contains the union of the variants of all four callers. Ambiguous calls (z.e., different calls at the same genomic coordinate) were resolved by a majority rule. In the absence of the majority, calls were rejected. High-confidence variants were defined as variants identified by two or more algorithms, and low-confidence variants were defined variants as identified by a single algorithm. Alternatively, when NeoDisc was operated in “sensitive” mode, the union of all variants identified by all variant calling algorithms, was considered, and annotated as high confidence. All identified variants were phased, using Whatshap 1.7 (Martin, M., et al. Methods in molecular biology 2590, 127-138 (2023)) to determine their detection on the same chromosome. This information was critical in determining whether neighboring variants affected the same transcript and necessitated their inclusion within the same peptide sequence. When gene-panel variants were used, a generic VCF was created and used for downstream analysis.
[0180] Clonality analysis
[0181] Only high-confidence SM were considered. The variant allele frequency (vaf ) was calculated for all high-confidence SMs detected in all samples from the bam files using samtools vl.17 mpileup (Li, H. et al. Bioinformatics 25, 2078-2079 (2009). CCF, clonal and sub-clonal assessments were calculated from the VAF, CNVs, and tumor content, as described in Letouze, E. etal. (Letouze, E. etal. Nature communications 8, 1315 (2017)). Zygosity annotation was derived from Sequenza analysis, where segments with one loss were annotated as LOH, and other segments were annotated as heterozygous (HET). LICHeE 1.0 (Popic, V. et al. Genome Biol 16, 91 (2015)) was used for phylogenetic reconstruction of cancer evolution, using all SMs identified. Trees were annotated with clonality information and with mutated genes annotated as drivers, based on IntoGen 2016.5 (Martinez-Jimenez, F. et al. Cancer 20, 555-572 (2020)).
[0182] RNA reads alignment and gene expression quantification
[0183] FastQ reads were aligned and counted per gene with STAR 2.7.10b (Dobin, A. et al. Bioinformatics 29, 15-21 (2013)) (— runMode alignReads, — twopassMode Basic, — Docket No.: 084276.00417 / LUD 6239 outReadsUnmapped Fastx, — quantMode GeneCounts, — outSAMtype BAM Unsorted), on the reference human genome (GRCh37) and the GENCODE v43 annotation. Resulting bam files were sorted and indexed with samtools 1.17 (sort, index) (Li, H. et al. Bioinformatics 25, 2078-2079 (2009)). RSeQC 5.0.1 (infer_experiment.py) (Wang, L., Wang, S. & Li, Bioinformatics 28, 2184- 2185 (2012)) was used to determine RNAseq data strandness, if the sequencing kit used was not known or not available in the list of pre-defined kits. TPM, RPKM, CPM, and UQN values were calculated from raw gene counts.
[0184] HLA Typing and HLA LOH analysis
[0185] HLA typing was performed on both DNA and RNA data with HLA-HD 1.7.0 (Kawaguchi, S., et al. Human mutation 38, 788-797 (2017)), that required paired-end sequencing data. HLA specific reads alignments derived from HLA-HD were used to derive HLA segments specific coverage. RNA coverage of each segment was calculated as the number of reads mapping uniquely to the segment. HLA allele-specific copy numbers were estimated following equations described in Van Loo et al. (Van Loo, P. et al. Proceedings of the National Academy of Sciences of the United States of America 107, 16910-16915 (2010)) and McGranaham et al. (McGranahan, N. et al. Cell 171, 1259-1271 el211 (2017)). LogR was calculated as:
[0186] Where Tumcovand Glcovare the tumor and germline segment specific coverages, respectively and M the multiplication factor, defined as the genome-wide tumor / germline ratio of uniquely mapped reads.
[0187] For each HLA allele pair, the segment baf was calculated as the coverage of the 1stallele segment divided by the sum of the coverage of the two alleles.
[0188] The segment copy number CN was calculated as:
[0189] > (p - 1) + baf x 2loaRx (2(1 - p) + p<p)
[0190] FA allele 1 —
[0191] P where p is the estimated tumor content and <p is the estimated tumor ploidy. Docket No.: 084276.00417 / LUD 6239
[0192] For heterozygous HLA alleles, p-values supporting loss events were calculated using paired t-test on CN, LogR, and RNAseq coverage of segments with the matched allele and as unpaired t-test on CN with segments having equal numbers of major and minor alleles, as determined in Sequenza segments file. A final pvalue supporting minor allele loss LOHpvcawas calculated by combining pvalues supporting loss events, using the Fisher method. Pvalues supporting allele presence were calculated using paired t-test on CN, LogR, and RNAseq coverage of segments with the matched allele and as unpaired t-test on CN with segments having lost the minor allele, as determined in Sequenza segments file. A final p-value supporting allele presence presencepvalwas calculated by combining pvalues supporting presence events, using the Fisher method.
[0193] A sample-specific naive Bayes classifier (A) was trained to predict CNaueie, as defined from Sequenza’ s sample-specific segments, from the tumor-germline coverage depth ratio and allele frequency at heterozygous sites. Model (A) was applied to the sample-specific heterozygous HLA alleles segments, and the CNaiieiepredictedwas defined as the segments median predicted CNs.
[0194] Finally, LOH events on heterozygous alleles were called if:
[0195] LOHpvai< 0.01
[0196] [and]
[0197] LO Hpvaipresencepvai
[0198] A second sample-specific naive Bayes classifier (B) was trained to predict CNtotaas defined from Sequenza’ s sample-specific segments, from the tumor-germline coverage depth ratio at all sites. Model (B) was applied to the sample-specific homozygous HLA alleles segments, and the CNaueiewas defined as the median predicted CN of its segments. The CNaueie confidencescore was calculated as the product of individual CNalieiesegments probabilities. LOH events were Docket No.: 084276.00417 / LUD 6239 defined for alleles for which CNallele= 0 [and] CNalleleconfidence< 0.01. The CNallelef nalwas set to 0 for LOH events and to max (1, round ,CNaueie')').
[0199] Both naive Bayes classifiers (A and B) were trained using 80% of the data used for training and 20% for testing, random_state was set to 312, priors were set to None, ski earn GridSearchCV algorithm was used to define the var_smoothing parameter using the following grid search parameters: param_grid= [‘var_smoothing’ : np.logspace(0,-9, num=50)], scoring=‘fl_weighted’ and refit = True.
[0200] A consensus HLA haplotype was built from the results of the germline and tumor typing from DNA and RNA sequencing data. Lost HLA-I and -II alleles are flagged by NeoDisc. By default, lost alleles are kept for downstream tumor-specific antigen prediction and prioritization. However, users have the option to choose whether to keep or discard these alleles via the command line. It’s important to note that this choice does not impact immunopeptidomic analysis. In this study, lost alleles were discarded from the predictions, where the effect of keeping or discarding lost HLA-I alleles on the prioritization was assessed.
[0201] Non-canonical genes
[0202] Non-canonical genes were defined as expressed genes annotated as lincRNA, IncRNA, miRNA, misc_RNA, Mt_rRNA, Mt_tRNA, processed_pseudogene, processed_transcript, pseudogene, retained_intron, rRNA, rRNA_pseudogene, scRNA, sense_intronic, sense_overlapping, snoRNA, snRNA, translated_processed_pseudogene, unprocessed_pseudogene, vault_RNA in GENCODE 38. Non-canonical genes were translated in 3 ORFs, from Methionine to stop and were added to the personalized references.
[0203] TAAs from canonical and non-canonical sources, and HC-TSAs
[0204] In general, gene expression level of >0.0 TPM was considered expressed in a given sample. Canonical and non-canonical genes (see definition above) with expression across all GTEx tissues, except testis and sun-exposed skin, with < 1.0 TPM at 90thpercentile was considered as TAAs. In addition, a list of immunogenic HC-TSAs was curated by mining the literature, IEDB (https : / / www.iedb.org / ) (Vita, R. et al. The Immune Epitope Database (IEDB): 2018 update. Nucleic acids research 47, D339-D343 (2019)) and CAPED
[0205] (https: / / caped.icp.ucl.ac.be / Peptide / list) databases. HC-TSA source genes were defined based on Docket No.: 084276.00417 / LUD 6239
[0206] GTEx expression (< 1 .0 TPM at 90thpercentile in all tissues, except testis and sun exposed skin), having expression >1.0 TPM in the sample, except for CTAG1B which was considered expressed if >0.0 TPM. All 8-19mer peptide sequences derived from expressed HC-TSAs and canonical TAAs were predicted for binding to the patient HLA-I and HLA-II alleles using PRIME 2.0 (Gfeller, D. et al. Cell Syst 14, 72-83 e75 (2023)) and MixMHC2pred v2.2 (Racle, J. et al. Immunity, Volume 56, Issue 6, 13 June 2023, Pages 1359-1375. el3 (2023)), respectively. HC- TSA and canonical TAAs derived peptides matching other sequences in the personalized healthy proteome reference were removed (at the exception of sequences derived from other HC-TSAs and canonical TAAs). Only HC-TSA and canonical TAAs derived peptides with a predicted %Rank< 0.5 or present in the curated database of immunogenic HC-TSA peptides were retained for further prioritization (see below). This restriction aims to limit the number of prioritized HC-TSA and canonical TAAs, but it does not impact the detection of HC-TSA and TAA derived peptides in the immunopeptidomic data.
[0207] Detection of viruses
[0208] NCBI Virus Genomes Resource database (Hatcher, E.L. et al. Nucleic acids research 45, D482-D490 (2017)) was used to retrieve all human-specific (filtering the database by host=human) and fetch their annotated genomes and proteomes. Non-human virus genomes and proteomes were retrieved from the list of all available viruses, selecting entries associated with non-human hosts, complete reference sequence available, and available accession identifier. ART 2.5.8 (Huang, W ., et al. Bioinformatics 28, 593-594 (2012)) was used to simulate illumina reads (art_illumina -ss HSXt -p -1 150 -f 4 -m 200 -s 10) from the human-specific and non-human viral transcript sequences. STAR 2.7.8a (Dobin, A. et al. Bioinformatics 29, 15-21 (2013)) was used for the alignment and counting of the simulated reads on the human-specific virus genomes. The Viral_Score of each virus was calculated as:
[0209] Z RPKMvg
[0210] Viral_Score = NVg Docket No.: 084276.00417 / LUD 6239
[0211] Where Rvgcorresponds to the number of reads mapped to a viral gene, RTotthe total number of mapped reads, Lvgthe length of the viral gene in base pairs, and Nvgthe number of viral genes.
[0212] Specificity and sensitivity of the identifications were calculated by comparing viral scores of human-specific and non-human viral reads. A threshold on the Viral_Score was selected, where it reached >97% sensitivity and specificity, to separate high-confidence from low confidence viral identifications.
[0213] Reads not mapped to the human reference genome were aligned and counted using STAR 2.7.10b (Huang, W., el al. Bioinformatics 28, 593-594 (2012)) to the human-specific viruses genomes database. Viral_Score, TPM, RPKM, CPM, and UQN values for each viral gene were calculated from reads counts.
[0214] Prediction of viral-peptides
[0215] A list of high-confidence immunogenic viral peptides was created from the literature and IEDB (https: / / www.iedb.org / ) database (Vita, R. et al. The Immune Epitope Database (IEDB): 2018 update. Nucleic acids research 47, D339-D343 (2019)). High-confidence peptides were defined as peptides not mapping to the human proteome. All 8-19 mers derived from high- confidence identified virus sequences, as determined from RNAseq, were predicted for binding to the patient HLA-I and HLA-II alleles using PRIME v2.0 and MixMHC2pred v2.0, respectively. NeoDisc screened all possible binding peptides across the entire protein sequence and excluded any peptide sequence matching with a personalized healthy proteome, resulting in a list of tumorspecific viral peptides.
[0216] Creation of sample-specific proteome references
[0217] GENCODE v43 reference proteome was combined with the tumor-specific three ORF translation products from methionine to stop of non-canonical transcripts and the identified viral proteomes. It was then personalized by addition of all VCF variants (SNPs and SMs) and annotated with gene expression, CNVs and variants-related features. This sample-specific proteome was subsequently used for the prediction of neoantigens, by including all high-confidence non- synonymous SMs and phased high-confidence SNPs, and for the MS search, by including all high and low confidence variants (SNPs and SMs). Additionally, a personalized healthy proteome Docket No.: 084276.00417 / LUD 6239 reference was created by including all high-confidence non-synonymous SNPs, and it was used to filter out sequences that are not tumor-specific.
[0218] Neoantigen prediction
[0219] All 8-19mers predicted or MS-detected peptides covering the SM were first filtered to remove any sequences mapping to the healthy personalized proteome to assure tumor-specificity. If multiple transcripts were affected by the SM, all unique sequences were considered. 8-12 mers peptides were predicted for binding to the patient HLA-I using PRIME v2.0, and all 12-19 mers peptides were predicted for binding to the patient HLA-II using MixMHC2pred v2.0. Predicted peptides were annotated with CScape (Rogers, M.F., et al. Sci Rep 7, 11597 (2017)) prediction of oncogenicity, median expression in the matched healthy tissue in GTEx, median expression of the matched tumor type in TCGA, intogen version 2016.5 (Martinez-Jimenez, F. etal. Cancer 20, 555- 572 (2020)) gene and mutation driver information, and with peptide presentation information from ipMSDB (Muller, M., et al. Frontiers in immunology 8, 1367 (2017). For each mutation, gene expression was derived from RNAseq analysis, and the mutation expression was determined as the percentage of reads carrying the mutation in the BAM file, using samtools mpileup (Li, H. et al. Bioinformatics 25, 2078-2079 (2009)). Long peptide sequences were designed by sorting annotated short peptides sequences based on their predicted binding to HLA-I and -II, their predicted binding rank (low to high), the number of different alleles predicted to bind the short peptide (high to low), ipMSDB presentation value on HLA-I and on HLA-II (high to low), and on the length of the peptide sequence (low to high). Optimal long peptide sequence design started with the best short peptide sequence. The sequence was sequentially expanded with the next short peptide sequences until reaching the maximum specified length. Short peptide sequences leading to a length longer than the requested length were ignored. To prevent synthesis issues, when the addition of short peptides led to Glu, Gin or His-Pro dimer at N-terminus or His, Pro or Cys at C- terminus, the long peptide sequence was further extended until reaching another amino-acid compatible with peptide synthesis.
[0220] Prioritization of predicted tumor-specific antigens
[0221] When RNAseq sequencing data was available, HLA-I SNVs neoantigens were prioritized with the machine learning based prioritization algorithm (Muller, M. et al. Machine learning methods and harmonized datasets improve immunogenic neoantigen prediction. Immunity 56, Docket No.: 084276.00417 / LUD 6239
[0222] 2650-2663. November 14, 2023). Other predicted neoantigens (including FS), HC-TSAs, TAAs, viral -peptides and HLA-II bound neoantigens were prioritized using specific rule-based algorithms.
[0223] For the machine learning based prioritization, HLA-I SNVs neoantigens were further annotated with NetMHCpan v4.1 (Reynisson, B., et al. Nucleic acids research 48, W449-W454 (2020)) for HLA class I binding affinity prediction, NetMHCstabpan vl.Oa (Rasmussen, M. et al. Journal of immunology 197, 1517-1524 (2016))) for HLA class I binding stability, NetChop v3.1 (Nielsen, M., et al. Immunogenetics 57, 33-41 (2005)) for C-terminal proteasomal cleavage, NetCTLpan vl. l (Stranzl, T., et al. Immunogenetics 62, 357-368 (2010)) for recognition by the TAP transporter complex, differential agretopicity indexes (DAI) calculated as the log (neoantigen%rank)-log(WT%rank), and HLA-binding anchor position derived from MixMHCpred sequence motifs. Those features, in combination with MixMHCpred, PRIME, CScape, GTEx (Carithers, L.J. & Moore, H.M. Biopreservation and biobanking 13, 307-308 (2015)), TCGA (https: / / www.cancer.gov / ccg / research / genome-sequencing / tcga), RNAseq and ipMSDB derived features were used to apply the machine learning model and prioritize predicted HLA-I SNVs neoantigens.
[0224] The rule-based approach was used for HLA-II neoantigen. First, long and short minimal epitope neoantigen sequences, are sorted separately based on a series of features as shown in FIGS. 3A-C. Second, the sorted neoantigens are classified into five categories (‘A’, ‘B’, ‘C’, ‘D’, and finally ‘Reject’) while maintaining the sequence order from the initial sorting step. Similar rulebased approaches were implemented for prioritization of antigens from viruses and TAAs (including both HC-TSAs and canonical TSAs), while they rely on different features as indicated in FIGS. 3A-C. Only positive T cell assays in the Immune Epitope Database were considered for immunogenicity.
[0225] Antigen processing and presentation machinery analysis
[0226] A list of genes involved in HLA-I and HLA-II presentation machinery pathways was selected from the KEGG PATHWAY Database (Kanehisa, M., et al. Nucleic acids research 44, D457-462 (2016)). All genes expression were quantified in TCGA and GTEx. Genes with sample RNAseq expression lower than the expression in the matched GTEx tissue at the 5thpercentile were considered downregulated and upregulated if their expression was greater than the 95th Docket No.: 084276.00417 / LUD 6239 percentile expression in the matched GTEx tissue. Tn addition, genes were annotated with LOH and SM events identified from DNA data.
[0227] Inflammation score calculation
[0228] From RNAseq quantifications data, NeoDisc derived the T-cell inflammation score by employing the T-cell inflammation signature as described in Danaher et al. (Luke, J. J., et al. Clinical cancer research: an official journal of the American Association for Cancer Research 25, 3074-3083 (2019); Spranger, S. et al. Proceedings of the National Academy of Sciences of the United States of America 113, E7759-E7768 (2016)). Nevertheless, users may employ alternative tools, as immunogenicity score has no impact on the prioritization of antigen. This score was used to determine the T-cell inflammation status, which could be non-inflamed, intermediate inflamed, or T-cell inflamed. Furthermore, the sample inflammation signature was represented within the context of the matched cancer type from TCGA.
[0229] Tandem mass spectrometry (MS / MS) spectra conversion
[0230] Proteowizard release 3.0.22132 (Chambers, M.C. etal. Nature biotechnology 30, 918-920 (2012)) msconvert tool was used to convert immunopeptidomic DDA and DIA MS / MS spectra acquired on Thermo scientific instruments (.raw) to MGF and mzML.
[0231] Peptides identification with Comet and NewAnce
[0232] DDA MS / MS data was converted into mgf and was searched against the personalized MS searchable proteome with comet 2022.01 (Eng, J.K., Jahan, T.A. & Hoopmann, M.R. Proteomics 13, 22-24 (2013)). Default comet search high-high parameters were modified as follow: decoy _search = 1; peff_format = 5; precursor_tolerance_type = 0; isotope_error = 1; isotope_error = 1; max_variable_mods_in_peptide = 2; max_variable_mods_in_peptide = 2; sample_enzyme_number = 0 ; activation_method = HCD; digest_mass_range = 500.0 3000.0; max_precursor_charge = 4; spectrum_batch_size = 30000. The following variable modifications were included in the search: Methionin oxidation (A mass=15.9949); N-term acetyl (A mass=42.010565); cysteins carbamyl (A mass=43.005814), carbamidomethylation of cystein (A mass=57.021464) and cysteinylation (A mass=l 19.004099); glutamine to pyroglutamatic acid (A mass=-17.026549); glutamic acid to pyroglutamic acid (A mass=-l 8.010565). No fixed modifications were considered. Peptides lengths were adjusted to 8-15 amino acids, and to 8-22 Docket No.: 084276.00417 / LUD 6239 amino acids for HLA-I and HLA-II searches, respectively, allowed missed cleavage variable was set to 15 and 22 for HLA-I and HLA-II, respectively.
[0233] HLA-I and HLA-II comet identifications were filtered at a 3% group-specific FDR using NewAnce 1.7.5 (Chong, C. et al. Nature communications 11, 1293 (2020)), where protein-coding (PC), non-canonical (NC) and virus derived peptides were separated for FDR calculation. Frameshift (FS) derived peptides are FDR filtered together with NC. PSMs were merged and grouped into unique peptide sequences, mapped to the most likely protein group (PC > FS > NC > VIRUS) and neoantigens were flagged.
[0234] Peptides identification with FragPipe
[0235] DDA and DIA mzML files were searched against the personalized MS searchable proteome with FragPipe 20.0 (https: / / fragpipe.nesvilab.org / ) pipeline, which contains DIA Umpire (Tsou, C.C. et al. Nature methods 12, 258-264, 257 p following 264 (2015)), philosopher 5.0.0 (da Veiga Leprevost, F. et al. Nature methods 17, 869-870 (2020)), MSFragger 3.8 (Kong, A.T., Leprevost, F.V., Avtonomov, D.M., Mellacheruvu, D. & Nesvizhskii, A.I. MSFragger: ultrafast and comprehensive peptide identification in mass spectrometry-based proteomics. Nature methods 14, 513-520 (2017)) and lonQuant 1.9.8 (Yu, F., Haynes, et al. Molecular & cellular proteomics: MCP 20, 100077 (2021). The default Nonspecific-HLA and Nonspecific-HLA DIA parameters available in FragPipe were edited as follow: addition of variable modifications (see comet section); removal of fixed modifications; setting philosopher ion, psm and peptide FDR to 0.03 and activation of group-specific FDR by setting msfragger.group_variable=2 and adding to the philosopher report filter the flag —group. lonQuant was turned on for peptide quantification when running in DDA only mode. Peptides lengths were adjusted to 8-15 amino acids and to 8-22 amino acids for HLA-I and HLA-II searches, respectively. HLA-I and HLA-II mzML files were searched separately. PSMs were merged and grouped into unique peptide sequences, mapped to the most likely protein group (PC > FS > NC > VIRUS) and neoantigens were flagged.
[0236] Combining Comet-New Ance and FragPipe identifications
[0237] Comet-NewAnce and FragPipe peptide identifications were combined by retaining the union of the identification. All peptides mapped to non-protein-coding groups were searched with Leu / Ile replacement against the blast (Coordinators, N.R. Nucleic acids research 44, D7-19 (2016)) non-redundant human database (downloaded using blast-plus 2.11.0 (Camacho, C. et al. BMC Docket No.: 084276.00417 / LUD 6239
[0238] Bioinformatics 10, 421 (2009)) on 06.04.2023). All peptides were annotated with available genomic and transcriptomic information and protein-coding peptides mapping to the list of high- confidence immunogenic HC-TSAs were annotated with literature-derived immunogenicity information. All identified peptides were predicted for binding to the patient’s HLA-I and HLA-II HLA alleles (including alleles subjects to LOH) with PRIME v2.0 and MixMHC2pred v2.0, respectively. MoDec 1.2 (Racle, J. et al. Nature biotechnology (2019)) was used for HLA-I and HLA-II motif deconvolution of peptides mapping to protein-coding, others (non-canonical) and viral groups. The number of motifs detected was based on the AIC value reported by MoDec. Each motif was assigned to the patient’s HLA with the closest motif, determined as the allele with the smallest Kullback-Leibler divergence KLDabetween the deconvoluted motif position weight matrices PWM™j and PWMpt(m: modec, i: the aminoacid, b binding predictor - MixMHCpred and MixMHC2pred, p : position, I : peptide length for HLA-I and core length for HLA-II):
[0239] Tumor specific peptides were defined as peptides not assigned to the protein-coding group predicted to bind with a rank < 2. PSMs spectrum images were generated with PDV 1.8.0 (Li, K., et al. Bioinformatics 35, 1249-1251 (2019)) and the best fragmentation pattern, as determined by NewAnce 1.7.5 (Chong, C. etal. Nature communications 11, 1293 (2020)) was reported. Identified tumor-specific peptides were sorted based on their protein group, evidence of immunogenicity in the literature (only for HC-TSA), source gene expression and their predicted binding affinity.
[0240] Tools used for NeoDisc benchmarking pTuneos vl.0.0 (Zhou, C. et al. Genome medicine 11, 67 (2019)) was run in WES mode, using matched germline and tumor WES and tumor RNA sequencing data. Default settings and hgl9 genome assembly were used to ensure a fair comparison. Immunogenic peptide ranks were retrieved from the fmal_neo_model.tsv file. pVACtools (Hundal, J. etal. pVACtools: Cancer immunology research 8, 409-420 (2020);
[0241] Hundal, J. et al. Genome medicine 8, 11 (2016)) starts with a VCF file. To allow a fair comparison Docket No.: 084276.00417 / LUD 6239 on the same set of mutations, a pVACtools compatible VCF containing all the high-confidence somatic mutations used for NeoDisc predictions and annotated with mutation and gene expression values was created. The official documentation for the VCF preparation was followed, and pVACtools v4.0.8 was run on patients’ HLA-I alleles with the default settings (all prediction algorithms, -el 8,9,10,11,12). Immunogenic peptide ranks were retrieved from the all_epitopes. aggregated. tsv file.
[0242] Gartner et al. immunogenic mutations and peptides ranks were retrieved from the original publication (Gartner, J. J. et al. Nature Cancer 2, 563-574 (2021)).
[0243] For LOHHLA vl.1.7 (McGranahan, N. et al. Allele-Specific HLA Loss and Immune Escape in Lung Cancer Evolution. Cell 171, 1259-1271 el211 (2017)), germline and tumor WES fastq files, HLA-I alleles, tumor purity and ploidy as inferred by NeoDisc were used as input. Default settings were used, HLA CN estimations were retrieved from columns HLA type IcopyNum vithBAFBin and HLA lype2copyNum wilhBA Bin in
[0244] DNA.HLAlossPrediction_CI.xls file. HLA loss events were called for the loss allele if COpyNu nwithBAFBin 0-5 [dUCi] PV O-lUnique 0.01.
[0245] Sample collection
[0246] Samples used in this study were collected from patients (Table 3) screened for inclusion in phase 1 trials approved by the institutional regulatory committee at Lausanne University Hospital (Ethics Committee, University Hospital of Lausanne-CHUV): NSCLC-1 (NCT05195619), MEL- 1 (Stranzl, T„ et al. Immunogenetics 62, 357-368 (2010)), MEL-2 (NCT03475134) (Stranzl, T., et al. Immunogenetics 62, 357-368 (2010)), and CESC-1, NPC-1, MEL-3 and MEL-4 (NCT04643574). Data included in this study is comprised of both published and unpublished data.
[0247] Tumor cell line generation
[0248] Tumor cell lines were derived from fresh tumor fragments, which underwent digestion using a combination of collagenase type I (Sigma-Aldrich) and DNAse (F. Hoffmann-La Roche) in R10 medium (RPMI 1640 GlutaMAX (Gibco) supplemented with 10% fetal bovine serum (FBS) (Gibco), 100 IU mL-1 of penicillin, 100 pg mL-1 of streptomycin (Bio-Concept)) at 37°C for 90 minutes. The resulting cell suspension was harvested, passed through a cell strainer to eliminate any residual tissue fragments, and then seeded in R10 medium for incubation at 37°C Docket No.: 084276.00417 / LUD 6239 with 5% C02. The culture medium was replenished every 2-3 days, and cells were subcultured upon reaching 80% confluence. For subculturing, tumor cells were gently detached using Accutase (ThermoFisher Scientific), split, and RIO medium was completely refreshed. Cells were harvested, washed in PBS and pellets of 20 million cells were stored at -80°C for downstream assays.
[0249] Samples preparation, libraries preparation and sequencing parameters
[0250] DNA was extracted with the commercially available DNeasy Blood & Tissue Kit (Qiagen) according to the manufacturers’ protocols. RNA was extracted by phase separation using Trizol / chloroform and centrifugation at 12000 g for 15 min. Then the aqueous phase containing the RNA was collected, mixed with 70% ethanol and loaded on RNeasy mini spin column from the Total RNA isolation RNeasy Mini Kit (Qiagen). The rest of the protocol was performed according to manufacturer’s instructions (including DNase I (Qiagen) on-column digestion). Whole exome sequencing (WES) and RNA sequencing were performed at Microsynth (CESC-1, MEL-4 tissues) or at the Health 2030 Genome Center (NSLC-1, NPC-1, MEL-1, MEL-2, MEL-3, MEL-4 cell lines). WES libraries were prepared using the Agilent SureSelect XT Human All exome V7 kit (Microsynth), the Twist Human Core Exome kit in combination with the Twist Human Refseq kit (Genome Center, MEL-1 and MEL-2), the Twist comprehensive exome panel kit (Genome Center-NSLC-1, NPC-1, MEL-3, MEL-4 cell lines). RNAseq libraries were prepared using the Illumina Truseq stranded mRNA reagents. WES and RNAseq samples were sequenced at Microsynth on the NextSeq 500 / 550 system or at the Genome Center on the Illumina HiSeq 4000 (MEL-1, MEL-2) or Novaseq 6000 (NSLC-1, NPC-1, MEL-3, MEL-4 cell lines) systems.
[0251] Flow cytometry analysis and isolation of tumor cells by fluorescence-activated cell sorting
[0252] Tumor cells from patient MEL-4 were stained with anti-HLA-ABC (AF700) antibody (Biolegend, 311438), washed and resuspended in PBS 10% FBS for FACS acquisition. Samples were acquired on a 5-laser LSR-II flow cytometer (BD Biosciences). The BD FACSDiva software was used for data acquisition. FlowJo vl0.9.0 (BD Biosciences) was used for data analysis.
[0253] The isolation of tumor cells from patient MEL-4 was performed using a 5-laser MoFlo Astrios EQ flow cytometer (Beckman Coulter), following cells staining with HLA-ABC AF700 (311438, BioLegend) and viability dye DAPI 422801 (BioLegend). HLA-ABC low and HLA- ABC high tumor cells were sorted and then directly expanded in RPMI 10% FBS 1% P / S for further analysis. Docket No.: 084276.00417 / LUD 6239
[0254] Peptides purification
[0255] The HLA-I and HLA-II peptides were purified according to the previously established protocols (Chong, C. el al. Molecular& cellular proteomics: MCP 17, 533-548 (2018)). Anti HLA- I W6 / 32 and anti HLA-II HB145 monoclonal antibodies were purified from the supernatants of HB95 (ATCC® HB-95™) and HB 145 cells (ATCC® HB-145™) using protein-A sepharose 4B (Pro- A) beads (Invitrogen), and antibodies were then cross-linked to Pro-A beads. Cell pellets and snap-frozen tissue samples were lysed in lysis buffer containing PBS with 0.25% sodium deoxycholate (Sigma Aldrich), 0.2 mM iodoacetamide (Sigma Aldrich), 1 mM EDTA, a 1 :200 protease inhibitors cocktail (Sigma Aldrich), 1 mM phenylmethylsulfonylfluoride (Roche), and 1% octyl-beta-D glucopyranoside (Sigma Aldrich). Tissues were homogenized on ice in 3-5 short intervals of 5 s each using an Ultra Turrax homogenizer (IKA). Samples were left for Ih at 4 °C then centrifuged for 50 min at 4°C at 75,600 x g in a high-speed centrifuge (Beckman Coulter, Avanti JXN-26 Series, JA-25.50 rotor) for tissues or in a table-top centrifuge (Eppendorf) at 21000 g for cell lysates. HLA immuno-purification was performed using the Waters Positive Pressure- 96 Processor (Waters) and 96-well single-use filter micro-plates with 3 pm glass fibers and 25 pm polyethylene membranes (Agilent, 204495-100). The cleared lysates were first loaded on a plate with wells containing Pro-A beads for endogenous antibody depletion, then on a second plate with wells containing anti-pan HLA-I antibody-cross-linked beads, and lastly through the third plate containing anti-pan HLA-II antibody-cross-linked beads. The HLA-I and HLA-II crosslinked beads were then washed separately in their plates with varying concentrations of salts using the processor. Finally, the beads were washed twice with 2 mL of 20 mM Tris-HCl pH 8.
[0256] Sep-Pak tC18 100 mg Sorbent 96-well plates (Waters, ref no: 186002321) were used for the purification and concentration of HLA-I and HLA-II peptides. The C18 sorbents were conditioned with ImL 80% acetonitrile (ACN; Chemie Brunschwig) in 0.1% trifluoroacetic acid (TFA; Sigma Aldrich) then twice with ImL 0.1% TFA, and the HLA complexes and bound peptides were directly eluted from the filter plates with 1% TFA. After washing the Cl 8 sorbents with 2 mL of 0.1% TFA, HLA-I peptides were eluted with 25 or 28% ACN in 0.1% TFA, and HLA-II peptides were eluted with 32% ACN in 0.1% TFA. Recovered HLA-I and -II peptides were dried using vacuum centrifugation (Concentrator plus, Eppendorf) and stored at -20 °C.
[0257] Liquid chromatography and mass spectrometry (LC-MS) Docket No.: 084276.00417 / LUD 6239
[0258] Immunopeptides were re-suspended in 2% ACN and 0.1% FA (formic acid). iRT peptides (Biognosis, Schlieren, Switzerland) were spiked into in the samples (Biognosis) as recommended by vendor and analyzed by LC-MS / MS.
[0259] The LC-MS system consisted of an Easy-nLC 1200 (ThermoFisher scientific, Bremen, Germany) coupled to Q Exactive HFX mass spectrometer (ThermoFisher scientific, Bremen, Germany) and or to Eclipse tribrid mass spectrometer (ThermoFisher Scientific, San Jose, USA). HLAp were eluted on a 450 - 500 mm analytical column (~8 pm tip, 75 pm ID) packed with ReproSil-Pur C18 (1.9 pm particle size, 120 A pore size, Dr Maisch, GmbH) and separated at a flow rate of 250 nL / min as described (Chong, C. et al. olecular & cellular proteomics : MCP 17, 533-548 (2018)).
[0260] For DDA measurements, the top-20 most abundant precursor ions selection was performed on the Q Exactive HFX as described (Chong, C. et al. olecular & cellular proteomics: MCP 17, 533-548 (2018)). On the Eclipse, full MS spectra were acquired in the Orbitrap m / z = 300 to 1650) with a resolution = 120’000 (m / z = 200), maximum ion accumulation time and automatic gain control (AGC) were set to auto. MS / MS spectra were acquired on the 20 most abundant precursor ions with a resolution of 30,000 (m / z = 200) with an isolation window = 1.2 m / z and ionaccumulation time = 54 ms. The AGC was set to 5e4 ions, the dynamic exclusion was set to 20 s, and a normalized collision energy (NCE) of 30 was used for fragmentation. No fragmentation was performed for HLA-I peptides with assigned precursor ion charge states of four and above, and the peptide match option was enabled.
[0261] For Data-independent acquisition (DIA), DIA cycle of acquisition was performed on the HFX as described (Pak, H. et al. Molecular & cellular proteomics: MCP 20, 100080 (2021)). On the Eclipse tribrid MS, the cycle of acquisitions consists of a full MS scan from 300 to 1650 m / z (R = 120,000), ion accumulation time = 60 ms, normalized AGC = 250% and, 22 DIA MS / MS scans in the orbitrap. For each DIA MS / MS scan, a resolution of 30,000, a normalized AGC of 2000%, and a stepped normalized collision energy (27, 30, and 32) were used. The maximum ion accumulation time was set to auto, the fixed first mass was set to 200 m / z, and the overlap between consecutive MS / MS scans was 1 m z.
[0262] For parallel reaction monitoring (PRM), synthetic peptides were ordered from Thermo as crude. After re-suspension in 2% ACN in 0.1% FA, peptides were analyzed by pools of peptides Docket No.: 084276.00417 / LUD 6239
[0263] (3 in total) for charge state of 2+ and 3+ The cycle of acquisition consisted of a full MS scan from 300-1650 m / z (resolution = 60’000, ion accumulation time = 100 ms) and consecutive acquisition of targeted peptide MS / MS scans (isolation window = 1.2 m / z' resolution = 30’000, accumulation time = 54 ms) in the Q Exactive HFX.
[0264] Immunogenicity assessment of TIL cultures and isolation of tumor antigen-specific TILs
[0265] Peptides were synthesized in-house (Peptide and Tetramer Core Facility, UNIL / CHUV Lausanne) or ordered externally (ThermoFisher Scientific). Set of peptides were then tested for immunogenicity by interrogating both clinical TILs (NCT03475134 & NCT04643574) and inhouse generated NeoScreen TILs, using IFNy Enzyme-Linked ImmunoSpot (ELISpot), as previously described (Arnaud, M. et al. Nat. Biotechnol. 40, 656-660 (2021)). NeoScreen method for TIL expansion / cultures involves the exposure of small tumor fragments to engineered autologous B cells and the neoantigenic peptides in the presence of IL -2 for a period of a 12 to 21 days (pre-REP). This is followed by further REP of TILs with OKT3 (anti-CD3) in presence of an excess of irradiated feeder cells and of IL-2 for 14 days. TILs were challenged with specific peptides at IpM peptides in pre-coated 96-well ELISpot plates (Mabtech). Following 16 to 20hrs co-culture, cells were gently harvested from ELISpot plates, which were then developed according to the manufacturer’s instructions and counted with a Bioreader-6000-E (BioSys). Positive conditions were defined as those with an average number of spots higher than the counts of the negative control (TILs alone) plus 3 times the standard deviation of the negative. The background level of IFNy Spot Forming Unit per 106cells by the negative control was subtracted from that of antigen re-challenged TILs in cumulative figures. Harvested TILs retrieved from plates were centrifuged and stained to assess the up-regulation of the activation marker CD137 on CD8+T- cells by flow cytometry. The background levels of CD137 expression by the negative controls (TILs alone) were subtracted to that of antigen re-challenged TILs in cumulative figures.
[0266] For the isolation of tumor antigen-specific cells for patients MEL-1 and MEL-4, IL-2 deprived TILs were stimulated with tumor-specific peptides at 1 pM for 24 hrs at 37°C, 5% CO2 and reactive TILs were sorted based on CD137 upregulation. After incubation, TILs were collected, washed, and stained with anti-human CD3 (Beckman Coulter), -CD8 (BD Biosciences), -CD4 (BioLegend), -CD137 (Miltenyi) and Aqua viability dye (Thermo Fisher Scientific) for 30 min at 4°C. Antigen-specific CD8+and CD4+TILs were FACS sorted using a BD FACS Melody cell Docket No.: 084276.00417 / LUD 6239 sorter (BD Biosciences) and used for downstream bulk TCR sequencing. In-house bulk TCR sequencing analyses were performed as described previously (Genolet, R. et al. Cell Rep Methods 3, 100459 (2023)).
[0267] Antigen and tumor reactivity validation by TCR cloning
[0268] Reactivities of TCRa / p pairs from samples MEL-1 and MEL-4 were validated by cloning into recipient T cells. Jurkat cell line (TCR / CD3 Jurkat-luc cells (NF AT), Promega, in-house stably transduced with human CD8a / p and TCRa / p CRISPR-KO) and activated peripheral T cells were used respectively to validate antigen and tumor specificity, as previously described (Stranzl, T., et al. Immunogenetics 62, 357-368 (2010))).
[0269] Jurkat cells were electroporated using the Neon electroporation system (Thermo Fisher Scientific) and co-cultured with peptide-pulsed HLA-matched presenting cells followed by a bioluminescence assay. For the luciferase assay, 2xl04Jurkat cells were cocultured with 4* 104PHA-activated CD4+T cells (CD4 blasts) in the presence of IpM specific peptides in 96-well plates. After overnight incubation, the assay was performed using the Bio-Gio Luciferase Assay System (Promega). Luminescence was measured with a Spark Multimode Microplate Reader (Tecan). To assess antitumor-reactivity of TCRs, 105TCR RNA-electroporated cells were cocultured in a 1 : 1 ratio with 200 ng / ml IFNy pre-treated tumor cell lines. After overnight incubation T cells were recovered and the upregulation of CD137 was evaluated by staining with anti-CD137 (Miltenyi), anti-CD3 (Biolegend or BD Biosciences), anti-CD4 (BD Biosciences), anti-CD8 (BD Biosciences) and anti-mouse TCRP-constant (Thermo Fisher Scientific) and with Aqua viability dye (Thermo Fisher Scientific). Flow cytometer LSRFortessaTM (BD Bioscience) or IntelliCyt iQue® Screener PLUS (Bucher Biotec) were used for the fluorescence readout and analyzed with FlowJo vlO (TreeStar). For both Jurkat and activated PBMCs, the following experimental controls of TCR transfection were used: Mock (transfection with water) and a control TCR (irrelevant crossmatch of a TCRa and P chain from a private TCR library).
[0270] (a) Multispectral immunofluorescence staining
[0271] Multiplexed staining was performed on 4-micrometer formalin-fixed paraffin-embedded (FFPE) tissue sections on automated Ventana Discovery Ultra staining module (Ventana, Roche). Slides were placed on the staining module for deparaffinization, epitope retrieval (64 minutes at 98°C) and endogenous peroxidase quenching (Discovery Inhibitor, 8 minutes, Ventana). Multiplex Docket No.: 084276.00417 / LUD 6239 staining consists of multiple rounds of staining. Each round includes non-specific sites blocking (Discovery Goat IgG and Discovery Inhibitor, Ventana), primary antibody incubation, secondary HRP-labeled antibody incubation for 16 minutes (Discovery OmniMap anti-rabbit HRP (Ventana, # 760-4311) or anti-mouse HRP (Ventana, #760-4310), OPAL™ reactive fluorophore detection (Akoya Biosciences, Marlborough, MS, USA) that covalently label the primary epitope (incubation: 12 minutes) and then antibodies heat denaturation. The panel was optimized to look at the HLA expression on tumor and immune cells. Sequence of antibodies used in the multiplex with the associated OPAL are the following : 1st: rabbit anti-SOXlO antibody (1.2 pg / ml, Clone EP268, Cellmarque, Ihour, 37°C), OPAL690 ; 2d: mouse anti-human HLA-DR antibody (1 pg / ml, Dako, Ihour, RT), OPAL570; 3rd: mouse anti-CDl lc antibody (1.5pg / mL, clone 5D11, Cellmarque, DAKO, Ihour, 37°C) , OPAL520 ; 4th: mouse anti-human HLA-ABC antibody (0.1 pg / ml, Clone EMR8-5, Abeam, Ihour, TA), OPAL480 ; 5th: rabbit anti-CD8 antibody (1 pg / mL, clone SP16, Cellmarque, Ihour, 37°C), OPAL620; 6th: rabbit anti-CD3 antibody (1.5pg / mL, DAKO, Ihour, 37°C), OPAL780. Nuclei were visualized by a final incubation with Spectral DAPI (1 / 10, FP1490, Akoya Biosciences) for 12 minutes. Multiplex IF images were acquired on PhenoImager allowing a whole slide multispectral imaging acquisition (Akoya Biosciences). IF signal extractions from qptiff were performed enabling a per-cell analysis of IF markers of multiplex stained tissue sections and counting of every cell population.
[0272] Digital profiling Slides Preparation
[0273] FFPE tissue sections of 5-pm were baked at 60 °C for 1 hour. Slides were deparaffinized in 3 xylol baths of 5 min, then rehydrated in ethanol gradient from 100% EtOH, 2 baths of 5 min followed by 95% EtOH, 5 min. Slides were then washed in PBS IX.
[0274] For protein detection, antigen retrieval was performed in Citrate buffer pH 6.0 buffer at 100°C for 15 min at high pressure. Jar containing slides were then removed and put at RT for 25 minutes, followed by one IX TBST wash of 5 minutes. In a humidity chamber, tissues were cover with blocking buffer W (Nanostring) for one hour at RT, followed by 4°C overnight incubation of antibodies mix containing oligonucleotide-conjugated antibodies and fluorochrome-conjugated morphologic antibodies, AF532-SOX10 (SPM607) and AF647-CD3 (CD3-12). The day after, slides were washed in IX TBST three times 10 minutes. Slides were then covered with 4% PFA and incubated 30 min at RT in a closed humidity chamber. This incubation was followed by two Docket No.: 084276.00417 / LUD 6239
[0275] 5 minutes washes in IX TBST. Next, 500 nM Sytol3 (Nuclear staining) in Buffer W was applied to each section for 15 minutes at room temperature in closed humidity chamber. Slides were washed twice in fresh IX TBST then loaded on the GeoMx Digital Spatial Profiler (DSP).
[0276] For RNA detection, antigen retrieval was done in Tris-EDTA pH 9.0 buffer at 100 °C for 15 min at low pressure. Slides were first deep into hot water 10 secondes to be then deep into Tris- EDTA buffer. Cooker vent stayed open during the procedure to ensure low pressure and reach 100 °C. Slides were then washed in PBS IX, and incubated into proteinase K in PBS (lug / ml) for 15 min at 37°C and washed again in PBS IX. Tissues were post-fixed in 10% neutral-buffered formalin 5 min, washed 2 times 5 min in NBF stop buffer (0.1M Tris Base, 0.1M Glycine) and finally once in PBS IX. The mix of WTA probes was dropped on each section and covered with HybriSlip Hybridization Covers. Slides were then put for hybridization overnight at 37 °C in a Hyb EZ II hybridization oven (Advanced cell Diagnostics). The day after, HybriSlip covers were gently removed and 25-min stringent washes were performed twice in 50% formamide and 2X SSC at 37 °C. Tissues were washed for 5 min in 2* SSC, then blocked in Buffer W (Nanostring Technologies) for 30 min at room temperature in a humidity chamber. Next, 500 nM Sytol3 and antibodies targeting AF532-SOX10 and AF647-CD3 in Buffer W were applied to each section for 1 h at room temperature. Slides were washed twice in fresh 2* SSC, and then loaded on the GeoMx Digital Spatial Profiler (DSP).
[0277] For slides collection, entire slides were imaged at *20 magnification and morphologic markers were used to select Region Of Interest (ROI) either using circle or organic shapes. Automatic segmentation of ROI based on SOX 10+ markers was used to defined Area of Illumination (AOIs). This allowed to separate tumor cells (SOX10+) from other cells (SOX10-). For RNA experiment, SOX10 staining was too faint to allow a good segmentation. Therefore, a negative segmentation based on CD3 was used. AOIs were exposed to 385 nm light (UV), releasing the indexing oligonucleotides which were collected with a microcapillary and deposited in a 96-well plate for subsequent processing. The indexing oligonucleotides were dried down overnight and resuspended in 10 pl of DEPC-treated water.
[0278] For protein profiling, nCounter® platform was used as readout. Photocleaved oligonucleotides from the spatially-resolved compartments were hybridized to 4-color, 6-spot Docket No.: 084276.00417 / LUD 6239 optical barcodes in the nCounter® platform, enabling labels counts per compartment of the protein targets representing the antibodies to which the tags were originally conjugated.
[0279] For RNA profding, NGS sequencing platform was used as readout. Sequencing libraries were generated by PCR from the photo-released indexing oligonucleotides and AOI-specific Illumina adapter sequences, and unique i5 and i7 sample indices were added. Each PCR reaction used 4 pl of indexing oligonucleotides, 4 pl of indexing PCR primers, 2 pl of Nanostring 5X PCR Master Mix. Thermocycling conditions were 37 °C for 30 min, 50 °C for 10 min, 95 °C for 3 min; 18 cycles of 95 °C for 15 s, 65 °C for 1 min, 68 °C for 30 s; and 68 °C for 5 min. PCR reactions were pooled and purified twice using AMPure XP beads (Beckman Coulter, A63881), according to the manufacturer’s protocol. Pooled libraries were paired-sequenced at 2 * 27 base pairs and with the single-index workflow on an Illumina Novaseq instrument.
[0280] CD8 infiltration assessment
[0281] To characterize the tissue heterogeneity for a specific phenotype, the entire multispectral image was segmented into Regions of Interest (ROIs). ROIs were determined by dividing the image into tiles, each evaluated for cellular metrics. The size of each tile was set by an area-based criterion, typically 640,000 square microns. A shift of 2 (the ROIs partially overlap at their centers) was applied to increase the granularity of tile positioning.
[0282] By iteratively shifting the tile, the tissue area and CD8+cell count per tile were calculated. These densities were plotted on a graph, displaying the density of CD8+cells in tumor islets on the x-axis and in the Tumor Microenvironment (TME) on the y-axis. Each graph point reflects the CD8+density in a particular tile, illustrating the spatial distribution of CD8+cells within the samples. A threshold density of 21 CD8 cells per mm2(in either tumor islets or TME) was set to determine infiltrated, excluded or desert tumor islets.
[0283] GeoMx digital spatial profiling (DSP) analysis
[0284] For GeoMx DSP protein data analysis, RCC files from Nanostring’s nCounter were imported back into the GeoMx DSP instrument for QC using GeoMx DSP analysis suite version 2.4.2.2. Negative Ctrl IgG normalization method was selected following normalization assessments using available custom script from Nanostring. Linear mixed-effects models were used to test for Docket No.: 084276.00417 / LUD 6239 differential protein expression between groups with R package ImerTest v3.1-3. Median expression of B2M across samples was used to stratified ROI into B2M -low or -high.
[0285] For GeoMx DSP RNA data analysis, Novaseq-derived FASTQ files for each sample were compiled for each compartment using the bcl2fastq program of Illumina, and then demultiplexed and converted to digital count conversion (DCC) files using the GeoMx DnD pipeline version 2.0.0.16 of NanoString according to manufacturer’s pipeline. QC and data analysis were performed using the R package, GeoMxTools v3.4.0. Probes were checked for outlier status. A probe was removed if the geometric mean of that probe’s counts from all segments divided by the geometric mean of all probe counts representing the target from all segments was less than 0.1 or if the probe was an outlier according to the Grubb’s test in at least 20% of segments. The counts for all remaining probes for a given target were then collapsed into a single metric by taking the geometric mean of probe counts. Segments and genes were filtered out with abnormally low signal. Segments with less than 1% of gene and genes with less than 5% of segments detected and below the limit of quantification (LOQ) set at 2 geometric standard deviations above the geometric mean were removed from the study.
[0286] Gene counts were normalized to the geometric mean of the 75thpercentile across all AOIs to give the upper quartile (Q3) normalization factors for each AOI.
[0287] Differential gene expression between groups of segments was assessed using Linear mixed-effects models (tool: mixedModelDE). Gene set enrichment analysis was performed with the R package, ClusterProfiler v4.8.1.
[0288] EXAMPLE 2
[0289] General overview of the NeoDisc pipeline
[0290] NeoDisc is a dedicated computational framework that combines genomic, transcriptomic and immunopeptidomic data for the prediction and direct identification of clinically relevant antigenic peptides presented exclusively on cancer cells, derived from multiple sources, including mutations, TSAs, TAAs, oncoviral and non-canonical transcripts (FIG. 1, FIG. 2 and Table 1). The framework integrates curated public databases of known immunogenic TSAs and viral antigens (Vita, R. et al. Nucleic acids research 47, D339-D343 (2019)) with machine learning and rulebased models for the prioritization and selection of clinically relevant antigens (FIGS. 3A-C). Docket No.: 084276.00417 / LUD 6239
[0291] NeoDisc utilizes matched tumor and germline genomic (WES or WGS) data for samplespecific variant characterization, tumor content estimation, and the identification of CNVs and somatic mutations (SMs). While WGS data can be employed in NeoDisc, it has not been implemented in this study. Bulk RNAseq data is employed for human gene and SMs expression quantification and estimation of T-cell inflammation. Unmapped RNAseq reads are further used for viral infection identification and viral gene expression quantification, primarily targeting oncoviruses. NeoDisc creates personalized proteome references, where single nucleotide polymorphisms (SNPs), SMs, as well as non-canonical expressed transcripts are annotated and used for downstream HLA binding and immunogenicity prediction of antigenic peptides. Immunopeptidomic data are searched against the personalized proteome to identify naturally presented antigenic peptides. HLA-I and -II typing is derived from germline and tumor WES and from RNAseq data. Importantly, defects in the tumor APPM and HLA loss of heterozygosity (LOH) are identified and highlighted. A machine learning algorithm, specifically trained on a complex matrix of tens of features, prioritizes likely immunogenic HLA-I neoantigens (minimal epitopes) and longer mutated sequences and supports the design of RNA- and peptide-based vaccines. Beyond HLA-I restricted neoantigens, ranking of other classes of antigens exclusively expressed in the tumor, such as HLA-II neoantigens, TAAs and oncoviruses, was performed by rule-based approaches, adjusted based on a decision scheme considering a limited set of features (FIGS. 3A-C ). NeoDisc supports data integration of multiple samples from the same patient providing insights into tumor heterogeneity and evolution (FIGS. 1 and 2).
[0292] NeoDisc runs on a Linux system and is packed in a Singularity container (Kurtzer, G.M., Sochat, V. & Bauer, M.W. PloS one 12, e0177459 (2017)). The runtime is 15.78h (+ / - 9.52h) (100G RAM, 24 CPUs) for a patient’s dataset that includes matched germline-tumor WES, tumor RNAseq and data dependent and independent acquisition (DDA and DI A, respectively) of immunopeptidome MS data. The runtime is determined by the sample’s sequencing depth, immunopeptidome depth and mutational load. Finally, NeoDisc generates processed genomic, transcriptomic and immunopeptidomic data, detailed characterization of the APPM, prioritized lists of tumor-specific antigens, and sample-specific reports, ensuring data traceability (FIGS. 1 and 2).
[0293] Detection of immunogenic tumor-reactive antigens from cancer-specific sources Docket No.: 084276.00417 / LUD 6239
[0294] NeoDisc applies four variant calling algorithms to WES / WGS data. Variants detected by only one caller are deemed low-confidence, whereas those identified by two or more callers are considered as high-confidence. In default settings, low and high confidence variants are included in the personalized proteome for immunopeptidomic searches, while only high confidence variants are considered for in-silico neoantigen predictions, ensuring a robust selection of predicted neoantigens. Highly mutated tumors tend to respond better to immunotherapy because they present more neoantigens. Yet, the selection of immunogenic neoantigens among numerous possibilities is challenging. Initially, rule-based approaches were applied for the prioritization of HLA-I and HLA-II restricted neoantigens and other antigenic peptides, exclusively expressed in the tumor (FIGS. 3A-C). Recently, large datasets of neoantigens in tumors of 112 patients (called NCI cohort) were systematically screened by Gartner et al. for autologous T cell responses, allowing training of machine learning tools for the prioritization. Minigenes expressing almost all the called mutations were transcribed in vitro and transfected into autologous antigens presenting cells (APCs) followed by a co-culture with TIL cultures and interferon (IFN)-g enzyme-linked immunospot (ELISpot) immunogenicity measurement. In most cases, additional immunogenicity screens were performed to identify the optimal epitopes and their HLA restrictions. Previously, machine learning classifiers were trained on a fraction of this dataset (NCI-train) and tested their performance on the remaining samples (NCI-test), as well as additional cohorts. With this machine learning tool, NeoDisc effectively prioritizes immunogenic neoantigens. The performance of NeoDisc’s rule-based and machine learning ranking approaches was evaluated against existing tools such as pTuneos (Zhou, C. et al. Genome medicine 11, 67 (2019)), pVACtools, and the machine learning tool introduced by Gartner et al. (Gartner, J. J. et al. Nature Cancer 2, 563-574 (2021)). NeoDisc machine learning prioritization algorithm surpassed all the evaluated tools (Table 2). Importantly, the rule-based approach was superior to pTuneos and pVACseq, motivating its implementation for ranking HLA-II neoantigens.
[0295] The NeoDisc’s efficient prioritization was illustrated on a cervix adenocarcinoma (CESC 1) characterized by an exceptionally high mutational burden (25 SM / Mb) (Table 3). Out of the 393 identified actionable mutations (i.e., non-synonymous in coding regions), representing a pool of 19’051 peptides with a predicted binding rank <2%, 66 HLA-I neoantigenic short peptides (minimal epitopes) were initially selected by the rule-based prioritization for T cell screening of autologous tumor-infiltrating lymphocytes (TILs) by IFNv ELISpot. 11 out of 66 peptides were Docket No.: 084276.00417 / LUD 6239 immunogenic, including two that ranked among the top-10 candidates. Reordering of the tested neoantigens with the NeoDisc machine learning model resulted in an impressive ranking of six immunogenic peptides within the top-10 (Table 3). Furthermore, NeoDisc successfully identified two confirmed immunogenic neoantigens in the CESC-1 tumor MS immunopeptidomic data, underscoring the potency of immunopeptidomics for pinpointing relevant targets. Importantly, the machine learning algorithm ranked the two immunogenic MS-detected neoantigens at positions 4 and 5 along with additional neoantigens in the top-50 candidates, which were ultimately not tested due to the limited T-cell availability. For comparison, they were ranked at positions 18 and 39 by pVACseq, and 230 and 667 with pTuneos (Table 3).
[0296] In viral-induced cancers, targeting viral antigens becomes an attractive approach as viral proteins often drive cancer initiation and progression. NeoDisc aligns unmapped RNAseq reads to a database of human viruses (from NCBI (Sayers, E.W. etal. Nucleic acids research 50, D20-D26 (2022))) to identify viral infections and consequently reports in-silico predicted HLA-I and -II bound peptides originating from actively expressed oncoviruses. Viral peptides are prioritized based on their source gene expression, evidence of immunogenicity in the Immune Epitope Database (IEDB) (Vita, R. etal. Nucleic acids research 47, D339-D343 (2019)) and their binding prediction, increasing their likelihood of eliciting an immune response (FIGS. 3A-C). NeoDisc successfully detected HPV18 infection in CESC-1. Out of the 1542 predicted (binding rank <2%) HPV18 HLA-I peptides, the third ranked peptide KLPDLCTEL (SEQ ID NO: 1) which derived from the highly expressed (34.045 TPM) E6 oncoprotein, and was confirmed immunogenic in IEDB, induced a spontaneous CD8+T cell response (Table 4).
[0297] In a patient diagnosed with epidermoid nasopharyngeal carcinoma (NPC-1), NeoDisc detected evidence of EBV infection in all five tumor samples. EB V was characterized by a long double-stranded DNA genome encoding a wide range of proteins and non-coding RNAs. Consistently high expression levels of RPMS1 across all samples were observed, in line with previous studies reporting higher RPMS1 circRNA expression in metastatic NPC (Liu, Q., et al. Cancer Manag Res 11, 8023-8031 (2019)). Expression of both latent and lytic genes revealed an active replication of EBV. T cell reactivity was assayed in TIL cultures collected at the end of the 1st pre rapid expansion protocol (preREP) and 2nd (REP) expansion phase. Remarkably, out of the 108 viral HLA-I and -II peptides tested in T-cell assays, 34 were immunogenic across the patient’s HLA alleles, of which 15 peptides were immunogenic in both preREP and REP Docket No.: 084276.00417 / LUD 6239 conditions. 16 and 3 peptides were found immunogenic only in the preREP and REP cultures, respectively, supporting previous studies demonstrating how REP protocol induces functional and phenotypic changes in TILs that may limit their responsiveness to REP (Li, Y. et al. Journal of immunology 184, 452-465 (2010); Parkhurst, M R. et al. Cancer discovery 9, 1022-1035 (2019)). 19 epitopes were reported immunogenic in IEDB, including six with the same HLA specificity as in NPC-1. Immunogenic peptides originated from genes associated with the latent (21 peptides) and lytic (13 peptides) phases of EBV. Moreover, while CD8+T cells responses were more prevalent, strong CD4+T cell responses were also observed against peptides derived from both the latent and lytic phase associated genes BARF1 and BALF2, respectively.
[0298] TAAs are minimally expressed in healthy tissues, except for the testis which is considered as an immune-privileged site. While some TAAs induce spontaneous responses in patients, they remain a concern for safety. To mitigate this risk, a list of high-confidence TSAs (HC-TSAs) that have no or extremely low evidence of expression in healthy tissues from the Genotype-Tissue Expression (GTEx) database (Consortium, G.T. Nature genetics 45, 580-585 (2013)) were curated, with the exception of testis and sun-exposed skin. NeoDisc systematically selects expressed HC- TSAs and ranks MS-detected or m-silico predicted HC-TSA derived peptides based on their binding affinity to the patient’s HLAs, their global HLA presentation hotspot coverage captured in the immunopeptidomics MS Database ipMSDB, comprising hundreds of healthy tissues, cell lines and tumor samples (Muller, M., et al. Frontiers in immunology 8, 1367 (2017)), and their annotated immunogenicity in IEDB (FIGS. 3A-C). Importantly, all non-unique peptides encoded by commonly expressed proteins are excluded. In a melanoma cell line derived from patient MEL- 1, NeoDisc identified 19 expressed HC-TSA genes. 161 HLA-bound peptides derived from 18 HC-TSAs were identified by MS, and among these, the immunogenicity of six HLA-I peptides by T-cell screening was confirmed. Following the cloning of antigen-specific T-cell receptor (TCR) candidates into recipient activated peripheral T cells, four HC-TSA peptides were confirmed as tumor rejection antigens (Table 3). A positive association for HC-TSAs between their coverage in ipMSDB and their detection by MS, immunogenicity, and ability to induce T cell recognition in MEL-1 was found. Importantly, globally, a significantly higher ipMSDB coverage was observed for confirmed immunogenic peptides from HC-TSA in IEDB (HLA-I n=91, HLA-II n=42) compared to all predicted peptides across HLAs, from HC-TSAs that were not reported as being immunogenic (HLA-I n=14,721, HLA-II n=16,769). This aligns with the previous findings, Docket No.: 084276.00417 / LUD 6239 indicating that mutations predicted to be covered by at least one ‘exact’ wild-type peptide in ipMSDB, are five times more likely to induce spontaneous CD8+T cell responses compared with all other mutations. Overall, these results demonstrate that the methods implemented in NeoDisc accurately identify and prioritize immunogenic antigenic peptides from multiple sources (Table 4).
[0299] Personalized vaccine design
[0300] While the default NeoDisc settings exhibit good performance, biopsies with very low tumor content and low mutation burden may result in detection of insufficient number of actionable high confidence expressed mutations, resulting in a suboptimal vaccine. To address this limitation, NeoDisc offers two additional modes; (i) the ‘sensitive’ mode, that considers the union of mutations called by the four mutation calling tools, to be used when insufficient number of mutations are detected with the default setting, and (ii) the ‘panel’ mode where mutations listed in available diagnostic clinical gene panel (GP) are utilized as input, allowing the design of vaccines for patients lacking dedicated biopsies. It is however important to acknowledge that GPs often provide insufficient number of mutations leading to suboptimal lists of neoantigens, or potentially none at all. The performance of the sensitive and default settings with a subset of the NCI-test cohort data (n=20) was compared, in which at least one immunogenic neoantigen was validated. With the sensitive mode, more mutations per sample were detected, while both the Variant Allele Frequency (VAF) and the fraction of RNAseq reads supporting mutations decreased compared to the default mode, indicating a decrease in the specificity of identified variants. Nevertheless, all immunogenic mutations were identified in both modes, and it was demonstrated that the ranking of immunogenic mutations was comparable, but better with the default setting. The three operational modes were compared on samples derived from the same non-small-cell lung cancer patient (NSCLC-l). A tumor tissue biopsy (NSCLC-l-Tissue) with an estimated -4% tumor content, and a pleural effusion punction sample (NSCLC-l -PEC) with a challenging estimated tumor content of -1.5%, were analyzed with the default and the sensitive modes, and a GP (NSCLC-l-GP) report of five SMs derived from a historical diagnostic sample, analyzed with the panel mode. No actionable SMs with RNAseq support were detected in NSCLC-l-PEC with the default mode, compared to 36 in the sensitive mode. Similar increase in the number of expressed actionable mutations was found in the NSCLC-l-Tissue in which 15 actionable mutations supported by RNAseq reads were identified in the default and 34 in the sensitive modes. Employing the sensitive mode enabled the detection of all five panel mutations in the tissue, Docket No.: 084276.00417 / LUD 6239 whereas only three were detected with the default mode. Likewise, the default mode resulted in the detection of three panel mutations in the PEC sample, whereas none were identified with the default mode. Sample-specific mutations likely reflect false positive identifications and differences in the samples’ composition, as those were collected from different sites and at different timepoints. Because the false positive rate can be high in such low tumor content samples, RNAseq data are required to accurately prioritize the mutations. In conclusion, the three modes in NeoDisc allow for the analysis of various sample sources ensuring effective identification of expressed mutations.
[0301] In cancer vaccines, long sequences, encoded as tandem minigenes in mRNA vaccines or as synthetic peptides, are often favored over minimal short peptides. This preference is motivated by the efficient uptake and processing by APCs, resulting in the presentation of multiple minimal HLA-I and -II neoantigens. The machine learning tool implemented in NeoDisc ranks the mutations based on their potential immunogenicity. NeoDisc optimally designs long sequences by maximizing coverage of high-quality predicted HLA-I and -II neoantigens, considering binding affinities and the location of the wild-type sequences in presentation hotspots in ipMSDB (Muller, M., et al. Frontiers in immunology 8, 1367 (2017)). Mutations tend to deviate from the long peptide’s midpoint, and in HLA-II neoantigens, mutations are restricted to those located within the predicted binding core. Furthermore, concerning peptide-based vaccines, NeoDisc supports the adjustment of the N’- and C’ termini to comprise amino-acids compatible with peptide synthesis. Importantly, NeoDisc provides lists of short minimal epitopes for immune-monitoring and long peptides, ready for vaccine manufacturing, ranked based on their potential immunogenicity.
[0302] Considering deficiencies in the antigen presentation machinery
[0303] Certain cancer types are particularly prone to CNVs, resulting in imbalances in gene expression levels of oncogenes and tumor suppressors. CNVs can disrupt the APPM, one notable mechanism being the HLA LOH, which results in reduced diversity in the presented antigens. NeoDisc annotates the expression levels, CNVs, and SMs, affecting essential components of the HLA-I and HLA-II APPMs, also relative to their expression on matched healthy tissues and cancer types from GTEx and TCGA, respectively. To capture HLA LOH, NeoDisc combines samplespecific HLA typing with HLA copy-number information. Briefly, NeoDisc uses Sequenza (Favero, F. et al. Annals of oncology / ESMO 26, 64-70 (2015)) to estimate tumor content and CNVs, and HLA-HD (Kawaguchi, S., etal. Human mutation 38, 788-797 (2017)) for HLA typing Docket No.: 084276.00417 / LUD 6239 of the germline and tumor WES and RNAseq. HLA allele-specific copy numbers (CN) are assessed using two methods, the first relies on tumor content and differential coverage ratios, similar to McGranahan et al. (McGranahan, N. et al. Cell 171, 1259-1271 el211 (2017)), and the second employs Naive Bayes models trained on sample-specific tumor-germline depth ratio and allele frequencies. Results from both approaches are combined to infer HLA LOH events. The performance of NeoDisc HLA-I / II LOH identification was compared with LOHHLA (McGranahan, N. etal. Cell 171, 1259-1271 el211 (2017)), which is the gold standard for HLA-I LOH detection, and with Sequenza that does not perform HLA LOH analysis, but instead provides copy number estimation at HLA-I and HLA-II loci. First, on 68 samples of the NCI cohort that NeoDisc’s HLA copy-number estimates strongly correlated with both LOHHLA for HLA-I (pearson corr = 0.871, pval < 0.0001) and with Sequenza for HLA-I (pearson corr = 0.863, pval < 0.0001) was demonstrated. 25 of the 37 HLA-I LOH events identified in 17 patients were common between NeoDisc and LOHHLA and were in agreement with Sequenza’ s estimated minor allele copy-number. The estimated tumor content in samples was examined, where LOH events were detected by both NeoDisc and LOHHLA, as well as those uniquely identified by each tool. NeoDisc’s estimation seems to be relatively uninfluenced by the tumor content. A similar analysis on HLA-II alleles showed relatively good agreement between the NeoDisc and Sequenza on HLA- II (pearson corr = 0.751, pval < 0.0001). Next, in five NCI cohort patients with HLA-I LOH events, eight immunogenic neoantigens were reported, of which five were restricted to alleles impacted by the LOH. How excluding or retaining lost HLA-I alleles affects immunogenic neoantigens ranking with the machine learning prioritization was evaluated. Three neoantigens, which were predicted to bind preferentially the lost HLA alleles, were negatively impacted when those alleles were excluded from the prioritization, while the ranking of the remaining neoantigens, that were predicted to bind equally well other alleles, was marginally altered. Consequently, as HLA LOH events are typically sub-clonal, to support comprehensive immuno-monitoring investigations in the autologous settings of the heterogenic antigenic landscape, alleles impacted by HLA LOH are retained by default. However, for vaccine design, excluding HLA-I alleles with LOH is recommended. HLA-II alleles impacted by LOH may be retained, as presentation of HLA-II neoantigens in tumors is commonly mediated by professional APCs. Importantly, retaining or discarding HLA-I or HLA-II alleles with LOH for subsequent predictions is configurable in NeoDisc and does not impact their potential identification by immunopeptidomics. Docket No.: 084276.00417 / LUD 6239
[0304] As demonstrated for MEL-2 melanoma tissue, as a case of functional APPM, all ETLA-I alleles were present and expressed in MEL-2 and due to CNVs in chromosome 6, duplications of the HLA-A*26:01, B*38:01 and C* 12:03, and other HLA-II alleles were found. Interestingly, experimentally validated immunogenic HLA-I peptides were predicted to bind HLA-A*26:01 and C* 12:03 (Table 3). In MEL-2, all six HLA-I alleles were represented in the tumor immunopeptidome where 34 unique HC-TSA peptides were identified. Widespread presentation of HLA-I was confirmed by multiplexed immunofluorescence (mIF) staining in SOX10+melanoma cells.
[0305] In the case of MEL-3, from which three samples were collected from a tumor bearing right external iliac lymph node, NeoDisc revealed significant impairment of HLA expression and heterogenous gains and losses in chromosomal segments in all three samples. The detection of clonal copy number gains of BRAF, KRAS, and MYC oncogenes, along with the loss of one copy of TP53 and a p.Leul94Arg mutation in the remaining copy, highlighted alterations that are critical for tumor fitness and progression. In addition, LOH in the entire HLA-I and HLA-DPA1 locus was observed. Notably, HLA-I allele expression levels, obtained from bulk RNA-seq, exhibited a strong association with their respective CN, underscoring the benefit of HLA expression information in the calculation of HLA LOH. Of note, evaluating HLA-II expression is generally ineffective for determining allele presence or absence since cancer cells frequently lack HLA-II expression.
[0306] Furthermore, NeoDisc detected the LOH of additional components of the APPM in sample MEL-3-C, including CIITA and CALR. A remarkable finding was the identification of B2M LOH and frameshift mutation (p.Glu67fs) with a cancer cell fraction (CCF) of 0.88, 0.79 and 0.02 in samples MEL-3-A, MEL-3-B and MEL-3-C, respectively. B2M expression in MEL-3 samples correlated with the estimated CCF, was downregulated compared to the matched healthy tissues (GTEx), and in the lower range of expression in TCGA melanoma samples. It was hypothesized that the B2M p.Glu67fs mutation, coupled with the B2M and HLA LOH, would abolish HLA-I presentation in most of the cancer cells. mIF staining confirmed that most tumor cells (SOX10+) lacked cell surface HLA-I expression.
[0307] Addressing tumor neoantigen heterogeneity Docket No.: 084276.00417 / LUD 6239
[0308] NeoDisc provides a deeper understanding of tumor heterogeneity by consolidating data from multiple lesions. To illustrate real-world challenges, MEL-4 was examined, from whom two distinct melanoma metastases were surgically removed on the same day, from the gastric (GI) and the pelvic (P) regions, each with three tissue samples (named GI1-3 and Pl-3, respectively). NeoDisc generated a tumor phylogenetic tree annotated with known driver mutations from Intogen (Martinez- Jimenez, F. et al. Cancer 20, 555-572 (2020)), effectively differentiating the samples based on their respective regions. Evidence of B2M LOH across all six samples was found while Pl-3 samples carried in addition a B2M frameshift mutation p.Ser31fs, causing a premature stop codon, with an estimated CCF of 0.85, 0.92 and 0.95 in samples Pl, P2, P3, respectively. Low T cell inflammation, determined by the calculation of immune-related gene expression signatures, was associated with the combined effect of B2M p.Ser3 Ifs and B2M LOH. An exception was the P-1 sample, which exhibited the lowest B2M p.Ser31fs CCF, likely due to greater heterogeneity. mIF staining of CD8+T cells in both tumor and stroma niches confirmed the heterogenous exclusion of CD81T cells in the pelvic region. In addition, a quarter of the tumor cells (SOX101) lacked HLA-I and -II expression in the pelvic tissues. Furthermore, TAAs and HC-TSAs with supporting evidence of expression were identified, including MAGEA1, MLANA, and TYR. NeoDisc further identified by MS six mutated neoantigens, of which two were immunogenic (Table 4), and multiple peptides derived from expressed non-canonical sources such as IncRNAs, including MAGEA4-AS1, ATF6-DT, DSCR8, and a processed pseudogene (GAPDHP40), not expressed in GTEx. Of note, MAGEA4-AS1 and DSCR8 were previously associated with high expression in several tumor types. While the expression levels of the above sources remained consistent across all six samples, their presentation on HLA-I and -II were significantly lower and almost undetectable in samples containing the p.Ser31fs mutation (p value=0.0042). Generation of a primary cell line from the MEL-4 P-1 tissue confirmed the presence of two melanoma cells populations, with about 30% of cells expressing very low HLA-I, corroborating the above mIF staining. HLA-low and HLA-high autologous tumor cell lines were isolated and expanded independently. Notably, while IFNy treatment increased the expression of IFNy responsive genes (including CIITA, HLA-II alleles, and the immunoproteasome subunits PSMB-8, -9, and -10) in both HLA-high and HLA-low cells, tumor-specific antigen presentation increased only in HLA- high cells. Importantly, the treatment had minimal impact on the HLA-low cells, which lost the ability to present peptides by HLA-I and -II. In line with these results, neoantigen-specific TCRs, Docket No.: 084276.00417 / LUD 6239 targeting MS-detected as well as predicted neoantigens, effectively recognized HLA-high tumor cells only, although the mutations were also detected and expressed in HLA-low cells. In summary, NeoDisc revealed significant tumor-intrinsic heterogeneity which is expected to hinder the effectiveness of anti -tumor immunity.
[0309] NeoDisc’s report and output datasets
[0310] Meeting the criteria set by clinical regulations, NeoDisc generates a folder containing output files and tables, lists of all MS-identified peptides and expressed genes, the VCFs, the personalized reference proteome fasta, as well as a comprehensive sample-specific pdf report containing metadata on the samples and on the version of the NeoDisc pipeline and the analysis date. The report provides graphical representations of the results and quality measures to highlight technical issues, including information about variant calling algorithm performance, TMB, HLA typing, and HLA LOH, which are useful for identifying potential sample mix-up. CNVs are annotated to highlight affected oncogenes and tumor-suppressors, thereby enhancing understanding of tumor evolution. It also offers a visual representation of gene expression levels, facilitating efficient evaluation of APPM integrity, expression profiles of mutated genes in the context of GTEx and TCGA samples, T-cell inflammation status, and lists of identified viruses and expressed and presented HC-TSAs and TAAs. A comprehensive summary of the immunopeptidomic analysis is also provided, including the number of canonical and non- canonical peptides identified and the respective percentages of peptides predicted as binders to the respective HLAs, peptide length distribution and deconvolved binding motifs. Overall, the report highlights key details concisely, facilitating an immediate interpretation of the results. Additionally, it offers valuable insights to support additional in-depth exploration for translational research, along with links to relevant publications to contextualize the findings.
[0311] EXAMPLE 3
[0312] This example describes the methods employed to generate the results in relation to EXAMPLE 4.
[0313] Patient Samples, Cell Lines, and Cell Culture
[0314] Written informed consent was obtained from participants in accordance with the requirements of the institutional review board of the local Ethics Committee (Commission Docket No.: 084276.00417 / LUD 6239 cantonale d'ethique de la recherche sur 1'etre humain (CER-VD), Centre hospitaller universitai re Vaudois, CHUV). Samples Til, Ti2, and Ti3 were collected from patients during screening for inclusion in a Phase 1 clinical trial (NCT04643574) under a research protocol approved by the local Ethics Committee (BASEC ID 2017-00305). Tumor cell lines were established as described in Huber et al. (Huber F, et al. Nature Biotechnology. 2024; 42(1): 108-120). The MH and M12 cell lines, along with associated immunogenicity assessments and next-generation sequencing (NGS) data, were previously reported in Muller et al. (Muller M, et al. Immunity 56, 2650-2663 e2656 (2023)), where they correspond to patientl and patient3, respectively. The M13 cell line, including immunogenicity data, corresponds to Mel-4 Pl HLA-I high cells stimulated with 100 U / mL Interferon-y (Miltenyi Biotec, ref. 130-096-482) as disclosed in Huber etal. HLA-restricted tumor-associated antigen (TAA) immunogenicity assessments were derived from CEDAR (Kosaloglu-Yalcin Z, et al. Nucleic Acids Res 51, D845-D852 (2023)).
[0315] All cell lines were cultured in medium (Gibco, ref. 61870-010) supplemented with 10% heat-inactivated fetal bovine serum (FBS; Gibco, ref. 10437-028) and 100 U / mL penicillinstreptomycin (P / S; BioConcept, cat No. 4-01F00-H). For harvesting JY (ATCC, ref. 77441) and RA957 cells, the suspension culture was collected, the medium removed, and the cells washed twice with phosphate-buffered saline (PBS; Bichsel, ref. 100 0 324). For adherent melanoma cell lines, medium was removed, cells were washed with PBS, detached using Trypsin (BioConcept, ref. 5-51F00-H), and the reaction was inactivated with culture medium. Cells were then washed twice in PBS and pelleted (108cells / pellet) prior to storage at -80°C.
[0316] Immunopeptide Enrichment
[0317] W6 / 32 antibodies were harvested from HB-95 hybridoma cells (ATCC HB-95) and crosslinked with DTT to Sepharose beads (Invitrogen, ref. 101042). Cell and tissue samples were lysed using a buffer containing 0.25% sodium deoxycholate (Sigma-Aldrich, ref. 30970), 0.2 mM iodoacetamide (IAA; Sigma, ref. I6125-5g), 1 mM EDTA (ThermoFisher Scientific, ref. 15575- 038), 1 :200 protease inhibitor cocktail (PIC; Roche, ref. 04693132001), 1 mM phenylmethyl sulfonyl fluoride (Roche, ref. 10837091001), and 1% octyl-P-D-glucopyranoside (Sigma-Aldrich, ref. 08001) in PBS, incubated for 1 hour at 4°C. For tissue samples, homogenization was performed using a bead homogenizer (Tissue Lyser II, Qiagen) at 30 Hz for 1 minute. Lysates were centrifuged for 50 minutes at 20,000 RCF and 4°C. Docket No.: 084276.00417 / LUD 6239
[0318] The HLA-peptide complexes were isolated following the procedure described by Chong et al. (Chong C, etal. Molecular & cellular proteomics: MCP 17, 533-548 (2018)). Briefly, 96-well plates were prepared with 75 pL of cross-linked beads. Sample lysates were added to each well, followed by salt washes. Elution of HLA-peptide complexes was carried out using 1% trifluoroacetic acid (TFA; Merck Millipore, ref. 1082620100), and peptides were desalted using Cl 8 columns with 28% acetonitrile (ACN; Biosolve, UN 1648) and 0.1% TFA as the elution buffer. Eluted peptides were dried in a vacuum concentrator (Concentrator plus, Eppendorf) and stored at -80°C. Prior to mass spectrometry (MS) analysis, peptides were reconstituted in 30 pL (for cell lines) or 18 pL (for tissue samples) of 0.1% TFA. MS internal retention time (iRT) kit peptides (Biognosys, ref. Ki-3002-2) were spiked into the samples according to the manufacturer's instructions.
[0319] Immunopeptides were isolated from 108cells per sample for cell line experiments and from approximately 60 mg of tissue per sample for tissue experiments. Injection volumes corresponding to 5 x 106or 107cells (cell lines) or approximately 10 mg of tissue (tissue samples) were used. In dilution experiments, the RA957 immunopeptidome was diluted into the JY immunopeptidome at defined ratios (1 / 1024, 1 / 256, 1 / 64, and 1 / 16), including a negative control composed exclusively of JY peptides. Equal quantities of JY peptides were maintained across all samples.
[0320] Whole-Exome Sequencing (WES) and RNA Sequencing (RNAseq) Library Preparation and Sequencing
[0321] Library preparation for WES and RNAseq was performed as previously described in Huber et al. (Huber F, et al. Nature Biotechnology. 2024; 42(2): 199-211). Genomic DNA was extracted using the DNeasy Blood and Tissue Kit (Qiagen, ref. 69504), following the manufacturer’s protocol. Total RNA was isolated by phase separation using TRIzol reagent (ThermoFisher Scientific, ref. 15596-026) and chloroform (Sigma-Aldrich, ref. C2432-25ML), followed by centrifugation at 12,000 * g for 15 minutes. The resulting aqueous phase was mixed with 70% ethanol and loaded onto an RNeasy Mini spin column (Qiagen, ref. 74104), with subsequent on- column DNase I digestion (Qiagen, ref. 79254) performed according to the manufacturer’s instructions.
[0322] WES libraries were prepared using the Agilent SureSelect XT Human All Exon V7 kit (Agilent, ref. 5191-4028), and RNAseq libraries were constructed with the Illumina TruSeq Docket No.: 084276.00417 / LUD 6239
[0323] Stranded mRNA reagents (Illumina, ref. 20020594). All sequencing was conducted at Microsynth using the Illumina NextSeq 500 / 550 system.
[0324] NeoDisc Pipeline Parameters
[0325] NeoDisc vl.7.0 was utilized for neoantigen discovery and prioritization as previously described (Huber F, et al. Nature Biotechnology. 2024; 42(2): 199-211)). The pipeline was executed in fastq mode using default parameters for single-sample analysis, which included paired germline and tumor WES data as well as matched tumor RNAseq data. Peptide lengths were defined as 8-12 amino acids for HLA class I and 12-15 amino acids for HLA class II, with binding affinity thresholds set at <2.0% for both classes.
[0326] Following single-sample analyses, the MergeSamples tool within the NeoDisc framework was used to aggregate results across samples, enabling prioritization of tumor-specific antigens.
[0327] NeoDiscMS Workflow
[0328] The NeoDiscMS method required the generation of a predefined list of target peptides. Approximately 1,500 peptide sequences were tested in the present study. Target sequences were derived from diverse sources and selected based on immunogenic potential.
[0329] For the JY-RA957 dilution experiment, RA957-specific HLA-I-bound peptides were identified from previously published raw files (Chong C, et al. Molecular & cellular proteomics: MCP 17, 533-548 (2018)). MSFragger-DDA+ was employed to process the data, and peptides were filtered based on the following criteria: 8-11 amino acids in length, detected across all eight replicates, predicted to bind HLA-A*68:01 (MixMHCpred v2.3) with a binding rank <2%, and not predicted to bind any other JY or RA957 alleles with a rank <10%. A total of 1,525 peptides were selected.
[0330] For melanoma cell lines Mil, M12, and M13, 500 predicted neoantigens with a binding rank <3% were selected using the NeoDisc prioritization results. An additional 1,000 peptides derived from three-frame translation of TAA genes were included. These peptides were fdtered to retain 8-14mers with predicted binding rank <2% for at least two HLA-A or HLA-B alleles corresponding to the haplotype of the respective melanoma cell line. Docket No.: 084276.00417 / LUD 6239
[0331] For melanoma tissues Til, Ti2, and Ti3, the top 500 neoantigens as prioritized by NeoDisc were included. Additionally, 518 TAA-derived HLA class I peptides predicted by NeoDisc (Expressed! AAs table) were added.
[0332] Retention time (RT) prediction was performed to enable efficient inclusion list scheduling and improve peptide detection. Required inputs included RTs of calibration peptides, the target peptide list, and a prediction tool (e.g, DeepLC (Bouwmeester R, et al. Nat Methods 18, 1363- 1369 (2021))). RTs were derived from peptide spectral matches (PSMs) obtained from a prior HeLa tryptic digest (ThermoFisher Scientific, ref. 88329) run using the same LC gradient. FragPipe (v22.0) was used to generate the spectral library, with methionine oxidation included as a variable modification. RT prediction was performed using DeepLC (v3.1.1). All methionine oxidation variants of target peptides were included. Inclusion windows of ±15 minutes around the predicted RT were applied, and inclusion lists were generated in a format compatible with Thermo Xcalibur Instrument Setup (v2.0), with peptide charge states of 1-3.
[0333] For real-time search (RTS), dedicated FASTA files were created for each set of target peptide sequences. For melanoma cell lines and tissues, FASTA headers annotated the peptides as neoantigens or TAAs. The RTS acquisition was setup within Thermo Xcalibur Instrument Setup, including inclusion tree design and parameter placement.
[0334] RTS settings were standardized across all NeoDiscMS experiments, aside from the inclusion list and FASTA file used. A digestion enzyme input was required despite the direct use of target peptide sequences. Static modifications were disallowed, and methionine oxidation was the sole variable modification (up to three modifications per peptide). The maximum search time was set at 40 ms. RTSf thresholds were set as follows: Xcorr > 0.4, dCn > 0, and precursor mass error < 5 ppm for charge states 1 to 3.
[0335] Data-independent Acquisition (DIA) and Mass Spectrometry Methods
[0336] For all DIA experiments, mass spectrometric acquisition was performed with an MSI scan range of 300-1650 Th at a resolution of 120,000. A total of 27 DIA windows were applied, with sizes as follows: [37, 30, 24, 24, 22, 23, 44, 21 , 24, 24, 25, 27, 27, 30, 35, 38, 43, 53, 72, 103, 594], The automatic gain control (AGC) target was set to 2000%, and the injection time was adjusted automatically. MS2 scans were acquired at a resolution of 30,000 using stepped collision energies of 27, 30, and 32. Docket No.: 084276.00417 / LUD 6239
[0337] Acquisition cycles for DDA, ilDDA, and NeoDiscMS were constructed as described in FIG. 4A, with cycle durations of 3 seconds. For ilDDA and NeoDiscMS runs, the appropriate inclusion list or FASTA file was included in the acquisition method.
[0338] Liquid chromatography was performed using an Easy-nLC 1200 system coupled to an Orbitrap Eclipse Tribrid mass spectrometer (Thermo Fisher Scientific). Analytical columns of 450 mm in length, 75 pm inner diameter, and 8 pm tip size (TSP-075375, BGB Analytik) were packed with C18 particles (1.9 pm, 120 A pore size; Dr. Maisch, rl 19. aq). A binary gradient was applied using solvent A (0.1% formic acid, Thermo Scientific, prod. nr. 85178) and solvent B (80% acetonitrile, 0.1% formic acid). The gradient was executed over 125 minutes at a flow rate of 250 nL / min: 2-25% B over 0-1 10 min, 25-35% B over 1 10-114 min, 35-100% B over 114-115 min, and 100% B over 115-125 min.
[0339] To generate a retention time calibration run for DeepLC prediction, 100 ng of tryptic HeLa digest was injected using the aforementioned LC-MS gradient. A MSFragger-DDA database search was performed using the human SwissProt FASTA database (release 28.3.2024, 20,419 entries, excluding isoforms), supplemented with a list of common contaminants provided by FragPipe. Default parameters for tryptic digestion were applied.
[0340] Immunopeptidomics Searches with MSFragger-DDA and MSFragger-DDA+
[0341] For immunopeptidomics searches using MSFragger-DDA, database searches were conducted using the human SwissProt FASTA database (release 28.3.2024, 20,419 entries), in combination with iRT peptides and common contaminant sequences. Digestion was set to “unspecific,” and peptides between 8-14 amino acids in length were included. Variable modifications included methionine oxidation, N-terminal acetylation, and cysteine carbamidomethylation. Precursor charge states of 1-3 were permitted. False discovery rates (FDRs) were controlled at 1% for proteins and 0.01 for peptides, ions, and PSMs.
[0342] For MSFragger-DDA+ searches, distinct FASTA files were used depending on the experiment. All files were concatenated with iRT peptides and common contaminant sequences. For JY- and RA957-based experiments, the human SwissProt FASTA (no isoforms) was used. For melanoma cell lines and tissues, NeoDisc-generated sample-specific FASTAs were applied. All searches were performed using “unspecific” digestion and included 8-14mer peptides. Modifications and precursor charge settings mirrored those of DDA searches. DDA+-specific Docket No.: 084276.00417 / LUD 6239 parameters were left at default. FDRs were controlled at 1 % for proteins and at 0.01 or 0.03 (groupspecific) for peptides, ions, and PSMs, depending on the FASTA database used (Nesvizhskii Al. Nat Methods 11, 1114-1125 (2014); Ferreira HJ, et al. Nat Common 15, 2357 (2024)).
[0343] Peptide-Centric Searches and Library Generation
[0344] For visualization of elution profiles, raw files and target peptide sequences were imported into Skyline (v22.2) (MacLean B, et al. Bioinformatics 26, 966-968 (2010)).
[0345] For spectrum-centric searches using DIA-Umpire, raw DIA files from the melanoma tissue experiment were searched against NeoDisc-generated FASTAs using unspecific digestion, 8- 14mer peptides, and the same variable modifications described above. DIA-Umpire-specific settings were retained at default. Charge states of 1-3 were allowed for PSMs. Protein FDR was set to 1%, and peptide / ion / PSM FDR was set to 3% with group-specificity.
[0346] Hybrid spectral libraries were generated by combining results from spectrum-centric DIA searches and corresponding NeoDiscMS MSFragger-DDA+ searches (Til, Ti2, and Ti3). These hybrid libraries were used to perform peptide-centric searches with DIA-NN. The FDR threshold was set to 1%.
[0347] Data Analysis and Visualization
[0348] For immunopeptidomics data analysis, psm.tsv files (spectrum-centric searches) and report.tsv files (peptide-centric searches) were analyzed. Peptide binding affinity was predicted using MixMHCpred (v2.3) (Bassani-Sternberg M, et al. PLoS computational biology 13, el005725 (2017)). Data analysis and wrangling were performed using R and Julia. Raw file metadata were extracted using the rawrr R package (vl .12.0) (Kockmann T, et al. J Proteome Res 20, 2028-2034 (2021)). For each precursor, the highest PSM intensity was used to calculate peptide-level intensities (sum of precursor intensities per peptide). Visualization was conducted using R in RStudio, with ggplot2 (v3.5.1) (H W. ggplot2: Elegant Graphics for Data Analysis. Springer-Verlag (2016)), ggseqlogo (v0.2) (Wagih O. Bioinformatics 33, 3645-3647 (2017)), ggside (v0.3.1), patchwork (vl.3.0), ggh4x (vO.2.8), and gt (vO.11.0).
[0349] NeoDiscMS Implementation
[0350] NeoDiscMS was introduced in NeoDisc vl.7.1 and is executed in three main steps: Docket No.: 084276.00417 / LUD 6239
[0351] 1 . Execution of NeoDisc in fastq mode to identify and prioritize immunogenic tumor-specific antigens from NGS data and predict chromatographic retention times (RTs).
[0352] 2. Acquisition of immunopeptidomic data using NeoDiscMS.
[0353] 3. Resumption of the NeoDisc workflow to search NeoDiscMS raw files against personalized peptide references, followed by prioritization of tumor-specific antigens.
[0354] For calibration, the MS raw file of a proteomics or immunopeptidomics sample was designated as NEODISCMS CALIBRATION in the configuration file. FragPipe v22.0 was used to generate a calibration library based on the GENCODE v43 protein-coding sequence reference, accounting for peptide modifications (MS_MODIFICA TIONS) and maximum modifications per peptide (MS MAXMODSPERP EP TIDE).
[0355] NeoDisc then partitioned prioritized peptides into HLA-I / II-restricted neoantigens, TAAs, and viral peptides. The number of peptides per group was controlled via NEODISCMS NEOCI, TAACI, VIRCI, NEOCII, TAACII, and VIRCII parameters. RTs were predicted using DeepLC v2.2.36 (Bouwmeester R, et al. Nat Methods 18, 1363-1369 (2021)), incorporating the calibration peptide library and peptide sequence features such as modifications, charge states (NEODISCMS CHARGES), collision energy NEODISCMS CE), gradient (NEODISCMS GRADIENT), and RT window width (NEODISCMS WINRT). Output files included target HLA-I / II peptide FASTAs and associated RT CSVs for NeoDiscMS.
[0356] Upon completion of NeoDiscMS acquisition, raw files were added to the configuration file under MS SPECTRA SAMPLES, MS SPECTRA HLA I, and MS SPECTRA HLA II. The NeoDisc analysis was resumed using the -neodiscms flag. Searches were conducted using Comet (v2024.02_0) with NewAnce (vl.7.5) and MSFragger-DDA+ against personalized references generated by NeoDisc. Final outputs were integrated into NeoDisc for tumor-specific antigen prioritization.
[0357] Statistics and Reproducibility
[0358] To ensure reproducibility, all experimental comparisons were conducted in triplicate. No data were excluded. Where different technical parameters (e.g., acquisition conditions) were assessed, sample runs were randomized to mitigate systematic bias.
[0359] EXAMPLE 4 Docket No.: 084276.00417 / LUD 6239
[0360] Targeting cancer-specific human leukocyte antigen (HLA)-peptide complexes has been recognized as a promising strategy in immunotherapy. Mutated neoantigens have been identified as optimal targets due to their inherent immunogenicity and cancer specificity. Immunopeptidomics based on mass spectrometry (MS) has been utilized to guide the identification and selection of naturally presented immunogenic targets within the immunopeptidome, thereby improving the accuracy of immunogenicity predictions. For clinical implementation, however, several challenges must be addressed, including the need to achieve comprehensive coverage of the immunopeptidome (global depth), to maintain high sensitivity in target detection, to accommodate limited sample input, and to enable a rapid turnaround time. Described herein is NeoDiscMS, an extension of the NeoDisc platform, which has been developed to facilitate personalized immunopeptidomics data acquisition. By employing real-time spectral acquisition guided by next-generation sequencing, NeoDiscMS has been configured to enhance sensitivity with minimal compromise to global immunopeptidome coverage. The system has been designed for operational efficiency and ease of implementation with minimal user intervention.
[0361] Using this approach, the detection of peptides derived from tumor-associated antigens has been improved by up to approximately 20% compared to conventional methods, and the confidence in neoantigen identification has been increased relative to the current gold standard. Accordingly, NeoDiscMS represents a significant advancement in the personalization of clinical antigen discovery by enabling more reliable neoantigen detection and simplified deployment in clinical workflows.
[0362] The accurate identification of immunogenic, cancer-specific antigens presented by HLA complexes is considered essential for the development of cancer immunotherapies. In most current approaches, neoantigens are predicted computationally from next-generation sequencing (NGS) data obtained from tumor and germline tissues. However, it has been repeatedly observed that only a small proportion (~1%) of somatic mutations are capable of inducing spontaneous or vaccine- elicited T cell responses. Although machine learning classifiers have been employed to improve neoantigen ranking, the selection of truly immunogenic candidates, particularly in tumors with high mutational burdens, remains a significant challenge.
[0363] MS-based immunopeptidomics has been adopted as a method to directly identify antigens presented on tumor cells, including neoantigens, and to characterize defects in antigen presentation. Docket No.: 084276.00417 / LUD 6239
[0364] Neoantigens that are detected through MS are generally regarded as more likely to be immunogenic, as their presentation on tumor cells has been experimentally confirmed. Nevertheless, MS-based identification is impeded by the typically low abundance of neoantigens, the limited availability of clinical sample material, and the need for rapid turnaround in therapeutic settings.
[0365] Traditionally, antigen discovery has been performed using data-dependent acquisition (DDA), a stochastic approach in which precursor ions are selected for fragmentation based on signal intensity and dynamic exclusion. While this method increases proteome coverage, it often lacks the sensitivity required to detect low-abundance targets. Although chimeric spectrum deconvolution methods have been applied in other proteomic workflows to improve detection sensitivity, such methods have not yet been implemented in immunopeptidomics. Targeted strategies involving heavy-labeled spike-in standards have also been explored to enhance detection sensitivity. However, these approaches are hindered by challenges in peptide synthesis, quality control, scalability, and the risk of sample contamination, particularly when unique, patientspecific neoantigens must be analyzed on short clinical timelines.
[0366] Several spike-in-free methods have been reported to improve sensitivity, but these have typically been limited by reliance on conventional TopN-based precursor selection, by their focus on reproducing known targets rather than discovering novel ones, or by requirements for specialized instrumentation and manual data processing. In many cases, an optimal balance between target sensitivity and comprehensive immunopeptidome coverage has not been achieved.
[0367] To overcome these limitations, a novel immunopeptidomic workflow, NeoDiscMS, has been developed and is described herein. NeoDiscMS is a spike-in-free, scalable, and targeted- DDA hybrid acquisition method designed to enhance the sensitivity and accuracy of peptide detection while minimizing loss of global immunopeptidome coverage. In this approach, in silico- prioritized antigenic peptides inferred from NGS data are used to guide MS acquisition through real-time peptide-to-spectrum matching fdters (RTSf), which selectively trigger high-sensitivity scans for precursor ions with target-like features. This capability has not previously been applied to immunopeptidomics. Additionally, chimeric spectrum deconvolution has been incorporated to further increase the depth of analysis. Docket No.: 084276.00417 / LUD 6239
[0368] The performance of NeoDiscMS has been benchmarked against conventional DDA methods using HLA-I immunopeptidome dilution series. The utility of the workflow has also been demonstrated through improved detection, reproducibility, and confidence in the identification of clinically relevant TAAs and mutated neoantigens from patient-derived melanoma cell lines and tissue samples. Importantly, the data generated using NeoDiscMS can be analyzed using standard MS search engines, allowing seamless integration into existing clinical pipelines such as NeoDisc. The method is also expected to be applicable for the identification of disease-related peptides in settings beyond cancer immunotherapy.
[0369] NeoDiscMS workflow layout
[0370] NeoDiscMS was implemented using a peptide database comprising an output of 1,500 predicted HLA class I-restricted neoantigens or tumor-associated antigens. These peptides were generated by the NeoDisc end-to-end clinical antigen discovery pipeline and provided in the form of an inclusion list and a FASTA file, wherein each peptide was represented as a separate entry (FIG. 4A).
[0371] MS data acquisition was performed using three-second duty cycles, with acquisition divided into three sequential tiers prioritized as follows: (i) MSI scans, (ii) targeted branch scans, and (iii) discovery data-dependent acquisition (DDA) branch scans (FIG. 4B). Each cycle was initiated with an MSI scan.
[0372] To qualify for high-sensitivity fragmentation in the targeted scan tier, a precursor ion detected in the MSI scan was required to match a precursor mass listed in the inclusion list. A matched precursor ion was then subjected to a first-stage scouting MS2 scan (sMS2). The fragment ion spectrum from the sMS2 scan was analyzed in real time by calculating the cross-correlation between the observed sMS2 fragments and predicted fragment ion spectra from the peptide database. This matching process was completed within milliseconds and produced a peptide- spectrum match (PSM) score.
[0373] If the PSM met predefined quality thresholds, the same precursor ion was subjected to a second-stage high-sensitivity MS2 scan (hMS2) (FIG. 4C). The hMS2 scan was time-intensive and optimized to enhance detection sensitivity, employing a higher automatic gain control (AGC) target, an extended maximum injection time, and stepped collision energies. The sMS2 scan Docket No.: 084276.00417 / LUD 6239 thereby served as a filter to reduce the number of unnecessary hMS2 scans, increasing overall efficiency.
[0374] Both sMS2 and hMS2 scans were intentionally exempted from dynamic exclusion and MSI precursor intensity threshold restrictions, thereby enabling repeated MS / MS acquisition of the same precursor ion to improve identification confidence.
[0375] The remaining time within the three-second acquisition cycle was allocated to the DDA branch. To increase overall peptide coverage, the DDA scans utilized wide MS2 isolation windows of 3.2 Th, thereby enhancing the likelihood of co-isolating additional precursor ions. The raw MS data were subsequently processed using chimeric spectra deconvolution with the MSFragger DDA+ mode, which facilitated PSMs involving precursor ions whose actual masses deviated from the nominally isolated precursor mass.
[0376] Increased depth with wide isolation windows
[0377] To benchmark the performance of wide isolation windows and chimeric spectra deconvolution in the context of immunopeptidomics data, experiments were initially conducted using different isolation window sizes with data-dependent acquisition (DDA). For each measurement, the equivalent of ten million JY cells was injected, and data were analyzed using MSFragger-DDA+. In each case, approximately 96% of the identified 8-14-mer peptides were predicted as binders (%-rank < 2) by MixMHCpred, thereby confirming consistent identification quality across varying conditions. Differences observed between isolation window widths were primarily related to the number of unique peptides identified. An isolation window size of 3.2 Th was selected, as it most consistently yielded the highest number of peptide identifications. This result is consistent with prior findings indicating that, for low-input wide window sizes that permit greater depth with chimeric spectral deconvolution tend to remain narrow.
[0378] To further extend the benchmarking analysis, the HLA-1 immunopeptidome of B-cell lines JY and RA957 36 was measured using isolation windows of either 1.2 Th (representing a conventional setting) or 3.2 Th (representing a wider setting). Measurements were performed using sample equivalents of five and ten million cells, and the resulting data were processed for peptide identification using either MSFragger-DDA or MSFragger-DDA+. Across all conditions and for both software tools, 96-97% of the identified 8-14-mer peptides were predicted as binders (%- rank < 2), and the average peptide length ranged from 9.4 to 9.6 amino acids. Docket No.: 084276.00417 / LUD 6239
[0379] When data were processed with MSFragger-DDA+ instead of MSFragger-DDA, an increase of 2% to 36% in the number of unique peptides identified was observed, with higher cell inputs yielding greater identification gains. Notably, the RA957 samples exhibited a greater increase in identified peptides relative to the JY samples. Upon comparison of replicates of the same cell type and input amount, the variation in peptide identifications between isolation windows of 1.2 Th and 3.2 Th, when processed using MSFragger-DDA+, ranged from 0% to 8%.
[0380] The overlap coefficients for peptides identified in the same raw file using both MSFragger- DDA and MSFragger-DDA+ (top three grey dots in each panel of the lower row) ranged from 88% to 96%. When using MSFragger-DDA+, an increase of 7-9% was observed in the fraction of scans producing one or more peptide-spectrum matches (PSMs) when the isolation window was increased from 1.2 Th to 3.2 Th. For data analyzed with MSFragger-DDA+, between 19% and 37% of the peptides were identified solely due to chimeric spectrum deconvolution, meaning that these peptides did not correspond to the nominal precursor mass selected for isolation. This fraction was found to increase consistently by 10-11% when a 3.2 Th isolation window was applied instead of 1.2 Th to the same sample.
[0381] To further evaluate the quality of PSMs obtained through chimeric spectrum deconvolution, identified precursors corresponding to isolated and co-isolated peptide masses were analyzed separately. Similar distributions were observed for both groups in terms of predicted binder percentages for the corresponding FTLA alleles and in peptide length profiles.
[0382] In chimeric spectra, the mass difference between co-isolated peptides is determined by combinations of amino acids whose cumulative mass allows for simultaneous isolation within the same isolation window. As expected, due to the narrow distribution (~0.1 Da range) of amino acid mass residuals, the absolute mass differences between co-isolated peptides (with identical charge states, accounting for 92% of cases) were frequently observed to approximate integer multiples of 1 Da.
[0383] In summary, the application of wide isolation windows, particularly in combination with chimeric spectrum deconvolution, resulted in an increased depth of the global immunopeptidome while maintaining high identification quality. Accordingly, all discovery -phase analyses described herein were conducted using a 3.2 Th isolation window and processed with MSFragger-DDA+. It is noted that the wide isolation window strategy was applied exclusively to the discovery branch. Docket No.: 084276.00417 / LUD 6239
[0384] Therefore, implementation of NeoDiscMS can be achieved independently of this feature, at the user’s discretion.
[0385] Scalable enhanced sensitivity for targets
[0386] To demonstrate the enhanced sensitivity of time-intensive scans for target discovery and the capability of RTSf to reduce the burden on global immunopeptidome depth, the performance of three acquisition modes was benchmarked: data-dependent acquisition (DDA), DDA with an inclusion list branch (ilDDA), and NeoDiscMS (FIG. 4B). DDA was employed as the reference standard for antigen discovery, with an inclusion of ilDDA configured to trigger high-sensitivity scans.
[0387] To accurately evaluate the performance of the target identification methodology, an inclusion list comprising approximately 1,500 unique RA957-specific, HLA-A68:01-restricted 8- 11-mer peptides was generated. These peptides were consistently identified across RA957 samples and were predicted to exhibit poor binding affinity to JY HLA-I alleles, with binding ranks exceeding 10% for HLA-A*02:01, -B*07:02, and -C*07:02. Methionine oxidation was included as a variable modification, and all charge states from 1+ to 3+ were incorporated, yielding a total of 5,046 precursor ions.
[0388] To prevent indiscriminate scanning across the chromatographic gradient, scheduled targeting was applied. Predicted retention times (RTs) were used to define a scheduling window of ±15 minutes per peptide. This approach was shown to capture approximately 75% of targets based on the comparison between predicted and measured retention times, albeit at the cost of a substantial precursor scheduling load.
[0389] An immunopeptide dilution series was prepared by mixing RA957-derived peptides with those from JY cells, using dilution ratios of 1: 16, 1 :64, 1 :256, and 1 :1024, along with a JY-only negative control. The 1 : 1024 dilution was designed to simulate a low-abundance target environment, whereas the 1 :16 dilution was intended to evaluate scalability in a high-target context. All dilution conditions were analyzed using DDA, ilDDA, and NeoDiscMS acquisition modes (FIG. 4B).
[0390] Across all conditions, the percentage of identified peptides binding to HLA-A*68:01 was maintained between 96% and 98%, with an average peptide length of approximately 9.4-9.5 amino Docket No.: 084276.00417 / LUD 6239 acids. The success of the dilution series was confirmed by analysis of the binder fractions. No target peptides were detected in the JY-only controls for any acquisition method. In all dilutions, NeoDiscMS yielded on average 0%-7% more peptide targets than ilDDA, and 16%- 51% more than DDA. The global depth achieved by NeoDiscMS was observed to exceed that of ilDDA by 5%— 21% while falling 5%- 14% below that of DDA.
[0391] In terms of unique peptide identifications, it was found that for every 1% decrease in global depth relative to DDA, ilDDA achieved a 2% increase in target sensitivity, and NeoDiscMS achieved a 5.6% increase. This demonstrates that the loss in global depth was minimized while substantially improving sensitivity through the RTSf-enabled targeted acquisition strategy.
[0392] The 1 : 1024 dilution condition, which most closely represents real -world tumor antigen discovery scenarios, enabled consistent identification of 18 peptide targets across all NeoDiscMS replicates, whereas only 9 such targets were consistently detected using DDA, a two-fold increase in reproducible target identification. In contrast, 9 peptides were uniquely detected by DDA, whereas 19 peptides were uniquely detected by NeoDiscMS, representing a 2.1 -fold improvement. In total, NeoDiscMS identified 24% more peptide targets than DDA across three replicates.
[0393] Peptides identified by NeoDiscMS also exhibited higher average peptide-spectrum match (PSM) counts per target compared to both DDA and ilDDA. Detailed analysis of the 1 : 1024 NeoDiscMS runs revealed that target peptides missed by DDA were distributed across various MS2 scan types and included peptides of low to medium intensity. This indicates that improved target identification in NeoDiscMS is attributable to both enhanced scan sensitivity and reduced ion sampling stochasticity.
[0394] In the 1 : 16 dilution condition, where target peptide signals were most abundant, higher hyperscore and spectral angle values were consistently observed for hMS2 PSMs relative to dMS2 PSMs, across the entire intensity spectrum, further supporting improved spectral quality and target detection.
[0395] To assess the reproducibility of quantification, MSI -based quantification results from FragPipe output were analyzed using data from the 1 : 1024 dilution. Comparable reproducibility between NeoDiscMS and DDA was observed, both for non-target peptides (gray; R2= 0.96-0.97) and target peptides (red; R2= 0.92-0.99). Although slightly lower correlation values were noted Docket No.: 084276.00417 / LUD 6239 in technical replicate TR2 vs. TR3 comparisons, this was attributed to the reduced number of target identifications rather than inherent variability.
[0396] Repeated fragmentation enabled by the targeted acquisition branch of NeoDiscMS is expected to facilitate the implementation of advanced MS2-based quantification strategies. However, it should be noted that hMS2 scans are acquired under different ion accumulation and collision energy settings than dMS2 and sMS2 scans, introducing variability in both absolute fragment ion intensities and relative ion ratios, as demonstrated for the peptide EVIL1DPFHK (SEQ ID NO: 2).
[0397] In conclusion, benchmarking with immunopeptidome dilution series demonstrated that the targeted acquisition branch enables more sensitive and reliable peptide target identification relative to traditional DDA approaches, while maintaining comparable quantitative performance.
[0398] Broader TAA and neoantigen discovery
[0399] To further characterize the utility of the disclosed method, immunopeptidomics data were acquired using data-dependent acquisition (DDA) and NeoDiscMS from three melanoma cell lines derived from patient tumors: MH, M12, and M13. For each cell line, sample-specific mutated neoantigen peptides prioritized by NeoDisc were selected as targets based on paired tumor-normal whole exome sequencing (WES) and tumor transcriptome data. In addition, the target list was expanded to include well-characterized TAAs, such as MAGEA3, MAGEA4, and ML ANA, which were confirmed to be expressed in the samples. TAA-derived peptides predicted from in silico three-frame translation products were also included to explore additional sources of antigenic peptides.
[0400] As expected, the deepest coverage of the immunopeptidome was achieved using DDA across all three cell lines, while NeoDiscMS yielded a mean reduction in depth of 21%, 13%, and 4% for MH, M12, and M13, respectively. Despite the reduction in depth, between 94% and 95% of peptide binders were retained across datasets, and the average peptide length remained within the range of 9.3 to 9.6 amino acids. A total of 49 TAA-derived peptides were identified, including several peptides previously confirmed as immunogenic epitopes. In addition, four mutated neoantigen peptides were detected, three of which had been previously demonstrated to be immunogenic in the corresponding patient samples. Docket No.: 084276.00417 / LUD 6239
[0401] Importantly, for each cell line, every target peptide identified by DDA was also detected using NeoDiscMS, with the exception of four targets — one of which was identified in three DDA replicates, and the remaining three only once across DDA replicates. Moreover, the fraction of target peptides detected by NeoDiscMS in at least two out of three replicates, but identified in only one or none of the DDA replicates, was calculated to be 16%, 24%, and 0% for Mil, M12, and M13, respectively.
[0402] At the protein level, NeoDiscMS was shown to substantially expand antigen coverage. For instance, in M12, three out of five SAGEl-derived peptides (THNVHEEKI (SEQ ID NO: 3), QHATIIHNL (SEQ ID NO: 4), and THNIREEKI (SEQ ID NO: 5)) were consistently identified using NeoDiscMS, whereas those same peptides were detected only once or not at all using DDA across all replicates.
[0403] Confident neoantigen detection
[0404] To further characterize the observed improvements in peptide detection confidence, the underlying characteristics of real-time search (RTS) events following activation of a targeted branch were examined. When the precursor mass was detected within the permitted retention time (RT) window, a scheduled MS2 (sMS2) scan was acquired. The resulting sMS2 spectrum was subsequently subjected to RTS processing. Upon passing the RTS filter (RTSf), the corresponding peptide-spectrum match (PSM) was designated as an RTS hit. Each RTS hit, in turn, triggered the acquisition of a higher-resolution MS2 (hMS2) scan.
[0405] When a peptide precursor triggered an RTS event and the same peptide was subsequently identified via MSFragger in either the sMS2 or hMS2 spectra, the identification was defined as a target identification (TID). On average, 12%, 8%, and 8% of RTS events resulted in RTS hits for Mil, M12, and M13, respectively. Notably, more than 85% of the inclusion list hits obtained during measurement failed to pass RTSf, indicating that these entries represented isobaric non-targets. Among RTS hit events, the average fractions that resulted in TIDs were 1%, 6%, and 2% for MH, M12, and M13, respectively. All observed TIDs occurred exclusively within RTS hit events.
[0406] Across analyzed samples, at least 90% of RTS hit events resulting in TIDs exhibited a delta parts-per-million (dPPM) deviation of less than 1. Among all RTS events with TIDs, the majority of peptides were identified in both sMS2 and hMS2 spectra. Notably, sample M12 particularly benefited from hMS2 acquisition, with the number of uniquely hMS2-derived TIDs exceeding the Docket No.: 084276.00417 / LUD 6239 number of uniquely sMS2-derived TIDs by approximately 60%. RTS results corresponding to TIDs uniquely identified through hMS2 scans exhibited cross-correlation (Xcorr) scores within the range of 0.4 to 2.7.
[0407] TIDs identified in both sMS2 and hMS2 scans from the same RTS events provided a direct measure of the improvement in identification accuracy between these scan types. For low-scoring sMS2 TIDs, both hyperscores and spectral angles associated with the same precursor PSM increased upon subsequent hMS2 scanning.
[0408] The MCTSlD130N-derived neoantigen YPAAVNTIVAI (SEQ ID NO: 6), detected in M12, exemplified the advantages of using the NeoDiscMS workflow for reliable discovery of clinically relevant mutant neoantigens. Based on predictive ranking, NeoDisc placed this neoantigen at position 186 among predicted HLA-I neoantigens, whereas an autologous tumorinfiltrating lymphocyte (TIL)-based immunogenicity assay confirmed its immunogenicity. In a hypothetical clinical scenario relying solely on mutanome-based predictions to select targets for personalized cancer vaccine development, a neoantigen ranked as low as position 186 would typically be excluded, since selection is often limited to the top 40 candidates. However, when immunopeptidomics data were integrated with mutanome predictions, the robust MS-based evidence provided by NeoDiscMS justified the inclusion of this lower-ranked, yet immunogenic, neoantigen. Without such proteomic validation, inclusion of a low-ranked neoantigen in place of a higher-ranked candidate would not be scientifically justified.
[0409] YPAAVNTIVAI was identified in all immunopeptidomic replicates using data-dependent acquisition (DDA), although only up to two PSMs were observed per replicate. In contrast, application of NeoDiscMS yielded repeated RTS hits across acquisition cycles, resulting in 9 to 13 PSMs per replicate. Among these, 5 to 7 PSMs per replicate were derived from hMS2 spectra. The best DDA-derived MS2 (dMS2) PSM for YPAAVNTIVAI exhibited a spectral angle of 0.9299 and matched only seven fragments (b5-b9, yl, y2), with no fragmentation coverage of the first four peptide bonds. Across NeoDiscMS runs, the best dMS2, sMS2, and hMS2 spectra for this peptide achieved spectral angles of 0.9635, 0.9766, and 0.995, respectively. The optimal hMS2 PSM matched twelve fragment ions (b2-bl0, yl-y3), thus enabling accurate and confident identification of the peptide. Docket No.: 084276.00417 / LUD 6239
[0410] Accordingly, NeoDiscMS provided consistent and high-confidence detection of the YPAAVNTIVAI neoantigen, as demonstrated by improved spectral quality metrics and increased fragment coverage.
[0411] Direct clinical application
[0412] To demonstrate clinical applicability, a clinical immunopeptidome pipeline was employed to extract the immunopeptidome from three distinct tumor lesions (Til, Ti2, and Ti3) of a patient diagnosed with uveal melanoma. For each lesion, NeoDiscMS analysis was performed using the equivalent of 10 mg of tissue.
[0413] A consolidated target list comprising the top 500 predicted neoantigens and the top 518 predicted TAA-derived peptides was generated and used for NeoDiscMS measurements across all three tissue samples. The clinical routine immunopeptidomics acquisition protocol included both data-dependent acquisition (DDA) and data-independent acquisition (DIA) measurements. These measurements enabled sensitive peptide-centric DIA searches against a sample-specific hybrid spectral library constructed from spectrum-centric analysis of DIA data in combination with a DDA-derived spectral library.
[0414] Accordingly, one technical DIA and NeoDiscMS acquisition was conducted for each tissue sample using a personalized, next-generation sequencing (NGS)-guided target database. Using NeoDiscMS, 14,797, 16,033, and 11 ,968 unique peptides were identified from Ti l, Ti2, and Ti3, respectively, with an average peptide length of 9.4 amino acids and 95% predicted binding affinity. Spectrum-centric searches of DIA data (e.g., DIA-Umpire) identified fewer than 50% of the peptides detected using the corresponding NeoDiscMS approach.
[0415] Subsequently, DIA files were searched against spectral hybrid libraries (generated from NeoDiscMS and DIA data), resulting in the identification and quantification of 12,917, 14,165, and 10,726 unique peptides in Til, Ti2, and Ti3, respectively. Comparative analysis of shared and lesion-specific peptide sets revealed variations in peptide presentation among the three samples. The proportion of peptides common to all three lesions was determined to be 31% using the NeoDiscMS search and 36% using the hybrid library DIA search.
[0416] As anticipated for uveal melanoma, the tumor mutation burden was low in the three lesions (2.601, 2.82, and 3.615 somatic mutations / Mb for Til, Ti2, and Ti3, respectively), and no Docket No.: 084276.00417 / LUD 6239 neoantigens were detected in any of the samples. Nevertheless, a total of nine TAA-derived peptides were identified, originating from MLANA, TYR, and SAGE1. Among these, the peptide MPREDAHF from MLANA was detected in the DIA files of Til, Ti2, and Ti3, while the peptide YRALMDKSL from MLANA was detected exclusively in Til, and only when analyzed against the hybrid library, but not through spectrum-centric analysis.
[0417] These results demonstrate that NeoDiscMS can be effectively integrated with DIA methodologies to generate sample-specific hybrid libraries enriched with high-confidence MS2 spectra. This integration enables sensitive detection and quantification of the global immunopeptidome as well as predefined target peptides.
[0418] Discussion
[0419] A novel immunopepti domic mass spectrometry acquisition method, referred to herein as NeoDiscMS, has been developed to address key challenges in the identification of clinically relevant tumor-specific antigens. Such antigens, typically present at low abundance, require rapid, sensitive, and robust detection to enable their utility in immunotherapeutic applications.
[0420] NeoDiscMS is implemented within the NeoDisc platform, an end-to-end computational pipeline that integrates genomic, transcriptomic, and MS-based immunopeptidomic data for the prediction and direct identification of HLA class I and II tumor-specific antigens. In prior implementations, next-generation sequencing (NGS) and MS-immunopepti domic data were integrated computationally using NeoDisc. In the present approach, matched tumor and germline genome data, together with tumor transcriptome data, are processed by NeoDisc to generate a personalized proteome reference annotated with single-nucleotide polymorphisms and somatic mutations. These annotations are subsequently used for HLA binding prediction and machine learning-based immunogenicity scoring of candidate antigenic peptides.
[0421] From this prioritized peptide list, a personalized inclusion list is generated by NeoDisc for use in NeoDiscMS acquisition. Mass spectrometric analysis of the matched tumor immunopeptidome sample is then performed using a data-independent acquisition (DIA) strategy. The resulting immunopeptidomic data are searched against the personalized proteome to identify naturally presented peptides, thereby refining the list to prioritize those most likely to be immunogenic and clinically actionable. This integrated workflow supports a two-week turnaround Docket No.: 084276.00417 / LUD 6239 time from biopsy receipt to the generation of a ranked antigen list suitable for downstream applications such as neoantigen vaccine design.
[0422] NeoDiscMS constitutes a scalable, spike-in-free immunopeptidomics workflow that may be used either in conjunction with NeoDisc or independently, using any predefined peptide target list. The method is particularly well-suited for low-input biopsy specimens with stringent timelines, offering both target sensitivity and comprehensive immunopeptidomic coverage. Sample preparation is conducted using standard immunopeptidomic protocols, and no prior MS measurements of the sample are required. This enables full sample consumption in a single acquisition, improving both sensitivity and turnaround time.
[0423] During acquisition, separation of the discovery and targeted branches, along with fully customizable scan parameters, allows for optimization of sensitivity-depth trade-offs. Compared to inclusion list-dependent data acquisition (ilDDA), NeoDiscMS demonstrates superior performance by minimizing acquisition time spent on off-target ions, thereby enabling greater sampling of target masses and improved discovery branch utilization. Moreover, NeoDiscMS enables up to three MS2 scans per precursor ion per cycle, whereas ilDDA permits only two. Each MS2 scan yields a peptide-spectrum match (PSM), enhancing identification confidence and increasing both global depth and target sensitivity.
[0424] The targeted branch of NeoDiscMS employs a simplified decision tree (sMS2 —> RTSf —> hMS2) which supports further enhancements in scan specificity and sensitivity through future improvements in real-time search (RTS) design and scan parameter interdependencies. Selective high-sensitivity scans may also incorporate optimizations in collisional energy, ion mobility, and isolation window width. Dynamic scan parameter definition using sMS2 output and refinement of RTS settings are anticipated to further improve peptide identification. Additionally, the ability to acquire multiple MS2 scans of the same precursor across acquisition cycles offers new capabilities for quantitative analysis and consensus spectrum generation.
[0425] The targeted branch of NeoDiscMS is specifically designed to enable direct identification of tumor-specific antigens for clinical applications. However, the discovery branch remains a critical component, providing global immunopeptidomic coverage and facilitating identification of antigen processing defects, presentation dynamics, and immunogenicity-relevant features. This branch also contributes additional spectra of both targeted and non-targeted antigens and supports Docket No.: 084276.00417 / LUD 6239 re-processing of data for new insights or database construction. Notably, the discovery branch yields a substantial number of high-quality spectra that are amenable to peptide identification using conventional search engines, thereby facilitating reliable false discovery rate (FDR) estimation and enhancing downstream bioinformatics analyses such as rescoring with MSBooster.
[0426] Compatibility with standard proteomics tools, including MSFragger and its derivatives (e.g., MSFragger-DDA+), is a central design objective of NeoDiscMS. The application of chimeric spectrum deconvolution and wide isolation windows in the discovery branch allows for high immunopeptidomic coverage without sacrificing identification confidence. Additionally, the capacity of NeoDiscMS to generate multiple MS2 spectra for a given precursor enhances the potential for quantification and identification.
[0427] Empirical validation of NeoDiscMS has demonstrated improved performance relative to conventional data-dependent acquisition (DDA) approaches. Enhanced immunopeptidome depth was observed in comparisons of wide versus standard isolation windows, and target identification was consistently improved in samples derived from melanoma and other tumor cell lines. Evaluation of real-time search (RTS) dynamics confirmed the contributions of sMS2 and hMS2 scans to increased sensitivity, with specific neoantigens, such as YPAAVNTIVAI, showing improved identification confidence.
[0428] Clinical feasibility was demonstrated through the analysis of low-input tumor biopsy samples, in which NeoDiscMS achieved significant depth and identification of tumor-associated antigen (TAA)-derived peptides. In comparison with DIA-based library searching, NeoDiscMS yielded substantially more peptide identifications. In contrast to DIA MS2 spectra, which are complex and challenging to interpret manually, NeoDiscMS produces low-complexity spectra, facilitating manual inspection of PSMs — an essential consideration for clinical prioritization of candidate antigens.
[0429] NeoDiscMS has been implemented on instruments equipped with vendor-supported acquisition method editors, including the Orbitrap Eclipse and Orbitrap Ascend mass spectrometers (Thermo Fisher Scientific, Tribrid series). NeoDiscMS constitutes the first personalized immunopeptidomics acquisition strategy explicitly designed to support clinical T- cell antigen discovery via proteogenomics. The system has been developed with emphasis on modularity, scalability, and ease of use, and is fully integrated within NeoDisc vl .7.1. NeoDiscMS Docket No.: 084276.00417 / LUD 6239 provides significant improvements in both analytical performance and clinical applicability for personalized cancer immunotherapy development.
[0430] The present disclosure is not to be limited in scope by the specific embodiments described herein. Indeed, various modifications of the invention in addition to those described herein will become apparent to those skilled in the art from the foregoing description and the accompanying figures. Such modifications are intended to fall within the scope of the appended claims.
[0431] Docket No.: 084276.00417 / LUD 6239
[0432] Table 1: Comparison of NeoDisc 1.7.0 with pTuneos vl.0.0 and pVACtools v4.0.8: Inputs, Outputs, and Supported Features. This table provides a comparative overview of the capabilities of NeoDisc, in comparison with other established bioinformatics tools in the field. The table details the input data formats, and the features supported by each tool. Docket No.: 084276.00417 / LUD 6239 Docket No.: 084276.00417 / LUD 6239
[0433] Docket No.: 084276.00417 / LUD 6239
[0434] Table 2: Comparative Analysis of Neoantigen Prediction Ranks of Immunogenic Mutations in the NCI Test Cohort. The table lists each mutation alongside the corresponding gene, amino acid change, and predicted peptide. For each tool — NeoDisc's rule-based and machine learning models, pTuneos, pVACseq, and Gartner — the ranks of the predicted neoantigens are provided, highlighting the performance and prioritization differences among the tools.
[0435] Docket No.: 084276.00417 / LUD 6239
[0436] Docket No.: 084276.00417 / LUD 6239
[0437] Docket No.: 084276.00417 / LUD 6239
[0438] * The original Gartner et al. publication did not link immunogenic peptides ranks to their sequence in samples with multiple immunogenic peptides identified. In such cases, the ranks were assigned randomly to the sequences.
[0439] ** The mutation was not identified
[0440] *** Another peptide sequence was reported for the mutation
[0441] **** pTuneos failed analysis
[0442] Docket No.: 084276.00417 / LUD 6239
[0443] Table 3: Summary of Patient Tumor Characteristics, Diagnosis, Material available, and Neoantigen Screening Results. This table provides a summary of the patient cohort described in the paper, including clinical and molecular characteristics. For each patient, the table details the diagnosis, tumor grade, gender, HLA typing (HLA-I and HLA-II), types of assays conducted, sample type and amount, estimated tumor content, tumor mutational burden, HLA loss of heterozygosity (LOH), and specific defects in the antigen processing and presentation machinery (APPM). Additionally, the table outlines the number of neoantigens, and viral peptides identified through screening, along with their immunogenic status. This comprehensive summary is essential for understanding the diversity and characteristics of the patient cohort, which informs the subsequent analysis of neoantigen identification and immunogenicity.
[0444] Docket No.: 084276.00417 / LUD 6239
[0445] Docket No.: 084276.00417 / LUD 6239
[0446] Docket No.: 084276.00417 / LUD 6239
[0447] Docket No.: 084276.00417 / LUD 6239
[0448] Docket No.: 084276.00417 / LUD 6239
[0449] * Manually corrected based on VAF ** Reported from gene panel report
[0450] Docket No.: 084276.00417 / LUD 6239
[0451] Table 4: Immunogenic Peptides Identified in Patients. This table reports all immunogenic peptides identified across the patient cohort described in the paper. For each patient, the table lists the ranking of the peptide sequence, the associated predicted HLA allele, and whether it was identified through mass spectrometry (MS). When tested, it reports whether the peptide induced tumor rejection or not, the type of T-cell response observed (CD4 or CD8), and results from Elispot assays. Additionally, information about the gene and mutation corresponding from which each peptide derives, as well as the wild-type (WT) peptide sequence, is provided.
[0452] Docket No.: 084276.00417 / LUD 6239
[0453] Docket No.: 084276.00417 / LUD 6239
[0454] Docket No.: 084276.00417 / LUD 6239
[0455] Docket No.: 084276.00417 / LUD 6239
[0456] Docket No.: 084276.00417 / LUD 6239
[0457] Docket No.: 084276.00417 / LUD 6239
[0458] *The rank is provided per predicted HLA class (I / II) type of antigen as different types of antigens (e.g. neo-antigens, viral antigens, HC-TSAs) are prioritized separately. Importantly, the rank provided reflects predictions prioritization and does not take MS identification into account.
[0459] ** Prioritization and design of class-II neo-antigens was updated, a longer or shorter peptides covering the mutation, predicted to bind the same allele, is favored by the algorithm (see sequence after the 7’) over the tested one
[0460] *** Original peptide sequence was selected by the rulebased algorithm, second sequence favored by the ML algorithm
Claims
1. Docket No.: 084276.00417 / LUD 6239CLAIMSWhat is claimed is:
1. A method of identifying and ranking neoantigens capable of inducing a tumor-specific immune response, comprising: selecting, from candidate amino acid sequences containing one or more mutations, a first set of amino acid sequences, each of which is free of a false-positive mutation; identifying, from the first set of amino acid sequences, a second set of amino acid sequences, each of which contains a tumor-specific mutant amino acid sequence; identifying, from the second set of amino acid sequences, a third set of amino acid sequences, each of which contains a mutation within a binding core region interacting with a binding groove of a human leukocyte antigen (HLA) molecule; selecting, from the third set of amino acid sequences, a fourth set of amino acid sequences, each of which exhibits high binding affinity to HLA-I and / or HLA-II, and identifying remaining unselected third set of amino acid sequences as a fifth set of amino acid sequences; selecting, from the fourth set of amino acid sequences, a sixth set of amino acid sequences, each of which has a high likelihood of presentation, and identifying remaining unselected fourth set of amino acid sequences as a seventh set of amino acid sequences; selecting, from the sixth set of amino acid sequences, an eighth set of amino acid sequences, each of which is from a gene having a transcripts-per-million expression value greater than 15, and adding remaining unselected sixth set of amino acid sequences as a seventh set of amino acid sequences; selecting, from the eighth set of amino acid sequences, a ninth set of amino acid sequences, each of which has a RNAseq mutation coverage between 0% and 30%; and selecting, from the eighth set of amino acid sequences, a tenth set of amino acid sequences, each of which has a RNAseq mutation coverage greater than 30%, as having a highest likelihood of inducing a tumor-specific immune response, and adding the remaining unselected eighth set of amino acid sequences to the seventh set of amino acid sequences.
2. The method of claim 1, comprising ranking a likelihood of inducing a tumor-specific immune response of the neoantigens in an order from high to low as: the tenth set of amino acidDocket No.: 084276.00417 / LUD 6239 sequences, the ninth set of amino acid sequences, the seventh set of amino acid sequences, and the fifth set of amino acid sequences.
3. The method of any one of the preceding claims, comprising sequentially sorting the fifth set of amino acid sequences, the seventh set of amino acid sequences, the ninth set of amino acid sequences, or the tenth set of amino acid sequences based on: a likelihood that a mutation within an amino acid sequence is oncogenic, binding affinity to HLA-I and / or HLA-II, a likelihood of presentation, a cancer cell fraction (CCF) value measuring a proportion of cancer cells within a tumor sample that contains a specific genetic mutation, number of predicted overlapping HLA-II binding amino acid sequences, number of predicted included HLA-I binding amino acid sequences, or a combination thereof.
4. The method of claim 3, wherein the likelihood that a mutation within an amino acid sequence is oncogenic is determined by CScape.
5. The method of claim 3, wherein the binding affinity to HLA-I is determined by mixMHCpred, and the binding affinity to HLA-II is determined by mixMHC2pred.
6. The method of claim 3, wherein the likelihood of presentation is represented as an ipMSDB presentation score.
7. The method of claim 3, wherein the cancer cell fraction (CCF) value equal to or greater than 1 indicates that a mutation is clonal and affects nearly all cancer cells, and the cancer cell fraction (CCF) value smaller than 1 indicates that the mutation is sub-clonal and affects only a fraction of the tumor cells.
8. The method of claim 3, wherein the number of predicted overlapping HLA-II binding amino acid sequences indicates a potential for broader immune recognition and is determined based on number of HLA-II restricted peptides that cover a given HLA class-I peptide.Docket No.: 084276.00417 / LUD 62399. The method of claim 3, wherein the number of predicted included HLA-I binding amino acid sequences indicates a potential of a HLA-II restricted peptide to simultaneously present multiple HLA-I peptides.
10. The method of any one of the preceding claims, wherein the candidate amino acid sequences are in form of peptides.
11. The method of any one of the preceding claims, wherein the false-positive mutation is repeatedly identified across samples and cancer types in highly polymorphic genes.
12. The method of any one of the preceding claims, wherein the tumor-specific mutant amino acid sequence does not exist in a reference proteome of a healthy subject.
13. The method of any one of the preceding claims, wherein the binding core region is determined by mixMHC2pred.
14. The method of any one of the preceding claims, wherein for a HLA-I binding amino acid sequence, the HLA-I binding amino acid sequence in whole is treated as the binding core region.
15. The method of any one of the preceding claims, wherein for a HLA-II binding amino acid sequence, the binding core region comprises about nine amino acids within the HLA-II binding amino acid sequence.
16. The method of any one of the preceding claims, wherein the fourth set of amino acid sequences that exhibit high binding affinity to HLA-I and / or HLA-II have a binding score equal to or less than 1 as determined by mixMHCpred for the HLA-I and mixMHC2pred for the HLA- II.
17. The method of claim 1, wherein the sixth set of amino acid sequences having the high likelihood of presentation have an ipMSDB presentation score higher than 5 as determined by bestWTPeptideCount l for the HLA-I and bestWTPeptideCount ll for the HLA-II.Docket No.: 084276.00417 / LUD 623918. The method of any one of the preceding claims, wherein the transcripts-per-million value indicates gene expression levels and is determined based on RNA sequencing data.
19. The method of any one of the preceding claims, wherein the RNAseq mutation coverage indicates prevalence of a specific genetic mutation within RNAseq data.
20. A method for identifying and prioritizing neoantigens in a subject based on mass spectrometry data, comprising:(a) identifying and ranking the neoantigens according to the method of any one of claims 1-19;(b) predicting chromatographic retention times for the neoantigens;(c) acquiring immunopeptidomic data using a mass spectrometry-based data acquisition system configured to perform real-time spectral acquisition based on the predicted retention times; and(d) searching raw mass spectrometry data files generated in step (c) against personalized peptide reference sequences to prioritize the neoantigens based on search results.
21. The method of claim 20, wherein the mass spectrometry-based data acquisition system comprises NeoDiscMS.
22. The method of claim 20, wherein the personalized peptide reference sequences are derived from somatic mutations identified from next generation sequencing data of the subject.
23. The method of claim 20, wherein the prioritizing of the neoantigens is based at least in part on predicted immunogenicity and expression levels.
24. The method of claim 20, wherein the predicted chromatographic retention times are used to guide real-time spectral acquisition during mass spectrometry analysis.
25. The method of claim 20, wherein the mass spectrometry-based data acquisition system configured to perform a data acquisition cycle comprising:Docket No.: 084276.00417 / LUD 6239 performing a full MSI scan to detect precursor ions; for each precursor ion detected in the MSI scan that matches a precursor mass in an inclusion list:(i) performing a first-stage scouting MS2 scan (sMS2);(ii) calculating a peptide-spectrum match (PSM) score in real time by comparing fragment ions from the sMS2 scan to predicted fragment ion spectra from a peptide database; and(iii) if the PSM score meets a predefined quality threshold, performing a second-stage high-sensitivity MS2 scan (hMS2) of the same precursor ion using a higher automatic gain control (AGC) target, an extended maximum injection time, and stepped collision energies; allocating any remaining time to a discovery data-dependent acquisition (DDA) scan branch using wide MS2 isolation windows; wherein the sMS2 and hMS2 scans are exempted from dynamic exclusion and MSI precursor intensity threshold restrictions.
26. The method of claim 25, wherein the PSM score is calculated using a cross-correlation algorithm between observed and predicted fragment ion spectra.
27. The method of claim 25, further comprising processing raw mass spectrometry data using chimeric spectra deconvolution with a deconvolution algorithm configured to assign peptide- spectrum matches to precursor ions having actual masses deviating from the nominal isolation mass.
28. The method of claim 25, wherein the deconvolution algorithm comprises MSFragger operating in DDA+ mode.
29. The method of claim 25, wherein the data acquisition cycle is repeated to enable repeated MS2 acquisition of the same precursor ion to improve identification confidence.
30. A method of ranking viral antigens, comprising:Docket No.: 084276.00417 / LUD 6239 identifying, from candidate amino acid sequences containing one or more mutations, a first set of amino acid sequences, each of which exhibits high binding affinity to HLA-I and / or HLA- II, and identifying remaining unselected candidate amino acid sequences as a second set of amino acid sequences; selecting, from the first set of amino acid sequences, a third set of amino acid sequences, each of which is immunogenic in Immune Epitope Database (IEDB) on an HLA allele which a patient expresses; selecting, from the first set of amino acid sequences, a fourth set of amino acid sequences, each of which is immunogenic in IEDB but not on an HLA allele which the patient expresses; identifying remaining unselected first set of amino acid sequences as a fifth set of amino acid sequences; and identifying the third set of amino acid sequences as having a highest likelihood as a viral antigen.
31. The method of claim 30, comprising ranking a likelihood as the viral antigen in an order from high to low as: the third set of amino acid sequences, the fourth set of amino acid sequences, the fifth set of amino acid sequences, and the second set of amino acid sequences.
32. The method of any one of claims 30-31, comprising sequentially sorting the second set of amino acid sequences, the third set of amino acid sequences, the fourth set of amino acid sequences, or the fifth set of amino acid sequences based on: IEDB immunogenicity, number of IEDB validations, viral gene expression, binding affinity to HLA-I and / or HLA-II, or a combination thereof.
33. The method of claim 32, wherein IEDB immunogenicity is determined based on whether an amino acid sequence has been reported as immunogenic in IEDB.
34. The method of claim 32, wherein the number of IEDB validations is determined based on number of times an amino acid sequence has been reported as immunogenic in IEDB.Docket No.: 084276.00417 / LUD 623935. The method of claim 32, wherein the viral gene expression is determined based on expression level of a viral gene.
36. The method of any one of claims 30-35, wherein the first set of amino acid sequences that exhibit high binding affinity to HLA-I and / or HLA-II have a binding score equal to or less than 1 as determined by mixMHCpred for the HLA-I and mixMHC2pred for the HLA-II.
37. A method of ranking tumor-associated antigens (TAAs) or high-confidence tumor-specific antigens (HC-TSAs) capable of inducing a tumor-specific immune response, comprising: selecting, from candidate amino acid sequences containing one or more mutations, a first set of amino acid sequences, each of which exhibits high binding affinity to HLA-I and / or HLA- II, and identifying remaining unselected third set of amino acid sequences as a second set of amino acid sequences; selecting, from the first set of amino acid sequences, a third set of amino acid sequences, each of which has a high likelihood of presentation, and identifying remaining unselected first set of amino acid sequences as a fourth set of amino acid sequences; selecting, from the third set of amino acid sequences, a fifth set of amino acid sequences, each of which is from a gene having a transcripts-per-million expression value greater than 15, and adding remaining unselected third set of amino acid sequences to the fourth set of amino acid sequences; selecting, from the fifth set of amino acid sequences, a sixth set of amino acid sequences, each of which is immunogenic in Immune Epitope Database (IEDB) on an HLA allele which a patient expresses; selecting, from the fifth set of amino acid sequences, a seventh set of amino acid sequences, each of which is immunogenic in IEDB but not on an HLA allele which the patient expresses; adding remaining unselected fifth set of amino acid sequences to the fourth set of amino acid sequences; and identifying the sixth set of amino acid sequences as having a highest likelihood as a tumor- associated antigen (TAA) or high-confidence tumor-specific antigen (HC-TSA).Docket No.: 084276.00417 / LUD 623938. The method of claim 37, comprising ranking a likelihood as the tumor-associated antigen (TAA) or high-confidence tumor-specific antigen (HC-TSA) in an order from high to low as: the sixth set of amino acid sequences, the seventh set of amino acid sequences, the fourth set of amino acid sequences, and the second set of amino acid sequences.
39. The method of claim 37, comprising sequentially sorting the second set of amino acid sequences, the fourth set of amino acid sequences, the sixth set of amino acid sequences, or the seventh set of amino acid sequences based on: binding affinity to HLA-I and / or HLA-II, number of HLA restrictions, likelihood of presentation, gene expression, or a combination thereof.
40. The method of claim 39, wherein the binding affinity to HLA-I is determined by mixMHCpred, and the binding affinity to HLA-II is determined by mixMHC2pred.
41. The method of claim 39, wherein the number of HLA restrictions is determined based on number of HLA-I and HLA-II alleles in a patient that presents an amino acid sequence.
42. The method of claim 39, wherein the number of HLA restrictions is determined by mixMHCpred for HLA-I and mixMHC2pred for HLA-II.
43. The method of claim 39, wherein the gene expression is determined based on samplespecific RNAseq data.
44. The method of any of claims 37-43, wherein the first set of amino acid sequences that exhibit high binding affinity to HLA-I and / or HLA-II have a binding score equal to or less than 1 as determined by mixMHCpred for the HLA-I and mixMHC2pred for the HLA-II.
45. The method of any of claims 37-44, wherein the likelihood of presentation is represented as an ipMSDB presentation score.Docket No.: 084276.00417 / LUD 623946. The method of claim 45, wherein the third set of amino acid sequences having the ipMSDB presentation score higher than 20 as determined by bestWTPeptideCount l for the HLA-I and bestWTPeptideCount II for the HLA-II.
47. The method of any of claims 37-44, wherein immunogenicity in IEDB is determined based on whether an amino acid sequence has been reported as immunogenic in IEDB.