Methods and systems for predicting peptide presentation by major histocompatibility complex molecules

A machine learning model processes peptide and MHC data to accurately predict neoantigen interactions and immunogenicity, improving the selection of tumor-specific antigens for personalized cancer vaccines, reducing false positives and enhancing therapeutic efficacy.

JP2026500160APending Publication Date: 2026-01-06GENENTECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025532477
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-05
Filing Date
2023-12-04
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

Existing methods fail to accurately identify tumor-specific neoantigens that can induce a strong immune response, leading to ineffective or harmful immune reactions from therapeutic proteins, and there is a need for personalized cancer vaccines that target these antigens effectively.

Method used

A machine learning model processes peptide and MHC data separately to predict interactions, affinities, and immunogenicity, using transformer stages and protein language models to generate composite representations for improved accuracy in selecting neoantigens for vaccines.

Benefits of technology

The model reduces false positives and improves the selection of neoantigens for personalized cancer vaccines, enhancing the immune response against tumors while minimizing harmful immune reactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026500160000001_ABST
    Figure 2026500160000001_ABST
Patent Text Reader

Abstract

The present disclosure relates to immunology, particularly to methods for predicting whether a therapeutic protein is likely to elicit an immunogenic response. An exemplary method for predicting amino acid-immune protein complex (IPC) interactions may include accessing a set of amino acid sequences, accessing immune protein complex (IPC) sequences identified for a subject IPC, processing the set of amino acid sequence representations to generate a set of transformed amino acid sequence representations based on a set of element concentration scores representing binding cores of the set of amino acid sequence representations, processing the IPC sequence representations to generate the transformed IPC sequence representations, generating a composite representation, and determining one or more predicted amino acid-IPC interactions based on the composite representations.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 430,297, filed December 5, 2022, entitled "PREDICTION OF PEPTIDE PRESENTION BY MAJOR HISTOCOMPATIBILITY COMPLEX MOLECULES," the entire contents of which are incorporated herein by reference.

[0002] The present disclosure relates generally to immunology, and in particular to methods for predicting whether a neoantigen or therapeutic protein is likely to elicit an immunogenic response. [Background technology]

[0003] Therapeutic proteins are a type of pharmaceutical (biologic) derived from biological sources (e.g., animal, plant, fungal, or microbial cells). Many therapeutic proteins, including monoclonal antibodies and soluble receptors, are now produced using recombinant DNA technology. This contrasts with small molecule drugs, which are simpler compounds typically produced by chemical synthesis.

[0004] Some therapeutic proteins have the same primary amino acid sequence as native human proteins, which typically do not elicit an immune response. However, it is frequently desirable to modify the amino acid sequence of therapeutic proteins to optimize various properties, such as potency, stability, and bioavailability. However, therapeutic proteins with significantly different amino acid sequences from those of the intended recipient (e.g., a human subject) may be recognized as foreign by the recipient's immune system, thereby eliciting an immune response similar to that elicited by an antigen (toxin or foreign substance). Thus, apparently foreign therapeutic proteins are recognized as vaccines rather than pharmaceutical compounds. While it is essential for a vaccine to effectively elicit an immune response, it is harmful for a therapeutic protein to elicit an immune response. If a therapeutic protein elicits an immune response, it may render the therapeutic protein ineffective and / or cause a harmful immune response in the recipient. In particular, repeated administration of immunogenic therapeutic proteins typically elicits anti-drug antibodies (ADAs). At best, immunogenic therapeutic proteins are neutralized by the recipient's ADAs, thereby reducing their pharmaceutical activity. At worst, immunogenic therapeutic proteins can induce hypersensitivity reactions in the recipient, which can be life-threatening.

[0005] Human leukocyte antigens (HLA) are expressed as cell surface receptors that present antigenic peptides to T cells in a restricted manner, allowing for the discrimination between self and foreign antigens. The HLA complex is a complex of genes on chromosome 6 that encode cell surface proteins involved in regulating the immune system. The human major histocompatibility complex (MHC) is a reservoir of HLA genes that play a fundamental role in transplant acceptance. The MHC contains many of the genes involved in cell-mediated immune defense. The MHC complex encodes the α chains of the MHC class I molecules HLA-A, HLA-B, and HLA-C (alleles) and the α and β chains of the MHC class II molecules HLA-DR, HLA-DP, and HLA-DQ (allotypes), all of which are expressed in a codominant manner.

[0006] Activation of helper T (Th) cells upon recognition of peptide fragments (epitopes) of protein antigens bound to MHC class II (MHC-II) molecules of antigen-presenting cells results in the development of antigen-reactive antibodies. When the antigen is a therapeutic protein, the antibodies are ADAs. Peptide binding to MHC molecules with sufficient affinity is a prerequisite for immunogenicity (the ability of a peptide to elicit an immune response). Therefore, it would be desirable to predict the MHC-II epitopes of candidate therapeutic proteins and identify amino acid residues in the protein that can be safely altered to eliminate the MHC-II epitopes of the candidate therapeutic protein and thereby reduce its immunogenicity. Peptide-MHC binding affinity is primarily determined by the core amino acid sequence (typically 9 amino acids long) of the peptide. However, the amino-terminal flanking (N-flank) and / or carboxy-terminal flanking (C-flank) sequences can also affect peptide-MHC binding.

[0007] Tumors, like the subjects they affect, are heterogeneous. In particular, the somatic mutations that cause cells to become cancerous vary even between tumors derived from the same cell type. Furthermore, although humans are predicted to share 99.9% of their genomes, even 0.1% differences are significant, especially with regard to the immune system. Therefore, therapeutic cancer vaccines are ideally designed as personalized cancer vaccines.

[0008] Neoantigens are tumor-specific antigens that arise from somatic mutations in the genome of tumor cells. Peptide fragments (epitopes) of protein neoantigens bind to major histocompatibility complex (MHC) molecules expressed on the surface of target cancer cells and antigen-presenting cells, where they can activate CD8+ cytotoxic T lymphocytes (CTLs) and CD4+ helper T (Th) cells, respectively. Neoantigen vaccines are a promising approach for personalized cancer therapy because they can direct target T cells to recognize and attack neoantigen-expressing cancer cells while sparing healthy cells.

[0009] A subject's tumor profile can be determined by determining DNA and / or RNA sequences from tumor cells obtained from a biopsy. From the subject-specific tumor profile, neoantigens of interest that are present in tumor cells but not in healthy cells can be identified. However, a significant portion of the mutant sequences detected in tumor cells correspond to neoantigens that are poorly expressed, do not contain epitopes presented by MHC molecules, and / or are otherwise not bound by T cell receptors (TCRs). Such neoantigens cannot induce an immune response by CD8+ CTLs in the case of MHC class I (MHC-I) molecules, or by CD4+ Th cells in the case of MHC class II (MHC-II) molecules. As a result, such neoantigens are poor candidates for inclusion in personalized cancer vaccines to generate tumor-specific immune responses.

[0010] There are methods for predicting peptide binding to MHC molecules.However, simply identifying the peptide fragments of neoantigens that can bind to MHC molecules is not sufficient to identify the neoantigens to be included in personalized cancer vaccines.This is because many peptide binders will be false positives in that they do not effectively prime cellular immune response.Therefore, what is needed in the art is a method for accurately identifying the epitopes of neoantigens that are presented on the surface of tumors, so that they can induce strong immune response, which will help select the peptide fragments of neoantigens to be included in therapeutic cancer vaccines. Summary of the Invention

[0011] Disclosed herein are systems, methods, and programming for using machine learning models to determine predictions of whether and / or to what extent a peptide interacts with an MHC molecule. The machine learning model can perform a combination of peptide data processing, nFlank / cFlank data processing, MHC data processing, TCR data processing, and / or protein data processing to obtain one or more interaction predictions (e.g., predicted peptide interactions with MHC molecules), one or more interaction affinity predictions (e.g., predicted binding affinity between a peptide and an MHC), and / or one or more immunogenicity predictions (predictions of the ability of a peptide to elicit an immune response in relation to an MHC) as described herein. In an exemplary workflow, MHC data processing can include processing BOS-tokenized MHC sequences using one or more transformer stages, or can include processing MHC sequence embeddings generated using a protein language model. In an exemplary workflow, peptide data processing can include processing BOS-tokenized peptide sequences using one or more transformer stages, or can include processing non-BOS-tokenized peptide sequences using a cross-attention module. The exemplary workflow may further incorporate protein data processing, or may not incorporate processed protein data. Exemplary workflows may or may not incorporate additional processing of TCR data.

[0012] In some embodiments, the machine learning model processes a set of amino acid sequence representations and immune protein complex (IPC) sequence representations in parallel using separate processing blocks. The sequence representations can be feature embeddings of the corresponding sequences. The machine learning model uses a set of element concentration scores representing the binding cores of the set of amino acid sequence representations to combine the BOS token embeddings of the converted amino acid sequence representations with the BOS token embeddings of the converted IPC sequence representations to generate a composite representation. The composite representation is used to determine one or more predicted amino acid-IPC interactions, such as an interaction affinity prediction that predicts the binding affinity between a peptide and an MHC, an interaction prediction that predicts whether the MHC will present the peptide on the cell surface, or an immunogenicity prediction that predicts the ability of a peptide to elicit an immune response in relation to the MHC.

[0013] Some embodiments include accessing a set of amino acid sequences. Each of the amino acid sequences in the set may have been identified from at least one protein. Identified immune protein complex (IPC) sequences for the IPC of interest may be accessed. One or more processing blocks of a processing subsystem of the machine learning model may be used to process the set of amino acid sequence representations to generate a set of transformed amino acid sequence representations based on a set of element concentration scores representing binding cores of the set of amino acid sequence representations. Each of the amino acid sequence representations may have been generated based on one of the amino acid sequences to which a start of sequence (BOS) token has been added. A second processing block of the processing subsystem may be used to process the IPC sequence representations to generate the transformed IPC sequence representations. The IPC sequence representations may have been generated based on the identified IPC sequences to which a BOS token has been added. The set of amino acid sequence representations and the IPC sequence representations may be processed in parallel. The system may generate a composite representation by combining each of the transformed BOS token representations of the set of transformed amino acid sequence representations with the transformed BOS token representation of the transformed IPC sequence representation. The system may determine one or more predicted amino acid-IPC interactions based on the composite representations.

[0014] In some embodiments, the set of converted amino acid sequence representations comprises a set of amino-terminal flanking (N-flank) representations or a set of carboxy-terminal flanking (C-flank) representations.

[0015] In some embodiments, the IPC of interest is a major histocompatibility complex (MHC).

[0016] In some embodiments, the set of converted amino acid sequence representations comprises a set of MHC binding representations.

[0017] In some embodiments, the MHC comprises MHC class II (MHC-II).

[0018] In some embodiments, the MHC comprises MHC class I (MHC-I).

[0019] In some embodiments, the IPC of interest is a T cell receptor (TCR).

[0020] In some embodiments, at least one protein can be a therapeutic protein.

[0021] In some embodiments, the at least one protein is present in a disease sample from the subject.

[0022] In some embodiments, the disease sample may be a tumor cell biopsy. In some embodiments, the disease sample comprises cancer. In some embodiments, the disease sample comprises tissue.

[0023] In some embodiments, generating the composite representation includes, for each of the set of converted amino acid sequence representations, element-wise multiplying a converted amino acid start of sequence (BOS) representation corresponding to the converted amino acid sequence representation by a converted IPC start of sequence (BOS) representation corresponding to the converted IPC sequence representation.

[0024] In some embodiments, processing the set of amino acid sequence representations comprises processing start of sequence (BOS) representations to generate transformed peptide sequence representations.

[0025] In some embodiments, processing the IPC sequence representation comprises processing an MHC sequence starting representation (BOS) to generate a converted MHC sequence representation.

[0026] Some additional embodiments include performing, for each of a set of IPC sequences, the steps of accessing the IPC sequences, processing the set of amino acid sequence representations, processing the IPC sequence representations with the respective IPC sequences, and generating a composite representation. The determined one or more predicted amino acid-IPC interactions can be based on the composite representation corresponding to the set of IPC sequences.

[0027] In some embodiments, the set of IPC sequences comprises 12 major histocompatibility complex (MHC) allotypes of the subject and / or 6 MHC alleles of the subject.

[0028] Some additional embodiments include selecting one or more peptide-IPC combinations from the set of amino acid sequences and the set of IPC sequences for inclusion as targets for immunotherapy based on one or more determined amino acid IPC interactions.

[0029] In some embodiments, processing the set of amino acid sequence representations includes converting the set of amino acid sequence representations into a set of converted amino acid sequence representations using one or more first processing blocks, each of which includes a set of processing sub-blocks.

[0030] Some additional aspects include embedding a set of amino acid sequences to generate a set of embedded amino acid sequence representations, and positionally encoding the set of embedded amino acid sequence representations.

[0031] In some embodiments, processing the IPC array representation includes converting the IPC array representation to a converted IPC array representation using a second processing block, the second processing block including a set of processing sub-blocks.

[0032] Some additional aspects include embedding an IPC array to generate an embedded IPC array representation, and positionally encoding the embedded IPC array representation.

[0033] In some embodiments, the machine learning model includes one or more transformer encoders, each of the one or more transformer encoders including a processing layer.

[0034] In some embodiments, each of the one or more first or second processing blocks includes a set of processing sub-blocks, and each of the set of processing sub-blocks includes a neural network including at least one processing layer.

[0035] In some embodiments, the set of amino acid sequence representations comprises an aggregate sequence representation, which comprises a set of peptide representations and one or more of a set of amino-terminal flanking (N-flank) representations or a set of carboxy-terminal flanking (C-flank) representations.

[0036] Some additional aspects include flattening the aggregated sequence representations into a single array and densifying the aggregated sequence representations by removing empty rows from the array before generating the set of transformed amino acid sequence representations, wherein the transformed amino acid sequence representations are generated based on the densified aggregated sequence representations.

[0037] In some embodiments, the IPC sequence representation comprises an aggregate sequence representation, and the primary aggregate sequence representation comprises one or more of a major histocompatibility complex (MHC) sequence representation or a T cell receptor (TCR) sequence representation.

[0038] In some embodiments, processing the set of amino acid sequence representations includes, for each amino acid sequence representation of the set, determining, for each element of the amino acid sequence representation, a plurality of vectors based on a set of weights associated with a processing layer of the machine learning model, and generating a set of element concentration scores based on the plurality of vectors and the set of weights.

[0039] In some embodiments, the plurality of vectors includes a key vector, a value vector, and a query vector, and the set of weights includes a set of key weights, a set of value weights, and a set of query weights.

[0040] In some embodiments, generating the set of element concentration scores includes determining a respective element concentration score from each pair of elements from the query vector and the key vector.

[0041] In some embodiments, one or more of the first processing blocks and the second processing blocks include an attention block, each attention block includes a set of attention sub-blocks, each attention sub-block includes a self-attention layer, and the machine learning model is an attention-based machine learning model.

[0042] Some additional aspects further include generating, by one or more of the attention blocks, an attention map that includes a mask that limits attention applied by the attention sub-block to the sequence length according to the mask.

[0043] Some additional aspects include using a fully connected block of an output subsystem of the machine learning model to process the synthetic representation to generate a first output, using a dropout block of the output subsystem of the machine learning model to apply dropout to the first output to generate a second output, and using a maximum layer of the output subsystem of the machine learning model to select a subset of the second output to generate a result, and determining one or more predicted amino acid-IPC interactions based on the result.

[0044] In some embodiments, the set of amino acid sequences comprises peptide sequences, the IPC sequences comprise major histocompatibility complex (MHC) sequences, and the one or more predicted amino acid-IPC interactions include one or more of an interaction affinity prediction for the peptide-IPC combination that predicts the binding affinity between the peptide and the MHC, an interaction prediction for the peptide-IPC combination that predicts whether the MHC will present the peptide on the cell surface, or an immunogenicity prediction for the peptide-IPC combination that predicts the ability of the peptide to elicit an immune response in relation to the MHC.

[0045] In some embodiments, determining one or more predicted amino acid-IPC interactions includes processing the composite representation to generate a set of results, and selecting an amino acid-IPC combination based on the highest result in the set of results.

[0046] In some embodiments, one or more predicted amino acid-IPC interactions comprise a prediction of tumor-specific immunogenicity of the peptide.

[0047] In some embodiments, the set of amino acid sequences comprises a set of peptide sequences, and a subset of peptide sequences is identified that has one or more predicted amino acid-IPC interactions that have increased tumor-specific immunogenicity or increased likelihood of presentation by an IPC compared to the set of peptide sequences.

[0048] Some additional embodiments include identifying a subset of peptides from the set of amino acid sequences for inclusion in a personalized vaccine based on the determined one or more predicted amino acid-IPC interactions.

[0049] Some additional aspects include generating treatment recommendations, including personalized vaccines.

[0050] Some additional embodiments include selecting a subset of peptides from the set of amino acid sequences for inclusion as immunotherapeutic targets based on the determined one or more predicted amino acid-IPC interactions.

[0051] In some embodiments, the immunotherapy comprises one or more of T cell therapy, personalized cancer therapy, antigen-specific immunotherapy, antigen-dependent immunotherapy, a vaccine, or natural killer (NK) cell therapy.

[0052] Some additional embodiments include selecting a subset of peptides from the set of amino acid sequences for exclusion as targets for immunotherapy based on the determined one or more predicted amino acid-IPC interactions.

[0053] In some embodiments, the immunotherapy comprises one or more of T cell therapy, personalized cancer therapy, antigen-specific immunotherapy, antigen-dependent immunotherapy, a vaccine, or natural killer (NK) cell therapy.

[0054] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods disclosed herein.

[0055] In some embodiments, a computer program product is provided that is tangibly embodied in a non-transitory machine-readable storage medium and that includes instructions configured to cause one or more data processors to perform some or all of one or more of the methods disclosed herein.

[0056] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium including instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium, the computer program product including instructions configured to cause one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein.

[0057] The terms and expressions which have been employed are used as terms of description rather than of limitation, and there is no intention in the use of such terms and expressions to exclude equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, although the invention as claimed has been specifically disclosed by embodiments and optional features, it is to be understood that modifications and variations of the concepts disclosed herein may be resorted to by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims.

[0058] The present disclosure is described in conjunction with the accompanying drawings, in which: [Brief explanation of the drawings]

[0059] [Figure 1] FIG. 1 is a block diagram of an exemplary forecasting system, according to some embodiments.

[0060] [Figure 2] 1 is a flowchart of an exemplary process for predicting amino acid-immune protein complex (IPC) interactions using a machine learning model, according to some embodiments.

[0061] [Figure 3] FIG. 1 is a schematic diagram of an exemplary configuration of a machine learning model, according to some embodiments.

[0062] [Figure 4A] FIG. 1 is a diagram of an exemplary workflow for obtaining one or more interaction predictions, one or more interaction affinity predictions, and / or one or more immunogenicity predictions, according to some embodiments.

[0063] [Figure 4B] FIG. 1 is a diagram of an exemplary workflow for obtaining one or more interaction predictions, one or more interaction affinity predictions, and / or one or more immunogenicity predictions, according to some embodiments.

[0064] [Figure 4C] FIG. 1 is a diagram of an exemplary workflow for obtaining one or more interaction predictions, one or more interaction affinity predictions, and / or one or more immunogenicity predictions, according to some embodiments.

[0065] [Figure 4D] FIG. 1 is a diagram of an exemplary workflow for obtaining one or more interaction predictions, one or more interaction affinity predictions, and / or one or more immunogenicity predictions, according to some embodiments.

[0066] [Figure 4E] 1 illustrates an exemplary protein language model for generating protein sequence embeddings, according to some embodiments.

[0067] [Figure 4F]1 illustrates an exemplary protein language model for generating MHC sequence embeddings, according to some embodiments.

[0068] [Figure 4G] FIG. 1 is a diagram of an exemplary workflow for predicting multiple peptide interactions with MHC molecules potentially expressed by multiple alleles or allotypes, according to some embodiments.

[0069] [Figure 4H] FIG. 1 is a diagram of an exemplary workflow for attention masking, according to some embodiments.

[0070] [Figure 4I] FIG. 1 illustrates an exemplary workflow for attention masking with a calibration step to reduce model bias, according to some embodiments.

[0071] [Figure 5A] FIG. 1 is a schematic diagram of an exemplary machine learning model, according to some embodiments. [Figure 5B] FIG. 1 is a schematic diagram of an exemplary machine learning model, according to some embodiments. [Figure 5C] FIG. 1 is a schematic diagram of an exemplary machine learning model, according to some embodiments.

[0072] [Figure 6] FIG. 2 is a schematic diagram of an exemplary processing block, according to some embodiments.

[0073] [Figure 7A] 1 is a flowchart of an exemplary process for processing an array representation using an exemplary processing layer, according to some embodiments.

[0074] [Figure 7B] FIG. 1 is a schematic diagram illustrating an exemplary process for processing an array representation using an exemplary processing layer, according to some embodiments.

[0075] [Figure 8] 1 is a flowchart of an exemplary process for generating information regarding the immunological activity of various peptides, according to some embodiments.

[0076] [Figure 9] 1 is a flowchart of an exemplary process for generating information regarding the immunological activity of various peptides, according to some embodiments.

[0077] [Figure 10] 1 is a flowchart of an exemplary process for training a machine learning model and using the trained machine learning model to generate predictions for amino acids and IPCs, according to some embodiments.

[0078] [Figure 11] 1 is an illustration including examples of training data, according to some embodiments, in which the IPCs are MHC class II allotypes (HLA-DR, HLA-DP, and HLA-DQ allotypes).

[0079] [Figure 12] 1 is an exemplary method for predicting which therapeutic antibodies are likely to increase immunogenicity risk.

[0080] [Figure 13] 1 is an illustration of exemplary neo-antigen candidates (mutated antigens) and corresponding potential neo-epitope candidates (mutated peptides), according to some embodiments.

[0081] [Figure 14A] 1 includes exemplary plots showing the performance of an exemplary P-MHC-II model, according to some embodiments.

[0082] [Figure 14B]1 includes exemplary plots showing the performance of a previously used approach, Model A, on its elution output, according to some embodiments.

[0083] [Figure 15] 1 is an exemplary plot comparing exemplary average accuracy values ​​of the elution-ligand output of a previously used approach, Model A, and an exemplary P-MHC-II model for each allotype in a test dataset, according to some embodiments.

[0084] [Figure 16A] 1 is an exemplary plot showing the performance of an exemplary P-MHC-II model, according to some embodiments. [Figure 16B] 1 is an exemplary plot showing the performance of an exemplary previously used approach, Model A, in accordance with some embodiments.

[0085] [Figure 17] 1 shows a plot of a latent space containing multiple peptide vectors, according to some embodiments.

[0086] [Figure 18] 1 shows a histogram illustrating counts of peptides with different levels of information content in an exemplary dataset, according to some embodiments.

[0087] [Figure 19A] 1 shows protein space colored by protein expression, according to some embodiments.

[0088] [Figure 19B] 1 illustrates the cellular compartmentalization of different proteins and where they appear in latent space, according to some embodiments.

[0089] [Figure 20] 1 illustrates exemplary performance data according to some embodiments.

[0090] [Figure 21] FIG. 1 is a block diagram of a computer system according to some embodiments.

[0091] [Figure 22] FIG. 17 is a block diagram of an exemplary artificial intelligence (AI) architecture included as part of the exemplary computing system of FIG. 16, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0092] In the accompanying drawings, similar components and / or features may have the same reference label. Furthermore, various components of the same type may be distinguished by following the reference label with a dash and a second label that distinguishes between the similar components. If only a first reference label is used herein, the description is applicable to any one of the similar components having the same first reference label, regardless of the second reference label.

[0093] Recognizing the importance of being able to predict which mutant peptides (e.g., neoantigens) to select as candidates for personalized vaccines, the embodiments described herein provide methodologies and systems for making more accurate predictions than currently available methods and systems. The embodiments described herein use machine learning methodologies and systems to improve predictive performance, for example, but not by way of limitation, by reducing the number of false positives generated when analyzing mutant peptide sequences to determine the viability of mutant peptides as vaccine candidates. The embodiments described herein also provide methodologies and systems for determining whether a particular therapeutic antibody may pose an immunogenic risk to a subject.

[0094] Disclosed herein are systems, methods, and programming for obtaining one or more interaction predictions (e.g., predicted peptide interactions with MHC molecules), one or more interaction affinity predictions (e.g., predicted binding affinity between a peptide and an MHC), and / or one or more immunogenicity predictions (predictions of the ability of a peptide to elicit an immune response in relation to an MHC) as described herein. The machine learning model can perform a combination of peptide data processing, nFlank / cFlank data processing, MHC data processing, TCR data processing, and / or protein data processing to predict peptide interactions with MHC molecules. In an exemplary workflow, processing of MHC data can include processing of BOS-tokenized MHC sequences using one or more transformer stages, or can include processing of MHC sequence embeddings generated using a protein language model. In an exemplary workflow, processing of peptide data can include processing of BOS-tokenized peptide sequences using one or more transformer stages, or can include processing of non-BOS-tokenized peptide sequences using a cross-attention module. An exemplary workflow may further incorporate processing of protein data, or may not incorporate processing protein data. Exemplary workflows may or may not incorporate additional processing of TCR data.

[0095] For example, embodiments described herein provide various methodologies, machine learning models, and / or outputs generated by machine learning models that use machine learning models to analyze sequences identified from disease samples from subjects. To predict whether and / or to what extent a mutant peptide identified from a disease sample interacts with an IPC, such as an MHC molecule (e.g., MHC-I or MHC-II), the machine learning model processes a set of amino acid sequence representations separately from processing the IPC sequence representations (e.g., MHC sequence representations). The sequence representations may be embeddings of corresponding sequence features. In some embodiments, the mutant peptide sequence representations may be referred to as variant coding sequences. The IPC sequence (e.g., MHC sequence) may include at least a portion of an MHC molecule, the complete sequence, a pseudosequence of the MHC molecule (including the portion that interacts with the mutant peptide (including the binding pocket, any other portion including the pseudosequence, etc.)).

[0096] A machine learning model includes various subsystems for processing. For example, a machine learning model may include a representation subsystem, a processing subsystem, a composite subsystem, and an output subsystem. Each "subsystem" may include one or more blocks, and each block may include one or more sub-blocks and / or layers. A sub-block may include any number of layers (or units).

[0097] The processing subsystem can be used to generate one or more transformed sequence representations, such as a set of transformed amino acid sequence representations (e.g., which may include variant coding sequences), a transformed IPC sequence representation, etc. In some embodiments, the processing subsystem can process one or more (e.g., a set of) amino acid sequence representations independently of or separately (e.g., in parallel) from the IPC sequence representation. For example, the set of amino acid sequence representations can be processed using one or more first processing blocks of the processing subsystem, and the IPC sequence representations can be processed using a second processing block of the processing subsystem. Processing the set of amino acid sequence representations and the IPC sequence representations through these parallel and / or separate processing blocks can improve the predictive performance of machine learning models. Separate processing engines can enable the system to learn separate representations of different biological structures. In contrast, processing different biological structures through a shared processing engine can lead to overfitting of the model.

[0098] Furthermore, the embodiments described herein recognize and take into account that training a model corresponding to a series of biological events may require significantly more data than training a model corresponding to a single biological event. Training a model for sequence analysis can be particularly complex due to the sheer number of potentially observable sequences. Not only are there millions of potential neoantigens, but the genes encoding proteins for, for example, MHC class II molecules are also highly polymorphic. In fact, nearly 6,000 variant alpha and beta chain proteins of HLA-DR, HLA-DQ, and HLA-DP (the three classical class II molecules in humans) currently exist in the IPD-IMGT / HLA database. Therefore, the embodiments described herein provide methodologies and systems for training machine learning models that reduce training complexity and improve training performance. For example, variant coding sequences used for training can be selected and / or trimmed so that training is performed using variant coding sequences having an amino acid length of a threshold amino acid length (e.g., 9 amino acids) or less. Generating a training dataset that includes variant coding sequences having lengths equal to or less than a threshold amino acid length can reduce the overall complexity of training and improve training and / or prediction performance (e.g., reduce the change in performance metrics from epoch to epoch, thereby improving prediction performance).

[0099] Thus, the technology disclosed herein includes a machine learning-based approach for determining predicted amino acid-IPC interactions associated with immunological activity associated with a peptide, such as a mutant peptide. The machine learning model may generate an output including one or more predicted amino acid-IPC interactions. For example, the output may include one or more interaction predictions, one or more interaction affinity predictions, or one or more immunogenicity predictions (i.e., predictions regarding the ability of the peptide to elicit an immune response). The interaction prediction may include a prediction regarding whether a peptide (e.g., a mutant peptide comprising a given ordered set of amino acids identified by a given variant coding sequence) will experience one or more target interactions. In some embodiments, the target interaction may be binding of the peptide to an IPC (e.g., an MHC molecule, a TCR), a peptide presented by an MHC molecule on a cell surface, or another type of target interaction. The interaction affinity prediction may include a prediction of affinity for one or more target interactions. For example, the interaction affinity prediction may indicate binding affinity for peptide-MHC binding. Interaction (e.g., binding) affinity can be determined based on the propensity, strength, and / or stability of the interaction (e.g., binding). Immunogenicity prediction refers to predicting the ability of a peptide to induce an immune response. The immune system recognizes peptides as non-self or foreign. Once recognized, the peptide stimulates the immune system to produce a response. This response can include the production of antibodies by B cells (humoral immunity) and the activation of T cells to eliminate cells presenting the peptide (cell-mediated immunity).

[0100] In some embodiments, the output may include or indicate the immunogenicity of the peptide or therapeutic antibody. For example, the output may predict whether the peptide will elicit an immune response in a particular subject or group of subjects. These immunogenicity predictions may be determined for each of a plurality of mutant peptides. In some embodiments, the immunogenicity predictions may be used to select or rank one or more mutant peptides and / or pharmaceutical compositions for inclusion in a vaccine and / or use in a treatment. For example, but not limited to, mutant peptides associated with high predicted binding affinity, a high probability of being presented on the surface of tumor cells, and / or high predicted immunogenicity may be selected for inclusion in a vaccine or use in a treatment.

[0101]

[0004] Embodiments described herein provide methods and systems for using machine learning models to determine predicted amino acid-IPC interactions that exhibit immunological activity associated with amino acids (e.g., peptides) and immune protein complexes (IPCs). The IPCs may include MHC or TCRs. A set of identified amino acid sequences from at least one protein may be accessed. In some embodiments, the at least one protein is a therapeutic protein. In some embodiments, the at least one protein is present in a disease sample from a subject. IPC sequences may be identified and then accessed for the subject's IPC. The set of amino acid sequence representations may be processed using one or more first processing blocks of a processing subsystem of the machine learning model. Processing of the set of amino acid sequence representations may generate a set of transformed amino acid sequence representations. The IPC sequence representations may be processed using a second processing block of the processing subsystem to generate transformed IPC sequence representations. In some embodiments, the BOS token embeddings of each of the transformed set of amino acid sequence representations may be combined with the BOS token embeddings of the transformed IPC sequence representations to generate a composite representation. The method may generate an output including predicted amino acid-IPC interactions determined based on the composite representation. The results (output) may include one or more interaction predictions, one or more interaction affinity predictions, or one or more immunogenicity predictions for the corresponding amino acid-IPC combinations. In some embodiments, a report is generated based on the output.

[0102] The techniques described herein offer numerous technical advantages. For example, the techniques described herein can avoid or reduce overfitting. Overfitting is an undesirable machine learning behavior that occurs when a machine learning model provides accurate outputs for training data but not for new data. When implementing a machine learning model such as a transformer, providing insufficient training data regarding peptide binding to MHC alleles can result in overfitting. A typical transformer can have a large number of parameters learned by the model from the training data. In particular, data regarding peptide binding to MHC alleles can be difficult to obtain for several reasons, including a) the high variability of MHC molecules, b) the complexity of peptide binding, c) the diversity of subjects expressing MHC, d) ethical constraints, and e) other technical limitations that require specialized equipment and expertise. Collecting training data for other IPC machine learning model predictions, such as TCRs, can face similar challenges.

[0103] Some embodiments describe computational techniques for training machine learning models in a way that eliminates overfitting caused by insufficient training data. Such techniques involve the use of protein language models (PLMs) to process IPC alleles, e.g., MHC alleles, and infer useful, generalizable amino acid or residue features that can be used in the training process of the machine learning model. This reduces the likelihood of overfitting and improves the performance of the machine learning model. For example, when IPC allele sequences are processed by a PLM (e.g., to generate training data), the trained machine learning model can learn which residues or amino acids in peptide-MHC binding complexes are associated with or affect such binding.

[0104] In some examples, the techniques described herein can advantageously incorporate information about the source protein from which the peptide originates. Without incorporating information about the source protein, the model may lack useful protein / peptide features for generating the predictions described herein. By incorporating information about the source protein (e.g., via protein sequence embedding generated by PLM), the system can capture features such as processing signals (i.e., frank signals generated when enzymes break down proteins into peptides), expression of genes associated with the source protein, cellular localization of the source protein, and other pertinent features that influence peptide presentation by MHC molecules and other interactions predicted by the implementations described herein, allowing the system to select peptides that are more likely to be presented by MHC molecules. In some examples, source protein expression can be important because a source protein that is not sufficiently expressed in a subject will not yield sufficient peptides, even if such peptides may induce immunogenic effects. Cellular localization of the source protein is another feature that can be obtained by processing the source protein with PLM. MHC1 presentation may occur more frequently when the source protein is located within or generated within the cell. On the other hand, proteins that are primarily extracellular (e.g., proteins present only in vesicles) may be less likely to be presented by MHC class I because they are located in different parts of the cell. Conversely, proteins that are primarily extracellular (e.g., proteins present only in vesicles) may be more likely to be presented by MHC class II. Thus, for MHC I, the source protein is preferably derived from an intracellular source, and for MHC II, vesicular or extracellular proteins are preferred. In summary, there are many characteristics associated with source proteins that affect peptide presentation, and these are advantageously incorporated into the technology described herein.

[0105] In some embodiments, the techniques described herein can advantageously utilize a cross-attention module to incorporate MHC allele-specific binding information. MHC allele information can be useful when processing peptide data. For example, a peptide may contain multiple binding cores (i.e., specific peptide regions where binding occurs) that can bind to different MHC alleles. The system can use the cross-attention module to incorporate MHC allele-specific vectors into the processing of peptide data. In this way, the system can predict binding core information from peptides with multiple binding cores that bind to different MHC alleles and generate multiple binding core predictions corresponding to multiple different alleles. In other words, the techniques presented herein are not limited to predicting one binding core per peptide. In some implementations, the machine learning model can predict two or more binding cores per peptide.

[0106] In some instances, the techniques described herein can advantageously prevent over-prediction of peptide amino acid positions that indicate the start of a binding core. For example, position 0 (i.e., the start position) is unlikely to be the start of a peptide's binding core because the binding core is part (e.g., a 9-mer) of a peptide (e.g., a 20-mer) and is expected to be surrounded by the flanks of the binding core within the peptide. To mitigate the problem of over-prediction, the techniques described herein can include a calibration step to remove bias toward position 0 or other peptide positions that are subject to over-prediction. Before performing the calibration step, the system can first calculate the model bias for any single position in the peptide. The system can perform the calibration step by modifying the attention map by subtracting the model bias to remove the model bias for any single position in the peptide.

[0107] In some embodiments, the techniques described herein can advantageously use a dimensionality reduction module to process MHC data (e.g., MHC sequence embeddings) to prevent overfitting. For example, overfitting can occur when different types of models, such as neural networks, are trained to reduce the dimensionality of the input embeddings. If a neural network is trained using the same set of alleles used to train a PLM model, typically a relatively small set of MHC alleles, the neural network model may generate correct outputs when processing input embeddings corresponding to MHC alleles in the training data, but may generate incorrect outputs when processing input embeddings corresponding to new MHC alleles not included in the training data. On the other hand, PCA techniques can be trained using a larger set of MHC alleles (e.g., the space of all alleles) and thus can efficiently and accurately reduce the dimensionality of input embeddings for new alleles not in the training dataset, thereby avoiding the overfitting problem.

[0108] In the following description, it should be noted that the embodiments presented herein are not limited to the specific advantages disclosed above. The present disclosure encompasses various technical advantages and enhancements detailed throughout this written description. The embodiments are presented in a non-limiting manner, with the understanding that numerous modifications, variations, and improvements can be made without departing from the spirit and scope of the present invention. The following description provides additional technical advantages and novel aspects inherent in the presented embodiments, thereby providing a broader perspective on the applicability and usefulness of the disclosed invention.

[0109] The following description provides exemplary embodiments of these methods and systems in which the output (e.g., predicted amino acid-IPC interactions) can be used to plan, design, and / or manufacture therapeutics.

[0110] Exemplary Prediction System 1 is a block diagram of an exemplary prediction system, according to some embodiments. The prediction system 100 is used to determine predicted amino acid-IPC interactions associated with the immunological activity of peptides, particularly mutant peptides. The prediction system 100 includes a computing platform 102, a data store 104, and a display system 106. The computing platform 102 can take a variety of forms. In some embodiments, the computing platform 102 includes a single computer (or computer system) or multiple computers in communication with each other. In some embodiments, the computing platform 102 can be a cloud computing platform.

[0111] The data store 104 and the display system 106 each communicate with the computing platform 102. In some examples, one or more of the data store 104 or the display system 106 may be considered part of or otherwise integrated with the computing platform 102. Thus, in some examples, the computing platform 102, the data store 104, and the display system 106 may be separate components that communicate with each other, while in other examples, some combination of these components may be integrated together. Communication between the different components may be implemented using any number of wired, wireless, or optical communication links, or a combination thereof.

[0112] The prediction system 100 includes a sequence analyzer 108, which may be implemented using hardware, software, firmware, or a combination thereof. In some embodiments, the sequence analyzer 108 is implemented on the computing platform 102. The sequence analyzer 108 receives sequence data 110 for processing. For example, the sequence data 110 may be sent as input to the sequence analyzer 108, retrieved from or accessed from the data store 104 or some other type of storage device (e.g., cloud storage), or obtained in some other manner. In some cases, the sequence data 110 may be retrieved from the data store 104 in response to receiving user input entered by a user via an input device.

[0113] The sequence data 110 may be generated from processing a set of samples 112. The set of samples 112 may take the form of one or more biological samples (e.g., diseased samples, healthy samples, or a combination thereof) from one or more subjects. The set of samples 112 may include samples obtained from a tumor in a subject. The tumor may be, for example, a symptom of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, stomach cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, T-cell lymphocytic leukemia, non-small cell lung cancer, small cell lung cancer, or a combination thereof.

[0114] The samples in the set of samples 112 may include, for example, different IPC molecules, different peptides, or a combination thereof. If the set of samples 112 includes disease samples, the peptides may include one or more mutant peptides (e.g., neoantigens). The IPC molecules may include, for example, different MHC molecules, different TCR molecules, or a combination thereof.

[0115] In some embodiments, the set of samples 112 includes immune protein complexes (IPCs) 114 (e.g., MHC class I molecules, MHC class II molecules, various TCR molecules, etc.). Additionally, the set of samples can include at least one protein 123 (i.e., a source protein). The amino acid chain 116 can be identified from the at least one protein 123 and can be a chain of amino acids including a peptide 118, an N-flank 120, and a C-flank 122. The amino acid chain 116 can include or exclude the N-terminus between the peptide 118 and the N-flank 120, or the C-terminus between the peptide 118 and the C-flank 122. The peptide 118 is considered a mutant peptide if the peptide 118 contains one or more variants (e.g., one or more sequence changes) when compared to a corresponding reference sequence. The protein 123 is a source protein for the amino acid chain 116, which can be generated by proteolysis, a process in which a protein (e.g., protein 123) is broken down into smaller polypeptides or amino acids. Protein 123 can be broken down into smaller polypeptides or amino acids by enzymatic cleavage, where certain enzymes called proteases cleave the peptide bonds between the amino acids of protein 123.

[0116] The set of samples 112 can be processed to generate sequence data 110. In some embodiments, multiple samples of the set of samples 112 can be processed at different times. In some embodiments, the prediction system 100 includes a sample analyzer used in processing the set of samples 112 to generate the sequence data 110. The sequence data 110 includes, for example, at least one amino acid sequence 129 and at least one immune protein complex (IPC) sequence 124 (e.g., one IPC sequence 124 corresponding to the IPC 114). The amino acid sequence 129 can include one or more of a peptide sequence 126 (e.g., one peptide sequence 126 corresponding to the peptide 118), an amino-terminal flanking (N-flank) sequence 128 (e.g., one N-flank sequence 128 corresponding to the N-flank 120), or a carboxy-terminal flanking (C-flank) sequence 130 (e.g., one C-flank sequence 130 corresponding to the C-flank 122). One or more subsequences of amino acid sequence 129 (eg, peptide sequence 126, N-flank sequence 128, and C-flank sequence 130) can be treated separately or as a single sequence.

[0117] If the immune protein complex 114 is an MHC, the IPC sequence 124 can be, for example, an MHC sequence 135 that characterizes at least a portion of the MHC. If the immune protein complex 114 is a TCR, the IPC sequence 124 can be, for example, a TCR sequence 131 that characterizes at least a portion of the TCR. In some embodiments, the IPC sequence 124 can include both an MHC sequence 135 that characterizes at least a portion of an MHC molecule and a TCR sequence 131 that characterizes at least a portion of a TCR molecule. In some embodiments, the sequence data 110 can include an IPC sequence 124 in the form of an MHC sequence 135 that characterizes at least a portion of an MHC molecule, as well as a separate TCR sequence 131 that characterizes at least a portion of a TCR.

[0118] Protein sequence 160 characterizes at least a portion of protein 123. In some embodiments, protein sequence 160 can be identified by performing a reverse lookup in a database (e.g., the UniProt database) based on variant peptide data (e.g., IPC sequence 124) obtained from a sample.

[0119] Peptide sequence 126 characterizes at least a portion of peptide 118. N-flank sequence 128 characterizes at least a portion of N-flank 120. In some embodiments, if there are a large number of amino acids (or amino acid residues) upstream from the N-terminus, the corresponding sequence of N-flank 120 can be trimmed to generate N-flank sequence 128. C-flank sequence 130 characterizes at least a portion of C-flank 122. In some embodiments, if there are a large number of amino acids (or amino acid residues) downstream from the C-terminus, the corresponding sequence of C-flank 122 can be trimmed to generate C-flank sequence 130.

[0120] The sequence analyzer 108 receives sequence data 110 as input for processing. The sequence analyzer 108 includes a machine learning model 132 that processes the sequence data 110. In some embodiments, the sequence data 110 is sent directly to the machine learning model 132 for processing. In some embodiments, the sequence analyzer 108 preprocesses the sequence data 110 before sending it to the machine learning model 132 for processing. Preprocessing the sequence data 110 may include adding a start of sequence (BOS) token to each of multiple sequences in the sequence data 110. The BOS token added to a peptide sequence may serve as an additional data structure that can be used to represent peptide characteristics that can be used to determine the presentation potential, binding affinity, or immunogenicity prediction of the corresponding peptide sequence 126. The BOS token may indicate interaction information such as whether the peptide is presented by an allele / allotype, binding affinity, immunogenicity, or any other suitable purpose for which the machine learning model 132 is trained.

[0121] The machine learning model 132 can be implemented in any of several different ways. In some embodiments, the machine learning model 132 can be any type of model that uses a set of element concentration scores that represent the binding core of a set of amino acid sequence representations. The machine learning model 132 can be used in either a training mode or a prediction mode. In the training mode, the machine learning model 132 is trained using training data 133. Examples of the training module 133 are described in more detail below. The machine learning model 132 is trained for use in a prediction mode.

[0122] The machine learning model 132 processes the IPC sequences 124 via an IPC processing engine 134 and the amino acid sequences 129 via an amino acid processing engine 139. Separate processing engines for the IPC and amino acids allow for improved predictive performance of the machine learning model 132. In some embodiments, the machine learning model 132 processes one or more of the MHC sequences 135 via an MHC processing engine 141 or the TCR sequences 131 via a TCR processing engine 142. In some embodiments, the machine learning model 132 processes one or more of the peptide sequences 126 via a peptide processing engine 136, the N-flank sequences 128 via an N-flank processing engine 138, or the C-flank sequences 130 via a C-flank processing engine 140. In some embodiments, the machine learning model 132 processes the protein sequences 160 via a protein processing engine 162. Examples of these different processing engines are described in more detail below.

[0123] As used herein, the terms "processing engine" and "engine" identify at least one software component and / or a combination of at least one software component and at least one hardware component that is designed / programmed / configured to interact and / or communicate data with other software and / or hardware components, including but not limited to other processing engines.

[0124] The machine learning model 132 processes the sequence data 110 to generate output that is used to generate a report 144. The report 144 may include the exact output of the machine learning model 132, a transformed or filtered version of the output, or both. In some cases, the report 144 may include notifications, recommendations, warnings, or other information generated by the sequence analyzer 108 based on the output of the machine learning model 132.

[0125] Report 144 may be an output including, for example, information regarding immunological activity of interest for one or more peptides (e.g., one or more mutant peptides). For example, report 144 may include information regarding immunological activity for amino acids 116 (e.g., peptide 118, N-flank 120, C-flank 122, etc.) and IPC 114 (e.g., MHC-I, MHC-II, TCR, etc.). Report 144 may include, for example, interaction information 146 (e.g., an interaction affinity prediction that predicts binding affinity between the peptide and an MHC, or an interaction prediction that predicts whether an MHC allele or allotype will present the peptide on a cell surface), immunogenicity information 148 (e.g., an immunogenicity prediction that predicts the ability of the peptide to elicit an immune response in relation to the MHC), or both. Interaction information 146 may provide predictions regarding a selected set of interactions between amino acids 116 and IPC 114. Immunogenicity information 148 may provide predictions regarding the immunogenicity of amino acids 116 (e.g., including the immunogenicity of peptide 118).

[0126] In some embodiments, the report 144 may be displayed on a graphical user interface (GUI) 150 of the display system 106. A user may view and / or interact with the report 144 via the graphical user interface 150. In some embodiments, a user may use the report 144 to make a decision regarding the treatment of a subject from whom at least one of the sets of samples 112 was obtained (or collected).

[0127] In some embodiments, the prediction system 100 sends (e.g., wirelessly) the report 144 to a remote system 152. The remote system 152 may be a cloud computing platform, cloud storage, another computer system, a user device (e.g., a smartphone, tablet, laptop, etc.), or some other type of platform. In some embodiments, the remote system 152 may be or be part of a treatment generation system (or machine).

[0128] 2 is a flowchart of an exemplary process (e.g., a computer-implemented method) for predicting amino acid-IPC interactions using a machine learning model, according to some embodiments. Process 200 can be implemented using the prediction system 100 described in FIG. 1. For example, process 200 can be implemented using the sequence analyzer 108 and machine learning model 132 of FIG. 1.

[0129] Process 200 may include, for example, step 202. Step 202 includes adding a BOS token to each amino acid sequence of a set of amino acid sequences. Each of the amino acid sequences of the set may be identified from at least one protein. In some embodiments, the at least one protein is a therapeutic protein. In some embodiments, the at least one protein is present in a disease sample from a subject. As one non-limiting example, the disease sample may be a tumor cell biopsy. Additionally or alternatively, the disease sample may include cancer, tissue, or both. In some embodiments, each of the amino acid sequences includes one or more of an amino-terminal flanking (N-flank) sequence or a carboxy-terminal flanking (C-flank) sequence.

[0130] Step 204 includes adding a BOS token to the IPC sequence identified for the IPC of interest. In some embodiments, the IPC of interest is an MHC. The MHC may include MHC class II (MHC-II) or MHC class I (MHC-I). In some embodiments, the IPC of interest is a TCR.

[0131] Step 206 includes generating a set of amino acid sequence representations (e.g., embeddings) for each of the amino acid sequences, and generating a set of IPC sequence representations (e.g., embeddings) from the set of IPC sequences. The embedding module may generate each of the sequence representations by creating an embedding for each element (e.g., a single amino acid or a single nucleic acid) of the sequence to represent the features of the element as a low-dimensional feature vector. The embedding module may also generate positional embeddings as part of the sequence representation, representing each absolute position (corresponding to one of the amino acids or one of the nucleic acids). The embeddings corresponding to the BOS tokens may have a length equal to the number of features corresponding to each individual sequence element represented in the respective sequence to which the BOS token is attached.

[0132] Step 208 includes processing the set of amino acid sequence representations using one or more processing blocks of the processing subsystem of the machine learning model. Processing the set of amino acid sequence representations through one or more converter stages may generate a set of converted amino acid sequence representations stored in embeddings corresponding to BOS tokens. Sequences with simple molecular structures (e.g., N-flank / C-flank sequences) may require only a single converter stage, while sequences with more complex molecular structures (e.g., peptide sequences) may require multiple converter stages. The converter stages calculate information about the frequency of correlation between each pair of amino acids in the input amino acid sequence representation and store information about the pairwise correlations in the converter output. The converter stages may thereby store information about the pairwise correlations between amino acids, as well as the absolute position of each amino acid within each amino acid sequence (e.g., from the positional embedding), in embeddings corresponding to BOS tokens in the converted amino acid sequence representations. The set of converted amino acid sequence representations may include a set of MHC-binding representations. The converted set of amino acid sequence representations may include one or more of a set of amino-terminal flanking (N-flank) sequence representations, a set of carboxy-terminal flanking (C-flank) sequence representations, or a set of combined N-flank / C-flank sequence representations. The N-flank / C-flank sequence representations may be processed in parallel with the amino acid sequence representations by using separate processing blocks. Each processing block may utilize an attention mechanism to determine an element concentration score.

[0133] Step 210 includes processing the IPC array representation using a second processing block of the processing subsystem, which may operate independently and in parallel with the first processing block, to generate a transformed IPC array representation stored in an embedding corresponding to the BOS token. This allows information about each IPC array to be stored in an embedding corresponding to the BOS token in the transformed IPC array representation. In some embodiments, the IPC array may include a TCR array in addition to an MHC array; in such embodiments, the TCR array representation may be generated separately from the MHC array representation, and the two array representations may be processed in parallel using independent processing blocks. Each processing block may utilize an attention mechanism to determine an element concentration score.

[0134] Step 212 involves generating a composite representation by combining each of the BOS token embeddings for each of the converted amino acid sequence representations with the BOS token embeddings for the converted IPC sequence representations. The combining process can be element-wise multiplication, element-wise addition, or calculating a dot product. In some embodiments, if the amino acid sequences include both peptide sequences and N-flank and / or C-flank sequences, the BOS token embeddings corresponding to the peptide sequences can be concatenated with the BOS token embeddings corresponding to the N-flank / C-flank sequence representations prior to the combining step.

[0135] Step 214 includes determining predicted amino acid-IPC interactions based on the composite representation. The predicted amino acid-IPC interactions may include one or more of one or more interaction predictions, one or more interaction affinity predictions, or one or more immunogenicity predictions for one or more corresponding amino acid-IPC combinations.

[0136] In some embodiments, an output (e.g., a report) can be generated. The output can be based on the predicted amino acid-IPC interactions. The output can be used to facilitate the design and / or manufacture of vaccines, treatments, and / or treatment regimens. For example, the report can identify a subset of the set of peptides (contained in the set of amino acid sequences), or can provide an indication of which peptides of the subset of peptides should be selected for use in devising a treatment for the subject. The treatment can be, for example, the subset of peptides, the precursors of each of the subset of peptides, or some other form.

[0137] Example architecture of a machine learning model The machine learning model 132 of the embodiments described herein can include multiple subsystems (or sub-networks). Each of the multiple subsystems can include an encoder, a transformer-encoder, and / or one or more processing layers. In some cases, the machine learning model 132 can be configured to learn alignments (e.g., between amino acid sequences and IPC sequences, between peptide sequences and MHC sequences, between MHC sequences and TCR sequences, between MHC-peptide complexes and TCRs, and peptide sequences and TCR sequences). The alignments can be learned and performed using alignment score functions, such as, for example, content-based functions, additive functions, position-based functions, dot product functions, and / or scaled dot product functions.

[0138] The machine learning model 132 may include, for example, one or more encoders configured to transform elements of the input sequence (e.g., amino acid sequences, nucleic acid sequences, codon sequences, etc.) based on other elements of the input sequence. The encoder may be a transformer encoder.

[0139] The machine learning model 132 may include one or more processing layers, such as a self-attention layer or a convolutional layer, or a neural network, such as a long-short-term memory unit (LSTM), a recurrent structure, or a recurrent component. The machine learning model 132 may, for example, implement one or more self-attention layers. The machine learning model 132 may use a self-attention mechanism, a global attention mechanism, a soft attention mechanism, a local attention mechanism, and / or a hard attention mechanism. In some cases, the machine learning model 132 does not include a convolutional layer, a recurrent structure, an LSTM unit, and / or a recurrent component. In some cases, the machine learning model 132 is not a regression machine learning model and / or does not include a recurrent neural network. In some cases, the machine learning model 132 includes a recurrent neural network and / or may use positional encoding to process a sliding window of array elements across one or more arrays. In some cases, the machine learning model 132 is not a convolutional machine learning model and / or does not include a convolutional neural network.

[0140] The machine learning model 132 may include processing blocks, such as one or more first processing blocks used to process one or more amino acid sequence representations, independent of a second processing block used to process the IPC sequence representations. In some embodiments, the second processing block may process some or all of the IPC sequence representations (e.g., MHC pseudosequences). The independence of these processing blocks can facilitate parallel processing when using the machine learning model 132. Furthermore, the independence can improve the performance (e.g., accuracy of predictions) of the machine learning model 132.

[0141] The machine learning model 132 can be configured such that the output value at any given layer depends not only on the corresponding input value but also on one or more (e.g., all) other input values. Thus, the machine learning model 132, loss function, and / or optimization function can be configured to optimize an output corresponding to a single position that represents the degree to which a given IPC (e.g., an MHC molecule) (represented by a corresponding input) binds to a given peptide (represented by another corresponding input) and / or elicits immunogenicity in response to the given peptide. In some cases, the loss function can include a supervised loss function, such as binary cross-entropy, or an unsupervised loss function. In some examples, the unsupervised loss function can include a contrastive loss or a regularization loss (e.g., an L1 / L2 loss applied to a peptide representation to make residue changes more continuous). In some cases, the loss function can include an auxiliary loss function. In such cases, the auxiliary loss function can be used together with the main loss function to train the machine learning model 132. Thus, in some cases, the auxiliary loss function can improve the learning process by adding additional information or constraints. In some cases, any of multiple outputs of the transformer encoder may represent such a probability of occurrence. The machine learning model 132 can be trained accordingly. In some cases, the endpoints (e.g., residual endpoints) may represent the probability or likelihood of binding affinity, presentation (eluted ligand or EL), and / or immunogenicity (in response to training). The aggregated output can be fed, for example, to another layer, subsystem, or processing block (e.g., including one or more processing layers, such as a self-attention layer, or encoders, such as a transformer encoder).

[0142] In some cases, one, two, or all dimensions of the output from another layer and / or another subsystem or processing block are the same size as the input provided to the other layer and / or other subsystem or processing block. In some cases, the input provided to this other layer and / or other subsystem or processing block has a length along one axis that is equal to or greater than the sum of one or more of the number of amino acids in the IPC sequence, the number of amino acids in the peptide sequence, the number of amino acids in the N-flank sequence, or the number of amino acids in the C-flank. In some cases, the length of the input is one longer than the total number of amino acids. For example, if additional feature vectors (e.g., feature vectors corresponding to BOS tokens or tokens representing IPC types) are added to the amino acid-specific feature values, the length of the input along one axis may exceed the total number of amino acids. Another dimension of the input can include several features (e.g., determined via hyperparameters). The output generated by the other layer and / or other subsystem or processing block can have the same size as the input.

[0143] A subset of the output values ​​generated by other layers and / or other sub-networks can be processed by another neural network (e.g., a fully connected feedforward network). The subset of values ​​can include a one-dimensional vector of values ​​that can correspond to one set of feature values. For example, the one-dimensional vector can correspond to feature values ​​associated with BOS tokens.

[0144] In some embodiments, the neural network in the machine learning model 132 can be configured to output one or more results. The one or more results can include, for example, a numeric result, a binary result, and / or a categorical result. Each of the one or more results can predict whether and / or to what extent an amino acid (e.g., a peptide) and an IPC undergo a particular type of reaction (e.g., combine together). The machine learning model 132 can include one or more activation layers to generate intermediate results (e.g., to convert real, provisional values ​​into binary and / or categorical outputs). The machine learning model 132 can be trained to generate multiple types of predictions (e.g., interaction predictions, interaction affinity predictions, and / or immunogenicity predictions). In some cases, the predictions can be binary or categorical. Other predictions can be non-binary or non-categorical. For example, the predictions can be scalar.

[0145] The machine learning model 132 may include and / or be contained within an ensemble model, which may include multiple (e.g., identical) sub-models that can be trained using different portions of a training dataset.

[0146] Example configuration of a machine learning model 3 is a schematic diagram of an exemplary configuration of the machine learning model 132 of FIG. 1 , according to some embodiments. With continued reference to FIG. 1 , the machine learning model 132 will be described. The machine learning model 132 may have a configuration 300 that includes a representation subsystem 302, a processing subsystem 304, a composition subsystem 306, and an output subsystem 310. One or more subsystems within the machine learning model 132 may include one or more blocks, one or more sub-blocks, one or more layers, or a combination thereof. One or more blocks within the machine learning model 132 may include one or more sub-blocks, one or more layers, or a combination thereof. One or more sub-blocks of the machine learning model 132 may include one or more layers (or units).

[0147] In some embodiments, the representation subsystem 302 receives the sequence data 110 as input and passes it through a tokenizer (not shown), which converts each character of the sequence into a token (e.g., a unique integer associated with the character and stored in a lookup table). The tokenizer can append a BOS token to one or more of the sequences in the sequence data 110. The representation subsystem 302 can generate a sequence representation for each of the sequences in the sequence data 110. The sequence representation can include, for example, a stack of feature vectors corresponding to a sequence of sequence elements, each sequence element representing or identifying one or more amino acids, one or more nucleic acids of the sequence corresponding to the sequence representation, or a BOS token. For example, each amino acid in a sequence can be represented by a unique feature vector, and a BOS token appended to an amino acid sequence can likewise be represented by a unique feature vector. The IPC sequence representation may include 6 to 12 MHC alleles / allotypes for a given subject, and the processing subsystem 304 may generate up to 12 transformed MHC sequence representations, corresponding to one MHC transformed sequence representation for each MHC allotype combined with a particular peptide sequence. The amino acid sequences 129 may include one or more of the peptide sequences 126, the n-flank sequences 128, or the c-flank sequences 130, and the processing subsystem 304 may generate one amino acid sequence representation for each of the amino acid sequences. To normalize across different sequence lengths, the stack of feature vectors can be padded to a standard sequence length (SL), for example, 39 vectors.

[0148] The processing subsystem 304 may receive one or more sequence representations (e.g., a set of amino acid sequence representations, an IPC sequence representation) as input, process these sequence representations through one or more converter stages, and generate converted sequence representations (e.g., a set of converted amino acid sequence representations, a converted IPC sequence representation) that are sent to the combining subsystem 306. The processing subsystem 304 includes one or more processing blocks. A processing block may include one or more processing layers (e.g., an attention layer). The various converter stages of the processing subsystem 304 may thereby store information about each sequence in embeddings that correspond to BOS tokens within the sequence representations.

[0149] In some embodiments, the representation subsystem 302 and / or the processing subsystem 304 can be configured to process subsequences of an amino acid sequence in parallel using independent processing engines. For example, the peptide processing engine 136 of Figure 1 may include (1) the representation subsystem 302 processing the peptide sequence 126 with the BOS tokens appended to it to generate a BOS+ (plus) peptide sequence representation 312, followed by (2) a processing block 314 of the processing subsystem 304 processing the BOS+ peptide sequence representation 312 to generate a converted peptide sequence representation 316. In this example, the N-flank processing engine 138 of Figure 1 can run independently in parallel by (1) the representation subsystem 302 processing the N-flank sequence 128 with the BOS tokens appended to it to generate a BOS+N-flank sequence representation 324, followed by (2) a processing block 326 of the processing subsystem 304 processing the BOS+N-flank sequence representation 324 to generate a converted N-flank sequence representation 328. In this example, C-flank processing engine 140 of FIG. 1 can execute independently and in parallel by (1) representation subsystem 302 processing C-flank array 130 with BOS tokens appended to it to generate BOS+C-flank array representation 330, followed by (2) processing block 332 of processing subsystem 304 processing BOS+C-flank array representation 330 to generate transformed C-flank array representation 334.

[0150] Similarly, the representation subsystem 302 and the processing subsystem 304 can be configured to process subsequences (e.g., MHC sequence, TCR sequence) of the IPC sequence in parallel using independent processing engines. For example, the MHC processing engine 141 of Figure 1 may include (1) the representation subsystem 302 processing the MHC sequence 135 with the BOS token appended to generate a BOS+MHC sequence representation 318, followed by (2) a processing block 320 of the processing subsystem 304 processing the BOS+MHC sequence representation 318 to generate a transformed MHC sequence representation 322. In this example, the TCR processing engine 142 of Figure 1 can execute independently in parallel by (1) the representation subsystem 302 processing the TCR sequence 131 with the BOS token appended to generate a BOS+TCR sequence representation 336, followed by (2) a processing block 338 of the processing subsystem 304 processing the BOS+TCR sequence representation 336 to generate a transformed TCR sequence representation 340.

[0151] In some embodiments, particular subarray representations generated by the representation subsystem 302 may be combined before adding a BOS token. For example, a C-flank representation may be added to an N-flank array representation before adding a single BOS token to the combined array representation. In such embodiments, the processing subsystem 304 that processes the combined array representations may generate a single transformed array representation accordingly.

[0152] Combining subsystem 306 generates composite representation 342 by combining embeddings corresponding to BOS tokens of the converted amino acid sequence representation with embeddings corresponding to each BOS token of the set of converted IPC sequence representations. Combining subsystem 306 may combine embeddings corresponding to BOS tokens of the converted peptide sequence representation 316 with a set of embeddings corresponding to BOS tokens of the converted MHC sequence representation 322 to generate composite representation 342. In some embodiments, the combining process may involve element-wise multiplication of the converted peptide BOS token embeddings by the set of converted MHC BOS token embeddings. Element-wise multiplication is beneficial because it causes the latent spaces of the various components to cluster into complementary regions of the latent space. One result is that peptides that bind to a particular MHC will result in embeddings in the same region (which has a similar effect on the TCR and peptide sequence). In some embodiments, the combining process may involve element-wise addition, or calculation of the dot product, of the converted peptide BOS token embeddings and the set of converted MHC BOS token embeddings.

[0153] Embodiments of the present disclosure may include a set of converted amino acid sequence representations including one or more of: set of converted N-flank sequence representations 328, set of converted peptide sequence representations 316, or set of converted C-flank sequence representations 334. In some embodiments, the set of converted IPC sequence representations may include one or more of: set of converted MHC sequence representations 322 or set of converted TCR sequence representations 340. Combining subsystem 306 combines the BOS token embeddings of the converted amino acid sequence representations with one or more (including all) of the BOS token embeddings of the converted IPC sequence representations. The combining step may include multiplication or addition of BOS token embeddings, such as element-wise multiplication of the converted amino acid BOS token embeddings (including one or more of the converted peptide BOS token embeddings, the converted N-flank BOS token embeddings, or the converted C-flank BOS token embeddings) by the converted IPC BOS token embeddings (including one or more of the converted peptide MHC BOS token embeddings and the converted TCR BOS token embeddings). In some embodiments, the combining step may involve computing dot products and / or element-wise additions instead of, or in addition to, element-wise multiplications.

[0154] The composite representations 342 can be processed by the output subsystem 310, and predicted amino acid-IPC interactions can be determined based on the composite representations 342. In some embodiments, the predicted amino acid-IPC interactions can be based on one or more composite representations selected from among the composite representations 342. In some embodiments, a report 144 can be generated that includes the selected one or more predicted amino acid-IPC interactions.

[0155] 4A-4D show example workflows illustrating various possible combinations of peptide data processing, nFlank / cFlank data processing, MHC data processing, TCR data processing, and / or protein data processing to obtain one or more interaction predictions, one or more interaction affinity predictions, and / or one or more immunogenicity predictions, according to some embodiments. The predictions may include, among other things, an interaction affinity prediction for a peptide-IPC combination that predicts the binding affinity between the peptide and an MHC molecule; an interaction prediction for the peptide-IPC combination, e.g., a prediction of whether the MHC molecule will present the peptide on the surface of a cell; or an immunogenicity prediction for the peptide-IPC combination that predicts the ability of the peptide to elicit an immune response. As described in more detail below, processing of MHC data may include processing of BOS-tokenized MHC sequences using one or more transformer stages (e.g., FIG. 4A ) or processing of MHC sequence embeddings generated using a protein language model (e.g., FIG. 4B , FIG. 4C , FIG. 4D ). Processing of peptide data can include processing of BOS-tokenized peptide sequences using one or more transformer stages (e.g., Figures 4A and 4B) or processing of non-BOS-tokenized peptide sequences using a cross-attention module (e.g., Figures 4C and 4D). Optionally, any workflow may further incorporate processing of protein data (e.g., Figure 4B) or may not incorporate processed protein data (e.g., Figures 4A, 4C, and 4D). Optionally, any workflow may further incorporate processing of TCR data (e.g., Figure 4A) or may not incorporate TCR data (e.g., Figures 4B, 4C, and 4D).

[0156] It should be understood that the workflows in Figures 4A, 4B, 4C, and 4D are merely examples, and the present disclosure encompasses any workflow that combines processing of MHC data with processing of peptide data to obtain one or more interaction predictions, one or more interaction affinity predictions, and / or one or more immunogenicity predictions. In exemplary workflows, processing of MHC data can include processing of BOS-tokenized MHC sequences using one or more transformer stages (e.g., Figure 4A) or processing of MHC sequence embeddings generated using a protein language model (e.g., Figures 4B, 4C, and 4D). In exemplary workflows, processing of peptide data can include processing of BOS-tokenized peptide sequences using one or more transformer stages (e.g., Figures 4A and 4B) or processing of non-BOS-tokenized peptide sequences using a cross-attention module (e.g., Figures 4C and 4D). Exemplary workflows may further incorporate processing of protein data (e.g., Figure 4B) or may not incorporate processed protein data (e.g., Figures 4A, 4C, and 4D). Exemplary workflows may further incorporate processing of TCR data (e.g., Figure 4A) or may not incorporate TCR data (e.g., Figures 4B, 4C, 4D).

[0157] 4A shows an exemplary workflow 400 for obtaining one or more interaction predictions, one or more interaction affinity predictions, and / or one or more immunogenicity predictions, according to some embodiments. A tokenizer (not shown) can create tokenized sequences including a BOS-tokenized peptide sequence 402a, a BOS-tokenized nFlank+cFlank sequence 402b, a BOS-tokenized MHC sequence 402c, and a BOS-tokenized TCR sequence 402d. An embedding module can generate sequence representations. Specifically, the embedding module can generate (e.g., via the representation subsystem 302) a BOS+peptide sequence representation 404a (which may correspond to 312 in FIG. 3 ) based on the BOS tokenized peptide sequence 402a, a BOS+nFlank+cFlank sequence representation 404b representing a combined version of the BOS+N-Flank sequence representation (which may correspond to 324 in FIG. 3 ) and the BOS+C-Flank sequence representation (which may correspond to 330 in FIG. 3 ) based on the BOS tokenized nFlank+cFlank sequence 402b, a BOS+MHC sequence representation 404c (which may correspond to 318 in FIG. 3 ) based on the BOS tokenized MHC sequence 402c, and a BOS+TCR sequence representation 404d (which may correspond to 336) based on the BOS tokenized TCR sequence 402d. Transformed versions 406a-d of each of the sequence representations 404a-d (e.g., transformed BOS+peptide sequence representations 316 and 406a, transformed BOS+nFlank+cFlank sequence representation 406b representing a combined version of BOS+N-Flank sequence representation 328 and BOS+C-Flank sequence representation 334, transformed BOS+MHC sequence representations 322 and 406c, or transformed BOS+TCR sequence representations 340 and 406d) can be generated using one or more transformer stages of the processing subsystem 304.During the transformer stage, BOS token embeddings 408a-d are also generated as part of the transformed sequence representations 406a-d (e.g., BOS token embedding 408a for the transformed BOS+peptide sequence representations 316 and 406a, BOS token embedding 408b for the transformed BOS+nFlank+cFlank sequence representations 406b, BOS token embedding 408c for the transformed BOS+MHC sequence representations 322 and 406c, or BOS token embedding 408d for the transformed BOS+TCR sequence representations 340 and 406d). The transformed BOS token embeddings 408a-d extract information about the sequence (e.g., pairwise correlation and positional information) within the embeddings corresponding to the BOS tokens added to the sequence. Each of the transformed BOS token embeddings 408a-d can represent the entire sequence via a single vector rather than multiple vectors, thus making the sequence easier to interpret. A composite of BOS token embeddings 408a-408d may then be generated by the composite subsystem 306, as described in Figure 3, and then the final output generated by the output subsystem 310, also described in Figure 3. Each of the sequence representations (e.g., embeddings) 404a-404d and the transformed sequence representations 406a-406d may be of uniform dimensions, including a subsequence length (SL), a vector length (VL), and a batch size (BS). As shown in Figure 4A, the dimensions of BOS token embeddings 408a-d have a subsequence length equal to 1 because they were generated from a single BOS token appended to each amino acid or IPC subsequence.

[0158] As shown in Figure 4A, in peptide processing engine 136 (see Figure 1), BOS tokenized peptide sequence 402a is created by a tokenizer (e.g., in representation subsystem 302) and then used by an embedding module (e.g., in representation subsystem 302) to generate peptide sequence representation 404a, from which several transformation stages result in transformed peptide sequence representation 406a, including transformed BOS token embedding 408a. In a parallel processing engine (e.g., combining N-flank processing engine 138 and C-flank processing engine 140), the N-flank and C-flank sequences can be appended with a single BOS token to create combined flank sequence 402b. After flank sequence representation 404b is generated by the embedding module, transformed flank sequence representation 406b is generated by processing subsystem 304, including transformed BOS token embedding 408b. As shown, transformed BOS token 408a is first combined with transformed BOS token embedding 408b, and then the combined result is combined with transformed BOS token 408c and transformed BOS token embedding 408d. In each of the two combining operations, element-wise addition, element-wise multiplication, concatenation, or dot product multiplication may be used.

[0159] 1 uses a BOS token-appended MHC sequence 402c (e.g., MHC sequence 135 of IPC sequence 124 corresponding to an allele for MHC class I or corresponding to an allotype for MHC class II) to generate an MHC sequence representation 404c. During the conversion stage, a converted MHC sequence representation 406c including a converted BOS token embedding 408c is then generated. Also in parallel, a BOS token-appended TCR sequence 402d (e.g., TCR sequence 131 of IPC sequence 124) can be used to generate a TCR sequence representation 404d, from which a converted TCR sequence representation 406d including a converted BOS token embedding 408d is generated.

[0160] The transformed BOS token embeddings can then be combined together by the combining subsystem 306 of FIG. 3 (e.g., by concatenating 408a and 408b and applying element-wise multiplication to the concatenated 408a+408b, 408c, and 408d) to yield a binary prediction of binding affinity, presentability, immunogenicity prediction, or another downstream evaluation, and the product vector is passed through a multi-layer perceptron network (e.g., output subsystem 310).

[0161] 4B shows an exemplary workflow 420 for obtaining one or more interaction predictions, one or more interaction affinity predictions, and / or one or more immunogenicity predictions, according to some embodiments. Similar to workflow 400 of FIG. 4A, a tokenizer (not shown) can create tokenized sequences including a BOS-tokenized peptide sequence 402a and a BOS-tokenized nFlank+cFlank sequence 402b. An embedding module (not shown) can generate (e.g., via representation subsystem 302 of FIG. 3 ) a BOS+peptide sequence representation 404a (which may correspond to 312 of FIG. 3 ) based on the BOS-tokenized peptide sequence 402a, and a BOS+nFlank+cFlank sequence representation 404b (which may correspond to 324 of FIG. 3 ) based on the BOS-tokenized nFlank+cFlank sequence 402b, representing a combined version of the BOS+N-Flank sequence representation (which may correspond to 324 of FIG. 3 ) and the BOS+C-Flank sequence representation (which may correspond to 330 of FIG. 3 ). One or more converter stages of the system (e.g., in the processing subsystem 304) can then create a converted BOS+peptide sequence representation 406a and a converted BOS+nFlank+cFlank sequence representation 406b. During the converter stages, the system can further generate BOS token embeddings as part of the converted sequence representations 406a and 406b, including a BOS token embedding 408a for the converted BOS+peptide sequence representation 406a and a BOS token embedding 408b for the converted BOS+nFlank+cFlank sequence representation 406b.

[0162] Workflow 400 in FIG. 4A does not incorporate processing of protein data (e.g., source protein 123 in FIG. 1 ) from which peptides (e.g., peptide 118 in FIG. 1 , which may be represented by 402a) and N-flank / C-flanks (e.g., N-flank 120 and C-flank 122, which may be represented by 402b) are identified. As described above with reference to FIG. 1 , protein 123 is a source protein from which amino acid chain 116 (including peptide 118 in FIG. 1 , N-flank 120 in FIG. 1 , and C-flank 122 in FIG. 1 ) can be generated by proteolysis, a process in which a protein (e.g., protein 123) is broken down into smaller polypeptides or amino acids. Protein 123 can be broken down into smaller polypeptides or amino acids by enzymatic cleavage, in which specific enzymes called proteases cleave peptide bonds between amino acids in protein 123. In contrast, the workflow 420 of Figure 4B may further include a protein processing engine (e.g., protein processing engine 162 of Figure 1) for generating and processing protein sequence embeddings 422 that encapsulate information about the protein. As shown in Figure 4B, a dimensionality reduction module 424 may receive the protein sequence embeddings 422, reduce the dimensionality of the protein sequence embeddings 422, and generate an output vector. As discussed below with reference to Figure 4E, the protein sequence embeddings 422 may be generated by a protein language model (PLM) based on the protein sequence (e.g., protein sequence 160 of Figure 1).

[0163] In some examples, the dimensionality reduction module 424 can be implemented as a fully connected layer ("FCN"). An FCN can include a neural network in which each neuron applies a transformation (e.g., a linear transformation) to an input vector via a weight matrix. As a result, all possible connections exist per layer, and therefore every input in the input vector (i.e., protein sequence embedding 422) affects every output in the output vector. In addition to reducing the dimensionality of the input vector, an FCN can include a relatively large number of learnable parameters, allowing the neural network to encode useful information into the output vector. The output vector (i.e., a dimensionally reduced version of the protein sequence embedding 422) can be further aggregated as 408a and 408b. The dimensionality reduction module 424 can encode, among other things, mappings from protein features to useful features such as cellular compartmentalization and gene expression. While FIG. 4B shows that element-wise addition is used on the aggregated 408a, 408b and the output vector of the dimensionality reduction module 424, other integration methods such as element-wise averaging, element-wise multiplication, concatenation, and dot-product multiplication can be used.

[0164] Incorporating protein information advantageously allows for the incorporation of useful protein features. Because workflow 420 can capture features such as processing signals (i.e., frank signals generated when enzymes degrade proteins into peptides), the expression of genes associated with the source protein, the cellular localization of the source protein, and other pertinent features, this workflow can select peptides that are more likely to be presented by MHC molecules. For example, for MHC1, the source protein is preferably derived from an intracellular source, while for MHCII, it is preferably a vesicular or extracellular protein. In summary, there are many features associated with source proteins that affect peptide presentation, and these are advantageously incorporated into workflow 420. Furthermore, the use of a PLM can be advantageous compared to the use of a transducer. Typical transducers can have a large number of parameters. Therefore, training a transducer using a relatively small training dataset can lead to overfitting. In contrast, the use of a PLM allows the workflow to learn useful, generalizable features, thus reducing the likelihood of overfitting and improving workflow performance.

[0165] In some embodiments, the system can enable or disable the incorporation of processing of protein data (i.e., protein processing engine 162 in FIG. 1 ) depending on the use case. For example, if the system is used to develop a personalized cancer vaccine and the subject generates peptides in a manner similar to how peptides in the presentation dataset are generated, the protein processing engine can be enabled. In contrast, if the system is used to develop antibody drugs, and the PLM provides information about endogenous proteins and the antibody drugs are not endogenous, the protein processing engine can be disabled.

[0166] Workflow 420 also differs from workflow 400 of FIG. 4A in its processing of MHC data. As described above, workflow 400 includes using BOS tokenized MHC sequence 402c to generate BOS+MHC sequence representation 404c and then generating a transformed BOS+MHC sequence representation using one or more transformer stages. Workflow 400 further includes obtaining BOS token embedding 408c and aggregating BOS token embedding 408c with the remaining BOS embedded tokens. In contrast, in workflow 420 of FIG. 4B, dimensionality reduction module 428 can receive MHC sequence embedding 426, reduce the dimensionality of MHC sequence embedding 426, and generate an output vector. As described below with reference to FIG. 4F, MHC sequence embedding 426 can be generated using PLM based on an MHC sequence (e.g., MHC sequence 135 of FIG. 1).

[0167] The purpose of the dimensionality reduction module 428 is to reduce the number of parameters used to represent information about the MHC sequences. The dimensionality reduction module 428 removes or down-prioritizes less useful parameters, such as parameters that do not change from one MHC allele to another (e.g., when the variance falls below a certain threshold). For example, similar binding behaviors can be removed or down-prioritized while different binding behaviors can be maintained. In some embodiments, the dimensionality reduction module 428 can be implemented as a principal component analysis (PCA) model. The PCA model can be configured to receive an MHC sequence embedding 426 having N vector values and generate a dimensionality-reduced version of the MHC sequence embedding 426 having M vector values (M < N). In some embodiments, to configure the PCA model, the PCA model first receives a plurality of MHC vectors corresponding to a plurality of MHC sequences (e.g., each having N vector values corresponding to N features), and ranks the N features based on how each feature varies across the plurality of MHC vectors (e.g., based on the variance associated with each feature). After the PCA model determines the ranking of the N features, the PCA model then receives an input MHC sequence embedding 426 having N vector values, rearranges the N vector values within the MHC sequence embedding 426 according to the ranking of the N features, and generates a dimensionality-reduced version of the MHC sequence embedding 426 having M vector values (M < N) by, for example, saving only the first M vector values of the rearranged N vector values. It should be understood that the dimensionality reduction module 428 can use other dimensionality reduction methods such as UMAP, t-distributed stochastic neighbor embedding, independent component analysis, multidimensional scaling, Isomap, and deep learning-based dimensionality reduction techniques.

[0168] The use of the dimensionality reduction module 428 to further process the MHC sequence embeddings 426 provides several technical advantages. For example, when different types of models, such as neural networks, are trained to reduce the dimensionality of the input embeddings, overfitting can occur. Overfitting is an undesirable machine learning behavior that occurs when a machine learning model provides accurate outputs for the training data but not for new data. If a neural network is trained using the same set of alleles used to train the PLM model, typically a relatively small set of alleles, the neural network model may generate correct outputs when processing input embeddings corresponding to alleles in the training data, but may generate incorrect outputs when processing input embeddings corresponding to new alleles not included in the training data. On the other hand, PCA techniques can be trained using a larger set of alleles (e.g., the space of all alleles) and therefore can efficiently and accurately reduce the dimensionality of input embeddings for new alleles not in the training dataset, thereby avoiding the overfitting problem.

[0169] Referring to FIG. 4B, the output of the dimensionality reduction module 428 (i.e., a reduced-dimensional version of the MHC sequence embedding 426) may optionally be provided to an FCN module 429. The FCN module 429 may include a neural network in which each neuron applies a transformation (e.g., a linear transformation) to the input vector via a weight matrix. As a result, all possible connections are present for each layer, and therefore every input in the input vector (i.e., the protein sequence embedding 422) affects every output in the output vector. The FCN module 429 may further reduce the dimensionality of the output vector of the dimensionality reduction module 428. Additionally, the FCN module 429 may encode the relationship between the MHC sequence and the type of peptide it presents. Furthermore, the FCN module 429 may encode information useful for downstream processing. For example, two MHC sequence embeddings represented by similar sequences may be close to each other in latent space due to the similarity of the sequence data, but the two MHCs may actually function differently (e.g., present different peptides). The FCN module 429 can reconcile discrepancies in the latent space so that the output vector is more useful for downstream processing (e.g., because the FCN module 429 allows encoding of information about how similar peptides and similar MHCs are present with each other). The output vector of the FCN module 429 can be aggregated with the output 410 as shown in FIG. 4B. While FIG. 4B shows that element-wise multiplication is used, other methods of integration can be used, such as element-wise averaging, element-wise addition, concatenation, and dot-product multiplication.

[0170] Figure 4C shows an exemplary workflow 470 for predicting peptide interactions with one or more MHC molecules expressed by one or more alleles or allotypes, according to some embodiments. Similar to workflow 400 of Figure 4A, a tokenizer (not shown) can create a BOS tokenized nFlank+cFlank sequence 402b. An embedding module (not shown) can generate (e.g., via representation subsystem 302 of Figure 3) a BOS+nFlank+cFlank sequence representation 404b based on the BOS tokenized nFlank+cFlank sequence 402b. One or more transformer stages of the system (e.g., in processing subsystem 304 of Figure 3) can then create a transformed BOS+nFlank+cFlank sequence representation 406b. During the transformer stage, the system can further generate a BOS token embedding 408b of the transformed BOS+nFlank+cFlank sequence representation 406b.

[0171] Unlike workflow 400 of FIG. 4A, workflow 470 does not receive a BOS-tokenized peptide sequence. Instead, workflow 470 receives a peptide sequence (e.g., peptide sequence 126 of FIG. 1 ) without BOS tokens and creates a tokenized peptide sequence 402e based on the peptide sequence. The embedding module can generate a peptide sequence representation 404e (e.g., via representation subsystem 302 of FIG. 3 ) based on the tokenized peptide sequence 402e. One or more transformer stages of the system (e.g., in processing subsystem 304 of FIG. 3 ) can then create a transformed peptide sequence representation 406e.

[0172] Additionally, workflow 470 receives BOS vector embeddings 472 (which may be generated using a PLM model based on the BOS vectors) and MHC sequence embeddings 426. The BOS vector embeddings 472 may be random vectors selected during training, e.g., intercepts representing one or more random bias terms. Dimensionality reduction component 428 may receive the MHC sequence embeddings 426 and obtain an output vector as described above with reference to FIG. 4B. The output vectors of the dimensionality reduction component 428 (e.g., with lower variance removed or deprioritized parameters) may be provided to an FCN module 473, which may further reduce the dimensionality of the vector. In an aggregator 471, workflow 470 aggregates the BOS vector embeddings 472 and the output vectors of the FCN module 473. While FIG. 4C illustrates element-wise addition in aggregator 471, other integration methods, such as element-wise averaging, element-wise multiplication, concatenation, and dot-product multiplication, may be used. Thus, the output vectors of aggregator 471 encode information from the BOS vectors and MHC sequences.

[0173] The workflow 470 further includes a cross-attention module 474. The cross-attention module 474 can be implemented as a self-attention transformer with three components: a query (Q), a key (K), and a value (V). As shown in FIG. 4C, both the K component of the cross-attention module 474 and the V component of the cross-attention module 474 are derived from the transformed peptide sequence representation 406e (i.e., implementing a self-attention mechanism). The Q component of the cross-attention module 474 is derived from the output vector of the aggregator 471 (i.e., the combined vector of the BOS vector embedding and the MHC embedding). The cross-attention module 474 can perform an attention function that includes mapping a query and a set of key-value pairs to an output, where the query, key, value, and output are all vectors. The output is calculated as a weighted sum of the values, and the weight assigned to each value is calculated by a compatibility function of the query with the corresponding key. Specifically, the cross-attention module 474 can perform scaled dot-product attention or other suitable types of integration techniques. Further details on scaled dot-product attention can be found in Ashish Vaswani et al., Attention is All You Need, Neural Information Processing Systems (2017). The configuration of workflow 470 is advantageous because it incorporates allele-specific binding information. Allele information can be useful when processing peptide data. For example, a peptide may contain multiple binding cores that may bind to different alleles. Workflow 470 uses the cross-attention module 474 to incorporate allele-specific vectors (i.e., component Q) into the processing of peptide data. In this way, the workflow can capture binding core information from peptides with multiple binding cores that bind to different alleles and generate predictions of multiple binding cores corresponding to multiple different alleles.In some instances, a peptide can have two or more element concentration scores, each element concentration score corresponding to a different allele.

[0174] 4D illustrates an alternative embodiment 476 of workflow 470. As shown, workflow 476 eliminates the BOS vector embedding 477 processing of workflow 470 and also eliminates aggregator 475, since allele information (i.e., MHC sequence embedding) is already taken into account by cross-attention module 474. Workflow 476 shares similar technical advantages as workflow 470, but by eliminating modules, it may be more computationally efficient. Additionally, workflow 476 of FIG. 4D can provide an allele-dependent latent space.

[0175] FIG. 4E shows an exemplary process for generating protein sequence embeddings, according to some embodiments. Referring to FIG. 4E, a PLM 444 can receive a protein sequence 160 and output a protein sequence embedding 422. The PLM is a machine learning model (e.g., a deep learning model) that can be based on natural language processing methods, such as attention and transformers, and can be trained on an ensemble of protein sequences. The protein language model is trained to understand and predict protein properties based on the amino acid sequences that form such proteins. In some examples, the protein language model can infer a set of features from amino acid sequences, including primary, secondary, tertiary, and quaternary structures. The PLM predicts how proteins fold, their domains, active sites, and stability. The PLM can also predict protein-protein and protein-nucleic acid interactions, post-translational modifications, and the effects of mutations. The PLM can identify localization signals within cells, understand evolutionary relationships, and predict protein function. Similarly, the PLM can provide insight into protein dynamic behavior and identify potential drug binding sites, which are beneficial for drug discovery and understanding the molecular basis of disease. In some embodiments, PLM 444 may include an evolutionary scale model (ESM model) or a variation of an ESM model, hi some embodiments, PLM 444 may include ProteinBERT, UniRep, or other suitable type of PLM.

[0176] In some embodiments, PLM 444 includes a pre-trained protein language model, such as a pre-trained ESM model. In some embodiments, input protein sequence 160 can include a sequence of amino acid residues, and PLM 444 can be configured to obtain multiple embeddings (i.e., vector representations) by obtaining a corresponding embedding for each amino acid residue. The model can be further configured to obtain a single embedding 422 by aggregating (e.g., by element-wise averaging) the multiple embeddings corresponding to the sequence of amino acid residues.

[0177] FIG. 4F shows an exemplary process for generating an MHC sequence embedding, according to some embodiments. Referring to FIG. 4F, PLM 445 can receive MHC sequence 135 and output MHC sequence embedding 426. As discussed herein, PLM is trained using a large number of proteins and can therefore encode useful information and context about the input sequence represented by the sequence embedding. In some embodiments, PLM 445 can include an evolutionary scale modeling (ESM) model or a variant of an ESM model. In some embodiments, PLM 445 can include ProteinBERT, UniRep, or the like. In some embodiments, PLM 445 includes a pre-trained protein language model, such as a pre-trained ESM model. In some embodiments, MHC sequence 135 can include multiple amino acids that make up corresponding alleles. PLM 445 can be configured to obtain multiple embeddings (i.e., vector representations) by obtaining a corresponding embedding for each amino acid. The model can further be configured to obtain a single embedding 426 by aggregating (e.g., by element-wise averaging) multiple embeddings corresponding to a sequence of amino acids.

[0178] FIG. 4G is a diagram 450 of an exemplary workflow for more efficiently predicting peptide interactions with MHC molecules that may be expressed by multiple alleles or allotypes, according to some embodiments. This technique can be applied to any of the workflows described herein. This technique can advantageously predict peptide interactions with MHC molecules when the probability of peptide binding or elution is known, but data is collected for multiple alleles or allotypes, and it is not known exactly which of the multiple alleles or allotypes bound the peptide. As shown in FIG. 4G, the set of alleles / allotypes 452 (e.g., HLA1, HLA2, HLA3, and HLA4) for which data was collected is tokenized and embedded (e.g., using the representation subsystem 302) to create one MHC sequence representation 404c per allele / allotype. Because the data for each sample for each allele / allotype may be sparse, steps can be taken (e.g., by the representation subsystem 302) to compress the MHC sequence representation. The MHC sequence representations of the detected binding interactions can be flattened into a single array 454, from which empty rows can be removed, resulting in a dense combined MHC sequence representation 456, which is then processed as usual by the processing subsystem 304 to generate binding affinity predictions 458 (e.g., during both the model training and inference stages). Once predictions 458 are generated, the model output can be resparsified (e.g., embedded 460) as needed for downstream tasks.

[0179] Figure 4H illustrates an exemplary workflow 480 for attention masking at each transformer stage of the transformed BOS+peptide sequence representation 406a, according to some embodiments. This technique can be applied to any workflow described herein, including processing of BOS-tokenized peptide sequences. The transformer stage (e.g., when performed by the processing subsystem 314 and processing subblocks 542, 594b, and 602) stores pairwise information in an attention map. To force the model to pay attention only to peptide-binding core sequences having a specified length (e.g., nine amino acids long, as shown in attention mask 482, or some other core length), masks 482 and 484 are applied to the attention map to limit the range of consecutive amino acids for which the model can record information about pairwise correlations between amino acids (e.g., considering up to nine positions ahead in the sequence, or some other core length). In the final transformer stage, the final attention map 484 may constrain the model to correspond only to BOS tokens 406a in the transformed BOS+peptide sequence representation 408a, recording the maximum value as the start of the binding core.

[0180] FIG. 4I shows exemplary attention maps used in the transformer stage, according to some embodiments. Each attention map visualizes the attention weights assigned by the transformer stage to different portions of the input. In FIG. 4I, the attention maps are represented as heat maps, with brighter colors corresponding to higher attention weights. Specifically, FIG. 4I shows exemplary attention map 483 used in the transformer stage to obtain a transformed BOS+peptide sequence representation 406a for a given peptide, and exemplary attention map 485 applied to obtain BOS tokens 408a for a given peptide. As described above with reference to FIG. 4H, attention mask 482 was applied to attention map 483 to direct the model to only peptide-binding core sequences having a specified length (e.g., 9 amino acids long, as shown in attention mask 482, or some other core length), and attention mask 484 was applied to attention map 485.

[0181] Generally, position 0 (i.e., the starting position) of a peptide is unlikely to be the binding core starting position because the binding core is a part (e.g., a 9-mer) of the peptide (e.g., a 20-mer) and is expected to be surrounded by the flanks of the binding core within the peptide. However, the attention mechanism of workflow 480 may result in over-prediction of position 0 within a peptide as the binding core starting position for a particular dataset and / or a particular workflow. To address the over-prediction problem, workflow 480 may include a calibration step to remove bias toward position 0. Before performing the calibration step, the system (e.g., computing platform 102 of FIG. 1) can first obtain a set of random peptides with a uniform length distribution. For a given peptide length, the system can calculate the average attention value at each position (e.g., position 0, position 1, position 2, etc.) across the set of random peptides (e.g., from the transformed BOS+ peptide sequence representation 406a of FIGS. 4A-4B, or the cross-attention module 474 of FIGS. 4C-4D). In other words, for a given peptide length N, the system can calculate the mean attention values ​​for various positions 0 through N-1 (i.e., the mean attention value for position 0, ..., and the mean attention value for position N-1). To the extent that these mean attention values ​​differ, they represent the bias of the model, since the probability of binding core initiation at any position should be equal and therefore the attention values ​​should be uniformly distributed (i.e., the mean attention values ​​should be identical).

[0182] Returning to Figure 4I, in workflow 480, the system can perform a calibration step by modifying attention map 485. For a given peptide with length N, the system can subtract the average attention values ​​of N positions from the N positions in attention map 485 (i.e., the average attention value of position 0, the average attention value of position 1, ..., and the average attention value of position N-1). After subtraction, the modified attention map 485 can be applied to obtain BOS token 408a for the given peptide in workflow 480. As mentioned above, the average attention value represents the bias of the model, and by subtracting the average attention value from the attention map, the calibration step removes the bias of the model toward any single position of the peptide.

[0183] 5A-5C are schematic diagrams of different configurations of a machine learning model 532, according to some embodiments. The machine learning model 532 is an example implementation of the machine learning model 132 of FIGS. 1 and 3 and the workflow 400 of FIG. 4A. The machine learning model 532 can be any type of machine learning model, including, but not limited to, an attention-based machine learning model. As shown in the schematic diagram of FIG. 5A, the machine learning model 532 includes a representation subsystem 501 (e.g., an embedding module that generates the array representation 404), a processing subsystem 503 (e.g., one or more transformer stages that generate the transformed array representation 406), a composition subsystem 505 (e.g., combines the BOS token embeddings 408 of the transformed array representation 406), and an output subsystem 509, which are example implementations of the representation subsystem 302, processing subsystem 304, composition subsystem 306, and output subsystem 310 of FIG. 3, respectively.

[0184] Representation subsystem 501 may include a BOS+peptide representation block 502 and a BOS+MHC representation block 504. In some embodiments, representation subsystem 501 further includes a BOS+N-flank representation block 506, a BOS+C-flank representation block 508, or both. In some embodiments, representation subsystem 501 further includes a BOS+TCR representation block 510. One or more (e.g., each) representation blocks may include at least one embedding layer (e.g., embedding layer 512, embedding layer 516, embedding layer 520, embedding layer 524, or embedding layer 528), e.g., at least one position encoder (e.g., position encoder 514, position encoder 518, position encoder 522, position encoder 526, or position encoder 530).

[0185] The embedding layer can embed a sequence, for example, by converting an initial non-numeric sequence representation (e.g., a series of amino acid identifiers) into a numeric sequence representation to generate an embedded representation. In some embodiments, the embedded amino acid sequence representation indicates, for each position in the sequence and for each of a set of (e.g., 21) amino acids, whether a particular amino acid is present at that position. Embedding can be performed using, for example, one-hot encoding, evolutionarily motivated encoding such as BLOSUM, randomly or pseudo-randomly initialized learned embeddings, or a combination thereof. The embedded representation can be positionally encoded to generate a coded representation. The representation generated by the representation block can be the coded representation or can be an aggregation (e.g., concatenation or union) of the coded representation and the embedded representation.

[0186] In some cases, the order of values ​​in the input data set may be useful. A positional encoder may be used and added to the embedded representation, where the positional encoding uses a learned or fixed encoding algorithm. For example, the fixed positional encoding may be determined using sine and / or cosine functions (e.g., with the position and / or dimension within the array as independent variables). The positional encoding may have the same dimensions as the coded representation. The positional encoding may be summed with the embedded representation to generate a position-directed embedded representation of the array, which is provided to the processing subsystem 503.

[0187] For example, the BOS+peptide representation block 502 may include an embedding layer 512 and a positional encoder 514. The embedding layer 512 embeds a peptide sequence (e.g., peptide sequence 126 in FIG. 1 ) to generate an embedded peptide representation, and the positional encoder 514 positionally encodes the embedded peptide representation to generate a peptide sequence representation (e.g., peptide sequence representation 312 in FIG. 3 ). The BOS+N-flank representation block 506 may include an embedding layer 520 and a positional encoder 522. The embedding layer 520 embeds an N-flank sequence (e.g., N-flank sequence 128 in FIG. 1 ) to generate an embedded N-flank representation, and the positional encoder 522 positionally encodes the embedded N-flank representation to generate an N-flank sequence representation (e.g., N-flank sequence representation 324 in FIG. 3 ). The BOS+C-flank representation block 508 may include an embedding layer 524 and a positional encoder 526. The embedding layer 524 embeds a C-flank array (e.g., C-flank array 130 of FIG. 1) to generate an embedded C-flank representation, and the positional encoder 526 position-encodes the embedded C-flank representation to generate a C-flank array representation (e.g., C-flank array representation 330 of FIG. 3).

[0188] The BOS+MHC representation block 504 may include an embedding layer 516 and a positional encoder 518. The embedding layer 516 embeds an MHC sequence (e.g., MHC sequence 135 in FIG. 1 ) to generate an embedded MHC representation, and the positional encoder 518 positionally encodes the embedded MHC representation to generate an MHC sequence representation (e.g., MHC sequence representation 318 in FIG. 3 ). The BOS+TCR representation block 510 may include an embedding layer 528 and a positional encoder 530. The embedding layer 528 embeds a TCR sequence (e.g., TCR sequence 131 in FIG. 1 ) to generate an embedded TCR representation, and the positional encoder 530 positionally encodes the embedded TCR representation to generate a TCR sequence representation (e.g., TCR sequence representation 336 in FIG. 3 ).

[0189] The sequence representations generated by representation subsystem 501 are sent as input to processing subsystem 503 for processing. In some embodiments, the sequence representations input to processing subsystem 503 may include embeddings corresponding to appended BOS tokens. For example, a peptide sequence representation may include embeddings for BOS tokens appended to the peptide sequence (e.g., BOS+peptide sequence representation 308 in FIG. 3 ), and an MHC sequence representation may include embeddings for BOS tokens appended to the MHC sequence (e.g., BOS+MHC sequence BOS representation 318 in FIG. 3 ).

[0190] The processing subsystem 503 may include various mechanisms for determining an element concentration score for each of one or more (e.g., all) positions of a sequence representation. The element concentration score may indicate a level of attention or importance. For example, the element concentration score for a set of amino acid sequence representations may indicate where the binding core of a peptide begins. The element concentration score may then be used to generate a transformed value for the position.

[0191] Processing subsystem 503 includes processing block 532 and processing block 534. In some embodiments, processing subsystem 501 may include processing block 536, processing block 538, processing block 540, or a combination thereof. Processing block 532 receives peptide sequence representations from BOS+peptide representation block 502 and processes the peptide sequence representations using a set of processing sub-blocks 542 to generate converted peptide sequence representations (e.g., converted peptide sequence representation 316 of FIG. 3 ). The converted amino acid sequence representations can be generated based on the amino acid sequence representations and one or more element concentration scores (representing the connecting core of the set of amino acid sequence representations). One exemplary implementation of processing sub-blocks that execute one or more converter stages to generate the converted sequence representations is described in more detail below in the context of FIG. 6 . Processing block 534 receives MHC sequence representations from BOS+MHC representation block 504 and processes the MHC sequence representations using a set of processing sub-blocks 544 to generate converted MHC sequence representations (e.g., converted MHC sequence representation 322 of FIG. 3 ). A converted MHC sequence representation can be generated based on the MHC sequence representation and one or more element concentration scores. In some embodiments, the element concentration score used to generate the converted amino acid sequence representation can be different from the element concentration score used to generate the converted MHC sequence representation.

[0192] Further, if included, processing block 536 receives the N-flank sequence representation from BOS+N-flank representation block 506 and processes the N-flank sequence representation using a set of processing sub-blocks 546 to generate a transformed N-flank sequence representation (e.g., transformed N-flank sequence representation 328 of FIG. 3 ). In some embodiments, processing block 538 receives the C-flank sequence representation from BOS+C-flank representation block 508 and processes the C-flank sequence representation using a set of processing sub-blocks 548 to generate a transformed C-flank sequence representation (e.g., transformed C-flank sequence 334 of FIG. 3 ). In some embodiments, processing block 540 receives the TCR sequence representation from BOS+TCR representation block 510 and processes the TCR sequence representation using a set of processing sub-blocks 550 to generate a transformed TCR sequence representation (e.g., transformed TCR sequence representation 340 of FIG. 3 ).

[0193] In some embodiments, one or more processing sub-blocks may process representations of different portions, or all, of an amino acid sequence and / or an IPC sequence separately. In some embodiments, one or more sequence representations (e.g., an NN-flank sequence representation, a CC-flank sequence representation, a peptide sequence representation, an MHC sequence representation, a TCR sequence representation) can be processed separately in different iterations of a processing sub-block. For example, an encoded representation of an amino acid sequence may be generated as a feature vector representing the amino acids, and the encoded representations of the sequence (e.g., all or part of the amino acid sequence, all or part of the IPC sequence, all or part of the amino acid sequence, all or part of the IPC sequence) may then be concatenated and fed to another iteration of a processing sub-block.

[0194] 5A, amino acid sequences can be processed separately and independently from the processing of IPC sequences. Having separate and independent processing engines for peptide sequences, N-flank sequences, C-flank sequences, MHC sequences, and / or TCR sequences prior to composite subsystem 505 can improve the predictive performance of machine learning model 532. For example, generating transformed peptide sequence representations using BOS+peptide representation block 502 and processing block 532 along a separate path from generating the transformed IPC sequence representations, and doing so before generating the composite representation, improves the accuracy of the output generated by output subsystem 509. Similarly, the predictive performance (e.g., accuracy) of the machine learning model 532 may be enhanced by generating a transformed N-flank sequence representation along a separate path using the BOS+N-flank representation block 506 and processing block 536, generating a transformed C-flank sequence representation along a separate path using the BOS+C-flank representation block 508 and processing block 538, generating a transformed MHC sequence representation using a separate path using the BOS+MHC representation block 504 and processing block 534, generating a transformed TCR sequence representation using a separate path using the BOS+TCR representation block 510 and processing block 540, or a combination thereof. In some embodiments, the separate paths may allow for efficient processing (e.g., using reduced computing resources, faster processing, etc.) because multiple amino acid-IPC (peptide-MHC, peptide-TCR) combinations can be considered modularly and / or processed in parallel.

[0195] The transformed sequence representations output from the processing subsystem 503 are sent to a compositing subsystem 505 for processing. The compositing subsystem 505 includes a compositing block 552. The compositing block 552 can form one or more composite representations (e.g., composite representation 342 in FIG. 3 ) using the transformed representations output from the processing subsystem 503. For example, the compositing block 552 can multiply a set of transformed sequence representations (e.g., a set of amino acid sequence representations, a set of IPC sequence representations) to form a composite representation.

[0196] In some embodiments, the size of the output generated by the combine block 552 can be equal to, for example, m×n, where m is equal to the total number of amino acids being considered plus 1 (e.g., for BOS tokens) plus any padding to fit the normalized sequence length (e.g., 39), and n is equal to the number of features (a predetermined value, e.g., 600). A single column (having n values) can be selected as the output of the combine block 552 for further processing. The single column can be the first column and / or the column associated with the BOS token. In some embodiments, the output from the combine block 552 can be aggregated to form a single vector, which can then be provided to the output subsystem 509.

[0197] The output subsystem 509 may include various blocks, sub-blocks, layers, or combinations thereof for generating a final output. In some embodiments, the output subsystem 509 includes a dropout block 560, a fully-connected block 562, and an output block 564. The dropout block 560 may include, for example, one or more dropout layers. The fully-connected block 562 may include, for example, one or more fully-connected layers. The output block 564 may include, for example, one or more layers for filtering, selecting, transforming, or generating a result. For example, the output block 564 may include at least one max layer 565 configured to select a subset of the input received at the output block 564, for example, based on a selected threshold or range.

[0198] In some cases, the composite representation is received and processed by dropout block 560 to generate a first output that is received by fully connected block 562. Fully connected block 562 can receive and process this first output to generate a second output, at least a portion of which is received by output block 564. Output block 564 receives and processes its input to generate a result, such as an interaction output 566, an immunogenicity output 568, or both.

[0199] In some embodiments, the fully connected block 562 can be configured to generate one or more outputs with a dimensionality that is less than the dimensionality of its inputs (e.g., fewer than a predetermined number of features fed into the fully connected block 562). For example, the output of the fully connected block 562 can include a single value, two values, or three values, each corresponding to a prediction regarding a target interaction or immune response. The fully connected block 562 can include, for example, a single hidden layer, two hidden layers, or three or more hidden layers. The number of nodes in an initial hidden layer can be greater than the number of nodes in a subsequent hidden layer. For example, the first hidden layer can include 256 nodes, and the second hidden layer can include 126 nodes. In some embodiments, each output from the fully connected block 562 can include a real score, which can be converted, for example, to a binary and / or categorical result (e.g., using a trained activation function) and / or converted to a scaled number. For example, the scaled number can include a probability on a scale from 0 to 1.

[0200] The interaction output 566 may include, for example, one or more of a set of interaction predictions 570 for one or more target interactions, or a set of interaction affinity predictions 572. The interaction predictions 570 may include, for example, a prediction of whether or to what extent an IPC (e.g., MHC, TCR) will bind to a peptide for a corresponding amino acid-IPC combination, such as a peptide-IPC (e.g., peptide-MHC, peptide-TCR) combination. In some embodiments, the interaction predictions 570 may include, for example, a prediction of a corresponding peptide-IPC (e.g., MHC) combination, for example, whether the IPC (e.g., peptide-MHC) will bind to the peptide. In some embodiments, the interaction affinity predictions may include, for example, a prediction 570 of affinity for the target interaction for a corresponding peptide-IPC (e.g., peptide-MHC, peptide-TCR) combination. The target interaction may be, for example, binding between the peptide and the IPC. The affinity for the target interaction may be, for example, binding affinity, indicating the strength, propensity, and / or stability of binding between the peptide and the IPC.

[0201] The immunogenicity output 568 includes a set of immunogenicity predictions. The immunogenicity predictions may include, for example, predictions of immunogenicity for corresponding amino acid-IPC combinations. For example, the immunogenicity predictions may indicate the ability of a peptide to elicit an immune response for a particular IPC (e.g., MHC, TCR) of interest. In some embodiments, the predicted amino acid-IPC interactions include predictions of tumor-specific immunogenicity of the peptide. In some embodiments, the predicted amino acid-IPC interactions identify a subset of peptide sequences with increased tumor-specific immunogenicity or increased likelihood of presentation by an IPC compared to the set of peptide sequences.

[0202] In some cases, a first portion of the output from the fully connected block 562 is sent to the output block 564, and a second portion of the output from the fully connected block 562 is in its final form and is used as a set of interaction affinity predictions 572.

[0203] In some embodiments, the transformed composite representation is received at the output subsystem 509 and processed by the fully connected block 562. The fully connected block 562 may process the transformed composite representation to generate a first output that is sent to the dropout block 560. The output of the dropout block 560, or a portion thereof, may then be sent to the output block 564 for processing.

[0204] In some cases, dropout may be applied to each fully connected sub-block within the fully connected block 562, followed by a batch normalization layer. In some embodiments, the output block 564 is used for deconvolution by applying an activation function to the representation predictions (e.g., via a max layer 565, which may include a softmax function, or simply using a maximum) so that the amino acid-IPC interactions (e.g., 12 pairs of peptide-MHCII interactions or 6 pairs of peptide-MHC1 interactions) correspond to a single selected MHCII allotype or MHC1 allele (respectively). During training, the selected peptide-MHC interaction outputs can be normalized as values ​​between 0 and 1 and compared to the true representation values ​​using a loss function (e.g., a binary loss function) to generate an error for adjusting the model parameters.

[0205] In some embodiments, output from output subsystem 509 may include a plurality of results for each IPC (e.g., MHC) allele, including predicted amino acid-IPC interactions indicating whether the peptide binds to the IPC allele and / or the probability that the peptide binds to the IPC allele. Allele-specific predictions may be output, or in some cases, a maximum tier 565 may be used to determine the maximum of the allele-specific predictions, and the maximum may be output.

[0206] In this manner, output subsystem 509 can be implemented in any of several different ways, using any number of different blocks, sub-blocks, and / or layers that enable the generation of interaction output 566, immunogenicity output 568, or both.

[0207] In some cases, the machine learning model 532 can facilitate an automated determination of which specific IPC alleles are predicted to bind to a peptide. For example, if an MHC molecule includes 12 MHC allotypes (as in humans), 12 iterations of at least a portion of the neural network process can be performed (e.g., in parallel), one for each allele. Each process can use as input an MHC sequence representation and a peptide representation of at least a portion of the peptide's sequence. Each process can generate a composite representation. Predicted amino acid-IPC interactions can be determined based on the composite representation. In some embodiments, the predicted amino acid-IPC interactions can include a prediction of whether or to what extent a peptide will bind to an MHC allele or allotype. The peptide associated with the highest predictive value across alleles (e.g., indicating the most likely binding prediction) can be inferred to be the peptide to which the peptide will bind.

[0208] In some examples, for six, up to 12 MHC alleles (e.g., MHC class I) or allotypes (e.g., MHC class II), corresponding MHC sequence representations can be generated by passing different MHC allotype sequences through the same BOS+MHC representation block 504 to generate an IPC sequence representation (e.g., MHC sequence representation) for each peptide-MHC combination. In some embodiments, the MHC sequence representations can be aggregated, with a single appended BOS token embedded in the embedding layer, by multiplying each BOS token of a set of converted IPC sequence representations (e.g., comprising each of 6-12 MHC sequence representations) with a BOS token of each of a set of converted amino acid sequence representations (e.g., comprising a set of converted peptide representations) to generate a composite representation.

[0209] In some embodiments, one or more of the processing blocks or sub-blocks included in the machine learning model 532 can be replaced with another type of network and / or processing unit to transform the representation of one or more sequences. The transformation can represent the extent to which different amino acids (at particular positions) are predicted to affect binding affinity and / or presentation probability, and / or the extent to which different particular combinations of amino acids (at particular positions) occurring across a single sequence or across sequences are predicted to affect binding affinity and / or presentation. For example, one or more processing sub-blocks can be replaced with one or more gated recurrent units.

[0210] 5B is a schematic diagram of an exemplary configuration of machine learning model 532, according to some embodiments. In the configuration shown in FIG. 5B, representation subsystem 501 includes amino acid sequence representation block 580. Amino acid sequence representation block 580 receives an amino acid sequence. For example, the amino acid sequence may include one or more of a peptide sequence (e.g., peptide sequence 126 of FIG. 1 ), an N-flank sequence (e.g., N-flank sequence 128 of FIG. 1 ), or a C-flank sequence (e.g., C-flank sequence 130 of FIG. 1 ).

[0211] Amino acid sequence representation block 580 may include, for example, an embedding layer 582 that processes the amino acid sequence with the appended BOS tokens to form an embedded amino acid sequence representation received by positional encoder 583. Positional encoder 583 positionally encodes the embedded amino acid sequence representation to generate BOS+ amino acid sequence representation 584. In some embodiments, BOS+ amino acid sequence representation 584 may include one or more of BOS token representation 581, peptide representation 585, N-flank representation 586, or C-flank representation 587.

[0212] BOS+ amino acid sequence representation 584 is output from amino acid sequence representation block 580 and sent to processing block 588 of processing subsystem 503. Processing block 588 includes a set of processing sub-blocks 589 that process BOS+ amino acid sequence representation 584 to generate a converted BOS+ amino acid sequence representation that is sent to composite block 552 for processing.

[0213] In some embodiments, if the amino acid sequence sent to amino acid sequence representation block 580 includes either an N-flank sequence or a C-flank sequence but not both, machine learning model 532 may also include a corresponding representation block (e.g., N-flank representation block 506 or C-flank representation block 508 in FIG. 5A ) and corresponding processing block (e.g., processing block 536 or processing block 538 in FIG. 5A , respectively) of the sequence included in the amino acid sequence.

[0214] FIG. 5C is a schematic diagram of an exemplary configuration of machine learning model 532, according to some embodiments. Representation subsystem 501 can be designed to process one or more subsequences (or combinations thereof) of each type (e.g., peptide, C-flank, N-flank, C-flank + N-flank, MHC, TCR) to accommodate concatenation of multiple subsequences by type, along with a single appended BOS token, before encoding the set of concatenated subsequences. As shown in the example of FIG. 5C, representation subsystem 501 can also include a single representation block 507 for a combined sequence including a single BOS token, a C-flank subsequence, and an N-flank subsequence. This representation block 507 can include embedding layer 521 and position encoder 523. The BOS + C-flank + N-flank sequence representation generated by representation block 507 can be transformed by processor block 590a, which includes a set of processing sub-blocks 592a.

[0215] Alternatively, the peptide representation, BOS+peptide sequence representation and BOS+C-flank+N-flank sequence representation generated by presentation subsystem 501 can be aggregated before the conversion stage to form an amino acid sequence representation that is sent to a single processing block 590. Processing block 590 includes a set of processing sub-blocks 592 that process the amino acid sequence representation to generate a converted amino acid sequence representation that is sent to composite block 552 for processing.

[0216] Similarly, the BOS+MHC and BOS+TCR sequence representations generated by representation subsystem 501 can be aggregated before the transformation stage to form an IPC sequence representation that is sent to processing block 594. Processing block 594 includes a set of processing sub-blocks 596 that process the IPC sequence representations to generate transformed IPC sequence representations that are sent to composite block 552 for processing. In some embodiments, the BOS+MHC sequence representation can be processed by a set of blocks (e.g., 594a, 596a) and the BOS+TCR sequence representation can be processed by another set of blocks (e.g., 594b, 596b).

[0217] 5A-5C, the machine learning model 532 can be implemented in any number or combination of blocks, sub-blocks, and / or layers within the various subsystems, and thus can be modular and customizable for a given task.

[0218] Exemplary Processing Blocks 6 is a schematic diagram of a processing block 600 for executing one or more transformer stages to generate a transformed array representation, according to some embodiments. Processing block 600 may be an example of an implementation of a processing block in processing subsystem 304 of FIG. 3 or processing subsystem 503 of FIGS. 5A-5C.

[0219] Processing block 600 includes one or more processing sub-blocks. For example, processing block 600 may include processing sub-block 1 602 and, optionally, one or more other processing sub-blocks up to processing sub-block n 604. When multiple processing sub-blocks are present in processing block 600, these processing sub-blocks may be connected in series (e.g., daisy-chained together to produce a final output).

[0220] Processing sub-block 1 602 may be implemented in a variety of ways. In some embodiments, processing sub-block 1 602 includes, for example, a processing layer 606, a summation and normalization layer 608, a feedforward layer 610, and a summation and normalization layer 612. In this configuration, sub-block 1 602 may also be referred to as a transformer encoder. In some embodiments, one or more processing sub-blocks 604 may be implemented in a similar manner as processing sub-block 1 602. In some embodiments, processing layer 606 may include one or more embedding components configured to perform positional and / or non-positional embedding.

[0221] In summation and normalization layer 608 or 612, the transformed representation may be added (via residual connections) to a position-directed embedded representation of the array, and the summed representation may be normalized. The normalized data may be fed to a corresponding feedforward layer 610 (e.g., a fully connected feedforward network). The feedforward network may affect one, two, three, or more linear transformations (for example) for each position, and / or may include activations (e.g., ReLU activations) between each of the linear transformations. For example, a feedforward layer may be represented by: TIFF2026500160000002.tif5170In the formula, x is the input to the layer, W1 and W2 are the gradients of the linear transformation, and b1 and b2 are the intercepts of the linear transformation.

[0222] The dimensionality of the output of the feedforward layer of a particular attention sub-block may be the same as the dimensionality of the input to the feedforward layer of the processing sub-block. In some cases, the inputs and outputs may be summed and normalized (e.g., via another residual connection through another summation and normalization layer) to preserve the representation of various types of information.

[0223] In some embodiments, the feedforward layer 610 may allow for processing of sequences of variable length. One or more additional feature vectors (e.g., assigned random or pseudorandom values) may be included in the concatenated representation, which is then encoded. This coded representation of the sequence combination may be processed by the feedforward layer 610 (e.g., a fully connected neural network), to which dropout and / or batch normalization may be applied. In some cases, the coded representation(s) of the additional feature vector(s) are selectively passed to the feedforward layer 610 (e.g., whereas feature vectors corresponding to individual amino acids of the MHC molecule and / or variant peptide are not). For example, assume that a subsequence of an MHC molecule contains x1 amino acids, a subsequence of a variant peptide (e.g., one or more additional flanks) contains x2 amino acids, and the feature transformation identifies y feature values ​​representing each amino acid. Thus, a concatenated representation including one additional feature vector may have a size of [(x1 + x2 + 1), y]. The input provided to the feedforward layer 610 may have a size of [1, y] if one feature vector is selected for processing by the feedforward layer 610.

[0224] The results generated by the feedforward layer 610 may correspond to a prediction regarding the binding affinity between the mutant peptide and an MHC molecule (e.g., an MHC molecule of interest) and / or whether the mutant peptide will be presented by the MHC molecule. The binding affinity prediction may be, for example, numerical (e.g., corresponding to a predicted probability, predicted binding strength, and / or predicted binding stability of the mutant peptide that it will bind to the MHC molecule), categorical (e.g., predicting no, low, or high binding stability between the mutant peptide and the MHC molecule), or binary (e.g., predicting whether the mutant peptide will bind to the MHC molecule).

[0225] The machine learning model 132 may include one or more processing layers, such as a self-attention layer or a convolutional layer, or a neural network, such as a long-short-term memory unit (LSTM), a recurrent structure, or a recurrent component. Figure 7A shows a flowchart of an exemplary process for processing a sequence representation using processing layers, according to some embodiments. Process 700 may be used, for example, by one or more processing blocks present in the machine learning model 132 of Figures 1 and 3, one or more attention blocks present in the machine learning model 532 of Figures 5A-5C, and / or processing block 600 of Figure 6.

[0226] Step 702 includes receiving a sequence representation including multiple elements. The sequence representation may be, for example, an amino acid sequence representation, an IPC sequence representation, an N-flank sequence representation, a C-flank sequence representation, an MHC sequence representation, a TCR sequence representation, an aggregate sequence representation, or another type of representation. For example, the sequence representation may represent a variant coding sequence, part or all of a wild-type or variant coding sequence, an epitope sequence (e.g., including a mutant), a candidate neoepitope sequence, part or all of a neoantigen sequence, a sequence beginning or ending at a peptide terminus (e.g., an N-flank or C-flank), or part or all of an MHC sequence (e.g., an MHC pseudosequence). The sequence representation may be generated, for example, using representation subsystem 302 of FIG. 3 or representation subsystem 501 of FIGS. 5A-5C. Each element of the sequence representation may be associated with a unique position in the sequence.

[0227] Step 704 includes determining, for each element of the array representation, a plurality of vectors, such as a key vector, a value vector, and a query vector, using a plurality of weights, such as a set of key weights, a set of value weights, and a set of query weights. For example, if the array representation includes, for example, 20 amino acids, 20 key vectors, 20 value vectors, and 20 query vectors can be generated. The elements of the array representation may correspond, for example, to rows or columns of a two-dimensional array representation (e.g., a first dimension representing different amino acids in the sequence and a second dimension representing, for example, different components that characterize each individual amino acid).

[0228] In some embodiments, the set of key weights is in the form of a key weight matrix. The key weight matrix for a particular element can have a size equal to the length of the element, minus some length of the key vector. For example, the elements can have a length of 20 (e.g., each value corresponds to a binary indicator of whether an amino acid in a sequence is the same as a particular one of 21 amino acids), and if the key vector has a length of 5 (e.g., representing five components or features), then the key weight matrix can have a size of [5, 21]. The key weight matrix can be learned during training (e.g., randomly initialized at the beginning of training).

[0229] The value vector of an element may have the same size as the key vector of the element. The value vector may be determined using a set of value weights that may be learned during training and included in a value weight matrix. The value weight matrix of a given element may have the size of the key weight matrix and / or may have a size defined based on the length of the element and the length of the value vector.

[0230] The query vector of an element may have the same size as the key vector and / or value vector of the element. The query vector may be learned during training and may be determined using a set of query weights contained in a query weight matrix. The query weight matrix of an element may have the size of the key weight matrix and / or the value weight matrix. In some embodiments, the query weight matrix may have a size based on the length of the element and the length of the query vector.

[0231] Step 706 includes generating, for each element in the array representation, a set of element concentration scores using the element's query vector (generated using the query weight and the array representation) and the multiple element key vectors (generated using the key weights and the array representation). For a given element, the set of element concentration scores can indicate weights to assign to the value vector of the given element. The elements whose key vectors are used in generating the set of element concentration scores for selected elements in the array representation can include some or all of the elements of the array representation (e.g., some or all of the amino acid sequence representation). The elements can include a focus element (e.g., a particular amino acid for which a set of element concentration scores is being determined).

[0232] A set of element focus scores is generated by generating a score for each pair of a focal element (first element) with a same or different element (second element) for each element of the array representation, where the score for this pair can be the product of the query vector of the first element and the key vector of the second element.

[0233] In some cases, step 706 may include performing an activation function and / or normalization. The normalization may be based on the dimensionality of the key vector (or query vector). For example, the normalization may be the square root of the length of the key vector. The activation function may include a softmax function. In some cases, the normalization is applied before the activation function.

[0234] Step 708 includes generating a transformed array representation. The transformed array representation can be determined by performing a transformation of the plurality of elements to form a plurality of modified elements. The transformation can be performed using the set of element concentration scores generated for each of the plurality of elements and the value vector determined for each of the plurality of elements. For example, if the array representation includes 11 elements (e.g., representing 11 amino acids) and scores are determined for all pairwise combinations of elements, a modified array representation including a plurality of modified elements is generated. In some embodiments, the modified element can be a weighted average of the value vectors of all elements (using the weighted scores).

[0235] Step 710 involves generating an encoding of the array using the transformed array representation, the initial array representation, and a feed-forward network. For example, the transformed array representation and the initial array representation can be summed. This result may still contain multiple elements (e.g., each update occurs via a translation, addition, and normalization). The feed-forward neural network can then process the summed representation (e.g., by performing one, two, or more linear transformations and / or implementing one or more activation functions). Summing the representations can reintroduce positional information that may be obscured in the transformed array representation (due to corresponding values ​​of other elements when generating a transformed value vector for a given element).

[0236] A feedforward neural network can be configured to process each of the updated elements separately (e.g., using the same technique and / or the same parameter set). Thus, the input to the feedforward network can include a vector corresponding to a single element, a single amino acid, and / or a single sequence position. The feedforward network can be configured so that the output of the feedforward network is the same size as the input to the feedforward network. In some cases, instead of using a feedforward network to process the transformed and initial sequence representations, a convolution (e.g., a one-dimensional convolution) is used to instead perform a local transformation that operates similarly (e.g., identically) across positions / elements. One-dimensional convolution can be used as an alternative way to interpret the function of a feedforward neural network.

[0237] The technique shown in Figure 7A relates to a process that uses a single set of key vectors, value vectors, and query vectors to calculate element concentration scores. Embodiments of the present disclosure may include using multiple sets of key weights, value weights, and query weights to generate separate key vectors, separate value vectors, and separate query vectors. These separate vectors can be used to generate processing scores and transformation values ​​for each element. The transformation values ​​can be concatenated and projected.

[0238] 7A refers to the calculation and use of various vectors, it should be further understood that a matrix representation may be used instead, which may facilitate efficiently performing calculations across elements, as opposed to iteratively calculating the various vectors individually.

[0239] Figure 7B is a schematic diagram illustrating the process 700 described in Figure 7A, according to some embodiments. In Figure 7B, process 750 receives as input a sequence 752. The sequence 752 may be, for example, an amino acid sequence. Another exemplary sequence may be an IPC sequence.

[0240] In the illustrative example of FIG. 7B, the sequence 752 includes a plurality of amino acids 754 (4 amino acids: x 1 ~x 4 ) contains multiple elements a 1 ~a 4 An array representation 756 including each element a is generated through embedding and, in some embodiments, positional encoding. i Array representation 756 may be an example of an array representation received in step 702 of FIG.

[0241] Multiple vectors 758 (e.g., a query vector q i , key vector k i , and the value vector v i ) to each element a of the array representation 752 i The plurality of vectors 758 may be an example of an implementation of the vectors generated in step 704 of FIG. 7A. The illustrated example is 1 Focusing on TIFF2026500160000003.tif5170. Element focus attention score 760 is an example of a set of element focus scores generated for a particular element in step 706 of Figure 7A. TIFF2026500160000004.tif5170q 1 and k i It can be taken as the dot product with TIFF2026500160000005.tif5170 Set value vector v i The weighted sum of the correction factors 762, b 1 A transformation is performed to generate a modified element 762. Modified element 762 is an example of a modified element generated in step 708 of FIG. 7A. Similar transformations may be performed on other elements of array representation 756. For example, further details of transformer architectures can be found in Ashish Vaswani et al., Attention is All You Need, Neural Information Processing Systems (2017).

[0242] Exemplary Methods for Using Machine Learning Models The machine learning model 132 of Figures 1 and 3, the workflow of Figures 4A-4D, and the machine learning model 532 of Figures 5A-5C can be used in various ways to generate predictions regarding immunological activities (e.g., predicted binding, binding affinity, predicted presentation occurrence, immunogenicity, etc.) associated with various peptides, including mutant peptides (e.g., neoantigens).

[0243] 8 is a flowchart of an exemplary process for generating information about the immunological activity of various peptides, according to some embodiments. At least a portion of process 800 can be implemented using, for example, but not limited to, prediction system 100 described in FIG. 1. For example, at least a portion of process 800 can be implemented using, for example, but not limited to, machine learning model 132 from FIGS. 1 and 3, machine learning model 532 from FIGS. 5A-5C, or the workflow of FIGS. 4A-4D.

[0244] Step 802 includes accessing an amino acid sequence that includes a peptide sequence characterizing a mutant peptide, where the peptide sequence may include variants relative to a corresponding reference sequence. The peptide sequence characterizes the mutant peptide by characterizing at least a portion of the mutant peptide. The mutant peptide may be, for example, a neoantigen. Step 802 may be performed, for example, by retrieving the peptide sequence from a data store (e.g., data store 104 of FIG. 1 , cloud storage, a server, or a server system, etc.). In some embodiments, the peptide sequence may be one of multiple peptide sequences processed by the machine learning model.

[0245] Step 804 includes receiving an IPC sequence identified for the IPC of interest. The IPC may be, for example, an MHC, a TCR, or an MHC-TCR complex. The IPC sequence characterizes the IPC by characterizing at least a portion of the IPC.

[0246] Step 806 includes processing the amino acid sequence and the IPC sequence using different processing engines within the machine learning model to generate an output, where the output provides information about immunological activity associated with both the mutant peptide and the IPC. Step 806 includes, for example, processing the amino acid sequence through corresponding representation blocks to generate an amino acid sequence representation. The amino acid sequence representation can be processed by the corresponding processing blocks to generate a transformed amino acid sequence representation. This amino acid processing engine is separate and independent from the IPC processing engine, where the IPC sequence is processed through the corresponding representation blocks to generate an IPC sequence representation (e.g., an MHC representation, a TCR representation, an MHC-TCR representation), and the IPC sequence representation is processed through the corresponding processing blocks to generate a transformed IPC sequence representation (e.g., a transformed MHC representation, a transformed TCR representation, an transformed MHC-TCR representation) representing the IPC sequence.

[0247] In some embodiments, the amino acid sequence representation is an aggregate representation that includes an NN-flank representation for an N-flank sequence and / or a C-flank representation for a C-flank sequence. In such embodiments, the aggregate processing engine (which may include an amino acid processing engine) remains separate from the IPC processing engine.

[0248] In various embodiments, in step 806, the converted amino acid sequence representation and the converted IPC sequence representation are used to form a composite representation, which is then further processed to generate an output, which may include, for example, but is not limited to, a set of interaction predictions, a set of interaction affinity predictions, a set of immunogenicity predictions, or a combination thereof.

[0249] Step 808 includes performing one or more actions based on the output. As an example, a report may be generated that includes the output. In some embodiments, the report includes a transformed or filtered version of the output. In some embodiments, the report includes a summary, outline, or visual representation of the output.

[0250] In some embodiments, step 808 includes other operations related to the design and / or manufacture of treatments based on the output. For example, pharmaceutical compositions can be selected or ranked based on the output. The output can include a prediction of which mutant peptides will bind to a subject's specific IPC (e.g., MHC allele or allotype). This binding prediction can indicate the likelihood that the subject's immune system will recognize, for example, cancerous cells. The binding prediction can be used to aid in the selection of candidate neoepitopes (mutant peptides) for a vaccine. In some embodiments, the composite representation with the highest result in the output (e.g., a predictive value indicating the most likely binding and / or presentation prediction) can be selected for the pharmaceutical composition. In some embodiments, the composite representations can be ranked according to their corresponding result in the output.

[0251] Embodiments of the present disclosure may include generating output based on a set of IPC sequences. For example, for a given subject, output can be generated based on six, or up to twelve, MHC alleles or allotypes. FIG. 9 is a flowchart of an exemplary process for generating information regarding the immunological activity of various peptides, according to some embodiments. At least a portion of process 900 can be implemented using, for example, but not limited to, the prediction system 100 described in FIG. 1. For example, at least a portion of process 900 can be implemented using, for example, but not limited to, the machine learning model 132 from FIGS. 1 and 3 or the machine learning model 532 from FIGS. 5A-5C.

[0252] Step 902 involves accessing sequence data that includes a set of amino acid sequences and a set of IPC sequences.

[0253] Step 904 involves generating a set of amino acid-IPC combinations using the set of amino acid sequences and the set of IPC sequences, each amino acid-IPC combination being unique.

[0254] Step 906 includes, for each amino acid-IPC combination, inputting the corresponding amino acid sequence into an amino acid processing engine of the machine learning model, and inputting the corresponding IPC sequence into an IPC processing engine of the machine learning model.

[0255] Step 908 includes, for each amino acid-IPC combination, processing the amino acid sequence representation using a first processing block and processing the IPC sequence representation using a second processing block to generate a converted amino acid sequence representation and a converted IPC sequence representation, respectively.

[0256] Step 910 involves generating a composite representation for each amino acid-IPC combination using the converted amino acid sequence representation and the converted IPC sequence representation.

[0257] Step 912 includes generating an output based on the composite representation. In some embodiments, predicted amino acid-IPC interactions can be determined based on the composite representation. The output can provide an indication of which peptide sequences can be used to generate treatments. For example, the output can provide an indication of which peptide sequences (and thereby peptides containing the peptide sequences) are likely to bind to MHC, be presented by MHC, have high interaction affinity for peptide-MHC binding, and / or are immunogenic, thereby eliciting an immune response.

[0258] Exemplary Method for Training a Machine Learning Model FIG. 10 is a flowchart of an exemplary process for training a machine learning model and using the trained machine learning model to generate predictions for amino acids (e.g., peptides) and IPCs (e.g., MHCs), according to some embodiments. Process 1000 can be performed using prediction system 100 of FIG. 1. For example, process 1000 can be implemented using machine learning model 132 of FIGS. 1 and 3, machine learning model 532 of FIGS. 5A-5C, or any of the workflows of FIGS. 4A-4D. In some cases, part or all of process 1000 can be performed on a remote computing system that is remote to a user's device and / or laboratory. The remote computing system can be a cloud computing system.

[0259] The machine learning model can be trained using at least a portion of a training dataset. Step 1002 includes accessing a training dataset having training elements identifying training amino acid sequence data, training IPC sequence data, and training immunological activity data. The training dataset may be an example of an implementation of training data 133 of FIG. 1. The training immunological activity data may include, for example, interaction indices.

[0260] The training dataset can include multiple training data elements. Each training data element can include a sequence representation and a result (e.g., indicating whether at least a portion of the peptide corresponding to the sequence is presented by an MHC molecule and / or whether it elicits immunogenicity). Training data elements for which no presentation or binding is detected can be computationally generated. For example, for each source protein in the positive set (corresponding to positively eluted ligand presentation data), one or more (e.g., all) possible peptide fragments (e.g., within a predetermined length range, such as 8 to 11) can be generated, potentially with uniform probability for each length. N-terminal and C-terminal flanking sequences can be retained (e.g., potentially with a maximum length, e.g., 10 amino acids). In some examples, peptide fragments (e.g., one or more (e.g., all) lengths between 8 and 11) can be generated for each allele represented in the positive examples of the training data. Generation and / or subsequent selection can be performed so that the probability of occurrence of a sequence of a given length is uniform across the length. N-terminal and C-terminal flanking sequences can be retained or can be retained at a specific maximum length (e.g., a maximum length of 10 amino acids). In certain embodiments, any other suitable sequence length range (eg, 9-30 for MHC class II) can be utilized.

[0261] The training dataset can be randomly parsed, shuffled, and / or split to train various models in the ensemble. The loss function can use an error term (e.g., mean squared error or median squared error) and / or an entropy term (e.g., cross-entropy or binary cross-entropy). Multi-task learning can be used so that models are simultaneously trained to predict each of two different types of outcomes (e.g., binding affinity and occurrence of presentation). Static or non-static learning rates can be used. For example, learning rate annealing (e.g., using stepwise annealing or cosine annealing) can be used to decrease the learning rate over iterations. Validation data evaluation can be used to terminate training early (e.g., upon determining that performance goals have been met).

[0262] The training amino acid sequence data may include, for example, one or more amino acid sequences for training (which may include variant coding sequences). The amino acid sequence may include a peptide sequence. The peptide sequence may identify an ordered set of amino acids within a peptide (e.g., a neoantigen). The peptide sequence may identify amino acids within an epitope (e.g., including variants, including neoepitopes, and / or being neoepitopes) of the peptide. In some embodiments, the peptide sequence is within an aggregate sequence that also includes an N-flank sequence (e.g., characterizing the stretch of amino acids at the N-terminus of the corresponding peptide) or a C-flank sequence (e.g., characterizing the stretch of amino acids at the C-terminus of the corresponding peptide). Neither the N-flank nor the C-flank binds to an MHC molecule, but each may influence whether it is presented by an MHC molecule.

[0263] In some instances, it is unknown how many amino acids from the flank (e.g., the N-flank) are used by the peptidase to determine when to trim long peptides to the presented peptide core. To address this unknown when generating training data, the flanks can then be trimmed to a length selected based on techniques (e.g., pseudorandom selection techniques), such as a length within a predetermined range (e.g., 1 to 10 amino acids). The selection technique may select the length using a distribution (e.g., a uniform or Gaussian distribution). In some instances, flanks below a threshold length (e.g., 10 amino acids) are not trimmed. In some instances, flank trimming can be such that the C-side of the N-flank is preserved.

[0264] The training MHC sequence data may include one or more MHC sequences for training. The MHC sequence may, for example, identify amino acids within part or all of an MHC molecule (e.g., an MHC-I molecule or an MHC-II molecule). The MHC sequence may include an MHC pseudo-sequence (e.g., including 34 amino acids). The MHC sequence may, for example, identify amino acids within 1, 2, 3, 4, 5, or 6 MHC alleles for MHC-I, or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 MHC allotypes for MHC-II. The MHC sequence may identify amino acids that make up part or all of an HLA molecule.

[0265] The MHC contains multiple alleles in vivo (e.g., six alleles and 12 allotypes per human). For a single MHC molecule, multiple sequence inputs can be generated (e.g., each representing a single allele of the multiple alleles). Each of the multiple sequence inputs can be processed separately using one or more neural networks (e.g., one or more transformer encoders) to generate predicted binding or presentation values ​​for the neo-antigens associated with each of the alleles. A function (e.g., a max function) can identify which allele among the multiple alleles is associated with the highest presentation prediction. During training, this maximum presentation prediction for this particular sequence input can be compared to the true presentation value using a binary loss function to generate an error for tuning parameters.

[0266] The training immunological activity data may include, for example, one or more interaction indices for one or more amino acid-IPC combinations. For example, a training dataset may include training elements, each of which includes an amino acid sequence and an IPC sequence for training, and one or more interaction indices for the corresponding peptide-IPC combination. The interaction indices may indicate whether a target interaction (e.g., binding of a peptide to an MHC, presentation of a peptide on the cell surface by an MHC) occurs between an amino acid (e.g., a peptide) and an IPC (e.g., an MHC), or indicate affinity for the target interaction and / or elicit an immunological response.

[0267] The interaction indicator may be, for example, a label. A negative interaction label may indicate that the peptide does not bind to an MHC molecule and / or is not presented by an IPC (e.g., an MHC molecule). A positive interaction label may indicate that the peptide binds to an MHC molecule and / or is presented by an MHC molecule. Furthermore, the interaction label may indicate the probability that the peptide binds to an MHC molecule, the binding affinity for the peptide-MHC combination, the strength of the binding between the peptide and the MHC molecule, the stability of the binding between the peptide and the MHC molecule, the tendency of the peptide to bind to MHC, or another metric or characteristic related to the interaction between the MHC and the peptide.

[0268] The training dataset can be generated, for example, by in vitro or in vivo experiments and / or based on medical records. In some embodiments, the machine learning model can be trained using binding affinity data and mass spectrometry elution data indicating which peptides are presented by MHC molecules. The binding affinity data can include qualitative data (e.g., as determined using ELISA, pull-down assays and / or gel shift assays, fluorescence resonance energy transfer assays and mass spectrometry assays) or quantitative data (e.g., using biosensor-based methodologies such as surface plasmon resonance, isothermal titration colorimetry, biolayer interferometry, or microscale thermophoresis). In some examples, the binding affinity data can include data from competitive binding assays, data from immune epitope databases, and / or types of data found in immune epitope databases. Elution data can be collected using peptide-MHC immunoprecipitation, followed by elution and detection of presented MHC ligands by mass spectrometry.

[0269] To collect training data, some of the sequences identified in the disease sample may be non-disease sequences corresponding to non-disease peptides. To identify disease-specific nucleic acid sequences and / or disease-specific amino acid sequences, for each sequence detected as a result of sequencing the disease-specific sample, it can be determined whether the sequence is also identified in the reference sequence dataset. The reference sequence dataset may include a set of reference sequences whose sequences are known, suspected, or hypothesized not to be indicative of or characteristic of a disease (e.g., any disease or a given disease). The reference sequence dataset may include, for example, sequences identified by sequencing one or more reference sample sequences collected from the same subject from which the disease-specific sample was collected, by sequencing one or more reference sample sequences collected from one or more other subjects who have not been diagnosed with the disease or the disease corresponding to the disease-specific sample, and / or by sequencing one or more cell lines not associated with a particular disease. In some cases, the reference sequence dataset may include sequences collected from one or more reference data repositories. Sequences that are detected in association with the disease-specific sample but not detected (or detected at a frequency below a predefined threshold) in the reference sequence dataset can be classified as variant coding sequences (e.g., in general, or for the subject from whom the disease-specific sample was collected).

[0270] In some cases, multiple variant coding sequences can be identified (e.g., each detected in a disease sample but not represented in a reference sample sequence). In some cases, a machine learning model disclosed herein can be used to process (e.g., individually, sequentially, and / or in parallel) representations of each of the multiple variant coding sequences to predict binding affinity and / or presentation predictions.

[0271] A disease sample may include, for example, tissue (e.g., a solid tumor), blood, and / or a set of cells (e.g., cancer cells that may be collected using fine needle aspiration or laparoscopy). A disease sample may include, for example, cancerous cells collected from a subject diagnosed with and / or having lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, stomach cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, and T-cell lymphocytic leukemia, non-small cell lung cancer, or small cell lung cancer.

[0272] In some cases, the initial sample is separated into a disease sample and another remaining sample (e.g., can be discarded or used as a reference sample). The reference sample can include a matched disease-free sample. Each of the disease sample and the reference sample can be collected from the same subject and / or can include or be the same or similar sample type (e.g., tissue type). In some cases, the disease sample is collected from a first subject (e.g., a person diagnosed with a disease or disease), and the reference sample is collected from a different second subject (e.g., a person not diagnosed with a disease or disease). In some cases, the reference sample sequence is searched from a database of known genes associated with the organism.

[0273] The training data may further include the sequence of one or more peptides, along with an indication of whether each peptide binds to, is presented by, and / or elicits an immunological response with an MHC molecule. To collect training data that correlates sequence data with the observed presentation and / or binding data, the disease sample (and potentially the reference sample) may be processed (separately) to isolate MHC / peptide complexes (e.g., by immunoprecipitation using an antibody specific for MHC) and / or to elute (and thereby sequence) peptides from MHC molecules (e.g., using chromatography and / or mass spectrometry). In some examples, reference sample sequences are identified for use in generating presentation data by sequencing one or more cell lines engineered to express one or more MHC alleles (e.g., those detected in the disease sample), which may include MHC class I alleles and / or MHC class II allotypes. The one or more cell lines may include one or more human cell lines obtained or derived from one or more subjects. For purposes herein, a peptide sequence that is identified using a disease sample but is not represented in the set of reference sample sequences may be identified as a variant coding sequence.

[0274] In some embodiments, collecting immunogenicity index metrics for use in training can be based on HLA typing analysis, which can identify a subject-specific MHC molecular profile. If the subject is human, this profile can be referred to as a human leukocyte antigen (HLA) profile, since the HLA complex is the gene complex that encodes MHC proteins in humans. HLA typing analysis can be performed using samples from the subject (e.g., normal tissue and / or non-disease samples). The profile can be determined using sequencing techniques such as PCR-based sequencing, direct sequencing, and / or next-generation sequencing. HLA typing analysis can include, for example, high-resolution typing (e.g., which excludes null alleles not expressed on the cell surface) or allele-level typing (e.g., referring to accurate nucleotide sequence HLA gene determination). HLA typing analysis can also include low-resolution typing and / or HLA supertyping, which identify a broader family of alleles.

[0275] For any type of sequencing (e.g., HLA typing, in which peptides bind to MHC molecules to identify sequences in a sample), the results may identify one or more nucleic acid sequences or one or more amino acid sequences. Once nucleic acid sequences have been identified and an attention-based model (or other process) has been configured to process amino acid sequences, techniques (e.g., lookup tables) can be used to convert individual codons within the nucleic acid sequence to individual amino acids.

[0276] Some embodiments involve synthesizing peptides (e.g., using a nucleic acid sequence encoding a peptide, such as a selected peptide) or precursors to the selected peptide. The synthesized peptides or precursors can then be used in experiments to identify corresponding presentation and / or binding data (e.g., to validate predicted presentation and / or binding or to generate results for use in training). For example, the experiment may involve assessing the binding affinity of the selected peptide with a particular MHC molecule using an ELISA pull-down assay, a gel shift assay, or a biosensor-based methodology. As another example, the experiment may involve using peptide-MHC immunoprecipitation to collect elution data indicating whether the selected peptide was presented by an MHC molecule, followed by elution and detection of the presented MHC ligand by mass spectrometry.

[0277] In addition to, or instead of, training or validation data indicating whether individual peptides bound to and / or were presented by individual MHC, the training or validation data may indicate whether individual peptides elicited immunogenicity. Immunogenicity results can be determined using in vivo or in vitro testing. Testing one or more selected peptides can be configured to examine one or more immunogenicity factors (e.g., to determine whether a given event occurs and / or the extent to which a given event occurs) and / or immunogenicity (e.g., to determine whether and / or the extent to which a peptide elicits an immune response). Testing can be configured to examine whether administration of a composition (e.g., a vaccine) containing one or more peptides to a given subject (e.g., in which the MHC sequence used during mutant peptide selection has been identified) is effective in preventing or treating a condition (e.g., tumor) or disease (e.g., cancer). The subject can be a human subject.

[0278] Accessing the training dataset may include, for example, retrieving the training dataset from local or remote storage, loading the training dataset, and / or requesting (and receiving) some or all of the training dataset from one or more data stores (e.g., cloud data storage, a server system, or some other data source).

[0279] The training data may contain "positive" instances (e.g., those for which mass spectrometry results indicate that the peptide was presented by an MHC molecule) and "negative" instances (e.g., corresponding to simulated length-matched n-mers (nmers)) from the same protein as the positive instance (e.g., but not detected by mass spectrometry assessment).

[0280] In some cases, the initial training dataset (which may include, for example, variant coding sequences) may contain primarily negative data, in that a relatively small portion of sequence combinations (e.g., peptide-MHC combinations) are found to be associated with actual target interactions. The training dataset may be designed to include negative training data elements. In some embodiments, the negative training data elements may be used to identify amino acids within pseudorandomly selected fragments of the source protein in the positive set (corresponding to observed presentations). For example, the negative training data elements may be simulated based on the positive set. The fragments may be selected to have lengths within a predetermined range (e.g., 8-14 amino acids for MHC-I and 8-30 amino acids for MHC-II, using uniform probability). N- and C-terminal flanking sequences may be retained within the negative training data elements, potentially imposing a maximum length (e.g., 10 amino acids). Any peptide fragments (e.g., at least 9-mers) that overlap with the positive peptides may be discarded from the negative training data.

[0281] In some embodiments, negative training data elements are simulated based on positive data elements. Furthermore, the training data is selected so that a different set of negative training data elements is used for each epoch of the training period. For example, for each epoch, a different "negative subset" of negative peptide sequences can be selected from the entire space of available negative peptide sequences identified based on the positive set of peptide sequences. The negative subset selected for each epoch can be unique in that no negative peptide sequence is repeated in any of the negative subsets for the total number of epochs. Thus, the training data used for each epoch of the training period contains the same positive set of peptide sequences but a completely different set of negative peptide sequences. This technique, sometimes referred to as "negative set switching," can provide overall robustness to training and help reduce the number of false negatives (e.g., false negative indicators / predictions) generated by the machine learning model or ensure that false negatives are not repeated multiple times. Furthermore, with this technique, the machine learning model can be trained with a total number of negative peptide sequences equal to the number of positive peptide sequences multiplied by the number of epochs during the training period.

[0282] In some examples, the number of positive instances in the training data is equal to the number of negative instances in the training data. In some examples, the number of positive instances is less than or greater than the number of negative instances. One or more (e.g., all) of the negative instances in the training data may each match in length with a positive instance in the training data. In some examples, all of the sequences in the training data have the same length.

[0283] Step 1004 includes training a machine learning model using the training dataset. The machine learning model may be, for example, machine learning model 132 of Figures 1 and 3, or the machine learning model may be, for example, machine learning model 532 of Figures 5A-5C.

[0284] Machine learning models can be trained using static or dynamic learning rates. Dynamic learning rates can be generated, for example, using learning rate annealing. Training can be performed, for example, using a classification loss function and / or a regression loss function. The loss function can be based, for example, on mean squared error, median squared error, mean absolute error, median absolute error, entropy-based error, cross-entropy error, and / or binary cross-entropy error. Validation data (e.g., a separate subset of the training dataset used to train the machine learning model) can be used to evaluate the performance of the machine learning model as it is being trained. Training can be terminated when a target performance is achieved and / or when a maximum number of training iterations have been completed and / or when target performance is achieved.

[0285] Step 1006 includes accessing a subject-specific set of variant coding sequences corresponding to the set of mutant peptides. As described above, a variant coding sequence is an example of a peptide sequence. The subject-specific set of variant coding sequences can correspond to a set of mutant peptides, such that each of the subject-specific set of variant coding sequences identifies an amino acid within a corresponding mutant peptide of the set of mutant peptides. In some embodiments, each of the subject-specific set of variant coding sequences identifies one or more amino acids in a mutation. Each of the subject-specific set of variant coding sequences can be associated with a particular subject (e.g., a human subject). The particular subject may have been diagnosed with a symptom, may experience a symptom, and / or may have received test results associated with a particular medical condition (e.g., cancer). For example, the subject-specific set of variant coding sequences can be identified by processing a sample from a tumor. The sample can be included, for example, in set of samples 112 of FIG. 1.

[0286] A subject-specific set of variant coding sequences can be identified using the techniques disclosed herein. For example, a subject-specific set of variant coding sequences can be identified by performing sequencing techniques to identify peptides in disease samples, and comparing the identified peptides with peptides detected in healthy samples or reference databases to identify unique sequences. In some embodiments, when the unique sequences are nucleic acid sequences, each unique nucleic acid sequence can be converted into an amino acid sequence.

[0287] Each subject-specific set of variant-encoding sequences can identify an amino acid within the peptide (which can be an amino acid within a neo-epitope of a neo-antigen). In some cases, one, more, or all of the subject-specific sets of variant-encoding sequences can be part of a corresponding aggregate sequence that further includes the sequence of the N-flank of the peptide and / or the sequence of the C-flank of the peptide.

[0288] Accessing the subject-specific set of variant coding sequences can include, for example, retrieving the subject-specific set of variant coding sequences from local or remote storage and / or requesting the subject-specific set of variant coding sequences from another device. Accessing the subject-specific set of variant coding sequences can include and / or can be performed in conjunction with determining the subject-specific set of variant coding sequences.

[0289] The subject-specific set of variant-encoding sequences can be obtained by identifying peptide sequences in a subject's disease sample and determining which peptide sequences are not represented in the reference, healthy sample, and / or wild-type sequence set. If a healthy sample is used for comparison, the healthy sample can be (but need not have been) collected from the subject.

[0290] Step 1008 includes accessing an IPC sequence corresponding to the IPC. In some embodiments, the IPC sequence can be an MHC sequence. The MHC sequence can include, for example, a pseudosequence of an MHC (e.g., an MHC molecule) in a sample collected from a subject. In some examples, a subject-specific set of MHC sequences and variant coding sequences is identified from the same sample from the subject or multiple samples from the subject (e.g., a diseased sample and a healthy sample). In some examples, a subject-specific set of MHC sequences and variant coding sequences is identified from samples from the subject and one or more other subjects. Thus, in some cases, the MHC sequence can be subject-specific. The MHC sequence can be or can be determined using, for example, sequencing and / or mass spectrometry techniques.

[0291] Accessing the MHC sequence can include, for example, retrieving the MHC sequence from a local or remote storage device and / or requesting the subject-specific MHC sequence from another device. Accessing the MHC sequence can include and / or be performed in conjunction with determining the MHC sequence.

[0292] Step 1010 may include, for example, processing the set of subject-specific variant coding sequences and MHC sequences using a trained machine learning model to generate an output. Step 1010 may include processing each unique combination of subject-specific variant coding sequence and MHC sequence (e.g., variant code-MHC combination or peptide-MHC combination) of the set of subject-specific variant coding sequences to generate an output.

[0293] The output generated by the machine learning model can include the same or similar types of data as included in the training immune activity data used to train the machine learning model. For each unique combination, the machine learning model generates an output that includes at least one of a set of interaction predictions or a set of interaction affinity predictions.

[0294] The interaction predictions in the set of interaction predictions include predictions regarding whether a target interaction will occur between a variant peptide (including a variant coding sequence) and an MHC (including an MHC sequence). For example, the interaction predictions may include binary or categorical predictions regarding whether a variant peptide having an amino acid structure (as indicated by the subject-specific variant coding sequence) will bind to an MHC molecule (presented by an MHC molecule having an amino acid structure as indicated by the MHC sequence). The interaction affinity predictions in the set of interaction affinity predictions include predictions regarding affinity for the target interaction. This affinity may be based, for example, on the strength, propensity, and / or stability of the target interaction. For example, the interaction affinity prediction may include a predicted real binding affinity associated with a variant peptide including an amino acid identified in the subject-specific variant coding sequence and an MHC molecule including an amino acid identified in the MHC sequence.

[0295] Step 1012 includes generating a report based on the output of the machine learning model. The report can be implemented, for example, as report 144 in Figures 1 and 3. The report can be the output or can include the output. In some cases, the report can be a transformed or filtered version of the output.

[0296] In some embodiments, the subject-specific set of variant coding sequences is filtered, ranked, and / or otherwise processed based on the output to generate information for inclusion in a report. For example, the subject-specific set of variant coding sequences can be filtered to exclude sequences whose predicted interaction affinity (e.g., binding affinity) is below a predetermined affinity threshold and / or sequences whose target interaction (e.g., binding to an MHC molecule) is predicted to occur but is unlikely to occur. In some examples, filtering is performed to identify a predetermined number and / or percentage of the subject-specific set of variant coding sequences. For example, filtering can be performed to identify 10, 20, 40, 60, 80, 100, 500, or 1,000 variant coding sequences associated with a relatively high predicted probability (e.g., relative to unselected variant coding sequences in the subject-specific set of variant coding sequences) of whether a mutant peptide will bind to an MHC molecule.

[0297] The report may identify one or more variant coding sequences (e.g., those not excluded from the set) and / or one or more mutant peptides (e.g., related to the selected variant coding sequences). The mutant peptides may be identified, for example, by their name, their sequence, and / or by identifying both the corresponding wild-type sequence and the mutant represented by the variant coding sequence.

[0298] The report may identify one or more predictions associated with one or more variant coding sequences or one or more mutant peptides. The report may include the subject's name. The report may be presented, for example, locally (e.g., sent as a notification on the user device for display on a display system of the user device, etc.) and / or sent to another device (e.g., sent to a cloud computing system, sent to cloud storage, sent to a user device associated with a medical professional or laboratory professional, sent as an email, etc.).

[0299] 11 is an illustration including an exemplary table of training data, according to some embodiments. Table 1100 includes training data 1102 (e.g., a training data set). Training data 1102 may be an example of a portion of training data 133 of FIG. 1. Training data 1102 may be an example of a portion of a training data set, such as the training data set described in step 1002 of FIG. 10.

[0300] The training data 1102 includes an allotype identifier 1106, training N-flank sequences 1108, training peptide sequences 1110, training C-flank sequences 1112, and training MHC sequences 1114 (e.g., MHC pseudosequences), binding affinity 1116 (e.g., normalized binding affinity scaled from 0 to 1), and presentation index 1118 (e.g., elution probability). The binding affinity 1116 indicates the detected (e.g., observed) binding affinity for the binding of the peptide characterized by the training peptide sequences 1110 and the respective MHC characterized by the training MHC sequences 1114. The presentation index 1118 indicates whether binding or presentation of the peptide by the MHC was detected (or observed).

[0301] Exemplary Forecasts Embodiments of the present disclosure may include determining one or more predictions including, but not limited to, immunogenicity, binding affinity, and potential interactions between a variant peptide and an MHC molecule.

[0302] Figure 12 is an exemplary method 1200 for predicting which therapeutic antibodies are likely to have increased immunogenicity risk. As shown in Figure 12, an exemplary sequence 1205 for a therapeutic antibody light chain can include various amino acid mutations relative to the germline (indicated by bold text with square brackets above), as well as various complementarity-determining regions (CDRs) indicated by carats below the letters. A set 1210 of all possible peptides can be generated using a sliding window (e.g., within a range of 9-30 amino acids for a given peptide sequence length), and candidate peptides 1215 can be identified within a specified length (e.g., 12-19 amino acids). For each of the candidate peptides 1215, a binding core can be identified (see 1220, as shown in Figure 12 using double underlining). The set of candidate peptides 1215 can then be filtered to retain only peptides whose binding core contains mutations (see 1225). By filtering peptides whose binding core does not contain any mutations, the method eliminates peptides that are not immunogenic due to their similarity to human peptides.

[0303] Next, for the set of candidate peptides having a binding core containing mutation 1230, the method determines the frequency 1235 with which the binding core appears in a database of B cell receptor binding cores found in healthy individuals (e.g., the frequency of a 9-mer derived from a B cell receptor). If the frequency is high, the candidate peptide is eliminated (again, to eliminate peptides unlikely to be immunogenic), leaving only candidate peptides with binding cores that do not appear or that appear only rarely in the database (see 1240). Next, the presentation probability (e.g., elution probability) 1245 is calculated for the remaining candidate peptides. Presentation probability can be calculated using the methods and systems described in the remainder of this description, for example, as described with respect to Figures 1-11 above. After eliminating any candidate peptides with negative presentation probability, the method identifies unique binding cores from the remaining candidate peptides, thereby arriving at a set 1250 of binding cores that are likely unique presenters. In some embodiments, the method may proceed to count the number of unique binding cores for each allele and calculate a total across all MHC1 alleles and / or MHCII allotypes. The results of this calculation can inform a determination as to whether a therapeutic antibody represents a risk of immunogenicity in a subject. This risk can take the form of a count of unique, likely presenting bound cores, or some other score, such as the number of unique presenting bound cores weighted by elution likelihood, optionally combined with other categorical or numerical information.

[0304] 13 is an illustration of exemplary neo-antigen candidates (mutated antigens) and corresponding potential neo-epitope candidates (mutated peptides), according to some embodiments. When a process such as process 1000 is performed, the mutant peptides can be neo-antigens.

[0305] For relatively long mutant peptides that are neoantigen candidates 1300, it is possible that multiple epitopes (called neoepitopes) that all contain the same mutation or variant may be presented by the MHC molecule. Thus, the immunogenicity of the neoantigen candidate can be predicted based on the predictions generated for each of the neoepitope candidates 1302.

[0306] Immunogenicity can be predicted, for example, by generating a list of all possible neoepitopes that can emerge from a given neoantigen and generating predictions for each of some or all of the neoepitope candidates in the list (flanks comprising the remaining amino acids upstream of the epitope's N-terminus and downstream of the C-terminus, up to 10 amino acids in length). From these presentation predictions, the neoepitope candidate with the greatest likelihood of presentation to the MHC candidate 1304 is selected to represent the entire neoantigen. Alternatively, a summarized representation of multiple candidate neoepitope-MHC pairs can be used to obtain a summarized score representing the neoantigen. Such summarization can be performed by considering all candidate neoepitope-MHC pairs, or by considering the best neoepitope per MHC and then summarizing across all MHC molecules. Summarization can be performed by several mathematical functions, including, for example, taking the arithmetic or harmonic mean of the presentation or binding affinity scores of each candidate neoepitope-HLA pair.

[0307] While Figure 13 is described with respect to neoantigens and neoepitopes, similar techniques can be used for other types of relatively long mutant peptides that contain mutations or variants and have multiple potential epitope candidates. In some embodiments, this technique can be used in conjunction with antibody drug sequences.

[0308] In some embodiments, if the machine learning model results predict that the mutant peptide has low binding affinity to an MHC molecule, it can predict that the neoantigen detected from the subject's disease sample will not induce immunogenicity or will have low immunogenicity. In some embodiments, it can predict that the MHC molecule will not present or is unlikely to present the mutant peptide. In some embodiments, it can predict that the mutant peptide will not elicit an immune response via a T cell receptor. The immunogenicity prediction generated in connection with the mutant peptide can be, for example, numerical (e.g., corresponding to the predicted probability that an immunogenic response will be elicited in response to the mutant peptide and / or corresponding to the predicted strength of any immunogenic response to the mutant peptide), categorical (e.g., predicting no, low, or high immune response), or binary (e.g., predicting whether a given mutant peptide will elicit an immune response in a subject).

[0309] The predicted immunogenicity may further be based on predicted and / or experimental indications of one or more immunogenicity factors. Factors determining immunogenicity may include one or more of the following: (i) protein levels of the mutant peptide precursor; (ii) expression levels of transcripts encoding the mutant peptide precursor; (iii) efficiency of processing of the mutant peptide precursor by the immunoproteasome; (iv) timing of expression of transcripts encoding the mutant peptide precursor; (v) binding affinity of the mutant peptide to a T cell receptor; (vi) position of the mutant amino acid within the variant peptide; (vii) solvent exposure of the mutant peptide when bound to an MHC molecule; (vii) solvent exposure of the variant amino acid when bound to an MHC molecule; (x) content of aromatic residues in the peptide; (xi) properties of the variant amino acid compared to the wild-type residue; (xii) properties of the mutant peptide precursor; (xiii) microbial similarity of the mutant peptide to learn about microbial peptides; (xiv) self-similarity or dissimilarity of the mutant peptide to the wild-type proteome; or (xv) thymic expression of the wild-type peptide. Immunogenic factors may also or additionally include the protein sequence of the variant peptide, the length of the variant peptide (e.g., as indicated by the number of amino acids identified within the variant coding sequence), and / or the expression level of the MHC allotype in the subject (e.g., as measured by RNA-Seq or mass spectrometry).

[0310] Binding affinity predictions and / or predictions regarding whether (or the probability of) mutant peptide presentation will occur (e.g., to one or more tumor cells and / or one or more MHC molecules in a subject) can be generated for each set of mutant peptides (e.g., those detected in a disease sample from a subject) according to the techniques disclosed herein (e.g., using an attention-based machine learning model). These predictions can be used to select an incomplete subset of the set (e.g., less than 50% of the set, less than 25% of the set, less than 10% of the set, less than 5% of the set, and / or less than 1% of the set). The incomplete subset can be selected using one or more relative thresholds (e.g., to identify mutant peptides in the set that have the most stable binding to MHC molecules and / or the highest likelihood of presentation compared to others in the group) or one or more absolute thresholds. For example, each selected mutant peptide can have a binding affinity with an MHC with a relatively strong affinity value (e.g., within the range of the best 50%, best 25%, best 10%, or best 5% affinity values ​​in the set) and / or an absolutely strong affinity value (e.g., having an affinity value better than a predetermined threshold / cutoff, such as 5000 nM, 1000 nM, or 500 nM). An incomplete subset of the set can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more mutant peptides, regardless of the predetermined affinity value threshold / cutoff. An incomplete subset of the set can include 20 or more neoantigens or 30 or more mutant peptides.

[0311] In some cases, the machine learning model generates predictions corresponding to one or more potential interactions between the mutant peptide and the MHC molecule. For example, the machine learning model may predict the binding affinity of the MHC molecule and the mutant peptide. Additionally or alternatively, the machine learning model may predict whether the MHC molecule will present the mutant peptide. The machine learning model may receive and process (e.g., using one or more processing layers) as input the sequence or subsequence of the MHC-I molecule and the variant coding sequence associated with the mutant peptide.

[0312] In some cases, the machine learning model generates predictions corresponding to one or more potential interactions between a mutant peptide, an MHC sequence or subsequence, and a T cell receptor (e.g., instead of, or in addition to, generating predictions corresponding to one or more potential interactions between the mutant peptide and an MHC molecule). The machine learning model can then predict, for example, the binding affinity between the mutant peptide and a T cell receptor and / or whether the mutant peptide will activate and / or elicit an immune response in a T cell. The machine learning model may receive and process (e.g., using one or more self-attention layers) as inputs the sequence or subsequence of the T cell receptor, the sequence or subsequence of the MHC, and the variant coding sequence of the mutant peptide.

[0313] The prediction generated related to a mutant peptide can be, for example, numerical (e.g., corresponding to the predicted probability that a subject's MHC molecule will present the mutant peptide on its cell surface, or the predicted fraction of the subject's tumor cells that will present the mutant peptide), categorical (e.g., predicting that presentation of the mutant peptide by the subject's MHC molecule will be absent, rare, or frequent), or binary (e.g., predicting whether the mutant peptide will be expressed by the subject's MHC molecule). Presentation predictions can (but need not) be normalized and / or represent conditional predictions. For example, a presentation prediction can correspond to a prediction regarding whether a subject's MHC molecule will present the mutant peptide if the mutant peptide is stably bound to the MHC molecule.

[0314] II.D. Exemplary Identification of Input Data for Machine Learning Models The exemplary methods and systems for identifying input data described herein can be used to identify input data for, for example, machine learning model 132 of FIGS. 1 and 3, any of the workflows of FIGS. 4A-4D, and / or machine learning model 532 described in FIGS. 5A-5C.

[0315] Each set of mutant peptides associated with a given subject can be analyzed using a machine learning model to generate one or more predictions regarding the binding affinity, presentation probability, and / or immunogenicity of the mutant peptides. To generate these predictions, the machine learning model can receive and process peptide (e.g., coding) sequences corresponding to the mutant peptides and one or more other sequences or subsequences (e.g., corresponding to MHC-I molecules, MHC-II molecules, or T cell receptors). In some examples, predictions are generated for each set of peptide sequences (e.g., a set of variant coding sequences corresponding to a set of mutant peptides). The set of mutant peptides can correspond to peptides present in disease samples collected from the subject but not observed in one or more non-disease samples (e.g., from the subject or another subject).

[0316] Various methods are available for identifying a set of mutant peptides associated with a given subject. Mutations may be present in the genome, transcript, proteome, or exome of the subject's diseased cells, but may not be present in non-disease samples, such as a non-disease sample from the subject or another subject. Mutations include, but are not limited to, (1) non-synonymous mutations resulting in different amino acids in the protein; (2) read-through mutations in which a stop codon is altered or deleted, resulting in translation of a longer protein with a novel tumor-specific sequence at the C-terminus; (3) splice site mutations resulting in the inclusion of an intron in the mature mRNA, thus resulting in a unique tumor-specific protein sequence; (4) chromosomal rearrangements (i.e., gene fusions) resulting in a chimeric protein with a tumor-specific sequence at the junction of two proteins; and (5) frameshift insertions or deletions resulting in a new open reading frame with a novel tumor-specific protein sequence. Mutations may also include one or more non-frameshift indels, missense or nonsense substitutions, splice site changes, genomic rearrangements or gene fusions, or any genomic or expression changes that result in neoORFs.

[0317] For example, peptides with mutations or mutant polypeptides resulting from splice site, frameshift, readthrough or gene fusion mutations in diseased cells can be identified by sequencing the DNA, RNA or protein in a diseased sample and comparing the resulting sequences with sequences from a non-diseased sample.

[0318] In some embodiments, whole genome sequencing (WGS) or whole exome sequencing (WES) data from disease and non-disease samples can be obtained and compared. Following alignment of the non-disease and disease sample reads to the human reference genome, somatic variants, including single nucleotide variants (SNVs), gene fusions, and insertion or deletion variants (indels), can be detected using a variant calling algorithm. One or more variant callers can be used to detect different somatic variants (e.g., SNVs, gene fusions, or indels).

[0319] In some examples, mutant peptides are identified based on transcriptome sequences in disease samples from individuals.For example, whole transcriptome sequences or partial transcriptome sequences (for example, by methods such as RNA-Seq) can be obtained from disease tissues of individuals and subjected to sequencing analysis.The sequences obtained from disease tissue samples can then be compared with sequences obtained from reference samples.Optionally, disease tissue samples are subjected to whole transcriptome RNA-Seq.Optionally, transcriptome sequences are "enriched" for specific sequences before being compared with reference samples.For example, specific probes can be designed to enrich for specific desired sequences (for example, disease-specific sequences) before being subjected to sequencing analysis.

[0320] In some embodiments, transcriptome sequencing techniques include, but are not limited to, RNA poly(A) library sequencing, microarray analysis, parallel sequencing, massively parallel sequencing, PCR, and RNA-Seq. RNA-Seq is a high-throughput technique for sequencing part or substantially all of a transcriptome. Briefly, an isolated population of transcriptome sequences is converted into a library of cDNA fragments with adapters attached to one or both ends. Each cDNA molecule is then analyzed, with or without amplification, to obtain short stretches of sequence information, typically 30-400 base pairs. These fragments of sequence information are then aligned to a reference genome, reference transcripts, or assembled de novo to reveal the structure (i.e., transcription boundaries) and / or expression levels of the transcripts.

[0321] Once obtained, the sequence in the diseased sample can be compared to the corresponding sequence in the reference sample. Sequence comparison can be performed at the nucleic acid level by aligning the nucleic acid sequence in the diseased tissue with the corresponding sequence in the reference sample. Gene sequence variations resulting in one or more changes in the encoded amino acids are then identified. Alternatively, sequence comparison can be performed at the amino acid level, i.e., the nucleic acid sequence is first converted in silico to an amino acid sequence before comparison. Either an amino acid-based approach or a nucleic acid-based approach can be used to identify one or more mutations (e.g., one or more point mutations) in a peptide. With regard to the nucleic acid-based approach, the discovered variants can be used to identify one or more nucleic acid sequences (e.g., DNA sequences, RNA sequences, or mRNA sequences) that give rise to a given observable mutant protein (e.g., via a lookup table that associates individual peptide mutations with multiple codon variants).

[0322] In some embodiments, comparison of sequences from disease samples to sequences of a reference sample can be completed by techniques such as manual alignment, FAST-All (FASTA), or Basic Local Alignment Search Tool (BLAST). In some embodiments, comparison of sequences from disease samples to sequences of a reference sample can be completed using short read aligners, such as GSNAP, BWA, and STAR.

[0323] In some embodiments, the reference sample is a matched disease-free sample. As used herein, a "matched" disease-free tissue sample is one selected from the same or similar samples, e.g., samples from the same or similar tissue type as the diseased sample. In some embodiments, the matched disease-free and diseased tissues may be derived from the same individual. The reference sample described herein may be a disease-free sample from the same individual. In some embodiments, the reference sample is a disease-free sample from a different individual (e.g., an individual without the disease). In some embodiments, the reference sample is obtained from a population of different individuals. In some embodiments, the reference sample is a database of known genes associated with an organism. In some embodiments, the reference sample may be derived from a cell line. In some embodiments, the reference sample may be a combination of known genes associated with an organism and genomic information from a matched disease-free sample. In some embodiments, the variant coding sequence may contain point mutations in the amino acid sequence. In some embodiments, the variant coding sequence may contain amino acid deletions or insertions.

[0324] In some embodiments, a set of variant coding sequences is first identified based on genome and / or nucleic acid sequences. This initial set is then further filtered to obtain a narrower set of expressed variant coding sequences based on the presence of the variant coding sequences in a transcriptome sequencing database (and thus, are said to be "expressed"). In some embodiments, the set of variant coding sequences is reduced by at least about 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, or more by filtering the transcriptome sequencing database.

[0325] Alternatively, protein mass spectrometry can be used to identify or verify the presence of mutant peptides, such as those bound to MHC proteins on tumor cells. Peptides can be acid-eluted from diseased cells, such as tumor cells, or from HLA molecules immunoprecipitated from tumors, and then identified using mass spectrometry.

[0326] The variant peptide can have, for example, 5 or more, 8 or more, 11 or more, 15 or more, 20 or more, 40 or more, 80 or more, 100 or more, 120 or less, 100 or less, 80 or less, 60 or less, 50 or less, 40 or less, 30 or less, 25 or less, 20 or less, 18 or less, 15 or less, or 13 or less amino acids.

[0327] Tumor-specific T cell receptor sequences can also be identified by, for example, single-cell T cell receptor sequencing.High-throughput sequencing of T cell repertoire can also or alternatively be carried out to identify the tumor-specific signature of specific disease.MHC-I sequence and / or MHC-II sequence can be determined by, for example, HLA genotyping or mass spectrometry.

[0328] II.E. Exemplary Identification of Training Data for Machine Learning Models The exemplary methods and systems for identifying training data described herein can be used to identify training data for, for example, machine learning model 132 of Figures 1 and 3, any of the workflows of Figures 4A-4D, and / or machine learning model 532 described in Figures 5A-5C. For example, these methods and systems can be used to identify training data 133 of Figure 1.

[0329] The training set can be generated using data collected from multiple other samples (e.g., potentially associated with one or more other subjects). Each of the multiple other samples can include, for example, tissue (e.g., a biopsy), a single cell, multiple cells, cell fragments, or an aliquot of bodily fluid. In some cases, the samples are collected from a different type of subject compared to the subject associated with the input data processed by the trained model. For example, a machine learning model can be trained using training data collected by processing samples from one or more cell lines, and the trained machine learning model can be used to process input data determined by processing one or more samples from a subject.

[0330] The training dataset can include multiple training elements. Each of the multiple training elements can include input data including a set of peptide sequences (including a set of wild-type or variant-encoding sequences) that encode and / or represent any variants in the corresponding peptide, as well as subsequences or pseudosequences of MHC molecules. The input data can be collected according to one or more of the techniques disclosed herein.

[0331] Each training element can also include one or more experimental results. The experimental results can indicate whether and / or to what extent each of one or more specific types of interactions occurs between a wild-type or mutant peptide (associated with a variant coding sequence in the training element) and an MHC molecule (associated with an MHC molecule subsequence in the training element). A specific type of interaction can include, for example, binding of the peptide to an MHC molecule and / or presentation of the peptide by an MHC molecule on the surface of a cell (e.g., a tumor cell).

[0332] The results can include the binding affinity between the peptide and the MHC molecule. The results can include or be based on qualitative and / or quantitative data characterizing whether a given peptide binds to a given MHC molecule, the strength of such binding, the stability of such binding, and / or the propensity for such binding to occur. For example, binary binding affinity indicators or qualitative binary affinity results can be generated using ELISA, pull-down assays, gel shift assays, biosensor-based methodologies such as surface plasmon resonance, isothermal titration colorimetry, biolayer interferometry, or microscale thermophoresis.

[0333] The results can, for example, additionally or alternatively, characterize whether and / or the probability that a given MHC molecule presents a given peptide. MHC ligands can be immunoprecipitated from the sample. Subsequent elution and mass spectrometry can be used to determine whether the MHC molecule presents the ligand.

[0334] Filtering example training data Figure 17 shows a plot of a latent space including multiple peptide vectors, according to some embodiments. Each peptide vector in Figure 17 corresponds to a BOS token embedding (e.g., BOS token embedding 408a in Figure 4A) of a peptide sequence for a given sample. In some examples, a sample refers to a row in a dataset, and the row represents a peptide, MHC, TCR, or a combination thereof. In the latent space, each peptide vector is reduced to two dimensions (e.g., using any dimensionality reduction technique described herein) and plotted as a single dot. The color of each dot represents the allele to which the peptide is bound.

[0335] 17, peptides of the same color (i.e., binding to the same allele) are generally close to each other in latent space and therefore form clusters such as cluster 1700. However, in region 1702, dots of the same color appear relatively scattered and do not form a clear cluster. Because peptides associated with various random alleles should generally not occupy the same region of latent space, the scattered distribution may indicate experimental error in the peptide vector data corresponding to region 1702.

[0336] The above experimental errors can be identified and removed from the training data described herein (e.g., training data for processing block 314). Specifically, for a given peptide, a system (e.g., computing platform 102 of FIG. 1) can identify the K closest neighbors of the peptide in latent space to generate a motif. For example, a motif can be generated by calculating the probability of each amino acid at each position of a peptide given a group of peptides with the same length (i.e., the K nearest neighbors) and converting the probability information into information entropy in bits. In other words, each motif indicates the probability of an amino acid occurring at a given position in the peptide. In some embodiments, an information metric can be extracted for each peptide based on the motif to quantify whether the position is related to the pattern of occurrence of amino acids (which can indicate experimental error). In some embodiments, the information content is calculated based on the number of bits of information for up to two positions. In some embodiments, data associated with low information content can be removed from the training data.

[0337] FIG. 17 shows an exemplary motif 1704 corresponding to the K nearest neighbors of cluster 1700. As shown in motif 1704, at position 0 (x-axis), most of the space is occupied by I and V, indicating that I and V occur frequently at this position. In contrast, at positions 1 or 2 (x-axis), no single amino acid occupies significantly more space than the other amino acids. Thus, position 0 is associated with high information content, while positions 1 and 2 are associated with low information content due to the lack of patterns in binding. The information content of the peptide can then be determined accordingly, in bits. In some examples, the position weight matrix (PWM) of the peptide with its nearest neighbor peptides is determined, and the information content is calculated using the KL divergence relative to a baseline PWM calculated from the human peptidome. In some examples, the Shannon entropy is calculated as the information content.

[0338] FIG. 18 shows a histogram illustrating counts of peptides with different levels of information content, according to some embodiments. In some examples, the input data includes a dataset of peptides and HMCs. For each peptide, information content is calculated (e.g., based on each peptide and neighboring peptides described herein). The X-axis refers to the peptide's information content, which can be quantified in bits. Specifically, for a peptide, the system (e.g., computing platform 102 in FIG. 1 ) can identify the peptide's K closest neighbors in the latent space and generate a motif. For example, a motif can be generated by calculating the probability of each amino acid at each position in the peptide given a group of peptides with the same length (i.e., K nearest neighbors) and converting the probability information into information entropy (in bits) as the peptide's information content. The Y-axis refers to the number of peptides with a particular level of information content. The histogram shows a large number of peptides (i.e., 1800) in the dataset with relatively low information content. A low information content may indicate a lack of pattern in binding (i.e., all positions in the peptide are equally random) and may indicate experimental error in the dataset. Therefore, those peptides with relatively low information content (eg, below a threshold) can be excluded from the training data.

[0339] FIG. 19A shows a protein space colored by protein expression, according to some embodiments. In FIG. 19A, each dot represents a protein vector (e.g., a dimensionally reduced version of the protein sequence embedding 422 in FIG. 4B). Furthermore, blue indicates relatively low-expressing proteins, and red indicates relatively high-expressing proteins. As shown in FIG. 19A, the major clusters have a continuous gradient from blue to red. FIG. 19A demonstrates that a protein language model (e.g., PLM444) learns a representation of this protein, even though it was not trained using explicit protein expression data. FIG. 19B shows the cellular compartmentalization of different proteins and where they appear in the latent space. FIG. 19B shows how the techniques disclosed herein can predict the space / location of source proteins by cellular compartment. In some instances, it is desirable to predict peptides that bind to MHCIs; such binding and presentation may occur more frequently if the peptides are derived from source proteins located intracellularly. Conversely, in some other instances, it is desirable to predict peptides that bind to MHCII, and such binding and presentation may occur more frequently if the peptides are derived from source proteins that are primarily extracellular.

[0340] Figure 20 shows exemplary performance data according to some embodiments. Specifically, Figure 20 shows the average accuracy values ​​of a baseline algorithm for the MHC Class I and MHC Class II datasets. The baseline algorithm does not incorporate processing of protein data (e.g., processing of protein sequence embedding 422 in Figure 4B) or processing of MHC data (e.g., processing of MHC sequence embedding 426 in Figure 4B). Instead, the baseline algorithm uses one or more transformer stages to generate a BOS+MHC sequence representation using BOS tokenized MHC sequences, and then uses one or more transformer stages to generate a transformed BOS+MHC sequence representation (e.g., processing of the BOS tokenized MHC sequences in Figure 4A). As shown in FIG. 20, incorporating protein information (e.g., processing protein sequence embedding 422 in FIG. 4B), incorporating MHC sequence embedding (e.g., processing MHC sequence embedding 426 in FIG. 4B), and incorporating both (e.g., workflow 420 in FIG. 4B incorporating both processing protein sequence embedding 422 and processing MHC sequence embedding 426) improves the performance of the baseline algorithm.

[0341] Exemplary Pharmaceutically Acceptable Compositions In some embodiments, for each of a set of mutant peptides (e.g., detected in a subject's sample), one or more techniques disclosed herein are used to predict whether the mutant peptide will bind to the subject's MHC molecule (or the strength, stability, and / or incidence of such binding) and / or whether the subject's MHC molecule will present the mutant peptide (and / or the incidence of such presentation). The prediction can be used to select an incomplete subset of mutant peptides (e.g., those predicted to have a high likelihood of MHC presentation of the mutant peptide). The selection can include, for each mutant peptide, comparing a metric corresponding to the predicted metric to an absolute threshold and / or comparing the predicted metric with metrics of other mutant peptides (e.g., thereby performing a relative comparison). Each selected mutant peptide can be identified as having one or more of the following: a high likelihood of being presented on the tumor cell surface; a high likelihood of being able to induce a tumor-specific immune response; a high likelihood of being presented to naive T cells by antigen-presenting cells (e.g., dendritic cells); a low likelihood of being inhibited by central or peripheral tolerance; or a low likelihood of being able to induce an autoimmune response against normal tissue in the subject.

[0342] As one non-limiting example, the selection may include identifying each of a subject-specific set of variant coding sequences for which the predicted binding affinity is less than 500 nM, and for which the MHC molecule is predicted to present the mutant peptide identified by the variant coding sequence and / or for which the mutant peptide is predicted to elicit an immune response. It will be understood that the output of the model may be on different scales, such that 500 nM can correspond to another value (e.g., 0.42) on a [0,1] scale.

[0343] Each selected mutant peptide can be produced, experimentally tested (e.g., to determine binding affinity, incidence of presentation, and / or other immunological factors), included in a composition (e.g., pharmaceutical composition such as a vaccine and / or treatment), and / or administered to a subject.

[0344] Each set of variant peptides from which binding affinity and presentation predictions are generated can include variant peptides associated with a particular subject (e.g., a particular human subject). Each set of variant peptides can be a disease-specific immunogenic variant peptide identified using a disease-specific sample from an individual. Individual variant-encoding sequences can be identified by sequencing genes and / or nucleic acid sequences (e.g., DNA, RNA, and / or mRNA sequences) in the disease sample and comparing each identified gene and / or nucleic acid sequence with a reference sample sequence. Codons within the genetic and / or nucleic acid sequences indicate the presence of corresponding amino acids in the peptide. Notably, since multiple codons can each encode a given amino acid, a nucleic acid sequence can represent an amino acid sequence (e.g., deterministically), but the same amino acid sequence can be encoded by other nucleic acid sequences.

[0345] Some embodiments include producing a composition based on one or more selected mutant peptides (or multiple nucleic acids encoding one or more selected mutant peptides). For example, each of the one or more selected mutant peptides can be predicted to bind to and be presented by an MHC molecule of interest (e.g., at least to a threshold extent). The composition can include one or more selected mutant peptides, one or more precursors to the one or more selected mutant peptides, one or more polypeptide sequences corresponding to the one or more selected mutant peptides, RNA (e.g., mRNA) corresponding to the one or more selected mutant peptides, DNA corresponding to the one or more selected mutant peptides, cells (e.g., antigen-presenting cells) containing nucleic acids encoding such peptides, plasmids corresponding to the one or more selected mutant peptides, and / or vectors corresponding to the one or more selected mutant peptides.

[0346] The composition may include mutant peptides corresponding to a single selected variant coding sequence. The composition may include mutant peptides and / or mutant peptide precursors corresponding to multiple selected variant coding sequences. A subset of peptide candidates (e.g., the highest-representing predictions associated with 5, 10, 15, 20, 30, or any number in between) may be used for further precursor development.

[0347] One, more, or all of the mutant peptides in the composition can each have a length of, for example, about 7 to about 40 amino acids (e.g., about 7, 8, 9, 10, 11, 12, 13, 14, 15, 17, 20, 22, 25, 30, 35, 40, 45, 50, 60, or 70 amino acids). In some embodiments, the length of one or more (e.g., all) of the mutant peptides in the composition is within a predetermined range (e.g., 8 to 11 amino acids, 8 to 12 amino acids, or 8 to 15 amino acids). In some embodiments, one or more (e.g., all) of the mutant peptides in the composition are each about 8 to 10 amino acids in length. One or more (e.g., all) of the mutant peptides in the composition can be in their isolated form. One or more (e.g., all) of the mutant peptides in the composition can each be a "long peptide" produced by adding one or more peptides to the end (or each end) of the mutant peptide. Each of one or more (eg, all) of the mutant peptides in the composition may be tagged and may be a fusion protein and / or a hybrid molecule.

[0348] In some embodiments, compositions can be developed by using one or more nucleic acids encoding the peptide. The nucleic acid can include DNA, RNA, and / or mRNA. Given that any of multiple codons can encode a given amino acid, the codons can be selected to optimize or enhance expression in a given type of organism, for example. Such selection can be based on the frequency with which each of the multiple potential codons is used by a given type of organism, the translation efficiency of each of the multiple potential codons in a given type of organism, and / or the degree of bias of a given type of organism toward each of the multiple potential codons.

[0349] The composition may include a polynucleotide construct (e.g., a DNA construct or an RNA construct). A polynucleotide construct is an artificially constructed segment of nucleic acid that can be "implanted" into a target tissue or cell. The polynucleotide construct includes a DNA or RNA (e.g., mRNA) insert containing a nucleotide sequence encoding one or more selected mutant peptides. To increase antigen presentation (e.g., presentation of one or more selected mutant peptides by MHC molecules), the polynucleotide construct may further include modifications developed for improved antigen presentation and therefore improved immunogenicity of one or more selected mutant peptides. In some examples, the modification is the incorporation of transmembrane and cytoplasmic regions of the chains of an MHC molecule into the polynucleotide construct.

[0350] To provide an RNA insert with increased stability and translation efficiency, the polynucleotide construct may further include modifications developed to improve stability and translation, and therefore improve the immunogenicity of one or more selected mutant peptides. In some examples, the modification includes incorporating into the polynucleotide construct a nucleic acid sequence having at least two copies of the 3' untranslated region of the human β-globin gene. In some examples, the modification includes incorporating a nucleic acid sequence encoding a 3' untranslated region, such as the F1 3' UTR.

[0351] In some examples, the composition may include a nucleic acid encoding the mutant peptide or a precursor of the mutant peptide. The nucleic acid may include sequences flanking the sequence encoding the mutant peptide (or its precursor). In some examples, the nucleic acid includes epitopes corresponding to two or more selected variant coding sequences. In some examples, the nucleic acid is DNA having a polynucleotide sequence encoding the mutant peptide or precursor.

[0352] In some examples, the nucleic acid is RNA. In some examples, the RNA is reverse transcribed from a DNA template having a polynucleotide sequence encoding the mutant peptide or precursor. In some examples, the RNA is mRNA. In some examples, the RNA is naked mRNA. In some examples, the RNA comprises modified mRNA (e.g., mRNA protected from degradation with protamine, mRNA comprising a modified 5'CAP structure, or mRNA comprising modified nucleotides). In some embodiments, the RNA comprises single-stranded mRNA.

[0353] To provide an RNA insert with increased stability and expression, the polynucleotide construct may further comprise modifications developed to improve stability and expression, and thus improve immunogenicity of one or more selected mutant peptides. In some instances, the modification is the incorporation of a cap (such as a 5'-cap structure) at the end of the RNA. The cap structure may be the D1 diastereomer of β-S-ARCA.

[0354] In order to deliver the polynucleotide construct to antigen-presenting cells with high selectivity, the composition may further comprise cationic liposomes or lipoplexes to improve the uptake of the polynucleotide construct and thus improve the immunogenicity of one or more selected mutant peptides. In some examples, the composition comprises nanoparticles containing the polynucleotide construct. The nanoparticles may be lipoplexes containing one or more lipids, such as DOTMA and DOPE.

[0355] The composition may include cells containing the mutant peptide and / or nucleic acid encoding the mutant peptide. The composition may further include one or more appropriate vectors and / or one or more delivery systems for the mutant peptide. In some embodiments, the composition may include a nucleic acid encoding the mutant peptide. In some examples, the mutant peptide and / or cell containing the nucleic acid encoding the mutant peptide is a non-human cell, such as a bacterial cell, a protozoan cell, a fungal cell, or a non-human animal cell. In some examples, the mutant peptide and / or cell containing the nucleic acid encoding the mutant peptide is a human cell. In some examples, the human cell is an immune cell. In some examples, the immune cell is an antigen-presenting cell (APC). In some examples, the APC is a professional APC such as a macrophage, monocyte, dendritic cell, B cell, or microglia. In other examples, the professional APC is a macrophage or dendritic cell. In some examples, the mutant peptide and / or APC containing a nucleic acid sequence encoding the mutant peptide is used as a cellular vaccine, thereby inducing a CD4+ or CD8+ immune response. In another example, a composition used as a cellular vaccine comprises mutant peptide-specific T cells primed by APCs comprising the mutant peptide and / or a nucleic acid sequence encoding the mutant peptide.

[0356] The composition may include a pharmaceutically acceptable adjuvant, a pharmaceutically acceptable excipient, an immunomodulator, a checkpoint protein, a PD-1 antagonist (e.g., an anti-PD-1 antibody) and / or a PD-L1 antagonist (e.g., an anti-PD-L1 antibody). An adjuvant refers to any substance that, when mixed into a composition, modifies the immune response to the mutant peptide. An adjuvant may be conjugated, for example, with an immunostimulatory agent. An excipient may increase the molecular weight of a particular mutant peptide to increase activity or immunogenicity, confer stability, increase biological activity, and / or increase serum half-life.

[0357] The pharmaceutically acceptable composition can be a vaccine, which can include a personalized vaccine specific to (e.g., and potentially developed for) a particular subject. For example, an MHC sequence can be identified using a sample from a particular subject, and a composition can be developed and / or used to treat the particular subject.

[0358] The vaccine can be a nucleic acid vaccine. The nucleic acid can encode a mutant peptide or a precursor of a mutant peptide. The nucleic acid vaccine can include sequences flanking the sequence encoding the mutant peptide (or its precursor). In some examples, the nucleic acid vaccine includes epitopes corresponding to more than one selected variant coding sequence. In some examples, the nucleic acid vaccine is a DNA-based vaccine. In some examples, the nucleic acid vaccine is an RNA-based vaccine. In some examples, the RNA-based vaccine includes mRNA. In some examples, the RNA-based vaccine includes naked mRNA. In some examples, the RNA-based vaccine includes modified mRNA (e.g., mRNA protected from degradation with protamine, mRNA including a modified 5'CAP structure, or mRNA including modified nucleotides). In some embodiments, the RNA-based vaccine includes single-stranded mRNA.

[0359] Nucleic acid vaccines can include personalized neoantigen-specific therapies tailored to specific subjects for use as part of next-generation immunotherapy. Personalized vaccines can be designed by first detecting mutant peptides in a sample from a specific subject, and then predicting, for each detected mutant peptide, whether and / or to what extent the peptide will bind to the specific subject's MHC, be presented by the MHC, bind to the specific subject's T cell receptor, and / or elicit an immune response. Based on these predictions, a subset of the detected mutant peptides can be selected (e.g., a subset having at least 1, at least 2, at least 3, at least 5, at least 8, at least 10, at least 12, at least 15, at least 18, up to 40, up to 30, up to 25, up to 20, up to 18, up to 15, and / or up to 10 mutant peptides). For each selected mutant peptide, a synthetic mRNA sequence encoding the mutant peptide can be identified. mRNA vaccines can include mRNA (encoding part or all of the mutant peptide) complexed with lipids to form an mRNA-lipoplex. Administration of a vaccine containing mRNA-lipoplex can result in mRNA stimulation of TLR7 and TLR8, leading to T cell activation by dendritic cells.Furthermore, administration can result in translation of the mRNA into mutant peptides, which can then bind to and be presented by MHC molecules, inducing T cell responses.

[0360] The composition may comprise a substantially pure mutant peptide, a substantially pure precursor thereof, and / or a substantially pure nucleic acid encoding the mutant peptide or precursor thereof. The composition may comprise one or more suitable vectors and / or one or more delivery systems for containing the mutant peptide, its precursor, and / or the nucleic acid encoding the mutant peptide or precursor thereof. Suitable vectors and delivery systems include viruses, such as systems based on adenovirus, vaccinia virus, retrovirus, herpesvirus, adeno-associated virus, or hybrids containing elements of more than one virus. Non-viral delivery systems include cationic lipids and cationic polymers (e.g., cationic liposomes). In some embodiments, physical delivery, such as using a "gene gun," may be used.

[0361] In certain embodiments, the RNA-based vaccine comprises, in a 5' to 3' direction, an RNA molecule comprising: (1) a 5' cap; (2) a 5' untranslated region (UTR); (3) a polynucleotide sequence encoding a secretory signal peptide; (4) a polynucleotide sequence encoding one or more mutant peptides resulting from cancer-specific somatic mutations present in a tumor specimen; (5) a polynucleotide sequence encoding at least a portion of the transmembrane and cytoplasmic domains of a major histocompatibility complex (MHC) molecule; (6) a 3' UTR comprising: (a) a 3' untranslated region of an Amino-Terminal Enhancer of Split (AES) mRNA or a fragment thereof; and (b) a 3' UTR comprising a non-coding RNA of a mitochondrially encoded 12S RNA or a fragment thereof; and (7) a poly(A) sequence.

[0362] In some embodiments, the RNA molecule comprises a polynucleotide sequence encoding an amino acid linker, wherein the polynucleotide sequence encoding the amino acid linker and a first peptide of the one or more variant peptides forms a first linker-neoepitope module, and the polynucleotide sequence forming the first linker-neoepitope module is located, in a 5' to 3' direction, between the polynucleotide sequence encoding the secretory signal peptide and the polynucleotide sequence encoding at least a portion of the transmembrane and cytoplasmic domains of the MHC molecule. In certain embodiments, the amino acid linker comprises the sequence GGSGGGGSGG. In certain embodiments, the polynucleotide sequence encoding the amino acid linker comprises the sequence GGCGGCUCUGGAGGAGGCGGCUCCGGAGGC.

[0363] In certain embodiments, the RNA molecule further comprises, in a 5' to 3' direction, at least a second linker-epitope module, the at least second linker-epitope module comprising a polynucleotide sequence encoding an amino acid linker and a polynucleotide sequence encoding a neoepitope, the polynucleotide sequence forming the second linker-neoepitope module being located, in a 5' to 3' direction, between the polynucleotide sequence encoding the neoepitope of the first linker-neoepitope module and the polynucleotide sequence encoding at least a portion of the transmembrane and cytoplasmic domains of an MHC molecule, and the neoepitope of the first linker-epitope module is different from the neoepitope of the second linker-epitope module. In certain embodiments, the RNA molecule comprises five linker-epitope modules, each of the five linker-epitope modules encoding a different neoepitope. In certain embodiments, the RNA molecule comprises 10 linker-epitope modules, each of which encodes a different neoepitope, hi certain embodiments, the RNA molecule comprises 20 linker-epitope modules, each of which encodes a different neoepitope.

[0364] In some embodiments, the RNA molecule further comprises a second polynucleotide sequence encoding an amino acid linker, wherein the second polynucleotide sequence encoding the amino acid linker is between the polynucleotide sequence encoding the distal-most neoepitope in the 3' direction and the polynucleotide sequence encoding at least a portion of the transmembrane and cytoplasmic domains of the MHC molecule.

[0365] In certain embodiments, the 5' cap comprises a D1 diastereoisomer of the following structure: TIFF2026500160000006.tif48170

[0366] In certain embodiments, the 5'UTR comprises the sequence UUCUUCUGGUCCCCACAGACUCAGAGAGAACCCGCCACC. In certain embodiments, the 5'UTR comprises the sequence GGCGAACUAGUAUUCUUCUGGUCCCCACAGACUCAGAGAGAACCCGCCACC.

[0367] In certain embodiments, the secretory signal peptide comprises the amino acid sequence MRVMAPRTLILLLSGALALTETWAGS. In certain embodiments, the polynucleotide sequence encoding the secretory signal peptide comprises the sequence AUGAGAGUGAUGGCCCCCAGAACCCUGAUCCUGCUGCUGUCUGGCGCCCUGGCCCUGACAGAGACAUGGGCCGGAAGC.

[0368] In certain embodiments, at least a portion of the transmembrane and cytoplasmic domains of the MHC molecule comprise the amino acid sequence IVGIVAGLAVLAVVVIGAVVATVMCRRKSSGGKGGSYSQAASSDSAQGSDVSLTA. In certain embodiments, the polynucleotide sequence encoding at least a portion of the transmembrane and cytoplasmic domains of the MHC molecule comprises the sequence AUCGUGGGAAUUGUGGCAGGACUGGCAGUGCUGGCCGUGGUGGUGAUCGGAGCCGUGGUGGCUACCGUGAUGUGCAGACGGAAGUCCAGCGGAGGCAAGGGCGGCAGCUACAGCCAGGCCGCCAGCUCUGAUAGCGCCCAGGGCAGCGACGUGUCACUGACAGCC.

[0369] In certain embodiments, the 3' untranslated region of the AES mRNA comprises the sequence CUGGUACUGCAUGCACGCAAUGCUAGCUGCCCCUUUCCCGUCCUGGGUACCCCGAGUCUCCCCCGACCUCGGGUCCCAGGUAUGCUCCCACCUCCACCUGCCCCACUCACCACCUCUGCUAGUUCCAGACACCUCC. In certain embodiments, the non-coding RNA of the mitochondrially encoded 12S RNA comprises the sequence CAAGCACGCAGCAAUGCAGCUCAAAACGCUUAGCCUAGCCACACCCCCACGGGAAACAGCAGUGAUUAACCUUUAGCAAUAAACGAAAGUUUAACUAAGCUAUACUAACCCCAGGGUUGGUCAAUUUCGUGCCAGCCACACCG. In certain embodiments, the 3'UTR comprises the sequence CUCGAGCUGGUACUGCAUGCACGCAAUGCUAGCUGCCCCUUUCCCGUCCUGGGUACCCCGAGUCUCCCCCGACCUCGGGUCCCAGGUAUGCUCCCACCUCCACCUGCCCCACUCACCACCUCUGCUAGUUCCAGACACCUCCCAAGCACGCAGCAAUGCAGCUCAAAACGCUUAGCCUAGCCACACCCCCACGGGAAACAGCAGUGAUUAACCUUUAGCAAUAAACGAAAGUUUAACUAAGCUAUACUAACCCCAGGGUUGGUCAAUUUCGUGCCAGCCACACCGAGACCUGGUCCAGAGUCGCUAGCCGCGUCGCU.

[0370] In certain embodiments, the poly(A) sequence comprises 120 adenine nucleotides.

[0371] In certain embodiments, the RNA-based vaccine comprises, in 5' to 3' direction, the polynucleotide sequence: GGCGAACUAGUAUUCUUCUGGUCCCCACAGACUCAGAGAGAACCCGCCACCAUGAGAGUGAUGGCCCCCAGAACCCUGAUCCUGCUGCUGUCUGGCGCCCUGGCCCUGACAGAGACAUGGGCCGGAAGC, a polynucleotide sequence encoding one or more mutant peptides resulting from cancer-specific somatic mutations present in a tumor specimen; and the polynucleotide sequence AUCGUGGGAAUUGUGGCAGGACUGGCAGUGCUGGCCGUGGUGGUGAUCGGAGCCGUGGUGGCUACCGUGAUGUGCAGACGGAAGUCCAGCGGAGGCAAGGGCGGCAGCUACAGCCAGGCCGCCAGCUCU GAUAGCGCCCAGGGCAGCGACGUGUCACUGACAGCCUAGUAACUCGAGCUGGUACUGCAUGCACGCAAUGCUAGCUGCCCCUUUCCCGUCCUGGGUACCCCGAGUCUCCCCCGACCUCGGGUCCCAGGUAUGCUCCCACCUCCACCUGCCCCACUCACCACCUCUGCUAGUUCCAGACACCUCCC Contains an RNA molecule containing AAGCACGCAGCAAUGCAGCUCAAAACGCUUAGCCUAGCCACACCCCCACGGGAAACAGCAGUGAUUAACCUUUAGCAAUAAACGAAAGUUUAACUAAGCUAUACUAACCCCAGGGUUGGUCAAUUUCGUGCCAGCCACACCGAGACCUGGUCCAGAGUCGCUAGCCGCGUCGCU.

[0372] In some embodiments, the mutant peptides described herein (e.g., comprising or consisting of an ordered set of amino acids identified by a variant coding sequence selected based on results from the machine learning techniques described herein) can be used to generate mutant peptide-specific therapeutics, such as antibody therapeutics. For example, the mutant peptides can be used to generate and / or identify antibodies that specifically recognize the mutant peptide. These antibodies can be used as therapeutics. Synthetic short peptides have been used to generate protein-reactive antibodies. The advantage of immunizing with synthetic peptides is that unlimited amounts of pure, stable antigens can be used. This approach involves synthesizing short peptide sequences, coupling them to large carrier molecules, and immunizing subjects with the peptide carrier molecules. The properties of the antibodies depend on primary sequence information. A good response to a desired peptide can usually be generated by careful selection of the sequence and coupling method. Most peptides can elicit a good response. The advantage of anti-peptide antibodies is that they can be prepared immediately after determining the amino acid sequence of the mutant peptide, allowing specific regions of the protein to be specifically targeted for antibody production. By selecting and / or screening mutant peptides predicted by machine learning models for immunogenicity, the resulting antibodies may be more likely to recognize the native protein in tumor conditions. The mutant peptides may be, for example, 15 or less, 18 or less, 20 or less, 25 or less, or 30 or less residues. The mutant peptides may be, for example, 9 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 50 or more, or 70 or more residues. Shorter peptides may improve antibody production.

[0373] Peptide-carrier protein coupling can be used to facilitate the production of high-titer antibodies. Coupling methods can include, for example, site-specific coupling and / or techniques that rely on reactive amino acid functional groups, such as -NH2, -COOH, -SH, and phenolic -OH. Any suitable method used for anti-peptide antibody production can be utilized with the mutant peptides identified by the methods of the present invention. Two such known methods are the multiple antigenic peptide system (MAP) and the lipid core peptide (LCP) method. The advantage of MAP is that no conjugation method is required. No carrier protein or conjugate is introduced into the immunized host. One disadvantage is that it is more difficult to control the purity of the peptide. Furthermore, MAP can bypass the immune response system in some hosts. The LCP method is known to provide higher titers than other anti-peptide vaccine systems and may therefore be advantageous.

[0374] Also provided herein are isolated MHC / peptide complexes containing one or more mutant peptides identified using the techniques disclosed herein. Such MHC / peptide complexes can be used, for example, to identify antibodies, soluble TCRs, or TCR analogs. One type of antibody is called a TCR mimetic because it binds to peptides from tumor-associated antigens in the context of a specific HLA environment. This type of antibody has been shown to mediate lysis of cells expressing the complex on their surface and protect mice from transplanted cancer cell lines expressing the complex. One advantage of TCR mimetics over IgG mAbs is that they can be affinity matured, allowing the molecule to be coupled to immune effector functions via the current Fc domain. These antibodies can also be used to target therapeutic molecules, such as toxins, cytokines, or pharmaceutical agents, to tumors.

[0375] Other types of molecules developed using mutant peptides such as those selected using the methods of the invention using non-hybridoma-based antibody production or production of binding-competent antibody fragments such as anti-peptide Fab molecules on bacteriophage. These fragments can also be conjugated to other therapeutic molecules for tumor delivery such as anti-peptide MHC Fab-immunotoxin conjugates, anti-peptide MHC Fab-cytokine conjugates, and anti-peptide MHC Fab-drug conjugates.

[0376] Exemplary Methods of Treatment Some embodiments include treating a condition (e.g., tumor) or disease (e.g., cancer) in an individual by administering to the individual an effective amount of a composition (e.g., vaccine) comprising one or more selected mutant peptides. The individual can be the same individual from whom the disease sample was collected. In some examples, the vaccine is administered to a different individual compared to the individual from whom the disease sample was collected. The different individual may, for example, be related to the individual from whom the disease sample was collected, may have a genetic risk of developing a particular type of cancer, and / or may have an MHC molecule with one or more (e.g., all) alleles corresponding to the same (or similar) sequences as one or more MHC alleles of the subject from whom the disease sample was collected.

[0377] Some embodiments provide methods of treatment including vaccines, which may be immunogenic vaccines. In some embodiments, methods of treating diseases (such as cancer) are provided that may include administering to an individual an effective amount of a composition described herein, a mutant peptide identified using the techniques disclosed herein, a precursor thereof, or a nucleic acid encoding a mutant peptide (or precursor) identified using the techniques described herein.

[0378] In some embodiments, a method for treating a disease (such as cancer) is provided. The method may include obtaining a sample (e.g., a blood sample) from a subject. T cells can be isolated and stimulated. Isolation can be performed, for example, using density gradient sedimentation (e.g., centrifugation), immunomagnetic selection, and / or antibody complex filtering. Stimulation can include antigen-independent stimulation, which can use, for example, a mitogen (e.g., PHA or ConA), an anti-CD3 antibody (e.g., to bind to CD3 and activate the T cell receptor complex), and / or an anti-CD28 antibody (e.g., to bind to CD28 and stimulate the T cell). One or more mutant peptides can be selected (or can be selected) for use in treating the subject (e.g., based on results generated by a machine learning model corresponding to a prediction regarding whether and / or the extent to which each of a set of mutant peptides will bind to or be presented by an individual's MHC molecule and / or elicit an immune response in the individual). The one or more mutant peptides can be selected based on the techniques disclosed herein, which involve identifying and processing one or more sequence representations (e.g., representations of MHC sequences, sets of variant coding sequences, and / or T cell receptor sequences) associated with a subject. The one or more sequences can be detected using the sample from which the T cells were isolated or a different sample.

[0379] In some examples, one or more mutant peptides (or precursors thereof) can be used to generate mutant peptide (e.g., neoantigen)-specific T cells. For example, peripheral blood T cells can be isolated from a subject and contacted with one or more mutant peptides to induce a mutant peptide-specific T cell population, which can be administered to the subject. In some examples, the T cell receptor sequences of mutant peptide-reactive T cells can be sequenced. If sequencing identifies an ordered set of nucleic acids, each codon of the nucleic acid can be translated into an amino acid (e.g., via probing techniques). Once the T cell receptor sequence (e.g., amino acid T cell receptor sequence) is obtained, T cells can be engineered to contain a T cell receptor that specifically recognizes the mutant peptide. These engineered T cells can then be administered to a subject. In any of the methods provided herein, T cells can be expanded in vitro and / or ex vivo before administration to a subject. The subject can then be administered (e.g., infused) with a composition comprising the expanded T cell population.

[0380] In some examples, methods are provided for treating diseases (such as cancer) that can include administering to an individual a composition comprising one or more mutant peptides (or one or more precursors thereof) in an amount effective to, for example, prime, activate, and expand T cells in vivo.

[0381] In some embodiments, methods for treating diseases (such as cancer) are provided that may include administering to an individual an effective amount of a composition comprising a precursor of a mutant peptide selected using the techniques described herein. In some embodiments, an immunogenic vaccine may include a pharmaceutically acceptable mutant peptide selected using the techniques described herein. In some embodiments, an immunogenic vaccine may include a pharmaceutically acceptable precursor of a mutant peptide selected using the techniques described herein (e.g., protein, peptide, DNA, and / or RNA, etc.). In some embodiments, methods for treating diseases (such as cancer) are provided that may include administering to an individual an effective amount of an antibody that specifically recognizes a mutant peptide selected using the techniques described herein. In some embodiments, methods for treating diseases (such as cancer) are provided that may include administering to an individual an effective amount of a soluble TCR or TCR analog that specifically recognizes a mutant peptide selected using the techniques described herein.

[0382] In some embodiments, the cancer is selected from the group consisting of carcinoma, lymphoma, blastoma, sarcoma, leukemia, squamous cell carcinoma, lung cancer (including small cell lung cancer, non-small cell lung cancer, adenocarcinoma of the lung, and squamous cell carcinoma of the lung), cancer of the peritoneum, hepatocellular carcinoma, gastric cancer or stomach cancer (including gastrointestinal cancer), pancreatic cancer, glioblastoma, cervical cancer, ovarian cancer, liver cancer, bladder cancer, hepatoma, breast cancer, colon cancer, melanoma, endometrial or uterine carcinoma, salivary gland carcinoma, kidney cancer or renal cancer. cancer), liver cancer, prostate cancer, vulvar cancer, thyroid cancer, hepatic carcinoma, head and neck cancer, colorectal cancer, rectal cancer, soft tissue sarcoma, Kaposi's sarcoma, B-cell lymphoma (low-grade / follicular non-Hodgkin's lymphoma (NHL), small lymphocytic (SL) NHL, intermediate-grade / follicular NHL, intermediate-grade diffuse NHL, high-grade immunoblastic NHL, high-grade lymphoblastic NHL, high-grade small noncleaved cell NHL, giant These include large-disease NHL, mantle cell lymphoma, AIDS-related lymphoma, and Waldenstrom's hypertumor globulinemia), chronic lymphocytic leukemia (CLL), acute lymphoblastic leukemia (ALL), melanoma, hairy cell leukemia, chronic myeloblastic leukemia, and post-transplant lymphoproliferative disorder (PTLD), as well as abnormal blood vessel growth associated with phacomatosis, edema (such as that associated with brain tumors), or Meigs' syndrome.

[0383] The embodiments disclosed herein can include identifying and / or implementing a part or all of a personalized medicine strategy. For example, one or more mutant peptides can be selected for use in a vaccine by using a sample from an individual to determine a set of MHC sequences and / or variant coding sequences; and processing the representation of the MHC sequences and variant coding sequences using a machine learning model (e.g., an attention-based machine learning model) disclosed herein. One or more mutant peptides (and / or their precursors) can then be administered to the same individual.

[0384] In some embodiments, methods are provided for treating a disease (such as cancer) in an individual, the methods including: (a) identifying one or more mutant peptides in the individual (e.g., based on results generated by a machine learning model corresponding to a prediction regarding whether and / or the extent to which each of a set of mutant peptides will bind to, be presented by, and / or elicit an immune response in the individual, according to one or more techniques disclosed herein); (b) synthesizing the identified mutant peptide(s) or one or more precursors of the mutant peptides or nucleic acid(s) (e.g., polynucleotides such as DNA or RNA) encoding the identified peptide(s) or peptide precursor(s); and c) administering the mutant peptide(s), mutant peptide precursor(s), or nucleic acid(s) to the individual.

[0385] In some embodiments, methods are provided for treating a disease (such as cancer) in an individual, the methods comprising: (a) identifying one or more mutant peptides in the individual (e.g., based on results generated by a machine learning model corresponding to a prediction regarding whether and / or to what extent each of a set of mutant peptides will bind to, be presented by, and / or elicit an immune response in the individual, in accordance with one or more techniques disclosed herein); and (b) optionally, identifying a set of nucleic acids (e.g., polynucleotides such as DNA or RNA) encoding the identified mutant peptide(s) or one or more precursors of the mutant peptide(s), synthesizing the set of nucleic acids, and administering the set of nucleic acids to the individual.

[0386] In some embodiments, methods are provided for treating a disease (such as cancer) in an individual, the methods including: (a) identifying one or more mutant peptides in the individual (e.g., based on results generated by a machine learning model corresponding to a prediction regarding whether and / or to what extent each of a set of mutant peptides will bind to, be presented by, and / or elicit an immune response in the individual, according to one or more techniques disclosed herein); (b) producing antibodies that specifically recognize the mutant peptides; and (c) administering the peptides to the individual.

[0387] The methods provided herein can be used to treat individuals (e.g., humans) diagnosed with or suspected of having cancer. In some embodiments, the individual can be human. In some embodiments, the individual can be at least about 18, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or 85 years of age. In some embodiments, the individual can be male. In some embodiments, the individual can be female. In some embodiments, the individual may have refused surgery. In some embodiments, the individual may be medically inoperable. In some embodiments, the individual may be at a clinical stage of Ta, Tis, T1, T2, T3a, T3b, or T4. In some embodiments, the cancer may be recurrent. In some embodiments, the individual can be human who exhibits one or more symptoms associated with cancer. In some embodiments, the individual may be genetically or otherwise predisposed to developing cancer (e.g., have risk factors).

[0388] The methods provided herein can be performed in an adjuvant setting. In some embodiments, the methods are performed in a neoadjuvant setting, i.e., the methods can be performed before primary / definitive therapy. In some embodiments, the methods are used to treat individuals who have been previously treated. Any of the treatment methods provided herein can be used to treat individuals who have not been previously treated. In some embodiments, the methods are used as first-line therapy. In some embodiments, the methods are used as second-line therapy.

[0389] In some embodiments, methods are provided for reducing the incidence or burden of metastasis of existing cancer tumors (such as lung metastasis or lymph node metastasis) in an individual, the methods comprising administering to the individual an effective amount of a composition disclosed herein. In some embodiments, methods are provided for extending the time to disease progression of cancer in an individual, the methods comprising administering to the individual an effective amount of a composition disclosed herein. In some embodiments, methods are provided for extending the survival of an individual with cancer, the methods comprising administering to the individual an effective amount of a composition disclosed herein.

[0390] In some embodiments, at least one or more chemotherapeutic agents may be administered in addition to the compositions disclosed herein. In some embodiments, the one or more chemotherapeutic agents may (but need not) belong to different classes of chemotherapeutic agents.

[0391] In some embodiments, methods of treating a disease (such as cancer) in an individual are provided, comprising administering (a) a vaccine disclosed herein (e.g., comprising a mutant peptide or precursor thereof selected based on the machine learning techniques disclosed herein), and (b) an immunomodulatory agent. In some embodiments, methods of treating a disease (such as cancer) in an individual are provided, comprising administering (a) a vaccine disclosed herein (e.g., comprising a mutant peptide or precursor thereof selected based on the machine learning techniques disclosed herein), and (b) an antagonist of a checkpoint protein. In some embodiments, methods of treating a disease (such as cancer) in an individual are provided, comprising (a) a vaccine disclosed herein (e.g., comprising a mutant peptide or precursor thereof selected based on the machine learning techniques disclosed herein), and (b) an antagonist of programmed cell death 1 (PD-1), such as anti-PD-1. In some embodiments, methods are provided for treating a disease (such as cancer) in an individual, the methods comprising administering (a) a vaccine disclosed herein (e.g., comprising a mutant peptide or precursor thereof selected based on the machine learning techniques disclosed herein), and (b) an antagonist of programmed death-ligand 1 (PD-1), such as anti-PD-L1. In some embodiments, methods are provided for treating a disease (such as cancer) in an individual, the methods comprising administering (a) a vaccine disclosed herein (e.g., comprising a mutant peptide or precursor thereof selected based on the machine learning techniques disclosed herein), and (b) an antagonist of cytotoxic T-lymphocyte-associated protein 4 (CTLA-4), such as anti-CTLA-4.

[0392] It will be understood that various disclosures refer to the use of amino acid sequences. Nucleic acid sequences may additionally or alternatively be used. For example, a disease-specific sample may be sequenced to identify a set of nucleic acid sequences that are not present in a corresponding non-disease-specific sample (e.g., from the same subject or a different subject). Similarly, nucleic acid sequences of MHC molecules and / or T cell receptors may be further identified. Each representation of the nucleic acid disease-specific sequence and the MHC molecule (or T cell receptor) may be processed by the attention-based model described herein (e.g., potentially trained using nucleic acid sequences).

[0393] Exemplary model performance An exemplary peptide-MHC (MHC class II) machine learning model (herein "P-MHC-II model") has been developed. This model is an exemplary implementation of machine learning model 132 of FIG. 1. The P-MHC-II model was implemented corresponding to the architecture shown in FIG. 5A. The P-MHC-II model is compared to other models currently available, such as NetMHC-Cpan-4.0 (herein referred to as "Model A"). The P-MHC-II model performed better than Model A for peptide presentation.

[0394] 14A and 14B are plots with exemplary precision-recall (PR) curves, according to some embodiments. Figures 14A and 14B show the performance of the P-MHC-II model compared to Model A. The eluted ligand (EL) test dataset was used to evaluate the proposed predictive performance between the EL output of the P-MHC-II model and that of Model A.

[0395] Figure 14A includes an exemplary plot 1300 showing the performance of the P-MHC-II model, according to some embodiments. Figure 14B includes an exemplary plot 1402 showing the performance of a previously used approach, Model A, on its elution output, according to some embodiments. The dots on the curves in each of plots 1400 and 1402 correspond to score thresholds for the top 10.00% and 9.64% quantiles of scores, respectively. Average precision (AP) represents threshold-independent performance. F1 score, precision, and recall values ​​are based on the respective thresholds.

[0396] Model A values ​​were percentile rank output from the previously used approach. P-MHC-II model values ​​were obtained from the output (of the final node) of the P-MHC-II model. Based on these PR curves, the results in Figures 14A and 14B show that the P-MHC-II model showed improved performance over Model A, with AP values ​​of 0.84 vs. 0.66 for Model A. The AP values ​​of the methods were compared allele-by-allele.

[0397] FIG. 15 is an exemplary plot 1500 comparing exemplary average precision values ​​of the elution-ligand output of Model A and the P-MHC-I model for each allele in a test dataset according to some embodiments.

[0398] 16A-16B are exemplary plots 1600 and 1602 showing P-MHCII Model (BA output) and Model A (BA output) performance, respectively, according to some embodiments.

[0399] Exemplary Computer System 21 is a block diagram of a computer system 2100, according to some embodiments. The computer system 2100 may be an example of one implementation of the computing platform 102 described above in FIG.

[0400] 21 illustrates an example of one or more computing devices 2100 that can be used to determine predicted amino acid-IPC predictions, according to some embodiments. In certain embodiments, the one or more computing devices 2100 may perform one or more steps of one or more methods described or illustrated herein. In certain embodiments, the one or more computing devices 2100 provide functionality described or illustrated herein. In certain embodiments, software running on the one or more computing devices 2100 performs one or more steps of one or more methods described or illustrated herein or provides functionality described or illustrated herein. Certain embodiments include one or more portions of the one or more computing devices 2100.

[0401] The present disclosure contemplates any suitable number of computing systems 2100. The present disclosure contemplates one or more computing devices 2100 taking any suitable physical form. By way of example and not limitation, the one or more computing devices 2100 may be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a mesh of computer systems, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these. Where appropriate, the one or more computing devices 2100 may be single or distributed, spanning multiple locations, spanning multiple machines, spanning multiple data centers, or in a cloud that may include one or more cloud components of one or more networks.

[0402] Where appropriate, one or more computing devices 2100 may perform one or more steps of one or more methods described or illustrated herein without substantial spatial or temporal limitations. By way of example and not limitation, one or more computing devices 2100 may perform one or more steps of one or more methods described or illustrated herein in real time or in batch mode. One or more computing devices 2100 may, where appropriate, perform one or more steps of one or more methods described or illustrated herein at different times or in different locations.

[0403] In particular embodiments, one or more computing devices 2100 include a processor 2102, a memory 2104, a database 2106, an input / output (I / O) interface 2108, a communication interface 2110, and a bus 2112. While this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement. In particular embodiments, the processor 2102 includes hardware for executing instructions, such as instructions making up a computer program. By way of example and not limitation, to execute instructions, the processor 2102 may retrieve (or fetch) instructions from an internal register, an internal cache, the memory 2104, or the database 2106, decode and execute those instructions, and then write one or more results to an internal register, an internal cache, the memory 2104, or the database 2106. In particular embodiments, the processor 2102 may include one or more internal caches for data, instructions, or addresses. This disclosure contemplates processor 2102 including any suitable number of any suitable internal caches, where appropriate. By way of example and not limitation, processor 2102 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). Instructions in an instruction cache may be copies of instructions in memory 2104 or database 2106, and the instruction cache may speed up retrieval of those instructions by processor 2102.

[0404] The data in the data cache may be a copy of data in memory 2104 or database 2106 upon which instructions executing in processor 2102 operate, results of previous instructions executed in processor 2102 for access by subsequent instructions executing in processor 2102 or for writing to memory 2104 or database 2106, or other suitable data. The data cache may speed up read or write operations by processor 2102. The TLB may speed up virtual address translation for processor 2102. In particular embodiments, processor 2102 may include one or more internal registers for data, instructions, or addresses. This disclosure contemplates processor 2102 including any suitable number of any suitable internal registers, where appropriate. Where appropriate, processor 2102 may include one or more arithmetic logic units (ALUs), may be a multi-core processor, or may include one or more processors 2102. While this disclosure describes and illustrates a particular processor, this disclosure contemplates any suitable processor.

[0405] In particular embodiments, memory 2104 includes main memory for storing instructions for processor 2102 to execute or data upon which processor 2102 operates. By way of example and not limitation, one or more computing devices 2100 may load instructions into memory 2104 from database 2106 or another source (such as, for example, another one or more computing devices 2100). Processor 2102 may then load the instructions from memory 2104 into an internal register or cache. To execute the instructions, processor 2102 may retrieve the instructions from the internal register or cache and decode them. During or after execution of the instructions, processor 2102 may write one or more results (which may be intermediate or final results) to an internal register or cache. Processor 2102 may then write one or more of those results to memory 2104.

[0406] In particular embodiments, processor 2102 executes instructions only from one or more internal registers, internal caches, or memory 2104 (as opposed to database 2106 or elsewhere) and operates only on data in one or more internal registers, internal caches, or memory 2104 (as opposed to database 2106 or elsewhere). One or more memory buses (each of which may include an address bus and a data bus) may connect processor 2102 to memory 2104. Bus 2112 may include one or more memory buses, as described below. In particular embodiments, one or more memory management units (MMUs) reside between processor 2102 and memory 2104 and facilitate accesses to memory 2104 requested by processor 2102. In particular embodiments, memory 2104 includes random access memory (RAM). This RAM may be volatile memory, where appropriate. Where appropriate, this RAM may be dynamic RAM (DRAM) or static RAM (SRAM). Moreover, where appropriate, this RAM may be single-ported or multi-ported RAM. This disclosure contemplates any suitable RAM. Memory 2104 may, where appropriate, include one or more memories 2104. Although this disclosure describes and illustrates particular memory, this disclosure contemplates any suitable memory.

[0407] In particular embodiments, database 2106 includes mass storage for data or instructions. By way of example and not limitation, database 2106 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk, a magneto-optical disk, magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Database 2106 may include removable or non-removable (or fixed) media, where appropriate. Database 2106 may be internal or external to one or more computing devices 2100, where appropriate. In particular embodiments, database 2106 is non-volatile solid-state memory. In particular embodiments, database 2106 includes read-only memory (ROM). Where appropriate, this ROM may be mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), flash memory, or a combination of two or more of these. The present disclosure contemplates mass database 2106 taking any suitable physical form. Database 2106 may include, where appropriate, one or more storage control units that facilitate communication between processor 2102 and database 2106. Where appropriate, database 2106 may include one or more databases 2106. Although this disclosure describes and illustrates particular storage devices, this disclosure contemplates any suitable storage device.

[0408] In particular embodiments, I / O interface 2108 includes hardware, software, or both that provide one or more interfaces for communication between one or more computing devices 2100 and one or more I / O devices. One or more computing devices 2100 may include one or more of these I / O devices, where appropriate. One or more of these I / O devices may enable communication between a person and one or more computing devices 2100. By way of example and not limitation, an I / O device may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet, touch screen, trackball, video camera, another suitable I / O device, or a combination of two or more of these. An I / O device may include one or more sensors. This disclosure contemplates any suitable I / O devices and any suitable I / O interface 2108 therefor. Where appropriate, I / O interface 2108 may include one or more device or software drivers that enable processor 2102 to drive one or more of these I / O devices. I / O interface 2108 may include, where appropriate, one or more I / O interfaces 2108. Although this disclosure describes and illustrates a particular I / O interface, this disclosure contemplates any suitable I / O interface.

[0409] In particular embodiments, communication interface 2110 includes hardware, software, or both that provide one or more interfaces for communication (e.g., packet-based communication) between one or more computing devices 2100 and one or more other computing devices 2100 or one or more networks. By way of example and not limitation, communication interface 2110 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network, or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI network. This disclosure contemplates any suitable network and any suitable communication interface 2110 for that network.

[0410] By way of example, and not by way of limitation, one or more computing devices 2100 may communicate with an ad hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), one or more portions of the Internet, or a combination of two or more of these. One or more portions of one or more of these networks may be wired or wireless. By way of example, one or more computing devices 2100 may communicate with a wireless PAN (WPAN) (e.g., a BLUETOOTH WPAN), a Wi-Fi network, a Wi-MAX network, a cellular telephone network (e.g., a Global System for Mobile Communications (GSM) network), other suitable wireless networks, or a combination of two or more of these. The one or more computing devices 2100 may include any suitable communication interface 2110 for any of these networks, where appropriate. The communication interface 2110 may include one or more communication interfaces, where appropriate. Although this disclosure describes and illustrates a particular communication interface, this disclosure contemplates any suitable communication interface.

[0411] In particular embodiments, bus 2112 includes hardware, software, or both that connects one or more components of computing device 2100 to one another. By way of example, and not limitation, bus 2112 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infiniband interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, another suitable bus, or a combination of two or more of these. Bus 2112 may include one or more buses 2112, where appropriate. Although this disclosure describes and illustrates a particular bus, this disclosure contemplates any suitable bus or interconnect.

[0412] As used herein, one or more computer-readable non-transitory storage media may comprise, where appropriate, one or more semiconductor-based or other integrated circuits (ICs) (such as, for example, field programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM drives, secure digital cards or drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these. Computer-readable non-transitory storage media may, where appropriate, be volatile, non-volatile, or a combination of volatile and non-volatile.

[0413] FIG. 22 shows a diagram 2200 of an exemplary artificial intelligence (AI) architecture 2202 (which may be included as part of one or more computing devices 2100, as described above in connection with FIG. 21) that may be utilized to determine one or more predicted amino acid-IPC interactions according to the disclosed embodiments. In particular embodiments, the AI ​​architecture 2202 may be implemented utilizing one or more processing devices, which may include, for example, hardware (e.g., a general-purpose processor, a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a system-on-chip (SoC), a microcontroller, a field programmable gate array (FPGA), a central processing unit (CPU), an application processor (AP), a vision processing unit (VPU), a neural processing unit (NPU), a neural decision processor (NDP), a deep learning processor (DLP), a tensor processing unit (TPU), a neuromorphic processing unit (NPU), and / or any other processing device that may be suitable for processing various molecular data and making one or more decisions based thereon), software (e.g., instructions operating / executing on one or more processing devices), firmware (e.g., microcode), or some combination thereof.

[0414] 22, AI architecture 2202 may include machine learning (ML) algorithms and functions 2204, natural language processing (NLP) algorithms and functions 2206, expert systems 2208, computer-based vision algorithms and functions 2210, speech recognition algorithms and functions 2212, planning algorithms and functions 2214, and robotics algorithms and functions 2216. In certain embodiments, ML algorithms and functions 2204 may include any statistical-based algorithms that may be suitable for finding patterns across large amounts of data (e.g., "big data" such as genomics data, proteomics data, metabolomics data, metagenomics data, transcriptomics data, or other omics data). For example, in certain embodiments, ML algorithms and functions 2204 may include deep learning algorithms 2218, supervised learning algorithms 2220, and unsupervised learning algorithms 2222.

[0415] In particular embodiments, deep learning algorithms 2218 may include any artificial neural network (ANN) that can be utilized to learn deep-level representations and abstractions from large amounts of data. For example, deep learning algorithms 2218 may include ANNs such as perceptrons, multi-layer perceptrons (MLPs), autoencoders (AEs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), long short-term memories (LSTMs), grated recurrent units (GRUs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), and deep Q-networks, neural autoregressive distribution estimation (NADEs), adversarial networks (ANs), attention models (AMs), spiking neural networks (SNNs), deep reinforcement learning, and the like.

[0416] In particular embodiments, supervised learning algorithm 2220 may include any algorithm that can be utilized to apply what has been learned in the past to new data, e.g., using labeled examples to predict future events. For example, starting from an analysis of a known training data set, supervised learning algorithm 2220 may create an inferred function to make a prediction about output values. Supervised learning algorithm 2220 may also compare its output with the correct intended output and find errors in order to correct supervised learning algorithm 2220 accordingly. On the other hand, unsupervised learning algorithm 2222 may include any algorithm that can be applied, for example, when the data used to train unsupervised learning algorithm 2222 is neither classified nor labeled. For example, unsupervised learning algorithm 2222 may study and analyze how a system can infer functions to describe hidden structure from unlabeled data.

[0417] In particular embodiments, NLP algorithms and functions 2206 may include any algorithms or functions that may be suitable for automatically manipulating natural language, such as speech and / or text. For example, NLP algorithms and functions 2206 may include content extraction algorithms or functions 2224, classification algorithms or functions 2226, machine translation algorithms or functions 2228, question answering (QA) algorithms or functions 2230, and text generation algorithms or functions 2232. In particular embodiments, content extraction algorithms or functions 2224 may include means for extracting text or images from electronic documents (e.g., web pages, text editor documents, etc.) for use in other applications, for example.

[0418] In particular embodiments, classification algorithm or function 2226 may include any algorithm that may utilize a supervised learning model (e.g., logistic regression, naive Bayes, stochastic gradient descent (SGD), k-nearest neighbors, decision trees, random forests, support vector machines (SVMs), etc.) to learn from and make new observations or classifications based on data input into the supervised learning model. Machine translation algorithm or function 2228 may include any algorithm or function that may be suitable for automatically converting source text in one language into text in another language, for example. QA algorithm or function 2230 may include any algorithm or function that may be suitable for automatically answering questions posed by humans in natural language, such as those performed by a voice-controlled personal assistant device, for example. Text generation algorithm or function 2232 may include any algorithm or function that may be suitable for automatically generating natural language text.

[0419] In particular embodiments, expert system 2208 may include any algorithms or functions that may be suitable for simulating the judgment and actions of a human or organization with expertise and experience in a particular field (e.g., stock trading, medicine, sports statistics, etc.). Computer-based vision algorithms and functions 2210 may include any algorithms or functions that may be suitable for automatically extracting information from images (e.g., photographic images, video images). For example, computer-based vision algorithms and functions 2210 may include image recognition algorithms 2234 and machine vision algorithms 2236. Image recognition algorithms 2234 may include any algorithms that may be suitable, for example, for automatically identifying and / or classifying objects, places, people, etc. that may be contained in one or more image frames or other display data. Machine vision algorithms 2236 may include any algorithms that may be suitable for enabling a computer to "see" or that may be suitable, for example, for relying on image sensor cameras with specialized optics to acquire images in order to process, analyze, and / or measure various data characteristics for decision-making purposes.

[0420] In particular embodiments, speech recognition algorithms and functions 2212 may include any algorithms or functions that may be suitable for recognizing and translating spoken language into text, such as through automatic speech recognition (ASR), computer speech recognition, speech-to-text (STT) 2238, or text-to-speech (TTS) 2240, for computing purposes to communicate with one or more users via voice. In particular embodiments, planning algorithms and functions 2214 may include any algorithms or functions that may be suitable for generating a series of actions, where each action may include its own set of preconditions to be satisfied before performing the action. Examples of AI planning may include classical planning, reduction to other problems, temporal planning, probabilistic planning, preference-based planning, conditional planning, etc. Finally, robotics algorithms and functions 2216 may include any algorithms, functions, or systems that may enable one or more devices to replicate human behavior, for example, through movements, gestures, performance tasks, decision-making...

Claims

1. 1. A computer-implemented method for predicting amino acid-immune protein complex (IPC) interactions, comprising: accessing a set of amino acid sequences, each of the amino acid sequences of the set being identified from at least one protein; accessing an identified immune protein complex (IPC) sequence for an IPC of interest; using one or more first processing blocks of a processing subsystem of the machine learning model, processing the set of amino acid sequence representations to generate a set of transformed amino acid sequence representations based on a set of element concentration scores representing a binding core of the set of amino acid sequence representations, each of the amino acid sequence representations generated based on one of the amino acid sequences with a start of sequence (BOS) token appended; processing IPC sequence representations using a second processing block of the processing subsystem to generate transformed IPC sequence representations, the IPC sequence representations being generated based on the identified IPC sequences with BOS tokens appended, the set of amino acid sequence representations and the IPC sequence representations being processed in parallel; generating a composite representation by combining each of the converted BOS token representations of the set of converted amino acid sequence representations with the converted BOS token representation of the converted IPC sequence representation; and determining one or more predicted amino acid-IPC interactions based on said composite representation; 11. A computer-implemented method comprising:

2. 2. The computer-implemented method of claim 1, wherein the set of converted amino acid sequence representations comprises a set of amino-terminal flanking (N-flank) representations or a set of carboxy-terminal flanking (C-flank) representations.

3. 2. The computer-implemented method of claim 1, wherein the IPC of the subject is a major histocompatibility complex (MHC), including MHC class I (MHC-I) and / or MHC class II (MHC-II), or a T-cell receptor (TCR), and the at least one protein is a therapeutic protein or is present in a disease sample from the subject.

4. Generating a composite representation is for each of the set of converted amino acid sequence representations, element-wise multiplying a converted amino acid start of sequence (BOS) representation corresponding to the converted amino acid sequence representation by a converted IPC start of sequence (BOS) representation corresponding to the converted IPC sequence representation; The computer-implemented method of claim 1 , comprising:

5. 2. The computer-implemented method of claim 1, wherein processing the set of amino acid sequence representations comprises processing peptide start of sequence (BOS) representations to generate converted peptide sequence representations, and wherein processing the IPC sequence representations comprises processing MHC start of sequence representations (BOS) to generate converted MHC sequence representations.

6. processing the set of amino acid sequence representations, converting the set of amino acid sequence representations into the set of converted amino acid sequence representations using the one or more first processing blocks, each of the one or more first processing blocks comprising a set of processing sub-blocks; The computer-implemented method of claim 1 , comprising:

7. processing the IPC array representation, converting the IPC array representation into a converted IPC array representation using the second processing block, the second processing block comprising a set of processing sub-blocks; The computer-implemented method of claim 1 , comprising:

8. The computer-implemented method of claim 1 , wherein the machine learning model includes one or more transformer encoders, each of the one or more transformer encoders including a processing layer.

9. 2. The computer-implemented method of claim 1, wherein the set of amino acid sequence representations comprises an aggregate sequence representation, the aggregate sequence representation comprising a set of peptide representations and one or more of a set of amino-terminal flanking (N-flank) representations or a set of carboxy-terminal flanking (C-flank) representations.

10. prior to generating said set of converted amino acid sequence representations, Flattening the aggregate array representation into a single array, and densifying the aggregated sequence representation by removing empty rows from the array, wherein the converted amino acid sequence representation is generated based on the densified aggregated sequence representation; The computer-implemented method of claim 1 , further comprising:

11. processing the set of amino acid sequence representations, For each amino acid sequence representation of the set: determining, for each element of the amino acid sequence representation, a plurality of vectors based on a set of weights associated with a processing layer of the machine learning model; and generating the set of element concentration scores based on the plurality of vectors and the set of weights; The computer-implemented method of claim 1 , comprising:

12. the one or more first processing blocks and the second processing block include attention blocks, each attention block including a set of attention sub-blocks, each attention sub-block including a self-attention layer; the machine learning model is an attention-based machine learning model; 2. The computer-implemented method of claim 1, wherein the method further comprises generating, by one or more of the attention blocks, an attention map including one or more masks that limit attention applied by the attention sub-blocks to sequence lengths according to the masks.

13. The method comprises: calculating a plurality of average attention values ​​corresponding to the plurality of peptide positions by calculating an average attention value at each peptide position of the plurality of peptide positions based on a set of peptides having a uniform length distribution; and subtracting the plurality of average attention values ​​from a mask of the one or more masks; The computer-implemented method of claim 12 further comprising:

14. A dataset for training the machine learning model, generating a plurality of transformed peptide representations for a plurality of training peptides; for each training peptide, obtaining a corresponding cluster of training peptides based on said plurality of transformed peptide representations; For each training peptide, calculating information content based on the cluster of the corresponding training peptide; and removing one or more training peptides from the training data based on the corresponding information content of the one or more training peptides; The computer-implemented method of claim 1 , further comprising obtaining by:

15. processing the composite representation to generate a first output using a fully connected block of an output subsystem of the machine learning model; applying dropout to the first output to generate a second output using a dropout block of the output subsystem of the machine learning model; and selecting a subset of the second outputs using a maximum layer of the output subsystem of the machine learning model to generate a result; further comprising The computer-implemented method of claim 1 , wherein the one or more predicted amino acid-IPC interactions are determined based on the results.

16. the set of amino acid sequences comprises peptide sequences, and the IPC sequences comprise major histocompatibility complex (MHC) sequences; the one or more predicted amino acid-IPC interactions are Interaction affinity prediction for peptide-IPC combinations, which predicts the binding affinity between peptides and MHC; an interaction prediction for the peptide-IPC combination that predicts whether the MHC will present the peptide on a cell surface; or an immunogenicity prediction for said peptide-IPC combination, which predicts the ability of said peptide to elicit an immune response in relation to said MHC; The computer-implemented method of claim 1 , comprising one or more of:

17. Determining the one or more predicted amino acid-IPC interactions comprises: processing the composite representation to generate a set of results; and selecting an amino acid-IPC combination based on the highest result in said set of results; The computer-implemented method of claim 1 , comprising:

18. 10. The computer-implemented method of claim 1, wherein the one or more predicted amino acid-IPC interactions comprise a prediction of tumor-specific immunogenicity of a peptide.

19. 2. The computer-implemented method of claim 1, wherein the set of amino acid sequences comprises a set of peptide sequences, and the one or more predicted amino acid-IPC interactions identify a subset of peptide sequences that have increased tumor-specific immunogenicity or increased likelihood of presentation by the IPC compared to the set of peptide sequences.

20. 2. The computer-implemented method of claim 1, further comprising identifying a subset of peptides from the set of amino acid sequences for inclusion in a personalized vaccine, for inclusion as targets for immunotherapy, and / or for exclusion as targets for immunotherapy based on the determined one or more predicted amino acid-IPC interactions.

21. accessing a protein sequence corresponding to said at least one protein; obtaining a protein sequence embedding based on said protein sequence; and determining the one or more predicted amino acid-IPC interactions based at least in part on the protein sequence embedding; The computer-implemented method of claim 1 , further comprising:

22. The computer-implemented method of claim 1 , wherein the protein language model comprises a pre-trained protein language model.

23. Reducing the dimensionality of protein sequence embeddings; and combining each of the converted BOS token representations of the set of converted amino acid sequence representations with the converted IPC sequence representation and the converted BOS token representation of a reduced-dimensionality protein sequence embedding; 23. The computer-implemented method of claim 22, further comprising:

24. 24. The computer-implemented method of claim 23, wherein the dimensionality of the protein sequence embedding is reduced via a neural network.

25. A system for predicting amino acid-immune protein complex (IPC) interactions, comprising: one or more non-transitory computer-readable storage media containing instructions; one or more processors coupled to the one or more storage media; Equipped with The one or more processors: accessing a set of amino acid sequences, each of the amino acid sequences in the set being identified from at least one protein; accessing an identified immune protein complex (IPC) sequence for an IPC of interest; using one or more first processing blocks of a processing subsystem of the machine learning model, processing the set of amino acid sequence representations to generate a set of transformed amino acid sequence representations based on a set of element concentration scores representing a binding core of the set of amino acid sequence representations, each of the amino acid sequence representations generated based on one of the amino acid sequences with a start of sequence (BOS) token appended; processing IPC sequence representations using a second processing block of the processing subsystem to generate transformed IPC sequence representations, the IPC sequence representations being generated based on the identified IPC sequences with BOS tokens appended, the set of amino acid sequence representations and the IPC sequence representations being processed in parallel; generating a composite representation by combining each of the converted BOS token representations of the set of converted amino acid sequence representations with the converted BOS token representation of the converted IPC sequence representation; and determining one or more predicted amino acid-IPC interactions based on said composite representation; A system configured to execute the instructions for:

26. 1. A non-transitory computer-readable medium for a system for predicting amino acid-immune protein complex (IPC) interactions, the system being configured, when executed by one or more processors of one or more computing devices, to cause the one or more processors to: accessing a set of amino acid sequences, each of the amino acid sequences in the set being identified from at least one protein; accessing an identified immune protein complex (IPC) sequence for an IPC of interest; using one or more first processing blocks of a processing subsystem of the machine learning model, processing the set of amino acid sequence representations to generate a set of transformed amino acid sequence representations based on a set of element concentration scores representing a binding core of the set of amino acid sequence representations, each of the amino acid sequence representations generated based on one of the amino acid sequences with a start of sequence (BOS) token appended; processing IPC sequence representations using a second processing block of the processing subsystem to generate transformed IPC sequence representations, the IPC sequence representations being generated based on the identified IPC sequences with BOS tokens appended, the set of amino acid sequence representations and the IPC sequence representations being processed in parallel; generating a composite representation by combining each of the converted BOS token representations of the set of converted amino acid sequence representations with the converted BOS token representation of the converted IPC sequence representation; and determining one or more predicted amino acid-IPC interactions based on said composite representation; A non-transitory computer-readable medium containing instructions to cause

27. one or more peptides, a plurality of nucleic acids encoding said one or more peptides; or a plurality of cells expressing said one or more peptides; 1. A vaccine comprising: The one or more peptides are selected from a set of peptides accessing a set of amino acid sequences, each of the amino acid sequences of the set being identified from at least one protein; accessing an identified immune protein complex (IPC) sequence for an IPC of interest; using one or more first processing blocks of a processing subsystem of the machine learning model, processing the set of amino acid sequence representations to generate a set of transformed amino acid sequence representations based on a set of element concentration scores representing a binding core of the set of amino acid sequence representations, each of the amino acid sequence representations generated based on one of the amino acid sequences with a start of sequence (BOS) token appended; processing IPC sequence representations using a second processing block of the processing subsystem to generate transformed IPC sequence representations, the IPC sequence representations being generated based on the identified IPC sequences with BOS tokens appended, the set of amino acid sequence representations and the IPC sequence representations being processed in parallel; generating a composite representation by combining each of the converted BOS token representations of the set of converted amino acid sequence representations with the converted BOS token representation of the converted IPC sequence representation; and determining one or more predicted amino acid-IPC interactions based on said composite representation; A vaccine selected by

28. 1. A method of producing a vaccine, comprising: one or more peptides, a plurality of nucleic acids encoding said one or more peptides; or a plurality of cells expressing said one or more peptides; and producing a vaccine comprising: The one or more peptides are selected from a set of peptides accessing a set of amino acid sequences, each of the amino acid sequences of the set being identified from at least one protein; accessing an identified immune protein complex (IPC) sequence for an IPC of interest; using one or more first processing blocks of a processing subsystem of the machine learning model, processing the set of amino acid sequence representations to generate a set of transformed amino acid sequence representations based on a set of element concentration scores representing a binding core of the set of amino acid sequence representations, each of the amino acid sequence representations generated based on one of the amino acid sequences with a start of sequence (BOS) token appended; processing IPC sequence representations using a second processing block of the processing subsystem to generate transformed IPC sequence representations, the IPC sequence representations being generated based on the identified IPC sequences with BOS tokens appended, the set of amino acid sequence representations and the IPC sequence representations being processed in parallel; generating a composite representation by combining each of the converted BOS token representations of the set of converted amino acid sequence representations with the converted BOS token representation of the converted IPC sequence representation; and determining one or more predicted amino acid-IPC interactions based on said composite representation; producing a vaccine selected by A method comprising:

29. A pharmaceutical composition comprising, from a set of peptides: accessing a set of amino acid sequences, each of the amino acid sequences of the set being identified from at least one protein; accessing an identified immune protein complex (IPC) sequence for an IPC of interest; using one or more first processing blocks of a processing subsystem of the machine learning model, processing the set of amino acid sequence representations to generate a set of transformed amino acid sequence representations based on a set of element concentration scores representing a binding core of the set of amino acid sequence representations, each of the amino acid sequence representations generated based on one of the amino acid sequences with a start of sequence (BOS) token appended; processing IPC sequence representations using a second processing block of the processing subsystem to generate transformed IPC sequence representations, the IPC sequence representations being generated based on the identified IPC sequences with BOS tokens appended, the set of amino acid sequence representations and the IPC sequence representations being processed in parallel; generating a composite representation by combining each of the converted BOS token representations of the set of converted amino acid sequence representations with the converted BOS token representation of the converted IPC sequence representation; and determining one or more predicted amino acid-IPC interactions based on said composite representation; A pharmaceutical composition comprising one or more peptides selected by

30. 1. A computer-implemented method for predicting amino acid-immune protein complex (IPC) interactions, comprising: accessing a set of amino acid sequences, each of the amino acid sequences of the set being identified from at least one protein; accessing an identified immune protein complex (IPC) sequence for an IPC of interest; using one or more first processing blocks of a processing subsystem of the machine learning model, processing the set of amino acid sequence representations to generate a set of transformed amino acid sequence representations based on a set of element concentration scores representing a binding core of the set of amino acid sequence representations, each of the amino acid sequence representations generated based on one of the amino acid sequences with a start of sequence (BOS) token appended; processing the IPC array to generate an IPC array embedding using a second processing block of the processing subsystem; generating a composite representation by aggregating each of the converted BOS token representations of the set of converted amino acid sequence representations with the IPC sequence embedding; and determining one or more predicted amino acid-IPC interactions based on said composite representation; 11. A computer-implemented method comprising:

31. 31. The computer-implemented method of claim 30, wherein the IPC of the subject is a major histocompatibility complex (MHC), including MHC class I (MHC-I) and / or MHC class II (MHC-II), or a T-cell receptor (TCR), and the at least one protein is a therapeutic protein or is present in a disease sample from the subject.

32. 31. The computer-implemented method of claim 30, wherein processing the set of amino acid sequence representations comprises processing start of peptide sequence (BOS) representations to generate transformed peptide sequence representations.

33. processing the set of amino acid sequence representations, converting the set of amino acid sequence representations into the set of converted amino acid sequence representations using the one or more first processing blocks, each of the one or more first processing blocks comprising a set of processing sub-blocks; 31. The computer-implemented method of claim 30, comprising:

34. 31. The computer-implemented method of claim 30, wherein the machine learning model includes one or more transformer encoders, each of the one or more transformer encoders including a processing layer.

35. 31. The computer-implemented method of claim 30, wherein the set of amino acid sequence representations comprises an aggregate sequence representation, the aggregate sequence representation comprising a set of peptide representations and one or more of a set of amino-terminal flanking (N-flank) representations or a set of carboxy-terminal flanking (C-flank) representations.

36. prior to generating said set of converted amino acid sequence representations, Flattening the aggregate array representation into a single array, and densifying the aggregated sequence representation by removing empty rows from the array, wherein the converted amino acid sequence representation is generated based on the densified aggregated sequence representation; 31. The computer-implemented method of claim 30, further comprising:

37. processing the set of amino acid sequence representations, For each amino acid sequence representation of the set: determining, for each element of the amino acid sequence representation, a plurality of vectors based on a set of weights associated with a processing layer of the machine learning model; and generating the set of element concentration scores based on the plurality of vectors and the set of weights; 31. The computer-implemented method of claim 30, comprising:

38. the one or more first processing blocks and the second processing block include attention blocks, each attention block including a set of attention sub-blocks, each attention sub-block including a self-attention layer; the machine learning model is an attention-based machine learning model; 31. The computer-implemented method of claim 30, wherein the method further comprises generating, by one or more of the attention blocks, an attention map including one or more masks that limit attention applied by the attention sub-blocks to sequence lengths according to the masks.

39. The method comprises: calculating a plurality of average attention values ​​corresponding to the plurality of peptide positions by calculating an average attention value at each peptide position of the plurality of peptide positions based on a set of peptides having a uniform length distribution; and subtracting the plurality of average attention values ​​from a mask of the one or more masks; 39. The computer-implemented method of claim 38, further comprising:

40. A dataset for training the machine learning model, generating a plurality of transformed peptide representations for a plurality of training peptides; for each training peptide, obtaining a corresponding cluster of training peptides based on said plurality of transformed peptide representations; For each training peptide, calculating information content based on the cluster of the corresponding training peptide; and removing one or more training peptides from the training data based on the corresponding information content of the one or more training peptides; 31. The computer-implemented method of claim 30, further comprising obtaining by:

41. the set of amino acid sequences comprises peptide sequences, and the IPC sequences comprise major histocompatibility complex (MHC) sequences; the one or more predicted amino acid-IPC interactions are Interaction affinity prediction for peptide-IPC combinations, which predicts the binding affinity between peptides and MHC; an interaction prediction for the peptide-IPC combination that predicts whether the MHC will present the peptide on a cell surface; or an immunogenicity prediction for said peptide-IPC combination, which predicts the ability of said peptide to elicit an immune response in relation to said MHC; 31. The computer-implemented method of claim 30, comprising one or more of:

42. Determining the one or more predicted amino acid-IPC interactions comprises: processing the composite representation to generate a set of results; and selecting an amino acid-IPC combination based on the highest result in said set of results; 31. The computer-implemented method of claim 30, comprising:

43. 31. The computer-implemented method of claim 30, wherein the one or more predicted amino acid-IPC interactions comprise a prediction of tumor-specific immunogenicity of a peptide.

44. 31. The computer-implemented method of claim 30, wherein the set of amino acid sequences comprises a set of peptide sequences, and the one or more predicted amino acid-IPC interactions identify a subset of peptide sequences that have increased tumor-specific immunogenicity or increased likelihood of presentation by the IPC compared to the set of peptide sequences.

45. 31. The computer-implemented method of claim 30, further comprising identifying a subset of peptides from the set of amino acid sequences for inclusion in a personalized vaccine, for inclusion as targets for immunotherapy, and / or for exclusion as targets for immunotherapy based on the determined one or more predicted amino acid-IPC interactions.

46. accessing a protein sequence corresponding to said at least one protein; obtaining a protein sequence embedding based on said protein sequence; and determining the one or more predicted amino acid-IPC interactions based at least in part on the protein sequence embedding; 31. The computer-implemented method of claim 30, further comprising:

47. 31. The computer-implemented method of claim 30, wherein the protein language model comprises a pre-trained protein language model.

48. Reducing the dimensionality of the protein sequence embedding; and combining each of the converted BOS token representations of the set of converted amino acid sequence representations with a converted IPC sequence representation and with the converted BOS token representation of a reduced-dimensionality protein sequence embedding; 47. The computer-implemented method of claim 46, further comprising:

49. 49. The computer-implemented method of claim 48, wherein the dimensionality of the protein sequence embedding is reduced via a neural network.

50. 31. The computer-implemented method of claim 30, wherein generating the IPC sequence embedding comprises inputting the IPC sequence into a protein language model.

51. 31. The computer-implemented method of claim 30, further comprising reducing the dimensionality of the IPC array embedding.

52. 52. The computer-implemented method of claim 51, wherein the dimensionality of the IPC array embedding is reduced by principal component analysis (PCA).

53. 52. The computer-implemented method of claim 51 , wherein generating the composite representation comprises, for each of the set of transformed amino acid sequence representations, element-wise multiplying a transformed starting amino acid sequence (BOS) representation corresponding to the transformed amino acid sequence representation by a reduced-dimensionality IPC sequence embedding.

54. A system for predicting amino acid-immune protein complex (IPC) interactions, comprising: one or more non-transitory computer-readable storage media containing instructions; one or more processors coupled to the one or more storage media; Including, accessing, by the one or more processors, a set of amino acid sequences, each of the amino acid sequences in the set being identified from at least one protein; accessing a set of amino acid sequences, each of the amino acid sequences in the set being identified from at least one protein; accessing an identified immune protein complex (IPC) sequence for an IPC of interest; using one or more first processing blocks of a processing subsystem of the machine learning model, processing the set of amino acid sequence representations to generate a set of transformed amino acid sequence representations based on a set of element concentration scores representing a binding core of the set of amino acid sequence representations, each of the amino acid sequence representations generated based on one of the amino acid sequences with a start of sequence (BOS) token appended; processing the IPC array to generate an IPC array embedding using a second processing block of the processing subsystem; generating a composite representation by aggregating each of the converted BOS token representations of the set of converted amino acid sequence representations with the IPC sequence embedding; and determining one or more predicted amino acid-IPC interactions based on said composite representation; A system configured to execute the instructions for:

55. 1. A non-transitory computer-readable medium for a system for predicting amino acid-immune protein complex (IPC) interactions, the system being configured, when executed by one or more processors of one or more computing devices, to cause the one or more processors to: accessing a set of amino acid sequences, each of the amino acid sequences in the set being identified from at least one protein; accessing an identified immune protein complex (IPC) sequence for an IPC of interest; using one or more first processing blocks of a processing subsystem of the machine learning model, processing the set of amino acid sequence representations to generate a set of transformed amino acid sequence representations based on a set of element concentration scores representing a binding core of the set of amino acid sequence representations, each of the amino acid sequence representations generated based on one of the amino acid sequences with a start of sequence (BOS) token appended; processing the IPC array to generate an IPC array embedding using a second processing block of the processing subsystem; generating a composite representation by aggregating each of the converted BOS token representations of the set of converted amino acid sequence representations with the IPC sequence embedding; and determining one or more predicted amino acid-IPC interactions based on said composite representation; A non-transitory computer-readable medium containing instructions to cause

56. one or more peptides, a plurality of nucleic acids encoding said one or more peptides; or a plurality of cells expressing said one or more peptides 1. A vaccine comprising: The one or more peptides are selected from a set of peptides accessing a set of amino acid sequences, each of the amino acid sequences of the set being identified from at least one protein; accessing an identified immune protein complex (IPC) sequence for an IPC of interest; using one or more first processing blocks of a processing subsystem of the machine learning model, processing the set of amino acid sequence representations to generate a set of transformed amino acid sequence representations based on a set of element concentration scores representing a binding core of the set of amino acid sequence representations, each of the amino acid sequence representations generated based on one of the amino acid sequences with a start of sequence (BOS) token appended; processing the IPC array to generate an IPC array embedding using a second processing block of the processing subsystem; generating a composite representation by aggregating each of the converted BOS token representations of the set of converted amino acid sequence representations with the IPC sequence embedding; and determining one or more predicted amino acid-IPC interactions based on said composite representation; A vaccine selected by

57. 1. A method of producing a vaccine, comprising: one or more peptides, a plurality of nucleic acids encoding said one or more peptides; or a plurality of cells expressing said one or more peptides and producing a vaccine comprising: The one or more peptides are selected from a set of peptides accessing a set of amino acid sequences, each of the amino acid sequences of the set being identified from at least one protein; accessing an identified immune protein complex (IPC) sequence for an IPC of interest; using one or more first processing blocks of a processing subsystem of the machine learning model, processing the set of amino acid sequence representations to generate a set of transformed amino acid sequence representations based on a set of element concentration scores representing a binding core of the set of amino acid sequence representations, each of the amino acid sequence representations generated based on one of the amino acid sequences with a start of sequence (BOS) token appended; processing the IPC array to generate an IPC array embedding using a second processing block of the processing subsystem; generating a composite representation by aggregating each of the converted BOS token representations of the set of converted amino acid sequence representations with the IPC sequence embedding; and determining one or more predicted amino acid-IPC interactions based on said composite representation; producing a vaccine selected by A method comprising:

58. A pharmaceutical composition comprising, from a set of peptides: accessing a set of amino acid sequences, each of the amino acid sequences of the set being identified from at least one protein; accessing an identified immune protein complex (IPC) sequence for an IPC of interest; using one or more first processing blocks of a processing subsystem of the machine learning model, processing the set of amino acid sequence representations to generate a set of transformed amino acid sequence representations based on a set of element concentration scores representing a binding core of the set of amino acid sequence representations, each of the amino acid sequence representations generated based on one of the amino acid sequences with a start of sequence (BOS) token appended; processing the IPC array to generate an IPC array embedding using a second processing block of the processing subsystem; generating a composite representation by aggregating each of the converted BOS token representations of the set of converted amino acid sequence representations with the IPC sequence embedding; and determining one or more predicted amino acid-IPC interactions based on said composite representation; A pharmaceutical composition comprising one or more peptides selected by

59. 1. A computer-implemented method for predicting amino acid-immune protein complex (IPC) interactions, comprising: accessing a set of amino acid sequences, each of the amino acid sequences of the set being identified from at least one protein; accessing an identified immune protein complex (IPC) sequence for an IPC of interest; generating an IPC array embedding based on the IPC array; processing the set of amino acid sequence representations using one or more processing blocks of a processing subsystem of the machine learning model to generate a set of transformed amino acid sequence representations based on a set of element concentration scores representing a binding core of the set of amino acid sequence representations; generating a set of composite representations based on the set of transformed amino acid sequence representations and the IPC sequence embeddings using a cross-attention machine learning module of the processing subsystem; and determining one or more predicted amino acid-IPC interactions based on said composite representation; 11. A computer-implemented method comprising:

60. 60. The computer-implemented method of claim 59, wherein generating the IPC sequence embedding comprises inputting the IPC sequence into a protein language model.

61. 60. The computer-implemented method of claim 59, further comprising reducing the dimensionality of the IPC array embedding.

62. 62. The computer-implemented method of claim 61, wherein the dimensionality of the IPC array embedding is reduced by principal component analysis (PCA).

63. 62. The computer-implemented method of claim 61, wherein the cross-attention module comprises a self-attention transformer having three components: a query (Q) component, a key (K) component, and a value (V) component.

64. each of the K components and the V components corresponds to a set of the transformed amino acid sequence representations; 64. The computer-implemented method of claim 63, wherein the Q components correspond to an aggregation of a start-of-sequence (BOS) vector embedding and a reduced-dimensionality IPC array embedding.

65. each of the K components and the V components corresponds to a set of the transformed amino acid sequence representations; 64. The computer-implemented method of claim 63, wherein the Q components correspond to reduced-dimensionality IPC array embeddings.

66. the set of amino acid sequences comprises at least one peptide sequence having multiple binding cores capable of binding to multiple alleles of the IPC; 60. The computer-implemented method of claim 59, wherein the one or more predicted amino acid-IPC interactions comprise at least a plurality of allele-specific and binding core-specific predicted amino acid-IPC interactions.

67. 60. The computer-implemented method of claim 59, wherein the IPC of the subject is a major histocompatibility complex (MHC), including MHC class I (MHC-I) and / or MHC class II (MHC-II), or a T-cell receptor (TCR), and the at least one protein is a therapeutic protein or is present in a disease sample from the subject.

68. 60. The computer-implemented method of claim 59, wherein processing the set of amino acid sequence representations comprises converting the set of amino acid sequence representations into the set of converted amino acid sequence representations using the one or more first processing blocks, each of the one or more first processing blocks comprising a set of processing sub-blocks.

69. 60. The computer-implemented method of claim 59, wherein the set of amino acid sequence representations comprises an aggregate sequence representation, the aggregate sequence representation comprising a set of peptide representations and one or more of a set of amino-terminal flanking (N-flank) representations or a set of carboxy-terminal flanking (C-flank) representations.

70. processing the set of amino acid sequence representations, For each amino acid sequence representation of the set: determining, for each element of the amino acid sequence representation, a plurality of vectors based on a set of weights associated with a processing layer of the machine learning model; and generating the set of element concentration scores based on the plurality of vectors and the set of weights; 60. The computer-implemented method of claim 59, comprising:

71. The one or more processing blocks include attention blocks, each attention block including a set of attention sub-blocks, each attention sub-block including a self-attention layer; the machine learning model is an attention-based machine learning model; 60. The computer-implemented method of claim 59, wherein the method further comprises generating, by one or more of the attention blocks, an attention map including one or more masks that limit attention applied by the attention sub-blocks to sequence lengths according to the masks.

72. The method comprises: calculating a plurality of average attention values ​​corresponding to the plurality of peptide positions by calculating an average attention value at each peptide position of the plurality of peptide positions based on a set of peptides having a uniform length distribution; and subtracting the plurality of average attention values ​​from a mask of the one or more masks; 72. The computer-implemented method of claim 71, further comprising:

73. A dataset for training the machine learning model, generating a plurality of transformed peptide representations for a plurality of training peptides; for each training peptide, obtaining a corresponding cluster of training peptides based on said plurality of transformed peptide representations; For each training peptide, calculating information content based on the cluster of the corresponding training peptide; and removing one or more training peptides from the training data based on the corresponding information content of the one or more training peptides; 60. The computer-implemented method of claim 59, further comprising obtaining by:

74. the set of amino acid sequences comprises peptide sequences, and the IPC sequences comprise major histocompatibility complex (MHC) sequences; the one or more predicted amino acid-IPC interactions are Interaction affinity prediction of peptide-IPC combinations, which predicts the binding affinity between peptides and MHC; an interaction prediction for the peptide-IPC combination that predicts whether the MHC will present the peptide on a cell surface; or an immunogenicity prediction for said peptide-IPC combination, which predicts the ability of said peptide to elicit an immune response in relation to said MHC; 60. The computer-implemented method of claim 59, comprising one or more of:

75. Determining the one or more predicted amino acid-IPC interactions comprises: processing the composite representation to generate a set of results; and selecting an amino acid-IPC combination based on the highest result in said set of results; 60. The computer-implemented method of claim 59, comprising:

76. 60. The computer-implemented method of claim 59, wherein the one or more predicted amino acid-IPC interactions comprise a prediction of tumor-specific immunogenicity of a peptide.

77. 60. The computer-implemented method of claim 59, wherein the set of amino acid sequences comprises a set of peptide sequences, and wherein the one or more predicted amino acid-IPC interactions identify a subset of peptide sequences that have increased tumor-specific immunogenicity or increased likelihood of presentation by the IPC compared to the set of peptide sequences.

78. 60. The computer-implemented method of claim 59, further comprising identifying a subset of peptides from the set of amino acid sequences for inclusion in a personalized vaccine, for inclusion as targets for immunotherapy, and / or for exclusion as targets for immunotherapy based on the determined one or more predicted amino acid-IPC interactions.

79. accessing a protein sequence corresponding to said at least one protein; obtaining a protein sequence embedding based on the protein sequence and the protein model; and determining the one or more predicted amino acid-IPC interactions based at least in part on the protein sequence embedding; 60. The computer-implemented method of claim 59, further comprising:

80. 80. The computer-implemented method of claim 79, wherein the protein language model comprises a pre-trained protein language model.

81. 81. The computer-implemented method of claim 80, wherein the composite representation is generated based at least in part on the protein sequence embedding.

82. A system for predicting amino acid-immune protein complex (IPC) interactions, comprising: one or more non-transitory computer-readable storage media containing instructions; one or more processors coupled to the one or more storage media; Equipped with The one or more processors: accessing a set of amino acid sequences, each of the amino acid sequences in the set being identified from at least one protein; accessing an identified immune protein complex (IPC) sequence for an IPC of interest; generating an IPC array embedding based on the IPC array; processing the set of amino acid sequence representations using one or more processing blocks of a processing subsystem of the machine learning model to generate a set of transformed amino acid sequence representations based on a set of element concentration scores representing a binding core of the set of amino acid sequence representations; generating a set of composite representations based on the set of transformed amino acid sequence representations and the IPC sequence embeddings using a cross-attention machine learning module of the processing subsystem; and determining one or more predicted amino acid-IPC interactions based on said composite representation; A system configured to execute the instructions for:

83. 1. A non-transitory computer-readable medium for a system for predicting amino acid-immune protein complex (IPC) interactions, the system being configured, when executed by one or more processors of one or more computing devices, to cause the one or more processors to: accessing a set of amino acid sequences, each of the amino acid sequences in the set being identified from at least one protein; accessing an identified immune protein complex (IPC) sequence for an IPC of interest; generating an IPC array embedding based on the IPC array; processing the set of amino acid sequence representations using one or more processing blocks of a processing subsystem of the machine learning model to generate a set of transformed amino acid sequence representations based on a set of element concentration scores representing a binding core of the set of amino acid sequence representations; generating a set of composite representations based on the set of transformed amino acid sequence representations and the IPC sequence embeddings using a cross-attention machine learning module of the processing subsystem; and determining one or more predicted amino acid-IPC interactions based on said composite representation; A non-transitory computer-readable medium containing instructions to cause

84. one or more peptides, a plurality of nucleic acids encoding said one or more peptides; or a plurality of cells expressing said one or more peptides 1. A vaccine comprising: The one or more peptides are selected from a set of peptides accessing a set of amino acid sequences, each of the amino acid sequences of the set being identified from at least one protein; accessing an identified immune protein complex (IPC) sequence for an IPC of interest; generating an IPC array embedding based on the IPC array; processing the set of amino acid sequence representations using one or more processing blocks of a processing subsystem of the machine learning model to generate a set of transformed amino acid sequence representations based on a set of element concentration scores representing a binding core of the set of amino acid sequence representations; generating a set of composite representations based on the set of transformed amino acid sequence representations and the IPC sequence embeddings using a cross-attention machine learning module of the processing subsystem; and determining one or more predicted amino acid-IPC interactions based on said composite representation; A vaccine selected by

85. 1. A method of producing a vaccine, comprising: one or more peptides, a plurality of nucleic acids encoding said one or more peptides; or a plurality of cells expressing said one or more peptides; and producing a vaccine comprising: The one or more peptides are selected from a set of peptides accessing a set of amino acid sequences, each of the amino acid sequences of the set being identified from at least one protein; accessing an identified immune protein complex (IPC) sequence for an IPC of interest; generating an IPC array embedding based on the IPC array; processing the set of amino acid sequence representations using one or more processing blocks of a processing subsystem of the machine learning model to generate a transformed set of amino acid sequence representations based on a set of element concentration scores representing a binding core of the set of amino acid sequence representations; generating a set of composite representations based on the set of transformed amino acid sequence representations and the IPC sequence embeddings using a cross-attention machine learning module of the processing subsystem; and determining one or more predicted amino acid-IPC interactions based on said composite representation; producing a vaccine selected by A method comprising:

86. A pharmaceutical composition comprising, from a set of peptides: accessing a set of amino acid sequences, each of the amino acid sequences of the set being identified from at least one protein; accessing an identified immune protein complex (IPC) sequence for an IPC of interest; generating an IPC array embedding based on the IPC array; processing the set of amino acid sequence representations using one or more processing blocks of a processing subsystem of the machine learning model to generate a set of transformed amino acid sequence representations based on a set of element concentration scores representing a binding core of the set of amino acid sequence representations; generating a set of composite representations based on the set of transformed amino acid sequence representations and the IPC sequence embeddings using a cross-attention machine learning module of the processing subsystem; and determining one or more predicted amino acid-IPC interactions based on said composite representation; A pharmaceutical composition comprising one or more peptides selected by