Methods and systems for predicting peptide presentation through major histocompatibility complex molecules
The processing of amino acid and immune protein complex sequences through machine learning models is achieved to predict the interaction and immunogenicity of peptides and MHC molecules, solving the problem of inaccurate identification of tumor neoantigen epitopes and therapeutic protein immune responses in the prior art, and achieving the effective design of individualized vaccines and the safety of therapeutic proteins.
Patent Information
- Application Number
- CN202380082616.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-05
- Filing Date
- 2023-12-04
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art is difficult to accurately identify neoantigen epitopes presented on the tumor surface, resulting in the inability of peptide fragments in individualized cancer vaccines to effectively trigger a robust immune response, and therapeutic proteins may trigger an immune response, affecting the efficacy.
The amino acid and immune protein complex sequences were treated using machine learning models, and through the focus score and synthesis representation of the core element, the affinity and immunogenicity of the interaction between the peptide and MHC molecules were predicted, and the appropriate peptide fragments were selected for individualized vaccines.
Improve the accuracy of predicting tumor-specific immunogenicity, reduce false positives, select appropriate peptide fragments for vaccines, avoid immunogenic risks, and ensure that therapeutic proteins do not trigger immune responses.
Smart Images

Figure CN120345029A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 430,297, filed on December 5, 2022, entitled "PREDICTION OF PEPTIDE PRESENTATION BY MAJOR HISTOCOMPATIBILITY COMPLEX MOLECULES", the entire disclosure of which is incorporated herein by reference. Technical Field
[0003] The present disclosure generally relates to immunology, and in particular, methods for predicting whether a neoantigen or a therapeutic protein is likely to trigger an immunogenic response. Background Art
[0004] Therapeutic proteins are a class of pharmaceuticals (biologics) obtained from living sources (e.g., animals, plants, fungi, or microbial cells). Many therapeutic proteins, including monoclonal antibodies and soluble receptors, are now produced using recombinant DNA technology. This contrasts with small - molecule drugs, which are typically simpler compounds manufactured by chemical synthesis.
[0005] Some therapeutic proteins have the same primary amino acid sequence as natural human proteins and generally do not trigger an immune response. However, it is often desirable to modify the amino acid sequence of a therapeutic protein to optimize various properties such as efficacy, stability, and bioavailability. However, a therapeutic protein with a significant amino acid sequence difference compared to the proteins of the intended recipient (e.g., a human subject) may be recognized as foreign by the recipient's immune system, thus triggering an immune response, as an antigen (toxin or foreign substance) might. In this way, a seemingly foreign therapeutic protein is regarded more like a vaccine rather than a medicinal compound. While a vaccine must trigger an immune response to be effective, an immune response triggered by a therapeutic protein is harmful. If a therapeutic protein triggers an immune response, it will render the therapeutic protein ineffective and / or will cause a harmful immune response in the recipient. In particular, repeated administration of an immunogenic therapeutic protein often triggers anti - drug antibodies (ADA). In the best case, the immunogenic therapeutic protein will be neutralized by the recipient's ADA, thus reducing its drug activity. In the worst case, the immunogenic therapeutic protein will induce a hypersensitivity reaction in the recipient, which can be life - threatening.
[0006] Human leukocyte antigen (HLA) is expressed as cell surface receptors that present antigenic peptides to T cells in a restricted manner, which allows for the discrimination of self-antigens from foreign antigens. The HLA complex is a gene complex on chromosome 6 that encodes cell surface proteins responsible for regulating the immune system. The human major histocompatibility complex (MHC) is the locus of the HLA genes and plays a fundamental role in the acceptance of transplanted tissues. The MHC contains many genes associated with cell-mediated immune defense. The MHC complex encodes the α chains of MHC class I molecules HLA-A, HLA-B, and HLA-C (alleles) and the α and β chains of MHC class II molecules HLA-DR, HLA-DP, and HLA-DQ (allotypes), all of which are expressed in a co-dominant manner.
[0007] Helper T (Th) cells are activated when they recognize peptide fragments (epitopes) of protein antigens that are bound to MHC class II (MHC-II) molecules on antigen-presenting cells, leading to the formation of antigen-reactive antibodies. If the antigen is a therapeutic protein, the antibodies will be ADA. Peptides that bind to MHC molecules with sufficient affinity are a prerequisite for immunogenicity (the ability of a peptide to trigger an immune response). Therefore, it is desirable to predict MHC-II epitopes of candidate therapeutic proteins to identify amino acid residues of the protein that can be safely altered to eliminate the MHC-II epitopes of the candidate therapeutic protein in order to reduce its immunogenicity. Peptide-MHC binding affinity is mainly determined by the amino acid sequence of the peptide-binding core (usually nine amino acids in length); however, the amino-terminal flanking (N-flanking) sequence and / or carboxyl-terminal flanking (C-flanking) sequence may also affect peptide-MHC binding.
[0008] Tumors, like the subjects they affect, are heterogeneous. Specifically, even in tumors derived from the same cell type, the somatic mutations that cause the cells to become cancerous may vary. In addition, although humans are expected to share 99.9% of their genomes, the 0.1% difference is significant, especially with respect to the immune system. Therefore, therapeutic cancer vaccines are ideally designed as personalized cancer vaccines.
[0009] Neoantigens are tumor-specific antigens generated by somatic mutations in the tumor cell genome. Peptide fragments (epitopes) of neoantigen proteins bind to major histocompatibility complex (MHC) molecules expressed on the surface of the subject's cancer cells and antigen-presenting cells, where they can activate CD8+ cytotoxic T lymphocytes (CTLs) and CD4+ helper T (Th) cells, respectively. Neoantigen vaccines are a promising personalized cancer therapy because they can trigger the subject's T cells to recognize and attack cancer cells expressing neoantigens without harming healthy cells.
[0010] The tumor characteristics of a subject can be defined by determining the DNA and / or RNA sequences of tumor cells from a biopsy. From the subject-specific tumor characteristics, novel antigens of interest that are present in tumor cells but not in healthy cells can be identified. However, the vast majority of mutated sequences detected in tumor cells correspond to poorly expressed novel antigens that do not contain epitopes presented by MHC molecules and / or are not otherwise bound by T cell receptors (TCRs). In the case of MHC class I (MHC-I) molecules, such novel antigens will not trigger an immune response by CD8+ CTLs, and in the case of MHC class II (MHC-II) molecules, such novel antigens will not trigger an immune response by CD4+ Th cells. Thus, such novel antigens are poor candidates for inclusion in personalized cancer vaccines for generating tumor-specific immune responses.
[0011] There are tools for predicting the binding of peptides to MHC molecules. However, simply identifying peptide fragments of novel antigens that can bind to MHC molecules is not sufficient to identify novel antigens for inclusion in personalized cancer vaccines. This is because many peptide binders will be false positives as they will not effectively elicit a cellular immune response. Thus, there is a need in the art for tools to accurately identify epitopes of novel antigens presented on the surface of tumors such that they elicit a robust immune response to aid in the selection of peptide fragments of novel antigens for inclusion in therapeutic cancer vaccines. Summary of the Invention
[0012] Systems, methods, and programming for determining predictions of whether a peptide interacts with an MHC molecule and / or the extent thereof using a machine learning model are disclosed herein. The machine learning model can perform a combination of peptide data processing, n-flank / c-flank data processing, MHC data processing, TCR data processing, and / or protein data processing to obtain one or more interaction predictions (e.g., prediction of interaction of a peptide with an MHC molecule), one or more interaction affinity predictions (e.g., binding affinity between a predicted peptide and an MHC), and / or one or more immunogenicity predictions (prediction of the ability of a peptide to elicit an immune response with respect to an MHC) as described herein. In one exemplary workflow, the processing of MHC data can involve processing of an MHC sequence appended with a BOS token using one or more transformer stages, or alternatively can involve processing of an MHC sequence embedding generated using a protein language model. In this exemplary workflow, the processing of peptide data can involve processing of a peptide sequence appended with a BOS token using one or more transformer stages, or alternatively can involve processing of a peptide sequence not appended with a BOS token using a cross-attention module. This exemplary workflow can additionally incorporate processing of protein data, or alternatively can not incorporate processing of protein data. This exemplary workflow can additionally incorporate processing of TCR data, or alternatively can not incorporate TCR data.
[0013] In some embodiments, the machine learning model uses separate processing blocks and processes a set of amino acid sequence representations and immunoprotein complex (IPC) sequence representations in parallel. The sequence representations may be embeddings of features of the corresponding sequences. The machine learning model uses a set of element focus scores representing the binding core of the set of amino acid sequence representations, and combines the BOS token embedding of the transformed amino acid sequence representation with the BOS token embedding of the transformed IPC sequence representation to generate a synthetic representation. These synthetic representations are used to determine one or more predicted amino acid-IPC interactions, such as interaction affinity predictions for predicting the binding affinity between a peptide and an MHC, interaction predictions for predicting whether the MHC will present a peptide at the cell surface, or immunogenicity predictions for predicting the ability of a peptide to elicit an immune response against the MHC.
[0014] Some aspects include accessing a set of amino acid sequences. Each of the amino acid sequences in the set may have been identified from at least one protein. An immunoprotein complex (IPC) sequence identified for the IPC of a subject may be accessed. One or more first processing blocks in the processing subsystem of the machine learning model may be used to process a set of amino acid sequence representations to generate a set of transformed amino acid sequence representations based on a set of element focus scores representing the binding core of the set of amino acid sequence representations. Each of the amino acid sequence representations may have been generated based on one of the amino acid sequences appended with a beginning-of-sequence (BOS) token. A second processing block in the processing subsystem may be used to process the IPC sequence representation to generate a transformed IPC sequence representation. The IPC sequence representation may be generated based on the identified IPC sequence appended with a BOS token. The set of amino acid sequence representations and the IPC sequence representation may be processed in parallel. The system may generate a synthetic representation by combining each of the transformed BOS token representations of the set of transformed amino acid sequence representations with the transformed BOS token representation of the transformed IPC sequence representation. The system may determine one or more predicted amino acid-IPC interactions based on the synthetic representation.
[0015] The set of transformed amino acid sequence representations includes a set of amino-terminal flank (N-flank) representations or a set of carboxyl-terminal flank (C-flank) representations.
[0016] In some embodiments, the IPC of the subject is a major histocompatibility complex (MHC).
[0017] In some embodiments, the set of transformed amino acid sequence representations includes a set of MHC binding representations.
[0018] In some embodiments, the MHC includes MHC class II (MHC-II).
[0019] In some embodiments, the MHC comprises major histocompatibility complex class I (MHC-I).
[0020] In some embodiments, the IPC of the subject is a T cell receptor (TCR).
[0021] In some embodiments, the at least one protein may be a therapeutic protein.
[0022] In some embodiments, the at least one protein is present in a disease sample from the subject.
[0023] In some embodiments, the disease sample may be a tumor cell biopsy. In some embodiments, the disease sample includes cancer. In some embodiments, the disease sample includes tissue.
[0024] In some embodiments, generating a synthetic representation includes, for each of the transformed amino acid sequence representations in the set of transformed amino acid sequence representations, element-wise multiplying the transformed amino acid sequence start (BOS) representation corresponding to the transformed amino acid sequence representation by the transformed IPC sequence start (BOS) representation corresponding to the transformed IPC sequence representation.
[0025] In some embodiments, processing a set of amino acid sequence representations includes processing a peptide sequence start (BOS) representation to generate a transformed peptide sequence representation.
[0026] In some embodiments, processing an IPC sequence representation includes processing an MHC sequence start representation (BOS) to generate a transformed MHC sequence representation.
[0027] Some additional aspects include, for each of a set of IPC sequences, the steps of: accessing the IPC sequence, processing the set of amino acid sequence representations, processing the IPC sequence representation using the corresponding IPC sequence, and generating a synthetic representation. The determined one or more predicted amino acid-IPC interactions may be based on the synthetic representation corresponding to the set of IPC sequences.
[0028] In some embodiments, the set of IPC sequences comprises twelve major histocompatibility complex (MHC) allotypes of the subject and / or six MHC alleles of the subject.
[0029] Some additional aspects include selecting one or more peptide-IPC combinations from the set of amino acid sequences and the set of IPC sequences to be included as targets for immunotherapy based on one or more determined amino acid-IPC interactions.
[0030] In some embodiments, processing an amino acid sequence representation group includes: converting the amino acid sequence representation group into a converted amino acid sequence representation group using the one or more first processing blocks. Each of the one or more first processing blocks includes a set of processing sub-blocks.
[0031] Some additional aspects include embedding a group of amino acid sequences to generate a group of embedded amino acid sequence representations; and performing positional encoding on the group of embedded amino acid sequence representations.
[0032] In some embodiments, processing an IPC sequence representation includes: converting the IPC sequence representation into a converted IPC sequence representation using a second processing block. The second processing block includes a set of processing sub-blocks.
[0033] Some additional aspects include embedding the IPC sequence to generate an embedded IPC sequence representation; and performing positional encoding on the embedded IPC sequence representation.
[0034] In some embodiments, a machine learning model includes one or more transformer encoders, and each of the one or more transformer encoders includes a processing layer.
[0035] In some embodiments, each of the one or more first processing blocks or the second processing block includes a set of processing sub-blocks; and each of the set of processing sub-blocks includes a neural network having at least one processing layer.
[0036] In some embodiments, the amino acid sequence representation group includes an aggregate sequence representation that includes a set of peptide representations and one or more of the following: a set of amino-terminal flanking (N-flanking) representations or a set of carboxy-terminal flanking (C-flanking) representations.
[0037] Some additional aspects include, prior to generating the converted amino acid sequence representation group: flattening the aggregate sequence representation into a single array; and densifying the aggregate sequence representation by removing blank lines from the array. The converted amino acid sequence representation is generated based on the densified aggregate sequence representation.
[0038] In some embodiments, the IPC sequence representation includes an aggregate sequence representation that includes one or more of the following: a major histocompatibility complex (MHC) sequence representation or a T cell receptor (TCR) sequence representation.
[0039] In some embodiments, processing an amino acid sequence representation group includes: for each amino acid sequence representation in the group: for each element of the amino acid sequence representation, determining a plurality of vectors based on a set of weights associated with the processing layer of the machine learning model; and generating a group of element focus scores based on the plurality of vectors and the set of weights.
[0040] In some embodiments, the plurality of vectors includes a key vector, a value vector, and a query vector, and the set of weights includes a set of key weights, a set of value weights, and a set of query weights.
[0041] In some embodiments, generating the element focus score group includes: determining each element focus score based on each pair of elements from the query vector and the key vector.
[0042] In some embodiments, the one or more first processing blocks and the second processing block include attention blocks, each attention block includes a set of attention sub-blocks, and each attention sub-block includes a self-attention layer; and the machine learning model is an attention-based machine learning model.
[0043] Some additional aspects include generating an attention map including masks through one or more of the attention blocks, where the masks restrict the attention applied by the attention sub-blocks to the sequence length according to the masks.
[0044] Some additional aspects include using a fully connected block in the output subsystem of the machine learning model to process the synthetic representation to generate a first output; using a dropout block in the output subsystem of the machine learning model to apply dropout to the first output to generate a second output; and using a max layer in the output subsystem of the machine learning model to select a subgroup of the second output to generate a result. The one or more predicted amino acid-IPC interactions are determined based on the result.
[0045] In some embodiments, the amino acid sequence group includes peptide sequences and the IPC sequence includes a major histocompatibility complex (MHC) sequence, and the one or more predicted amino acid-IPC interactions include one or more of the following: an interaction affinity prediction for a peptide-IPC combination, which predicts the binding affinity between a peptide and an MHC; an interaction prediction for a peptide-IPC combination, which predicts whether the MHC will present the peptide at the cell surface; or an immunogenicity prediction for a peptide-IPC combination, which predicts the ability of a peptide to induce an immune response regarding the MHC.
[0046] In some embodiments, determining the one or more predicted amino acid-IPC interactions includes: processing the synthetic representation to generate a set of results; and selecting an amino acid-IPC combination based on the highest result among the set of results.
[0047] In some embodiments, the one or more predicted amino acid-IPC interactions include a prediction of the tumor-specific immunogenicity of a peptide.
[0048] In some embodiments, the amino acid sequence set comprises a set of peptide sequences. The one or more predicted amino acid-IPC interactions identify a subset of peptide sequences having increased tumor-specific immunogenicity or increased likelihood of presentation by IPC relative to the set of peptide sequences.
[0049] Some additional aspects include identifying a subset of peptides from the amino acid sequence set for inclusion in an individualized vaccine based on the determined one or more predicted amino acid-IPC interactions.
[0050] Some additional aspects include generating a treatment recommendation comprising the individualized vaccine.
[0051] Some additional aspects include selecting a subset of peptides from the amino acid sequence set for inclusion as targets for immunotherapy based on the determined one or more predicted amino acid-IPC interactions.
[0052] In some embodiments, the immunotherapy comprises one or more of the following: T cell therapy, personalized cancer therapy, antigen-specific immunotherapy, antigen-dependent immunotherapy, a vaccine, or natural killer (NK) cell therapy.
[0053] Some additional aspects include selecting a subset of peptides from the amino acid sequence set for exclusion as targets for immunotherapy based on the determined one or more predicted amino acid-IPC interactions.
[0054] In some embodiments, the immunotherapy comprises one or more of the following: T cell therapy, personalized cancer therapy, antigen-specific immunotherapy, antigen-dependent immunotherapy, a vaccine, or natural killer (NK) cell therapy.
[0055] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium that contains instructions that, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more of the methods disclosed herein.
[0056] In some embodiments, a computer program product is provided that is tangibly embodied in a non-transitory machine-readable storage medium and includes instructions configured to cause one or more data processors to execute part or all of one or more of the methods disclosed herein.
[0057] Some embodiments of the present disclosure include a system that includes one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods disclosed herein and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium that includes instructions configured to cause one or more data processors to perform some or all of one or more methods disclosed herein and / or some or all of one or more processes disclosed herein.
[0058] The terms and expressions employed are used in a descriptive rather than a restrictive sense, and in using such terms and expressions, no intention is made to exclude any equivalents of the features shown and described or portions thereof, but it should be recognized that various modifications are possible within the scope of the invention as claimed. Accordingly, it should be understood that although the invention as claimed has been specifically disclosed by way of embodiments and optional features, modifications and variations of the concepts disclosed herein may be employed by those skilled in the art, and such modifications and variations are considered to be within the scope of the invention as defined by the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The present disclosure is described in conjunction with the accompanying drawings:
[0060] Figure 1 FIG. is a block diagram of an exemplary prediction system according to some embodiments.
[0061] Figure 2 FIG. is a flowchart of an exemplary method for predicting amino acid-immune protein complex (IPC) interactions using a machine learning model according to some embodiments.
[0062] Figure 3 FIG. is a schematic diagram of an exemplary configuration of a machine learning model according to some embodiments.
[0063] Figure 4A FIG. is an exemplary workflow diagram for obtaining one or more interaction predictions, one or more interaction affinity predictions, and / or one or more immunogenicity predictions according to some embodiments.
[0064] Figure 4B FIG. is an exemplary workflow diagram for obtaining one or more interaction predictions, one or more interaction affinity predictions, and / or one or more immunogenicity predictions according to some embodiments.
[0065] Figure 4CAn exemplary workflow diagram for obtaining one or more interaction predictions, one or more interaction affinity predictions, and / or one or more immunogenicity predictions according to some embodiments.
[0066] Figure 4D An exemplary workflow diagram for obtaining one or more interaction predictions, one or more interaction affinity predictions, and / or one or more immunogenicity predictions according to some embodiments.
[0067] Figure 4E An exemplary protein language model for generating protein sequence embeddings according to some embodiments.
[0068] Figure 4F An exemplary protein language model for generating MHC sequence embeddings according to some embodiments.
[0069] Figure 4G An exemplary workflow diagram for predicting the interactions of multiple peptides with MHC molecules that may have been expressed by multiple alleles or allotypes according to some embodiments.
[0070] Figure 4H An exemplary workflow diagram for an attention mask according to some embodiments.
[0071] Fig. 4I An exemplary workflow diagram for an attention mask with a calibration step for reducing model bias according to some embodiments.
[0072] Figures 5A to 5C A schematic diagram of an exemplary machine learning model according to some embodiments.
[0073] Figure 6 A schematic diagram of an exemplary processing block according to some embodiments.
[0074] Fig. 7A A flowchart of an exemplary method for processing sequence representations using an exemplary processing layer according to some embodiments.
[0075] Figure 7B A schematic diagram showing an exemplary method for processing sequence representations using an exemplary processing layer according to some embodiments.
[0076] Figure 8 A flowchart of an exemplary method for generating information about the immunological activity of various peptides according to some embodiments.
[0077] Fig. 9 A flowchart of an exemplary method for generating information about the immunological activity of various peptides according to some embodiments.
[0078] Fig.10Flowchart of an exemplary method for training a machine learning model and using the trained machine learning model to generate predictions related to amino acids and IPC according to some embodiments.
[0079] Fig.11 Illustration of an example of training data according to some embodiments. In this figure, the IPC is MHC class II allotypes (HLA-DR, HLA-DP, and HLA-DQ allotypes).
[0080] Fig.12 Exemplary method for predicting which therapeutic antibodies are likely to increase the risk of immunogenicity.
[0081] Fig.13 Illustration of exemplary neoantigen candidates (mutated antigens) and corresponding potential neoantigen candidates (mutated peptides) according to some embodiments.
[0082] Fig.14A Includes an exemplary graph indicating the performance of an exemplary P-MHC-II model according to some embodiments.
[0083] Fig. 14B Includes an exemplary graph indicating the performance of a previously used method (Model A) regarding its elution output according to some embodiments.
[0084] Fig.15 Exemplary graph of exemplary mean precision values of the elution-ligand outputs of a previously used method (Model A) and an exemplary P-MHC-II model for each allotype in a comparative test dataset according to some embodiments.
[0085] Fig.16A and 16B Exemplary graphs respectively showing the performance of an exemplary P-MHC-II model and a previously used method (Model A) according to some embodiments.
[0086] Fig.17 Graph showing a latent space including multiple peptide vectors according to some embodiments.
[0087] Fig.18 Histogram showing the count of peptides with different information content levels in an exemplary dataset according to some embodiments.
[0088] Fig.19A Protein space colored by protein expression according to some embodiments.
[0089] Fig.19B Shows the cellular compartmentalization of different proteins and their positions in the latent space according to some embodiments.
[0090] Fig. 20 Shows exemplary performance data according to some embodiments.
[0091] Fig.21 Is a block diagram of a computer system according to some embodiments.
[0092] Fig. 22 Is a block diagram of an exemplary artificial intelligence (AI) architecture included as part of the exemplary computing system of FIG. 16 according to some embodiments.
[0093] In the figures, similar components and / or features may have the same reference numerals. Additionally, various parts of the same type can be distinguished by following the reference numeral with a dash and a second label that differentiates the similar parts. If only the first reference numeral is used in the specification, the description applies to any one of the similar parts having the same first reference numeral, regardless of the second reference numeral. Detailed Description
[0094] Recognizing the importance of being able to predict which mutant peptides (e.g., neoantigens) are selected as candidates for personalized vaccines, the embodiments described herein provide methods and systems for making more accurate predictions than currently available methods and systems. The embodiments described herein use machine learning methods and systems to improve prediction performance by, for example but not limited to, reducing the number of false positives generated when analyzing mutant peptide sequences to determine the viability of those mutant peptides as candidate vaccines. The embodiments described herein also provide methods and systems for determining whether certain therapeutic antibodies are likely to pose an immunogenic risk to a subject.
[0095] The present disclosure provides systems, methods, and programming for obtaining the following as described herein: one or more interaction predictions (e.g., interaction of a predicted peptide with an MHC molecule), one or more interaction affinity predictions (e.g., binding affinity between a predicted peptide and an MHC), and / or one or more immunogenicity predictions (prediction of the ability of a peptide to elicit an immune response with respect to an MHC). A machine learning model can perform a combination of peptide data processing, n-flank / c-flank data processing, MHC data processing, TCR data processing, and / or protein data processing for predicting the interaction of a peptide with an MHC molecule. In one exemplary workflow, the processing of MHC data can involve processing of an MHC sequence appended with a BOS token using one or more transformer stages, or alternatively can involve processing of an MHC sequence embedding generated using a protein language model. In this exemplary workflow, the processing of peptide data can involve processing of a peptide sequence appended with a BOS token using one or more transformer stages, or alternatively can involve processing of a peptide sequence not appended with a BOS token using a cross-attention module. This exemplary workflow can additionally incorporate the processing of protein data, or alternatively can not incorporate the processing of protein data. This exemplary workflow can additionally incorporate the processing of TCR data, or alternatively can not incorporate the processing of TCR data.
[0096] For example, the embodiments described herein provide a machine learning model, various methods using the machine learning model, and / or an output generated by analyzing sequences identified from a disease sample from a subject via the machine learning model. To predict whether a mutant peptide identified from a disease sample interacts with an IPC such as an MHC molecule (e.g., MHC-I or MHC-II) and / or the extent of the interaction, the machine learning model separately processes a set of amino acid sequence representations in addition to the processing of the IPC sequence representation (e.g., MHC sequence representation). The sequence representation can be an embedding of the features of the corresponding sequence. In some embodiments, the mutant peptide sequence representation can be referred to as a variant-encoded sequence. The IPC sequence (e.g., MHC sequence) can include at least a portion of an MHC molecule, a complete sequence, a pseudo-sequence of an MHC molecule (the portion that interacts with the mutant peptide (including the binding pocket, including some portions of the pseudo-sequence)).
[0097] The machine learning model includes various subsystems for processing. The machine learning model can include, for example, a representation subsystem, a processing subsystem, a synthesis subsystem, and an output subsystem. Each "subsystem" can include one or more blocks, where each block includes one or more sub-blocks and / or layers. A sub-block can include any number of layers (or units).
[0098] A processing subsystem can be used to generate one or more transformed sequence representations, such as a set of transformed amino acid sequence representations (e.g., which may include variant coding sequences), a transformed IPC sequence representation, etc. In some embodiments, the processing subsystem can process one or more (e.g., a set of) amino acid sequence representations independently of or separately (e.g., in parallel) from the IPC sequence representation. For example, one or more first processing blocks in the processing subsystem can be used to process a set of amino acid sequence representations, and a second processing block in the processing subsystem can be used to process the IPC sequence representation. Processing the amino acid sequence representation set and the IPC sequence representation via these parallel processing engines and / or separate processing blocks can improve the prediction performance of the machine learning model. Separate processing engines can force the system to learn separate representations of different biological structures. In contrast, processing different biological structures via a common processing engine may lead to model overfitting.
[0099] In addition, the embodiments described herein recognize and account for the fact that training a model corresponding to a series of biological events may significantly require more data than training a model corresponding to a single biological event. Due to the vast number of potentially observable sequences, training a model for sequence analysis can be particularly complex. For example, not only are there millions of potential new antigens, but the genes encoding proteins for MHC class II molecules also have a high degree of polymorphism. In fact, nearly 6,000 variant α and β chain proteins of HLA-DR, HLA-DQ, and HLA-DP (the three classical class II molecules in humans) are currently present in the IPD-IMGT / HLA database. Thus, the embodiments described herein provide methods and systems for training machine learning models that both reduce training complexity and improve training performance. For example, variant coding sequences for training can be selected and / or pruned such that training is performed using variant coding sequences having an amino acid length equal to or below a threshold amino acid length (e.g., nine (9) amino acids). Generating a training data set that includes variant coding sequences having a length equal to or shorter than the threshold amino acid length can reduce overall training complexity as well as improve training and / or prediction performance (e.g., reduce the variation in performance metrics per epoch, thereby improving prediction performance).
[0100] Accordingly, the technology disclosed herein includes machine learning-based methods for determining predicted amino acid-IPC interactions related to immune activity associated with a peptide, such as a mutant peptide. A machine learning model can generate an output that includes one or more predicted amino acid-IPC interactions. For example, the output can include one or more of the following: one or more interaction predictions, one or more interaction affinity predictions, or one or more immunogenicity predictions (i.e., predictions related to the ability of a peptide to elicit an immune response). An interaction prediction can include a prediction related to whether one or more target interactions occur with a peptide (e.g., a mutant peptide, including a given ordered set of amino acids identified by a given variant coding sequence). In some embodiments, the target interaction can be a binding of the peptide to an IPC (e.g., an MHC molecule, a TCR), a peptide presented by an MHC molecule at the cell surface, or another type of target interaction. An interaction affinity prediction can include a prediction of the affinity of one or more target interactions. For example, an interaction affinity prediction can indicate a binding affinity relative to peptide-MHC binding. An interaction (e.g., binding) affinity can be determined based on the trend, strength, and / or stability of the interaction (e.g., binding). An immunogenicity prediction refers to predicting the ability of a peptide to elicit an immune response. The immune system discriminates a peptide as non-self or foreign. Once recognized, the peptide stimulates the immune system to generate a response. The response can include the generation of antibodies by B cells (humoral immunity) and the activation of T cells (cell-mediated immunity) to eliminate the cell presenting the peptide.
[0101] In some embodiments, the output can include or indicate the immunogenicity of a peptide or a therapeutic antibody. For example, the output can predict whether a peptide will trigger an immune response in a particular subject or group of subjects. These immunogenicity predictions can be determined for each of a plurality of mutant peptides. In some embodiments, the immunogenicity predictions can be used to select and / or rank one or more mutant peptides and / or pharmaceutical compositions to be included in a vaccine and / or for treatment. For example, but not limited to, mutant peptides associated with high predicted binding affinity, a high probability of being presented on the surface of tumor cells, and / or high predicted immunogenicity can be selected for inclusion in a vaccine or for treatment.
[0102] The embodiments described herein provide methods and systems for using a machine learning model to determine predicted amino acid-IPC interactions indicative of immunological activity associated with amino acids (e.g., peptides) and immunological protein complexes (IPCs). The IPC may comprise MHC or TCR. A set of amino acid sequences identified from at least one protein may be accessed. In some embodiments, the at least one protein is a therapeutic protein. In some embodiments, the at least one protein is present in a disease sample from a subject. IPC sequences may be identified for the subject's IPCs and then accessed. One or more first processing blocks in a processing subsystem of the machine learning model are used to process a set of amino acid sequence representations. Processing the set of amino acid sequence representations may generate a set of transformed amino acid sequence representations. A second processing block in the processing subsystem may be used to process the IPC sequence representations to generate transformed IPC sequence representations. In some embodiments, the BOS token embedding of each of the set of transformed amino acid sequence representations may be combined with the BOS token embedding of the transformed IPC sequence representation to generate a synthetic representation. The method may generate an output comprising predicted amino acid-IPC interactions determined based on the synthetic representation. The results (e.g., the output) may comprise one or more of the following: one or more interaction predictions for a corresponding amino acid-IPC combination, one or more interaction affinity predictions, or one or more immunogenicity predictions. In some embodiments, a report is generated based on the output.
[0103] The techniques described herein provide a number of technical advantages. For example, the techniques described herein may avoid or reduce overfitting. Overfitting is an undesirable machine learning behavior that occurs when a machine learning model provides accurate output for training data rather than for new data. Insufficient training data for the implementation of a machine learning model such as a transformer for peptide-MHC allele binding may result in overfitting. A typical transformer may have a large number of parameters that the model learns from the training data. In particular, obtaining data on peptide-MHC allele binding may be challenging for several reasons, including a) the high variability of MHC molecules, b) the complexity of peptide binding, c) subject variability in MHC expression, d) ethical constraints, and e) other technical limitations that require specialized equipment and expertise. Collection of training data for other IPC machine learning model predictions (e.g., TCR) may encounter similar challenges.
[0104] Some embodiments illustrate computational techniques for training machine learning models in a manner that eliminates overfitting caused by insufficient training data. Such techniques include using a protein language model (PLM) to process IPC alleles (e.g., MHC alleles) and infer useful and generalizable amino acid or residue features that can be used in the training process of a machine learning model. Thereby reducing the likelihood of overfitting and improving the performance of the machine learning model. For example, when the PLM processes an IPC allele sequence (e.g., to generate training data), the trained machine learning model can learn which residues or amino acids in the peptide-MHC binding complex are relevant or affect such binding.
[0105] In some instances, the techniques described herein can advantageously incorporate information about the source protein from which the peptide is derived. Without incorporating information about the source protein, the model may lack useful protein / peptide features to make the predictions described herein. By incorporating information about the source protein (e.g., protein sequence embeddings generated via the PLM), the system can select peptides that are more likely to be presented by MHC molecules because the system can capture features such as: processing signals (i.e., flanking signals generated when an enzyme breaks down a protein into peptides), expression of genes associated with the source protein, cellular localization of the source protein, and other suitable features that affect the presentation of peptides by MHC molecules and other interactions predicted by the embodiments described herein. In some cases, source protein expression may be important because a source protein that is not sufficiently expressed in a subject will not generate enough peptides, even if such peptides are capable of eliciting an immunogenic effect. Cellular localization of the source protein is another feature that can be obtained by processing the source protein using the PLM. When the source protein is intracellular or generated intracellularly, MHC I presentation can occur more frequently. On the other hand, proteins that are predominantly extracellular (e.g., proteins that reside only in vesicles) may be less likely to be presented by MHC class I because they are located in different parts of the cell. Instead, proteins that are predominantly extracellular (e.g., proteins that reside only in vesicles) are more likely to be presented by MHC II. Thus, for MHC I, it is preferred that the source protein is derived intracellularly, and for MHC II, it is preferred that the protein is vesicular or extracellular. In summary, there are many features associated with the source protein that affect peptide presentation, and these features are advantageously incorporated into the techniques described herein.
[0106] In some embodiments, the techniques described herein can advantageously utilize a cross-attention module to incorporate MHC allele-specific binding information. When processing peptide data, MHC allele information can be valuable. For example, a peptide can contain multiple binding cores (i.e., specific peptide regions where binding occurs), and the multiple binding cores can bind to different MHC alleles. The system can use the cross-attention module to incorporate MHC allele-specific vectors into the processing of peptide data. In this way, the system can predict binding core information from peptides with multiple binding cores that can bind to different MHC alleles and generate multiple binding core predictions corresponding to the multiple different alleles. In other words, the techniques proposed herein are not limited to predicting one binding core per peptide. In some embodiments, a machine learning model can predict more than one binding core per peptide.
[0107] In some instances, the techniques described herein can advantageously prevent over-prediction of the start of peptide amino acid positions indicating binding cores. For example, position zero (i.e., the starting position) is unlikely to be the starting position of a peptide's binding core because it is expected that the binding core is part of the peptide (e.g., a 20-mer) (e.g., a 9-mer) and is flanked by binding cores within the peptide. To mitigate the over-prediction problem, the techniques described herein can include a calibration step to remove the bias towards position zero or other peptide positions that suffer from over-prediction. Before performing the calibration step, the system can first calculate the model bias for any single position in the peptide. The system can modify the attention map by subtracting the model bias to remove the model bias for any single position in the peptide to perform the calibration step.
[0108] In some embodiments, the techniques described herein can advantageously use a dimensionality reduction module to process MHC data (e.g., MHC sequence embeddings) to prevent overfitting. For example, overfitting can occur if different types of models (such as neural networks) are trained to reduce the dimensionality of the input embeddings. If the same set of MHC alleles (usually a relatively small set of MHC alleles) used to train the PLM model is used to train a neural network, the neural network model can generate correct outputs when processing input embeddings corresponding to the MHC alleles in the training data but generate incorrect outputs when processing input embeddings corresponding to new MHC alleles not included in the training data. On the other hand, PCA techniques can be trained using a larger set of MHC alleles (e.g., the space of all alleles) and can therefore efficiently and accurately reduce the dimensionality of input embeddings for new alleles not in the training dataset, thus avoiding the overfitting problem.
[0109] In the following description, it should be noted that the embodiments shown herein are not limited to the specific advantages disclosed above. The present disclosure encompasses various technical benefits and enhancements, which are detailed throughout this written description. These embodiments are presented in a non-limiting manner, and it is understood that many modifications, variations, and refinements can be made without departing from the spirit and scope of the present invention. The following description provides further technical advantages and novel aspects inherent in the presented embodiments, thus providing a broader perspective on the applicability and utility of the disclosed invention.
[0110] The following description provides exemplary implementations of these methods and systems, where the output (e.g., predicted amino acid-IPC interactions) can be used to plan, design, and / or manufacture treatments.
[0111] Exemplary Prediction System
[0112] Figure 1 FIG. is a block diagram of an exemplary prediction system according to some embodiments. The prediction system 100 is for determining predicted amino acid-IPC interactions related to the immunological activity of a peptide and specifically a mutant peptide. The prediction system 100 includes a computing platform 102, a data repository 104, and a display system 106. The computing platform 102 can take various forms. In some embodiments, the computing platform 102 includes a single computer (or computer system) or multiple computers that communicate with each other. In some embodiments, the computing platform 102 can be a cloud computing platform.
[0113] The data storage 104 and the display system 106 each communicate with the computing platform 102. In some instances, one or more of the data repository 104 or the display system 106 can be considered part of the computing platform 102 or otherwise integrated with the computing platform. Thus, in some instances, the computing platform 102, the data repository 104, and the display system 106 can be independent components that communicate with each other, but in other instances, some combination of these components can be integrated together. Communication between the different components can be performed using any number of wired communication links, wireless communication links, optical communication links, or combinations thereof.
[0114] The prediction system 100 includes a sequence analyzer 108, which can be implemented using hardware, software, firmware, or a combination thereof. In some embodiments, the sequence analyzer 108 is implemented in the computing platform 102. The sequence analyzer 108 receives sequence data 110 for processing. For example, the sequence data 110 can be sent as input into the sequence analyzer 108, retrieved from the data repository 104 or some other type of storage (e.g., cloud storage), accessed from cloud storage, or obtained in some other way. In some cases, the sequence data 110 can be retrieved from the data repository 104 in response to receiving user input entered by the user via an input device.
[0115] Sequence data 110 can be generated by processing a set of samples 112. The set of samples 112 can take the form of one or more biological samples from one or more subjects (e.g., diseased samples, healthy samples, or combinations thereof). The set of samples 112 can include samples obtained from tumors of the subjects. The tumors can be, for example, lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, T-cell lymphocytic leukemia, non-small cell lung cancer, small cell lung cancer, or combinations thereof.
[0116] The samples in the set of samples 112 can include, for example, various IPC molecules, various peptides, or combinations thereof. When the set of samples 112 includes diseased samples, the peptides can include one or more mutant peptides (e.g., neoantigens). The IPC molecules can include, for example, various MHC molecules, various TCR molecules, or combinations thereof.
[0117] In some embodiments, the set of samples 112 includes immune protein complexes (IPCs) 114 (e.g., MHC class I molecules, MHC class II molecules, various TCR molecules, etc.). Further, the set of samples can include at least one protein 123 (i.e., the source protein). An amino acid chain 116 can be identified from at least one protein 123 and can be an amino acid chain including a peptide 118, an N-flank 120, and a C-flank 122. The amino acid chain 116 can include or exclude the N-terminus between the peptide 118 and the N-flank 120, or the C-terminus between the peptide 118 and the C-flank 122. When the peptide 118 includes one or more variants (e.g., one or more sequence variants) compared to a corresponding reference sequence, the peptide is considered a mutant peptide. The protein 123 is the source protein of the amino acid chain 116, which can be generated by proteolysis, the process by which a protein (e.g., protein 123) is broken down into smaller polypeptides or amino acids. The protein 123 can be broken down into smaller polypeptides or amino acids by enzymatic cleavage, where a specific enzyme called a protease cleaves the peptide bonds between the amino acids in the protein 123.
[0118] A sample group 112 can be processed to generate sequence data 110. In some embodiments, multiple samples in the sample group 112 can be processed at different times. In some embodiments, the prediction system 100 includes a sample analyzer that is configured to process the sample group 112 to generate the sequence data 110. The sequence data 110 includes, for example, at least one amino acid sequence 129 and at least one immune protein complex (IPC) sequence 124 (e.g., an IPC sequence 124 corresponding to the IPC 114). The amino acid sequence 129 can include one or more of the following: a peptide sequence 126 (e.g., a peptide sequence 126 corresponding to a peptide 118), an amino-terminal flanking (N-flanking) sequence 128 (e.g., an N-flanking sequence 128 corresponding to an N-flanking 120), or a carboxyl-terminal flanking (C-flanking) sequence 130 (e.g., a C-flanking sequence 130 corresponding to a C-flanking 122). One or more subsequences of the amino acid sequence 129 (e.g., the peptide sequence 126, the N-flanking sequence 128, and the C-flanking sequence 130) can be processed individually or as a single sequence.
[0119] When the immune protein complex 114 is MHC, the IPC sequence 124 can be, for example, an MHC sequence 135 that characterizes at least a portion of the MHC. When the immune protein complex 114 is TCR, the IPC sequence 124 can be, for example, a TCR sequence 131 that characterizes at least a portion of the TCR. In some embodiments, the IPC sequence 124 can include an MHC sequence 135 that characterizes at least a portion of the MHC molecule and a TCR sequence 131 that characterizes at least a portion of the TCR molecule. In some embodiments, the sequence data 110 can include an IPC sequence 124 in the form of an MHC sequence 135 that characterizes at least a portion of the MHC molecule, and a separate TCR sequence 131 that characterizes at least a portion of the TCR.
[0120] The protein sequence 160 characterizes at least a portion of the protein 123. In some embodiments, the protein sequence 160 can be identified by performing a reverse lookup in a database (e.g., the UniProt database) based on mutant peptide data obtained from the sample (e.g., the IPC sequence 124).
[0121] The peptide sequence 126 characterizes at least a portion of the peptide 118. The N-flanking sequence 128 characterizes at least a portion of the N-flanking 120. In some embodiments, when the number of amino acids (or amino acid residues) upstream of the N-terminus is large, the corresponding sequence of the N-flanking 120 can be trimmed to generate the N-flanking sequence 128. The C-flanking sequence 130 characterizes at least a portion of the C-flanking 122. In some embodiments, when the number of amino acids (or amino acid residues) downstream of the C-terminus is large, the corresponding sequence of the C-flanking 122 can be trimmed to generate the C-flanking sequence 130.
[0122] The sequence analyzer 108 receives sequence data 110 as input for processing. The sequence analyzer 108 includes a machine learning model 132 that processes the sequence data 110. In some embodiments, the sequence data 110 is sent directly to the machine learning model 132 for processing. In some embodiments, the sequence analyzer 108 preprocesses the sequence data 110 before sending the sequence data 110 to the machine learning model 132 for processing. The preprocessing of the sequence data 110 may include appending a sequence start (BOS) token to each of the plurality of sequences in the sequence data 110. The BOS token appended to the peptide sequence can be used as an additional data structure that can be used to represent peptide characteristics that can be used to determine the presentation likelihood, binding affinity, or immunogenicity prediction of the corresponding peptide sequence 126. The BOS token can indicate interaction information, such as whether the peptide will be presented by an allele / allotype, binding affinity, immunogenicity, or any other suitable purpose for which the machine learning model 132 has been trained.
[0123] The machine learning model 132 can be implemented in any of many different ways. In some embodiments, the machine learning model 132 can be any type of model that uses an element-focused scoring of a binding core represented by a set of amino acid sequence representations. The machine learning model 132 can be used in a training mode or a prediction mode. In the training mode, the machine learning model 132 is trained using training data 133. Examples of the training data 133 are described in more detail below. The machine learning model 132 is trained to be used in the prediction mode.
[0124] The machine learning model 132 processes the IPC sequence 124 via the IPC processing engine 134 and processes the amino acid sequence 129 via the amino acid processing engine 139. Separate processing engines for IPC and amino acids can improve the prediction performance of the machine learning model 132. In some embodiments, the machine learning model 132 processes one or more of the following: processes the MHC sequence 135 via the MHC processing engine 141, or processes the TCR sequence 131 via the TCR processing engine 142. In some embodiments, the machine learning model 132 processes one or more of the following: processes the peptide sequence 126 via the peptide processing engine 136, processes the N-flank sequence 128 via the N-flank processing engine 138, or processes the C-flank sequence 130 via the C-flank processing engine 140. In some embodiments, the machine learning model 132 processes the protein sequence 160 via the protein processing engine 162. Examples of the implementation for these different processing engines are described in more detail below.
[0125] As used herein, the terms "processing engine" and "engine" identify at least one software component and / or a combination of at least one software component and at least one hardware component that is designed / programmed / configured to interact and / or communicate with other software and / or hardware components, including but not limited to other processing engines.
[0126] Machine learning model 132 processes sequence data 110 to generate an output for generating report 144. Report 144 may include the accurate output of machine learning model 132, a transformed or filtered version of the output, or both. In some cases, report 144 may include a notification, a recommendation, a warning, or other information generated by sequence analyzer 108 based on the output of machine learning model 132.
[0127] Report 144 may be an output that includes information, for example, about a target immunological activity associated with one or more peptides (e.g., one or more mutant peptides). For example, report 144 may include information about an immunological activity associated with amino acid 116 (e.g., peptide 118, N-flank 120, C-flank 122, etc.) and IPC 114 (e.g., MHC-I, MHC-II, TCR, etc.). Report 144 may include, for example, interaction information 146 (e.g., an interaction affinity prediction that predicts the binding affinity between a peptide and an MHC, or an interaction prediction that predicts whether an MHC allele or allotype will present a peptide at the cell surface), immunogenicity information 148 (e.g., an immunogenicity prediction that predicts the ability of a peptide to elicit an immune response against an MHC), or both. Interaction information 146 may provide a prediction of a selected set of interactions between amino acid 116 and IPC 114. Immunogenicity information 148 may provide a prediction of the immunogenicity of amino acid 116 (e.g., including the immunogenicity of peptide 118).
[0128] In some embodiments, report 144 may be displayed on a graphical user interface (GUI) 150 on display system 106. A user may view report 144 and / or interact with the report via graphical user interface 150. In some embodiments, a user may use report 144 to make a decision about the treatment of a subject from whom at least one of a set of samples 112 was obtained (or collected).
[0129] In some embodiments, prediction system 100 sends report 144 to a remote system 152 (e.g., wirelessly). Remote system 152 may be a cloud computing platform, cloud storage, another computer system, a user device (e.g., a smartphone, a tablet, a laptop, etc.), or some other type of platform. In some embodiments, remote system 152 may be a treatment manufacturing system (or machine) or a part thereof.
[0130] Figure 2Flowchart of an exemplary method (e.g., a computer-implemented method) for predicting amino acid-IPC interactions using a machine learning model according to some embodiments. Process 200 may be implemented using the prediction system 100 described in Figure 1 For example, process 200 may be implemented using the sequence analyzer 108 and the machine learning model 132 in Figure 1 .
[0131] Process 200 may include, for example, step 202. Step 202 includes appending a BOS token to each amino acid sequence in a set of amino acid sequences. Each of the amino acid sequences in the set may have been identified from at least one protein. In some embodiments, the at least one protein is a therapeutic protein. In some embodiments, the at least one protein is present in a disease sample from a subject. As a non-limiting example, the disease sample may be a tumor cell biopsy. Additionally or alternatively, the disease sample may include cancer, tissue, or both. In some embodiments, each amino acid sequence includes one or more of the following: an amino-terminal flanking (N-flanking) sequence or a carboxyl-terminal flanking (C-flanking) sequence.
[0132] Step 204 includes appending a BOS token to an IPC sequence identified for the IPC of a subject. In some embodiments, the IPC of the subject is MHC. MHC may include MHC class II (MHC-II) or MHC class I (MHC-I). In some embodiments, the IPC of the subject is TCR.
[0133] Step 206 includes generating a set of amino acid sequence representations (e.g., embeddings) for each of the amino acid sequences and a set of IPC sequence representations (e.g., embeddings) from the set of IPC sequences. The embedding module may generate each of the sequence representations by forming an embedding for each element of the sequence (e.g., a single amino acid or a single nucleic acid) to represent the features of the element as a low-dimensional feature vector. The embedding module may also generate a positional embedding representing each absolute position (corresponding to one of the amino acids or one of the nucleic acids) as part of the sequence representation. The embedding corresponding to the BOS token may have a length equal to the number of features corresponding to each individual sequence element represented in the corresponding sequence to which the BOS token is appended.
[0134] Step 208 includes processing the set of amino acid sequence representations using one or more processing blocks in a processing subsystem of a machine learning model. Processing the set of amino acid sequence representations through each of one or more transformer stages may generate a set of transformed amino acid sequence representations stored in an embedding corresponding to the BOS token. Sequences with simple molecular structures (e.g., N-flanking / C-flanking sequences) may require only a single transformer stage, while sequences with more complex molecular structures (e.g., peptide sequences) may require multiple transformer stages. A transformer stage computes information about the frequency of correlations between each pair of amino acids in the input amino acid sequence representation and stores information about the pairwise correlations in the transformer output. Thus, a transformer stage may store information about the pairwise correlations between amino acids and the absolute position of each amino acid within each amino acid sequence (e.g., from positional embeddings) in the embedding corresponding to the BOS token within the transformed amino acid sequence representation. The set of transformed amino acid sequence representations may include a set of MHC binding representations. The set of transformed amino acid sequence representations may include one or more of the following: a set of amino-terminal flanking (N-flanking) sequence representations, a set of carboxyl-terminal flanking (C-flanking) sequence representations, or a set of combined N-flanking / C-flanking sequence representations. The N-flanking / C-flanking sequence representations may be processed in parallel with the amino acid sequence representations using separate processing blocks. Each of the processing blocks may utilize an attention mechanism to determine an element focus score.
[0135] Step 210 includes processing the IPC sequence representation using a second processing block in the processing subsystem to generate a transformed IPC sequence representation stored in an embedding corresponding to the BOS token, where the second processing block may operate independently of and in parallel with the first processing block. Thus, information about each IPC sequence may be stored in the embedding corresponding to the BOS token of the transformed IPC sequence representation. In some embodiments, in addition to MHC sequences, IPC sequences may also include TCR sequences, in which case TCR sequence representations may be generated separately from MHC sequence representations, and the two sequence representations may be processed in parallel using separate processing blocks. Each of the processing blocks may utilize an attention mechanism to determine an element focus score.
[0136] Step 212 includes generating a synthetic representation by combining each of the BOS token embeddings of each of the transformed amino acid sequence representations with the BOS token embedding for the transformed IPC sequence representation. The combination method may be element-wise multiplication, element-wise addition, or dot product calculation. In some embodiments, when the amino acid sequence includes a peptide sequence and an N-flanking sequence and / or a C-flanking sequence, before the combination step, the BOS token embedding corresponding to the peptide sequence may be concatenated with the BOS token embedding corresponding to the N-flanking / C-flanking sequence representation.
[0137] Step 214 includes determining predicted amino acid-IPC interactions based on the synthetic representation. The predicted amino acid-IPC interactions can include one or more of the following: one or more interaction predictions for one or more corresponding amino acid-IPC combinations, one or more interaction affinity predictions, or one or more immunogenicity predictions.
[0138] In some embodiments, an output (e.g., a report) can be generated. The output can be based on the predicted amino acid-IPC interactions. The output can be used to facilitate the design and / or manufacture of vaccines, therapeutics, and / or treatment regimens. For example, the report can identify a subgroup of the peptidome (comprised within a set of amino acid sequences), or provide an indication of which peptides to select as a subset for use in forming a treatment for a subject. The treatment can be, for example, a subset of peptides, precursors for each of the subset of peptides, or some other form.
[0139] Example architecture of a machine learning model
[0140] The machine learning model 132 of the embodiments described herein can include a plurality of subsystems (or subnets). Each of the plurality of subsystems can include an encoder, a transformer encoder, and / or one or more processing layers. In some cases, the machine learning model 132 can be configured to learn alignments (e.g., between an amino acid sequence and an IPC sequence, between a peptide sequence and an MHC sequence, between an MHC sequence and a TCR sequence, between an MHC-peptide complex and a TCR, between a peptide sequence and a TCR sequence). Alignment scoring functions (such as content-based functions, additive functions, position-based functions, dot product functions, and / or scaled dot product functions) can be used to learn and perform the alignments.
[0141] The machine learning model 132 can include one or more encoders configured to transform elements of an input sequence (e.g., an amino acid sequence, a nucleic acid sequence, a codon sequence, etc.) based on other elements of the input sequence, for example. The encoder can be a transformer encoder.
[0142] The machine learning model 132 may include one or more processing layers, such as self-attention layers or convolutional layers, or neural networks such as long short-term memory units (LSTMs), recursive structures, or recursive components. The machine learning model 132 may execute, for example, one or more self-attention layers. The machine learning model 132 may use a self-attention mechanism, a global attention mechanism, a soft attention mechanism, a local attention mechanism, and / or a hard attention mechanism. In some cases, the machine learning model 132 does not include any convolutional layers, any recursive structures, any LSTM units, and / or any recursive components. In some cases, the machine learning model 132 is not a recursive machine learning model and / or does not include a recurrent neural network. In some cases, the machine learning model 132 includes a recurrent neural network and / or may use positional encoding to process a sliding window of sequence elements across one or more sequences. In some cases, the machine learning model 132 is not a convolutional machine learning model and / or does not include a convolutional neural network.
[0143] The machine learning model 132 may include processing blocks, such as one or more first processing blocks, that are used to process one or more amino acid sequence representations independently of a second processing block that is used to process an IPC sequence representation. In some embodiments, the second processing block may process some or all of the IPC sequence representation (e.g., an MHC pseudo-sequence). The independence of these processing blocks may facilitate parallel processing when the machine learning model 132 is used. Further, the independence may improve the performance (e.g., prediction accuracy) of the machine learning model 132.
[0144] The machine learning model 132 can be configured such that the output value at any given layer depends not only on the corresponding input value but also on one or more (e.g., all) other input values. Thus, the machine learning model 132, the loss function, and / or the optimization function can be configured to optimize the output corresponding to a single position, which represents the degree to which a given IPC (e.g., MHC molecule) (represented by the corresponding input) will bind to a given peptide (represented by another corresponding input) and / or trigger immunogenicity in response to the given peptide. In some cases, the loss function can include a supervised loss function, such as binary cross - entropy or an unsupervised loss function. In some cases, the unsupervised loss function can include a contrastive loss or a regularization loss (e.g., L1 / L2 loss applied to the peptide representation to make the residual variation more continuous). In some cases, the loss function can include an auxiliary loss function. In such cases, the auxiliary loss function can be used together with the main loss function to train the machine learning model 132. Thus, in some cases, the auxiliary loss function can improve the learning process by adding additional information or constraints. In some cases, any one of the multiple outputs of the transformer encoder can represent such occurrence probabilities. The machine learning model 132 can be trained accordingly. In some cases, the end point (e.g., the remaining end point) can represent (in response to training) the binding affinity, the presentation (eluted ligand or EL), and / or the immunogenicity probability or likelihood. For example, the aggregated output can be fed to another layer, subsystem, or processing block (e.g., including one or more of the following: a processing layer such as a self - attention layer, or an encoder such as a transformer encoder).
[0145] In some cases, one, two, or all dimensions of the output from another layer and / or another subsystem or processing block have the same size as the input fed to other layers and / or other subsystems or processing blocks. In some cases, the input fed to other layers and / or other subsystems or processing blocks has a length along one axis that is greater than or equal to the sum of one or more of the following: the number of amino acids in the IPC sequence, the number of amino acids in the peptide sequence, the number of amino acids in the N - flanking sequence, or the number of amino acids in the C - flanking sequence. In some cases, the length of the input is one longer than the total number of amino acids. When, for example, additional feature vectors (e.g., a feature vector corresponding to a BOS token, or a token representing the IPC type) are appended to the amino - acid - specific feature values, the length of the input along one axis can exceed the total count of amino acids. Another dimension of the input can include a number of features (e.g., determined via hyperparameters). The output generated by other layers and / or other subsystems or processing blocks can have the same size as the input.
[0146] A subset of output values generated by other layers and / or other subnetworks may be processed by another neural network (e.g., a fully connected feed-forward network). The subset of values may include a 1-dimensional vector that may correspond to a set of feature values. For example, a 1-dimensional vector may correspond to a feature value associated with a BOS token.
[0147] In some embodiments, the neural network within the machine learning model 132 may be configured to output one or more results. The one or more results may include, for example, numerical results, binary results, and / or classified results. Each of the one or more results may predict whether an amino acid (e.g., a peptide) and an IPC undergo a specific type of reaction (e.g., combined together) and / or degree. The machine learning model 132 may include one or more activation layers to generate intermediate results (e.g., converting real intermediate values into binary and / or classified outputs). The machine learning model 132 may be trained to generate multiple types of predictions (e.g., interaction predictions, interaction affinity predictions, and / or immunogenicity predictions). In some cases, the prediction may be binary or classified. Other predictions may be non-binary or non-classified. For example, the prediction may be a scalar.
[0148] The machine learning model 132 may include an integrated model and / or may be included within an integrated model. The integrated model may include multiple (e.g., identical) sub-models that may be trained using different portions of the training data set.
[0149] Example configuration for machine learning models
[0150] Figure 3 According to some embodiments, Figure 1 A schematic diagram of an exemplary configuration of the machine learning model 132. Figure 1 300, which includes a representation subsystem 302, a processing subsystem 304, a synthesis subsystem 306, and an output subsystem 310. One or more subsystems within the machine learning model 132 may include one or more blocks, one or more sub-blocks, one or more layers, or a combination thereof. One or more blocks within the machine learning model 132 may include one or more sub-blocks, one or more layers, or a combination thereof. One or more sub-blocks of the machine learning model 132 may include one or more layers (or units).
[0151] In some embodiments, the representation subsystem 302 receives sequence data 110 as input and passes it through a tokenizer (not shown) that converts each letter of the sequence into a token (e.g., a unique integer stored in a lookup table associated with that letter). The tokenizer may append a BOS token to one or more of the sequences in the sequence data 110. The representation subsystem 302 may generate a sequence representation for each of the sequences in the sequence data 110. The sequence representation may include, for example, a stack of feature vectors corresponding to sequence elements (each sequence element representing or identifying one or more amino acids, one or more nucleic acids in the sequence corresponding to the sequence representation), or a BOS token. For example, each amino acid in the sequence may be represented by a unique feature vector, and the BOS token appended to the amino acid sequence may also be represented by a unique feature vector. The IPC sequence representation may contain six to twelve MHC alleles / allotypes of a given subject, and the processing subsystem 304 may generate up to twelve transformed MHC sequence representations, the transformed MHC sequence representations corresponding to one MHC-transformed sequence representation for each MHC allotype in combination with a particular peptide sequence. The amino acid sequence 129 may include one or more of the following: peptide sequence 126, N-flanking sequence 128, or C-flanking sequence 130, and the processing subsystem 304 may generate one amino acid sequence representation for each of the amino acid sequences. To normalize across different sequence lengths, the stack of feature vectors may be padded to a standard sequence length (SL), e.g., 39 vectors.
[0152] The processing subsystem 304 may receive one or more sequence representations (e.g., a set of amino acid sequence representations, IPC sequence representations) as input, process these sequence representations in one or more transformer stages, and generate transformed sequence representations (e.g., a set of transformed amino acid sequence representations, transformed IPC sequence representations) to be sent to the synthesis subsystem 306. The processing subsystem 304 includes one or more processing blocks. The processing blocks may include one or more processing layers (e.g., attention layers). The various transformer stages of the processing subsystem 304 may thus store information about each sequence into an embedding corresponding to the BOS token within the sequence representation.
[0153] In some embodiments, the representation subsystem 302 and / or the processing subsystem 304 may be configured to process subsequences of the amino acid sequence in parallel using separate processing engines. For example, Figure 1 the peptide processing engine 136 in may include (1) the representation subsystem 302, which processes the peptide sequence 126 appended with a BOS token to generate a BOS+(plus) peptide sequence representation 312, followed by (2) a processing block 314 in the processing subsystem 304, which processes the BOS+ peptide sequence representation 312 to generate a transformed peptide sequence representation 316. In this example, Figure 1 The N-flank processing engine 138 in Figure 1 can be executed independently and in parallel by: (1) a representation subsystem 302 that processes the N-flank sequence 128 appended with a BOS token to generate a BOS+N-flank sequence representation 324, followed by (2) a processing block 326 in a processing subsystem 304 that processes the BOS+N-flank sequence representation 324 to generate a transformed N-flank sequence representation 328. In this example, Figure 1 The C-flank processing engine 140 in Figure 1 can be executed independently and in parallel by: (1) a representation subsystem 302 that processes the C-flank sequence 130 appended with a BOS token to generate a BOS+C-flank sequence representation 330, followed by (2) a processing block 332 in a processing subsystem 304 that processes the BOS+C-flank sequence representation 330 to generate a transformed C-flank sequence representation 334.
[0154] Similarly, the representation subsystem 302 and the processing subsystem 304 can be configured to process subsequences of IPC sequences (e.g., MHC sequences, TCR sequences) in parallel using independent processing engines. For example, Figure 1 The MHC processing engine 141 in Figure 1 can include: (1) a representation subsystem 302 that processes the MHC sequence 135 appended with a BOS token to generate a BOS+MHC sequence representation 318, followed by (2) a processing block 320 in a processing subsystem 304 that processes the BOS+MHC sequence representation 318 to generate a transformed MHC sequence representation 322. In this example, Figure 1 The TCR processing engine 142 in Figure 1 can be executed independently and in parallel by: (1) a representation subsystem 302 that processes the TCR sequence 131 appended with a BOS token to generate a BOS+TCR sequence representation 336, followed by (2) a processing block 338 in a processing subsystem 304 that processes the BOS+TCR sequence representation 336 to generate a transformed TCR sequence representation 340.
[0155] In some embodiments, certain subsequence representations generated by the representation subsystem 302 can be combined before appending a BOS token. For example, the N-flank sequence representation can be appended with the C-flank representation before appending a single BOS token to the combined sequence representation. In such embodiments, the processing subsystem 304 that processes the combined sequence representation can correspondingly generate a single transformed sequence representation.
[0156] The synthesis subsystem 306 generates a synthetic representation 342 by combining the embedding of the BOS token corresponding to the transformed amino acid sequence representation with the embeddings of the BOS tokens corresponding to each of the set of transformed IPC sequence representations. The synthesis subsystem 306 may combine the embedding of the BOS token corresponding to the transformed peptide sequence representation 316 with a set of embeddings of the BOS tokens corresponding to the transformed MHC sequence representation 322 to generate the synthetic representation 342. In some embodiments, the combining method may include element-wise multiplying the transformed peptide BOS token embedding with the set of transformed MHC BOS token embeddings. Element-wise multiplication is beneficial because it forces the latent spaces of the individual components to converge into complementary regions of the latent space. As a result, peptides that bind to a particular MHC cause embeddings in the same region (similar to the effect of TCR and peptide sequences). In some embodiments, the combining method may include element-wise adding the transformed peptide BOS token embedding to the set of transformed MHC BOS token embeddings, or calculating their dot product.
[0157] Embodiments of the present disclosure may include a set of transformed amino acid sequence representations, the set of transformed amino acid sequence representations including one or more of the following: a set of transformed N-flanking sequence representations 328, a set of transformed peptide sequence representations 316, or a set of transformed C-flanking sequence representations 334. In some embodiments, the set of transformed IPC sequence representations may include one or more of the following: a set of transformed MHC sequence representations 322 or a set of transformed TCR sequence representations 340. The synthesis subsystem 306 combines the embedding of the BOS token of the transformed amino acid sequence representation with one or more (including all) of the embeddings of the BOS tokens represented by the transformed IPC sequences. The combining step may include multiplying or adding the BOS token embeddings, such as element-wise multiplying the transformed amino acid BOS token embedding (including one or more of the following: the transformed peptide BOS token embedding, the transformed N-flanking BOS token embedding, or the transformed C-flanking BOS token embedding) with the transformed IPC BOS token embedding (including one or more of the following: the transformed MHC BOS token embedding and the transformed TCR BOS token embedding). In some embodiments, instead of or in addition to element-wise multiplication, the combining step may include dot product calculation and / or element-wise addition.
[0158] The synthetic representation 342 may be processed by the output subsystem 310, and a predicted amino acid-IPC interaction may be determined based on the synthetic representation 342. In some embodiments, the predicted amino acid-IPC interaction may be based on one or more synthetic representations selected from the synthetic representation 342. In some embodiments, a report 144 including the selected one or more predicted amino acid-IPC interactions may be generated.
[0159] Figures 4A to 4D An example of a workflow is shown that illustrates various possible combinations of peptide data processing, n-flank / c-flank data processing, MHC data processing, TCR data processing, and / or protein data processing for obtaining one or more interaction predictions, one or more interaction affinity predictions, and / or one or more immunogenicity predictions according to some embodiments. In addition, the predictions can include: an interaction affinity prediction for a peptide-IPC combination that predicts the binding affinity between a peptide and an MHC molecule; an interaction prediction for a peptide-IPC combination, e.g., predicting whether an MHC molecule will present a peptide at the cell surface; or an immunogenicity prediction for a peptide-IPC combination that predicts the ability of a peptide to elicit an immune response. As detailed below, the processing of MHC data can involve processing an MHC sequence appended with a BOS token using one or more transformer stages (e.g., Figure 4A ), or alternatively involve processing an MHC sequence embedding generated using a protein language model (e.g., Figure 4B 、 4C 、4D). The processing of peptide data can involve processing a peptide sequence appended with a BOS token using one or more transformer stages (e.g., Figure 4A and 4B ), or alternatively involve processing a peptide sequence not appended with a BOS token using a cross-attention module (e.g., Figure 4C and 4D ). Optionally, any workflow can additionally incorporate the processing of protein data (e.g., Figure 4B ), or alternatively may not incorporate the processing of protein data (e.g., Figure 4A 、 4C 、4D). Optionally, any workflow can additionally incorporate the processing of TCR data (e.g., Figure 4A ), or alternatively may not incorporate TCR data (e.g., Figure 4B 、 4C 、4D).
[0160] It should be understood that Figure 4A 、 4B 、4C and 4D are only examples, and the present disclosure encompasses any workflow involving a combination of the processing of MHC data and the processing of peptide data for obtaining one or more interaction predictions, one or more interaction affinity predictions, and / or one or more immunogenicity predictions. In one exemplary workflow, the processing of MHC data can involve processing an MHC sequence appended with a BOS token using one or more transformer stages (e.g., Figure 4A ), or alternatively involve processing an MHC sequence embedding generated using a protein language model (e.g., Figure 4B 、4C , 4D). In one exemplary workflow, the processing of peptide data may involve processing a peptide sequence appended with a BOS token using one or more transducer stages (e.g., Figure 4A and 4B ), or alternatively may involve processing a peptide sequence without an appended BOS token using a cross-attention module (e.g., Figure 4C and 4D ). The exemplary workflow may additionally incorporate the processing of protein data (e.g., Figure 4B ), or alternatively may not incorporate the processing of protein data (e.g., Figure 4A , 4C , 4D). The exemplary workflow may additionally incorporate the processing of TCR data (e.g., Figure 4A ), or alternatively may not incorporate TCR data (e.g., Figure 4B , 4C , 4D).
[0161] Figure 4A FIG. 400 shows an exemplary workflow for obtaining one or more interaction predictions, one or more interaction affinity predictions, and / or one or more immunogenicity predictions according to some embodiments. A tokenizer (not depicted) may create a tokenized sequence that includes a peptide sequence 402a appended with a BOS token, an n-flank + c-flank sequence 402b appended with a BOS token, an MHC sequence 402c appended with a BOS token, and a TCR sequence 402d appended with a BOS token. An embedding module may generate sequence representations. Specifically, the embedding module may generate (e.g., via the representation subsystem 302): a BOS + peptide sequence representation 404a (which may correspond to Figure 3 312 in Figure 3 ) (based on the peptide sequence 402a appended with a BOS token); a BOS + n-flank + c-flank sequence representation 404b, which represents a combined version of the BOS + N-flank sequence representation (which may correspond to Figure 3 324 in Figure 3(318) in (based on the MHC sequence 402c appended with the BOS token); and the BOS+TCR sequence representation 404d (which may correspond to 336) (based on the TCR sequence 402d appended with the BOS token). The converted versions 406a-d of each of the sequence representations 404a-d (e.g., the converted BOS+peptide sequence representation 316 and 406a, the converted BOS+n-flank+c-flank sequence representation 406b (which represents a combined version of the BOS+N-flank sequence representation 328 and the BOS+C-flank sequence representation 334), the converted BOS+MHC sequence representation 322 and 406c, or the converted BOS+TCR sequence representation 340 and 406d) can be generated using one or more converter stages in the processing subsystem 304. During the converter stages, BOS token embeddings 408a-d are also generated as part of the converted sequence representations 406a-d (e.g., the BOS token embedding 408a of the converted BOS+peptide sequence representation 316 and 406a, the BOS token embedding 408b of the converted BOS+n-flank+c-flank sequence representation 406b, the BOS token embedding 408c of the converted BOS+MHC sequence representation 322 and 406c, or the BOS token embedding 408d of the converted BOS+TCR sequence representation 340 and 406d). The converted BOS token embeddings 408a-d extract information about the sequence (e.g., information about pairwise correlations and positions) into the embeddings corresponding to the BOS tokens appended to the sequence. Each of the converted BOS token embeddings 408a-d can represent the entire sequence via a single vector rather than multiple vectors, making the sequence easier to interpret. Then, as Figure 3 described, complexes of the BOS token embeddings 408a-d can be generated by the synthesis subsystem 306, and the final output generated by the output subsystem 310 is also described in Figure 3 . Each of the sequence representations (e.g., embeddings) 404a-d and the converted sequence representations 406a-d can have uniform dimensions, including subsequence length (SL), vector length (VL), and batch (BS). As Figure 4A shown, the dimension of the BOS token embeddings 408a-d has a subsequence length equal to 1 because they are generated from a single BOS token appended to each amino acid or IPC subsequence.
[0162] As Figure 4A shown, in the peptide processing engine 136 (see Figure 1) In it, the peptide sequence 402a appended with the BOS token is formed by a tokenizer (e.g., in the representation subsystem 302), and then used by an embedding module (e.g., in the representation subsystem 302) to generate a peptide sequence representation 404a. A number of transformer stages generate a transformed peptide sequence representation 406a (including the transformed BOS token embedding 408a) from this peptide sequence representation. In a parallel processing engine (e.g., combining the N-flank processing engine 138 and the C-flank processing engine 140), the N-flank and C-flank sequences can be appended together with a single BOS token to form a combined flank sequence 402b. After generating a flank sequence representation 404b through the embedding module, a transformed flank sequence representation 406b (including the transformed BOS token embedding 408b) is generated by the processing subsystem 304. As shown, first, the transformed BOS token 408a is combined with the transformed BOS token embedding 408b, and then the combined result is combined with the transformed BOS token 408c and the transformed BOS token embedding 408d. In each of these two combination operations, element-wise addition, element-wise multiplication, concatenation, or dot product multiplication can be used.
[0163] In parallel, in Figure 1 the MHC processing engine 141, the MHC sequence 402c appended with the BOS token (e.g., the MHC sequence 135 of the IPC sequence 124, corresponding to an allele of MHC class I, or an allotype corresponding to MHC class II) is used to generate an MHC sequence representation 404c. During the transformation stage, then a transformed MHC sequence representation 406c (including the transformed BOS token embedding 408c) is generated. Additionally, in parallel, the TCR sequence 402d appended with the BOS token (e.g., the TCR sequence 131 of the IPC sequence 124) can be used to generate a TCR sequence representation 404d, and a transformed TCR sequence representation 406d (including the transformed BOS token embedding 408d) is generated from this TCR sequence representation.
[0164] Then, the transformed BOS token embeddings can be combined together (e.g., by concatenating 408a and 408b, and applying element-wise multiplication to the concatenated 408a + 408b, 408c, and 408d) through Figure 3 the synthetization subsystem 306 in, and the product vector is passed through a multi-layer perceptron network (e.g., the output subsystem 310) to provide a binary prediction for binding affinity, presentation likelihood, immunogenicity prediction, or other downstream evaluations.
[0165] Figure 4B Shows an exemplary workflow 420 for obtaining one or more interaction predictions, one or more interaction affinity predictions, and / or one or more immunogenicity predictions according to some embodiments. Similar to Figure 4A In the workflow 400, a tokenizer (not depicted) can create a tokenized sequence that includes a peptide sequence 402a appended with a BOS token and an n-flank + c-flank sequence 402b appended with a BOS token. An embedding module (not depicted) can generate (e.g., via the Figure 3 representation subsystem 302 in): a BOS+peptide sequence representation 404a (which can correspond to Figure 3 312 in Figure 3 )(based on the peptide sequence 402a appended with a BOS token); and a BOS+n-flank + c-flank sequence representation 404b, which represents a combined version of the BOS+N-flank sequence representation (which can correspond to Figure 3 324 in
[0166] Figure 4A The workflow 400 in does not incorporate the processing of protein data (e.g., the source protein 123 in Figure 1 ), from which peptides (e.g., the peptide 118 that can be represented by 402a in Figure 1 ) and N-flank / C-flank (e.g., the N-flank 120 and C-flank 122 that can be represented by 402b) are identified. As discussed above with reference to Figure 1 , the protein 123 is the source protein of the amino acid chain 116 (including the peptide 118 in Figure 1 , the N-flank 120 in Figure 1 , the C-flank 122 in Figure 1 ), which can be generated by proteolysis, which is the process of breaking down a protein (e.g., the protein 123) into smaller polypeptides or amino acids. The protein 123 can be broken down into smaller polypeptides or amino acids by enzymatic cleavage, where a specific enzyme called a protease cleaves the peptide bonds between amino acids in the protein 123. In contrast, Figure 4B the workflow 420 in can further include a protein processing engine (e.g., the protein processing engine 162 in Figure 1 ) to generate and process a protein sequence embedding 422 that encapsulates information about the protein. As in Figure 4BAs shown, the dimensionality reduction module 424 can receive the protein sequence embedding 422 to reduce the dimensionality of the protein sequence embedding 422 and generate an output vector. As discussed below with reference to Figure 4E the protein sequence embedding 422 can be generated by a protein language model (PLM) based on a protein sequence (e.g., Figure 1 the protein sequence 160 in).
[0167] In some instances, the dimensionality reduction module 424 can be implemented as a fully connected layer ("FCN"). The FCN can include a neural network where each neuron applies a transformation (e.g., a linear transformation) to the input vector via a weight matrix. As a result, there are all possible layer-to-layer connections, and thus each input of the input vector (i.e., the protein sequence embedding 422) affects each output of the output vector. In addition to reducing the dimensionality of the input vector, the FCN can include a relatively large number of learnable parameters, enabling the neural network to encode useful information into the output vector. The output vector (i.e., the dimensionality-reduced version of the protein sequence embedding 422) can be further aggregated with 408a and 408b. The dimensionality reduction module 424 can encode a mapping from protein features to useful features such as subcellular compartmentalization and gene expression. Although Figure 4B element-wise addition is shown to aggregate the output vectors of 408a, 408b, and the dimensionality reduction module 424, other integration methods such as element-wise averaging, element-wise multiplication, concatenation, and dot product multiplication can also be used.
[0168] The incorporation of protein information advantageously allows for the incorporation of useful protein features. The workflow 420 can select peptides that are more likely to be presented by MHC molecules because the workflow can capture features such as: processing signals (i.e., flanking signals generated when an enzyme breaks down a protein into peptides), the expression of the gene associated with the source protein, the cellular localization of the source protein, and other suitable features. For example, for MHC I, it is preferred that the source protein is intracellular, while for MHC II, it is preferred that the protein is vesicular or extracellular. In summary, there are many features associated with the source protein that affect peptide presentation, and these features are advantageously incorporated into the workflow 420. In addition, using a PLM may be advantageous compared to using a Transformer. A typical Transformer can have a large number of parameters. Therefore, training a Transformer using a relatively small training dataset may lead to overfitting. In contrast, using a PLM can allow the workflow to learn useful generalizable features, thereby reducing the likelihood of overfitting and improving the performance of the workflow.
[0169] In some embodiments, the system can enable or disable the incorporation of protein data processing (i.e., Figure 1The protein processing engine 162) in depends on the use case. For example, when the system is used to develop a personalized cancer vaccine and the subject generates peptides in a manner similar to the peptides in the generated presentation dataset, the protein processing engine can be enabled. In contrast, when the system is used to develop antibody drugs, the protein processing engine can be disabled when the PLM provides information about endogenous proteins and the antibody drugs are not endogenous.
[0170] The workflow 420 also differs in the processing of MHC data from Figure 4A the workflow 400 in. As described above, the workflow 400 involves using the MHC sequence 402c appended with the BOS token to generate the BOS+MHC sequence representation 404c, and then using one or more transformer stages to generate the transformed BOS+MHC sequence representation. The workflow 400 further involves obtaining the BOS token embedding 408c and aggregating the BOS token embedding 408c with the remaining BOS embedding tokens. In contrast, in Figure 4B the workflow 420 in, the dimensionality reduction module 428 can receive the MHC sequence embedding 426 to reduce the dimensionality of the MHC sequence embedding 426 and generate an output vector. As discussed below with reference to Figure 4F , the PLM can be used to generate the MHC sequence embedding 426 based on the MHC sequence (e.g., Figure 1 the MHC sequence 135 in).
[0171] The goal of the dimensionality reduction module 428 is to reduce the number of parameters used to represent information about MHC sequences. The dimensionality reduction module 428 removes or down-prioritizes less useful parameters (such as parameters that are invariant between one MHC allele and another MHC allele (e.g., having a variance below a certain threshold)). For example, similar binding behaviors may be removed or down-prioritized while different binding behaviors are retained. In some embodiments, the dimensionality reduction module 428 may be implemented as a principal component analysis (PCA) model. The PCA model may be configured to receive an MHC sequence embedding 426 having N vector values and generate a reduced-dimensional version of the MHC sequence embedding 426 having M vector values (M < N). In some embodiments, to configure the PCA model, the PCA model first receives a plurality of MHC vectors corresponding to a plurality of MHC sequences (e.g., each having N vector values corresponding to N features), and sorts the N features based on how each feature varies across the plurality of MHC vectors (e.g., based on the variance associated with each feature). After the PCA model determines the sorting of the N features, the PCA model may then receive an input MHC sequence embedding 426 having N vector values, reorder the N vector values in the MHC sequence embedding 426 according to the sorting of the N features, and generate a reduced-dimensional version of the MHC sequence embedding 426 having M vector values (M < N) by, for example, retaining only the first M vector values of the reordered N vector values. It should be understood that the dimensionality reduction module 428 may use other dimensionality reduction methods, such as UMAP, t-distributed stochastic neighbor embedding, independent component analysis, multidimensional scaling, Isomap, and deep learning-based dimensionality reduction techniques.
[0172] Using the dimensionality reduction module 428 to further process the MHC sequence embedding 426 provides many technical advantages. For example, if different types of models (such as neural networks) are trained to reduce the dimension of the input embedding, overfitting may occur. Overfitting is an undesirable machine learning behavior that occurs when a machine learning model provides accurate outputs for the training data but not for new data. If a neural network is trained using the same set of alleles (usually a relatively small set of alleles) used to train the PLM model, the neural network model may generate correct outputs when processing input embeddings corresponding to the alleles in the training data, but generate incorrect outputs when processing input embeddings corresponding to new alleles not included in the training data. On the other hand, the PCA technique can be trained using a larger set of alleles (e.g., the space of all alleles) and can therefore efficiently and accurately reduce the dimension of input embeddings for new alleles not in the training dataset, thus avoiding the overfitting problem.
[0173] Reference Figure 4B, optionally, the output of the dimensionality reduction module 428 (i.e., the dimensionality-reduced version of the MHC sequence embedding 426) can be provided to the FCN module 429. The FCN module 429 can include a neural network where each neuron applies a transformation (e.g., a linear transformation) to the input vector via a weight matrix. As a result, there are all possible layer-to-layer connections, and thus each input of the input vector (i.e., the protein sequence embedding 422) affects each output of the output vector. The FCN module 429 can further reduce the dimensionality of the output vector of the dimensionality reduction module 428. Additionally, the FCN module 429 can encode the relationship between the MHC sequence and the type of peptide presented by the MHC sequence. Additionally, the FCN module 429 can encode information that can be used for downstream processing. For example, due to the similarity of sequence data, two MHC sequence embeddings represented by similar sequences may be close to each other in the latent space, but in fact the functions of these two MHCs may be different (e.g., presenting different peptides). The FCN module 429 can adjust for the differences in the latent space such that the output vector is more useful for downstream processing (e.g., since the FCN module 429 allows encoding of information about how similar peptides and similar MHCs present each other). The output vector of the FCN module 429 can be aggregated with the output 410, as Figure 4B shown. Although Figure 4B element-wise multiplication is shown, other integration methods can be used, such as element-wise averaging, element-wise addition, concatenation, and dot product multiplication.
[0174] Figure 4C shows an example workflow 470 for predicting the interaction of a peptide with one or more MHC molecules expressed by one or more alleles or allotypes. Similar to Figure 4A the workflow 400 in, a tokenizer (not depicted) can create an n-flank + c-flank sequence 402b appended with a BOS token. An embedding module (not depicted) can generate (e.g., via Figure 3 the representation subsystem 302 in): a BOS + n-flank + c-flank sequence representation 404b (based on the n-flank + c-flank sequence 402b appended with a BOS token). Then, one or more transformer stages of the system (e.g., in Figure 3 the processing subsystem 304 in) can create a transformed BOS + n-flank + c-flank sequence representation 406b. During the transformer stage, the system can further generate a BOS token embedding 408b of the transformed BOS + n-flank + c-flank sequence representation 406b.
[0175] Unlike Figure 4A the workflow 400 in, the workflow 470 does not receive a peptide sequence appended with a BOS token. Instead, the workflow 470 receives a peptide sequence not appended with a BOS token (e.g., Figure 1the peptide sequence 126), and create a tokenized peptide sequence 402e based on the peptide sequence. The embedding module can generate (e.g., via Figure 3 the representation subsystem 302 in) a peptide sequence representation 404e (based on the tokenized peptide sequence 402e). Then, one or more transformer stages of the system (e.g., in Figure 3 the processing subsystem 304 in) can create a transformed peptide sequence representation 406e.
[0176] Further, the workflow 470 receives a BOS vector embedding 472 (which can be generated using a PLM model based on the BOS vector) and an MHC sequence embedding 426. The BOS vector embedding 472 can be a random vector selected during training, e.g., representing an intercept for one or more random bias terms. The dimensionality reduction component 428 can receive the MHC sequence embedding 426 and obtain an output vector, as discussed above with reference to Figure 4B The output vector of the dimensionality reduction component 428 (e.g., with parameters having lower variance removed or deprioritized) can be provided to the FCN module 473, which can further reduce the dimensionality of the vector. At the aggregator 471, the workflow 470 aggregates the BOS vector embedding 472 with the output vector of the FCN module 473. Although Figure 4C element-wise addition at the aggregator 471 is shown, other integration methods can be used, such as element-wise averaging, element-wise multiplication, concatenation, and dot product multiplication. Thus, the output vector of the aggregator 471 encodes information from the BOS vector and the MHC sequence.
[0177] The workflow 470 further includes a cross-attention module 474. The cross-attention module 474 can be implemented as a self-attention transformer with three components: query (Q), key (K), and value (V). As Figure 4CAs shown, both the K component of the cross-attention module 474 and the V component of the cross-attention module 474 are from the transformed peptide sequence representation 406e (i.e., implementing the self-attention mechanism). The Q component of the cross-attention module 474 is from the output vector of the aggregator 471 (i.e., the combined vector of the BOS vector embedding and the MHC embedding). The cross-attention module 474 can execute an attention function that involves mapping a query and a set of key-value pairs to an output, where the query, key, value, and output are all vectors. The output is calculated as a weighted sum of the values, where the weight assigned to each value is calculated by using a compatibility function of the query with the corresponding key. Specifically, the cross-attention module 474 can execute scaled dot-product attention or other suitable types of integration techniques. For more details on scaled dot-product attention, see Ashish Vaswani et al., Attention is All You Need, Neural Information Processing Systems (2017). The configuration of the workflow 470 is advantageous because it incorporates allele-specific binding information. Allele information can be valuable when processing peptide data. For example, a peptide may contain multiple binding cores that can bind to different alleles. The workflow 470 uses the cross-attention module 474 to incorporate allele-specific vectors (i.e., component Q) into the processing of peptide data. In this way, the workflow can capture binding core information from peptides with multiple binding cores that can bind to different alleles and generate multiple binding core predictions corresponding to the multiple different alleles. In some instances, a peptide may have more than one element focus score, each corresponding to a different allele.
[0178] Figure 4D An alternative embodiment 476 of the workflow 470 is shown. As shown, the workflow 476 eliminates the processing of the BOS vector embedding 477 in the workflow 470 and also eliminates the aggregator 475 because the allele information (i.e., the MHC sequence embedding) has been taken into account by the cross-attention module 474. The workflow 476 has similar technical advantages to the workflow 470 but may be computationally more efficient due to the elimination of modules. Further, Figure 4D the workflow 476 in [reference] can provide an allele-dependent latent space.
[0179] Figure 4E An exemplary method for generating protein sequence embeddings according to some embodiments is depicted. Refer to Figure 4E, the PLM 444 can receive a protein sequence 160 and output a protein sequence embedding 422. The PLM is a machine learning model (e.g., a deep learning model) that can be based on natural language processing methods such as attention and transformers and can be trained on the overall protein sequences. Protein language models are trained to understand and predict the properties of proteins based on the amino acid sequences that form such proteins. In some cases, protein language models can infer a range of features from the amino acid sequences, including primary, secondary, tertiary, and quaternary structures. The PLM can predict how proteins fold, their domains, active sites, and stability. The PLM can also predict protein-protein and protein-nucleic acid interactions, the effects of post-translational modifications, and mutations. The PLM can identify localization signals within cells, understand evolutionary relationships, and predict protein functions. Similarly, the PLM can provide insights into the dynamic behavior of proteins and identify potential drug-binding sites, which are valuable for drug discovery and understanding the molecular basis of diseases. In some embodiments, the PLM 444 can include an evolutionary scale modeling (ESM model) or a variant of the ESM model. In some embodiments, the PLM 444 can include ProteinBERT, UniRep, or other suitable types of PLMs.
[0180] In some embodiments, the PLM 444 includes a pre-trained protein language model, such as a pre-trained ESM model. In some embodiments, the input protein sequence 160 can include a sequence of amino acid residues, and the PLM 444 can be configured to obtain a plurality of embeddings (i.e., vector representations) by obtaining a corresponding embedding for each amino acid residue. The model can further be configured to obtain a single embedding 422 by aggregating the plurality of embeddings corresponding to the amino acid residue sequence (e.g., by element-wise averaging).
[0181] Figure 4F An exemplary method for generating MHC sequence embeddings according to some embodiments is depicted. Refer to Figure 4F, the PLM 445 can receive the MHC sequence 135 and output the MHC sequence embedding 426. As discussed herein, the PLM is trained using a large number of proteins and can thus encode useful information and context about the input sequence represented by the sequence embedding. In some embodiments, the PLM 445 can include an Evolutionary Scale Modeling (ESM model) or a variant of the ESM model. In some embodiments, the PLM 445 can include ProteinBERT, UniRep, etc. In some embodiments, the PLM 445 includes a pre-trained protein language model, such as a pre-trained ESM model. In some embodiments, the MHC sequence 135 can include multiple amino acids that make up the corresponding allele. The PLM 445 can be configured to obtain multiple embeddings (i.e., vector representations) by obtaining a corresponding embedding for each amino acid. The model can further be configured to obtain a single embedding 426 by aggregating the multiple embeddings corresponding to the amino acid sequence (e.g., by element-wise averaging).
[0182] Figure 4G Exemplary workflow diagram 450 for more efficiently predicting the interaction of a peptide with an MHC molecule that may have been expressed by multiple alleles or allotypes according to some embodiments. This technique can be applied to any workflow described herein. When the binding or elution likelihood of a peptide is known, this technique can advantageously predict the interaction of the peptide with the MHC molecule, but data is collected for multiple alleles or allotypes, and it is not known exactly which allele or allotype among the multiple alleles or allotypes binds to the peptide. As Figure 4G shown, a set of alleles / allotypes 452 for which data is collected (e.g., HLA 1, HLA 2, HLA 3, and HLA 4) have been tokenized and embedded (e.g., using the representation subsystem 302) to create an MHC sequence representation 404c for each allele / allotype. Since the data for each sample for each allele / allotype may be sparse, steps can be taken (e.g., by the representation subsystem 302) to compress the MHC sequence representation. The MHC sequence representations for the detected binding interactions can be flattened into a single array 454, from which blank rows can be removed, resulting in a dense combined MHC sequence representation 456, which is then processed by the processing subsystem 304 as usual to generate a binding affinity prediction 458 (e.g., in both the model training and inference phases). Once the prediction 458 has been generated, the model output can be re-sparsified (e.g., embedding 460) for the needs of downstream tasks.
[0183] Figure 4HExemplary workflow 480 for an attention mask at each transformer stage of the converted BOS+ peptide sequence representation 406a. This technique can be applied to any workflow described herein that involves processing of peptide sequences appended with the BOS token. The transformer stages (e.g., as performed by processing subsystem 314 and processing sub-blocks 542, 594b, and 602) store pairwise information in the attention map. To force the model to focus only on a peptide binding core sequence of a specified length (e.g., nine amino acids long, as shown in attention mask 482, or some other core length), masks 482 and 484 are applied to the attention map to limit the range of consecutive amino acids (e.g., consider at most the nine positions preceding in the sequence, or some other core length), within which the model can record information about pairwise correlations between amino acids. At the final transformer stage, the final attention map 484 can limit the model to focus only on the BOS token 408a of the converted BOS+ peptide sequence representation 406a and record the maximum value as the start of the binding core.
[0184] Fig. 4I Illustrates exemplary attention maps used at the transformer stage. Each attention map visualizes the attention weights assigned to different parts of the input by the transformer stage. In Fig. 4I it, the attention maps are represented as heatmaps, where lighter colors correspond to higher attention weights. Specifically, Fig. 4I illustrates an exemplary attention map 483 used at the transformer stage to obtain the converted BOS+ peptide sequence representation 406a of a given peptide, and an exemplary attention map 485 applied to obtain the BOS token 408a of a given peptide. As discussed above with reference to Figure 4H the attention mask 482 has been applied to the attention map 483 to force the model to focus only on a peptide binding core sequence of a specified length (e.g., nine amino acids long, as shown in attention mask 482, or some other core length), and the attention mask 484 has been applied to the attention map 485.
[0185] Generally, position zero (i.e., the starting position) of a peptide is unlikely to be the binding core starting position, as it is expected that the binding core is part of a peptide (e.g., a 20-mer) (e.g., a 9-mer) and is flanked by the binding core within the peptide. However, the attention mechanism in workflow 480 may lead to over-prediction of position zero in the peptide as the binding core starting position for certain datasets and / or certain workflows. To address this over-prediction problem, workflow 480 may include a calibration step to remove the bias towards position zero. Before performing the calibration step, the system (e.g., Figure 1The computing platform 102) in can first obtain a set of random peptides with a uniform length distribution. For each given peptide length, the system can calculate the average attention value at each position in the set of random peptides (e.g., position zero, position one, position two, etc.) (e.g., according to Figure 4A to the transformed BOS+ peptide sequence representation 406a in B, Figures 4C to 4D the cross-attention module 474) in. In other words, for each given peptide length N, the system can calculate the average attention value for each position from 0 to N-1 (i.e., the average attention value for position zero, the average attention value for position one,..., and the average attention value for position N-1). In the range where these average attention values are different, the average attention value represents a model bias because the probability of the binding core starting from any position should be equal, and thus the attention values should be uniformly distributed (i.e., the average attention values should be the same).
[0186] Return Fig. 4I , in the workflow 480, the system can perform a calibration step by modifying the attention map 485. For a given peptide with length N, the system can subtract the average attention value for the N positions (i.e., the average attention value for position zero, the average attention value for position one,..., and the average attention value for position N-1) from the N positions in the attention map 485. After the subtraction, the modified attention map 485 can be applied to obtain the BOS token 408a for the given peptide in the workflow 480. As described above, the average attention value represents a model bias, and by subtracting the average attention value from the attention map, the calibration step removes the model bias towards any single position in the peptide.
[0187] Figures 5A to 5C is a schematic diagram of different configurations of a machine learning model 532 according to some embodiments. The machine learning model 532 is for Figure 1 and Figure 3 the machine learning model 132 in and Figure 4A an example of the implementation of the workflow 400 in. The machine learning model 532 can be any type of machine learning model, including but not limited to an attention-based machine learning model. As shown in the schematic diagram of Figure 5A , the machine learning model 532 includes a representation subsystem 501 (e.g., an embedding module that generates the sequence representation 404), a processing subsystem 503 (e.g., one or more transformer stages that generate the transformed sequence representation 406), a synthesis subsystem 505 (e.g., which combines the BOS token embedding 408 of the transformed sequence representation 406), and an output subsystem 509, which are respectively Figure 3 examples of the implementation of the representation subsystem 302, the processing subsystem 304, the synthesis subsystem 306, and the output subsystem 310 in.
[0188] The representation subsystem 501 may include a BOS+ peptide representation block 502 and a BOS+ MHC representation block 504. In some embodiments, the representation subsystem 501 further includes a BOS+ N-flank representation block 506, a BOS+ C-flank representation block 508, or both. In some embodiments, the representation subsystem 501 further includes a BOS+ TCR representation block 510. One or more (e.g., each) representation blocks include at least one embedding layer (e.g., embedding layer 512, embedding layer 516, embedding layer 520, embedding layer 524, or embedding layer 528), and may include, for example, at least one positional encoder (e.g., positional encoder 514, positional encoder 518, positional encoder 522, positional encoder 526, or positional encoder 530).
[0189] The embedding layer may embed a sequence, for example, by converting an initial non-numerical sequence representation (e.g., a string of amino acid identifiers) into a numerical sequence representation to generate an embedded representation. In some embodiments, the embedded amino acid sequence indicates whether a particular amino acid is present at each position of the sequence and for each of a set (e.g., 21) of amino acids. The embedding may be performed using, for example, one-hot encoding, evolution-driven encoding (such as BLOSUM), randomly or pseudo-randomly initialized learned embeddings, or a combination thereof. The embedded representation may be positionally encoded to generate an encoded representation. The representation generated by the representation block may be an encoded representation, or an aggregation (e.g., concatenation or summation) of the encoded representation and the embedded representation.
[0190] In some cases, the order of the values in the input dataset may be useful. A positional encoder may be used and added to the embedded representation, where the positional encoding uses a learned or fixed encoding algorithm. For example, sine and / or cosine functions (e.g., having the in-sequence position and / or dimension as arguments) may be used to determine the fixed positional encoding. The positional encoding may have the same dimension as the encoded representation. The positional encoding may be summed with the embedded representation to generate a position-indicating embedded representation of the sequence that is fed into the processing subsystem 503.
[0191] For example, the BOS+ peptide representation block 502 may include an embedding layer 512 and a positional encoder 514. The embedding layer 512 embeds a peptide sequence (e.g., Figure 1 the peptide sequence 126 in Figure 3 to generate an embedded peptide representation, and the positional encoder 514 positionally encodes the embedded peptide representation to generate a peptide sequence representation (e.g., Figure 1The N-flanking sequence 128) in to generate an embedded N-flanking representation, and the positional encoder 522 performs positional encoding on the embedded N-flanking representation to generate an N-flanking sequence representation (e.g., Figure 3 The N-flanking sequence representation 324) in. The BOS + C-flanking representation block 508 may include an embedding layer 524 and a positional encoder 526. The embedding layer 524 embeds the C-flanking sequence (e.g., Figure 1 The C-flanking sequence 130) in to generate an embedded C-flanking representation, and the positional encoder 526 performs positional encoding on the embedded C-flanking representation to generate a C-flanking sequence representation (e.g., Figure 3 The C-flanking sequence representation 330) in.
[0192] The BOS + MHC representation block 504 may include an embedding layer 516 and a positional encoder 518. The MHC sequence of the embedding layer 516 (e.g., Figure 1 The MHC sequence 135) in to generate an embedded MHC representation, and the positional encoder 518 performs positional encoding on the embedded MHC representation to generate an MHC sequence representation (e.g., Figure 3 The MHC sequence representation 318) in. The BOS + TCR representation block 510 may include an embedding layer 528 and a positional encoder 530. The TCR sequence of the embedding layer 528 (e.g., Figure 1 The TCR sequence 131) in to generate an embedded TCR representation, and the positional encoder 530 performs positional encoding on the embedded TCR representation to generate a TCR sequence representation (e.g., Figure 3 The TCR sequence representation 336) in.
[0193] The sequence representations generated by the representation subsystem 501 are sent as inputs to the processing subsystem 503 for processing. In some embodiments, the sequence representations input to the processing subsystem 503 may include embeddings corresponding to additional BOS tokens. For example, the peptide sequence representation may include an embedding of the BOS token appended to the peptide sequence (e.g., Figure 3 The BOS + peptide sequence representation 308) in, and the MHC sequence representation may include an embedding of the BOS token appended to the MHC sequence (e.g., Figure 3 The BOS + MHC sequence BOS representation 318) in.
[0194] The processing subsystem 503 may include various mechanisms for determining element focus scores for each of one or more (e.g., all) positions in the sequence representation. The element focus scores may indicate an attention or importance level. For example, the element focus scores of a set of amino acid sequence representations may indicate the position where the binding core of the peptide begins. Then, the transformed values for the positions may be generated using the element focus scores.
[0195] The processing subsystem 503 includes processing block 532 and processing block 534. In some embodiments, the processing subsystem 501 may include processing block 536, processing block 538, processing block 540, or a combination thereof. Processing block 532 receives a peptide sequence representation from the BOS+ peptide representation block 502 and processes the peptide sequence representation using a set of processing sub-blocks 542 to generate a transformed peptide sequence representation (e.g., the transformed peptide sequence representation 316 in Figure 3 ). The transformed amino acid sequence representation may be generated based on the amino acid sequence representation and one or more element focus scores (representing the binding core of the amino acid sequence representation group). An exemplary implementation of the processing sub-blocks that perform one or more transformer stages to generate the transformed sequence representation is described in more detail below in the context of Figure 6 . Processing block 534 receives an MHC sequence representation from the BOS+ MHC representation block 504 and processes the MHC sequence representation using a set of processing sub-blocks 544 to generate a transformed MHC sequence representation (e.g., the transformed MHC sequence representation 322 in Figure 3 ). The transformed MHC sequence representation may be generated based on the MHC sequence representation and one or more element focus scores. In some embodiments, the element focus scores used to generate the transformed amino acid sequence representation may be different from the element focus scores used to generate the transformed MHC sequence representation.
[0196] Further, when included, processing block 536 receives an N-flank sequence representation from the BOS+ N-flank representation block 506 and processes the N-flank sequence representation using a set of processing sub-blocks 546 to generate a transformed N-flank sequence representation (e.g., the transformed N-flank sequence representation 328 in Figure 3 ). In some embodiments, processing block 538 receives a C-flank sequence representation from the BOS+ C-flank representation block 508 and processes the C-flank sequence representation using a set of processing sub-blocks 548 to generate a transformed C-flank sequence representation (e.g., the transformed C-flank sequence representation 334 in Figure 3 ). In some embodiments, processing block 540 receives a TCR sequence representation from the BOS+ TCR representation block 510 and processes the TCR sequence representation using a set of processing sub-blocks 550 to generate a transformed TCR sequence representation (e.g., the transformed TCR sequence representation 340 in Figure 3 ).
[0197] In some embodiments, one or more processing sub - blocks may process different parts or all of the representations of the amino acid sequence and / or the IPC sequence separately. In some embodiments, one or more sequence representations (e.g., N - flanking sequence representation, C - flanking sequence representation, peptide sequence representation, MHC sequence representation, TCR sequence representation) may be processed separately in different iterations of the processing sub - block. For example, the encoded representation of an amino acid sequence may include a feature vector representing an amino acid, and the encoded representation of the sequence (e.g., all or part of the amino acid sequence, all or part of the IPC sequence) may then be concatenated and fed into another iteration of the processing sub - block.
[0198] As Figure 5A shown, the amino acid sequence can be processed separately and independently from the processing of the IPC sequence. By having separate and independent processing engines 505 for peptide sequences, N - flanking sequences, C - flanking sequences, MHC sequences, and / or TCR sequences before the synthesis subsystem, the prediction performance of the machine learning model 532 can be enhanced. For example, using the BOS + peptide representation block 502 and the processing block 532 to generate a transformed peptide sequence representation along a path separate from the generation of the transformed IPC sequence representation, and doing so before generating the synthetic representation, improves the accuracy of the output generated by the output subsystem 509. Similarly, the prediction performance (e.g., accuracy) of the machine learning model 532 can be enhanced by using the BOS + N - flanking representation block 506 and the processing block 536 to generate a transformed N - flanking sequence representation along a separate path, using the BOS + C - flanking representation block 508 and the processing block 538 to generate a transformed C - flanking sequence representation along a separate path, using the BOS + MHC representation block 504 and the processing block 534 to generate a transformed MHC sequence representation along a separate path, using the BOS + TCR representation block 510 and the processing block 540 to generate a transformed TCR sequence representation using a separate path, or a combination thereof. In some embodiments, the separate paths can enable efficient processing (e.g., using reduced computational resources, faster processing, etc.) because multiple amino acid - IPC (peptide - MHC, peptide - TCR) combinations and / or parallel processing can be considered in a modular way.
[0199] The output of the transformed sequence representation from the processing subsystem 503 is sent into the synthesis subsystem 505 for processing. The synthesis subsystem 505 includes a synthesis block 552. The synthesis block 552 may use the output of the transformed representation from the processing subsystem 503 to form one or more synthetic representations (e.g., Figure 3 the synthetic representation 342 in). For example, the synthesis block 552 may multiply groups of transformed sequence representations (e.g., groups of amino acid sequence representations, groups of IPC sequence representations) to form a synthetic representation.
[0200] In some embodiments, the size of the output generated by the synthesis block 552 can be equal to, for example, m x n, where m is equal to the total number of amino acids under consideration plus 1 (e.g., for the BOS token) plus any padding to conform to the normalized sequence length (e.g., 39), and n is equal to the number of features (a predetermined value, e.g., 600). A single row (with n values) can be selected for further processing as the output of the synthesis block 552. This single row can be the first row and / or the row associated with the BOS token. In some embodiments, the output from the composite block 552 can be aggregated to form a single vector, which can then be fed to the output subsystem 509.
[0201] The output subsystem 509 can include various blocks, sub - blocks, layers, or combinations thereof for generating the final output. In some embodiments, the output subsystem 509 includes a dropout block 560, a fully - connected block 562, and an output block 564. The dropout block 560 can include, for example, one or more dropout layers. The fully - connected block 562 can include, for example, one or more fully - connected layers. The output block 564 can include, for example, one or more layers for filtering, selecting, transforming, or otherwise generating the result. For example, the output block 564 can include at least one max layer 565 configured to select a subgroup of the inputs received by the output block 564 based on, for example, a selected threshold or range.
[0202] In some cases, the synthetic representation is received and processed by the dropout block 560 to generate a first output received by the fully - connected block 562. The fully - connected block 562 can receive and process the first output to generate a second output, at least a portion of which is received by the output block 564. The output block 564 receives and processes its input to generate results such as an interaction output 566, an immunogenicity output 568, or both.
[0203] In some embodiments, the fully - connected block 562 can be configured to generate one or more outputs with dimensions smaller than the dimension of its input (fed into the fully - connected block 562, e.g., less than a predetermined number of features). For example, the output of the fully - connected block 562 can include a single value, two values, or three values, each corresponding to a prediction related to interaction with a target or an immune response. The fully - connected block 562 can include, for example, a single hidden layer, two hidden layers, or three or more hidden layers. The number of nodes in the initial hidden layer can be greater than the number of nodes in subsequent hidden layers. For example, the first hidden layer can include 256 nodes, while the second hidden layer can include 126 nodes. In some embodiments, each output from the fully - connected block 562 can include a real - valued score, which can be, for example, converted to a binary and / or categorical result (e.g., using a trained activation function) and / or converted to a scaled number. For example, the scaled number can include a probability on a scale from 0 to 1.
[0204] The interaction output 566 may include, for example, one or more of the following: a set of interaction predictions 570 that interact with respect to one or more targets or a set of interaction affinity predictions 572. The interaction predictions 570 may include, for example, predictions as to whether an IPC (e.g., MHC, TCR) in a corresponding amino acid-IPC combination such as a peptide-IPC (e.g., peptide-MHC, peptide-TCR) combination will bind to the peptide or the degree of binding. In some embodiments, the interaction predictions 570 may include, for example, predictions as to whether an IPC (e.g., MHC) in a corresponding peptide-IPC (e.g., peptide-MHC) combination will bind to the peptide. In some embodiments, the interaction affinity predictions 570 may include, for example, predictions of the affinity of the target interaction of the corresponding peptide-IPC (e.g., peptide-MHC, peptide-TCR) combination. The target interaction may be, for example, the binding of a peptide to an IPC. The affinity for the target interaction (which may be, for example, the binding affinity) indicates the strength, tendency, and / or stability of the binding between the peptide and the IPC.
[0205] The immunogenicity output 568 includes a set of immunogenicity predictions. The immunogenicity predictions may include, for example, immunogenicity predictions with respect to corresponding amino acid-IPC combinations. For example, the immunogenicity predictions may indicate the ability of a peptide to elicit an immune response with respect to a particular IPC of interest (e.g., MHC, TCR). In some embodiments, the predicted amino acid-IPC interactions include predictions of the tumor-specific immunogenicity of the peptide. In some embodiments, the predicted amino acid-IPC interactions identify a subset of peptide sequences that have increased tumor-specific immunogenicity or an increased likelihood of presentation by the IPC with respect to a group of peptide sequences.
[0206] In some cases, a first portion of the output from the fully connected block 562 is sent into the output block 564, while a second portion of the output from the fully connected block 562 is in its final form and serves as a set of interaction affinity predictions 572.
[0207] In some embodiments, the transformed synthetic representation is received at the output subsystem 509 and processed by the fully connected block 562. The fully connected block 562 may process the transformed synthetic representation to generate a first output that is sent into the discard block 560. Then, the output of the discard block 560 or a portion thereof may be sent to the output block 564 for processing.
[0208] In some cases, dropout may have been applied to each fully-connected block within the fully-connected block 562, followed by a batch normalization layer. In some embodiments, the output block 564 is used for deconvolution such that the amino acid-IPC interactions (e.g., twelve pairs of peptide-MHCII interactions or six pairs of peptide-MHCI interactions) (respectively) correspond to a single selected MHC II allotype or MHC I allele by applying an activation function to the presentation prediction (e.g., via a max layer 565 that may include a softmax function or simply use the maximum value). During training, the selected peptide-MHC interaction output can be normalized to a value between 0 and 1, and a loss function (e.g., a binary loss function) can be used to compare it with the true presentation value to generate an error for adjusting the model parameters.
[0209] In some embodiments, the output from the output subsystem 509 can include multiple results, the results including predicted amino acid-IPC interactions for each IPC (e.g., MHC) allele, which indicate whether the peptide binds to the IPC allele and / or the probability. Allele-specific predictions can be output, or in some cases, the max layer 565 can be used to determine the maximum value of the allele-specific predictions, and that maximum value can be output.
[0210] In this way, the output subsystem 509 can be implemented in any number of different ways using any number of different blocks, sub-blocks, and / or layers capable of generating an interaction output 566, an immunogenicity output 568, or both.
[0211] In some cases, the machine learning model 532 can facilitate the automated determination of which specific IPC allele binds to a peptide. For example, if the MHC molecule includes twelve MHC allotypes (as in the case of humans), then at least twelve iterations (e.g., in parallel) of neural network processing can be performed, one for each allele. Each processing can use the MHC sequence representation and a peptide representation of at least a portion of the peptide sequence as inputs. Each processing can generate a synthetic representation. The predicted amino acid-IPC interaction can be determined based on the synthetic representation. In some embodiments, the predicted amino acid-IPC interaction can include a prediction regarding whether the peptide will bind to the MHC allele or allotype and the degree of binding. It can be inferred that the peptide associated with the highest predicted value in the allele (e.g., indicating the most likely binding prediction) is the peptide that the peptide will bind to.
[0212] In some cases, for six to twelve MHC alleles (e.g., MHC class I) or allotypes (e.g., MHC class II), different MHC allotype sequences can be run through the same BOS+MHC representation block 504 and corresponding MHC sequence representations (such as MHC sequence representations) can be generated by generating IPC sequence representations for each peptide-MHC combination. In some embodiments, the MHC sequence representations, along with a single additional BOS token that has been embedded with an embedding layer, can be aggregated by multiplying the BOS token of each of the transformed amino acid sequence representations (e.g., containing a set of transformed peptide representations) with the BOS token of the set of transformed IPC sequence representations (e.g., containing each of the six to twelve MHC sequence representations) to generate a synthetic representation.
[0213] In some embodiments, one or more of the processing blocks or sub-blocks included in the machine learning model 532 can be replaced with another type of network and / or processing unit to transform the representation of one or more sequences. This transformation can represent predicting the degree to which various amino acids (at specific positions) affect binding affinity and / or presentation probability and / or predicting the degree to which various specific amino acid combinations (at specific positions) occurring on a single sequence or across sequences affect binding affinity and / or presentation. For example, one or more processing sub-blocks can be replaced with one or more gated recurrent units.
[0214] Figure 5B Schematic diagram of an exemplary configuration of a machine learning model 532 according to some embodiments. For Figure 5B the configuration depicted in, the representation subsystem 501 includes an amino acid sequence representation block 580. The amino acid sequence representation block 580 receives an amino acid sequence. For example, the amino acid sequence can include one or more of the following: a peptide sequence (e.g., Figure 1 peptide sequence 126 in), an N-flanking sequence (e.g., Figure 1 N-flanking sequence 128 in), or a C-flanking sequence (e.g., Figure 1 C-flanking sequence 130 in).
[0215] The amino acid sequence representation block 580 can include, for example, an embedding layer 582 that processes the amino acid sequence appended with a BOS token to form an embedded amino acid sequence representation received by a position encoder 583. The position encoder 583 performs position encoding on the embedded amino acid sequence representation to generate a BOS+amino acid sequence representation 584. In some embodiments, the BOS+amino acid sequence representation 584 can include one or more of the following: a BOS token representation 581, a peptide representation 585, an N-flanking representation 586, or a C-flanking representation 587.
[0216] The BOS+ amino acid sequence representation 584 is the output from the amino acid sequence representation block 580 and is sent to the processing block 588 in the processing subsystem 503. The processing block 588 includes a set of processing sub-blocks 589 that process the BOS+ amino acid sequence representation 584 to generate a transformed BOS+ amino acid sequence representation, and send the transformed BOS+ amino acid sequence representation to the synthesis block 552 for processing.
[0217] In some embodiments, if the amino acid sequence sent to the amino acid sequence representation block 580 includes an N-flanking sequence or a C-flanking sequence, but not both, the machine learning model 532 may also include a corresponding representation block (e.g., Figure 5A the N-flanking representation block 506 or the C-flanking representation block 508) and a corresponding processing block (e.g., the processing block 536 or the processing block 538 of Figure 5A respectively) for the sequence included in the amino acid sequence.
[0218] Figure 5C is a schematic diagram of an exemplary configuration of the machine learning model 532 according to some embodiments. The representation subsystem 501 can be designed to process one or more subsequences (or combinations thereof) of each type (e.g., peptide, C-flanking, N-flanking, C-flanking+N-flanking, MHC, TCR) in order to accommodate the splicing of multiple subsequences and a single additional BOS token by type before encoding the spliced group of subsequences. As Figure 5C shown in the example of, the representation subsystem 501 may also include a single representation block 507 for a combined sequence including a single BOS token, a C-flanking subsequence, and an N-flanking subsequence. The representation block 507 can include an embedding layer 521 and a positional encoder 523. The BOS+C-flanking+N-flanking sequence representation generated by the representation block 507 can be transformed by a processor block 590a that includes a set of processing sub-blocks 592a.
[0219] Alternatively, the peptide representation, the BOS+ peptide sequence representation, and the BOS+C-flanking+N-flanking sequence representation generated by the representation subsystem 501 can be aggregated before the transformation stage to form an amino acid sequence representation, and the amino acid sequence representation is sent to a single processing block 590. The processing block 590 includes a set of processing sub-blocks 592 that process the amino acid sequence representation to generate a transformed amino acid sequence representation, and send the transformed amino acid sequence representation to the synthesis block 552 for processing.
[0220] Similarly, the BOS+MHC sequence representation and the BOS+TCR sequence representation generated by the presentation subsystem 501 can be aggregated before the transformation stage to form an IPC sequence representation, which is sent to the processing block 594. The processing block 594 includes a set of processing sub-blocks 596, and the set of processing sub-blocks processes the IPC sequence representation to generate a transformed IPC sequence representation, which is sent to the synthesis block 552 for processing. In some embodiments, the BOS+MHC sequence representation can be processed by a set of blocks (e.g., 594a, 596a), while the BOS+TCR sequence representation can be processed by another set of blocks (e.g., 594b, 596b).
[0221] As Figures 5A to 5C shown, the machine learning model 532 can be implemented in any number of ways using any number of blocks, sub-blocks, and / or layers or combinations thereof within various subsystems. Thus, the machine learning model 532 is modular and can be customized for a given task.
[0222] Exemplary Processing Blocks
[0223] Figure 6 FIG. is a schematic diagram of a processing block 600 for performing one or more transformer stages to generate a transformed sequence representation according to some embodiments. The processing block 600 can be an instance implemented for the processing block in the processing subsystem 304 in Figure 3 or for the processing subsystem 503 in Figures 5A to 5C .
[0224] The processing block 600 includes one or more processing sub-blocks. For example, the processing block 600 can include a processing sub-block 1 602, and optionally, one or more other processing sub-blocks up to a processing sub-block n 604. When multiple processing sub-blocks are present in the processing block 600, these processing sub-blocks can be connected in series (e.g., daisy-chained together to generate a final output).
[0225] The processing sub-block 1 602 can be implemented in various ways. In some embodiments, the processing sub-block 1 602 includes, for example, a processing layer 606, a residual connection and a normalization layer 608, a feed-forward layer 610, and a residual connection and a normalization layer 612. With this configuration, the sub-block 1 602 can also be referred to as a transformer encoder. In some embodiments, one or more of the processing sub-blocks 604 can be implemented in a manner similar to the processing sub-block 1 602. In some embodiments, the processing layer 606 can include one or more embedding components configured to perform positional and / or non-positional embeddings.
[0226] In the residual connection and normalization layers 608 or 612, the transformed representation can be added to the positional embedding representation of the sequence (via the residual connection), and the summed representation can be normalized. The normalized data can be fed to the corresponding feed-forward layer 610 (e.g., a fully-connected feed-forward network). For each position, the feed-forward network can affect (e.g.) one, two, three, or more linear transformations and / or can include activations (e.g., ReLU activations) between each linear transformation. For example, the feed-forward layer can be represented as:
[0227] FF(x) = max(0, xW1 + b1)W2 + b2,
[0228] where x is the input to the layer, W1 and W2 are the linear transformation slopes, and b1 and b2 are the linear transformation intercepts.
[0229] The dimension of the output of the feed-forward layer of a particular processing sub-block can be the same as the dimension of the input of the feed-forward layer of the processing sub-block. In some cases, to preserve the representation of various types of information, the input and output can be summed and normalized (e.g., via another residual connection and normalization layer through another residual connection).
[0230] In some embodiments, the feed-forward layer 610 can allow processing of variable-length sequences. One or more additional feature vectors (e.g., assigned random or pseudo-random values) can be included in the concatenated representation, which is then encoded. This encoded representation of the sequence combination can be processed by the feed-forward layer 610 (e.g., a fully-connected neural network), where dropout and / or batch normalization can be applied. In some cases, the encoded representation of the additional feature vectors is selectively passed to the feed-forward layer 610 (e.g., the feature vectors corresponding to the individual amino acids of the MHC molecule and / or the mutant peptide are not passed). For example, assume that a subsequence of the MHC molecule includes x1 amino acids, a subsequence of the mutant peptide (e.g., and one or more flanks) includes x2 amino acids, and the feature transformation identifies y feature values to represent each amino acid. Thus, the concatenated representation including one additional feature vector can have a size of [(x1 + x2 + 1), y]. In the case of selecting one feature vector for processing by the feed-forward layer 610, the input fed to the feed-forward layer 610 can have a size of [1, y].
[0231] The result generated by the feed-forward layer 610 may correspond to a prediction regarding the binding affinity between a mutant peptide and an MHC molecule (e.g., the MHC molecule of a subject) and / or whether the mutant peptide will be presented by the MHC molecule. The binding affinity prediction may be, for example, numerical (e.g., corresponding to the predicted probability that the mutant peptide will bind to the MHC molecule, the predicted binding strength, and / or the predicted binding stability), categorical (e.g., predicting no binding, low binding stability, or high binding stability between the mutant peptide and the MHC molecule), or binary (e.g., predicting whether the mutant peptide binds to the MHC molecule).
[0232] The machine learning model 132 may include one or more processing layers, such as a self-attention layer or a convolutional layer, or a neural network such as a long short-term memory unit (LSTM), a recursive structure, or a recursive component. Fig. 7A A flowchart illustrating an exemplary method of using a processing layer to process a sequence representation according to some embodiments. The method 700 may be performed by, for example, one or more of the processing blocks in the machine learning model 132 present in Figure 1 and Figure 3 one or more of the processing blocks in the machine learning model 532 present in Figures 5A to 5C and / or Figure 6 the processing block 600 in
[0233] Step 702 includes receiving a sequence representation including a plurality of elements. The sequence representation may be, for example, an amino acid sequence representation, an IPC sequence representation, an N-flanking sequence representation, a C-flanking sequence representation, an MHC sequence representation, a TCR sequence representation, an aggregate sequence representation, or another type of representation. For example, the sequence representation may represent some or all of the following: a variant coding sequence, part or all of a sequence encoding a wild-type or mutant peptide, an epitope sequence (e.g., which includes a variant), a candidate new epitope sequence, part or all of a neoantigen sequence, a sequence starting or ending at the end of a peptide (e.g., N-flanking or C-flanking), or an MHC sequence (e.g., an MHC pseudo-sequence). For example, the representation subsystem 302 in Figure 3 or the representation subsystem 501 in Figures 5A to 5C may be used to generate the sequence representation. Each element in the sequence representation may be associated with a unique position in the sequence.
[0234] Step 704 includes determining a plurality of vectors (such as key vectors, value vectors, and query vectors) for each element in the sequence representation, respectively, using a plurality of weights (such as a set of key weights, a set of value weights, and a set of query weights). For example, if the sequence representation includes, for example, 20 amino acids, 20 key vectors, 20 value vectors, and 20 query vectors can be generated. The elements in the sequence representation can correspond to, for example, rows or columns in a two-dimensional sequence representation (e.g., where the first dimension represents different amino acids in the sequence and the second dimension represents, for example, different components characterizing a single amino acid).
[0235] In some embodiments, the set of key weights is in the form of a key weight matrix. The key weight matrix for a particular element can have a size equal to the length of the element multiplied by the length of the key vector. For example, the element can have a length of 20 (e.g., each value corresponds to a binary indication of whether an amino acid in the sequence is the same as a particular one of 21 amino acids), and if the length of the key vector is 5 (e.g., representing 5 components or features), the key weight matrix can have a size of [5, 21]. The key weight matrix can be learned during training and is, for example, randomly initialized at the start of training.
[0236] The value vector of an element may have the same size as the key vector of the element. A set of value weights can be used to determine the value vector, and the set of value weights can be learned during training and included within a value weight matrix. The value weight matrix for a given element can have the size of the key weight matrix and / or a size based on the length of the element and the length of the value vector.
[0237] The query vector for an element can have the same size as the key vector and / or value vector for that element. A set of query weights can be used to determine the query vector, and the set of query weights can be learned during training and included within a query weight matrix. The query weight matrix for an element can have the size of the key weight matrix and / or the value weight matrix. In some embodiments, the query weight matrix can have a size based on the length of the element and the length of the query vector.
[0238] Step 706 includes, for each element in the sequence representation, generating a set of element focus scores using the query vector of the element (generated using query weights and the sequence representation) and the key vectors of a plurality of elements (generated using key weights and the sequence). For a given element, the set of element focus scores can indicate the weight of the value vector of the given element. The elements for which a set of element focus scores is generated for selected elements in the sequence representation using key vectors can include a part or all of the elements in the sequence representation (e.g., a part or all of an amino acid sequence representation). The elements can include a focus element (e.g., a particular amino acid for which the set of element focus scores is being determined).
[0239] An element focus score group is generated by generating a score for each pair of focus elements (first elements) having the same or different elements (second elements) for each element of the sequence representation. The score for the pair can be the product of the query vector of the first element and the key vector of the second element.
[0240] In some cases, step 706 can include implementing an activation function and / or normalization. The normalization can be based on the dimension of the key vector (or query vector). For example, the normalization can be the square root of the length of the key vector. The activation function can include the softmax function. In some cases, normalization is applied before the activation function.
[0241] Step 708 includes generating a transformed sequence representation. The transformed sequence representation can be determined by transforming a plurality of elements to form a plurality of modified elements. The transformation can be performed using the element focus score group generated for each of the plurality of elements and the value vector determined for each of the plurality of elements. For example, if the sequence representation includes 11 elements (e.g., representing 11 amino acids), and if scores are determined for all pairwise combinations of the elements, a modified sequence representation including a plurality of modified elements is generated. In some embodiments, the modified element can be a weighted average of the value vectors of all elements (weighted using the scores).
[0242] Step 710 includes generating an encoding of the sequence using the transformed sequence representation, the initial sequence representation, and a feed-forward network. For example, the transformed sequence representation and the initial sequence representation can be summed. This result may still include a plurality of elements (e.g., each element is updated via transformation, summation, and normalization). Then the feed-forward neural network can process the summed representation (e.g., by performing one, two, or more linear transformations and / or implementing one or more activation functions). Summing the representations can reintroduce positional information that may have been confounded in the transformed sequence representation (due to focusing on the values of other elements when generating the transformed value vector for a given element).
[0243] The feed-forward neural network can be configured to process each of the updated plurality of elements separately (e.g., using the same technique and / or the same set of parameters). Thus, the input to the feed-forward network can include a vector corresponding to a single element, a single amino acid, and / or a single sequence position. The feed-forward network can be configured such that the output of the feed-forward network has the same size as the input to the feed-forward network. In some cases, instead of using a feed-forward network to process the transformed sequence representation and the initial sequence representation, convolution (e.g., 1D convolution) is used to perform local transformations that operate similarly (e.g., identically) across positions / elements. 1D convolution can be used as an alternative way to interpret the functionality of the feed-forward neural network.
[0244] Fig. 7AThe techniques shown involve a process of calculating element focus scores using a single set of key vectors, value vectors, and query vectors. Embodiments of the present disclosure may include using multiple sets of key weights, value weights, and query weights to generate different key vectors, different value vectors, and different query vectors. These different vectors can be used to generate processing scores and transformed values for each element. The transformed values can be concatenated and projected.
[0245] It should be further understood that although Fig. 7A it involves the calculation and use of various vectors, a matrix representation can alternatively be used. Instead of calculating various vectors iteratively one by one, the matrix representation can effectively facilitate the execution of cross-element calculations.
[0246] Figure 7B For an illustration according to some embodiments Fig. 7A of the method 700 described. In Figure 7B this, the method 750 receives a sequence 752 as input. The sequence 752 can be, for example, an amino acid sequence. Another exemplary sequence can be an IPC sequence
[0247] In Figure 7B the illustrative example, the sequence 752 includes a plurality of amino acids 754 (4 amino acids: x 1 to x 4 ). A sequence representation 756 including a plurality of elements a 1 -a 4 is generated via embedding and, in some embodiments, positional encoding. Each element a i can be a numerical vector. The sequence representation 756 can be an instance of the sequence representation received in step 702 in Fig. 7A .
[0248] Multiple vectors 758 (e.g., query vector q i , key vector k i , and value vector v i ) can be generated for each element a i in the sequence representation 752. The multiple vectors 758 can be an instance of the implementation of the vectors generated in step 704 in Fig. 7A . The shown example corresponds to generating a selected element focus score 760, where the focus is on the first element a 1 . The element focus score 760 is an instance of a set of element focus scores generated for a specific element in step 706 in Fig. 7A . Each of the element focus scores can be the dot product of q 1 and k i . The calculation weight is set to the value vector v iThe weighted sum to perform the generation of the modified element 762,b 1 The transformation of. The modified element 762 is an instance of the modified element generated in step 708 in Fig. 7A Similar transformations can be performed on other elements of the sequence representation 756. For more details on the converter architecture, see Ashish Vaswani et al., Attention is All You Need, Neural Information Processing Systems (2017).
[0249] Example Methods of Using Machine Learning Models
[0250] Figure 1 and Figure 3 The machine learning model 132 in Figures 4A to 4D The workflow in Figures 5A to 5C The machine learning model 532 in can be used in various ways to generate predictions about the immunological activity associated with various peptides (including mutant peptides (e.g., neoantigens)), such as predicted binding, binding affinity, predicted presentation occurrence, immunogenicity, etc.
[0251] Figure 8 Is a flowchart of an exemplary method for generating information about the immunological activity of various peptides according to some embodiments. At least a portion of method 800 can be implemented using, for example, but not limited to Figure 1 The prediction system 100 described. For example, at least a portion of method 800 can be implemented using, for example, but not limited to the machine learning model 132 from Figure 1 and Figure 3 The machine learning model 532 from Figures 5A to 5C Or the workflow in Figures 4A to 4D To perform.
[0252] Step 802 includes accessing an amino acid sequence that includes a peptide sequence characterizing a mutant peptide, which may include variants with respect to a corresponding reference sequence. The peptide sequence characterizes the mutant peptide by characterizing at least a portion of the mutant peptide. The mutant peptide can be, for example, a neoantigen. Step 802 can be performed, for example, by retrieving the peptide sequence from a data repository (e.g., Figure 1 The data repository 104 in, cloud storage, a server, or a server system, etc.). In some embodiments, the peptide sequence can be one of a plurality of peptide sequences processed by a machine learning model.
[0253] Step 804 includes receiving an IPC sequence identified for the IPC of a subject. The IPC can be, for example, MHC, TCR, or an MHC-TCR complex. The IPC sequence characterizes the IPC by characterizing at least a portion of the IPC.
[0254] Step 806 includes processing the amino acid sequence and the IPC sequence using different processing engines within a machine learning model to generate an output, where the output provides information regarding immunological activity associated with both the mutant peptide and the IPC. Step 806 includes, for example, processing the amino acid sequence through a corresponding representation block to generate an amino acid sequence representation. The amino acid sequence representation can be processed through a corresponding processing block to generate a transformed amino acid sequence representation. The amino acid processing engine and the IPC processing engine are separate and independent, where in the IPC processing engine, the IPC sequence is processed through a corresponding representation block to generate an IPC sequence representation (e.g., MHC representation, TCR representation, MHC-TCR representation), and the IPC sequence representation is processed through a corresponding processing block to generate a transformed IPC sequence representation representing the IPC sequence (e.g., transformed MHC representation, transformed TCR representation, transformed MHC-TCR representation).
[0255] In some embodiments, the amino acid sequence representation is an aggregate representation that includes an N-flank representation for an N-flank sequence and / or a C-flank representation for a C-flank sequence. In such embodiments, the aggregate processing engine (which may include the amino acid processing engine) remains separate from the IPC processing engine.
[0256] In various embodiments, in step 806, the transformed amino acid sequence representation and the transformed IPC sequence representation are used to form a synthetic representation, which is then further processed to generate an output. The output can include, for example but not limited to, an interaction prediction set, an interaction affinity prediction set, an immunogenicity prediction set, or a combination thereof.
[0257] Step 808 includes performing one or more operations based on the output. As an example, a report including the output can be generated. In some embodiments, the report includes a transformed or filtered version of the output. In some embodiments, the report includes a summary, abstract, or visual representation of the output.
[0258] In some embodiments, step 808 includes other actions related to the design and / or manufacture of a treatment based on the output. For example, a drug composition can be selected or ranked based on the output. The output can include predictions of which mutant peptides bind to a specific IPC of a subject (e.g., MHC allele or allotype). The binding prediction can indicate the likelihood that the subject's immune system can recognize (e.g., cancer cells). The binding prediction can be used to help select candidate neoepitopes (mutant peptides) for a vaccine. In some embodiments, the synthetic representation with the highest result (e.g., a prediction value indicating the most likely binding and / or presentation prediction) in the output can be selected for the drug composition. In some embodiments, the synthetic representations can be ranked according to the corresponding results in the output.
[0259] Embodiments of the present disclosure may include generating an output based on a set of IPC sequences. For example, for a given subject, the output may be generated based on six to up to twelve MHC alleles or allotypes. Fig. 9 FIG. is a flowchart of an exemplary method for generating information regarding the immunological activity of various peptides according to some embodiments. At least a portion of method 900 may be implemented using, for example but not limited to Figure 1 the prediction system 100 described above. For example, at least a portion of method 900 may be implemented using, for example but not limited to, a machine learning model 132 from Figure 1 and Figure 3 or a machine learning model 532 from Figures 5A to 5C
[0260] Step 902 includes accessing sequence data including a set of amino acid sequences and a set of IPC sequences.
[0261] Step 904 includes generating a set of amino acid-IPC combinations using the set of amino acid sequences and the set of IPC sequences. Each amino acid-IPC combination is a unique combination.
[0262] Step 906 includes, for each amino acid-IPC combination, inputting the corresponding amino acid sequence into the amino acid processing engine of the machine learning model and inputting the corresponding IPC sequence into the IPC processing engine of the machine learning model.
[0263] Step 908 includes, for each amino acid-IPC combination, processing the amino acid sequence representation using a first processing block and processing the IPC sequence representation using a second processing block to generate a transformed amino acid sequence representation and a transformed IPC sequence representation, respectively.
[0264] Step 910 includes, for each amino acid-IPC combination, generating a synthetic representation using the transformed amino acid sequence representation and the transformed IPC sequence representation.
[0265] Step 912 includes generating an output based on the synthetic representation. In some embodiments, a predicted amino acid-IPC interaction may be determined based on the synthetic representation. The output may provide an indication of which peptide sequences may be used for a treatment. For example, the output may provide an indication of which peptide sequences (and thus the peptides containing such peptide sequences) have a high likelihood of binding to MHC, a high likelihood of being presented by MHC, a high interaction affinity with peptide-MHC binding, and / or a high likelihood of being immunogenic to trigger an immune response.
[0266] Example method for training a machine learning model
[0267] Fig.10Flowchart of an exemplary method for training a machine learning model and using the trained machine learning model to generate predictions related to amino acids (e.g., peptides) and IPC (e.g., MHC). Method 1000 can be performed using the Figure 1 prediction system 100 therein. For example, method 1000 can use the Figure 1 and Figure 3 machine learning model 132 therein, the Figures 5A to 5C machine learning model 532 therein, or any workflow in Figures 4A to 4D to perform. In some cases, part or all of method 1000 can be performed at a remote computing system that is remote relative to the user device and / or laboratory. The remote computing system can be a cloud computing system.
[0268] At least a portion of the training dataset can be used to train the machine learning model. Step 1002 includes accessing a training dataset having training elements that identify training amino acid sequence data, training IPC sequence data, and training immunological activity data. The training dataset can be an instance of the implementation of the training data 133 in Figure 1 . The training immunological activity data can include, for example, interaction indicators.
[0269] The training dataset can include a plurality of training elements. Each training data element can include a sequence representation and a result (e.g., indicating whether at least a portion of the peptide corresponding to the sequence is presented by an MHC molecule and / or triggers immunogenicity). Training data elements for which presentation or binding is not detected can be generated computationally. For example, for each source protein in the positive set (corresponding to positive elution ligand presentation data), one or more (e.g., all) possible peptide fragments (e.g., within a predetermined length range, such as 8 to 11) can be generated, where for each length there can be a uniform probability. The N-terminal and C-terminal flanking sequences can be retained (e.g., may have a maximum length, such as 10 amino acids). In some cases, for each allele represented in the positive instances of the training data, peptide fragments (e.g., one or more (e.g., all) lengths of 8:11) can be generated. Generation and / or subsequent selection can be performed such that the occurrence probability of sequences of a given length is uniform over the length. The N-terminal and C-terminal flanking sequences can be or may have retained a specific maximum length (e.g., a maximum length of 10 amino acids). In a particular embodiment, any other suitable sequence length range can be utilized (e.g., 9 to 30 for MHC class II).
[0270] The training dataset can be randomly parsed, shuffled, and / or split to train various models within the population. The loss function can use error terms (e.g., mean squared error or median squared error) and / or entropy terms (e.g., cross entropy or binary cross entropy). Multitask learning can be used to train the model to predict each of two different types of outcomes simultaneously (e.g., combining affinity and presentation occurrence). A static or non-static learning rate can be used. For example, learning rate annealing (e.g., using step annealing or cosine annealing) can be used to reduce the learning rate in iterations. Validation data evaluation can be used to potentially terminate training early (e.g., when it is determined that a performance goal has been reached).
[0271] The training amino acid sequence data can include, for example, one or more amino acid sequences for training (which can include variant coding sequences). The amino acid sequence can contain a peptide sequence. The peptide sequence can identify an ordered set of amino acids within a peptide (e.g., a neoantigen). The peptide sequence can discriminate the amino acids within an epitope of the peptide (e.g., including variants, including neoepitopes, and / or being a neoepitope). In some embodiments, the peptide sequence is within a polymeric sequence that also includes an N-flanking sequence (e.g., characterizing the amino acid chain at the N-terminus of the corresponding peptide) or a C-flanking sequence (e.g., characterizing the amino acid chain at the C-terminus of the corresponding peptide). Neither the N-flanking nor the C-flanking binds to the MHC molecule, although each may affect whether it is presented by the MHC molecule.
[0272] In some cases, it is not known how many amino acids from the flanking region (e.g., N-flanking) are used by the peptidase to determine when to trim a long peptide into the presented peptide core. To address this unknown issue when generating training data, the flanking region can then be trimmed to a length selected based on a technique (e.g., a pseudo-random selection technique), such as a length within a predetermined range (e.g., 1 to 10 amino acids). The selection technique can use a distribution (e.g., a uniform distribution or a Gaussian distribution) to select the length. In some cases, flanking regions below a threshold length (e.g., 10 amino acids) are not trimmed. In some cases, flanking region trimming can be such that the C-side on the N-flanking is retained.
[0273] The training MHC sequence data can include one or more MHC sequences for training. For example, the MHC sequence can identify the amino acids within part or all of an MHC molecule (e.g., an MHC-I molecule or an MHC-II molecule). The MHC sequence can include an MHC pseudo-sequence (e.g., including 34 amino acids). The MHC sequence can discriminate the amino acids within, for example, 1, 2, 3, 4, 5, or 6 MHC alleles for MHC-I, or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 MHC allotypes for MHC-II. The MHC sequence can identify the amino acids that make up part or all of an HLA molecule.
[0274] MHC includes multiple alleles in the body (e.g., each person has six alleles and twelve allotypes). For a single MHC molecule, multiple sequence inputs can be generated (e.g., each sequence input represents a single allele among the multiple alleles). One or more neural networks (e.g., one or more transformer encoders) can be used to process each of the multiple sequence inputs separately to generate predicted binding or presentation values of neoantigens associated with each allele. A function (e.g., max function) can identify which allele among the multiple alleles is associated with the highest presentation prediction. During training, a binary loss function can then be used to compare this maximum presentation prediction for the particular sequence input with the true presentation value to generate an error for tuning the parameters.
[0275] Training immunological activity data can include, for example, one or more interaction indications for one or more amino acid-IPC combinations. For example, a training data set can include training elements, where each training element includes an amino acid sequence and an IPC sequence for training, and one or more interaction indications for the corresponding amino acid-IPC combination. The interaction indication can indicate whether a target interaction (e.g., binding of a peptide to an MHC, presentation of a peptide by an MHC on the cell surface) occurs between the amino acid (e.g., peptide) and the IPC (e.g., MHC), or the affinity for the target interaction and / or the triggering of an immune response.
[0276] The interaction indication can be, for example, a label. A negative interaction label can indicate that a peptide does not bind to and / or is not presented by the IPC (e.g., MHC molecule). A positive interaction label can indicate that a peptide binds to and / or is presented by the MHC molecule. Further, the interaction label can indicate the probability that a peptide binds to the MHC molecule, the binding affinity for the peptide-MHC combination, the binding strength between the peptide and the MHC molecule, the stability of the binding between the peptide and the MHC molecule, the tendency of the peptide to bind to the MHC, and other metrics or characteristics related to the interaction between the MHC and the peptide.
[0277] The training dataset may have been generated, for example, by in vitro or in vivo experiments and / or based on medical records. In some embodiments, binding affinity data indicating which peptides are presented by MHC molecules and mass spectrometry elution data may be used to train a machine learning model. The binding affinity data may include qualitative data (e.g., using ELISA, pull-down assays and / or gel shift assays, fluorescence resonance energy transfer assays, and mass spectrometry assays) or quantitative data (e.g., using biosensor-based methods such as surface plasmon resonance, isothermal titration calorimetry, biolayer interferometry, or microscale thermophoresis). In some cases, the binding affinity data may include data from competitive binding assays, data from immune epitope databases, and / or data of types in immune epitope databases. Elution data may be collected using peptide-MHC immunoprecipitation, and then the MHC ligands presented are eluted and detected by mass spectrometry.
[0278] To collect training data, a portion of the sequences identified in a disease sample may be non-disease sequences corresponding to non-disease peptides. To identify disease-specific nucleic acid sequences and / or disease-specific amino acid sequences, for each sequence detected as a result of sequencing a disease-specific sample, it may be determined whether the sequence is also identified in a reference sequence dataset. The reference sequence dataset may include a set of reference sequences for which it is known, inferred, or assumed that the sequence does not indicate or characterize a disease (e.g., any disease or a given disease). The reference sequence dataset may, for example, include sequences identified by sequencing one or more reference sample sequences collected from the same subject from whom the disease-specific sample was collected, sequencing one or more reference sample sequences collected from one or more other subjects who have never been diagnosed with any disease or the disease corresponding to the disease-specific sample, and / or sequencing one or more cell lines not related to a particular disease. In some cases, the reference sequence dataset may include sequences collected from one or more reference data repositories. Sequences detected in association with a disease-specific sample but not detected (or detected at a frequency below a predetermined threshold) in the reference sequence dataset may be classified as variant coding sequences (e.g., generally or for the subject from whom the disease-specific sample was collected).
[0279] In some cases, multiple variant coding sequences may be identified (e.g., each variant coding sequence has been detected in a disease sample but not represented in the reference sample sequences). In some cases, the machine learning model disclosed herein may be used to process (e.g., individually, sequentially, and / or in parallel) the representations of each of the multiple variant coding sequences to predict binding affinity and / or presentation prediction.
[0280] Disease samples can include, for example, tissue (e.g., solid tumors), blood, and / or cell aggregates (e.g., cancer cells, which can be collected using fine needle aspiration or laparoscopy). Disease samples can include cancerous cells collected from subjects diagnosed with and / or suffering from, for example, lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, and T-cell lymphocytic leukemia, non-small cell lung cancer or small cell lung cancer.
[0281] In some cases, an initial sample is split into a disease sample and another remaining sample (e.g., which can be discarded or used as a reference sample). The reference sample can include a matched disease-free sample. The disease sample and the reference sample can each be collected from the same subject and / or can include or have the same or similar sample type (e.g., tissue type). In some cases, the disease sample is collected from a first subject (e.g., diagnosed with a medical condition or disease), and the reference sample is collected from a different second subject (e.g., not diagnosed with a medical condition or disease). In some cases, the reference sample sequence is retrieved from a database of known genes associated with an organism.
[0282] The training data can further include the sequences of one or more peptides, and an indication as to whether each of the peptides binds to an MHC molecule, is presented by an MHC molecule, and / or triggers an immune response. To collect training data correlating sequence data with observed presentation and / or binding data, the disease sample (and possibly the reference sample) can be processed (separately) to isolate MHC / peptide complexes (e.g., by immunoprecipitation using MHC-specific antibodies) and / or elute (and thereby sequence) peptides from MHC molecules (e.g., using mass spectrometry and / or mass spectroscopy). In some cases, the reference sample sequence for generating presentation data is identified by sequencing one or more cell lines engineered to express one or more MHC alleles (e.g., the alleles detected in the disease sample), wherein the one or more MHC alleles can include MHC class I alleles and / or MHC class II allotypes. The one or more cell lines can include one or more human cell lines obtained from or derived from one or more subjects. For the purposes of this description, peptide sequences identified using the disease sample but not represented in the reference sample sequence set can be identified as variant coding sequences.
[0283] In some embodiments, collecting immunogenicity indicative metrics for training can be based on HLA typing analysis, which can identify a subject-specific MHC molecular profile. When the subject is human, this profile can be referred to as a human leukocyte antigen (HLA) profile, since the HLA complex is the gene complex encoding human MHC proteins. HLA typing analysis can be performed using a sample from the subject (e.g., normal tissue and / or non-disease sample). Sequencing techniques such as PCR-based sequencing, direct sequencing, and / or next-generation sequencing can be used to determine the profile. HLA typing analysis can include, for example, high-resolution typing (e.g., which excludes null alleles indicative of non-expression on the cell surface) or allele-level typing (e.g., which refers to the determination of the precise nucleotide sequence of HLA genes). HLA typing analysis can include low-resolution typing that identifies broader allele families and / or HLA supertyping.
[0284] Regarding any type of sequencing (e.g., for identifying sequences in a sample, peptides that bind to MHC molecules, HLA typing), the results can identify one or more nucleic acid sequences or one or more amino acid sequences. When a nucleic acid sequence is identified and an attention-based model (or other processing) is configured to process amino acid sequences, techniques (e.g., lookup tables) can be used to convert individual codons within the nucleic acid sequence into individual amino acids.
[0285] Some embodiments include synthetic peptides (e.g., using nucleic acid sequences encoding peptides such as selected peptides) or precursors of selected peptides. The synthetic peptides or precursors can then be used in experiments to identify corresponding presentation and / or binding data (e.g., to validate predicted presentation and / or binding or to generate results for training). For example, the experiments can include evaluating the binding affinity of selected peptides to specific MHC molecules using ELISA pull-down assays, gel shift assays, or biosensor-based methods. As another example, the experiments can include collecting elution data indicative of whether a selected peptide is presented by MHC molecules by using peptide-MHC immunoprecipitation, followed by elution and detection of the presented MHC ligands by mass spectrometry.
[0286] In addition to training or validation data that indicates whether a single peptide binds to and / or is presented by a single MHC, the training or validation data can indicate whether a single peptide triggers immunogenicity. Immunogenicity results can be determined using in vivo or in vitro tests. Testing one or more selected peptides can be configured to study one or more immunogenicity factors (e.g., determining whether a given event occurs and / or determining its extent) and / or immunogenicity (e.g., determining whether a peptide triggers an immune response and / or determining its extent). The testing can be configured to study whether a composition (e.g., a vaccine) that includes one or more peptides is effective in preventing or treating a medical condition (e.g., a tumor) or a disease (e.g., cancer) when administered to a given subject (e.g., a subject whose MHC sequence was identified during mutant peptide selection). The subject can be a human subject.
[0287] Accessing the training data set can include, for example, retrieving the training data set from local or remote memory, loading the training data set, and / or requesting (and receiving) a portion or all of the training data set from one or more data stores (e.g., cloud data storage, server systems, or some other data source).
[0288] The training data can include "positive" instances (e.g., whose mass spectrometry results indicate that the peptide is presented by an MHC molecule) and "negative" instances (corresponding to, e.g., mock length-matched n-mers (nmers)) from the same protein as the positive instances (e.g., but not detected in the mass spectrometry assessment).
[0289] In some cases, the initial training data set (e.g., which can include variant coding sequences) can primarily include negative data because a relatively small portion of the sequence combinations (e.g., peptide-MHC combinations) are associated with actual target interactions. The training data set can be designed to include negative training data elements. In some embodiments, the negative training data elements can be used to discriminate amino acids within pseudo-randomly selected fragments of the source protein in the positive set (corresponding to observed presentation). For example, the negative training data elements can be simulated based on the positive set. The fragments can be selected to have a length within a predetermined range (e.g., between 8 and 14 amino acids for MHC-I and between 8 and 30 amino acids for MHC-II using uniform probability). The N-terminal and C-terminal flanking sequences can be retained within the negative training data elements, possibly imposing a maximum length (e.g., 10 amino acids). Any peptide fragments that overlap with the positive peptide (e.g., at least 9-mer) can be discarded from the negative training data.
[0290] In some embodiments, negative training data elements are simulated based on positive data elements. Additionally, training data is selected such that a different set of negative training data elements is used for each epoch of a training cycle. For example, for each epoch, a different "negative subset" of negative peptide sequences can be selected from the overall space of available negative peptide sequences identified from a positive set based on peptide sequences. The negative subset selected for each epoch can be unique in that, for the total number of epochs, no negative peptide sequence is repeated in any of the negative subsets. Thus, the training data for each epoch of the training cycle includes the same set of positive peptide sequences but a completely different set of negative peptide sequences. This technique, which can be referred to as "negative set switching", can provide overall robustness to the training and help ensure that the number of false negatives (e.g., false negative indications / predictions) is reduced by the machine learning model or does not occur more than once. Further, using this technique, the machine learning model can be trained on the total number of negative peptide sequences, which is equal to the number of positive peptide sequences multiplied by the number of epochs in the training cycle.
[0291] In some instances, the number of positive instances in the training data is equal to the number of negative instances in the training data. In some instances, the number of positive instances is less than or greater than the number of negative instances. Each of one or more (e.g., all) of the negative instances in the training data can be length-matched to the positive instances in the training data. In some instances, all sequences in the training data have the same length.
[0292] Step 1004 includes training a machine learning model using a training data set. The machine learning model can be, for example, Figure 1 and Figure 3 the machine learning model 132 in Figures 5A to 5C or the machine learning model can be, for example,
[0293] the machine learning model 532 in
[0294] Step 1006 includes accessing a subject - specific set of variant coding sequences corresponding to a set of mutant peptides. As described above, a variant coding sequence is an example of a peptide sequence. The subject - specific set of variant coding sequences can correspond to the set of mutant peptides such that each variant coding sequence in the subject - specific set of variant coding sequences identifies an amino acid within the corresponding mutant peptide of the set of mutant peptides. In some embodiments, each in the subject - specific set of variant coding sequences identifies one or more amino acids in the mutation. Each variant coding sequence in the subject - specific set of variant coding sequences can be associated with a particular subject (e.g., a human subject). The particular subject may have been diagnosed with, have symptoms of, and / or received test results associated with a particular medical condition (e.g., cancer). For example, the subject - specific set of variant coding sequences can be identified by processing a sample from a tumor. The sample can be included within, for example Figure 1 sample set 112.
[0295] The techniques disclosed herein can be used to identify the subject - specific set of variant coding sequences. For example, the subject - specific set of variant coding sequences can be identified by performing sequencing techniques to identify peptides in a disease sample and comparing the identified peptides to peptides detected in a healthy sample or a reference database to identify unique sequences. In some embodiments, if the unique sequences are nucleic acid sequences, each unique nucleic acid sequence can be translated into an amino acid sequence.
[0296] Each in the subject - specific set of variant coding sequences can identify an amino acid within a peptide (which can be an amino acid within a neo - epitope of a neo - antigen). In some cases, each of one, more, or all in the subject - specific set of variant coding sequences can be part of a corresponding polymeric sequence that further includes a sequence N - flanking the peptide and / or a sequence C - flanking the peptide.
[0297] Accessing the subject - specific set of variant coding sequences can include, for example, retrieving the subject - specific set of variant coding sequences from local or remote memory and / or requesting the subject - specific set of variant coding sequences from another device. Accessing the subject - specific set of variant coding sequences can include and / or can be combined with determining the subject - specific set of variant coding sequences to perform.
[0298] The subject - specific set of variant coding sequences may have been obtained by identifying peptide sequences in a disease sample of a subject and determining which peptide sequences are not represented in a reference sample, a healthy sample, and / or a set of wild - type sequences. In cases where a healthy sample is used for comparison, a healthy sample may have (but need not have) been collected from the subject.
[0299] Step 1008 includes accessing an IPC sequence corresponding to the IPC. In some embodiments, the IPC sequence can be an MHC sequence. The MHC sequence can include, for example, a pseudo-sequence of MHC (e.g., MHC molecules) within a sample collected from a subject. In some cases, the MHC sequence and a set of subject-specific variant coding sequences are identified from the same sample from a subject or from multiple samples from a subject (e.g., a diseased sample and a healthy sample). In some cases, the MHC sequence and a set of subject-specific variant coding sequences are identified from samples from a subject and one or more other subjects. Thus, in some cases, the MHC sequence can be subject-specific. The MHC sequence can be determined using, for example, sequencing and / or mass spectrometry techniques or may have been determined using, for example, sequencing and / or mass spectrometry techniques.
[0300] Accessing the MHC sequence can include, for example, retrieving the MHC sequence from local or remote memory and / or requesting a subject-specific MHC sequence from another device. Accessing the MHC sequence can include determining an MHC sequence combination and / or performing operations therewith.
[0301] Step 1010 includes, for example, processing a set of subject-specific variant coding sequences and a set of MHC sequences using a trained machine learning model to generate an output. Step 1010 can include processing each unique combination of a subject-specific variant coding sequence in the set of subject-specific variant coding sequences and an MHC sequence (e.g., a variant coding-MHC combination or a peptide-MHC combination) to generate an output.
[0302] The output generated by the machine learning model can include data of the same or a similar type as the data included in the training immunological activity data used to train the machine learning model. For each unique combination, the machine learning model generates an output that includes at least one of a set of interaction predictions or a set of interaction affinity predictions.
[0303] Interaction predictions in the interaction prediction group include predictions regarding whether a target interaction will occur between a mutant peptide (including variant coding sequences) and an MHC (including MHC sequences). For example, the interaction prediction may include a binary or categorical prediction regarding whether a mutant peptide having an amino acid structure (as indicated by a subject-specific variant coding sequence) will be presented by an MHC molecule (having an amino acid structure as indicated by the MHC sequence) and / or bind to the MHC molecule. Interaction affinity predictions in the interaction affinity prediction set include predictions regarding the affinity for a target interaction. The affinity may be based on, for example, the strength, trend, and / or stability of the target interaction. For example, the interaction affinity prediction may include a predicted real-valued binding affinity associated with the mutant peptide and the MHC molecule, the mutant peptide including amino acids identified within a subject-specific variant coding sequence, and the MHC molecule including amino acids identified within the MHC sequence.
[0304] Step 1012 includes generating a report based on the output of the machine learning model. The report may be implemented as, for example, Figure 1 and Figure 3 Report 144 therein. The report may be or include the output. In some cases, the report may be a transformed or filtered version of the output.
[0305] In some embodiments, the subject-specific variant coding sequence set is filtered, sorted, and / or otherwise processed based on the output to generate information for inclusion in the report. For example, the subject-specific variant coding sequence set may be filtered to exclude sequences for which the predicted interaction affinity (e.g., binding affinity) is below a predetermined affinity threshold and / or for which the predicted target interaction (e.g., binding to an MHC molecule) will not or is unlikely to occur. In some cases, the filtering is performed to identify a predetermined number and / or scoring of the subject-specific variant coding sequence set. For example, the filtering may be performed to identify 10, 20, 40, 60, 80, 100, 500, or 1,000 variant coding sequences associated with a relatively high predicted probability (e.g., relative to non-selected variant coding sequences in the subject-specific variant coding sequence set) regarding whether a mutant peptide will bind to an MHC molecule.
[0306] The report may identify one or more variant coding sequences (e.g., those not filtered out of the set) and / or one or more mutant peptides (e.g., those associated with the selected variant coding sequences). The mutant peptide may be identified, for example, by its name, its sequence, and / or the variant represented in the corresponding wild-type sequence and the variant coding sequence.
[0307] The report may identify one or more predictions associated with one or more variant coding sequences or one or more mutant peptides. The report may include the name of the subject. For example, the report may be presented locally (e.g., for display on a display system of a user device, sent as a notification on a user device, etc.) and / or transmitted to another device (e.g., sent to a cloud computing system, sent to cloud storage, sent to a user device associated with a medical professional or laboratory professional, transmitted as an email, etc.).
[0308] Fig.11 Illustration of an exemplary table including training data according to some embodiments. Table 1100 includes training data 1102 (e.g., a training data set). The training data 1102 may be Figure 1 an instance that is part of the training data 133 in Fig.10 . The training data 1102 may be an instance that is part of a training data set (such as the training data set described in step 1002 in
[0309] The training data 1102 includes an allotype identifier 1106, a training N-flanking sequence 1108, a training peptide sequence 1110, a training C-flanking sequence 1112, and a training MHC sequence 1114 (e.g., an MHC pseudo-sequence), a binding affinity 1116 (e.g., a normalized binding affinity scaled between 0 and 1), and a presentation indication (e.g., elution likelihood) 1118. The binding affinity 1116 indicates the (e.g., observed) binding affinity for the detection of the binding of the peptide characterized by the training peptide sequence 1110 and the corresponding MHC characterized by the training MHC sequence 1114. The presentation indication 1118 indicates whether the binding or presentation of the MHC to the peptide is detected (or observed).
[0310] Example predictions
[0311] Embodiments of the present disclosure may include determining one or more predictions, including but not limited to immunogenicity, binding affinity, and potential interactions between mutant peptides and MHC molecules.
[0312] Fig.12 An exemplary method 1200 for predicting which therapeutic antibodies are likely to increase the risk of immunogenicity. As Fig.12As shown, an exemplary sequence 1205 of the light chain of a therapeutic antibody may include various amino acid mutations relative to the germline (as indicated by bold letters in square brackets above), as well as various complementarity determining regions (CDRs) (as indicated by carets below the letters). A set of all possible peptides 1210 may be generated using a sliding window (e.g., in the range of 9 to 30 amino acids for a given peptide sequence length), and candidate peptides 1215 within a specified length (e.g., 12 to 19 amino acids) may be identified. For each of the candidate peptides 1215, a binding core may be identified (as indicated by double underlines as used in Fig.12 – see 1220). The set of candidate peptides 1215 may then be filtered to retain only those whose binding cores include mutations (see 1225). By filtering out peptides whose binding cores do not contain any mutations, the method eliminates those peptides that would not be immunogenic due to similarity to human peptides.
[0313] Next, for the set of candidate peptides having binding cores that include mutations 1230, the method determines the frequency 1235 with which the binding cores occur in a database of B cell receptor binding cores that occur in healthy human subjects (e.g., the frequency of obtaining 9-mers from B cell receptors). If the frequency is high, the candidate peptide is filtered out (again eliminating peptides that are unlikely to be immunogenic) so as to retain only those candidate peptides having binding cores that do not occur or occur rarely in the database (see 1240). Next, the presentation likelihood (e.g., elution likelihood) 1245 of those remaining candidate peptides is calculated. The methods and systems discussed in the remainder of this specification (e.g., with respect to Figures 1 to 11 discussed above) may be used to calculate the presentation likelihood. After filtering out any candidate peptides having a negative presentation likelihood, the method identifies the set of unique binding cores from the remaining candidate peptides, resulting in a set 1250 of unique potential presenter binding cores. In some embodiments, the method may continue to count the number of unique binding cores for each allele and calculate the sum for all MHC I alleles and / or MHC II allotypes. The result of this calculation may guide the decision as to whether the therapeutic antibody represents an immunogenic risk in a subject. The risk may take the form of a count of unique, potentially presenting binding cores, or some other scoring, such as the number of uniquely presented binding cores weighted by elution likelihood, optionally in combination with other categorical or numerical information.
[0314] Fig.13 Illustration of exemplary neoantigen candidates (mutated antigens) and corresponding potential neoantigen candidates (mutated peptides) according to some embodiments. When implementing methods such as method 1000, the mutated peptides may be neoantigens.
[0315] For relatively long mutant peptides such as neoantigen candidate 1300, MHC molecules can present multiple epitopes (referred to as neoepitopes) all containing the same mutation or variant. Thus, the immunogenicity of the neoantigen candidate can be predicted based on predictions generated for each of the neoepitope candidates 1302.
[0316] Immunogenicity can be predicted, for example, by generating a list of all possible neoepitopes that might arise from a given neoantigen and generating predictions for each neoepitope candidate in the list (flanked by the remaining amino acids upstream of the N-terminus and downstream of the C-terminus of the epitope that makes up the longest 10 amino acids). Based on these presentation predictions, neoepitope candidates with the greatest presentation likelihood for MHC candidate 1304 are selected to represent the entire neoantigen. Alternatively, a summary score representing the neoantigen can be obtained using a summary representation of multiple candidate neoepitope-MHC pairs. Such a summary can be performed by considering all candidate neoepitope-MHC pairs or by considering the best neoepitopes for each MHC and then summing over all MHC molecules. This summary can be accomplished by several mathematical functions, including, for example, taking the arithmetic mean or harmonic mean of the presentation or binding affinity scores of each candidate neoepitope-HLA pair.
[0317] Although Fig.13 neoantigens and neoepitopes have been described, similar techniques can be used for other types of relatively long mutant peptides that contain mutations or variants and have multiple possible epitope candidates. In some embodiments, the technique can be used in conjunction with antibody drug sequences.
[0318] In some embodiments, when the machine learning model results predict that a mutant peptide will have a low binding affinity for an MHC molecule, it can be predicted that a neoantigen detected from a disease sample of a subject will not trigger immunogenicity or will have low immunogenicity. In some embodiments, it can be predicted that the MHC molecule will not or is unlikely to present the mutant peptide. In some embodiments, it can be predicted that the mutant peptide will not trigger an immune response of a T cell receptor. The immunogenicity prediction generated in association with the mutant peptide can be, for example, numerical (e.g., corresponding to the probability of predicting an immunogenic response triggered in response to the mutant peptide and / or the intensity of the prediction of any immunogenic response to the mutant peptide), categorical (e.g., predicting no immune response, low or high immune response), or binary (e.g., predicting whether a given mutant peptide triggers an immune response in a subject).
[0319] Predicted immunogenicity can be further based on the prediction and / or experimental indication of one or more immunogenicity factors. Factors determining immunogenicity can include one or more of the following: (i) the protein content of the mutant peptide precursor; (ii) the expression level of the transcript encoding the mutant peptide precursor; (iii) the processing efficiency of the mutant peptide precursor by the immunoproteasome; (iv) the expression time of the transcript encoding the mutant peptide precursor; (v) the binding affinity of the mutant peptide to the T cell receptor; (vi) the position of the variant amino acid within the mutant peptide; (vii) the solvent exposure of the mutant peptide when bound to the MHC molecule; (vii) the solvent exposure of the variant amino acid when bound to the MHC molecule; (x) the content of aromatic residues in the peptide; (xi) the properties of the variant amino acid compared to the wild-type residue; (xii) the nature of the mutant peptide precursor; (xiii) the microbial similarity of the mutant peptide to known microbial peptides; (xiv) the self-similarity or dissimilarity of the mutant peptide to the wild-type proteome; or (xv) the thymic expression of the wild-type peptide. Immunogenicity factors can further or alternatively include the protein sequence of the mutant peptide, the length of the mutant peptide (e.g., as indicated by the number of amino acids identified within the variant coding sequence), and / or the expression level of MHC allotypes in the subject (e.g., as measured by RNA-Seq or mass spectrometry).
[0320] For each of a set of mutant peptides (e.g., detected in a disease sample from a subject), a binding affinity prediction and / or a prediction regarding whether (or the probability that) the mutant peptide will be presented (e.g., by one or more tumor cells and / or one or more MHC molecules in the subject) can be generated according to the techniques disclosed herein (e.g., using an attention-based machine learning model). These predictions can be used to select an incomplete subgroup of the set (e.g., less than 50% of the set, less than 25% of the set, less than 10% of the set, less than 5% of the set, and / or less than 1% of the set). One or more relative thresholds (e.g., to identify mutant peptides within the set that have the most stable binding to the MHC molecule and / or the highest presentation likelihood relative to other molecules in the set) or one or more absolute thresholds can be used to select the incomplete subgroup. For example, each selected mutant peptide can have a relatively strong affinity value for the MHC (e.g., within the best 50%, best 25%, best 10%, or best 5% of the affinity values within the set) and / or an absolutely strong affinity value (e.g., having an affinity value better than a predetermined threshold / cutoff value such as 5000 nM, 1000 nM, or 500 nM). The incomplete subgroup of the set can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more mutant peptides that are independent of the predetermined affinity value threshold / cutoff value. The incomplete subset of the set can include 20 or more neoantigens or 30 or more mutant peptides.
[0321] In some cases, a machine learning model generates predictions corresponding to one or more potential interactions between a mutant peptide and an MHC molecule. For example, the machine learning model can predict the binding affinity of an MHC molecule to a mutant peptide. Additionally or alternatively, the machine learning model can predict whether an MHC molecule will present a mutant peptide. The machine learning model can receive as input the sequence or subsequence of an MHC molecule and the variant coding sequence associated with the mutant peptide, and can process them (e.g., using one or more processing layers).
[0322] In some cases, a machine learning model generates predictions corresponding to one or more potential interactions between a mutant peptide, an MHC sequence or subsequence, and a T cell receptor (e.g., instead of or in addition to generating predictions corresponding to one or more potential interactions between a mutant peptide and an MHC molecule). The machine learning model can then predict, for example, the binding affinity between a mutant peptide and a T cell receptor and / or whether the mutant peptide activates and / or triggers an immune response in a T cell. The machine learning model can receive as input, and can process (e.g., using one or more self-attention layers), the sequence or subsequence of a T cell receptor, the sequence or subsequence of an MHC, and the variant coding sequence of a mutant peptide.
[0323] Predictions generated in association with a mutant peptide can be, for example, numerical (e.g., a probability corresponding to a prediction that an MHC molecule of a subject presents the mutant peptide at the cell surface or a score for a prediction that a tumor cell of a subject presents the mutant peptide), categorical (e.g., predicting that an MHC molecule of a subject does not present, infrequently presents, or frequently presents a mutant peptide), or binary (e.g., predicting whether an MHC molecule of a subject expresses a mutant peptide). The presentation prediction can (but need not) be normalized and / or represent a conditional prediction. For example, if the mutant peptide has stably bound to an MHC molecule, the presentation prediction can correspond to a prediction regarding whether the MHC molecule of the subject presents the mutant peptide.
[0324] Example Identification of Input Data for Machine Learning Models
[0325] The exemplary methods and systems described herein for identifying input data can be used to identify input data for, for example, Figure 1 and Figure 3 the machine learning model 132 in Figures 4A to 4D any workflow in Figures 5A to 5C and / or the input data for the machine learning model 532 described in
[0326] A machine learning model can be used to analyze each of a set of mutant peptides associated with a given subject to generate one or more predictions regarding the binding affinity, presentation probability, and / or immunogenicity of the mutant peptides. To generate these predictions, the machine learning model can receive and process the peptide (e.g., encoded) sequence corresponding to the mutant peptide as well as one or more other sequences or subsequences (e.g., corresponding to an MHC-I molecule, an MHC-II molecule, or a T cell receptor). In some cases, predictions are generated for each peptide sequence in a set of peptide sequences (e.g., a set of variant encoded sequences corresponding to a set of mutant peptides). The set of mutant peptides can correspond to peptides present in a disease sample collected from a subject but not observed in one or more non-disease samples (e.g., from the subject or another subject).
[0327] There are a variety of methods for identifying a set of mutant peptides associated with a given subject. Mutations can be present in the genome, transcriptome, proteome, or exome of a subject's diseased cells but not in non-disease samples (e.g., non-disease samples from the subject or from another subject). Mutations include, but are not limited to: (1) non-synonymous mutations that result in a different amino acid in the protein; (2) read-through mutations, where a stop codon is modified or deleted, resulting in a longer protein that is translated with a new tumor-specific sequence at the C-terminus; (3) splice-site mutations that result in the inclusion of an intron in the mature mRNA and thus the formation of a unique tumor-specific protein sequence; (4) chromosomal rearrangements that generate a chimeric protein (i.e., gene fusion) with a tumor-specific sequence at the junction of two proteins; (5) frameshift insertions or deletions that result in a new open reading frame with a new tumor-specific protein sequence. Mutations can also include one or more non-frameshift indels, missense or nonsense substitutions, splice-site alterations, genomic rearrangements, or gene fusions, or any genomic or expression alteration that results in a new ORF.
[0328] Mutant peptides or mutant polypeptides generated by, for example, splice-site, frameshift, read-through, or gene fusion mutations in diseased cells can be identified by sequencing the DNA, RNA, or protein in the disease sample and comparing the obtained sequences to sequences from non-disease samples.
[0329] In some embodiments, whole-genome sequencing (WGS) or whole-exome sequencing (WES) data from disease and non-disease samples can be obtained and compared. After aligning the non-disease and disease sample reads to the human reference genome, somatic variants, including single nucleotide variants (SNVs), gene fusions, and insertion or deletion variants (indels), can be detected using variant calling algorithms. One or more variant callers can be used to detect different types of somatic variants (e.g., SNVs, gene fusions, or insertions or deletions).
[0330] In some examples, mutant peptides are identified based on transcriptome sequences in disease samples from an individual. For example, all or part of the transcriptome sequence can be obtained from a diseased tissue of a subject (e.g., by methods such as RNA-Seq) and subjected to sequencing analysis. The sequence obtained from the diseased tissue sample can then be compared with the sequence obtained from a reference sample. Optionally, whole transcriptome RNA-Seq is performed on the diseased tissue sample. Optionally, the transcriptome sequence is "enriched" with specific sequences before comparison with the reference sample. For example, specific probes can be designed to enrich certain desired sequences (e.g., disease-specific sequences) before performing sequencing analysis.
[0331] In some embodiments, transcriptome sequencing techniques include, but are not limited to, RNA poly(A) libraries, microarray analysis, parallel sequencing, massively parallel sequencing, PCR, and RNA-Seq. RNA-Seq is a high-throughput technique for sequencing part or substantially all of the transcriptome. Briefly, a population of isolated transcriptome sequences is converted into a library of cDNA fragments with adaptors attached to one or both ends. Then, with or without amplification, each cDNA molecule is analyzed to obtain short sequence information, typically 30 to 400 base pairs in length. These sequence fragment information is then aligned with a reference genome, reference transcript, or assembled de novo to reveal the structure (e.g., transcriptional boundaries) and / or expression levels of the transcript.
[0332] Once obtained, the sequence in the diseased sample can be compared with the corresponding sequence in the reference sample. Sequence comparison can be performed at the nucleic acid level by aligning the nucleic acid sequence in the diseased tissue with the corresponding sequence in the reference sample. Genetic sequence variations that result in one or more changes in the encoded amino acids are then identified. Alternatively, sequence comparison can be performed at the amino acid level, i.e., the nucleic acid sequence is first converted to an amino acid sequence via computer simulation before comparison. Amino acid-based methods or nucleic acid-based methods can be used to identify one or more mutations in the peptide (e.g., one or more point mutations). With respect to nucleic acid-based methods, the variants found can be used to identify one or more nucleic acid sequences (e.g., DNA sequence, RNA sequence, or mRNA sequence) that will generate a given observable mutant protein (e.g., via a lookup table that correlates a single peptide mutation with multiple codon variants).
[0333] In some embodiments, comparison of the sequence from the disease sample with the sequence from the reference sample can be done by techniques such as manual alignment, FAST-All (FASTA), or Basic Local Alignment Search Tool (BLAST). In some embodiments, comparison of the sequence from the disease sample with the sequence from the reference sample can be done using a short read aligner (e.g., GSNAP, BWA, or STAR).
[0334] In some embodiments, the reference sample is a matched disease-free sample. As used herein, a "matched" disease-free tissue sample is one selected from the same or a similar sample (e.g., a sample from the same or a similar tissue type as the disease sample). In some embodiments, the matched disease-free tissue and the diseased tissue can be from the same individual. The reference sample described herein can be a disease-free sample from the same individual. In some embodiments, the reference sample is a disease-free sample from a different individual (e.g., an individual not suffering from the disease). In some embodiments, the reference sample is obtained from a population of different individuals. In some embodiments, the reference sample is a database of known genes associated with an organism. In some embodiments, the reference sample can be from a cell line. In some embodiments, the reference sample can be a combination of known genes associated with an organism and genomic information from a matched disease-free sample. In some embodiments, the variant coding sequence can include a point mutation in the amino acid sequence. In some embodiments, the variant coding sequence can include an amino acid deletion or insertion.
[0335] In some embodiments, a set of variant coding sequences is first identified based on genomic and / or nucleic acid sequences. Then, the initial set is further filtered based on the presence (and thus considered to be "expressed") of the variant coding sequences in a transcriptome sequencing database to obtain a narrower set of expressed variant coding sequences. In some embodiments, by filtering through the transcriptome sequencing database, the set of variant coding sequences is reduced by at least about 10, 20, 30, 40, 50 or more fold.
[0336] Alternatively, protein mass spectrometry can be used to identify or verify the presence of mutant (e.g., mutants that bind to MHC proteins on tumor cells) peptides. The peptides can be eluted from diseased cells (e.g., tumor cells) or from HLA molecules from tumor immunoprecipitates by acid washing and then identified using mass spectrometry.
[0337] The mutant peptides can have, for example, 5 or more, 8 or more, 11 or more, 15 or more, 20 or more, 40 or more, 80 or more, 100 or more, 120 or fewer, 100 or fewer, 80 or fewer, 60 or fewer, 50 or fewer, 40 or fewer, 30 or fewer, 25 or fewer, 20 or fewer, 18 or fewer, 15 or fewer, or 13 or fewer amino acids.
[0338] Tumor-specific T cell receptor sequences can also be identified, for example, by single-cell T cell receptor sequencing. High-throughput sequencing of the T cell repertoire can additionally or alternatively be performed to identify tumor-specific features of a particular disease. MHC-I sequences and / or MHC-II sequences can be determined, for example, via HLA genotyping or mass spectrometry.
[0339] Example Identification of Training Data for Machine Learning Models
[0340] The exemplary methods and systems for identifying training data described herein can be used to identify training data for, for example, Figure 1 and Figure 3 the machine learning model 132 in Figures 4A to 4D any workflow in Figures 5A to 5C and / or the training data for the machine learning model 532 described in Figure 1 For example, these methods and systems can be used to identify the training data 133 in
[0341] A training set can be generated using data collected from multiple other samples (e.g., that may be associated with one or more other subjects). Each of the multiple other samples can include, for example, tissue (e.g., a biopsy), single cells, multiple cells, cell debris, or aliquots of body fluids. In some cases, samples are collected from different types of subjects compared to the subject associated with the input data to be processed by the trained model. For example, training data collected by processing samples from one or more cell lines can be used to train a machine learning model, and the trained machine learning model can be used to process input data determined by processing one or more samples from a human subject.
[0342] The training data set includes a plurality of training elements. Each element of the plurality of training elements can include input data that includes a set of peptide sequences (which includes a set of wild-type or variant coding sequences), each sequence encoding and / or representing any variant in the corresponding peptide, and a subsequence or pseudo-sequence of an MHC molecule. The input data can be collected according to one or more techniques disclosed herein.
[0343] Each training element can also include one or more experiment-based results. The experiment-based results can indicate whether and / or to what extent one or more specific types of interactions occur between a wild-type peptide or a mutant peptide (associated with the variant coding sequence in the training element) and an MHC molecule (associated with the MHC molecule subsequence in the training element). Specific types of interactions can include, for example, binding of the peptide to the MHC molecule and / or presentation of the peptide by the MHC molecule on the surface of a cell (e.g., a tumor cell).
[0344] Results can include the binding affinity between a peptide and an MHC molecule. The results can include or be based on qualitative data and / or quantitative data that characterize whether a given peptide binds to a given MHC molecule, the strength of such a bond, the stability of such a bond, and / or the tendency for such a bond to occur. For example, ELISA, pull-down assays, gel shift assays, biosensor-based methods (e.g., surface plasmon resonance, isothermal titration calorimetry, biolayer interferometry, or microscale thermophoresis) can be used to generate binary binding affinity indicators or qualitative binary affinity results.
[0345] The results can, for example, further or alternatively characterize whether a given MHC molecule presents a given peptide and / or the probability of presentation. MHC ligands can be immunoprecipitated from a sample. Subsequent elution and mass spectrometry can be used to determine whether the MHC molecule presents the ligand.
[0346] Example Training Data Filtering
[0347] Fig.17 A plot showing a latent space that includes a plurality of peptide vectors according to some embodiments. Fig.17 Each peptide vector therein corresponds to a BOS token embedding of a peptide sequence in a given sample (e.g., Figure 4A the BOS token embedding 408a therein). In some instances, a sample refers to a row in a dataset and the row specifies a peptide, MHC, TCR, or a combination thereof. In the latent space, each peptide vector has been reduced to two dimensions (e.g., using any dimensionality reduction technique described herein) and is plotted as a single point. The color of each point represents the allele to which the peptide binds.
[0348] As Fig.17 shown, peptides having the same color (i.e., binding to the same allele) are generally close to each other in the latent space and thus form clusters such as cluster 1700. However, in region 1702, many points having the same color appear relatively dispersed and do not form a clear cluster. The dispersed distribution may indicate experimental error in the peptide vector data corresponding to region 1702, because peptides associated with various random alleles should generally not occupy the same region in the latent space.
[0349] The above experimental error can be identified and excluded from the training data described herein (e.g., the training data for processing block 314). Specifically, the system (e.g., Figure 1The computing platform 102) can identify the K nearest neighbors of a given peptide in the latent space and generate a motif. For example, given a set of peptides of the same length (i.e., the K nearest neighbors), the probability of each amino acid at each position in the peptide is calculated, and the probability information is converted into information entropy represented in bits, thereby generating a motif. In other words, each motif indicates the probability of an amino acid occurring at a given position in the peptide. In some embodiments, an information content metric can be extracted for each peptide based on the motif to quantify whether the position is associated with an amino acid occurrence pattern (which can indicate experimental error). In some embodiments, the information content is calculated based on the number of bits of information in the top 2 positions. In some embodiments, data associated with low information content can be filtered out from the training data.
[0350] Fig.17 An exemplary motif 1704 corresponding to the K nearest neighbors in cluster 1700 is shown. As shown in motif 1704, at position 0 (x-axis), most of the space is occupied by I and V, indicating that I and V occur frequently at this position. In contrast, at position 1 or position 2 (x-axis), no single amino acid occupies significantly more space than the others. Thus, position 0 is associated with high information content, while positions 1 and 2 are associated with low information content due to the lack of a binding pattern. The information content of the peptide can then be determined in bits accordingly. In some instances, the position weight matrix (PWM) of the peptide and its nearest neighbor peptides is determined, and the Kullback-Leibler divergence with respect to the baseline PWM calculated from the human peptidome is used to calculate the information content. In some instances, the Shannon entropy is calculated as the information content.
[0351] Fig.18 A histogram according to some embodiments is shown, which shows the count of peptides with different information content levels. In some instances, the input data includes a dataset of peptides and HMC. For each peptide, the information content is calculated (e.g., based on the corresponding peptide and adjacent peptides as described herein). The x-axis refers to the information content of the peptide, which can be quantified in bits. Specifically, for a peptide, the system (e.g., Figure 1 the computing platform 102) can identify the K nearest neighbors of the peptide in the latent space and generate a motif. For example, given a set of peptides of the same length (i.e., the K nearest neighbors), the probability of each amino acid at each position in the peptide is calculated, and the probability information is converted into information entropy (represented in bits) as the information content of the peptide, thereby generating a motif. The y-axis refers to the number of peptides with a specific information content level. The histogram shows a large number of peptides (i.e., 1800) in the dataset with relatively low information content. Low information content indicates a lack of binding pattern (i.e., all positions in the peptide are equally random), and can indicate the presence of experimental errors in the dataset. Thus, those peptides with relatively low information content (e.g., below a threshold) can be filtered out from the training data.
[0352] Fig.19A Shows a protein space colored by protein expression according to some embodiments. In Fig.19A , each point represents a protein vector (e.g., Figure 4B a reduced-dimensional version of the protein sequence embedding 422 in Fig.19A ). Further, blue indicates proteins with lower expression, while red indicates proteins with higher expression. As Fig.19A shown, there is a continuous gradient from blue to red in the main cluster. Fig.19B Demonstrates that although the protein language model (e.g., PLM 444) is not trained using explicit protein expression data, the model still learns this representation of proteins. Fig.19B Shows the cellular compartmentalization of different proteins and their positions in the latent space. Fig.19B Shows how the techniques disclosed herein can predict the source protein space / location by cellular compartments. In some cases, it is desirable to predict peptides that bind to MHC I, and such binding and presentation can occur more frequently when the peptides are derived from source proteins located intracellularly. Conversely, in some other cases, it is desirable to predict peptides that bind to MHC II, and such binding and presentation can occur more frequently when the peptides are derived from source proteins that are predominantly extracellular.
[0353] Fig. 20 Shows exemplary performance data according to some embodiments. Specifically, Fig. 20 shows the average precision values of the baseline algorithm for MHC class I datasets and MHC class II datasets. The baseline algorithm does not incorporate the processing of protein data (e.g., Figure 4B the processing of the protein sequence embedding 422 in Figure 4B ) or the processing of MHC data (e.g., Figure 4B the processing of the MHC sequence embedding 426 in Figure 4A ). Instead, the baseline algorithm uses one or more transformer stages to generate a BOS+MHC sequence representation using the MHC sequence appended with the BOS token, and then uses one or more transformer stages to generate a transformed BOS+MHC sequence representation (e.g., Figure 4A the processing of the MHC sequence appended with the BOS token in Fig. 20 ). As Fig. 20 shown, the incorporation of protein information (e.g., Figure 4B the processing of the protein sequence embedding 422 in Figure 4B ), the incorporation of MHC sequence embedding (e.g., Figure 4B the processing of the MHC sequence embedding 426 in Figure 4B ), and the incorporation of both (e.g., Figure 4B the workflow 420 in
[0354] Exemplary Pharmaceutical Compositions which incorporates both the processing of the protein sequence embedding 422 and the processing of the MHC sequence embedding 426) improves the performance of the baseline algorithm.
[0355] In some embodiments, for each of a set of mutant peptides (e.g., detected in a sample from a subject), one or more of the techniques disclosed herein are used to predict whether the mutant peptide will bind to the subject's MHC molecules (or the strength, stability, and / or incidence of such binding) and / or to predict whether the subject's MHC molecules will present the mutant peptide (and / or the incidence of such presentation). The prediction can be used to select an incomplete subset of the mutant peptides (e.g., for which MHC presentation of the mutant peptide is predicted to be possible). The selection can include, for each mutant peptide, comparing a metric corresponding to the prediction metric to an absolute threshold and / or to the prediction metrics of other mutant peptide metrics (e.g., to make a relative comparison). Each selected mutant peptide can be identified as having one or more of the following: a high likelihood of being present on the surface of a tumor cell, a high likelihood of being able to induce a tumor-specific immune response, a high likelihood of being able to be presented to naive T cells by antigen-presenting cells (e.g., dendritic cells), a low likelihood of being inhibited via central or peripheral tolerance, or a low likelihood of being able to induce an autoimmune response against the subject's normal tissues.
[0356] As a non-limiting example, the selection can include identifying each of the subject-specific variant coding sequences for which the predicted binding affinity is less than 500 nM, for which the predicted MHC molecules will present the mutant peptide identified by the variant coding sequence, and / or for which the predicted mutant peptide will trigger an immune response. It should be understood that the output of the model can be on different scales such that 500 nM can correspond to another value (e.g., 0.42) on a scale such as [0,1].
[0357] Each selected mutant peptide can be manufactured, experimentally tested (e.g., to determine binding affinity, presentation incidence, and / or other immunological factors), included in a composition (e.g., a pharmaceutical composition such as a vaccine and / or a therapy), and / or administered to a subject.
[0358] Each mutant peptide in the set of mutant peptides for which binding affinity and presentation predictions are generated can include a mutant peptide associated with a particular subject (e.g., a particular human subject). Each mutant peptide in the set of mutant peptides can be a disease-specific, immunogenic mutant peptide identified using a disease-specific sample from an individual. The subject variant coding sequences can be identified by sequencing the genetic and / or nucleic acid sequences (e.g., DNA, RNA, and / or mRNA sequences) in the disease sample and comparing each identified genetic and / or nucleic acid sequence to a reference sample sequence. Codons within the genetic and / or nucleic acid sequence indicate the presence of the corresponding amino acid in the peptide. Notably, each of multiple codons can encode a given amino acid, so while the nucleic acid sequence can indicate (e.g., deterministically) the amino acid sequence, the same amino acid sequence can be encoded by other nucleic acid sequences.
[0359] Some embodiments include making a composition based on one or more selected mutant peptides (or multiple nucleic acids encoding one or more selected mutant peptides). For example, each mutant peptide among the one or more selected mutant peptides may have been predicted to bind to and be presented by (e.g., at least to a threshold degree) the MHC molecules of a subject. The composition can include each of the one or more selected mutant peptides, one or more precursors of the one or more selected mutant peptides, one or more polypeptide sequences corresponding to the one or more selected mutant peptides, RNA (e.g., mRNA) corresponding to the one or more mutant peptides, DNA corresponding to the one or more selected mutant peptides, cells (e.g., antigen-presenting cells) containing the one or more selected mutant peptides and / or nucleic acids encoding such peptides, plastids corresponding to the one or more mutant peptides, and / or vectors corresponding to the one or more mutant peptides.
[0360] The composition can include mutant peptides corresponding to a single selected variant coding sequence. The composition can include mutant peptides and / or mutant peptide precursors corresponding to multiple selected variant coding sequences. A subgroup of candidate peptides (e.g., associated with the highest presentation prediction of 5, 10, 15, 20, 30, or any number between them) can be used for further precursor development.
[0361] Each of one or more (e.g., all) of the mutant peptides in the composition can have a length of, for example, about 7 to about 40 amino acids (e.g., any one of about 7, 8, 9, 10, 11, 12, 13, 14, 15, 17, 20, 22, 25, 30, 35, 40, 45, 50, 60, or 70 amino acids). In some embodiments, each of one or more (e.g., all) of the mutant peptides in the composition has a length within a predetermined range (e.g., 8 to 11 amino acids, 8 to 12 amino acids, or 8 to 15 amino acids). In some embodiments, each of one or more (e.g., all) of the mutant peptides in the composition has a length of about 8 to 10 amino acids. Each of one or more (e.g., all) of the mutant peptides in the composition can be in its isolated form. Each of one or more (e.g., all) of the mutant peptides in the composition can be a "long peptide" generated by adding one or more peptides to the ends (or to each end) of the mutant peptide. Each of one or more (e.g., all) of the mutant peptides in the composition can be a labeled fusion protein and / or a hybrid molecule.
[0362] In some embodiments, compositions can be developed by using one or more nucleic acids encoding a peptide. The nucleic acids can include DNA, RNA, and / or mRNA. Given that any one of a variety of codons can encode a given amino acid, codons can be selected, for example, to optimize or facilitate expression in a given type of organism. Such selection can be based on the frequency with which each of a variety of potential codons is used by a given type of organism, the translation efficiency of each of a variety of potential codons in a given type of organism, and / or the degree of bias of a given type of organism towards each of a variety of potential codons.
[0363] The composition can include a polynucleotide construct (e.g., a DNA construct or an RNA construct). A polynucleotide construct is an artificially constructed nucleic acid fragment that can be "transplanted" into a target tissue or cell. The polynucleotide construct contains a DNA or RNA (e.g., mRNA) insert that contains a nucleotide sequence encoding one or more selected mutant peptides. To increase antigen presentation (e.g., presentation of one or more selected mutant peptides by MHC molecules), the polynucleotide construct can further contain modifications developed to improve antigen presentation, thereby improving the immunogenicity of one or more selected mutant peptides. In some cases, the modification is incorporation of the transmembrane and cytoplasmic regions of an MHC molecule chain into the polynucleotide construct.
[0364] To provide an RNA insert with enhanced stability and translation efficiency, the polynucleotide construct can further contain modifications developed for improved stability and translation, thereby improving the immunogenicity of one or more selected mutant peptides. In some cases, the modification includes incorporation of a nucleic acid sequence having at least two copies of the 3'-untranslated region of the human β-globin gene into the polynucleotide construct. In some cases, the modification includes incorporation of a nucleic acid sequence encoding a 3'-untranslated region (such as F1 3'UTR).
[0365] In some cases, the composition can include a nucleic acid encoding the above mutant peptide or a precursor of the mutant peptide. The nucleic acid can include sequences flanking the sequence encoding the mutant peptide (or its precursor). In some cases, the nucleic acid includes epitopes corresponding to more than one selected variant coding sequence. In some cases, the nucleic acid is a DNA having a polynucleotide sequence encoding the above mutant peptide or precursor.
[0366] In some cases, the nucleic acid is RNA. In some cases, the RNA is reverse transcribed from a DNA template having a polynucleotide sequence encoding the mutant peptide or precursor described above. In some cases, the RNA is mRNA. In some cases, the RNA is a modified mRNA. In some cases, the RNA is a modified mRNA (e.g., using protamine mRNA containing a modified 5' cap structure or mRNA containing modified nucleotides to protect the mRNA from degradation). In some embodiments, the RNA is single-stranded mRNA.
[0367] To provide RNA inserts with enhanced stability and expression, the polynucleotide construct can further include modifications developed for improved stability and expression, thereby improving the immunogenicity against one or more selected mutant peptides. In some cases, the modification is the incorporation of a cap, such as a 5'-cap structure, at the ends of the RNA. The cap structure can be the D1 enantiomer of β-S-ARCA.
[0368] To deliver the polynucleotide construct to antigen-presenting cells with high selectivity, the composition can further include cationic liposomes or lipid complexes to improve the uptake of the polynucleotide construct, thereby improving the immunogenicity against one or more selected mutant peptides. In some cases, the composition includes nanoparticles containing the polynucleotide construct. The nanoparticles can be lipid complexes containing one or more lipids such as DOTMA and DOPE.
[0369] The composition can include cells containing the mutant peptide and / or the nucleic acid encoding the mutant peptide described above. The composition can further contain one or more suitable carriers and / or one or more delivery systems for the mutant peptide. In some embodiments, the composition can contain the nucleic acid encoding the mutant peptide. In some cases, the cells containing the mutant peptide and / or the nucleic acid encoding the mutant peptide are non-human cells, such as bacterial cells, protozoan cells, fungal cells, or non-human animal cells. In some cases, the cells containing the mutant peptide and / or the nucleic acid encoding the mutant peptide are human cells. In some cases, the human cells are immune cells. In some cases, the immune cells are antigen-presenting cells (APCs). In some cases, the APCs are professional APCs, such as macrophages, monocytes, dendritic cells, B cells, and microglia. In other cases, the professional APCs are macrophages or dendritic cells. In some cases, the APCs containing the mutant peptide and / or the nucleic acid sequence encoding the mutant peptide are used as cell vaccines to induce CD4+ or CD8+ immune responses. In other cases, the composition used as a cell vaccine includes mutant peptide-specific T cells primed by APCs containing the mutant peptide and / or the nucleic acid sequence encoding the mutant peptide.
[0370] The composition may comprise a pharmaceutical adjuvant, a pharmaceutical excipient, an immunomodulator, a checkpoint protein, an antagonist of PD-1 (e.g., an anti-PD-1 antibody) and / or an antagonist of PD-L1 (e.g., an anti-PD-L1 antibody). An adjuvant refers to any substance that, when incorporated into the composition, can alter the immune response to the mutant peptide. Immunostimulants, for example, can be used to conjugate the adjuvant. An excipient can increase the molecular weight of a specific mutant peptide to enhance activity or immunogenicity, confer stability, increase biological activity and / or increase the serum half-life.
[0371] The pharmaceutical composition can be a vaccine, which can include an individualized vaccine specific to a particular subject (e.g., and potentially developed for a particular subject). For example, a sample from a particular subject can be used to identify MHC sequences, and the composition can be developed for and / or used to treat a particular subject.
[0372] The vaccine can be a nucleic acid vaccine. The nucleic acid can encode a mutant peptide or a precursor of the mutant peptide. The nucleic acid vaccine can include sequences flanking the sequence encoding the mutant peptide (or its precursor). In some cases, the nucleic acid vaccine includes epitopes corresponding to more than one selected variant coding sequence. In some cases, the nucleic acid vaccine is a DNA-based vaccine. In some cases, the nucleic acid vaccine is an RNA-based vaccine. In some cases, the RNA-based vaccine contains mRNA. In some cases, the RNA-based vaccine contains naked mRNA. In some cases, the RNA-based vaccine contains modified mRNA (e.g., using protamine mRNA containing a modified 5' cap structure or mRNA containing modified nucleotides to protect the mRNA from degradation). In some embodiments, the RNA-based vaccine contains single-stranded mRNA.
[0373] Nucleic acid vaccines may include individualized neoantigen-specific therapies manufactured for a particular subject for use as part of a second-generation immunotherapy. The individualized vaccine may have been designed by first detecting mutant peptides in a sample from a particular subject and then predicting for each detected mutant peptide whether the peptide will bind to the MHC of the particular subject, be presented by the MHC, bind to the T cell receptor of the particular subject, and / or trigger an immune response and / or the extent thereof. Based on these predictions, a subgroup of the detected mutant peptides may be selected (e.g., having at least 1, at least 2, at least 3, at least 5, at least 8, at least 10, at least 12, at least 15, at least 18, up to 40, up to 30, up to 25, up to 20, up to 18, up to 15, and / or up to 10 mutant peptides). For each selected mutant peptide, a synthetic mRNA sequence encoding the mutant peptide may be identified. The mRNA vaccine may include mRNA complexed with lipids to form an mRNA-lipid complex (which encodes some or all of the mutant peptides). Administration of the vaccine comprising the mRNA-lipid complex may result in the mRNA stimulating TLR7 and TLR8, triggering activation of T cells by dendritic cells. Additionally, administration may result in the mRNA being translated into mutant peptides, which may then bind to and be presented by MHC molecules and induce a T cell response.
[0374] The composition may include substantially pure mutant peptides, substantially pure precursors thereof, and / or substantially pure nucleic acids encoding mutant peptides or precursors thereof. The composition may include one or more suitable carriers and / or one or more delivery systems to contain the mutant peptides, precursors thereof, and / or nucleic acids encoding mutant peptides or precursors thereof. Suitable carriers and delivery systems include viruses such as systems based on adenovirus, vaccinia virus, retrovirus, herpes virus, adeno-associated virus, or hybrids containing more than one viral element. Non-viral delivery systems include cationic lipids and cationic polymers (e.g., cationic liposomes). In some embodiments, physical delivery such as with a "gene gun" may be used.
[0375] In certain embodiments, the RNA-based vaccine includes an RNA molecule that, in the 5'→3' direction, comprises: (1) a 5' cap; (2) a 5' untranslated region (UTR); (3) a polynucleotide sequence encoding a secretory signal peptide; (4) a polynucleotide sequence encoding one or more mutant peptides generated by cancer-specific somatic mutations present in a tumor sample; (5) a polynucleotide sequence encoding at least a portion of the transmembrane and cytoplasmic domains of a major histocompatibility complex (MHC) molecule; (6) a 3' UTR that includes: (a) the 3' untranslated region of the split aminoterminal enhancer (AES) mRNA or a fragment thereof; and (b) a non-coding RNA of mitochondrially encoded 12S RNA or a fragment thereof; and (7) a poly(A) sequence.
[0376] In certain embodiments, the RNA molecule further comprises a polynucleotide sequence encoding an amino acid linker; wherein the polynucleotide sequence encoding the amino acid linker forms a first linker-epitope module with the first of one or more mutant peptides; and wherein, in the 5'→3' direction, the polynucleotide sequence forming the first linker-epitope module is between the polynucleotide sequence encoding the secretion signal peptide and the polynucleotide sequence encoding at least a portion of the transmembrane and cytoplasmic domains of the MHC molecule. In certain embodiments, the amino acid linker comprises the sequence GGSGGGGSGG. In certain embodiments, the polynucleotide sequence encoding the amino acid linker comprises the sequence GGCGGCUCUGGAGGAGGCGGCUCCGGAGGC.
[0377] In certain embodiments, the RNA molecule further comprises, in the 5'→3' direction: at least a second linker-epitope module, wherein the at least second linker-epitope module comprises a polynucleotide sequence encoding an amino acid linker and a polynucleotide sequence encoding a neoepitope; wherein, in the 5'→3' direction, the polynucleotide sequence forming the second linker-neoepitope module is between the polynucleotide sequence encoding the neoepitope of the first linker-epitope module and the polynucleotide sequence encoding at least a portion of the transmembrane and cytoplasmic domains of the MHC molecule; and wherein the neoepitope of the first linker-epitope module is different from the neoepitope of the second linker-epitope module. In certain embodiments, the RNA molecule comprises 5 linker-epitope modules, wherein the 5 linker-epitope modules each encode a different neoepitope. In certain embodiments, the RNA molecule comprises 10 linker-epitope modules, wherein the 10 linker-epitope modules each encode a different neoepitope. In certain embodiments, the RNA molecule comprises 20 linker-epitope modules, wherein the 20 linker-epitope modules each encode a different neoepitope.
[0378] In certain embodiments, the RNA molecule further comprises a second polynucleotide sequence encoding an amino acid linker, wherein the second polynucleotide sequence encoding the amino acid linker is between the polynucleotide sequence encoding the most distal neoepitope in the 3' direction and the polynucleotide sequence encoding at least a portion of the transmembrane and cytoplasmic domains of the MHC molecule.
[0379] In certain embodiments, the 5' cap comprises the D1 diastereoisomer of the following structure:
[0380]
[0381] In certain embodiments, the 5'UTR comprises the sequence UUCUUCUGGUCCCCACAGACUCAGAGAGAACCCGCCACC. In certain embodiments, the 5'UTR comprises the sequence GGCGAACUAGUAUUCUUCUGGUCCCCACAGACUCAGAGAGAACCCGCCACC.
[0382] In certain embodiments, the secretory signal peptide comprises the amino acid sequence MRVMAPRTLILLLSGALALTETWAGS. In certain embodiments, the polynucleotide sequence encoding the secretory signal peptide comprises the sequence AUGAGAGUGAUGGCCCCCAGAACCCUGAUCCUGCUGCUGUCUGGCGCCCUGGCCCUGACAGAGACAUGGGCCGGAAGC.
[0383] In certain embodiments, at least a portion of the transmembrane and cytoplasmic domains of the MHC molecule comprises the amino acid sequence IVGIVAGLAVLAVVVIGAVVATVMCRRKSSGGKGGSYSQAASSDSAQGSDVSLTA. In certain embodiments, the polynucleotide sequence encoding at least a portion of the transmembrane and cytoplasmic domains of the MHC molecule comprises the sequence AUCGUGGGAAUUGUGGCAGGACUGGCAGUGCUGGCCGUGGUGGUGAUCGGAGCCGUGGUGGCUACCGUGAUGUGCAGACGGAAGUCCAGCGGAGGCAAGGGCGGCAGCUACAGCCAGGCCGCCAGCUCUGAUAGCGCCCAGGGCAGCGACGUGUCACUGACAGCC.
[0384] In certain embodiments, the 3' untranslated region of the AES mRNA contains the sequence CUGGUACUGCAUGCACGCAAUGCUAGCUGCCCCUUUCCCGUCCUGGGUACCCCGAGUCUCCCCCGACCUCGGGUCCCAGGUAUGCUCCCACCUCCACCUGCCCCACUCACCACCUCUGCUAGUUCCAGACACCUCC. In certain embodiments, the non-coding RNA of the mitochondrially encoded 12S RNA contains the sequence CAAGCACGCAGCAAUGCAGCUCAAAACGCUUAGCCUAGCCACACCCCCACGGGAAACAGCAGUGAUUAACCUUUAGCAAUAAACGAAAGUUUAACUAAGCUAUACUAACCCCAGGGUUGGUCAAUUUCGUGCCAGCCACACCG. In certain embodiments, the 3' UTR contains the sequence CUCGAGCUGGUACUGCAUGCACGCAAUGCUAGCUGCCCCUUUCCCGUCCUGGGUACCCCGAGUCUCCCCCGACCUCGGGUCCCAGGUAUGCUCCCACCUCCACCUGCCCCACUCACCACCUCUGCUAGUUCCAGACACCUCCCAAGCACGCAGCAAUGCAGCUCAAAACGCUUAGCCUAGCCACACCCCCACGGGAAACAGCAGUGAUUAACCUUUAGCAAUAAACGAAAGUUUAACUAAGCUAUACUAACCCCAGGGUUGGUCAAUUUCGUGCCAGCCACACCGAGACCUGGUCCAGAGUCGCUAGCCGCGUCGCU.
[0385] In certain embodiments, the poly(A) sequence contains 120 adenine nucleotides.
[0386] In certain embodiments, the RNA vaccine comprises an RNA molecule that, in the 5'→3' direction, comprises: the polynucleotide sequence GGCGAACUAGUAUUCUUCUGGUCCCCACAGACUCAGAGAGAACCCGCCACCAUGAGAGUGAUGGCCCCCAGAACCCUGAUCCUGCUGCUGUCUGGCGCCCUGGCCCUGACAGAGACAUGGGCCGGAAGC; a polynucleotide sequence encoding one or more neoepitopes generated by cancer-specific somatic mutations present in a tumor sample; and the polynucleotide sequence AUCGUGGGAAUUGUGGCAGGACUGGCAGUGCUGGCCGUGGUGGUGAUCGGAGCCGUGGUGGCUACCGUGAUGUGCAGACGGAAGUCCAGCGGAGGCAAGGGCGGCAGCUACAGCCAGGCCGCCAGCUCUGAUAGCGCCCAGGGCAGCGACGUGUCACUGACAGCCUAGUAACUCGAGCUGGUACUGCAUGCACGCAAUGCUAGCUGCCCCUUUCCCGUCCUGGGUACCCCGAGUCUCCCCCGACCUCGGGUCCCAGGUAUGCUCCCACCUCCACCUGCCCCACUCACCACCUCUGCUAGUUCCAGACACCUCCCAAGCACGCAGCAAUGCAGCUCAAAACGCUUAGCCUAGCCACACCCCCACGGGAAACAGCAGUGAUUAACCUUUAGCAAUAAACGAAAGUUUAACUAAGCUAUACUAACCCCAGGGUUGGUCAAUUUCGUGCCAGCCACACCGAGACCUGGUCCAGAGUCGCUAGCCGCGUCGCU。
[0387] In some embodiments, the mutant peptides described herein (e.g., comprising or consisting of an ordered set of amino acids, such as identified by variant coding sequences selected based on the results of the machine learning techniques described herein) can be used to prepare mutant peptide-specific therapeutic agents, such as for antibody therapy. For example, the mutant peptides can be used to generate and / or identify antibodies that specifically recognize the mutant peptides. These antibodies can be used as therapeutic agents. Synthetic short peptides have been used to generate protein-reactive antibodies. One advantage of immunizing with synthetic peptides is that an unlimited amount of pure and stable antigen can be used. This method involves synthesizing short peptide sequences, conjugating them to large carrier molecules, and immunizing a subject with the peptide-carrier molecule. The properties of the antibody depend on the primary sequence information. By carefully selecting the sequence and conjugation method, a good response to the desired peptide can generally be generated. Most peptides can elicit a good response. The advantage of anti-peptide antibodies is that they can be prepared immediately after the amino acid sequence of the mutant peptide is determined and can specifically target a particular region of the protein for antibody production. Selecting and / or screening mutant peptides predicted by a machine learning model for their immunogenicity is likely to result in antibodies that may recognize the native protein in the tumor environment. The mutant peptide can be, for example, 15 or fewer, 18 or fewer, or 20 or fewer, 25 or fewer, or 30 or fewer residues. The mutant peptide can be, for example, 9 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 50 or more, or 70 or more residues. Shorter peptides can improve antibody production.
[0388] Peptide-carrier protein conjugation can be used to facilitate the production of high-titer antibodies. The conjugation methods can include, for example, site-directed conjugation and / or techniques that rely on reactive functional groups in the amino acids, such as -NH2, -COOH, -SH, and the -OH of phenols. Any suitable method used in anti-peptide antibody production can be used with the mutant peptides identified by the methods of the present invention. Two such known methods are the multiple antigen peptide system (MAPs) and the lipid core peptide (LCP method). The advantage of MAPs is that a conjugation method is not required. No carrier protein or linker bond is introduced into the immunized host. A disadvantage is that it is more difficult to control the purity of the peptide. In addition, MAPs can bypass the immune response system in certain hosts. The LCP method is known to provide higher titers than other anti-peptide vaccine systems and may therefore be advantageous.
[0389] The present invention also provides isolated MHC / peptide complexes that comprise one or more mutant peptides identified using the techniques disclosed herein. Such MHC / peptide complexes can be used, for example, to identify antibodies, soluble TCRs or TCR analogs. One type of antibody has been termed a TCR mimic because they are antibodies that bind peptides from tumor-associated antigens in a particular HLA context. This type of antibody has been shown to mediate lysis of cells expressing the complex on their surface and protect mice from implanted cancer cell lines expressing the complex. An advantage of TCR mimics as IgG mAbs is that they can be affinity matured and these molecules are coupled to immune effector functions via the current Fc domain. These antibodies can also be used to target therapeutic molecules to tumors, such as toxins, cytokines or pharmaceutical products.
[0390] Other types of molecules have been developed using mutant peptides such as those selected using the methods of the present invention, the production of antibodies or binding-capable antibody fragments using non-hybridoma-based methods, such as anti-peptide Fab molecules on phage. These fragments can also be conjugated to other therapeutic molecules for tumor delivery, such as anti-peptide MHC Fab-immunotoxin conjugates, anti-peptide MHC Fab-cytokine conjugates and anti-peptide MHC Fab-drug conjugates.
[0391] Exemplary Treatment Methods
[0392] Some embodiments include treating a medical condition (e.g., a tumor) or a disease (e.g., cancer) in an individual by administering to the individual an effective amount of a composition (e.g., a vaccine) comprising one or more selected mutant peptides. The individual can be the same individual from whom the disease sample was collected. In some cases, the vaccine is administered to an individual different from the individual from whom the disease sample was collected. For example, the different individual can be related to the individual from whom the disease sample was collected, have a genetic risk of developing a particular type of cancer, and / or have MHC molecules that have one, several (e.g., all) alleles corresponding to sequences that are the same (or similar) to one or more MHC alleles of the subject from whom the disease sample was collected.
[0393] Some embodiments provide methods of treatment comprising a vaccine, which can be an immunogenic vaccine. In some embodiments, methods for treating a disease such as cancer are provided, which can include administering to an individual an effective amount of a composition described herein, a mutant peptide identified using the techniques disclosed herein, a precursor thereof or a nucleic acid encoding a mutant peptide (or precursor) identified using the techniques described herein.
[0394] In some embodiments, a method for treating a disease (such as cancer) is provided. The method can include collecting a sample (e.g., a blood sample) from a subject. T cells can be isolated and stimulated. Isolation can be performed using, for example, density gradient sedimentation (e.g., and centrifugation), immunomagnetic selection, and / or antibody complex filtration. Stimulation can include, for example, antigen-independent stimulation, which can use mitogens (e.g., PHA or Con A), anti-CD3 antibodies (e.g., to bind to CD3 and activate the T cell receptor complex), and / or anti-CD28 antibodies (e.g., to bind to CD28 and stimulate T cells). One or more mutant peptides can be selected (or may have been selected) for treating the subject (e.g., based on results generated by a machine learning model corresponding to predictions regarding whether each peptide of a set of mutant peptides will bind to an individual's MHC molecules and / or the extent thereof, whether it will be presented by the individual's MHC molecules and / or the extent thereof, and / or whether it will trigger an immune response in the individual and / or the extent thereof). One or more mutant peptides can be selected based on techniques disclosed herein, which include identifying and processing one or more sequence representations associated with the subject (e.g., representations: MHC sequences, sets of variant coding sequences, and / or T cell receptor sequences). One or more of the sequences may have been detected using the sample from which the T cells were isolated or a different sample.
[0395] In some cases, one or more mutant peptides (or precursors thereof) can be used to generate mutant peptide (e.g., neoantigen)-specific T cells. For example, peripheral blood T cells can be isolated from a subject and contacted with one or more mutant peptides to induce a population of mutant peptide-specific T cells that can be administered to the subject. In some examples, the T cell receptor sequences of mutant peptide-reactive T cells can be sequenced. If the sequencing identifies an ordered set of nucleic acids, the codons of each nucleic acid can be translated into amino acids (e.g., via a lookup technique). Once the T cell receptor sequences (e.g., amino acid T cell receptor sequences) are obtained, the T cells can be engineered to include T cell receptors that specifically recognize the mutant peptides. These engineered T cells can then be administered to the subject. In any of the methods provided herein, the T cells can be expanded in vitro or ex vivo prior to administration to the subject. A composition comprising the expanded population of T cells can then be administered (e.g., infused) to the subject.
[0396] In some cases, a method for treating a disease (such as cancer) is provided, which can include administering to a subject, in an effective amount to initiate, activate, and expand T cells in vivo, a composition comprising one or more mutant peptides (or one or more precursors thereof).
[0397] In some embodiments, a method for treating a disease, such as cancer, is provided, which may include administering to an individual an effective amount of a composition comprising a precursor of a mutant peptide selected using the techniques described herein. In some embodiments, an immunogenic vaccine may include a pharmaceutically mutant peptide selected using the techniques described herein. In some embodiments, an immunogenic vaccine may include a pharmaceutically precursor (such as a protein, peptide, DNA, and / or RNA) of a mutant peptide selected using the techniques described herein. In some embodiments, a method for treating a disease, such as cancer, is provided, which may include administering to an individual an effective amount of an antibody that specifically recognizes a mutant peptide selected using the techniques described herein. In some embodiments, a method for treating a disease, such as cancer, is provided, which may include administering to an individual an effective amount of a soluble TCR or TCR analogue that specifically recognizes a mutant peptide selected using the techniques described herein.
[0398] In some embodiments, the cancer is any of the following: cancer, lymphoma, blastoma, sarcoma, leukemia, squamous cell carcinoma, lung cancer (including small cell lung cancer, non-small cell lung cancer, lung adenocarcinoma, and lung squamous cell carcinoma), melanoma, renal cell carcinoma, peritoneal cancer, hepatocellular carcinoma, gastric cancer or stomach cancer (including gastrointestinal cancer), pancreatic cancer, glioblastoma, cervical cancer, ovarian cancer, liver cancer, bladder cancer, hepatoma, breast cancer, colon cancer, colorectal cancer, endometrial cancer or uterine cancer, salivary gland cancer, kidney cancer or renal carcinoma, liver cancer, prostate cancer, vulvar cancer, thyroid cancer, hepatocellular carcinoma, and various types of head and neck cancer, as well as B-cell lymphoma (including low-grade / follicular non-Hodgkin lymphoma (NHL), small lymphocyte (SL) NHL, intermediate-grade / follicular NHL, intermediate-grade diffuse NHL, high-grade immunogenic NHL, high-grade lymphoblastic NHL, high-grade small non-cleaved cell NHL, large mass NHL, mantle cell lymphoma, AIDS-related lymphoma, Waldenström macroglobulinemia, chronic lymphocytic leukemia (CLL), acute lymphocytic leukemia (ALL), hairy cell leukemia, chronic myelogenous leukemia, and post-transplant lymphoproliferative disorder (PTLD), as well as abnormal angiogenesis associated with phakomatosis, edema (such as that associated with a brain tumor), and Meigs syndrome.
[0399] The embodiments disclosed herein may include identifying some or all and / or implementing some or all of an individualized medical strategy. For example, one or more mutant peptides may be selected for use in a vaccine by: determining an MHC sequence and / or a set of variant coding sequences using a sample from an individual; processing a representation of the MHC sequence and the variant coding sequences using a machine learning model disclosed herein (such as an attention-based machine learning model). Then one or more mutant peptides (and / or their precursors) may be administered to the same individual.
[0400] In some embodiments, a method of treating a disease (such as cancer) in an individual is provided, the method comprising: (a) identifying one or more mutant peptides in the individual (e.g., based on the results generated by a machine learning model, which correspond to predictions based on one or more techniques disclosed herein regarding whether each peptide in a set of mutant peptides will bind to the MHC molecules of the individual and / or the extent thereof, whether it will be presented by the MHC molecules of the individual and / or the extent thereof, and / or whether it will trigger an immune response in the individual and / or the extent thereof); (b) synthesizing the identified mutant peptides or one or more precursors of the mutant peptides or nucleic acids encoding the identified peptides or peptide precursors (e.g., polynucleotides such as DNA or RNA); (c) administering to the individual the mutant peptides, mutant peptide precursors, or nucleic acids.
[0401] In some embodiments, a method of treating a disease (such as cancer) in an individual is provided, the method comprising: (a) identifying one or more mutant peptides in the individual (e.g., based on the results generated by a machine learning model, which correspond to predictions based on one or more techniques disclosed herein regarding whether each peptide in a set of mutant peptides will bind to the MHC molecules of the individual and / or the extent thereof, whether it will be presented by the MHC molecules of the individual and / or the extent thereof, and / or whether it will trigger an immune response in the individual and / or the extent thereof); (b) optionally, identifying a set of nucleic acids (e.g., polynucleotides such as DNA or RNA) encoding the identified mutant peptides or one or more precursors of the mutant peptides, synthesizing the set of nucleic acids, and administering the set of nucleic acids to the subject.
[0402] In some embodiments, a method of treating a disease (such as cancer) in an individual is provided, the method comprising: (a) identifying one or more mutant peptides in the individual (e.g., based on the results generated by a machine learning model, which correspond to predictions based on one or more techniques disclosed herein regarding whether each peptide in a set of mutant peptides will bind to the MHC molecules of the individual and / or the extent thereof, whether it will be presented by the MHC molecules of the individual and / or the extent thereof, and / or whether it will trigger an immune response in the individual and / or the extent thereof); (b) generating antibodies that specifically identify the mutant peptides; (c) administering the peptides to the subject.
[0403] The methods provided herein can be used to treat an individual (e.g., a human) diagnosed with or suspected of having cancer. In some embodiments, the individual can be a human. In some embodiments, the individual can be at least about 18, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or 85 years old. In some embodiments, the individual can be male. In some embodiments, the individual can be female. In some embodiments, the individual may have refused surgery. In some embodiments, the individual may be medically inoperable. In some embodiments, the individual can be at a clinical stage of Ta, Tis, T1, T2, T3a, T3b, or T4. In some embodiments, the cancer can be recurrent. In some embodiments, the individual can be a person exhibiting one or more symptoms associated with cancer. In some embodiments, the subject can be genetically or otherwise predisposed (e.g., having risk factors) to develop cancer.
[0404] The methods provided herein can be implemented in an adjuvant setting. In some embodiments, the method is implemented in a neoadjuvant setting, i.e., the method can be performed prior to primary / definitive therapy. In some embodiments, the method is used to treat an individual who has been previously treated. Any of the treatment methods provided herein can be used to treat an individual who has not been previously treated. In some embodiments, the method is used as a first-line therapy. In some embodiments, the method is used as a second-line therapy.
[0405] In some embodiments, methods are provided for reducing the incidence or burden of pre-existing cancer tumor metastases (such as lung metastases or lymph node metastases) in an individual, comprising administering to the individual an effective amount of a composition disclosed herein. In some embodiments, a method is provided for prolonging the time to cancer disease progression in an individual, comprising administering to the individual an effective amount of a composition disclosed herein. In some embodiments, a method is provided for prolonging the survival of an individual with cancer, comprising administering to the individual an effective amount of a composition disclosed herein.
[0406] In some embodiments, in addition to the compositions disclosed herein, at least one or more chemotherapeutic agents can also be administered. In some embodiments, the one or more chemotherapeutic agents can (but do not necessarily) belong to different classes of chemotherapeutic agents.
[0407] In some embodiments, a method of treating a disease (such as cancer) in a subject is provided, the method comprising administering: (a) a vaccine as disclosed herein (e.g., which comprises a mutant peptide or a precursor thereof selected based on the machine learning techniques disclosed herein); and (b) an immunomodulator. In some embodiments, a method of treating a disease (such as cancer) in a subject is provided, the method comprising administering: (a) a vaccine as disclosed herein (e.g., which comprises a mutant peptide or a precursor thereof selected based on the machine learning techniques disclosed herein); and (b) an antagonist of a checkpoint protein. In some embodiments, a method of treating a disease (such as cancer) in a subject is provided, the method comprising administering: (a) a vaccine as disclosed herein (e.g., which comprises a mutant peptide or a precursor thereof selected based on the machine learning techniques disclosed herein); and (b) an antagonist of programmed cell death 1 (PD-1), such as anti-PD-1. In some embodiments, a method of treating a disease (such as cancer) in a subject is provided, the method comprising administering: (a) a vaccine as disclosed herein (e.g., which comprises a mutant peptide or a precursor thereof selected based on the machine learning techniques disclosed herein); and (b) an antagonist of programmed death ligand 1 (PD-L1), such as anti-PD-L1. In some embodiments, a method of treating a disease (such as cancer) in a subject is provided, the method comprising administering: (a) a vaccine as disclosed herein (e.g., which comprises a mutant peptide or a precursor thereof selected based on the machine learning techniques disclosed herein); and (b) an antagonist of cytotoxic T lymphocyte-associated protein 4 (CTLA-4), such as anti-CTLA-4.
[0408] It should be understood that the various disclosures relate to the use of amino acid sequences. Nucleic acid sequences may be used additionally or alternatively. For example, a disease-specific sample may be sequenced to identify a set of nucleic acid sequences that are not present in a corresponding non-disease-specific sample (e.g., from the same subject or a different subject). Similarly, nucleic acid sequences of MHC molecules and / or T cell receptors may be further identified. Representations of either the nucleic acid disease-specific sequences and MHC molecules (or T cell receptors) may be processed by an attention-based model as described herein (e.g., and may have been trained using nucleic acid sequence representations).
[0409] Exemplary Model Performance
[0410] An exemplary peptide-MHC (MHC class II) machine learning model (the "P-MHC-II model" herein) has been developed. This model is an exemplary implementation of Figure 1 the machine learning model 132 in Figure 5Aimplemented using the architecture shown in. The P-MHC-II model is compared with other previously available models (e.g., NetMHCpan-4.0) (referred to herein as "Model A"). The P-MHC-II model performs better than Model A in peptide presentation.
[0411] Fig.14A and 14B is a graph with an exemplary precision-recall (PR) curve according to some embodiments. Fig.14A and 14B shows the performance of the P-MHC-II model compared to Model A. The eluted ligand (EL) test dataset is used to evaluate the presentation prediction performance between the EL output of the P-MHC-II model and the EL output of Model A.
[0412] Fig.14A includes exemplary graph 1300 indicating the performance of the exemplary P-MHC-II model according to some embodiments. Fig. 14B includes exemplary graph 1402 indicating the performance of a previously used method (Model A) regarding its elution output according to some embodiments. The points on the curves in each of FIGS. 1400 and 1402 correspond to scoring thresholds for the top 10.00% and 9.64% percentiles of the scores, respectively. The average precision (AP) represents the performance independent of the threshold. The F1 score, precision, and recall values are based on the corresponding thresholds.
[0413] The Model A values are percentile rank outputs from the previously used method. The P-MHC-II model values are taken from the output of the P-MHC-II model (of the final node). Based on these PR curves, Fig.14A and 14B the results in indicate that the P-MHC-II model shows improved performance compared to Model A, with an AP value of 0.84 for this model while the AP value of Model A is 0.66. The AP value of the method is compared based on each allele.
[0414] Fig.15 is exemplary graph 1500 of exemplary average precision values comparing the eluted ligand outputs of Model A and the P-MHC-II model for each allele in the test dataset according to some embodiments.
[0415] Figures 16A to 16B are exemplary graphs 1600 and 1602 respectively showing the performance of the P-MHC-II model (BA output) and Model A (BA output) according to some embodiments.
[0416] Exemplary Computer System
[0417] Fig.21A block diagram of a computer system according to some embodiments. The computer system 2100 may be Figure 1 An example of an implementation of the computing platform 102 described above in
[0418] Fig.21 An example of one or more computing devices 2100 that can be used to determine a predicted amino acid-IPC prediction according to some embodiments is shown. In certain embodiments, one or more computing devices 2100 may perform one or more steps of one or more of the methods described or shown herein. In certain embodiments, one or more computing devices 2100 provide the functionality described or shown herein. In certain embodiments, software running on one or more computing devices 2100 performs one or more steps of one or more of the methods described or shown herein, or provides the functionality described or shown herein. Certain embodiments include one or more portions of one or more computing devices 2100.
[0419] The present disclosure contemplates any suitable number of computing systems 2100. The present disclosure contemplates one or more computing devices 2100 in any suitable physical form. By way of example and not limitation, one or more computing devices 2100 may be an embedded computing system, a system-on-chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-module (COM) or a system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a computer system network, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more thereof. In appropriate cases, one or more computing devices 2100 may: be integrated or distributed; span multiple locations; span multiple machines; span multiple data centers; or reside in the cloud, which may include one or more cloud components in one or more networks.
[0420] In appropriate cases, one or more computing devices 2100 may perform one or more steps of one or more of the methods described or shown herein without substantial spatial or temporal limitations. By way of example and not limitation, one or more computing devices 2100 may perform one or more steps of one or more of the methods described or shown herein in real time or in batch mode. In appropriate cases, one or more computing devices 2100 may perform one or more steps of one or more of the methods described or shown herein at different times or in different locations.
[0421] In certain embodiments, one or more computing devices 2100 include a processor 2102, a memory 2104, a database 2106, an input / output (I / O) interface 2108, a communication interface 2110, and a bus 2112. Although this disclosure describes and shows a particular computer system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement. In certain embodiments, the processor 2102 includes hardware for executing instructions, such as those that make up a computer program. By way of example and not limitation, to execute instructions, the processor 2102 may: retrieve (or fetch) instructions from an internal register, an internal cache, the memory 2104, or the database 2106; decode and execute those instructions; and then write one or more results to an internal register, an internal cache, the memory 2104, or the database 2106. In certain embodiments, the processor 2102 may include one or more internal caches for data, instructions, or addresses. In appropriate instances, this disclosure contemplates a processor 2102 including any suitable number of any suitable internal caches. By way of illustration and not limitation, the processor 2102 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). The instructions in the instruction cache may be copies of the instructions in the memory 2104 or the database 2106, and the instruction cache may accelerate the retrieval of these instructions by the processor 2102.
[0422] The data in the data cache may be: a copy of the data in the memory 2104 or the database 2106 for operations by instructions executed at the processor 2102; the result of a previous instruction executed at the processor 2102 for access or writing to the memory 2104 or the database 2106 by a subsequent instruction executed at the processor 2102; or other suitable data. The data cache may accelerate the read or write operations of the processor 2102. The TLB may accelerate the virtual address translation of the processor 2102. In certain embodiments, the processor 2102 may include one or more internal registers for data, instructions, or addresses. In appropriate instances, this disclosure contemplates a processor 2102 including any suitable number of any suitable internal registers. In appropriate instances, the processor 2102 may include one or more arithmetic logic units (ALUs); may be a multi-core processor; or may include one or more processors 2102. Although this disclosure describes and shows a particular processor, this disclosure contemplates any suitable processor.
[0423] In some embodiments, the memory 2104 includes a main memory that stores instructions for execution by the processor 2102 or data for the processor 2102 to operate on. By way of example and not limitation, one or more computing devices 2100 may load instructions from a database 2106 or another source (such as, for example, another one or more computing devices 2100) into the memory 2104. Then, the processor 2102 may load instructions from the memory 2104 into internal registers or an internal cache. To execute the instructions, the processor 2102 may retrieve the instructions from the internal registers or the internal cache and decode the instructions. During or after the execution of the instructions, the processor 2102 may write one or more results (which may be intermediate results or final results) to the internal registers or the internal cache. Then, the processor 2102 may write one or more of those results to the memory 2104.
[0424] In some embodiments, the processor 2102 executes instructions only in one or more internal registers, the internal cache, or the memory 2104 (and not in the database 2106 or elsewhere) and operates on data only in one or more internal registers, the internal cache, or the memory 2104 (and not in the database 2106 or elsewhere). One or more memory buses (which may each include an address bus and a data bus) may couple the processor 2102 to the memory 2104. The bus 2112 may include one or more memory buses, as described below. In some embodiments, one or more memory management units (MMUs) reside between the processor 2102 and the memory 2104 and facilitate access to the memory 2104 requested by the processor 2102. In some embodiments, the memory 2104 includes random access memory (RAM). In appropriate cases, the RAM may be volatile memory. In appropriate cases, the RAM may be dynamic RAM (DRAM) or static RAM (SRAM). Additionally, in appropriate cases, the RAM may be single-port or multi-port RAM. The present disclosure contemplates any suitable RAM. In appropriate cases, the memory 2104 may include one or more memory devices 2104. Although the present disclosure describes and illustrates particular memories, the present disclosure contemplates any suitable memory.
[0425] In some embodiments, database 2106 includes a mass storage device for data or instructions. By way of example and not limitation, database 2106 can include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of them. In appropriate cases, database 2106 can include removable or non-removable (or fixed) media. In appropriate cases, database 2106 can be internal or external to one or more computing devices 2100. In some embodiments, database 2106 is non-volatile solid-state memory. In some embodiments, database 2106 includes read-only memory (ROM). In appropriate cases, the ROM can be mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically rewritable ROM (EAROM), flash memory, or a combination of two or more of them. The present disclosure contemplates a mass database 2106 in any suitable physical form. In appropriate cases, database 2106 can include one or more storage control units that facilitate communication between processor 2102 and database 2106. In appropriate cases, database 2106 can include one or more databases 2106. Although the present disclosure describes and illustrates particular storage devices, the present disclosure contemplates any suitable storage device.
[0426] In some embodiments, I / O interface 2108 includes hardware, software, or both that provide one or more interfaces for communication between one or more computing devices 2100 and one or more I / O devices. In appropriate cases, one or more of the computing devices 2100 can include one or more of these I / O devices. One or more of these I / O devices can enable communication between a person and one or more computing devices 2100. By way of example and not limitation, I / O devices can include a keyboard, a keypad, a microphone, a monitor, a mouse, a printer, a scanner, a speaker, a still camera, a stylus, a tablet computer, a touch screen, a trackball, a video camera, another suitable I / O device, or a combination of two or more of them. The I / O devices can include one or more sensors. The present disclosure contemplates any suitable I / O devices and any suitable I / O interface 2108 for them. In appropriate cases, I / O interface 2108 can include one or more device or software drivers such that processor 2102 can drive one or more of these I / O devices. In appropriate cases, I / O interface 2108 can include one or more I / O interfaces 2108. Although the present disclosure describes and illustrates particular I / O interfaces, the present disclosure encompasses any suitable I / O interface.
[0427] In some embodiments, communication interface 2110 includes hardware, software, or both that provide one or more interfaces for communication (such as, for example, packet-based communication) between one or more computing devices 2100 and one or more other computing devices 2100 or one or more networks. By way of example and not limitation, communication interface 2110 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network, or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network (such as a WI-FI network). The present disclosure contemplates any suitable network and any suitable communication interface 2110 therefor.
[0428] By way of example and not limitation, one or more computing devices 2100 may communicate with an ad hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), one or more portions of the Internet, or a combination of two or more of these. One or more portions of one or more of these networks may be wired or wireless. By way of example, one or more computing devices 2100 may communicate with a wireless PAN (WPAN) (such as, for example, a BLUETOOTH WPAN), a WI-FI network, a WI-MAX network, a cellular telephone network (such as, for example, a Global System for Mobile Communications (GSM) network), other suitable wireless networks, or a combination of two or more of these. In appropriate instances, one or more computing devices 2100 may include any suitable communication interface 2110 for any of these networks. In appropriate instances, communication interface 2110 may include one or more communication interfaces. Although the present disclosure describes and shows particular communication interfaces, the present disclosure contemplates any suitable communication interface.
[0429] In some embodiments, bus 2112 includes hardware, software, or both that couple components of one or more computing devices 2100 to each other. By way of example and not limitation, bus 2112 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an INFINIBAND interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, another suitable bus, or a combination of two or more of these. In appropriate instances, bus 2112 may include one or more buses 2112. Although the present disclosure describes and shows particular buses, the present disclosure contemplates any suitable bus.
[0430] In this document, one or more computer-readable non-transitory storage media may include one or more semiconductor-based or other integrated circuits (ICs) (such as, for example, field-programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard disk drives (HHDs), optical discs, optical disc drives (ODDs), magneto-optical discs, magneto-optical disc drives, floppy disks, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM drives, Secure Digital cards or drives, any other suitable computer-readable non-transitory storage medium, or any suitable combination of two or more thereof. In appropriate cases, the computer-readable non-transitory storage medium may be a volatile storage medium, a non-volatile storage medium, or a combination of a volatile storage medium and a non-volatile storage medium.
[0431] Fig. 22 FIG. 2200 shows an exemplary artificial intelligence (AI) architecture 2202 according to the disclosed embodiments, which may be included as part of one or more computing devices 2100 as discussed above Fig.21 and may be used to determine one or more predicted amino acid-IPC interactions. In certain embodiments, the AI architecture 2202 may be implemented using, for example, one or more processing devices, which may include hardware (e.g., general-purpose processors, graphics processing units (GPUs), application-specific integrated circuits (ASICs), systems-on-a-chip (SoCs), microcontrollers, field-programmable logic gate arrays (FPGAs), central processing units (CPUs), application processors (APs), vision processing units (VPUs), neural processing units (NPUs), neural decision processors (NDPs), deep learning processors (DLPs), tensor processing units (TPUs), neuromorphic processing units (NPUs), and / or other processing devices suitable for processing various molecular data and making one or more decisions based thereon), software (e.g., instructions running / executing on one or more processing devices), firmware (e.g., microcode), or some combination thereof.
[0432] In certain embodiments, as Fig. 22As depicted, the AI architecture 2202 may include machine learning (ML) algorithms and functions 2204, natural language processing (NLP) algorithms and functions 2206, expert systems 2208, computer-based vision algorithms and functions 2210, speech recognition algorithms and functions 2212, planning algorithms and functions 2214, and robotic algorithms and functions 2216. In some embodiments, the ML algorithms and functions 2204 may include any statistics-based algorithms that may be suitable for finding patterns in large amounts of data (e.g., “big data” such as genomic data, proteomic data, metabolomic data, metagenomic data, transcriptomic data, or other omics data). For example, in some embodiments, the ML algorithms and functions 2204 may include deep learning algorithms 2218, supervised learning algorithms 2220, and unsupervised learning algorithms 2222.
[0433] In some embodiments, the deep learning algorithms 2218 may include any artificial neural network (ANN) that may be used to learn deep representations and abstract concepts from large amounts of data. For example, the deep learning algorithms 2218 may include ANNs such as perceptrons, multi-layer perceptrons (MLPs), autoencoders (AEs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), long short-term memory (LSTMs), gated recurrent units (GRUs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), and deep Q-networks, neural autoregressive distribution estimators (NADEs), adversarial networks (ANs), attention models (AMs), spiking neural networks (SNNs), deep reinforcement learning, etc.
[0434] In some embodiments, the supervised learning algorithms 2220 may include any algorithms that may be used to apply, for example, what has been learned in the past using labeled examples to new data to predict future events. For example, starting from the analysis of a known training data set, the supervised learning algorithms 2220 may produce an inference function to make predictions of output values. The supervised learning algorithms 2220 may also compare their output with the correct and expected output and find errors in order to modify the supervised learning algorithms 2220 accordingly. On the other hand, the unsupervised learning algorithms 2222 may include, for example, any algorithms that may be applied when the data used to train the unsupervised learning algorithms 2222 is neither classified nor labeled. For example, the unsupervised learning algorithms 2222 may study and analyze how a system infers a function that describes a hidden structure from unlabeled data.
[0435] In certain embodiments, the NLP algorithms and functions 2206 may include any algorithms or functions suitable for automatically manipulating natural language, such as speech and / or text. For example, the NLP algorithms and functions 2206 may include a content extraction algorithm or function 2224, a classification alg...
Claims
1. A computer-implemented method for predicting amino acid-immune protein complex (IPC) interactions, the computer-implemented method comprising: Accessing a set of amino acid sequences, each of the amino acid sequences in the set having been identified from at least one protein; Accessing an IPC sequence identified for an immune protein complex (IPC) of a subject; Using one or more first processing blocks in a processing subsystem of a machine learning model to process a set of amino acid sequence representations to generate a set of transformed amino acid sequence representations based on a set of element focus scores representing a binding core of the set of amino acid sequence representations, wherein each of the amino acid sequence representations is generated based on one of the amino acid sequences appended with a beginning-of-sequence (BOS) token; Using a second processing block in the processing subsystem to process the IPC sequence representation to generate a transformed IPC sequence representation, wherein the IPC sequence representation is generated based on the identified IPC sequence appended with a BOS token, and wherein the set of amino acid sequence representations and the IPC sequence representation are processed in parallel; Generating a synthetic representation by combining each of the transformed BOS token representations of the set of transformed amino acid sequence representations with the transformed BOS token representation of the transformed IPC sequence representation; and Determining one or more predicted amino acid-IPC interactions based on the synthetic representation.
2. The computer-implemented method according to claim 1, wherein the set of transformed amino acid sequence representations comprises a set of amino-terminal flank (N-flank) representations or a set of carboxyl-terminal flank (C-flank) representations.
3. The computer-implemented method according to claim 1, wherein the IPC of the subject is a major histocompatibility complex (MHC) comprising MHC class I (MHC-I) and / or MHC class II (MHC-II) or is a T cell receptor (TCR), and wherein the at least one protein is a therapeutic protein or is present in a disease sample from the subject.
4. The computer-implemented method according to claim 1, wherein generating the synthetic representation comprises: For each of the set of transformed amino acid sequence representations, element-wise multiplying the transformed amino acid sequence beginning (BOS) representation corresponding to the transformed amino acid sequence representation by the transformed IPC sequence beginning (BOS) representation corresponding to the transformed IPC sequence representation.
5. The computer-implemented method according to claim 1, wherein processing the set of amino acid sequence representations comprises processing a peptide sequence beginning (BOS) representation to generate a transformed peptide sequence representation, and wherein processing the IPC sequence representation comprises processing an MHC sequence beginning representation (BOS) to generate a transformed MHC sequence representation.
6. The computer-implemented method according to claim 1, wherein processing the set of amino acid sequence representations comprises: Convert the amino acid sequence representation set using the one or more first processing blocks, wherein each of the one or more first processing blocks includes a set of processing sub - blocks.
7. The computer - implemented method according to claim 1, wherein processing the IPC sequence representation includes: Convert the IPC sequence representation using the second processing block, wherein the second processing block includes a set of processing sub - blocks.
8. The computer - implemented method according to claim 1, wherein the machine - learning model includes one or more Transformer encoders, and each of the one or more Transformer encoders includes a processing layer.
9. The computer - implemented method according to claim 1, wherein the amino acid sequence representation set includes an aggregated sequence representation, and the aggregated sequence representation includes a set of peptide representations and one or more of the following: a set of amino - terminal flanking (N - flanking) representations or a set of carboxyl - terminal flanking (C - flanking) representations.
10. The computer - implemented method according to claim 1, further comprising, before generating the converted amino acid sequence representation set: Flatten the aggregated sequence representation into a single array; and Densify the aggregated sequence representation by removing blank lines from the array, wherein the converted amino acid sequence representation is generated based on the densified aggregated sequence representation.
11. The computer - implemented method according to claim 1, wherein processing the amino acid sequence representation set includes: For each amino acid sequence representation in the set: For each element of the amino acid sequence representation, determine a plurality of vectors based on a set of weights associated with the processing layer of the machine - learning model; and Generate the set of element - focused scores based on the plurality of vectors and the set of weights.
12. The computer - implemented method according to claim 1, wherein the one or more first processing blocks and the second processing block include attention blocks, each attention block includes a set of attention sub - blocks, and each attention sub - block includes a self - attention layer; and wherein the machine - learning model is an attention - based machine - learning model, and The method further includes: Generate an attention map including one or more masks through one or more of the attention blocks, and the one or more masks limit the attention applied by the attention sub - blocks to the sequence lengths according to the masks.
13. The computer-implemented method according to claim 12, the method further comprising: Calculate a plurality of average attention values corresponding to the plurality of peptide positions by calculating the average attention value at each peptide position among the plurality of peptide positions based on a set of peptides with a uniform length distribution; and Subtract the plurality of average attention values from the masks in the one or more masks.
14. The computer - implemented method according to claim 1, further comprising obtaining a data set for training the machine - learning model by: Generating a plurality of converted peptide representations for a plurality of training peptides; Obtaining a corresponding training peptide cluster for each training peptide based on the plurality of converted peptide representations; Calculating the information content for each training peptide based on the corresponding training peptide cluster; and Exclude the one or more training peptides from the training data based on the corresponding information content of the one or more training peptides.
15. The computer-implemented method according to claim 1, further comprising: Processing the combined representation using a fully-connected block in the output subsystem of the machine learning model to generate a first output; Applying dropout to the first output using a dropout block in the output subsystem of the machine learning model to generate a second output; And Selecting a subgroup of the second output using a max layer in the output subsystem of the machine learning model to generate a result, Wherein the one or more predicted amino acid-IPC interactions are determined based on the result.
16. The computer-implemented method according to claim 1, wherein the amino acid sequence group comprises peptide sequences and the IPC sequence comprises a major histocompatibility complex (MHC) sequence, and Wherein the one or more predicted amino acid-IPC interactions comprise one or more of the following: An interaction affinity prediction for a peptide-IPC combination, which predicts the binding affinity between a peptide and an MHC; An interaction prediction for the peptide-IPC combination, which predicts whether the MHC will present the peptide at the cell surface; or An immunogenicity prediction for the peptide-IPC combination, which predicts the ability of the peptide to elicit an immune response against the MHC.
17. The computer-implemented method according to claim 1, wherein determining the one or more predicted amino acid-IPC interactions comprises: Processing the synthetic representation to generate a set of results; And Selecting an amino acid-IPC combination based on the highest result in the set of results.
18. The computer-implemented method according to claim 1, wherein the one or more predicted amino acid-IPC interactions comprise a prediction of the tumor-specific immunogenicity of a peptide.
19. The computer-implemented method according to claim 1, wherein the amino acid sequence group comprises a set of peptide sequences, and wherein the one or more predicted amino acid-IPC interactions identify a subgroup of peptide sequences having increased tumor-specific immunogenicity or increased likelihood of being presented by the IPC relative to the set of peptide sequences.
20. The computer-implemented method according to claim 1, further comprising: Identifying a subgroup of peptides from the amino acid sequence group based on the determined one or more predicted amino acid-IPC interactions for inclusion in an individualized vaccine, for inclusion as a target for immunotherapy, and / or for exclusion as a target for immunotherapy.
21. The computer-implemented method according to claim 1, further comprising: Accessing a protein sequence corresponding to the at least one protein; Obtaining a protein sequence embedding based on the protein sequence; And Determining the one or more predicted amino acid-IPC interactions at least in part based on the protein sequence embedding.
22. The computer-implemented method according to claim 1, wherein the protein language model comprises a pre-trained protein language model.
23. The computer-implemented method according to claim 22, further comprising: Reducing the dimension of the protein sequence embedding; And Combining each of the transformed BOS token representations in the transformed amino acid sequence representation group with the transformed BOS token representation of the transformed IPC sequence representation and the dimension-reduced protein sequence embedding.
24. The computer-implemented method according to claim 23, wherein the dimension of the protein sequence embedding is reduced via a neural network.
25. A system for predicting amino acid-immunoprotein complex (IPC) interactions, the system comprising: One or more non-transitory computer-readable storage media, which comprise instructions; And One or more processors coupled to the one or more storage media, the one or more processors being configured to execute the instructions to: Access a set of amino acid sequences, each of the amino acid sequences in the set having been identified from at least one protein; Access an IPC sequence identified for an immunoprotein complex (IPC) of a subject; Use one or more first processing blocks in a processing subsystem of a machine learning model to process a set of amino acid sequence representations to generate a set of transformed amino acid sequence representations based on a set of element focus scores representing a binding core of the amino acid sequence representation group, wherein each of the amino acid sequence representations is generated based on one of the amino acid sequences based on an additional ordered sequence start (BOS) token; Use a second processing block in the processing subsystem to process the IPC sequence representation to generate a transformed IPC sequence representation, wherein the IPC sequence representation is generated based on the identified IPC sequence with a BOS token appended, and wherein the set of amino acid sequence representations and the IPC sequence representation are processed in parallel; Generate a synthetic representation by combining each of the transformed BOS token representations in the set of transformed amino acid sequence representations with the transformed BOS token representation of the transformed IPC sequence representation; and Determine one or more predicted amino acid-IPC interactions based on the synthetic representation.
26. A non-transitory computer-readable medium for a system for predicting amino acid-immunoprotein complex (IPC) interactions, the non-transitory computer-readable medium comprising instructions that, when executed by one or more processors of one or more computing devices, cause the one or more processors to: Access a set of amino acid sequences, each of the amino acid sequences in the set having been identified from at least one protein; Access an IPC sequence identified for an immunoprotein complex (IPC) of a subject; Processing, by one or more first processing blocks in a processing subsystem using a machine learning model, a set of amino acid sequence representations to generate a set of transformed amino acid sequence representations based on a set of element focus scores representing a binding core of the set of amino acid sequence representations, wherein each of the amino acid sequence representations is generated based on one of the amino acid sequences with an appended beginning-of-sequence (BOS) token; Processing, by a second processing block in the processing subsystem, an IPC sequence representation to generate a transformed IPC sequence representation, wherein the IPC sequence representation is generated based on the identified IPC sequences with appended BOS tokens, and wherein the set of amino acid sequence representations and the IPC sequence representation are processed in parallel; Generating a synthetic representation by combining each of the transformed BOS token representations of the set of transformed amino acid sequence representations with the transformed BOS token representation of the transformed IPC sequence representation; and Determining one or more predicted amino acid-IPC interactions based on the synthetic representation.
27. A vaccine comprising: One or more peptides; Multiple nucleic acids encoding the one or more peptides; or Multiple cells expressing the one or more peptides, wherein the one or more peptides are selected from a set of peptides by: Accessing a set of amino acid sequences, each of which in the set has been identified from at least one protein; Accessing IPC sequences identified for an immune protein complex (IPC) of a subject; Processing, by one or more first processing blocks in a processing subsystem using a machine learning model, a set of amino acid sequence representations to generate a set of transformed amino acid sequence representations based on a set of element focus scores representing a binding core of the set of amino acid sequence representations, wherein each of the amino acid sequence representations is generated based on one of the amino acid sequences with an appended beginning-of-sequence (BOS) token; Processing, by a second processing block in the processing subsystem, an IPC sequence representation to generate a transformed IPC sequence representation, wherein the IPC sequence representation is generated based on the identified IPC sequences with appended BOS tokens, and wherein the set of amino acid sequence representations and the IPC sequence representation are processed in parallel; Generating a synthetic representation by combining each of the transformed BOS token representations of the set of transformed amino acid sequence representations with the transformed BOS token representation of the transformed IPC sequence representation; and Determining one or more predicted amino acid-IPC interactions based on the synthetic representation.
28. A method for manufacturing a vaccine, the method comprising: Producing a vaccine comprising: One or more peptides; Multiple nucleic acids encoding the one or more peptides; or Multiple cells expressing the one or more peptides, wherein the one or more peptides are selected from a set of peptides by: Accessing a set of amino acid sequences, each of which in the set has been identified from at least one protein; Access the IPC sequences identified for the immune protein complexes (IPCs) of a subject; Use one or more first processing blocks in a processing subsystem of a machine learning model to process a set of amino acid sequence representations to generate a set of transformed amino acid sequence representations based on a set of element focus scores representing the binding cores of the set of amino acid sequence representations, wherein each of the amino acid sequence representations is generated based on one of the amino acid sequences appended with a sequence start (BOS) token; Use a second processing block in the processing subsystem to process the IPC sequence representation to generate a transformed IPC sequence representation, wherein the IPC sequence representation is generated based on the identified IPC sequence appended with a BOS token, and wherein the set of amino acid sequence representations and the IPC sequence representation are processed in parallel; Generate a synthetic representation by combining each of the transformed BOS token representations of the set of transformed amino acid sequence representations with the transformed BOS token representation of the transformed IPC sequence representation; and Determine one or more predicted amino acid-IPC interactions based on the synthetic representation.
29. A pharmaceutical composition comprising one or more peptides selected from a group of peptides by: Access a set of amino acid sequences, each of the amino acid sequences in the set having been identified from at least one protein; Access the IPC sequences identified for the immune protein complexes (IPCs) of a subject; Use one or more first processing blocks in a processing subsystem of a machine learning model to process a set of amino acid sequence representations to generate a set of transformed amino acid sequence representations based on a set of element focus scores representing the binding cores of the set of amino acid sequence representations, wherein each of the amino acid sequence representations is generated based on one of the amino acid sequences appended with a sequence start (BOS) token; Use a second processing block in the processing subsystem to process the IPC sequence representation to generate a transformed IPC sequence representation, wherein the IPC sequence representation is generated based on the identified IPC sequence appended with a BOS token, and wherein the set of amino acid sequence representations and the IPC sequence representation are processed in parallel; Generate a synthetic representation by combining each of the transformed BOS token representations of the set of transformed amino acid sequence representations with the transformed BOS token representation of the transformed IPC sequence representation; and Determine one or more predicted amino acid-IPC interactions based on the synthetic representation.
30. A computer-implemented method for predicting amino acid-immune protein complex (IPC) interactions, the computer-implemented method comprising: Access a set of amino acid sequences, each of the amino acid sequences in the set having been identified from at least one protein; Access the IPC sequences identified for the immune protein complexes (IPCs) of a subject; Process a set of amino acid sequence representations using one or more first processing blocks in a processing subsystem to generate a set of transformed amino acid sequence representations based on a set of element focus scores representing the binding core of the set of amino acid sequence representations, wherein each of the amino acid sequence representations is generated based on one of the amino acid sequences with an additional start-of-sequence (BOS) token; Process the IPC sequence using a second processing block in the processing subsystem to generate an IPC sequence embedding; Generate a synthetic representation by aggregating each of the transformed BOS token representations of the set of transformed amino acid sequence representations with the IPC sequence embedding; and Determine one or more predicted amino acid-IPC interactions based on the synthetic representation.
31. The computer-implemented method according to claim 30, wherein the IPC of the subject is a major histocompatibility complex (MHC) comprising MHC class I (MHC-I) and / or MHC class II (MHC-II), or a T cell receptor (TCR), and wherein the at least one protein is a therapeutic protein or is present in a disease sample from the subject.
32. The computer-implemented method according to claim 30, wherein processing a set of amino acid sequence representations includes processing a peptide sequence start (BOS) representation to generate a transformed peptide sequence representation.
33. The computer-implemented method according to claim 30, wherein processing the set of amino acid sequence representations includes: Using the one or more first processing blocks to transform the set of amino acid sequence representations into the set of transformed amino acid sequence representations, wherein each of the one or more first processing blocks includes a set of processing sub-blocks.
34. The computer-implemented method according to claim 30, wherein the machine learning model includes one or more Transformer encoders, and each of the one or more Transformer encoders includes a processing layer.
35. The computer-implemented method according to claim 30, wherein the set of amino acid sequence representations includes an aggregated sequence representation, and the aggregated sequence representation includes a set of peptide representations and one or more of the following: a set of amino-terminal flank (N-flank) representations or a set of carboxyl-terminal flank (C-flank) representations.
36. The computer-implemented method according to claim 30, further comprising, before generating the set of transformed amino acid sequence representations: Flattening the aggregated sequence representation into a single array; and Densifying the aggregated sequence representation by removing blank lines from the array, wherein the transformed amino acid sequence representations are generated based on the densified aggregated sequence representation.
37. The computer-implemented method according to claim 30, wherein processing the set of amino acid sequence representations includes: For each amino acid sequence representation in the set: For each element of the amino acid sequence representation, determine a plurality of vectors based on a set of weights associated with the processing layer of the machine learning model; Generate the element focus score group based on the plurality of vectors and the weight group.
38. The computer-implemented method according to claim 30, wherein the one or more first processing blocks and the second processing block include attention blocks, each attention block includes a set of attention sub-blocks, and each attention sub-block includes a self-attention layer; and wherein the machine learning model is an attention-based machine learning model, and The method further includes: generate an attention map including one or more masks through one or more of the attention blocks, the one or more masks restricting the attention applied by the attention sub-blocks to a sequence length according to the mask.
39. The computer-implemented method according to claim 38, the method further comprising: Calculate a plurality of average attention values corresponding to the plurality of peptide positions by calculating an average attention value at each peptide position among the plurality of peptide positions based on a set of peptides having a uniform length distribution. and Subtract the plurality of average attention values from the masks in the one or more masks.
40. The computer-implemented method according to claim 30, further comprising obtaining a data set for training the machine learning model by: generating a plurality of transformed peptide representations for a plurality of training peptides; obtaining a corresponding training peptide cluster for each training peptide based on the plurality of transformed peptide representations; calculating an information content for each training peptide based on the corresponding training peptide cluster; and excluding the one or more training peptides from the training data based on the corresponding information content of the one or more training peptides.
41. The computer-implemented method according to claim 30, wherein the amino acid sequence group includes peptide sequences and the IPC sequence includes a major histocompatibility complex (MHC) sequence, and wherein the one or more predicted amino acid-IPC interactions include one or more of the following: Interaction affinity prediction for a peptide-IPC combination, which predicts the binding affinity between a peptide and an MHC; Interaction prediction for the peptide-IPC combination, which predicts whether the MHC will present the peptide at the cell surface; or Immunogenicity prediction for the peptide-IPC combination, which predicts the ability of the peptide to elicit an immune response against the MHC.
42. The computer-implemented method according to claim 30, wherein determining the one or more predicted amino acid-IPC interactions includes: processing the synthetic representation to generate a set of results; and selecting an amino acid-IPC combination based on the highest result in the set of results.
43. The computer-implemented method according to claim 30, wherein the one or more predicted amino acid-IPC interactions include a prediction of the tumor-specific immunogenicity of a peptide.
44. The computer-implemented method according to claim 30, wherein the amino acid sequence group includes a set of peptide sequences, and the one or more predicted amino acid-IPC interactions identify a subgroup of peptide sequences having increased tumor-specific immunogenicity or increased likelihood of being presented by the IPC relative to the set of peptide sequences.
45. The computer-implemented method according to claim 30, further comprising: Based on the determined one or more predicted amino acid-IPC interactions, identifying a subgroup of peptides from the group of amino acid sequences to be included in an individualized vaccine, to be included as a target for immunotherapy, and / or to be excluded as a target for immunotherapy.
46. The computer-implemented method according to claim 30, further comprising: Accessing a protein sequence corresponding to the at least one protein; Obtaining a protein sequence embedding based on the protein sequence; And Determining the one or more predicted amino acid-IPC interactions at least in part based on the protein sequence embedding.
47. The computer-implemented method according to claim 30, wherein the protein language model comprises a pre-trained protein language model.
48. The computer-implemented method according to claim 46, further comprising: Reducing the dimension of the protein sequence embedding; And Combining each of the transformed BOS token representations in the transformed amino acid sequence representations of the group of transformed amino acid sequence representations with the transformed BOS token representation of the transformed IPC sequence representation and the dimension-reduced protein sequence embedding.
49. The computer-implemented method according to claim 48, wherein the dimension of the protein sequence embedding is reduced via a neural network.
50. The computer-implemented method according to claim 30, wherein generating the IPC sequence embedding comprises: Inputting the IPC sequence into a protein language model.
51. The computer-implemented method according to claim 30, further comprising: Reducing the dimension of the IPC sequence embedding.
52. The computer-implemented method according to claim 51, wherein the dimension of the IPC sequence embedding is reduced via principal component analysis (PCA).
53. The computer-implemented method according to claim 51, wherein generating the synthetic representation comprises: For each of the group of transformed amino acid sequence representations, element-wise multiplying the transformed amino acid sequence start (BOS) representation corresponding to the transformed amino acid sequence representation by the dimension-reduced IPC sequence embedding.
54. A system for predicting amino acid-immunoprotein complex (IPC) interactions, the system comprising: One or more non-transitory computer-readable storage media, which comprise instructions; And One or more processors coupled to the one or more storage media, the one or more processors being configured to execute the instructions to: access a group of amino acid sequences, each of the amino acid sequences in the group having been identified from at least one protein; Access a group of amino acid sequences, each of the amino acid sequences in the group having been identified from at least one protein; Accessing an IPC sequence identified for an immunoprotein complex (IPC) of a subject; Using one or more first processing blocks in a processing subsystem of a machine learning model to process a group of amino acid sequence representations to generate a group of transformed amino acid sequence representations based on a set of element-wise focus scores representing the binding core of the group of amino acid sequence representations, wherein each of the amino acid sequence representations is generated based on one of the amino acid sequences with an additional start-of-sequence (BOS) token. Process the IPC sequence using a second processing block in the processing subsystem to generate an IPC sequence embedding; Generate a synthetic representation by aggregating each of the transformed BOS token representations of the group of transformed amino acid sequence representations with the IPC sequence embedding; and Determine one or more predicted amino acid-IPC interactions based on the synthetic representation.
55. A non-transitory computer-readable medium for a system for predicting amino acid-immunoprotein complex (IPC) interactions, the non-transitory computer-readable medium comprising instructions that, when executed by one or more processors of one or more computing devices, cause the one or more processors to: Access a set of amino acid sequences, each of which has been identified from at least one protein; Access an IPC sequence identified for an immunoprotein complex (IPC) of a subject; Process a set of amino acid sequence representations using one or more first processing blocks in a processing subsystem of a machine learning model to generate a set of transformed amino acid sequence representations based on a set of element focus scores representing a binding core of the group of amino acid sequence representations, wherein each of the amino acid sequence representations is generated based on one of the amino acid sequences with an appended start-of-sequence (BOS) token; Process the IPC sequence using a second processing block in the processing subsystem to generate an IPC sequence embedding; Generate a synthetic representation by aggregating each of the transformed BOS token representations of the group of transformed amino acid sequence representations with the IPC sequence embedding; and Determine one or more predicted amino acid-IPC interactions based on the synthetic representation.
56. A vaccine comprising: One or more peptides; Multiple nucleic acids encoding the one or more peptides; or Multiple cells expressing the one or more peptides, wherein the one or more peptides are selected from a group of peptides by: Access a set of amino acid sequences, each of which has been identified from at least one protein; Access an IPC sequence identified for an immunoprotein complex (IPC) of a subject; Process a set of amino acid sequence representations using one or more first processing blocks in a processing subsystem of a machine learning model to generate a set of transformed amino acid sequence representations based on a set of element focus scores representing a binding core of the group of amino acid sequence representations, wherein each of the amino acid sequence representations is generated based on one of the amino acid sequences with an appended start-of-sequence (BOS) token; Process the IPC sequence using a second processing block in the processing subsystem to generate an IPC sequence embedding; Generate a synthetic representation by aggregating each of the transformed BOS token representations of the group of transformed amino acid sequence representations with the IPC sequence embedding; and Determine one or more predicted amino acid-IPC interactions based on the synthetic representation.
57. A method for manufacturing a vaccine, the method comprising: Producing a vaccine comprising: One or more peptides; Multiple nucleic acids encoding the one or more peptides; or Multiple cells expressing the one or more peptides, wherein the one or more peptides are selected from a group of peptides by: Accessing a group of amino acid sequences, each of which in the group has been identified from at least one protein; Accessing IPC sequences identified for an immune protein complex (IPC) of a subject; Using one or more first processing blocks in a processing subsystem of a machine learning model to process a group of amino acid sequence representations to generate a group of transformed amino acid sequence representations based on a set of element focus scores representing a binding core of the group of amino acid sequence representations, wherein each of the amino acid sequence representations is generated based on one of the amino acid sequences appended with a sequence start (BOS) token; Using a second processing block in the processing subsystem to process the IPC sequence to generate an IPC sequence embedding; Generating a synthetic representation by aggregating each of the transformed BOS token representations of the group of transformed amino acid sequence representations with the IPC sequence embedding; and Determining one or more predicted amino acid-IPC interactions based on the synthetic representation.
58. A pharmaceutical composition comprising one or more peptides selected from a group of peptides by: Accessing a group of amino acid sequences, each of which in the group has been identified from at least one protein; Accessing IPC sequences identified for an immune protein complex (IPC) of a subject; Using one or more first processing blocks in a processing subsystem of a machine learning model to process a group of amino acid sequence representations to generate a group of transformed amino acid sequence representations based on a set of element focus scores representing a binding core of the group of amino acid sequence representations, wherein each of the amino acid sequence representations is generated based on one of the amino acid sequences appended with a sequence start (BOS) token; Using a second processing block in the processing subsystem to process the IPC sequence to generate an IPC sequence embedding; Generating a synthetic representation by aggregating each of the transformed BOS token representations of the group of transformed amino acid sequence representations with the IPC sequence embedding; and Determining one or more predicted amino acid-IPC interactions based on the synthetic representation.
59. A computer-implemented method for predicting amino acid-immune protein complex (IPC) interactions, the computer-implemented method comprising: Accessing a group of amino acid sequences, each of which in the group has been identified from at least one protein; Accessing IPC sequences identified for an immune protein complex (IPC) of a subject; Generating an IPC sequence embedding based on the IPC sequence; Using one or more processing blocks in a processing subsystem of a machine learning model to process the group of amino acid sequences to generate a group of transformed amino acid sequence representations based on a set of element focus scores representing a binding core of the group of amino acid sequence representations; Using the cross-attention machine learning module in the processing subsystem, generate a set of synthetic representations based on the group of transformed amino acid sequence representations and the IPC sequence embedding; and Determine one or more predicted amino acid-IPC interactions based on the synthetic representations.
60. The computer-implemented method according to claim 59, wherein generating the IPC sequence embedding comprises: Input the IPC sequence into a protein language model.
61. The computer-implemented method according to claim 59, further comprising: Reduce the dimension of the IPC sequence embedding.
62. The computer-implemented method according to claim 61, wherein the dimension of the IPC sequence embedding is reduced via principal component analysis (PCA).
63. The computer-implemented method according to claim 61, wherein the cross-attention module includes a self-attention transformer, and the self-attention transformer has three components: a query (Q) component, a key (K) component, and a value (V) component.
64. The computer-implemented method according to claim 63, wherein: Each of the K component and the V component corresponds to the group of transformed amino acid sequence representations; and The Q component corresponds to the aggregation of the sequence start (BOS) vector embedding and the dimension-reduced IPC sequence embedding.
65. The computer-implemented method according to claim 63, wherein: Each of the K component and the V component corresponds to the group of transformed amino acid sequence representations; and The Q component corresponds to the dimension-reduced IPC sequence embedding.
66. The computer-implemented method according to claim 59, wherein: The group of amino acid sequences includes at least one peptide sequence having multiple binding cores, the multiple binding cores being capable of binding to multiple alleles of the IPC, and wherein the one or more predicted amino acid-IPC interactions include at least multiple allele-specific and binding core-specific predicted amino acid-IPC interactions.
67. The computer-implemented method according to claim 59, wherein the IPC of the subject is a major histocompatibility complex (MHC) comprising MHC class I (MHC-I) and / or MHC class II (MHC-II) or is a T cell receptor (TCR), and wherein the at least one protein is a therapeutic protein or is present in a disease sample from the subject.
68. The computer-implemented method according to claim 59, wherein processing the set of amino acid sequence representations comprises: Use the one or more first processing blocks to transform the group of amino acid sequence representations into the group of transformed amino acid sequence representations, wherein each of the one or more first processing blocks includes a set of processing sub-blocks.
69. The computer-implemented method according to claim 59, wherein the group of amino acid sequence representations includes an aggregated sequence representation, the aggregated sequence representation including a set of peptide representations and one or more of the following: a set of amino-terminal flank (N-flank) representations or a set of carboxyl-terminal flank (C-flank) representations.
70. The computer-implemented method according to claim 59, wherein processing the group of amino acid sequence representations includes: For each amino acid sequence representation in the group: For each element represented by the amino acid sequence, determine a plurality of vectors based on a set of weights associated with a processing layer of the machine learning model; and Generate the set of element focus scores based on the plurality of vectors and the set of weights.
71. The computer-implemented method according to claim 59, wherein the one or more processing blocks include attention blocks, each attention block includes a set of attention sub-blocks, and each attention sub-block includes a self-attention layer; and wherein the machine learning model is an attention-based machine learning model, and The method further includes: generate an attention map including one or more masks through one or more of the attention blocks, the one or more masks restricting the attention applied by the attention sub-blocks to the sequence length according to the mask.
72. The computer-implemented method according to claim 71, the method further comprising: Calculate a plurality of average attention values corresponding to the plurality of peptide positions by calculating the average attention value at each peptide position in the plurality of peptide positions based on a set of peptides having a uniform length distribution; and Subtract the plurality of average attention values from the mask in the one or more masks.
73. The computer-implemented method according to claim 59, further comprising obtaining a data set for training the machine learning model by: Generating a plurality of transformed peptide representations for a plurality of training peptides; Obtaining a corresponding training peptide cluster for each training peptide based on the plurality of transformed peptide representations; Calculating the information content for each training peptide based on the corresponding training peptide cluster; and Excluding the one or more training peptides from the training data based on the corresponding information content of the one or more training peptides.
74. The computer-implemented method according to claim 59, wherein the amino acid sequence group contains peptide sequences and the IPC sequence contains a major histocompatibility complex (MHC) sequence, and wherein the one or more predicted amino acid-IPC interactions include one or more of the following: Interaction affinity prediction for a peptide-IPC combination, which predicts the binding affinity between a peptide and an MHC; Interaction prediction for the peptide-IPC combination, which predicts whether the MHC will present the peptide at the cell surface; or Immunogenicity prediction for the peptide-IPC combination, which predicts the ability of the peptide to induce an immune response against the MHC.
75. The computer-implemented method according to claim 59, wherein determining the one or more predicted amino acid-IPC interactions includes: Processing the synthetic representation to generate a set of results; and Selecting an amino acid-IPC combination based on the highest result in the set of results.
76. The computer-implemented method according to claim 59, wherein the one or more predicted amino acid-IPC interactions include a prediction of the tumor-specific immunogenicity of a peptide.
77. The computer-implemented method according to claim 59, wherein the amino acid sequence group comprises a set of peptide sequences, and wherein the one or more predicted amino acid-IPC interactions identify a subgroup of peptide sequences having increased tumor-specific immunogenicity or increased likelihood of presentation by the IPC relative to the set of peptide sequences.
78. The computer-implemented method according to claim 59, further comprising: identifying, based on the determined one or more predicted amino acid-IPC interactions, a subgroup of peptides from the amino acid sequence group to be included in an individualized vaccine, to be included as a target for immunotherapy, and / or to be excluded as a target for immunotherapy.
79. The computer-implemented method according to claim 59, further comprising: accessing a protein sequence corresponding to the at least one protein; obtaining a protein sequence embedding based on the protein sequence and a protein model; and determining the one or more predicted amino acid-IPC interactions at least in part based on the protein sequence embedding.
80. The computer-implemented method according to claim 79, wherein the protein language model comprises a pre-trained protein language model.
81. The computer-implemented method according to claim 80, wherein the synthetic representation is generated at least in part based on the protein sequence embedding.
82. A system for predicting amino acid-immune protein complex (IPC) interactions, the system comprising: one or more non-transitory computer-readable storage media, which comprise instructions; and one or more processors coupled to the one or more storage media, the one or more processors being configured to execute the instructions to: access a group of amino acid sequences, each of which has been identified from at least one protein; access an IPC sequence identified for an immune protein complex (IPC) of a subject; generating an IPC sequence embedding based on the IPC sequence; processing the group of amino acid sequences using one or more processing blocks in a processing subsystem of a machine learning model to generate a group of transformed amino acid sequence representations based on a set of element focus scores representing a binding core of the amino acid sequence group; using a cross-attention machine learning module in the processing subsystem to generate a set of synthetic representations based on the group of transformed amino acid sequence representations and the IPC sequence embedding; and determining one or more predicted amino acid-IPC interactions based on the synthetic representations.
83. A non-transitory computer-readable medium for a system for predicting amino acid-immune protein complex (IPC) interactions, the non-transitory computer-readable medium comprising instructions that, when executed by one or more processors of one or more computing devices, cause the one or more processors to: access a group of amino acid sequences, each of which has been identified from at least one protein; Access the IPC sequences identified for the immune protein complexes (IPCs) of a subject; Generate IPC sequence embeddings based on the IPC sequences; Use one or more processing blocks in a processing subsystem of a machine learning model to process the set of amino acid sequences to generate a set of transformed amino acid sequence representations based on a set of element focus scores representing the binding cores of the set of amino acid sequence representations; Use a cross-attention machine learning module in the processing subsystem to generate a set of synthetic representations based on the set of transformed amino acid sequence representations and the IPC sequence embeddings; And Determine one or more predicted amino acid-IPC interactions based on the synthetic representations.
84. A vaccine comprising: One or more peptides; Multiple nucleic acids encoding the one or more peptides; or Multiple cells expressing the one or more peptides, wherein the one or more peptides are selected from a set of peptides by: Accessing a set of amino acid sequences, each of which in the set has been identified from at least one protein; Accessing the IPC sequences identified for the immune protein complexes (IPCs) of a subject; Generating IPC sequence embeddings based on the IPC sequences; Using one or more processing blocks in a processing subsystem of a machine learning model to process the set of amino acid sequences to generate a set of transformed amino acid sequence representations based on a set of element focus scores representing the binding cores of the set of amino acid sequence representations; Using a cross-attention machine learning module in the processing subsystem to generate a set of synthetic representations based on the set of transformed amino acid sequence representations and the IPC sequence embeddings; And Determine one or more predicted amino acid-IPC interactions based on the synthetic representations.
85. A method for manufacturing a vaccine, the method comprising: Producing a vaccine comprising: One or more peptides; Multiple nucleic acids encoding the one or more peptides; or Multiple cells expressing the one or more peptides, wherein the one or more peptides are selected from a set of peptides by: Accessing a set of amino acid sequences, each of which in the set has been identified from at least one protein; Accessing the IPC sequences identified for the immune protein complexes (IPCs) of a subject; Generating IPC sequence embeddings based on the IPC sequences; Using one or more processing blocks in a processing subsystem of a machine learning model to process the set of amino acid sequences to generate a set of transformed amino acid sequence representations based on a set of element focus scores representing the binding cores of the set of amino acid sequence representations; Using a cross-attention machine learning module in the processing subsystem to generate a set of synthetic representations based on the set of transformed amino acid sequence representations and the IPC sequence embeddings; And Determine one or more predicted amino acid-IPC interactions based on the synthetic representations.
86. A pharmaceutical composition comprising one or more peptides selected from a set of peptides by: Access a set of amino acid sequences, each of which in the set has been identified from at least one protein; Access IPC sequences identified for an immune protein complex (IPC) of a subject; Generate an IPC sequence embedding based on the IPC sequences; Use one or more processing blocks in a processing subsystem of a machine learning model to process the set of amino acid sequences to generate a set of transformed amino acid sequence representations based on a set of element focus scores representing the binding cores of the set of amino acid sequence representations; Use a cross-attention machine learning module in the processing subsystem to generate a set of synthetic representations based on the set of transformed amino acid sequence representations and the IPC sequence embedding; And Determine one or more predicted amino acid-IPC interactions based on the synthetic representations.