Selection of diverse candidate peptides for peptide therapeutics
A machine learning-based method using metric learning algorithms trains models to generate peptide sequence vectors, addressing the challenge of selecting diverse candidate peptides for personalized immunotherapy, thereby improving therapeutic efficacy by ensuring peptides with different binding motifs are selected, enhancing the effectiveness of peptide vaccines and cell therapies.
Patent Information
- Application Number
- JP2025528166
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-15
- Filing Date
- 2023-11-14
- Publication Date
- 2025-12-10
AI Technical Summary
Current methods for selecting candidate peptides for personalized immunotherapy are challenging due to the large number of peptide sequences detected in a given sample, and determining peptide similarity is crucial for improving therapeutic efficacy.
A machine learning-based approach using a metric learning algorithm trains a model to generate peptide sequence vectors in an n-dimensional space, ensuring that peptides presented by the same MHC allele are closer together and those presented by different MHC alleles are farther apart, facilitating the selection of a diverse group of candidate peptides with different binding motifs.
This method increases the likelihood of peptide presentation and elicits a desired immunological response by selecting peptides with diverse binding motifs, enhancing the effectiveness of personalized immunotherapies such as peptide vaccines and cell therapies.
Smart Images

Figure 2025539935000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 425,647, entitled "Selection of Diverse Candidate Peptides for Peptide Therapeutics," filed November 15, 2022, which is incorporated herein by reference in its entirety.
[0002] The present disclosure relates generally to the selection of candidate peptides for therapeutic development. More particularly, the present disclosure relates to machine learning-based methods and systems for selecting a diverse group of candidate peptides with respect to binding motifs for use in the development of peptide immunotherapies (e.g., peptide therapeutics such as peptide vaccines, cell therapies, etc.) with improved efficacy. [Background technology]
[0003] Personalized immunotherapy involves the use of an individual's own immune system to fight diseases such as cancer. Immunotherapies can include, for example, peptide therapeutics (e.g., peptide vaccines) and cell therapy. Peptide vaccines are created using one or more peptides that mimic epitopes of antigens that elicit an immune response. Peptide vaccines can be used to induce protection against infectious pathogens and non-infectious diseases and can be used as therapeutic cancer vaccines (e.g., neoantigen vaccines). Neoantigen vaccines are a relatively new approach to providing personalized cancer treatment, using peptides derived from tumor-associated antigens to induce effective anti-tumor T cell responses. Cell therapy can involve injecting cells (e.g., T cells or tumor cells) into an individual to generate or induce an immune response. For example, T cells can be harvested from an individual's blood and modified to more aggressively attack tumor cells. These T cells can then be infused into the individual to generate the desired immune response. As another example, tumor cells can be collected from an individual and reengineered to elicit attack by the immune system.
[0004] Selecting candidate peptides for use in developing personalized immunotherapy can be challenging.Currently available methods for identifying and ranking candidate peptides for use in developing personalized immunotherapy can be challenging in that a large number of peptide sequences may be detected in a given sample.Knowing whether a given peptide is similar to another peptide can be useful for identifying and ranking candidate peptides.Therefore, it may be desirable to have a method and / or system for evaluating the similarity between peptides to assist in selecting candidate peptides for use in developing personalized immunotherapy with improved efficacy. Summary of the Invention
[0005] In one or more embodiments, a method for developing a therapeutic agent is provided. Peptide sequence data identifying a plurality of peptide sequences corresponding to a plurality of peptides is received. A plurality of peptide sequence vectors are generated for each of the peptide sequences, and thereby for each of the peptides, via a trained machine learning model in n-dimensional space. The machine learning model has been trained using a metric learning algorithm, training peptide sequence data, and training allele presentation data. The training peptide sequence data identifies a training peptide sequence corresponding to each training peptide of the plurality of training peptides. The training allele presentation data identifies, for each training peptide sequence in the training peptide sequence data, one or more major histocompatibility complex (MHC) alleles that are expected to present the training peptide corresponding to the training peptide sequence. The metric learning algorithm has been used to train the machine learning model such that a first distance between a first pair of peptide sequence vectors generated in the n-dimensional space for each first pair of peptides presented by the same MHC allele is smaller than a second distance between a second pair of peptide sequence vectors generated in the n-dimensional space for a second pair of peptides presented by different MHC alleles. Output is generated using the peptide sequence vectors. The output provides an indication of similarity between peptide sequences for use in selecting a diverse group of candidate peptides from the peptides for therapeutic development.
[0006] In one or more embodiments, a method for developing a peptide vaccine is provided. A machine learning model is trained using a metric learning algorithm, training peptide sequence data, and training allele representation data corresponding to the training peptide sequence data. Peptide sequence data identifying a plurality of peptide sequences corresponding to a plurality of peptides is received. A peptide sequence vector is generated for each peptide sequence of the plurality of peptide sequences using the peptide sequence data via the machine learning model to form a plurality of peptide sequence vectors. Output is generated using the plurality of peptide sequence vectors. The output provides an indication of similarity between peptide sequences of the plurality of peptide sequences. A group of candidate peptides is selected from the plurality of peptides for development of the peptide vaccine based on the output, such that the diverse group of candidate peptides includes at least two dissimilar candidate peptides.
[0007] In one or more embodiments, a method is provided that includes receiving training peptide sequence data including a plurality of training peptide sequences, generating training allele representation data for the training peptide sequence data, where the training allele representation data identifies MHC alleles that are predicted to represent the training peptide sequences for the training peptide sequences among the plurality of training peptide sequences, and training a machine learning model using the training peptide sequence data, the training allele representation data, and a metric learning algorithm. The machine learning model is trained to generate a peptide sequence vector for a given peptide sequence. The peptide sequence vector is a vector in n-dimensional space that provides an indication of the similarity of the given peptide sequence to other peptide sequences.
[0008] In one or more embodiments, the vaccine comprises a set of multiple peptides, multiple precursors of multiple peptides, or nucleic acids encoding multiple peptides or multiple precursors. The multiple peptides include at least two peptides with different binding motifs. The multiple peptides are selected from a group of diverse candidate peptides selected based on some or all of one or more of the methods described herein.
[0009] In one or more embodiments, a method for producing a vaccine is provided. The vaccine comprises a plurality of peptides, a plurality of precursors of a plurality of peptides, or a set of nucleic acids encoding a plurality of peptides or a plurality of precursors. The plurality of peptides comprises at least two peptides with different binding motifs. The plurality of peptides is selected from a group of diverse candidate peptides selected based on some or all of one or more of the methods described herein.
[0010] In one or more embodiments, the pharmaceutical composition comprises two or more peptides selected from a diverse group of candidate peptides selected based on some or all of one or more of the methods described herein.
[0011] In one or more embodiments, the pharmaceutical composition comprises two or more nucleic acid sequences encoding two or more respective peptides selected from a group of diverse candidate peptides selected based on some or all of one or more of the methods described herein.
[0012] In one or more embodiments, a method of treating a subject is provided, the method comprising administering at least one of two or more peptides, two or more pharmaceutical compositions, or two or more nucleic acid sequences identified based on output generated by some or all of one or more of the methods described herein.
[0013] In one or more embodiments, the engineered T cells are generated using some or all of one or more of the methods described herein.
[0014] In one or more embodiments, the population of engineered T cells is generated using some or all of one or more of the methods described herein.
[0015] In one or more embodiments, a method of treating a subject with cancer is provided. A population of T cells is provided. At least a subset of the population of T cells is engineered to express an exogenous T cell receptor (TCR) and knock out endogenous TCR-β, thereby forming a population of engineered T cells. The exogenous TCR binds to an antigen expressed by the cancer and selected using some or all of one or more of the methods described herein. The population of engineered T cells is expanded. The expanded population of engineered T cells is administered to a subject.
[0016] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods disclosed herein.
[0017] In some embodiments, a computer program product is provided that is tangibly embodied in a non-transitory machine-readable storage medium and includes instructions configured to cause one or more data processors to perform some or all of one or more of the methods disclosed herein.
[0018] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium including instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium, the computer program product including instructions configured to cause one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein.
[0019] The terms and expressions which have been employed are used as terms of description rather than of limitation, and there is no intention in the use of such terms and expressions to exclude equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, although the claimed invention has been specifically disclosed by embodiments and optional features, it will be understood that modifications and variations of the concepts disclosed herein may be employed by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims. [Brief explanation of the drawings]
[0020] The present disclosure will be described with reference to the accompanying drawings, in which:
[0021] [Figure 1] FIG. 1 is a block diagram of a therapeutic drug development system 100, according to various embodiments.
[0022] [Figure 2] FIG. 2 is a schematic diagram of an example of the configuration of the model of FIG. 1, according to one or more embodiments.
[0023] [Figure 3] 1 is a flowchart of a process for use in therapeutic drug development, according to one or more embodiments.
[0024] [Figure 4] 1 is a flowchart of a process for training a machine learning model to generate a peptide sequence vector for a peptide sequence, according to one or more embodiments.
[0025] [Figure 5] 1 is a flowchart of a process for training a machine learning model to generate a peptide sequence vector for a peptide sequence, according to one or more embodiments.
[0026] [Figure 6] 1 is a flowchart of a process for treating a subject, according to one or more embodiments.
[0027] [Figure 7] FIG. 1 is a plot of a peptide sequence vector in reduced dimensional space, according to one or more embodiments.
[0028] [Figure 8] 1 is a list of MHC alleles according to one or more embodiments.
[0029] [Figure 9] FIG. 1 is a block diagram of a computer system according to various embodiments.
[0030] In the accompanying drawings, similar components and / or features may have the same reference label. Furthermore, various components of the same type may be distinguished by following the reference label with a dash and a second label that distinguishes between the similar components. If only a first reference label is used herein, the description is applicable to any one of the similar components having the same first reference label, regardless of the second reference label. DETAILED DESCRIPTION OF THE INVENTION
[0031] I. Overview Recognizing the importance of being able to determine which peptides to select as candidates for use in the development of personalized immunotherapies (e.g., cell therapy, peptide therapeutics such as peptide vaccines, etc.), the embodiments described herein provide methodologies and systems for making such decisions in a manner that results in improved therapeutic efficacy compared to various currently available methods and systems. The embodiments described herein use machine learning methodologies and systems to improve peptide selection performance, for example, but not limited to, by increasing the diversity of binding motifs of candidate peptides selected for use in the development of therapeutics. Increasing the diversity of binding motifs can increase the likelihood of presentation of one or more of these peptides. A peptide's binding motif can be a particular configuration of amino acids (or "motif") that allows the peptide to bind to and be presented by a corresponding major histocompatibility complex (MHC) allele. In humans, these MHC alleles are called human leukocyte antigen (HLA) alleles.
[0032] References to MHC alleles or HLA alleles herein refer to proteins defined by these alleles. For example, an MHC allele described as presenting or capable of presenting a peptide may refer to a protein defined by the MHC allele that has a binding pocket (or binding groove) that can bind to the corresponding binding motif of the peptide. The binding pocket may have a specific sequence configuration that can bind to the corresponding binding motif. A peptide may have one or more binding motifs that allow the peptide to bind to one or more respective MHC alleles. Furthermore, two or more peptides may have a common binding motif. Therefore, an MHC allele that is capable of presenting or can present a peptide may also be described as presenting or capable of presenting a peptide sequence corresponding to that peptide. In other words, a protein defined by that MHC allele can bind to a peptide having that peptide sequence.
[0033] The embodiments described herein provide various methodologies for using machine learning models and / or outputs generated by the machine learning models to analyze peptide sequences identified from one or more samples from one or more subjects. A peptide sequence is a sequence corresponding to at least a portion of a peptide, and the sequence may be an amino acid sequence, a codon sequence, or a nucleic acid sequence. The sample may be, for example, but not limited to, a disease sample (e.g., diseased tissue, tumor tissue). The peptide sequences detected or identified from one or more samples may be processed using a machine learning model to generate mathematical representations of the peptide sequences that capture information about the peptide sequences and, therefore, about the peptides that contain these peptide sequences. Furthermore, because the binding motifs are sequence-based, these mathematical representations also capture information about one or more binding motifs contained within a given peptide sequence.
[0034] In particular, the machine learning model processes the peptide sequence to generate a mathematical representation in the form of a peptide sequence vector, also called a peptide sequence embedding. The peptide sequence vector (or PS vector) created for a peptide sequence can also be referred to as being for a peptide having the peptide sequence or corresponding to a peptide having the peptide sequence. The peptide sequence vector is a vector with n dimensions (or n discrete elements) in the embedding space. That is, these peptide sequence vectors are vectors in an n-dimensional (embedded) space.
[0035] The distance between any two peptide sequence vectors in this n-dimensional space can provide some indication of how similar or dissimilar the corresponding peptides are to each other. Peptide sequence vectors representing similar peptide sequences are embedded closer to each other in the n-dimensional space, while peptide sequence vectors representing dissimilar peptide sequences are embedded further apart in the n-dimensional space. If two peptide sequences are similar, the corresponding peptides can also be referred to as similar. Furthermore, because binding motifs are sequence-based, two peptides that are similar to each other as determined by their peptide sequence vectors can be considered to have similar binding motifs. Two different peptides as determined by their peptide sequence vectors can be considered to have different binding motifs. In some cases, a peptide can have more than one binding motif.
[0036] The embodiments described herein recognize that metric learning can be used to improve the training and performance of machine learning models in distinguishing between binding motifs of different peptide sequences. Thus, in one or more embodiments, a machine learning model is trained to generate peptide sequence vectors for peptide sequences using a metric learning algorithm. The metric learning algorithm uses one or more loss functions, which can include, but are not limited to, at least one of a contrast loss function, a triplet loss function, a quartet loss function, a circle loss function, a multi-class n-pairs loss function, a lifted structure loss function, an angle loss function, a divergence loss function, or a constellation loss function, or any other type of distance-based loss function.
[0037] Using information in the metric learning algorithm about MHC alleles that are expected (or predicted) to present the peptide improves training performance. Thus, the metric learning algorithm is used to train a machine learning model to generate peptide sequence vectors for peptides presented by the same MHC allele that are closer to each other in n-dimensional space, and to generate peptide sequence vectors for peptides presented by different MHC alleles that are farther apart from each other in n-dimensional space.
[0038] The peptide sequence vectors generated by the machine learning model trained using the metric learning algorithm can then be used to generate an output that provides an indication of the similarity between peptides. The output can directly represent the peptide sequence vectors generated by the machine learning model, or can represent the peptide sequence vectors in a form that can be more easily understood or interpreted by humans. For example, the output can take the form of a visual representation, a spreadsheet, a list, or some other type of output.
[0039] In one or more embodiments, the output takes the form of a graphical representation of the peptide sequence vector in k-dimensional space. The k-dimensional space may contain the same or fewer dimensions as the n-dimensional space of the peptide sequence vector (i.e., k≦n). Presenting the peptide sequence vector in a reduced dimensional space may allow a human to more quickly and / or easily understand or ascertain similarity relationships between corresponding peptides. For example, the n-dimensional space may contain 20 to 200 dimensions. In some embodiments, the output presents the peptide sequence vector in two or three dimensions for ease of visualization and understanding.
[0040] The generated output can be used to select a group of diverse candidate peptides for the development of personalized immunotherapy. For example, personalized immunotherapy (e.g., peptide vaccine, cell therapy, etc.) can be created using the entire group of diverse candidate peptides, or can be created using at least two peptides from the group of diverse candidate peptides, where the at least two peptides are dissimilar. The diversity in the group of candidate peptides can include, for example, but is not limited to, diversity in binding motifs. For example, at least two peptides in the group of diverse candidate peptides can have different binding motifs. Personalized immunotherapy can be created using at least two peptides with different binding motifs.
[0041] Individualized immunotherapies (e.g., peptide vaccines, cell therapies, etc.) developed using a diverse group of candidate peptides may be more effective because this diversity allows immunotherapy to be applied when the biological context results in a particular binding motif being more favorable than expected. For example, the microenvironment of a tumor or diseased tissue (e.g., pH, temperature, etc.) may preferentially favor one or more different binding motifs over those identified by processing various samples with mass spectrometry. Furthermore, individual subjects may have biological differences that lead to one or more different binding motifs being preferred over those identified by processing various samples with mass spectrometry. Developing peptide vaccines using a group of candidate peptides with diverse binding motifs increases the likelihood of eliciting a desired immunological response. For example, having diverse peptides in a peptide vaccine may increase the likelihood of presentation via MHC alleles, thereby increasing the likelihood of eliciting a desired immunological response and increasing the strength of the elicited immunological response.
[0042] The embodiments described herein further recognize and take into account that training models for sequence analysis can be particularly complex due to the overwhelming number of potentially observable peptide sequences. Not only are there millions of potentially presented peptides (e.g., neoantigens), but the genes encoding proteins, for example, of MHC class I molecules, are also highly polymorphic. There are approximately 20,000 human MHC class I alleles. Accordingly, the embodiments described herein provide methodologies and systems for training machine learning models in a manner that improves the overall performance and efficiency of selecting groups of candidate peptides for peptide vaccines with improved chances of success. The machine learning model is trained to embed peptide sequences of peptides into n-dimensional space, such that the further apart two embeddings are, the more dissimilar they are. Candidate peptides are selected to maximize the distance between embeddings representing these candidate peptides, or to ensure a distance above a certain distance threshold to ensure diversity. Candidate peptides are selected to ensure diverse binding motifs, but may also be selected to ensure well-defined and / or well-known binding motifs.
[0043] II. Exemplary Systems for Diversifying Peptides Used in the Development of Peptide Therapeutics (e.g., Peptide Vaccines) II.A. Exemplary Therapeutic Drug Development System Referring now to the drawings, Figure 1 is a block diagram of a therapeutic agent development system 100 according to various embodiments. The therapeutic agent development system 100 includes a computing platform 102, a data store 104, and a display system 106. The computing platform 102 may take a variety of forms. In one or more embodiments, the computing platform 102 includes a single computer (or computer system) or multiple computers in communication with each other. In other examples, the computing platform 102 takes the form of a cloud computing platform.
[0044] The data store 104 and the display system 106 each communicate with the computing platform 102. In some examples, the data store 104, the display system 106, or both may be considered part of the computing platform 102 or may be otherwise integrated. Thus, in some examples, the computing platform 102, the data store 104, and the display system 106 may be separate components that communicate with each other, while in other examples, some combination of these components may be integrated together. Communication between the different components may be implemented using any number of wired, wireless, or optical communication links, or a combination thereof.
[0045] The therapeutic development system 100, also referred to as a peptide therapeutic development system, is used to develop peptide therapeutics 108. The peptide therapeutics 108 can be, for example, a peptide vaccine that includes multiple peptides, precursors of multiple peptides, or one or more nucleic acids encoding multiple peptides or their precursors. In one or more embodiments, the peptide vaccine can be a personalized vaccine. The peptide vaccine can be, for example, a neo-antigen vaccine that includes multiple neo-antigens selected to treat cancer. The neo-antigen vaccine can be engineered or selected based on the subject-specific tumor profile of the peptide.
[0046] The therapeutic development system 100 is used to select a diverse set of candidate peptides 110 for use in generating a peptide therapeutic 108. The diverse set of candidate peptides 110 includes at least two peptides that are dissimilar. In one or more embodiments, the at least two peptides may be dissimilar with respect to their binding motifs. For example, at least two peptides in the diverse set of candidate peptides 110 may have different binding motifs. The peptide therapeutic 108 can be generated using at least two peptides from the diverse set of candidate peptides 110, where the at least two peptides are dissimilar (e.g., have different binding motifs).
[0047] Binding motifs can be dissimilar if the amino acids (or sequences of amino acids) in the binding motifs differ, if the spacing (or intervals) between amino acids differ, or a combination thereof. For example, binding motifs can be dissimilar by differing by more than a selected number of amino acids (e.g., two, three, four, five, or more amino acids). Thus, the dissimilarity between binding motifs of peptides can vary by degree. Two binding motifs that differ by a single amino acid are more dissimilar than two binding motifs that differ by three or four amino acids.
[0048] Therapeutic drug development system 100 includes a data analyzer 111. Data analyzer 111 may be implemented using hardware, software, firmware, or a combination thereof. In one or more embodiments, data analyzer 111 is implemented on computing platform 102.
[0049] The data analyzer 111 may include, for example, without limitation, a sequence analyzer 112 and a candidate selector 114, each of which may be implemented using hardware, software, firmware, or a combination thereof. In some embodiments, the sequence analyzer 112 and the candidate selector 114 are integrated together within the same module of the data analyzer 111. The sequence analyzer 112 is used to evaluate the similarities and / or dissimilarities of different peptides. The candidate selector 114 is used to select a group 110 of diverse candidate peptides for a peptide therapeutic 108 based on the analysis performed by the sequence analyzer 112.
[0050] Data analyzer 111 may receive peptide sequence data 116 (e.g., via one or more wired, wireless, and / or optical communication links), retrieve peptide sequence data 116 from data store 104 or some other type of storage device (e.g., cloud storage), access peptide sequence data 116 from multiple types of storage devices, generate peptide sequence data 116 based on the results of mass spectrometry, and / or obtain peptide sequence data 116 in some other manner. In one or more embodiments, peptide sequence data 116 may be retrieved from data store 104 in response to receiving user input entered by a user via an input device.
[0051] In one or more embodiments, peptide sequence data 116 is generated from processing a set of samples 118. The set of samples 118 may take the form of one or more biological samples from one or more subjects (e.g., diseased samples, healthy samples, combinations thereof). In one or more embodiments, the set of samples 118 includes samples obtained from a tumor in a subject. The tumor may be, for example, a symptom of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, stomach cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, T-cell lymphocytic leukemia, non-small cell lung cancer, small cell lung cancer, another type of cancer, or a combination thereof.
[0052] The set of samples 118 may be processed together to generate peptide sequence data 116. In some examples, multiple samples within the set of samples 118 may be processed at different times to generate the peptide sequence data 116. In some embodiments, the therapeutic development system 100 includes a sample analyzer that is used in the processing of the set of samples 118 to generate the peptide sequence data 116. The sample analyzer may include, for example, but is not limited to, a mass spectrometry system.
[0053] The peptide sequence data 116 identifies a plurality of peptide sequences 122 detected within the set of samples 118. Each peptide sequence of the detected peptide sequences 122 characterizes at least a portion of the corresponding peptide. In other words, each peptide sequence forms at least a portion of a peptide. A peptide sequence may be, for example, without limitation, an amino acid sequence, a nucleic acid sequence, or a codon sequence. A peptide, such as a peptide in the plurality of peptides 120, may be a mutant peptide (e.g., a neoantigen) if the peptide contains one or more variants (e.g., one or more sequence mutations) compared to a corresponding reference sequence. In other words, a mutant peptide has a peptide sequence that includes a variant coding sequence that includes at least one variant relative to a corresponding reference sequence.
[0054] In one or more embodiments, the sequence analyzer 112 of the data analyzer 111 receives peptide sequence data 116 as input for processing. The sequence analyzer 112 includes a model 124 that processes the peptide sequence data 116 and a loss evaluator 125 that is used to train the model 124. In some embodiments, the peptide sequence data 116 is sent directly to the model 124 for processing. In other embodiments, the sequence analyzer 112 preprocesses the peptide sequence data 116 before sending the peptide sequence data 116 to the model 124 for processing.
[0055] Model 124 may be comprised of any number or combination of models, algorithms, functions, etc. In one or more embodiments, model 124 includes a machine learning model, which may be implemented in any of several different ways. For example, model 124 may include a deep learning model. A deep learning model may include, for example, but is not limited to, one or more convolutional neural networks, one or more recurrent neural networks, one or more feedforward neural networks, one or more other types of neural networks, or a combination thereof.
[0056] In one or more embodiments, model 124 includes various subsystems of processing. Each "subsystem" may be composed of one or more blocks, and each block may be composed of one or more sub-blocks and / or layers. A sub-block may be composed of any number of layers (or units).
[0057] In one or more embodiments, model 124 includes an encoder-decoder model. For example, model 124 may include a sequence-to-sequence (seq2seq) learning model, which may be implemented in any of several different ways. For example, the sequence-to-sequence learning model may take the form of an attention-based machine learning model (e.g., including one or more attention layers). The sequence-to-sequence learning model may include, but is not limited to, one or more recurrent neural networks. For example, the sequence-to-sequence learning model may include a multi-layer Long Short-Term Memory model.
[0058] The model 124 may include multiple subsystems (or subnetworks). Each of the multiple subsystems may include an encoder, a transformer, a transformer-encoder, one or more attention layers, and / or one or more self-attention layers. For example, the model 124 may include one or more encoders configured to transform an input (e.g., a sequence representation representing, for example, an amino acid sequence, a nucleic acid sequence, a codon sequence, etc.) into a higher-dimensional space. The encoder may be a transformer encoder. The encoder may be configured to implement an attention-based technique and / or include one or more attention layers (e.g., one or more self-attention layers). The model 124 may use a self-attention mechanism, a global attention mechanism, a soft attention mechanism, a local attention mechanism, and / or a hard attention mechanism. The model 124 may include one or more functions, such as, but not limited to, at least one of a content-based function, an adjunction function, a position-based function, a dot-product function, a scaled dot-product function, or another function.
[0059] In one or more embodiments, the model 124 is trained using metric learning to learn a representation function that maps input peptide sequences to peptide sequence vectors in an embedding space. Thus, the model 124 may be referred to as a metric learning model (e.g., a deep metric learning model). The embedding space may be an n-dimensional space. Thus, the peptide sequence vector may be a vector of n dimensions (or n discrete elements). Such a vector may also be referred to as an embedding. The n-dimensional space may include, but is not limited to, for example, 2, 5, 10, 20, 30, 50, 75, 100, 128, 200, 256, 300, 400, 450, 500, 800, 1024, 1600, 2048, 2500, or some other dimension (e.g., up to 3000 dimensions).
[0060] The distance between peptide sequence vectors (or embeddings) for various peptide sequences within the embedded space preserves the similarity of the peptide sequences (and thereby the corresponding peptides). For example, the distance (calculated distance metric) between any two peptide sequence vectors within the embedded space provides an indication of how similar the corresponding peptide sequences (and thereby the corresponding peptides) are to each other. This distance may be, for example, but is not limited to, Euclidean distance, cosine distance, or some other type of distance metric. A shorter distance between two peptide sequence vectors indicates that the corresponding peptides are more similar, while a longer distance means that the corresponding peptides are less similar (or more dissimilar).
[0061] For example, the position of a peptide sequence vector within the embedded space can provide information about the binding properties and / or abilities of the corresponding peptide. For example, two peptide sequence vectors that are close to each other within the embedded space can correspond to peptides with the same or similar binding motifs. Two peptide sequence vectors that are far apart within the embedded space can correspond to peptides with different binding motifs.
[0062] The model 124 is trained using metric learning via a loss evaluator 125. The loss evaluator 125 may use a metric learning algorithm 126 to adjust the parameters of the model 124. The loss evaluator 125 is used to train the model 124 so that peptide sequences of the same class (and thereby corresponding peptides) are mapped close to each other in the embedding space, and peptide sequences of different classes (and thereby corresponding peptides) are mapped farther apart in the embedding space. The class of a peptide sequence may be, for example, an individual MHC allele that is expected (or predicted) to present the peptide corresponding to the peptide sequence. Thus, the model 124 is trained so that peptide sequences presented by the same MHC allele are embedded closer to each other in the n-dimensional space compared to peptide sequences presented by different MHC alleles.
[0063] The MHC gene family is divided into three subgroups: MHC class I, MHC class II, and MHC class III. The genes in these subgroups can be highly polymorphic and can each contain thousands of different individual MHC alleles, each identifiable via an allelic identifier. Two different MHC alleles can be both MHC class I, both MHC class II, or can contain a first allele of MHC class I and a second allele of MHC class II. Thus, two peptide sequences belong to different classes for purposes of the metric learning algorithm 126 if they correspond to peptides that are presented or expected to be presented by different MHC alleles with different allelic identifiers (whether these MHC alleles are both MHC class I, both MHC class II, or MHC class I and MHC class II).
[0064] Therefore, whether the MHC allele that presents a first peptide is the same as the MHC allele that presents a second peptide can be determined by the allelic identifier of the MHC allele. Each MHC allele can be identified by an allelic identifier consisting of any number of digits (e.g., 4, 6, 8, or any other number of digits). When dealing with MHC alleles in humans (i.e., HLA alleles), the allelic identifier can be composed of various letters and / or numbers that form one or more fields for representing different information about the allele. The allelic identifier can include one or more letters that indicate the corresponding MHC (HLA) gene, the expression level, or both.
[0065] The allele identifier for an HLA allele may include a four-digit, six-digit, or eight-digit identifier for the HLA allele. In a four-digit identifier, the first and second digits identify the allele group, and the third and fourth digits identify the specific allele protein. The specific allele protein is determined based on differences in the DNA sequence and the amino acid sequence of the encoded protein. A six-digit identifier adds fifth and sixth digits identifying exon region information to the four-digit identifier. The exon region information captures changes in one or more exon regions of an HLA allele, such as, for example, synonymous nucleotide substitutions. An eight-digit allele identifier adds seventh and eighth digits identifying intron region information to the six-digit identifier as described above. The intron region information captures changes in one or more intron regions of an HLA allele, such as, for example, polymorphisms in intron regions.
[0066] The metric learning algorithm 126 may include a set of loss functions 128 used to tune the parameters of the model 124. The set of loss functions 128 may include, for example, but not limited to, at least one of a contrast loss function, a triplet loss function, a quartet loss function, a circular loss function, a multi-class n-pairs loss function, a lifted structure loss function, an angular loss function, a divergence loss function, a constellation loss function, or another type of loss function depending on the learned distance metric. The distance metric may be, for example, Euclidean distance, cosine distance, or some other type of distance metric.
[0067] The metric learning algorithm 126 trains the model 124 using a set of loss functions 128 such that peptide sequence vectors for peptide sequences of the same class (i.e., corresponding to peptides presented by the same MHC allele with respect to their allelic identifiers) are moved closer to each other in the embedded space, and peptide sequence vectors for peptide sequences of different classes (i.e., corresponding to peptides presented by different MHC alleles with respect to their allelic identifiers) are moved further away from each other in the embedded space. Examples of various types of loss functions that can be included in the metric learning algorithm 126 are described in more detail below in Section III.C.
[0068] In one or more embodiments, model 124 is trained using training peptide sequence data 130 and training allele representation data 132. In one or more embodiments, training peptide sequence data 130 includes data generated based on processing of set of samples 118 (e.g., may include peptide sequence data 116). In other embodiments, training peptide sequence data 130 includes other data generated from processing of different sample sets that may differ with respect to at least one sample. In one or more embodiments, training peptide sequence data 130 may include one or more of peptide sequence 122, one or more other peptide sequences, or a combination thereof.
[0069] The training allele presentation data 132 includes information regarding which MHC alleles are expected (or predicted) to present which peptides. For example, the training allele presentation data 132 may identify multiple MHC (e.g., HLA) alleles that are expected to present various peptides. MHC alleles that are expected to present a given peptide are those that have been confirmed to present the peptide, predicted to present the peptide, or otherwise determined to be capable of presenting the peptide (e.g., based on the binding pocket of the MHC allele). In one example, for a given peptide sequence, the training allele presentation data 132 identifies one or more MHC alleles that have a binding pocket configured to bind to a binding motif present in the peptide sequence. In one or more embodiments, the training allele presentation data 132 identifies a set of MHC alleles for each peptide sequence (or training peptide sequence) in the training peptide sequence data 130. The set of MHC alleles includes one or more MHC alleles for each corresponding peptide sequence. In some embodiments, the training allele presentation data 132 identifies, for a given peptide sequence, a single MHC allele that is predicted to present the peptide (or training peptide) corresponding to the given peptide sequence.
[0070] The training allele representation data 132 may be generated based on, for example, without limitation, the training peptide sequence data 130. For example, the data analyzer 111 may include a representation model 134 that outputs the training allele representation data 132 based on input peptide sequence data (e.g., the training peptide sequence data 130). In one or more embodiments, the representation model 134 may be implemented within the sequence analyzer 112 (e.g., separate from the model or within the model 124). As an example, the representation model 134 may be implemented as a separate model that receives the training peptide sequence data 130 and outputs the training allele representation data 132.
[0071] The presentation model 134 may take the form of a machine learning model. For example, the presentation model 134 may include a deep learning model (e.g., one or more neural networks). In one or more embodiments, the presentation model 134 includes a softmax function. The presentation model 134 may be implemented, for example, using NetMHC (e.g., NetMHC 4.0, NetMHCpan, etc.). In other embodiments, the presentation model 134 may be integrated with and trained simultaneously with the model 124 to identify, for a given peptide sequence, a set of MHC alleles that are expected (or predicted) to present the peptide sequence. For example, the model 124 may be trained to generate, for a given peptide sequence, a peptide sequence vector in embedded space and predict a set of MHC alleles to present the corresponding peptide.
[0072] In one or more embodiments, the presentation model 134 receives as input a peptide sequence and outputs an allele set identifier that identifies a respective MHC allele set that is expected (or predicted) to present the peptide corresponding to the peptide sequence.
[0073] In one or more embodiments, the data analyzer 111 may search for, access, or obtain training allele representation data 132 from the data store 104, one or more other types of storage devices (e.g., databases, servers, cloud storage, etc.), another source, or a combination thereof. In this manner, the training allele representation data 132 may be pre-generated data.
[0074] Additionally, in one or more embodiments, loss evaluator 125 includes miner 135 (also referred to as sampler 135), which may be implemented using hardware, software, firmware, or a combination thereof. Miner 135 may be used to implement a mining strategy (or sampling strategy) for evaluating the loss during training. The mining strategy is selected to help prevent freezes in the training of model 124 and / or to help move model 124 more quickly toward convergence. Such mining strategies are described in more detail below in Section III.D.
[0075] After training, the model 124 may be used in a predictive mode to generate multiple peptide sequence vectors 136 for the peptide sequence 122. Each of the peptide sequence vectors 136 may be a vector within an embedded space (e.g., an n-dimensional space). The model 124 may generate these peptide sequence vectors 136 in a manner that is agnostic to the length of the peptide sequence. The positions of the peptide sequence vectors 136 relative to one another within the embedded space provide an indication of the similarity and / or dissimilarity of the corresponding peptides relative to one another.
[0076] The data analyzer 111 may generate output 140 based on the peptide sequence vectors 136. The output 140 may include the peptide sequence vectors 136, information generated using the peptide sequence vectors 136, or both. The output 140 may be generated in a variety of forms, such as, but not limited to, a visual representation, a spreadsheet, a list, and / or some other type of output. In some embodiments, the output 140 includes a list of the peptide sequence vectors 136. In some embodiments, the output 140 includes a spreadsheet identifying the peptide sequence vectors 136 as well as other information (e.g., the corresponding peptide sequence for each peptide sequence vector, the MHC alleles predicted to present the corresponding peptide, other information, or a combination thereof).
[0077] In one or more embodiments, the output 140 includes a visual (e.g., graphic) representation of the peptide sequence vector 136 in k-dimensional space. The k-dimensional space may include the same or fewer dimensions (i.e., k≦n) as the n-dimensional space of the peptide sequence vector 136. As an example, the peptide sequence vector 136 corresponding to the peptide 120 may be a vector having 32 dimensions, while the graphical representation may show the peptide 120 represented in a two- or three-dimensional space. By presenting the peptide sequence vector 136 in a reduced dimensional space, a human may be able to more quickly and / or easily understand or ascertain the similarity relationships between corresponding peptides.
[0078] The output 140 may, in some cases, classify the peptides 120 into clusters based on the position of the peptide sequence vectors 136 within the embedded space. For example, each of the peptides 120 may be assigned to a different cluster (or group or category) based on the position of its corresponding peptide sequence vector within the n-dimensional space. In one or more embodiments, peptide sequence vectors (and thus the peptide sequence and corresponding peptide) assigned to the same cluster (or group or category) may generally have the same binding motif or a set of similar binding motifs. The data analyzer 111 may use one or more clustering algorithms to identify these clusters. Such clustering algorithms include, but are not limited to, a K-means clustering algorithm, an affinity propagation clustering algorithm, an agglomerative clustering algorithm, a mini-batch K-means clustering algorithm, a mean-shift clustering algorithm, a spectral clustering algorithm, a Gaussian mixture clustering algorithm, a balanced iterative reduction and clustering (BIRCH) algorithm, a density-based spatial application and noise (DBSCAN) clustering algorithm, and an ordering points for identifying clustering structure (OPTICS) algorithm.
[0079] In one or more embodiments, the output 140 is sent to the candidate selector 114 for processing. The candidate selector 114 may include a model (e.g., a machine learning model or another type of model) trained to select a diverse set of candidate peptides 110 for development of peptide therapeutics 108 based on the output 140. For example, the candidate selector 114 may use the output 140 to rank and select the top x candidate peptides for inclusion in the diverse set of candidate peptides 110. The x candidate peptides may be, for example, but not limited to, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or some other number of candidate peptides.
[0080] The candidate selector 114 may select candidate peptides to form a diverse group of candidate peptides 110 to ensure that dissimilar candidate peptides are selected for the peptide therapeutic 108 to improve the likelihood of MHC presentation and eliciting an immune response. The candidate selector 114 may select a diverse group of candidate peptides 110 such that the peptides included are diverse with respect to, for example, binding motifs and / or MHC alleles that present the peptides.
[0081] Candidate peptides can be selected in a variety of ways. The candidate selector 114 can, for example, use the output 140 to identify multiple clusters of peptide sequence vectors 136. The output 140 can explicitly identify these clusters, or the candidate selector 114 can identify these clusters using the information in the output 140 and one or more clustering algorithms, as described above. The candidate selector 114 can then, for example, select at least one peptide sequence vector and its corresponding peptide from each of the different clusters to ensure sufficient diversity among the candidate peptides. For example, for a given cluster, the candidate selector 114 can select the peptide closest to the center of that cluster. The center can be a centroid (e.g., a centroid based on Euclidean distance), a mean center, a median center, a density-based center, or another type of center.
[0082] In some examples, the candidate selector 114 identifies a subset of clusters that are greater than a selected threshold distance from each other (e.g., with respect to their centers), and then selects at least one peptide from each cluster of this subset. For example, each cluster of the subset may have a distance from all other clusters of the subset that is greater than a threshold.
[0083] When multiple peptides are selected from a given cluster, the candidate selector 114 may select these peptides in different ways to ensure local diversity. As one example, the candidate selector 114 may identify the center of the selected cluster (e.g., centroid, mean center, median center, density-based center, etc.) and then select two or more peptides that are midway between the center and the edge of the cluster, while also maximizing the distance between the two or more peptides. In another example, the candidate selector 114 may identify two or more peptides along (e.g., or near) the edge of the selected cluster, while also maximizing the distance between two or more peptides. In yet another example, the candidate selector 114 may select multiple peptides from the selected cluster using density estimates and / or calculations. For example, the candidate selector 114 may use density estimates to identify multiple "local" centers (e.g., centroids) within a given cluster and then select the peptide closest to each of these local centers.
[0084] In one or more embodiments, the candidate selector 114 uses an algorithm to select X (e.g., 4, 5, 8, 10, 12, 15, 20, 25, 30, 50, 75, etc.) peptides such that the distance between each possible pairing of the X peptides is maximized. In some cases, the candidate selector 114 selects the X peptides such that the distance between any two of the X peptides is greater than a distance threshold. This type of selection ensures a minimum level of dissimilarity between the peptides.
[0085] The candidate selector 114 may use one or more filters to reduce the pool of peptides from which candidate peptides are selected. For example, the candidate selector 114 may reduce the pool of peptides to highly evaluated / ranked peptides that have well-known or well-defined binding motifs. The candidate selector 114 may use one or more filters to reduce "noise" so that clusters can be more clearly defined.
[0086] In one or more embodiments, the candidate peptides of the diverse group of candidate peptides 110 can be, for example, mutant peptides. For example, the candidate peptides can include peptides having peptide sequences that include variant coding sequences.
[0087] In one or more embodiments, the data analyzer 111 generates a report 142 based on the output 140. The report 142 may be generated by the data analyzer 111 using information output from the sequence analyzer 112, information output from the candidate selector 114, or both. For example, the report 142 may include at least a portion of the output 140, the identification of the group of diverse candidate peptides 110, or both. The identification of the group of diverse candidate peptides 110 may include, for example, the identification of peptide sequences corresponding to the group of diverse candidate peptides 110. In some embodiments, the report 142 may include the exact output of the model 124, a transformed or filtered version of the output 140, or both. In some cases, the data analyzer 111 may generate notifications, recommendations, alerts, or other information based on at least a portion of the output 140, the identification of the group of candidate peptides 110, or both, and this additional information is included in the report 142. In some embodiments, the report 142 may be generated by the sequence analyzer 112 and / or the candidate selector 114.
[0088] The report 142 may include, for example, a recommendation regarding which candidate peptides from the diverse group of candidate peptides 110 to select for inclusion in the peptide therapeutic 108. For example, the report 142 may identify multiple therapeutic peptides 144 for inclusion in the peptide therapeutic 108. The therapeutic peptides 144 include two or more peptides selected from the diverse group of candidate peptides 110. Further, the therapeutic peptides 144 include at least two peptides that are dissimilar (e.g., dissimilar with respect to binding motifs). In one or more embodiments, the peptide therapeutic 108 is developed to include a therapeutic peptide 144, a precursor of the therapeutic peptide 144, or one or more nucleic acids (e.g., DNA and / or RNA) encoding the therapeutic peptide 144 or a precursor of the therapeutic peptide 144. The report 142 may identify the therapeutic peptide 144, a precursor of the therapeutic peptide 144, or one or more nucleic acids (e.g., DNA and / or RNA) encoding the therapeutic peptide 144 or a precursor thereof.
[0089] The report 142 may include, for example, instructions for facilitating the manufacture of a peptide therapeutic 108 based on the recommended candidate peptides. In one or more embodiments, the report 142 may include alerts for triggering computerized processes involved in the manufacture of the peptide therapeutic 108.
[0090] The report 142, in one or more embodiments, may be displayed on a graphical user interface 150 of the display system 106. A user may, for example, view and / or interact with the report 142 via the graphical user interface 150 and use the report 142 to make decisions regarding the design and development of the peptide therapeutic 108. In some embodiments, the data analyzer 111 sends (e.g., wirelessly) the report 142 to a remote system 152. The remote system 152 may be a cloud computing platform, cloud storage, another computer system, a user device (e.g., a smartphone, tablet, laptop, etc.), or some other type of platform. In some embodiments, the remote system 152 may be, or be part of, a therapeutic manufacturing system (or machine) that may use the report 142 to manufacture the peptide therapeutic 108.
[0091] II.B. Example of a Machine Learning Model – Attention-Based Figure 2 is a schematic diagram of an example configuration of model 124 of Figure 1, according to one or more embodiments. Model 124 can be configured to receive peptide sequence 202 as input and generate peptide sequence vector 204 as output. Peptide sequence 202 can be an example of an implementation of one of peptide sequences 122 described in Figure 1 or one of the peptide sequences in training peptide sequence data 130 of Figure 1. Peptide sequence vector 204 can be an example of an implementation of one of peptide sequence vectors 136 described in Figure 1.
[0092] In one or more embodiments, model 124 takes the form of an attention-based machine learning model 201. In one or more embodiments, attention-based machine learning model 201 may be implemented as described in U.S. Patent Application Publication No. 2022 / 0122690, or WO 2022016125, each of which is incorporated by reference in its entirety.
[0093] In one or more embodiments, the attention-based machine learning model 201 includes a peptide representation block 206, an attention block 208, and a vector output subsystem 210. Each of the peptide representation block 206, the attention block 208, and the vector output subsystem 210 may include one or more sub-blocks and / or layers. A sub-block may be composed of any number of layers (or units).
[0094] The peptide representation block 206 may include at least one embedding layer 212, which may optionally include, for example, a position encoder 214. The embedding layer 212 receives the peptide sequence 202 as input and embeds the peptide sequence 202 to generate an embedded peptide representation. The peptide sequence 202 may be embedded, for example, by converting the peptide sequence 202, which is a non-numeric representation (e.g., a series of amino acid identifiers, a series of nucleic acid identifiers), into a numerical representation that results in the embedded peptide representation. The embedding may be performed using, for example, one-hot encoding, evolutionarily motivated encoding such as BLOcks Substitution Matrix (BLOSUM), randomly or pseudo-randomly initialized learned embeddings, or a combination thereof.
[0095] In some cases, various attention mechanisms may fail to detect the underlying information conveyed by the order of values in an input dataset, such as peptide sequence 202. Therefore, a positional encoder, such as positional encoder 214, may be used to apply the embedded peptide representations generated by embedding layer 212. Positional encoder 214 may perform the positional encoding using a trained or fixed encoding algorithm. Positional encoder 214 receives the embedded peptide representations from embedding layer 212 and positionally encodes the embedded peptide representations to generate peptide representation 216, which represents the peptide sequence.
[0096] Peptide representation 216 can be, for example, a multidimensional vector that represents or otherwise corresponds to each peptide element (e.g., each amino acid, nucleic acid, codon, etc.) of peptide sequence 202. For example, peptide representation 216 can be a matrix that includes a vector (e.g., having a dimension from 20 to 1000) for each peptide element of peptide sequence 202. Thus, for a peptide sequence that includes e peptide elements (e.g., amino acids, nucleic acids, or codons), the matrix can include e vectors with d dimensions (e.g., where 20≦d≦1000).
[0097] In one or more embodiments, the fixed position encoding may be defined using sine and / or cosine functions (e.g., with intra-sequence position and / or dimension as independent variables). The position encoding output by position encoder 214 may have the same dimensions as the embedded peptide representation output by embedding layer 212. In one or more embodiments, the position encoding may be summed with the embedded representation to generate a position-indicated embedded representation of the sequence. Thus, or in multiple embodiments, the peptide representation 216 generated by peptide representation block 206 may be an encoded representation or an aggregation (e.g., concatenation or sum) of the encoded representation formed by position encoder 214 and the embedded peptide representation formed by embedding layer 212.
[0098] In other embodiments, position encoder 214 may generate a unique learned embedding for each possible position in the peptide sequence. These embeddings are then added to the embedded peptide representation formed by embedding layer 212 to form peptide representation 216.
[0099] The peptide representation 216 is sent as input to the attention block 208. The attention block 208 may include one or more sub-blocks and / or layers. For example, the attention block 208 may include an attention sub-block 1218 and, optionally, one or more other attention sub-blocks up to attention sub-block n 220. If there are multiple attention sub-blocks in the attention block 208, these attention sub-blocks may be connected in series (e.g., daisy-chained together to generate the final output). In one or more embodiments, the attention block 208 may use a set of query weights, a set of key weights, and a set of value weights to determine, for a given peptide element (e.g., an amino acid) of a peptide sequence, the degree to which each of one or more other peptide elements should be “attentioned” when processing the given peptide element.
[0100] The attention sub-block 1218 may be implemented in a variety of ways. In one or more embodiments, the attention sub-block 1218 includes, but is not limited to, a self-attention layer 222, a summation and normalization layer 224, a feedforward layer 226, and a summation and normalization layer 228. Due to this configuration of the attention sub-block 1218, the attention sub-block 1218 may also be referred to as a transformer encoder. If present, one or more other attention sub-blocks, from attention block 208 to attention sub-block n 220, may be implemented in a manner similar to the attention sub-block 1218.
[0101] The self-attention layer 222 may be implemented using, for example, a one-head attention unit or a multi-head attention unit. The self-attention layer 222 converts the peptide representation 216 into a transformed representation. In the summation and normalization layer 224, the transformed representation may be added to the position-directed embedded representation of the sequence (e.g., peptide representation 216), for example, via residual concatenation, and the summed representation may be normalized.
[0102] The normalized data can be fed to a corresponding feedforward layer 226 (e.g., a fully connected feedforward network). The feedforward layer 226 can affect (for example) one, two, three, or more linear transformations for each location and / or can include activations (e.g., ReLU activations) between each of the linear transformations. For example, the feedforward layer 226 can be represented by: TIFF2025539935000002.tif5170 where x is the input to the layer, W1 and W2 are the gradients of the linear transformation, and b1 and b2 are the intercepts of the linear transformation. The dimensionality of the output of a particular attention sub-block's feedforward layer may be the same as the dimensionality of the input to the attention sub-block's feedforward layer. Thus, in some cases, the inputs and outputs can be summed and normalized (e.g., via another residual connection through another summation and normalization layer, such as summation and normalization layer 228) to preserve representations of various types of information.
[0103] The attention block 208 receives and processes the peptide representation 216 using a set of attention sub-blocks to generate as output a transformed peptide representation 230. The transformed peptide representation 230 may be a matrix containing a vector (e.g., with dimensions from 20 to 1000) for each peptide element in the peptide sequence 202. The transformed peptide representation 230 may be sent to the vector output subsystem 210 for processing.
[0104] The vector output subsystem 210 may include various blocks, sub-blocks, layers, or combinations thereof for generating the final output of the model 124, including the peptide sequence vector 204. In one or more embodiments, the vector output subsystem 210 includes an averaging block 232, a fully connected block 234, a dropout block 236, an activation block 238, and a fully connected block 240. Each of the fully connected block 234 and the fully connected block 240 may include, for example, one or more fully connected layers. The dropout block 236 may include, for example, one or more dropout layers. The activation block 238 may include one or more activation layers (e.g., linear or nonlinear functions), such as, for example, a nonlinear function, such as a rectified linear function.
[0105] The transformed peptide representation 230 may be further processed before the information is sent to the all-join block 234. In one or more embodiments, an averaging block 232 is used to average the various vectors of the transformed peptide representation 230. For example, the vectors of the transformed peptide representation 230 for different peptide elements of the peptide sequence 202 may be averaged together to form an average representation of the peptide sequence 202 that includes a single numerical value for each peptide element (e.g., amino acid, nucleic acid, or codon) of the peptide sequence 202. Thus, the averaged representation may be a single vector. This averaged representation may be sent as an input to the all-join block 234.
[0106] In some embodiments, the averaging block 232 may be replaced by a block that concatenates different vectors together to form an aggregate, which may then add a start of sequencing (BoS) token before aggregation to form a new representation that is sent to the join all block 234.
[0107] In some embodiments, the fully connect block 234 is configured to output a vector having fewer dimensions than the average representation provided as input to the fully connect block 234. The fully connect block 234 may include, for example, a single hidden layer, two hidden layers, or three or more hidden layers. The number of nodes in an initial hidden layer may be greater than the number of nodes in a subsequent hidden layer. For example, the first hidden layer may include 256 nodes, and the second hidden layer may include 126 nodes.
[0108] The dropout block 236 may be used to apply dropout regularization to one or more layers of the fully connected block 234. In particular, the dropout block 236 may be used to disable some portion of neurons in one or more of the hidden layers within the fully connected block 234.
[0109] The activation block 238 may be used to apply a non-linear activation function to the fully connected block 234. For example, the activation block 238 may use one or more rectified linear units (ReLUs) to convert any negative values to zero.
[0110] The fully connect block 240 may receive as input the output produced after applying dropout regularization and the activation function of the fully connect block 234, and generate an output that is the peptide sequence vector 204. Similar to the fully connect block 234, the fully connect block 240 may be configured to output a vector having fewer dimensions than the input provided to the fully connect block 240. The fully connect block 240 may include, for example, a single hidden layer, two hidden layers, or three or more hidden layers. The number of nodes in an initial hidden layer may be greater than the number of nodes in a subsequent hidden layer.
[0111] Peptide sequence vector 204 may be a vector having n dimensions. In some cases, the n dimensions may be selected so that training (or learning) of model 124 is sufficiently robust and rich. For example, in some instances, it may be desirable for model 124 to generate peptide sequence vector 204 in a low-dimensional space (e.g., n is less than 250) to promote more robust or substantial learning by model 124. In one or more embodiments, peptide sequence vector 204 is a vector in an n-dimensional space having 16 dimensions, 32 dimensions, 64 dimensions, 128 dimensions, or some other number of dimensions.
[0112] Peptide sequence vector 204 captures, represents, or provides information about peptide sequence 202. For example, peptide sequence vector 204 may capture information about peptide sequence 202 such that a peptide sequence vector 204 relative to another peptide sequence provides an indication of the similarity or dissimilarity of the two peptide sequences.
[0113] 2 is just one example of an implementation of attention-based machine learning model 201 (and thereby model 124). Other embodiments may use other configurations for attention-based machine learning model 201. For example, in one or more embodiments, vector output subsystem 210 may include one or more other layers for filtering, selecting, transforming, or otherwise modifying the output of any one or more of the blocks or layers of vector output subsystem 210 to ultimately generate peptide sequence vector 204.
[0114] III. Examples of Methods Used to Select Diverse Peptides for Peptide Therapeutics III.A. Selection of a diverse set of candidate peptides using machine learning models 3 is a flowchart of a process for use in developing a therapeutic agent, according to one or more embodiments. Process 300 may be implemented, for example, using therapeutic agent development system 100 described in FIG. 1. Process 300 may be implemented for use in developing a peptide therapeutic agent. For example, process 300 may be used to select a group of diverse candidate peptides, such as group of diverse candidate peptides 110 described in FIG. 1, for use in developing a peptide therapeutic agent, such as peptide therapeutic agent 108 of FIG. 1.
[0115] Process 300 may include step 302. Step 302 includes training a machine learning model using a metric learning algorithm, training peptide sequence data, and training allele representation data corresponding to the training peptide sequence data. The machine learning model may be trained to generate a peptide sequence vector (or embedding) for a given peptide sequence. The peptide sequence may take the form of, for example, an amino acid sequence, a nucleic acid sequence, a codon sequence, or another type of sequence that defines at least a portion of the corresponding peptide. The machine learning model may be, for example, model 124 of FIG. 1 or FIG. 2. The deep learning model may be, for example, attention-based machine learning model 201 of FIG. 2. The machine learning model may be, for example, a deep learning model and may include, but is not limited to, one or more neural networks.
[0116] The training peptide sequence data may be, for example, training peptide sequence data 130 in Figure 1. The training peptide sequence data includes peptide sequences for training (also referred to as training peptide sequences).
[0117] The training allele representation data used in training the machine learning model may identify, for each peptide sequence in the training peptide sequence data, one or more MHC alleles having a binding pocket (or binding groove) configured to bind to the binding motif of the peptide sequence. The training allele representation data may be, for example, training allele representation data 132 of FIG. 1.
[0118] In one or more embodiments, the training allele representation data is generated independent of the training in step 302. For example, the training allele representation data may be generated by another model (e.g., representation model 134 of FIG. 1 ) prior to the training in step 302. The training allele representation data may then be stored for later use in step 302. In some cases, the training allele representation data is stored in a data store (e.g., data store 104 of FIG. 1 ) or some other type of data storage device or source.
[0119] In other embodiments, the training allele representation data is generated as part of the training in step 302. For example, the machine learning model may include a first system for generating a peptide sequence vector for a given peptide sequence and a second system for identifying a set of MHC alleles predicted to represent the given peptide sequence. The second subsystem may be trained before the first subsystem is trained.
[0120] The metric learning algorithm used to train the machine learning model may be, for example, metric learning algorithm 126 of FIG. 1 . The metric learning algorithm includes one or more loss functions used to evaluate learned distance metrics of peptide sequence vectors for various peptide sequence groups (e.g., two, three, four, or more peptide sequences) to influence how parameters (e.g., weights) of the machine learning model are adjusted after each training batch and / or each epoch. For example, the learned distance metric between pairwise peptide sequence vectors may be Euclidean distance, cosine distance, or some other type of distance metric. In one or more embodiments, the metric learning algorithm includes at least one of a contrast loss function, a triplet loss function, a quartet loss function, a circular loss function, a multi-class n-pairs loss function, a lifted structure loss function, an angle loss function, a divergence loss function, a constellation loss function, or another type of loss function that depends on the learned distance metric.
[0121] The training allele representation data is used to determine how the metric learning algorithm trains the machine learning model. For example, in one or more embodiments, the metric learning algorithm can use the training allele representation data to train the machine learning model to generate peptide sequence vectors for peptides represented by the same MHC allele that are closer to each other in n-dimensional space, and to generate peptide sequence vectors for peptides represented by different MHC alleles that are farther apart in n-dimensional space. Examples of methods that can be used for training are described in Section III.B below.
[0122] Step 304 of process 300 includes receiving peptide sequence data identifying a plurality of peptide sequences for a plurality of peptides. The peptide sequence data may be, for example, peptide sequence data 116 of FIG. 1 . The peptide sequence data may be formed, for example, by at least a portion of training sequence data. In other embodiments, the peptide sequence data may be data generated from processing a set of samples (e.g., sample set 118 of FIG. 1 ). In one or more embodiments, the peptide sequence data may include peptide sequences that are also included in the training sequence data. In one or more embodiments, the peptide sequence data may include peptide sequences that were not included in the training peptide sequence data. The peptide sequence of a peptide may be, for example, peptide sequence 122 of peptide 120 of FIG. 1 . Peptide sequence 202 of FIG. 2 may be an example of one of the plurality of peptide sequences in step 304.
[0123] Step 306 includes generating, via the trained machine learning model, a plurality of peptide sequence vectors in n-dimensional space for each of the plurality of peptide sequences, and thereby for each of the plurality of peptides. The machine learning model is trained via step 302 using a metric learning algorithm, training peptide sequence data, and training allele presentation data corresponding to the training peptide sequence data. The training peptide sequence data identifies a training peptide sequence corresponding to each training peptide of the plurality of training peptides. The training allele presentation data identifies, for each training peptide sequence in the training peptide sequence data, a set of major histocompatibility complex (MHC) alleles that are expected to present the training peptide corresponding to the training peptide sequence.
[0124] A metric learning algorithm is used to train a machine learning model so that peptide sequence vectors for peptides presented by the same MHC allele are embedded closer to each other in n-dimensional space, and peptide sequence vectors for peptides presented by different MHC alleles are embedded farther from each other in n-dimensional space. Thus, the position of any peptide sequence vector relative to n-dimensional space can provide an indication of its similarity or dissimilarity to other peptide sequence vectors in n-dimensional space, and thus to their corresponding peptides.
[0125] For example, a metric learning algorithm is used to train a machine learning model so that a first distance between a first pair of peptide sequence vectors generated in n-dimensional space for each first pair of peptides presented by the same MHC allele is smaller than a second distance between a second pair of peptide sequence vectors generated in n-dimensional space for a second pair of peptides presented by different MHC alleles. The first pair of peptide sequence vectors and the second pair of peptide sequence vectors may contain a common peptide sequence vector or completely different peptide sequence vectors. For example, the first pair may contain peptide sequence vectors A (PSV-A) and PSV-B, and the second pair may contain PSV-A and PSV-C. In another example, the first pair may contain PSV-A and PSV-B, and the second pair may contain PSV-C and PSV-D.
[0126] In one or more embodiments, step 306 includes converting each peptide sequence (e.g., peptide sequence 202 of FIG. 2 ) to a peptide representation (e.g., peptide representation 216 of FIG. 2 ). The peptide representation may include multiple vectors, with a vector for each peptide element (e.g., amino acid, nucleic acid, or codon) of the peptide sequence. Each vector in the peptide representation may include e elements (e.g., 20≦e≦1000). Step 306 further includes converting each peptide representation to a peptide sequence vector (e.g., peptide sequence vector 204 of FIG. 2 ). In one or more embodiments, each of the n elements of the peptide sequence vector includes fewer elements than the e elements of the vector of the peptide representation. However, in some cases, the n elements may include more elements than the peptide elements that made up the original peptide sequence.
[0127] Step 308 includes generating an output using the plurality of peptide sequence vectors, the output providing an indication of inter-peptide similarity of the plurality of peptides for use in selecting a diverse group of candidate peptides from the plurality of peptides for therapeutic development. In one or more embodiments, the output may be, for example, output 140 of Figure 1. The output may include the plurality of peptide sequence vectors, information generated based on the plurality of peptide sequence vectors, or both.
[0128] The output may include, for example, but is not limited to, a visual (e.g., graphic) representation of the peptide sequence vector in k-dimensional space. The k-dimensional space may contain the same or fewer dimensions as the n-dimensional space of the peptide sequence vector. As an example, the peptide sequence vector of a peptide may be a vector having 32 dimensions, but the graphical representation may show the peptide represented in two- or three-dimensional space. This reduction in dimension may allow a person to more easily understand and quickly confirm the distance relationships between different peptide sequence vectors.
[0129] In one or more embodiments, the output may result in a classification of peptides into clusters. For example, each peptide may be assigned to a different cluster (or group or category) based on its position in n-dimensional space relative to the positions of other peptides. Each cluster (or group or category) may generally correspond to a unique binding motif or a set of similar binding motifs.
[0130] Process 300 may optionally, in one or more embodiments, include step 310. Step 310 includes selecting a group of candidate peptides from the plurality of peptides for therapeutic development based on the output. For example, the output generated in step 306 may be used to select a diverse group of candidate peptides (e.g., group of diverse candidate peptides 110 in FIG. 1 ) for development of a peptide therapeutic (e.g., peptide therapeutic 108 in FIG. 1 ). The diverse group of candidate peptides may include at least two dissimilar candidate peptides. The peptide therapeutic may be, for example, a peptide vaccine.
[0131] A diverse group of candidate peptides can be selected such that the peptides included are diverse with respect to, for example, but not limited to, binding motifs. For example, a diverse group of candidate peptides can include at least two candidate peptides with different binding motifs. In some cases, the selection in step 310 can be performed such that bias toward any single binding motif is reduced.
[0132] In one or more embodiments, the output may be used to rank and select the top x number of candidate peptides for inclusion in a diverse group of candidate peptides for development of a peptide vaccine (e.g., a neo-antigen vaccine). The x number of candidate peptides may be, for example, but not limited to, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or some other number of candidate peptides.
[0133] The selection of a diverse group of candidate peptides in step 310 can be performed in various ways. In one or more embodiments, if the output generated in step 308 identifies or presents an indication of clusters corresponding to different binding motifs and / or different sets of similar binding motifs, one, two, or three peptides may be selected from each cluster to ensure diversity of binding motifs. For example, peptides closest to the center of each cluster may be selected as candidate peptides. The centers of the clusters may be centroids (e.g., centroids based on Euclidean distance), mean centers, median centers, density-based centers, or centers of different types of clusters.
[0134] In other embodiments, two or more peptides may be selected such that they are the most dissimilar peptides. For example, the output generated in step 308 may provide a ranking of pairs of peptides based on distance. A diverse group of candidate peptides may then be selected by selecting some number (e.g., 1, 2, 3, 4, etc.) of peptide pairs that have the greatest distance from each other.
[0135] In yet other embodiments, where the output takes the form of a visual representation in k-dimensional space, step 310 may be performed by dividing the k-dimensional space into multiple regions and selecting one or more candidate peptides from each of the multiple regions. In one example, the k-dimensional space may be a two- or three-dimensional space divided into quadrants, with one or more candidate peptides selected from each different quadrant. In some examples, one candidate peptide is selected from each quadrant, the candidate peptide being furthest from the other quadrants.
[0136] In some embodiments, the information provided by the output can be used in combination with other information to select a diverse group of candidate peptides. This other information can include, for example, quantitative information specifying the quantification of each of the peptide sequences detected in a given sample. In some cases, the other information can include, for example, biological information about the subject to be administered the peptide therapeutic.
[0137] The various methods for selecting a diverse group of candidate peptides described above are merely examples of ways in which step 310 can be performed to ensure peptide (e.g., binding motif) diversity. Selecting a group of candidate peptides with such peptide (e.g., binding motif) diversity can help improve the overall likelihood of presentation of one or more peptides included in a peptide therapeutic by one or more MHC alleles in a subject. For example, having dissimilar peptides in a peptide therapeutic can help account for biological circumstances that may result in different binding motifs actually being preferred in a subject compared to those expected. Such biological circumstances may include, for example, the microenvironment of the tumor or diseased tissue (e.g., pH, temperature, etc.), the biological makeup of a particular subject, one or more comorbidities, etc. Thus, selecting a group of candidate peptides with binding motif diversity can help improve the overall efficacy of a peptide therapeutic.
[0138] Process 300 may optionally include step 312. Step 312 includes generating a report based on the group of diverse candidate peptides for use in therapeutic development. The report may be, for example, similar to report 142 of FIG. 1. The report may include, for example, at least a portion of the output generated in step 308, an identification of the group of diverse candidate peptides selected in step 310 (e.g., by identifying peptide sequences corresponding to the group of diverse candidate peptides), a peptide sequence vector generated in step 306, or a combination thereof. In some embodiments, the report may include a transformed or filtered version of the output generated in step 308. The report may include, for example, one or more notifications, recommendations, alerts, and / or other information.
[0139] In one or more embodiments, the report may include a recommendation as to which candidate peptides from the diverse group of candidate peptides 110 to select as therapeutic peptides for inclusion in a therapeutic drug. The report may identify precursors of the therapeutic peptides and / or nucleic acid sequences encoding the therapeutic peptides and / or their precursors.
[0140] The report may include, for example, instructions to facilitate manufacturing of the therapeutic. In one or more embodiments, the report may include alerts to trigger computerized processes involved in manufacturing the therapeutic. The report may be used to make decisions regarding the design and development of the therapeutic and / or to manufacture the therapeutic.
[0141] III.B. Training Machine Learning Models III.B.1. Examples of Training Methods - General 4 is a flowchart of a process for training a machine learning model to generate a peptide sequence vector for a peptide sequence, according to one or more embodiments. Process 400 may be implemented, for example, using therapeutic drug development system 100 described in FIG. 1. Process 400 may be implemented, for example, to train a machine learning model, such as, but not limited to, model 124 of FIG. 1 and / or attention-based machine learning model 201 of FIG. 2. Process 400 may be an example of an implementation of a process that may be used to perform step 302 of FIG. 3.
[0142] Step 402 includes receiving training peptide sequence data. The training peptide sequence data may be, for example, training peptide sequence data 130 of FIG. 1. The training peptide sequence data includes peptide sequences (referred to as training peptide sequences), each of which forms at least a portion of a peptide. The training peptide sequences in the training peptide sequence data may take the form of, for example, amino acid sequences, nucleic acid sequences, codon sequences, or another type of sequence that defines at least a portion of a corresponding peptide.
[0143] Step 404 includes generating training allele representation data for the training peptide sequence data. The training allele representation data may be, for example, allele representation data 132 of Figure 1. In some embodiments, the training allele representation data is generated by a model, such as representation model 134 of Figure 1.
[0144] The training allele presentation data identifies a set of MHC alleles that are expected to present a peptide corresponding to each peptide sequence in the training peptide sequence data. In one or more embodiments, if multiple MHC alleles are capable of presenting a given peptide having a given peptide sequence, the MHC allele most likely to present the given peptide is included in the training allele presentation data.
[0145] When dealing with human training peptide sequence data, the training allele representation data includes a label for each peptide sequence in the training peptide sequence data, where the label identifies the HLA allele predicted to represent the peptide corresponding to the peptide sequence. The label may be, for example, an allele identifier composed of various letters and / or digits forming one or more fields for representing different information about the HLA allele. As described above in Section II.A., the allele identifier may include one or more letters indicating the corresponding HLA gene, the expression level, or both. The allele identifier for an HLA allele may include, for example, but is not limited to, a four-digit, six-digit, or eight-digit identifier for the HLA allele.
[0146] Step 406 includes training a machine learning model using training peptide sequence data, training allele representation data, a metric learning algorithm, and at least one mining strategy selected based on the metric learning algorithm. The metric learning algorithm includes at least one of a contrast loss function, a triplet loss function, a quartet loss function, a circle loss function, a multi-class n-pairs loss function, a lifted structure loss function, an angle loss function, a divergence loss function, or a constellation loss function. The metric learning algorithm may be, for example, metric learning algorithm 126 of FIG. 1 .
[0147] The at least one mining strategy used in step 406 may be selected based on the type of loss function included in the metric learning algorithm. In one or more embodiments, step 406 includes using a hard negative mining strategy, a semi-hard negative mining strategy, or both. Such mining strategies are described in more detail in Section III.D below.
[0148] III.B.2. Training Method Example - Batch Processing Figure 5 is a flowchart of a process for training a machine learning model to generate a peptide sequence vector for a peptide sequence, according to one or more embodiments. Process 500 may be implemented, for example, using therapeutic drug development system 100 described in Figure 1. Process 500 may be implemented, for example, to train a machine learning model, such as, but not limited to, model 124 of Figure 1 and / or attention-based machine learning model 201 of Figure 2. Method 500 may be an example of a process that may be used to implement step 302 of Figure 3. In some cases, process 500 may be an example of an implementation of a process used to perform step 406 described in Figure 4.
[0149] Step 502 involves selecting a batch of training peptide sequences from the training peptide sequence data for processing. The training peptide sequence data may be, for example, training peptide sequence data 130 of FIG. 1. The training peptide sequence data may be the training peptide sequence data received in step 402 of FIG. 4. The batch of training peptide sequences may include all of the training peptide sequences or a subset of the training peptide sequences. For example, but not limited to, the batch may include 5, 10, 20, 30, 50, 60, 70, 80, 90, 100, 150, 200, 250, or some other number of training peptide sequences.
[0150] Step 504 includes generating a batch of training peptide sequence vectors for the batch of training peptide sequences via a machine learning model. The machine learning model may be, for example, model 124 of FIG. 1. The deep learning model may be, for example, attention-based machine learning model 201 of FIG. 2. The machine learning model may be, for example, a deep learning model and may include, for example, but is not limited to, one or more neural networks.
[0151] The batch of training peptide sequence vectors includes a training peptide sequence vector for each training peptide sequence in the batch of training peptide sequences. Each training peptide sequence vector in this batch of training peptide sequence vectors can be an n-dimensional vector (or embedding) representing the corresponding training peptide sequence.
[0152] Step 506 includes calculating a distance metric for pairs of training peptide sequence vectors in the batch of training peptide sequence vectors. In one or more embodiments, step 506 includes calculating a distance metric (e.g., Euclidean distance) for each pairing of training peptide sequence vectors in the batch of training peptide sequence vectors. In this manner, a distance metric can be calculated for the distance between each training peptide sequence vector and every other training peptide sequence vector.
[0153] Step 508 may include identifying a mining strategy for the batch of training peptide sequence vectors. In one or more embodiments, step 508 may be performed by, for example, but not limited to, loss evaluator 125 of FIG. 1. For example, step 508 may be performed by miner 135 of loss evaluator 125 of FIG. 1. In step 508, identifying a mining strategy may include selecting a mining strategy from a set of mining strategies based on the batch number of the current batch of training peptide sequence vectors. The mining strategy identified in step 508 is the strategy used in generating a group of training peptide sequence vectors from the current batch of training peptide sequence vectors for evaluating loss. For example, the selected mining strategy may determine which pairs, triplets, quartets, or other types of multiplets of training peptide sequence vectors the loss is calculated for.
[0154] The mining strategy may be, for example, an all-in strategy, a hard negative mining strategy, a semi-hard negative mining strategy, another type of strategy, or a combination thereof. In one or more embodiments, the all-in strategy refers to using all possible unique pairs of training peptide sequence vectors in a batch of training peptide sequence vectors to evaluate the loss. Hard negative mining and semi-hard negative mining are described in more detail in Section III.D below.
[0155] In one or more embodiments, an all-in strategy is selected for a first portion of the batches (e.g., batches up to batch number 10). A semi-hard negative mining strategy may be used for a second portion of the processed batches (e.g., batches between batch number 10 and batch number 40). A hard negative mining strategy may be used for a third portion of the processed batches (e.g., batches after batch number 40). Other embodiments may use different combinations of mining strategies and / or different batch cutoffs.
[0156] Step 510 includes forming an evaluation bundle from the batch of training peptide sequence vectors based on the identified mining strategy and distance metric. Step 510 may be performed, for example, but not limited to, by the loss evaluator 125 of FIG. 1. For example, step 510 may be performed by the miner 135 of the loss evaluator 125 of FIG. 1.
[0157] In one or more embodiments, an evaluation bundle is formed using each peptide sequence vector in a batch of training peptide sequence vectors. In other embodiments, an evaluation bundle is formed using a portion or subset of the batch of training peptide sequence vectors. The evaluation bundle includes groupings of training peptide sequence vectors (e.g., groupings formed by all or a subset of the training peptide sequence vectors in the batch). Each grouping includes at least two training peptide sequence vectors. For example, the groupings can be pairs, triplets, quartets, or some other multiplet or training peptide sequence vectors.
[0158] Step 510 may include forming a group of training peptide sequence vectors from the batch of training peptide sequence vectors based on the identified mining strategy and a metric learning algorithm used to evaluate losses. The metric learning algorithm may include at least one of a contrast loss function, a triplet loss function, a quartet loss function, a circle loss function, a multi-class n-pairs loss function, a lifted structure loss function, an angle loss function, a divergence loss function, or a constellation loss function. The metric learning algorithm may be, for example, metric learning algorithm 126 of FIG. 1.
[0159] When the metric learning algorithm takes the form of a triplet loss function, the groupings formed in the evaluation bundle in step 510 are triplets. The identified mining strategy determines how these triplets are formed.
[0160] Step 512 includes evaluating the loss of the evaluation bundle using a metric learning algorithm. In one or more embodiments, step 512 may be performed by, for example, but not limited to, loss evaluator 125 of FIG. 1 . The loss may be evaluated in various ways in step 512. For example, the loss may be evaluated for each grouping (e.g., pair, triplet, etc.) of the evaluation bundle, where various losses may be evaluated collectively. In some embodiments, the loss may be calculated as the sum or average of the losses calculated for each grouping in the evaluation bundle. In one or more embodiments, the loss may be evaluated such that the loss is lower if (1) peptides presented by the same MHC allele are embedded closer to each other in n-dimensional space (as training peptide sequence vectors) and (2) peptides presented by different MHC alleles are embedded farther from each other in n-dimensional space. Thus, the loss is higher if (1) peptides presented by the same MHC allele are embedded farther from each other in n-dimensional space and (2) peptides presented by different MHC alleles are embedded closer to each other in n-dimensional space. Examples of loss functions that may be used in step 512 are described in more detail below in Section III.C.
[0161] Step 514 includes updating parameters of the machine learning model based on the loss. For example, step 514 may be performed to reduce or minimize the loss calculated in step 512. The machine learning model parameters may also be referred to as weights in some cases. Updating the parameters in step 514 may include, for example, changing at least one parameter of the machine learning model. Thus, in some cases, some of the parameters may be changed and other portions may not be changed, and in other cases, all of the parameters may be changed. In one or more embodiments, step 512 and step 514 are integrated.
[0162] Step 518 includes determining whether any unprocessed training peptide sequence vectors remain. If any unprocessed training peptide sequence vectors remain, process 500 returns to step 502 above. Otherwise, if no unprocessed training peptide sequence vectors remain, process 500 ends, completing one epoch of processing. One epoch comprises a complete pass of processing through the entire training peptide sequence data. Any number of epochs may be processed as part of process 500 to train a machine learning model. For example, process 500 may be repeated any number of times to train a machine learning model. In one or more embodiments, process 500 may be repeated until the loss evaluated in step 410 falls within a selected tolerance, until the machine learning model reaches convergence, or until the machine learning model reaches a selected tolerance for convergence.
[0163] Different implementations of process 500 may be used. In some embodiments, the parameters of the machine learning model may be updated after every epoch rather than after every batch of processing. For example, step 514 may be performed after an entire epoch is completed (e.g., after step 518).
[0164] For example, using a mining strategy such as hard negative mining or semi-hard negative mining in process 500 may reduce the overall computational resources and time that may be required to evaluate the loss and perform training of the machine learning model. Additionally, using a mining strategy such as hard negative mining or semi-hard negative mining in process 500 may help the machine learning model move more quickly towards convergence and / or prevent training freezes.
[0165] III.C. Loss Evaluation: An Overview of Loss Functions Used in Training Machine Learning Models As described above with respect to process 300 of Figure 3, process 400 of Figure 4, and process 500 of Figure 5, a metric learning algorithm may be used to train a machine learning model to generate peptide sequence vectors for the peptide sequences. The machine learning model may be, for example, deep learning model 124 of Figure 1 or attention-based machine learning model 201 of Figure 2. The metric learning algorithm, which may be, for example, metric learning algorithm 126 of Figure 1, may include at least one of a contrast loss function, a triplet loss function, a quartet loss function, a circle loss function, a multi-class n-pairs loss function, a lifted structure loss function, an angle loss function, a divergence loss function, a constellation loss function, or some other type of distance-based loss function.
[0166] Various loss functions evaluate the loss of grouping peptide sequence vectors based on the class to which these peptide sequence vectors belong. Here, the class can be determined by the specific MHC alleles that present or are expected to present a given peptide, as identified by an allelic identifier (e.g., a 4-digit, 6-digit, or 8-digit allelic identifier). For example, two peptide sequence vectors representing peptide sequences (and thereby peptides) presented by the same MHC allele (e.g., MHC alleles with the same allelic identifier) belong to the same class. Two peptide sequence vectors representing peptide sequences (and thereby peptides) presented by two different MHC alleles (e.g., MHC alleles with different allelic identifiers) belong to different classes.
[0167] III.C.1. Contrast loss The contrast loss function focuses on evaluating a pair of peptide sequence vectors at a time. If two peptide sequence vectors belong to the same class (i.e., represent peptides presented by the same MHC allele), the loss is higher if the peptide sequence vectors are further apart and lower if the peptide sequence vectors are closer to each other. If two peptide sequence vectors belong to different classes (i.e., represent peptides presented by different MHC alleles), the loss is higher if the peptide sequence vectors are closer to each other and lower if the peptide sequence vectors are further apart.
[0168] In one or more embodiments, an example of a contrast loss function that can be used to evaluate a single pair of peptide sequence vectors is defined as follows: TIFF2025539935000003.tif17170 formula, Y = 0 for pairs of peptide sequence vectors belonging to different classes; Y = 1 for pairs of peptide sequence vectors belonging to the same class; D W is the distance between the two peptide sequence vectors of a pair (e.g., Euclidean distance); m is the margin.
[0169] In one or more embodiments, the margin is a constant set such that the loss function penalizes the model if the distance between peptide sequence vectors belonging to different classes is less than m. However, if this distance is equal to or less than m, the loss function is set to 0. This ensures that peptide sequence vectors of different classes are not separated more than necessary.
[0170] Referring back to Figure 5, in some embodiments, the evaluation of loss in step 512 of Figure 5 may be performed by evaluating the loss for each pair of peptide sequence vectors in the evaluation bundle. In other embodiments, the evaluation of loss in step 512 may be performed by summing, averaging, or otherwise combining or aggregating the losses calculated for each pair of peptide sequence vectors in the evaluation bundle.
[0171] III.C.2. Triplet Loss Function The triplet loss function evaluates loss based on triplets, each of which includes an anchor, a positive example, and a negative example. The anchor peptide sequence vector is a peptide sequence vector in which positive and negative are defined. For example, the positive example may be a peptide sequence vector of the same class as the anchor. In other words, the positive example and the anchor may represent peptide sequences presented by the same MHC allele. The negative example may be a peptide sequence vector of a different class from the anchor. In other words, the negative example and the anchor may represent peptide sequences presented by different MHC alleles. Using the triplet loss function, the machine learning model is trained to increase (e.g., maximize) the distance between the anchor and the negative example and to decrease (e.g., minimize) the distance between the anchor and the positive example.
[0172] In one or more embodiments, an example triplet loss function is defined as follows: TIFF2025539935000004.tif7170, r a is the anchor representation (e.g., anchor peptide sequence vector), and r p is a positive expression (e.g., a positive peptide sequence vector that belongs to the same class as the anchor peptide sequence vector), and r nis the representation of the negative (e.g., a peptide sequence vector that belongs to a different class from the anchor peptide sequence vector), d() is the distance function, and m is the margin. The margin may be a constant set based on the objective that the distance between the anchor and the negative must be greater than the margin m.
[0173] 5, in some embodiments, the evaluation of loss in step 512 of FIG. 5 may be performed by evaluating the loss of each triplet in the evaluation bundle. In other embodiments, the evaluation of loss in step 512 may be performed by summing, averaging, or otherwise combining or aggregating the losses calculated for each triplet in the evaluation bundle.
[0174] III.C.3. Other Loss Function Examples Any number of other loss functions may be used by the metric learning algorithm. Regardless of the loss function selected, the metric learning algorithm trains the machine learning algorithm to increase the distance between peptide sequence vectors belonging to different classes and decrease the distance between peptide sequence vectors belonging to the same class.
[0175] The quartet loss function is based on the concept involved in the triplet loss function. It evaluates the loss for a quartet containing an anchor peptide sequence vector, a positive peptide sequence vector (belonging to the same class as the anchor), and two negative peptide sequence vectors (belonging to one or more different classes with respect to the anchor).
[0176] In one or more embodiments, an example quartet loss function is defined as follows: TIFF2025539935000005.tif14170, r a is the anchor representation (e.g., anchor peptide sequence vector), and r pis a positive expression (e.g., a positive peptide sequence vector that belongs to the same class as the anchor peptide sequence vector), and r n1 is the first negative representation (e.g., the first negative peptide sequence vector belonging to a different class from the anchor peptide sequence vector), and r n2 is the second negative representation (e.g., the anchor peptide sequence vector, the positive peptide sequence vector, and the second negative peptide sequence vector that belongs to a different class from the first negative sequence vector), d() is the distance function, and m is the margin. The margin may be a constant set based on the objective that the distance between the anchor and the negative must be greater than the margin m.
[0177] The lifted structure loss function also builds on the concepts behind the triplet and quartet loss functions. The lifted structure loss function uses multiple negative examples in a minibatch. The multiple negative examples include not only the negatives of the anchors but also their positives. Thus, the lifted structure loss function analyzes the distance between all possible pairs in the minibatch. For example, a minibatch may contain x1, x2, x3, x4, x5, and x6, where x1 and x2 belong to the same class, x3 and x4 belong to the same class, and x5 and x6 belong to the same class. All other pairings may be of different classes, as follows: x1 belongs to a different class than x3, x4, x5, and x6; x2 belongs to a different class than x3, x4, x5, and x6; x3 belongs to a different class than x1, x2, x5, and x6; x4 belongs to a different class than x1, x2, x5, and x6; x5 belongs to a different class than x1, x2, x3, and x4; x6 belongs to a different class than x1, x2, x3, and x4.
[0178] In one or more embodiments, an example of a lifted structure loss function is as follows: TIFF2025539935000006.tif23170where Dij = ||f(xi)-f(xj)||2; P are all positive pairs in the minibatch; N are all negative pairs in the minibatch.
[0179] Multi-class n-pair loss randomly selects negative examples in each class for grouping. For example, grouping may be performed using an anchor peptide sequence vector, a positive peptide sequence vector, and multiple negative peptide sequence vectors. The anchor peptide sequence vector may represent peptides presented by a specific MHC allele (e.g., allele A), and the multiple negative peptide sequence vectors may include randomly selected peptide sequence vectors for different MHC alleles (e.g., peptide sequence vectors corresponding to all non-A alleles).
[0180] The circular loss function attempts to provide different levels of optimization for different types of triplets. For example, a first triplet T may include an anchor A, a positive P, and a negative N. A second triplet T' may include an anchor A, a positive P', and a negative N'. If P' is much farther away than P, the circular loss function places more emphasis on reducing the distance between A and P' for the second triplet T' and more emphasis on increasing the distance between A and N for the first triplet T. Thus, different penalties may be associated with different triplets based on the distance between the anchor and the positive and the anchor and the negative. These distances may be weighted independently and may be optimized at different paces.
[0181] The angle loss function considers the cosine distance for triplets. Triplets of peptide sequence vectors may form a triangle. The angle loss function evaluates the angle between the first edge formed by the anchor and the positive pole and the second edge formed by the anchor and the negative pole. The angle loss function seeks to move the negative away from both the anchor and the positive and to move the anchor and the positive closer to each other. The angle loss function may be used with other loss functions (e.g., triplet loss, multi-class n-pairs loss, etc.) to improve overall performance.
[0182] A divergence loss function may be used for regularization when an ensemble of learners (e.g., models and / or loss functions) is used. For example, in the case of training peptide sequences sent as input through multiple models (e.g., implemented using model 124 of FIG. 1 or attention-based machine learning model 201 of FIG. 2, respectively), the divergence loss function embeds the training peptide sequences in different models focusing on different features, such that diverse embeddings are generated. In this way, the models may have diverse embedding spaces, but all may still satisfy the constraint of embedding similar peptide sequences closer to each other and dissimilar peptide sequences further apart.
[0183] The constellation loss function integrates the concepts of the triplet loss function and the multi-class n-pairs loss function. The constellation loss function simultaneously learns the distances between different class combinations. For example, similar to the multi-class n-pairs loss function, a grouping may include an anchor, positives, and negatives from each possible class. The constellation loss evaluates the distance between the anchor and the positives, the distance between the anchor and each negative, and the distance between each negative and all negatives.
[0184] The loss functions described above are just some of the different types of loss functions that may be included in a metric learning algorithm, such as metric learning algorithm 126 of Figure 1. Other types of loss functions may also be utilized. Furthermore, hybrid loss functions that include two or more of the loss functions described above and / or other loss functions may be utilized.
[0185] III.D. Mining (Sampling) Strategy III.D.1. Mining Strategies, General Different types of mining strategies can be used when determining how to evaluate the loss of a machine learning model. A mining strategy, also called a sampling strategy, is a strategy for selecting a grouping (or embedding) of peptide sequence vectors for evaluating the loss. A grouping can include at least two peptide sequence vectors. For example, a grouping can be a pair, triplet, quartet, or other multiplet of peptide sequence vectors.
[0186] In general, a miner, such as miner 135 of FIG. 1, may take the form of a subset batch miner, a tuple miner, or another type of miner. A subset batch miner may take a batch of training peptide sequence vectors (e.g., the batch of training peptide sequence vectors generated in step 504 of FIG. 5) and return a subset to be used by a tuple miner or loss function (e.g., a metric learning algorithm). A tuple miner may take a batch of training peptide sequence vectors and return a fixed number of tuples (e.g., pairs, triplets, quartets, etc.) to be used to calculate a loss.
[0187] As used herein, peptide sequence vectors can be mined based on their class, which refers to specific MHC alleles that are expected (or predicted) to present the peptide corresponding to the peptide sequence vector. MHC alleles can be identified through allelic identifiers (e.g., 4-digit, 6-digit, or 8-digit allelic identifiers). A tuple (e.g., pair, triplet, quartet, etc.) can include an anchor peptide sequence vector (or simply anchor) and at least one positive peptide sequence vector (or simply positive) or negative peptide sequence vector (or simply negative). A positive is a peptide sequence vector that belongs to the same class as the anchor. A negative is a peptide sequence vector that belongs to a different class from the anchor or positive.
[0188] Two examples of tuple miners include pair miners and triplet miners. In one or more embodiments, pair miners may be used, for example, with a contrast loss function. A pair miner may receive M peptide sequence vectors (embeddings) and output T tuples of size 4, each tuple including an anchor positive pair and an anchor negative pair. In other embodiments, a pair miner may receive M peptide sequence vectors (embeddings) and output P anchor positive pairs and P anchor negative pairs for processing through the loss function. In the absence of a pair miner, the contrast loss function may, by default, evaluate all possible pairs in the training batch.
[0189] In one or more embodiments, the triplet miner may be used with, for example, a triplet loss function, as well as other types of loss functions that evaluate triplets. The triplet miner may receive M peptide sequence vectors (embeddings) and output T triplets, each triplet including an anchor, a positive, and a negative. In the absence of a triplet miner, the contrast loss function may, by default, use all possible triplets in the training batch.
[0190] However, not all positives and negatives may be considered of equal difficulty or interest for training purposes. Some positives may be more difficult or challenging in the sense that they are less similar to the anchors (e.g., further apart in distance) than other positives. Some negatives may be more difficult or challenging in the sense that they are more similar to the anchors (e.g., closer to each other in distance) than other negatives.
[0191] In one or more embodiments, the selected positives and / or negatives for a tuple are selected based on a selected positive strategy and / or a selected negative strategy, respectively. The selected positive strategy may include using hard positives, semi-hard positives, easy positives, and / or all positives. For example, a hard positive strategy may include returning the most difficult positive per anchor, or in other words, the positive that is most dissimilar (farthest) from the anchor. An easy positive strategy may include returning the easiest positive per anchor, or in other words, the positive that is most similar (closest) to the anchor. A semi-hard positive strategy, typically used by triplet miners, may include returning a semi-hard positive per anchor. A semi-hard positive may be, for example, the most difficult positive that is even easier than the selected negative, or in other words, the positive that is most dissimilar (farthest) from the anchor but more similar (closer) than the selected negative. In some cases, the selected positive strategy used in one or more batches may differ from the positive strategy used in other batches.
[0192] The selected negative strategy may include using hard negatives, semi-hard negatives, easy negatives, and / or all negatives. For example, a hard negative strategy may include returning the most difficult negative for each anchor, in other words, returning the negative that is most similar (closest) to the anchor. An easy negative strategy may include returning the easiest negative for each anchor, in other words, returning the negative that is most dissimilar (furthest away) to the anchor. A semi-hard negative strategy, typically used by triplet miners, may include returning a semi-hard negative for each anchor. A semi-hard negative may, for example, be the most difficult negative that is even easier than the selected positive, or in other words, the negative that is most similar (closest) to the anchor but more dissimilar (furthest away) than the selected positive. In some cases, the selected negative strategy used in one or more batches may differ from the negative strategy used in other batches.
[0193] When a hard negative strategy is used, regardless of the type of positive strategy used, the overall mining strategy may be referred to as hard negative mining (or online hard negative mining). When a semi-hard negative strategy is used, regardless of the type of positive strategy used, the overall mining strategy may be referred to as semi-hard negative mining.
[0194] In other embodiments, one or more other types of miners may be used to prioritize specific pairs or triplets of peptide sequence vectors based on loss for training. Other types of miners include, but are not limited to, angle miners, base miners, base tuple miners, base subset batch miners, batch easy hard miners, batch hard miners, distance weighted miners, miners for embedding what is already packaged as triplets, hardware-aware deep cascade embedding miners, maximum loss miners, multiple similarity miners, pairwise margin miners, triplet margin miners, and uniform histogram miners. In some cases, users may create customized miners. III.D.2. Hard Negative and Semi-Hard Negative Mining for Use with Triplet Loss Functions
[0195] The triplet loss function forms triplets that include an anchor, a positive, and a negative. Hard negative mining for triplet loss functions. Semi-hard negative mining can be determined for the positives selected for the triplet.
[0196] Both hard negative mining and semi-hard negative mining are performed to exclude the evaluation of triplets that contain easy negatives. An easy negative can be defined with respect to the distance and margin m between the anchor and the selected positive. For example, an easy negative satisfies the following constraint: TIFF2025539935000007.tif7170
[0197] In hard negative mining, hard triplets are formed using hard negatives. A hard negative can be a peptide sequence vector that is closer to the anchor than the selected positive. In other words, the distance between the anchor and the negative is smaller than the distance between the anchor and the positive. Therefore, a hard triplet satisfies the following constraint: TIFF2025539935000008.tif5170
[0198] In hard negative mining, each hard triplet is formed using the hardest negative of that triplet's corresponding anchor, in other words, the negative closest to the anchor is selected.
[0199] For semi-hard negative mining, each triplet is formed using an anchor, a positive, and a semi-hard negative. The selected semi-hard negative may be farther from the anchor than the selected positive, but closer to the anchor than either easy negative. Thus, the semi-hard triplet satisfies the following constraint: TIFF2025539935000009.tif5170
[0200] In one or more embodiments, semi-hard negative mining may be used to select the most challenging semi-hard negative. For example, the negative that is closest to the anchor but meets the above constraints may be selected. In other examples, negatives may be selected randomly from the possible semi-hard negatives.
[0201] Using hard negative mining, semi-hard negative mining, or both can help prevent or reduce the likelihood of collapse during training, where training freezes or stalls. Such mining strategies enable faster convergence to machine learning models trained with improved accuracy. For example, given a batch of randomly selected peptide sequence vectors, there may be a large number of anchor-negative pairs that can be selected. Randomly selecting which anchor-negative pairs to use to evaluate loss may result in pairs that are too easy, and the machine learning model may not learn well from these pairs, resulting in poor performance.
[0202] III.E. Generation of Peptide Sequence Data III.E.1. Exemplary Processing of Samples Embodiments described herein provide machine learning models that can be used to generate peptide sequence vectors of peptide sequences that can be used to select a diverse group of candidate peptides for peptide therapeutics. These peptide sequences can be identified from one or more samples (e.g., set of samples 118 of FIG. 1 ) from one or more subjects.
[0203] In one or more embodiments, peptide sequences of interest for the development of peptide vaccines (e.g., neo-antigen vaccines) are disease-specific peptide sequences. However, some of the peptide sequences identified in disease samples may be non-disease peptide sequences corresponding to non-disease peptides. To identify whether each sequence detected as a result of sequencing a disease-specific sample is a disease-specific peptide sequence (e.g., a disease-specific nucleic acid sequence and / or a disease-specific amino-acid sequence), it can be determined whether the sequence is also identified in a reference sequence dataset. A reference sequence dataset can include a set of peptide reference sequences whose peptide sequences are known, suspected, or hypothesized not to be indicative of or characteristic of a disease (e.g., any disease or a given disease). A reference sequence dataset can include peptide sequences identified, for example, by sequencing one or more reference sample sequences collected from the same subject from which the disease-specific sample was collected, by sequencing one or more reference sample sequences collected from one or more other subjects not diagnosed with the disease or the disease corresponding to the disease-specific sample, and / or by sequencing one or more cell lines not associated with a particular disease. In some cases, the reference sequence dataset may include peptide sequences collected from one or more reference data repositories. Peptide sequences that are detected in association with disease-specific samples but are not detected in the reference sequence dataset (or are not detected at a frequency below a predefined threshold) can be classified as variant-encoding peptide sequences (e.g., generally, or for the subject from whom the disease-specific sample was collected).
[0204] In some examples, multiple variant-encoding peptide sequences may be identified (e.g., each detected in a disease sample but not represented in a reference sample sequence), and a machine learning model disclosed herein (e.g., an attention-based machine learning model such as attention-based machine learning model 201 of FIG. 2) may be used to process each representation of the multiple variant-encoding sequences (e.g., individually, sequentially, and / or in parallel).
[0205] Disease samples may include, for example, tissue (e.g., solid tumors), blood, and / or collections of cells (e.g., cancer cells that may be collected using fine needle aspiration or laparoscopy). Disease samples may include, for example, cancerous cells collected from a subject diagnosed with and / or having lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, stomach cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, and T-cell lymphocytic leukemia, non-small cell lung cancer, or small cell lung cancer.
[0206] In some cases, the initial sample is separated into a disease sample and a remaining sample (e.g., which can be discarded or used as a reference sample). The reference sample can include a matched disease-free sample. The disease sample and the reference sample can each be collected from the same subject and / or can include the same or similar sample type (e.g., tissue type). In some cases, the disease sample is collected from a first subject (e.g., a person diagnosed with a condition or disease), and the reference sample is collected from a different second subject (e.g., a person not diagnosed with a condition or disease). In some cases, the reference sample peptide sequence is searched from a database of known genes associated with the organism.
[0207] For any type of sequencing (e.g., HLA typing, where peptides bind to MHC molecules to identify sequences in a sample), the results may identify one or more nucleic acid sequences or one or more amino acid sequences. Once a nucleic acid sequence has been identified and an attention-based model (or other process) has been configured to process the amino acid sequence, techniques (e.g., lookup tables) can be used to convert individual codons within the nucleic acid sequence to individual amino acids.
[0208] III.E.2. Example of Identification of Peptide Sequence Data: Variant Peptide Sequences When developing peptide therapeutics, mutant peptides are selected to elicit a desired immunological response. The mutant peptides may be disease-specific peptides, e.g., specific to an individual subject with a disease. Thus, peptide sequence data 116 in FIG. 1 may be mutant peptide sequence data. In particular, peptide sequence 122 of peptide 120 in FIG. 1 may be a mutant peptide sequence of a mutant peptide that is detected in a disease sample collected from a subject but is not observed in one or more non-disease samples (e.g., from the subject or another subject).
[0209] Various methods can be used to identify mutant peptide sequences, thereby identifying a given subject and its corresponding mutant peptide. Mutations can be present in the genome, transcript, proteome, or exome of the subject's diseased cells, but not in non-disease samples, such as those from the subject or another subject. Mutations include, but are not limited to: (1) non-synonymous mutations that result in different amino acids in the protein; (2) read-through mutations that alter or delete stop codons, resulting in the translation of longer proteins with novel tumor-specific sequences at the C-terminus; (3) splice site mutations that result in the inclusion of introns in mature mRNA, thus resulting in unique tumor-specific protein sequences; (4) chromosomal rearrangements (i.e., gene fusions) that result in chimeric proteins with tumor-specific sequences at the junction of two proteins; and (5) frameshift insertions or frameshift deletions that result in new open reading frames with novel tumor-specific protein sequences. Mutations may also include one or more non-frameshift insertions / deletions (indels), missense or nonsense substitutions, splice site alterations, genomic rearrangements or gene fusions, or any genomic or expression alterations that give rise to a neoORF. For example, peptides with mutations or mutant polypeptides resulting from splice site, frameshift, readthrough or gene fusion mutations in diseased cells can be identified by sequencing the DNA, RNA or protein in a diseased sample and comparing the resulting sequences with sequences from a non-diseased sample.
[0210] In some embodiments, whole genome sequencing (WGS) or whole exome sequencing (WES) data from disease samples and non-disease samples can be obtained and compared. After aligning non-disease sample and disease sample reads to the human reference genome, somatic variants, including single nucleotide variants (SNVs), gene fusions, and insertion or deletion variants (indels), can be detected using a variant calling algorithm. One or more variant callers can be used to detect different somatic variants (i.e., SNVs, gene fusions, or indels).
[0211] In some examples, mutant peptides are identified based on transcriptome sequences in disease samples from individuals.For example, whole transcriptome sequences or partial transcriptome sequences (for example, by methods such as RNA-Seq) can be obtained from disease tissues of individuals and subjected to sequencing analysis.The sequences obtained from disease tissue samples can then be compared with the sequences obtained from reference samples.Optionally, the disease tissue samples are subjected to whole transcriptome RNA-Seq.Optionally, the transcriptome sequences are "enriched" for specific sequences before being compared with reference samples.For example, specific probes can be designed to enrich for specific desired sequences (for example, disease-specific sequences) before being subjected to sequencing analysis.
[0212] In some embodiments, transcriptome sequencing techniques include, but are not limited to, RNA poly(A) library sequencing, microarray analysis, parallel sequencing, massively parallel sequencing, PCR, and RNA-Seq. RNA-Seq is a high-throughput technique for sequencing part or substantially all of a transcriptome. Briefly, an isolated population of transcriptome sequences is converted into a library of cDNA fragments with adapters attached to one or both ends. Each cDNA molecule is then analyzed, with or without amplification, to obtain short stretches of sequence information, typically 30-400 base pairs. These fragments of sequence information are then aligned to a reference genome, reference transcripts, or assembled de novo to reveal the structure (i.e., transcription boundaries) and / or expression levels of the transcripts.
[0213] Once obtained, the peptide sequence in the disease sample can be compared to the corresponding peptide sequence in the reference sample. Sequence comparison can be performed at the nucleic acid level by aligning the nucleic acid sequence in the disease tissue with the corresponding sequence in the reference sample. Gene sequence variations resulting in one or more changes in the encoded amino acids are then identified. Alternatively, sequence comparison can be performed at the amino acid level, i.e., the nucleic acid sequence is first converted in silico to an amino acid sequence before comparison. Either an amino acid-based approach or a nucleic acid-based approach can be used to identify one or more mutations (e.g., one or more point mutations) in the peptide. With regard to the nucleic acid-based approach, the discovered variants can be used to identify one or more nucleic acid sequences (e.g., DNA, RNA, or mRNA sequences) that give rise to a given observable mutant protein (e.g., via a lookup table that associates individual peptide mutations with multiple codon variants).
[0214] In some embodiments, comparison of peptide sequences from disease samples to sequences of a reference sample can be completed by techniques known in the art, such as manual alignment, FAST-All (FASTA), and Basic Local Alignment Search Tool (BLAST). In some embodiments, comparison of sequences from disease samples to sequences of a reference sample can be completed using short read aligners, such as GSNAP, BWA, and STAR.
[0215] In some embodiments, the reference sample is a matched disease-free sample. As used herein, a "matched" disease-free tissue sample is one selected from the same or similar samples, e.g., samples from the same or similar tissue type as the diseased sample. In some embodiments, the matched disease-free and diseased tissues may be derived from the same individual. In some embodiments, the reference sample described herein is a disease-free sample from the same individual. In some embodiments, the reference sample is a disease-free sample from a different individual (e.g., an individual without the disease). In some embodiments, the reference sample is obtained from a population of different individuals. In some embodiments, the reference sample is a database of known genes associated with an organism. In some embodiments, the reference sample may be derived from a cell line. In some embodiments, the reference sample may be a combination of known genes associated with an organism and genomic information from a matched disease-free sample. In some embodiments, the variant coding sequence may contain point mutations in the amino acid sequence. In some embodiments, the variant coding sequence may contain amino acid deletions or insertions.
[0216] In some embodiments, a set of variant coding sequences is first identified based on genome and / or nucleic acid sequences. This initial set is then further filtered to obtain a narrower set of expressed variant coding sequences based on the presence of the variant coding sequences in a transcriptome sequencing database (and thus, are said to be "expressed"). In some embodiments, the set of variant coding sequences is reduced by at least about 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, or more by filtering the transcriptome sequencing database.
[0217] Alternatively, protein mass spectrometry can be used to identify or verify the presence of mutant peptides, such as those bound to MHC proteins on tumor cells. Peptides can be acid-eluted from diseased cells, such as tumor cells, or from HLA molecules immunoprecipitated from tumors, and then identified using mass spectrometry.
[0218] The variant peptide can have, for example, 5 or more, 8 or more, 11 or more, 15 or more, 20 or more, 40 or more, 80 or more, 100 or more, 120 or less, 100 or less, 80 or less, 60 or less, 50 or less, 40 or less, 30 or less, 25 or less, 20 or less, 18 or less, 15 or less, or 13 or less amino acids.
[0219] Tumor-specific T cell receptor sequences can also be identified, for example, by single-cell T cell receptor sequencing. High-throughput sequencing of T cell repertoires can also, or alternatively, be performed to identify tumor-specific signatures for specific diseases.
[0220] III.E.3. Examples of Identification of Training Peptide Sequence Data Training of a machine learning model, such as model 124 of Figure 1 and / or attention-based machine learning model 201 of Figure 2, may be performed using training peptide sequence data, such as training peptide sequence data 130 of Figure 1. This training peptide sequence data may be generated in a variety of ways, including those described above in Sections III.E.1 and III.E.2.
[0221] In one or more embodiments, the peptide sequence data described above in Sections III.E.1 and III.E.2 may form at least a portion of the training peptide sequence data. In other embodiments, the training peptide sequence data may be generated using data collected from multiple other samples (e.g., potentially associated with one or more other subjects). Each of the multiple other samples may include, for example, tissue (e.g., a biopsy), a single cell, multiple cells, cell fragments, or an aliquot of bodily fluid. In some cases, the multiple other samples are collected from a different type of subject compared to the subject associated with the input data processed by the trained model. For example, a machine learning model may be trained using training peptide sequence data collected by processing samples from one or more cell lines, and the trained machine learning model may be used to process input data determined by processing one or more samples from a subject.
[0222] IV. Pharmaceutically Acceptable Compositions and Manufacturing Pharmaceutically acceptable compositions can be developed and / or manufactured based on a group of diverse candidate peptides (e.g., group of diverse candidate peptides 110 in FIG. 1 ) identified using, for example, a machine learning model described herein (e.g., model 124 in FIG. 1 , attention-based machine learning model 201 in FIG. 2 ). For example, a group of diverse candidate peptides can be selected using a peptide sequence vector generated by a machine learning model. A therapeutic peptide (e.g., therapeutic peptide 144 in FIG. 1 ) can be formed from the group of diverse candidate peptides for the development of a pharmaceutically acceptable composition. The therapeutic peptide includes two or more candidate peptides from the group of diverse candidate peptides. The pharmaceutically acceptable composition can be a peptide therapeutic (e.g., peptide therapeutic 108 in FIG. 1 ). For example, the pharmaceutically acceptable composition can be a peptide vaccine (e.g., a tumor vaccine). The two or more candidate peptides selected from the group of diverse candidate peptides include at least two dissimilar peptides (e.g., having different binding motifs).
[0223] The therapeutic peptide can be, for example, a mutant peptide that can be identified by a corresponding mutant coding sequence. The composition can include the mutant peptide, a precursor of the mutant peptide, a polypeptide sequence corresponding to the mutant peptide, RNA (e.g., mRNA) corresponding to the mutant peptide, DNA corresponding to the mutant peptide, cells containing the mutant peptide and / or one or more nucleic acids encoding such peptides, a plasmid corresponding to the mutant peptide, and / or a vector corresponding to the mutant peptide.
[0224] In one or more embodiments, a pharmaceutically acceptable composition may be developed and / or manufactured using selected variant coding sequences for the mutant peptides. The composition may contain mutant peptides corresponding to a single selected variant coding sequence. The composition may contain mutant peptides and / or mutant peptide precursors corresponding to multiple selected variant coding sequences.
[0225] One, more, or all of the mutant peptides in the composition can each have a length of, for example, about 7 to about 40 amino acids (e.g., about 7, 8, 9, 10, 11, 12, 13, 14, 15, 17, 20, 22, 25, 30, 35, 40, 45, 50, 60, or 70 amino acids). In some embodiments, the length of one, more, or all of the mutant peptides in the composition is within a predetermined range (e.g., 8 to 11 amino acids, 8 to 12 amino acids, or 8 to 15 amino acids). In some embodiments, one, more, or all of the mutant peptides in the composition are each about 8 to 10 amino acids in length. One, more, or all of the mutant peptides in the composition can each be in their isolated form. Each of one or more of all of the mutant peptides in the composition can be a "long peptide" produced by adding one or more peptides to the end (or each end) of the mutant peptide. One, more or all of the mutant peptides in the composition may each be tagged, may be a fusion protein, and / or may be a hybrid molecule.
[0226] For one, several, or all of the selected variant coding sequences, a pharmaceutically acceptable composition can be developed and / or manufactured that includes or uses one or more nucleic acids encoding peptides containing or consisting of the amino acids identified in the variant coding sequence. The nucleic acid(s) can include DNA, RNA, and / or mRNA. Given that any of multiple codons can encode a given amino acid, the codons can be selected to optimize or enhance expression in a given type of organism, for example. Such selection can be based on the frequency with which each of the multiple potential codons is used by a given type of organism, the translation efficiency of each of the multiple potential codons in a given type of organism, and / or the degree of bias of a given type of organism toward each of the multiple potential codons.
[0227] In some examples, a pharmaceutically acceptable composition may include one or more nucleic acids encoding the mutant peptides or precursors of the mutant peptides. For example, the pharmaceutically acceptable composition may be a nucleic acid vaccine. The nucleic acid vaccine may be, for example, a personalized vaccine specific to a particular subject (e.g., potentially developed for a particular subject). The nucleic acid vaccine may include a nucleic acid encoding a mutant peptide or a precursor of the mutant peptide. The nucleic acid vaccine may include sequences flanking the sequence encoding the mutant peptide (or its precursor). In some examples, the nucleic acid vaccine includes an epitope corresponding to a selected variant coding sequence. The nucleic acid vaccine may be a DNA-based vaccine, an RNA-based vaccine, an mRNA-based vaccine, or a modified mRNA vaccine (e.g., including a modified mRNA protected from degradation using protamine, an mRNA containing a modified 5' cap structure, or an mRNA containing modified nucleotides). In some embodiments, the RNA-based vaccine includes a single-stranded mRNA.
[0228] Nucleic acid vaccines can include personalized neoantigen-specific therapies tailored to specific subjects for use as part of next-generation immunotherapy. Personalized vaccines can be designed by first detecting mutant peptides in a sample from a specific subject and then selecting dissimilar peptides from a group of peptides identified as most likely to be presented. For each selected mutant peptide, a synthetic mRNA sequence encoding the mutant peptide can be identified. mRNA vaccines can contain mRNA (encoding part or all of the mutant peptide) complexed with lipids to form mRNA-lipoplexes. Administration of a vaccine containing an mRNA-lipoplex can result in mRNA stimulation of TLR7 and TLR8, triggering T cell activation by dendritic cells. Furthermore, administration can result in translation of the mRNA into the mutant peptide, which can then bind to and be presented by MHC molecules, inducing a T cell response.
[0229] In one or more embodiments, the composition may include multiple polynucleotide constructs (e.g., DNA constructs or RNA constructs). Polynucleotide constructs are artificially constructed segments of nucleic acid that can be "implanted" into target tissues or cells. The polynucleotide constructs include DNA or RNA (e.g., mRNA) inserts containing nucleotide sequences encoding mutant peptides. To increase antigen presentation (e.g., presentation of one or more selected mutant peptides by MHC molecules), the polynucleotide constructs may further include modifications developed for improved antigen presentation and, therefore, improved immunogenicity of the mutant peptides. In some examples, the modification is the incorporation into the polynucleotide construct of transmembrane and cytoplasmic regions of a chain of an MHC molecule, as described in International Publication No. WO2005038030A1, which is incorporated herein by reference in its entirety for all purposes.
[0230] To provide an RNA insert with increased stability and translation efficiency, the polynucleotide construct may further include modifications developed to improve stability and translation, and therefore improve immunogenicity of the mutant peptide. In some examples, the modification is the incorporation into the polynucleotide construct of a nucleic acid sequence having at least two copies of the 3' untranslated region (UTR) of the human β-globin gene, as described in WO2007036366A2, which is incorporated herein by reference in its entirety for all purposes. In other examples, the modification is the incorporation of a nucleic acid sequence encoding a 3' UTR, such as the F1 3' UTR described in WO2017060314A3, which is incorporated herein by reference in its entirety for all purposes.
[0231] To provide an RNA insert with increased stability and expression, the polynucleotide construct may further include modifications developed to improve stability and expression, and therefore immunogenicity, of the selected mutant peptide. In some instances, the modification is the incorporation of a cap (e.g., a 5'-cap structure) at the end of the RNA. The cap structure may be the D1 diastereomer of beta-S-ARCA, as described in International Publication No. WO2011015347A1, which is incorporated herein by reference in its entirety for all purposes.
[0232] To deliver the polynucleotide construct to antigen-presenting cells with high selectivity, the composition may further comprise a cationic liposome or lipoplex to improve uptake of the polynucleotide construct and thus improve immunogenicity against the selected mutant peptide. In some examples, the composition comprises nanoparticles containing the polynucleotide construct. The nanoparticles may be lipoplexes containing one or more lipids, such as DOTMA and DOPE, as described in International Publication No. WO 2013143683 A1, the entire contents of which are incorporated herein by reference for all purposes.
[0233] The composition may comprise a substantially pure mutant peptide, a substantially pure precursor thereof, and / or a substantially pure nucleic acid encoding the mutant peptide or precursor thereof. The composition may comprise one or more suitable vectors and / or one or more delivery systems for containing the mutant peptide, its precursor, and / or the nucleic acid encoding the mutant peptide or precursor thereof. Suitable vectors and delivery systems include viruses, such as adenovirus, vaccinia virus, retrovirus, herpesvirus, adeno-associated virus, or hybrid systems containing elements of more than one virus. Non-viral delivery systems include cationic lipids and cationic polymers (e.g., cationic liposomes). In some embodiments, physical delivery, such as using a "gene gun," may be used.
[0234] The composition may include cells containing the mutant peptides and / or nucleic acid(s) encoding the mutant peptides. The composition may further include one or more suitable vectors and / or one or more delivery systems for the mutant peptides and / or nucleic acid(s) encoding the mutant peptides. In some examples, the cells containing the mutant peptides and / or nucleic acids encoding the mutant peptides are non-human cells, such as bacterial cells, protozoan cells, fungal cells, or non-human animal cells. In some examples, the cells containing the mutant peptides and / or nucleic acids encoding the mutant peptides are human cells. In some examples, the human cells are immune cells. In some examples, the immune cells are antigen-presenting cells (APCs). In some examples, the APCs are professional APCs such as macrophages, monocytes, dendritic cells, B cells, and microglia. In other examples, the professional APCs are macrophages or dendritic cells. In some examples, the APCs containing the nucleic acid sequence(s) encoding the mutant peptides and / or mutant peptides are used as cellular vaccines, thereby inducing a CD4+ or CD8+ immune response. In other examples, compositions used as cellular vaccines include mutant peptide-specific T cells primed by APCs containing the mutant peptide and / or nucleic acid sequence(s) encoding the mutant peptide.
[0235] The composition may include a pharmaceutically acceptable adjuvant and / or a pharmaceutically acceptable excipient. An adjuvant refers to any substance whose incorporation into the composition modifies the immune response to the mutant peptide. The adjuvant may be conjugated, for example, with an immunostimulatory agent. The excipient may increase the molecular weight of a particular mutant peptide to increase activity or immunogenicity, confer stability, increase biological activity, and / or increase serum half-life. In one or more embodiments, the composition may include an adjuvant, excipient, immunomodulator, checkpoint protein, PD-1 antagonist (e.g., anti-PD-1 antibody) and / or PD-L1 antagonist (e.g., anti-PD-L1 antibody).
[0236] In one or more embodiments, a pharmaceutically acceptable composition developed and / or manufactured based on a diverse group of candidate peptides (e.g., diverse group of candidate peptides 110 in FIG. 1 ) identified using a machine learning model described herein (e.g., model 124 in FIG. 1 , attention-based machine learning model 201 in FIG. 2 ) can be a T cell therapy. For example, the pharmaceutically acceptable composition can include one or more engineered T cells. For example, the pharmaceutically acceptable composition can include a population of engineered T cells. V.IV. Treatment Methods Involving Immunogenic Vaccines or T Cells
[0237] The embodiments described herein provide methods for use in the development of therapeutics for medical conditions (e.g., diseases such as, but not limited to, cancer) and / or in the treatment of individuals with medical conditions.
[0238] 6 is a flowchart of a process 600 for treating a subject, according to one or more embodiments. Process 600 may be used to treat a subject via a therapeutic agent developed and / or manufactured based on a diverse set of candidate peptides identified via a machine learning model (e.g., model 124 of FIG. 1 , attention-based machine learning model 201 of FIG. 2 ).
[0239] Step 602 includes processing one or more samples to detect peptide sequences. The subject may have a medical condition, such as, but not limited to, cancer. The one or more samples may be one or more disease samples.
[0240] Step 604 includes generating a peptide sequence vector for the peptide sequence using a machine learning model trained using a metric learning algorithm. The machine learning model may be, for example, deep learning model 124 of Figure 1 or attention-based machine learning model 201 of Figure 2. The metric learning algorithm may be, for example, metric learning algorithm 126 of Figure 1. The machine learning model may be trained using, for example, process 400 of Figure 4 or process 500 of Figure 5.
[0241] Step 606 involves selecting a group of diverse candidate peptides for use in developing a therapeutic to treat the subject. The therapeutic can be, for example, a peptide therapeutic (e.g., peptide therapeutic 108 of FIG. 1 ) that includes or is based on a therapeutic peptide (e.g., therapeutic peptide 144 of FIG. 1 ) selected from the group of diverse candidate peptides (group of diverse candidate peptides 110 of FIG. 1 ). The peptide therapeutic can include the selected therapeutic peptide, a precursor of the selected therapeutic peptide, or one or more nucleic acids encoding the therapeutic peptide or precursor thereof. The individual can be treated by administering an effective amount of a composition (such as the composition described above in Section IV) that includes the selected therapeutic peptide, a precursor of the selected therapeutic peptide, or one or more nucleic acids encoding the therapeutic peptide or precursor thereof. The therapeutic can be, for example, a peptide vaccine (e.g., a tumor vaccine) for treating a disease (e.g., cancer).
[0242] Process 600 may optionally include step 608 and may optionally include step 610. Step 608 includes developing a therapeutic agent based on the group of diverse candidate peptides. Step 610 includes administering the therapeutic agent to a subject.
[0243] The individual being treated may be the same individual from whom one or more samples (e.g., a disease sample) were collected in step 602. In some examples, the therapeutic agent is administered to a different individual compared to the individual from whom the disease sample was collected. The different individual may, for example, be related to the individual from whom the disease sample was collected, may have a genetic risk for developing a particular type of cancer, and / or may have an MHC molecule with one, more, or all alleles corresponding to the same (or similar) sequences as one or more MHC alleles of the subject from whom the disease sample was collected.
[0244] Therapeutic drugs are used to treat various forms of cancer: carcinoma, lymphoma, blastoma, sarcoma, leukemia, squamous cell carcinoma, lung cancer (including small cell lung cancer, non-small cell lung cancer, adenocarcinoma of the lung, and squamous cell carcinoma of the lung), cancer of the peritoneum, hepatocellular carcinoma, gastric cancer or stomach cancer (including gastrointestinal cancer), pancreatic cancer, glioblastoma, cervical cancer, ovarian cancer, bladder cancer, hepatocellular carcinoma, breast cancer, colon cancer, melanoma, endometrial or uterine carcinoma, salivary gland carcinoma, kidney cancer or renal cancer. cancer), liver cancer, prostate cancer, vulvar cancer, thyroid cancer, hepatic carcinoma, head and neck cancer, colorectal cancer, rectal cancer, soft tissue sarcoma, Kaposi's sarcoma, B-cell lymphoma (low-grade / follicular non-Hodgkin's lymphoma (NHL), small lymphocytic (SL) NHL, intermediate-grade / follicular NHL, intermediate-grade diffuse NHL, high-grade immunoblastic NHL, high-grade lymphoblastic NHL, high-grade small noncleaved cell NHL, bulky disease NHL) , mantle cell lymphoma, AIDS-related lymphoma, and Waldenstrom's hypertumor globulinemia), chronic lymphocytic leukemia (CLL), acute lymphoblastic leukemia (ALL), melanoma, hairy cell leukemia, chronic myeloblastic leukemia, and post-transplant lymphoproliferative disorder (PTLD), as well as abnormal blood vessel growth associated with phacomatosis, edema (such as that associated with brain tumors), and Meigs' syndrome.
[0245] The therapeutic developed in step 608 can be, for example, a vaccine (e.g., a tumor vaccine). The vaccine can include multiple peptides, multiple precursors of multiple peptides, or a set of nucleic acids encoding multiple peptides or multiple precursors. The multiple peptides can be selected from the group of diverse candidate peptides identified in step 606. The multiple peptides include at least two peptides with different binding motifs.
[0246] The vaccine may comprise DNA comprising a set of nucleic acids (i.e., one or more nucleic acids), RNA comprising a set of nucleic acids, or mRNA comprising a set of nucleic acids. The set of nucleic acids may be identified based on the amino acids in the plurality of peptides. The set of nucleic acids may encode a plurality of peptides. In some embodiments, for each peptide of the plurality of peptides, the tumor vaccine comprises at least one of a nucleotide sequence encoding each peptide, an amino acid sequence corresponding to each peptide, RNA corresponding to each peptide, DNA corresponding to each peptide, a cell corresponding to each peptide, a plasmid corresponding to each peptide, or a vector corresponding to each peptide. The vaccine may comprise at least one excipient or adjuvant. The vaccine may comprise an RNA molecule. In the 5' to 3' direction, the RNA molecule can include a 5' cap:5' untranslated region (UTR); a polynucleotide sequence encoding a secretory signal peptide; a polynucleotide sequence encoding multiple peptides; a polynucleotide sequence encoding at least a portion of the transmembrane and cytoplasmic domains of a major histocompatibility complex (MHC) molecule; a 3' UTR comprising a 3' untranslated region of an Amino-Terminal Enhancer of Split (AES) mRNA or a fragment thereof and a non-coding RNA of a mitochondrially encoded 12S RNA or a fragment thereof; and a poly(A) sequence.
[0247] In one or more embodiments, the therapeutic developed in step 608 is a T cell therapy. The T cell therapy may include a single engineered T cell or multiple engineered T cells (e.g., a population thereof). In these embodiments, developing a T cell therapy in step 608 may include, but is not limited to, providing a population of T cells. The cells may be autologous or allogeneic to the subject. Step 608 may further include engineering at least a subset of the population of T cells to express an exogenous T cell receptor (TCR) and knock out endogenous TCR-β, thereby forming a population of engineered T cells. The exogenous TCR may bind to an antigen expressed by the cancer and selected based on a diverse group of peptides. The antigen may be, for example, a neoantigen or a TAA (tumor-associated antigen). The presence of the antigen may have been determined, for example, by sequencing at least a portion of the cancer genome and / or transcriptome. The TCR may bind to an antigen presented on an MHC class I (MHCI) molecule containing an MHCI allele expressed by the subject.
[0248] Step 608 may further include expanding the population of engineered T cells. The expanded population of engineered T cells may be 1×10 5 ~1×10 11 In these embodiments, step 610 may then include administering the expanded population of engineered T cells to the subject.
[0249]
[0013] Embodiments disclosed herein may include identifying and / or implementing some or all of a personalized medicine strategy. For example, multiple dissimilar mutant peptides (e.g., having diverse binding motifs) may be selected based on processing of peptide sequences detected in a sample from an individual. Processing may be performed using a machine learning model, such as model 124 of FIG. 1 or attention-based machine learning model 201 of FIG. 2. The mutant peptides (and / or their precursors) may then be administered to the same individual.
[0250] In some embodiments, methods of treating a disease such as cancer may include selecting a therapeutic peptide, identifying a precursor of the therapeutic peptide and / or one or more nucleic acid sequences encoding the therapeutic peptide or its precursor, and / or synthesizing the therapeutic peptide, a precursor of the therapeutic peptide, or one or more nucleic acids encoding the therapeutic peptide or peptide precursor. The synthesized products may then be administered. VI. [Example]
[0251] 7 is an illustration of a plot of peptide sequence vectors in a reduced dimensional space, according to one or more embodiments. Plot 700 may be an example of an implementation of the visual representation included in output 140 of FIG. 1 . Plot 700 shows the similarity relationships between various peptide sequence vectors and the corresponding peptide sequences represented by these peptide sequence vectors. In one or more embodiments, the peptide sequence vectors are embedded in an n-dimensional space and visualized in plot 700 in a k-dimensional space. K-dimensional space is a two-dimensional space having fewer dimensions than n-dimensional space. This reduced dimensional space makes it easier to interpret the similarity relationships between various peptide sequences.
[0252] Peptide sequence vectors that are closer to each other in plot 700 may be more similar and / or may be presented by the same or similar MHC alleles. More specifically, peptide sequence vectors that are closer share the same or similar binding motifs. Peptide sequence vectors that are further apart in plot 700 may be more dissimilar and / or may be presented by different MHC alleles. More specifically, peptide sequence vectors that are further apart have different binding motifs. Binding motifs may be dissimilar in that the amino acids included in the binding motifs are different, the sequence of amino acids is different, the spacing (or interval) between amino acids is different, or a combination thereof. For example, binding motifs may be dissimilar by differing by more than a selected number of amino acids (e.g., 2, 3, 4, 5, or more amino acids).
[0253] 7, peptide sequence vector 702 and peptide sequence vector 704 have the same or similar binding motifs. Peptide sequence vector 702 and peptide sequence vector 706 may have different binding motifs. Peptide sequence vector 708 may have a binding motif that is more similar to the binding motif of peptide sequence vector 706 than to the binding motif of peptide sequence vector 702.
[0254] In some cases, the peptide sequence for a given peptide may contain multiple binding motifs for binding to a group of MHC alleles. The peptide sequence vector generated for that peptide sequence captures this information. For example, this peptide sequence vector may be located in n-dimensional space near the peptide sequence vectors of other peptides that also bind to one or more of the MHC allele groups. Thus, these peptide sequence vectors may appear closer to each other, or even overlap, in the k-dimensional space of the plot 700.
[0255] 8 is a list of MHC alleles according to one or more embodiments. List 800 in FIG. 8 includes various MHC alleles that present peptide sequences represented by peptide sequence vectors included in plot 700. VII. Computer-Implemented Systems
[0256] 9 is a block diagram of a computer system according to various embodiments. The computer system 900 may be an example of one implementation of the computing platform 102 described above in FIG.
[0257] In one or more examples, computer system 900 may include a bus 902 or other communication mechanism for communicating information, and a processor 904 coupled with bus 902 for processing information. In various embodiments, computer system 900 may also include memory, which may be a random access memory (RAM) 906 or other dynamic storage device, coupled to bus 902 for determining instructions to be executed by processor 904. The memory may also be used for storing temporary variables or other intermediate information during execution of instructions to be executed by processor 904. In various embodiments, computer system 900 may further include a read-only memory (ROM) 908 or other static storage device coupled to bus 902 for storing static information and instructions for processor 904. A storage device 910, such as a magnetic disk or optical disk, may be provided and coupled to bus 902 for storing information and instructions.
[0258] In various embodiments, computer system 900 may be coupled via bus 902 to a display 912, such as a cathode ray tube (CRT) or liquid crystal display (LCD), for displaying information to a computer user. An input device 914, including alphanumeric and other keys, may be coupled to bus 902 for communicating information and command selections to processor 904. Another type of user input device is a cursor control device 916, such as a mouse, joystick, trackball, gesture input device, eye-gaze-based input device, or cursor direction keys, for communicating directional information and command selections to processor 904 and for controlling cursor movement on display 912. This input device 914 typically has two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), that allow the device to specify a position in a plane. However, it should be understood that input devices 914 that allow three-dimensional (e.g., x, y, and z) cursor movement are also contemplated herein.
[0259] Consistent with certain implementations of the present teachings, results may be provided by computer system 900 in response to processor 904 executing one or more sequences of one or more instructions stored in RAM 906. Such instructions may be read into RAM 906 from another computer-readable medium or computer-readable storage medium, such as storage device 910. Execution of the sequences of instructions stored in RAM 906 may cause processor 904 to perform the processes described herein. Alternatively, hard-wired circuitry may be used in place of or in combination with software instructions to implement the present teachings. Thus, implementations of the present teachings are not limited to any specific combination of hardware circuitry and software.
[0260] As used herein, the terms “computer-readable medium” (e.g., data store, data storage, storage device, data storage device, etc.) or “computer-readable storage medium” refer to any medium that participates in providing instructions to processor 904 for execution. Such a medium may take many forms, including, but not limited to, non-volatile media, volatile media, and transmission media. Examples of non-volatile media include, but are not limited to, optical, solid-state, and magnetic disks, such as storage device(s) 910. Examples of volatile media include, but are not limited to, dynamic memory, such as RAM 906. Examples of transmission media include, but are not limited to, coaxial cables, copper wire, and fiber optics, including the wires that comprise bus 902.
[0261] Common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape or any other magnetic medium, CD-ROMs, any other optical medium, punch cards, paper tape, any other physical medium with a pattern of holes, RAM, PROMs, and EPROMs, flash EPROMs, any other memory chip or cartridge, or any other tangible medium from which a computer can read.
[0262] In addition to computer-readable media, instructions or data may be provided as signals on a transmission medium included in a communication device or system to provide sequences of one or more instructions to the processor 904 of the computer system 900 for execution. For example, a communication device may include a transceiver having signals indicative of instructions and data. The instructions and data are configured to cause one or more processors to implement the functions outlined in the disclosure herein. Representative examples of data communication transmission connections may include, but are not limited to, a telephone modem connection, a wide area network (WAN), a local area network (LAN), an infrared data connection, an NFC connection, an optical communication connection, etc.
[0263] It should be understood that the methodologies, flowcharts, diagrams, and accompanying disclosure described herein can be implemented using computer system 900 as a standalone device or over a distributed network of shared computer processing resources, such as a cloud computing network.
[0264] The methodologies described herein may be implemented by various means depending on the application. For example, the methodologies may be implemented in hardware, firmware, software, or any combination thereof. In the case of a hardware implementation, the processing unit may be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described herein, or combinations thereof.
[0265] In various embodiments, the methods of the present teachings may be implemented as firmware and / or software programs and applications written in conventional programming languages such as C, C++, Python, etc. When implemented as firmware and / or software, the embodiments described herein may be implemented in a non-transitory computer-readable medium having stored thereon a program for causing a computer to perform the above-described methods. It should be understood that the various engines described herein may be provided on a computer system such as computer system 900, whereby processor 904 performs the analyses and decisions provided by these engines according to instructions provided by any one or combination of memory components RAM 906, ROM 908, or storage device 910, and user input provided via input device 914.
[0266] VIII. Examples of Terminology Explanations Unless otherwise defined, scientific and technical terms used in connection with the present teachings described herein shall have the meanings commonly understood by those of ordinary skill in the art. Furthermore, unless the context requires otherwise, singular terms shall include the plural and plural terms shall include the singular. Generally, the nomenclature utilized in connection with, and techniques of, chemistry, biochemistry, molecular biology, pharmacology, and toxicology described herein are those well known and commonly used in the art.
[0267] As used herein, "substantially" means sufficient to function for its intended purpose. Thus, the term "substantially" allows for slight, insignificant variations from an absolute or perfect state, dimension, measurement, result, etc., as would be expected by one of ordinary skill in the art, but does not noticeably affect overall performance. With respect to a parameter or characteristic that is a number or can be expressed as a number, "substantially" means within 10%.
[0268] The term "ones" means two or more.
[0269] As used herein, the term "plurality" can be 2, 3, 4, 5, 6, 7, 8, 9, 10 or more.
[0270] As used herein, the term "set" means one or more. For example, a set of items includes one or more items.
[0271] As used herein, the phrase "at least one of," when used in conjunction with a list of items, means that different combinations of one or more of the listed items may be used, or that only one of the items in the list may be required. An item may be a specific object, thing, step, action, process, or category. In other words, "at least one of" means that any combination or number of items from the list may be used, but not all of the items in the list may be required. For example, without limitation, "at least one of item A, item B, or item C" means item A; item A and item B; item B; item A, item B, and item C; item B and item C; or items A and C. In some cases, "at least one of item A, item B, or item C" means, but is not limited to, two of item A, one of item B, and ten of item C; four of item B and seven of item C; or some other suitable combination.
[0272] When a reference is made to a list of elements (e.g., elements a, b, c), such reference is intended to include any one of the listed elements alone, any combination of fewer than all of the listed elements, and / or all combinations of the listed elements.
[0273] As used herein, a "model" includes at least one of an algorithm, a formula, a mathematical technique, a machine algorithm, a probability distribution or model, or another type of mathematical or statistical representation.
[0274] As used herein, a "subject" may refer to or include one or more cells, tissues, or organisms. A subject may be in vivo, ex vivo, or in vitro, male or female, human or non-human. A subject may be a mammal, such as a human. A subject may refer to a mammal being evaluated and / or treated for therapy, a mammal participating in a clinical trial, a mammal receiving anti-cancer therapy, or any other mammal of interest. In various embodiments, the terms "subject," "individual," and "patient" are used interchangeably herein. A subject may be a healthy or asymptomatic individual, an individual having or suspected of having a disease (e.g., cancer) or predisposition to a disease, an individual in need of therapy, or a combination thereof. A subject may be, for example, but not limited to, an individual with cancer or an individual with an autoimmune disease. A subject may be a human. In other cases, a subject may be any other type of mammal. For example, a subject may be a mammal used in creating a laboratory model of a human disease. Such mammals include, but are not limited to, mice, rats, primates (e.g., cynomolgus monkeys), and the like.
[0275] As used herein, a "sample" can refer to a "biological sample" of a subject. As used herein, a "sample" can include tissue (e.g., a biopsy), a single cell, multiple cells, cell fragments, or an aliquot of bodily fluid. A sample can be obtained from a subject by means such as, but not limited to, venipuncture, excretion, ejaculation, massage, biopsy, needle aspirate, lavage sample, scraping, surgical incision, intervention, another type of sampling means, or a combination thereof.
[0276] As used herein, "nucleotide" includes a nucleoside and a phosphate group. As used herein, "nucleoside" includes a nucleic acid base and a five-carbon sugar (e.g., ribose, deoxyribose, or an analog thereof). When a nucleic acid base is bound to ribose, the nucleoside can be referred to as a ribonucleoside. When a nucleic acid base is bound to deoxyribose, the nucleoside can be referred to as a deoxyribonucleoside. The "nucleobase," also referred to as a "nitrogenous base," can take the form of one of five types: adenine (A), guanine (G), thymine (T), uracil (U), and cytosine (C).
[0277] As used herein, "polynucleotide," "nucleic acid," or "oligonucleotide" refers to a linear polymer of nucleotides (or nucleosides linked by internucleoside linkages). Generally, a polynucleotide contains at least three nucleosides. Generally, an oligonucleotide is composed of nucleotides ranging in number from a few nucleotides (or monomer units) to several hundred nucleotides (monomer units). Whenever a polynucleotide, such as an oligonucleotide, is represented by a sequence of letters such as "ATGCCTG," it is understood that, unless otherwise noted, the nucleotides are in a 5'→3' order or orientation from left to right, "A" means adenosine, "C" means cytosine, "G" means guanosine, and "T" means thymidine. The letters A, C, G, and T may be used, as standard in the art, to refer to the nucleobases themselves, nucleosides comprising those nucleobases, or nucleotides comprising those bases, as described above.
[0278] Deoxyribonucleic acid (DNA) is a chain of nucleotides made up of four types of nucleotides: adenine (A), thymine (T), cytosine (C), and guanine (G). Ribonucleic acid (RNA) is composed of four types of nucleotides: A, C, G, and uracil (U). Certain pairs of nucleotides specifically bind to each other in a complementary manner, which is called complementary base pairing. For example, C pairs with G and A pairs with T. However, in RNA, A pairs with U. When a first nucleic acid strand binds to a second nucleic acid strand made up of nucleotides complementary to those in the first strand, the two strands combine to form a duplex. As used herein, "nucleic acid sequencing data," "nucleic acid sequencing information," "nucleic acid sequence," "genomic sequence," "gene sequence," "fragment sequence," or "nucleic acid sequencing read" refers to any information or data that indicates the order of nucleotide bases (e.g., A, C, G, T / U) in a DNA or RNA molecule (e.g., a whole genome, a whole transcriptome, an exome, an oligonucleotide, a polynucleotide, a fragment, etc.). It is understood that the present disclosure contemplates that this sequence information may be obtained using any of a variety of available techniques, platforms, or technologies, including, but not limited to, capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion- or pH-based detection systems, electronic-based systems, etc., or combinations thereof.
[0279] The term "genome," as used herein, refers to the genetic material of a cell or organism, including animals such as mammals (e.g., humans), and includes nucleic acids such as DNA. A genome is stored in one or more chromosomes, which are composed of DNA sequences. In humans, DNA includes genes, non-coding DNA, mitochondrial DNA, and the like. The human genome typically contains 23 pairs of chromosomes: 22 pairs of autosomal chromosomes (autosomes) plus sex-determining X and Y chromosomes. The 23 pairs of chromosomes include one copy from each parent. The DNA that makes up the chromosomes is called chromosomal DNA and is present in the nucleus of human cells (nuclear DNA).
[0280] As used herein, a "gene" is a distinct portion of an inherited genomic sequence that influences a subject's traits by being expressed as a functional product or by regulating gene expression. The total complement of genes in a subject or cell is known as the subject's or cell's genome. The region of a chromosome where a particular gene is located is called its locus. Each locus contains one allele of the gene. Thus, a pair of chromosomes share two loci, each containing an allele of the gene to form an allele pair. The two alleles may be the same or different (e.g., have slightly varying gene sequences).
[0281] As used herein, " allele " is a variant of gene.One allele of gene can be different from another allele of the same gene in various ways.For example, two alleles of the same gene can be different by, for example, protein (for example, difference in the amino acid sequence of encoded protein), other (silent or synonymous) differences in exon region that do not affect amino acid sequence, differences in intron region, or any combination of these differences.
[0282] As used herein, a "peptide sequence" can refer to an ordered sequence that identifies at least some amino acids of a peptide, for example, via amino acid identifiers, codon identifiers, or nucleotide identifiers. In some cases, a peptide sequence includes a variant coding sequence that includes variants not observed in the corresponding reference sequence.
[0283] If the peptide contains a mutant peptide, the variant coding sequence will identify the mutant or variant amino acid. However, if the peptide does not contain a mutation or variant, the variant coding sequence will not identify the mutant or variant amino acid (in which case it will be the same as the reference sequence). A variant coding sequence can be determined by collecting a disease and / or tumor sample (e.g., containing tumor cells) and performing a sequencing analysis to identify one or more sequences corresponding to the disease and / or tumor cells in the sample. In some cases, the sequencing analysis outputs an amino acid sequence. In some instances, the sequencing analysis outputs a nucleic acid sequence, which can then be processed to convert codons into amino acid identifiers and thus generate the amino acid sequence. A variant coding sequence can include a sequence of a neoantigen. A variant coding sequence may, but need not, include one or more termini (e.g., the C-terminus and / or N-terminus) of the peptide. A variant coding sequence can include an epitope of the peptide. A variant coding sequence can identify amino acids within a peptide that have one or more variants (e.g., one or more amino acid differences) compared to the corresponding reference sequence. In some examples, the variant coding sequence comprises an ordered set of amino acids. In some examples, the variant coding sequence identifies a reference peptide (e.g., by identifying a genetic reference sequence by gene, start position and / or end position, etc.; or by gene, start position and / or length) and one or more point mutations relative to the reference peptide.
[0284] As used herein, a "reference sequence" may refer to a sequence that identifies amino acids within at least a portion of a non-mutant or wild-type peptide (e.g., a wild-type parent sequence). A non-mutant or wild-type peptide may contain no variants or fewer variants than contained in a mutant peptide. A reference sequence may include an amino acid sequence encoded by a gene sequence within the same gene as a gene containing a corresponding variant coding sequence. A reference sequence may include an amino acid sequence encoded by a gene sequence spanning the same start and stop within a gene relative to the intragenic position associated with the gene sequence associated with the corresponding variant coding sequence. A reference sequence may be identified by collecting non-disease and / or non-tumor samples from one or more subjects (which may, but need not, include subjects from whom disease samples were collected to determine variant coding sequences) and performing sequencing analysis using the samples.
[0285] As used herein, "MHC" refers to the major histocompatibility complex, a system, complex, or group of cell surface proteins involved in regulating the immune system. Human MHC is also called the human leukocyte antigen (HLA) complex. The HLA system or complex is encoded by the MHC gene complex in humans. MHC molecules that present antigens to cells are classified as belonging to one of three classes of MHC molecules: MHC class I, MHC class II, and MHC class III. For example, certain HLA genes, including HLA-A, HLA-B, and HLA-C, correspond to MHC class I. For example, certain HLA genes, including HLA-DP, HLA-DM, HLA-DO, HLA-DQ, and HLA-DR, correspond to MHC class II. Known HLA genes include, for example, HLA-A, HLA-B, HLA-C, HLA-E, HLA-F, HLA-G, HLA-H, HLA-J, HLA-K, HLA-L, HLA-N, HLA-P, HLA-S, HLA-T, HLA-U, HLA-V, HLA-W, HLA-X, HLA-Y, HLA-Z, HLA-DRA, HLA-DRB, HLA-DQ, HLA-DOA, HLA-DOB, HLA-DMA, HLA-DMB, HLA-DPA, HLA-DPB, and HFE. Other genes found in the HLA region include, for example, TAP1, TAP2, PSMB9, PSMB8, MICA, MICB, MICC, MICD, and MICE.
[0286] As used herein, "immunotherapy" refers to a treatment or class of treatments that use one or more parts of a subject's immune system to fight diseases such as, for example, but not limited to, cancer. Immunotherapy can use substances made by the body or synthesized outside the body to improve how the immune system works to find and destroy cancer cells.
[0287] As used herein, a "neoantigen" refers to a tumor-specific antigen derived from a somatic mutation in a tumor and presented by a subject's cancer cells and antigen-presenting cells. Neoantigen therapy, including but not limited to neoantigen vaccines, is a relatively new approach to providing personalized cancer treatment. Neoantigen vaccines can prime a subject's T cells to recognize and attack cancer cells expressing one or more specific tumor neoantigens. This approach results in a tumor-specific immune response that targets tumor cells while sparing healthy cells. Personalized vaccines can be engineered or selected based on a subject's specific tumor profile. The tumor profile can be defined by determining DNA and / or RNA sequences from the subject's tumor cells and using the sequences to identify neoantigens present in tumor cells but absent in normal cells.
[0288] As used herein, the terms "peptide," "polypeptide," and "protein" may be used interchangeably and refer to a polymer of amino acid residues. The term encompasses amino acid chains of any length, including full-length proteins having amino acid residues linked by covalent peptide bonds.
[0289] As used herein, a "mutant peptide" may refer to a peptide that is not present in an individual subject's normal tissue (e.g., the wild-type amino acid sequence of the normal tissue). A mutant peptide contains at least one mutant amino acid and may be present in diseased tissue (e.g., collected from a particular subject) but not in normal tissue (e.g., collected from a particular subject, collected from a different subject, and / or identified in a database as corresponding to normal tissue). A mutant peptide may contain an epitope. An epitope is the portion of a mutant peptide that is bound by an MHC molecule or a T cell receptor (TCR). Thus, this binding between the epitope of the mutant peptide and the MHC molecule or TCR can induce an immune response (as a result of the mutant peptide not being associated with the subject's "self"). A mutant peptide may contain or be a neoantigen. Mutant peptides can result, for example, from nonsynonymous mutations (e.g., point mutations) that result in different amino acids in the protein; read-through mutations in which a stop codon is altered or deleted, resulting in translation of a longer protein with a new tumor-specific sequence at the C-terminus; splice site mutations that result in a unique tumor-specific protein sequence; and chromosomal rearrangements that result in frameshift insertions or deletions that result in chimeric proteins with tumor-specific sequences at the junction of two proteins (i.e., gene fusions) and / or new open reading frames with tumor-specific protein sequences. Mutant peptides can comprise polypeptides (characterized by a polypeptide sequence) and / or can be encoded by a nucleotide sequence.
[0290] As used herein, the "epitope" of a peptide refers to the region of the peptide between the C flank and the N flank that can be recognized by a TCR. The epitope of a peptide is the portion of the peptide that is recognized by a TCR on a T cell and an MHC I on an antigen-presenting cell. For example, the epitope can be a peptide that is bound by a TCR, for example, a peptide that is bound by a TCR when the peptide is bound to an MHC I on an antigen-presenting cell.
[0291] As used herein, a "representation" of a sequence can include a set of values representing or identifying amino acids in the sequence and / or a set of values representing or identifying nucleic acids encoding the sequence. For example, each amino acid can be represented by a binary string and / or vector of values that differ from each other binary string and / or vector representing each other amino acid. This representation can be generated, for example, using one-hot encoding or a BLOcks SUbstitution Matrix (BLOSUM) matrix. For example, a multidimensional (e.g., 20- or 21-dimensional) array is initialized (e.g., randomly or pseudorandomly). The initialized array can include, for each amino acid, a unique vector corresponding to that amino acid. The values can be fixed so that the use of such a unique vector can be assumed to represent the corresponding amino acid. Given that any of multiple codons can encode a single amino acid, there can be multiple possible nucleic acid representations of a given sequence.
[0292] As used herein, "presentation" of a peptide refers to at least a portion of the peptide being presented on the surface of a cell by binding to an MHC molecule in a specific manner. The presented peptide may then be accessible to other cells, such as nearby T cells.
[0293] As used herein, "immunogenic" can refer to the ability to elicit an immune response (e.g., via T cells and / or B cells). A peptide that is "immunogenic" can be a peptide that is capable of eliciting an immune response. IX. Enumeration of Exemplary Embodiments
[0294] Embodiment 1. A method for developing a therapeutic agent, the method comprising: receiving peptide sequence data identifying a plurality of peptide sequences corresponding to a plurality of peptides; generating, via a trained machine learning model, a plurality of peptide sequence vectors in n-dimensional space for each of the plurality of peptide sequences, and thereby for each of the plurality of peptides; the machine learning model being trained using a metric learning algorithm, training peptide sequence data, and training allele presentation data; the training peptide sequence data identifying a training peptide sequence corresponding to each training peptide of the plurality of training peptides; and the training allele presentation data identifying, for each training peptide sequence of the plurality of training peptides in the training peptide sequence data, one or more major histocompatibility complex (MHC) alleles that are expected to present the training peptide corresponding to each training peptide sequence. wherein a metric learning algorithm has been used to train a machine learning model to output a plurality of peptide sequence vectors in n-dimensional space such that a first distance between a first pair of peptide sequence vectors in the plurality of peptide sequence vectors generated in the n-dimensional space for a first respective pair of peptides presented by the same MHC allele is less than a second distance between a second pair of peptide sequence vectors in the plurality of peptide sequence vectors generated in the n-dimensional space for a second pair of peptides presented by different MHC alleles. A method comprising: generating a plurality of peptide sequence vectors; and generating an output using the plurality of peptide sequence vectors, wherein the output provides a measure of similarity between the peptide sequences for use in selecting a group of candidate peptides from the plurality of peptides for development of a therapeutic agent.
[0295] Embodiment 2. The method of embodiment 1, further comprising selecting a group of candidate peptides from the plurality of peptides for development of therapeutic agents based on the output, such that at least two candidate peptides in the group of candidate peptides have different binding motifs.
[0296] Embodiment 3. The method of embodiment 1 or 2, further comprising selecting a group of candidate peptides from the plurality of peptides for development of therapeutic agents based on the output, such that bias towards any single binding motif is reduced.
[0297] Embodiment 4. The method of embodiment 2 or 3, wherein selecting a group of candidate peptides from the plurality of peptides comprises using the output to identify a plurality of clusters of the plurality of peptide sequence vectors, and selecting at least one peptide sequence vector from each of the clusters for use in developing a therapeutic agent.
[0298] Embodiment 5. The method of embodiment 4, wherein selecting at least one peptide sequence vector for a selected cluster of the plurality of clusters comprises selecting one peptide sequence vector from the selected cluster that is closest to a center of the selected cluster, wherein the center is selected from one of a centroid, a mean center, a median center, and a density-based center.
[0299] Embodiment 6. The method of embodiment 4 or 5, wherein selecting at least one peptide sequence vector for a selected cluster of clusters comprises selecting at least two peptide sequence vectors from the selected cluster, wherein each of the at least two peptide sequence vectors is midway between the center of the cluster and the edge of the cluster, or each of the at least two peptide sequence vectors is located along the edge of the cluster.
[0300] Embodiment 7 The method of any one of embodiments 1 to 6, wherein the therapeutic agent comprises at least two candidate peptides of a group of candidate peptides, and wherein the at least two candidate peptides have different binding motifs.
[0301] Embodiment 8. The method of any one of embodiments 1 to 7, wherein training the machine learning model comprises training the machine learning model using a metric learning algorithm, training peptide sequence data, and training allele representation data.
[0302] Embodiment 9. The method of embodiment 8, wherein training the machine learning model includes calculating a distance metric for pairs of training peptide sequence vectors in the batch of training peptide sequence vectors, forming an evaluation bundle from the batch of training peptide sequence vectors based on the distance metric and the mining strategy, evaluating a loss of the evaluation bundle using a metric learning algorithm, and updating parameters of the machine learning model based on the loss.
[0303] Embodiment 10. The method of embodiment 9, wherein training the machine learning model comprises repeating the calculating, forming, and evaluating steps for multiple batches of training peptide sequence vectors formed from training peptide sequence data.
[0304] Embodiment 11. The method of embodiment 10, wherein the updating step is performed after the batch is processed.
[0305] Embodiment 12. The method of embodiment 10, wherein the updating step is performed for each of the batches.
[0306] Embodiment 13. The method of any one of embodiments 9 to 12, wherein the mining strategy used for a first portion of the batch is semi-hard negative mining, and the mining strategy used for a second portion of the batch, which is processed after the first portion, is hard negative mining.
[0307] Embodiment 14. The method of any one of embodiments 1 to 13, wherein the metric learning algorithm includes at least one of a contrast loss function, a triplet loss function, a quartet loss function, a circle loss function, a multi-class n-pairs loss function, a lifted structure loss function, an angle loss function, a divergence loss function, or a constellation loss function.
[0308] Embodiment 15. The method of any one of embodiments 1 to 14, further comprising training a machine learning model using a sampling strategy and a metric learning algorithm that prioritizes pairs of more similar peptide sequence vectors with different presenting major histocompatibility complex (MHC) alleles compared to more dissimilar peptide sequence vectors with different presenting MHC alleles.
[0309] Embodiment 16. The method of any one of embodiments 1 to 15, wherein generating a peptide sequence vector via a trained machine learning model comprises generating an n-dimensional vector of one of the peptides via the trained machine learning model using at least one of an embedding layer, a positional encoder, a transformer encoder, a self-attention layer, a summation and normalization layer, a feedforward layer, a fully connected layer, an activation layer, or a dropout layer.
[0310] Embodiment 17. The method of any one of embodiments 1 to 16, wherein generating a plurality of peptide sequence vectors via a trained machine learning model comprises converting peptide sequences of the plurality of peptide sequences into peptide representations that represent the peptide sequences, and converting the peptide representations into peptide sequence vectors for the peptide sequences.
[0311] Embodiment 18. The method of any one of embodiments 1 to 17, wherein the trained machine learning model comprises at least one of a convolutional neural network, a recurrent neural network, or a feedforward neural network.
[0312] Embodiment 19. The method of any one of embodiments 1 to 18, wherein the trained machine learning model comprises an attention-based machine learning model.
[0313] Embodiment 20. The method of any one of embodiments 1 to 19, further comprising generating training allele presentation data via a presentation model trained to identify one or more MHC alleles that are expected to present the peptide based on the peptide sequence identified for the peptide.
[0314] Embodiment 21 The method of any one of embodiments 1 to 20, wherein the plurality of peptide sequences is detected by processing a disease sample comprising tissue.
[0315] Embodiment 22 The method of any one of embodiments 1 to 21, further comprising generating a treatment recommendation for the subject, wherein the peptide vaccine comprises at least two candidate peptides from the group of candidate peptides.
[0316] Embodiment 23 The method of any one of embodiments 1 to 22, wherein the therapeutic agent is a peptide vaccine, and further comprising generating a report based on the output, wherein the report identifies a group of candidate peptides.
[0317] Embodiment 24 The method of embodiment 23, further comprising initiating action based on the report to facilitate production of a peptide vaccine.
[0318] Embodiment 25. The method of embodiment 24, wherein initiating an action comprises generating an alert that triggers a computerized process involved in the production of a peptide vaccine.
[0319] Embodiment 26. The method of any one of embodiments 1 to 25, further comprising generating a report based on the output, the report identifying the group of candidate peptides, and sending the report to a computing platform via a set of communication links, including at least one of a wired communication link or a wireless communication link.
[0320] Embodiment 27. The method of any one of embodiments 1 to 26, wherein the therapeutic agent is selected from the group consisting of T cell therapy, personalized cancer therapy, antigen-specific immunotherapy, antigen-dependent immunotherapy, a vaccine, and natural killer (NK) cell therapy.
[0321] Embodiment 28. The method of any one of embodiments 1 to 27, further comprising generating a report based on the output, wherein the report identifies a group of candidate peptides, and manufacturing the therapeutic agent to comprise a plurality of therapeutic peptides selected from the group of candidate peptides, a plurality of precursors of the therapeutic peptides, or at least one nucleic acid encoding a plurality of therapeutic peptides or precursors, wherein the plurality of therapeutic peptides comprises at least two different peptides.
[0322] Embodiment 29. The method of any one of embodiments 1 to 28, further comprising sequencing a disease sample from the subject, defining a plurality of peptide sequences based on the sequencing of the disease sample from the subject, synthesizing mRNA encoding at least two candidate peptides included in the group of candidate peptides, complexing the mRNA with lipids to produce an mRNA-lipoplex treatment, and administering the mRNA-lipoplex treatment to the subject.
[0323] Embodiment 30. A method for developing a peptide vaccine, the method including: training a machine learning model using a metric learning algorithm, training peptide sequence data, and training allele representation data corresponding to the training peptide sequence data; receiving peptide sequence data identifying a plurality of peptide sequences corresponding to a plurality of peptides; generating a peptide sequence vector for each peptide sequence of the plurality of peptide sequences using the peptide sequence data via the machine learning model to form a plurality of peptide sequence vectors; generating an output using the plurality of peptide sequence vectors, the output providing an indication of similarity between peptide sequences of the plurality of peptide sequences; and selecting a group of candidate peptides from the plurality of peptides for development of the peptide vaccine based on the output, such that the group of candidate peptides includes at least two different candidate peptides.
[0324] Embodiment 31. A method comprising: receiving training peptide sequence data comprising a plurality of training peptide sequences; generating training allele presentation data for the training peptide sequence data, wherein the training allele presentation data identifies, for a training peptide sequence among the plurality of training peptide sequences, MHC alleles that are predicted to present the training peptide sequence; and training a machine learning model using the training peptide sequence data, the training allele presentation data, and a metric learning algorithm, wherein the machine learning model is trained to generate a peptide sequence vector for a given peptide sequence, wherein the peptide sequence vector is a vector in n-dimensional space that provides an indication of similarity of the given peptide sequence to other peptide sequences.
[0325] Embodiment 32. The method of embodiment 31, further comprising generating a plurality of peptide sequence vectors for a plurality of peptide sequences detected in the disease sample via the trained machine learning model, and using the plurality of peptide sequence vectors to generate an output, wherein the output provides an indication of similarity between the peptide sequences of the plurality of peptide sequences.
[0326] Embodiment 33. The method of embodiment 32, further comprising selecting a group of candidate peptides from the plurality of peptides for development of a peptide vaccine based on the output, such that the group of candidate peptides comprises at least two candidate peptides with different binding motifs.
[0327] Embodiment 34. The method of embodiment 32 or 33, further comprising selecting a group of candidate peptides from the plurality of peptides for development of a peptide vaccine based on the output, such that bias towards any single binding motif is reduced in the group of candidate peptides.
[0328] Embodiment 35. The method of any one of embodiments 31 to 34, wherein training comprises calculating a distance metric for pairs of training peptide sequence vectors in the batch of training peptide sequence vectors, forming an evaluation bundle from the batch of training peptide sequence vectors based on the distance metric and the mining strategy, evaluating the loss of the evaluation bundle using a metric learning algorithm, and updating parameters of the machine learning model based on the loss.
[0329] Embodiment 36. The method of any one of embodiments 31 to 35, wherein training comprises training the machine learning model using a sampling strategy that prioritizes pairs of more similar peptide sequence vectors with different major histocompatibility complex (MHC) alleles compared to more dissimilar peptide sequence vectors with different MHC alleles, and a metric learning algorithm.
[0330] Embodiment 37. The method of any one of embodiments 31 to 36, wherein the metric learning algorithm includes at least one of a contrast loss function, a triplet loss function, a quartet loss function, a circle loss function, a multi-class n-pairs loss function, a lifted structure loss function, an angle loss function, a divergence loss function, or a constellation loss function.
[0331] Embodiment 38. The method of any one of embodiments 31 to 37, further comprising generating training allele presentation data via a presentation model trained to identify one or more MHC alleles that are expected to present the peptide based on the peptide sequence identified for the peptide.
[0332] Embodiment 39. A vaccine comprising a plurality of peptides, a plurality of precursors of a plurality of peptides, or a set of nucleic acids encoding a plurality of peptides or a plurality of precursors, wherein the plurality of peptides are selected from a group of candidate peptides selected based on the method of any one of embodiments 1 to 38, and the plurality of peptides comprises at least two peptides having different binding motifs.
[0333] Embodiment 40. The vaccine of embodiment 39, wherein the vaccine comprises DNA comprising the set of nucleic acids or RNA comprising the set of nucleic acids.
[0334] Embodiment 41. The vaccine of embodiment 39 or 40, wherein the vaccine comprises mRNA comprising the set of nucleic acids.
[0335] Embodiment 42 The vaccine of any one of embodiments 39 to 41, wherein the vaccine is a tumor vaccine.
[0336] Embodiment 43. A method for producing a vaccine, comprising: producing a vaccine comprising a plurality of peptides, a plurality of precursors of a plurality of peptides, or a set of nucleic acids encoding a plurality of peptides or a plurality of precursors, wherein the plurality of peptides are selected from among a group of candidate peptides selected according to the method of any one of embodiments 1 to 38, and wherein the plurality of peptides comprises at least two peptides having different binding motifs.
[0337] Embodiment 44. The method of embodiment 43, wherein the vaccine comprises DNA comprising the set of nucleic acids, RNA comprising the set of nucleic acids, or mRNA comprising the set of nucleic acids.
[0338] Embodiment 45. The method of embodiment 43 or 44, further comprising identifying a set of nucleic acids encoding the plurality of peptides based on the amino acids in the plurality of peptides, wherein the vaccine comprises the set of nucleic acids.
[0339] Embodiment 46 The method of any one of embodiments 43 to 45, wherein the vaccine is a tumor vaccine.
[0340] Embodiment 47. The method of embodiment 46, wherein, for each peptide of the plurality of peptides, the tumor vaccine comprises at least one of a nucleotide sequence encoding each peptide, an amino acid sequence corresponding to each peptide, RNA corresponding to each peptide, DNA corresponding to each peptide, mRNA corresponding to each peptide, a cell corresponding to each peptide, a plasmid corresponding to each peptide, or a vector corresponding to each peptide.
[0341] Embodiment 48. The method of any one of embodiments 43 to 47, wherein the vaccine further comprises at least one of an excipient or an adjuvant.
[0342] Embodiment 49. The method of any one of embodiments 43 to 48, wherein the vaccine comprises an RNA molecule comprising, in a 5' to 3' direction, a 5' cap, a 5' untranslated region (UTR), a polynucleotide sequence encoding a secretory signal peptide, a polynucleotide sequence encoding a plurality of peptides, a polynucleotide sequence encoding at least a portion of the transmembrane and cytoplasmic domains of a major histocompatibility complex (MHC) molecule, a 3' UTR comprising the 3' untranslated region of an Amino-Terminal Enhancer of Split (AES) mRNA or a fragment thereof, and a non-coding RNA of mitochondrially encoded 12S RNA or a fragment thereof, and a poly(A) sequence.
[0343] Embodiment 50. A pharmaceutical composition comprising two or more peptides selected from a group of candidate peptides selected according to the method of any one of embodiments 1 to 38.
[0344] Embodiment 51. A pharmaceutical composition comprising two or more nucleic acid sequences encoding two or more respective peptides selected from a group of candidate peptides selected based on the method of any one of embodiments 1 to 38.
[0345] Embodiment 52. A method of treating a subject, comprising administering at least one of two or more peptides, two or more pharmaceutical compositions, or two or more nucleic acid sequences identified based on output generated by the method of any one of embodiments 1 to 38.
[0346] Embodiment 53. An engineered T cell produced using the method of any one of embodiments 1 to 38.
[0347] Embodiment 54. A population of engineered T cells produced using the method of any one of embodiments 1 to 38.
[0348] Embodiment 55. A method for treating a subject having cancer, comprising: providing a population of T cells; engineering at least a subset of the population of T cells to express an exogenous T cell receptor (TCR) and knock out endogenous TCR-β, thereby forming a population of engineered T cells, wherein the exogenous TCR is expressed by the cancer and binds to an antigen selected using the method of any one of embodiments 1 to 38; expanding the population of engineered T cells; and administering the expanded population of engineered T cells to the subject.
[0349] Embodiment 56. The method of embodiment 55, wherein the antibody is a neoantigen or TAA.
[0350] Embodiment 57. The method of embodiment 55 or 56, wherein at least a portion of the cancer genome and / or transcriptome has been sequenced to determine the presence of the antigen.
[0351] Embodiment 58 The method of any one of embodiments 55 to 57, wherein the engineered T cells are produced using a method of any one of embodiments 1 to 38.
[0352] Embodiment 59. The method of any one of embodiments 55 to 58, wherein the TCR binds to an antigen presented on a major histocompatibility complex class I (MHCI) molecule.
[0353] Embodiment 60. The method of embodiment 59, wherein the MHCI comprises an MHCI allele expressed by the subject.
[0354] Embodiment 61. The expanded population of engineered T cells comprises 1 x 10 5 pieces~1×10 11 61. The method of any one of embodiments 55 to 60, comprising engineered T cells.
[0355] Embodiment 62. The method of any one of embodiments 55 to 61, wherein the T cells are autologous to the subject.
[0356] Embodiment 63 The method of any one of embodiments 55 to 61, wherein the T cells are allogeneic to the subject.
[0357] Embodiment 64. A method of treating cancer, comprising administering to a patient having cancer a T cell, composition or pharmaceutical composition of any of embodiments 55 to 63.
[0358] Embodiment 65. A system including one or more data processors and a non-transitory computer-readable storage medium containing instructions, which, when executed on the one or more data processors, cause the one or more data processors to perform a method according to any one of embodiments 1 to 38.
[0359] Embodiment 66. A computer program product tangibly embodied in a non-transitory machine-readable storage medium comprising instructions configured to cause one or more data processors to perform the method of any one of embodiments 1 to 38.
[0360] X. Further Considerations Headings and subheadings between sections and subsections herein are included merely to improve readability and do not imply that features may be combined across sections and subsections, and therefore, the sections and subsections do not describe separate embodiments.
[0361] The terms and expressions which have been employed are used as terms of description rather than of limitation, and there is no intention in the use of such terms and expressions to exclude equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, although the claimed invention has been specifically disclosed by embodiments and optional features, it will be understood that modifications and variations of the concepts disclosed herein may be employed by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims.
[0362] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium including instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium, the computer program product including instructions configured to cause one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein.
[0363] The description presents only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the description of preferred exemplary embodiments provides those skilled in the art with an enabling description for implementing various embodiments. It will be understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope as set forth in the appended claims.
[0364] In the following description, specific details are given to provide a comprehensive understanding of the embodiments. However, it will be understood that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order to avoid obscuring the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.
Claims
1. 1. A method for developing a therapeutic agent, said method comprising: receiving peptide sequence data identifying a plurality of peptide sequences corresponding to a plurality of peptides; generating, via a trained machine learning model, a plurality of peptide sequence vectors in n-dimensional space for each of said plurality of peptide sequences, and thereby for each of said plurality of peptides; the machine learning model is trained using a metric learning algorithm, training peptide sequence data, and training allele representation data; the training peptide sequence data identifying a training peptide sequence corresponding to each training peptide of a plurality of training peptides; the training allele presentation data identifies, for each training peptide sequence of the plurality of training peptides in the training peptide sequence data, one or more major histocompatibility complex (MHC) alleles that are predicted to present the training peptide corresponding to each training peptide sequence; the metric learning algorithm has been used to train the machine learning model to output the plurality of peptide sequence vectors in the n-dimensional space such that a first distance between a first pair of the peptide sequence vectors in the plurality of peptide sequence vectors generated in the n-dimensional space for each first pair of the peptides presented by the same MHC allele is less than a second distance between a second pair of the peptide sequence vectors in the plurality of peptide sequence vectors generated in the n-dimensional space for a second pair of the peptides presented by different MHC alleles. generating a plurality of peptide sequence vectors in n-dimensional space; and generating an output using the plurality of peptide sequence vectors, the output providing an indication of similarity between the peptide sequences for use in selecting a group of candidate peptides from the plurality of peptides for development of the therapeutic agent; A method comprising:
2. 2. The method of claim 1, further comprising selecting a group of candidate peptides from the plurality of peptides for the development of the therapeutic agent based on the output such that at least two candidate peptides in the group have different binding motifs.
3. 3. The method of claim 1 or 2, further comprising selecting the group of candidate peptides from the plurality of peptides for the development of the therapeutic agent based on the output such that bias towards any single binding motif is reduced.
4. selecting the group of candidate peptides from the plurality of peptides, using the output to identify a plurality of clusters of the peptide sequence vectors; and selecting at least one peptide sequence vector from each of said plurality of clusters for use in said development of said therapeutic agent; The method of claim 2 or 3, comprising:
5. selecting the at least one peptide sequence vector for a selected cluster of the plurality of clusters, selecting one peptide sequence vector from the selected cluster that is closest to a center of the selected cluster, wherein the center is selected from one of a centroid, a mean center, a median center, and a density-based center; The method of claim 4, comprising:
6. selecting said at least one peptide sequence vector for a selected one of said clusters; selecting at least two peptide sequence vectors from the selected clusters, each of the at least two peptide sequence vectors is midway between the center of the cluster and the edge of the cluster; or each of the at least two peptide sequence vectors is arranged along the edge of the cluster; selecting at least two peptide sequence vectors; The method of claim 4 or 5, comprising:
7. The method of claim 1 , wherein the therapeutic agent comprises at least two candidate peptides from the group of candidate peptides, and the at least two candidate peptides have different binding motifs.
8. Training the machine learning model includes: training the machine learning model using the metric learning algorithm, the training peptide sequence data, and the training allele representation data; 8. The method of claim 1, comprising:
9. Training the machine learning model includes: calculating a distance metric for pairs of training peptide sequence vectors in the batch of training peptide sequence vectors; forming an evaluation bundle from the batch of training peptide sequence vectors based on the distance metric and the mining strategy; evaluating the loss of the evaluation bundle using the metric learning algorithm; and updating parameters of the machine learning model based on the loss; The method of claim 8, comprising:
10. Training the machine learning model includes: repeating the calculating, forming, and evaluating steps for multiple batches of training peptide sequence vectors formed from the training peptide sequence data; 10. The method of claim 9, comprising:
11. The method of claim 10 , wherein the updating step is performed after the batch is processed.
12. The method of claim 10 , wherein the updating step is performed for each of the batches.
13. 13. The method of claim 9, wherein the mining strategy used for a first portion of the batch is semi-hard negative mining, and the mining strategy used for a second portion of the batch, which is processed after the first portion, is hard negative mining.
14. 14. The method of any one of claims 1 to 13, wherein the metric learning algorithm comprises at least one of a contrast loss function, a triplet loss function, a quartet loss function, a circular loss function, a multi-class n-pairs loss function, a lifted structure loss function, an angular loss function, a divergence loss function, or a constellation loss function.
15. 15. The method of any one of claims 1 to 14, further comprising training a machine learning model using a sampling strategy that prioritizes pairs of more similar peptide sequence vectors with different presenting major histocompatibility complex (MHC) alleles compared to more dissimilar peptide sequence vectors with different presenting MHC alleles, and the metric learning algorithm.
16. generating the plurality of peptide sequence vectors via the trained machine learning model, generating an n-dimensional vector of one of the peptides via the trained machine learning model using at least one of an embedding layer, a positional encoder, a transformer encoder, a self-attention layer, a summation and normalization layer, a feedforward layer, a fully connected layer, an activation layer, or a dropout layer; 16. The method of any one of claims 1 to 15, comprising:
17. generating the plurality of peptide sequence vectors via the trained machine learning model, converting peptide sequences of the plurality of peptide sequences into peptide representations that represent the peptide sequences; and converting said peptide representation to said peptide sequence vector for said peptide sequence; 17. The method of any one of claims 1 to 16, comprising:
18. 18. The method of any one of claims 1 to 17, wherein the trained machine learning model comprises at least one of a convolutional neural network, a recurrent neural network, or a feedforward neural network.
19. 19. The method of claim 1, wherein the trained machine learning model comprises an attention-based machine learning model.
20. 20. The method of any one of claims 1 to 19, further comprising generating the training allele presentation data via a presentation model trained to identify the one or more MHC alleles that are expected to present the peptide based on a peptide sequence identified for the peptide.
21. 21. The method of any one of claims 1 to 20, wherein the plurality of peptide sequences is detected by processing a disease sample comprising tissue.
22. 22. The method of claim 1, further comprising generating a treatment recommendation for the subject that identifies a peptide vaccine comprising at least two candidate peptides from the group of candidate peptides.
23. the therapeutic agent is a peptide vaccine, generating a report based on the output, the report identifying the group of candidate peptides; 23. The method of any one of claims 1 to 22, further comprising:
24. 24. The method of claim 23, further comprising initiating an action based on said report that facilitates the production of said peptide vaccine.
25. Initiating said measures generating an alert that triggers a computerized process involved in the production of the peptide vaccine; 25. The method of claim 24, comprising:
26. generating a report based on the output, the report identifying the group of candidate peptides; and sending said report to a computing platform over a set of communication links, the set including at least one of a wired communication link or a wireless communication link; 26. The method of any one of claims 1 to 25, further comprising:
27. 27. The method of any one of claims 1 to 26, wherein the therapeutic agent is selected from the group consisting of T cell therapy, personalized cancer therapy, antigen-specific immunotherapy, antigen-dependent immunotherapy, a vaccine, and natural killer (NK) cell therapy.
28. generating a report based on the output, the report identifying the group of candidate peptides; and manufacturing the therapeutic agent to include a plurality of therapeutic peptides selected from the group of candidate peptides, a plurality of precursors of the plurality of therapeutic peptides, or at least one nucleic acid encoding the plurality of therapeutic peptides or the plurality of precursors; 28. The method of any one of claims 1 to 27, further comprising:
29. sequencing a disease sample from the subject; defining the plurality of peptide sequences based on the sequencing of the disease sample from the subject; synthesizing mRNA encoding at least two candidate peptides included in the group of candidate peptides; complexing the mRNA with a lipid to produce an mRNA-lipoplex treatment, and administering the mRNA-lipoplex treatment to the subject; 29. The method of any one of claims 1 to 28, further comprising:
30. 1. A method for developing a peptide vaccine, said method comprising: training a machine learning model using a metric learning algorithm, training peptide sequence data, and training allele representation data corresponding to said training peptide sequence data; receiving peptide sequence data identifying a plurality of peptide sequences corresponding to a plurality of peptides; generating a peptide sequence vector for each peptide sequence of the plurality of peptide sequences using the peptide sequence data via the machine learning model to form a plurality of peptide sequence vectors; generating an output using the plurality of peptide sequence vectors, the output providing an indication of similarity between peptide sequences of the plurality of peptide sequences; and selecting a group of candidate peptides from the plurality of peptides for development of the peptide vaccine based on the output, such that the group of candidate peptides comprises at least two different candidate peptides; A method comprising:
31. receiving training peptide sequence data comprising a plurality of training peptide sequences; generating training allele presentation data for the training peptide sequence data, the training allele presentation data identifying, for a training peptide sequence among the plurality of training peptide sequences, an MHC allele that is predicted to present the training peptide sequence; and training a machine learning model using the training peptide sequence data, the training allele representation data, and a metric learning algorithm; Including, the machine learning model is trained to generate a peptide sequence vector for a given peptide sequence; The method, wherein said peptide sequence vector is a vector in n-dimensional space that provides a measure of the similarity of said given peptide sequence to other peptide sequences.
32. generating a plurality of peptide sequence vectors for a plurality of peptide sequences detected in the disease sample via the trained machine learning model; and generating an output using the plurality of peptide sequence vectors, the output providing an indication of similarity between peptide sequences of the plurality of peptide sequences; 32. The method of claim 31 , further comprising:
33. 33. The method of claim 32, further comprising selecting a group of candidate peptides from the plurality of peptides for development of a peptide vaccine based on the output, such that the group of candidate peptides includes at least two candidate peptides having different binding motifs.
34. The method of claim 32 or 33, further comprising selecting a group of candidate peptides from the plurality of peptides for development of a peptide vaccine based on the output, such that bias towards any single binding motif is reduced in the group of candidate peptides.
35. The training includes: calculating a distance metric for pairs of training peptide sequence vectors in the batch of training peptide sequence vectors; forming an evaluation bundle from the batch of training peptide sequence vectors based on the distance metric and the mining strategy; evaluating the loss of the evaluation bundle using the metric learning algorithm; and updating parameters of the machine learning model based on the loss; 35. The method of any one of claims 31 to 34, comprising:
36. The training includes: training the machine learning model using a sampling strategy that prioritizes pairs of more similar peptide sequence vectors with different presenting major histocompatibility complex (MHC) alleles compared to more dissimilar peptide sequence vectors with different presenting MHC alleles, and the metric learning algorithm; 36. The method of any one of claims 31 to 35, comprising:
37. 37. The method of any one of claims 31 to 36, wherein the metric learning algorithm comprises at least one of a contrast loss function, a triplet loss function, a quartet loss function, a circular loss function, a multi-class n-pairs loss function, a lifted structure loss function, an angular loss function, a divergence loss function, or a constellation loss function.
38. 38. The method of any one of claims 31 to 37, further comprising generating the training allele presentation data via a presentation model trained to identify the one or more MHC alleles expected to present peptides based on peptide sequences identified for the peptides.
39. Multiple peptides, a plurality of precursors of said plurality of peptides; or A set of nucleic acids encoding said plurality of peptides or said plurality of precursors Including, The plurality of peptides is selected from the group of candidate peptides selected according to the method of any one of claims 1 to 38, the plurality of peptides comprises at least two peptides having different binding motifs; vaccine.
40. 40. The vaccine of claim 39, wherein the vaccine comprises DNA comprising the set of nucleic acids or RNA comprising the set of nucleic acids.
41. 41. The vaccine of claim 39 or 40, wherein the vaccine comprises mRNA comprising the set of nucleic acids.
42. 42. The vaccine of any one of claims 39 to 41, wherein the vaccine is a tumor vaccine.
43. 1. A method of producing a vaccine, comprising: Multiple peptides, a plurality of precursors of said plurality of peptides; or A set of nucleic acids encoding said plurality of peptides or said plurality of precursors 1. A vaccine comprising: The plurality of peptides is selected from the group of candidate peptides selected according to the method of any one of claims 1 to 38, producing a vaccine wherein the plurality of peptides comprises at least two peptides having different binding motifs; method.
44. 44. The method of claim 43, wherein the vaccine comprises DNA comprising the set of nucleic acids, RNA comprising the set of nucleic acids, or mRNA comprising the set of nucleic acids.
45. 45. The method of claim 43 or 44, further comprising identifying the set of nucleic acids encoding the plurality of peptides based on amino acids in the plurality of peptides, wherein the vaccine comprises the set of nucleic acids.
46. 46. The method of any one of claims 43 to 45, wherein the vaccine is a tumor vaccine.
47. 47. The method of claim 46, wherein for each peptide of the plurality of peptides, the tumor vaccine comprises at least one of a nucleotide sequence encoding each peptide, an amino acid sequence corresponding to each peptide, RNA corresponding to each peptide, DNA corresponding to each peptide, mRNA corresponding to each peptide, a cell corresponding to each peptide, a plasmid corresponding to each peptide, or a vector corresponding to each peptide.
48. 48. The method of any one of claims 43 to 47, wherein the vaccine further comprises at least one of an excipient or an adjuvant.
49. The vaccine comprises an RNA molecule, the RNA molecule comprising, in a 5' to 3' direction: 5' cap, 5' untranslated region (UTR), a polynucleotide sequence encoding a secretory signal peptide; a polynucleotide sequence encoding said plurality of peptides; a polynucleotide sequence encoding at least a portion of the transmembrane and cytoplasmic domains of a major histocompatibility complex (MHC) molecule; a 3′UTR, Amino-Terminal Enhancer of Split (AES) mRNA 3' untranslated region or a fragment thereof, and a 3'UTR comprising a non-coding RNA of mitochondrially encoded 12S RNA or a fragment thereof; and poly(A) sequence, 49. The method of any one of claims 43 to 48, comprising:
50. 39. A pharmaceutical composition comprising two or more peptides selected from the group of candidate peptides selected based on the method of any one of claims 1 to 38.
51. 39. A pharmaceutical composition comprising two or more nucleic acid sequences encoding two or more respective peptides selected from the group of candidate peptides selected based on the method of any one of claims 1 to 38.
52. 39. A method of treating a subject, comprising administering at least one of two or more peptides, two or more pharmaceutical compositions, or two or more nucleic acid sequences identified based on the output generated by the method of any of claims 1 to 38.
53. 39. An engineered T cell produced using the method of any one of claims 1 to 38.
54. 39. A population of engineered T cells produced using the method of any one of claims 1 to 38.
55. 1. A method for treating a subject having cancer, comprising: providing a population of T cells; engineering at least a subset of the population of T cells to express an exogenous T cell receptor (TCR) and knock out endogenous TCR-β, thereby forming a population of engineered T cells, wherein the exogenous TCR binds to an antigen expressed by the cancer and selected using the method of any one of claims 1 to 38; expanding said population of engineered T cells; and administering the expanded population of engineered T cells to the subject; A method comprising:
56. 56. The method of claim 55, wherein the antibody is a neoantigen or a TAA.
57. 57. The method of claim 55 or 56, wherein at least a portion of the genome and / or transcriptome of the cancer has been sequenced to determine the presence of the antigen.
58. 58. The method of any one of claims 55 to 57, wherein the engineered T cells are made using a method of any one of claims 1 to 38.
59. 59. The method of any one of claims 55 to 58, wherein the TCR binds to the antigen presented on a major histocompatibility complex class I (MHCI) molecule.
60. 60. The method of claim 59, wherein the MHCI comprises MHCI alleles expressed by the subject.
61. The expanded population of engineered T cells is 1 x 10 5 pieces ~ 1×10 11 61. The method of any one of claims 55 to 60, comprising engineered T cells.
62. 62. The method of any one of claims 55 to 61, wherein the T cells are autologous to the subject.
63. 62. The method of any one of claims 55 to 61, wherein the T cells are allogeneic to the subject.
64. 64. A method of treating cancer comprising administering to a patient having cancer a T cell, composition or pharmaceutical composition of any of claims 55 to 63.
65. 1. A system comprising: one or more data processors; and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform the method of any one of claims 1 to 38.
66. 39. A computer program product tangibly embodied in a non-transitory machine-readable storage medium comprising instructions configured to cause one or more data processors to perform the method of any one of claims 1 to 38.